Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
125,979 characters · 10 sections · 140 citation commands
Efficient Semiparametric Estimation of Average Treatment Effects Under Covariate Adaptive Randomization
{\it Keywords:} Efficient semiparametric estimation, average treatment effect, randomized experiments, covariate adaptive randomization
{\it JEL Classification:} C14, C90
Experiments that use covariate adaptive randomization (CAR) are commonplace in applied economics and other fields. In such experiments, the experimenter first groups sample units according to observed baseline covariates (stratification) and then assigns treatment randomly to achieve balance within these groups (strata). The term balance here means that the experimenter additionally specifies stratum-specific target treatment proportions and assigns treatment so that the proportion of units assigned to treatment reaches the corresponding target as the sample size grows across strata. A textbook treatment of covariate adaptive and stratified randomization in clinical trials can be found in 2015rosenbergerRandomizationClinicalTrials. For review articles on their use in development economics, see 2007dufloChapter61Using, 2009bruhnPursuitBalanceRandomization and 2017atheyChapterEconometricsRandomized. In experiments comparing outcomes from binary treatments, the quantity of interest is often the average treatment effect (ATE). In this paper, we are concerned with efficient semiparametric estimation of the ATE while allowing for a broad class of CAR procedures motivated by those used in practice. We have two main questions. The first is: is there a well-defined semiparametric efficiency bound (SPEB) for this class of procedures? The second is: does there exist a semiparametric estimator that achieves this bound and under what conditions will this happen? Our answers to both are affirmative. For the first, we show that a version of the bound in 1998hahnRolePropensityScore is valid. For the second, we show that under the same (weak) conditions used for derivation of the bound, there is a feasible estimator that achieves the bound asymptotically.
In randomized experiments, correctly implemented randomization ensures that in expectation, confounding factors are equally distributed across treatment arms so that differences in outcomes are solely due to the differences in treatment. Stratified randomization additionally aims to ensure that along observable dimesions, this also remains true in practice. In CAR experiments, both discrete and continuously distributed covariates are combined to form strata. As shown in 2017atheyChapterEconometricsRandomized, the main statistical benefit of stratified randomization for ATE estimation is improved precision. The standard recommendation for ATE estimation in stratified experiments is to regress the observed outcomes on indicators for strata and their interactions with treatment status in a linear regression equation (a fully saturated linear regression model). The coefficients on the interaction terms from this regression are then combined with sample stratum proportions to construct the ATE estimate. The resulting estimator of the ATE is analogous to the Horvitz-Thompson estimator (1952horvitzGeneralizationSamplingReplacement) and does not use information beyond the strata. However, experimenters concerned with estimation precision may want to use information in the baseline covariates not contained in the strata.
An additional complication in the CAR context comes from the choice of treatment assignment mechanism by the experimenter. Many popular treatment assignment mechanisms used in practice aim for faster targeting of the target assignment proportions than simple independent and identically distributed (i.i.d.) assignment and in doing so, induce dependence in the observed outcomes through dependence in the treatment assignments. One such example is stratified permuted block randomization (SPBR, \th(ref)). 2018bugniInferenceCovariateAdaptiveRandomization provide additional examples. The dependence induced by CAR designs can affect the behavior of ATE estimators in surprising ways. Analyzing the case where target assignment proportions are constant across strata, 2018bugniInferenceCovariateAdaptiveRandomization show that the standard difference in treatment and control group means can have a limit variance that depends explicitly on the choice of treatment assignment mechanism. For example, all else held equal, the same estimator exhibits a larger limit variance under i.i.d. treatment assignments than when treatments are assigned according to SPBR, even when both mechanisms have the same target proportions. 2018bugniInferenceCovariateAdaptiveRandomization also show that the same phenomenon holds true of the “stratum fixed effects” estimator which is recommended by 2009bruhnPursuitBalanceRandomization. 2019bugniInferenceCovariateAdaptiveRandomization extends this work and consider both multiple treatments and target assignment proportions that vary by strata. They show that the estimator of the ATE from a fully saturated regression has a limit variance that depends on the target proportions (among other things), but not on the particular choice of assignment mechanism. They also show that for the stratum fixed effects estimator however, the limit variance is still affected explicitly by the choice of assignment mechanism. 2018bugniInferenceCovariateAdaptiveRandomization, 2019bugniInferenceCovariateAdaptiveRandomization do not consider efficiency questions.
There is a literature that considers efficiency gains in experiments from using baseline covariates via linear regression adjustments. However, whether there are gains at all depends on the linear regression specification. In the case without stratification, 2008freedmanRegressionAdjustmentsExperimental shows that linear regression adjustments (without interactions between treatment status and baseline covariates) can hurt asymptotic precision of the ATE estimates though the estimates remain consistent. However, work by 2001yangEfficiencyStudyEstimators and 2013linAgnosticNotesRegression show that appropriately formed regression adjustments (i.e. with the correct interaction terms) cannot hurt (and indeed can improve) asymptotic precision in ATE estimation. For CAR experiments, 2022maRegressionAnalysisCovariate build on the results of 2018bugniInferenceCovariateAdaptiveRandomization, 2019bugniInferenceCovariateAdaptiveRandomization and show that correctly formed linear regression adjustments cannot hurt (and can improve) asymptotic precision under CAR. These works do not treat semiparametrically efficient estimation.
In this paper, we are concerned with the semiparametrically efficient estimation of the ATE in the CAR framework under the minimal set of assumptions laid out by 2019bugniInferenceCovariateAdaptiveRandomization. There is a large statistics literature around semiparametric efficiency, starting with the seminal work of 1956steinEfficientNonparametricTesting. Most of these are developed assuming i.i.d. data. 1998bickelEfficientAdaptiveEstimation offers a comprehensive textbook treatment and 1990neweySemiparametricEfficiencyBounds provides an approachable review. Notable examples with non-i.i.d. data can be found in 2001bickelInferenceSemiparametricModels, 2004greenwoodIntroductionEfficientEstimation and 2010komunjerSemiparametricEfficiencyBound as well in references therein. For treatment effects, 1998hahnRolePropensityScore derives the SPEB for the ATE in observational studies with i.i.d. data when treatment assignment is ignorable conditional on observable covariates (ignorable as in 1983rosenbaumCentralRolePropensity). The 1998hahnRolePropensityScore bound does not apply immediately to our setting since the treatment assignment rule is allowed to depend on the entire profile of sample strata. For instance, it is not clear a priori if the choice of assignment mechanism will affect the efficiency bound since it can clearly influence limit behavior of estimators as in 2018bugniInferenceCovariateAdaptiveRandomization, 2019bugniInferenceCovariateAdaptiveRandomization. Furthermore, even when treatment assignments are i.i.d. in a CAR context, the covariates being conditioned on during assignment are the strata, so it is again not immediate from the bound in 1998hahnRolePropensityScore what role the additional baseline covariates can play in providing efficiency gains. This paper shows that a version of the 1998hahnRolePropensityScore bound accounting for both stratum-specific target proportions and all baseline covariates is the SPEB across all CAR experimental procedures. The target proportions play the role of the propensity score conditional on baseline covariates. Additionally, the choice of assignment mechanism does not affect this bound, only the target proportions do. We derive this by using the partial sums arguments of 2018bugniInferenceCovariateAdaptiveRandomization, 2019bugniInferenceCovariateAdaptiveRandomization to show that the log likelihood ratios under \(1 / \sqrt{n}\) local alternatives has the local asymptotic normality (LAN) property of 1960lecamLocallyAsymptoticallyNormal.
In concurrent work, 2022armstrongAsymptoticEfficiencyBounds proves the LAN property under a larger class of experimental designs using martingale methods. The experiments considered there include the CAR framework and additionally allows for arbitrary dependence of the treatment rule on the covariates as well as observations of past outcomes (to account for sequential assignment). 2022armstrongAsymptoticEfficiencyBounds's main goal is to show that an optimized form of the 1998hahnRolePropensityScore bound is a lower bound on limit variance of the ATE across all experimental designs in that class. This is motivated by recent papers on optimal design of experiments by using past waves to either allocate treatment sequentially (e.g. 2011hahnAdaptiveExperimentalDesign) or to form optimal strata in a main experiment (e.g. 2022tabord-meehanStratificationTreesAdaptive, 2022baiOptimalityMatchedPairDesigns and 2021cytrynbaumDesigningRepresentativeBalanced). The optimized bound is analogous to implementing a “conditional on covariates” Neyman allocation (1934neymanTwoDifferentAspects). We differ from 2022armstrongAsymptoticEfficiencyBounds work on two counts. First, while their lower bound holds across the designs the aforementioned class, they do not provide efficiency bounds in specific instances within the class. We provide efficiency bounds for a given fixed stratification scheme and we do not consider the question of optimizing the bound. As a result our efficiency bound as a variance lower bound is higher than the optimized one in 2022armstrongAsymptoticEfficiencyBounds and hence sharper for a given fixed stratification scheme. Second, 2022armstrongAsymptoticEfficiencyBounds does not consider the question of when the bounds are achievable. We do this for the CAR framework explicitly as explained in the subsequent two paragraphs.
Once an efficiency bound is established, the question of its sharpness remains, in the sense of whether or not it is achievable. A well known phenomenon in the literature is that conditions under which a SPEB is achievable are typically much stronger than those required to derive the bound. 1990ritovAchievingInformationBounds provide counterexamples where a finite and non-singular SPEB exists but in certain non-trivial submodels, even consistent (let alone efficient) estimation of the parameter of interest is impossible. One of their examples is the partially linear model (1986engleSemiparametricEstimatesRelation, 1988robinsonRootNConsistent) which is commonly used in applied economics. A well known condition in the literature is that \(1 / \sqrt{n}\)-consistent and asymptotically normal (\(1 / \sqrt{n}\)-CAN) two-step semiparametric estimators require any first step infinite-dimensional (henceforth nonparametric) nuisance parameters to be estimated at a rate faster than \(n^{- \frac{1}{4}}\) where \(n\) is the sample size (see for instance 1994neweyAsymptoticVarianceSemiparametric and 2003chenEstimationSemiparametricModels). This is mainly because nonparametric estimators exhibit considerable bias and this condition limits the effect of this bias on the second estimation step. The \(n^{- \frac{1}{4}}\) rate condition however is especially restrictive when the covariates are of higher dimension, due to the curse of dimensionality. Achieving this rate condition typically requires imposing smoothness and/or Donsker conditions (often dimension dependent) on unknown nuisance parameters and the estimator in question. For the ATE, 1998hahnRolePropensityScore shows that semiparametrically efficient estimation can be done under regularity conditions through nonparametric regression adjustments or imputation. In their first step, series estimators of conditional means and the propensity score are used. Estimation by kernel methods can also be done, see 2004imbensNonparametricEstimationAverage for a review. In all of these, the use of smoothness restrictions on nonparametric population unknowns is ubiquitous.
An additional contribution of this paper is to show that the SPEB derived for the ATE is achievable under the same conditions used for its derivation. We do this by leveraging the efficient influence function to form an estimating equation (or moment condition) for the ATE. The resulting estimator is the familiar (and famed) augmented inverse probability weighted (AIPW) estimators due to 1994robinsEstimationRegressionCoefficients, 1995robinsAnalysisSemiparametricRegression, and 1999scharfsteinAdjustingNonignorableDrop. We further note that since the propensity score is known in these experiments, the estimating equation is linear in the unknown nonparametric component and has the local robustness property of 2018chernozhukovDoubleDebiasedMachine and 2022chernozhukovLocallyRobustSemiparametric. This reduces the first order impact of bias in the first stage nonparametric estimates on the resulting ATE estimator considerably and allows for efficient estimation under much weaker conditions. The nonparametric unknowns here are conditional means of the potential outcomes given baseline covariates. We show first that for a generic estimator of conditional means, a weak \(L_{2}\) consistency requirement combined with cross-fitting as in 2018chernozhukovDoubleDebiasedMachine is sufficient for efficient estimation of the ATE under CAR. The conditions under which the particular choice of estimator achieves this \(L_{2}\) consistency are first left abstract - different conditions apply to kernel estimators, random forests, series estimators and neural networks for instance. For a given choice of nonparametric estimator, smoothness conditions or functional form restrictions may be unavoidable in achieving the \(L_{2}\) consistency property. Next, we establish that if the particular nonparametric estimator is a cross-fitted Nadaraya-Watson kernel regression estimator, then efficient estimation is possible under no additional restrictions on the semiparametric model. Our motivation for using the cross-fitted Nadaraya-Watson estimator is two-fold. First, the results of 1980devroyeDistributionFreeConsistencyResults and 1980spiegelmanConsistentWindowEstimation show that the Nadaraya-Watson estimator is universally \(L_{2}\) consistent if the outcome in the regression has a finite second moment. Second, the algebra of the cross-fitted Nadaraya-Watson estimator is convenient in showing negligiblility of remainder terms.
We provide simulation evidence of the finite sample performance of the feasible efficient estimator. Our simulations compare this estimator to an infeasible efficient estimator as well as an estimator that uses information only from the stratum labels but not the additional baseline covariates. A comparison is also provided with an imputation estimator of 1994chengNonparametricEstimationMean and 1998hahnRolePropensityScore. Across all simulations designs, we find that using the feasible efficient estimator produces an efficiency gain of 13% in terms of mean squared error reduction in comparison to the estimator that discards information from baseline covariates beyond strata. The maximal gain in this comparison over all simulations is a 40% reduction in mean squared error. Compared to the imputation estimator, our estimator exhibits considerably less bias across most simulation designs.
The remainder of the paper is organized as follows. Section (ref) sets up the assumptions about the underlying population from which the sample is drawn as well as the sampling process and the treatment assignment scheme. Section (ref) presents the main result concerning the semiparametric efficiency bound (\th(ref)) after a discussion about parametric submodels in CAR experiment context. Section (ref) shows that the efficiency bound is achievable under the same assumptions as used for its derivation via a two-step semiparametric estimator using the efficient influence function. Section (ref) provides Monte Carlo evidence of the performance of the efficient estimator in finite samples. Section (ref) concludes the paper.
This section describes assumptions on the populations of interest as well as the sampling process which produces the observed data. This is done in two subsections. Subsection (ref) first describes the unobserved study population of interest through a Neyman-Rubin causal model. Then, a target observed population is described for all CAR experiments considered in this paper. This target observed population is useful for deriving the efficiency bound in Section (ref). Next, Subsection (ref) describes sampling assumptions and provides the two main examples of sampling schemes in this paper.
Throughout this paper, all random variables and vectors will be defined on a sufficiently rich underlying probability space \((\Omega, \mathscr{F}, \Pr)\). Independence of random elements is denoted by \(\raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}}\) and expectations computed against the probability measure \(\Pr\) are denoted by \(\mathbb{E} [\cdot]\). We will seldom refer to the underlying probability space but define it nonetheless for clarity. Furthermore, absent any subscripts, \(\mathbb{E} [\cdot]\) will also denote expected values assuming the “true” distribution of underlying data to be defined later on. The \(d\)-dimensional multivariate normal distribution with mean vector \(\mathbf{m} \in \mathbb{R}^{d}\) and covariance matrix \(\Sigma \in \mathbb{R}^{d \times d}\) is denoted \(\mathcal{N} \left( \mathbf{m}, \Sigma \right)\). We denote the matrix/vector transpose operation by \((\cdot)^{\prime}\). The natural numbers are denoted by \(\mathbb{N} = \{1, 2, 3, 4, \dots\}\) and for a given \(\mathcal{S} \in \mathbb{N}\), denote \(\mathbb{N}_{\mathcal{S}} = \{1, \dots, \mathcal{S}\} = \{s \in \mathbb{N} : 1 \leq s \leq \mathcal{S}\}\).
We employ a standard binary treatment potential outcomes framework and assume that samples are drawn from an infinite super-population. The population of interest is described by a random vector \(W\) taking values in \(\mathbb{R}^{2 + k}\) for some \(k \in \mathbb{N}\), with \(W^{\prime} = \left(Y (0), Y (1), Z^{\prime} \right)\). For each \(a \in \{0, 1\}\), \(Y (a)\) is a scalar random variable representing the potential outcome from receiving treatment \(a\). The treatment \(a = 1\) can be interpreted as an “innovation” whereas \(a = 0\) is a “status quo” or “control”. \(Z\) is a random \(\mathbb{R}^{k}\)-vector of baseline covariates. We denote the true distribution of \(W\) by \(Q_{0}\). Prior to the assignment of treatment, only baseline covariates are observable. Furthermore, after treatment has been assigned, only the outcome associated with received treatment is observable so that \((Y (0), Y (1))\) is not jointly observable. We maintain the following assumptions about the population distribution \(Q_{0}\).
The dominance assumption is made for mathematical convenience and is standard in the literature on semiparametric efficiency. Additionally, the choices of dominating measures \(\mu_{0}, \mu_{1}\) and \(\mu_{Z}\) are irrelevant and do not affect the derivation of the efficiency bound. The finite second moment condition is used for deriving Gaussian limiting distributions in later sections. The family of distributions \(\mathbf{Q}\) is a nonparametric family since it is infinite dimensional. The parameter of interest is the average treatment effect (ATE), \(\beta_{\ast} : \mathbf{Q} \to \mathbb{R}\) defined by
The true value of the ATE will be denoted \(\beta_{0} = \beta_{\ast} \left( Q_{0} \right)\).
In both observational and experimental studies on the average effect of treatment, observed outcomes result from the receipt of treatment. That is, observed outcomes are given by \(Y\) defined by
where \(A\) is a Bernoulli random variable denoting the treatment received. In experimental settings, the conditional distribution of \(A\) given the baseline covariates \(Z\) is assumed to be fully known to the experimenter. The target observed population in the special case of a CAR experiment is the distribution of the vector \(X^{\prime} = \left( Y, A, Z^{\prime} \right)\) which satisfies \th(ref) below.
The function \(\mathbb{S}\) is used to construct a (measurable) finite partition of the support of the covariates \(Z\) and thus, \(\mathcal{S}\) is an upper bound on the number of strata. The condition in (ref) requires that assignment to treatment be exogenous to both potential outcomes and the remaining variation in covariates given stratum labels. The condition in (ref) requires that treatment assignment proportions correspond to the pre-specified target probabilities in \(\pi\). In finite samples, the covariate adaptive randomization schemes that will be described in the next subsection all provide different ways to target the experiment in \th(ref).
In this subsection, we describe the assumptions maintained for the sampling process that produces observed data in a CAR experiment.
\th(ref) states that potential outcome and covariate values for observations in the sample are drawn at random from the population distribution \(Q_{0}\). That is, if both potential outcomes and covariates were all observable, the experimenter would have a representative sample from the underlying population. This assumption is maintained in the recent literature on CAR experiments as well as optimal experimental designs, for instance in 2018bugniInferenceCovariateAdaptiveRandomization, 2019bugniInferenceCovariateAdaptiveRandomization, 2022tabord-meehanStratificationTreesAdaptive, 2022baiOptimalityMatchedPairDesigns and 2021cytrynbaumDesigningRepresentativeBalanced. Given the baseline covariates and corresponding strata, the experimenter chooses a vector of treatment assignments \(\mathbf{A}_{n}^{\prime} = \left( A_{n 1}, \dots, A_{nn} \right)\) which is a random vector supported in \({\{0, 1\}}^{n}\). The experimenter has full control over the distribution of treatment assignments and treatment assignments do not have to be i.i.d. We introduce the following notation for convenience. For each \(s \in \mathbb{N}_{\mathcal{S}}\) and \(a \in \{0, 1\}\) the stratum size and stratum treatment group size are respectively
We maintain the following assumptions about the observed data in a covariate adaptive randomized experiment.
Denote the distribution of \(\mathbf{X}_{n}\) by \(P_{0, n}\). Note that \(P_{0, n}\) is determined by the population distribution \(Q_{0}\), the stratification scheme \(\mathbb{S}\), equation (ref), and the choice of randomization scheme. \th(ref) places the same set of restrictions on \(P_{0, n}\) as in 2019bugniInferenceCovariateAdaptiveRandomization on the relationship between the randomization scheme, the underlying population and the strata. \th(ref) (ref) requires that the observed outcome for a given observation is exactly the potential outcome associated with the assigned treatment (as in equation (ref) for the target experiment population \(\mathbf{P}\)). \th(ref) (ref) requires that the treatment assignment be ignorable (or exogenous) given the strata. In particular, the assignment scheme can only be a function of the stratum labels and a randomization device exogenous to the information contained in the potential outcomes and the baseline covariates beyond that afforded by the strata. This assumption is analogous to (ref) in \th(ref). Its main use is to guarantee identification of the ATE within each stratum and it is additionally a sufficient condition to identify the overall ATE. \th(ref) (ref) requires that the randomization procedure be fully known to the experimenter. This assumption plays a role in providing a simple characterization of the joint distribution of the observed data and allows us to avoid mathematical complications when talking about parametric sub-models during the discussion of semiparametric efficiency. Finally, \th(ref) (ref) requires the randomization procedure to reach the target treatment proportion within a given stratum at least asymptotically in the sense of convergence in probability. This is analogous to (ref) in \th(ref). There are a number of examples of randomization schemes which will satisfy the requirements imposed by \th(ref). We provide two examples of such randomization schemes.
Both examples satisfy the requirement imposed by \th(ref) (ref) for approaching the stratum-specific target proportions, but they do so at different rates of convergence. The SSRA method provides convergence to the stratum-specific targets at a \(n^{- \frac{1}{2}}\) rate due to the Lindeberg-L\'evy Central Limit Theorem. That is,
It can be shown using the Strong Law of Large Numbers that SPBR achieves the same targeting property at a faster \(n^{- 1}\) rate, so that
When a randomization scheme achieves this faster convergence property we say that it achieves “strong balance”. Numerous other kinds of stratified randomization techniques have been developed and analyzed in the literature on randomized control trials --- for instance Efron's biased coin design and Wei's urn --- see 2018bugniInferenceCovariateAdaptiveRandomization and references therein for detailed descriptions. Furthermore, the treatment assignments are not required to be i.i.d. and hence, the observed outcomes are all potentially non-i.i.d. For instance, with the SPBR procedure illustrated in \th(ref), assignments are independent across strata, but they exhibit dependence within strata. Indeed, most (if not all) assignment schemes that achieve strong balance will result in treatment assignments and observed outcomes that are non-i.i.d. However, we will show that the same efficiency bound as that for the i.i.d. sampling scheme in \th(ref) holds for all CAR schemes that satisfy \th(ref). Finally while the assumptions are satisfied by a fairly broad class of randomization schemes, there are ones that violate them that are used in practice, e.g. the minimization method of 1975pocockSequentialTreatmentAssignment. Developing efficiency theory for these methods may be of interest but they are not covered in this work.
In this section, we first state the semiparametric efficiency bound for the ATE and discuss it briefly. We also discuss the derivation of the bound in two subsections. As stated in the introduction, of the efficiency bounds established in the literature, the closest one to our setting is that of 1998hahnRolePropensityScore. This is derived under i.i.d. observational data under the assumption of ignorable treatment assignment conditional on observable covariates (1983rosenbaumCentralRolePropensity). The 1998hahnRolePropensityScore bound does not apply immediately to our setting since the treatment assignment rule is allowed to depend on the entire profile of sample strata. For instance, it is not clear a priori if the choice of assignment mechanism will affect the efficiency bound since it can clearly influence limit behavior of estimators as in 2018bugniInferenceCovariateAdaptiveRandomization, 2019bugniInferenceCovariateAdaptiveRandomization. Furthermore, even when treatment assignments are i.i.d. in a CAR context, the covariates being conditioned on during assignment are the strata, so it is again not immediate from the bound in 1998hahnRolePropensityScore what role the additional baseline covariates can play in providing efficiency gains. The SPEB derived here provides an answer to these questions. Before stating our bound, let \(m \left( a, z; Q \right) = \mathbb{E}_{Q} [Y (a) | Z = z]\) for each \(a \in \{0, 1\}\), and for brevity, let \(m_{\ast} (a, z) = m \left( a, z; Q_{0} \right)\) Furthermore, define
The following theorem establishes that \(\mathbb{V}_{\ast}\) above is the SPEB.
The proof of \th(ref) provides formal justification for \(\mathbb{V}_{\ast}\) as an efficiency bound by appealing to extensions of the famed convolution (1970hajekCharacterizationLimitingDistributions) and local asymptotic minimax theorems (1972hajekLocalAsymptoticMinimax) to semiparametric problems. The efficiency bound in (ref) is that of 1998hahnRolePropensityScore for i.i.d. data, if treatment is assigned conditional on baseline covariates according to the propensity score \(\pi (\mathbb{S} (\cdot))\). The intuition for this is that even though randomization happens conditional on strata, the baseline covariates are independent to treatment assignment after conditioning on strata (\th(ref) (ref)). Thus, the minimum possible limit variance in estimating the ATE can be achieved by utilizing any additional information the baseline covariates may have about potential outcomes. This can be better understood by considering what happens if only information from the strata are used during estimation. For instance, consider the fully saturated regression estimator of the ATE from 2019bugniInferenceCovariateAdaptiveRandomization. This estimator is constructed from regressing observed outcomes \(Y_{n i}\) on the stratum indicators and their interactions with treatment status \(A_{n i}\). The coefficients from this regression are combined to then form the ATE estimate. The resulting estimator has a “weighted difference in means” form:
2019bugniInferenceCovariateAdaptiveRandomization show that \(\widehat{\beta}_{n, \mathrm{SAT}}\) is asymptotically normal with mean-zero and limit variance given by
Our results thus establish that the fully saturated regression estimator of the ATE achieves the SPEB among all estimators that only use information from the stratum labels. This follows from comparing (ref) and (ref) and noting that we have \(\mathbb{V}_{\mathrm{SAT}} = \mathbb{V}_{\ast}\) if \(Z = \mathbb{S} (Z)\) (up to information preserving relabelling of strata). The latter condition amounts to only using information given in stratum labels. An additional implication is that when conditional means or conditional variances of the potential outcomes given baseline covariates exhibit variation within strata, \(\mathbb{V}_{\ast}\) can be a strict improvement over \(\mathbb{V}_{\mathrm{SAT}}\) - an observation also made by 1998hahnRolePropensityScore. This is confirmed by Monte Carlo simulation evidence in Section (ref) in which \(\mathbb{V}_{\ast}\) can be approximately half of \(\mathbb{V}_{\mathrm{SAT}}\) in some simulation designs (i.e. there is up to a 50% possible reduction in asymptotic variance).
We provide a few additional comments on the SPEB in (ref). Note that the choice of assignment mechanism only affects the bound through the choice of the target assignment proportions in \(\pi (\cdot)\). All else held equal, the SPBR assignment mechanism in \th(ref), or any other procedure satisfying \th(ref), will produce the same SPEB as i.i.d. draws from \(\mathbf{P}\) (\th(ref)). Furthermore, 1998hahnRolePropensityScore notes that knowledge of the propensity score is ancillary to the SPEB for the ATE. In the context of CAR, we will show that this knowledge is useful for achieving the SPEB during estimation. An important intermediate product of our derivation for this purpose is the efficient influence function, \(\varphi_{0}\), which is defined by
\(\varphi_{0}\) is called the efficient influence function because \(\mathbb{E} \left[ \varphi_{0} (Y, A, Z) \right] = 0\) and its variance is the SPEB in (ref), i.e. \(\mathbb{E} \left[ \varphi_{0} (Y, A, Z)^{2} \right] = \mathbb{V}_{\ast}\). This function is a key ingredient to achieving the efficiency bound as illustrated in Section (ref). In the subsequent two subsections, we provide some additional discussion of how (ref) is derived, but mathematical details are left to the appendix.
We establish here that the joint distribution of the observed data \(\mathbf{X}_{n}\), \(P_{0, n}\), has a product structure even though outcomes and treatment assignments are potentially non-i.i.d. This product structure will then imply that \(P_{0, n}\) has a density that also has a product form. To that end, we first define a dominating measure for a single observation. For Borel sets \(\mathcal{Y} \subseteq \mathbb{R}\) and \(\mathcal{Z} \subseteq \mathbb{R}^{k}\) and for \(\mathcal{A} \subseteq \{0, 1\}\), define the measures \(\nu_{a}\) with \(a \in \{0, 1\}\) and \(\nu\) by
Let \(q_{a} (y , z; Q)\) denote the marginal density of \((Y (a), Z)\) if the underlying population has distribution \(Q \in \mathbf{Q}\). Furthermore, let \(P_{n} (\cdot; Q)\) denote the distribution of the observed data \(\mathbf{X}_{n}\) if \(\mathbf{W}_{n}\) is an i.i.d. sample from \(Q \in \mathbf{Q}\). Under this notation, we have \(P_{0, n} \equiv P_{n} \left( \cdot; Q_{0} \right)\).
The main consequence of \th(ref) is as follows. Let \(\mathcal{P} = \{\mathbf{P}_{n} : n \in \mathbb{N}\}\) be the sequence of families of distributions for the observed data \(\mathbf{X}_{n}\) such that for each \(n \in \mathbb{N}\), every member of \(\mathbf{P}_{n}\) is determined by an i.i.d. sample from some distribution in \(\mathbf{Q}\) (as defined in \th(ref)), the observed outcome equation (ref) and a treatment assignment mechanism that satisfies \th(ref). Any member of \(\mathbf{P}_{n}\) must then satisfy the product structure in (ref) and (ref). Furthermore, different covariate adaptive randomization schemes give rise to different sequences of families \(\mathcal{P}\) that only vary according to the choice of the sequence of conditional treatment assignment distributions \(\left\{ \alpha_{n} (\cdot) : n \in \mathbb{N} \right\}\).
In this subsection, we first introduce and discuss definitions of parametric submodels and regular parametric submodels which are essential concepts to the derivation of the semiparametric efficiency bound. It is then shown in \th(ref) that the log likelihood ratios in regular parametric submodels under \(1 / \sqrt{n}\) local alternatives exhibit the local asymptotic normality (LAN) property of 1960lecamLocallyAsymptoticallyNormal. An efficiency bound is established for parametric submodels (\th(ref)). Then, via a discussion on differentiability of the ATE in regular parametric submodels we provide an intuitive description of how the parametric efficiency bound is extended to the semiparametric efficiency bound presented in \th(ref).
In \th(ref), condition (ref) requires that a parametric submodel be a subset of the overall semiparametric model. Condition (ref) requires that the parametrization \(\theta \mapsto p_{n} (\cdot; \theta)\) produces densities of the same form as those in \th(ref). The additional \(Z\)-marginal restriction ensures that both densities \(q_{0} (\cdot; \theta)\) and \(q_{1} (\cdot; \theta)\) can be derived from a common distribution \(Q_{\theta} \in \mathbf{Q}^{0}\). We do not require the submodel to specify the joint density of \((Y (0), Y (1), Z)\) since only the marginals with respect to \((Y (0), Z)\) and \((Y (1), Z)\) are identified by the randomized experiment (essentially due to (ref) in \th(ref)). Finally, (ref) requires the parametric submodel to contain the true distribution \(P_{0, n}\). Note that a parametric submodel does not parameterize the conditional mass function of the treatment assignments, \(\alpha_{n}\), since it is known and does not depend on any population unknowns. This is different to the case of observational data considered in 1998hahnRolePropensityScore where the propensity score is unknown and has to be parameterized.
The idea behind the use of parametric submodels is as follows. Since \(\Theta\) (and therefore \(\mathbf{P}_{n}\)) in \th(ref) is finite dimensional, estimation of the ATE restricted to \(\mathcal{P}^{0}\) cannot be more difficult than estimation within \(\mathcal{P}\). One can imagine estimation of the ATE by first estimating the nuisance parameter \(\theta_{0}\) by maximum likelihood and then forming a plug-in estimate of the ATE. If \(\mathcal{P}^{0}\) is a parametric submodel with a well-defined efficiency bound, any semiparametric estimator for the ATE that is consistent and asymptotically normal under the distributions in \(\mathcal{P}\) cannot have asymptotic variance lower than the efficiency bound under \(\mathcal{P}^{0}\). Taking the supremum over all parametric submodels, we get a variance lower bound for the semiparametric model. The statement about parametric submodels having well-defined efficiency bounds requires some qualification. Establishing efficiency requires restricting our analysis to regular parametric submodels (in which efficiency bounds are well-defined) and to regular estimators (to rule out the phenomenon of super-efficiency). We provide the definition of a regular parametric submodel in our case below. The definition of a regular estimator can be found in 1998vandervaartAsymptoticStatistics or 1996vandervaartWeakConvergenceEmpirical.
The log-likelihood in any parametric submodel is
When the parametric submodel is regular, the associated score function at the truth, \(\theta_{0}\), is
Note that the score in (ref) is expressed in “root density” terms since it is technically defined by the quadratic mean differentiability assumption in \th(ref).\footnote{When the parametric submodels are such that the densities are continuously differentiable with respect to \(\theta\), then the usual log-derivative form and (ref) coincide. That is, we can write \(\dot{\ell}_{a} (\cdot) = \nabla_{\theta} \left[ \log q_{a} \right] \left( \cdot; \theta_{0} \right)\).} Let \(X^{\prime} = (Y, A, Z)\) be as in the CAR target experiment defined by \th(ref). Under CAR, the associated Fisher information matrix will be shown to be
Note that the information matrix, \(\mathcal{I}\) in (ref) is related to the marginal information matrices \(\mathcal{I}_{a}\) for \(a \in \{0, 1\}\) in (ref) but adjusts these for stratification. Furthermore the score \(\dot{\ell}\) and information matrix \(\mathcal{I}\) both depend on the choice of parametric submodel since the quadratic mean derivatives \(D_{a}\) depend on the choice of parametric submodel. \th(ref) below shows that the main consequence of \th(ref) is that the likelihood ratio in any regular parametric submodel has the usual quadratic expansion and exhibits the local asymptotic normality (LAN) property of 1960lecamLocallyAsymptoticallyNormal under general CAR procedures.
The proof of \th(ref) in the appendix of this paper uses the tools developed 2018bugniInferenceCovariateAdaptiveRandomization and 2019bugniInferenceCovariateAdaptiveRandomization to adapt the proof of Proposition 2.1.2 in the appendix of 1998bickelEfficientAdaptiveEstimation to the context of CAR. In particular, a partial-sums empirical process argument is used to establish conclusion (ref) in \th(ref). Using similar arguments, the sample information matrix is shown to converge in probability to \(\mathcal{I}\). The remainder \(R_{n} \left( \theta_{0}, t \right)\) summarizes both the error replacing the sample information matrix with its population (or limiting) analogue as well as the error inherent in a second-order Taylor expansion. The latter is shown to be negligible as well. Under the LAN phenomenon, \th(ref) below establishes an efficiency result for estimation of the nuisance parameter \(\theta_{0}\).
\th(ref) establishes that the usual Cram\'er-Rao parametric information bound holds for estimation of \(\theta_{0}\) in regular parametric submodels. This is true even under general CAR procedures that may result in dependence within observed data. \th(ref) is a direct consequence of the LAN property established in \th(ref). In \th(ref), conclusion (ref) follows from the celebrated convolution theorem of 1970hajekCharacterizationLimitingDistributions and conclusion (ref) follows from the local asymptotic minimax theorem of 1972hajekLocalAsymptoticMinimax. These are standard tools used to justify the parametric Cram\'er-Rao bound. Note that in \th(ref) the random vectors \(G\) and \(U\) are again specific to the parametric submodel chosen and in the case of \(U\), also specific to the choice of estimator sequence.
\th(ref) establishes an efficiency bound for estimation of the nuisance parameter \(\theta\) whereas the content of \th(ref) is about estimation of the ATE, \(\beta\) in (ref). To relate \th(ref) to a version of \th(ref), we have to first show that \(\beta\) is a pathwise differentiable parameter (see 1998bickelEfficientAdaptiveEstimation or 1990neweySemiparametricEfficiencyBounds). We do this by considering regular parametric submodels of the target CAR experiment, \(\mathbf{P}\), as defined by \th(ref) and \th(ref). For a regular parametric submodel \(\mathbf{P}_{0} = \left\{ P_{\theta} : \theta \in \Theta \right\}\) of \(\mathbf{P}\), the average treatment effect parameter is
Adapting the pathwise differentiability arguments in 1998hahnRolePropensityScore to the context of a CAR experiment, the derivative (gradient) of \(\gamma (\theta)\) at \(\theta_{0}\) in any given regular parametric submodel can be written as
where \(\varphi_{0}\) is the efficient influence function from (ref). The full derivation of this adaptation is included the appendix for completeness (\th(ref)). Combining \th(ref) and (ref), it follows via the usual “delta method style” argument that the information bound for estimating the ATE in the regular parametric submodel \(\mathcal{P}^{0}\) is
The variant of Cauchy-Schwarz inequality for random vectors established in 1999tripathiMatrixExtensionCauchySchwarz can be used to conclude that \(\mathbb{V} \left( \mathcal{P}^{0} \right) \leq \mathbb{E} \left[\varphi_{0} (Y, A, Z)^{2} \right] = \mathbb{V}_{\ast}\). This establishes that \(\mathbb{V}_{\ast}\) in (ref) is an upper bound over all parametric information bounds for estimation of the ATE. However, one still needs to show that \(\mathbb{V}_{\ast}\) is in fact a supremum. This final step is done by establishing that \(\varphi_{0}\) can be approximated by the score functions of regular parametric submodels.
The efficiency bound derived in \th(ref) is a lower bound on asymptotic variances of regular semiparametric estimators for the ATE. However, the question remains as to whether it can be achieved. Typically, the conditions under which estimators can achieve semiparametric efficiency bounds are stronger than those required to derive the bounds. 1990ritovAchievingInformationBounds provide examples for cases when an efficiency bound can be established and shown to be finite as well as non-singular, but even consistent estimation (let alone efficient) is impossible without restricting the semiparametric model. One of their examples is the partially linear model (1986engleSemiparametricEstimatesRelation, 1988robinsonRootNConsistent). The restrictions they require are in the form of additional smoothness conditions for the nonparametric nuisance parameters that appear in the efficient influence function. These smoothness conditions are to ensure that the nonparametric nuisance parameters can be estimated consistently with a rate of convergence of at least \(n^{- 1 / 4}\) (see also 1994neweyAsymptoticVarianceSemiparametric). Furthermore, achieving \(n^{- 1 / 4}\)-consistency for a non-parametric estimator becomes more difficult when there are many covariates due to the curse of dimensionality, and thus higher order (i.e. stronger) smoothness conditions are required.
In the context of estimating average treatment effects with the assumption of selection on observables (or ignorability conditional on covariates) and i.i.d. data, 1998hahnRolePropensityScore proposes nonparametric imputation estimators. These estimators impute the potential outcome \(Y_{i} (a)\) whenever it is unobserved by using an estimate of the predicted value \(m \left( a, Z_{i} \right)\), say \(\widehat{m}_{n} \left( a, Z_{i} \right)\). The difference of the imputed potential outcomes is then averaged to estimate the ATE. 1998hahnRolePropensityScore suggests the use of series estimators for the propensity score \(\Pr (A = 1 | Z = z)\) and the conditional expectations \(\mathbb{E} [Y \mathbb{I} \{A = a\} | Z = z]\) which then get combined to construct estimators for \(m_{\ast} (a, \cdot)\). Similar estimators appear in various parts of the literature on semiparametric estimation of treatment effects under treatment ignorability - see 2004imbensNonparametricEstimationAverage for a review and examples with kernel estimators for the aforementioned propensity score and conditional expectations. The requirement of higher order smoothness conditions on the propensity score and the conditional expectations \(m_{\ast} (a, \cdot)\) are ubiquitous in these examples.
In our case, the propensity score does not need to be estimated since the target proportions by strata, \((\pi (1), \dots, \pi (\mathcal{S}))\), are chosen by the experimenter. Furthermore, we should be able to use sample treatment proportions by strata in place of true target proportions without any issue since these form a finite-dimensional additional nuisance parameter. This feature makes the efficient influence function linear in the nonparametric nuisance parameters \(m_{\ast} (a, \cdot)\). It is also straightforward to verify that \(\varphi\) has the Neyman orthogonality or local robustness property of 2018chernozhukovDoubleDebiasedMachine and 2020chernozhukovLocallyRobustSemiparametric with respect to estimation of \(m _{\ast}(a, \cdot)\). To see this, denote
so that the efficient influence function satisfies \(\varphi_{0} (\cdot) = \varphi \left( \cdot; m_{\ast} (0, \cdot), m_{\ast} (1, \cdot), \beta_{0} \right)\). If \(m_{\ast} (0, \cdot), m_{\ast} (1, \cdot)\) were known, then we could estimate \(\beta\) by solving the sample equivalent of the moment condition
The local robustness property stems from noting that at \(\beta_{0}\),
This local robustness property allows for \(n^{- \frac{1}{2}}\) consistent and efficient estimation of the ATE by using plug-in estimates of \(m_{\ast} (a, \cdot)\) under much weaker conditions. In particular, we do not require smoothness conditions for \(m_{\ast} (a, \cdot)\) to achieve the efficiency bound in (ref).
In the i.i.d. case (\th(ref)), an “ideal” efficient estimator of the ATE would be \(\widetilde{\beta}^{\ast}_{n} = \beta_{0} + n^{- 1} \sum_{i = 1}^{n} \varphi_{0} \left( Y_{n i}, A_{n i}, Z_{i} \right)\). Since adding \(b\) to \(\varphi\) in (ref) removes it as an argument, this estimator is
\(\widetilde{\beta}_{n}^{\ast}\) is the familiar AIPW estimator of 1994robinsEstimationRegressionCoefficients, 1995robinsAnalysisSemiparametricRegression, and 1999scharfsteinAdjustingNonignorableDrop if \(m_{\ast} (a, \cdot)\) were known. A feasible estimator must replace these with first stage estimates that can be computed from the sample. Since it will be convenient later on, we allow for the estimates of \(m_{\ast} (a, \cdot)\) to vary across observations. We will also allow for the experimenter to use estimated treatment proportions rather than the known true ones. That is, let \(\widehat{\pi}_{n} (s)\) be a consistent estimator for \(\pi (s)\) for each \(s \in \mathbb{N}_{\mathcal{S}}\). A feasible estimator for \(\beta_{0}\) would be
It is straightforward to show that \(\widetilde{\beta}^{\ast}_{n}\) achieves the efficiency bound \(\mathbb{V}_{\ast}\) in (ref) under the CAR procedures satisfying our assumptions. Our plug-in estimators \(\widehat{m}_{n i}\) will be cross-fitted estimators as described in 2018chernozhukovDoubleDebiasedMachine (see their DML2 estimators). We now provide conditions on the estimators \(\widehat{m}_{n i}\), and \(\widehat{\pi}_{n, a} (s)\) as well as the overall model under which the difference between \(\widehat{\beta}^{\ast}_{n}\) to \(\widetilde{\beta}^{\ast}_{n}\) in (ref) are asymptotically negligible after scaling by \(\sqrt{n}\). We start with assumptions on a sequence of nonparametric estimators \(\widehat{m} (n, a, \cdot; \cdot)\) from which \(\widehat{m}_{n i} (a, \cdot)\) will be constructed.
The estimator \(\widehat{m} (n, a \cdot; \cdot)\) is allowed to use the entire sample of potential outcomes for \(a\) as well as the covariates. Of course, all of these are not available - \(\widehat{m} (n, a \cdot; \cdot)\) can only be computed from the treatment group for treatment \(a\). This will be accounted for later on when we describe the cross-fitting procedure that leads to \(\widehat{m}_{n i}\). \th(ref) requires that \(\widehat{m}\) be \(L_{2}\) consistent for the true conditional expectation \(m (\cdot; Q)\) when the data are generated according to \(Q\) which belongs to an appropriately chosen subset \(\mathbf{Q} \left( \widehat{m} \right) \subseteq \mathbf{Q}\) which also contains \(Q_{0}\). Note that the \(L_{2}\) convergence rate of \(\widehat{m}\) can be arbitrarily slow. For our purposes, simple consistency of this kind is enough. However, \(\mathbf{Q} \left( \widehat{m} \right)\) can be formed by restrictions that impose a rate condition on \(\widehat{m}\). For instance, the experimenter might impose smoothness conditions required in existing literature for the particular \(\widehat{m}\) to obtain particular rates - e.g. a H\"older or Sobolev ball. If \(\widehat{m}\) is a series estimator, these can be found for example in 1997neweyConvergenceRatesAsymptotic and 2015belloniSomeNewAsymptotic. For nonlinear sieves such as neural networks, one can see for 1994barronApproximationEstimationBounds, 1999chenImprovedRatesAsymptotic, 2020schmidt-hieberNonparametricRegressionUsing and 2021farrellDeepNeuralNetworks. For kernel estimators with uniform rates, see 1996masryMultivariateLocalPolynomial and 2008hansenUniformConvergenceRates. Other restrictions may be functional form restrictions, e.g., random forests are known to be \(L_{2}\) consistent for additive models - see 2015scornetConsistencyRandomForests. When \(\mathbf{Q} \left( \widehat{m} \right) = \mathbf{Q}\), we will say that \(\widehat{m}\) is universally \(L_{2}\) consistent. This is a property enjoyed by the Nadaraya-Watson Kernel estimator - see for instance 1980devroyeDistributionFreeConsistencyResults and 1980spiegelmanConsistentWindowEstimation. Next, we state assumptions that describe the cross-fitting approach to be used here.
\th(ref) first requires the experimenter to split each treatment group within a given stratum into \(J\) folds in a manner that is independent to the overall sample using information only about the group size. These folds are further required to satisfy an “equal size” requirement specific to the treatment group within the stratum. For a given stratum, the two treatment groups corresponding to fold \(j\) are then combined which produces \(J\) folds for the overall stratum. We first split treatment groups by strata into folds to ensure that a given combined fold has units in both treatment and control groups. \th(ref) requires that the estimator \(\widehat{m}_{n i}\) for \(i\)\textsuperscript{th} observation be constructed from \(\widehat{m}\) in \th(ref) using observations in the same stratum as \(i\) but not in the same fold as \(i\). Finally, \th(ref) requires that the treatment proportion estimator \(\widehat{\pi}_{n} (s)\) will be either the true treatment proportions or the sample treatment proportions. For the latter, \th(ref) requires require consistency at a \(1 / \sqrt{n}\) rate. This corresponds to requiring that \(\alpha_{n} (\cdot | \cdot)\) in \th(ref) at least satisfy “weak balance” as defined in the paragraph after \th(ref). The key result is \th(ref) below.
\th(ref) first establishes in (ref) that the infeasible ideal estimator \(\widetilde{\beta}^{\ast}_{n}\) achieves the efficiency bound in (ref) under general CAR procedures. This implies that any estimator of the ATE that is asymptotically linear with influence function \(\varphi_{0}\) in (ref) achieves the semiparametric efficiency bound. \th(ref) then establishes in (ref) that \(\widehat{\beta}^{\ast}_{n}\) as defined by (ref) has this aforementioned property (under the additional assumption of this section). The only additional restrictions placed on the overall semiparametric model \(\mathbf{Q}\) are ones defining \(\mathbf{Q} \left( \widehat{m} \right)\) specific to the choice of estimator \(\widehat{m}\) in \th(ref). As a further corollary (\th(ref) below), we can show that \(\mathbf{Q} \left( \widehat{m} \right) = \mathbf{Q}\) when \(\widehat{m}\) is the Nadaraya-Watson kernel regression estimator. Note that the significance of \(\mathbf{Q} \left( \widehat{m} \right) = \mathbf{Q}\) is that efficient estimation is possible (pointwise) over the whole of \(\mathbf{Q}\). That is, the efficiency bound can be achieved under the same conditions used for its derivation. This is a novel result since the overall semiparametric problem is non-trivial and involves estimation of (possibly) infinite-dimensional conditional means. We now first describe the Nadaraya-Watson estimator, state assumptions on its tuning parameters and then provide the result. In what follows, we use the normalization that \(0/0 = 0\). Let \(\widehat{m}\) in \th(ref) be defined by
\th(ref) below states the assumptions required for the bandwidth sequence \(\left\{ h_{n} : n \in \mathbb{N} \right\}\) and the kernel function \(\kappa\).
\th(ref) maintains exactly the same conditions required by 1980devroyeDistributionFreeConsistencyResults and 1980spiegelmanConsistentWindowEstimation for universal \(L_{2}\) consistency. Note that \th(ref) (ref) requires the kernel function \(\kappa\) to be non-negative, to have compact support and to be strictly positive in a closed neighborhood of the origin with non-empty interior. The following result shows that efficient estimation is possible under exactly the same conditions required to derive the efficiency bound.
The only requirement for the phenomenon in \th(ref) to occur is the universal consistency requirement, i.e. that \(\mathbf{Q} \left( \widehat{m} \right) = \mathbf{Q}\). This is not special to the Nadaraya-Watson estimator - one can replace this with a nearest neighbor estimator and replace the conditions in \th(ref) with those of 1977stoneConsistentNonparametricRegression. As for local polynomial kernel estimators, these are not universally \(L_{2}\) consistent in their typical form (see 2002gyoerfiDistributionFreeTheory), but adjusted versions can have the universal \(L_{2}\) consistency property (see 2002kohlerUniversalConsistencyLocal). These adjustments introduce a number of additional tuning parameters beyond the bandwidth. For universal \(L_{2}\) consistency conditions for series estimators, see 2002gyoerfiDistributionFreeTheory.
We now provide a brief sketch of the main arguments that lead to (ref) in \th(ref). The difference between \(\sqrt{n} \left( \widehat{\beta}_{n}^{\ast} - \beta_{0} \right)\) and \(\sqrt{n} \left( \widetilde{\beta}_{n}^{\ast} - \beta_{0} \right)\) and can be summarized as
In the above, \(\pi_{a} (s) = \pi (s)^{a} [1 - \pi (s)]^{1 - a}\) and \(\widehat{\pi}_{n, a} (s)\) is defined analogously for each \(s \in \mathbb{N}_{\mathcal{S}}\) and \(a \in \{0, 1\}\). Showing (ref) is a matter of showing that \(R_{a, n} = o_{\mathrm{p}} (1)\) and \(\widetilde{R}_{a, n} = o_{\mathrm{p}} (1)\). The remainder term \(\widetilde{R}_{a, n}\) is straightforward to deal with - a central limit theorem applies to
and the difference between estimated and true inverse treatment probabilities is constant across observations within a given stratum and converging to zero. For \(R_{a, n}\), we deal with these by first splitting them into a sum of \(\mathcal{S} \cdot J\) terms each corresponding to a given stratum and a given fold within that stratum. Each of these terms can be treated separately since we are trying to prove convergence in probability to zero. The argument then proceeds by noting that within the stratum-fold remainder each component of the form \(\widehat{\pi}_{n, a} (s)^{- 1} \left( \mathbb{I} \left\{ A_{n i} = a \right\} - \widehat{\pi}_{n, a} (s) \right)\) is “approximately” mean zero and has finite variance. Furthermore, for the term \(\widehat{m}_{n i} \left(a, Z_{i} \right) - m_{\ast} \left( a, Z_{i} \right)\) has second moment tending to zero. By careful use of the independence between the evaluation point \(Z_{i}\) and the estimation sample (observations outside the corresponding fold), the stratum-fold specific version of \(R_{a, n}\) is shown to be \(o_{\mathrm{p}} (1)\) via Chebychev's inequality. Since the overall \(R_{a, n}\) is the sum of its stratum-fold specific counterparts and there are \(\mathcal{S} \cdot J\) (i.e. finitely many) of these, the conclusion follows from Slutsky's theorem. A rigorous version of these arguments is left to the appendix.
The key finding here is that knowledge of the propensity score and its finite support structure reduces the burden for efficient estimation considerably. The fact that knowledge of the propensity score can reduce this burden was also noticed by 2021aronowNonparametricIdentificationNot and 2018rotheFlexibleCovariateAdjustments. 2021aronowNonparametricIdentificationNot show that potentially non-linear parametric regression adjustments can improve estimation accuracy, generalizing 2001yangEfficiencyStudyEstimators and 2013linAgnosticNotesRegression to the nonlinear case. 2018rotheFlexibleCovariateAdjustments shows that semiparametrically efficient estimation is possible under a fairly broad class of estimators using Donsker conditions. 2018rotheFlexibleCovariateAdjustments also shows that efficient estimation is possible under slightly weaker conditions by using a leave-one-out locally linear kernel regression estimator by leveraging the results of 2019rothePropertiesDoublyRobust. However, both still require dimension-dependent smoothness conditions on the conditional means to be estimated. To the best of our knowledge, existing results using nonparametric estimators in this context all require existence of densities for baseline covariates that are bounded away from zero on their support and higher order differentiability on conditional means and possibly also densities. Our results do not require any of these for efficient estimation.
It should be noted that while \(\widehat{\beta}^{\ast}_{n}\) in (ref) achieves the efficiency bound a few issues prevent it from being usable immediately in the context of statistical inference. In particular, we have not provided consistent estimators of the asymptotic variance \(\mathbb{V}_{\ast}\) in (ref). It may be possible to construct these directly from the sample and the nonparametric estimators \(\widehat{m}_{n i}\). Another possibility might be a bootstrap procedure to produce confidence intervals for hypothesis tests. Another issue we have not addressed is data-based choice of tuning parameters for the estimators \(\widehat{m}_{n i}\) when these are based on the Nadaraya-Watson kernel estimator. Since our main questions are the efficiency bound and the conditions under which it can be achieved, we leave the construction of consistent inference procedures from \(\widehat{\beta}^{\ast}_{n}\) and data-based tuning parameter selection for future research.
In this section, we examine the finite sample performance of feasible efficient estimator \(\widehat{\beta}_{n}^{\ast}\) from (ref). We start by describing the data generating processes (DGPs) for our simulations. We consider values of \(n \in \{500, 1000, 2000, 4000, 8000\}\) and \(k \in \{1, 5\}\). The DGPs are all of the form
In all cases, we have \(Z_{i} \sim \mathrm{Uniform} \left([-1, 1]^{k} \right)\) and \(\varepsilon_{i} \sim \mathcal{N} (0, 1)\). For each DGP, we conduct 5000 Monte Carlo simulations. In each simulation round, we compute \(\widetilde{\beta}_{n}^{\ast}\) in (ref) and \(\widehat{\beta}_{n}^{\ast}\) in (ref) as well as additional estimators of the ATE to be described below. We report mean squared errors as well as biases of these estimators relative to the true value of the ATE. For stratification, we stratify on the basis of the first component of the covariates \(Z_{i}\). We consider \(\mathcal{S} \in \{5, 20\}\) and for each value of \(\mathcal{S}\), strata are constructed by splitting the interval \([-1, 1]\) into adjacent segments of length \(2 / \mathcal{S}\). We report results with constant assignment proportions across all strata as well as proportions that vary by strata. We present results with SPBR here, and defer results with SSRA to the appendix. Inspecting both, the reader can see that the choice of assignment mechanism has no visible impact on the performance of any of the estimators considered here. For constant proportions, we set assignment proportions to 1/2 and for varying proportions we use
We use the boldfaced \(\bm{\pi}\) here to distinguish assignment probabilities from the mathematical constant \(\pi = 3.14159\dots\), since our choice of conditional mean functions involve the latter through trigonometric functions. We consider four different DGP's by choice of \(k\), \(m_{\ast} (a, \cdot)\) and \(\sigma (a, \cdot)\). In what follows, DGP's 1 and 2 have \(k = 1\) whereas DGP's 3 and 4 have \(k = 5\). In all DGP's, the true value of the ATE is \(\beta_{0} = 0\). This is done for convenience.
In both DGP 1 and DGP 3, \(m_{\ast} (1, \cdot)\) and \(m_{\ast} (0, \cdot)\) are analytic and exhibit considerable non-linear variation within strata with either \(\mathcal{S} = 5\) or \(\mathcal{S} = 20\). In DGP 2, \(m_{\ast} (1, \cdot)\) and \(m_{\ast} (0, \cdot)\) have 19 discontinuities and are exactly fully saturated regressions with \(\mathcal{S} = 20\) but not \(\mathcal{S} = 5\). DGP 4 adds a jump discontinuity at \(z_{1} = 0\) to the conditional mean functions of DGP 3.
For comparison, we also examine the finite sample performance of the infeasible efficient estimator \(\widetilde{\beta}_{n}^{\ast}\), the fully saturated regression estimator of the ATE \(\widehat{\beta}_{n, \mathrm{SAT}}\) in (ref), as well as the imputation estimator of 1998hahnRolePropensityScore using the cross-fitted Nadaraya-Watson estimator:
The estimator above is an imputation estimator since it imputes the treatment \(a\) potential outcome for observation \(i\) with the predicted value \(\widehat{m}_{n i} \left( a, Z_{i} \right)\) whenever that potential outcome is unobserved. This estimator is also asymptotically linear and with influence function \(\varphi_{0}\) when the functions \(m_{\ast} (a, \cdot)\) are sufficiently smooth - see for instance 1994chengNonparametricEstimationMean. For our nonparametric estimators \(\widehat{m}_{n, i} \left( a, \cdot \right)\), we use bandwidths of the form \(c_{k} n^{- 1 / (4 + k)}\) and the uniform kernel so that \th(ref) is satisfied. For DGP 1 and DGP 3, the bandwidths are rate-optimal in terms of integrated mean-squared error (see 1982stoneOptimalGlobalRates) since the conditional mean functions \(m_{\ast} (a, \cdot)\) are twice continuously differentiable. For \(k = 1\), we set \(c_{k} = 1 / \sqrt{3}\) and for \(k = 5\), we set \(c_{k} = 3\). In our results, estimated mean squared errors for \(\widetilde{\beta}_{n}^{\ast}\) act as an estimate of the semiparametric efficiency bound. The comparison between \(\widehat{\beta}_{n}^{\ast}\) and \(\widetilde{\beta}_{n}^{\ast}\) illustrates how far from optimal the feasible estimator can be due to noise in the estimation of the conditional mean functions. Following the discussion in Section (ref), the fully saturated regression estimator \(\widehat{\beta}_{n, \mathrm{SAT}}\) in (ref) achieves the semiparametric efficiency bound among all estimators that use only information contained in the strata. This is \(\mathbb{V}_{\mathrm{SAT}}\), defined in (ref). Thus, comparing \(\widehat{\beta}_{n, \mathrm{SAT}}\) to \(\widetilde{\beta}_{n}^{\ast}\) illustrates the magnitude of efficiency loss from ignoring information from the covariates (except in DGP 2 with \(\mathcal{S} = 20\)). Additionally, comparing \(\widehat{\beta}_{n, \mathrm{SAT}}\) to \(\widehat{\beta}_{n}^{\ast}\) illustrates the possible tradeoff between finite sample and asymptotic accuracy when estimates of the nonparametric components are potentially far from the truth. The additional nonparametric estimator \(\widehat{\beta}_{n, \mathrm{IMP}}\) has the same influence function as \(\widehat{\beta}_{n}^{\ast}\) but does not directly use the structure of the influence function. Thus, comparisons between these estimators illustrates the debiasing effect of using the efficient influence function directly.
Tables (ref) and (ref) present results from simulations with constant target treatment proportions across strata with \(\mathcal{S} = 5\) and \(\mathcal{S} = 20\) respectively. Tables (ref) and (ref) present results with varying target treatment proportions across strata with \(\mathcal{S} = 5\) and \(\mathcal{S} = 20\) respectively. Some common patterns appear across examples. In the univariate case (DGP's 1 and 2), \(\widehat{\beta}^{\ast}_{n}\) is fairly close to the infeasible estimator \(\widetilde{\beta}^{\ast}_{n}\) in terms of mean squared error. Furthermore, while \(\widehat{\beta}_{n, \mathrm{SAT}}\) is far from optimal under the nonlinearities in DGP 1 and \(\mathcal{S} = 5\), the distance is greatly reduced by using finer strata (i.e. \(\mathcal{S} = 20\)). However, in higher dimensions, both estimators suffer albeit in different ways. Under DGP's 3 and 4 the feasible optimal estimator \(\widehat{\beta}^{\ast}_{n}\), despite still being consistent and asymptotically unbiased, is much slower in its tendency towards the SPEB. This is mainly due to the curse of dimensionality. In simulations with \(n \in \{10000, 20000\}\) for instance, \(\widehat{\beta}^{\ast}_{n}\) is closer to achieving the SPEB, but these sample sizes may be unrealistic for most experiments. As we can see, \(\widehat{\beta}_{n}^{\ast}\) is also sensitive to the number of strata - having more strata reduces the effective estimation sample used for estimation of the conditional mean. This is particularly stark in the examples with varying treatment proportions - in these, some treatment groups are going to have small sample sizes within strata by design and the aforementioned phenomenon is exacerbated. In the multivariate case, increasing the number of strata does not help reduce the MSE of \(\widehat{\beta}_{n, \mathrm{SAT}}\) since stratification is done on the basis of a single component of a continuously distributed covariate. On the other hand, though its performance in some instances (e.g. (ref), DGP 3 and 4) leaves much to be desired, in most cases with DGP 3 and 4, \(\widehat{\beta}^{\ast}_{n}\) picks up on the nonlinear variation in the conditional means across all dimensions. To summarize the comparison between \(\widehat{\beta}^{\ast}_{n}\) and \(\widehat{\beta}_{n, \mathrm{SAT}}\) we find that across all simulation designs, \(\widehat{\beta}^{\ast}_{n}\) is on average 13% more efficient (averaging the quotient of the MSEs). The maximal gain from using \(\widehat{\beta}^{\ast}_{n}\) is a 40% reduction in MSE - Table (ref), DGP 4 with \(n = 8000\).
For \(\widehat{\beta}_{n, \mathrm{IMP}}\), while it seems to perform well under \(k = 1\) and \(\mathcal{S} = 20\) the impact of the curse of dimensionality and its interaction with the lack of debiasing is quite stark as can be seen from the other cases. Furthermore, the need for undersmoothing to reduce bias for this estimator is apparent since it seems to sometimes do worse with more effective observations. For instance, comparing results for DGP 1 in Tables (ref) and (ref), we see that increasing the number of strata (i.e. decreasing the effective number of observations per strata) produces bias reductions without seriously affecting variance for \(\widehat{\beta}_{n, \mathrm{SAT}}\) or \(\widehat{\beta}_{n, \mathrm{IMP}}\). However, reducing bandwidth to undersmooth will reduce bias and inflate the variance of this estimator.
In this paper, we have characterized the maximal gains in efficiency from using baseline covariates beyond strata in RCT's covariate adaptive randomization is used. To do this, we have established a semiparametric efficiency bound under CAR for such experiments. This is the minimum estimation variance achievable if one uses all the relevant information the data have to offer about the parameter of interest (the ATE here). If baseline covariates are used, we use the information they contain about the ATE fully through differences in conditional means, averaged over the support of the covariates. With continuous covariates, or more covariates present than used in stratification, a strict improvement over fully saturated regression estimates of the ATE is possible. The efficiency bound established is shown to be achievable under the same conditions used for its derivation by using a leave one out Nadaraya-Watson kernel regression estimator. Simulation evidence presented shows that this estimator can indeed reach the efficiency bound in large samples. However, such improvements are not guaranteed in finite samples, and reaching the efficiency bound can be slow especially in the presence of many covariates due to the curse of dimensionality.
{\printbibliography}