EconBase
← Back to paper

Robust and Efficient Estimation of Potential Outcome Means under Random Assignment

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

78,595 characters · 19 sections · 49 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Robust and Efficient Estimation of Potential Outcome Means Under Random Assignment

abstract\singlespacing We study efficiency improvements in randomized experiments for estimating a vector of potential outcome means using regression adjustment (RA) when there are more than two treatment levels. We show that linear RA which estimates separate slopes for each assignment level is never worse, asymptotically, than using the subsample averages. We also show that separate RA improves over pooled RA except in the obvious case where slope parameters in the linear projections are identical across the different assignment levels. We further characterize the class of nonlinear RA methods that preserve consistency of the potential outcome means despite arbitrary misspecification of the conditional mean functions. Finally, we apply these regression adjustment techniques to efficiently estimate the lower bound mean willingness to pay for an oil spill prevention program in California.

JEL Codes: C21, C25

Keywords: Multivalued Treatments, Experiment, Regression Adjustment, Heterogeneous Effects

\begingroup \footnote{$^\dag$Department of Econometrics and Business Statistics, Monash University. Email: [email removed], $^\ddag$Department of Economics, Michigan State University. Email: [email removed]. We would like to thank the Associate Editor and three anonymous referees for their helpful comments and suggestions.} \addtocounter{footnote}{-1} \endgroup

Introduction

In the past several decades, the potential outcomes framework has become a staple of causal inference in statistics, econometrics, and related fields. Envisioning each unit in a population under different states of intervention or treatment allows one to define treatment or causal effects without reference to a model. One merely needs the potential outcome (PO) means for subpopulations.

When interventions are randomized -- as in a clinical trial [hirano2001estimation], assignment to participate in a job training program [calonico2017women], receiving a private school voucher [angrist2006long], or contingent valuation studies where different bid values are randomized among individuals [carson2004valuing] -- one can simply use the subsample means for each treatment level in order to obtain unbiased and consistent estimators of the PO means. In some cases, the precision of the subsample means will be sufficient. Nevertheless, with the availability of good predictors of the outcome or response, it is natural to think that the precision can be improved, thereby shrinking confidence intervals and making conclusions about interventions more reliable.

In this paper we build on negi2021revisiting, NW (2021) hereafter, who studied the problem of estimating the average treatment effect (ATE) under random assignment with one control and one treatment group. In the context of random sampling, NW (2021) showed that performing separate linear regressions for the two groups while estimating the ATE never does worse, asymptotically, than the simple difference in means estimator or a pooled regression adjustment estimator. These findings are complementary to lin2013agnostic, who used a finite population framework to study linear regression adjustment and arrived at essentially the same conclusions. NW (2021) also characterized the class of nonlinear regression adjustment methods that produce consistent estimators of the ATE without any additional assumptions (except regularity conditions).

In the current paper, we allow for $G\geq 2$ treatment levels and study the problem of joint estimation of the vector of PO means. We assume that the assignment to treatment is random -- independent of both potential outcomes and observed predictors of the POs. In terms of the assignment mechanism, we assume a Bernoulli scheme whereby each unit has an equally likely probability of being assigned to a given treatment level, $g$. Importantly, other than substantive restrictions such as random assignment, $i.i.d$ sampling, and stable unit treatment value assumption, we impose only standard regularity conditions (such as finite second moments of the covariates). In other words, the regression adjustment (RA) estimators are consistent under essentially the same assumptions as the subsample means estimator with, generally, smaller asymptotic variances. Interestingly, even if the predictors are unhelpful, or the slopes in the linear projections are the same across all groups, no asymptotic efficiency is lost from using the most general RA method.

The extension from the binary treatment case to more than two treatment levels is nontrivial, both in terms of derivations and scope of applications. In the binary case, NW (2021) compared variances of estimators of the scalar ATE. In the current paper, we show that differences in asymptotic variance matrices of the PO estimators are positive semi-definite. Naturally, this result implies that any linear functions of the PO means are more efficiently estimated using separate RA. In addition, nonlinear functions of PO means --such as ratios -- are also more efficiently estimated. Even in the binary treatment case, our results significantly improve over NW (2021).\footnote{It is not enough, even in the $G=2$ case, to use NW (2021) for each of the pairwise difference in means. Knowing that each pairwise ATE is more efficiently estimated does not imply that all linear combinations are.}

We also extend the nonlinear RA results in NW (2021) to this general PO framework with $G$ treatment levels. gail1984biased, who study nonlinear regression for the binary treatment case, have cautioned against using certain nonlinear regression models as being biased for the ATE. imbens2015causal also leave the impression that nonlinear RA will be more efficient but at the cost of consistency, and therefore should be avoided. We show that for particular kinds of responses-- the leading cases being binary, fractional, and nonnegative -- it is possible to consistently estimate PO means using pooled and separate RA by combining results from Quasi Maximum Likelihood Estimation (QMLE) in the linear exponential family (LEF) with features of the canonical link function.\footnote{In the generalized linear model literature, a link function $m^{-1}(\cdot)$ relates the conditional mean of the outcome to a linear predictor such that $\mathbb{E}[Y(g)|\mathbf{X}]=m(\mathbf{X}\bm{\beta})$. A canonical link is a special link function which guarantees that the model $m(\mathbf{X}\bm{\beta})$ fits the overall mean of the outcome without necessarily being correctly specified.} While gail1984biased discussed maximum likelihood estimation in the exponential family, they do not distinguish between canonical and non-canonical link functions. As shown in NW (2021), and extended to the multiple treatment case here, one needs to use the mean function associated with the canonical link to ensure consistency without imposing additional restrictions. guo2023generalized reach a similar conclusion in the binary treatment case studied in NW (2021), showing that generalized linear models with canonical links are examples of imputation models that are prediction unbiased for the sample average treatment effect under the design-based framework. Unlike the linear RA case, we do not have a general result that shows separate nonlinear RA is asymptotically more efficient than the subsample means estimator. However, if we add the assumption that the conditional mean functions are correctly specified, we show that doing separate RA is (weakly) more efficient than the subsample means estimator.\footnote{cohen2024no propose a two-step calibrated Generalized Oaxaca Blinder estimator that is efficient even under misspecification.} We further establish that separate RA attains the semiparametric efficiency bound under correct conditional mean specification.

A Monte Carlo exercise substantiates our theoretical results.\footnote{We consider three different population models. The outcomes are generated to be either fractional, non-negative, or continuous with an unrestricted support. The results for fractional and non-negative outcomes can be found in section B of the online appendix.} We find that RA estimators, both linear and nonlinear, improve over subsample means in terms of standard deviations and have small biases, even with fairly small sample sizes. Not surprisingly, the magnitude of the precision gains with RA methods generally depends on how well the covariates predict the potential outcomes. Finally, we conclude with an empirical application that uses the RA methods to estimate the lower bound mean willingness to pay for an oil spill prevention program along California's coast.

In terms of contribution, this paper provides a unified framework for studying regression adjustment in experiments, allowing for multiple treatments under the infinite population (or superpopulation) setting. Besides studying linear and nonlinear RA methods and making efficiency comparisons, we also characterize the semiparametric efficiency bound for all consistent and asymptotically linear estimators of the PO means and show that separate RA attains this bound under correct conditional mean specification. Because our results compare asymptotic variance matrices for different PO mean estimators, the efficiency results derived here extend to linear and (smooth) nonlinear functions of the PO means. We also fill an important gap in the regression adjustment literature by studying pooled nonlinear methods which have not been discussed elsewhere. We show how nonlinear pooled RA is consistent whenever nonlinear separate RA is, making it an important alternative when the researcher lacks sufficient degrees of freedom for estimating separate slopes. On the practical side, the proposed RA estimators are easy to implement with common statistical software like Stata (see online appendix D for reference).

Related Literature: This paper contributes to a vast and rich literature on regression adjustment in experiments. In the design-based framework, freedman2008regression, freeedman2008regression2 highlighted that adjusting for covariates additively in a regression is not guaranteed to produce efficiency gains in experiments compared to the simple difference in means estimator. For a binary treatment, this debate was followed up in lin2013agnostic which showed that separate RA is weakly more efficient than both the pooled RA and simple difference in means estimators of the ATE.\footnote{More recently, \citet*{chang2021exact} propose exact bias-corrected estimators to alleviate the finite sample concerns of RA. This is different from the focus of our paper which is interested in improving efficiency with RA and assumes that sample sizes are large enough to ensure consistent estimation of the PO means. A similar argument for bias correction in high dimensional settings can be found in lei2021regression and \citet*{chiang2023regression}.}

Among those adopting a superpopulation approach with a binary treatment, yang2001efficiency and \citet*{leon2003semiparametric} focus on a pre-test post-test trial with no baseline covariates. Much like the current paper, specialized to the binary treatment case with ATE as the parameter of interest, yang2001efficiency show that the interacted estimator is the most efficient among simple difference in means and pooled RA estimators. leon2003semiparametric identify the most efficient ATE estimator using the semiparametric theory of robins1994estimation but do not provide any efficiency comparisons between commonly employed estimators of the ATE. tsiatis2008covariate characterize the class of all consistent and asymptotically normal (CAN) estimators of the ATE. They further show that such estimators have a common form which can be used to rank the different linear estimators in terms of efficiency. The current paper (with $G=2$) and NW (2021) arrive at the same efficiency results with respect to linear RA as tsiatis2008covariate. In addition, the former explicitly discuss consistency with nonlinear RA (separate and pooled) involving quasi-log-likelihoods and canonical links from the generalized linear model literature. Furthermore, this paper establishes the asymptotic efficiency of separate RA relative to subsample averages when the conditional mean functions are correctly specified. pitkin2013improved compare the asymptotic variances of simple difference in means and linear RA (pooled and separate) under essentially a similar assumption as this paper wherein the conditional mean functions are not assumed to be linear. However, pitkin2013improved neither discuss nonlinear RA nor conditions under which it is efficient relative to simple difference in means. Similar to leon2003semiparametric, \citet*{zhang2008improving} also study regression adjustment from a semiparametric perspective. However, their focus is on CAN estimators of pairwise treatment comparisons in a multiarmed trial. As mentioned above, the efficiency results derived in this paper extend to linear and nonlinear functions of the PO means, including pairwise ATEs. The current paper also explicitly ranks the different linear and nonlinear RA estimators, which is not explored in zhang2008improving. We further expand our set of results to include the semiparametric efficiency bound for all CAN estimators of the PO means and establish that SRA of the QML variety attains the efficiency bound when the conditional means are correctly specified. Another paper that studies nonlinear RA for estimating the ATE is rosenblum2010simple. The efficiency results discussed there are nested within this paper, which additionally considers pooled nonlinear methods that have not been discussed elsewhere. An added advantage of the nonlinear estimators proposed here is that they are computationally simple and easy to implement.

In the design-based setting, guo2023generalized and cohen2024no also discuss nonlinear RA in binary experiments. The former proposes ATE estimators which impute the missing potential outcomes using nonlinear models that are prediction unbiased. This is similar to the “mean-fitting” property discussed here and in NW (2021) which is crucial for consistent estimation of the PO means. This property is guaranteed by choosing appropriate combinations of canonical link and quasi-log-likelihood functions in the LEF estimated using QMLE. cohen2024no extend guo2023generalized's efficiency result to the case of mean misspecification.\footnote{They argue that a second-step calibration of the Generalized Oaxaca Blinder estimator produces an estimator of the ATE that is efficient even under misspecification of the nonlinear model.}

A recent strand also looks at multiarmed treatments in the context of factorial experiments under the design-based framework [zhao2022reconciling, zhao2022regression, zhao2023covariate, pashley2023causal]. Such experiments are designed to accommodate multiple factors of interest in a single experiment and define treatment levels as all possible combinations of the factors involved. The typical concern in this literature is degree of freedom conservation while estimating all main and interaction effects of these factors simultaneously. An exception is zhao2023covariate which discusses linear RA for multiarmed experiments (factorial experiments being a special case). The efficiency results concerning linear RA in zhao2023covariate have a direct analogue in the current paper under the superpoulation framework. For more complex experimental designs, see cytrynbaum2023covariate, jiang2023regression, roth2023efficient, and chang2023design.

The rest of the paper is organized as follows. Section (ref) describes the potential outcomes framework with $G$ treatment levels along with a discussion of the main assumptions. Section (ref) presents the asymptotic variances of the linear RA estimators, whereas section (ref) ranks them in terms of asymptotic efficiency. Section (ref) considers a class of nonlinear RA estimators that ensure consistency for the PO means without imposing additional assumptions. Section (ref) characterizes the semiparametric efficiency bound for all regular estimators of the PO means and establishes local efficiency of separate RA. Section (ref) constructs a simulation exercise for studying the finite sample behavior of the RA estimators. Section (ref) applies the RA methods to study willingness to pay for an oil spill prevention program, and section (ref) concludes.

Potential Outcomes Framework and Assumptions

We use the standard potential outcomes framework, also known as the Neyman-Rubin causal model. The goal is to estimate the population means of $G$ potential (counterfactual) outcomes, $Y(g)$, $ g=1,...,G$. Define $ \mu _{g}=\mathbb{E}\left[ Y(g)\right]$ for each $g=1,...,G.$

The vector of assignment indicators is $\mathbf{W}=(W_{1},...,W_{G})$, where each $W_{g}$ is binary and $W_{1}+W_{2}+\cdots +W_{G}=1$. In other words, the groups are exhaustive and mutually exclusive. The setup applies to many situations, including the standard treatment-control setup with $G=2$, multiple treatment levels (with $g=1$ as the control group), and in contingent valuation studies where subjects are presented with a set of $G$ prices or bid values.

Next, let $\mathbf{X}=(X_{1},X_{2},...,X_{K})$ be a vector of observed covariates of dimension $K$ which is assumed to be fixed in our asymptotic analysis. With respect to the assignment process, we make the following assumption.

assumption[Random Assignment] Assignment is independent of the potential outcomes and observed covariates: $\mathbf{W}\perp \left[ Y(1),Y(2),...,Y(G),\mathbf{X}\right]$. Further, $\rho _{g}\equiv \mathbb{P}(W_{g}=1)>0$ holds.

Assumption 1 puts us in the framework of experimental interventions. We also assume that each group, $g$, has a positive probability of being assigned such that $\rho _{1}+\rho _{2}+\cdots +\rho _{G}=1$.

assumption[Random Sampling] For a nonrandom integer $N$, $\big\{ \big[\mathbf{W}_{i},Y_{i}(1),Y_{i}(2),...,Y_{i}(G),$ $\mathbf{X}_{i} \big]$ $:i=1,2,...,N\big\}$ is independent and identically distributed.

The i.i.d assumption is not the only one we can make. For example, we could allow for a sampling-without-replacement scheme given a fixed sample size $N$. This would complicate the analysis because it generates a slight correlation across draws from the population. As discussed in NW (2021), Assumption (ref) is traditional in studying the asymptotic properties of estimators and is realistic as an approximation in cases where a sample is taken from a large population. It also forces us to account for the sampling error in $\mathbf{\bar{X}}$, as an estimator of $\bm{\mu }_{\mathbf{X}}=\mathbb{E}\left( \mathbf{X} \right) $.

A different approach is to assume that there is no sampling uncertainty and instead assume we observe the entire population. In the design--based approach, all uncertainty is in the assignment of the treatment indicators, $ \mathbf{W}=(W_{1},...,W_{G})$. Among others, freedman2008regression and lin2013agnostic take this approach in studying linear regression adjustment with a binary treatment. In many examples in economics (including the empirical example in our supplement), one does not observe the entire population and so it seems natural to study the random sampling framework. Our results for the random sampling case can be used to further study the efficiency issue in a framework that allows both sampling and assignment uncertainty, as in abadie2020sampling. Such a framework will affect inference only in cases where the sample is a substantial fraction of the population. If the sample size is small relative to the population size, the finite-population adjustments of the kind derived in abadie2020sampling would have a trivial effect.

Given Assumption (ref), for each draw $i$ from the population we only observe

equation[equation omitted — 98 chars of source]

and so the data are $ \left\{ \left( \mathbf{W}_{i},Y_{i},\mathbf{X}_{i}\right):i=1,2,...,N\right\}$. Definition of population quantities only requires us to use the random vector $\left( \mathbf{W},Y,\mathbf{X}\right)$, which represents the population.

Along with assumptions (ref) and (ref), we also assume that the stable unit treatment value assumption (SUTVA) holds. This implies that there are no spillovers or hidden variations of the treatments. Random assignment, $i.i.d$ sampling, and SUTVA are the only substantive restrictions used in this paper and will be maintained throughout. Subsequently, we assume that linear projections exist and that the central limit theorem holds for properly standardized sample averages of i.i.d random vectors. Therefore, we are implicitly imposing at least finite second moment assumptions on each $Y(g)$ and $X_{j}$. We do not make this explicit in what follows.

Completely Randomized Experiment

In this paper, we consider a randomized experiment where each unit has an equally likely probability, $\rho_g$, of being assigned to treatment level $g$. An assignment scheme that is common in practice is where a fixed number of units are assigned to each treatment level. This is known as a completely randomized experiment and creates correlation between draws. Common analyses of a completely randomized experiment occur in the design-based framework where the sample coincides with the population (which is, naturally, assumed to be finite). Our framework is based on the availability of i.i.d draws, which includes the infinite superpopulation setting and also the finite population setting where sampling is done with replacement. An interesting question is whether the results obtained here apply to the setting of a completely randomized experiment, or settings where the population is finite and the assignments are i.i.d draws from a multinomial distribution. Other settings of interest include when the assignment might be determined in two (or more) stages. For example, clusters of units are defined, and then clusters are randomly assigned to be treated or control clusters. For the treated clusters, units within the cluster are randomly assigned to control or treatment-- as in abadie2023should for the case of binary treatment without covariates. We leave these topics for future research.

Subsample Means and Linear Regression Adjustment

In this section we discuss the asymptotic variances of three estimators:\ the subsample means, separate regression adjustment, and pooled regression adjustment.\footnote{The asymptotic representation proofs of the three estimators can be found in section A of the online appendix.}

Subsample Means (SM)

The simplest estimator of $\mu _{g}$ is the sample average within treatment group $g$: $ \bar{Y}_{g}=N_{g}^{-1}\sum_{i=1}^{N}W_{ig}Y_{i}=N_{g}^{-1} \sum_{i=1}^{N}W_{ig}Y_{i}(g),$ where $ N_{g}=\sum_{i=1}^{N}W_{ig}$ is a random variable in our setting. In expressing $\bar{Y}_{g}$ as a function of the $Y_{i}(g)$ we use $W_{ih}W_{ig}=0$ for $h\neq g$. Under random assignment and random sampling, {

align*[align* omitted — 308 chars of source]

} and so $\bar{Y}_{g}$ is unbiased conditional on observing a positive number of units in group $g$. By the law of large numbers, a consistent estimator of $\rho _{g}$ is $ \hat{\rho}_{g}=N_{g}/N$ which is the sample share of units in group $g$. Therefore, by the law of large numbers and Slutsky's Theorem, {

align*[align* omitted — 273 chars of source]

} and so $\bar{Y}_{g}$ is consistent for $\mu _{g}$. By the central limit theorem, $\sqrt{N}\left( \bar{Y}_{g}-\mu _{g}\right) $ is asymptotically normal. We need an asymptotic representation of $\sqrt{N}\left( \bar{Y}_{g}-\mu _{g}\right) $ that allows us to compare its asymptotic variance with those from regression adjustment estimators. To this end, write

equation*[equation* omitted — 111 chars of source]

where $\mathbf{\dot{X}}$ is demeaned using the population mean, $\bm{\mu }_{\mathbf{X}}$. The demeaning ensures that the intercept in the equation above can be directly interpreted as the PO mean. Now project each $V(g)$ linearly onto $\mathbf{\dot{X}}$: $V(g)=\mathbf{\dot{X}}\bm{\beta_g}+U(g)$. By construction, the population projection errors $U(g)$ have the properties: $ \mathbb{E}\left[ U(g)\right] =0 \text{ and } \mathbb{E}[\mathbf{\dot{X}}^{\prime }U(g)] =\mathbf{0}\text{, }$ for each $g$. Plugging in gives

equation*[equation* omitted — 92 chars of source]

Importantly, by random assignment, $\mathbf{W}$ is independent of $[U(1),...,U(G),\mathbf{\dot{X}}]$. The observed outcome can be written as

equation*[equation* omitted — 110 chars of source]

Our goal is to be able to make efficiency statements about both linear and nonlinear functions of the vector of means $\bm{\mu }=\left( \mu_{1},\mu _{2},...,\mu _{G}\right) ^{\prime }$, and so we stack the subsample means into the $G\times 1$ vector $\mathbf{\bar{Y}}$. For later comparison, it is helpful to remember that $\mathbf{\bar{Y}}$ is the vector of OLS coefficients in the regression (without an intercept): $Y_{i}\text{ on }W_{i1}\text{, }W_{i2}\text{, ..., }W_{iG}\text{, } i=1,2,...,N.$ The asymptotic representation for the subsample means estimator is given as {

equation*[equation* omitted — 157 chars of source]

} where {

equation[equation omitted — 488 chars of source]

}

For the vectors defined above, we know that by random assignment and the linear projection property, $\mathbb{E}\left( \mathbf{L}_{i}\right) =\mathbb{E}\left( \mathbf{Q}_{i}\right) =\mathbf{0}$, and $\mathbb{E}\left( \mathbf{L}_{i}\mathbf{Q}_{i}^{\prime }\right) =\mathbf{0}$. Also, because $W_{ig}W_{ih}=0$, $g\neq h$, the elements of $\mathbf{L}_{i}$ are pairwise uncorrelated; the same is true of the elements of $\mathbf{Q}_{i}$.

Separate Regression Adjustment (SRA)

To motivate SRA, write the linear projection for each $g$ as

equation[equation omitted — 196 chars of source]

It follows immediately that $\mu _{g}=\alpha _{g}+ \bm{\mu}_{\mathbf{X}} \bm{\beta_g}.$ Consistent estimators of $\alpha _{g}$ and $\bm{\beta_g}$ are obtained from the regression: $Y_{i}\text{ on }1\text{, }\mathbf{X}_{i}\text{,\ if }W_{ig}=1,$ which produces intercept and slopes $\hat{\alpha}_{g}$ and $\bm{\hat{\beta}_g}$. Therefore, the SRA estimator of PO mean, $\mu_g$ is given by $\hat{\mu}_g = \hat{\alpha}_g+\mathbf{\bar{X}}\bm{\hat{\beta}}_g$. Stacking all the SRA estimators for PO means into the vector, $\bm{\hat{\mu}}_{SRA}$, we obtain the following asymptotic representation {

equation*[equation* omitted — 135 chars of source]

} where $\mathbf{Q}_{i}$ is given in ((ref)) and {

equation[equation omitted — 214 chars of source]

}

Again, for the vectors defined above, both $\mathbf{K}_{i}$ and $\mathbf{Q}_{i}$ have zero means, the latter by random assignment. Further, $\mathbb{E}\left( \mathbf{K} _{i}\mathbf{Q}_{i}^{\prime }\right) =\mathbf{0}$ because $ \mathbb{E}[\mathbf{\dot{X}}_{i}^{\prime }W_{ig}U_{i}(g)] = \mathbb{E}(W_{ig})\mathbb{E}[ \mathbf{\dot{X}}_{i}^{\prime }U_{i}(g)] =\mathbf{0}$. However, unlike the elements of $\mathbf{L}_{i}$, we must recognize that the elements of $\mathbf{K}_{i}$ are correlated except in the trivial case that all but one of the $\bm{\beta_g}$ are zero. It is important to note here that if assignment is unconfounded given $\mathbf{X}$, SRA will be consistent provided that the conditional means are linear.

Pooled Regression Adjustment (PRA)

Now consider the pooled estimator, $\bm{\hat{\mu}}_{PRA}$, which can be obtained as the vector of coefficients on $\mathbf{W}_{i}=\left(W_{i1},W_{i2},...,W_{iG}\right) $ from the regression: $Y_{i}\text{ on }\mathbf{W}_{i}\text{, }\mathbf{\ddot{X}}_{i},\text{ } i=1,2,...,N$ where $\mathbf{\ddot{X}}_i= \mathbf{X}_i-\mathbf{\bar{X}}$ is the sample demeaned covariate vector. We refer to this as a pooled method because the coefficients on $\mathbf{\ \ddot{X}}_{i}$, say, $\bm{\check{\beta}}$, are assumed to be the same for all groups. Compared with subsample means, we add the controls $\mathbf{\ \ddot{X}}_{i}$, but unlike SRA, the pooled method imposes the same coefficients across all $g$. The asymptotic representation is given as {

equation*[equation* omitted — 178 chars of source]

} where $\mathbf{K}_{i}$ and $\mathbf{Q}_{i}$ are defined as before and {

equation[equation omitted — 397 chars of source]

} and $\bm{\delta_g}=\bm{\beta_g}-\bm{\beta}$, where $\bm{\beta}$ is the linear projection of \ $Y$ on $\mathbf{\dot{X}}$. Again, by random assignment and the linear projection property, $ \mathbb{E}\left( \mathbf{F}_{i}\mathbf{K}_{i}^{\prime }\right) =\mathbb{E}\left(\mathbf{F}_{i}\mathbf{Q}_{i}^{\prime }\right) =\mathbf{0}$.

Comparing the Asymptotic Variances

We now take the representations given above and use them to compare the asymptotic variances of the three estimators. It is helpful to summarize the conclusions reached in Section 3: {

align[align omitted — 519 chars of source]

} where $\mathbf{L}_{i}$, $\mathbf{Q}_{i}$, $\mathbf{K}_{i}$, and $\mathbf{F} _{i}$ are defined in ((ref)), ((ref)) and ((ref)), respectively.

Comparing Separate RA to SM

We now show that, asymptotically, $\bm{\hat{\mu}}_{SRA}$ is no worse than $\bm{\hat{\mu}}_{SM}$.

theorem[Efficiency of SRA relative to SM] Under Assumptions (ref), (ref), SUTVA, and finite second moments used to obtain the asymptotic representations in (ref) and (ref), \begin{equation*} \mathrm{Avar}\left[ \sqrt{N}\left( \bm{\hat{\mu}}_{SM}-\bm{\mu } \right) \right] -\mathrm{Avar}\left[ \sqrt{N}\left( \bm{\hat{\mu}}_{SRA}- \bm{\mu }\right) \right] =\mathbf{\Omega _{L}}-\mathbf{\Omega _{K}} \end{equation*} is PSD where $\bm{\Omega }_{\mathbf{L}}\equiv \mathbb{E}(\mathbf{L}_{i}\mathbf{L }_{i}^{\prime})$ and $\bm{\Omega }_{\mathbf{K}}\equiv \mathbb{E}(\mathbf{K}_{i}\mathbf{K }_{i}^{\prime})$.

The proof can be found in Appendix (ref). To get an intuition for this efficiency gain, note that SRA uses covariate values across the entire sample for imputing the missing potential outcomes. This requires estimating the population mean of $\mathbf{X}$. Therefore, the asymptotic variance of $\bm{\hat{\mu}}_{SRA}$ is influenced not only by the linear projection error but also $\mathbf{\bar{X}}$. In contrast, SM only averages observed outcomes for each treatment level which means that its asymptotic variance ends up being a function of the variance of the subsample means, $\mathbf{\bar{X}}_g$, and the linear projection errors. So a comparison of the two estimators boils down to comparing $\mathbf{L}_i$ and $\mathbf{K}_i$ whose $g$-th elements are given by $W_{ig}\mathbf{\dot{X}}_i\bm{\beta_g}/\rho_g$ and $\mathbf{\dot{X}}_i\bm{\beta_g}$, respectively. To put it simply, SM is inefficient because it uses $\mathbf{\bar{X}}_g$ which is a consistent (due to random assignment) but less efficient estimator of the population mean compared to $\mathbf{\bar{X}}$, that is used by SRA. Adding orthogonal controls that reduce the error variance without inducing collinearity is relatively intuitive. SRA does that without using a “misspecified” model in the sense that it does not impose incorrect restrictions: the slopes are the same. Imposing zero slopes (SM) and common slopes (PRA) can both be viewed as forms of misspecification. They don’t cause inconsistency, but it is not generally possible to compare efficiencies when using two “misspecified” models.

The one case where there is no gain in asymptotic efficiency in using SRA is when $\bm{\beta_g}=\mathbf{0}$, $g=1,...,G$, in which case the covariates do not help predict any of the potential outcomes. Importantly, there is no gain in asymptotic efficiency in imposing $\bm{\beta_g}=\mathbf{0}$ when it is true. From an asymptotic perspective, it is harmless to separately estimate the $\bm{\beta_g}$ even when they are zero. When they are not all zero, estimating them leads to asymptotic efficiency gains.

Theorem (ref) also implies that any smooth nonlinear function of $\bm{\mu}$ is estimated more efficiently using $\bm{\hat{\mu}}_{SRA}$ by the delta method. For example, in estimating a percentage difference in means, we would be interested in $\mu _{2}/\mu _{1}$, and using SRA is asymptotically more efficient than using SM.

Separate RA versus Pooled RA

The comparison between SRA and PRA is simple given the expressions in ((ref)) and ((ref)) because, as stated earlier, $\mathbf{F}_{i}$, $\mathbf{K}_{i}$, and $\mathbf{Q}_{i}$ are pairwise uncorrelated.

theorem[Efficiency of SRA relative to PRA] Under Assumptions (ref), (ref), SUTVA, and finite second moments used to obtain the asymptotic representations in (ref) and (ref), \begin{equation*} \mathrm{Avar}\left[ \sqrt{N}\left( \bm{\hat{\mu}}_{PRA}-\bm{\mu } \right) \right] -\mathrm{Avar}\left[ \sqrt{N}\left( \bm{\hat{\mu}}_{SRA}- \bm{\mu }\right) \right] =\mathbf{\Omega _{F}} \end{equation*} where $\bm{\Omega}_{\mathbf{F}}\equiv \mathbb{E}(\mathbf{F}_{i}\mathbf{F }_{i}^{\prime})$ is PSD.

The proof can be found in Appendix (ref). It follows immediately from Theorem (ref) that $\bm{\hat{\mu}}_{SRA}$ is never less asymptotically efficient than $\bm{\hat{\mu}}_{PRA}$. There are some special cases where the estimators achieve the same asymptotic variance, the most obvious being when the slopes in the linear projections are homogeneous: $ \bm{\beta_1}=\bm{\beta_2}=\cdots =\bm{\beta_g}$. As with comparing SRA with subsample means, there is no gain in efficiency from imposing this restriction when it is true. This is another fact that makes SRA attractive if the sample sizes within each treatment level are not small.

Other situations where there is no asymptotic efficiency gain in using SRA are more subtle. In general, suppose we are interested in linear combinations $\tau =\mathbf{a}^{\prime } \bm{\mu }$ for a given $G\times 1 $ vector $\mathbf{a}$. If $\mathbf{a}^{\prime } \mathbf{\Omega }_{\mathbf{F}}\mathbf{a}=0$, then $\mathbf{a}^{\prime } \bm{\hat{\mu}}_{PRA}$ is asymptotically as efficient as $\mathbf{a}^{\prime } \bm{\hat{\mu}}_{SRA}$ for estimating $ \tau $. Generally, the diagonal elements of $ \mathbf{\Omega }_{F}=\mathbb{E}\left( \mathbf{F}_{i}\mathbf{F}_{i}^{\prime }\right)$ are $ (1-\rho _{g})\bm{\delta_g}^{\prime } \mathbf{ \Omega }_{\mathbf{X}} \bm{\delta_g}/\rho _{g}$ because $\mathbb{E}\left[ \left( W_{ig}-\rho _{g}\right) ^{2}\right] =\rho_{g}(1-\rho _{g})$. The off diagonal terms of $\mathbf{\Omega }_{\mathbf{F}}$ are $- \bm{\delta_g}^{\prime } \mathbf{\Omega }_{\mathbf{X}} \bm{\delta_h}$ because $\mathbb{E}\left[ \left( W_{ig}-\rho _{g}\right) \left( W_{ih}-\rho_{h}\right) \right] =-\rho _{g}\rho _{h}$. Now consider the case covered in NW (2021), where $G=2$ and $\mathbf{a}^{\prime }=\left(-1,1\right) $, so the parameter of interest is $\tau =\mu _{2}-\mu _{1}$ (the ATE). If $\rho _{1}=\rho _{2}=1/2$ then {

equation*[equation* omitted — 419 chars of source]

} Now $\bm{\delta_2}=- \bm{\delta_1}$ because $\bm{\delta_1}= \bm{\beta_1}-( \bm{\beta_1}+ \bm{\beta_2})/2=(\bm{\beta_1}- \bm{\beta_2})/2=- \bm{\delta}_{2}$. Therefore, {

equation*[equation* omitted — 914 chars of source]

} Interestingly, the asymptotic equivalence of the separate and pooled OLS methods does not extend to the case $G\geq 3$. For any $G\geq 3$ with $\rho_{g}=1/G$ for all $g$, $ 1-\rho _{g}=1-\frac{1}{G}=\frac{\left( G-1\right) }{G}$ and so $\frac{1-\rho _{g}}{\rho _{g}}=G-1$. Note that $ \bm{\delta_g}=\bm{\beta_g}-\left( \bm{\beta_1}+\bm{\beta_2}+\cdots +\bm{\beta_g}\right) /G$ and it is less clear when a degeneracy occurs in the difference in asymptotic variances. So, for example, in our application which estimates a lower bound mean willingness to pay, SRA is generally more efficient than PRA even though the assignment probabilities are identical.

Nonlinear Regression Adjustment

We now discuss a class of nonlinear regression adjustment methods that preserve consistency without adding additional assumptions other than the ones discussed in section (ref) (and weak regularity conditions). In particular, we extend the setup in NW (2021) to more than two treatment levels.

Typically, when the outcome is discrete or has limited support, nonlinear functions of $\mathbb{E}[Y(g)|\mathbf{X}]$ may be more appropriate and may provide a better approximation to the true conditional mean. However, consistency with nonlinear models is closely tied to functional form assumptions. In this section, we discuss nonlinear methods that are consistent for $\mathbb{E}[Y(g)]$ despite possible misspecification in $\mathbb{E}[Y(g)|\mathbf{X}]$. Not surprisingly, using a canonical link function in the context of QMLE in the linear exponential family plays a key role. We show that both separate and pooled nonlinear methods are consistent provided we choose the mean functions and objective functions appropriately. Unlike in the linear case, we can only show that SRA improves over the subsample means estimator when the conditional means are correctly specified.

Separate Regression Adjustment

We model the conditional means, $\mathbb{E}\left[ Y(g)|\mathbf{X}\right]$, for each $g=1,2,...G$ to be $m(\alpha _{g}+\mathbf{X} \bm{\beta_g})$, where $m\left( \cdot \right)$ is a smooth function defined on $\mathbb{R}$.\footnote{The range of $m\left( \cdot \right) $ is chosen to reflect the nature of $Y(g)$. Given that the nature of $Y(g)$ does not change across treatment levels, we choose a common function $m\left( \cdot \right) $ across all $g$. Also, as usual, the vector $\mathbf{X}$ can include nonlinear functions (typically squares, interactions, and so on) of underlying covariates.} In the generalized linear model literature, $m^{-1}(\cdot)$ is known as the link function and relates the conditional mean of the outcome to a linear predictor. A canonical link is tied to a specific quasi-log-likelihood (QLL) in the LEF such that if one chooses certain combinations of QLLs and canonical link functions, we obtain

equation[equation omitted — 148 chars of source]

where $\alpha _{g}^{\ast }$ and $\bm{\beta_g}^{\ast }$ are the probability limits of the QMLE whether or not the conditional mean function, chosen to be $m(\cdot)$, is correctly specified. Table (ref) gives the pairs of mean function (or canonical links) and QLLs that ensure consistent estimation of PO means. To ensure consistency, the mean should have the index form common in the generalized linear models literature.

table[table omitted — 1,129 chars of source]

The binomial QMLE is a good choice for counts with a known upper bound, even if it is individual-specific ($B_{i}$ need not be a positive integer for each $i$). It can also be applied to corner solution outcomes in the interval $[0,B_{i}]$ where the outcome is continuous on $(0,B_{i})$ but perhaps has mass at zero or $B_{i}$. The leading case is $B_{i}=B$. Note that we do not recommend a Tobit model in such cases because the Tobit estimates of $\mu_{g}$ are not robust to failure of the Tobit distributional assumptions or the Tobit form of the conditional mean.

This \textquotedblleft mean fitting property\textquotedblright\ given in (ref) can be established by studying the population first order conditions (FOCs) with the choice of canonical link function, {

align[align omitted — 303 chars of source]

} It follows from equations ((ref)) and ((ref)) that {

equation[equation omitted — 200 chars of source]

} Also, we can rewrite the right hand side using the law of iterated expectations as {

equation[equation omitted — 354 chars of source]

} where the second equality holds by random assignment. Using ((ref)) and ((ref)) along with the fact that $\mathbb{E}\left[ W_{g}Y(g)\right] =\rho_{g}\mu _{g}$, we obtain the mean fitting result in equation ((ref)). Applying nonlinear RA, therefore, with multiple treatment levels is straightforward. One can obtain $\hat{\alpha}_{g}$ and $\bm{\hat{\beta}_g}$ by solving the sample analogues of conditions ((ref)) and ((ref)). For treatment level $g$, after obtaining $\hat{\alpha}_{g}$, $\bm{\hat{\beta}_g}$ by QMLE, the mean, $\mu _{g}$, is estimated as $\hat{\mu}_{g}=N^{-1}\sum_{i=1}^{N}m(\hat{\alpha}_{g}+\mathbf{X}_{i}\bm{ \hat{\beta}_g})$, which includes linear RA as a special case. This estimator is consistent by a standard application of the uniform law of large numbers; see, for example, wooldridge2010econometric (Chapter 12, Lemma 12.1). As in the linear case, the FOCs of the QMLEs using a canonical link in the LEF allow us to write the subsample averages as {

equation*[equation* omitted — 145 chars of source]

} Given that $\hat{\mu}_{g}$ averages across all of the observations rather than just the subset of units at treatment level $g$, it seems that $\hat{\mu}_{g}$ should be asymptotically more efficient than $\bar{Y}_{g}$. Unfortunately, the proof used in the linear case does not generally go through in the nonlinear case. Nevertheless, when the conditional mean function is correctly specified, it is possible to show that nonlinear SRA improves over the subsample means estimator.

theorem[Efficiency of SRA relative to SM] Assume (ref), (ref), SUTVA, and finite second moments used to obtain the asymptotic representation for the subsample means estimator and add that the SRA estimator uses canonical link function, $m(\cdot )$, with a QMLE in the linear exponential family. If for some $\alpha _{g}^{\ast }$, $\bm{\beta _g}^{\ast }$, $\mathbb{E}\left[ Y(g)|\mathbf{X}\right]=m(\alpha _{g}^{\ast }+\mathbf{X}\bm{\beta_g}^{\ast })\text{, }g=1,...,G$ with probability one, then \begin{equation*} \mathrm{Avar}\left[ \sqrt{N}(\bm{\hat{\mu}}_{SM}-\bm{\mu })\right]-\mathrm{Avar}\left[ \sqrt{N}(\bm{\hat{\mu}}_{SRA}-\bm{\mu })\right] is PSD. \end{equation*}

The proof can be found in Appendix (ref). Even when the means are not correctly specified, it seems that separate nonlinear estimation should improve over the subsample means estimator if the $\mathbf{X}$ are sufficiently predictive. We see that in the simulation results in Section (ref). Finally, as with the linear case, if assignment is only unconfounded conditional on X, nonlinear SRA is consistent under correct specification of the conditional mean, whereas SM is not.

Pooled Regression Adjustment

In cases where $N$ is not especially large, one might, just as in the linear case, resort to a pooled method. Provided the mean/QLL\ combinations are chosen as in Table (ref), PRA is still consistent under arbitrary misspecification of the mean function. To see why, write the mean function with common slopes $\bm{\beta}$, and without an intercept in the index, as $m(\gamma _{1}W_{1}+\gamma _{2}W_{2}+\cdots +\gamma _{G}W_{G}+\mathbf{X}\bm{\beta})$. Using a similar argument as in the previous section, the first-order conditions of the pooled QMLE include the $G$ equations {

equation[equation omitted — 226 chars of source]

} Therefore, assuming no degeneracies, the probability limits of the estimators, indicated with a \textquotedblleft *\textquotedblright\ superscript, solve the population analogs of eq ((ref)): {

equation[equation omitted — 201 chars of source]

} where $\mathbf{W}=(W_{1},W_{2},...,W_{G})$. By random assignment, $\mathbb{E}\left[ W_{g}Y(g) \right] =\rho _{g}\mu _{g}$ and by iterated expectations {

equation[equation omitted — 270 chars of source]

} Using random assignment again, {

equation[equation omitted — 286 chars of source]

} Combining equations ((ref)) and ((ref)) gives $\mathbb{E}\left[ W_{g}m(\mathbf{W}\bm{\gamma}^{\ast }+\mathbf{X}\bm{\beta }^{\ast })\right] =\rho _{g}\mathbb{E}\left[ m(\gamma _{g}^{\ast }+\mathbf{X}\bm{\beta}^{\ast })\right]$. Now, using $\rho _{g}>0$ and $\mathbb{E}[W_{g}Y(g)]=\rho _{g}\mu _{g}$, we have shown that equation ((ref)) gives us {

equation[equation omitted — 125 chars of source]

} which shows that the $\mu _{g}$ are identified by the population mean of $m(\gamma _{g}^{\ast }+\mathbf{X}\bm{\beta}^{\ast })$ even when the conditional mean function is misspecified. Under weak regularity conditions, $\check{\gamma}_{g}$ is consistent for $\gamma _{g}^{\ast }$ and $\bm{\check{\beta}}$ is consistent for $\bm{\beta}^{\ast }$. Therefore, after the pooled QMLE estimation, we obtain the estimated means as $\check{\mu}_{g}=N^{-1}\sum_{i=1}^{N}m(\check{\gamma}_{g}+\mathbf{X}_{i}\bm{\check{\beta}}),$ and these are consistent [wooldridge2010econometric (Chapter 12, Lemma 12.1)]. As in the case of comparing nonlinear SRA to the SM, we have no general asymptotic efficiency results comparing nonlinear SRA to nonlinear PRA.

Semiparametric Efficiency Bound

In this section, we generalize the efficiency result with regression adjustment in experiments for multiple treatments by providing an efficiency benchmark for estimating the vector of PO means. We characterize the efficient influence function and the semiparametric efficiency bound (SEB) for all regular and asymptotically linear estimators of $\bm{\mu}$. The SEB is useful as it provides the semiparametric analogue of the Cramer-Rao lower bound for parametric models. It gives us a standard against which to compare the asymptotic variance of any regular estimator of the potential outcome means, including the RA estimators studied in the previous sections.

In our case, the most efficient estimator for $\bm{\mu}$ has the following asymptotically linear form

equation*[equation* omitted — 114 chars of source]

such that $\mathbb{E}[\bm{\psi}_i] = \bm{0}$ and $\mathbb{E}[\bm{\psi}_i\bm{\psi}_i^\prime]$ is finite and non-singular. Then, the SEB is given by $\mathbf{V_{\ast}} = \mathbb{E}[\bm{\psi}_i\bm{\psi}_i^\prime]$ where the influence function is

align[align omitted — 416 chars of source]

This result follows from cattaneo2010efficient, who studies efficient semiparametric estimation of treatment effects under the assumption of unconfoundedness. Given that we have the stronger assumption of randomized treatment, this essentially leads us to the same moment conditions as cattaneo2010efficient with the exception that the propensity score here is a constant.

theorem[Semiparametric efficiency of SRA] If the conditional means of the potential outcomes, $\mathbb{E}[Y(g)|\mathbf{X}]$ for each $g$, are correctly specified then SRA is semiparametrically efficient.

The proof can be found in Appendix (ref). The above result establishes that the separate regression adjustment estimator achieves the semiparametric efficiency bound if the true conditional mean of the POs is correctly specified. This implies that running separate linear regressions for different treatment level, as advocated in section (ref), is optimally efficient if the mean function, $m(\cdot)$, is truly linear. However, if $m(\cdot)$ is truly logistic or exponential, then nonlinear SRA discussed in section (ref), is optimally efficient for estimating the PO means.

Monte Carlo Simulations

This section studies the finite sample properties of the different estimators of the PO means, namely, SM, PRA, SRA, and their nonlinear counterparts. For the simulations, we generate a population of one million observations and mimic the asymptotic setting of random sampling from an “infinite” population. The empirical distributions of the RA estimators are simulated for sample sizes $N\in\{500, 1000, 5000\}$ by randomly drawing the data vector $\left\{ \left(\mathbf{W}_{i},Y_{i},\mathbf{X}_{i}\right) :i=1,2,...,N\right\}$, without replacement, ten thousand times from the population. We consider three different population models, each of which uses a unique data generating process for the potential outcomes, but we report results only for the linear model. Results for fractional and non-negative outcomes can be found in online appendix B.\footnote{See Tables B.1-B.4 for results on fractional and non-negative outcomes.}

To simulate multiple treatments, we consider potential outcomes, $Y(g)$, corresponding to three treatment states, $g=1, 2, 3$. Hence, $G=3$ for all population models. In each of the populations, the treatment vector $\mathbf{W}= \left(W_1, W_2, W_3\right)$ is generated with probability mass function defined by $\rho_g$ such that there is an equally likely probability of being assigned to a particular treatment group.

Population Models

To compare the empirical distributions of the RA estimators, we consider the first population model which simulates $Y(g)$ for each $g$ to be continuous with an unrestricted range using a linear specification. We consider two covariates, $\mathbf{X} =\left(X_1, X_2\right)$, where $X_1$ is continuously distributed whereas $X_2$ is binary. The covariates are generated as follows:

align*[align* omitted — 177 chars of source]

where $K_1, K_2 \sim N(0,2)$, and $V\sim N(0,1)$.

\paragraph{Population 1: }For each $g$, $Y(g) =\mathbf{\breve{X}}\bm{\gamma _{g}}+R(g)$ where $\mathbf{\breve{X}}=\left( 1,X_{1},X_{2},X_{1}\cdot X_{2}\right) $ and $\bm{\gamma}_{g}=\left( \gamma _{0g},\gamma _{1g},\gamma _{2g},\gamma_{3g}\right) ^{\prime }$. Moreover, $R(g)\lvert X_{1},X_{2}\sim N(0,\sigma_{g}^{2})$ and $R(1),R(2),\text{ and }R(3)$ are correlated.\footnote{$R(1)\sim N(0,1),R(2)=3/\sqrt{2}\cdot \lbrack R(1)+V_{1}],R(3)=1/\sqrt{2}\cdot \lbrack R(1)+V_{2}]$ where $V_{1}$ and $V_{2}$ are distributed $N(0,1)$.} The parameter vector, $\bm{\gamma}_{g}$, is chosen to depict settings where covariates are either mildly predictive ($R^{2}(g)<0.5$) or highly predictive ($R^{2}(g)>0.5$) of the potential outcomes. For example, the parameter vectors, $\bm{\gamma}_{g}^{L}$ and $\bm{\gamma}_{g}^{H}$ are chosen such that they lead to low and high $R$-squared settings, respectively.

equation*[equation* omitted — 462 chars of source]

While estimating the PO means, we assume that the true population model is unknown. With linear RA methods, we simply regress the observed outcome on a constant and the covariates. For nonlinear RA, based on the nature of the outcomes, we use a particular combination of log-likelihood and canonical link function from the LEF.

Discussion

Tables (ref) and (ref) report the bias, standard deviation, standard error, and 95% coverage probabilities for SM, PRA, and SRA across both high and low R-squared settings for population 1. Each table considers three different sample sizes. Note that in most cases, the bias of RA methods is comparable to SM, which is unbiased for the PO means. However, one may be willing to forego the bias in RA estimates in favor of efficiency and so we turn our attention to efficiency comparisons.

Across both low and high $R^2$ configurations, we see that regression adjustment typically improves over subsample means. The magnitude of such a precision gain with PRA depends on how predictive the covariates are of the potential outcomes. We see that this gain is generally smaller at low $R^2$ than high $R^2$ values. For instance, in Table 2, the standard deviations of the PRA estimates are between 0.2%-7.7% smaller, across the different sample sizes, compared to the SM estimator at low $R^2$ settings, whereas these differences are of order 13.6%-31.8% for the high $R^2$ setting (see Table 3). The difference in precision between PRA and SRA depends on how heterogeneous are the slope coefficients in the true linear projections. While the standard deviation of the SRA estimates are between 2.3%-6.9% smaller than those of PRA in the low $R^2$ case, these differences are of order 5.7%-19.4% for the high $R^2$ case. Note that comparisons made with standard error estimates give us similar ranges, since large sample approximations are quite good. For fractional and non-negative outcomes, nonlinear RA generally seems to improve over linear RA methods and SM (see Tables B.1-B.4 in the online appendix).

table[table omitted — 3,143 chars of source]
table[table omitted — 3,117 chars of source]

Application to California Oil Spill Study

This section applies regression adjustment to estimate the lower bound mean willingness-to-pay (WTP) using survey data from the California oil spill study that can be found in carson2004valuing.\footnote{See online appendix C on how RA methods can be applied to estimate the lower bound on mean WTP.} The study implements a contingent valuation survey to assess the value of damages to natural resources from future oil spills along California's central coast. The survey provides respondents with the choice of voting for or against a governmental program that would prevent natural resource injuries to shorelines and wildlife along the coast over the next decade. In return, the public is asked to pay a one-time lump-sum income tax surcharge for setting up the program.

The main survey sample used to elicit the yes or no votes was conducted by Westat, Inc. The data are a random sample of 1,085 interviews conducted with English speaking households where the respondent is 18 years or older and lives in private residence that is either owned or rented. Each respondent is randomly assigned one of five tax amounts: \$5, \$25, \$65, \$120, or \$220 and a choice of \textquotedblleft yes\textquotedblright\ or \textquotedblleft no\textquotedblright\ is recorded at the assigned amount.

The survey also collects data on five major characteristics of the respondents' household which are economic, demographic, preferences and attitudes towards the environment, interest in and use of the affected natural resources, evaluations of the expected harm and prevention program, and interpretations of the payment mechanism. The economic and demographic variables include log-income of the household, whether household pays state taxes in 1994, and resides in the central coast primary sampling unit. Preference towards environment include indicator variables reflecting the importance of preventing oil spills in coastal areas, spending to protect wildlife, and whether respondents consider themselves to be environmentalists or environmental activists. The attitudinal characteristics relate to respondents' attitudes towards government programs such as whether people consider spending to be important and whether respondents consider taxes to be the appropriate payment method for protecting the environment. Variables reflecting interest and use of the affected natural resources are: driving along the central coast on highway 1 and familiarity with at least one of the five species of birds most often harmed by past oil spills. Variables measuring respondents' evaluations of the expected harm and prevention program include identifying people who think oil spills over the next decade would cause more harm than mentioned in survey, who think oil spills would cause less harm, those who believe in the program's effectiveness in achieving the desired goal, and those who have concerns regarding program's effectiveness. The final set of variables relate to respondents' interpretation of the payment mechanism, such as whether respondents believe the tax to not be limited to one year and those who protested that either the oil companies should pay for the program or those who thought that the companies would pass the costs to consumers in the form of higher gas and oil prices.

Table (ref) reports the proportion of respondents randomly assigned to the different bid amounts where we see approximately the same number of people at each bid value. Finally, Table (ref) provides estimates of the PO means using linear and nonlinear RA estimators. These control for the respondent characteristics described above. We then use the PO mean estimates to obtain a lower bound on mean WTP.

In terms of the standard errors, we see a ranking among the linear RA estimators. Standard errors for the linear SRA estimator are between 0.3%-3% smaller than those for linear PRA estimates. Except in one case, nonlinear SRA standard errors are between 0.5%-1.5% smaller than for the linear SRA ones. Finally, nonlinear PRA standard errors are between 0%-5% smaller than those for the nonlinear SRA estimates. In this application, using pooled logistic regression produces the most efficiency gains over the usual SM estimator, also known as ABERS estimator introduced by ayer1955empirical in contingent valuation studies--see the online appendix for reference.

table[table omitted — 561 chars of source]
table[table omitted — 2,383 chars of source]

Conclusion

Building on the binary treatment case in NW (2021), we study efficiency improvements with RA when there are more than two treatment levels. In particular, we consider the case of random assignment of $G$ treatment levels. We show that jointly estimating the vector of potential outcome means using linear SRA, which allows for separate slopes for the different assignment levels, is asymptotically never worse than just using subsample averages; this result improves on the earlier work even when $G=2$. One case when there is no gain in asymptotic efficiency from using SRA is when the slopes are all zero. In other words, when the covariates are not predictive of the potential outcomes, then using separate slopes does not produce more precise estimates than subsample averages. We also show that SRA is generally more efficient compared to PRA, unless the slopes in true linear projections are homogeneous. In this case, using SRA to estimate the vector of PO means is harmless.

In addition, we also extend the nonlinear RA results in NW (2021) to multiple treatment levels. We show that nonlinear RA of the QML variety is consistent if one chooses the conditional mean and objective functions appropriately from the linear exponential family of quasi-likelihoods. Furthermore, we also characterize the semiparametric efficiency bound for estimating the vector of PO means and show that SRA attains this bound, when the conditional means are correctly specified.

In simulations, we find that RA can provide substantial improvements over subsample means. Naturally, the magnitude of this precision gain generally depends on how strongly covariates predict the potential outcomes. As an illustration, we apply the different RA estimators to estimate the lower bound mean willingness to pay for a contingent valuation study which randomized different bid amounts to measure the cost of oil spills along California's coast. We find that the lower bound is estimated more efficiently when we use SRA rather than the commonly used ABERS estimator, which uses subsample averages for the PO means.

\singlespacing