EconBase
← Back to paper

Semiparametrically efficient estimation of the average linear regression function

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

97,893 characters · 17 sections · 119 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Semiparametrically efficient estimation of the average linear regression function

singlespacing\begin{center} Bryan S. Graham$^{+}$ and Cristine Campos de Xavier Pinto$^{\#}$\footnote{{$^{+}$Department of Economics, University of California - Berkeley, 530 Evans Hall \#3380, Berkeley, CA 94720-3888 and National Bureau of Economic Research, }{\uline{e-mail:}}{ \href{http://[email removed]}{[email removed]}, }{\uline{web:}}{ \url{http://bryangraham.github.io/econometrics/}. }\\ {$^{\#}$Escola de Economia de Sao Paulo, FGV, Rua Itapeva 474, sala 1010, CEP: 01332-000. }{\uline{e-mail:}}{ $\mathtt{[email removed]}$. }{\uline{web:}}{ $\mathtt{http://sites.google.com/site/cristinepinto/}$.}\\ {We thank Guido Imbens, Pat Kline, Tony Strittmatter and seminar participants at University College London, UC Berkeley and University of St. Gallen for helpful discussion. Financial support from NSF grant SES \#1357499 is gratefully acknowledged. The initial draft of this paper was prepared in October of 2016. All the usual disclaimers apply.}} \end{center} \begin{abstract} Let $Y$ be an outcome of interest, $X$ a vector of treatment measures, and $W$ a vector of pre-treatment control variables. Here $X$ may include (combinations of) continuous, discrete, and/or non-mutually exclusive “treatments”. Consider the linear regression of $Y$ onto $X$ in a subpopulation homogenous in $W=w$ (formally a conditional linear predictor). Let $b_{0}\left(w\right)$ be the coefficient vector on $X$ in this regression. We introduce a semiparametrically efficient estimate of the average $\beta_{0}=\mathbb{E}\left[b_{0}\left(W\right)\right]$. When $X$ is binary-valued (multi-valued) our procedure recovers the (a vector of) average treatment effect(s). When $X$ is continuously-valued, or consists of multiple non-exclusive treatments, our estimand coincides with the average partial effect (APE) of $X$ on $Y$ when the underlying potential response function is linear in $X$, but otherwise heterogenous across agents. When the potential response function takes a general nonlinear/heterogenous form, and $X$ is continuously-valued, our procedure recovers a weighted average of the gradient of this response across individuals and values of $X$. We provide a simple, and semiparametrically efficient, method of covariate adjustment for settings with complicated treatment regimes. Our method generalizes familiar methods of covariate adjustment used for program evaluation as well as methods of semiparametric regression (e.g., the partially linear regression model). \uline{JEL Codes:} C14, C21, C31 \uline{Keywords:} Conditional Linear Predictor, Causal Inference, Average Treatment Effect, Propensity Score, Semiparametric Efficiency, Semiparametric Regression \end{abstract}

\thispagestyle{empty}

\setcounter{page}{1}

\sloppy

Let $Y$ be a scalar-valued outcome of interest, $X$ a $K\times1$ vector of policy variables, and $W$ a $J\times1$ vector of additional controls. For example $Y$ might equal hours worked, $X$ include the real wage rate and total unearned income ($K=2$), and $W$ be a vector of demographic measures capturing heterogeneity in preferences for work Pencavel_HLE86. The goal is to summarize how $Y$ \textendash labor supply \textendash covaries with $X$ \textendash the wage rate and unearned income \textendash “holding the controls $W$ fixed”. In a second example, $Y$ might be an end-of-year student mathematics achievement measure, $X$ a vector containing (i) number of days absent from school, (ii) class size and (iii) an indicator for whether the student received supplemental tutoring. Here the vector $W$ might include beginning of school year joint predictors of $Y$ and $X$ (e.g., prior mathematics achievement, socioeconomic background, health indicators, and known determinants of class size and tutoring assignment used by the school). The goal is to summarize how math achievement covaries with attendance, class size and supplemental tutoring conditional on $W$ Gottfried_Kirksey_ER17.

Following the prototype established by Yule_JRSS1899 over one hundred years ago, social scientists typically report the coefficient on $X$ in the (long) least squares fit of $Y$ onto a constant, $X$, and $W$ for this purpose.

When $X$ is a scalar binary variable, the econometrician can choose from \textendash in addition to least squares \textendash an ever more elaborate menu of covariate adjustment methods (see Imbens_Rubin_CIBook15 for a recent textbook introduction). Many of these methods extend naturally to settings where $X$ is multi-valued Cattaneo_JOE10.

When $X$ is continuously-valued, and/or consists of multiple distinct policy variables ($K\geq2$), options are fewer Wooldridge_EACSPDBook10. The partially linear regression (PLM) model

equation[equation omitted — 139 chars of source]

represents one semiparametric generalization of (long) linear regression. Chamberlain_WP86, in an influential but never published paper, introduced an estimator for $\beta_{0}$ in ((ref)) Robinson_EM88. In later work he characterized its semiparametric efficiency bound (SEB) Chamberlain_EM92.

Partially linear regression is widely, albeit heuristically, used in empirical work. Typically researchers proceed by (i) choosing $W$ to be a rich vector of basis functions in the underlying controls (e.g., a vector of polynomial or piecewise polynomial terms) and then (ii) estimate $\beta_{0}$ by least squares. With discretely-valued control variables a saturated specification for $h_{0}\left(W\right)$ is possible, at least when utilizing a very large dataset Angrist_Krueger_HLE99. A principled variant of this general approach is embodied in the E-Estimation algorithm of Newey_JAE90 and Robins_Mark_Newey_BM92.

In this paper we propose a different approach to covariate adjustment. Consider a subpopulation homogenous in $W=w$. Within this subpopulation we compute the linear regression of $Y$ onto a constant and $X$ (formally a conditional linear predictor as in Wooldridge_JE99). Let $b_{0}\left(w\right)$ be the coefficient on $X$ in the conditional linear regression for the subpopulation homogenous in $W=w$. We propose a method for identifying and efficiently estimating the average regression coefficient

equation[equation omitted — 93 chars of source]

The average is over the marginal distribution of controls, $W$.

In the absence of controls, the relationship between the linear predictor slope coefficient and the gradient of the (possibly nonlinear) conditional expectation function (CEF) of $Y$ given $X=x$ is well-understood Goldberger_ACE91,Yitzhaki_JBES96. In the presence of controls, this relationship is rather more complicated Angrist_EM98,Sloczynski_WP18. Our focus on averages of conditional linear predictor coefficients allows for conditioning on $W$, while also preserving the interpretative transparency of unconditional linear analyses. That is, $\beta_{\mathrm{0}}$, as we demonstrate below, is easy to interpret.

When $X$ is binary-valued (multi-valued) $\beta_{0}$ coincides with the (a vector of) average treatment effect(s); estimands familiar from the program evaluation literature Hahn_EM98,Imbens_BM00. These estimands have causal interpretations under certain conditions. Modestly extending the analysis of Wooldridge_CWP04, we show that this causal interpretation generalizes under a (i) heterogenous random coefficients potential outcome structure and (ii) an unconfoundedness-type assumption. These assumptions coincide with their program evaluation counterparts when $X$ is binary- or multi-valued. Our semiparametric model includes both the program evaluation model and the partially linear regression model as special cases.

Our work is also connected to the varying coefficient model of Hastie_Tibshirani_JRSS93. Hastie_Tibshirani_JRSS93 focus on pointwise estimation of $b_{0}\left(w\right)$, while we focus on (efficient) estimation of the average $\beta_{0}=\mathbb{E}\left[b_{0}\left(W\right)\right]$.

The relationship of our work with that of Wooldridge_CWP04 is as follows.\footnote{The Wooldridge_CWP04 paper remains unpublished, but a textbook treatment of the material in it can be found in Chapter 21.6.3 of Wooldridge_EACSPDBook10.} We both study the same functional of the joint distribution of $W$, $X$ and $Y$ (see Equation ((ref)) below). Relative to Wooldridge_CWP04 we provide an average partial effect interpretation of this estimand under (i) weaker assumptions when maintaining a correlated random coefficient potential outcome structure and (ii) a new weighted average partial effect interpretation under a general potential response function structure. These are useful, but relatively modest generalizations. More significantly we (i) provide distribution theory for the estimator proposed by Wooldridge_CWP04, (ii) characterize the semiparametric efficiency bound (SEB) for $\beta_{0}$, and (iii) introduce a new locally efficient estimator. The procedure proposed by Wooldridge_CWP04 is inefficient.

Another feature of our estimator is computational simplicity. Let $\hat{\mu}_{W}=\frac{1}{N}\sum_{i=1}^{N}W_{i}$ be the sample mean of $W$. A common approach to modeling heterogeneous effects in applied work is to compute the least squares fit of $Y$ onto a constant, $W-\hat{\mu}_{W}$, $\left(W-\hat{\mu}_{W}\right)\otimes X$, and $X$. As is well-known from textbook treatments on interaction terms in linear regression analysis, centering the control variable vector, $W$, about is mean in this way ensures that the coefficient on $X$ captures an average effect. This approach essentially coincides with Oaxaca-Blinder type methods of covariate adjustment popular in labor economics Kline_EL14. One variant of our procedure involves computing the exact same regression, but where $X$ is instead instrumented with a particular function of its conditional distribution given $W$ (i.e., of the “generalized” propensity score). Theorems (ref) and (ref) below show that this small modification to a familiar estimation procedure delivers considerable gains.

The next section introduces our average linear regression model. We provide a statistical definition of $\beta_{\mathrm{0}}$ as well as sets of assumptions under which it has a causal \textendash average partial effect (APE) \textendash interpretation. Section (ref) presents the semiparametric efficiency bound for $\beta_{0}$. Section (ref) studies the large sample properties of the Wooldridge_CWP04 estimator. We also introduce our new estimator and present its large sample properties. Finally, in Section (ref), we connect our results with prior work on efficient estimation of average treatments effects as well as the partially linear semiparametric regression model. We end our paper with a small simulation study in Section (ref). All proofs are collected in the Appendix or the supplemental materials.

Average linear regression model

We begin with a conventional sampling assumption.

assumption(Random Sampling) Let $\left\{ \left(W_{i}',X_{i}',Y_{i}\right)'\right\} _{i=1}^{\infty}$ be a sequence of independent and identically distributed random draws from some population $F_{W,X,Y}$ with $\mathbb{E}\left[\left.Y^{2}\right|W=w\right]<\infty$ and $\mathbb{E}\left[\left.\left\Vert X\right\Vert ^{2}\right|W=w\right]<\infty$ for all $w\in\mathbb{W}.$

The finite moment restrictions included in Assumption (ref) ensure that a conditional linear predictor (CLP) is well-defined for all $w\in\mathbb{W}$.

Let

equation[equation omitted — 92 chars of source]

be the conditional mean of $X$ given $W=w$ and

equation[equation omitted — 92 chars of source]

the corresponding conditional variance. We also require that $X$ vary conditional on $W=w$.

assumption(Overlap) For all $w\in\mathbb{W}$ and any non-zero column vector $t$, $t'v_{0}\left(w\right)t\geq\kappa>0.$

Assumption (ref) ensures that the CLP is uniquely defined. In the absence of conditioning it is equivalent to linear independence of the elements of $X$. When $X$ is binary $v_{0}\left(W\right)=e_{0}\left(W\right)\left(1-e_{0}\left(W\right)\right)$ with $e_{0}\left(W\right)=\Pr\left(\left.X=1\right|W\right)$ equal to the propensity score; in this case Assumption (ref) coincides with the familiar strong overlap assumption from the program evaluation literature. More generally Assumption (ref) implies that $X$ varies conditional on $W=w$ for all $w\in\mathbb{W}.$

Under Assumptions (ref) and (ref) the conditional linear predictor is well-defined for all $w\in\mathbb{W}$. Wooldridge_JE99 provides a self-contained introduction to conditional linear predictors. The following definition and lemma is taken from Wooldridge_JE99.

defn(Conditional Linear Predictor) The mean squared error minimizing linear predictor of $Y$ given $X$ conditional on $W=w$, henceforth the conditional linear predictor (CLP), equals \begin{equation} \mathbb{E}^{*}\left[\left.Y\right|X;W=w\right]\overset{def}{\equiv}a_{0}\left(w\right)+X'b_{0}\left(w\right) \end{equation} with \begin{align} a_{0}\left(w\right) & \overset{def}{\equiv}\mathbb{E}\left[\left.Y\right|W=w\right]-e_{0}\left(w\right)'b_{0}\left(w\right)\\ b_{0}\left(w\right) & \overset{def}{\equiv}v_{0}\left(w\right)^{-1}\mathbb{C}\left(\left.X,Y\right|W=w\right).\nonumber \end{align} It is straightforward to show that the prediction error $U=Y-\mathbb{E}^{*}\left[\left.Y\right|X;W\right]$ is conditionally mean zero and conditionally uncorrelated with $X$. This property of $\mathbb{E}^{*}\left[\left.Y\right|X;W\right]$ will prove useful for what follows.
lemWooldridge_JE99. Let $U\overset{def}{\equiv}Y-a_{0}\left(W\right)-X'b_{0}\left(W\right)$, then $\mathbb{E}\left[\left.U\right|W=w\right]=0$ and $\mathbb{E}\left[\left.XU\right|W=w\right]=0$ for all $w\in\mathbb{W}.$

Identification of the average regression slope

We begin by presenting a convenient representation of the average slope coefficient $\beta_{\mathrm{0}}=\mathbb{E}\left[b_{0}\left(W\right)\right]$ in terms of the joint distribution of $\left(W',X',Y\right)'$. The most direct representation follows directly from ((ref)): \[ \beta_{\mathrm{0}}=\mathbb{E}\left[v_{0}\left(W\right)^{-1}\mathbb{C}\left(\left.X,Y\right|W\right)\right]. \] For our purposes, however, an alternative representation of $\beta_{\mathrm{0}}$ is more convenient; both for our semiparametric efficiency bound (SEB) analysis and for the approach to estimation developed below. Using the law of iterated expectations and the definition of conditional covariance we get, under Assumptions (ref) and (ref),

align*[align* omitted — 314 chars of source]

Applying definition ((ref)) then gives our preferred estimand representation:

align[align omitted — 136 chars of source]

Wooldridge_CWP04 emphasizes the coincidence between ((ref)) and the average partial effect of $X$ on $Y$ associated with a particular correlated random coefficients (CRC) potential outcomes structure. This endows $\beta_{0}$ with causal meaning. While we also develop this connection below, we wish to initially emphasize that ((ref)) is also just one way of representing a population average of conditional linear predictor coefficients. Under Assumptions (ref) and (ref) the expectation in ((ref)) is well-defined and $\beta_{0}$ is simply a “statistical” estimand. We are interested in estimating it as precisely as possible.

Causal interpretation

In this subsection we show that ((ref)) admits a causal interpretation under a particular treatment response model and selection on observables type assumption. As noted earlier, this interpretation was previously emphasized by Wooldridge_CWP04, but under stronger conditions than we maintain here.

Associated with each agent in the target population is an individual-specific potential response function, $Y\left(x\right)$, which maps counterfactual values of the input vector $X$ into their corresponding (potential) outcomes. The observed outcome coincides with the value of the potential response function at the observed input level $X$: $Y=Y\left(X\right)$. We assume that $Y\left(x\right)$ is linear in $x$, but otherwise heterogeneous across individuals:

equation[equation omitted — 72 chars of source]

where $A$ and $B$ are an individual-specific intercept and slope vector respectively.

Equation ((ref)) allows for each individual to have their own potential response function, but restricts them to be linear in $X$. When $X$ is binary, or multi-valued, linearity is unrestrictive. For example, in the binary case, we have the potential outcome under control ($X=0$) and active $(X=1$) treatment equal to $Y\left(0\right)=A$ and $Y\left(1\right)=A+B$. In the multi-valued treatment setting of Imbens_BM00 and Cattaneo_JOE10, with $X$ a vector of treatment indicators for $K$ mutually exclusive treatments, we have $Y\left(0\right)=A$ and $Y\left(k\right)=A+B_{k}$ for $k=1,\ldots,K$. In contrast, when $X$ is ordered, continuously valued, or includes multiple treatments/policies, linearity is restrictive.

Consider the following thought experiment: draw a unit at random and (exogenously) increase the value of the $k^{th}$ component of $X$ by one unit. The expected effect of this intervention is $\mathbb{E}\left[B_{k}\right]$. In the binary- and multi-valued treatment setting $\mathbb{E}\left[B_{k}\right]$ corresponds to an average treatment effect (ATE) \[ \mathbb{E}\left[B_{k}\right]=\mathbb{E}\left[Y\left(k\right)-Y\left(0\right)\right]. \] More generally $\mathbb{E}\left[B_{k}\right]$ equals the average partial effect (APE) of a unit increase in $X_{k}$. This estimand was introduced in a panel data setting by Chamberlain_HBE84; general expositions, with additional results, are available in Blundel_Powell_WC03 and Wooldridge_IIEM05.

Under the following assumption, in addition to those introduced above, we can show that $\beta_{0}$ coincides with the APE vector, $\mathbb{E}\left[B\right]$.

assumption(Conditional Exogeneity) For all $w\in\mathbb{W}$ and $k,l=1,\ldots,K$, and under potential responses of the form given in ((ref)) \begin{equation} \mathbb{C}\left(\left.A,X_{k}\right|W=w\right)=\mathbb{C}\left(\left.B,X_{k}\right|W=w\right)=\mathbb{C}\left(\left.B,X_{k}X_{l}\right|W=w\right)=0. \end{equation}

Assumption (ref) restricts the form of any dependence between the potential response function, $Y\left(x\right)=A+x'B$, and the treatment vector actually chosen by the respondent, $X$. It is a conditional exogeneity or selection on observables type assumption. To see this observe that when $X$ is binary Assumption (ref) coincides with the standard mean independence assumption familiar from the program evaluation literature, implying that \[ \mathbb{E}[Y(x)|X,W]=\mathbb{E}[Y(x)|W]. \] In the multi-valued treatment setting Assumption (ref) also coincides with standard generalizations of the mean independence assumption Imbens_BM00. See also Section (ref) below.

When the linearity of ((ref)) is restrictive, as occurs when $X$ includes continuously-valued components, or non-mutually exclusive binary inputs, Assumption (ref) is less restrictive than other possible formulations of conditional exogeneity. For example, Wooldridge_CWP04,Wooldridge_EACSPDBook10 works with the identifying restrictions

equation[equation omitted — 244 chars of source]

which imply ((ref)), but are generally stronger. An even stronger notion of conditional exogeneity is

equation[equation omitted — 223 chars of source]

Assumption (ref) is (apparently) the weakest assumption necessary to equate $\beta_{0}$ with the average partial effect of $X$ on $Y$ when the potential response function takes form ((ref)). The estimator we introduce below will remain consistent under the stronger restrictions, ((ref)) and ((ref)), but will generally not be semiparametrically efficient in those cases. We elaborate further on this observation below.

prop(Average Partial Effect Identification) Under Assumptions (ref), (ref) and (ref) the average of the CLP coefficients, $\beta_{0}=\mathbb{E}\left[b_{0}\left(W\right)\right]$, and the average partial effect (APE), $\mathbb{E}\left[B\right]$, coincide: \[ \beta_{0}=\mathbb{E}\left[B\right]. \]
proofWooldridge_CWP04 demonstrates the equality under the stronger restriction ((ref)). Under Assumption (ref), however, the proof proceeds differently. Given the linear potential response ((ref)) and by lemma ((ref)), we have the $1+K$ conditional moment restrictions \begin{eqnarray} \mathbb{E}\left[\left.U\right|W=w\right] & = & \mathbb{E}\left[\left.A-a_{0}\left(W\right)\right|w\right]+\mathbb{E}\left[\left.X'\left(B-b_{0}\left(W\right)\right)\right|w\right]=0\nonumber \\ \mathbb{E}\left[\left.XU\right|W=w\right] & = & \mathbb{E}\left[\left.X\left(A-a_{0}\left(W\right)\right)\right|w\right]+\mathbb{E}\left[\left.XX'\left(B-b_{0}\left(W\right)\right)\right|w\right]=0. \end{eqnarray} Under Assumption (ref) conditions ((ref)) simplify to \begin{align*} \left\{ \mathbb{E}\left[\left.A\right|w\right]-a_{0}\left(w\right)\right\} +e_{0}\left(w\right)'\left\{ \mathbb{E}\left[\left.B\right|w\right]-b_{0}\left(w\right)\right\} & =0\\ e_{0}\left(w\right)\left\{ \mathbb{E}\left[\left.A\right|w\right]-a_{0}\left(w\right)\right\} +\mathbb{E}\left[\left.XX'\right|w\right]\left\{ \mathbb{E}\left[\left.B\right|w\right]-b_{0}\left(w\right)\right\} & =0 \end{align*} or, in matrix form, \[ \left[\begin{array}{cc} 1 & e_{0}\left(w\right)'\\ e_{0}\left(w\right) & \mathbb{E}\left[\left.XX'\right|w\right] \end{array}\right]\left(\begin{array}{c} \mathbb{E}\left[\left.A\right|w\right]-a_{0}\left(w\right)\\ \mathbb{E}\left[\left.B\right|w\right]-b_{0}\left(w\right) \end{array}\right)=\left(\begin{array}{c} 0\\ 0 \end{array}\right). \] Under the Assumption (ref) the first matrix to the left of the equality is invertible for all $w\in\mathbb{W}$. This implies that $\mathbb{E}\left[\left.A\right|W=w\right]=a_{0}\left(w\right)$ and $\mathbb{E}\left[\left.B\right|W=w\right]=b_{0}\left(w\right)$ for all $w\in\mathbb{W}$. The result follows by iterated expectations.

Causal interpretation under misspecification

Angrist_Krueger_HLE99 and Angrist_Pischke_MHE09 emphasize that when $X$ is a continuously-valued random variable its slope coefficient in the linear predictor of $Y$ onto a constant, $X$ and the vector of “saturated” controls admits a weighted average derivative interpretation when the potential response function takes a general nonlinear form Angrist_Graddy_Imbens_ReStud00. Angrist and Krueger's Angrist_Krueger_HLE99 expression is also isomorphic to the probability limit of the E-Estimator of Newey_JAE90 and Robins_Mark_Newey_BM92

equation[equation omitted — 182 chars of source]

when the partially linear regression structure, equation ((ref)) above, is incorrect.

In this section, using similar arguments to those appearing in Angrist_Graddy_Imbens_ReStud00 and Graham_Imbens_Ridder_NBER10, we provide a representation result for $\beta_{0}$ under a general potential response function.

Assume that the potential response function is nonlinear and heterogeneous such that $Y\left(x\right)=h\left(x,U\right)$. Further assume, stronger than Assumption (ref) above, that $U$ is conditionally independent of $X$ given $W=w$ for all $w\in\mathbb{W}$. Blundel_Powell_WC03 show that the partial mean $\mathbb{E}_{W}\left[\mathbb{E}\left[\left.Y\right|W,X=x\right]\right]$ identifies the average structural function (ASF) $m\left(x\right)=\mathbb{E}_{U}\left[h\left(x,U\right)\right]$ when the support of $W$ given $X=x$ coincides with its marginal support. Newey_ET94b provides an explicit partial mean estimator and derives in asymptotic properties.

Here we show that our average regression slope estimand, $\beta_{0}$, can be expressed as a weighted average of the gradient of $h\left(X,U\right)$. This provides a causal interpretation of $\beta_{0}$ under a general potential response function. To present this result we replace Assumption (ref) with:

assumption(Nonlinear Potential Response Function) (i) $X$ is a continuous scalar random variable with bounded support $\mathbb{X}=\left[\underline{x},\overline{x}\right]$, (ii) the conditional density function of $X$ given $W=w$ is bounded and bounded away from zero for all $\left(w,x\right)\in\mathbb{W}\times\mathbb{X}$, (iii) $Y=h\left(X,U\right)$ with $h\left(x,u\right)$ a continuously differentiable function of $x$ for all $\left(x,u\right)\in\mathbb{X}\times\mathbb{U}$ and $\underline{h}\left(u\right)=h\left(\underline{x},u\right)$ finite for all $u\in\mathbb{U}$, and (iv) $U$ is conditionally independent of $X$ given $W=w$ for all $w\in\mathbb{W}$.
prop(Weighted Average Derivative Representation) Under Assumptions (ref), (ref) and (ref) \[ \beta_{0}=\mathbb{E}\left[\omega\left(W,X\right)\frac{\partial h\left(X,U\right)}{\partial x}\right] \] where \[ \omega\left(w,x\right)=\frac{1}{f_{\left.X\right|W}\left(\left.x\right|w\right)}\frac{\mathbb{E}\left[\left.X-e_{0}\left(W\right)\right|W=w,X\geq x\right]\left(1-F_{\left.X\right|W}\left(\left.x\right|w\right)\right)}{\int_{\underline{x}}^{\bar{x}}\mathbb{E}\left[\left.X-e_{0}\left(W\right)\right|W=w,X\geq v\right]\left(1-F_{\left.X\right|W}\left(\left.v\right|w\right)\right)\mathrm{d}v}. \]
proofSee the Supplemental Web Appendix.

A key feature of the weighting function $\omega\left(w,x\right)$ is that its conditional mean, $\mathbb{E}\left[\left.\omega\left(W,X\right)\right|W=w\right]$, equals $1$ for every value of $w\in\mathbb{W}$. Furthermore, Lemma A.1 of Graham_Imbens_Ridder_NBER10 implies that, conditional on $W=w$, the weight given to $\frac{\partial h\left(X,U\right)}{\partial x}$ is highest for those values of $X$ near its conditional mean, $\mathbb{E}\left[\left.X\right|W=w\right]$, and lowest for those at the boundary of its support, $\underline{x}$ and $\overline{x}$.

These features of the weights appearing in Proposition (ref) imply the following intuitive interpretation: (i) for each value of $w\in\mathbb{W}$ compute a weighted average of $\frac{\partial h\left(X,U\right)}{\partial x}$, where the average emphasizes values of $X$ near its conditional mean given $W=w$, (ii) average these (weighted average) gradients over the marginal distribution of $W$. This indicates that $\beta_{0}$ only differs from the unweighted average $\mathbb{E}\left[\frac{\partial h\left(X,U\right)}{\partial x}\right]$ due to variation in $\omega\left(W,X\right)$ within $W=w$ cells. The contribution of each subpopulation, defined in terms of the control, $W$, mirrors its density in the sampled population. Since $W$ proxies for $U$ in this set-up we are averaging over the correct heterogeneity distribution.

More precisely, since $\mathbb{E}\left[\left.\omega\left(W,X\right)\right|W=w\right]=1$, we have that, using the definition of conditional covariance,

equation[equation omitted — 246 chars of source]

The bias of $\beta_{0}$ for $\mathbb{E}\left[\frac{\partial h\left(x,U\right)}{\partial x}\right]$ is therefore solely due to conditional covariance between the weight function and the gradient of interest within subpopulations homogenous in $W$.

In contrast to the one for $\beta_{0}$, the weight function appearing in the weighted average derivative representation result of Angrist_Krueger_HLE99 or Angrist_Pischke_MHE09 for $\beta_{\mathrm{E}}$ is only unconditionally mean zero. This implies that $\beta_{\mathrm{E}}$ averages over the incorrect heterogeneity distribution as well as the incorrect policy variable distribution.

align[align omitted — 433 chars of source]

If the ultimate object of interest is the average derivative $\mathbb{E}\left[\frac{\partial h\left(X,U\right)}{\partial x}\right]$, then, relative to ((ref)), a focus on $\beta_{0}$ eliminates one source of potential bias. Namely that the weight function may over- or under-emphasize various subpopulations defined in terms of their value of the control variable vector $W$. In this case $\mathbb{E}\left[\left.\omega\left(W,X\right)\right|W\right]$ may not equal one and the second term to the right of the equality in ((ref)) may be non-zero.\footnote{To be clear $\omega\left(W,X\right)$ are different functions in expressions ((ref)) and ((ref)); for its form in the latter case see Angrist_Krueger_HLE99 or Angrist_Pischke_MHE09.}

Motivating $\beta_{0}$

Our focus on averages of conditional linear predictor slope coefficients is motivated by a combination of principled and pragmatic reasons.

First, the kitchen sink long regression remains a workhorse of everyday empirical social science research. Our model extends kitchen sink regression in an easy to understand way. Relative to the partially linear regression model, our model allows for heterogenous responses of $Y$ to variation in $X$; a feature likely to be both empirically relevant and a priori attractive to researchers.

Second, $\beta_{0}$ has a causal interpretation under additional assumptions. When the potential response function is linear, but heterogeneous across agents, it coincides with an average partial effect (APE) under a selection on observables type assumption. When $X$ is binary- or multi-valued, as in the program evaluation literature, it coincides with the well-known average treatment effect (ATE). Our causal model nests the usual one as a special case, but accommodates continuous and/or multiple treatments as well (albeit under restrictions).

Third, in the presence of misspecification $\beta_{0}$ coincides with a weighted average of the derivative of a general non-linear potential response function. This weighted average derivative is more interpretable than existing representation results; for example those of Angrist_Krueger_HLE99 for $\beta_{\mathrm{E}}$.

Fourth, as we show next, $\beta_{0}$ is $\sqrt{N}$ estimable (or regularly identified). This is not the case for, say, a partial mean with a continuous policy variable Newey_ET94b. Regular identification suggests that estimation is practically feasible and we present one such feasible estimator below.

Ultimately the balance between ease of interpretation under various population assumptions and, as we show below, ease of estimation, provides the strongest case for focusing on $\beta_{0}$.

Semiparametric efficiency bound

Using the method of calculation outlined by Bickel_et_al_Bk93 and Newey_JAE90, we derive the semiparametric variance bound for $\beta_{0}$ of,

equation[equation omitted — 137 chars of source]

where \[ \Omega_{0}(w)=\mathbb{E}\left[\left.v_{0}(W)^{-1}\left(X-e_{0}\left(W\right)\right)UU'\left\{ v_{0}(W)^{-1}\left(X-e_{0}\left(W\right)\right)\right\} '\right|W=w\right]. \]

The corresponding efficient influence function equals

align[align omitted — 340 chars of source]

with $Z=\left(W',X',Y\right)'$, $g\left(W\right)=\left(e(W),v(W)\right)$ and $h\left(W\right)=\left(a\left(W\right),b\left(W\right)\right)$.

thm(Semiparametric Efficiency Bound) The efficient influence function for $\beta_{\mathrm{0}}=\mathbb{E}\left[b_{0}\left(W\right)\right]$ in the semiparametric problem established by Definition (ref) and Assumptions (ref) and (ref) equals ((ref)).
proofSee Appendix (ref). We also have the following corollary, which is similar to a result for the binary case due to Robins_Rotnitzky_Zhao_JASA94, Hahn_EM98 and Chen_Hong_Tarozz_AS08. This corollary will be useful when we discuss locally efficient estimation in Section ((ref)).
cor(Redundancy) Let $f\left(\left.x\right|w;\phi\right)$ be a parametric family of conditional distributions for $X$ given $W$ with $f_{0}\left(\left.x\right|w\right)=f\left(\left.x\right|w;\phi\right)$ at some unique $\phi=\phi_{0}$. The knowledge that $f_{0}\left(\left.x\right|w\right)$ is a member of the family $f\left(\left.x\right|w;\phi\right)$ does not change the semiparametric efficiency bound for $\beta_{0}.$
proofSee the Supplemental Web Appendix.

See Frolich_ER04 and Graham_Pinto_Egel_JBES16 for additional intuition about results like Corollary (ref).

Double robustness property of the efficient influence function

Before introducing our estimator in the next section we highlight an important property of the efficient influence function for $\beta_{\mathrm{0}}$.

Consider replacing $h_{0}\left(W\right)=\left(a_{0}\left(W\right),b_{0}\left(W\right)\right)$ in ((ref)) with the incorrect conditional linear predictor coefficients $h_{*}\left(W\right)=\left(a_{*}\left(W\right),b_{*}\left(W\right)\right)$. Use the notation $U_{*}=\left(Y-a_{*}\left(W\right)-X'b_{*}\left(W\right)\right)$ to emphasize that $U_{*}$ is the prediction error associated with an arbitrary conditional linear predictor (which need not be the mean squared error minimizing one). Note that $U_{*}$ will not be conditionally mean zero or conditionally uncorrelated with $X$ (i.e., $\mathbb{E}\left[\left.U_{*}\right|W\right]\neq0$ and $\mathbb{E}\left[\left.XU_{*}\right|W\right]\neq0$). Nevertheless, as long as $e_{0}\left(X\right)$ and $v_{0}\left(W\right)$ equal the true conditional mean and variance of $X$ given $W$, we have the pair of equalities, using iterated expectations,

align*[align* omitted — 270 chars of source]

(the second equality follows from the fact that $\mathbb{E}\left[\left.\left(X-e_{0}\left(W\right)\right)X'\right|W\right]=v_{0}\left(W\right)$).

Therefore ((ref)) remains mean zero even if the nuisance functions $h_{0}\left(W\right)=\left(a_{0}\left(W\right),b_{0}\left(W\right)\right)$ are replaced by arbitrary functions of $W$:

equation[equation omitted — 167 chars of source]

One special choice of $h_{*}(W)$ is the zero vector. This choice directly recovers the representation of $\beta_{\mathrm{0}}$ derived earlier (Equation ((ref)) above). In moment condition form \[ \mathbb{E}\left[v_{0}\left(W\right)^{-1}\left(X-e_{0}\left(W\right)\right)Y-\beta_{0}\right]=0. \]

Next consider replacing $g_{0}\left(W\right)=\left(e_{0}\left(W\right),v_{0}\left(W\right)\right)$ in ((ref)) with the incorrect conditional mean and variance functions $g_{*}\left(W\right)=\left(e_{*}\left(W\right),v_{*}\left(W\right)\right)$. Use the notation $U_{0}=\left(Y-a_{0}\left(W\right)-X'b_{0}\left(W\right)\right)$ to emphasize that $U_{0}$ is the prediction error associated with the mean squared error minimizing conditional linear prediction of $Y$ given $X$ conditional on $W$. By Lemma (ref) $\mathbb{E}\left[\left.U_{0}\right|W\right]=0$ and $\mathbb{E}\left[\left.XU_{0}\right|W\right]=0$. Therefore

align*[align* omitted — 565 chars of source]

Hence ((ref)) also remains mean zero even if the nuisance functions $g_{0}\left(W\right)=\left(e_{0}\left(W\right),v_{0}\left(W\right)\right)$ are replaced by arbitrary functions of $W$.

Moment ((ref)) has the so-called doubly robust property of Scharfstein_Rotnitzky_Robins_JASA99b Ruud_JOE86. It is mean zero as long as one of the two nuisance functions, $g\left(W\right)$ or $h\left(W\right)$, coincides with its population one. We exploit this property when constructing our estimator in the next section.

Estimation

In this section we present a locally semiparametrically efficient estimate of $\beta_{0}$. To motivate the precise form of our estimator we also discuss the estimator proposed by Wooldridge_CWP04. A textbook presentation of this estimator is available in Chapter 21.6.3 of Wooldridge_EACSPDBook10.

For the purposes of estimation we impose a parametric restriction on the conditional distribution of X given W. Since the distribution of X given W is ancillary for $\beta_{0},$ this parametric restriction does not change the semiparametric efficiency bound (cf., Corollary (ref)). We call, borrowing nomenclature from related settings Hirano_Imbens_BM04, the resulting model for $X$ the generalized propensity score.

assumption(Generalized Propensity Score) $f\left(\left.x\right|w;\phi\right)$ is a parametric family of densities indexed by $\phi\in\Phi\subset\mathbb{R}^{L}$ with (i) $f_{0}\left(\left.x\right|w\right)=\mbox{f\ensuremath{\left(\left.x\right|w;\phi_{0}\right)}}$ at some unique $\phi_{0}\in\mathrm{int}\left(\Phi\right)$, (ii) a maximum likelihood estimate (MLE) of $\phi_{0}$ equal to \[ \hat{\phi}=\arg\underset{\phi\in\Phi}{\max}\sum_{i=1}^{N}\ln f\left(\left.X_{i}\right|W_{i};\phi\right) \] with a score vectors of $\mathbb{S}_{\phi}\left(\left.X\right|W;\phi\right)=\nabla_{\phi}f\left(\left.X\right|W;\phi\right)/f\left(\left.X\right|W;\phi\right)$, (iii) $\hat{\phi}\stackrel{p}{\rightarrow}\phi_{0}$ with $\mathbb{E}\left[\mathbb{S}_{i}\mathbb{S}_{i}'\right]$ non-singular and the asymptotically linear representation \begin{equation} \sqrt{N}\left(\hat{\phi}-\phi_{0}\right)=\mathbb{E}\left[\mathbb{S}_{i}\mathbb{S}_{i}'\right]^{-1}\frac{1}{\sqrt{N}}\sum_{i=1}^{N}\mathbb{S}_{i}+o_{p}\left(1\right) \end{equation} where $\mathbb{S}_{i}=\mathbb{S}_{\phi}\left(\left.X_{i}\right|W_{i};\phi_{0}\right).$

Assumption (ref) corresponds to a parametric model for the propensity score when $X$ is a binary treatment indicator. More generally Assumption (ref) requires the researcher to model the distribution of the policy given controls. Consider a researcher interested in the relationship between regular school attendance and student achievement. In this case $Y$ could be a measure of end-of-school-year achievement, $X$ number of school days absent, and $W$ a vector of joint determinants of achievement and attendance (e.g., family background measures, prior achievement, pre-existing health conditions etc.). In this case the researcher might assume that the distribution of $X$ given $W$ is Poisson with \[ \mathbb{E}\left[\left.X\right|W\right]=\exp\left(k\left(W\right)'\phi_{0}\right),\ \mathbb{V}\left(\left.X\right|W\right)=\exp\left(k\left(W\right)'\phi_{0}\right), \] where $k\left(W\right)$ is a known $L\times1$ vector of functions of $W$. Estimation of $\phi_{0}$ , and hence $e\left(W;\phi_{0}\right)$ and $v\left(W;\phi_{0}\right)$, is by maximum likelihood. In most cases the conditional distribution of $X$ given $W$ can be conveniently modeled by, depending on the nature of $X$, the appropriate generalized linear model (GLM). When $X$ is multivariate, the outcome of censoring, or has mixed discrete/continuous components, then specifying $f\left(\left.x\right|w;\phi\right)$ may involve considerable work. For complicated likelihoods $e\left(W;\hat{\phi}\right)$ and $v\left(W;\hat{\phi}\right)$ may need to be approximated numerically or by simulation.

The Wooldridge_CWP04 estimator

Wooldridge_CWP04 introduced a two-step estimator for $\beta_{0}$. A textbook exposition appears in Wooldridge_EACSPDBook10. His procedure is summarized in Algorithm (ref).

algorithm[algorithm omitted — 547 chars of source]

Wooldridge_CWP04 does not characterize the asymptotic sampling properties of $\hat{\beta}_{W}$. In this section, we show that Wooldridge's estimator is not efficient under Assumptions (ref), (ref) and (ref). Furthermore it requires the generalized propensity score to be correctly specified. The structure of this inefficiency and lack of robustness, as well as the form of the efficient influence function derived earlier, guides the construction of our new, locally efficient and doubly robust estimator.

The second step of Algorithm (ref) corresponds to finding the $\hat{\beta}_{\mathrm{W}}$ which solves the sample moment

equation[equation omitted — 131 chars of source]

for $\rho\left(Z,\phi,\beta\right)=v\left(W;\phi\right)^{-1}\left(X-e\left(W;\phi\right)\right)\left(Y-X'\beta\right).$ Here $\hat{\phi}$ corresponds to the MLE of $\phi_{0}$ computed in the first step of the procedure. A mean value expansion of ((ref)) in $\hat{\beta}_{W}$ about $\beta_{0}$ yields

eqnarray*[eqnarray* omitted — 148 chars of source]

Rearrangement of terms and a second mean value expansion in $\hat{\phi}$ about $\phi_{0}$ gives

eqnarray*[eqnarray* omitted — 338 chars of source]

Observe that under Assumptions (ref) and (ref)

eqnarray*[eqnarray* omitted — 272 chars of source]

since $\mathbb{E}\left[\left.v\left(W;\phi_{0}\right)^{-1}\left(X-e\left(W;\phi_{0}\right)\right)X'\right|W=w\right]=I_{K}$. In integral form:

equation[equation omitted — 218 chars of source]

Differentiating ((ref)) through the integral with respect to $\phi$ gives:

equation[equation omitted — 227 chars of source]

which is a Generalized Information Matrix Equality (GIME) result Newey_JAE90.

Using ((ref)) and ((ref)) we have

eqnarray[eqnarray omitted — 522 chars of source]

for $\rho_{i}=\rho\left(Z_{i},\phi_{0},\beta_{0}\right).$

Similar to the result of Wooldridge_JE07 for the binary $X$ case, this asymptotically linear representation of $\hat{\beta}_{\mathrm{W}}$ implies that if practitioners ignore sampling error in $\hat{\phi}$, they can get conservative confidence intervals. In addition, this expression shows that over-parameterizing the conditional distribution of $X$ given $W$ will not decrease the asymptotic precision $\hat{\beta}_{\mathrm{W}}$.

We show next that $\hat{\beta}_{\mathrm{W}}$ is inefficient for $\beta_{0}$ in the semiparametric model defined by Assumptions (ref), (ref) and (ref). This demonstration of inefficiency usefully provides insight into how to construct a more efficient estimator. We begin by decomposing Wooldridge's Wooldridge_CWP04 identifying moment into the efficient influence function and a remainder: $\rho\left(Z,\phi_{0},\beta_{0}\right)=\mathrm{\psi}_{\beta}^{\mathrm{eff}}\left(Z,\beta_{0},\phi_{0},h_{0}\left(W\right)\right)+r\left(W,X,\beta_{0},\phi_{0},h_{0}\left(W\right)\right)$ with

align[align omitted — 294 chars of source]

Let $r_{0}\left(W,X\right)=r\left(W,X,\beta_{0},\phi_{0},h_{0}\left(W\right)\right).$ Note that $\mathbb{E}\left[\left.r_{0}\left(W,X\right)\right|W\right]=0.$ Note further that $\mathbb{S}$ is also conditionally mean zero given $W$.

Now observe that for $l=1,\ldots,\dim\left(\phi\right)$

eqnarray*[eqnarray* omitted — 349 chars of source]

and hence that

eqnarray[eqnarray omitted — 518 chars of source]

by Lemma (ref) above.

Next start with the fact that \[ \int\mathrm{\psi}_{\beta}^{\mathrm{eff}}f_{0}\left(\left.y\right|x,w\right)f\left(\left.x\right|w;\phi_{0}\right)f_{0}\left(w\right)=0. \] Differentiating through the integral gives the equality \[ \int\frac{\partial\mathrm{\psi}_{\beta}^{\mathrm{eff}}}{\partial\phi'}f_{0}\left(\left.y\right|x,w\right)f\left(\left.x\right|w;\phi_{0}\right)f_{0}\left(w\right)=-\int\left\{ \mathrm{\psi}_{\beta}^{\mathrm{eff}}\mathbb{S}'\right\} f_{0}\left(\left.y\right|x,w\right)f\left(\left.x\right|w;\phi_{0}\right)f_{0}\left(w\right) \] and hence that, using the decomposition of $\rho\left(Z,\phi_{0},\beta_{0}\right)$ introduced above and equation ((ref)), \[ \mathbb{E}\left[\rho\mathbb{S}'\right]=\mathbb{E}\left[\psi_{\beta}^{\mathrm{eff}}\mathbb{S}'\right]+\mathbb{E}\left[r\mathbb{S}'\right]=\mathbb{E}\left[r\mathbb{S}'\right]. \] Plugging this into our influence function we get

eqnarray*[eqnarray* omitted — 485 chars of source]

and hence an asymptotic distribution of

equation[equation omitted — 296 chars of source]

with $\Pi_{r\mathbb{S}}=\mathbb{E}\left[r\mathbb{S}'\right]\times\mathbb{E}\left[\mathbb{S}\mathbb{S}'\right]^{-1}.$

The form of the the limit distribution ((ref)) is similar to that of the familiar inverse probability weighting (IPW) estimator for binary treatments Graham_Pinto_Egel_ReStud12. In that context it is well-known that replacing a known propensity score with an estimated one increases precision Hirano_et_al_EM03,Hitomi_et_al_ET08,Graham_EM11. In principle the degree of precision increase is increasing in the complexity/richness of the fitted propensity score model. Expression ((ref)) indicates that a similar phenomena operates in our setting. If the portion of the efficient influence function that is omitted by the Wooldridge_CWP04 procedure is well-approximated by a linear combination of the scores used to estimate the propensity score, then the $\hat{\beta}_{\mathrm{W}}$ will be precisely determined. In practice, instead of relying on a possibly overfitted propensity score to yield efficient estimates, it is better to redesign the estimation procedure with efficiency in mind at the outset.

A locally efficient, doubly robust estimator

Our estimator for $\beta_{0}$ requires a working parametric model for the CLP coefficients $a_{0}\left(W\right)$ and $b_{0}\left(W\right)$. Consistency and asymptotic normality of our estimate, $\hat{\beta}$, will not depend on the correctness of this working model, but its limiting variance will. A convenient working model is provided by Assumption (ref).

assumption(CLP Coefficients): $a_{0}\left(W\right)=\alpha_{0}+\left(W-\mu_{W}\right)'\gamma_{0}$ and $b_{0}\left(W\right)=\beta_{0}+\Delta_{0}\left(W-\mu_{W}\right)$.

In practice these models for $a_{0}\left(W\right)$ and $b_{0}\left(W\right)$ can be made arbitrarily flexible since $W$ can include a rich set of basis functions (e.g., squares, cross-products etc.) in the underlying controls.

Under Assumption (ref) we have that

eqnarray[eqnarray omitted — 331 chars of source]

where $\delta_{0}=\mathrm{vec}\left(\Delta_{0}\right).$

Equation ((ref)) implies that, maintaining Assumption (ref), one approach to estimating $\beta_{0}$ would be to compute the least squares fit of $Y_{i}$ onto a constant, $W_{i}-\mu_{W}$, all interactions of $W_{i}-\mu_{i}$ and $X_{i}$, and $X_{i}$ itself. For the special case where $X$ is a binary treatment indicator, this estimator is familiar to labor economists as a Oaxaca-Blinder average treatment effect (ATE) estimator Sloczynski_OBES15.\footnote{In this literature researchers typically center $W$ around $\mathbb{E}\left[\left.W\right|X=1\right]$, not the unconditional mean $\mu_{W}=\mathbb{E}\left[W\right]$ as is done here. With this alternative centering the coefficient on $X_{i}$ will coincide with the average treatment effect on the treated (ATT).} Consistency of of this estimator hinges upon Assumption (ref) accurately characterizing the sampled population.

In our setting Assumption (ref) plays a different role. Unlike in the Oaxaca-Blinder procedure, its validity is not required for consistency, but if it does accurately described the sampled population our estimator will be highly efficient. These benefits come at the cost of assuming that prior knowledge regarding the form of the generalized propensity score is available (i.e., maintaining Assumption (ref)).

To describe our procedure we require some additional notation. Let $\lambda=\left(\alpha,\gamma',\delta'\right)'$, $R\left(\mu_{W}\right)=\left(1,\left(W-\mu_{W}\right)',\left(\left(W_{i}-\mu_{W}\right)\varotimes X_{i}\right)'\right)'$ and \[ U_{i}\left(\mu_{W},\lambda,\beta\right)=\left(Y_{i}-R\left(\mu_{W}\right)'\lambda-X_{i}'\beta\right). \] When $R\left(\mu_{W}\right)$ is evaluated at the correct population mean of $W$, we often simply write $R$.

Our estimator is based upon the $\left(L+J+1+J+JK+K\right)\times1$ vector of moment conditions, $m\left(Z_{i},\theta\right)$, partitioned into the three ordered sub-vectors:

align[align omitted — 474 chars of source]

where $\theta=\left(\phi,\mu_{W},\lambda',\beta\right)'$ with $\dim\left(\theta\right)=L+J+1+J+JK+K$.

Equations ((ref)), ((ref)) and ((ref)) constitute a just-identified system. The corresponding method-of-moments estimate of $\beta_{0}$ can be computed in the three simple steps listed in Algorithm (ref).

algorithm[algorithm omitted — 781 chars of source]

In many cases of interest Algorithm (ref) is easily implemented using standard software. Standard errors may be constructed in the usual way for GMM estimators Newey_McFadden_HBE94,Wooldridge_EACSPDBook10 or using a bootstrap.

In step 3, if instead we let $X_{i}$ serve as its own instrument, we get an “Oaxaca-Blinder” type estimator.

The next theorem summarizes the large sample properties of $\hat{\beta}$. In the statement of the Theorem, $\Delta_{*}$ denotes the limiting pseudo-true value of $\hat{\Delta}$. If Assumption (ref) additionally holds then $\Delta_{*}=\Delta_{0}$. We also define $\tilde{\epsilon}=v\left(W\right)_{0}^{-1}\left(X-e_{0}\left(W\right)\right)\epsilon$ where $\epsilon=\left\{ a_{0}\left(W\right)+X'\left(b_{0}\left(W\right)-\beta_{0}\right)-R'\lambda_{*}\right\} $ (with $\lambda_{*}$ denoting a pseudo-true parameter value). Finally we let $\Pi_{\tilde{\epsilon}\mathbb{S}}=\mathbb{E}\left[\tilde{\epsilon}_{i}\mathbb{S}'\right]\mathbb{E}\left[\mathbb{S}\mathbb{S}'\right]^{-1}$ denote the coefficient matrix associated with the multivariate regression of $\tilde{\epsilon}$ onto the score vector associated with $\phi_{0}$ (the parameter indexing the generalized propensity score).

thm(Large Sample Distribution) Consider the semiparametric problem established by Definition (ref) and Assumptions (ref), (ref), and (ref). Let $\hat{\beta}$ be the method of moments estimate of $\beta_{0}$ based upon restrictions ((ref)) to ((ref)). Under regularity conditions Newey_McFadden_HBE94 $\hat{\beta}$ is (i) asymptotically normal with a limiting distribution of \begin{equation} \sqrt{N}\left(\hat{\beta}-\beta_{0}\right)\overset{D}{\rightarrow}\mathcal{N}\left(0,\mathbb{E}\left[\Omega_{0}\left(W\right)\right]+\Delta_{*}\mathbb{V}\left(W\right)\Delta_{*}'+\mathbb{E}\left[\left(\tilde{\epsilon}-\Pi_{\tilde{\epsilon}\mathbb{S}}\mathbb{S}\right)\left(\tilde{\epsilon}-\Pi_{\tilde{\epsilon}\mathbb{S}}\mathbb{S}\right)'\right]\right), \end{equation} and (ii) locally efficient for $\beta_{0}$ at Assumption (ref) with \begin{equation} \sqrt{N}\left(\hat{\beta}-\beta_{0}\right)\overset{D}{\rightarrow}\mathcal{N}\left(0,\mathcal{I}(\beta_{0})^{-1}\right). \end{equation}
proofSee Appendix (ref).

Part (ii) of Theorem ((ref)) follows easily from part (i). In the proof we show that $\epsilon$ equals the prediction error associated with the mean squared error minimizing linear prediction of $a_{0}\left(W\right)+X'\left(b_{0}\left(W\right)-\beta_{0}\right)$ given $R\left(\mu_{W}\right)$. When Assumption (ref) additional holds this prediction error will be identically equal to zero and the third term in the variance expressing appearing in part (i) drops out. Similarly when Assumption (ref) holds we have $\Delta_{*}\mathbb{V}\left(W\right)\Delta_{*}'=\mathbb{V}\left(b_{0}\left(W\right)\right)$. Together these two observations give part (ii).

Our efficiency bound calculation, Theorem (ref), gives the information bound for $\beta_{0}$ without imposing the additional auxiliary Assumption (ref). This assumption imposes restrictions on the joint distribution of the data not implied by the baseline model. If these restrictions are added to the prior used to calculate the efficiency bound, then it will generally be possible to estimate $\beta_{0}$ more precisely. Our estimator is not efficient with respect to this augmented model. Rather it attains the bound provided by Theorem (ref) if Assumption (ref) “happens to be true” in the sampled population, but is not part of the prior restriction used to calculate the bound. Newey_JAE90 discusses the concept of local efficiency in detail. In what follows we will, for brevity, say $\hat{\beta}$ is locally efficient at Assumption (ref).

Even if Assumption (ref) does not hold precisely, our procedure will be “nearly” efficient when it is approximately true (in which case variability in $\epsilon$ about zero is small). A caveat to this claim is that the third variance term in ((ref)) may still be large in this case if $v_{0}\left(w\right)$ is nearly zero for enough values of $w$. This occurs when overlap is poor, or there exists a lack of variation in the policy variable for some subpopulations defined in terms of $W=w$. Graham_Pinto_Egel_JBES16 develop this observation more extensively for the special case where $X$ is binary, but similar issues apply in the more general setting considered here.

Our next result formalizes the above observation. It extends our local efficiency result to “near” global efficiency. The basic argument mirrors that given by Chamberlain_JE87 for approximately efficient estimation of conditional moment problems. Presenting this result requires defining a sequence of estimators based upon Algorithm (ref).

Let $\mathcal{L}_{2}$ be the space of functions $f\thinspace:\thinspace\mathbb{W}\rightarrow\mathbb{R}$ with finite second moment $\mathbb{E}\left[f\left(W\right)^{2}\right]<\infty$. Under Assumptions (ref) and (ref) the set of feasible conditional linear predictor coefficients lies within this space such that $a\thinspace:\thinspace W\rightarrow\mathbb{R}^{1}$ and $b\thinspace:\thinspace W\rightarrow\mathbb{R}^{K}$ with $\mathbb{E}\left[a\left(W\right)^{2}\right]<\infty$ and $\mathbb{E}\left[\left\Vert b\left(W\right)\right\Vert ^{2}\right]<\infty.$ Let $\left\{ k_{j}\left(W\right)\right\} _{j=1}^{\infty}$ be a sequence of linearly independent functions of the control variables, each with finite variance. Similar to Chamberlain_JE87 we call this sequence complete if, (i) for any $\zeta>0$ and (ii) any feasible conditional linear predictor coefficients $a\left(W\right)$ and $b\left(W\right)$ in $\mathcal{L}_{2}$, there are the real numbers $\alpha,\gamma_{1},\ldots,\gamma_{J}$ and $\delta_{k1},\ldots,\delta_{kJ}$ for $k=1,\ldots,K$ such that

equation[equation omitted — 137 chars of source]

with $\delta^{\left(J\right)}\left(W\right)$ defined as

equation[equation omitted — 413 chars of source]

Let $k^{\left(J\right)}\left(W\right)$ be the $J\times1$ vector of functions of $W$ with $j^{th}$ element $k_{j}\left(W\right)$. We can construct a sequence of estimators, indexed by $J$, based upon Algorithm (ref) with $k^{\left(J\right)}\left(W\right)$ replacing $W$. To do this let $\mu^{\left(J\right)}=\mathbb{E}\left[k^{\left(J\right)}\left(W\right)\right]$ and additionally define \[ R^{\left(J\right)}=\left(1,\left(k^{\left(J\right)}\left(W\right)-\mu^{\left(J\right)}\right)',\left(\left(k^{\left(J\right)}\left(W\right)-\mu^{\left(J\right)}\right)\otimes X\right)'\right)'. \] We can then estimate $\beta_{0}$ by Algorithm (ref) with $k^{\left(J\right)}\left(W\right)$, $\mu^{\left(J\right)}$ and $R^{\left(J\right)}$ respectively replacing $W$, $\mu_{W}$, and $R\left(\mu_{W}\right)$.

Consider the asymptotic precision matrix of this method of moments estimator; from Theorem (ref) we get

align*[align* omitted — 488 chars of source]

with $\mathcal{I}^{\left(J\right)}\left(\beta_{0}\right)^{-1}\geq\mathcal{I}\left(\beta_{0}\right)^{-1}$ (here “$A\geq B$” denotes “$A-B$ is positive semi-definite”). Recall that $\mathcal{I}\left(\beta_{0}\right)$ is the semiparametric efficiency bound given in Theorem (ref). Let $\hat{\beta}^{\left(J\right)}$ denote the estimate of $\beta_{0}$ based upon $k^{\left(J\right)}\left(W\right)$.

prop(Near Efficiency) If, maintaining the Assumptions of part (i) of Theorem (ref), $\left\{ \hat{\beta}^{\left(J\right)}\right\} $ is based upon a linearly independent, complete sequence $\left\{ k_{j}\left(W\right)\right\} _{j=1}^{\infty},$ then, for $\mathbb{X}\times\mathbb{W}$ a compact subset of $\mathbb{R}^{K+\dim\left(W\right)}$, \[ \underset{J\rightarrow\infty}{\lim}\mathcal{I}^{\left(J\right)}\left(\beta_{0}\right)^{-1}=\mathcal{I}\left(\beta_{0}\right)^{-1}. \]
proofSee Appendix (ref).

The compact support assumption invoked in the statement of the theorem is used in the proof, but does not appear to be essential.

Proposition (ref) leaves unanswered important practical questions, such as how quickly $J$ should increase with $N$. More generally the question of exactly how to choose the elements of $k^{\left(J\right)}\left(W\right)$ in order to achieve good precision in practice remains unanswered. However we expect that many insights from related settings could be applied here Belloni_et_al_ReStud14.

We conclude this section by demonstrating double robustness in the sense of Scharfstein_Rotnitzky_Robins_JASA99b. If the specification of the generalized propensity score is not correct (i.e., Assumption (ref) does not hold), but Assumption (ref) is true, then our estimator remains consistent for $\beta_{0}$. Recall that Assumption (ref) was initially invoked to ensure local efficiency of our procedure. It turns out that modeling the form of the conditional linear predictor coefficients has the added benefit of ensuring that our estimator remains consistent even if our generalized propensity score model is incorrect. Double robustness results are familiar from the literature on missing data and program evaluation Scharfstein_Rotnitzky_Robins_JASA99b,Cattaneo_JOE10,Graham_EM11. In these settings $X$ is binary or a vector of mutually-exclusive treatment indicators. Double robustness in our more general setting is perhaps unsurprising, but nevertheless a new result.

To understand this result observe that step 3 of Algorithm (ref) corresponds to solving the sample analog of \[ \mathbb{E}\left[\left(

array[array omitted — 108 chars of source]

\right)U\left(\mu_{W},\lambda_{0},\beta_{0}\right)\right]=0 \] for $\lambda_{0}$ and $\beta_{0}$. Here we use the notation $\phi_{*}$ to denote that our generalized propensity score model may be miss-specified.

If Assumption (ref) holds in the population, then $U_{0}=U_{i}\left(\mu_{W},\lambda_{0},\beta_{0}\right)$ is a conditional linear predictor (CLP) error and Lemma (ref) above applies. Recall that $R\left(\mu_{W}\right)=\left(1,\left(W-\mu_{W}\right)',\left(\left(W_{i}-\mu_{W}\right)\varotimes X_{i}\right)'\right)'$; by Lemma (ref) $U_{0}$ is uncorrelated with all components of this vector. Likewise, because $U_{0}$ is mean independent of $W$ and conditionally uncorrelated with $X$, we also have that $\mathbb{E}\left[v\left(W;\phi_{*}\right)^{-1}\left(X-e\left(W;\phi_{*}\right)\right)U\right]$ is mean zero as well. Hence step 3 of Algorithm (ref) involves the computation of a correctly specified method-of-moments estimator under Assumption (ref); irrespective of whether Assumption (ref) additionally holds. Double robustness follows, more or less, directly.

The above discussion also clarifies why, as is sometimes true in practice, sampling variability in our estimator can theoretically be lower than the semiparametric variance bound in Theorem (ref) when the generalized propensity score is misspecified, but the form of the CLP coefficients are not. First, recall that the variance bound is computed without making any a priori assumptions about the form of the CLP coefficients. It turns out that in our setting such assumptions generally increase the precision with which $\beta_{0}$ may be estimated. When we invoke the double robustness property of our procedure to ensure consistency we are in a situation where the veracity of Assumption (ref) is central. Whereas is the setting covered by Theorem (ref), Assumption (ref) “may happen to be true”, but need not be.

It is instructive to compare our estimator with the “Oaxaca-Blinder-type” one described earlier. The Oaxaca-Blinder procedure necessarily maintains Assumption (ref). Since this restriction is part of the prior, it would not be surprising to find that, under correct specification, that the Oaxaca-Blinder estimator is more efficient than ours. For the purposes of developing this point, additionally assume that $U_{0}$ is homoscedastic in $X$ and $W$ (but that this is not part of the prior), then \textendash maintaining Assumption (ref) \textendash replacing $v\left(W;\phi_{*}\right)^{-1}\left(X-e\left(W;\phi_{*}\right)\right)$ with $X$ in the above moment would be natural. This replacement leads the researcher to the Oaxaca-Blinder estimator (which will also be efficient in this case). Hence, when Assumption (ref) does hold in the sampled population, our procedure will be less efficient that the Oaxaca-Blinder one (at least under homoscedasticity of $U_{0}$). Of course, when Assumption (ref) does not characterize the sampled population, our procedure remains consistent, while the Oaxaca-Blinder one does not.

thm(Double Robustness) Under Assumptions (ref) and (ref) , $\hat{\beta}\overset{p}{\rightarrow}\beta_{0}$ if either Assumption (ref) or (ref) holds.

The proof is straightforward and omitted (see Graham_Pinto_Egel_ReStud12 and Graham_Pinto_Egel_JBES16 for proofs of related results). As a practical matter using the standard method-of-moments sandwich variance-covariance matrix estimator associated with the moment problem defined by ((ref)), ((ref)) and ((ref)) above will support asymptotically valid inference under the conditions of both Theorems (ref) and (ref).

Examples and special cases

In this section we demonstrate that our semiparametric regression model encompasses several other well-known models.

Example 1: Binary Treatment Effect

Following the program evaluation literature let $Y_{0}$ denote the potential outcome under control and $Y_{1}$ the potential outcome under active treatment treatment. For each sampled unit we observe either $Y_{0}$ or $Y_{1}$ but not both. The observed outcome, $Y$, therefore equals \[ Y=XY_{1}+(1-X)Y_{0} \] where $X$ equals 1 if the unit is treated and zero otherwise. Rewriting yields a random coefficients model of \[ Y=A+BX \] with $A=Y_{0}$ and $B=Y_{1}-Y_{0}$. The average treatment effect (ATE) equals \[ \beta_{0}=\mathbb{E}[Y_{1}-Y_{0}]=\mathbb{E}[B]. \]

Rosenbaum_Rubin_BM83 show that the ATE is identified when $(Y_{0},Y_{1})\bot X|W$ (unconfoundedness) and $0<\Pr\left(\left.X=1\right|W=w\right)<1$ for all $w\in\mathbb{W}$ (overlap).

When $X$ is binary our Assumption (ref) implies unconfoundedness. Assumption (ref) implies that $X$ is conditionally uncorrelated with the two potential outcomes. When $X$ is binary this also corresponds to mean and full conditional independence. Next observe that $e_{0}(W)=\Pr\left(\left.X=1\right|W=w\right)$ and $v_{0}(W)=e_{0}(W)\left[1-e_{0}(W)\right]$. Hence our Assumption (ref) implies that $0<\kappa\leq e_{0}\left(W\right)\leq1-\kappa<1$ or so called strong overlap.

Now consider Algorithm (ref). When $X$ is binary we have that \[ v\left(W,\hat{\phi}\right)^{-1}\left(X-e\left(W,\hat{\phi}\right)\right)=\frac{X}{e\left(W,\hat{\phi}\right)}-\frac{1-X}{1-e\left(W,\hat{\phi}\right)}. \] Our ATE estimate is the coefficient on $X$ associated with the linear instrumental variables fit of $Y$ onto a constant, $(W-\hat{\mu}_{W})$, $(W-\hat{\mu}_{W})\cdot X$, and $X$ where $\frac{X}{e\left(W,\hat{\phi}\right)}-\frac{1-X}{1-e\left(W,\hat{\phi}\right)}$ serves as an instrument for $X$. This estimator is similar to, but distinct from, the weighted least squares (WLS) one introduced by Hirano_Imbens_HSORM01.

Wooldridge_CWP04 shows, for $X$ binary, that equation ((ref)) coincides with

align*[align* omitted — 202 chars of source]

which is the familiar inverse probability weighting (IPW) representation of the average treatment effect (ATE) in, for example, Hirano_et_al_EM03.

The general form of the efficient influence function given in Theorem (ref) above corresponds to the specialized one for the ATE when $X$ is binary derived by, for example, Hahn_EM98 and Hirano_et_al_EM03. Hence our general procedure, as summarized by Algorithm (ref), provides a locally efficient and double robust estimator of the ATE. To the best of our knowledge, our proposed estimator is a new one, even in the special case where it identifies the ATE of a binary treatment. Bang_Robins_BM05 and Tsiatis_Book06 provide introductions to double robust causal inference.

Example 2: Multiple Treatment Effects

Following Imbens_BM00, Wooldridge_CWP04, and Cattaneo_JOE10 consider finite collection of mutually exclusive treatments indexed by $k\in\{0,1,2,...,K\}$ with $K\in\mathbb{N}$. Associated with these treatments are the $K+1$ potential outcomes, $Y(0),Y(1),\ldots,Y(K)$. The observed outcome is \[ Y=Y(0)+\sum_{k=1}^{K}X_{k}\left\{ Y(k)-Y(0)\right\} \] where $X_{k}$ is a binary random variable that equals $1$ if treatment $k=0,\ldots,K$ is assigned to the unit and zero otherwise. In this case, we work with the following random coefficient model: \[ Y=A+X^{'}B \] where $X=\left(X_{1},\ldots,X_{K}\right)'$ is a $K\times1$ vector of treatment indicators and $B$ a corresponding vector of individual treatment effects.

In this setup $X$ is multinomial with a conditional mean of \[ e_{0}\left(W\right)=\left(

array[array omitted — 102 chars of source]

\right) \] and an inverse conditional variance of Henderson_Searle_SIAM81

align*[align* omitted — 255 chars of source]

A little bit of tedious algebra then gives \[ \beta_{0}=\mathbb{E}\left[v_{0}\left(W\right)^{-1}\left(X-e_{0}\left(W\right)\right)Y\right]=\mathbb{E}\left[

array[array omitted — 238 chars of source]

\right], \] which corresponds to the IPW representation of the ATEs \[ \beta_{0}=\left(

array[array omitted — 136 chars of source]

\right), \] in the multiple treatment setting.

As in the case where $X$ is binary, the general form of the efficient influence function given in Theorem (ref) above corresponds to the specialized one derived by Cattaneo_JOE10. Hence our general procedure also provides a locally efficient and doubly robust estimate of ATEs in the multiple, mutually exclusive, treatments setting.

Example 3: Partially linear model

Chamberlain_WP86,Chamberlain_EM92 and Robinson_EM88 studied the semiparametric partially linear regression model (PLM) \[ Y=X'\beta_{0}+h_{0}(W)+U \] with $\mathbb{E}[U|W,X]=0.$ This model can be represented by the random coefficient model \[ Y=A+X'B \] where $\mathbb{E}[A|W,X]=a_{0}\left(W\right)=h_{0}(W)$ and $\mathbb{E}[B|X,W]=b_{0}\left(W\right)=\beta_{0}$. These assumptions are stronger than those implied by Assumption (ref).

To fit this into our framework we replace Assumption (ref) with a working CLP model of \[ a_{0}\left(W\right)=h_{0}\left(W\right)=\alpha_{0}+\left(W-\mu_{W}\right)'\lambda_{0},\thinspace\thinspace\thinspace b_{0}\left(W\right)=\beta_{0}. \] This implies a constant additive treatment effect structure.

Estimation follows Algorithm (ref). First compute the MLE of $\phi_{0}$. Second compute the sample means $\hat{\mu}_{W}=\frac{1}{N}\sum_{i=1}^{N}W_{i}$. Finally compute the linear instrumental variables fit of $Y$ onto a constant, $\left(W-\hat{\mu}_{W}\right)$ and $X$, using $v\left(W,\hat{\phi}\right)^{-1}\left(X-e\left(W,\hat{\phi}\right)\right)$ as an instrument for $X$. Because of the constant additive treatment effect structure of the PLM we exclude the interactions $\left(W-\hat{\mu}_{W}\right)\otimes X$ from the IV fit computed in the third step.

It is important to recognize that although our procedure invokes the working assumption that the treatment effect is constant in $W$ (i.e., $b_{0}\left(w\right)=\beta_{0}$ for all $w\in\mathbb{W}$), this assumption is not required for consistency as long as our generalized propensity score model is correct. Put differently although our procedure incorporates the PLM structure, this structure is not part of the maintained prior (albeit the form of the generalized propensity score is part of the prior).

If $b_{0}\left(w\right)=\beta_{0}$ for all $w\in\mathbb{W}$ and $U$ is conditionally mean zero given both $W$ and $X$ (and also has a constant variance), but these are not part of the prior restriction used to calculate the bound, then ((ref)) evaluates to \[ \mathcal{I}(\beta_{0})^{-1}=\sigma^{2}\mathbb{E}\left[v_{0}(W)^{-1}\right]. \] The modified PLM estimator described above, and based on our Algorithm (ref), will attain this bound when the true model is a partially linear one.

Chamberlain_EM92 gives a bound for $\beta_{0}$ \textendash where the partially linear regression structure is part of the prior restriction (but the homoscedasticity assumption is not) \textendash of \[ \mathcal{I}_{\mathrm{plm}}(\beta_{0})^{-1}=\sigma^{2}\mathbb{E}\left[v_{0}(W)\right]^{-1}. \] The difference $\mathcal{I}(\beta_{0})^{-1}-\mathcal{I}_{\mathrm{plm}}(\beta_{0})^{-1}$ is positive semi-definite. This follows directly from, for example, the Theorem in Groves_Rothenberg_BM69 on the expectations of inverse matrices. Hence although our approach to estimation remains consistent for $\beta_{0}$ when the true regression function takes a partially linear form, it will generally be less efficient than methods which exploit this structure at the outset Robinson_EM88,Robins_Mark_Newey_BM92.

Finite sample properties

In order to assess the approximation accuracy of Theorems (ref) and (ref) in finite samples we conducted a small simulation experiment, the results of which we report here. We considered four designs. The outcome was generated according to \[ Y=a_{0}\left(W\right)+b_{0}\left(W\right)X+U \] with $W$ and $U$ independent standard normal random variables and $a_{0}\left(W\right)$ and $b_{0}\left(W\right)$ either linear (designs 1 and 2) or quadratic (designs 3 and 4) in $W$. The conditional distribution of $X$ given $W$ was specified as Poisson with parameter $\exp\left(k\left(W\right)'\phi\right)$ and $k\left(W\right)=\left(1,W\right)'$ in designs 1 and 3 and $k\left(W\right)=\left(1,W,W^{2}\right)'$ in designs 2 and 4. Complete details on the data generating process are given in Table (ref).

table[table omitted — 1,226 chars of source]

We evaluate the performance of three estimators. First we consider a simple “Oaxaca-Blinder” type estimator. Specifically we estimate $\beta_{0}$ by the coefficient on $X$ in the least squares fit of $Y$ onto a constant, $W-\hat{\mu}_{W}$, $\left(W-\hat{\mu}_{W}\right)\times X$ and $X$. As in Kline_EL14 we appropriately account for the effect of estimating $\mu_{W}$ when constructing standard errors and confidence intervals. This estimator is consistent for the true average partial effect in both designs 1 and 2. It is also, since $U$ is Gaussian and homoscedastic, efficient in these two designs. Efficiency is in the semiparametric model which, in addition to Assumptions (ref) and (ref), maintains Assumption (ref). The variance of the Oaxaca-Blinder estimate therefore lies (weakly) below the bound given by Theorem (ref) in these two designs. In designs 3 and 4, where $a_{0}\left(W\right)$ and $b_{0}\left(W\right)$ are quadratic in $W$, the “Oaxaca-Blinder” estimator is inconsistent.

The second estimator is the generalized inverse probability weighting (GIPW) one due to Wooldridge_CWP04,Wooldridge_EACSPDBook10. Our implementation tracks our analysis in Section (ref). For estimation we correctly assume that the conditional distribution of $X$ given $W$ is Poisson, but set the parameter to $\exp\left(k\left(W\right)'\phi\right)$ with $k\left(W\right)=\left(1,W\right)'$. This is correct in designs 1 and 3, but not 2 and 4. Hence the GIPW estimate of $\beta_{0}$ is consistent in designs 1 and 3, but not 2 and 4. The GIPW is never efficient. Standard errors are constructed using the sample analog of the influence function given in ((ref)) above.

Finally we consider the properties of our locally efficient, doubly robust estimator. Implementing this procedure requires assumptions on both the CLP and the generalized propensity score. We make the same assumptions used to implement the Oaxaca-Blinder and GIPW procedures. Consequently this last estimator is efficient \textendash in the sense of Theorem (ref) \textendash in design 1 and consistent in designs 1, 2 and 3. Like all the estimators it is inconsistent in design 4. We construct standard errors using the (sample analog) of the influence function given in Theorem (ref); consequently our intervals are conservative in design 2 (where our propensity score model is misspecified).\footnote{We use Python 3.6 to conduct our experiments. Replication code is available in the supplemental materials. }

Each of the four designs are calibrated such that $\sqrt{\mathcal{I}\left(\beta_{0}\right)^{-1}/N}=0.05$ ($0.025$) when $N=1,000$ ($4,000$). In an asymptotic sense inference on $\beta_{0}$ is equally hard across all the designs considered. We focus on the $N=1,000$ experiments in our discussion (the quality of the asymptotic approximations predictably improve in the larger sample).

Results from the four designs are reported in Table (ref). As expected our DR estimator is median unbiased across Designs 1, 2 and 3. In contrast the Oaxaca-Blinder estimator only performs acceptably in designs 1 and 2, and the GIPW estimator in design 1 and 3. In designs 1 and 2 the variability of the DR estimator is nearly as small as that of the Oaxaca-Blinder one. Neither the DR, nor the GIPW, estimators are expected to be efficient in design 3 but, interestingly, GIPW is more efficient than DR in this case. In design 1, where the DR estimator is locally efficient, its standard deviation is substantially smaller than that of the GIPW estimator (as expected).

landscape\begin{table} \caption{Simulation Results} \begin{centering} \begin{tabular}{|c|c|c|c|c|c|c|c|c|} \hline & \multicolumn{4}{c|}{Panel A, $N=1,000$} & \multicolumn{4}{c|}{Panel B, $N=4,000$}\tabularnewline \hline & Bias & Std. Dev. & Std. Err. & Coverage & Bias & Std. Dev. & Std. Err. & Coverage\tabularnewline \hline \hline Design 1 & & & & & & & & \tabularnewline \hline Oaxaca-Blinder & -0.0003 & 0.0500 & 0.0496 & 0.9480 & 0.0006 & 0.0252 & 0.0248 & 0.9438\tabularnewline \hline GIPW & -0.0008 & 0.0853 & 0.0809 & 0.9438 & 0.0016 & 0.0429 & 0.0419 & 0.9494\tabularnewline \hline DR & 0.0001 & 0.0507 & 0.0499 & 0.9450 & 0.0009 & 0.0254 & 0.0250 & 0.9448\tabularnewline \hline Design 2 & & & & & & & & \tabularnewline \hline Oaxaca-Blinder & 0.0009 & 0.0504 & 0.0497 & 0.9442 & 0.0000 & 0.0251 & 0.0248 & 0.9456\tabularnewline \hline GIPW & -0.2597 & 0.1331 & 0.1198 & 0.4634 & -0.2613 & 0.0710 & 0.0666 & 0.0258\tabularnewline \hline DR & 0.0014 & 0.0518 & 0.0561 & 0.9620 & -0.0001 & 0.0258 & 0.0283 & 0.9694\tabularnewline \hline Design 3 & & & & & & & & \tabularnewline \hline Oaxaca-Blinder & -0.3268 & 0.0993 & 0.0899 & 0.0772 & -0.3319 & 0.0531 & 0.0481 & 0.0002\tabularnewline \hline GIPW & -0.0018 & 0.0830 & 0.0794 & 0.9436 & -0.0005 & 0.0428 & 0.0413 & 0.9452\tabularnewline \hline DR & -0.0113 & 0.1099 & 0.0981 & 0.9276 & -0.0036 & 0.0570 & 0.0529 & 0.9368\tabularnewline \hline Design 4 & & & & & & & & \tabularnewline \hline Oaxaca-Blinder & -0.2010 & 0.1571 & 0.1087 & 0.5148 & -0.2391 & 0.1255 & 0.0674 & 0.0168\tabularnewline \hline GIPW & 0.2717 & 0.1897 & 0.1216 & 0.4006 & 0.3052 & 0.1335 & 0.0748 & 0.0034\tabularnewline \hline DR & 0.4123 & 0.2035 & 0.1687 & 0.3076 & 0.4362 & 0.1058 & 0.1013 & 0.0986\tabularnewline \hline \hline $\sqrt{\mathcal{I}\left(\beta_{0}\right)^{-1}/N}$ & \multicolumn{4}{c|}{0.05} & \multicolumn{4}{c|}{0.0250}\tabularnewline \hline \end{tabular} \end{centering} \uline{Notes:} The bias column reports median bias across all $B=5,000$ simulations. The Std. Dev. column reports the standard deviation of the point estimates across these simulations, Std. Err. the median estimated standard error, and Coverage the actual coverage of a nominal 95 percent confidence interval (constructed using the estimated point estimate and standard error in the normal way). The standard error associated with a Monte Carlo coverage estimate is $\sqrt{\alpha\left(1-\alpha\right)/B}$. With $B=5,000$ simulations and $\alpha=0.05$, this results in a standard error of approximately 0.003 or a 95 percent confidence interval of $[0.944,0.956]$. \end{table}

Overall the simulation results track our theoretical predictions remarkably closely. Of course exploring the performance of these estimators in the context of real world empirical applications and other, more realistic, simulation designs would be of interest.

Conclusion

We have introduced a locally efficient, doubly robust, semiparametric method of estimating averages of conditional linear predictor coefficients. Our estimand, and semiparametric efficiency bound, specialize to familiar counterparts found in the program evaluation literature Hahn_EM98,Cattaneo_JOE10. While encompassing well-known program evaluation settings, our framework allows for (semiparametric) covariate adjustment in many other settings as well (including ones with few extant alternative methods of such adjustment).

Researchers interested in estimating the average treatment effect (ATE) associated with a binary treatment can apply our methods. While we believe the precise form of our procedure is new even to this familiar setting, it is a variant of the class of augmented inverse probability weighting (AIPW) estimators introduced by Robins_Rotnitzky_Zhao_JASA94 in the missing data context over 20 years ago. The real attraction of Algorithm (ref), and the corresponding Theorems (ref) and (ref) (as well as Proposition (ref)), is that they apply to models beyond the “classic” program evaluation one of Rosenbaum_Rubin_BM83. Multiple, mutually exclusive treatments, as in Imbens_BM00 and Cattaneo_JOE10 are easily handled as a special case. Similarly, maintaining a linear, but heterogeneous, potential response function structure, Algorithm (ref) recovers average partial effects (APE) for continuous treatments, multiple non-exclusive treatments, mixtures of binary, discrete and continuous treatments and so on. A weighted average derivative interpretation of our estimand is also available for settings where the linear potential response function structure may not hold (Proposition (ref)).

We also wish to emphasize that averages of conditional linear predictor coefficients represent a natural, but substantial, generalization of linear predictor coefficients as estimated by the method of least squares. Hence Algorithm (ref) also provides a method of flexible covariate adjustment that may be of independent interest even in settings where formal causal inference is not warranted; similar to how least squares is sometimes used for descriptive purposes.

Our work leaves several questions unanswered. First, although the flexible parametric modeling embodied in Assumptions (ref) and (ref) closely mirrors empirical practice, it would be useful to development methods that leave the generalized propensity score and CLP coefficients non-parametric. It seems likely that results from the binary and multiple treatments case could be extended to apply here Hirano_et_al_EM03,Cattaneo_JOE10,Belloni_et_al_ReStud14.

In other work we have shown that first order equivalent estimators may have appreciably different higher order properties in program evaluation settings Graham_Pinto_Egel_ReStud12. We expect that other locally efficient, doubly robust approaches to estimation for the class of problems considered in this paper are feasible. These approaches may exhibit superior or inferior higher order bias.

Third, maintaining the correlated random coefficient structure, different notions of conditional exogeneity will imply different semiparametric efficiency bounds (when linearity is restrictive). Our decision to work with a weak notion of exogeneity maintains a connection with conditional linear predictors. If a researcher was comfortable with the correlated random coefficient structure, then it would generally be possible to construct more efficient estimates of $\beta_{0}=\mathbb{E}\left[B\right]$ if she was willing to assume, for example, that $\left.\left(A,B'\right)'\perp X\right|W=w$ for all $w\in\mathbb{W}.$ Such estimators would likely be quite complicated and may have poor finite sample properties.