Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
81,749 characters · 12 sections · 111 citation commands
Estimation and Inference for Policy Relevant Treatment Effects
The policy relevant treatment effect heckman/vytlacil:2001,heckman/vytlacil:2005,heckman/vytlacil:2007 measures the average effect of switching from a status-quo policy to a counterfactual policy.\footnote{Also see stock:1989 and ichimura/taber:2000 for parameters related to the PRTE.} Among a number of measures for treatment effects, the PRTE has the advantage of directly evaluating alternative policy scenarios under consideration. A rich set of identification results have been established in the literature for this treatment effect parameter heckman/vytlacil:2001,heckman/vytlacil:2005,heckman/vytlacil:2007 based on the marginal treatment effects bjorklund/moffitt:1987.
Estimation of the PRTE involves nonparametric estimation for multiple preliminary parameters. First, we nonparametrically estimate the propensity score, the conditional probability of being treated given an instrumental variable. Next, when the outcome model involves additive controls as in carneiro/lee:2009 and carneiro/heckman/vytlacil:2010, we estimate the conditional expectation functions of the outcome and covariates given the propensity score through the partially linear regression estimation procedure of robinson1988root. Third, with the residual from the partially linear regression, we estimate the marginal treatment effect via a nonparametric derivative estimation.
These first-, second-, and third-stage estimates can affect the asymptotic distribution of the PRTE estimator in complicated and intractable manners, especially when their convergence rates are slower than the parametric rate of $n^{-1/2}$. Estimation of the PRTE based on an orthogonal score can mitigate and asymptotically vanish the estimation errors of these first, second, and third preliminary parameters. In this light, we propose an orthogonal score for double debiased estimation of the PRTE, whereby these first-, second-, and third-stage estimators will not affect the first-order asymptotic distribution of the PRTE estimator. Consequently, we characterize the asymptotic distribution for the PRTE estimator based on the proposed orthogonal score. Even conventionally in the absence of orthogonal scores, there has been no semi- or non-parametric inference method for the PRTE available in the existing literature, perhaps because of the complicated dependence of this parameter on multiple preliminary functions. Therefore, to our knowledge, our work is the first to develop limit distribution theories for inference about the PRTE.
Orthogonal scores for double debiased and doubly robust estimation appear in a few contexts in the literature of econometrics and statistics. One instance is about partial linear models,\footnote{Estimation of partial linear models is studied by robinson1988root under exogeneity, and is studied by okui2012doubly under endogeneity. For a robust minimum-distance approach under endogeneity, see ai2007estimation. See belloni2013honest, belloni2014high, and farrell2015robust for machine learning approaches.} and another is about weighted averages\footnote{Weighted estimation is studied by rosenbaum1987model, robins1992estimating, imbens1992efficient, robins1994estimation, robins1995semiparametric, lawless1999, wooldridge1999,wooldridge2001asymptotic, lee/okui/whang:2017 and sloczynski_wooldridge_2018 under parametric weight function, is studied by Hahn:1998, hirano2003efficient, suzukawa2004unbiased, NeweyRuud2005, Firpo:2007, lewbel2007simple, magnac2007identification, and chen2008semiparametric, graham2011efficiency, graham2012inverse, graham2016efficient, sant2018doubly, and rothe2019properties under nonparametric weight function, and is studied by BCFH:2017 and Wager/Athey:2018 with machine learning. Particularly, rothe2019properties analyzes the orthogonal scores for the policy effect parameter defined in stock:1989. } such as inverse propensity score weighting. Estimation of the PRTE consists of both of these two types of estimation procedures: (i) weighted estimation with the propensity score, estimation of a partial linear model to account for additive covariates as in carneiro/lee:2009 and carneiro/heckman/vytlacil:2010, and (ii) weighted estimation with a counterfactual policy. As such, it is a natural idea to develop an orthogonal score for estimation and inference for the PRTE by taking advantage of and combining these techniques. With this said, the existing literature has not proposed an orthogonal score for the PRTE, and hence we aim to contribute to this literature by proposing one in this paper.
There is a large literature on orthogonal scores, and it characterizes desired properties of orthogonal scores such as double robustness against misspecification, local robustness in semiparametric estimation, and even semiparametric efficiency for certain orthogonal scores -- hitomi2008puzzling discuss these features of orthogonal scores in general from geometric perspectives. Among these characteristics, our purpose of developing an orthogonal score is, as mentioned earlier, to remove the influence of the first, second, and third-stage estimators on the asymptotic distribution of the PRTE estimator, which concerns the local robustness. There are a number of general procedures for estimation and inference based on orthogonal scores which we rely on in the present paper. The seminal paper by newey1994asymptotic proposes orthogonal scores in semiparametric methods and provides forms of adjustment terms to to obtain orthogonal scores from moment functions. The way in which we derive our orthogonal score is based on his prescription newey1994asymptotic. belloni2014uniform and belloni2018uniformly propose a generic Z-estimation framework based on orthogonal scores. chernozhukov2016locally propose a general procedure for construction of orthogonal scores from moment restriction models. chernozhukov2018double combine the use of orthogonal scores and cross-fitting as a generic semiparametric strategy. Our proposed method of estimation and inference takes advantages of this existing body of knowledge.
To the best of our knowledge, no limit distribution has been established for the PRTE in the existing literature. Hence, this work is the first to develop asymptotic theories for inference about the PRTE. The closest is carneiro/lee:2009 who develop point-wise limit distributions for the MTE, although they do not seem to imply or lead to a limit distribution for the PRTE. Our orthogonal score does not only pave the way for a limit distribution for the PRTE for the first time in the literature, but also allows for a flexibility in types of preliminary estimators, e.g., kernel, sieve, and lasso. In order to make our method and theory more accessible in practice, we also provide specific estimation procedures and lower-level sufficient conditions.\footnote{We consider a generalized linear model for the propensity score function allowing for possibly high-dimensional covariates and propose primitive conditions based on nonparametric estimation of other preliminary functional parameters. This framework involves both shrinkage and kernel estimation, and is related to a branch of the recent literature including kennedy:2017, fan:2019, su/ura/zhang:2019, zimmert/lechner:2019 and colangelo/lee:2020.}
Our proposed method is applicable to empirical studies that identify and estimate marginal treatment effects and/or policy relevant treatment effects. Examples include, but are not limited to, auld:2005, basu/heckman/navarro/urzua:2007, doyle:2008, moffitt:2008, chuang/lai:2010, carneiro/heckman/vytlacil:2011, galasso/schankerman/serrano:2013, basu/jena/goldman/philipson/dubois:2014, belskaya/peter/posso:2014, johar/maruyama:2014, lindquist/santavirta:2014, moffitt:2014, dobbie/song:2015, jensen/nielsen:2016, kasahara/liang/rodrigue:2016, carneiro/lokshin/umapathi:2017, cornelissen/dustmann/raute/schonberg:2018, felfe/lalive:2018, and kamhofer/schmitz/westphal:2018.
The rest of this paper is organized as follows. Section (ref) introduces a model based on carneiro/lee:2009. Section (ref) presents an informal overview of our proposed method. Section (ref) presents an orthogonal score for double debiased estimation of the PRTE. Section (ref) discusses large sample properties of the double debiased estimator for inference. Section (ref) presents Monte Carlo simulation studies. Section (ref) presents an empirical illustration. The paper is summarized in Section (ref). The appendix contains proofs and additional details that are important but relegated there due to their lengths.
Following the model and notations in carneiro/lee:2009, consider the structure
of outcome production, where $Y_1$ and $Y_0$ denote the potential outcomes under treatment ($S=1$) and no treatment ($S=0$), respectively. With covariates $X$, the potential outcomes are modeled by $Y_1=\mu_1\left(X,U_1\right)$ and $Y_0=\mu_0\left(X,U_0\right)$ with unobserved variables $\left(U_0,U_1\right)$. The binary treatment assignment status $S$ is in turn determined by the threshold crossing model $S=1\{\mu_S\left(Z\right)-U_S>0\}$, where $Z$ is a vector of exogenous variables and $U_S$ is an error term. $Z$ can contain a sub-vector of $X$ as included exogenous variables, and the remaining excluded exogenous variables serve as instruments. In this model, the dependence between $U_S$ and $\left(U_0,U_1\right)$ is the source of endogeneity in the treatment selection. Researchers observe the random vector $W = \left(Y,S,X',Z'\right)'$, but do not observe $Y_1$, $Y_0$, $U_1$, $U_0$ or $U_S$.
Following carneiro/lee:2009, we further introduce the following notations for convenience. Let $V=F_{U_S}\left(U_S\right)$ denote unobserved innate propensity to select into treatment, and let $P=F_{U_S}\left(\mu_S\left(Z\right)\right)$ denote the treatment selection probability or the propensity score. Under these setup and notations, we adopt the following conventional assumption.
We can write the treatment selection model in this setting by
This representation provides the interpretation that those individuals with $V=P$ are at the margin of indifference between the two treatment statuses, and motivates the marginal treatment effect bjorklund/moffitt:1987 defined by
This treatment parameter serves as a building block for many treatment parameters heckman/vytlacil:1999,heckman/vytlacil:2001,heckman/vytlacil:2005. We consider a counterfactual propensity score $P^\ast=P^\ast\left(P,Z\right)$ with a known function $P^\ast\left(\cdot,\cdot\right)$ satisfying Assumption (ref) to be stated ahead.\footnote{carneiro/heckman/vytlacil:2010 consider alternative specifications for counterfactual policies $P^\ast$. One scenario takes the form of $P^\ast=P^\ast\left(P,Z\right)$, in which a policy maker directly affects the probability of being in the treatment group. Another scenario takes the form of $P^\ast=P\left(Z^\ast\left(Z\right)\right)$, in which the counterfactual policy affects the distribution of $Z$. We consider the former scenario in the main text, whereas we consider the latter scenario in Appendix (ref).} Under a counterfactual propensity score $P^\ast$, the counterfactual treatment and outcome are
The parameter of our interest is the policy relevant treatment effect heckman/vytlacil:1999,heckman/vytlacil:2001,heckman/vytlacil:2005:
This treatment parameter is of policy interest because it allows to compare alternative virtual policies $P^\ast$ under consideration. heckman/vytlacil:1999,heckman/vytlacil:2001,heckman/vytlacil:2005 has demonstrated that $PRTE$ can be represented as the weighted average of $MTE\left(x,p\right)$: $$ PRTE=\int\int_0^1MTE\left(x,p\right)\frac{F_{P\mid X}\left(p\mid x\right)-F_{P^\ast\mid X}\left(p\mid x\right)}{E[P^\ast]-E[P]}dpf_X\left(x\right)dx. $$ Note that $MTE\left(x,p\right)$ measures the average treatment effects for the subpopulation of those individuals with $V=p$ and $F_{P|X}\left(p|x\right)-F_{P^\ast|X}\left(p|x\right)$ measures the probability of compliance with the counterfactual policy for this subpopulation. Therefore, setting aside the denominator $E[P^\ast]-E[P]$, this integral measures the average treatment effect of the counterfactual policy for the population. A division of quantity by $E[P^\ast]-E[P]$ in turn yields the average treatment effect among those who complied with the policy.
In the rest of this paper, we propose a method of double debiased estimation and inference for the PRTE. Since the identification of the PRTE hinges on that of the marginal treatment effect, the theories of our estimation and inference methods rely on the prior result of the identification of the marginal treatment effect, which is formally stated in the following theorem.
In actual empirical research, researchers often use observed controls $X$. Without imposing additional structural restrictions, they would suffer from the curse of dimensionality of $X$ in estimation of the marginal treatment effect. The existing literature proposes alternative suggestions to address this issue. Following carneiro/lee:2009 and carneiro/heckman/vytlacil:2010, we employ the following structural restriction on $\mu_1$ and $\mu_0$.
Under Assumptions (ref) and (ref), we can write the conditional mean of $Y$ given $\left(X,P\right)$ as $$ E[Y\mid X,P]=P\mu_1\left(X\right)'\beta_1+\left(1-P\right)\mu_0\left(X\right)' \beta_0+E[U\mid P], $$ where $U=SU_1+\left(1-S\right)U_0$. Thus the marginal treatment effect takes the partial linear form of
where $\Delta_{U\mid P}\left(p\right) = \frac{d}{dp} E[U|P=p]$. In the rest of the paper, we shall use this form ((ref)) of the MTE expression for analysis of the PRTE:
where $\boldsymbol{g}_{S\mid Z}\left(z\right) = E[S|Z=z]$. Appendix (ref) provides a derivation of (ref).
We close this section by discussing feasible and infeasible extensions and variants of our parameter, $PRTE$, of interest. First, the PRTE may be defined as $E[\omega\left(X\right)PRTE\left(X\right)]$ with a known weight function $\omega$ for heterogeneous policy designs. Our method and theory presented ahead extends to this variant of the PRTE by replacing $Y$ by $\omega\left(X\right)Y$. Second, our framework also extends to analysis of $PRTE\left(x\right)$ when $X$ is discrete by focusing on the subpopulation with $X=x$. When $X$ is continuous, however, it is generally infeasible to estimate $PRTE\left(x\right)$ at the root-$n$ convergence rate.
While we present a full-fledged framework and formal asymptotic theories in Sections (ref) and (ref), respectively, we first provide an informal overview in this section focusing on a simplified model without covariates.
Define $\theta = \left(\theta_N,\theta_D\right)$ by \begingroup \allowdisplaybreaks
\endgroup where $\Delta_{U\mid P}\left(p\right) = \frac{d}{dp} E[U|P=p]$ and $\boldsymbol{g}_{S\mid Z}\left(z\right) = E[S|Z=z]$ from Section (ref). In the current setting without $X$, the PRTE consists of the last term in (ref) and can be written as
We collect the possibly infinite-dimensional nuisance parameters as
Let $L>1$ be a predetermined natural number of folds that is fixed as the sample size increases. Randomly split the sample into sub-samples $I_1,...,I_L$ of (approximately) equal size, i.e., $|I_\ell| = \lfloor n/L \rfloor$ or $\lfloor n/L \rfloor+1$. For each $\ell \in L$, estimate $\gamma$ by using the sub-sample $I_\ell^c$, and denote the estimator by $\hat\gamma_\ell$.\footnote{See Appendix (ref) for concrete estimators along with tuning parameter choice rules used for them in our empirical application.} We then define our double debiased estimator $\hat\theta=\left(\hat\theta_N,\hat\theta_D\right)$ of $\theta$ as the solution to
where $\left(m_N,m_D\right)$ is the orthogonal score defined by \begingroup \allowdisplaybreaks
\endgroup
Plug $\hat\theta$ into (ref) to in turn obtain an estimator of the PRTE: $$ \widehat{PRTE} = \hat\theta_N / \hat\theta_D. $$ This double debiased estimator for the PRTE is root-$n$ asymptotically normal as $$ \sqrt{n}\left(\widehat{PRTE}-PRTE\right) \rightarrow_d N\left(0,D' \Omega D\right), $$ where $ D = \left(1/\theta_D, \ -\theta_N/\theta_D^2\right)' $ and $ \Omega = \text{Var}\left(\left( m_N\left(Y,Z;\theta_N,\gamma\right), \ m_D\left(Y;\theta_d,\gamma\right)\right)'\right). $
Having outlined the step-by-step procedure of our proposed method, we next present intuitions behind this proposed method of double debiased estimation and inference for the PRTE as well as some heuristic explanations of why it works.
{\bf On the orthogonal score:} If we knew the nuisance parameters $(\boldsymbol{g}_{S|Z}, \boldsymbol{g}_{U|P})$, then we could estimate $\theta$ by using the moment conditions $E[M_N\left(Y,Z;\theta,\gamma\right)]=E[M_D\left(Y,Z;\theta,\gamma\right)]=0$, where
Since we estimate $(\boldsymbol{g}_{S|Z}, \boldsymbol{g}_{U|P})$, however, influence function adjustments should be added as in (ref), (ref) and (ref). Lines (ref) and (ref) precisely correspond to the na\"ive moment functions (ref) and (ref), respectively. An influence function adjustment for the estimation of ${\boldsymbol{g}}_{U\mid P}$ appears in (ref). An estimation of ${\boldsymbol{g}}_{U\mid P}\left(P\right)$ entails an influence function adjustment by $Y-{\boldsymbol{g}}_{U\mid P}\left({\boldsymbol{g}}_{S\mid Z}\left(Z\right)\right)$. Since the function ${\boldsymbol{g}}_{U\mid P}$ is evaluated at $P^\ast\left({\boldsymbol{g}}_{S\mid Z}\left(Z\right),Z\right)$ in line (ref), the adjustment is scaled by the coefficient ${{f}_{P^\ast}\left({\boldsymbol{g}}_{S\mid Z}\left(Z\right)\right)}/{{f}_{P}\left({\boldsymbol{g}}_{S\mid Z}\left(Z\right)\right)}$ as in (ref). Likewise, an estimation of ${\boldsymbol{g}}_{S\mid Z}\left(Z\right)$ entails an influence function adjustment by $S-{\boldsymbol{g}}_{S\mid Z}\left(Z\right)$, and it appears in lines (ref) and (ref). Since the function ${\boldsymbol{g}}_{S\mid Z}\left(Z\right)$ shows up inside other functions in lines (ref) and (ref), the adjustments are scaled by their derivatives in lines (ref) and (ref). Thanks to these influence function adjustments, these moment functions $m_N$ and $m_D$ are robust against local perturbations in $\gamma$, i.e., $m_N$ and $m_D$ constitute an orthogonal score. This property, together with the cross fitting discussed below, allows the effects of an estimation of $\gamma$ on an estimation of $\theta$ to be asymptotically negligible.
{\bf On the cross fitting:} Using the same sample to estimate both $\gamma$ and $\theta$ would result in an over-fitting bias. To circumvent such a bias, we use the sub-sample $I_\ell^c$ to estimate $\gamma$ by $\hat\gamma_\ell$ and use the complementary sub-sample $I_\ell$ to evaluate the orthogonal score, $|I_\ell|^{-1}\sum_{i\in I_\ell} m_N\left(Y_i,Z_i;\theta_N,\hat\gamma_\ell\right)$ and $|I_\ell|^{-1}\sum_{i\in I_\ell} m_D\left(Y_i;\theta_D,\hat\gamma_\ell\right)$, for each $\ell = 1,...,L$. These lead to the double debiased estimator $\hat\theta$ introduced in Section (ref).
{\bf On the root-$n$ asymptotic normality:} Because of the orthogonality property and the cross fitting that allow the effects of the estimation error of $\gamma$ on an estimation of $\theta$ to be asymptotically negligible, $\hat\theta$ enjoys the root-$n$ asymptotic normality
as if $\gamma$ were known. The delta method thus yields the root-$n$ asymptotic normality for the PRTE estimator: $$ \sqrt{n}\left(\widehat{PRTE}-PRTE\right) \rightarrow_d N\left(0,D' \Omega D\right), $$ again as if $\gamma$ were known.
{\bf Orthogonal Score and Double Robustness:} The orthogonality property and double robustness are closely related concepts, but neither implies the other -- see chernozhukov2016locally for example. A natural question is whether our orthogonal score also satisfies the double robustness with respect to the nuisance parameters, $f_{P^\ast}\left({\boldsymbol{g}}_{S\mid Z}\right)/f_P\left({\boldsymbol{g}}_{S\mid Z}\right)$, ${\boldsymbol{g}}_{S\mid Z}$, and $\left({\boldsymbol{g}}_{U\mid P},\Delta_{U\mid P}\right)$. It turns out that our orthogonal score does not possess the double robustness property against the propensity score function ${\boldsymbol{g}}_{S\mid Z}$ in general. This is because ${\boldsymbol{g}}_{S\mid Z}$ appears in our score inside possibly nonlinear functions -- we show a concrete case in point in Appendix (ref). On the other hand, given a fixed propensity score function ${\boldsymbol{g}}_{S\mid Z}$, our orthogonal score has double robustness between $f_{P^\ast}\left({\boldsymbol{g}}_{S\mid Z}\right)/f_P\left({\boldsymbol{g}}_{S\mid Z}\right)$ and $\left({\boldsymbol{g}}_{U\mid P},\Delta_{U\mid P}\right)$. This is because the score takes forms of products of affine functions of these nuisance parameters -- see Appendix (ref) for details.
{\bf More intuitions:} Using this discussion on the relation to the double robustness, we now provide more intuitions behind why the orthogonal score works to our goal. One may wonder why our PRTE estimator converges at the rate of root-$n$ while possibly all the preliminary estimators converge at slower rates. As demonstrated in Appendix (ref), our orthogonal score after some rewritings takes the form of a product of two estimation errors as in the right-hand side of
See Appendix (ref) for its derivation. This shows that, if each of $ \tilde{f}_{P^\ast}\left( \cdot \right)/\tilde{f}_{P}\left( \cdot \right)-{f}_{P^\ast}\left( \cdot \right)/{f}_{P}\left( \cdot \right) $ and $ {\boldsymbol{g}}_{U\mid P}\left( \cdot \right)-\tilde{\boldsymbol{g}}_{U\mid P}\left( \cdot \right) $ converges to zero at a rate faster than $n^{1/4}$ (yet slower than $n^{1/2}$), then the product converges at a rate faster than $n^{1/4} \cdot n^{1/4} = n^{1/2}$. This property of the orthogonality (taking the form of a product in this case) solves the puzzle that the PRTE estimator of our interest can converge at the rate of root-$n$ while convergence rates of the preliminary estimators are often slower.
From Theorem (ref), we can express the PRTE as a function of estimable moments. Consider the vector, $\theta=\left(\theta_1',\theta_2',\theta_3\right)'$, of estimable moments defined by \begingroup \allowdisplaybreaks
\endgroup where $\boldsymbol{g}_{S\mid Z}\left(z\right) = E[S|Z=z]$, $\Delta_{U\mid P}\left(p\right) = \frac{d}{dp} E[U|P=p]$,
$\boldsymbol{g}_{\mu_0\left(X\right)\mid P}\left(p\right) = E[\mu_0\left(X\right) | P=p]$, $\boldsymbol{g}_{\mu_1\left(X\right)\mid P}\left(p\right) = E[\mu_1\left(X\right) | P=p]$, and $\boldsymbol{g}_{Y\mid P}\left(p\right) = E[Y | P=p]$.
To express $PRTE$ in ((ref)) as a function of $\theta$, we write $$ PRTE = \Lambda\left(\theta\right) \equiv \frac{\theta_{2,1}' \boldsymbol{d}_1\left(\theta_1\right) - \theta_{2,0}' \boldsymbol{d}_0\left(\theta_1\right)+ \theta_3}{\theta_{2,2}}, $$ where $\boldsymbol{d}=\left(\boldsymbol{d}_0',\boldsymbol{d}_1'\right)'$ is the function defined by $\boldsymbol{d}\left(\mathrm{vec}\left(\mathbf{B},\mathbf{A}\right)\right)={\mathbf{B}^{-1}} {\mathbf{A}}$ for a $2p \times 1$ vector $\mathbf{A}$ and a $2p \times 2p$ matrix $\mathbf{B}$, cf. Equation (ref) below for the function $\boldsymbol{d}$.
Note that the definition of $\xi_1$ comes from the moment condition for $\left(\beta_0',\beta_1'\right)'$. Specifically, as in robinson1988root, the parameter vector $\left(\beta_0',\beta_1'\right)'$ can be written as
Recall the notation $W = \left(Y,S,X',Z'\right)'$ from Section (ref). With these definitions and notations, we now propose an orthogonal score function of the form
with all the preliminary parameters collected in the concise notation
where $\zeta\left(z\right)=E\left[\left.\left.\frac{\partial}{\partial p} \xi_1\left(X,Y,p\right)\right|_{p=\boldsymbol{g}_{S\mid Z}\left(z\right)} \right\vert Z=z\right]$. The components of the orthogonal score function are defined by \begingroup \allowdisplaybreaks
\endgroup where $\mathcal{U}\left(W,\theta\right)=Y-\left(1-S\right)\mu_0\left(X\right)'\beta_0-S\mu_1\left(X\right)'\beta_1$ and $\partial P^\ast\left(p,z\right)=\frac{\partial}{\partial p}P^\ast\left(p,z\right)$. See Appendix (ref) for an alternative representation of $m\left(W;\theta,\gamma\right)$
This score function $m$ is orthogonal in the sense that
See Appendix (ref) for a derivation of this property as well as the definition of $\Gamma$. To arrive at the orthogonal score, we take advantage of two convenient features of $\theta$ and $\gamma$. First, all the nuisance parameters $\gamma$ take forms of mean-square projections or densities. Second, the nuisance parameters $\gamma$ enter the equations for $\theta$ through integrals and derivatives, and therefore the Gateaux derivative operator (and thus the orthogonalization operator) can directly act on $\gamma$. These two features allow us to obtain the adjustment terms using the result of newey1994asymptotic.
The orthogonality (ref) allows for the score to be insensitive to local perturbations $\breve{\gamma}$ of $\gamma$. In particular, the first-order effect of the estimation error $\hat\gamma-\gamma$ is zero, and the remaining effects are of a smaller order: $$ \int m\left(w;\theta,\hat{\gamma}\right)F_W\left(dw\right)=o_p\left(n^{-1/2}\right). $$ In other words, the effects of the estimation error $\hat{\gamma}_\ell - \gamma$ of preliminary parameters are asymptotically negligible relative to the convergence rate of the empirical mean, which is of order $n^{-1/2}$. Although we postpone rigorous discussions until Section (ref) and proofs of the main theorem therein, we here discuss intuitions behind the orthogonality property. If we knew the true functions ${\xi}_1$, ${\boldsymbol{g}}_{U\mid P}$, and ${\boldsymbol{g}}_{S\mid Z}$, then we would use the moment conditions $$ E\left[ \left(
\right) \right] =0, $$ which correspond to line \eqref{eq:m_1_a} for $m_1$, line \eqref{eq:m_2_a} for $m_2$, and line \eqref{eq:m_3_a} for $m_3$. However, since we estimate ${\xi}_1$, ${\boldsymbol{g}}_{U\mid P}$, and ${\boldsymbol{g}}_{S\mid Z}$, influence function adjustments need to be added for these estimators.
First, the influence function adjustment for the estimated ${\xi}_1$ is zero. This is because, as is well known in the literature on partial linear models robinson1988root, the moment condition $E[{\xi}_1\left(X,Y,{\boldsymbol{g}}_{S\mid Z}\left(Z\right)\right)-\tilde{\theta}_1]=0$ satisfies the orthogonality property with respect to ${\xi}_1$. Second, the influence function adjustment to account for an estimation of ${\boldsymbol{g}}_{U\mid P}$ appears in (ref). An estimation of ${\boldsymbol{g}}_{U\mid P}\left(P\right)$ entails an influence function adjustment by $\mathcal{U}\left(W,{\theta}\right)-{\boldsymbol{g}}_{U\mid P}\left({\boldsymbol{g}}_{S\mid Z}\left(Z\right)\right)$. Since the function ${\boldsymbol{g}}_{U\mid P}$ is evaluated at $P^\ast\left({\boldsymbol{g}}_{S\mid Z}\left(Z\right),Z\right)$ in line (ref) for $m_3$, this adjustment is scaled by the coefficient ${{f}_{P^\ast}\left({\boldsymbol{g}}_{S\mid Z}\left(Z\right)\right)}/{{f}_{P}\left({\boldsymbol{g}}_{S\mid Z}\left(Z\right)\right)}$ as in (ref). Third, the influence function adjustments for the estimated ${\boldsymbol{g}}_{S\mid Z}$ appear in line (ref) for $m_1$, line (ref) for $m_2$, and line (ref) for $m_3$. An estimation of ${\boldsymbol{g}}_{S\mid Z}\left(Z\right)$ entails an influence function adjustment by $S-{\boldsymbol{g}}_{S\mid Z}\left(Z\right)$. Since the function ${\boldsymbol{g}}_{S\mid Z}$ appears inside other functions in lines (ref), (ref), (ref) and (ref), these adjustments are scaled by their derivatives of the respective functions, which result in the coefficients in front of $S-{\boldsymbol{g}}_{S\mid Z}\left(Z\right)$ in lines (ref), (ref) and (ref). Note that all these adjustments take the form prescribed by newey1994asymptotic. This is because, as mentioned earlier, our moment functions share two convenient properties. First, $\gamma$ takes forms of mean-square projections or densities. Second, $\gamma$ enter the equations for $\theta$ so that the Gateaux derivative operator (and thus the orthogonalization operator) can directly act on $\gamma$. Formal mathematical analyses rationalizing these intuitions are found in Appendix (ref).
Given a random sample $\{W_i\}_{i=1}^n$ of size $n$, we now use the above orthogonal score to construct a double debiased estimator with cross fitting or sample splitting. Let $L > 1$ be a natural number of folds, and randomly partition the sample index set $\{1,...,n\}$ into $L > 1$ subsets $I_1,...,I_L$ of approximately equal size. For every subsample index $\ell \in \{1,...,L\}$, let $$\hat{\gamma}_\ell= \left( \frac{\hat{f}_{P^\ast}\left(\hat{\boldsymbol{g}}_{S\mid Z}\right)}{\hat{f}_{P}\left(\hat{\boldsymbol{g}}_{S\mid Z}\right)}, \hat{\boldsymbol{g}}_{S\mid Z}, \hat{\xi}_1, \hat{\zeta}, \hat{\boldsymbol{g}}_{U\mid P}, \hat{\Delta}_{U\mid P} \right)$$ denote a preliminary parameter estimate obtained by using all observations $i \in \{1,...,n\} \backslash I_\ell$.\footnote{For the components of the preliminary parameter estimate $\hat{\gamma}_\ell$, we omit $\ell$ for the sake of notational simplicity.} We define our double debiased estimator $\hat\theta$ of $\theta$ as the solution to
Accordingly, our double debiased estimator $\widehat{PRTE}$ of the PRTE is defined by
In this section, we investigate the large sample theory for the double debiased estimators, $\hat\theta$ and $\widehat{PRTE}$, defined in (ref) and (ref), respectively. To this end, we formally present and discuss relevant assumptions. First, we consider a counterfactual policy $P^\ast$ satisfying the following conditions.
Recall that many treatment parameters can be represented in terms of the MTE in similar manners to the PRTE -- see heckman/vytlacil:2005. Assumption (ref) rules out some of the parameters such as $ATT$, $ATU$, $ATE$, $LATE\left(x\right)$, and $TUT\left(x\right)$.
Second, assume that all the preliminary parameter components of the orthogonal moment function are identified in the following sense.
The invertibility condition serves to identify the parameter vector $\left(\beta_0',\beta_1'\right)'$ in the partial linear model as in carneiro/lee:2009. The requirement for $\gamma$ to be identified essentially boils down to the identification of the density functions and the conditional expectation functions, which is satisfied for a large class of data generating processes, and is usually taken for granted in the nonparametrics literature.
We next make three high-level conditions (Assumptions (ref)--(ref)) regarding large sample behaviors of the preliminary parameter estimators. While we stress that each of these three assumptions is a mild requirement for most common nonparametric estimators, we supplement each of the three non-primitive assumptions by lower-level sufficient conditions in Appendix (ref).
Assumption (ref) requires that some preliminary parameter estimators to have the $n^{1/4}$ rate of convergence. This slow rate of convergence suffices because of the orthogonal score we developed and proposed in Section (ref) -- see the intuitions and discussions in Section (ref). As far as this condition is satisfied, the asymptotic distribution of our parameters of interest, namely $\hat\theta$ and $\widehat{PRTE}$, will not be affected by estimation errors of these preliminary parameter estimates. Furthermore, this condition can be satisfied by common nonparametric estimators of density and conditional expectation functions. For completeness, we show lower-level sufficient conditions for Assumption (ref) in Appendix (ref).
We show lower-level sufficient conditions for Assumption (ref) in Appendix (ref).
The final piece of assumptions is to control higher-order effects of the propensity score estimation on the orthogonal score. Note that the orthogonality property (ref) forces the first-order effects of nuisance parameter estimation to be zero. Since all the nuisance parameters except for the propensity score function $\boldsymbol{g}_{S|Z}$ appear in our score in a linear manner, the orthogonality property precisely vanishes their estimation effects. On the other hand, since the propensity score function $\boldsymbol{g}_{S|Z}$ appears in our score nonlinearly in general, we need to make sure that the remaining higher-order effects are small enough. Specifically, we require the following condition.
Since higher-order effects are usually of smaller magnitudes, Assumption (ref) is a quite mild condition. We present lower-level sufficient conditions for Assumption (ref) in Appendix (ref).
In addition to these conditions, we also present regularity conditions (Assumptions (ref), (ref) and (ref)) in Appendix (ref). These additional conditions are relegated to the appendix for the sake of readability, as they are even more standard, even milder, and are even easier to verify with common nonparametric estimators than Assumptions (ref)--(ref).
We obtain the following asymptotic normality result for the double debiased estimator $\hat\theta$ of the intermediate parameter vector.
A proof is provided in Appendix (ref). An immediate consequence of this result through the delta method is the following asymptotic normality result for the double debiased estimator $\widehat{PRTE}$ of the PRTE.
As we emphasized throughout this paper, this root-$n$ convergence without any influence of preliminary estimation errors $\hat\gamma-\gamma$ is possible as a result of the orthogonality that implies (ref). This method can accommodate a wide array of preliminary estimation techniques including kernel-smoothing, sieve estimation, and shrinkage methods among others. This flexibility in the choice of preliminary estimation approaches at the current high-level theory is another advantage of using the orthogonal score, unlike conventional methods in absence of orthogonal scores. A drawback of this root-$n$ asymptotic normality result is that it is not guaranteed to be efficient.
We can estimate every component of the asymptotic variance for $\widehat{PRTE}$ as follows. First, we can estimate $\lambda\left(\theta\right)$ by $\lambda\left(\hat\theta\right)$ since $\lambda$ is a known function. Second, estimate $\mathcal{M}$ by $$ \hat{\mathcal{M}}=\left(
\right). $$ Last, estimate $E[m\left(W;{\theta},{\gamma}\right)m\left(W;{\theta},{\gamma}\right)']$ by
The following theorem guarantees the asymptotic validity of the resultant variance estimator.
A proof is provided in Appendix (ref).
In this section, we use Monte Carlo simulations to evaluate finite sample performance of the proposed method of double debiased estimation and inference. We first consider a benchmark design from the literature in Section (ref), and demonstrate that even completely nonparametric preliminary estimation can produce desirable finite-sample performance thanks to the orthogonal score. We second extend the design with higher dimensions of $Z$ in Section (ref).
We generate artificial datasets from the same distribution as in carneiro/lokshin/umapathi:2017. Specifically, we generate $$ \left(U_0,U_1,U_S\right)=\left(-0.050 \varepsilon_1 + 0.020 \varepsilon_3, 0.012 \varepsilon_1 + 0.010 \varepsilon_2,-1.000 \varepsilon_1\right), $$ where $\varepsilon_1$ and $\varepsilon_2$ are independent standard normal random variables. Then, we generate \begingroup \allowdisplaybreaks
\endgroup where $X_1 \sim N\left(-2,2^2\right)$, $X_2 \sim N\left(2,2^2\right)$, $Z_1 \sim N\left(-1,3^2\right)$, and $Z_2 \sim N\left(1,3^2\right)$ are mutually independent and are also independent of $\left(U_1,U_0,U_S\right)'$. The observed outcome is generated in turn as $Y = SY_1 + \left(1-S\right)Y_0$. We consider two forms of policy changes. The first takes the form of $P^\ast = P + a\left(1-P\right)$, where $a \in \{0.1,...,0.9\}$.\footnote{Note that $PRTE \rightarrow MPRTE$ carneiro/heckman/vytlacil:2010 as $a \rightarrow 0$ and $PRTE \rightarrow ATU$ as $a \rightarrow 1$.} The second takes the form of $Z^\ast = Z + a$, where $a \in \{-0.5,...,-0.1,0.1,...,0.5\}$. See Appendix (ref) for additional details about this simulation setting, including implied analytic expressions for the counterfactual distribution, marginal treatment effects, and the PRTEs.
For each instance of artificial datasets, we estimate the PRTE and its estimated standard error using our proposed method of double debiased estimation and inference with completely nonparametric preliminary estimation. See Appendix (ref) and (ref) for additional details about concrete estimation and inference procedures. We experiment with alternative numbers, $L=5$ and $10$, of folds in cross fitting, where the number of observations in each fold is equal. We report the simulated bias, root mean square error, and coverage frequencies for the nominal probability of 95%. Simulated coverage frequencies are computed based on the standard symmetric confidence interval generated by the PRTE estimate and its estimated standard error according to Corollary (ref). The number of Monte Carlo iterations is set to 1000 following carneiro/lokshin/umapathi:2017. Since there is no existing limit distribution theory for inference about the PRTE to our best knowledge, we do not have any benchmark in the literature against which to make comparisons of our results. We therefore present simulation results only for our proposed method.
Tables (ref) and (ref) summarize simulation results for policy changes of the forms $P^\ast = P + a\left(1-P\right)$ and $Z^\ast = Z + a$, respectively. The results concerning the mean and bias demonstrate that the proposed double debiased estimator indeed produces small biases (relative to the root mean square error) even in small samples. In light of these relatively small biases, the results for the root mean square error demonstrate that the estimator converges approximately at the rate of $n^{-1/2}$, consistently with our theory. The results concerning the coverage frequencies demonstrate that our limit normal distribution results are useful to construct asymptotically valid confidence intervals for the PRTE, except when $a$ is infinitesimal. The alternative numbers, $L=5$ and $10$, of folds for sample splitting entail quite similar simulation results in terms of all of the bias, root mean square error, and coverage frequencies. In summary, the proposed estimator enjoys the main properties suggested in this paper even in small samples; namely, it is debiased, converges at the parametric rate, and is asymptotically normal with the proposed asymptotic variance formula, regardless of the number of folds in sample splitting.
In the previous subsection, we demonstrate that even completely nonparametric preliminary estimation yields desirable finite-sample performances. We next consider extended designs with high dimensions, employ parametric propensity score estimation with absolute shrinkage penalization, and present finite-sample performances of our estimation and inference procedures in this setting.
Suppose that the selection model is specified by \begingroup \allowdisplaybreaks
\endgroup where $Z_1 \sim N\left(-1,3^2\right)$ and $Z_2, ..., Z_{\text{dim}\left(Z\right)} \sim N\left(1,3^2\right)$. This design includes that of the benchmark by carneiro/lokshin/umapathi:2017 studied in the previous subsection as a special case where $\text{dim}\left(Z\right) = 2$. Furthermore, this design preserves the same value of PRTE as that in the benchmark by carneiro/lokshin/umapathi:2017 for each $a$, regardless of the dimension $\text{dim}\left(Z\right) \ge 2$. In this manner, we devise this extended design to allow for comparisons of simulation results with those presented in Section (ref). We set the high-dimensional setting with $\text{dim}\left(Z\right) = 100$ in this design.
Estimates for the PRTE and their estimated standard errors are obtained using our proposed method of double debiased estimation and inference with preliminary estimation of propensity scores by lasso logit. See Appendix (ref) for additional details. Table (ref) summarizes simulation results. Note that we obtain qualitatively similar results to those presented in Table (ref), and hence similar remarks follow. It is worth noting that the magnitude of the statistics (such as the bias the RMSE) displayed in this table is similar to that in Table (ref) despite the difference in $\text{dim}\left(Z\right)$ and the preliminary estimators. This observation is also consistent with the property of the orthogonal score that it reduces and asymptotically vanishes effects of preliminary estimation errors.
In this section, we apply the proposed method to an analysis of the effects of counterfactual policies that expose various fractions of the population to upper secondary schooling in Indonesia. Following the study by carneiro/lokshin/umapathi:2017, we use the data from the third wave of the Indonesia Family Life Survey (IFLS) fielded from June through November 2000 -- we refer readers to their supplementary appendix for further details of this data set. We also use the same subsample as that of carneiro/lokshin/umapathi:2017, that consists of employed males aged 25--60, who have reported non-missing wage and schooling information. This subsample consists of 2,608 individuals.
We set our variables following carneiro/lokshin/umapathi:2017. The outcome variable $Y$ denotes the log of hourly wages constructed from self-reported monthly wages and hours worked per week. The binary treatment variable $S$ indicates attendance of upper secondary school or higher. Control variables $X$ include age, age squared, an indicator for whether the individual was living in a village at age 12, indicators for the province of residence, an indicator of rural residence, distance from the office of the head of the community of residence to the nearest community health post, and indicators for the level of schooling by each parent. The excluded instrument is the distance from the office of the head of the community of residence in kilometers to the nearest secondary school. carneiro/lokshin/umapathi:2017 provide detailed analysis to support the validity of this instrument. In this paper, we take advantage of their analysis and discussions, and directly adopt their empirical approach in our estimation and inference framework. See Appendix (ref) for details of this model specification.
We consider the counterfactual policies of the form $P^\ast = P + a \left(1-P\right)$ with various levels of $a \in \{0.05,0.10,...,0.45,0.50\}$. Note that such a counterfactual treatment probability $P^\ast$ arises from the counterfactual policy that exposes fraction $a$ of the population to upper secondary schooling. The estimation and inference approaches that we take in this analysis are the same as those used for our Monte Carlo simulation studies except that we use the probit propensity score estimation following the benchmark study by carneiro/lokshin/umapathi:2017 -- see Appendix (ref) for a step-by-step procedure of estimation and inference. We use $L=5$ as the number of folds in cross fitting.\footnote{We also ran another set of estimation with $L=10$ but the results are almost the same both in terms of estimates and their standard errors, similarly to what we observed in our Monte Carlo simulation studies.} Finally, to mitigate the finite-sample randomness in estimates due to sample splitting, we use the robust re-randomization method following chernozhukov2018double.
Table (ref) summarizes the results. For each $a \in \{0.05,0.10,...,0.45,0.50\}$, displayed in this table are the point estimate, standard error, and the 95% confidence interval. Observe that the policy that assigns a fraction $a=0.05$ of the population to upper secondary schooling is expected to increase the log of hourly wages by 0.169 (with the standard error of 0.064 and the 95% confidence interval of $[0.044,0.294]$) per treated individual on average. As the fraction $a$ of the population assigned to treatment increases, these per-individual average policy relevant treatment effects tend to increase. Specifically, the policy that assigns a fraction $a=0.50$ of the population to upper secondary schooling is expected to increase the log of hourly wages by 0.225 (with the standard error of 0.076 and the 95% confidence interval of $[0.076,0.375]$) per treated individual on average.
Estimation of the PRTE involves estimation of multiple preliminary parameters. These preliminary parameters include propensity scores, conditional expectation functions of the outcome and covariates given the propensity score, and marginal treatment effects. These preliminary estimators can affect the asymptotic distribution of the PRTE estimator in complicated and intractable manners. To solve this issue, we propose an orthogonal score for double debiased estimation of the PRTE. Our proposed orthogonal score allows for the asymptotic distribution of the PRTE estimator to be obtained without any influence of preliminary parameter estimators as far as they satisfy mild convergence rate conditions. Simulation results confirm our theoretical properties, and demonstrate that the method of estimation and inference works well even in small samples. Our empirical application demonstrates that the proposed method indeed works with real data. To our knowledge, our work is the first to develop asymptotic distribution theories for inference about the PRTE. We hope that our proposed method contributes to empirical analyses of policy relevant treatment effects.