Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
210,504 characters · 14 sections · 0 citation commands
Identification and Auto-debiased Machine Learning for Outcome Conditioned Average Structural Derivatives
\baselineskip=20pt
\largeKeywords: Heterogeneity, Local average structural derivative, Debiased machine learning, Doubly/locally robust score, Unconditional quantile partial effect, Counterfactual policy effect;
\thispagestyle{empty}
When determining the causal effect of an intervention or a treatment of interest $D$ on an outcome $Y$, applied researchers more often focus on mean quantities such as average treatment effect (ATE) than distributional parameters such as quantile treatment effect (QTE), because the formers are easier to interpret. Since unobserved heterogeneity is so pervasive in microdata, distributional parameters such as quantile regression (QR) coefficients or QTE still play an important role in summarizing heterogeneous impacts of variables on different points of an outcome distribution. For example, in understanding the effect of a tax credit reform, one might be more interested in the effect of the tax rate change on the lower tail of the labor supply or savings distribution conditional on individual characteristics than the mean effect. This is the question answered exactly by QR. However, to understand heterogenous effects of such reform, one may ask a different question from the above one: what is the average effect of the tax rate change on labor supply or savings for the individuals at the lower tail of the labor supply (or savings) distribution, irrespective of (by integrating out) individual characteristics?
Formally, assume that there is an outcome variable $Y$ (labor supply, savings or consumptions) with a continuous support $\mathcal{S}_{Y}\subset \mathbb{R}$, a continuous treatment variable $D$ (tax rate) and a $K_{X}$-dimensional vector of covariates $X$, which are related through a general nonseparable structural model \[ Y=m\big(D,X,U\big)\eqno(1.1) \] with $U$ an unobservable random vector, capturing omitted factors and all types of unobserved heterogeneity. This model (1.1) is very general as it imposes on $m(\cdot)$ neither additivity structure nor monotonicity with respect to the error term. Such general model has been studied by Imbens and Newey (2009), Altonji and Matzkin (2005), and Hoderlein and Mammen (2007, 2009), Firpo (2009), Rothe (2010, 2012), etc.
Let $\partial_{D}$ be the derivative of $m(\cdot)$ with respect to its first argument. In this paper, we focus on the following class of parameters \[ \theta(y_{1},y_{2})=E\left(\partial_{D} m\big(D,X,U\big)\bigg|Y\in(y_{1},y_{2})\right)\eqno(1.2) \] indexed by $(y_{1},y_{2})\subset \mathcal{S}_{Y}$, which are called the outcome conditioned average structural derivatives (OASD). OASD combines both features of ATE and QTE: it is interpreted as straightforwardly as ATE while at the same time more granular than ATE by breaking the entire population up according to the rank of the outcome distribution.
The OASDs defined by (1.2) provide direct answers to a wide range of questions in applied economic analysis. To give a simple example, let $Y$ be log wage, $D$ be years of schooling and $X$ be other individual characteristics. Applied researchers mainly focus on the average partial effect or $E\left[\partial_{D} m(D,X,U)\right]$, which measures the mean gain of receiving one more year of schooling for all individuals. To explore the heterogenous effect of receiving more education, one may run quantile regression of $Y$ on $D$ and $X$ (Buchinsky (1994), Chamberlain (1994), Angrist, Chernozhukov and Fernandez-Val (2006)), and the coefficient on $D$ describes the impact of receiving one more year of schooling on the conditional quantiles of wage distribution. Sometimes policymakers may be interested in a related, but more straightforward question: what is the average effect of a small increase in schooling for low/middle/high wage individuals ? Let $y_{0.1}, y_{0.2}, \cdots, y_{0.9}$ equal to the $10\%, 20\%, \cdots, 90\%$-th empirical quantiles of the wage distribution, the latter question can be answered by estimating $\theta(y_{0.1},y_{0.2})$, $\cdots$, and $\theta(y_{0.8},y_{0.9})$ directly. If the analysis shows that high-earning individuals, on average, earn more than low-earning individuals, for instance, $\theta(y_{0.8},y_{0.9})>\theta(y_{0.2},y_{0.3})$, then one can tell that the dispersion in earnings is likely to go up with schooling and vice versa.
Our OASD and the subsequent identification strategy partially build on the insights of Hoderlein and Mammen (2007, 2009), who consider identification and estimation in a nonseparable model of the quantity
which they call the Local Average Structural Derivative (LASD). At first glance, OASD makes only one step forward given LASD by integrating $X$ out conditional on $Y\in(y_{1},y_{2})$. However, as we shall demonstrate later, such modification produces several desirable properties and sheds light on some interesting connections between OASD and other distributional parameters. Specifically, we establish two new identification results for OASD. The first one (Proposition 2.1) points out some important connection between OASD and the unconditional quantile partial effect (UQPE) proposed by Firpo et al. (2009). This relationship has two implications: it offers an alternative economic interpretation to UQPE, which can be understood as a mean causal effect on an identifiable subpopulation located at some level of the outcome distribution. Moreover, many desirable properties of OASD may carry over to the UQPE. For example, the semiparametric efficiency bound of UQPE may be learnt from that of OASD, which we shall derive in Section 4. Our second identification result (Proposition 2.2) derives an orthogonal score for OASD, which paves the way for developing subsequent automatic debiased machine learning estimator and facilitates establishing some desirable properties of the estimator. For example, we prove that our estimator is root-$n$ consistent and semi-parametrically efficient. We also mention that OASDs have some robustness property against censoring data. In many applications using micro dataset, censoring is a pervasive phenomenon. It is important that the estimator still works under censoring. All these results are new relative to Hoderlein and Mammen (2007, 2009).
Main contributions to the Literature. This paper makes three contributions to the literature. First, we propose OASDs as novel and parsimonious quantities to measure impacts of a continuous treatment that are heterogeneous across the unconditional distribution of an outcome. There are two major differences between OASD and the QR approach. QR aims to estimate impacts of explanatory variables at different points of the outcome distribution conditional on a large number of covariates, while OASD measures the mean effect of a given explanatory variable on the subpopulations located at different parts of the unconditional outcome distribution. \footnote{The difference between conditional vs unconditional outcome distribution can be understood by a simple example relating wages to years of education. The 0.9 quantile of wage distribution conditional on education, which is the subpopulation targeted by QR, refers to the high wage workers within each education class, who however may not necessarily be high earners overall. However, the unconditional 0.9 quantile of wage distribution, which is the subpopulation targeted by OASD, refers straightforwardly to the high wage workers.} OASDs are the right estimands to consider when the ultimate object of interest is the people located at specific parts (lower or upper tail) of the unconditional outcome distribution whatever individual characteristics (education, race, age) they have. In this sense, OASD can be viewed as generalizing the conception of unconditional quantile treatment effect (Firpo (2007), Frolich and Melly (2013)) for a binary treatment to a continuous treatment. The second difference between OASD and QR is more technical: OASDs are indexed by intervals $(y_{1},y_{2})\subset\mathbb{R}$ instead of $y\in \mathcal{S}_{Y}$. As will be clear later, this technical subtlety ensures the resulting estimator has several desirable properties such as $\sqrt{n}$ convergence and semiparametric efficiency. We provide two identification results for OASDs in a general nonseparable model (1.1) without monotonicity. When the treatment is binary, we show that the outcome conditioned average treatment effect is identified with an additional monotonicity condition (see Remark 2.2.1).
As a second contribution, we establish some close relationships between two classes of causal quantities which apparently have different interpretation: one is the outcome conditioned average partial effects (including OASD) studied in the current paper, the other is parameters measuring the effect of counterfactually changing the distribution of a single covariate on the unconditional distribution of the outcome. Examples of the latter class include the unconditional partial quantile effect (UPQE, Firpo et al. (2009)) and the marginal partial policy effect (MPPE, Rothe 2012). We show there is a close connection between OASD and UPQE, also between MPPE and a corresponding outcome conditioned average effect parameter. These results have an important implication: many desirable properties of OASD may well carry over to the UQPE. For example, to the best of our knowledge, the literature has not obtained any efficiency bound of UQPE. In Section 4, we derive the semiparametric efficiency bound of OASD, which can be used to learn about the efficiency bound of UQPE. As another example, Firpo et al. (2009, Proposition 1, page 959) has shown that the UQPE can be written as a weighted average of a family of conditional quantile partial effects (CQPE), under the monotonicity assumption. Using the equivalence result between UQPE and OASD, we show this result (Proposition 2.3-iii) still holds under much weaker assumptions without monotonicity. Similar arguments apply to MPPE.
The third contribution is that we propose a novel automatic debiased machine learning (ADML) estimator for OASD, by taking advantage of the debiased estimation approach recently developed by Belloni et al. (2017), Chernozhukov et al. (2022 a,b,c). For LASD, Hoderlein and Mammen (2009) have introduced a local polynomial kernel estimator and derived its large sample properties. Although this kernel based approach can be taken to estimate OASD, it cannot accommodate high dimensional controls and has no robustness against local perturbations in the nuisance functions. Since identification of OASD is attained under a conditional exogeneity assumption, by controlling for a rich information about covariates, a researcher may ideally use high-dimensional controls in data. Motivated by this, we contribute to the literature by proposing a first orthogonal score based estimator for OASD (and meanwhile for the LASD), that is shown to be root-$n$ consistent and semiparametrically efficient, and allowing for a flexibility in types of preliminary estimators, e.g., kernel, sieve, and Lasso.
Like QTE, our estimator for OASD falls under the framework where there are a continuum of finite-dimensional parameters of interest, identified via a continuum of moment conditions that involve a continuum of nuisance functions. We prove the uniform Gaussianity of the OASD process and the uniform validity of a multiplier bootstrap, by taking advantage of the general theory for the Lasso and post-Lasso estimators for functional response data established by Belloni et al. (2017).
Relationship to the Literature. This paper is related to two branches of the literature. The first branch is about identifying causal parameters, particularly those measuring heterogenous distributional impacts in nonseparable models. This branch can be broadly divided into two categories according to whether the treatment variable is discrete or continuous. Contributions to the category dealing with a binary treatment variable include Heckman and Vytlacil (2001) about the policy relevant treatment effects, and Heckman and Vytlacil (2005) about the marginal treatment effects, Firpo (2007) about the unconditional QTE, Frolich and Melly (2013) about the local QTE for compliers, Donald and Hsu (2014) about the distributional treatment effects, to name only a few.
Our work is more relevant to the second category about nonparametric identification and estimation of causal parameters or counterfactual policy effect of a continuous treatment, particularly without assuming the error term entering monotonically. Contributions to this category include Altonji and Matzkin (2005), Chernozhukov et al. (2013), Florens, Heckman, Meghir, and Vytlacil (2008), Imbens and Newey (2009) and Hoderlein and Mammen (2007, 2009), Rothe (2010, 2012), Firpo et al. (2009), and Ai et al. (2022). Less closely related to our work are Chesher (2003, 2005), Chernozhukov and Hansen (2005), Chernozhukov, Imbens and Newey (2007) who assume that the error term, at least at some stage, enters the model monotonically. Among them, Hoderlein and Mammen (2007, 2009), Firpo et al. (2009) are perhaps the most relevant to our work. We contribute to this branch of the literature from two aspects: we propose a new class of quantities, namely OASD to measure heterogeneous impacts of a continuous treatment; we provide insights into the relationship between OASD and a class of counterfactual policy effect parameters; and we obtain the semi-parametric efficiency bound for OASD.
Our paper is also related to another fast growing branch of the literature on estimation and inference of causal or structural parameters based on orthogonal scores. The seminal paper by Newey (1994) proposes orthogonal scores for many semiparametric models and provides forms of adjustment terms to obtain orthogonal scores from moment functions. Chernozhukov et al. (2022a) propose a general procedure for construction of orthogonal scores from moment restriction models. We derive the orthogonal score for OASDs following these general prescriptions. Usually, the orthogonal score depends on another unknown function, denoted as $\alpha$ ($\alpha=\partial_{D} \ln f(D,X)$ in OASD) in addition to the nonparametric components in the original moment. Chernozhukov et al. (2022 a,b,c) develop a Lasso minimum distance learner of $\alpha$, that is automatic in the sense that it depends only on the identifying moment function and not on the functional form of $\alpha$. Our proposed method of estimation takes advantages of these knowledge. A technical novelty in our proof is that we estimate CDF derivatives by a high order partial difference approach, with which we do not need to impose substantially more restrictive approximate sparsity conditions than before while preserving the good rates of convergence for high dimensional CDF and well as its derivatives.
In work related but independent from ours, Sasaki et al. (2022) propose a doubly robust score for debiased estimation of the UQPE (Firpo et al. 2009). Their results complement ours, though the motivation of the two papers are different: Sasaki et al. only consider estimation and inference of the UQPE as a measure of heterogeneous counterfactual marginal effects while we focus on outcome conditioned partial effect of continuous treatment. In the absence of the connection between UQPE and OASD (see Section 2.2), the parameters considered by the two papers are entirely different. Moreover, Sasaki et al. do not consider the semiparametric efficiency bound, nor derive the uniform Gaussian distribution of their estimator, while we establish the uniform limiting distribution of the OASD process.
Organization of the paper. The rest of this paper is organized as follows. Section 2 presents the setting, the main identification results and discusses the relationship between the OASD and other counterfactual policy effect parameters. Sections 3 develops an automatic debiased learning estimator for OASD. Section 4 provides theoretical guarantees of the estimator. Section 5 presents Monte Carlo simulation studies. The paper is summarized in Section 6. The appendix contains proofs and additional details that are important but relegated there due to their lengths.
Notations. We work with the i.i.d. data $\{W_{i}\}_{i=1}^{n}$ which is defined on the probability space $\left(\mathcal{W},\mathcal{A}_{\mathcal{W}},P\right)$. We denote by $\mathbb{P}_{n}$ the empirical probability measure that assigns probability $n^{-1}$ to each $W_{i}\in\{W_{i}\}_{i=1}^{n}$. $\mathbb{E}_{n}$ denotes the expectation with respect to the empirical measure, and $\mathbb{G}_{n}$ denotes the empirical process, that is \[ \mathbb{G}_{n}B(W)=\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\bigg[B(W_{i})-EB(W)\bigg] \] indexed by a measurable class of functions $\mathcal{B}:\mathcal{W}\mapsto\mathbb{R}$. In what follows, we use $\Vert\cdot\Vert_{P,q}$ to denote the $L^{q}(P)$ norm. $\left\Vert{A}\right\Vert_{\infty}=\max_{i,j}|a_{ij}|$, $\left\Vert{A}\right\Vert_{1}=\sum_{i,j}|a_{ij}|$, and $\Vert{A}\Vert_{0}$ equals the number of nonzero components of $A$ for a matrix $A=[a_{ij}]$.
Suppose that the outcome variable $Y$ is determined by \[ Y=m\big(D,X,U\big),\eqno(2.1.1) \] where $D$ is a continuous treatment variable, $X$ is a $K_X$-dimensional vector of covariates, $m(\cdot)$ is a smooth measurable function, and $U$ is an unobservable random vector that captures omitted factors and all types of unobserved heterogeneity. Given $(y_{1},y_{2})\subset \mathcal{S}_{Y}$, the object of interest is
which measures the average partial effect of a marginal change in $D$ on the individuals with $Y\in (y_1, y_2)$.
Let $Q_{Y}(\tau|d,x)$ denote the $\tau$-th quantile of $Y$ conditional on $D=d$, $X=x$. Under the assumption that the random variables $U$ and $D$ are independent conditional on $X$ and some other technical assumptions (See Appendix A), Hoderlein and Mammen (2007, Theorem 2.1, page 1515) have established that for any $(d,x)\in\mathcal{S}_{D}\times\mathcal{S}_{X}$, $0<\tau<1$, \[ E\left(\partial_{D} m(D,X,U)\bigg|D=d,X=x,Y=Q_{Y}(\tau|d,x)\right)=\frac{\partial Q_{Y}(\tau|d,x)}{\partial d}. \eqno(2.1.2) \]
Let $F_{Y}(\cdot|d,x)$ be the CDF of $Y$ conditional on $(D,X)=(d,x)$. (2.1.2) can be equivalently expressed as \[ E\left(\partial_{D} m(D,X,U)\bigg|D=d,X=x,Y=y\right)=\frac{\partial Q_{Y}(u|d,x)}{\partial d}\bigg|_{u=F_{Y}(y|d,x)}. \eqno(2.1.3) \]
Thus $\theta(y_{1},y_{2})$ can be identified straightforwardly by \[ \theta(y_{1},y_{2})=\int_{y_{1}}^{y_{2}}\int\frac{\partial Q_{Y}(u|d,x)}{\partial d}\big|_{u=F_{Y}(y|d,x)}f(d,x,y)dddxdy\bigg/ P(y_{1}<Y<y_{2}). \eqno(2.1.4) \]
Below we provide an alternative expression of $\theta(y_{1},y_{2})$, under slightly weaker assumptions than Hoderlein and Mammen (2007). Note that the error term $U$ in (2.1.1) is invariant with respect to realizations of $D$. More generally, we can assume that $U=U_{D}$, that is, $U$ may change across $d$.
\noindentAssumption 2.1. For each $d\in\mathcal{S}_{D}$, $U_{d}$'s are identically distributed across $d$ conditional on $X$.
Assumption 2.1 is called “rank similarity” (Chernozhukov and Hansen (2005)). It permits that the realizations of error term may vary with the treatment intensity, but they should have the same distribution conditional on covariates. Assumption 2.1 incorporates $U_{d}\equiv U$ as a special case.
\noindentAssumption 2.2. For each $d\in\mathcal{S}_{D}$, $U_{d}$ is independent of $D$ conditional on $X$.
This conditional independence assumption is weaker than full joint independence of $U_{d}$ and $(D,X)$. Other examples of identifying counterfactual or causal parameters based on the conditional exogeneity condition are Firpo et al. (2009), Chernozhukov, Fernandez-Val, and Melly (2013), Rothe (2010, 2012).
\noindentProposition 2.1. Under Assumptions 2.1-2.2 and some regularity conditions (listed in Appendix A),
Compared with (2.1.4), Proposition 2.1 provides three novel insights. First, it shows $\theta(y_{1},y_{2})$ can be expressed as an integral of the average derivative of $F_{Y}(y|d,x)$ over $(y_{1},y_{2})$ with some rescaling. Because the influence function of average derivatives has been well studied in the existing literature (Newey, 1994), it is possible to derive the orthogonal score for $\theta(y_{1},y_{2})$ (see Proposition 2.2). Second, it shows that like the distributional treatment effect, OASD is robust against censoring data. For example, let $Y=\max\{0,Y^{*}\}$. Then $\theta(y_{1},y_{2})$ is identifiable for any $(y_{1},y_{2})\subset (0,+\infty)\cap \mathcal{S}_{Y^{*}}$. Third, let \[ \theta(y)=E\left(\partial_{D} m(D,X,U)\bigg|Y=y\right).\eqno(2.1.5) \] Fix $y_{1}=y$ in Proposition 2.1 and let $y_{2}$ approach $y_{1}=y$,
with $\Delta y=y_{2}-y_{1}$. It is straightforward to see \[ \theta(y)=\frac{-E\left(\partial_{D}F_{Y}(y|D,X)\right)}{f_{Y}(y)}\eqno(2.1.6) \] which is equivalent to the unconditional quantile partial effect (UQPE) proposed in Firpo et al. (2009). For a detailed discussion on the connection between OASD and UQPE, see Section 2.2.
Below we provide another identification result based on an orthogonal score, which is necessary for estimation with high dimensional controls. Let $W=(Y,D,X)$. Let $\eta=\eta(w; y_{1},y_{2})$ collect the (possibly infinite-dimensional) nuisance parameters: \[ \eta\big(W;y_1,y_2\big)=\left(P\big(y_1<Y<y_2\big),\int_{y_1}^{y_2}F_Y\left(y\big|D,X\right)dy,\frac{\partial_Df\big(D,X\big)}{f\big(D,X\big)}\right) \eqno(2.1.7) \] with $f\big(d,x\big)$ the joint density of $(D,X)$. Let \[
\eqno{(2.1.8)}\]
\noindentProposition 2.2. (Identification based on orthogonal score) Under the same assumptions as Proposition 2.1, we have (i) $\theta(y_{1},y_{2})$ satisfies $E\psi\bigg(W,\theta(y_{1},y_{2}),\eta;y_1,y_2\bigg)=0$.
(ii) $\psi\bigg(W,\theta,\eta;y_1,y_2\bigg)$ satisfies the Neyman orthogonality property:
Proposition 2.2 has three implications. First, result (ii) means the functional Gateaux derivative of the moment function $\psi(\cdot)$ with respect to the nonparametric component vanishes when evaluated at the true parameters. This orthogonality property is crucial to establishing good behaviour (such as root $n$-consistency and asymptotically Gaussian distribution) of the subsequent automatic debiased machine learning estimator for OASD. Second, we have shown (in Appendix A) that the orthogonal score $\psi(\cdot)$ has some double robustness property in the sense that one may still identify $\theta(y_{1},y_{2})$ from the moment condition, even if one (but not all) component of $\eta$ is incorrectly specified. Third, we further prove (Theorem 4.2) that the orthogonal score is also semiparametrically efficient.
\noindentRemark 2.1.1. When the treatment is binary, the quantity in parallel with OASD can be written as
which measures the average treatment effect of participating in a program on the individuals with outcome variable $Y\in(y_{1},y_{2})$. In Appendix A, we show $\vartheta(y_{1},y_{2})$ is identifiable under Assumptions 2.1-2.2, with a monotonicity condition. Moreover, we also derive an orthogonal score for $\vartheta(y_{1},y_{2})$.
By definition, OASDs characterize mean impacts on the subpopulations located at different parts of the unconditional distribution of $Y$. Another related and well known approach to estimate counterfactual effects that are heterogeneous across the unconditional outcome distribution $F_{Y}$ is the unconditional quantile regression (UQR) proposed by Firpo et al. (2009). This subsection establishes the relationship between OASD and UQR.
Let $Y$ again be generated by a general nonseparable model $Y=m(D,X,U_D)$, where $D$ is a scalar continuous treatment variable of interest and $X$ consists of controls. The causal parameter UQR aims to identify is called unconditional quantile partial effect (UQPE), which measures the marginal effect of counterfactually shifting the distribution of a coordinate of the explanatory variables on unconditional quantiles of $Y$. The counterfactual distribution of $Y$ after shifting the distribution of $D$ infinitesimally while holding $X$ fixed can be written as
Let $Q_{Y^{\epsilon}}(\cdot)$ be the inverse of $F_{Y^{\epsilon}}(y)$. The $\tau$-th UQPE with respect to $D$ is defined as \[ UQPE(\tau)=\frac{\partial Q_{Y^{\epsilon}}(\tau)}{\partial\epsilon}\bigg|_{\epsilon=0}.\eqno(2.2.1) \]
Under the assumption of conditional exogeneity, as shown by Firpo et al. (2009, Proposition 1, page 959), the UQPE can be interpreted as the causal effect of changing the distribution of $D$ infinitesimally. Without such an assumption, UQPE may still be of interest as a summary statistic of the counterfactual distributional relationship between $Y$ and $D$.
Let $\theta(y)$ be the average structural derivative of $D$ conditional on $Y=y$, namely,
The relationship between $\theta(y_{1},y_{2})$ and $\theta(y)$ is
Let $Q_{Y}\big(\tau\big)$ denote the $\tau$-th quantile of $Y$. Like Firpo et al. (2009, page 959), we define the CQPE, which is the effect of a small change of $D$ on the conditional quantile of $Y$: \[ CQPE_{\tau}(d,x)=\frac{\partial{Q}_{Y}(\tau|d,x)}{\partial{d}}. \] Let $\zeta_{\tau}(d,x)=F_{Y}\left(Q_{Y}(\tau)|d,x\right)$ be the matching function in Firpo et al. (2009).
\noindentProposition 2.3. Under the same assumptions as Proposition 2.1, we have (i)
(ii)
(iii)
Results (i)-(ii) indicate some interesting connections between OASD and UQPE, that is, both quantities actually possess the identical economic interpretation, and such equivalence holds under a very general nonseparable data generating process without monotonicity. Based on this equivalence, we obtain two additional insights. First, UQPE can be alternatively interpreted as a mean causal effect on an identifiable subpopulation located at some part of the outcome distribution. Second, many desirable properties of OASD may carry over to the UQPE. For example, until now, the literature has not obtained any efficiency bound of UQPE. By Proposition 2.3, the semiparametric efficiency bound of UQPE may be learnt from that of OASD, which we shall derive in Section 4. Last but not least, result (iii) has been proved by Firpo et al. (2009, Proposition 1, page 959). There, the proof needs two assumptions: (i) joint independence of $(D,X)$ and $U_D$ and (ii) $m(d,x,u)$ is monotonic in $u$. Here we show the result still holds under much weaker assumptions, that is, we only need conditional exogeneity and without monotonicity.
The preceding subsection demonstrates there is a close connection between two different classes of quantities: one is the outcome conditioned average partial effect of a continuous treatment (OASD) and the other class incorporates parameters measuring the effect of a counterfactual change in the distribution of a single covariate on unconditional quantiles of the outcome (UQPE). This subsection provides an additional example in support of this insight, that is, we show another counterfactual policy effect parameter, called the marginal partial policy effects (Rothe 2012) can also be represented as some outcome conditioned average partial effect.
Rothe (2012) proposes a class of quantities to evaluate the effect of a counterfactual change in the unconditional distribution of a single covariate on the unconditional distribution of an outcome variable of interest, holding everything else, in particular the dependence structures of the covariates constant. Using the notations in this paper, the parameters in Rothe (2012) can be described as follows. An outcome variable $Y$ is related to a continuously distributed covariate $D$, and $K$ dimensional vector of covariates $X=(X_{1},\cdots,X_{K})$ through a general nonseparable structural model
Let $Q_{0}(\cdot), Q_{1}(\cdot),\cdots, Q_{K}(\cdot)$ be the unconditional quantile function of $D$, $X_{1},\cdots, X_{K}$ respectively. Then $(D,X)$ can be equivalently expressed in terms of their unconditional quantile functions and a rank vector $R=\left(R_{0},R_{1},\cdots, R_{K}\right)$ of standard uniformly distributed latent variables, that is, \[ Y=m\left(Q_{0}\big(R_{0}\big),Q_{1}\big(R_{1}\big),\cdots, Q_{K}\big(R_{K}\big),U_D\right),\eqno(2.3.1) \] with $R_{k}\sim^{d}\mbox{Uniform} (0,1)$, $k=0,\cdots, K$. The joint distribution of $R$, also the copula function of $(D,X)$, measures the dependence structure between $D$ and $X$. Because $D$ is continuously distributed, the latent rank variable $R_{0}$ constitutes a one-to-one transformation of $D$. Define the outcome $Y_{H}$ of the counterfactual experiment in which the unconditional distribution of $D$ has been changed to some CDF $H(\cdot)$, but everything else has been held constant, \[ Y_{H}=m\left(H^{-1}\big(R_{0}\big),Q_{1}\big(R_{1}\big),\cdots, Q_{K}\big(R_{K}\big),U_D\right).\eqno(2.3.2) \]
Let $Q_{A}(\cdot)$ denote the quantile function of a generic random variable $A$. It is natural to define the $\tau$-th partial quantile policy effect of changing the marginal distribution of $D$ from $F_{0}(\cdot)=Q_{0}^{-1}(\cdot)$ to $H(\cdot)$ as
In practice, most policies are contracted, expanded or adjusted gradually, thus one may naturally consider the effect of an infinitesimal change of the distribution of $D$ at a certain given direction. Let $G_{0}$ be any fixed CDF, representing the direction of policy change. Let $H=H_{t}$ be an element of a continuum of CDFs indexed by $t\in[0,1]$ such that \[ H_{t}=F_{0}+t\left(G_{0}-F_{0}\right).\eqno(2.3.3) \] Then the $\tau$-th marginal quantile partial policy effect (MQPE) at the direction of $G_{0}$ is given by\footnote{In applications, the policy effect of changing the unconditional distribution of $D$ from $F_{0}$ to any given fixed CDF $H$ can be well approximated by the effect of an appropriately designed infinitesimal change. To see this, let $t$ be a sufficiently small real number, say, $t=0.01$. To learn about the effect of changing $F_{0}$ to a fixed $H$, one may solve $G_{0}$ from $H=F_{0}+t\left(G_{0}-F_{0}\right)$, that is, $G_{0}=\frac{H-(1-t)F_{0}}{t}$. Then $Q_{Y_{H}}(\tau)-Q_{Y}(\tau)\simeq MQPE(\tau,G_{0}) \cdot t$.} \[ MQPE(\tau,G_{0})=\frac{\partial Q_{Y_{H_{t}}}(\tau)}{\partial t}\bigg|_{t=0}.\eqno(2.3.4) \]
The MQPE described above is very similar to the UQPE defined by (2.2.1). Given the equivalence between UQPE and OASD, an interesting question is whether there exists some outcome conditioned average partial effect quantity, which is equivalent to MQPE? The next proposition answers this question.
\noindentProposition 2.4. Define
with $H_{t}$ given by (2.3.3). $\varsigma(y,G_{0})$ is called outcome conditioned average partial policy effect, which measures the average effect of changing the unconditional distribution of $D$ infinitesimally towards the direction of $G_{0}$ on the individuals with $Y=y$. Similarly, we can define
Under the same assumptions in Proposition 2.1, (i)
(ii)
Like OASD and UQPE, the above proposition indicates MQPE can be interpreted as a mean effect on an identifiable subpopulation located at some part of the outcome distribution. As Rothe (2012) has shown MQPE is identified under a conditional exogeneity condition, a corollary of Proposition 2.4 is that both $\varsigma(y,G_{0})$ and $\varsigma(y_{1},y_{2},G_{0})$ are also identifiable under the same conditions. Further investigation of these parameters is beyond the scope of this paper.
The preceding section shows that OASD is identified under a conditional exogeneity assumption, by controlling for a rich information about covariates $X$, thus it is ideal to consider an estimation procedure using high-dimensional controls in data. We propose an automatic debiaed/double machine learning (ADML) procedure for estimating OASDs with high dimensional covariates. The procedure is easily implemented and semiparametrically efficient. The estimation method consists of three steps:
(i) Estimate the CDF $F_{Y}\left(y|d,x\right)$, its integral and derivatives using high-dimensional nonparametric methods with model selection.
(ii) Using the orthogonal score to estimate $\frac{\partial_Df(D,X)}{f(D,X)}$ automatically.
(iii) Estimate $\theta(y_{1},y_{2})$ based on the orthogonal score via the plug-in rule.
We now describe the estimation procedure in detail.
Step 1. (Estimate CDF) Let $b(d,x)=\{b_{k}(d,x)\}_{k=1}^{p}$ be the basis functions used to approximate $F_{Y}(y|d,x)$, and $\Lambda(\cdot)$ be logistic link function. Then $F_{Y}(y|d,x)$ can be estimated by \[ \widehat{F}_{Y}(y|d,x)=\Lambda\left(b^{\prime}(d,x)\widehat{\beta}(y)\right). \] To obtain $\widehat{\beta}(y)$, we first estimate $\widetilde{\beta}(y)$ by the Lasso penalized distribution regression \[ \widetilde{\beta}(y)=\arg\min_{\beta}\frac{1}{n}\sum_{i=1}^{n}\bigg[1\bigg\{Y_{i}\leq{y}\bigg\}\ln\Lambda\bigg(b^{\prime}(D_{i},X_{i})\beta\bigg)+1\bigg\{Y_{i}>{y}\bigg\}\ln\left(1-\Lambda\bigg(b^{\prime}(D_{i},X_{i})\beta\bigg)\right)\bigg]+\frac{\lambda}{n}\left\Vert\widehat{\Psi}_{y}\beta\right\Vert_{1}, \] where $\lambda$ denotes the penalty level to guarantee good theoretical properties of the lasso estimator, and $\widehat{\Psi}_y=\text{diag}(\widehat{\psi}_{y1}^q,\cdots, \widehat{\psi}_{yp}^q)$ denotes the diagonal matrix of penalty loadings. According to Belloni et al. (2017), we set the penalty level $\lambda$ as \[ \lambda=1.1\sqrt{n}\Phi^{-1}\left(1-\frac{0.1/\ln(n)}{2p n}\right). \] The penalty loadings can be obtained by the following algorithm, proposed by Belloni et al. (2017, Algorithm 6.1, page 261)
Given $y\in\mathcal{S}_{Y}$, define $\widehat{I}(y)=\text{supp}\left(\widehat{\beta}(y)\right)$. The post-lasso estimator $\widehat{\beta}(y)$ is a solution to \[
\]
(Estimate the integral of CDF) Let $\Delta{y}=\left(y_{2}-y_{1}\right)/J$ for some positive integer $J$. Notice that by definition of integration, \[ IF(y_{1},y_{2},d,x)\doteq\int_{y_{1}}^{y_{2}}F_{Y}(y|d,x)dy=\lim_{J\to\infty}\sum_{j=1}^{J}F_{Y}(y_{1}+j\Delta{y}|d,x)\Delta{y}. \] Thus, the estimator of the integral of CDF can be constructed by \[ \widehat{IF}(y_{1},y_{2},d,x)=\sum_{j=1}^{J}\widehat{F}_{Y}(y_{1}+j\Delta{y}|d,x)\Delta{y}=\sum_{j=1}^{J}\Lambda\left(b^{\prime}(d,x)\widehat{\beta}\left(y_{1}+j\Delta{y}\right)\right)\Delta{y}. \eqno{(3.1)} \]
(Estimate the derivative of the integral of CDF) Let
We estimate $DIF(y_{1},y_{2},d,x)$ by a high order partial difference approach, by borrowing the idea from Belloni et al. (2019) in dealing with the estimation of conditional density function. Let $\ell$ be some positive integer. A partial difference estimator of $DIF(y_{1},y_{2},d,x)$ with a bias of general order $O\left(h_{n}^{2\ell}\right)$ is given by \[ \widehat{DIF}(y_{1},y_{2},d,x)=\frac{1}{2h_{n}}\sum_{l=1}^{\ell}\eta_{l}\left(\widehat{IF}(y_{1},y_{2},d+lh_{n},x)-\widehat{IF}(y_{1},y_{2},d-lh_{n},x)\right), \eqno{(3.2)} \] with $h_{n}$ the bandwidth satisfying $h_{n}\to{0}$. The constants $\eta_{l}$ are determined by\footnote{For example, we have $\eta_{1}=1$ for $\ell=1$; $\eta_{1}=4/3$ and $\eta_{2}=-1/6$ for $\ell=2$; $\eta_{1}=3/2$, $\eta_{2}=-3/10$ and $\eta_{3}=1/30$ for $\ell=3$.} \[ \sum_{l=1}^{\ell}l\cdot\eta_{l}=1 \eqno{(3.3)} \] and for $v=3,5,\dots,2\ell-1$, \[ \sum_{l=1}^{\ell}l^{v}\cdot\eta_{l}=0. \eqno{(3.4)} \] As will be clear in the next section, the bandwidth should satisfy the following two conditions \[ h_{n}^{4}n\to\infty \quad\text{and}\quad h_{n}^{2\ell}n^{1/4}\to{0}. \] Thus, we suggest to select the bandwidth $h_{n}=n^{-1/\left(4\ell+2\right)}$ by maximizing the convergence rate of $\widehat{DIF}$. We also compare the finite sample performance of DIF in Eq.(3.2) with Sasaki et al. (2022)'s estimator. Simulation results are reported in Appendix B, which show that our estimator performs better in terms of bias ratio.
\noindentRemark 3.1. By Eq.(3.2), we employ an estimator in the spirit of the high order bias reduction kernel smoothing to estimate the CDF derivative instead of directly differentiating the lasso CDF estimator. By using a partial difference estimator with a bias of sufficiently high order, we do not need to impose substantially more restrictive approximate sparsity conditions than before while preserving the good rates of convergence for both CDF and its derivatives. To understand how estimator (3.2) works, we first take Taylor expansion of $\widehat{IF}$ with respect to $d\pm{lh_{n}}$ at $d$. By the construction of (3.3), all the coefficients of first order derivatives sum up to one, while the coefficients of remaining orders sum up to zero by construction of Equation (3.4). Thus, only the first order derivative of CDF and the terms of order $O\left(h_{n}^{2\ell}\right)$ are left.
Step 2. (Automatic estimation of $\frac{\partial_Df(D,X)}{f(D,X)}$.) According to Chernozhukov et al. (2022a,b) and Singh and Sun (2021), we estimate $\frac{\partial_Df(D,X)}{f(D,X)}$ automatically based on the double robustness property of orthogonal score function $\psi$ (defined below Proposition 2.2). For any real function $\delta(d,x)$, notice that \[ \frac{\partial}{\partial\tau}E\left[\psi\bigg(W,\theta,\eta+\tau\delta;y_1,y_2\bigg)\right]\bigg|_{\tau=0}=0. \eqno{(3.5)}\] Let $\delta(d,x)=\left(0,\widetilde{\delta}(d,x),0\right)^{\prime}$. $(3.5)$ is equal to \[ E\left[\partial_{D}\widetilde{\delta}(D,X)+\frac{\partial_Df(D,X)}{f(D,X)}\widetilde{\delta}(D,X)\right]=0. \eqno{(3.6)} \] The automatic estimator for $\frac{\partial_Df(D,X)}{f(D,X)}$ is constructed on the basis of $(3.6)$. Suppose $L(D,X)=\frac{\partial_Df(D,X)}{f(D,X)}$ is replaced by a linear combination $b^{\prime}(D,X)\gamma$ and let $\widetilde{\delta}(D,X)$ be one element $b_{k}(D,X)$ for $k=1,\dots,p$. Then $L(D,X)=\frac{\partial_Df(D,X)}{f(D,X)}$ can be estimated by \[ \widehat{L}(d,x)=b^{\prime}(d,x)\widehat{\gamma}. \] $\widehat{\gamma}$ is the Lasso estimator which is constructed by\footnote{We can also use post-lasso estimator to replace $\gamma$.} \[ \widehat{\gamma}=\arg\min_{\gamma}-2\widehat{M}^{\prime}\gamma+\gamma^{\prime}\widehat{G}\gamma+2\widetilde{\lambda}\left\Vert\gamma\right\Vert_{1}, \eqno{(3.7)} \] where $\widetilde{\lambda}>0$ is a positive scalar to control for the degree of penalty, and \[
\] We estimate (3.7) by the iterative tuning procedure for data-driven regularization parameter $\tilde{\lambda}$ , proposed by Chernozhukov et al. (2022b, Appendix A, page 1000). \\
Step 3. (Estimate $\theta(y_{1},y_{2})$ by plug-in) $P(y_{1},y_{2})=P(y_{1}<Y<y_{2})$ can be directly estimated by \[ \widehat{P}(y_{1},y_{2})=\frac{1}{n}\sum_{i=1}^{n}1\bigg\{y_{1}<Y_{i}<y_{2}\bigg\}. \eqno{(3.8)} \] Based on the estimator of $ \eta=\left(P\big(y_1<Y<y_2\big),\int_{y_1}^{y_2}F_Y\left(y\big|D,X\right)dy,\frac{\partial_Df\big(D,X\big)}{f\big(D,X\big)}\right)$, $\theta(y_{1},y_{2})$ is straightforwardly estimated via a plug-in rule, such that \[ \frac{1}{n}\sum_{i=1}^{n}\psi\bigg(W_{i},\widehat{\theta},\widehat{\eta};y_{1},y_{2}\bigg)=0. \] Or \[
\eqno{(3.9)}\] where \[ \int_{y_{1}}^{y_{2}}1\bigg\{Y_{i}<y\bigg\}dy=1\bigg\{Y_{i}\leq y_{1}\bigg\}(y_{2}-y_{1})+1\bigg\{y_{1}<Y_{i}<y_{2}\bigg\}(y_{2}-Y_{i}). \]
\noindentRemark 3.2. As is standard in the literature (Chernozhukov et al. (2018)), we can use various data splitting methods to further relax the entropy condition required in the subsequent asymptotic analysis. There is no asymptotic efficiency loss from sample splitting under cross fitting. See Chernozhukov et al. (2018) for more details. For simplicity, we do not describe the estimation procedure using the data splitting method in this paper. One reason is that the entropy of the function classes can be easily verified (which can be found in the next section), e.g., differentiability of density function. Second, sample splitting facilitates the statistical inference, at the cost of making estimation more involved. For example, researchers need to choose the number of fold $K$, and repeat running Steps 1-3 $K$ times. \\
\noindentAlgorithm 1. (Auto Double Machine Learning Estimator )
\noindentStep 1. Pick a finite set $\mathcal{Y}\subset\mathcal{S}_{Y}$ of grid points of outcome values.
\noindentStep 2. For any $y\in\mathcal{Y}$, compute $\widetilde{\beta}(y)$ from $\ell_{1}$-penalized logistic regression of $1\{Y<y\}$ on $b(D,X)$.
\noindentStep 3. For any $y\in\mathcal{Y}$, compute $\widehat{\beta}(y)$ from logistic regression of $1\{Y<y\}$ on $\left\{b_k(D,X):\widetilde{\beta}_{k}\neq{0}\right\}$.
\noindentStep 4. Estimate the integral of CDFs and its derivative via (3.1) and (3.2).
\noindentStep 5. Compute $\widehat{\gamma}$ from $\ell_{1}$-penalized GMM via (3.7).
\noindentStep 6. Estimate the unconditional probability $P(y_{1},y_{2})$ via (3.8).
\noindentStep 7. For a pair of $(y_{1},y_{2})\in\mathcal{Y}\times\mathcal{Y}$ and $y_{1}<y_{2}$, compute $\widehat{\theta}(y_{1},y_{2})$ via (3.9).
In this section, we establish the asymptotic properties for the ADML estimator of $\theta(y_{1},y_{2})$. Overall, the ADML estimator, which can be viewed as a stochastic process of $(y_{1},y_{2})$, is proved to be uniformly Gaussian based on some high level conditions. Following, we provide a set of sufficient conditions for these high level conditions, which are convenient to hold in practice. We show the ADML estimator is semiparametrically efficient, that is, it achieves the semiparametric efficiency bound. Finally, we derive the uniform validity of the multiplier bootstrap used to construct uniform confidence bands.
Consider fixed sequence of numbers $\Delta_{n}\to{0^{+}}$ at a speed at most polynomial in $n$ (e.g., $\Delta_{n}\geq{1/n^{\bar{c}}}$ for some $\bar{c}>0$), and positive constants $c$, $C$, $H$ and $T$.
We introduce a set of assumptions used to prove the uniform Gaussianity of the ADML estimator. Some of these assumptions are high level. Sufficient conditions for these high level conditions are provided in the next subsection.
\noindentAssumption 4.1. The random element $W$ takes values in a compact measure space $(\mathcal{W},\mathcal{A}_{\mathcal{W}})$ and its law is determined by a probability measure $P$. The observed data $\{W_{i}\}_{i=1}^{n}$ consist of $n$ i.i.d. copies of a random element $W=(Y,D,X)\in\mathbb{R}^{2+d_{X}}$.
\noindentAssumption 4.2. Let $u=(y_{1},y_{2})\in\mathcal{U}\subset\mathbb{R}^{2}$ be the index of target parameter $\theta$. $\mathcal{U}$ is a totally bounded metric space equipped with a semi-metric $d_{\mathcal{U}}$.\footnote{Our OASDs are defined on $(y_1,y_2)\in \mathcal{S}_Y \times \mathcal{S}_Y$ with $y_1<y_2$. Let $c_{0}>0$, $c_{1}$ and $c_{2}$ be three constants. The metric space $\mathcal{U}$ can be defined as $\mathcal{U}=\left\{(y_{1},y_{2}):c_{1}\leq{y_{1}+c_{0}}\leq{y_{2}}\leq{c_{2}}\right\}$, which is a bounded upper triangular.} Denote $Y(u)$ as a measurable transform $t(Y,u)$ of $Y$ and $u$.\footnote{Specifically, $Y(u)\in\left\{\displaystyle\int_{y_{1}}^{y_{2}}1\{Y\leq{y}\}dy,1\{y_{1}\leq{Y}\leq{y_{2}}\}\right\}$ in this paper.} The map $u\mapsto{Y(u)}$ obeys the following uniform continuity property: \[ \lim\limits_{\epsilon\to{0^{+}}} \sup_{d_{\mathcal{U}}(u,\bar{u})\leq\epsilon}\left\Vert Y(u)-Y(\bar{u})\right\Vert_{P,2}=0, \quad E\sup_{u\in\mathcal{U}}|Y(u)|^{2+c}<\infty, \] where the supremum in the first expression is taken over $u,\bar{u}\in\mathcal{U}$.
Assumption 4.2 defines a valid metric space for OASDs, and restricts the continuity and boundedness of $Y$. According to Assumption 4.2, there exists a positive constant $H$, which ensures that $\mathcal{U}\subset[-H,H]\times[-H,H]$. Denote the space $\mathcal{H}=\{y: |y|\leq{H}\}$.
\noindentAssumption 4.3. Assume the functions $F_{Y}(y|d,x)$ and $L(d,x)$ can be approximated by \[ F_{Y}(y|d,x)=\Lambda\bigg(b(d,x)^{\prime}\beta(y)\bigg)+r_{F}(y,d,x) \] and \[ L(d,x)=b(d,x)^{\prime}\gamma+r_{L}(d,x), \] where $r_{F}(y,d,x)$ and $r_{L}(d,x)$ are the approximation errors. Then uniformly over $y\in\mathcal{H}$ and $u\in\mathcal{U}$,
(i) The sparsity condition $\Vert\beta(y)\Vert_{0}\leq{s_{\beta}}$ and $\Vert\gamma\Vert_{0}\leq{s_{\gamma}}$ holds, and the approximation errors satisfy $\Vert r_{F}\Vert_{P,2}=o_{p}\left(h_{n}n^{-1/4}\right)$, $\Vert r_{F}\Vert_{P,\infty}=o_{p}\left(h_{n}\right)$, and $\Vert r_{L}\Vert_{P,2}=o_{p}\left(n^{-1/4}\right)$, $\Vert r_{L}\Vert_{P,\infty}=o_{p}(1)$. The sparsity indices $s_{\beta}$, $s_{\gamma}$ and the number of terms $p$ in the vector $b(d,x)$ obeying $s_{\beta}^{2}\log^{2}(p\vee n)=o\left(h_{n}^{4}n\right)$, and $s_{\gamma}^{2}\log^{2}(p\vee n)=o(n)$. The bandwidth $h_{n}$ satisfies $h_{n}^{2\ell}=o\left(n^{-1/4}\right)$ for some positive integer $\ell$.
(ii) There are estimators $\widehat{\beta}(y)$ and $\widehat{\gamma}$ such that, with probability no less than $1-\Delta_{n}$, the estimation errors satisfy $\left\Vert b(D,X)^{\prime}\left(\widehat{\beta}(y)-\beta(y)\right)\right\Vert_{\mathbb{P}_{n},2}=o_{p}\left(h_{n}n^{-1/4}\right)$, $K_{n}\left\Vert \widehat{\beta}(y)-\beta(y)\right\Vert_{1}=o_{p}\left(h_{n}\right)$, and $\left\Vert b(D,X)^{\prime}\left(\widehat{\gamma}-\gamma\right)\right\Vert_{\mathbb{P}_{n},2}=o_{p}\left(n^{-1/4}\right)$, $K_{n}\left\Vert \widehat{\gamma}-\gamma\right\Vert_{1}=o_{p}(1)$; the estimators are sparse such that $\left\Vert\widehat{\beta}(y)\right\Vert_{0}\leq{s_{\beta}}$ and $\Vert\widehat{\gamma}\Vert_{0}\leq{s_{\gamma}}$.
(iii) The empirical and population norms induced by the Gram matrix formed by $\{b_{j}(d,x)\}_{j=1}^{p}$ are equivalent on sparse subsets, such as \[ \sup_{\Vert\upsilon\Vert_{0}\leq{s\log{n}}}\left|\frac{\Vert b(D,X)^{\prime}\upsilon\Vert_{\mathbb{P}_{n},2}}{\Vert b(D,X)^{\prime}\upsilon\Vert_{P,2}}-1\right|\leq{\epsilon_{n}}. \] The boundedness conditions hold: $\left\Vert\left\Vert\partial^{l}_{D} b(D,X)\right\Vert_{\infty}\right\Vert_{P,\infty}\leq{K_{n}}$ for $l=0,1$, $\Vert Y(u)\Vert_{P,\infty}\leq{C}$ and $\left\Vert \partial^{2}_{D}F(y|D,X)\right\Vert_{P,\infty}\leq{C}$, $\left\Vert\partial_{D}F(y|D,X)\right\Vert_{P,2}\leq{C}$.
(iv) $\partial_{d}F(y|d,x)$ is bounded and $\sigma$-th continuously differentiable with respect to $(d,x)$, and satisfies $2\sigma>\max\{1+d_{X_{c}},4(\ell-1)\}$, where $d_{X_{c}}$ denotes the dimension of continuous components in $X$.
\noindentRemark 4.1. Assumption 4.3-(i) includes conditions on the approximate sparsity of the model and bandwidths used to estimate CDFs, their derivatives and $L(D,X)=\frac{\partial_Df(D,X)}{f(D,X)}$. To ensure that the CDF and its derivative have a desirable convergence rate, i.e., $o_{p}\left(n^{-1/4}\right)$, the growth rate of the approximate sparsity and the order of the approximation error are restricted. The sparsity condition of $s_{\beta}$ and the convergence rate condition of $h_{n}$ together require that $s_{\beta}$ should grow slower than $n^{1/2-1/(4\ell)}$. This growth rate can be improved if the order of bias $2\ell$ increases, namely that the CDF derivative is estimated with a higher order bias, accompanied with additional smoothness assumptions (Assumption 4.3(iv)). Sasaki et al. (2022) estimate the CDF derivative by directly differentiating the CDF function. In practice, such estimation procedure might not be satisfactory. The uniform convergence rate of the derivative estimator mainly depends on the level of sparsity, and is usually slower than the CDF estimator. To achieve the faster uniform convergence rate, particularly faster than $n^{-1/4}$, the divergence rate of sparsity should be restricted to grow slower than $n^{1/2-c}$ for some constant $c$, which cannot be improved anymore. As previously discussed, the divergence rate of sparsity $s_{\beta}$ in this paper can be improved to $n^{1/2}$ as close as possible by choosing a partially difference estimator with higher order bias.
\noindentRemark 4.2. Assumption 4.3(ii) imposes some high-level conditions on the estimators of nuisance functions. In the next subsection, we provide a set of regular and sufficient conditions for both (Post-)Lasso and automatic estimators to satisfy the uniform bounds. Assumption 4.3(iii) first presents the equivalence between empirical and population norms. Sufficient conditions and primitive examples of functions admitting sparse approximations are given in Belloni et al. (2014). The boundedness conditions in Assumption 4.3(iii) are made to simplify arguments, and they could be removed at the cost of more regular conditions and complicated proofs. Assumption 4.3(iv) is majorly used to establish the Gaussian process of the ADML estimator, which restricts the set of functions $\left\{\partial_{D}\displaystyle\int_{y_{1}}^{y_{2}}F_{Y}(y|D,X)dy:(y_{1},y_{2})\in\mathcal{U}\right\}$ to be a Donsker class. Furthermore, another composition of the smoothness restriction, such that $\sigma>2(\ell-1)$, is assumed to achieve a partially difference estimator with higher order bias in Section 3.
The next theorem establishes the uniform Gaussianity of the empirical reduced-form process $\widehat{Z}_{n}(u)=\sqrt{n}\left(\widehat{\theta}(u)-\theta(u)\right)$ defined by Equation (3.9).
\noindentTheorem 4.1. Under Assumptions 2.1-2.2 and 4.1-4.3, the reduced-form empirical process admits a linearization; namely \[ \widehat{Z}_{n}(u)=\sqrt{n}\left(\widehat{\theta}(u)-\theta(u)\right)=Z_{n}(u)+o_{p}(1) \quad in \quad \mathbb{D}=\ell^{\infty}\left(\mathcal{U}\right) \] where $Z_{n}(u)=\mathbb{G}_{n}\psi(W,\theta,\eta;u)$. The process $\widehat{Z}_{n}(u)$ is asymptotically Gaussian, namely \[ \widehat{Z}_{n}(u)\leadsto Z(u) \quad in \quad \mathbb{D}=\ell^{\infty}\left(\mathcal{U}\right) \] where $Z(u)=\mathbb{G}\psi(W,\theta,\eta;u)$ with $\mathbb{G}$ denoting Brownian bridge and with $Z(u)$ having bounded, uniformly continuous paths: \[ E\sup_{u\in\mathcal{U}}\Vert{Z}(u)\Vert<\infty, \quad \lim_{\epsilon\to{0}^{+}}E\sup_{d_{\mathcal{U}}\left(u,\widetilde{u}\right)\leq{\epsilon}}\Vert{Z}(u)-{Z}\left(\widetilde{u}\right)\Vert=0. \]
This subsection describes the asymptotic properties for the estimators of the nuisance functions (like the CDFs, their derivatives and $L(D,X)=\frac{\partial_Df(D,X)}{f(D,X)}$). These asymptotic results can be viewed as the sufficient conditions for Assumption 4.3(ii) to hold in practice. We first discuss the results for Lasso and Post-Lasso estimators with function valued outcomes and logistic link. Belloni et al. (2017) establish the general results for both linear and logistic link (e.g., Theorem 6.1 and 6.2). They explore that uniform consistency and convergence rate are mainly depends on the rate of sparsity. We invoke the same assumptions of Assumption 6.2 in Belloni et al. (2017).
\noindentAssumption 4.4. The conditions of Assumption 6.2 in Belloni et al. (2017) hold.
\noindentLemma 4.1. If Assumption 4.4 together with Assumption 4.3(i) hold, $K_{n}^{2}s_{\beta}^{2}\log(p\vee{n})=o\left(h_{n}^{2}n\right)$, then \[ \sup_{y\in\mathcal{H}}\left\Vert b(D,X)^{\prime}\left(\widehat{\beta}(y)-\beta(y)\right)\right\Vert_{\mathbb{P}_{n},2}=o_{p}\left(h_{n}n^{-1/4}\right) \] and \[ K_{n}\sup_{y\in\mathcal{H}}\left\Vert \widehat{\beta}(y)-\beta(y)\right\Vert_{1}=o_{p}\left(h_{n}\right). \]
\noindentRemark 4.3. Here we restrict the divergence rate of basis functions $K_{n}$, sparsity $s_{\beta}$ and $p\vee{n}$ to satisfy $K_{n}^{2}s_{\beta}^{2}\log(p\vee{n})=o\left(h_{n}^{2}n\right)$, which is quiet strong than one required in Belloni et al. (2017) for standard (Post-)Lasso estimation, but is consistent with one in Belloni et al. (2019, Condition D(P)). The major reason is that in this paper, as well as Belloni et al. (2019), we need to estimate the derivative of (Post-)Lasso estimator. Thus, a faster convergence rate of the original estimator is required to ensure the estimator of derivative function achieving the convergence rate $o_{p}\left(n^{-1/4}\right)$. \\
The following lemma establishes the bound rates for the derivative of (Post-)Lasso estimator. \noindentLemma 4.2. If Assumption 4.4 together with Assumption 4.3(i) and (iii) hold, $K_{n}^{2}s_{\beta}^{2}\log(p\vee{n})=o\left(h_{n}^{2}n\right)$, then \[ \sup_{u\in\mathcal{U}}\left\Vert\widehat{DIF}(u,D,X)-{DIF}(u,D,X)\right\Vert_{\mathbb{P}_{n},2}=o_{p}\left(n^{-1/4}\right) \] and \[ \sup_{u\in\mathcal{U}}\left\Vert\widehat{DIF}(u,D,X)-{DIF}(u,D,X)\right\Vert_{\mathbb{P}_{n},\infty}=o_{p}\left(1\right). \]
\\
Now we establish the bounds for automatic estimator. Recall that \[ \widehat{M}=-\frac{1}{n}\sum_{i=1}^{n}\partial_{D_{i}}b(D_{i},X_{i}) \quad\text{and}\quad \widehat{G}=\frac{1}{n}\sum_{i=1}^{n}b(D_{i},X_{i})b^{\prime}(D_{i},X_{i}). \] Also define \[ M=-E\partial_{D}b(D,X) \quad\text{and}\quad G=Eb(D,X)b^{\prime}(D,X). \] The following boundedness condition is assumed to achieve convergence rates for $\widehat{M}$ and $\widehat{G}$.
\noindentAssumption 4.5. There exists $K^{*}>0$ such that, with probability 1, $\left\Vert\partial^{l}_{D} b(D,X)\right\Vert_{\infty}\leq{K^{*}}$ for $l=0,1$.
Assumption 4.5 is quiet different with the boundedness condition provided in Assumption 4.3(iii). The boundary of basis functions with probability $K^*$ is constant, while the boundary with supremum norm $K_{n}$ can diverge to infinity as $n$ increases. Next lemma investigates the convergence rates for $\widehat{M}$ and $\widehat{G}$.
\noindentLemma 4.3. If Assumption 4.5 holds, then \[ \left\Vert{\widehat{M}-M}\right\Vert_{\infty}=O_{p}\left(\sqrt{\frac{\ln{p}}{n}}\right), \quad \left\Vert{\widehat{G}-G}\right\Vert_{\infty}=O_{p}\left(\sqrt{\frac{\ln{p}}{n}}\right). \]
A set of additional assumptions, for example, the quality of approximation and the level of penalty, are assumed to derive the bounds for automatic estimator.
\noindentAssumption 4.6. (i) There exist $c^{*}>1$, $\xi>0$ such that for all $\bar{s}\leq{c}^{*}\left(\frac{\ln{p}}{n}\right)^{-\frac{1}{1+2\xi}}$. There exists some $\widetilde{\gamma}\in\mathbb{R}^{p}$ with $\left\Vert\widetilde{\gamma}\right\Vert_{1}\leq{c^{*}}$ and $\bar{s}$ nonzero elements such that $\left\Vert L-b^{\prime}\widetilde{\gamma}\right\Vert_{P,2}^{2}\leq{c^{*}}\bar{s}^{-\xi}$.
(ii) $G$ is nonsingular with largest eigenvalue uniformly bounded in $n$.
(iii) Denote $\mathcal{S}_{\bar{\bar{\gamma}}}$ as the support of $\bar{\bar{\gamma}}$. There exist $k>3$ such that for $\bar{\bar{\gamma}}\in\left\{\widetilde{\gamma},\bar{\gamma}\right\}$, \[ RE(k)=\inf_{\upsilon\neq{0},\sum_{j\in\mathcal{S}_{\bar{\bar{\gamma}}}^{c}}|\upsilon_{j}|\leq{k}\sum_{j\in\mathcal{S}_{\bar{\bar{\gamma}}}}|\upsilon_{j}|}\frac{\upsilon^{\prime}G\upsilon}{\sum_{j\in\mathcal{S}_{\bar{\bar{\gamma}}}}\upsilon_{j}^{2}}>0, \] where \[ \bar{\gamma}=\arg\min_{\dot{\gamma}}\left\Vert L-b^{\prime}\dot{\gamma}\right\Vert_{P,2}^{2}+2\widetilde{\lambda}\left\Vert\dot{\gamma}\right\Vert_{1}. \]
(iv) $\ln{p}=O(\ln{n})$ and $\widetilde{\lambda}=\kappa_{n}\sqrt{\frac{\ln{p}}{n}}$ for some $\kappa_{n}\to\infty$.
\noindentLemma 4.4. If Assumptions 4.5 and 4.6 hold, then \[ \left\Vert{b}(D,X)^{\prime}\widehat{\gamma}-L(D,X)\right\Vert_{P,2}=O_{p}\left(\kappa_{n}\left(\frac{\ln{p}}{n}\right)^{\frac{\xi}{1+2\xi}}\right). \]
Lemma 4.4 shows that the convergence rate of automatic estimator is faster than $n^{-1/4}$ if $\xi>1/2$, which requires the sparsity and the approximation error to grow with some restricted rate. Based on the sparsity and approximation error conditions in Assumption 4.3(i), we show that the bounds for the automatic estimator required in Assumption 4.3(ii) are satisfied.
\noindentLemma 4.5. Let $\phi_{\min}(k)$ be the minimum sparse eigenvalue. If Assumptions 4.5-4.6 together with 4.3(i) and (iii) hold, $\phi_{\min}(k)$ is bounded from zero and $K_{n}\sqrt{s_{\gamma}}=o(n^{1/4})$, then $\widehat{\gamma}$ is sparse, $\left\Vert\widehat{\gamma}\right\Vert_{0}\leq{s_{\gamma}}$, and the following performance bounds hold: \[ \left\Vert b(D,X)^{\prime}\left(\widehat{\gamma}-\gamma\right)\right\Vert_{\mathbb{P}_{n},2}=o_{p}\left(n^{-1/4}\right) \] and \[ K_{n}\left\Vert \widehat{\gamma}-\gamma\right\Vert_{1}=o_{p}(1). \]
Condition $K_{n}\sqrt{s_{\gamma}}=o(n^{1/4})$ is also required in a similar way with Belloni et al. (2017) to develop uniform Gaussianity, for example, Assumption 6.1(iv) in Belloni et al. (2017).
This next theorem shows that $\psi\left(W,\theta,\eta;y_1,y_2\right)$ in Eq.(2.1.8) is also an efficient score, and thus our ADML estimator achieves the semiparametric efficiency bound.
\noindentTheorem 4.2. Under the assumptions of Proposition 4.2, Assumptions 4.1 and 4.2, for any $(y_1,y_2)\in\mathcal{U}$, the semiparametric efficiency bound of $\theta(y_1,y_2)$ is $E\left[\psi\big(W,\theta_0,\eta;y_1,y_2\big)\right]^2$.
In practice, inference based on directly estimating the asymptotic variance of the limit process can be overly complicated. In such cases, bootstrap methods can be effectively applied to construct the confidence bands. Let $\{\xi\}_{i=1}^{n}$ be a random sample drawn from the distribution with zero-mean and unit-variance. We then define the estimated multiplier process for $Z(u)$ by \[ \widehat{Z}_{n}^{*}(u)=\sqrt{n}\left(\widehat{\theta}^{*}(u)-\widehat{\theta}(u)\right)=\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\xi_{i}\psi\bigg(W_{i},\widehat{\theta},\widehat{\eta};u\bigg). \] The main result of this section shows that the bootstrap law of the process $\widehat{Z}_{n}^{*}(u)$ provides a valid approximation to the large sample law of $\widehat{Z}_{n}(u)$. We develop such validity by imposing the following regular assumption.
\noindentAssumption 4.7. A random element $\xi$ with values in a measure space $(\Omega,\mathcal{A}_{\Omega})$ that is independent of $(\mathcal{W},\mathcal{A}_{\mathcal{W}})$, and law determined by a probability measure $P_{\xi}$, with zero-mean and unit-variance. The observed data $\{\xi_{i}\}_{i=1}^{n}$ consist of $n$ i.i.d. copies of a random element $\xi$.
We introduce some useful notations to describe the following results. We define the conditional weak convergence of the bootstrap law in probability, denoted by $\widetilde{Z}_{n}(u)\leadsto_{B}\widetilde{Z}(u)$ in $\ell^{\infty}\left(\mathcal{U}\right)$, by \[ \sup_{T\in{BL_{1}\left(\ell^{\infty}\left(\mathcal{U}\right)\right)}}\left|E_{\xi|P}T\left(\widetilde{Z}_{n}(u)\right)-ET\left(\widetilde{Z}(u)\right)\right|=o_{p}(1), \] where $BL_{1}\left(\mathbb{D}\right)$ denotes the space of functions mapping $\mathbb{D}$ to [0,1] with Lipschitz norm at most 1, and $E_{\xi|P}$ denote the expectation over the multiplier weights $\{\xi_{i}\}_{i=1}^{n}$ holding the data $\{W_{i}\}_{i=1}^{n}$ fixed.
\noindentTheorem 4.3. If Assumptions 2.1-2.2, 4.1-4.3 and 4.7 hold, the bootstrap law consistently approximates the large sample law $Z(u)$ of $Z_{n}(u)$, namely \[ \widehat{Z}_{n}^{*}(u)\leadsto_{B} Z(u) \quad in \quad \mathbb{D}=\ell^{\infty}\left(\mathcal{U}\right) \]
One of the most relevant practical applications of this result is to test the null hypothesis of treatment homogeneity across $u$: \[ H_{0}: \theta(u)=\theta(u^{\prime}) \ \ \text{for all}\ \ u, u^{\prime}\in\mathcal{U}. \] To test this hypothesis, we can construct the test statistic $\sup_{u\in\mathcal{U}}\sqrt{n}\left|\widehat{\theta}(u)-\left(\#\mathcal{U}\right)^{-1}\displaystyle\int_{\mathcal{U}}\widehat{\theta}\left(\widetilde{u}\right)d\widetilde{u}\right|$, and use \[ \sup_{u\in\mathcal{U}}\left|\widehat{Z}_{n}^{*}(u)-\left(\#\mathcal{U}\right)^{-1}\displaystyle\int_{\mathcal{U}}\widehat{Z}_{n}^{*}\left(\widetilde{u}\right)d\widetilde{u}\right| \] to simulate its asymptotic distribution, in which $\#\mathcal{U}$ denotes the area of space $\mathcal{U}$. For example, suppose one is interested in whether the treatment is homogeneous across any fixed length of intervals, such that $c_{1}\leq{y_1}+c_{0}=y_{2}\leq{c_{2}}$ for some constants $c_{0}$, $c_{1}$ and $c_{2}$ satisfying $c_{0}>0$ and $c_{2}-c_{1}>0$. In such case, we can define $\mathcal{U}=\left\{(y_{1},y_{2}):c_{1}\leq{y_1}+c_{0}=y_{2}\leq{c_{2}}\right\}$ and then $\#\mathcal{U}=c_{2}-c_{1}$.
In this section, we study the finite sample performance of the naive estimator based on moment condition in Proposition 2.1 and the ADML estimator (based on orthogonal score). Let $y_\tau$ be the $\tau$-th empirical quantile of $Y$. We consider the estimation of $(y_{0.05},y_{0.15})$, $(y_{0.15},y_{0.25})$, $(y_{0.25},y_{0.35})$, $(y_{0.35},y_{0.45})$,$(y_{0.45},y_{0.55})$, $(y_{0.55},y_{0.65})$, $(y_{0.65},y_{0.75})$, $(y_{0.75},y_{0.85})$, $(y_{0.85},y_{0.95})$.
The data generating process is
where $V_1,V_2,V_3,V_4\sim N(0,1)$, $X=(X_1,\cdots,X_{p_x})^\intercal\sim N(0,\Sigma)$ with $\Sigma=(0.5)^{|j-k|}$. Note that our DGP allows for the error term $U$ relying on realizations of $D$ and a dependence between error term $U$ and the treatment variable $D$. $\delta_0$ is a $p_x\times 1$ vector with elements $\delta_{0,j}=(1/j)^2$ for $j\in\{1,2,...,p_x\}$. $c_d$ and $c_y$ are scalars to control the strength of the relationship between the covarites, the outcome, and the treatment variable. We use 16 combinations of $c_d$ and $c_y$, respectively, with $c_d=\sqrt{\frac{(\pi^2/3)R_d^2}{(1-R_d^2)\delta_0'\Sigma\delta_0}}$ and $c_y=\sqrt{\frac{R_y^2}{(1-R_y^2)\delta_0'\Sigma\delta_0}}$. $R_d^2\in\{0.1,0.2,0.3,0.4\}$ and $R_y^2\in\{0.1,0.2,0.3,0.4\}$, in which $R_d^2$ reflects the sparsity level of the effect of $X$ on $D$ while $R_y^2$ reflects the sparsity level of the effect of $X$ on $Y$.
We consider $N=500$, $p_x=30$ throughout the simulation. To form the basis function, we include all first order, second order and interaction terms among $(D,X)$. The basis function can be written as $b(D,X)=\left(D,X_1,\cdots,X_{30},D^2,X_1^2,\cdots,X_{30}^2,DX_1,\cdots,DX_{30},X_1X_2,\cdots,X_{29}X_{30}\right)$ with $p=dim\left(b(D,X)\right)=527>N$.
For each design, we calculate the naive estimator and ADML estimator. We do 500 iterations to compute bias ratio, std, mean square error (MSE) and the probability of 500 estimators lies in nominal 95% confidence interval (Cvg). These results are reported in Tables 5.1-5.4. We find that ADML estimators outperform naive estimators in terms of MSE and Cvg in all cases.
This paper proposes a new class of heterogeneous causal quantities, named outcome conditioned average structural derivatives (OASD) to measure the average partial effect of a marginal change in a continuous treatment on the individuals located at different parts of the outcome distribution, irrespective of individuals' characteristics. OASD combines both features of ATE and QTE: it is interpreted as straightforwardly as ATE while at the same time more granular than ATE by breaking the entire population up according to the rank of the outcome distribution.
In addition to providing identification results for OASD, we show there is a close relationship between the outcome conditioned average partial effects and a class of parameters measuring the effect of counterfactually changing the distribution of a single covariate on the unconditional outcome quantiles. We illustrate this point by two examples: equivalence between OASD and the unconditional partial quantile effect (Firpo et al. (2009)), and equivalence between the marginal partial distribution policy effect (Rothe (2012)) and a corresponding outcome conditioned parameter.
Because identification of OASD is attained under a conditional exogeneity assumption, by controlling for a rich information about covariates, a researcher may ideally use high-dimensional controls in data. We propose for OASD a novel automatic debiased machine learning estimator, and present asymptotic statistical guarantees for it. We prove our estimator is root-$n$ consistent, asymptotically normal, and semiparametrically efficient. We also prove the validity of the bootstrap procedure for uniform inference on the OASD process. Simulation studies support our theories.
\noindentAppendix A. Proofs
{A.1 Notations and Assumptions}
{Denote} three nuisance parameters in Proposition 2.2 by $\eta_1(\cdot)$, $\eta_2(\cdot)$, and $\eta_3(\cdot)$, respectively, i.e.,
{The} following regularity conditions are needed to show Proposition 2.1.
{Assumption A.1.} Conditional CDF $F_Y(y|d,x)$ is absolutely continuous with respect to the Lebesgue measure for in a neighborhood of $d\in\mathcal{S}_D$ given $x$. The density $f_Y(y|d,x)$ is continuous at $(y,d)=\left(Q_Y(\tau|d,x),d\right)$ and bounded in $y\in\mathbb{R}$.
{Assumption A.2.} $Q_Y(\tau|d,x)$ is partially differentiable with respect to $d$. There exists a measurable function $\Delta$ that satisfies
for $\delta\rightarrow 0^+$ and any fixed $\epsilon>0$. We write $\partial_dm(d,x,u)$ for $\Delta(u)$ and $\partial_dm\left(d,x,U_{d}\right)$ for $\Delta(U_{d})$.
{Assumption A.3.} The conditional distribution of $\left(Y,\partial_dm\left(d,x,U_{d}\right)\right)$ given $D=d$ and $X=x$ is absolutely continuous with respect to the Lebesgue measure. For the conditional density $f_{Y,\partial_dm\left(d,x,U_{d}\right)|D,X}$ of $(Y;\partial_dm\left(d,x,U_{d}\right))$ given $D$ and $X$, we require that $f_{Y,\partial_dm\left(d,x,U_{d}\right)|D,X}(y,y'|d,x)\leq Cg(y')$, where $C$ is a constant and $g$ is a positive density on $\mathbb{R}$ with finite mean (i.e., $\int |y'|g(y')dy'<\infty$).\\ \\ {A.2 Proofs for Propositions 2.1--2.4}.\qquad
{Proof of Proposition 2.1.} Note that by definition of $Q_Y(\tau|d,x)$
Similarly, for $\delta>0$
Thus
where
For $A_1$ we get, for $\delta\rightarrow0^+$
where the first equality follows from definition of $A_1$, the second from data generating process $Y=m(D,X,U_D)$, the third from simple algebra and Assumptions A1--A2, and the last from Taylor expansion with Peano remainder and $\delta\rightarrow0^+$. For $A_2$ we get
where the first equality follows from definition of $A_2$, the second from Assumption 2.2, the third from Assumption 2.1, and the last from simple algebra. For $A_3$ we get
where the first equality follows from definition of $A_3$, the second from Assumption 2.2, and the last from simple algebra. For $A_4$ we get, for $\delta\rightarrow0^+$
where the first equality follows from definition of $A_4$, the second from data generating process, the third and fourth from simple algebra, the fifth from data generating process, the sixth from Assumptions A2--A3 and $\delta\rightarrow0^+$, the seventh and last from simple algebra. Let $-\frac{y-Q_Y(\tau|d,x)}{\delta}=u$, then $A_4$ can be simplified to
where the first equality follows from substitution method for definite integral, the second and third from property of the integral, the fourth from $\delta\rightarrow0^+$ and Taylor expansion, the fifth and sixth from simple algebra, and the last from definition of conditional expectation.
From (1)--(5) we get
Thus
The following process provides the identification of $\partial_dQ_Y(\tau|d,x)\big|_{\tau=F(y|d,x)}$ from the perspective of the definition of $\tau$-th conditional quantile
Taking derivative with respect to $d$ on both sides
By simple algebra
Evaluating at $\tau=F_Y(y|d,x)$ on both sides
Then $\theta(y_{1},y_{2})$ can be identified by
where the first equality follows from definition of $\theta_0(y_1,y_2)$, the second from simple algebra, the third from the law of iterated expectation, the fourth from property of conditional expectation and $f(d,x,y)=f(d,x|y)f(y)$, the fifth from equation (6), the sixth from equation (7), and the remaining from simple algebra.$\blacksquare$\\
{Proof of Proposition 2.2.} We rewrite the score function in Proposition 2.2 in terms of $\eta=(\eta_1,\eta_2,\eta_3)$ as
Note that
where the first equality follows from definition of $\psi(\cdot)$, Proposition 2.1, and simple algebra, the second from law of iterated expectation, the third from definition of conditional expectation, the fourth from property of double integral, and the remaining from simple algebra.
Thus, $\theta(y_{1},y_{2})$ satisfies $ E\left( \psi\left(y_1,y_2,W;\theta(y_1,y_2),\eta\right)\right)=0$.
Note that $\eta$ is the true value of the nuisance parameter $\tilde{\eta}\in\mathcal{T}$. Then the pathwise (or the Gateaux) derivative against the nuisance parameters at $0$ is
where the first equality follows from simple algebra, and the second from the law of iterated expectation.
Note that
where the first equality follows from definition of $\eta_3(D,X)$, the second from definition of expectation, the third from simple algebra, the fourth from integration by parts and the regularity condition that $f(d,x)$ is zero on the boundary of $\mathcal{S}_D$, and the last from definition of expectation.
Thus
$\blacksquare$\\ \\ \noindentProof of Proposition 2.3. For (i), it follows from (2.1.6) that
By Firpo et al. (2009, Corollary 1, page 958), $UQPE(\tau)$ can be expressed as
which is equal to $\theta\left(Q_{Y}\big(\tau\big)\right)$ by replacing $y=Q_{Y}(\tau)$. Result (ii) follows from (i) and
For (iii), we only need to show that
The desired result follows by applying (i). Notice that by (2.1.4), \[
\] Replacing $y_{1}$ by $Q_{Y}(\tau)$ yields \[
\] which completes the proof. $\blacksquare$\\ \\ \noindentProof of Proposition 2.4.
{Assumption A.4.} The support of $H_t$ is a subset of the support of $D$ conditional on $X$, that is, $\text{supp}(H_t)\subset \text{supp}(D|X=x)$ for all $x\in \text{supp}(X)$.
Under regularity condition Assumption A.4 and Assumptions 2.1--2.2, similar to Rothe (2012, Lemma 1, page 2274), we get
where the first equality follows from definition of unconditional distribution function, the second from data generating process of $Y_{H_t}$, the third from $D_{H_t}=H_t^{-1}(R_0)$, which implies by continuity of $D$, the fourth from the law of iterated expectation, the fifth from Assumption 2.2 and $D=F_0^{-1}(R_0)$, the sixth from Assumption 2.1, the seventh from change of variable and Assumption A.4, a regularity condition, the eighth from Assumption 2.2, the ninth from property of conditional probability, the tenth from data generating process $Y=m(D,X,U_D)$, the eleventh from change of variable and Assumption A.4, the twelfth from definition of expectation, and the last from $R_0=F_0(D)$, which implies by continuity of $D$ and $D=Q_0(R_0)$.
Thus $f_{Y_{H_t}}(\cdot)$, the corresponding PDF of $Y$ under distribution $F_{Y_{H_t}}(\cdot)$, can be written as
Taking the derivative of the both sides with respect to $t$
Evaluating at $t=0$
where the last equality follows from $R_0=F_0(D)$.
We refer PDFs of continuously distributed random variable $D$ under distribution $F_0$, $G_0$ and $H_t$ as the corresponding lowercase notations $f_0$, $g_0$ and $h_t$, respectively. From $H_t(d)=F_0(d)+t\left(G_0(d)-F_0(d)\right)$, we get $h_t(d)=f_0(d)+t\left(g_0(d)-f_0(d)\right)$.
By definition, we get
Taking the derivative of both sides with respect to $t$
Evaluating at $t=0$
By simple algebra and property of CDFs
Thus
where the second equality from $R_0=F_0(D)$.
By definition
Taking the derivative of both sides with respect to $t$
Evaluating at $t=0$
Thus
where the first equality follows from definition of $MQPE(\tau,G_0)$, the second from simple algebra, the third from equation (9), the fourth from equation (10), and the last from property of CDF.
Moreover, $\varsigma\left(Q_Y(\tau),G_0\right)$, the average treatment effect of changing the unconditional distribution of $D$ infinitesimally towards the direction of $G_0$ on the individuals with $Y=Q_Y(\tau)$, simplifies to
where the first equality follows definition of $\varsigma(\cdot,G_0)$, the second from the chain rule, the third from equation (10), the fourth from the law of iterated expectation, the fifth from equations (6)--(7), which hold under assumption in Proposition 2.1, the sixth from $f\left(d,x,Q_Y(\tau)\right)=f_Y\left(Q_Y(\tau)|d,x\right)f_{DX}(d,x)$, and the last from definition of expectation.
Thus, we get
which completes the proof of (i) in Proposition 2.4.
For (ii),
where the first equality follows from definition of $\varsigma(y_1,y_2,G_0)$ and $\varsigma(y,G_0)$, the second from property of definite integral, and the last from part (i) of Proposition 2.4.
This completes the proof. $\blacksquare$\\ \\ {A.3 Double Robustness Property}
\noindentProof. If $\eta_2(D,X;y_1,y_2)$ is correctly specified and $\eta_3(D,X)$ is specified as $\tilde{\eta}_3(D,X)$, we get
where the first equality follows from definition of $\psi(\cdot)$, Proposition 2.1, and simple algebra, the second from law of iterated expectation, the third from definition of conditional expectation, the fourth from property of double integral, and the remaining from simple algebra.
If $\eta_3(D,X)$ is correctly specified and $\eta_2(D,X;y_1,y_2)$ is specified as $\tilde{\eta}_2(D,X;y_1,y_2)$, we get
where the first equality follows from definition of $\psi(\cdot)$ and simple algebra, the second from definition of expectation, the third from integration by parts, and the last from Proposition 2.1.$\blacksquare$ \\ \\ {A.4 Binary treatment variable}.
When the treatment is binary, the quantity in parallel with OASD is
\noindentAssumption A.5 (Monotonicity condition) $m(d,x,u)$ is strictly increasing with respect to $u$ for each $d$ and $x$.
\noindentProposition A.1. Under Assumptions 2.1-2.2 and Assumption A.5,
{Proof.} Define
Thus
where the first equality follows from definition of $\vartheta(d,x,y)$, the second from property of conditional expectation, the third from data generating process, the fourth from Assumption A.5, and the last from property of conditional expectation.
Denote the distribution of $U_0$ conditional on $X=x$ as $F_{U_0}(\cdot|x):=P(U_0\leq u_0|X=x)$, then $F_{U_0}\left(m^{-1}(d,x,y)|x\right)$ can be identified for that
where the first equality follows from the definition of $F_{U_0}(\cdot|x)$, the second from Assumption 2.1, the third from Assumption 2.2, the fourth and fifth from simple algebra, and the last from data generating process and definition of $F_Y(\cdot|d,x)$.
Note that under mild conditions, $F_{U_0}(\cdot|x)$ is strictly increasing with respect to $u$, then we get, from (12)
Combing the above equality with the identity $m\left(d,x,m^{-1}(d,x,y)\right)\equiv y$, which holds under Assumption A.5, gives
Evaluating at $y=Q_Y(\tau |d,x)$ gives
Thus
Combining (8), (13) and (14) gives
This implies that when treatment variable is binary, the relationship shown in (2.1.3) given in the main text still hold under Assumptions 2.1, 2.2 and A.5.
We get
Thus, $\vartheta(y_1,y_2)$ can be identified by
where the first equality follows from definition of $\vartheta(y_1,y_2)$, the second and third from law of iterated expectation, and the remaining from simple algebra. $\blacksquare$ \\
To clarify the orthogonal score, we first define the corresponding nuisance parameter
where $F_Y^{-1}(\cdot|1,x)$ and $F_Y^{-1}(\cdot|0,x)$ are equivalent to $Q_Y(\cdot|1,x)$ and $Q_Y(\cdot|0,x)$, respectively.
Let
The expressions of $\phi_1(\cdot)$, $\phi_2(\cdot)$, and $\phi_3(\cdot)$ are
\noindentProposition A.2. (Identification based on orthogonal score with binary treatment variable) Under the same assumptions as Proposition A.1, we have (i) $\vartheta(y_{1},y_{2})$ satisfies $E\Psi\left(y_1,y_2,W;\vartheta(y_{1},y_{2}),\Phi\right)=0$.
(ii) $\Psi\left(W,\theta,\eta;y_1,y_2\right)$ satisfies the Neyman orthogonality property:
{Proof.} To simplify the notation, define
Rewrite $\Psi(\cdot)$ in terms of $\Phi(\cdot)$
where $\phi_1^{-1}(\cdot,x)$ and $\phi_2^{-1}(\cdot,x)$ denote the left-inverse transform of $\phi_1(y,x)$ and $\phi_2(y,x)$ with respect to $y$.
Moreover, $\phi_1(\cdot)$, $\phi_2(\cdot)$, and $\phi_3(\cdot)$ can be written as
Note that
Thus
Similarly,
Thus
From (12) and (16), the proof of (i) is apparent.
Before starting the proof of (ii), we define some useful notations. For $k=1,2,\cdots,6$, $\Phi_{k,r}(\cdot)=\Phi_k(\cdot)+r\left(\tilde{\Phi}_k(\cdot)-\Phi_k(\cdot)\right)$ with $\tilde{\Phi}_k(\cdot)$ being an alternative nuisance function.
Consider that $\Phi_1(y,x)$ is misspecified.
For $\dag_1$, we get
For $\dag_2$, we get
For $\dag_3$, we get
The following identity holds
Taking derivative with respect to $\tau$ gives
By simple algebra
Thus
We get
For $\dag_4$, we get
For $\dag_5$, we get
For $\dag_6$, we get
Then
Consider that $\Phi_2(y,x)$ is misspecified.
For $\ddagger_1$, we get
Similar to the simplification of $\dag_3$, we get
For $\ddagger_2$, we get
For $\ddagger_3$, we get
For $\ddagger_4$, we get
For $\ddagger_5$, we get
For $\ddagger_6$, we get
Then
From (15)--(16), we get
Moreover, by (15) and (16), we get
Then, we can conclude that
$\blacksquare$\\
\noindentProof of Theorem 4.1. In the proof main theorems, $a\lesssim{b}$ means that $a\leq{Ab}$, where the constant $A$ depends on the constants in Assumptions 4.1-4.3 only, but not on $n$. We suppress the claim “uniformly over $u\in\mathcal{U}$" throughout the proof.
\noindentStep 1. (Linearization) In this step, we establish the claim that the pre-estimator has no first order effects, namely \[ \sqrt{n}\left(\widehat{\theta}(u)-\theta(u)\right)=Z_{n}(u)+o_{p}(1) \quad in \quad \mathbb{D}=\ell^{\infty}\left(\mathcal{U}\right), \] where $Z_{n}(u)=\mathbb{G}_{n}\psi(W,\theta,\eta;u)$.
Define the following spaces of functions: \[ \mathcal{F}_{1}=\left\{
\right\}, \] \[ \mathcal{F}_{2}=\left\{
\right\}, \] \[ \mathcal{F}_{3}=\left\{
\right\}, \] \[ \mathcal{L}=\left\{
\right\}. \] We observe that with probability no less than $1-\Delta_{n}$, \[ \widehat{IF}(u,d,x)\in\mathcal{F}_{1}, \quad \widehat{DIF}(u,d,x)\in\mathcal{F}_{2}, \quad \mathbb{E}_{n}\widehat{DIF}(u,D_{i},X_{i})\in\mathcal{F}_{3} \quad \text{and} \quad \widehat{L}(d,x)\in\mathcal{L}. \] To see this, note that under Assumption 4.2 and 4.3, \[
\] and \[
\] for $\widetilde{\beta}(y)=\widehat{\beta}(y)$, with evaluation after computing the norms, and for $\Vert\partial\Lambda\Vert_{\infty}$ denoting $\sup_{k\in\mathbb{R}}|\partial\Lambda(k)|$ here and below. Similarly, \[
\] and \[
\] for $\widetilde{\beta}(y)=\widehat{\beta}(y)$, with evaluation after computing the norms. The verification of the remaining terms are identical and omitted. Moreover, let \[ \mathcal{P}=\left\{
\right\}. \] Obviously, with probability no less than $1-\Delta_{n}$, we can observe that $\widehat{P}(u)\in\mathcal{P}$.
We have that \[
\] with $\widetilde{\eta}$ evaluated at $\widetilde{\eta}=\widehat{\eta}$.
Firstly, we consider $\dagger_{4.1.3}$. Note that for \[ \Delta_{\widetilde{\eta}}=\left(\Delta_{\widetilde{\eta}}^{1},\Delta_{\widetilde{\eta}}^{2},\Delta_{\widetilde{\eta}}^{3}\right)^{\prime}=\widetilde{\eta}-\eta. \] For all $k=\{k_{j}\}_{j=1}^{3}\in\mathbb{N}^{3}$: $0\leq|k|\leq3$, $|k|=\sum_{j=1}^{3}k_{j}$ and $\partial_{\widetilde{\eta}}^{k}=\partial_{\widetilde{\eta}_{1}}^{k_{1}}\partial_{\widetilde{\eta}_{2}}^{k_{2}}\partial_{\widetilde{\eta}_{3}}^{k_{3}}$. After applying Taylor expansion, \[
\] with $\widetilde{\eta}$ evaluated at $\widetilde{\eta}=\widehat{\eta}$ after computing the expectations. By the law of iterated expectations and the orthogonality property of the moment function for $\eta$, $\forall k\in\mathbb{N}^{3}:|k|=1$, \[ E\bigg[\partial_{\widetilde{\eta}}^{k}{\psi}\left(W,\theta,{\eta};u\right)\Delta_{\widetilde{\eta}}^{k}\bigg]=0, \] and thus $\dagger_{4.1.3a}=0$. Moreover, uniformly over $\widetilde{\eta}\in\mathcal{R}=\mathcal{P}\times\left(\mathcal{F}_{1}\cup\mathcal{F}_{2}\cup\mathcal{F}_{3}\right)\times\mathcal{L}$, we have \[
\] Since $\widehat{\eta}\in\mathcal{R}$, with probability $1-\Delta_{n}$, \[ P\bigg(\left|\dagger_{4.1.3}\right|\lesssim\delta_{n}\bigg)\geq1-\Delta_{n}. \]
Then we consider $\dagger_{4.1.2}$. With probability $1-\Delta_{n}$, \[
\] Applying similar arguments as the proceeding one, uniformly over $\widetilde{\eta}\in\mathcal{R}$, we have \[ E\bigg[\psi\left(W,\theta,\widetilde{\eta};u\right)-\psi\left(W,\theta,{\eta};u\right)\bigg]^{2}\lesssim\Vert\widetilde{\eta}-\eta\Vert_{P,2}^{2}+\Vert\widetilde{\eta}-\eta\Vert_{P,2}^{2}\Vert\widetilde{\eta}-\eta\Vert_{P,\infty}^{2}=o_{p}\left(n^{-1/2}\right). \] Thus, with probability $1-\Delta_{n}$, \[ P\bigg(\left|\dagger_{4.1.2}\right|\lesssim\delta_{n}n^{-1/4}\bigg)\geq1-\Delta_{n}. \]
\noindentStep 2. (Uniform Donskerness) Here we claim that Assumptions 4.1-4.3 imply that the set of functions $\{\psi\left(W,\theta,{\eta};u\right)\}_{u\in\mathcal{U}}$ is $P$-Donsker, namely \[ Z_{n}(u)\leadsto Z(u) \quad in \quad \mathbb{D}=\ell^{\infty}\left(\mathcal{U}\right), \] where $Z(u)=\mathbb{G}\psi\left(W,\theta,{\eta};u\right)$.
We apply Theorem B.1 in Belloni et al. (2017) to verify this claim. The classes of functions \[ \mathcal{V}_{1}=\left\{\displaystyle\int_{y_{1}}^{y_{2}}1\{Y\leq{y}\}dy:(y_{1},y_{2})\in\mathcal{U}\right\} \] and \[ \mathcal{V}_{2}=\bigg\{1\{y_{1}\leq{Y}\leq{y_{2}}\}:(y_{1},y_{2})\in\mathcal{U}\bigg\} \] viewed as maps from the sample space $\mathcal{W}$ to the real line, are bounded by constant envelops and have finite VC dimensions. According to Theorem 2.6.7 in van der Vaart and Wellner (1996), we can deduce that \[
\] with the supremum taken over all finitely discrete probability measures $Q$ on $(\mathcal{W},\mathcal{A}_{\mathcal{W}})$. According to Lemma L.2 in Belloni et al. (2017), the following class of functions \[ \mathcal{V}_{3}=\left\{\displaystyle\int_{y_{1}}^{y_{2}}F_{Y}(y|D,X)dy:(y_{1},y_{2})\in\mathcal{U}\right\} \] is bounded by a constant envelop and obeys \[ \sup_{Q}\log N\left(\epsilon,\mathcal{V}_{3},\Vert\cdot\Vert_{Q,2}\right)\lesssim\log(1/\epsilon)\vee{0}.\\ \] The class of functions \[ \mathcal{V}_{4}=\left\{\partial_{D}\displaystyle\int_{y_{1}}^{y_{2}}F_{Y}(y|D,X)dy:(y_{1},y_{2})\in\mathcal{U}\right\} \] is bounded by a measurable envelop $T\geq\sup_{(y_{1},y_{2})\in\mathcal{U}}\left|\partial_{D}\displaystyle\int_{y_{1}}^{y_{2}}F_{Y}(y|D,X)dy\right|$ with $\Vert{T}\Vert_{P,2}\leq{\infty}$ and obeys \[ \sup_{Q}\log N\left(\epsilon\Vert{T}\Vert_{Q,2},\mathcal{V}_{4},\Vert\cdot\Vert_{Q,2}\right)\lesssim(1/\epsilon)^{(1+d_{X_{c}})/\sigma} \\ \] by Corollary 2.7.2 in van der Vaart and Wellner (1996) and the relationship between covering numbers and bracketing numbers. The classes of functions \[ \mathcal{V}_{5}=\bigg\{P\bigg(y_{1}\leq{Y}\leq{y_{2}}\bigg):(y_{1},y_{2})\in\mathcal{U}\bigg\} \] and \[ \mathcal{V}_{6}=\left\{E\left[\partial_{D}\displaystyle\int_{y_{1}}^{y_{2}}F_{Y}(y|D,X)dy\right]:(y_{1},y_{2})\in\mathcal{U}\right\} \] are bounded by constant envelops and obey \[
\] which holds by Lemma L.2 in Belloni et al. (2017) or Lemma A.2 in Ghosal et al. (2000). Moreover, the VC dimension of the measurable function set \[ \mathcal{V}_{7}=\left\{g:(D,X)\mapsto\frac{\partial_{D}f(D,X)}{f(D,X)}\right\} \] is finite according to Lemma 2.6.15 in van der Vaart and Wellner (1996). Thus, $\mathcal{V}_{7}$ is bounded by a measurable envelop $T^{*}(W)=\left|\dfrac{\partial_{D}f(D,X)}{f(D,X)}\right|$ with $\Vert{T^{*}}\Vert_{P,2}\leq{\infty}$ and obeys \[ \sup_{Q}\log N\left(\epsilon\Vert{T^{*}}\Vert_{Q,2},\mathcal{V}_{7},\Vert\cdot\Vert_{Q,2}\right)\lesssim\log(1/\epsilon)\vee{0}. \]
By a slight abuse of notation, let $\eta$ also contains $\displaystyle\int_{y_{1}}^{y_{2}}1\{Y\leq{y}\}dy$ and $1\{y_{1}\leq{Y}\leq{y_{2}}\}$, such that \[ \eta(W;u)=\left(\displaystyle\int_{y_{1}}^{y_{2}}1\{Y\leq{y}\}dy,1\{y_{1}\leq{Y}\leq{y_{2}}\},P\big(y_1<Y<y_2\big),\int_{y_1}^{y_2}F_Y\left(y\big|D,X\right)dy,\frac{\partial_Df\big(D,X\big)}{f\big(D,X\big)}\right). \] Note that \[ \mathcal{Z}=\big\{ \psi\left(W,\theta,{\eta};u\right):u\in\mathcal{U} \big\} \] is formed as a uniform Lipschitz transform of the function sets $\mathcal{V}_{1}$, $\mathcal{V}_{2}$, $\mathcal{V}_{3}$, $\mathcal{V}_{4}$, $\mathcal{V}_{5}$, $\mathcal{V}_{6}$ and $\mathcal{V}_{7}$, and is bounded by a measurable envelop $\bar{T}$ with $\left\Vert\bar{T}\right\Vert_{P,2}\leq\infty$.\footnote{The envelop $\bar{T}$ of $\mathcal{Z}$ can be constructed as $\bar{T}=2\left(\sum_{k=1}^{5}L_{k}^{2}R_{k}^{2}\right)^{1/2}$, where $R_{k}$ denotes the envelop of $\mathcal{V}_{k}$, and $L_{k}$ satisfies \[ |\psi\left(W,\theta,\widetilde{\eta};u\right)-\psi\left(W,\theta,\bar{\eta};u\right)|^{2}\leq\sum_{k=1}^{5}L_{k}^{2}(W;u)|\widetilde{\eta}_{k}(W;u)-\bar{\eta}_{k}(W;u)|^{2}, \] for all $\widetilde{\eta}$ and $\bar{\eta}$ in $\mathcal{V}_{1}\times\mathcal{V}_{2}\times\mathcal{V}_{5}\times\left(\mathcal{V}_{3}\cup\mathcal{V}_{4}\cup\mathcal{V}_{6}\right)\times\mathcal{V}_{7}$. } According to Theorem 2.10.20 of van der Vaart and Wellner (1996), the class of functions $\mathcal{Z}$ obeys \[ \sup_{Q}\log N\left(\epsilon\left\Vert{\bar{T}}\right\Vert_{Q,2},\mathcal{Z},\Vert\cdot\Vert_{Q,2}\right)\lesssim(1/\epsilon)^{(1+d_{X_{c}})/\sigma}. \] Since \[ \lim_{\delta\to{0}}\int_{0}^{\delta}\sqrt{(1/\epsilon)^{(1+d_{X_{c}})/\sigma}} d\epsilon\to{0}, \] by Assumption 4.3(iv), the entropy condition (B.2) in Theorem B.1 of Belloni et al. (2017) holds.
The first condition in (B.1) is trivially satisfied. We demonstrate the second condition in (B.1). Consider a sequence of positive constants $\epsilon$ approaching zero, and it suffice to verify that \[ \lim_{\epsilon\to{0}^{+}}\sup_{d_{\mathcal{U}}\left(u,\widetilde{u}\right)\leq{\epsilon}}\left\Vert\psi\left(W,\theta,{\eta};u\right)-\psi\left(W,\theta,{\eta};\widetilde{u}\right)\right\Vert_{P,2}=0. \] Notice that \[
\] Under Assumption 4.2, $\dagger_{4.1.4}$ and $\dagger_{4.1.5}$ converges to $0$ as $d_{\mathcal{U}}\left(u,\widetilde{u}\right)\to{0}$. Note that \[ \displaystyle\int_{y_{1}}^{y_{2}}F_{Y}(y|D,X)dy=E\left[\displaystyle\int_{y_{1}}^{y_{2}}1\{Y\leq{y}\}dy\bigg|D,X\right] \] By the contradiction property of the conditional expectation, \[ \dagger_{4.1.6}\leq\dagger_{4.1.4}\to{0} \] as $d_{\mathcal{U}}\left(u,\widetilde{u}\right)\to{0}$. For $\dagger_{4.1.7}$, \[
\] as $d_{\mathcal{U}}\left(u,\widetilde{u}\right)\to{0}$ by Assumption 4.3(iii). Consequently, we have that $\dagger_{4.1.8}\to{0}$ and $\dagger_{4.1.9}\to{0}$ as $d_{\mathcal{U}}\left(u,\widetilde{u}\right)\to{0}$. $\blacksquare$\\
\noindentProof of Theorem 4.2. The joint density of the observed variables $W=(Y,D,X)$ can be written as
Consider a regular parametric submodel indexed by $\epsilon$ with $\epsilon_0$ corresponding to the true model:\\$f(y,d,x;\epsilon_0)\equiv f(y,d,x)$. The density of $f(y,d,x;\epsilon)$ can be written as
We will assume that all terms of the previous equation admit an interchange of the order of integration and differention, which will hold under sufficient condition given by Theorem 1.3.2 of Amemiya (1985) such that \[ \int\frac{\partial f(y,d,x;\epsilon)}{\partial\epsilon}dxdddy=\frac{\partial}{\partial\epsilon}\underbrace{\int f(y,d,x;\epsilon)dxdddy}_1=0 \eqno{(A.4.2.1)} \] The corresponding score of $f(y,d,x;\epsilon)$ is \[ s(y,d,x;\epsilon)=\frac{\partial \ln f(y,d,x;\epsilon)}{\partial\epsilon}=\check{f}_Y(y|d,x;\epsilon)+\check{f}_D(d|x;\epsilon)+\check{f}_{X}(x;\epsilon) \] where $\check{f}$ defines a derivative of the log, that is, $\check{f}_Y(y|d,x;\epsilon)\equiv\partial \ln f_Y(y|d,x;\epsilon)/\partial\epsilon$, $\check{f}_D(d|x;\epsilon)\equiv\partial \ln f_D(d|x;\epsilon)/\partial\epsilon$ and $\check{f}_{X}(x;\epsilon)\equiv\partial \ln f_{X}(x;\epsilon)/\partial\epsilon$. Notice that the expectation of the score is zero if $\epsilon$ is evaluated at the true value $\epsilon_{0}$. According to Proposition 2.1, we have \[ \theta(u)=-\dfrac{\displaystyle\int\bigg(\int1\{y_{1}<y<y_{2}\}1\{t\leq y\}\partial_df_Y(t|d,x)dtdy\bigg)f_{D}(d|x)f_{X}(x)dddx}{\displaystyle\int1\{y_{1}<y<y_{2}\}f(y,d,x)dxdddy}. \] Therefore, the parameter $\theta(u;\epsilon)$ induced by the submodel $f(d,y,x;\epsilon)$ satisfies \[ \theta(u;\epsilon)=-\dfrac{\displaystyle\int\bigg(\int1\{y_{1}<y<y_{2}\}1\{t\leq y\}\partial_df_Y(t|d,x;\epsilon)dtdy\bigg)f_{D}(d|x;\epsilon)f_{X}(x;\epsilon)dddx}{\displaystyle\int1\{y_{1}<y<y_{2}\}f(y,d,x;\epsilon)dxdddy}. \]
The tangent space of the model is the set of functions that are mean zero and satisfy the additive structure of the score:
for any functions $s_y$, $s_{d}$ and $s_{x}$ satisfying the mean zero property \[ E[s_y(Y|D,X)|D,X]=E[s_d(D|X)|X]=Es_x(X)=0. \] Then the semiparametric variance bound of $\theta(u)$ is the variance of the projection on $\Im$ of a function $\Gamma(W;u)$(with $E\Gamma(\cdot;u)=0$ and $E\left\|\Gamma^2(\cdot;u)\right\|<\infty$ for any $u\in\mathcal{U}$) that satisfies for all regular parametric submodels
If $\Gamma(W;u)$ itself already lies in the tangent space, the variance bound is given by $E\Gamma^2(W;u)$ for any $u\in\mathcal{U}$.
We first calculate $\dfrac{\partial\theta(u;\epsilon)}{\partial\epsilon}\bigg|_{\epsilon=\epsilon_0}$. By calculation, \[
\] Recall that the Neyman-orthogonal score is \[
\] To justify our theorem, it suffice to show that (i) \[ \dfrac{\partial\theta(u;\epsilon)}{\partial\epsilon}\bigg|_{\epsilon=\epsilon_0}=E\left[\psi\left(W,\theta,\eta;u\right)\cdot s(W;\epsilon_0)\right] \] and (ii) $\psi\left(W,\theta,\eta;u\right)$ lies in the tangent space $\Im$ for any $u\in\mathcal{U}$.
The second argument can be easily verified. For (i), substituting the representation $\psi\left(W,\theta,\eta;u\right)$ into $E\left[\psi\left(W,\theta,\eta;u\right)\cdot s(W;\epsilon_0)\right]$ implies \[ E\left[\psi\left(W,\theta,\eta;u\right)\cdot s(W;\epsilon_0)\right]=-\frac{1}{P(u)}\bigg(\dagger_{4.2.1}+\dagger_{4.2.2}\bigg)+\frac{E\left(\partial_D\displaystyle\int_{y_1}^{y_2}F_Y(y|D,X)dy\right)}{P^2(u)}\dagger_{4.2.3}, \] where \[ \dagger_{4.2.1}=E\left[\left(\partial_D\int_{y_1}^{y_2}F_Y(y|D,X)dy+\theta(u)\right)\cdot\bigg(\check{f}_Y(Y|D,X;\epsilon_{0})+\check{f}_D(D|X;\epsilon_{0})+\check{f}_{X}(X;\epsilon_{0})\bigg)\right], \] \[ \dagger_{4.2.2}=E\left[\frac{\partial_Df(D,X)}{f(D,X)}\int_{y_1}^{y_2}\bigg(F_Y\left(y\big|D,X\right)-1\big\{Y<y\big\}\bigg)dy\cdot\bigg(\check{f}_Y(Y|D,X;\epsilon_{0})+\check{f}_D(D|X;\epsilon_{0})+\check{f}_{X}(X;\epsilon_{0})\bigg)\right], \] and \[ \dagger_{4.2.3}=E\left[\bigg(1\big\{y_1<Y<y_2\big\}-P(u)\bigg)\cdot\bigg(\check{f}_Y(Y|D,X;\epsilon_{0})+\check{f}_D(D|X;\epsilon_{0})+\check{f}_{X}(X;\epsilon_{0})\bigg)\right]. \] By $(A.4.2.1)$, we have \[ E\left[\check{f}_Y(Y|D,X;\epsilon_{0})|D,X\right]=E\left[\check{f}_D(D|X;\epsilon_{0})|X\right]=E\check{f}_{X}(X;\epsilon_{0})=0. \] For $\dagger_{4.2.1}$, we get \[
\] Similarly, we can derive \[
\] and \[
\] which completes the proof. $\blacksquare$ \\
\noindentProof of Theorem 4.3. Let $P^{*}=P\times{P}_{\xi}$. Then the operator $E_{P^{*}}$ then denotes the expectation with respect to $P^{*}=P\times{P_{\xi}}$ and $\mathbb{G}_{n}$ denotes the corresponding empirical process, that is \[ \mathbb{G}_{n}B(\xi,W)=\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\bigg[B(\xi_{i},W_{i})-E_{P^{*}}B(\xi,W)\bigg]. \]
Recall that we define the bootstrap draw as \[ \widehat{Z}_{n}^{*}(u)=\sqrt{n}\left(\widehat{\theta}^{*}(u)-\widehat{\theta}(u)\right)=\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\xi_{i}\psi\bigg(W_{i},\widehat{\theta},\widehat{\eta};u\bigg)=\mathbb{G}_{n}\xi\psi\bigg(W,\widehat{\theta},\widehat{\eta};u\bigg), \] since $E_{P^{*}}\left[\xi\psi\bigg(W,\widehat{\theta},\widehat{\eta};u\bigg)\right]=0$ because $\xi$ is independent of $W$ and has zero mean. The proof also consists of two steps.
\noindentStep 1. In this step, we establish that \[ \widehat{Z}_{n}^{*}(u)={Z}_{n}^{*}(u)+o_{p^{*}}(1), \quad in \quad \mathbb{D}=\ell^{\infty}\left(\mathcal{U}\right), \] where $Z_{n}^{*}(u)=\mathbb{G}_{n}\xi\psi(W,\theta,\eta;u)$.
Define $\widetilde{\psi}(W,\eta;u)=\psi(W,\theta,\eta;u)+\theta$. We then have the representation \[
\] According to Theorem 4.1, $\left(\widehat{\theta}(u)-{\theta}(u)\right)\mathbb{G}_{n}\xi=O_{p^{*}}\left(n^{-1/2}\right)=o_{p^{*}}(1)$. With probability $1-\Delta_{n}$, \[
\] Similarly, uniformly over $\widetilde{\eta}\in\mathcal{R}$, \[ E\bigg[\widetilde{\psi}\left(W,\widetilde{\eta};u\right)-\widetilde{\psi}\left(W,\eta;u\right)\bigg]^{2}\lesssim\Vert\widetilde{\eta}-\eta\Vert_{P,2}^{2}+\Vert\widetilde{\eta}-\eta\Vert_{P,2}^{2}\Vert\widetilde{\eta}-\eta\Vert_{P,\infty}^{2}=o_{p}\left(n^{-1/2}\right). \] Thus, we can conclude that with probability $1-\Delta_{n}$, $\dagger_{4.3.1}=o_{p^{*}}(1)$.
\noindentStep 2. Here we claim that \[ \widehat{Z}_{n}^{*}(u)\leadsto_{B} Z(u) \quad in \quad \mathbb{D}=\ell^{\infty}\left(\mathcal{U}\right). \] Applying Theorem B.2 in Belloni et al. (2017) or equivalently, Theorem 2 in Kosorok (2003), we have ${Z}_{n}^{*}(u)\leadsto_{B} Z(u)$ in $\mathbb{D}=\ell^{\infty}\left(\mathcal{U}\right)$. Then by Lemma 2 in Chiang et al. (2019) and the result in Step 1, we have $\widehat{Z}_{n}^{*}(u)\leadsto_{B} Z(u)$ in $\mathbb{D}=\ell^{\infty}\left(\mathcal{U}\right)$. $\blacksquare$\\ \\ \noindentAppendix B. Monte Carlo Simulations for CDF derivative
In this appendix, we compare the performance of CDF derivative based on our estimator and Sasaki et al. (2022). We consider the same data-generating process as Sasaki et al. (2022). The outcome variable is generated according to the partial linear high-dimensional modem
where the function $g(d)$ is defined in the following three ways
The treatment variable $D$ is generated by
where $(X_1,\cdots,X_p)\sim N(0,\Sigma_p)$ with $\Sigma_p$ be the $p\times p$ variance-covariance matrix whose $(r-c)$ element is $0.5^{2\left(|r-c|+1\right)}$. The following four cases of varying sparsity level are considered
Across sets of Monte Carlo simulations, we vary DGP$\in$\{DGP1, DGP2, DGP3\}, and the sparsity design $\in$\{(i), (ii), (iii), (iv)\}. We set $b(d,x)$ by including powers of $D$ and $X$ up to the third degree, i.e., $b(d,x)=\left(d,x,d^2,(x^2)^\intercal,d^3,(x^3)^\intercal\right)^\intercal$. We fix the sample size $n=500$ and the dimension $p=99$ throughout.
We estimate $\partial_DF_Y(y_\tau|D,X)$ with $\tau\in\{0.1,0.2,\cdots,0.9\}$. $y_{0.1},y_{0.2},\cdots,y_{0.9}$ equal to the 10%, 20%, $\cdots$, $90\%$-th quantiles of $Y$ distribution.
Sasaki et al. (2022) propose to estimate $\partial_DF_Y(y_\tau|D,X)$ by
where $\Lambda'(\cdot)$ is the derivative function of logistic function $\Lambda(\cdot)$, $\bar{\beta}_\tau$ is based on the estimation procedure in Sasaki et al. (2022, page 959--960). By contrast, we propose to estimate $\partial_DF_Y(y_\tau|D,X)$ by
where $h=n^{-1/6}$. $\hat{\beta}_\tau$ is the post-lasso estimator with penalty level and penalty loading described in Step 1 of Section 3. This is the special case of our estimator based on Eq. (3.2) by setting $\ell=1$.
From data-generating process the true function form of $\partial_DF(y_\tau|D,X)$ is
where $\phi$ is the probability density function of standard normal distribution, $g'(\cdot)$ is the derivative function of $g(\cdot)$. In each simulation, we calculate $L^2$ distance between estimator and true function by
We do 500 iterations to compute mean $L^2$ distance (MeanDist).
Tables B.1--2 summarize the simulation results under the sparsity designs (i), (ii), (iii), and (iv). We can conclude that our estimation procedure outperforms the one proposed in Sasaki et al. (2022) in all of data-generating processes, especially at the tails of the unconditional distribution of $Y$.