EconBase
← Back to paper

Identification and Estimation of Unconditional Policy Effects of an Endogenous Binary Treatment: An Unconditional MTE Approach

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

123,082 characters · 24 sections · 67 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identification and Estimation of Unconditional Policy Effects of an Endogenous Binary Treatment: An Unconditional MTE Approach

singlespace\begin{abstract} This paper studies the identification and estimation of policy effects when treatment status is binary and endogenous. We introduce a new class of marginal treatment effects (MTEs) based on the influence function of the functional underlying the policy target. We show that an unconditional policy effect can be represented as a weighted average of the newly defined MTEs over the individuals who are indifferent about their treatment status. We provide conditions for point identification of the unconditional policy effects. When a quantile is the functional of interest, we introduce the UNconditional Instrumental Quantile Estimator (UNIQUE) and establish its consistency and asymptotic distribution. In the empirical application, we estimate the effect of changing college enrollment status, induced by higher tuition subsidy, on the quantiles of the wage distribution. \end{abstract} \thispagestyle{empty} Keywords: marginal treatment effect, marginal policy-relevant treatment effect, selection model, instrumental variable, unconditional policy effect, unconditional quantile regression. JEL: C14, C31, C36.

\normalem

\pagenumbering{arabic}

Introduction

An unconditional policy effect is the effect of a change in a target covariate on the unconditional distribution of an outcome variable of interest.\footnote{ There are several groups of variables in this framework. The target covariate is the variable a policy maker aims to change. We often refer to the target covariate as the treatment or treatment variable. The outcome variable is the variable that a policy maker ultimately cares about. A policy maker hopes to change the target covariate in order to achieve a desired effect on the distribution of the outcome variable. Sometimes, a policy maker can not change the treatment variable directly and has to change it by\ intervening some other covariates, which may be referred to as the policy covariates. There may also be covariates that will not be intervened.} When the target covariate has a continuous distribution, we may be interested in shifting its location and evaluating the effect of such a shift on the distribution of the outcome variable. For example, we may consider increasing the number of years of education for every worker in order to improve the median of the wage distribution. When the change in the covariate distribution is small, such an effect may be referred to as the marginal unconditional policy effect.

In this paper, we consider a binary target covariate that indicates the treatment status. In this case, a location shift is not possible, and the only way to change its distribution is to change the proportion of treated individuals. We analyze the impact of such a marginal change on a general functional of the distribution of the outcome. For example, when the functional of interest is the mean of the outcome, this corresponds to the marginal policy-relevant treatment effect (MPRTE) of Carneiro2010, Carneiro2011. For the case of quantiles, we obtain an unconditional quantile effect (UQE). Previously, in a seminal contribution, Firpo2009 proposed using\ an unconditional quantile regression (UQR) to estimate the UQE.\footnote{mukhin2019 generalizes Firpo2009 to allow for non-marginal changes in continuous covariates. Other recent studies in the case of continuous covariates include SasakiUraZhang20 who allow for a high-dimensional setting, InoueLiXu21 who tackle a two-sample problem, MontesRojas who analyze location-scale and compensated shifts, and alejo2022 who propose an alternative estimation method based on the slopes of (conditional) quantile regression.} However, we show that their identification strategy can break down under endogeneity. An extensive analysis of the resulting asymptotic bias of the UQR estimator is provided.

The first contribution of this paper is to introduce a new class of unconditional marginal treatment effects (MTEs) and show that the corresponding unconditional policy effect can be represented as a weighted average of these unconditional MTEs. The novel MTEs are derived from the influence function of the functional of the outcome distribution that we care about. This framework allows us to show that the MPRTE and UQE belong to the same family of parameters. To the best of our knowledge, this was not previously recognized in either the literature on MTEs or the literature on unconditional policy effects.

To illustrate the usefulness of this general approach, we provide an extensive analysis of the unconditional quantile effects. This is empirically important since the UQR estimator proposed by Firpo2009 is consistent for the UQE only if a certain distributional invariance assumption holds. Such an assumption is unlikely to hold when the treatment status is endogenous. We note that treatment endogeneity is the rule rather than the exception in economic applications.

The second contribution of this paper is to provide a closed-form expression for the asymptotic bias of the UQR estimator when endogeneity is overlooked. We show that when selection into treatment follows a threshold-crossing model, the UQR estimator can be inconsistent, even if the treatment status is exogenous. This intriguing result underscores the need for caution in using the UQR without careful consideration. The asymptotic bias can be traced back to two sources. First, the subpopulation of the individuals at the margin of indifference might have different characteristics than the whole population. We refer to this source of bias as the marginal heterogeneity bias. Second, the treatment effect for a marginal individual might be different from an apparent effect obtained by comparing the treatment group with the control group. It is the marginal subpopulation, not the whole population or any other subpopulation, that contributes to the UQE. We refer to the second source of bias as the marginal selection bias.

The third contribution of this paper is to show that if assumptions similar to instrument validity are imposed on the policy variables under intervention, the resulting UQE can be point identified using the local instrumental variable approach as in Carneiro2009. Building on this, we introduce the UNconditional Instrumental QUantile Estimator (UNIQUE) and develop methods for statistical inference based on the UNIQUE when the binary treatment is endogeneous. We take a nonparametric approach but allow for the propensity score function to be either parametric or nonparametric. We establish the asymptotic distribution of the UNIQUE. This is a formidable task, as the UNIQUE is a four-step estimator, and we have to pin down estimation errors from each step.

Related Literature. This paper is related to the literature on marginal treatment effects. Introduced by Bjorklund1987, MTEs can be used as a building block for many different causal parameters of interest as shown in Heckman2001. For an excellent review, the reader is referred to mogstad2018. As a main departure from this literature, this paper introduces unconditional MTEs targeted at studying unconditional policy effects. As mentioned above, if we focus on the mean, the unconditional marginal policy effect we study corresponds to the MPRTE of Carneiro2010. In this case, the unconditional MTE is the same as the usual MTE. For quantiles, the unconditional marginal policy effect is studied by Firpo2009, but in a setting that does not allow for endogeneity. \footnote{ For the case of continuous endogenous covariates, Rothe2010b shows that the control function approach of Imbens2009 can be used to achieve identification. This paper considers a binary endogenous covariate.} The unconditional MTE is novel in this case, as well as in the case with a more general nonlinear functional of interest. In our setting that allows some covariates to enter both the outcome equation and the selection equation, conditioning on the propensity score is not enough, but we show that the propensity score plays a key role in averaging the unconditional MTEs to obtain the unconditional policy effect. This is in contrast with Zhou2019 where the marginal treatment effect parameter is defined based on conditioning on the propensity score. Among other contributions, torgo2020 provide an extensive list of applications for MTEs, many of which can also be applied to the unconditional MTEs introduced in this paper.

Our general treatment of the problem using functionals is closely related to that of Rothe2012. Rothe2012 analyzes the effect of an arbitrary change in the distribution of a target covariate, either continuous or discrete, on some feature of the distribution of the outcome variable. By assuming a form of conditional independence, for the case of continuous target covariates, Rothe2012 generalizes the approach of Firpo2009. However, for the case of a discrete treatment, instead of point identifying the effect as we do here, bounds are obtained by assuming that either the highest-ranked or lowest-ranked individuals enter the program under the new policy.

We are not the first to consider unconditional quantile regressions under endogeneity.\footnote{pereda2023 studies the unconditional quantile treatment effect of a binary endogenous regressor. However, the effect there differs from the unconditional quantile effect considered in this paper. Here, unconditional refers to the marginal distribution of the observed $Y$, while in pereda2023 unconditional means that the distribution of the potential outcomes after the covariates are integrated out.} Kasy2016 focuses on ranking counterfactual policies and, for the case of discrete regressors, allows for endogeneity. However, a key difference from our approach is that the counterfactual policies analyzed in Kasy2016 are randomly assigned conditional on a covariate vector. In our setting, selection into treatment follows a threshold-crossing model, where we use the exogenous variation of an instrument to obtain different counterfactual scenarios. martinez2020 introduces the quantile breakdown frontier in order to perform a sensitivity analysis on departures from the distributional invariance assumption employed by Firpo2009.

Estimation of the MTE and parameters derived from the MTE curve is discussed in urzua2006, Carneiro2009, Carneiro2010, Carneiro2011, and ura2021. All of these studies make a linear-in-parameters assumption regarding the conditional means of the potential outcomes, which yields a tractable partially linear model for the MTE curve. This strategy is not very helpful in our case because our newly defined unconditional MTE might involve a nonlinear function of the potential outcomes. For example, in the case of quantiles, there is an indicator function involved. Using our expression for the weights, we can write the UQE as a quotient of two average derivatives. One of them, however, involves as a regressor the estimated propensity score as in the setting of hahn2013. We provide conditions, different from those in hahn2013, under which the error from estimating the propensity score function, either parametrically or nonparametrically, does not affect the asymptotic variance of the UNIQUE. This may be of independent interest.

Outline. Section (ref) introduces the new MTE curve and shows how it relates to the unconditional policy effect. Section (ref) presents a model for studying the UQE under endogeneity. Section (ref) considers intervening an instrumental variable in order to change the treatment status and establishes the identification of the corresponding UQE. Section (ref) introduces and studies the UNIQUE under a parametric specification of the propensity score. Section (ref) provides simulation evidence. In Section (ref) we revisit the empirical application of Carneiro2011 and focus on the unconditional quantile effect. Section (ref) concludes. An appendix contains the proof of the main results. A supplementary appendix provides the technical conditions for two lemmas, the proof of all lemmas and a proposition, the estimation of the asymptotic variance for the UNIQUE, and its asymptotic properties under a nonparametric specification of the propensity score.

Notation. For any generic random variable $W_{1}$, we denote its CDF and pdf by $F_{W_{1}}\left( \cdot \right) $ and $f_{W_{1}}\left( \cdot \right) ,$ respectively. We denote its conditional CDF and pdf conditional on a second random variable $W_{2}$ by $F_{W_{1}|W_{2}}\left( \cdot |\cdot \right) $ and $f_{W_{1}|W_{2}}\left( \cdot |\cdot \right) ,$ respectively.

Unconditional Policy Effects under Endogeneity

Policy Intervention

We employ the potential outcomes framework. For each individual, there are two potential outcomes: $Y(0)$ and $Y(1)$, where $Y(0)$ is the outcome had she received no treatment and $Y(1)$ is the outcome had she received treatment. We assume that the potential outcomes are given by

equation*[equation* omitted — 129 chars of source]

for a pair of unknown functions $r_{0}$ and $r_{1}$. The vector $X\in \mathbb{R}^{d_{X}}$ consists of observables and $U:=\left( U_{0}^{\prime },U_{1}^{\prime }\right) ^{\prime }$ consists of unobservables. Depending on the individual's actual choice of treatment, denoted by $D,$ we observe either $Y(0)$ or $Y(1)$, but we can never observe both. The observed outcome is denoted by $Y$:

equation[equation omitted — 113 chars of source]

Following Heckman1999, Heckman2001, Heckman2005, we assume that selection into treatment is determined by a threshold-crossing equation

equation[equation omitted — 103 chars of source]

where $W:=(Z,X)$ and $Z\in \mathbb{R}^{d_{Z}}$ consists of covariates that do not affect the potential outcomes directly. In the above, the unknown function ${\Greekmath 0116} \left( W\right) $ can be regarded as the benefit from the treatment and $V$ as the cost of the treatment. Individuals decide to take up the treatment if and only if its benefit outweighs its cost. While we observe $\left( D,W,Y\right) $, we observe neither $U$ nor $V.$ Also, we do not restrict the dependence among $U,W,$ and $V$. Hence, they can be mutually dependent and $D$ could be endogenous. Moreover, the dimension of $U$ is left unrestricted, while $V$ is a real-valued random variable.

The propensity score is $P(w):=\Pr \left[ D=1|W=w\right] $. In view of (ref), we can represent it as

equation[equation omitted — 133 chars of source]

If the conditional CDF $F_{V|W}(\cdot |w)$ is a strictly increasing function for all $w\in \mathcal{W}$, the support of $W$, we have

equation*[equation* omitted — 219 chars of source]

where $U_{D}:=F_{V|W}(V|W)$ measures an individual's relative resistance to the treatment, and it can be shown that $U_{D}$ is uniform on $[0,1]$ and is independent of $W.$

To change the treatment take-up rate, we manipulate $Z,$ a subvector of $W$. \footnote{ When we induce the covariate $Z$ to change, the distribution of this covariate will change. However, we do not specify the new distribution a priori. Instead, we specify the policy rule that dictates how the value of the covariate will change for each individual in the population. Our intervention may then be regarded as a value intervention. This is in contrast to a distribution intervention that stipulates a new covariate distribution directly. An advantage of our policy rule is that it is directly implementable in practice while a hypothetical distribution intervention is not. The latter intervention may still have to be implemented via a value intervention, which is our focus here. For a similar comment in the continuous case, see Section 3 in MontesRojas.} More specifically, we consider a policy intervention that changes $Z$ into $ Z_{{\Greekmath 010E} }=\mathcal{G}\left( W,{\Greekmath 010E} \right) $ for a vector of smooth functions $\mathcal{G}\left( \cdot ,\cdot \right) \in \mathbb{R}^{d_{Z}}$. We assume that $\mathcal{G}(W,0)=Z$ so that the status quo policy corresponds to ${\Greekmath 010E} =0.$\footnote{ For notational convenience, when ${\Greekmath 010E} =0$, we drop the subscript and denote $Y_{0}$ and $D_{0}$ as $Y$ and $D,$ respectively.} With the induced change in $Z,$ the selection equation becomes

equation[equation omitted — 249 chars of source]

which can be written as $D_{{\Greekmath 010E} }=\mathds{1}\left\{ U_{D}\leq P_{{\Greekmath 010E} }(W)\right\} $, for $P_{{\Greekmath 010E} }(W)=F_{V|W}({\Greekmath 0116} \left( \mathcal{G}(W,{\Greekmath 010E} ),X\right) |W)$.\footnote{ Note that $U_{D}$ is still defined as $F_{V|W}(V|W)$, and so it does not change under the counterfactual policy regime. This is to say that, relative to others, an individual's resistance to the treatment is preserved across the two policy regimes. In particular, $U_{D}$ is still uniform on $[0,1]$ and independent of $W.$} The outcome equation, in turn, is now

equation[equation omitted — 244 chars of source]

Equations ((ref)) and ((ref)) are the same as the status quo equations; the only exception is that $Z$ has been replaced by $ Z_{{\Greekmath 010E} }.$ We have maintained the structural forms of the outcome equation and the treatment selection equation. Importantly, we have also maintained the stochastic dependence among $U,W,$ and $V,$ which is manifested through the use of the same notation $U,W,$ and $V$ in equations ( (ref)) and ((ref)) as in equations ((ref)) and ((ref)). Our policy intervention has a ceteris paribus interpretation at the population level: we apply the same form of intervention on $Z$ for all individuals in the population but hold all else, including the causal mechanism and the stochastic dependence among the status quo variables, constant. In particular, the conditional distribution of $\left( U,V\right) $ given $W$ is invariant to the value of ${\Greekmath 010E} .$

We allow $\mathcal{G}(\cdot ,{\Greekmath 010E} )$ to take a general form but some examples may be helpful. Consider the case that ${\Greekmath 0116} \left( W\right) =Z$ for a univariate $Z.$ We may take $\mathcal{G}\left( W,{\Greekmath 010E} \right) =Z+{\Greekmath 010E} $ . Under such a policy, we change the value of $Z$ by ${\Greekmath 010E} $ for each individual in the population so that the location of the distribution of $Z$ is shifted by ${\Greekmath 010E} .$ We may also take $\mathcal{G}\left( W,{\Greekmath 010E} \right) =Z\left( 1+{\Greekmath 010E} \right) $, in which case, there is a proportional or scale change in the value of $Z$ for all individuals. Both cases have been studied by Carneiro2010; see also table 2 in mogstad2018. More general location and scale changes, such as those given in MontesRojas, are allowed. In fact, any policy function that satisfies Assumption (ref) in the next subsection is permitted.

Unconditional Policy Effects

To define an unconditional policy effect, we first describe the space of distributions and the functional of interest on this space. Let $\mathcal{F} ^{\ast }$ be the space of finite signed measures ${\Greekmath 0117} $ on $\mathcal{Y} \subseteq \mathbb{R}$ with distribution function $F_{{\Greekmath 0117} }\left( y\right) ={\Greekmath 0117} (-\infty ,y]$ for $y\in \mathcal{Y}$. We endow $\mathcal{F}^{\ast }$ with the usual supremum norm: for two distribution functions $F_{{\Greekmath 0117} _{1}}$ and $F_{{\Greekmath 0117} _{2}}$ associated with the respective signed measures ${\Greekmath 0117} _{1}$ and ${\Greekmath 0117} _{2}$ on $\mathcal{Y}$, we define $\left\Vert F_{{\Greekmath 0117} _{1}}-F_{{\Greekmath 0117} _{2}}\right\Vert _{\infty }:=\sup_{y\in \mathcal{Y}}\left\vert F_{{\Greekmath 0117} _{1}}\left( y\right) -F_{{\Greekmath 0117} _{2}}\left( y\right) \right\vert .$ With some abuse of notation, denote $F_{Y}$ as the distribution of $Y:$ $F_{Y}\left( y\right) ={\Greekmath 0117} _{Y}(-\infty ,y]$ where ${\Greekmath 0117} _{Y}$ is the measure induced by the distribution of $Y.$ Define $F_{Y_{{\Greekmath 010E} }}$ similarly. Clearly, both $ F_{Y}$ and $F_{Y_{{\Greekmath 010E} }}$ belong to $\mathcal{F}^{\ast }.$ We consider a general functional ${\Greekmath 011A} :\mathcal{F}^{\ast }\rightarrow \mathbb{R}$ and study the general unconditional policy effect.

definitionGeneral Unconditional Policy Effect • The general unconditional policy effect for the functional ${\Greekmath 011A} $ is defined as \begin{equation*} \Pi _{{\Greekmath 011A} }:=\lim_{{\Greekmath 010E} \rightarrow 0}\frac{{\Greekmath 011A} [F_{Y_{{\Greekmath 010E} }}]- {\Greekmath 011A}[ F_{Y}]}{E[D_{{\Greekmath 010E} }]-E[D]}, \end{equation*} whenever this limit exists.

In this paper, we consider a Hadamard differentiable functional and its associated influence function.\footnote{ An earlier working paper sun2021 considers the mean functional under different assumptions since the mean functional is not Hadamard differentiable.} For completeness, we provide the definitions of Hadamard differentiability and the influence function below.

definition${\Greekmath 011A} :\mathcal{F}^{\ast }\rightarrow \mathbb{R}$ is Hadamard differentiable at $F\in \mathcal{F}^{\ast }$ if there exists a linear and continuous functional $\dot{{\Greekmath 011A}}_{F}:\mathcal{F}^{\ast }\rightarrow \mathbb{R}$ such that for any $G\in \mathcal{F}^{\ast }$ and $ G_{{\Greekmath 010E} }\in \mathcal{F}^{\ast }$ with $\lim_{{\Greekmath 010E} \rightarrow 0}\left\Vert G_{{\Greekmath 010E} }-G\right\Vert _{\infty }=0,$ we have \begin{equation*} \lim_{{\Greekmath 010E} \rightarrow 0}\frac{{\Greekmath 011A} \lbrack F+{\Greekmath 010E} G_{{\Greekmath 010E} }]-{\Greekmath 011A} \lbrack F]}{{\Greekmath 010E} }=\dot{{\Greekmath 011A}}_{F}\left[ G\right] . \end{equation*}
definitionThe influence function of ${\Greekmath 011A} :\mathcal{F}^{\ast }\rightarrow \mathbb{R}$ at $F\in \mathcal{F}^{\ast }$ is given by \begin{equation*} {\Greekmath 0120} (y,{\Greekmath 011A} ,F):=\lim_{{\Greekmath 010F} \rightarrow 0+}\frac{{\Greekmath 011A} \left[ \left( 1-{\Greekmath 010F} \right) F+{\Greekmath 010F} \Delta _{y}\right] -{\Greekmath 011A} \left[ F\right] }{ {\Greekmath 010F} }, \end{equation*} where $\Delta _{y}$ is the distribution function that assigns all probability mass to the single point $\left\{ y\right\} ,$ that is, $\Delta _{y}\left( x\right) =1\left\{ x\geq y\right\} .$

To see how we can use the Hadamard differentiability to obtain $\Pi _{{\Greekmath 011A} }$ , we write

equation*[equation* omitted — 271 chars of source]

for

equation[equation omitted — 116 chars of source]

As long as we can show that $\lim_{{\Greekmath 010E} \rightarrow 0}\left\Vert G_{{\Greekmath 010E} }-G\right\Vert _{\infty }=0$ for some $G,$ then, we obtain

equation*[equation* omitted — 436 chars of source]

In the proof of Theorem (ref) below, we show that we can use the influence function to represent $\dot{{\Greekmath 011A}}_{F_{Y}}\left[ G\right] $ as $ \dot{{\Greekmath 011A}}_{F_{Y}}\left[ G\right] =\int_{\mathcal{Y}}{\Greekmath 0120} (y,{\Greekmath 011A} ,F_{Y})dG\left( y\right) $.

Next, we provide sufficient conditions for $\lim_{{\Greekmath 010E} \rightarrow 0}\left\Vert G_{{\Greekmath 010E} }-G\right\Vert _{\infty }=0$. We first formalize two primary assumptions.\footnote{ In Assumptions (ref) and (ref), \textquotedblleft for all $w\in \mathcal{W}$\textquotedblright\ can be replaced by \textquotedblleft for almost all $w\in \mathcal{W}$ \textquotedblright , and the supremum over $w\in \mathcal{W} $ can be replaced by the essential supremum over $w\in \mathcal{W}$.}

assumptionPrimary Assumptions \begin{enumerate}[(a)] • $U_{D}$ is independent of $W$ and is uniformly distributed on $\left[ 0,1\right] .$ • For all $w=(z^{\prime },x^{\prime })^{\prime }\in \mathcal{W}$, $\mathcal{G}(w,0)=z,$ and for sufficiently small ${\Greekmath 010E} $, $ \mathcal{G}(w,{\Greekmath 010E} )\in \mathcal{Z}$, the support of $Z.$ \end{enumerate}

As discussed earlier, Assumption (ref)((ref)) holds if the conditional CDF $F_{V|W}(\cdot |w)$ is a strictly increasing function for all $w\in \mathcal{W}$. Assumption (ref)((ref)) requires that the policy function $\mathcal{G}(w,{\Greekmath 010E} )$ be feasible. This assumption, along with Assumption (ref)( (ref).ii) below, requires that the subvector of $Z$ undergoing nontrivial intervention consists of continuous random variables. The variables in $X$ and the part of $Z$ not subject to intervention do not need to be continuous random variables.

assumptionRegularity Conditions \begin{enumerate}[(a)] • For $d=0,1,$ the conditional distribution of $ (Y(d),U_{D})$ conditional on $W=w\in \mathcal{W}$ is absolutely continuous with conditional density function given by $f_{Y(d),U_{D}|W}(y,u|w).$ • (i) For $d=0,1,$ $u\mapsto f_{Y(d)|U_{D},W}(y|u,w)$ is continuous for all $y\in \mathcal{Y}\left( d\right) $ and all $w\in \mathcal{ W}.$ (ii) For $d=0,1,$ $\sup_{y\in \mathcal{Y}\left( d\right) }\sup_{w\in \mathcal{W}}\sup_{{\Greekmath 010E} \in N_{{\Greekmath 0122} }}f_{Y(d)|U_{D},W}(y|P_{{\Greekmath 010E} }(w),w)<\infty $ where $N_{{\Greekmath 0122} }:=\left\{ {\Greekmath 010E} :\left\vert {\Greekmath 010E} \right\vert \leq {\Greekmath 0122} \right\} $ for some ${\Greekmath 0122} >0.$ • (i) For all $w\in \mathcal{W}$, $P(w)\in (0,1).$ (ii) For all $w\in \mathcal{W}$, the map ${\Greekmath 010E} \mapsto P_{{\Greekmath 010E} }(w)$ is continuously differentiable on $N_{{\Greekmath 0122} }.$ (iii) $\sup_{w\in \mathcal{W}}\sup_{{\Greekmath 010E} \in N_{{\Greekmath 0122} }}\left\vert \frac{\partial P_{{\Greekmath 010E} }(w)}{\partial {\Greekmath 010E} }\right\vert <\infty .$ \end{enumerate}
assumptionDomination Conditions For $d=0,1,$ \begin{equation*} \int_{\mathcal{Y}(d)}\sup_{{\Greekmath 010E} \in N_{{\Greekmath 0122} }}f_{Y(d)|D_{{\Greekmath 010E} }}(y|d)dy<\infty \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ and }\int_{\mathcal{Y}(d)}\sup_{{\Greekmath 010E} \in N_{{\Greekmath 0122} }}\left\vert \frac{\partial f_{Y(d)|D_{{\Greekmath 010E} }}(y|d)}{ \partial {\Greekmath 010E} }\right\vert dy<\infty . \end{equation*}

Under Assumption (ref)((ref), (ref)), we have $E\left[ P_{{\Greekmath 010E} }(W)\right] =E\left[ P(W)\right] +E\left[ \dot{P} \left( W\right) \right] {\Greekmath 010E} +o({\Greekmath 010E} ),$ as ${\Greekmath 010E} \rightarrow 0$, where

equation*[equation* omitted — 168 chars of source]

So, under the new policy regime $\mathcal{G}(\cdot ,{\Greekmath 010E} )$, the participation rate in the treatment is increased by {approximately }$E\left[ \dot{P}\left( W\right) \right] {\Greekmath 010E} .$ Locally at ${\Greekmath 010E} =0$, the policy function $\mathcal{G}(\cdot ,{\Greekmath 010E} )$ affects the participation rate via $ \dot{P}\left( \cdot \right) $, which is the rate of change in the propensity score at ${\Greekmath 010E} =0$. The function $\dot{P}\left( \cdot \right) $ depends on the policy function $\mathcal{G}(\cdot ,{\Greekmath 010E} )$ used, and a more cumbersome notation for $\dot{P}\left( \cdot \right) $ is $\dot{P}_{\mathcal{ G}}\left( \cdot \right) .$ For notational simplicity, we suppress such dependence. Section (ref) provides further analysis on $\dot{P} \left( \cdot \right) $.

theoremLet Assumptions (ref)--(ref) hold. Assume further that ${\Greekmath 011A} :\mathcal{F}^{\ast }\rightarrow \mathbb{R}$ is Hadamard differentiable with influence function $ {\Greekmath 0120} $, and $E\left[ \dot{P}\left( W\right) \right] \neq 0$. Then, \begin{equation*} \Pi _{{\Greekmath 011A} }=\int_{\mathcal{W}}\mathrm{MTE}_{{\Greekmath 011A} }\left( P(w),w\right) \mathcal{\dot{P}}\left( w\right) dF_{W}(w), \end{equation*} where \begin{equation} \mathrm{MTE}_{{\Greekmath 011A} }\left( u_{D},w\right) :=E\left[ {\Greekmath 0120} (Y\left( 1\right) ,{\Greekmath 011A} ,F_{Y})-{\Greekmath 0120} (Y\left( 0\right) ,{\Greekmath 011A} ,F_{Y})|U_{D}=u_{D},W=w\right] , \end{equation} is the unconditional marginal treatment effect for the ${\Greekmath 011A} $ functional, and \begin{equation*} \mathcal{\dot{P}}\left( w\right) =\frac{\dot{P}\left( w\right) }{E\left[ \dot{P}\left( W\right) \right] }. \end{equation*}

Theorem (ref) reveals that the general unconditional policy effect $\Pi _{{\Greekmath 011A} }$ is composed of two key elements: the unconditional $ \mathrm{MTE}_{{\Greekmath 011A} }\left( u_{D},w\right) $ evaluated at $u_{D}=P(w)$ and the weighting function $\mathcal{\dot{P}}\left( w\right) .$ To understand the first element, consider the group of individuals with the same value $w$ of $W.$ Within this group, those for whom $u_{D}=P\left( w\right) $ are indifferent between participating and not participating. A small incentive will induce a change in the treatment status for and only for this subgroup of individuals. It is the change in their treatment status, and hence the change in the composition of $Y(1)$ and $Y(0)$ in the observed outcome $Y,$ that changes its unconditional characteristics, such as the quantiles. We refer to the individuals for whom\ $u_{D}=P\left( w\right) $ as the marginal subpopulation. As for the second element, we defer the discussion to Section (ref).

Unconditional quantile effect

Throughout the rest of this paper, we consider the case that ${\Greekmath 011A} $ is a quantile functional at the quantile level ${\Greekmath 011C} \in \left( 0,1\right) ,$ that is, ${\Greekmath 011A} _{{\Greekmath 011C} }[F_{Y}]=F_{Y}^{-1}({\Greekmath 011C} ):=\inf_{y}\left\{ y\in \mathcal{Y}:F_{Y}(y)\geq {\Greekmath 011C} \right\} .$ Here we have added a subscript $ {\Greekmath 011C} $ to ${\Greekmath 011A} $ to signify the quantile level under consideration. We are interested in how an improvement in the treatment take-up rate affects the $ {\Greekmath 011C} $-quantile of the (unconditional) outcome distribution.

definitionUnconditional Quantile Effect • The unconditional quantile effect (UQE) is defined as \begin{equation*} \Pi _{{\Greekmath 011C} }:=\lim_{{\Greekmath 010E} \rightarrow 0}\frac{{\Greekmath 011A} _{{\Greekmath 011C} }[F_{Y_{{\Greekmath 010E} }}]-{\Greekmath 011A} _{{\Greekmath 011C} }[F_{Y}]}{E[D_{{\Greekmath 010E} }]-E[D]} \end{equation*} whenever this limit exists.

Let $y_{{\Greekmath 011C} }\equiv {\Greekmath 011A} _{{\Greekmath 011C} }[F_{Y}]$ be the ${\Greekmath 011C} $-quantile of $Y.$ If $f_{Y}(y_{{\Greekmath 011C} })>0,$ then, under Assumption (ref)( (ref)), the influence function of the ${\Greekmath 011C} $-quantile functional is\footnote{ Strictly speaking, ${\Greekmath 0120} \left( y,{\Greekmath 011A} _{{\Greekmath 011C} },F_{Y}\right) =0$ for $ y=y_{{\Greekmath 011C} }$. However, redefining ${\Greekmath 0120} \left( y,{\Greekmath 011A} _{{\Greekmath 011C} },F_{Y}\right) $ at one point has no consequence on our results, as $F_{Y}\left( \cdot \right) $ is absolutely continuous under Assumption (ref) ((ref)).}

equation*[equation* omitted — 240 chars of source]

Plugging this influence function into ((ref)), we obtain the unconditional marginal treatment effect for the ${\Greekmath 011C} $-quantile.

definitionThe unconditional marginal treatment effect for the ${\Greekmath 011C} $ -quantile is \begin{equation*} \mathrm{MTE}_{{\Greekmath 011C} }\left( u,w\right) =\frac{1}{f_{Y}\left( y_{{\Greekmath 011C} }\right) }E\left[ \mathds{1}\left\{ Y(0)\leq y_{{\Greekmath 011C} }\right\} -\mathds{1}\left\{ Y(1)\leq y_{{\Greekmath 011C} }\right\} \mid U_{D}=u,W=w\right] . \end{equation*}

The $\mathrm{MTE}_{{\Greekmath 011C} }$ defined above is a basic building block for the unconditional quantile effect. It is different from the quantile analogue of the marginal treatment effect of Carneiro2009 and ping2014, which is defined as $F_{Y(1)|U_{D},W}^{-1}({\Greekmath 011C} |u,w)-F_{Y(0)|U_{D},W}^{-1}({\Greekmath 011C} |u,w)$. An unconditional quantile effect can not be represented as an integrated version of the latter. The next corollary follows from applying Theorem (ref) to ${\Greekmath 011A}_{\Greekmath 011C}$.

corollaryLet Assumptions (ref)--(ref) hold. Assume further that $f_{Y}(y_{{\Greekmath 011C} })>0$. Then, \begin{eqnarray} \Pi _{{\Greekmath 011C} } &=&\int_{\mathcal{W}}\mathrm{MTE}_{{\Greekmath 011C} }\left( P(w),w\right) \mathcal{\dot{P}}\left( w\right) dF_{W}(w) \notag \\ &=&\frac{1}{f_{Y}(y_{{\Greekmath 011C} })}\int_{\mathcal{W}}E\left[ \mathds{1}\left\{ Y(0)\leq y_{{\Greekmath 011C} }\right\} |U_{D}=P\left( w\right) ,W=w\right] \mathcal{\dot{ P}}\left( w\right) dF_{W}\left( w\right) \notag \\ &-&\frac{1}{f_{Y}(y_{{\Greekmath 011C} })}\int_{\mathcal{W}}E\left[ \mathds{1}\left\{ Y(1)\leq y_{{\Greekmath 011C} }\right\} |U_{D}=P\left( w\right) ,W=w\right] \mathcal{\dot{ P}}\left( w\right) dF_{W}\left( w\right) . \end{eqnarray}

Understanding $\mathrm{MTE}_{\protect{\Greekmath 011C} }\left( u,w\right)$

To understand $\mathrm{MTE}_{{\Greekmath 011C} }\left( u,w\right) ,$ we define $\Delta (y_{{\Greekmath 011C} }):=\left( \mathds{1}\left\{ Y(0)\leq y_{{\Greekmath 011C} }\right\} -\mathds{1} \left\{ Y(1)\leq y_{{\Greekmath 011C} }\right\} \right) /f_{Y}(y_{{\Greekmath 011C} }),$ which underlies the above definition of $\mathrm{MTE}_{{\Greekmath 011C} }\left( u,w\right) $. The random variable $\Delta (y_{{\Greekmath 011C} })$ can take three values:

equation*[equation* omitted — 1,009 chars of source]

For a given individual, $\Delta (y_{{\Greekmath 011C} })=1/f_{Y}(y_{{\Greekmath 011C} })$ when the treatment induces the individual to \textquotedblleft cross\textquotedblright\ the ${\Greekmath 011C} $-quantile $y_{{\Greekmath 011C} }$ of $Y$ from below, and $\Delta (y_{{\Greekmath 011C} })=-1/f_{Y}(y_{{\Greekmath 011C} })$ when the treatment induces the individual to \textquotedblleft cross\textquotedblright\ the ${\Greekmath 011C} $ -quantile $y_{{\Greekmath 011C} }$ of $Y$ from above. In the first case, the individual benefits from the treatment, while in the second case, the treatment harms her. The intermediate case, $\Delta (y_{{\Greekmath 011C} })=0,$ occurs when the treatment induces no quantile crossing of any type. Thus, the unconditional expected value $f_{Y}(y_{{\Greekmath 011C} })\times E[\Delta (y_{{\Greekmath 011C} })]$ equals the difference between the proportion of individuals who benefit from the treatment and the proportion of individuals who are harmed by it. For the UQE, whether the treatment is beneficial or harmful is measured in terms of quantile crossing. Among the individuals with characteristics $U_{D}=u$ and $ W=w$, $\mathrm{MTE}_{{\Greekmath 011C} }\left( u,w\right) $ is then equal to the rescaled (by $1/f_{Y}(y_{{\Greekmath 011C} })$) difference between the proportion of individuals who benefit from the treatment and the proportion of individuals who are harmed by it. Thus, $\mathrm{MTE}_{{\Greekmath 011C} }\left( u,w\right) $ is positive if the treatment leads to a greater number of individuals improving their outcome above the threshold $y_{{\Greekmath 011C} }$, compared to those whose outcome falls below $y_{{\Greekmath 011C} }$. Conversely, $\mathrm{MTE}_{{\Greekmath 011C} }\left( u,w\right) $ is negative if the treatment leads to a greater number of individuals experiencing a decline in their outcome, falling below $y_{{\Greekmath 011C} }$, compared to those experiencing an increase above $y_{{\Greekmath 011C} }.$

Understanding $\mathcal{\dot{P}}$

To understand the weighting function $\mathcal{\dot{P}}\left( w\right) ,$ consider the case ${\Greekmath 010E} >0$, $P_{{\Greekmath 010E} }(w)\geq P(w)$ for all $w\in W$, and $F_{W}\left( \cdot \right) $ is absolutely continuous with density $ f_{W}\left( \cdot \right) .$ Let ${\Greekmath 010F} $ be a small positive number. Then, $f_{W}\left( w\right) {\Greekmath 010F} $ measures the proportion of individuals for whom $W$ is in $\left[ w-{\Greekmath 010F} /2,w+{\Greekmath 010F} /2\right] .$ Note that for $W\in \left[ w-{\Greekmath 010F} /2,w+{\Greekmath 010F} /2\right] ,$ the propensity scores under $D$ and $D_{{\Greekmath 010E} }$ are approximately $P(w)$ and $ P_{{\Greekmath 010E} }\left( w\right) $. The proportion of the individuals for whom $ W\in \left[ w-{\Greekmath 010F} /2,w+{\Greekmath 010F} /2\right] $ and who have switched their treatment status from $0$ to $1$ is then equal to $\left[ P_{{\Greekmath 010E} }(w)-P(w) \right] f_{W}\left( w\right) \cdot {\Greekmath 010F} .$ Scaling this by $E[D_{{\Greekmath 010E} }]-E[D]$, which is the overall proportion of the individuals who have switched the treatment status, we obtain

equation*[equation* omitted — 147 chars of source]

Thus, we can regard $\left[ P_{{\Greekmath 010E} }(w)-P(w)\right] f_{W}\left( w\right) / \left[ E[D_{{\Greekmath 010E} }]-E[D]\right] $ as the density function of $W$ among those who have switched their treatment status from $0$ to $1$ as a result of the policy intervention. On the one hand, when ${\Greekmath 010E} \rightarrow 0,$ the set of individuals who change their treatment status are the individuals on the margin, namely, those for whom $u_{D}=P(w).$ On the other hand, when $ {\Greekmath 010E} \rightarrow 0,$ the density function approaches $\mathcal{\dot{P}} \left( w\right) f_{W}\left( w\right) .$ So $\mathcal{\dot{P}}\left( w\right) f_{W}\left( w\right) $ is the probability density function (with respect to the Lebesgue measure) of the distribution of $W$ over the marginal subpopulation. Also, note that by construction,\ $\int_{\mathcal{W}}\mathcal{ \dot{P}}\left( w\right) dF(w)=1$, and thus $\mathcal{\dot{P}}\left( w\right) $ can be interpreted as the density of the distribution of $W$ for the marginal subpopulation with respect to the distribution of $W$ for the entire population.

It is now clear that the general unconditional policy effect is equal to the average of the unconditional $\mathrm{MTE}$ over the marginal subpopulation. Such an interpretation is still valid even if $\mathcal{\dot{P}}\left( w\right) $ is not positive for all $w\in \mathcal{W}$. In this case, we only need to view the distribution with density $\mathcal{\dot{P}}\left( w\right) $ (with respect to the distribution of $W$ for the entire population) as a signed measure.

Understanding the Unconditional MTE

To gain a deeper understanding of the unconditional MTE, both in the general case and the special quantile case, we will explore another perspective here. Note that this subsection will only provide heuristics, as the formal developments have already been covered in the previous subsections.

heckman_prte, Heckman2005 focus on the mean functional and consider the policy-relevant treatment effect defined as

equation[equation omitted — 203 chars of source]

Taking the limit ${\Greekmath 010E} \rightarrow 0$ yields the marginal policy-relevant treatment effect (MPRTE) of Carneiro2010: $ \mathrm{MPRTE}=\lim_{{\Greekmath 010E} \rightarrow 0}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{\textrm{PRTE}}_{{\Greekmath 010E} }.$ Carneiro2010, Carneiro2011 show that $\mathrm{MPRTE}$ can be represented in terms of the conventional marginal treatment effect defined by $\mathrm{MTE}(u):=E\left[ Y(1)-Y(0)|U_{D}=u\right] $. These results are applicable to the mean functional only.

For a general functional ${\Greekmath 011A} $ that is Hadamard differentiable, we have

equation*[equation* omitted — 337 chars of source]

where $\dot{{\Greekmath 011A}}_{F_{Y}}$ is the derivative of ${\Greekmath 011A} $ at $F_{Y}$. One way to represent the limit on the right-hand side is to replace $E\left( Y_{{\Greekmath 010E} }\right) $ and $E\left( Y\right) $ in ((ref)) by $E[\mathds 1\left\{ Y_{{\Greekmath 010E} }\leq y\right\} ]$ and $E[\mathds1\left\{ Y\leq y\right\} ],$ respectively. So, for a given $y\in \mathcal{Y}$,

equation[equation omitted — 252 chars of source]

In the spirit of heckman_prte, Heckman2005, we may then regard the above as a policy-relevant treatment effect: it is the treatment effect of the policy on the percentage of individuals whose value of $Y$ is less than $ y.$ The effect is tied to a particular value $y$, and we obtain a continuum of policy-relevant effects indexed by $y\in \mathcal{Y}$ if $y$ is allowed to vary over $\mathcal{Y}$. The limit

equation*[equation* omitted — 158 chars of source]

can then be regarded as a continuum of marginal policy-relevant treatment effects indexed by $y\in \mathcal{Y}$. By the results of Carneiro2009 , for each $y\in \mathcal{Y}$, the marginal policy-relevant treatment effect can be represented as a weighted integral of the following policy-relevant \textquotedblleft distributional\textquotedblright \ MTE:

eqnarray*[eqnarray* omitted — 250 chars of source]

The composition $\dot{{\Greekmath 011A}}_{F_{Y}}\circ \mathrm{MTE}_{d}$, defined as $ \int_{\mathcal{Y}}{\Greekmath 0120} (y,{\Greekmath 011A} ,F_{Y})\mathrm{MTE}_{d}(u,w;dy),$ is then

eqnarray*[eqnarray* omitted — 633 chars of source]

This shows that $\dot{{\Greekmath 011A}}_{F_{Y}}\circ \mathrm{MTE}_{d}$ is exactly the unconditional MTE defined in ((ref)). The unconditional MTE is therefore a composition of the underlying influence function with the policy-relevant distributional MTE.

UQR with a Threshold-crossing Model

In this section, we study whether the UQR proposed by Firpo2009 can provide a consistent estimator of UQE. When it is inconsistent, we investigate the sources of asymptotic bias in the UQR estimator and show that it is asymptotically biased, even when $D$ is exogenous. As a result, the UQR may not be suitable for estimating unconditional policy effects when the treatment variable follows a binary threshold-crossing model. In Section (ref), we will present a consistent estimator of the unconditional policy effect for such a model.

UQR with a Binary Regressor

We provide a quick review of the UQR. As before, let $y_{{\Greekmath 011C} }$ be the $ {\Greekmath 011C} $-quantile of $Y,$ and let $y_{{\Greekmath 011C} ,{\Greekmath 010E} }$ be the ${\Greekmath 011C} $-quantile of $Y_{{\Greekmath 010E} }$. That is, $\Pr [Y\leq y_{{\Greekmath 011C} }]=\Pr [Y_{{\Greekmath 010E} }\leq y_{{\Greekmath 011C} ,{\Greekmath 010E} }]={\Greekmath 011C} .$ By definition, we have

equation*[equation* omitted — 298 chars of source]

When $W$ is not present, Corollary 3 in the working paper Firpo2007 makes the following assumption to achieve identification:

equation[equation omitted — 293 chars of source]

We refer to this assumption as distributional invariance, and it readily identifies the counterfactual distribution:

equation*[equation* omitted — 213 chars of source]

Under some mild conditions, we can follow Firpo2007 to show that

eqnarray[eqnarray omitted — 393 chars of source]

Hence, under the distributional invariance assumption, the UQE can be consistently estimated by regressing $1\left\{ Y\geq y_{{\Greekmath 011C} }\right\} /f_{Y}(y_{{\Greekmath 011C} })$ on a constant and $D.$ Such a regression with no additional regressor $W$ is a special case of more general unconditional quantile regressions.

Asymptotic Bias of the UQR Estimator

The distributional invariance assumption given in equation (ref) is crucial for achieving\ the identification result in (ref). It states that the conditional distribution of the outcome variable given the treatment status remains the same across the two policy regimes. If treatments are randomly assigned under both policy regimes (e.g., $D_{{\Greekmath 010E} }=\mathds{1}\left\{ U_{D}\leq P_{{\Greekmath 010E} }(W)\right\} $ and $\left( U_{D},W\right) $ is independent of $\left( U_{0},U_{1}\right) ),$ then $\left( U_{0},U_{1}\right) $ is clearly independent of $D_{{\Greekmath 010E} }$. In this case, both $\Pr \left[ Y_{{\Greekmath 010E} }\leq y|D_{{\Greekmath 010E} }=d\right] $ and $\Pr \left[ Y\leq y|D=d\right] $ are equal to $ \Pr \left[ Y\left( d\right) \leq y\right] ,$ and the distributional invariance assumption is satisfied. However, when $D_{{\Greekmath 010E} }$ is allowed to be correlated with $U=\left( U_{0},U_{1}\right) ^{\prime }$, the distributional invariance assumption does not hold in general. For example, when $d=1,$

equation*[equation* omitted — 260 chars of source]

and $\Pr [Y\leq y|D=1]=\Pr \left[ r_{1}\left( X,U_{1}\right) \leq y|U_{D}\leq P\left( W\right) \right] .$ These two conditional probabilities are different under the general dependence of $(W,U,U_{D}).$

To allow for the endogeneity of $D$, we have dropped the distributional invariance assumption and assumed a threshold-crossing model as in ((ref)). The next corollary decomposes the unconditional quantile effect given in Corollary (ref) into two components. The decomposition reveals that the UQR estimator of Firpo2009 is asymptotically biased under a wide range of conditions, including when $D$ is exogenous.

corollaryLet Assumptions (ref)--(ref) hold. Assume further that $f_{Y}(y_{{\Greekmath 011C} })>0$. Then \footnote{ We use \textquotedblleft A\textquotedblright\ to denote the A pparent component and use \textquotedblleft B\textquotedblright\ to denote the Bias component.} \begin{equation*} \Pi _{{\Greekmath 011C} }=A_{{\Greekmath 011C} }-B_{{\Greekmath 011C} }, \end{equation*} where \begin{eqnarray} A_{{\Greekmath 011C} } &=&\frac{1}{f_{Y}(y_{{\Greekmath 011C} })}\int_{\mathcal{W}}E\left[ \mathds{1} \left\{ Y\leq y_{{\Greekmath 011C} }\right\} |D=0,W=w\right] dF_{W}\left( w\right) \notag \\ &-&\frac{1}{f_{Y}(y_{{\Greekmath 011C} })}\int_{\mathcal{W}}E\left[ \mathds{1}\left\{ Y\leq y_{{\Greekmath 011C} }\right\} |D=1,W=w\right] dF_{W}\left( w\right) , \end{eqnarray} and $B_{{\Greekmath 011C} }=B_{1{\Greekmath 011C} }+B_{2{\Greekmath 011C} }$, for \begin{eqnarray*} B_{1{\Greekmath 011C} } &=&\frac{1}{f_{Y}(y_{{\Greekmath 011C} })}\int_{\mathcal{W}}\left[ F_{Y|D,W}\left( y_{{\Greekmath 011C} }|1,w\right) -F_{Y|D,W}\left( y_{{\Greekmath 011C} }|0,w\right) \right] \mathcal{\dot{P}}\left( w\right) dF_{W}\left( w\right) \\ &-&\frac{1}{f_{Y}(y_{{\Greekmath 011C} })}\int_{\mathcal{W}}\left[ F_{Y|D,W}\left( y_{{\Greekmath 011C} }|1,w\right) -F_{Y|D,W}\left( y_{{\Greekmath 011C} }|0,w\right) \right] dF_{W}\left( w\right) \end{eqnarray*} and \begin{eqnarray*} B_{2{\Greekmath 011C} } &=&\frac{1}{f_{Y}(y_{{\Greekmath 011C} })}\int_{\mathcal{W}}\left[ F_{Y\left( 0\right) |D,W}\left( y_{{\Greekmath 011C} }|0,w\right) -F_{Y\left( 0\right) |U_{D},W}\left( y_{{\Greekmath 011C} }|P\left( w\right) ,w\right) \right] \mathcal{\dot{P} }\left( w\right) dF_{W}\left( w\right) \\ &-&\frac{1}{f_{Y}(y_{{\Greekmath 011C} })}\int_{\mathcal{W}}\left[ F_{Y\left( 1\right) |D,W}\left( y_{{\Greekmath 011C} }|1,w\right) -F_{Y\left( 1\right) |U_{D},W}\left( y_{{\Greekmath 011C} }|P\left( w\right) ,w\right) \right] \mathcal{\dot{P}}\left( w\right) dF_{W}\left( w\right) . \\ && \end{eqnarray*}

To facilitate understanding of Corollary (ref), we define and organize the average influence functions (AIF) in a table:

equation*[equation* omitted — 549 chars of source]

where ${\Greekmath 0120} _{{\Greekmath 011C} }\left( \cdot \right) $ is short for ${\Greekmath 0120} (\cdot ,{\Greekmath 011A} _{{\Greekmath 011C} },F_{Y})$, the influence function of the quantile functional. In the above, $E_{w}\left[ \cdot \right] $ stands for the conditional mean operator given $W=w.$ For example, $E_{w}\left[ {\Greekmath 0120} _{{\Greekmath 011C} }\left( Y\left( 0\right) \right) |D=0\right] $ stands for $E\left[ {\Greekmath 0120} _{{\Greekmath 011C} }\left( Y\left( 0\right) \right) |D=0,W=w\right] .$ Let

eqnarray*[eqnarray* omitted — 507 chars of source]

The unconditional quantile effect $\Pi _{{\Greekmath 011C} }$ is the average of the difference ${\Greekmath 0120} _{\Delta ,U_{D}}\left( w\right) $ with respect to the distribution of $W$ over the marginal subpopulation. The average apparent effect $A_{{\Greekmath 011C} }$ is the average of the difference ${\Greekmath 0120} _{\Delta ,D}\left( w\right) $ with respect to the distribution of $W$ over the whole population distribution. It is also equal to the limit of the UQR estimator of Firpo2009, where the endogeneity of the treatment selection is ignored.\footnote{ To see why this is the case, we note that, in its simplest form, the UQR involves regressing the \textquotedblleft influence function\textquotedblright\ $\hat{f}_{Y}(\hat{y}_{{\Greekmath 011C} })^{-1}\left( {\Greekmath 011C} - \mathds{1}\left\{ Y_{i}\leq \hat{y}_{{\Greekmath 011C} }\right\} \right) $ on $D_{i}$ and $W_{i}$ by OLS and using the estimated coefficient on $D_{i}$ as the estimator of the unconditional quantile effect. Here, $\hat{f}_{Y}(y_{{\Greekmath 011C} }) $ is a consistent estimator of $f_{Y}(y_{{\Greekmath 011C} })$ and $\hat{y}_{{\Greekmath 011C} }$ is a consistent estimator of $y_{{\Greekmath 011C} }.$ It is now easy to see that the UQR estimator converges in probability to $A_{{\Greekmath 011C} }$ if the conditional expectations in ((ref)) are linear in $W.$} Note that if our model contains no covariate $W,$ then

eqnarray[eqnarray omitted — 410 chars of source]

This is identical to the unconditional quantile effect given in ((ref)).

The discrepancy between $\Pi _{{\Greekmath 011C} }$ and $A_{{\Greekmath 011C} }$ gives rise to the asymptotic bias $B_{{\Greekmath 011C} }$ of the UQR estimator:

eqnarray[eqnarray omitted — 671 chars of source]

It is easy to see that $B_{1{\Greekmath 011C} }$ and $B_{2{\Greekmath 011C} }$ given above are identical to those given in Corollary (ref).

The decomposition in Equation ((ref)) traces the asymptotic bias back to two sources. The first one, $B_{1{\Greekmath 011C} }$, captures\ the heterogeneity of the averaged apparent effects averaged over two different subpopulations. For every $w,$ ${\Greekmath 0120} _{\Delta ,D}\left( w\right) $ is the average effect of $D$ on $\left[ {\Greekmath 011C} -\mathds{1}\left\{ Y\leq y_{{\Greekmath 011C} }\right\} \right] /f_{Y}\left( y_{{\Greekmath 011C} }\right) $ for the individuals with $ W=w.$ These effects are averaged over two different distributions of $W$: the distribution of $W$ for the marginal subpopulation (i.e., $\mathcal{\dot{ P}}\left( w\right) f_{W}\left( w\right) )$ and the distribution of $W$ for the whole population (i.e., $f_{W}\left( w\right) )$. $B_{1{\Greekmath 011C} }$ is equal to the difference between these two average effects.\ If the effect $ {\Greekmath 0120} _{\Delta ,D}\left( w\right) $ does not depend on $w$, then $B_{1{\Greekmath 011C} }=0 $. If $\mathcal{\dot{P}}\left( \cdot \right) $ is a constant function that always equals 1, then the distribution of $W$ over the whole population is the same as that over the marginal subpopulation, and hence $B_{1{\Greekmath 011C} }=0$ as well. For $B_{1{\Greekmath 011C} }\neq 0,$ it is necessary that there is an effect heterogeneity (i.e., ${\Greekmath 0120} _{\Delta ,D}\left( w\right) $ depends on $w$) and a distributional heterogeneity (i.e., $\mathcal{\dot{P}}\left( \cdot \right) $ does not always equal $1$, and as a result, the distribution of $W$ over the marginal subpopulation is different from that over the whole population). To highlight the necessary conditions for a nonzero $B_{1{\Greekmath 011C} }, $ we refer to $B_{1{\Greekmath 011C} }$ as the marginal heterogeneity bias.

The second bias component, $B_{2{\Greekmath 011C} },$ embodies the second source of the bias and has a difference-in-differences interpretation.\ Each of ${\Greekmath 0120} _{\Delta ,D}\left( \cdot \right) $ and ${\Greekmath 0120} _{\Delta ,U_{D}}\left( \cdot \right) $ is the difference in the average influence functions associated with the counterfactual outcomes $Y\left( 1\right) $ and $Y\left( 0\right) .$ However, ${\Greekmath 0120} _{\Delta ,D}\left( \cdot \right) $ is the difference over the two subpopulations who actually choose $D=1$ and $D=0$, while ${\Greekmath 0120} _{\Delta ,U_{D}}\left( \cdot \right) $ is the difference over the marginal subpopulation. So ${\Greekmath 0120} _{\Delta ,D}\left( \cdot \right) -{\Greekmath 0120} _{\Delta ,U_{D}}\left( \cdot \right) $ is a difference in differences. $B_{2{\Greekmath 011C} }$ is simply the average of this difference in differences with respect to the distribution of $W$ over the marginal subpopulation. This term arises because the change in the distributions of $Y$ for those with $D=1$ and those with $D=0$ is different from that for those whose $U_{D}$ is just above $P\left( w\right) $ and those whose $U_{D}$ is just below $P\left( w\right) $. Thus, we can label $B_{2{\Greekmath 011C} }$ as a marginal selection bias.

If ${\Greekmath 0120} _{\Delta ,D}\left( w\right) ={\Greekmath 0120} _{\Delta ,U_{D}}\left( w\right) $ for almost all $w\in \mathcal{W},$ then $B_{2{\Greekmath 011C} }=0.$ The condition ${\Greekmath 0120} _{\Delta ,D}\left( w\right) ={\Greekmath 0120} _{\Delta ,U_{D}}\left( w\right) $ is the same as

equation*[equation* omitted — 397 chars of source]

Equivalently,

eqnarray*[eqnarray* omitted — 448 chars of source]

The condition resembles the parallel-paths assumption or the constant-bias assumption in a difference-in-differences analysis. If $U_{D}$ is independent of $\left( U_{0},U_{1}\right) $ given $W,$ then this condition holds, and $B_{2{\Greekmath 011C} }=0.$

In general, when $U_{D}$ is not independent of $\left( U_{0},U_{1}\right) $ given $W$, and $W$ enters the selection equation, we have $B_{1{\Greekmath 011C} }\neq 0$ and $B_{2{\Greekmath 011C} }\neq 0,$ hence $\Pi _{{\Greekmath 011C} }\neq A_{{\Greekmath 011C} }.$ If $\mathcal{ \dot{P}}\left( w\right) $ is not identified, then $B_{1{\Greekmath 011C} }$ is not identified. In general, $B_{2{\Greekmath 011C} }$ is not identified without additional assumptions. Therefore, in the absence of additional assumptions, the asymptotic bias can not be eliminated, and $\Pi _{{\Greekmath 011C} }$ is not identified.

It is not surprising that in the presence of endogeneity, the UQR estimator is asymptotically biased. The virtue of Corollary (ref) is that it provides a closed-form characterization and clear interpretations of the asymptotic bias. To the best of our knowledge, this bias formula is new in the literature. If point identification can not be achieved, then the bias formula can be used in a bound analysis or sensitivity analysis. From a broad perspective, the asymptotic bias $B_{{\Greekmath 011C} }$ is the unconditional quantile counterpart of the endogenous bias of the OLS estimator in a linear regression framework.\footnote{ The bias decomposition is not unique. Corollary (ref) gives only one possibility. We can also write

equation[equation omitted — 452 chars of source]

The interpretations of $\tilde{B}_{1{\Greekmath 011C} }$ and $\tilde{B}_{2{\Greekmath 011C} }$ are similar to those of $B_{1{\Greekmath 011C} }$ and $B_{2{\Greekmath 011C} }$ with obvious and minor modifications. In this case, it may be more revealing to call $\tilde{B} _{1{\Greekmath 011C} }$ the marginal heterogeneity bias, because a necessary condition for a nonzero $\tilde{B}_{1{\Greekmath 011C} }$ is that there is an effect heterogeneity among the marginal subpopulation.}

Unconditional Quantile Effect under Instrumental Intervention

In this section, we consider two types of interventions: one on a valid instrumental variable and the other on an invalid instrumental variable. We show how we may recover the effect of the latter by using a valid instrument.

Instrumental Intervention with a valid instrument

To introduce the instrumental intervention, we partition $Z$ into two parts and write $Z=\left( Z_{1},Z_{-1}\right) $\ where $Z_{1}$\ is a univariate continuous policy variable and $Z_{-1}$ consists of other variables. If there is only one variable in $Z,$ then $Z=Z_{1},$ and $Z_{-1}$ is not present. We consider the intervention

equation[equation omitted — 176 chars of source]

where $g\left( \cdot \right) $ is a measurable function and $s\left( {\Greekmath 010E} \right) $ is a smooth function satisfying $s\left( 0\right) =0$. That is, we intervene to change the first component of $Z.$ Note that while $s({\Greekmath 010E} )$ is the same for all individuals, $g\left( W\right) $ depends on the value of $W,$ and hence it is individual-specific. Thus, we allow the intervention to be heterogeneous. In empirical applications, $Z$ may consist of a few variables, and $Z_{1}$ is the continuous target variable that we attempt to change. There are no continuity requirements for the other elements of $Z.$

To present our identification assumptions, we write $W=(W_{1},W_{-1})$ where $W_{1}$ is the first element of $W$ and $W_{-1}$ consists of other elements of $W.$ Since $W=(Z,X),$ the first element of $W$ is also the first element of $Z$, and hence $W_{1}=Z_{1}$. By definition, we have $W_{-1}=\left( Z_{-1},X\right) .$ We maintain the following assumptions, which are similar to the corresponding assumptions in Heckman1999, Heckman2001, Heckman2005.

assumptionRelevance and Exogeneity \begin{enumerate}[(a)] • Conditional on $W_{-1}=\left( Z_{-1},X\right) ,$ ${\Greekmath 0116} (Z,X)$ is a non-degenerate random variable. • Conditional on $W_{-1}=\left( Z_{-1},X\right) ,$ $ Z_{1}$ is independent of $(U_{0},U_{1},V).$ \end{enumerate}

Assumption (ref)((ref)) is a relevance assumption: for any given level of $W_{-1}$, $Z_{1}$ can induce some variation in $D$. Assumption (ref)((ref)) is a conditional exogeneity assumption: for any given level of $W_{-1}$, $Z_{1}$ is independent of the unobservables. These two assumptions are essentially the conditions for $Z_{1}$ to be a valid instrumental variable, hence we will refer to $Z_{1}$ as the instrumental variable, and the intervention in (ref) as the instrumental intervention.\footnote{ If all variables in $Z$ satisfy the conditional exogeneity assumption, we may replace Assumption (ref) by the following: (a) Conditional on $X,$ ${\Greekmath 0116} \left( Z,X\right) $\ is a non-degenerate random variable, and (b) Conditional on $X,$ $Z$ \ is independent of $(U_{0},U_{1},V)$. With minor modifications, our results will remain valid. More specifically, in the statement of each result, we only need to replace \textquotedblleft conditioning on $W_{-1}$\textquotedblright\ by \textquotedblleft conditioning on $X$\textquotedblright . The working paper sun2021 is based on this alternative assumption. Here we will work with Assumption (ref), which appears to be more plausible.}

Let $w:=\left( w_{1},w_{-1}\right) =(z_{1},w_{-1}).$ Under Assumption (ref)((ref)), the unconditional MTE for the quantile functional ${\Greekmath 011A} _{{\Greekmath 011C} }$ becomes

eqnarray[eqnarray omitted — 599 chars of source]

where in the second line above, conditioning on $Z_{1}$ is not necessary and has been dropped.

Next, we present a lemma that characterizes the weighting function for the instrumental intervention in (ref).

lemmaAssume that (i) for almost all $w\in \mathcal{W}$, the conditional distribution of$\ V$ conditional on $W=w$ is absolutely continuous conditional density $f_{V|W}\left( v|w\right) $;\ (ii) ${\Greekmath 0116} \left( w\right) $\ is differentiable in $z_{1}$\ for almost all $w\in \mathcal{W}$ with derivative ${\Greekmath 0116} _{z_{1}}^{\prime }\left( \cdot \right) $ such that $E\left[ f_{V|W}\left( {\Greekmath 0116} (W)|W\right) {\Greekmath 0116} _{z_{1}}^{\prime }\left( W\right) g\left( W\right) \right] $ is well defined and is not equal to zero; (iii) $s\left( {\Greekmath 010E} \right) $\ is a differentiable function in a neighborhood of zero and $\left. \partial s\left( {\Greekmath 010E} \right) /\partial {\Greekmath 010E} \right\vert _{{\Greekmath 010E} =0}\neq 0$. Then \begin{equation} \mathcal{\dot{P}}\left( w\right) =\frac{f_{V|W}\left( {\Greekmath 0116} (w)|w\right) {\Greekmath 0116} _{z_{1}}^{\prime }\left( w\right) g(w)}{E\left[ f_{V|W}\left( {\Greekmath 0116} \left( W\right) |W\right) {\Greekmath 0116} _{z_{1}}^{\prime }\left( W\right) g(W)\right] }. \end{equation}

The lemma shows that the weighting function $\mathcal{\dot{P}}\left( w\right) $ does not depend on the function form of $s\left( \cdot \right) .$ Hence, the unconditional policy effect does not depend on $s\left( \cdot \right) .$

corollaryLet Assumptions (ref)--(ref) and (ref), and the assumptions of Lemma (ref) hold. Assume further that $f_{Y}(y_{{\Greekmath 011C} })>0$. Then, the unconditional quantile effect of the instrumental intervention given in (ref) is \begin{equation} \Pi _{{\Greekmath 011C} }=\int_{\mathcal{W}}\widetilde{\mathrm{MTE}}_{{\Greekmath 011C} }\left( P(w),w_{-1}\right) \mathcal{\dot{P}}\left( w\right) dF_{W}\left( w\right). \notag \end{equation}

The proposition below establishes the identifiability of $\widetilde{\mathrm{ MTE}}_{{\Greekmath 011C} }$ and of the weighting function $\mathcal{\dot{P}}\left( w\right) $ given in Corollary (ref).

propositionLet Assumptions (ref)( (ref)), (ref)((ref)), and (ref)((ref)), and the assumptions in Lemma (ref) hold. Then, for every $u=P(w)$ for some $w=(w_{1},w_{-1})\in \mathcal{W}$, we have\footnote{ For a more general functional ${\Greekmath 011A} ,$ we can show that \begin{equation*} \widetilde{\mathrm{MTE}}_{{\Greekmath 011A} }(u,w_{-1})=\frac{\partial E\left[ {\Greekmath 0120} (Y,{\Greekmath 011A} ,F_{Y})|P(W)=u,W_{-1}=w_{-1}\right] }{\partial u} \end{equation*} for every $u$ equal to $P(w)$ for some $w=(w_{1},w_{-1})\in \mathcal{W}$.} \begin{equation*} \widetilde{\mathrm{MTE}}_{{\Greekmath 011C} }\left( u,w_{-1}\right) =-\frac{1}{ f_{Y}\left( y_{{\Greekmath 011C} }\right) }\frac{\partial E\left[ \mathds{1}\left\{ Y\leq y_{{\Greekmath 011C} }\right\} |P(W)=u,W_{-1}=w_{-1}\right] }{\partial u}, \end{equation*} and \begin{equation} \mathcal{\dot{P}}\left( w\right) =\frac{\frac{\partial P(w)}{\partial z_{1}} g(w)}{E\left[ \frac{\partial P(W)}{\partial z_{1}}g\left( W\right) \right] } \end{equation} where $\frac{\partial P(W)}{\partial z_{1}}$ is short for $\left. \frac{ \partial P(w)}{\partial z_{1}}\right\vert _{w=W}.$

Using Proposition (ref), and the fact that $g(w)$ is known, we can represent $\Pi _{{\Greekmath 011C} }$ as

equation[equation omitted — 399 chars of source]

All objects in the above are point identified, hence $\Pi _{{\Greekmath 011C} }$ is point identified.\footnote{ We note that Assumption (ref)((ref)) plays a key role in identifying $\mathcal{\dot{P}}\left( w\right) .$ Without the assumption that $V$ is independent of $Z_{1}$ conditional on $\left( Z_{-1},X\right) ,$ we can have only that

equation*[equation* omitted — 248 chars of source]

The presence of the second term in the above equation invalidates the identification result in ((ref)).}

Instrumental Intervention with an invalid instrument

While identifying the unconditional effect of an instrumental intervention is of interest in its own right, the identification result can be further leveraged to identify the unconditional policy effect of another intervention. Theorem (ref) has shown that the unconditional effects of two interventions will be the same if their weighting functions coincide. Consider a counterfactual policy with a target weighting function $ \mathcal{\dot{P}}^{\circ }\left( \cdot \right) .$ By strategically choosing the instrument function $g\left( \cdot \right) $, we can ensure that the weighting function under the policy intervention in (ref) is the same as $\mathcal{\dot{P}}^{\circ }\left( \cdot \right) .$ In other words, with appropriate choices of $g\left( \cdot \right) ,$ the unconditional effect of intervening $Z_{1}$ is the same as the unconditional effect of another counterfactual policy. If the former is identified, then the latter is also identified.

As an example, suppose we intervene on the second element $Z_{2}$ of $Z$ with

equation[equation omitted — 210 chars of source]

for some $g^{\circ }\left( \cdot \right) $ and $s^{\circ }\left( {\Greekmath 010E} \right) .$ Under conditions similar to those in Lemma (ref), the weighting function for this intervention is

equation*[equation* omitted — 319 chars of source]

Letting $\mathcal{\dot{P}}^{\circ }\left( w\right) =\mathcal{\dot{P}}\left( w\right) $ for $\mathcal{\dot{P}}\left( w\right) $ given in Lemma (ref) and solving for $g(\cdot )$ yields

equation*[equation* omitted — 149 chars of source]

So, the unconditional effect of the intervention given in ((ref)) is the same as that of the intervention given in (ref) when $g(w)$ is chosen appropriately. It is important to point out that $Z_{2}$ may not be a valid instrument, and its unconditional effect is identified via \textquotedblleft intervention matching.\textquotedblright

In general, identifying an unconditional effect of an invalid instrument via intervention matching\ is feasible only if we have prior knowledge of the ratio ${\Greekmath 0116} _{z_{2}}^{\prime }\left( w\right) /{\Greekmath 0116} _{z_{1}}^{\prime }\left( w\right) .$ This information may be available from economic theory. When such knowledge is unavailable, an alternative approach can be employed.

The alternative approach hinges on the assumption that $\left( Z_{1},Z_{2}\right) $ is independent of $V$ conditional on the remaining elements in $W$, denoted as $W_{-(1,2)}$. In this case, for $ W=(Z_{1},Z_{2},W_{-(1,2)})$ and $w=(z_{1},z_{2},w_{-\left( 1,2\right) }),$ we have

eqnarray*[eqnarray* omitted — 396 chars of source]

and

equation*[equation* omitted — 589 chars of source]

Thus, the ratio ${\Greekmath 0116} _{z_{2}}^{\prime }\left( w\right) /{\Greekmath 0116} _{z_{1}}^{\prime }\left( w\right) $ can be identified via the ratio of two partial derivatives of the propensity score function.

It is important to note that the conditional independence of $Z_{2}$ from $V$ given $W_{-(1,2)}$ does not rule out the possibility that $Z_{2}$ may still be dependent on $(U_{0},U_{1})$, and consequently, $Z_{2}$ could still be an invalid instrument. This shows that intervention matching may be used to identify the effect of intervening on an invalid instrument.

Unconditional Instrumental Quantile Estimation

This section is devoted to the estimation and inference of the UQE under the instrumental intervention in ((ref)). We assume that the propensity score function is parametric, and we leave the case with a nonparametric propensity score to Section (ref) of the supplementary appendix. In order to simplify the notation, we set $g(\cdot )\equiv 1$ for the remainder of this paper.\footnote{ The presence of $g\left( \cdot \right) $ amounts to a change of measure: from a measure with density $f_{W}(w)$ to a measure with density $ g(w)f_{W}\left( w\right) $. When $g\left( \cdot \right) $ is not equal to a constant function, we only need to change the population expectation operator $E\left[ h(W)\right] $ that involves the distribution of $W$ into $E \left[ h(W)g(W)\right] $ and the empirical average operator $\mathbb{P}_{n} \left[ h\left( W\right) \right] $ into $\mathbb{P}_{n}\left[ h(W)g\left( W\right) \right] .$ All of our results will remain valid.}

Letting

equation*[equation* omitted — 154 chars of source]

and using ((ref)), we have

equation[equation omitted — 250 chars of source]

$\Pi _{{\Greekmath 011C} }$ consists of two average derivatives and a density evaluated at a point, some of which depend on the unconditional ${\Greekmath 011C} $-quantile $ y_{{\Greekmath 011C} }$. Altogether $\Pi _{{\Greekmath 011C} }$ depends on four unknown quantities.

The method of unconditional instrumental quantile estimation involves first estimating the four quantities separately and then plugging these estimates into $\Pi _{{\Greekmath 011C} }$ to obtain the estimator $\hat{\Pi}_{{\Greekmath 011C} }.$ See ((ref)) in Subsection (ref) for the formula of $\hat{ \Pi}_{{\Greekmath 011C} }.$

We consider estimating the four quantities in the next few subsections. For a given sample $\left\{ O_{i}=(Y_{i},Z_{i},X_{i},D_{i})\right\} _{i=1}^{n}$, we will use $\mathbb{P}_{n}$ to denote the empirical measure. The expectation of a function ${\Greekmath 011F} \left( O\right) $ with respect to $\mathbb{P} _{n}$ is then $\mathbb{P}_{n}{\Greekmath 011F} =n^{-1}\sum_{i=1}^{n}{\Greekmath 011F} (O_{i})$. Similarly, we use $\mathbb{P}$ to denote the population measure, and so $ \mathbb{P}{\Greekmath 011F} \left( O\right) =E{\Greekmath 011F} \left( O\right) .$

Estimating the Quantile and Density

For a given ${\Greekmath 011C} $, we estimate $y_{{\Greekmath 011C} }$ using the (generalized) inverse of the empirical distribution function of $Y$: $\hat{y}_{{\Greekmath 011C} }=\inf \left\{ y:\mathbb{F}_{n}(y)\geq {\Greekmath 011C} \right\} $, where

equation*[equation* omitted — 100 chars of source]

By Lemma (ref) in the appendix, we can write $\hat{y}_{{\Greekmath 011C} }-y_{{\Greekmath 011C} }=\mathbb{P}_{n}{\Greekmath 0120} _{Q}(Y,y_{{\Greekmath 011C} })+o_{p}(n^{-1/2})$ where

equation*[equation* omitted — 180 chars of source]

Here the subscript \textquotedblleft $Q$\textquotedblright\ on ${\Greekmath 0120} _{Q}$ signifies that it is the influence function for a Quantile functional.

We use a kernel density estimator to estimate $f_{Y}(y)$. We maintain the following assumptions on the kernel function and the bandwidth.

assumptionKernel Assumption • The kernel function $K(\cdot )$ satisfies (i) $\int_{-\infty }^{\infty }K(u)du=1$, (ii) $\int_{-\infty }^{\infty }u^{2}K(u)du<\infty $, and (iii) $ K(u)=K(-u)$, and it is twice differentiable with Lipschitz continuous second-order derivative $K^{\prime \prime }\left( u\right) $ satisfying (i) $ \int_{-\infty }^{\infty }K^{\prime \prime }(u)udu<\infty $ and $\left( ii\right) $ there exist positive constants $C_{1}$ and $C_{2}$ such that $ \left\vert K^{\prime \prime }\left( u_{1}\right) -K^{\prime \prime }\left( u_{2}\right) \right\vert \leq C_{2}\left\vert u_{1}-u_{2}\right\vert ^{2}$ for $\left\vert u_{1}-u_{2}\right\vert \geq C_{1}.$
assumptionRate Assumption : $n\uparrow \infty $ and $ h\downarrow 0$ such that $nh^{3}\uparrow \infty $ but $nh^{5}=O(1)$.

The non-standard condition $nh^{3}\uparrow \infty $ is due to the estimation of $y_{{\Greekmath 011C} }$. Since we need to expand $\hat{f}_{Y}(\hat{y}_{{\Greekmath 011C} })-\hat{f} _{Y}(y_{{\Greekmath 011C} })$, which involves the derivative of $\hat{f}_{Y}(y),$ we have to impose a slower rate of decay for $h$ to control the remainder. The details can be found in the proof of Lemma (ref). We note, however, that $nh^{3}\uparrow \infty $ implies the usual rate condition $ nh\uparrow \infty $.

The estimator of $f_{Y}(y)$ is then given by

equation*[equation* omitted — 153 chars of source]

where $K_{h}\left( u\right) :=K(u/h)/h.$

By Lemma (ref) in the appendix, we can isolate the estimation errors from estimating the density and quantile and write:

align[align omitted — 525 chars of source]

Here, ${\Greekmath 0120} _{f_{Y}}\left( Y,y_{{\Greekmath 011C} }\right) :=K_{h}\left( Y-y_{{\Greekmath 011C} }\right) -E\left[ K_{h}\left( Y-y_{{\Greekmath 011C} }\right) \right] $ and $ B_{f_{Y}}(y_{{\Greekmath 011C} })$ is the bias term, which is $O(h^{2})$. The first two terms on the right-hand side of ((ref)) represent the dominating terms in the error from estimating $f_{Y}$. The third term reflects the error from estimating $y_{{\Greekmath 011C} }$.

Estimating the Average Derivatives

To estimate the two average derivatives, we make a parametric assumption on the propensity score, leaving the nonparametric specification to Section (ref) of the supplementary appendix.

assumptionThe propensity score $P(Z,X,{\Greekmath 010B} _{0})$ is known up to a finite-dimensional vector ${\Greekmath 010B} _{0}\in \mathbb{R} ^{d_{{\Greekmath 010B} }}$.

Under Assumption (ref), the UQE $\Pi _{{\Greekmath 011C} }$ can be written as

equation[equation omitted — 140 chars of source]

where

equation*[equation* omitted — 360 chars of source]

First, we estimate $T_{1},$ which is the mean of the derivative of the propensity score, by

equation*[equation* omitted — 193 chars of source]

where $\hat{{\Greekmath 010B}}$ is an estimator of ${\Greekmath 010B} _{0}$ satisfying $\hat{{\Greekmath 010B} }-{\Greekmath 010B} _{0}=\mathbb{P}_{n}{\Greekmath 0120} _{{\Greekmath 010B} _{0}}\left( D,W\right) +o_{p}(n^{-1/2})$ for some measurable function ${\Greekmath 0120} _{{\Greekmath 010B} _{0}}\left( \cdot ,\cdot \right) .$ To save space, we slightly abuse notation and write

equation*[equation* omitted — 181 chars of source]

We adopt this convention in the rest of the paper. Under Lemma (ref) in the appendix, we have

eqnarray[eqnarray omitted — 382 chars of source]

where

equation*[equation* omitted — 210 chars of source]

Equation ((ref)) has a similar interpretation to equation ((ref)). It consists of a term that ignores the estimation uncertainty in $\hat{{\Greekmath 010B}}$ but accounts for the variability of the sample mean, and another term that accounts for the uncertainty in $\hat{{\Greekmath 010B}}$ but ignores the variability of the sample mean.

We estimate the second average derivative $T_{2}$ by

equation[equation omitted — 244 chars of source]

See ((ref)) for an explicit construction. We can regard $T_{2n}(\hat{y} _{{\Greekmath 011C} },\hat{m},\hat{{\Greekmath 010B}})$ as a four-step estimator. The first step estimates $y_{{\Greekmath 011C} }$, the second step estimates ${\Greekmath 010B} _{0}$, the third step estimates the conditional expectation $m_{0}(y,P(Z,X,{\Greekmath 010B} _{0}),W_{-1})$ using the generated regressor $P(Z,X,\hat{{\Greekmath 010B}})$, and the fourth step averages the derivative (with respect to $Z_{1})$ over the generated regressor $P(Z,X,\hat{{\Greekmath 010B}})$ and $W_{-1}$.

We use the series method to estimate $m_{0}$. To alleviate notation, define the vector $\tilde{w}({\Greekmath 010B} ):=(P(z,x,{\Greekmath 010B} ),w_{-1})^{\prime }\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ and } \tilde{W}_{i}({\Greekmath 010B} ):=(P(Z_{i},X_{i},{\Greekmath 010B} ),W_{-1,i})^{\prime }.$ We write $\tilde{w}=\tilde{w}\left( {\Greekmath 010B} _{0}\right) $ and $\tilde{W}_{i}= \tilde{W}_{i}({\Greekmath 010B} _{0})$ to suppress their dependence on the true parameter value ${\Greekmath 010B} _{0}.$ Both $\tilde{w}({\Greekmath 010B} )$ and $\tilde{W}_{i}({\Greekmath 010B} )$ are in $\mathbb{R} ^{d_{W}}.$ Let ${\Greekmath 011E} ^{J}(\tilde{w}({\Greekmath 010B} ))=({\Greekmath 011E} _{1J}(\tilde{w}({\Greekmath 010B} )),\ldots ,{\Greekmath 011E} _{JJ}(\tilde{w}({\Greekmath 010B} )))^{\prime }$ be a vector of $J$ basis functions of $\tilde{w}({\Greekmath 010B} )$ with finite second moments\footnote{ If any variable in $W_{-1}$ is discrete with a small number of possible values, we can exclude it in the basis functions. Instead, we can construct the basis functions using only the remaining variables and apply the series method to each subsample defined by the values of the discrete variable.}. Here, each ${\Greekmath 011E} _{jJ}\left( \cdot \right) $ is a differentiable basis function. Then, the series estimator of $m_{0}(y_{{\Greekmath 011C} },\tilde{w}({\Greekmath 010B} ))$ is $\hat{m}(\hat{y}_{{\Greekmath 011C} },\tilde{w}(\hat{{\Greekmath 010B}}))={\Greekmath 011E} ^{J}(\tilde{w}( \hat{{\Greekmath 010B}}))^{\prime }\hat{b}(\hat{{\Greekmath 010B}},\hat{y}_{{\Greekmath 011C} }),$ where $\hat{b}(\hat{{\Greekmath 010B}},\hat{y}_{{\Greekmath 011C} })$ is:

equation*[equation* omitted — 390 chars of source]

The estimator of the average derivative $T_{2}$ is then

equation[equation omitted — 293 chars of source]

We use the path derivative approach of newey1994 to obtain a decomposition of $T_{2n}(\hat{y}_{{\Greekmath 011C} },\hat{m},\hat{{\Greekmath 010B}})-T_{2},$ which is similar to that in Section 2.1 of hahn2013. To describe the idea, let $\left\{ F_{{\Greekmath 0112} }\right\} $ be a path of distributions indexed by $ {\Greekmath 0112} \in \mathbb{R}$ such that $F_{{\Greekmath 0112} _{0}}$ is the true distribution of $O:=(Y,Z,X,D)$. The parametric assumption on the propensity score does not need to be imposed on the path.\footnote{ As we show later, the error from estimating the propensity score does not affect the asymptotic variance of $T_{2n}(\hat{y}_{{\Greekmath 011C} },\hat{m},\hat{{\Greekmath 010B} }).$} The score of the parametric submodel is $S(O)=\frac{\partial \log dF_{{\Greekmath 0112} }(O)}{\partial {\Greekmath 0112} }\big|_{{\Greekmath 0112} ={\Greekmath 0112} _{0}}.$ For any ${\Greekmath 0112} ,$ we define

equation*[equation* omitted — 240 chars of source]

where $m_{{\Greekmath 0112} },$ $y_{{\Greekmath 011C} ,{\Greekmath 0112} },$ and ${\Greekmath 010B} _{{\Greekmath 0112} }$ are the probability limits of $\hat{m},$ $\hat{y}_{{\Greekmath 011C} },$ and $\hat{{\Greekmath 010B}},$ respectively, when the distribution of $O$ is $F_{{\Greekmath 0112} }$. Note that when $ {\Greekmath 0112} ={\Greekmath 0112} _{0},$ we have ${\Greekmath 010B} _{{\Greekmath 0112} _{0}}={\Greekmath 010B} _{0},m_{{\Greekmath 0112} _{0}}=m_{0}$ and $T_{2,{\Greekmath 0112} _{0}}=T_{2}.$ Suppose the set of scores $ \left\{ S(O)\right\} $ for all parametric submodels $\left\{ F_{{\Greekmath 0112} }\right\} $ can approximate any zero-mean, finite-variance function of $O$ in the mean square sense.\footnote{ This is the \textquotedblleft generality\textquotedblright\ requirement of the family of distributions in newey1994.} If the function ${\Greekmath 0112} \rightarrow T_{2,{\Greekmath 0112} }$ is differentiable at ${\Greekmath 0112} _{0}$ and we can write

equation[equation omitted — 191 chars of source]

for some mean-zero and finite second-moment function $\Gamma (\cdot )$ and any path $F_{{\Greekmath 0112} },$ then, by Theorem 2.1 of newey1994, the asymptotic variance of $T_{2n}(\hat{y}_{{\Greekmath 011C} },\hat{m},\hat{{\Greekmath 010B}})$ is\ $ E[\Gamma (O)^{2}]$.

In Lemma (ref) in the appendix, we show that ${\Greekmath 0112} \rightarrow T_{2,{\Greekmath 0112} }$ is differentiable at ${\Greekmath 0112} _{0}$. Then, by the chain rule, we can write

eqnarray*[eqnarray* omitted — 1,232 chars of source]

To use Theorem 2.1 of newey1994, we need to write all these terms in an outer-product form, namely the form of the right-hand side of ((ref)). To search for the required function $\Gamma (\cdot )$, we follow newey1994 and examine the components of $T_{2,{\Greekmath 0112} }$ one at a time, treating the remaining components as known.

Lemma (ref) in the appendix shows that under some conditions

equation*[equation* omitted — 291 chars of source]

that is, we can ignore the error from estimating the propensity score in our asymptotic analysis. Lemma (ref) in the appendix characterizes the influence function of $T_{2n}(\hat{y}_{{\Greekmath 011C} },\hat{m},\hat{ {\Greekmath 010B}})$ and establishes a stochastic approximation of $T_{2n}(\hat{y} _{{\Greekmath 011C} },\hat{m},\hat{{\Greekmath 010B}})-T_{2}$ as follows:

equation[equation omitted — 375 chars of source]

where

align*[align* omitted — 441 chars of source]

and

equation*[equation* omitted — 306 chars of source]

This characterizes the contribution of each stage to the influence function of $T_{2n}(\hat{y}_{{\Greekmath 011C} },\hat{m},\hat{{\Greekmath 010B}})$. The contribution from estimating $m_{0}$, given by $\mathbb{P}_{n}{\Greekmath 0120} _{m_{0}}$, corresponds to the one in Proposition 5 of newey1994 (p. 1362).

Estimating the UQE

With $\hat{y}_{{\Greekmath 011C} },\hat{f}_{Y},\hat{m},\hat{{\Greekmath 010B}}$ given in the previous subsections, we estimate the UQE by

equation[equation omitted — 311 chars of source]

This is our unconditional instrumental quantile estimator (UNIQUE). With the asymptotic linear representations of all three components $\hat{f}_{Y}(\hat{y }_{{\Greekmath 011C} }),$ $T_{1n}(\hat{{\Greekmath 010B}}),$ and $T_{2n}(\hat{y}_{{\Greekmath 011C} },\hat{m}, \hat{{\Greekmath 010B}}),$ we can obtain the asymptotic linear representation of $\hat{ \Pi}_{{\Greekmath 011C} }(\hat{y}_{{\Greekmath 011C} },\hat{f}_{Y},\hat{m},\hat{{\Greekmath 010B}}).$ The next theorem follows from combining Lemmas (ref), (ref), (ref), and (ref).

theoremUnder the assumptions of Lemmas (ref), (ref), (ref), and (ref), we have \begin{eqnarray} \hat{\Pi}_{{\Greekmath 011C} }-\Pi _{{\Greekmath 011C} } &=&\frac{T_{2}}{f_{Y}(y_{{\Greekmath 011C} })^{2}T_{1}} \left[ \mathbb{P}_{n}{\Greekmath 0120} _{f_{Y}}(Y,y_{{\Greekmath 011C} })+B_{f_{Y}}(y_{{\Greekmath 011C} })\right] + \frac{T_{2}}{f_{Y}(y_{{\Greekmath 011C} })^{2}T_{1}}f_{Y}^{\prime }(y_{{\Greekmath 011C} })\mathbb{P} _{n}{\Greekmath 0120} _{Q}(Y,y_{{\Greekmath 011C} }) \notag \\ &+&\frac{T_{2}}{f_{Y}(y_{{\Greekmath 011C} })T_{1}^{2}}\mathbb{P}_{n}{\Greekmath 0120} _{\partial P}\left( W\right) +\frac{T_{2}}{f_{Y}(y_{{\Greekmath 011C} })T_{1}^{2}}E\left[ \frac{ \partial ^{2}P(Z,X,{\Greekmath 010B} _{0})}{\partial z_{1}\partial {\Greekmath 010B} _{0}^{\prime } }\right] \mathbb{P}_{n}{\Greekmath 0120} _{{\Greekmath 010B} _{0}}\left( D,W\right) \\ &-&\frac{1}{f_{Y}(y_{{\Greekmath 011C} })T_{1}}\mathbb{P}_{n}{\Greekmath 0120} _{\partial m_{0}}\left( W,y_{{\Greekmath 011C} }\right) -\frac{1}{f_{Y}(y_{{\Greekmath 011C} })T_{1}}\mathbb{P}_{n}{\Greekmath 0120} _{m_{0}}\left( Y,W,y_{{\Greekmath 011C} }\right) \notag \\ &&-\frac{1}{f_{Y}(y_{{\Greekmath 011C} })T_{1}}\mathbb{P}_{n}\tilde{{\Greekmath 0120}}_{Q}(Y,y_{{\Greekmath 011C} })+R_{\Pi }, \notag \end{eqnarray} where \begin{eqnarray*} R_{\Pi } &=&O_{p}\left( |\hat{f}_{Y}(\hat{y}_{{\Greekmath 011C} })-f_{Y}(y_{{\Greekmath 011C} })|^{2}\right) +O_{p}\left( n^{-1}\right) +O_{p}\left( n^{-1/2}|\hat{f}_{Y}( \hat{y}_{{\Greekmath 011C} })-f_{Y}(y_{{\Greekmath 011C} })|\right) \\ &+&o_{p}\left( n^{-1/2}h^{-1/2}\right) +o_{p}(h^{2}). \end{eqnarray*} Furthermore, under Assumption (ref), $\sqrt {nh} R_\Pi=o_p(1).$

Equation ((ref)) consists of six influence functions and a bias term. The bias term $B_{f_{Y}}(y_{{\Greekmath 011C} })$ arises from estimating the density and is of order $O(h^{2})$. The six influence functions reflect the impact of each estimation stage. The rate of convergence of $\hat{\Pi} _{{\Greekmath 011C} }$ is slowed down through $\mathbb{P}_{n}{\Greekmath 0120} _{f_{Y}}(Y)$, which is of order $O_{p}(n^{-1/2}h^{-1/2})$. We can summarize the results of Theorem (ref) in a single equation:

equation*[equation* omitted — 217 chars of source]

where ${\Greekmath 0120} _{\Pi _{{\Greekmath 011C} }}$ collects all the influence functions in ((ref)) except for the bias, and

equation*[equation* omitted — 147 chars of source]

If $nh^{5}\rightarrow 0,$ then the bias term is $o(n^{-1/2}h^{-1/2})$. The following corollary provides the asymptotic distribution of $\hat{\Pi}_{{\Greekmath 011C} }$.

corollaryUnder the assumptions of Theorem (ref) and the assumption that $nh^{5}\rightarrow 0,$ \begin{equation*} \sqrt{nh}\left( \hat{\Pi}_{{\Greekmath 011C} }-\Pi _{{\Greekmath 011C} }\right) =\sqrt{n}\mathbb{P}_{n} \sqrt{h}{\Greekmath 0120} _{\Pi _{{\Greekmath 011C} }}\left( O\right) +o_{p}(1)\Rightarrow \mathcal{N} (0,V_{{\Greekmath 011C} }), \end{equation*} where \begin{equation} V_{{\Greekmath 011C} }=\lim_{h\downarrow 0}E\left\{ h\left[ {\Greekmath 0120} _{\Pi _{{\Greekmath 011C} }}\left( O\right) \right] ^{2}\right\} . \end{equation}

From the perspective of asymptotic theory, all of the following terms are all of order $O_{p}\left( h\right) =o_{p}\left( 1\right) $ and hence can be ignored in large samples: $\sqrt{nh}\mathbb{P}_{n}{\Greekmath 0120} _{Q},$ $\sqrt{nh} \mathbb{P}_{n}{\Greekmath 0120} _{\partial P},$ $\sqrt{nh}\mathbb{P}_{n}{\Greekmath 0120} _{{\Greekmath 010B} _{0}},$ $\sqrt{nh}\mathbb{P}_{n}{\Greekmath 0120} _{\partial m_{0}},$ $\sqrt{nh}\mathbb{P} _{n}{\Greekmath 0120} _{m_{0}},$ and $\sqrt{nh}\mathbb{P}_{n}\tilde{{\Greekmath 0120}}_{Q}.$ The asymptotic variance is then given by

equation*[equation* omitted — 302 chars of source]

However, $V_{{\Greekmath 011C} }$ ignores all estimation uncertainties except that in $ \hat{f}_{Y}(y_{{\Greekmath 011C} })$, and we do not expect it to reflect the finite-sample variability of $\sqrt{nh}(\hat{\Pi}_{{\Greekmath 011C} }-\Pi _{{\Greekmath 011C} })$ well. To improve the finite-sample performances, we keep the dominating term from each source of estimation errors and employ a sample counterpart of $ E[h{\Greekmath 0120} _{\Pi _{{\Greekmath 011C} }}^{2}]$ to estimate $V_{{\Greekmath 011C} }.$ The details can be found in Section (ref) of the supplementary appendix.

Testing the Null of No Effect

We can apply Corollary (ref) to conduct hypothesis testing on $\Pi _{{\Greekmath 011C} }$. Since $\hat{\Pi}_{{\Greekmath 011C} }$ converges to $\Pi _{{\Greekmath 011C} }$ at a nonparametric rate, in general, the test will have power only against a local departure of a nonparametric rate. However, if we are interested in testing the null of a zero effect, that is, $H_{0}:\Pi _{{\Greekmath 011C} }=0$ vs. $ H_{1}:\Pi _{{\Greekmath 011C} }\neq 0,$ we can detect a parametric rate of departure from the null. The reason is that, by ((ref)), $\Pi _{{\Greekmath 011C} }=0$ if and only if $T_{2}=0,$ and $T_{2}$ can be estimated at the usual parametric rate. Hence, instead of testing $H_{0}:\Pi _{{\Greekmath 011C} }=0$ vs. $ H_{1}:\Pi _{{\Greekmath 011C} }\neq 0,$ we can test the equivalent hypotheses $ H_{0}:T_{2}=0$ vs. $H_{1}:T_{2}\neq 0.$

Our test is based on the estimator $T_{2n}(\hat{y}_{{\Greekmath 011C} },\hat{m},\hat{ {\Greekmath 010B}})$ of $T_{2}$. In view of its influence function given in Lemma (ref), we can estimate the asymptotic variance of $T_{2n}( \hat{y}_{{\Greekmath 011C} },\hat{m},\hat{{\Greekmath 010B}})$ by

equation*[equation* omitted — 357 chars of source]

where $\hat{{\Greekmath 0120}}_{\partial m,i},$ $\hat{{\Greekmath 0120}}_{m,i}$ and $\hat{{\Greekmath 0120}}_{Q,i}$ are plug-in estimates of ${\Greekmath 0120} _{\partial m_{0}}\left( W_{i},y_{{\Greekmath 011C} }\right) ,{\Greekmath 0120} _{m_{0}}\left( Y_{i},W_{i},y_{{\Greekmath 011C} }\right) ,$ and $\hat{{\Greekmath 0120}} _{Q}\left( Y_{i},y_{{\Greekmath 011C} }\right) ,$ respectively.

We can then form the test statistic:

equation[equation omitted — 160 chars of source]

By Lemma (ref) and using standard arguments, we can show that $T_{2n}^{o}\Rightarrow \mathcal{N}(0,1).$ To save space, we omit the details here.

Simulation Evidence

In the simulation study, we consider the following structural equations:

align*[align* omitted — 165 chars of source]

Here, $(U_{0},U_{1},Z_{1},Z_{2},Z_{3},X_{1},X_{2})$ are jointly normal, independent variables, with mean $0$ and unit variances, and

equation*[equation* omitted — 135 chars of source]

where $e$ is standard normal, independent of $ (U_{0},U_{1},Z_{1},Z_{2},Z_{3},X_{1},X_{2})$.\footnote{ Note that $V$ has also unit variance.} The parameter ${\Greekmath 0125} $ governs the endogeneity of $D$.

The target of interest is the UQE $\Pi _{{\Greekmath 011C} }$, as defined in ((ref)), involving $y_{{\Greekmath 011C} },f_{Y},$ $T_{1},$ and $T_{2}$.\footnote{ Detailed calculations of $\Pi _{{\Greekmath 011C} }$ are available from Appendix B of an earlier working paper sun2021,} We estimate $\Pi _{{\Greekmath 011C} }$ using the UNIQUE outlined in Section (ref). More specifically, for estimating the (population) quantile $y_{{\Greekmath 011C} },$ we utilize the sample quantile. To estimate $f_{Y}$, we use a Gaussian kernel with bandwidth $ h=1.06\times \hat{{\Greekmath 011B}}_{Y}\times n^{-1/4}$, where $\hat{{\Greekmath 011B}}_{Y}$ is the sample standard deviation of $Y$. For estimating $T_{1}$, we employ a probit model, and for $T_{2}$, we run a cubic series regression.

Testing the Null Hypothesis of No Effect

If we set ${\Greekmath 010C} =0$, $\Pi _{{\Greekmath 011C} }=0$ and so the null hypothesis of a zero effect holds. The test statistic $T_{2n}^{o}$ is constructed following equation (ref). Because the test statistic does not involve estimating the density (or $T_{1}$), the test has nontrivial power again $1/\sqrt{n}$-departures (i.e., ${\Greekmath 010C} =c/\sqrt{n}$ for some $c\neq 0)$ from the null.

To simulate the power function of the nominal 5% test, we consider a range of 25 values of ${\Greekmath 010C} $ between $-1$ and $1$. The endogeneity, governed by the parameter ${\Greekmath 0125} $, takes five values: 0, 0.25, 0.5, 0.75, and 0.9. We perform 1,000 simulations with 1,000 observations. For different values of $ {\Greekmath 011C} $, the power functions are shown below in Figure (ref). \ The test has the desired null rejection probability, except for the extreme quantile ${\Greekmath 011C} =0.1$, where under high endogeneity, the rejection probability does not increase fast enough. The power function for the median, ${\Greekmath 011C} =0.5$, is not shown here as it is almost identical to that of $ {\Greekmath 011C} =0.4$. Furthermore, simulation results not reported here show that the power functions for ${\Greekmath 011C} =0.6,0.7,0.8,0.9$ are very similar to those of $ {\Greekmath 011C} =0.4,0.3,0.2,0.1,$ respectively.

figure[figure omitted — 206 chars of source]

Empirical Coverage of Confidence Intervals

In this subsection, we investigate the empirical coverage of confidence intervals built using $\hat{V}_{{\Greekmath 011C} }$, the variance estimator given in (ref). Since $\sqrt{nh}\left( \hat{\Pi}_{{\Greekmath 011C} }-\Pi _{{\Greekmath 011C} }\right) \approx \mathcal{N}(0,\hat{V}_{{\Greekmath 011C} }),$ a 95% confidence interval for $\Pi _{{\Greekmath 011C} }$ can be constructed using

equation*[equation* omitted — 116 chars of source]

We use a grid of ${\Greekmath 010C} $ that takes values $-1,-0.5,-0.25,0,0.25,0.5,$ and $ 1$. For the endogeneity parameter, ${\Greekmath 0125} $, we take the values $ 0,0.25,0.5,0.75,$ and $0.9$. Finally, ${\Greekmath 011C} $ takes the values from $0.1$ to $0.9$ with an increment of 0.1. We note that, for values of ${\Greekmath 010C} \neq 0$, where the effect is not 0, we need to numerically compute the value of $\Pi _{{\Greekmath 011C} }$. We perform 10,000 simulations with 1,000 observations each. The results are reported in the tables below for ${\Greekmath 011C} =0.1$ and ${\Greekmath 011C} =0.5$. It is clear that the confidence intervals have reasonable coverage accuracy in almost all cases.

table[table omitted — 626 chars of source]
table[table omitted — 626 chars of source]

Empirical Application

We estimate the unconditional quantile effect of expanding college enrollment on (log) wages. The outcome variable $Y$ is the log wage, and the binary treatment is the college enrollment status. Thus, $p=\Pr [D=1]$ is the proportion of individuals who ever enrolled in a college. Arguably, the cost of tuition $(Z_{1})$, assumed to be continuous, is an important factor that affects the college enrollment status but not the wage directly. In order to alter the proportion of enrolled individuals, we consider a policy that subsidizes tuition by a certain amount. The UQE is the effect of this policy on the different quantiles of the unconditional distribution of wages when the subsidy is small. This policy shifts $Z_{1}$, the tuition, to $Z_{1{\Greekmath 010E} }=Z_{1}+s({\Greekmath 010E} )$ for some $s({\Greekmath 010E} )$, which is the same for all individuals, and induces a small change in college enrollment. Note that we do not need to specify $ s({\Greekmath 010E} )$ because we look at the limiting version as ${\Greekmath 010E} \rightarrow 0$ . In practice, we may set $s({\Greekmath 010E} )$ equal to a number that is relatively small compared to the total tuition.

We use the same data as in Carneiro2010 and Carneiro2011: a sample of white males from the 1979 National Longitudinal Survey of Youth (NLSY1979). The web appendix to Carneiro2011 contains a detailed description of the variables. The outcome variable $Y$ is the log wage in 1991. The treatment indicator $D$ is equal to $1$ if the individual ever enrolled in college by 1991, and $0$ otherwise. The other covariates are AFQT score, mother's education, number of siblings, average log earnings 1979--2000 in the county of residence at age 17, average unemployment 1979--2000 in the state of residence at age 17, urban residence dummy at age 14, cohort dummies, years of experience in 1991, average local log earnings in 1991, and local unemployment in 1991. We collect these variables into a vector and denote it by $X$.

We assume that the following four variables (denoted by $ Z_{1},Z_{2},Z_{3},Z_{4}$) enter the selection equation but not the outcome equation: tuition at local public four-year colleges at age 17, presence of a four-year college in the county of residence at age 14, local earnings at age 17, and local unemployment at age 17. The total sample size is 1747, of which 882 individuals had never enrolled in a college ($D=0$) by 1991, and 865 individuals had enrolled in a college by 1991 ($D=1)$. We estimate the UQE of a marginal shift in the tuition at local public four-year colleges at age 17 ($Z_{1})$ using the UNIQUE.

Here are some details of the UNIQUE. To estimate the propensity score, we use a parametric logistic specification. To estimate the conditional expectation function $m_{0}$,$\ $we run a series regression using the estimated propensity score and the covariates $Z_{2},Z_{3},Z_{4},$ and $X$\ as the regressors. Due to the large number of variables involved, a penalization of ${\Greekmath 0115} =10^{-4}$ was imposed on the $L_{2}$-norm of the coefficients, excluding the constant term as in ridge regressions. We estimate the UQE at the quantile level ${\Greekmath 011C} =0.1,0.15,\ldots ,0.9$. For each ${\Greekmath 011C} ,$ we also construct the 95% (pointwise) confidence interval.

Figure (ref) presents the results. The estimated UQE ranges between 0.22 and 0.47 across the quantiles with an average of 0.37. When we estimate the unconditional mean effect, we obtain an estimate of 0.21, which is somewhat consistent with the quantile cases. We interpret these estimates in the following way: the effect of a ${\Greekmath 010E} $ (small) increase in college enrollment induced by an additive change in tuition increases (log) wages between $0.22\times {\Greekmath 010E} $ and $0.47\times {\Greekmath 010E} $ across quantiles. For example, for ${\Greekmath 010E} =0.01$, we obtain an increase in the quantiles of the wage distribution between $0.22\%$ and $0.47\%$.

figure[figure omitted — 186 chars of source]

Conclusion

In this paper we study the unconditional policy effect with an endogenous binary treatment. Framing the selection equation as a threshold-crossing model allows us to introduce a novel class of unconditional marginal treatment effects and represent the unconditional effect as a weighted average of these unconditional marginal treatment effects. When the policy variable used to change the participation rate satisfies a conditional exogeneity condition, it is possible to recover the unconditional policy effect using the proposed UNIQUE method.

To illustrate the usefulness of unconditional MTEs, we focus on the unconditional quantile effect. We find that the unconditional quantile regression estimator that neglects endogeneity can be severely biased. Moreover, the bias may not be uniform across quantiles. Any attempt to sign the bias a priori requires very strong assumptions on the data-generating process. Intriguingly, the unconditional quantile regression estimator can be inconsistent even if the treatment status is exogenously determined. This happens when the treatment selection is partly determined by some covariates that also influence the outcome variable.

Our findings reveal that the unconditional quantile effect (UQE) and the marginal policy-relevant treatment effect (MPRTE) can be seen as part of the same family of policy effects. From a purely robustness perspective, the unconditional median effect---a special UQE---can be considered a more robust version of the MPRTE, in the same way that the median is the robust counterpart of the mean. Both represent specific examples of a general unconditional policy effect. To the best of our knowledge, this connection has not been established in the literature on either UQE or MPRTE.