EconBase
← Back to paper

Unconditional Effects of General Policy Interventions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

100,795 characters · 17 sections · 44 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Unconditional Effects of General Policy Interventions

abstractThis paper studies the unconditional effects of a general policy intervention, which includes location-scale shifts and simultaneous shifts as special cases. The location-scale shift is intended to study a counterfactual policy aimed at changing not only the mean or location of a covariate but also its dispersion or scale. The simultaneous shift refers to the situation where shifts in two or more covariates take place simultaneously. For example, a shift in one covariate is compensated at a certain rate by a shift in another covariate. Not accounting for these possible scale or simultaneous shifts will result in an incorrect assessment of the potential policy effects on an outcome variable of interest. The unconditional policy parameters are estimated with simple semiparametric estimators, for which asymptotic properties are studied. Monte Carlo simulations are implemented to study their finite sample performances. The proposed approach is applied to a Mincer equation to study the effects of changing years of education on wages and to study the effect of smoking during pregnancy on birth weight. Keywords: Location-scale shift, quantile regression, simultaneous shift, unconditional policy effect, unconditional regression. JEL: J01, J31.

\normalem

\onehalfspacing

Introduction

In many research areas, it is important to assess the distributional effects of covariates on an outcome variable. Several methods have been implemented in the literature to study this. A prolific line of research is a combination of conditional mean and quantile regression models together with micro simulation exercises, as in AutorKatzKearney05, MachadoMata05, and Melly05 (see FortinLemieuxFirpo11 for a review). A more recent and popular method is the recentered influence function (RIF) regression of FirpoFortinLemieux09, which directly estimates the effect of a change in the covariate distribution on a functional of the unconditional distribution of the outcome variable. The functional of interest can be the mean, quantile, or any other aspect of the unconditional distribution.

Consider, as an example, the unconditional quantile of the outcome variable $Y$. Let $F_{Y}$ be the unconditional distribution function of $Y,$ then the $\tau$-quantile of $F_{Y}$ is defined by \[ Q_{\tau}[Y]:=\arg\min\{q:\tau\leq F_{Y}(q)\}\ \ \mbox{for}\ \ \tau\in(0,1). \] In this paper, we seek to study how $Q_{\tau}[Y]$ changes when we induce an infinitesimal change in a covariate $X\in\mathbb{R}$, \ allowing the presence of other observable covariates $W$ and unobservable covariates collected in $U.$ These covariates and the outcome variable are related via a structural or causal function $h$ so that $Y=h(X,W,U)$. We consider a sequence of policy experiments that change $X$ into $X_{\delta}=\mathcal{G}(X;\delta)$ for a smooth function $\mathcal{G}(\cdot;\cdot)$. The policy experiments are indexed by $\delta$ satisfying $\mathcal{G}(X;0)=X.$ That is, $\delta=0$ corresponds to the status quo policy. With this induced change in $X$, the outcome variable becomes $Y_{\delta}=h(X_{\delta},W,U)=h(\mathcal{G}(X;\delta),W,U)$ where the distribution of $\left( X,W,U\right) $ is held constant. Our policy experiment has a ceteris paribus interpretation at the population level: we change $X$ into $X_{\delta}$ while holding the stochastic dependence among $X,W,$ and $U$ constant. Such a policy experiment is implementable if the covariate $X$ is not a causal factor for either $W$ or $U.$ In this case, when we intervene $X$ and change it into $X_{\delta}$, $W$ and $U$ will not change. This does not rule out the stochastic dependence among $X,$ $W,$ and $U.$ In the meanwhile, the structural function $h(\cdot,\cdot,\cdot)$ is also held constant. The main parameter of interest is the marginal effect of the change on the unconditional quantile of the outcome variable: \[ \Pi_{\tau}:=\lim_{\delta\rightarrow0}\frac{Q_{\tau}[Y_{\delta}]-Q_{\tau} [Y]}{\delta}. \]

FirpoFortinLemieux09 develop methods to study what corresponds to a location shift $X_{\delta}=X+\delta$. This shift affects the entire unconditional distribution of $Y=h(X,W,U)$, moving it towards a counterfactual distribution of $Y_{\delta}=h(X+\delta,W,U)$. One of the main results in FirpoFortinLemieux09 is that $\Pi_{\tau}$ can be represented as an average derivative: \[ \Pi_{\tau}=E\left[ \dot{\psi}_{x}\left( X,W\right) \right] , \] where \[ \dot{\psi}_{x}\left( x,w\right) =\frac{\partial E\left[ \psi\left( Y,\tau,F_{Y}\right) |X=x,W=w\right] }{\partial x}, \] $\psi\left( y,\tau,F_{Y}\right) =\left[ \tau-1\left\{ y\leq Q_{\tau }[Y]\right\} \right] /f_{Y}(Q_{\tau}[Y])$ is the influence function of the quantile functional, and $f_{Y}(Q_{\tau}[Y])$ is the unconditional density of $Y$ evaluated at the $\tau$-quantile $Q_{\tau}[Y].$ The unconditional quantile effect $\Pi_{\tau}$ can then be estimated by first running an unconditional quantile regression (henceforth, UQR), which involves regressing the influence function $\psi\left( Y_{i},\tau,F_{Y}\right) $ on the covariates $(X_{i},W_{i})$ and then taking an average of the partial derivatives of the regression function with respect to $X.$

The same method is applicable to other functionals of interest --- we only need to replace $\psi\left( y,\tau,F_{Y}\right) $ by the influence function underlying the functional we care about. This leads to the general RIF regression of FirpoFortinLemieux09. The potential simplicity and flexibility that the methodology offers motivates subsequent research to expand the use of RIF regressions. On the empirical side, after its introduction, RIF regressions became a popular method for analyzing and identifying the distributional effects on outcomes in terms of changes in observed characteristics in areas such as labor economics, income and inequality, health economics, and public policy. On the theoretical side, Rothe2012 provides a generalization of FirpoFortinLemieux09 for the case of location shifts, and, more recently, SasakiUraZhang20 study the high-dimensional setting while InoueLiXu21 focus on the two-sample problem. An alternative estimation procedure is proposed in cquq2023.

This paper extends the UQR and RIF regression in several ways. First, we study general counterfactual policy changes, of which the location shift is a special case. Our framework allows for any smooth and invertible intervention of the target covariates. As a complement to the existing literature that focuses on changing the marginal distribution of the target covariates, we consider changing the values of the target covariates directly. An advantage of our approach is that the changes under consideration are directly implementable. We note that it may not be easy to induce a desired shift in the marginal distribution, and when possible, such a shift is often achieved via transforming the target covariates, which is what we consider here.

Second, we provide extensive discussions of a counterfactual policy that, in addition to the location shift, affects the scale of a covariate. For example, we may consider $X_{\delta}=X/(1+\delta)+\delta.$ We find that in this case, the marginal effect can be decomposed as the sum of two effects: one related to the location shift and the other related to the scale shift. In order to interpret the scale effect, we introduce the quantile-standard deviation elasticity: the percentage change in the unconditional quantiles of the outcome variable induced by a 1% change in the standard deviation of the target covariate.

Third, we allow the target covariates to be endogenous, and we characterize the asymptotic bias of the unconditional effect estimator when the endogeneity is not appropriately accounted for. We eliminate the endogeneity bias using a control variable/function approach. Such an approach is analogous to the method of causal inference under the unconfoundedness assumption.

Fourth, by letting the policy function depend on covariates so that $X_{\delta}=\mathcal{G}(X,W;\delta)$, we allow the interventions to vary across covariate-specific strata. In the Supplemental Appendix, we also consider the case of simultaneous shifts in different covariates. We focus on the case of simultaneous location shifts in two covariates. This happens when a location shift in one covariate induces a location shift in another covariate at the same time. For example, $Y=h\left( X_{1},X_{2},W,U\right) $ for two scalar target covariates $X_{1}$ and $X_{2}$, and the policy induces $X_{1\delta}=X_{1}+\delta$ and $X_{2\delta}=X_{2}-\delta$. Our approach can easily accommodate this case, and we show that the simultaneous effect can be obtained as a linear combination of individual effects obtained by considering one change at a time.

Finally, we propose consistent and asymptotically normal semiparametric estimators of the location-scale effect and the simultaneous effect. The estimators can be easily implemented in empirical work using either a probit or logit specification of the conditional distribution function. We conduct an extensive Monte Carlo study evaluating the finite sample performances of the location-scale effect estimator and the accuracy of the normal approximation. Simulation results show that the estimator works reasonably well under different specifications and that the standard normal distribution provides a good approximation to the finite sample distribution of a studentized test statistic introduced in this paper.

As potential applications of our proposed approach, consider the following empirical examples to motivate its use.

exampleEffect of increasing education on wage inequality. In a Mincer equation, log wages are modeled as a function of certain observable covariates such as years of education. A study of the effect of a shift in education on wage inequality could be implemented using our proposed framework. We can accommodate a counterfactual policy experiment where there may be not only a general increase in the education level but also a change in its dispersion.
exampleSmoking and birth weight. Consider a tax levied on the consumption of cigarettes. It is reasonable to think that the consumption $X$ will be reduced to $X/(1+\delta)$, where $\delta$ is the tax burden on the consumer. Thus, the tax induces a reduction in the level and dispersion of cigarette consumption. We will use the proposed method to assess its effect on the distribution of birth weights.
exampleWage controls and earnings distribution During War World II, the National War Labor Board imposed wage controls in the form of brackets: wages below the bracket were allowed to rise, while wages above the bracket were not allowed to rise. Importantly, these brackets differed across industries, occupations, and regions. ziebarth2022 use the tools developed in this paper to analyze the effect of a more uniform (less dispersion) distribution of brackets on the distributions of earnings.
exampleTrade integration and skill distribution Gu2020 document the impact of trade integration on both the mean and the standard deviation of the skill distribution across municipalities in Denmark. Moreover, as argued by Hanushek2008, skills are related to income distribution. Thus, a quantification of the impact of a scale effect in the skills distribution on the quantiles of the income distribution appears to be relevant.
exampleDays in a job training program. SasakiUraZhang20 develop high-dimensional UQR to analyze the effect on wages of counterfactual increase in: $(i)$ the days of participation in a job training program; and $(ii)$ the days actually taking classes in the same job training program. Our simultaneous effect analysis can consider, for example, a reduction in $(i)$ with a simultaneous increase in $(ii)$. Thus, our paper can be used to study the effect of a more concentrated job training program.

We illustrate the proposed method with two empirical applications. The first one is related to Example (ref): the effect of changing education on wage inequality, decomposing it into location and scale effects. Empirical results reveal the contrasting nature of the two effects. The location effects are seen to be positive and relatively similar across quantiles. On the other hand, the scale effects are highly heterogeneous and monotonically decreasing across quantiles. Hence, the scale effects can more than offset the location effects. This shows that not accounting for both shifts may result in a biased assessment of the policy effects on the quantiles of the outcome variable. The second application is related to Example (ref) where we estimate the unconditional effects of smoking during pregnancy on the birth weight. The effects from reducing the mean and variance of the\ number of cigarettes smoked are positive and are different for different quantiles of the birth weight distribution.

The paper is organized as follows. Section (ref) studies the unconditional effects of general policy interventions with the location-scale shift as the main example. Section (ref) provides some further discussion on the methodological contribution of this paper relative to FirpoFortinLemieux09. Section (ref) describes the estimator of the location-scale effect and studies its asymptotic properties. Section (ref) reports the finite sample performance of the location-scale effect estimator and the associated tests. Section (ref) presents the empirical applications. Section (ref) concludes. The proofs are in the Appendix. The case of simultaneous changes and the details for a theoretical example are given in the Supplementary Appendix.

A word on notation: we use $F_{Y|X}(y|x)$ and $f_{Y|X}(y|x)$ to denote the cumulative distribution function and the probability density function of $Y,$ respectively, conditional on $X=x$. For a random variable $Z$, the unconditional $\tau$-quantile is denoted by $Q_{\tau}[Z]$, i.e., $\Pr(Z\leq Q_{\tau}[Z])=\tau$, and its variance is denoted by $var(Z).$ For a pair of random variables $Z_{1}$ and $Z_{2}$, the conditional quantile is denoted by $Q_{\tau}[Z_{1}|z_{2}]$, i.e., $\Pr(Z_{1}\leq Q_{\tau }[Z_{1}|z_{2}]|Z_{2}=z_{2})=\tau$. We adopt the following notational conventions: \[ \frac{\partial E(Z|X)}{\partial X}=\left. \frac{\partial E\left( Z|X=x\right) }{\partial x}\right\vert _{x=X},\text{ }\frac{\partial F_{Z|X}(z|X)}{\partial X}=\left. \frac{\partial F_{Z|X}(z|X=x)}{\partial x}\right\vert _{x=X}. \] For a column vector $v,$ $d_{v}$ stands for the number of elements in $v.$

Unconditional effects of general policy interventions

Introducing location-scale shifts

We start with a general structural model $Y=h(X,W,U)$, where the function $h$ is unknown, and we only observe $(X,W)$ and $Y$. Here $X$ is univariate but the dimension of $W$ is left unrestricted. All the unobserved causal factors of $Y$ are collected in $U$. We are concerned with the effect on the distribution of $Y$ of general (infinitesimal) changes in $X$, the target variable.

Perhaps the simplest example of a counterfactual change in $X$ is a location shift: $X_{\delta}=X+\delta$. The popular method of UQR of FirpoFortinLemieux09 can be used to assess the effect of such changes in the unconditional quantiles of $Y$.\footnote{See Section (ref) for a discussion about how this paper relates to FirpoFortinLemieux09.} In this paper we provide results for the general case where $X_{\delta}=\mathcal{G}(X;\delta)$ for some (suitable) policy function $\mathcal{G}$ chosen by the researcher or policy maker. A counterfactual change in $X$ to $X_{\delta}=\mathcal{G}(X;\delta)$ induces a counterfactual outcome $Y_{\delta}=h(\mathcal{G}(X;\delta),W,U)=h(X_{\delta },W,U)$. Our parameter of interest, the marginal effect for the $\tau$-quantile, is an infinitesimal contrast of unconditional quantiles and is defined as

equation[equation omitted — 120 chars of source]

whenever this limit exists.

A particular policy function that we analyze in detail is the following location-scale shift in $X$

equation[equation omitted — 125 chars of source]

Here, $\mu$ is a known policy parameter, and we refer to $\ell (\delta)$ as the location shift and to $s(\delta)>0$ as the scale shift. In order to take limits to find $\Pi_{\tau}$, we assume that $\ell(\delta)$ and $s(\delta)$ are continuously differentiable functions of the scalar $\delta$. Both $\ell(\delta)$ and $s(\delta)$ are chosen by the researcher or policy maker subject to the restriction that $s(0)=1$ and $\ell(0)=0$. Note that this choice of $\mathcal{G}$ nests the case $X_{\delta}=X+\delta$ by choosing $\ell(\delta)=\delta$ and $s(\delta)\equiv1$.

A distinctive feature of $X_{\delta}$ in (ref) is that \[ var[X_{\delta}]=s(\delta)^{2}var[X], \] and so it allows for the study of counterfactual changes in the dispersion of the target variable. To see this, suppose that $s(\delta)<1$, then, realizations of $X$ that are above/below $\mu$ are \textquotedblleft moved\textquotedblright\ towards $\mu$, followed by a location shift of $\ell(\delta)$. Therefore, we have a constant location shift, given by $\ell(\delta)$, and a relative location shift induced by the scale shift, which tends to bunch observations near $\mu$. The result is a reduction of the variance of $X$. If, on the other hand, $s(\delta)>1$, then the counterfactual policy moves $X$ away from $\mu$ and consequently increases its variance.

Under some regularity assumptions spelled below, the marginal effect $\Pi_{\tau}$ corresponding to the policy function $\mathcal{G}$ given in (ref) can be decomposed into the sum of two effects: one associated with the location shift governed by $\ell(\delta)$, and the other associated with the scale shift $s(\delta)$. The former corresponds to a version of the estimand studied by FirpoFortinLemieux09. The latter effect is, to the best of our knowledge, new.

Subsection (ref) contains a rigorous development of our main results for a general policy function. Readers interested in the location-scale shift only can skip subsection (ref) and focus on subsections (ref) and (ref) where we provide the specific results for the location-scale shift, discuss their interpretations, and offer examples.

Results for a general policy function

Central to our results is the counterfactual policy function $\mathcal{G}$, which maps $X$ to $X_{\delta}$ and generates a counterfactual outcome $Y_{\delta}$. As mentioned before, our parameter of interest $\Pi_{\tau}$ given in (ref), compares the quantiles of

equation[equation omitted — 45 chars of source]

to the quantiles of

equation[equation omitted — 93 chars of source]

An important assumption is that the distribution of $\left( X,W,U\right) $ in ((ref)) is held the same as that in ((ref)). To understand the latter condition, we can consider two parallel worlds: the worlds before and after the intervention. For each given $\delta,$ let $\mathcal{G}^{-1}(x;\delta)$ be the inverse function of $\mathcal{G} (x;\delta)$ such that $\mathcal{G}(\mathcal{G}^{-1}(x;\delta);\delta)=x.$ After applying the inverse transform to the target covariate in the post-intervention world, the distribution of $\left( \mathcal{G} ^{-1}\mathcal{(}X^{\delta};\delta),W^{\delta},U^{\delta}\right) $ in the post-intervention world is assumed to be the same as that of $\left( X,W,U\right) $ in the pre-intervention world. Here, no change is induced on $W$ and $U$ and so $\left( W^{\delta},U^{\delta}\right) $ is actually the same as $(W,U)$ for every individual in the population. In essence, we keep the structural function $h\left( \cdot,\cdot ,\cdot\right) $ and the distribution of $\left( X,W,U\right) $ intact during the policy intervention. The effect under consideration is then the policy effect due to the policy intervention only and thus has a ceteris paribus causal interpretation.

For notational economy, we write $x^{\delta}=\mathcal{G}^{-1}(x;\delta)$. Then $X_{\delta}=x$ if and only if $X=x^{\delta}.$ Define the Jacobian of the inverse transform $x\mapsto x^{\delta}:=\mathcal{G}^{-1}(x;\delta)$ as \[ J(x^{\delta};\delta):=\frac{\partial x^{\delta}}{\partial x}=\left[ \frac{\partial\mathcal{G}\left( x;\delta\right) }{\partial x}\right] ^{-1}\bigg|_{x=x^{\delta}}. \] Then, the joint probability density functions\ of the covariate vector before and after the intervention satisfy \[ f_{X_{\delta},W}(x,w)=J(x^{\delta};\delta)\cdot f_{X,W}(x^{\delta},w). \]

For $\varepsilon>0$, define $\mathcal{N}_{\varepsilon}:=\left\{ \delta:\left\vert \delta\right\vert \leq\varepsilon\right\} $. We maintain the following assumption.

assumption(i.a) For some $\varepsilon>0,$ $\mathcal{G}\left( x;\delta\right) $ is continuously differentiable on $\mathcal{X\otimes N}_{\varepsilon}$, where $\mathcal{X}$ is the support of $X.$ (i.b) $\mathcal{G}\left( x;\delta\right) $ is strictly increasing in $x$ for each $\delta\in\mathcal{N}_{\varepsilon}.$ (i.c) $\mathcal{G}\left( x;0\right) =x$ for all $x\in\mathcal{X}$. (ii) for $\delta\in\mathcal{N}_{\varepsilon}$, the conditional density of $U$ satisfies $f_{U|X_{\delta},W}(u|x,w)=f_{U|X,W}(u|x^{\delta},w)$, and the support $\mathcal{U}$ of $U$ given $X$ and $W$ does not depend on $\left( X,W\right) .$ (iii.a) $x\mapsto f_{X,W}(x,w)$ is continuously differentiable for all $w\in\mathcal{W}$ and \[ \int_{\mathcal{W}}\int_{\mathcal{X}}\sup_{\delta\in\mathcal{N}_{\varepsilon} }\left\vert \frac{\partial\left[ J\left( x^{\delta};\delta\right) f_{X,W}(x^{\delta},w)\right] }{\partial\delta}\right\vert dxdw<\infty \] where $\mathcal{W}$ is the support of $W.$ (iii.b) $x\mapsto f_{U|X,W}(u|x,w)$ is continuously differentiable for all $\left( u,w\right) $ and \begin{align*} \int_{\mathcal{W}}\int_{\mathcal{X}}\int_{\mathcal{U}}\sup_{\delta \in\mathcal{N}_{\varepsilon}}\left\vert \frac{\partial}{\partial\delta}\left[ f_{U|X,W}(u|x^{\delta},w)f_{X,W}(x^{\delta},w)\right] \right\vert dudxdw & <\infty,\\ \int_{\mathcal{W}}\int_{\mathcal{X}}\int_{\mathcal{U}}\sup_{\delta \in\mathcal{N}_{\varepsilon}}\left\vert \frac{\partial f_{X,W}(x^{\delta} ,w)}{\partial\delta}\right\vert f_{U|X,W}(u|x,w)dudxdw & <\infty. \end{align*} (iv) $f_{X,W}(x,w)$ is equal to $0$ on the boundary of the support of $X$ given $W=w$ for all $w\in\mathcal{W}.$ (v) $f_{Y}(Q_{\tau}[Y])>0.$
remarkAssumption (ref)(i) imposes some restrictions on the policy function $\mathcal{G}\left( x;\delta\right) .$ It is reasonable that $\mathcal{G}\left( x;\delta\right) $ is strictly increasing in $x,$ as a non-monotonic and non-invertible function does not seem to be practically relevant. The strictly increasing property implies that $J\left( x;\delta\right) >0$ for all $x\in\mathcal{X}$ and $\delta\in\mathcal{N} _{\varepsilon}.$ The condition that $\mathcal{G}\left( x;0\right) =x$ says that there is no intervention when $\delta=0,$ and it implies that $J\left( x;0\right) =1$ for all $x\in\mathcal{X}.$ Assumption (ref) (ii) assumes that how $U$ depends on the covariate vector is maintained when we induce a change in the covariate vector. Note that Assumption (ref)(ii) is different from $f_{U|X_{\delta},W} (u|x,w)=f_{U|X,W}(u|x,w)$, which in general cannot hold when $U$ depends on $X$ and $W$. The counterfactual model in ((ref)) says that we maintain the structure of the causal system. Assumption (ref) (ii) says that we also maintain how the unobservable depends on the observables. As discussed above, we also implicitly assume that $\left( \mathcal{G}^{-1}\mathcal{(}X^{\delta};\delta),W^{\delta}\right) $ has the same distribution as $\left( X,W\right) .$ The rest of Assumption (ref) consists of regularity conditions.
remarkAssumption (ref) does not assume that $U$ is independent of $\left( X,W\right) .$ It does not assume that $U$ is conditionally independent of $X$ given $W$ either. Assumption (ref) below will impose identification assumptions.

The following theorem characterizes the effects of the policy change on the distribution of $Y_{\delta}$ and its quantiles.

theoremLet Assumption (ref) hold. (i) For each $\left( x,w\right) \in\mathcal{X}\otimes\mathcal{W},$ \[ \lim_{\delta\rightarrow0}\frac{f_{X_{\delta},W}(x,w)-f_{X,W}(x,w)}{\delta }=-\frac{\partial}{\partial x}\left[ \kappa\left( x\right) f_{X,W} (x,w)\right] , \] where \[ \kappa\left( x\right) :=\frac{\partial\mathcal{G}(x;\delta)}{\partial\delta }\bigg|_{\delta=0}. \] (ii) As $\delta\rightarrow0$, we have \begin{align*} & \frac{F_{Y_{\delta}}(y)-F_{Y}\left( y\right) }{\delta}\\ & \rightarrow E\left[ \left( \frac{\partial F_{Y|X,W}(y|X,W)}{\partial X}-\mathds1\left\{ h(X,W,U)\leq y\right\} \frac{\partial\ln f_{U|X,W} (U|X,W)}{\partial X}\right) \kappa\left( X\right) \right] \end{align*} uniformly in $y\in\mathcal{Y}$, the support of $Y$. (iii) The marginal effect of the intervention $X_{\delta}=\mathcal{G} (X;\delta)$ on the $\tau$-quantile of the outcome variable $Y$ can be represented by \begin{equation} \Pi_{\tau}=A_{\tau}-B_{\tau} \end{equation} where \begin{align*} A_{\tau} & =E\left[ \frac{\partial E\left[ \psi\left( Y,\tau ,F_{Y}\right) |X,W\right] }{\partial X}\kappa\left( X\right) \right] ,\\ B_{\tau} & =E\left[ \psi\left( Y,\tau,F_{Y}\right) \frac{\partial\ln f_{U|X,W}(U|X,W)}{\partial X}\kappa\left( X\right) \right] , \end{align*} and \[ \psi\left( y,\tau,F_{Y}\right) =\frac{\tau-1\left( y<Q_{\tau}[Y]\right) }{f_{Y}(Q_{\tau}[Y])}. \]
remarkTo understand Theorem (ref)(i), we can write \[ f_{X_{\delta},W}(x,w)-f_{X,W}(x,w)=f_{X_{\delta},W}(x,w)-f_{X,W}(x^{\delta },w)+f_{X,W}(x^{\delta},w)-f_{X,W}(x,w). \] It is quite intuitive that the second term is approximately $\delta\cdot \frac{\partial x^{\delta}}{\partial\delta}|_{\delta=0}\cdot\frac{\partial f_{X,W}(x,w)}{\partial x}=-\delta\cdot\kappa\left( x\right) \cdot \frac{\partial f_{X,W}(x,w)}{\partial x}$ when $\delta$ is small. Here we have used the result that $\kappa\left( x\right) $ also equals $-\frac{\partial x^{\delta}}{\partial\delta}|_{\delta=0}$ (see the proof of Theorem (ref) in the appendix). The first term reflects the effect from the Jacobian of the transformation. Indeed, $f_{X_{\delta},W}(x,w)-f_{X,W} (x^{\delta},w)=\left[ J(x^{\delta};\delta)-J(x^{\delta};0)\right] f_{X,W}(x^{\delta},w)$ as $J(x^{\delta};0)=1.$ The first term is then approximately equal to $\delta\cdot f_{X,W}(x,w)\cdot\frac{\partial J\left( x,\delta\right) }{\partial\delta}\big|_{\delta=0}.$ But \[ \frac{\partial J\left( x,\delta\right) }{\partial\delta}\bigg|_{\delta =0}=\frac{\partial}{\partial\delta}J\left( x,\delta\right) \bigg|_{\delta =0}=\frac{\partial}{\partial\delta}\frac{\partial x^{\delta}}{\partial x}\bigg|_{\delta=0}=\frac{\partial}{\partial x}\frac{\partial x^{\delta} }{\partial\delta}\bigg|_{\delta=0}=-\frac{\partial\kappa\left( x\right) }{\partial x}, \] and hence the first term is approximately $-\delta\cdot f_{X,W}(x,w)\cdot \frac{\partial\kappa\left( x\right) }{\partial x}.$ Combining these two approximations yields Theorem (ref)(i).
remarkBy definition, $\kappa\left( x\right) $ measures the marginal change of $\mathcal{G}(x;\delta)$ as we increase $\delta$ from zero infinitesimally. Theorems (ref) (ii) and (iii) show that only $\kappa\left( x\right) $ appears in the marginal effect and the Jacobian does not. This is not surprising, as what matters for the marginal effect is the marginal change in the policy function.
remarkTheorem (ref)(iii) represents the structural parameter $\Pi_{\tau}$ in terms of statistical objects. While the first term $A_{\tau}$ is identifiable, the second term $B_{\tau}$, which involves the conditional density of $U$ given $X$ and $W,$ is not. If we use $\hat{A}_{\tau},$ a consistent estimator of $A_{\tau},$ as an estimator of $\Pi_{\tau},$ then the second term $B_{\tau}$ is the asymptotic bias of $\hat{A}_{\tau}.$ This bias is an endogeneity bias, as it is in general not equal to zero when $X$ is not independent of $U$ (conditioning on $W$). Similar results have been established in yixiao2021 but only for location shifts. If we do not have the identification condition such as what is given in Assumption (ref) below, Theorem (ref)(iii) allows us to use a bound approach to bound $B_{\tau}$ and infer the range of the policy effect or conduct a sensitivity analysis similar to that in martinez2020.
remarkWhile the paper focuses on the quantile functional, Theorem (ref)(iii) is formulated in a general way. The result holds for any Hadamard differentiable functional and for the mean functional. We only need to replace $\psi\left( y,\tau,F_{Y}\right) $ by the influence function of the functional that we are interested in. For example, for the mean functional, we can replace $\psi\left( y,\tau,F_{Y}\right) $ by $y-E(Y)$, and Theorem (ref)(iii) remains valid.

To identify $\Pi_{\tau},$ we make the following independence or conditional independence assumption.

assumptionFor $\delta\in\mathcal{N}_{\varepsilon}$, the unobservable $U$ satisfies either $f_{U|X,W}(u|x,w)=f_{U|X,W}(u|x^{\delta },w)=f_{U}(u)$ or $f_{U|X,W}(u|x,w)=f_{U|X,W}(u|x^{\delta},w)=f_{U|W}(u|w).$

Under the above assumption, $\partial\ln f_{U|X,W}(u|x,w)/\partial x=0$ and the second term $B_{\tau}$ in ((ref)) vanishes. In this case, $\Pi_{\tau}=A_{\tau}$ and hence is identified. The corollary below then follows directly from Theorem (ref)(iii).

corollaryLet Assumption (ref) hold with Assumption (ref) (ii) strengthened to Assumption (ref). Then \ \begin{equation} \Pi_{\tau}=E\left[ \frac{\partial E\left[ \psi\left( Y,\tau,F_{Y}\right) |X,W\right] }{\partial X}\kappa\left( X\right) \right] =\frac{1} {f_{Y}(Q_{\tau}[Y])}E\left[ \frac{\partial\mathcal{S}_{Y|X,W}\left( Q_{\tau }[Y]|X,W\right) }{\partial X}\kappa\left( X\right) \right] \end{equation} where $\mathcal{S}_{Y|X,W}\left( \cdot|x,w\right) :=1-F_{Y|X,W}\left( \cdot|x,w\right) $ is the\ conditional survival function.
remarkBoth conditions in Assumption (ref) require that $f_{U|X,W} (u|x,w)=f_{U|X,W}(u|x^{\delta},w)$. This is related to the assumption in FirpoFortinLemieux09, framed as \textquotedblleft maintaining the conditional distribution of Y given X unaffected.\textquotedblright\ In essence, FirpoFortinLemieux09 requires\ $f_{U|X}(u|x)=f_{U|X}(u|x^{\delta}).$ When this condition fails, we may still have $f_{U|X,W}(u|x,w)=f_{U|X,W}(u|x^{\delta},w).$ Such a condition has also been used in Lieli2020 and Spini2021 in a context of extrapolation to populations with different distributions of the covariates.
remarkThe first condition in Assumption (ref) is satisfied if $U$ is independent of $(X,W)$. In our view, this condition is hard to achieve in empirical applications. The second condition in Assumption (ref), which is commonly used to achieve identification in applied work, is a conditional independence assumption. Such a condition is often referred to as the unconfoundedness condition in the causal inference literature. {The assumption is more general than $Y(x)\perp X|W$ for any }$x\in\mathcal{X}.${ It is a \textquotedblleft local\textquotedblright\ unconfoundedness condition in the sense that $Y(x)|\left( X,W\right) =_{d}Y(x^{\delta})|$}$\left( {X,W}\right) ${ for $\delta\in\mathcal{N}_{\varepsilon}$, a small $\varepsilon$-radius neighborhood around 0. A more stringent condition would require $Y(x)|\left( X,W\right) =_{d}Y(\tilde{x})|$}$\left( {X,W}\right) ${ for any $x,\tilde {x}\in$}$\mathcal{X}${. }

We note in passing that Corollary (ref) has the following alternative representation: \[ \Pi_{\tau}=\left\langle E\left[ \frac{\partial E\left[ \psi\left( Y,\tau,F_{Y}\right) |X,W\right] }{\partial X}\bigg| X\right] ,\frac {\partial\mathcal{G}(X;\delta)}{\partial\delta}\bigg|_{\delta=0}\right\rangle , \] where $\left\langle \cdot,\cdot\right\rangle $ is the inner product defined by $\left\langle h\left( X\right) ,g\left( X\right) \right\rangle :=E\left[ h\left( X\right) g\left( X\right) \right] $ in the space $\mathcal{L} _{2}\left( X\right) .$ By the Cauchy-Schwarz inequality, $|\langle h\left( X\right) ,g\left( X\right) \rangle|\leq\left\Vert h\left( X\right) \right\Vert \left\Vert g\left( X\right) \right\Vert $, where $\left\Vert \cdot\right\Vert $ is the norm defined by $\left\Vert h\left( X\right) \right\Vert :=\sqrt{\left\langle h\left( X\right) , h\left( X\right) \right\rangle }$. Consider the class of policy functions with a unit norm, namely $\left\Vert \frac{\partial\mathcal{G}(X;\delta)}{\partial\delta }\bigg|_{\delta=0}\right\Vert =1.$ Then \[ |\Pi_{\tau}|\leq\left\Vert E\left[ \frac{\partial E\left[ \psi\left( Y,\tau,F_{Y}\right) |X,W\right] }{\partial X}\bigg|X\right] \right\Vert . \] Thus, if a policy function satisfies \[ \frac{\partial\mathcal{G}(X;\delta)}{\partial\delta}\bigg|_{\delta=0}=E\left[ \frac{\partial E\left[ \psi\left( Y,\tau,F_{Y}\right) |X,W\right] }{\partial X}\bigg|X\right] \cdot\left\Vert E\left[ \frac{\partial E\left[ \psi\left( Y,\tau,F_{Y}\right) |X,W\right] }{\partial X}\bigg|X\right] \right\Vert ^{-1}, \] then it achieves the highest $\Pi_{\tau}$ (in magnitude) in this class. We leave optimal policy designs based on a cost-benefit analysis for future research.

Results for the location-scale shift

In this subsection we obtain a representation for $\Pi_{\tau}$ for the particular case of the location-scale shift given in ((ref) ): \[ X_{\delta}=\mathcal{G}(X;\delta) = \left( X-\mu\right) s(\delta)+\mu +\ell(\delta). \] The corollary below also follows directly from Theorem (ref)(iii).

corollaryLet Assumption (ref) hold with Assumption (ref) (ii) strengthened to Assumption (ref). Then, for the location-scale shift in ((ref)) with $\ell(0)=0,$ $s(0)=1,$ and $s(\delta)>0,$ the marginal effect can be decomposed as \begin{equation} \Pi_{\tau}=\Pi_{\tau,L}+\Pi_{\tau,S}, \end{equation} where \begin{align*} \Pi_{\tau,L} & =\frac{\dot{\ell}\left( 0\right) }{f_{Y}(Q_{\tau}[Y])} \int_{\mathcal{W}}\int_{\mathcal{X}}\frac{\partial\mathcal{S}_{Y|X,W}(Q_{\tau }[Y]|x,w)}{\partial x}f_{X,W}(x,w)dxdw,\\ \Pi_{\tau,S} & =\frac{\dot{s}\left( 0\right) }{f_{Y}(Q_{\tau}[Y])} \int_{\mathcal{W}}\int_{\mathcal{X}}\frac{\partial\mathcal{S}_{Y|X,W}(Q_{\tau }[Y]|x,w)}{\partial x}\left( x-\mu\right) f_{X,W}(x,w)dxdw, \end{align*} and $\mathcal{S}_{Y|X,W}\left( \cdot|x,w\right) :=1-F_{Y|X,W}\left( \cdot|x,w\right) $ is the\ conditional survival function.

Corollary (ref) shows that the overall effect $\Pi_{\tau}$ can be decomposed into the sum of $\Pi_{\tau,L}$ and $\Pi _{\tau,S}.$ Here $\Pi_{\tau,L}$ is the location effect and is the estimand in FirpoFortinLemieux09 when we set $\dot{\ell}(0)=1$ and $s\left( \delta\right) \equiv1$. $\Pi_{\tau,S}$ is the scale effect and is present whenever $s(\delta)$ is not identically 1 and $\dot{s}\left( 0\right) \neq 0$.\footnote{It can be seen that $\Pi_{\tau,S}$ depends on $\mu$. However, we suppress this dependence from the notation for simplicity.}

To better understand the location and scale effects in Corollary (ref), consider the case that $X$ and $U$ are independent and there is no $W.$ Then

align[align omitted — 398 chars of source]

To sign the location effect $\Pi_{\tau,L}$, we can assess whether $\mathcal{S}_{Y|X}(Q_{\tau}[Y]|x)$ is increasing in $x$ or not. If $\dot{\ell }\left( 0\right) \geq0$ and $\mathcal{S}_{Y|X}(Q_{\tau}[Y]|x)$ is increasing in $x$ on average, more precisely, $\int_{\mathcal{X}}\frac{\partial \mathcal{S}_{Y|X}(Q_{\tau}[Y]|x)}{\partial x}f_{X}(x)dx\geq0,$ then $\Pi _{\tau,L}\geq0.$ As an example, consider the case that $h\left( x,u\right) $ is increasing in $x$ for each $u.$ Then, $\mathcal{S}_{Y|X}(Q_{\tau}[Y]|x)$ is increasing in $x$ for all $x\in\mathcal{X}$, and so $\Pi_{\tau,L}\geq0$ if $\dot{\ell}\left( 0\right) \geq0.$

It is a bit more challenging to determine the sign of the scale effect $\Pi_{\tau,S}$, which depends on, not only the function form of $\frac {\partial\mathcal{S}_{Y|X}(Q_{\tau}[Y]|x)}{\partial x}$, but also the distribution of $X.$ The next example provides some insight into the scale effect.

exampleNormal Covariate. Consider the linear model $Y=\lambda+X\gamma+U$ where $X$ and $U$ are independent and $X\sim N(\mu _{X},\sigma_{X}^{2})$. We can use Stein's lemma (see, for example, casella and references therein) to gain some insight into the scale effect. The lemma states that for a differentiable function $m\left( \cdot\right) $ such that $E[|m^{\prime}(X)|]<\infty$, $E[m(X)(X-\mu_{X})]=\sigma^{2}E[m^{\prime}(X)]$ whenever $X\sim N(\mu _{X},\sigma_{X}^{2})$. Taking $m(x)=\partial\mathcal{S}_{Y|X}(Q_{\tau }[Y]|x)/\partial x$ and using Stein's lemma, we can express the scale effect for $\mu=\mu_{X}$ as \[ \Pi_{\tau,S}=\frac{\dot{s}\left( 0\right) }{f_{Y}(Q_{\tau}[Y])}E\left[ \frac{\partial\mathcal{S}_{Y|X}(Q_{\tau}[Y]|X)}{\partial X}\left( X-\mu _{X}\right) \right] =\frac{\dot{s}\left( 0\right) \sigma_{X}^{2}} {f_{Y}(Q_{\tau}[Y])}E\left[ \frac{\partial^{2}\mathcal{S}_{Y|X}(Q_{\tau }[Y]|X)}{\partial X^{2}}\right] . \] Therefore, when $X$ is normal and $\dot{s}(0)>0,$ the scale effect is non-negative (non-positive) if $\mathcal{S}_{Y|X}(Q_{\tau}[Y]|x)$ is a convex (concave) function of $x$. It is interesting to see that the location effect depends on the first order derivative of $\mathcal{S}_{Y|X}(Q_{\tau}[Y]|x)$ (see equation ((ref))) while the scale effect depends on its second-order derivative.

In the next example, we simplify $\Pi_{\tau,S}$ under the additional assumption that $U$ is also normal.

exampleNormal Covariate and Normal Noise. Consider a linear model $Y=\lambda +X\gamma+U$ where $X$ and $U$ are independent. We have: $\Pi_{\tau,L} =\dot{\ell}\left( 0\right) \gamma$. In addition to the normal covariate assumption $X\sim N(\mu_{X},\sigma^{2}_{X})$, suppose $U$ is also normal $U\sim N(0,\sigma_{U}^{2})$. Then, for $\mu=\mu_{X},$ \[ \Pi_{\tau,S}=\dot{s}\left( 0\right) \sqrt{R_{YX}^{2}}Q_{\tau}[X_{\gamma }^{\circ}] \] where $R_{YX}^{2}$ is the population R-squared defined by $R_{YX} ^{2}:=var(\lambda+X\gamma)/var(Y)$ and $X_{\gamma}^{\circ}=\left( X-\mu _{X}\right) \gamma.$ While the location effect is constant across quantiles, the scale effect varies across quantiles.

In Example (ref), the scale effect, when $\mu=\mu_{X}$, does not depend on $\mu_{X}$ or sign$\left( \gamma\right) $. To understand this and obtain a more general result, we note that $\Pi_{\tau,S}$ is proportional to the following covariance:

align[align omitted — 288 chars of source]

where $U^{\circ}:=U+X_{\gamma}^{\circ}$ and we have used

align*[align* omitted — 184 chars of source]

Now, if $X-\mu_{X}$ is symmetrically distributed around zero, then $X_{\gamma }^{\circ}$ also shares this property. In this case, the covariance in ((ref)) does not depend on sign$\left( \gamma\right) $ as the distributions of $X_{\gamma}^{\circ}$ and $U^{\circ}$ remain the same if we flip the sign of $\gamma.$ Also, since the distributions of $X_{\gamma}^{\circ}$ and $U^{\circ}$ do not depend on $\mu_{X},$ the covariance in ((ref)) does not depend on $\mu_{X}.$ On the other hand, for the denominator of the scale effect, we have \[ f_{Y}(Q_{\tau}[Y])=f_{Y}(Q_{\tau}[U^{\circ}]+\delta+\mu_{X}\gamma )=f_{U^{\circ}}\left( Q_{\tau}[U^{\circ}]\right) . \] If $X-\mu_{X}$ is symmetrically distributed around zero, then the distribution of $U^{\circ}$ does not depend on $\mu_{X}$ or sign$\left( \gamma\right) .$ Hence, $f_{Y}(Q_{\tau}[Y])$ does not depend on $\mu_{X}$ or sign$\left( \gamma\right) .$

Since both the numerator and the denominator of $\Pi_{\tau,S}$ are invariant to $\mu_{X}$ and sign$\left( \gamma\right) $, we obtain the following proposition immediately.

propositionConsider the linear model $Y=\lambda+X\gamma+U$ where $X$ and $U$ are independent. If $X-E\left[ X\right] $ is symmetrically distributed around zero, then the scale effect computed for $\mu=E\left[ X\right] $ does not depend on either $E\left[ X\right] $ or sign$\left( \gamma\right) .$

Interpretation of the scale effects

Consider a situation where we only care about the scale effect, that is, we set $\ell(\delta)\equiv0$. Then, we have $X_{\delta}=\mu+\left( X-\mu\right) s(\delta)$. If we denote by $\sigma_{X}$ and $\sigma_{X_{\delta}}$ the standard deviations of $X$ and $X_{\delta}$, respectively, then $\sigma _{X_{\delta}}=\sigma_{X}s\left( \delta\right) $. To interpret $\Pi_{\tau,S} $, we assume $Q_{\tau}[Y_{\delta}]\neq0$ and consider the following quantile-standard deviation elasticity \[ \mathcal{E}_{\tau,\delta}:=\frac{dQ_{\tau}[Y_{\delta}]}{Q_{\tau}[Y_{\delta} ]}\left( \frac{d\sigma_{X_{\delta}}}{\sigma_{X_{\delta}}}\right) ^{-1}. \] By straightforward calculations, we have \[ \mathcal{E}_{\tau,\delta}=\frac{1}{Q_{\tau}[Y]}\frac{dQ_{\tau}[Y_{\delta} ]}{d\delta}\left( \frac{1}{s\left( \delta\right) }\frac{ds\left( \delta\right) }{d\delta}\right) ^{-1}. \] When $s(0)=1$ and $\dot{s}\left( 0\right) \neq0$, the elasticity at $\delta=0$ is

equation[equation omitted — 125 chars of source]

Therefore, a $1\%$ increase in the standard deviation of $X$ results in a $\Pi_{\tau,S}/\{\dot{s}\left( 0\right) Q_{\tau}[Y]\}\%$ change in the $\tau $-quantile of $Y$.

\addtocounter{example}{-1}

example[Continued]Plugging $\Pi_{\tau,S}=\dot{s}\left( 0\right) \sqrt{R_{YX}^{2} }Q_{\tau}[X_{\gamma}^{\circ}]$, we obtain the quantile-standard deviation elasticity at $\delta=0$ as \[ \mathcal{E}_{\tau,\delta=0}=\sqrt{R_{YX}^{2}}\frac{Q_{\tau}[X_{\gamma}^{\circ }]}{Q_{\tau}[Y]}. \] So, $\mathcal{E}_{\tau,\delta=0}$ is positive if $Q_{\tau}[X_{\gamma}^{\circ }]$ and $Q_{\tau}[Y]$ have the same sign. When $\alpha=0$, $\mu_{X}=0$, and $X$ and $U$ are independent normals, we have $Q_{\tau}[X_{\gamma}^{\circ }]/Q_{\tau}[Y]=\sqrt{R_{YX}^{2}}$ and so $\mathcal{E}_{\tau,\delta=0} =R_{YX}^{2}$. Interestingly, the quantile-standard deviation elasticity is equal to the population R-squared for all quantile levels.

Often times, when the outcome of interest (e.g., price and wage) is strictly positive, we are interested in $\log Y$. In such a case, we denote the scale effect by $\tilde{\Pi}_{\tau,S}$. Since we set $\ell(\delta)\equiv0$ and there is no location effect, the new scale effect is given by \[ \tilde{\Pi}_{\tau,S}:=\lim_{\delta\rightarrow0}\frac{Q_{\tau}[\log Y_{\delta }]-Q_{\tau}[\log Y]}{\delta}. \] Since $\log\left( \cdot\right) $ is a strictly increasing transformation, we have \[ \tilde{\Pi}_{\tau,S}=\lim_{\delta\rightarrow0}\frac{\log Q_{\tau}[Y_{\delta }]-\log Q_{\tau}[Y]}{\delta}, \] and we can relate $\tilde{\Pi}_{\tau,S}$ to ${\Pi}_{\tau,S}$ by \[ \tilde{\Pi}_{\tau,S}=\frac{1}{Q_{\tau}[Y]}{\Pi}_{\tau,S}. \] Comparing this to (ref), we obtain that the elasticity at $\delta=0$ is \[ \mathcal{E}_{\tau,\delta=0}=\frac{\tilde{\Pi}_{\tau,S}}{\dot{s}(0)}. \] This says that a $1\%$ increase in the standard deviation of $X$ results in a $\tilde{\Pi}_{\tau,S}/\dot{s}\left( 0\right) \%$ change in the $\tau $-quantile of $Y$. When $\dot{s}\left( 0\right) =1$ (e.g., $s\left( \delta\right) =1+\delta),$ the scale effect $\tilde{\Pi}_{\tau,S}$ (based on $\log\left( Y)\right) $ can be interpreted directly as the quantile-standard deviation elasticity. When $\dot{s}\left( 0\right) =-1$ (e.g., $s\left( \delta\right) =1/(1+\delta)),$ the scale effect $\tilde{\Pi} _{\tau,S}$ has the same magnitude as the quantile-standard deviation elasticity but with an opposite sign.

{\color{red} }

Other potential applications

The framework developed here can be extended in several directions. In the Supplementary Appendix (ref), we consider a case where a location shift in one covariate is compensated or amplified by a location shift in another covariate, allowing for simultaneous changes in different covariates.

Our framework is also useful for evaluating heterogeneous interventions. Specifically, we can accommodate cases where interventions vary across covariate-specific strata.\footnote{We thank an anonymous referee for suggesting this possibility.} For instance, a plausible intervention could involve increasing $X$ among units with $W\in\mathcal{W}_{1}$ and decreasing $X$ among units with $W\in\mathcal{W}_{2}$ where $\mathcal{W}_{1}$ and $\mathcal{W}_{2}$ are non-overlapping subsets of $\mathcal{W}.$ One possible implementation of this is through the following $\mathcal{G}$ function, which now depends on $W$: \[ X_{\delta}=\mathcal{G}(X,W,\delta)=(X+\delta)\mathds1\{W\in\mathcal{W} _{1}\}+(X-\delta)\mathds1\{W\in\mathcal{W}_{2}\}. \] In this case, individuals with characteristics $W\in\mathcal{W}_{1}$ experience an upperward shift in $X$, while those with $W\in\mathcal{W}_{2}$ experience a downward shift.

For a general $\mathcal{G}$ function that depends on $W,$ we need to replace Assumption (ref)(i) by the following:

assumption(i.a) For some $\varepsilon>0$ and for each $w\in\mathcal{W},$ $\mathcal{G}\left( x,w;\delta\right) $ is continuously differentiable in $\left( x,\delta\right) $ on $\mathcal{X\otimes N}_{\varepsilon}$, where $\mathcal{X}$ represents the support of $X$. (i.b) $\mathcal{G}\left( x,w;\delta\right) $ is strictly increasing in $x$ for each $\delta\in\mathcal{N}_{\varepsilon}$ and $w\in\mathcal{W}$. (i.c) $\mathcal{G}\left( x,w;0\right) =x$ for all $x\in\mathcal{X}$ and $w\in\mathcal{W}$.

With the above assumption in place, we redefine $\kappa\left( \cdot\right) $ as \[ \kappa\left( x,w\right) =\frac{\partial\mathcal{G}(x,w;\delta)} {\partial\delta}\bigg|_{\delta=0}. \] Then Theorem (ref) remains valid with $\kappa\left( x\right) $ replaced by $\kappa\left( x,w\right) $ and Assumption (ref) (i) by Assumption\ (ref). It is noteworthy that the differentiability of $\mathcal{G}(x,w;\delta)$ with respect to $w$ is not required for the theorem to hold. Rather, the previously stated assumptions need to hold for each $w\in\mathcal{W}$.

Since $G$ can take a general form, our framework is not only applicable to the scenarios mentioned above but can also be further extended in other directions.

Distribution intervention vs. value intervention

The seminal paper by FirpoFortinLemieux09 (FFL hereafter in this section) considers the effect of a change in the marginal distribution of $X$ from $F_{X}$ to either $(i)$ a fixed $G_{X}$ or $(ii)$ a \textquotedblleft variable\textquotedblright\ $G_{X,\delta}$ which depends on $\delta$. Rothe2012 also focuses on these two cases.

In the first case, FFL considers a change from $F_{X}$ to a fixed $G_{X}$. Keeping $F_{Y|X}$ the same, a counterfactual distribution can be obtained by $F_{Y}^{\ast}(y)=\int_{\mathcal{X}}F_{Y|X}(y|x)dG_{X}(x)$. For $\delta\in\lbrack0,1]$, the convex combination $F_{Y,\delta}:=(1-\delta )F_{Y}+\delta F_{Y}^{\ast}$ is a cdf and can be interpreted as a perturbation of $F_{Y}$ in the direction of $F_{Y}^{\ast}-F_{Y}$. For a certain statistic $\rho(F)$ of interest, such as a particular quantile of $Y$, we have

equation[equation omitted — 150 chars of source]

where $\psi_{\rho}(y,F_{Y})$ is the influence function of $\rho$ at $F_{Y}$. See Chapter 20 in vanderVaart98 or Section 2.1 in NeweyIchimura2022. Since $F_{Y}^{\ast}(y)-F_{Y}(y)=\int_{\mathcal{X} }F_{Y|X}(y|x)d(G_{X}-F_{X})(x)$, we have

align[align omitted — 255 chars of source]

This is essentially Theorem 1 in FFL, which provides a characterization of a directional derivative of the functional $\rho\left( \cdot\right) $ in the direction induced by a change in the marginal distribution of $X.$ The theorem is silent on how the change in the marginal distribution is implemented.

The second case, covered in Corollary 1 in FFL, is closer to what we consider here. In this case, $G_{X,\delta}$ is the distribution induced by the location shift $X+\delta$. The counterfactual distribution is $F_{Y,\delta}^{\ast }(y)=\int_{\mathcal{X}}F_{Y|X}(y|x)dG_{X,\delta}(x)$. The parameter of interest is $\lim_{\delta\rightarrow0}\left[ \rho(F_{Y,\delta}^{\ast} )-\rho(F_{Y})\right] /\delta.$ Corollary 1 in FFL shows that

equation[equation omitted — 215 chars of source]

Our general intervention $X_{\delta}=\mathcal{G}(X;\delta)$ includes the above location shift as a special case. To see this, we assume that $W$ is not present and set $\mathcal{G}(X;\delta)=X+\delta$, in which case $\kappa\left( x\right) =1,$ and it follows from Remark (ref) and Corollary (ref) that $\Pi_{\rho}:=E\left[ \frac{\partial E\left[ \psi_{\rho}\left( Y,F_{Y}\right) |X\right] }{\partial X}\right] =\int_{\mathcal{X}}\frac{\partial E[\psi_{\rho}(Y,F_{Y})|X=x]}{\partial x}dF_{X}(x)$, which is identical to the right-hand side of ((ref)). This shows that our approach is strictly more general than the second case considered by FFL.

There is another main difference between FFL and our paper. From a broad point of view, FFL considers the scenario where the conditional distribution of $Y$ given $X$ is fixed, and ask how the unconditional distribution of $Y$ would change if the marginal distribution of $X$ had changed. This is largely a predictive exercise unless the conditional distribution of $Y$ given $X$ has a structural or causal interpretation, that is, $X$ is exogenous. In our paper, we allow for an endogenous $X$ in the sense that $X$ and $U$ may be correlated. This could arise, for example, when a common factor causes both $X$ and $U$. As discussed in Remark (ref), $X$ and $U$ may be dependent even after conditioning on the causal variable $W$. In such a case, we need to find additional control variables that do not necessarily enter the structural function $h$ such that $X$ and $U$ become conditionally independent conditional on $W$ and these additional control variables. The endogeneity problem is then addressed by using the control variable approach.

At the conceptual level, we consider the policy experiment where both the structural function $h$ and the distribution of ($X,W,U$) are kept intact. Given that $h$ is the same, we can say that the effect is causal and have a ceteris paribus interpretation. Given that the distribution of ($X,W,U$) is the same, the policy experiment applies to the current population under consideration and is fully implementable. Hence the effect is what a policy maker can achieve under the current environment and is therefore fully policy-relevant.

Furthermore, our counterfactual exercise focuses on manipulating the value of the target covariate, while the bulk of the literature focuses more on manipulating its marginal distribution and often uses a value intervention as an example of how the marginal distribution may be shifted. The advantage of using a value intervention is that the policy function $\mathcal{G} (\cdot;\delta)$ defines clearly how the policy can be implemented. This is in contrast to the intervention of the marginal distribution where the policy maker is not given a clear recipe to achieve such an intervention. In addition, it seems to be easier to attach a cost implication to the value intervention. A policy maker may want to trade off the cost with the policy goal they hope to achieve. A marginal distribution intervention seems to be more of theoretical interest unless it can be implemented empirically via a value intervention as considered in this paper.\footnote{An important example of value interventions is the literature on policy relevant treatment effects where an instrumental variable is manipulated in order to shift the program participation rate. See, for example, Heckman2005.}

Estimation and asymptotic results

In this section, we focus on the estimation of $\Pi_{\tau}$ given in (ref). The estimator involves several preliminary steps. Firstly, for a given quantile, we need to estimate $Q_{\tau}[Y]$. This is given by

equation[equation omitted — 144 chars of source]

Next, we need to estimate the density of $Y$ evaluated at $Q_{\tau}[Y]$. This can be estimated by

equation[equation omitted — 153 chars of source]

where $\mathcal{K}_{h}(u)=h^{-1}\mathcal{K}(h^{-1}u)$ for a given kernel $\mathcal{K}$ and a bandwidth $h$. For the average derivative of the conditional cdf, we propose either a logit model as in FirpoFortinLemieux09 or a probit model. Note that $\mathcal{S} _{Y|X,W}(Q_{\tau}[Y]|x,w)=1-F_{Y|X,W}(Q_{\tau}[Y]|x,w).$ We model $\mathcal{S}_{Y|X,W}(Q_{\tau}[Y]|x,w)$ via $F_{Y|X,W}(Q_{\tau}[Y]|x,w)$ by assuming that

equation[equation omitted — 170 chars of source]

where $\phi_{\mathrm{x}}\left( \cdot\right) $ and $\phi_{\mathrm{w}}\left( \cdot\right) $ are column vectors of smooth basis functions and $G(\cdot)$ is either the cdf of a logistic random variable (logit) or a standard normal random variable (probit). Note that the subscripts \textquotedblleft $\mathrm{x}$\textquotedblright\ and \textquotedblleft$\mathrm{w} $\textquotedblright\ serve only to distinguish $\phi_{\mathrm{x}}\left( \cdot\right) $ from $\phi_{\mathrm{w}}\left( \cdot\right) .$ They are not related to the arguments of these functions. For the choices of $\phi _{\mathrm{x}}\left( \cdot\right) $ and $\phi_{\mathrm{w}}\left( \cdot\right) ,$ we may take $\phi_{\mathrm{x}}\left( x\right) =x$ or ($x,x^{2})^{\prime}$ and $\phi_{\mathrm{w}}\left( w\right) =\left( 1,w\right) ^{\prime}.$ By default, we include the constant in the vector $w.$ Other more flexible choices are possible, but it is beyond the scope of this paper to consider a fully nonparametric specification.

Let $Z_{i}=[\phi_{\mathrm{x}}\left( X_{i}\right) ^{\prime},\phi_{\mathrm{w} }(W_{i})^{\prime}]^{\prime}$ and $\theta_{\tau}=(\alpha_{\tau}^{\prime} ,\beta_{\tau}^{\prime})^{\prime}.$ We estimate $\theta_{\tau}$ by the maximum likelihood estimator:

align[align omitted — 450 chars of source]

where $\Theta$ is a compact parameter space that contains $\theta_{\tau}$ as an interior point. The estimator of $\Pi_{\tau}$ is then \[ \hat{\Pi}_{\tau}=\hat{\Pi}_{\tau,L}+\hat{\Pi}_{\tau,S} \] where

align[align omitted — 532 chars of source]

In the above, $g$ is the derivative of $G$, that is, the logistic density or the standard normal density and $\dot{\phi}_{\mathrm{x}}\left( x\right) =\partial\phi_{\mathrm{x}}\left( x\right) /\partial x$, which has the same dimension as $\phi_{\mathrm{x}}\left( x\right) $. In order to establish the asymptotic distribution of $\hat{\Pi}_{\tau}$, we need the following three sets of assumptions, one for each preliminary estimation step.

assumptionQuantile. The density of $Y$ is positive, continuous, and differentiable at $Q_{\tau}[Y]$.
assumptionLogit/Probit. For $G$ either the cdf of a logistic or a standard normal random variable, we have \begin{enumerate} • $F_{Y|Z}(Q_{\tau}[Y]|z)=G(z^{\prime}\theta_{\tau})$ for an interior point $\theta_{\tau}\in\Theta$ and $\hat{\theta}_{\tau}=\theta_{\tau} +o_{p}\left( 1\right) .$ • For \[ H_{i}\left( \theta;q\right) =\frac{\partial^{2}l_{i}\left( \theta;q\right) }{\partial\theta\partial\theta^{\prime}}, \] which is the Hessian of observation $i$, the following holds \[ \sup_{(\theta,q)\in\mathcal{N}}\left\Vert \frac{1}{n}\sum_{i=1}^{n} H_{i}(\theta;q)-E[H_{i}(\theta;q)]\right\Vert \overset{p}{\rightarrow}0, \] where $\mathcal{N}$ is a neighborhood of $(\theta_{\tau}^{\prime},Q_{\tau }[Y]^{\prime})^{\prime}$, and $H:=E[H_{i}(\theta_{\tau};Q_{\tau}[Y])]$ is negative definite. • For the score $s_{i}$ defined by \[ s_{i}\left( \theta,q\right) =\frac{\partial l_{i}\left( \theta;q\right) }{\partial\theta}, \] the following stochastic equicontinuity assumption holds: \[ \frac{1}{n}\sum_{i=1}^{n}\left\{ s_{i}(\theta_{\tau};\hat{q}_{\tau})-E\left[ s_{i}(\theta_{\tau};q)\right] |_{q=\hat{q}_{\tau}}\right\} =\frac{1}{n} \sum_{i=1}^{n}s_{i}(\theta_{\tau};Q_{\tau}[Y])+o_{p}(n^{-1/2}), \] and the map $q\mapsto E\left[ s_{i}(\theta_{\tau};q)\right] $ is continuously differentiable at $Q_{\tau}[Y]$ with \[ \frac{\partial E\left[ s_{i}(\theta_{\tau};q)\right] }{\partial q}\bigg\vert_{q=Q_{\tau}[Y]}=:H_{Q}. \] • For $\tilde{X}_{i}=(1,X_{i})^{\prime}$, \begin{align*} M_{1}\left( \theta\right) & :=E\left\{ [\dot{g}(Z_{i}^{\prime}\theta )\dot{\phi}_{\mathrm{x}}\left( X_{i}\right) ^{\prime}\alpha]\tilde{X} _{i}Z_{i}^{\prime}\right\} \in\mathbb{R}^{2\times d_{Z}}\\ M_{2}\left( \theta\right) & :=E\left[ g(Z_{i}^{\prime}\theta)\tilde {X}_{i}\dot{\phi}_{\mathrm{x}}(X_{i})^{\prime}\right] \in\mathbb{R}^{2\times d_{\phi_{\mathrm{x}}}} \end{align*} are well defined for any $\theta\in\mathcal{N}_{\theta_{\tau}}$, a neighborhood of $\theta_{\tau}$; and the following uniform law of large numbers holds: \begin{align*} & \sup_{\theta\in\mathcal{N}_{\theta_{\tau}}}\bigg\|\frac{1}{n}\sum_{i=1} ^{n}[\dot{g}(Z_{i}^{\prime}\theta)\dot{\phi}_{\mathrm{x}}\left( X_{i}\right) ^{\prime}\alpha]\tilde{X}_{i}Z_{i}^{\prime}-M_{1}\left( \theta\right) \bigg\|\overset{p}{\rightarrow}0,\\ & \sup_{\theta\in\mathcal{N}_{\theta_{\tau}}}\bigg\|\frac{1}{n}\sum_{i=1} ^{n}g(Z_{i}^{\prime}\theta)\tilde{X}_{i}\dot{\phi}_{\mathrm{x}}(X_{i} )^{\prime}-M_{2}\left( \theta\right) \bigg\|\overset{p}{\rightarrow}0, \end{align*} where $\dot{g}$ is the derivative of $g$. \end{enumerate}

In the above assumption, we assume that $F_{Y|Z}(Q_{\tau}[Y]|z)=G(z^{\prime }\theta_{\tau})$ with $G$ being either the cdf of a logistic or a standard normal random variable. It is important to note that other cdfs can also be utilized. For instance, when the interest lies in the lowest quantiles with $\tau$ very close to 0, the cdf of a Gumbel distribution (also known as a Type I extreme value distribution) can be employed. This choice leads to a complementary log-log model, wherein $F_{Y|Z}(Q_{\tau}[Y]|z)$ is modeled by $1-\exp(-\exp(z^{\prime}\theta_{\tau}))$ and the index $z^{\prime}\theta _{\tau}$ can be written in the complementary log-log form $\log(-\log (1-F_{Y|Z}(Q_{\tau}[Y]|z)).$\footnote{We thank an anonymous referee for suggesting the complementary log-log or log-log link when our focus is on extreme quantiles.}

assumptionDensity. \begin{enumerate} • The kernel function $K(\cdot)$ satisfies (i) $\int_{-\infty }^{\infty}K(u)du=1$, (ii) $\int_{-\infty}^{\infty}u^{2}K(u)du<\infty$, and (iii) $K(u)=K(-u)$, and it is twice differentiable with Lipschitz continuous second-order derivative $K^{\prime\prime}\left( u\right) $ satisfying (i) $\int_{-\infty}^{\infty}K^{\prime\prime}(u)udu<\infty$ and $\left( ii\right) $ there exist positive constants $C_{1}$ and $C_{2}$ such that $\left\vert K^{\prime\prime}\left( u_{1}\right) -K^{\prime\prime}\left( u_{2}\right) \right\vert \leq C_{2}\left\vert u_{1}-u_{2}\right\vert ^{2}$ for $\left\vert u_{1}-u_{2}\right\vert \geq C_{1}.$ • As $n\uparrow\infty$, the bandwidth satisfies: $h\downarrow0$, $nh^{3}\uparrow\infty$, and $nh^{5}=O(1)$. \end{enumerate}

Under Assumption (ref), $\hat{q}_{\tau}$ given in (ref) is asymptotically linear with \[ \hat{q}_{\tau}-Q_{\tau}[Y]=\frac{1}{n}\sum_{i=1}^{n}\frac{\tau -\mathds{1}\left\{ Y_{i}\leq Q_{\tau}[Y]\right\} }{f_{Y}(Q_{\tau}[Y])} +o_{p}(n^{-1/2})=\frac{1}{n}\sum_{i=1}^{n}\psi(Y_{i},\tau,F_{Y})+o_{p} (n^{-1/2}). \] See, for example, serfling1980. Assumption (ref) is mostly necessary to deal with the preliminary estimator $\hat{q}_{\tau}$ that enters the likelihood in (ref). Assumption (ref) is taken from yixiao2020.

The following lemma contains the influence function for the maximum likelihood estimator $\hat{\theta}_{\tau}$.

lemmaUnder Assumptions (ref) and (ref), we have \[ \hat{\theta}_{\tau}-\theta_{\tau}=-H^{-1}\frac{1}{n}\sum_{i=1}^{n}s_{i} (\theta_{\tau};Q_{\tau}[Y])-H^{-1}H_{Q}\frac{1}{n}\sum_{i=1}^{n}\psi (Y_{i},\tau,F_{Y})+o_{p}(n^{-1/2}). \]
theoremUnder Assumptions (ref), (ref), and (ref), the estimators given in (ref) and (ref) satisfy \[ \begin{pmatrix} \hat{\Pi}_{\tau,L}\\ \hat{\Pi}_{\tau,S} \end{pmatrix} - \begin{pmatrix} \Pi_{\tau,L}\\ \Pi_{\tau,S} \end{pmatrix} =\frac{1}{n}\sum_{i=1}^{n}\Phi_{i,\tau}+O\left( h^{2}\right) +o_{p} (n^{-1/2})+o_{p}(n^{-1/2}h^{-1/2}), \] where \begin{align*} \Phi_{i,\tau} & =\frac{1}{f_{Y}\left( Q_{\tau}[Y]\right) }D_{\mu}\left\{ g(Z_{i};\theta_{\tau})\dot{\phi}_{\mathrm{x}}\left( X_{i}\right) ^{\prime }\alpha_{\tau}\tilde{X}_{i}-E\left[ g(Z_{i};\theta_{\tau})\dot{\phi }_{\mathrm{x}}\left( X_{i}\right) ^{\prime}\alpha_{\tau}\tilde{X} _{i}\right] \right\} \\ & -\frac{1}{f_{Y}\left( Q_{\tau}[Y]\right) }D_{\mu}MH^{-1}s_{i} (\theta_{\tau};Q_{\tau}[Y])\\ & -\left[ \begin{pmatrix} \Pi_{\tau,L}\\ \Pi_{\tau,S} \end{pmatrix} \frac{\dot{f}_{Y}(Q_{\tau}[Y])}{f_{Y}\left( Q_{\tau}[Y]\right) }+\frac {1}{f_{Y}\left( Q_{\tau}[Y]\right) }D_{\mu}MH^{-1}H_{Q}\right] \psi (Y_{i},\tau,F_{Y})\\ & - \begin{pmatrix} \Pi_{\tau,L}\\ \Pi_{\tau,S} \end{pmatrix} \frac{1}{f_{Y}\left( Q_{\tau}[Y]\right) }\left\{ \mathcal{K}_{h}\left( Y_{i}-Q_{\tau}[Y]\right) -E\mathcal{K}_{h}\left( Y_{i}-Q_{\tau}[Y]\right) \right\} , \end{align*} $\dot{f}_{Y}\left( \cdot\right) $ is the derivative of $f_{Y}\left( \cdot\right) ,$ \[ D_{\mu}= \begin{pmatrix} D_{L}^{\prime}\\ D_{\mu,S}^{\prime} \end{pmatrix} = \begin{pmatrix} -\dot{\ell}(0) & 0\\ \mu\dot{s}(0) & -\dot{s}(0) \end{pmatrix} , \] \[ M=M_{1}\left( \theta_{\tau}\right) + \begin{pmatrix} M_{2}\left( \theta_{\tau}\right) , & O \end{pmatrix} \in\mathbb{R}^{2\times d_{Z}}, \] and $O\in\mathbb{R}^{2\times d_{\phi_{\mathrm{w}}}}$ is a matrix of zeros.

Theorem (ref) establishes the contribution from each estimation step. In particular, the last term in $n^{-1}\sum_{i=1}^{n}\Phi_{i,\tau}$ is the contribution from estimating the density of $Y$ non-parametrically. This term converges at a non-parametric rate, which is slower than other terms. As a result, the asymptotic distribution of the location-scale effect estimator is determined by the last term in $n^{-1}\sum_{i=1}^{n}\Phi_{i,\tau}$. However, we do not recommend dropping all other terms. Instead, we write the asymptotic normality result in the form

equation[equation omitted — 325 chars of source]

as $n\uparrow\infty$, $nh^{3}\uparrow\infty$, and $nh^{5}\downarrow0$ where $\hat{\Phi}_{i,\tau}$ is a plug-in estimator of $\Phi_{i,\tau}.$ In particular,

align[align omitted — 389 chars of source]

as $n\uparrow\infty$, $nh^{3}\uparrow\infty$, and $nh^{5}\downarrow0$ where $l_{1}=\left( 1,0\right) ^{\prime}$ and $l_{2}=\left( 0,1\right) ^{\prime }$. Note that Theorem (ref) has shown that the estimation error in $\hat{\Pi}_{\tau,L}$ or $\hat{\Pi}_{\tau,S}$ is an average of independent observations. The above asymptotic normality results can be proved using a Lyapunov CLT under the following conditions (see the proof of Theorem 2.9 in pagan_ullah_1999):

(i) $n^{-1}h\sum_{i=1}^{n}E\left[ \Phi_{i,\tau}\Phi_{i,\tau}^{\prime}\right] $ is nonsingular for all large enough $n.$

(ii) $\left( n^{-1}h\sum_{i=1}^{n}E\Phi_{i,\tau}\Phi_{i,\tau}^{\prime }\right) ^{-1/2}\left( n^{-1}h\sum_{i=1}^{n}\hat{\Phi}_{i,\tau}\hat{\Phi }_{i,\tau}^{\prime}\right) \left( n^{-1}h\sum_{i=1}^{n}E\Phi_{i,\tau} \Phi_{i,\tau}^{\prime}\right) ^{-1/2}=I_{2}+o_{p}\left( 1\right) .$

(ii) Assumption (ref) holds, $\int_{-\infty}^{\infty }\left\vert K(u)\right\vert ^{2+\Delta}du<\infty$ for some $\Delta>0,$ and $\left\vert f_{Y}^{\prime\prime}(Q_{\tau}[Y]\right\vert <C$ for some constant $C.$

Inferences based on our asymptotic results account for the estimation errors from all estimation steps and are more reliable in finite samples. This is supported by simulation evidence not reported here, but available upon request. On the other hand, if we parametrize the density of $Y$ and estimate it at the parametric $\sqrt{n}$-rate, then the last term in $n^{-1}\sum _{i=1}^{n}\Phi_{i,\tau}$ will take a different form and will be of the same order as the other terms. In this case, the location-scale effect estimator is $\sqrt{n}$-asymptotically normal, and all the terms in Theorem (ref) will contribute to the asymptotic variance. With an obvious modification of the last term in $\Phi_{i,\tau},$ the asymptotic normality can be presented in the same way as in ((ref)).

Let \[ \Gamma_{\tau,S}=D_{\mu,S}^{\prime}E\left[ \frac{\partial F_{Y|X,W}(Q_{\tau }[Y]|X,W)}{\partial X}\tilde{X}\right] \] be the numerator of $\Pi_{\tau,S}.$ Then the scale effect $\Pi_{\tau,S}$ is zero if and only if $\Gamma_{\tau,S}=0.$ To test the null hypothesis $H_{0}:\Pi_{\tau,S}=0,$ we can equivalently test the null hypothesis $H_{0}:\Gamma_{\tau,S}=0.$ Unlike $\Pi_{\tau,S},$ $\Gamma_{\tau,S}$ can be estimated at the parametric rate even if $f_{Y}\left( \cdot\right) $ is not parametrically specified. More specifically, under Assumption (ref), we can estimate $\Gamma_{\tau,S}$ by \[ \hat{\Gamma}_{\tau,S}:=D_{\mu,S}^{\prime}\frac{1}{n}\sum_{i=1}^{n} [g(Z_{i}^{\prime}\hat{\theta}_{\tau})\dot{\phi}_{\mathrm{x}}\left( X_{i}\right) ^{\prime}\hat{\alpha}_{\tau}]\tilde{X}_{i}, \] where $D_{\mu,S}^{\prime}=\left( \mu,-1\right) $ upon setting $\dot{s}(0)=1$ without loss of generality.

Under the assumptions of Theorem (ref), we can show that \[ \hat{\Gamma}_{\tau,S}-\Gamma_{\tau,S}=D_{\mu,S}^{\prime}\frac{1}{n}\sum _{i=1}^{n}\Phi_{i,\tau}^{\Gamma}+o_{p}\left( \frac{1}{\sqrt{n}}\right) , \] where

align*[align* omitted — 372 chars of source]

Define \[ V_{\tau}=\lim_{n\rightarrow\infty}\frac{1}{n}\sum_{i=1}^{n}E(D_{\mu,S} ^{\prime}\Phi_{i,\tau}^{\Gamma})^{2}. \] If $D_{\mu,S}^{\prime}\Phi_{i,\tau}^{\Gamma}$ has a finite second moment and $V_{\tau}>0,$ then a standard CLT yields $V_{\tau}^{-1/2}\sqrt{n}(\hat{\Gamma }_{\tau,S}-\Gamma_{\tau,S}^{\mu})\overset{d}{\rightarrow}N\left( 0,1\right) $. To test $H_{0}:\Gamma_{\tau,S}=0,$ we construct the test statistic \[ t_{\tau,S}:=\frac{\sqrt{n}\hat{\Gamma}_{\tau,S}}{\sqrt{\hat{V}_{\tau}}}\text{ for }\hat{V}_{\tau}=\frac{1}{n}\sum_{i=1}^{n}(D_{\mu,S}^{\prime}\hat{\Phi }_{i,\tau}^{\Gamma})^{2}, \] where

align[align omitted — 466 chars of source]

In the above, $\hat{\psi}(Y_{i},\tau,F_{Y})=\left[ \tau-1\left\{ Y_{i} \leq\hat{q}_{\tau}\right\} \right] /\hat{f}_{Y}\left( \hat{q}_{\tau }\right) $ and the score $s_{i}(\hat{\theta}_{\tau};\hat{q}_{\tau})$ is obtained by evaluating the expression given in (ref) at $\theta=\hat{\theta}_{\tau}$ and $q=\hat{q}_{\tau}$. ${\hat{M}},$ ${\hat{H},}$ and $\hat{H}_{Q}$ are the sample versions of ${M},$ ${H,}$ and $H_{Q},$ respectively. Details are given in the proof of the corollary below.

corollaryLet the assumptions of Theorem (ref) hold. Assume that $D_{\mu,S}^{\prime}\Phi_{i,\tau}^{\Gamma}$ has a finite second moment and $\hat{V}_{\tau}/V_{\tau}\overset{p}{\rightarrow}1$ for some $V_{\tau}>0.$ Then, under the null hypothesis $H_{0}:\Pi_{\tau,S}=0,$ \[ t_{\tau,S}\overset{d}{\rightarrow}N(0,1). \]

Monte Carlo experiments

In this section, we use Monte Carlo simulations to evaluate the finite sample performances of the proposed estimators and tests of location and scale effects. We employ the same data generating process as in Example (ref) for which we have derived the closed-form expressions for the location and scale effects. In particular, we let \[ Y=\lambda+X\gamma+U, \] where $X\sim N(\mu_{X},\sigma_{X}^{2})$ and $U\sim N(0,\sigma_{U}^{2})$. We set $\lambda=0,$ $\sigma_{U}^{2}=1,$ $\dot{\ell}(0)=1$ and $\dot{s}(0)=-1$. The last derivative corresponds to, for example, $s(\delta)=(1+\delta)^{-1}$. Then, from the results in Example (ref), the true location effect is $\Pi_{\tau,L}=\gamma,$ and the true scale effect is \[ \Pi_{\tau,S}^{\mu_{X}}=-\sqrt{R_{YX}^{2}}Q_{\tau}[X_{\gamma}^{\circ} ]=-\sqrt{R_{YX}^{2}}\sqrt{var(X_{\gamma}^{\circ})}Q_{\tau}[\varepsilon ]=-\sqrt{R_{YX}^{2}}\cdot\sigma_{X}\cdot\left\vert \gamma\right\vert \cdot Q_{\tau}[\varepsilon] \] where $\varepsilon$ is standard normal.

We consider quantiles $\tau\in\{0.10,0.25,0.50,0.75,0.90\}$ and sample sizes $n=500$ and $n=1000$. The number of simulations is set to $10,000$ for each experiment.

We implement our estimators in Matlab. The unconditional quantile estimator in equation (ref) is easily computed as an order statistic. The density function is estimated as a kernel density estimator as in equation (ref) using a standard normal kernel. For the bandwidth choice in the kernel density estimation, we use a modified version of Silverman's rule of thumb. More specifically, since we require $nh^{3} \uparrow\infty$ and $nh^{5}\downarrow0$ as $n\uparrow\infty$, we take $h=1.06\hat{\sigma}_{Y}n^{-1/4}$, where $\hat{\sigma}_{Y}$ is the sample standard deviation of $Y$.

Bias, variance, and mean squared error

In this subsection, we consider the bias, variance, and mean-squared error (MSE) of the proposed location and scale effects estimators. For each effect estimator, we consider either a probit or a logit specification for the conditional cdf $F_{Y|X}(Q_{\tau}[Y]|X).$ Under our data generating process, the probit with $F_{Y|X}(Q_{\tau}[Y]|X)=\boldsymbol{\Phi}(X\alpha_{\tau} +\beta_{\tau})$ for the standard normal CDF $\boldsymbol{\Phi}$ is correctly specified while the logit with $F_{Y|X}(Q_{\tau}[Y]|X)=\left[ 1+\exp\left( X\alpha_{\tau}+\beta_{\tau}\right) \right] ^{-1}$ is misspecified.

The bias, variance, and MSE are reported in Table (ref) when $\mu _{X}=0$, $\gamma=1$ and $\sigma_{X}^{2}=1$ so that the true location effect is $1$ for any $\tau$ and the true scale effect is $-\sqrt{0.5}Q_{\tau }[\varepsilon]\approx-0.707Q_{\tau}[\varepsilon].$ To save space, simulation results for other values of $\gamma$ and $\sigma_{X}^{2}$ are omitted.

Table (ref) shows that the estimator based on the probit specification outperforms that based on the logit one. This is consistent with the correct specification of probit. For each estimator, the bias decreases as the sample size $n$ increases. The variance also decreases as the sample size $n$ increase, and as a result, the MSE also becomes smaller when the sample size grows. For our purposes, the scale effect estimator performs well. For non-central quantiles, the difference in the scale effect estimates under the probit and logit specifications is in general larger than the difference in the location effect estimates. For central quantiles, the probit and logit specifications lead to more or less the same estimates for both the scale effect and the location effect.

center[center omitted — 2,119 chars of source]

Accuracy of the normal approximation

In this subsection, we investigate the finite sample accuracy of the normal approximation given in ((ref)). Using the same data generating process as in the previous subsection and employing the probit specification, we simulate the distributions of the studentized statistics \[ \left[ n^{-2}\sum_{i=1}^{n}(l_{1}^{\prime}\hat{\Phi}_{i,\tau})^{2}\right] ^{-1/2}(\hat{\Pi}_{\tau,L}-\Pi_{\tau,L}) \] and \[ \left[ n^{-2}\sum_{i=1}^{n}(l_{2}^{\prime}\hat{\Phi}_{i,\tau})^{2}\right] ^{-1/2}(\hat{\Pi}_{\tau,S}-\Pi_{\tau,S})., \] for the location and scale effects, respectively. We plot each distribution and compare it with the standard normal distribution. We consider $\gamma \in\left\{ 0.25,0.50,0.75,1\right\} $ and use the same $\tau$ values as in the previous subsection. Simulation results for the two sample sizes $n=500$ and $n=1000$ are qualitatively similar, and we report only the case when $n=1000$ here. Figures (ref)--(ref) report the (simulated) finite sample distributions when $\sigma_{X}^{2}=1$ and $n=1000$ for some selected values of $\gamma$ and $\tau$ together with a standard normal density that is superimposed on each figure. It is clear from these figures that the standard normal distribution provides an accurate approximation to the distribution of the studentized test statistic for both the location and scale effects.

figure[figure omitted — 300 chars of source]
figure[figure omitted — 269 chars of source]
figure[figure omitted — 263 chars of source]
figure[figure omitted — 291 chars of source]

Table (ref) reports the empirical coverage of 95% confidence intervals for the location and scale effects. The empirical coverage is close to the nominal coverage in all cases. This is consistent with Figures (ref)--(ref). We may then conclude that the normal approximation can be reliably used for inference on the location and scale effects.

table[table omitted — 1,465 chars of source]

Power of the t-test of a zero scale effect

To investigate the power of the t-test proposed in Corollary (ref), we simulate the following model: \[ Y=\lambda+X\gamma+U, \] where \[

pmatrix[pmatrix omitted — 20 chars of source]

\sim N\left(

pmatrix[pmatrix omitted — 20 chars of source]

,

pmatrix[pmatrix omitted — 28 chars of source]

\right) . \] Here we set $\lambda=0,$ $\mu_{X}=1$ and $\dot{s}(0)=-1$. When $\gamma=0$, $X$ is excluded from the outcome equation and thus the scale effect is 0. The null hypothesis of a zero scale effect corresponds to the case that $\gamma=0$. The power of the test is obtained by varying $\gamma$ around 0 in a grid from $-0.4$ to $0.4$ with an increment of $0.01$.

Figure (ref) graphs the size-adjusted power of the t-test for different quantile levels when $n=500$ and when $n=1000$. The power is calculated using the probit specification, namely $F_{Y|X}(Q_{\tau }[Y]|X)=\boldsymbol{\Phi}(X\alpha_{\tau}+\beta_{\tau})$. The size adjustment is based on the empirical critical value such that the test rejects the null 5% of the time. Figure (ref) shows that the power increases as $\gamma$ deviates more from its null value of zero, and that for a given nonzero value of $\gamma,$ the power increases with the sample size. Results not reported here show that the test has a quite accurate size in that the empirical rejection probability under the null is close to 5%, the nominal level of the test.

figure[figure omitted — 179 chars of source]

Empirical application

In this section, we consider two applications: education and wages, and smoking and birth weights.

Education and wages

Our first application is based on a household labor survey from wooldridge that can be accessed online for replication.\footnote{See \url{http://fmwww.bc.edu/ec-p/data/wooldridge/wage1.des} and \url{http://fmwww.bc.edu/ec-p/data/wooldridge/wage1.dta} for the data in the Stata data file format.} The idea is to evaluate the effects of education on the quantile of the unconditional distribution of log wages. In this application, $Y=lwage,$ which is log hourly wage, and $X=educ$, which is years of education is our target variable. The controls are: $W=[exper\ tenure\ nonwhite\ female]$, where $exper$ is years of working experience, $tenure$ is years with current employer, $nonwhite$ is a dummy that equals 1 if the individual is non-white, and $female$ is a dummy that equals 1 if the individual is female. We assume that Assumption (ref) holds for this choice of $W.$

While the main goal is to study the scale effect, we also present results for the location effect. We set $\dot{\ell}(0)=1$ and $\dot{s}(0)=-1$. Note that when $\dot{s}(0)=-1,$ the estimated effects we present below are the unconditional scale effects when the variance of the covariate is reduced by a small amount. For the mean of years of education $\mu_{X} $, we let $\mu_{X}=12.29$ based on the Barro-Lee Data on Educational Attainment.\footnote{The dataset is available from \url{https://databank.worldbank.org/reports.aspx?source=Education Statistics} We use the series \textquotedblleft Barro-Lee: Average years of total schooling, age 25+, total\textquotedblright\ for the US between 1970-2010 and find that the average years of schooling is 12.29.} We set $\mu=\mu_{X}=12.29$ to study the location and scale effects. In a similar fashion to the Monte Carlo analysis, we consider $\tau\in\{0.10,0.25,0.50,0.75,0.90\}$. The sample size for the household labor survey is $n=526$, which is comparable to $n=500$ in the simulation exercises. We compute the standard errors using the approximation in (ref).

table[table omitted — 1,470 chars of source]
figure[figure omitted — 327 chars of source]

The most interesting results in Table (ref) appear in the unconditional scale effects. As discussed in Section (ref), the scale effects can be interpreted as percentage changes of the unconditional quantiles. Consider the scale effect for $\tau=0.10$. Both the probit and logit specifications suggest an effect of about .045. Then, using the quantile-standard deviation elasticity, a $1\%$ decrease in the standard deviation of education would produce a positive effect of $.045\%$ on the unconditional quantile at the quantile level $\tau=0.10$. Given that the sample standard deviation of $educ$ is $2.77$, the $1\%$ decrease is approximately a change in the standard deviation from $2.77$ to $2.74$. Consider now the scale effect for $\tau=0.50$. In this case, both probit and logit specifications provide a statistically insignificant effect (at the $5\%$ level). Confront this with the results of Example (ref) where in the linear model $Y=\lambda +X\gamma+U$, the scale effect $\Pi_{0.50,S}=0$ if both $X$ and $U$ are symmetrically distributed around 0. Thus, $\hat{\Pi}_{0.50,S}\approx0$ is consistent with a linear model with symmetrically distributed $X$ and $U.$ Finally, consider the scale effect for $\tau=0.90$, again using both probit and logit specifications. In this case, the effects are negative, suggesting a $1\%$ decrease in the standard deviation would reduce the upper $\tau=0.90$ quantile by $.20\%$ (probit) and $.23\%$ (logit). Overall this analysis shows that the scale effects are monotonically decreasing in $\tau$. This can be seen in Figure (ref) that plots, for a finer grid of $\tau $,\footnote{For Figure (ref) we use $\tau =0.10,0.11,...,0.89,0.90$.} the probit estimates for both the location (dashed blue) and scale (solid red) effects.

How can this be interpreted? The location effects suggest that the marginal contribution of one more year of education benefits more the upper parts of the unconditional distribution of wages. The scale effects suggest the contrary. Reducing the overall dispersion of education would increase the lower quantile wages, but reduce the upper ones.

table[table omitted — 1,470 chars of source]

Smoking and birth weight

This second application considers the relationship between smoking during pregnancy and the child's birth weight. This was previously studied by Abrevaya2001, KoenkerHallock01, Rothe2010, and ChernozhukovFernandezVal11. We use the natality data from the National Vital Statistics System for the year 2018.\footnote{Available here: \url{https://www.nber.org/research/data/vital-statistics-natality-birth-data}.} The outcome variable is birth weight in grams, while the target variable is the average number of cigarettes smoked daily during pregnancy. We focus on the sample of mothers who are smokers. The sample consists of 219,667 observations.

For this model $Y$ is birth weight in grams and $X$ is the mother's reported average number of cigarettes smoked per day during pregnancy. We use the same covariates as Abrevaya2001:\footnote{We omit the dummy of whether the mother smoked during pregnancy because we focus on the sample of smoking mothers.} $(i)$ a dummy for whether the mother is black; $(ii)$ a dummy for marital status; $(iii)$ age and age squared; $(iv)$ a set of dummies for education attainment: high school graduate, some college, and college graduate; $(v)$ weight gain during pregnancy, $(vi)$ a set of dummies for prenatal visit: visit during the second trimester, visit during the third trimester, and no visit at all; and $(vii)$ a dummy for the sex of the child.

For this application, we set $\mu=0$, $l(\delta)\equiv0$, and $s(\delta )=1/\left( 1+\delta\right) $, so that according to (ref), counterfactual cigarette consumption is now $X_{\delta}=X/(1+\delta)$, which has a smaller mean and variance than $X.$ Note, again, that $\dot s(0)=-1$. To motivate such a counterfactual policy, we can think of a tax on the price of cigarettes, which induces the consumer to reduce cigarette consumption from $X$ to $X/(1+\delta)$. .\footnote{Suppose that $\alpha_{x}$ is the exponent of $X$ in the Cobb-Douglas utility function. Suppose further that the exponents are normalized to sum to 1. Then, if $M$ is the income, and $p_{x}$ is the price of $X$, we have that $\alpha_{x}M=p_{x} X$. Similarly, under the proposed counterfactual tax $\alpha_{x} M=p_{x}(1+\delta)X_{\delta}$. It follows that $X_{\delta}=X/(1+\delta)$.}

Table (ref) and Figure (ref) show the results. The effects are positive and monotonically increasing across quantiles. This means that the marginal impact on the birth weight of a tax on cigarettes is positive. The effects are stronger for upper quantiles of the distribution of birth weight. In order to interpret the magnitudes, we use the quantile-standard deviation elasticity. According to (ref), the elasticity can be calculated as $\mathcal{E} _{\tau,\delta=0}=-\Pi_{\tau,S}/Q_{\tau}[Y]$ as $\dot{s}\left( 0\right) =-1.$ For example, for $\tau=0.50$, $\mathcal{E}_{0.50,\delta=0}=-0.0128$. This means that a $1\%$ decrease in the standard deviation of the consumption of cigarettes increases the median birth weight by $0.0128\%$.

table[table omitted — 1,023 chars of source]
figure[figure omitted — 291 chars of source]

Conclusion

This paper has provided a general procedure to analyze the distributional impact of changes in covariates on an outcome variable. The standard unconditional quantile regression analysis focuses on a particular impact coming from a location shift. We have provided a framework to study the unconditional policy effects generated by a smooth and invertible intervention of one or more target variables, allowing them to be possibly endogeneous. We focus particularly on a location-scale shift and show how to additively decompose the total effect into a location effect and a scale effect. They can be analyzed and estimated separately. Additionally, we consider the case of simultaneous changes in different covariates. We show how this can be obtained from the usual vector-valued unconditional quantile regressions.