EconBase
← Back to paper

Structural Representations and Identification of Marginal Policy Effects

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

49,228 characters · 12 sections · 76 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Structural Representations and Identification of Marginal Policy Effects

\allowdisplaybreaks

titlepage\thispagestyle{empty} \hypersetup{linkcolor=black} (Y. Zhang);\, [email removed]} (Z. Zhang). }} \affil[a,c]{School of Economics, Shanghai University of Finance and Economics} \affil[b]{School of International Business, Shanghai University of International Business and Economics} \begin{abstract} This paper investigates the structural interpretation of the marginal policy effect (MPE) within nonseparable models. We demonstrate that, for a smooth functional of the outcome distribution, the MPE equals its functional derivative evaluated at the outcome-conditioned weighted average structural derivative. This equivalence is definitional rather than identification-based. Building on this theoretical result, we propose an alternative identification strategy for the MPE that complements existing methods. {\bfseriesKeywords}: Counterfactual distribution; Policy effect; Nonseparable model; Unconditional quantile regression. {\bfseriesJEL Classification}: C01, C21, C50. \end{abstract} \thispagestyle{empty}

Introduction

This paper investigates the structural interpretation of the marginal policy effect (MPE) within nonseparable models. The MPE quantifies the marginal impact of a counterfactual change in the distribution of a policy variable on specific features of the outcome distribution. While the identification of the MPE has been studied by firpo2009, rothe2010el,rothe2012 and mms2024, its general connection to underlying structural effects remains underexplored. Specifically, firpo2009 and rothe2010el examine this issue for the quantile functional.

This paper contributes to the literature by providing a structural interpretation of the MPE for general smooth functionals, thereby extending the analysis beyond the specific case of quantiles. Our structural analysis commences with the definition of the parameter itself, as opposed to its identification results. We demonstrate that, under regularity conditions, the MPE for a Hadamard-differentiable functional equals its functional derivative, evaluated at the outcome-conditioned weighted average structural derivative. For the quantile functional, we derive a structural decomposition of unconditional quantile regression (UQR) estimand defined by firpo2009. We further extend their proposition by showing that, under conditional independence and without requiring monotonicity in the structural error, the \(\tau\)-th UQR estimand equals the average structural derivative for individuals at the \(\tau\)-th quantile of the outcome distribution.

As a second contribution, this paper presents an alternative strategy for identifying the MPE. This strategy is applicable when the functional of interest is Hadamard-differentiable and the local average structural derivative, as proposed by hm2007,hm2009, is identifiable. We present two general identification propositions: one for the single-function model under conditional independence and the other for the triangular system using the control variable method of imbensNewey2009. These propositions can be used to derive or generalize existing identification results for MPEs.

The literature has extensively studied the structural interpretations of parameters and estimands. For instance, cherno2013 provide treatment effect interpretations for counterfactual distribution effects under unconfoundedness. sasaki2015 examines structural interpretations of conditional quantile regression estimands. The relationship between continuous quantile treatment effects and structural partial derivatives is also explored by su2019 for single-equation models, and by chernoPanel2015 for panel data models.

This paper is organized as follows. Section (ref) defines the target parameter. Section (ref) introduces the structural model and presents the main results. Section (ref) presents the general identification propositions and their applications. Section (ref) discusses the estimation approaches briefly. Section (ref) concludes. Complete proofs of all results are available in the Supplementary Appendix.

Target Parameter

Let \(Y\) denote the outcome variable and \(D\) the policy variable subject to intervention, with both variables taking values in $\mathbb{R}$.

definition[Policy Intervention and Counterfactual Outcome] Let $\pi_t:\mathbb{R} \to \mathbb{R}$ denote a policy function indexed by \(t\in[0,1]\). Policymakers can implement a marginal intervention through \(\pi_t\), perturbing the policy variable \(D\) to create the counterfactual variable \(D^t:=\pi_t(D)\). The counterfactual outcome corresponding to \(D^t\) is denoted by \(Y^t\). This outcome remains unobserved because the structural relationship between \(D^t\) and \(Y^t\) is unspecified.

The above policy intervention corresponds to a counterfactual change in the distribution of the policy variable, from its original cumulative distribution function (CDF) $F_D$ to $F_{D^t}$. For instance, an intervention implemented by $\pi_t(D)=D+t$ yields the counterfactual CDF $F_{D^t}(d)=\mathbb{P}[D+t\leq d]=F_D(d-t)$. Unlike the counterfactual change considered here, standard treatment effect models define potential outcomes \(Y(d)\) and evaluate the impact of changing the value of \(D\) from $d$ to $d'$ on functionals of $F_{Y(d)}$, independent of changes in the distribution of \(D\) itself. To further clarify the distinction, consider a linear random coefficient model \(Y(d)=\beta_0(\varepsilon) + \beta_1(\varepsilon)d\) with \(Y=Y(D)\) and \(Y^t=Y(D^t)\), where \(\varepsilon\) captures unobserved heterogeneity. In this context, the average policy effect is \(\mathbb{E}[Y^t-Y]=\mathbb{E}[(\pi_t(D)-D)\beta_1(\varepsilon)]\), while the average treatment effect is \(\mathbb{E}[Y(d')-Y(d)]=(d'-d)\mathbb{E}[\beta_1(\varepsilon)]\).

We now present several concrete examples of policy functions \(\pi_t\).

examplefirpo2009 and rothe2010el study the location shift intervention \(\pi_t(D)=D+t\), which represents a special case of the more general \(\pi_t(D)=\mu + l(t) + (D-\mu)s(t)\) proposed by mms2024. Here, \(\mu\) is a known parameter, while \(l(t)\) and \(s(t)>0\) represent the location and scale transformations, respectively. rothe2012 analyzes the rank-preserving transformation $\pi_t(D)= H_t^{-1}(F_D(D))$, where \(H_t^{-1}(\tau):=\inf\{d:H_t(d) \geq \tau\}\) is the \(\tau\)-th quantile of a perturbed distribution \(H_t\). Other examples include the mean-preserving transformation \(\pi_t(D) = \mathbb{E}[D] + (1+\alpha t)(D-\mathbb{E}[D])\), where \(\alpha \in [-1,1]\) governs the extent of dispersion.

Let \(\mathcal{F}\) denote the space of one-dimensional CDFs and \(\ell^\infty(\mathbb{T})\) the space of all bounded functions defined on a set \(\mathbb{T}\). Let \(F_{Y^t}(\cdot):=\mathbb{P}[Y^t\leq \cdot]\) denote the CDF of the counterfactual outcome \(Y^t\). We study the following parameter.

definitionFor a functional $\Gamma:\mathcal{F} \to \ell^\infty(\mathbb{R})$, the Marginal Policy Effect (MPE) of a counterfactual change in $D$ on $\Gamma(F_Y)$ is defined as \begin{align*} \theta_\Gamma:=\partial_t\Gamma(F_{Y^t})|_{t=0} :=\lim_{t \downarrow 0} \frac{\Gamma(F_{Y^t})-\Gamma(F_Y)}{t} \end{align*} provided that the limit exists.

We illustrate the above parameter through several representative examples of functionals \(\Gamma\).

exampleA theoretically useful special case is the identity mapping \(id_y:= F\mapsto F(y)\), which yields the distributional MPE: \begin{align*} \theta_{id}(y):=\lim_{t\downarrow 0} \frac{F_{Y^t}(y)-F_Y(y)}{t}. \end{align*} Another fundamental example is the quantile functional \(Q_\tau:= F \mapsto \inf\{y:F(y) \geq \tau\}\) for \(\tau \in (0,1)\), which generates the quantile MPE studied by firpo2009 and mms2024: \begin{align*} \theta_{Q}(\tau):= \lim_{t \downarrow 0} \frac{Q_\tau(F_{Y^t})-Q_\tau(F_Y)}{t}. \end{align*} For the mean functional \(\mu: F \mapsto \oldint\nolimits y \,dF(y)\), the average MPE is defined as \begin{align*} \theta_\mu := \lim_{t \downarrow 0}\frac{\mu(F_{Y^t})-\mu(F_Y)}{t}. \end{align*} Furthermore, \(\Gamma\) can be specified as various inequality measures. For instance, consider the Gini coefficient \(GC:F \mapsto 1 - 2\oldint\nolimits_{0}^{1}L_p(F)\,dp\), where \(L_p(F):=\frac{\oldint\nolimits_{0}^{p}Q_\tau(F)\,d\tau}{\oldint\nolimits y\,dF(y)}\) denotes the Lorenz curve. The Gini MPE is defined as \begin{align*} \theta_{GC}:=\lim_{t \downarrow 0}\frac{GC(F_{Y^t}) - GC(F_{Y})}{t}, \end{align*} which quantifies the marginal effect of a perturbation in \(F_D\) on the inequality of outcome $Y$.

Structural Interpretations

Nonseparable Models and Related Assumptions

We introduce the following nonseparable model to characterize the dependence structure.

assumptionSLet $m(\cdot,\cdot,\cdot)$ be a measurable function of unknown form. \begin{itemize} • The status quo outcome $Y$ is generated by \begin{align} Y=m(D,X,\varepsilon), \end{align} where $D$ is the continuous policy variable, $X$ is a vector of covariates, and $\varepsilon$ represents the unobserved disturbance. Suppose that $m(d,x,e)$ is continuously differentiable in $d$ for each $(x,e)$. The partial derivative $\partial_d m(\cdot,\cdot,\cdot)$ is termed the structural derivative. • The counterfactual outcome $Y^t$ is generated by \begin{align} Y^t=m(\pi_t(D),X,\varepsilon)=:m(D^t,X,\varepsilon), \end{align} where $\pi_0(d)\equiv d$ (status quo policy). For each $d$, assume $\pi_t(d)$ is continuously differentiable in $t$ at $t=0$, with bounded derivative $\dot{\pi}(d):=\partial_t\pi_t(d)|_{t=0}$. \end{itemize}

To derive the structural interpretation, the following conditions are required.

assumptionR[Regularity] \begin{itemize} • The CDF $F_Y$ is continuously differentiable on its support \(\mathcal{S}_{Y}\), with density $f_Y$ satisfying \(0<f_Y(y)<\infty\) for every \(y\in \mathcal{S}_{Y}\). • Let $\dot{m}(d,x,e):=\partial_t m(\pi_t(d),x,e)|_{t=0}=\dot{\pi}(d)\partial_d m(d,x,e)$. Suppose that $\dot{m}(\cdot,\cdot,\cdot)$ is measurable and satisfies \begin{align*} \mathbb{P}\Big[ \big| m(\pi_t(D),X,\varepsilon) - m(D,X,\varepsilon) - t \,\dot{m}(D,X,\varepsilon) \big| \geq \nu t \Big] = o(t) \end{align*} as $t \downarrow 0$ for every $\nu > 0$. • The distribution of $(Y,\dot{m}(D,X,\varepsilon))$ is absolutely continuous with density $f_{Y,\dot{m}}(y,y')$ that is continuous in $y$ for each $y'$. Furthermore, there exists a Lebesgue integrable function $g:\mathbb{R} \to \mathbb{R}$ satisfying $\oldint\nolimits |y' g(y')|\,dy' < \infty$ and a positive constant \(C\) such that $f_{Y,\dot{m}}(y,y') \leq C |g(y')|$ for all \((y,y')\). \end{itemize}

Assumption (ref) does not impose structural conditions like the invariance of \(F_{Y|D,X}\) under small manipulations, as assumed in firpo2009. Assumption (ref) (a) implies that $Y$ is continuously distributed with a regular density. Assumptions (ref)(b)--(c) are analogous to hm2007's (hm2007) Assumptions A3--A4, implying smooth variation in the structural response $\dot{m}$.

The directional derivative structure of \(\theta_\Gamma\) motivates our focus on smooth functionals \(\Gamma\). Hadamard-differentiable functionals encompass key causal statistics for policy analysis, including means, quantiles, and common inequality measures (e.g., the Gini coefficient).\footnote{For additional Hadamard-differentiable inequality measures, see, e.g., rothe2010joe and firpo2016.} This differentiability is necessary for both the functional Delta method (vaart2000) and standard bootstrap inference (fang2019).

Formally, let $(\mathbb{D},\|\cdot\|_{\mathbb{D}})$ and $(\mathbb{B},\|\cdot\|_{\mathbb{B}})$ be normed spaces. A functional $\Gamma:\mathbb{D}_\Gamma \subseteq \mathbb{D} \to \mathbb{B}$ is Hadamard-differentiable at $F\in\mathbb{D}_\Gamma$ tangentially to a set $\mathbb{D}_0 \subseteq \mathbb{D}$, if there exists a continuous linear functional $\Gamma'_{F}: \mathbb{D}_0 \to \mathbb{B}$ such that

align*[align* omitted — 136 chars of source]

for every sequence $\{h_t\}\subset \mathbb{D}$ satisfying $h_t \to h\in\mathbb{D}_0$ and $F+th_t\in \mathbb{D}_\Gamma$, where $\Gamma'_F$ is called the Hadamard derivative of $\Gamma$ at $F$. See, for example, vaart2000 for more details.

Structural Representation of the MPE

theorem[Structural Representation of $\theta_\Gamma$] Under Assumptions (ref) and (ref), we obtain \begin{align*} \theta_{id}(y) :=\lim_{t\downarrow 0} \frac{F_{Y^t}(y)-F_Y(y)}{t} = \mathbb{E}\big[ \omega^f(y,D)\,\partial_d m(D,X,\varepsilon) \big| Y=y \big] \end{align*} for every $y\in \mathcal{S}_{Y}$, where $\omega^f(y,d):=-f_Y(y)\dot{\pi}(d)$. If $\Gamma$ is Hadamard-differentiable at $F_Y$, we have \begin{align*} \theta_\Gamma = \Gamma'_{F_Y} \Big( \mathbb{E}\big[ \omega^f(\cdot,D)\,\partial_d m(D,X,\varepsilon) \big| Y=\cdot \big]\Big), \end{align*} where \(\Gamma'_{F_Y}\) denotes the Hadamard derivative of \(\Gamma\) at \(F_Y\).

Theorem (ref) states that: (i) the marginal effect of counterfactual changes in \(F_D\) on the outcome CDF \(F_Y(y)\), equals the weighted average structural effect for individuals with \(Y=y\). (ii) For any Hadamard-differentiable transformation \(\Gamma\), the marginal impact on \(\Gamma(F_Y)\) is given by the Hadamard derivative of \(\Gamma\) at \(F_Y\), applied to the weighted average structural effect function identified in (i).

Theorem (ref) extends the literature in two principal ways: (i) When considering \(\pi_t(\cdot) = H^{-1}_t(F_D(\cdot))\), where \(H_t\) is the counterfactual CDF of \(F_D\), Theorem (ref) offers a structural interpretation of rothe2012's (rothe2012) marginal partial distributional policy effect. (ii) For \(\Gamma=Q_\tau\), it provides a structural interpretation for the unconditional quantile effect studied by mms2024, as formalized below.

corollary[Structural Representation of $\theta_Q$] For \(\Gamma=Q_\tau\), under the assumptions of Theorem (ref) and assuming the support \(\mathcal{S}_{Y}\) is compact, we obtain \begin{align*} \theta_Q(\tau)=\mathbb{E}\big[ \dot{\pi}(D)\,\partial_d m(D,X,\varepsilon) \big| Y=q_\tau \big] \end{align*} for every \(\tau \in (0,1)\), where \(q_\tau := Q_Y(\tau)\).

Corollary (ref) states that the MPE of a counterfactual change in \(F_D\) on the \(\tau\)-th quantile of the outcome distribution, \(Q_\tau(F_Y)\), equals the weighted average structural derivative for individuals with outcome value \(q_\tau\), with weights determined by the policy function \(\pi_t\).

Structural Representation of the UQR

This section extends the structural interpretation of the unconditional quantile partial effect (UQPE) proposed by firpo2009. For comparison purposes, we focus on the location shift \(D^t=D+t\) and denote \(\theta_Q\) as \(\theta_Q^L\) for this specific case. The UQPE of firpo2009 is defined as

align*[align* omitted — 123 chars of source]

We denote this UQPE as \(\beta^{\mathrm{UQR}}(\tau)\), which corresponds to their unconditional quantile regression (UQR) estimand rather than an unobserved parameter.

Note that \(\theta_Q^L\) is not equal to \(\beta^{\mathrm{UQR}}\) under the standard regularity assumptions mentioned above. Consequently, Corollary (ref) cannot be used to structurally interpret \(\beta^{\mathrm{UQR}}\) as \( \mathbb{E}[\partial_dm(D,X,\varepsilon)|Y=q_\tau]\) under these assumptions. firpo2009, however, demonstrate that \(\theta_Q^L(\tau)=\beta^{\mathrm{UQR}}(\tau)\) holds under the structural assumption \(F_{Y^t|D+t,X} = F_{Y|D,X}\). A broader question thus emerges: What does the UQR estimand $\beta^{\mathrm{UQR}}$ identify in nonseparable models when structural assumptions are not imposed? To address this question formally, we introduce the following conditions.

assumptionR[Regularity] For each $(d,x)\in\mathcal{S}_{D}\times\mathcal{S}_{X}$: \begin{itemize} • The conditional distribution of $Y$ given $(D,X)$ is absolutely continuous, with density $f_{Y|D,X}(\cdot|d,x)$ that is strictly positive, bounded and continuous on its support. Furthermore, \(F_{Y|D,X}(y|d,x)\) is continuously differentiable in \(d\) for each $(y,x)$. • $\partial_d m(d,x,\cdot)$ is measurable and satisfies \[ \mathbb{P}\Big[ \big| m(d+\delta,x,\varepsilon) - m(d,x,\varepsilon) - \delta\,\partial_d m(d,x,\varepsilon) \big| \geq \nu \delta \Big| D=d,X=x \Big] = o(\delta) \] as $\delta \downarrow 0$ for every $\nu > 0$. • The conditional distribution of \((Y,\partial_d m(d,x,\varepsilon))\) given \((D,X)\) is absolutely continuous, with density $f_{Y,\partial_dm^{d,x}|D,X}(\cdot,y'|d,x)$ that is continuous for each $(y',d,x)$. Furthermore, there exists a Lebesgue integrable function $g:\mathbb{R}\to\mathbb{R}$ satisfying $\oldint\nolimits |y'g(y')|\,dy'<\infty$ and a positive constant $C$ such that $f_{Y,\partial_dm^{d,x}|D,X}(y,y'|d,x) \leq C |g(y')|$ for every \((y,y',d,x)\). • $\mathbb{P}[m(d',x,\varepsilon)=Q_{Y|D,X}(\alpha|d,x)|D=d,X=x]=0$ for $d'$ in a neighborhood of $d$. The conditional distribution of \(\varepsilon\) given \((D,X)\) is absolutely continuous with the strictly positive density $f_{\varepsilon|D,X}(\cdot|d,x)$. Furthermore, $f_{\varepsilon|D,X}(e|d,x)$ is partially differentiable in \(d\) for each \((e,x)\), and satisfies \(\sup_{|\delta| \leq c} \left| \frac{f_{\varepsilon|D,X}(e|d+\delta,x)-f_{\varepsilon|D,X}(e|d,x)}{\delta}\right| < \infty\) for some positive constant $c$. \end{itemize}

Assumptions (ref)(a)--(b) correspond to hm2007's (hm2007) Assumptions A2--A3, where the continuous differentiability of \(F_{Y|D,X}(y|\cdot,x)\) is equivalent to that of \(Q_{Y|D,X}(\alpha|\cdot,x)\). Assumptions (ref)(c) parallels Assumption A4 in hm2007, ensuring smooth variation in the structural effect \(\partial_dm\). Finally, Assumption (ref)(d) aligns with Assumption A1 in hm2009, which facilitates decomposition of the endogeneity-induced bias in the UQR estimand.

theorem[Structural Representation of \(\beta^{\mathrm{UQR}}\)] Under Assumptions (ref) and (ref), we obtain \begin{align*} \beta^{\mathrm{UQR}}(\tau) =\mathbb{E}\big[ \partial_dm(D,X,\varepsilon) \big| Y=q_\tau \big] - \frac{\mathbb{E}\big[ \usefont{U}{bbold}{m}{n}1\{ Y\leq q_\tau \}\,\partial_d{\ln\big(f_{\varepsilon|D,X}(\varepsilon|D,X)\big)} \big]}{f_Y(q_\tau)} \end{align*} for every $\tau \in (0,1)$.

Theorem (ref) generalizes Proposition 1 of firpo2009. Note that if conditional independence $\varepsilon \perp \!\!\! \perp D|X$ are satisfied, we have \(\partial_d\ln(f_{\varepsilon|D,X}(\varepsilon|d,x)) = 0\) and

align[align omitted — 132 chars of source]

Therefore, under conditional independence, the estimand \(\beta^{\mathrm{UQR}}(\tau)\) is interpreted as the average structural effect for individuals with outcome value \(q_\tau\). Furthermore, if $e\mapsto m(d,x,e)$ is strictly monotonic, as assumed by firpo2009's (firpo2009) Proposition 1, we can derive their result using Equation ((ref)).

remarkFor identifying \(\theta_Q^L\) via \(\theta_Q^L(\tau)=\beta^{\mathrm{UQR}}(\tau)\), the conditional independence assumption $\varepsilon \perp \!\!\! \perp D|X$ is stronger than the distributional invariance condition \(F_{Y^t|D+t,X}=F_{Y|D,X}\) used by firpo2009. This is because \[F_{Y^t|D+t,X}(\cdot|d,x)=\oldint\nolimits \text{\usefont{U}{bbold}{m}{n}1}\{ m(d,x,e)\leq \cdot \}\,dF_{\varepsilon|D+t,X}(e|d,x),\] and under $\varepsilon \perp \!\!\! \perp D|X$, we have \(F_{\varepsilon|D+t,X}(e|d,x)=F_{\varepsilon|X}(e|x) = F_{\varepsilon|D,X}(e|d,x)\). Nonetheless, the conditional independence remains essential for the structural interpretation of $\beta^{\mathrm{UQR}}(\tau)$, as established in Theorem (ref) or initially demonstrated by firpo2009.

To conclude this section, Table (ref) summarizes the structural interpretations of the various MPEs. All reported results follow directly from Theorem (ref) or Theorem (ref).

table[table omitted — 1,519 chars of source]

Applying Theorem (ref) for Identification

Theorem (ref) suggests a practical identification strategy for \(\theta_\Gamma\), which we term the LASD method. This approach builds on the identification of the local average structural derivative (LASD) proposed by hm2007. This quantity is defined as \[\theta_{\mathrm{LASD}}(d,x,y):=\mathbb{E}\big[ \partial_d m(d,x,\varepsilon) \big| D=d,X=x,Y=y \big]. \] We note that, by the law of iterated expectations, Theorem (ref) implies the following relationship:

align[align omitted — 182 chars of source]

where the weight $\omega^f(y,d):=-f_Y(y)\dot{\pi}(d)$ is estimable from data. Consequently, identifying \(\theta_\Gamma\) reduces to identifying \(\theta_{\mathrm{LASD}}\).

This section establishes two general identification results for \(\theta_\Gamma\): one applicable to single-equation models and another for triangular systems. Prior to formal analysis, we compare the key identification step of the LASD method (Equation ((ref))) with existing approaches in the literature.

Comparison of the MPE Identification Strategies

firpo2009 and mms2024 employ influence function (IF) approximations for identification purposes. firpo2009 analyze the counterfactual distribution \(F_{Y^t} =\oldint\nolimits F_{Y|D,X}\,dF_{D+t,X}\), assuming distributional invariance \( F_{Y^t|D+t,X}=F_{Y|D,X}\). Their MPE identification builds on the G\^ateaux derivative characterization of the IF:

align*[align* omitted — 93 chars of source]

where \(\mathrm{RIF}(y;\Gamma,F_Y):=\Gamma(F_Y) + \mathrm{IF}(y;\Gamma,F_Y)\) denotes the recentered influence function (RIF) and the distribution \(G^*_Y\) is constructed to satisfy the first-order local approximation \(F_{Y^t} = F_Y + t(G^*_Y-F_Y) + \mathcal{O}(t^2)\) as \(t \downarrow 0\). This methodology applies to G\^ateaux-differentiable functionals \(\Gamma\), a weaker condition than Hadamard differentiability, while specifically designed to evaluate counterfactual effects stemming solely from the location shift \(D^t=D+t\) in \(F_D\).

mms2024 examine general counterfactual transformations \(\pi_t(D)\) and employ the IF approximation of the Hadamard-differentiable functionals \(\Gamma\) for identification. Their approach builds on

align*[align* omitted — 91 chars of source]

where \(\theta_{id}\) is identified under the conditional independence. This methodology differs from the LASD approach in two key respects: (i) It uses the analytic form of the influence function \(\mathrm{IF}(\cdot;\Gamma,F_Y)\) rather than the Hadamard derivative \(\Gamma'_{F_Y}\). (ii) It requires different smoothness conditions than Assumption (ref), including continuous differentiability of both \(f_{D,X}(\cdot,x)\) and \(f_{Y|D,X}(y|\cdot,x)\) for each \(x\in \mathcal{S}_{X}\) and \(y\in \mathcal{S}_{Y} \).

Differing from the IF approximation approach, rothe2010el,rothe2012 establishes the identification building on the Hadamard derivative characterization

align*[align* omitted — 66 chars of source]

While our LASD method also employs this form, the key distinction lies in the implementation: Rothe's method directly identifies \(\theta_\Gamma\) through \(\theta_{id}\) for specific marginal perturbations of \(F_D\) (e.g. \(D^t=D+t\) or \(H_t^{-1}(F_D(D))\)), whereas our approach first establishes the relationship between \(\theta_\Gamma\) and \(\theta_{\mathrm{LASD}}\) for general perturbations \(D^t=\pi_t(D)\), and then identifies \(\theta_\Gamma\) via \(\theta_{\mathrm{LASD}}\). This leads to the general identification formulas for \(\theta_\Gamma\) that include existing results as special cases, as we demonstrate below.

Single-Equation Model

propositionUnder the assumptions of Theorem (ref), Assumption (ref)(a)--(c), and the conditional independence $\varepsilon \perp \!\!\! \perp D \,|X$, we obtain \begin{align*} \theta_\Gamma = \Gamma'_{F_Y} \Big(\mathbb{E}\big[ \omega^f(\cdot,D) \beta(D,X,\cdot) \big| Y=\cdot \big] \Big) \end{align*} for every Hadamard-differentiable functional \(\Gamma\), where \(\beta(d,x,y):=-\frac{\partial_d{F_{Y|D,X}(y|d,x)}}{f_{Y|D,X}(y|d,x)}\).

Proposition (ref) provides a general identification formula of \(\theta_\Gamma\) in the single-function model ((ref)), where \(\beta(d,x,y)\) identifies \(\theta_{\mathrm{LASD}}(d,x,y)\) under the conditional independence (hm2007).

Turning to the quantile case, we define the conditional $\alpha$-quantile derivative (CQD) as $\beta^{\mathrm{CQD}}(\alpha,d,x):=\partial_d{Q_{Y|D,X}(\alpha|d,x)}$, corresponding to what firpo2009 call the conditional quantile partial effect (CQPE).

corollaryFor \(\Gamma=Q_\tau\), under the assumptions of Proposition (ref) and assuming the support \(\mathcal{S}_{Y}\) is compact, we obtain \begin{align} \theta_Q(\tau) &= \frac{-1}{f_{Y}(q_\tau)}\oldint\nolimits \dot{\pi}(d)\,\partial_d F_{Y|D,X}(q_\tau|d,x) \,dF_{D,X}(d,x) \\ &= \oldint\nolimits \dot{\pi}(d)\,\left(\omega_{q_\tau}(d,x)\cdot\beta^{\mathrm{CQD}}(\zeta_\tau(d,x),d,x)\right)\,dF_{D,X}(d,x)\\ &=\oldint\nolimits \dot{\pi}(d)\,\beta^{\mathrm{CQD}}(\zeta_\tau(d,x),d,x)\,dF_{D,X|Y}(d,x|q_\tau), \end{align} for every \(\tau \in (0,1)\), where \(\omega_y(d,x):=f_{Y|D,X}(y|d,x)/f_Y(y)\), and \(\zeta_\tau(d,x):=\{\alpha\in(0,1):Q_{Y|D,X}(\alpha|d,x)=q_\tau\}\) denotes the `matching' function.

Equation ((ref)) coincides with Corollary 1 of mms2024. Furthermore, under location shifts \(\pi_t(d)=d+t\) (where \(\dot{\pi}(d)\equiv 1\)), Equation ((ref)) recovers Proposition 1(ii) of firpo2009, and Equation ((ref)) aligns with Lemma 1 of alejo2024.

Proposition (ref) can also be used to establish rothe2012's (rothe2012) Theorem 4 under the additional regularity conditions introduced therein.

corollaryLet the assumptions of Proposition (ref) hold. For \(D^t=H_t^{-1}(F_D(D))\), suppose that for each \( t\in [0,1]\), (a) \(\mathcal{S}_{D^t} \subseteq \mathcal{S}_{D}\); (b) \(H_t\) is continuous, and $f_D > 0$ on $\mathcal{S}_{D}$; (c) the copula \(C(\cdot,\cdot)\) of \((D^t,X)\) is partially differentiable with respect to its first component. Then: \begin{itemize} • For a marginal perturbation \(H_t(d) = F_D(d) + t\,(G_D(d) - F_D(d))\), where \(G_D\) denotes a CDF of \(D\), we have \(\dot{\pi}(d) = -(G_D(d)-F_D(d))/f_D(d)\), and \begin{align*} \theta_\Gamma &= \oldint\nolimits \Gamma'_{F_Y} \big(F_{Y|D,X}(\cdot|d,x)\big)\,d\left(F_{X|D}(x|d)(G_D(d)-F_D(d))\right) \\ &= \oldint\nolimits \Gamma'_{F_Y}\big(F_{Y|D,X}(\cdot|d,x)\big)\,d\left(\partial_t C\big(H_t(d),F_X(x)\big)|_{t=0}\right). \end{align*} • For a location shift \(F_{D^t}(d)=F_{D}(d-t)\), we have \(\theta_\Gamma=\Gamma'_{F_Y}\big(\mathbb{E}[\partial_d F_{Y|D,X}(\cdot|D,X)]\big)\). \end{itemize}

Note that rothe2012 adopts the conditional independence assumption \( \varepsilon \perp \!\!\! \perp U|X\) for identification, where \(U:=F_D(D)\) represents the rank of $D$ in its distribution. This condition is equivalent to \(\varepsilon \perp \!\!\! \perp D|X\) when the policy variable $D$ is continuously distributed.

Triangular System

assumptionSMaintain Assumption (ref), but replace Equation ((ref)) for the status quo outcome $Y$ with the following system: \begin{align} \begin{split} Y&= m(D,X,\varepsilon), \\ D&= h(Z,X,\eta), \end{split} \end{align} where $Z$ is an observed variable taking values in $\mathbb{R}$, and $\eta$ denotes the unobserved disturbances in the selection equation.

We establish identification using a control variable approach, following an analogous construction to that in Theorem 1 of imbensNewey2009.

lemmaC[Control Variable] In the triangular model defined by equations ((ref)), assume that (a) \((\varepsilon,\eta) \perp \!\!\! \perp Z|X\); (b) \(\eta\) is one-dimensional and \(p \mapsto h(Z,X,p)\) is strictly monotonic in \(p\) with probability 1; (c) for each \(x\), \(F_{\eta|X}(\cdot|x)\) is continuous and strictly increasing on its support. Then, there exists a control variable \(V=F_{D|Z,X}(D|Z,X)=F_{\eta|X}(\eta|X)\) such that $\varepsilon \perp \!\!\! \perp D \,|V, X$.
propositionUnder the assumptions of Theorem (ref), Lemma (ref), and Assumption R3 (stated in Supplementary Appendix A), we obtain \begin{align*} \theta_\Gamma = \Gamma'_{F_Y} \Big(\mathbb{E}\big[ \omega^f(\cdot,D)\,\beta^{\mathrm{CV}}(D,V,X,\cdot) \big| Y=\cdot \big] \Big) \end{align*} for every Hadamard-differentiable functional \(\Gamma\), where $\beta^{\mathrm{CV}}(d,v,x,y):=-\frac{\partial_d{F_{Y|D,V,X}(y|d,v,x)}}{f_{Y|D,V,X}(y|d,v,x)}$.

Proposition (ref) presents a general identification formula for \(\theta_\Gamma\) in the triangular system ((ref)), where \(\mathbb{E}[\beta^{\mathrm{CV}}(d,V,x,y)|D=d,X=x,Y=y] \) identifies \(\theta_{\mathrm{LASD}}(d,x,y)\) under $\varepsilon \perp \!\!\! \perp D \,|V, X$.

Let $\beta^{\mathrm{CQD}}(\alpha,d,v,x):=\partial_d Q_{Y|D,V,X}(\alpha|d,v,x)$. The following result extends Corollary (ref) to the triangular system.

corollaryFor \(\Gamma=Q_\tau\), under the assumptions of Proposition (ref) and assuming the support \(\mathcal{S}_{Y}\) is compact, we obtain \begin{align} \theta_Q(\tau) &= \frac{-1}{f_Y(q_\tau)}\oldint\nolimits \dot{\pi}(d)\,\partial_d F_{Y|D,V,X}(q_\tau|d,v,x)\,dF_{D,V,X}(d,v,x) \\ &= \oldint\nolimits \dot{\pi}(d) \Big( \omega_{q_\tau}(d,v,x)\, \beta^{\mathrm{CQD}}\big(\zeta_\tau(d,v,x),d,v,x\big)\Big)\,dF_{D,V,X}(d,v,x) \\ &= \oldint\nolimits \dot{\pi}(d)\, \beta^{\mathrm{CQD}}\big(\zeta_\tau(d,v,x),d,v,x\big)\,dF_{D,V,X|Y}(d,v,x|q_\tau) \end{align} for every $\tau\in(0,1)$, where $\omega_{y}(d,v,x):=\frac{f_{Y|D,V,X}(y|d,v,x)}{f_Y(y)}$, and $\zeta_\tau(d,v,x):=\{\alpha\in(0,1):Q_{Y|D,V,X}(\alpha|d,v,x)=q_\tau\}$.

The following result extends rothe2012's (rothe2012) Theorem 4 to the triangular system, with part (ii) aligning with Theorem 1 of rothe2010el.

corollaryLet the assumptions of Proposition (ref) and Conditions (a)--(b) in Corollary (ref) hold. For \(D^t=H_t^{-1}(F_D(D))\), suppose the copula \(C(\cdot,\cdot,\cdot)\) of \((D^t,V,X)\) is partially differentiable with respect to its first component. Then: \begin{itemize} • For a marginal perturbation \(H_t(d) = F_D(d) + t\,(G_D(d) - F_D(d))\), where \(G_D\) denotes a CDF of \(D\), we have \begin{align*} \theta_\Gamma &= \oldint\nolimits \Gamma'_{F_Y} \big(F_{Y|D,V,X}(\cdot|d,v,x)\big)\,d\left(F_{V,X|D}(v,x|d)(G_D(d)-F_D(d))\right) \\ &= \oldint\nolimits \Gamma'_{F_Y}\big(F_{Y|D,V,X}(\cdot|d,v,x)\big)\,d\left(\partial_t C\big(H_t(d),F_V(v),F_X(x)\big)|_{t=0}\right). \end{align*} • For a location shift \(H_t(d)=F_{D}(d-t)\), we have \(\theta_\Gamma=\Gamma'_{F_Y}\big(\mathbb{E}[\partial_d F_{Y|D,V,X}(\cdot|D,V,X)]\big)\). \end{itemize}

Notably, baltagi2019 analyze counterfactual changes in the conditional distribution \(F_{D|V}\) within a triangular system, rather than changes in the marginal distribution \(F_D\). Consequently, their results do not follow from Proposition (ref) in our framework.

We conclude this section by presenting in Table (ref) a summary of our identification results.

table[table omitted — 1,917 chars of source]

Overview of Estimation Approaches

The general identification results in Section (ref) motivate the following estimation procedure for \(\theta_\Gamma\):

itemize• For specified \(\pi_t\) and \(\Gamma\), compute the Hadamard derivative \(\Gamma'_{F_Y}\) and simplify the identification formula according to Proposition (ref) or (ref). • Implement the plug-in estimator based on the simplified formula. For high-dimensional covariates \(X\), employ the Neyman orthogonal score approach instead.

As an illustration, Corollary (ref) yields the following nonparametric estimators for the quantile MPE:

align[align omitted — 420 chars of source]

where \(K(\cdot)\) denotes a kernel function. When \(\pi_t(d)=d+t\), Equation ((ref)) corresponds to the RIF-NP estimator of firpo2009, and Equation ((ref)) is the reweighting estimator proposed by alejo2024. See their respective papers for detailed computational procedures of each estimator.

When the covariate vector \(X\) is high-dimensional, the nonparametric first steps (e.g., \(\partial_dF_{Y|D,X}\)) typically lead to large biases in the plug-in estimator \(\hat{\theta}_\Gamma\) (ddml2018,localRobust2022). Using the orthogonal scores helps reduce the sensitivity with respect to the first steps to estimate \(\theta_\Gamma\). We continue with the example of \(\theta_{Q}(\tau)\). Let \(w:=(y,d,x)\) and \(\gamma_0(d,x,\tau):=F_{Y|D,X}(q_\tau|d,x)\). The orthogonal score of \(\theta_Q\) is given by

align*[align* omitted — 218 chars of source]

where \(\alpha(d,x):= \frac{\partial_d[\dot{\pi}(d)f_{D,X}(d,x)]}{f_{D,X}(d,x)}\) denotes the Riesz representer. The function $\psi_{Q}$ satisfies the moment condition: \(\mathbb{E}[\psi_{Q}(W,\theta_Q(\tau),\gamma_0,\tau)]=0\), and the Neyman orthogonality condition: \( \partial_\delta\mathbb{E}[\psi_{Q}(W,\theta_Q(\tau),\gamma+\delta(\gamma-\gamma_0),\tau)]|_{\delta=0}=0\) for every \(\gamma\) in a convex subset of some normed vector space. Then, the debiased estimator of \(\theta_Q(\tau)\) is given by

align*[align* omitted — 312 chars of source]

where \(\hat{\alpha}\) can be computed using the automatic estimation approach (adml2022), and \(\hat{F}_{Y|D,X}\) can be estimated via the post-Lasso penalized logistic regression (sasaki2022). When \(\pi_t(d)=d+t\), \(\check{\theta}_Q(\tau)\) corresponds to the high dimensional UQR estimator proposed by sasaki2022.

In summary, our identification framework delivers plug-in estimators for general MPEs \(\theta_\Gamma\) in low-dimensional settings. In high-dimensional settings, constructing debiased estimators requires \(\Gamma\)-specific Neyman orthogonal scores. To our knowledge, developed orthogonal score functions are lacking even for commonly used inequality-related MPEs,\footnote{In the binary treatment setting, firpo2016 develop semiparametrically efficient estimators for inequality treatment effects.} such as the Gini coefficient and the Theil index.

Conclusion

This paper demonstrates that, under regularity conditions, the marginal policy effect (MPE) for a Hadamard-differentiable functional equals its functional derivative evaluated at the outcome-conditioned weighted average structural derivative. This equivalence is definitional rather than dependent on specific identification assumptions. The main theorem offers an alternative identification strategy for the MPE, using the identification of the Local Average Structural Derivative (LASD) as an intermediate step.

This study points to several avenues for future research. First, the LASD identification framework can be extended to other nonseparable structural models, such as panel data models, where identifying the relevant LASD objects poses the central theoretical challenge. Second, developing debiased estimators for commonly used inequality-related MPEs (e.g., the Gini coefficient) in high-dimensional settings remains an important avenue for future work.

thebibliography\bibitem[Alejo et al.(2024)]{alejo2024} Alejo, J., Galvao, A. F., Mart\'inez-Iriarte, J., and Montes-Rojas, G. (2024). “Unconditional quantile partial effects via conditional quantile regression.” Journal of Econometrics, 105678. \bibitem[Baltagi and Ghosh(2019)]{baltagi2019} Baltagi, B. H., and Ghosh, P. K. (2019). “Partial distributional policy effects under endogeneity.” Sankhya B, 81(Suppl 1), 123-145. \bibitem[Chernozhukov et al.(2013)]{cherno2013} Chernozhukov, V., Fern\'andez-Val, I., and Melly, B. (2013). ”Inference on counterfactual distributions.” Econometrica, 81(6), 2205-2268. \bibitem[Chernozhukov et al.(2015)]{chernoPanel2015} Chernozhukov, V., Fern\'andez-Val, I., Hoderlein, S., Holzmann, H., and Newey, W. (2015). “Nonparametric identification in panels using quantiles.” Journal of Econometrics, 188(2), 378-392. \bibitem[Chernozhukov et al.(2018)]{ddml2018} Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W., and Robins, J. (2018). “Double/debiased machine learning for treatment and structural parameters.” The Econometrics Journal, 21(1), C1-C68. \bibitem[Chernozhukov et al.(2022a)]{adml2022} Chernozhukov, V., Newey, W. K., and Singh, R. (2022a). “Automatic debiased machine learning of causal and structural effects.” Econometrica, 90(3), 967-1027. \bibitem[Chernozhukov et al.(2022b)]{localRobust2022} Chernozhukov, V., Escanciano, J. C., Ichimura, H., Newey, W. K., and Robins, J. M. (2022b). “Locally robust semiparametric estimation.” \textit{Econometrica}, 90(4), 1501-1535. \bibitem[Fang and Santos(2019)]{fang2019} Fang, Z., and Santos, A. (2019). “Inference on directionally differentiable functions.” \textit{The Review of Economic Studies}, 86(1), 377-412. \bibitem[Firpo et al.(2009)]{firpo2009} Firpo, S., Fortin, N. M., and Lemieux, T. (2009), “Unconditional quantile regressions.” \textit{Econometrica}, 77 (3), 953--973. \bibitem[Firpo and Pinto(2016)]{firpo2016} Firpo, S., and Pinto, C. (2016). “Identification and estimation of distributional impacts of interventions using changes in inequality measures.” \textit{Journal of Applied Econometrics}, 31(3), 457-486. \bibitem[Hoderlein and Mammen(2007)]{hm2007} Hoderlein, S., and Mammen, E. (2007). ”Identification of marginal effects in nonseparable models without monotonicity.” \textit{Econometrica}, 75(5), 1513-1518. \bibitem[Hoderlein and Mammen(2009)]{hm2009} Hoderlein, S., and Mammen, E. (2009). “Identification and estimation of local average derivatives in nonseparable models without monotonicity.” \textit{The Econometrics Journal}, 12(1), 1-25. \bibitem[Imbens and Newey(2009)]{imbensNewey2009} Imbens, G. W., and Newey, W. K. (2009). “Identification and estimation of triangular simultaneous equations models without additivity.” \textit{Econometrica}, 77 (5), 1481--1512. \bibitem[Mart\'inez-Iriarte et al.(2024)]{mms2024} Mart\'inez-Iriarte J, Montes-Rojas G, Sun Y (2024), “Unconditional effects of general policy interventions.” \textit{Journal of Econometrics}, 238(2): 105570. \bibitem[Rothe(2010a)]{rothe2010joe} Rothe, C. (2010a). “Nonparametric estimation of distributional policy effects.” \textit{Journal of Econometrics}, 155(1), 56-70. \bibitem[Rothe(2010b)]{rothe2010el} Rothe, C. (2010b), “Identification of unconditional partial effects in nonseparable models.” \textit{Economics Letters}, 109 (3), 171--174. \bibitem[Rothe(2012)]{rothe2012} Rothe, C. (2012), “Partial distributional policy effects.” \textit{Econometrica}, 80 (5), 2269--2301. \bibitem[Sasaki(2015)]{sasaki2015} Sasaki, Y. (2015). “What do quantile regressions identify for general structural functions?” \textit{Econometric Theory}, 31(5), 1102-1116. \bibitem[Sasaki et al.(2022)]{sasaki2022} Sasaki, Y., Ura, T., and Zhang, Y. (2022). “Unconditional quantile regression with high-dimensional data.” \textit{Quantitative Economics}, 13(3), 955-978. \bibitem[Su et al.(2019)]{su2019} Su, L., Ura, T., and Zhang, Y. (2019). “Nonseparable models with high-dimensional data.” \textit{Journal of Econometrics}, 212(2), 646-677. \bibitem[van der Vaart(2000)]{vaart2000} van der Vaart, AW (2000), “Asymptotic Statistics.” Cambridge University Press \bibitem[Wooldridge(2004)]{wooldridge2004} Wooldridge, J. M. (2004). “Estimating average partial effects under conditional moment independence assumptions (No. CWP03/04).” cemmap working paper.
titlepage\onehalfspacing \hypersetup{linkcolor=Black} \begin{center} Supplementary Appendix to “Structural Representations and Identification of Marginal Policy Effects” Zhixin Wang, Yu Zhang, and Zhengyu Zhang \end{center} \begin{abstract} In this supplementary appendix, Section (ref) provides the proofs of the main results, while Section (ref) presents additional results on the average and Gini marginal policy effects that were omitted from the main text. \end{abstract} \ToC \thispagestyle{empty}