EconBase
← Back to paper

Identification and Inference of Partial Effects in Sharp Regression Kink Designs

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

120,598 characters · 17 sections · 86 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Unified Framework for Identification and Inference of Local Treatment Effects in Sharp Regression Kink Designs

\allowdisplaybreaks

titlepage\onehalfspacing } \begin{abstract} This paper develops a unified framework for the identification, estimation, and uniform inference of local treatment effects (LTEs) in sharp regression kink designs (RKDs). These LTEs quantify the effect of a marginal change in the treatment at the kink point on various features of the outcome distribution. The identification strategy applies to Hadamard-differentiable functionals of the outcome distribution---including means, quantiles, and inequality measures---and encompasses several existing RKD estimands as special cases. For estimation, we categorize the corresponding estimands into two general classes and implement their estimation via local polynomial constrained regression. We establish the asymptotic theory for this framework and provide a valid resampling procedure for uniform inference. The method is applied to examine the effect of unemployment insurance on unemployment durations, focusing on the policy's impact on the distribution and inequality of durations, as a complement to existing empirical evidence. {Keywords}: Continuous treatments; regression kink design; nonparametric identification; nonseparable models; multiplier bootstrap. \end{abstract} \thispagestyle{empty}

Introduction

The regression kink design (RKD) exploits the change in the slope of the relationship between an endogenous treatment variable and a covariate. This quasi-experimental method has become a popular approach in empirical economics, particularly for studying the labor supply effects of social insurance (landais2015assessing,card2015effect,kolsrud2018optimal,landais2021value). Formal econometric frameworks for the RKD are provided by card2015inference and chiang2019causal. These frameworks establish that conventional RKD estimands, such as those for means and quantiles, identify a weighted average of structural derivatives, with weights that vary across (sub)populations. However, the causal interpretation of these estimands within the potential outcomes framework (rubin2005causal,imbens2015causal) remains unclear. Furthermore, the identification and inference of more general causal effects in this setting remain relatively underexplored.

This paper develops a unified framework for the identification and estimation of local treatment effects in sharp regression kink designs. Unlike the prior literature, which primarily seeks causal interpretations for specific RKD estimands, our approach begins by defining a broad class of causal parameters of interest and then derives the corresponding strategies for their identification and estimation. In a sharp RKD, the treatment \(B\) is a deterministic function of the running variable \(X\), given by \(B=b(X)\). The general causal effect at the treatment level \(b_0:=b(x_0)\) is captured by the following parameter:

align*[align* omitted — 191 chars of source]

where \(\phi\) is a functional on the space of one-dimensional distribution functions, and \(Y(b)\) denotes the potential outcome under treatment level \(b\in\mathbb{R}\). Typically, $x_0$ represents the kink point in this design, defined as a threshold where the derivative of $b(x)$ is discontinuous: $\lim_{x\to x_0^+}\frac{d}{dx}b(x) \ne \lim_{x\to x_0^-}\frac{d}{dx}b(x)$. The parameter \(\Delta_{\phi}\) measures the effect of an infinitesimal change in the treatment variable on a feature of the conditional potential outcome distribution $F_{Y(b)|X=x_0}$, evaluated at $b=b_0$. For instance, when \(\phi\) is the mean functional, i.e., \(\phi(F) = \mu(F):=\oldint\nolimits w\,dF(w)\), \(\Delta_{\phi}\) corresponds to the “local average response” parameter identified in card2015inference.

We refer to the parameter \(\Delta_\phi\) as the local treatment effect (LTE). This terminology is analogous to that of florens2008identification, who study the identification of the average treatment effect (ATE) in models with a continuous endogenous treatment. In their framework, the ATE at treatment level \(b \in \mathbb{R}\) is defined as

align*[align* omitted — 153 chars of source]

In standard settings, the unconditional distribution of a potential outcome, \(F_{Y(b)}\), can often be identified for any \(b\in\mathbb{R}\) using methods like instrumental variables or control functions (florens2008identification,imbens2009identification). This distribution then serves as a foundation for identifying causal parameters such as \(\Delta^\mathrm{ATE}(b)\). However, such approaches are not suitable for the sharp RKD, where the treatment is a deterministic function of the running variable, precluding the existence of a valid external instrument. Instead, identification in a sharp RKD relies on the change in the slope of the treatment assignment function \(b(x)\) at the kink $x_0$. While the deterministic nature of treatment assignment is a known restriction, a challenge in identifying the derivative effect $\Delta_\phi$ arises because the conditional distribution of the potential outcome $F_{Y(b_0+\delta)|X=x_0}$ remains unidentified under the kink assumption alone. Consequently, functionals of this conditional distribution, such as its mean or quantiles, cannot be directly recovered.

This paper proposes an identification strategy for the general local treatment effect \(\Delta_{\phi}\) in the sharp RKD, applicable to any Hadamard differentiable functional \(\phi\). The strategy utilizes standard assumptions from the RKD literature, primarily smoothness conditions on relevant structural functions and on the conditional distributions of the unobserved disturbance $\varepsilon$ and the outcome $Y$ given the running variable $X$ (card2015inference,chiang2019causal,qu2019uniform). As a result, we derive a unified identification formula for $\Delta_{\phi}$ and formally establish the relationship between this general class of LTEs and their corresponding conventional RKD estimands.

This paper develops a unified estimation and inference framework for two new classes of RKD estimands, which correspond to the LTEs for various smooth functionals. Motivated by the local polynomial constrained quantile smoother of chiang2019causal, we employ a local polynomial constrained regression RKD estimator and establish its asymptotic properties. This approach has two notable features. First, it explicitly imposes the continuity of the conditional distribution $F_{Y|X=x}$ at the kink point $x_0$ as a constraint. Second, unlike methods that require separate local polynomial regressions on each side of the threshold (calonico2014robust), this estimator computes the required left- and right-derivatives in a single step, which can improve computational efficiency. Moreover, we develop uniform inference procedures for the proposed distributional RKD estimators using a multiplier bootstrap method. This inference framework applies to various LTEs. By integrating pivotal methods developed for quantiles by chiang2019causal, our framework covers effects on the mean, the distribution, quantiles, and inequality measures such as the Lorenz curve. To complete the methodology, we also provide a bandwidth selector based on mean squared error (MSE) minimization.

To illustrate the applicability of our framework, we analyze the effect of unemployment insurance (UI) benefits on the distribution of unemployment durations using the Continuous Wage and Benefit History (CWBH) data. This analysis extends previous empirical studies, which primarily focused on mean or quantile effects. In particular, we estimate the local distributional treatment effect (LDTE) to assess the marginal impact on the probability distribution, and the local Lorenz treatment effect (LLTE) to examine the effects on the dispersion and inequality of unemployment durations.

The remainder of this paper is organized as follows. The next subsection outlines our contributions to the literature. Section (ref) introduces the model and our parameters of interest. Section (ref) establishes our main identification results. Section (ref) provides further discussion and a detailed comparison with related literature. Section (ref) develops the estimation and inference framework and derives its asymptotic properties. Section (ref) presents an empirical application and evaluates the finite-sample performance of the estimators via Monte Carlo simulations. Section (ref) concludes. Appendix (ref) contains the technical assumptions and auxiliary lemmas, while Appendix (ref) details the bandwidth selection procedure. All mathematical proofs are provided in the online Supplementary Material.

Contribution to the Literature

This paper contributes to the theoretical literature on regression kink designs (card2015inference,dong2018jump,chiang2019causal,chen2019identification,chen2020quantile). The main contribution of this paper is the development of a unified framework that clarifies the causal interpretation of a broad class of sharp RKD estimands within the potential outcomes model. This framework serves two key purposes. First, it provides a unified identification formula for the general causal parameter $\Delta_\phi$. Second, by nesting existing estimands as special cases, it refines and elucidates their causal interpretations. To illustrate the latter, consider the conventional quantile regression kink design (QRKD) estimand:

align*[align* omitted — 430 chars of source]

It is known that $\mathrm{QRKD}(\tau)$ lacks a direct interpretation as a standard quantile treatment effect (QTE). chiang2019causal show that $\mathrm{QRKD}(\tau)$ identifies a weighted average of marginal effects across a subpopulation on the boundary $\partial \mathcal{V}(y_\tau,x_0):=\{e\in\mathbb{R}^{d_\varepsilon}:g(b_0,x_0,e)=y_\tau\}$, where $y_\tau:=Q_{Y|X}(\tau|x_0)$. This paper shows that, under a continuity condition on the conditional density of the outcome $Y$ given the running variable $X$ (Assumption (ref)), the QRKD estimand identifies the local quantile treatment effect (LQTE) for the subpopulation at the kink point $X=x_0$:

align*[align* omitted — 145 chars of source]

This continuity assumption, while commonly used for developing asymptotic theory in the RKD and related literature (e.g., chiang2019causal,qu2019uniform), has not previously been applied for identification purposes. The identifying conditions are discussed in greater detail in Section (ref).

Second, our estimation and inference framework, based on local polynomial constrained regression, extends the framework of chiang2019robust for local Wald estimands. Their approach cannot directly accommodate certain causal parameters, particularly those defined by smooth inequality functionals such as the Lorenz curve and the Gini coefficient, because these parameters lack a local Wald structure. In contrast, our framework addresses this limitation. We show that such LTEs can be expressed as smooth transformations of foundational distributional and quantile RKD estimands. Based on this result, they can be estimated by applying the corresponding transformations to the foundational estimators. Furthermore, the uniform Bahadur representations established for the foundational estimators, combined with the functional Delta method, provide the basis for valid uniform inference on the class of resulting LTEs.

Finally, this paper contributes to the literature on identifying continuous causal effects with endogenous treatments (florens2008identification,imbens2009identification,kasy2014instrumental,hoderlein2017corrigendum). Within this literature, identification often requires recovering the full conditional distribution of potential outcomes. However, this step typically relies on structural assumptions, such as separability (florens2008identification), the existence of control functions (imbens2009identification), or monotonicity conditions (kasy2014instrumental,hoderlein2017corrigendum). Our framework provides an alternative identification approach. We establish a structural equivalence between the LTE and the sharp RKD analogue of the local average structural derivative (LASD) proposed by hoderlein2007identification,hoderlein2009identification, as formalized in Lemma (ref). This approach is analogous to that of chernozhukov2015nonparametric, who exploit a similar connection to a panel version of the LASD to identify quantile derivative effects in nonseparable panel models. Taken together, our work and these related studies suggest that even when identification of the full distribution of potential outcomes is not readily available, causal derivative effects can often be identified by establishing a link to an appropriate version of the LASD.

Model and Target Parameters

We observe a random sample $\{(Y_i,B_i,X_i)\}_{i=1}^n$, where $Y_i$ is the outcome, $B_i$ is the treatment, and $X_i$ is the running variable. We assume the data are drawn independently and identically distributed (i.i.d.) from a population where these variables are continuous. For notational simplicity, we drop the subscript $i$ when discussing the population model. Let $\mathcal{Y} \subset\mathbb{R}$, $\mathcal{B} \subset \mathbb{R}$, and $\mathcal{X} \subset \mathbb{R}$ denote the supports of $Y$, $B$, and $X$, respectively.

assumption[Nonseparable Model] Let $g: \mathbb{R} \times \mathcal{X} \times \mathcal{E} \to \mathbb{R}$ be a measurable function. \begin{itemize} • The potential outcome $Y(b)$ for any treatment level $b \in \mathbb{R}$ is given by \begin{align*} Y(b) &= g(b,X,\varepsilon) \end{align*} where $\varepsilon \in \mathcal{E} \subset \mathbb{R}^{d_\varepsilon}$ is a vector representing unobserved individual heterogeneity. • Let the Stable Unit Treatment Value Assumption (SUTVA) hold, which implies $Y = Y(B)$. Given the deterministic treatment assignment rule $B = b(X)$, the observed outcome is thus determined by \begin{align*} Y = g(b(X),X,\varepsilon). \end{align*} \end{itemize}

Let $I_{x_0}$ be a closed interval containing the kink point $x_0$, and define $I_{x_0}^o:=I_{x_0} \backslash \{x_0\}$.

assumption[Sharp Kink Characterization in the First Stage] The treatment assignment rule $b(\cdot)$ is continuous on $I_{x_0}$ and continuously differentiable on $I_{x_0}^o$. This implies continuity at the kink point, $b(x_0^+) = b(x_0^-) = b(x_0)$, but a discontinuity in the first-order derivative, $b'(x_0^+) \ne b'(x_0^-)$, where $b'(x_0^\pm):=\lim_{x\to x_0^\pm} \frac{d}{dx}b(x)$.

Assumptions (ref) and (ref) define the nonparametric, nonseparable structural model that forms the basis of our analysis. Assumption (ref)(i) specifies the general nonseparable form of the potential outcome. Assumption (ref)(ii) imposes a standard condition from the causal inference literature: the observed outcome $Y$ is determined by the treatment $B$, which itself is generated by a deterministic treatment selection rule. This setup is consistent with the canonical RKD framework (card2015inference,chiang2019causal). Assumption (ref) formally imposes the sharp kink on the treatment selection function, $b(\cdot)$. This requires that $b(x)$ is continuous at the kink point $x_0$, while its first derivative, $b'(x)$, exhibits a jump discontinuity at that point.

Within this model, we focus on the following causal parameter.

definition[Local Treatment Effect] Let $\mathcal{F}$ denote the space of all one-dimensional distribution functions. For any functional $\phi: \mathcal{F} \to \mathbb{R}$, the local treatment effect at the kink (LTE) is defined as \begin{align*} \Delta_\phi := \frac{\partial}{\partial b}\phi\big(F_{Y(b)|X=x_0}\big)\big|_{b=b_0} :=\lim_{\delta\to 0}\frac{\phi(F_{Y(b_0+\delta)|X=x_0})-\phi(F_{Y(b_0)|X=x_0})}{\delta} \end{align*} provided the limit exists, where \(b_0 :=b(x_0)\).

The LTE parameter $\Delta_{\phi}$ has a clear causal interpretation. Under the structural model defined by Assumptions (ref)--(ref), its definition is equivalent to $$\lim_{\delta \to 0}\frac{\phi(F_{g(b_0+\delta,x_0,\varepsilon)|X=x_0})-\phi(F_{g(b_0,x_0,\varepsilon)|X=x_0})}{\delta}.$$ This derivative captures the marginal effect of the treatment $B$ at its baseline level $b_0$ on the feature of the potential outcome distribution characterized by the functional $\phi$. The fundamental identification problem, however, is that the key component of this definition—the counterfactual distribution $F_{Y(b_0+\delta)|X=x_0}$—is unobserved, since the structural function $g$ is unknown.

We conclude this section with several key examples of the LTE $\Delta_\phi$.

example[Average Effect] When $\phi$ is the mean functional, $\phi(F) = \mu(F):=\oldint\nolimits w\,dF(w)$, $\Delta_\phi$ becomes the local average treatment effect (LATE): \begin{align*} \Delta_\mu :=\frac{\partial}{\partial b}\mu\big(F_{Y(b)|X=x_0}\big)\big|_{b=b_0} :=\lim_{\delta\to 0}\frac{E[Y(b_0+\delta)|X=x_0] - E[Y(b_0)|X=x_0]}{\delta}. \end{align*}
example[Distributional and Quantile Effects] When $\phi$ is the identity mapping at a point $y$, which we denote by $\phi(F) = Id_y(F) := F(y)$, $\Delta_\phi$ becomes the local distributional treatment effect (LDTE). This measures the marginal effect on the cumulative distribution function: \begin{align*} \Delta_{Id}(y):= \frac{\partial}{\partial b} Id_y\big(F_{Y(b)|X=x_0}\big)\big|_{b=b_0}:= \lim_{\delta \to 0}\frac{F_{Y(b_0+\delta)|X}(y|x_0) - F_{Y(b_0)|X}(y|x_0)}{\delta}. \end{align*} When $\phi$ is the $\tau$th quantile functional, $\phi(F)= Q_\tau(F):=\inf\{w:F(w)\geq \tau\}$ for $\tau \in (0,1)$, $\Delta_\phi$ becomes the local quantile treatment effect (LQTE): \begin{align*} \Delta_{Q}(\tau) := \frac{\partial}{\partial b} Q_\tau\big(F_{Y(b)|X=x_0}\big)\big|_{b=b_0}:= \lim_{\delta \to 0} \frac{Q_{Y(b_0+\delta)|X}(\tau|x_0) - Q_{Y(b_0)|X}(\tau|x_0)}{\delta}, \end{align*} where $Q_{Y(b)|X}(\tau|x)$ is a shorthand for $Q_\tau(F_{Y(b)|X=x})$.
example[Lorenz Effect] Our framework also covers inequality measures. For instance, if $\phi$ corresponds to the Lorenz curve, $L_\tau(F):=\oldint\nolimits_0^\tau Q_u(F)\,du/\mu(F)$, then $\Delta_\phi$ is the local Lorenz treatment effect (LLTE): \begin{align*} \Delta_{L}(\tau) := \frac{\partial}{\partial b}L_\tau\big(F_{Y(b)|X=x_0}\big)\big|_{b=b_0} := \lim_{\delta \to 0}\frac{L_\tau\big(F_{Y(b_0+\delta)|X=x_0}\big) - L_\tau\big(F_{Y(b_0)|X=x_0}\big)}{\delta}. \end{align*}

Identification

The definition of the LTE $\Delta_\phi$ as a derivative makes the class of Hadamard differentiable functionals a natural choice for our analysis. This class is particularly relevant as it includes most statistical measures used in policy evaluation, such as means, quantiles, and various inequality indices. To formalize this concept, let $\ell^\infty(\mathbb{T})$ denote the space of all bounded functions defined on a set $\mathbb{T}$, and let $\mathbb{B}$ be a Banach space. A functional $\phi: \mathcal{F}\subset \ell^\infty(\mathbb{R}) \to \mathbb{B}$ is called Hadamard differentiable at $F \in \mathcal{F}$ tangentially to a set $\mathbb{D}_0\subseteq \ell^\infty(\mathbb{R})$, if there exists a continuous linear map $\phi'_F:\ell^\infty(\mathbb{R})\to \mathbb{B}$ such that

align*[align* omitted — 132 chars of source]

holds for all sequences $\{h_\delta\}$ satisfying $h_\delta \to h\in\mathbb{D}_0$ and $F + \delta h_\delta \in \mathcal{F}$ for all sufficiently small $\delta$. The map $\phi'_F$ is the Hadamard derivative of $\phi$ at $F$ tangentially to $\mathbb{D}_0\subseteq\ell^\infty(\mathbb{R})$. Building on this definition, we impose the following assumption on the functional $\phi$.

assumptionS[Smooth Functional] The functional \(\phi:\mathcal{F}\to \mathbb{R}\) is Hadamard differentiable at $F_{Y|X=x_0}$, with its Hadamard derivative denoted by $\phi'_{F_{Y|X=x_0}}$.

The smooth structural model is specified by the following assumption.

assumptionS[Smooth Structural Functions] The structural function $g(\cdot,\cdot,e)$ is continuously differentiable for all $e\in\mathcal{E}$, with its partial derivatives with respect to the first and second arguments denoted by $g_1(\cdot,\cdot,e)$ and $g_2(\cdot,\cdot,e)$, respectively.

Assumption (ref), which aligns with Assumptions 1(ii) and 2 in card2015inference, imposes smoothness on the structural function, requiring that the outcome varies continuously with both the treatment and the running variable. Notably, this assumption is weaker than Assumption 2(i) of chiang2019causal, as we do not impose continuous differentiability with respect to the unobserved heterogeneity $\varepsilon$. Under this smoothness condition, the partial derivative of the function $h(x,e):=g(b(x),x,e)$ with respect to $x$ is given by

align*[align* omitted — 88 chars of source]

A key implication follows: although the structural partial derivatives, $x\mapsto g_1(b(x),x,e)$ and $x\mapsto g_2(b(x),x,e)$, are continuous at $x_0$, the function $x\mapsto h(x,e)$ is not continuously differentiable at this point because $b'(x)$ has a jump discontinuity.

Our identification strategy requires the following smoothness conditions on the conditional distributions given the running variable.

assumptionS[Smooth Disturbance Distributions] The conditional distribution of \(\varepsilon\) given \(X\) is absolutely continuous with respect to Lebesgue measure. Its conditional density, \(f_{\varepsilon|X}(e|\cdot)\), is continuously differentiable on the interval \(I_{x_0}\) for all \(e\in \mathcal{E}\).
assumptionS[Smooth Outcome Distribution] The conditional distribution of \(Y\) given \(X\) is absolutely continuous with respect to Lebesgue measure. Furthermore, for each \(\tau \in (0,1) \), the conditional density \(f_{Y|X}(Q_{Y|X}(\tau|\cdot)|\cdot)\) is continuous and strictly positive on $I_{x_0}$.

Assumption (ref) is used to address selection bias in the identification strategy and is analogous to Assumption 2(iii) in chiang2019causal. A set of sufficient conditions for Assumption (ref) is provided in Assumption (ref) (Section (ref)), which follows the formulation used by card2015inference.

Assumption (ref) facilitates the derivation of the unified identification formula. A contribution of this paper relates to the use of this assumption for identification purposes. In the existing regression discontinuity design (RDD) and RKD literature, the continuity of the conditional density $f_{Y|X}$ has been used primarily for inference---specifically, to derive the asymptotic properties of estimators (e.g., qu2019uniform,chiang2019causal). In contrast, our framework is the first to exploit this condition to achieve point identification. Finally, it is important to note that these two assumptions are not redundant, as the continuity of $f_{\varepsilon|X}$ (Assumption (ref)) does not, in general, imply the continuity of $f_{Y|X}$.

The following lemma forms the cornerstone of our identification strategy, establishing a key link between the LTE and the causal parameter defined by the structural model. All regularity conditions, prefixed with `R', are collected in Appendix (ref).

lemma[Structural Representation] Suppose Assumptions (ref)--(ref), (ref)--(ref), (ref), and (ref)(i)--(ii) hold. Then, \begin{align*} \Delta_{\phi} = \phi'_{F_{Y|X=x_0}}\left(E\left[-f_{Y|X}(\cdot|x_0)\,g_1(b_0,x_0,\varepsilon) \middle| Y=\cdot, X=x_0\right] \right). \end{align*}

Lemma (ref) establishes that the LTE $\Delta_\phi$ can be expressed as the Hadamard derivative of $\phi$ applied to a conditional expectation. The term inside this expectation is the structural derivative of interest, $g_1$, weighted by the negative conditional density of the outcome, $-f_{Y|X}$. This representation provides the link to identification. While the LTE is defined using unobserved potential outcomes, the lemma expresses it in terms of two key components: (i) an estimable weight function, and (ii) the conditional expectation of the structural derivative, $E[g_1(b_0,x_0,\varepsilon)|Y=\cdot,X=x_0]$. The second component, which is the object of identification, is analogous to the local average structural derivative (LASD) proposed by hoderlein2007identification,hoderlein2009identification.

Building on Lemma (ref), the following theorem establishes our main identification result for the LTE.

theorem[Identification of the LTE] Under Assumptions (ref)--(ref), (ref)--(ref), and (ref), the local treatment effect is identified as: \begin{align*} \Delta_{\phi} = \phi'_{F_{Y|X=x_0}}\left(\mathrm{DRKD}(\cdot)\right), \end{align*} where $\mathrm{DRKD}(\cdot)$ is the distributional RKD (DRKD) estimand, defined by: \begin{align*} \mathrm{DRKD}(y):= \frac{\frac{\partial}{\partial x} F_{Y|X}(y|x_0^+) - \frac{\partial}{\partial x} F_{Y|X}(y|x_0^-)}{b'(x_0^+) - b'(x_0^-)}. \end{align*}

Theorem (ref) shows that the LTE, $\Delta_{\phi}$, is identified by applying the Hadamard derivative of the functional $\phi$ in the direction of the DRKD estimand. This DRKD estimand is a local Wald-type ratio constructed from the derivatives of the observable conditional distribution on either side of the kink. Furthermore, this result provides a causal interpretation for a specific case of the generalized local Wald ratio estimand proposed by chiang2019robust, defined as

align*[align* omitted — 542 chars of source]

To establish the connection, setting the tuning parameters to $v = 1$, $(\varUpsilon, \varphi ,\psi)=(\phi'_{F_{Y|X=x_0}}, Id, Id)$, $h_1(Y,\cdot) = \mathbbm{1}(Y\leq \cdot)$, and $h_2(B,\cdot)=B=b(X)$ yields the following form for their generalized estimand:

align*[align* omitted — 250 chars of source]

Therefore, under the conditions of our theorem, the estimand $\tau^{(1)}(\theta''; \phi'_{F_{Y|X=x_0}}, Id, Id)$ admits a causal interpretation as the local treatment effect, $\Delta_{\phi}$.

We conclude this section by applying Theorem (ref) to the functionals introduced earlier. For detailed derivations, we refer interested readers to Appendix S.1 in the supplementary material. \setcounter{example}{0}

example[Average Effect, Continued] The mean functional, $\mu(F) = \oldint\nolimits w\,dF(w)$, is linear. Consequently, its Hadamard derivative is the functional itself: $\mu'_F(h) = \oldint\nolimits w\,dh(w)$. Substituting this into the general formula from Theorem (ref) and applying integration by parts yields the familiar RKD estimand for the mean effect: \begin{align} \Delta_\mu = \frac{\frac{d}{dx}E[Y|X=x_0^+] - \frac{d}{dx}E[Y|X=x_0^-]}{b'(x_0^+) - b'(x_0^-)} =:\mathrm{MRKD}. \end{align} This result is consistent with the causal interpretation for the MRKD estimand in card2015inference, which we compare in detail in Section (ref).
example[Distributional and Quantile Effects, Continued] For the identity mapping \(\phi = Id_y\), Theorem (ref) directly yields the identification of the local distributional treatment effect (LDTE): \begin{align*} \Delta_{Id}(y) = \frac{\frac{\partial}{\partial x} F_{Y|X}(y|x_0^+) - \frac{\partial}{\partial x} F_{Y|X}(y|x_0^-)}{b'(x_0^+) - b'(x_0^-)} =:\mathrm{DRKD}(y) \end{align*} for all \( y\in \mathcal{Y}_{x_0}\), where $\mathcal{Y}_{x}$ denotes the support of the conditional distribution $F_{Y|X=x}$. This result provides a direct causal interpretation for the DRKD estimand: it identifies the marginal effect of the treatment on the potential outcome's conditional distribution $F_{Y(b)|X=x_0}$ at the baseline treatment level $b_0$. For the quantile functional $\phi = Q_\tau$, Hadamard differentiability is guaranteed under certain regularity conditions (see, e.g., Lemma 21.4 of van2000asymptotic). Applying Theorem (ref) then yields the identification of the local quantile treatment effect (LQTE): \begin{align} \Delta_Q(\tau) =\frac{\frac{\partial}{\partial x} Q_{Y|X}(\tau|x_0^+) - \frac{\partial}{\partial x} Q_{Y|X}(\tau|x_0^-)}{b'(x_0^+) - b'(x_0^-)} =:\mathrm{QRKD}(\tau) \notag \end{align} for all \(\tau \in (0,1)\). This result is a key contribution of our paper. It establishes that $\mathrm{QRKD}(\tau)$ has a direct and intuitive causal interpretation as the marginal effect on the potential outcome quantile. This contrasts with the complex weighted average interpretation in chiang2019causal, which we discuss in detail in Section (ref).
example[Lorenz Effect, Continued] Our framework extends naturally to inequality measures. Consider the Lorenz curve functional, $L_\tau(F):=\oldint\nolimits_0^\tau Q_u(F)\,du/\mu(F)$. Under standard regularity conditions (e.g., Proposition 2 in bhattacharya2007inference), its Hadamard derivative is given by: \begin{align*} [L_\tau]'_{F}(h) & = \frac{\oldint\nolimits_{0}^{\tau}[Q_u]'_F(h)\,du}{\mu(F)} - \frac{L_\tau(F)}{\mu(F)}\cdot\mu(h). \end{align*} Substituting this derivative into the general identification result from Theorem (ref) yields the local Lorenz treatment effect (LLTE): \begin{align} \Delta_{L}(\tau) &= \frac{1}{\mu_0} \left(\oldint\nolimits_{0}^{\tau} \mathrm{QRKD}(u)\,du - L_{Y|X}(\tau|x_0) \cdot \mathrm{MRKD} \right) \end{align} for all $\tau \in (0,1)$, where $\mu_0:=E[Y|X=x_0]$ and $L_{Y|X}(\tau|x_0):=L_{\tau}(F_{Y|X=x_0})$ is the baseline conditional Lorenz curve. This outcome illustrates the structure of the framework: the treatment effect on the Lorenz curve (LLTE) is identified as a function involving the effects on the mean (MRKD) and quantiles (the integrated QRKD).

Comparison and Further Discussion

This section presents a detailed comparison between our main identification results and the existing literature on RKD designs. We focus in particular on the foundational contributions of card2015inference and chiang2019causal, and conclude by discussing several extensions of our identification framework.

Average Effect

card2015inference propose a set of sufficient conditions for the causal interpretation of the mean RKD estimand. To facilitate a comparison with our framework, we now summarize their key identifying assumptions—specifically, those required in addition to our shared Assumptions (ref)--(ref) and (ref)—along with their main identification result. \setcounter{assumptionS}{3}

assumptionSPrime\begin{itemize} • The support \(\mathcal{E}\subset I_\varepsilon\) is bounded for some large compact set \(I_\varepsilon\subset \mathbb{R}^{d_\varepsilon}\). • The set \(\mathcal{A}_\varepsilon:=\{e\in\mathbb{R}^{d_\varepsilon}:f_{X|\varepsilon}(x|e) > 0,\, \forall x \in I_{x_0}\}\) has a positive measure under $\varepsilon$: \(\oldint\nolimits_{\mathcal{A}_\varepsilon}\,dF_{\varepsilon}(e) > 0 \). • The conditional density \(f_{X|\varepsilon}(\cdot|e)\) is continuously differentiable on \(I_{x_0}\) for all \(e \in I_\varepsilon\). \end{itemize}
lemma[card2015inference, Proposition 1] Suppose Assumptions (ref)--(ref), (ref), and (ref) hold. Then, \begin{align*} \mathrm{MRKD} = E\left[g_1(b_0,x_0,\varepsilon)\middle|X=x_0\right] = \oldint\nolimits \omega_{x_0}(e) \cdot g_1(b_0,x_0,\varepsilon) \,dF_\varepsilon(e), \end{align*} where \(\omega_x(e):=f_{X|\varepsilon}(x|e)/f_X(x)\).

We first note that our main conclusion is consistent with card2015inference. Under standard regularity conditions that permit interchanging limits and expectations, our LTE for the mean simplifies to:

align*[align* omitted — 82 chars of source]

This is precisely the parameter identified by card2015inference, confirming that both frameworks target the same causal object for the mean effect.

Despite this consistency, our identifying assumptions differ in two key respects. First, embedding the mean case within our unified framework requires Assumptions (ref) and (ref). While not strictly necessary for identifying the mean effect in isolation, these assumptions are needed for the generality of a framework covering all Hadamard differentiable functionals. Second, our framework utilizes Assumption (ref)---which directly requires the continuous differentiability of the conditional density $x \mapsto f_{\varepsilon|X}(e|x)$---to establish the causal interpretation of $\mathrm{MRKD}$ as the LATE, $\Delta_\mu$. While Assumption (ref) provides sufficient conditions for (ref) (as shown in their Lemma 2), it is specifically employed by card2015inference to enable the population-weighted average interpretation by addressing the measure-zero conditioning problem. In particular, (ref) allows one to rigorously handle conditioning on the measure-zero event $\{X=x_0\}$ and to express $\Delta_\mu$ as a weighted average of the structural derivative over the entire distribution of unobserved heterogeneity. The resulting weight function, $\omega_{x_0}(e)$, is thus well defined and reflects the relative likelihood that an individual with type $\varepsilon=e$ is located at the kink point (card2015inference, Remark 1).

Quantile Effect

We now provide a detailed comparison between our framework and that of chiang2019causal. The two approaches differ significantly in their starting point. Our analysis addresses a “forward problem”: we first define a causal effect of interest ($\Delta_Q$) and then derive an identification strategy for it. In contrast, chiang2019causal address an “inverse problem”: they start with the conventional estimand $\mathrm{QRKD}(\tau)$, and investigate what causal interpretation, if any, it admits. To facilitate this comparison, we first summarize their main assumptions and results

The identification strategy of chiang2019causal relies on geometric concepts. Let $h(x,e):=g(b(x),x,e)$ be the reduced-form outcome function. They define a volume in the space of unobservables as $\mathcal{V}(y,x):=\{e\in\mathbb{R}^{d_\varepsilon}:h(x,e) \leq y\}$, with its boundary denoted by $\partial \mathcal{V}(y,x):=\{e\in\mathbb{R}^{d_\varepsilon}:h(x,e)=y\}$. Using the Hausdorff probability measure developed in sasaki2015quantile, they derive a reduced-form expression for the key object of their analysis: the conditional quantile partial derivative, $\frac{\partial}{\partial x}Q_{Y|X}(\tau|x)$. The Hausdorff probability measure on the boundary $\partial\mathcal{V}(y,x)$ is defined differently depending on the dimension $M$ of the unobserved heterogeneity. For \(M > 1\), the measure is defined as the ratio of weighted surface areas:

align*[align* omitted — 237 chars of source]

for any set \(S\) in the collection of Borel subsets of \(\partial\mathcal{V}(y,x)\). For \(M = 1 \), the boundary $\partial\mathcal{V}(y,x)$ is typically a discrete set of points. The Hausdorff measure $H^0$ becomes the counting measure, yielding the discrete probability measure:

align*[align* omitted — 182 chars of source]

for each point $e\in \partial\mathcal{V}(y,x)$.

\setcounter{assumptionS}{4}

assumptionSPrime[chiang2019causal, Assumptions 2--3] \begin{itemize} • The boundary \(\partial\mathcal{V}(y,x)\) is a smooth manifold that can be parameterized by a mapping \(\varPi_{y,x}: \varSigma \to \partial\mathcal{V}(y,x)\) satisfying Assumption 3 of chiang2019causal, where \(\varSigma\) is an \((M-1)\)-dimensional rectangle. • \(e\mapsto h(x,e)\) is continuously differentiable for all \(x\in I_{x_0}\), and \(\|\nabla_e h(x,e)\| \ne 0\) on \(\partial \mathcal{V}(y,x)\) for all \((y,x)\in \mathcal{Y}_x \times I_{x_0}\)\(\oldint\nolimits_{\partial \mathcal{V}(y,x)} f_{\varepsilon|X}(e|x)\,dH^{M-1}(e) > 0\) for all \((y,x)\in \mathcal{Y}_x \times I_{x_0}\). \end{itemize}

We now relate Assumption (ref) to the conditions in chiang2019causal. Assumption (ref)(i) is identical to their Assumption 3, while Assumptions (ref)(ii) and (iii) correspond to parts of their Assumption 2. The remaining conditions in their Assumption 2 are already encompassed by our earlier Assumptions (ref) and (ref). It is worth noting that (ref)(ii) is stronger than Assumption (ref) in one respect, as it implies the existence and continuous differentiability of the function $e\mapsto g(b(x),x,e)$---a condition not required under (ref).

Under these conditions, chiang2019causal exploit the sharp kink to derive a reduced-form expression for the QRKD estimand that is free from selection bias. This result is stated in the following lemma.

lemma[chiang2019causal, Theorem 1 and Corollary 1] Suppose Assumptions (ref)--(ref), (ref)--(ref), and (ref) hold, and that the probability measure \(P^{M-1}_{y,x}\) on \(\partial \mathcal{V}(y,x)\) is well-defined for \((y,x)\in \mathcal{Y}_x \times I_{x_0}\).\footnote{We omit the regularity assumptions that ensure \(P^{M-1}_{y,x}\) is a well-defined probability measure and that justify the use of the dominated convergence theorem. Interested readers may refer to chiang2019causal and sasaki2015quantile for details.} Then, \begin{align*} \mathrm{QRKD}(\tau) = E_{P^{M-1}_{y_\tau,x_0}}\left[g_1(b(x_0),x_0,\varepsilon)\right] = \oldint\nolimits_{\partial \mathcal{V}(y_\tau,x_0)} g_1(b(x_0),x_0,e)\,dP^{M-1}_{y_\tau,x_0}(e), \end{align*} where \(y_\tau:=Q_{Y|X}(\tau|x_0)\). If \(\varepsilon\) is a random scalar and \(\partial \mathcal{V}(y_\tau,x_0)\) a singleton, then \begin{align*} \mathrm{QRKD}(\tau) = g_1\big(b(x_0),x_0,\varepsilon(y_\tau,x_0)\big), \end{align*} where \(\varepsilon(y_\tau,x_0)\) denotes the sole element of \(\partial \mathcal{V}(y_\tau,x_0)\).

The causal interpretation provided by chiang2019causal differs significantly from ours. Their result shows that the QRKD estimand identifies a complex weighted average of the structural derivative, $E_{P^{M-1}_{y_\tau,x_0}}[g_1(b_0,x_0,\varepsilon)]$, where the weighting depends on the geometry of the structural function. In contrast, our Theorem (ref) establishes that the same estimand identifies the local quantile treatment effect, $\Delta_Q(\tau)$, a direct marginal effect of the treatment on the potential outcome's conditional quantile $Q_{Y(b)|X}(\tau|x_0)$ at the baseline level $b_0$. The divergence in results stems from a difference in the underlying identification assumptions. Interpreting $\mathrm{QRKD}(\tau)$ as the local quantile treatment effect $\Delta_Q(\tau)$ requires Assumption (ref). This assumption is less restrictive than Assumption (ref), which is used in approaches addressing the measure-zero conditioning problem. Notably, Assumption (ref) does not require certain conditions employed in those contexts, such as smoothness of the function $e \mapsto h(x,e)$ or conditions ensuring the Hausdorff probability measure is well-defined. The ability to obtain a direct causal interpretation under these less restrictive conditions is an advantage of our framework.

To clarify the relationship between Assumption (ref) and more primitive conditions such as Assumption (ref), we provide two sets of sufficient conditions under which (ref) holds. These conditions involve smoothness restrictions on the conditional density $f_{\varepsilon|X}$ and the structural function $g$.

assumptionSDefine \[\tilde{f}(y,x):=\oldint\nolimits_{\partial\mathcal{V}(y,x)} \frac{1}{\|\nabla_eh(x,e)\|}f_{\varepsilon|X}(e|x)\,dH^{M-1}(e). \] Suppose \(\tilde{f}(y_\tau(\cdot),\cdot)\) is continuous on \(I_{x_0}\) for all \(\tau \in (0,1)\).
assumptionS\begin{itemize} • The unobserved heterogeneity $\varepsilon$ can be decomposed into two components, $\varepsilon=(\varepsilon_1,A)$, satisfying the following conditions: (a) The scalar component $\varepsilon_1$ is absolutely continuous given $(A,X)$. Its conditional density, $f_{\varepsilon_1|A,X}(e_1|a,x)$ is continuous in both $e_1$ and $x \in I_{x_0}$, and is strictly positive for all $a$. (b) The vector component $A$ is absolutely continuous given $X$. Its conditional density $f_{A|X}(a|x)$ is continuous in $x \in I_{x_0}$, and is strictly positive for all $a$. • The function $h(x,e_1,a):=g(b(x),x,e_1,a)$ is continuously differentiable with respect to its second argument $e_1$ for all $a$ and $x \in I_{x_0}$. Its partial derivative $\partial_{e_1}h$ satisfies $\partial_{e_1}h(x,e_1,a) \geq c > 0$ for some constant $c>0$. Furthermore, $x\mapsto \partial_{e_1}h(x,e_1,a)$ is continuous for all $e_1$ and $a$. \end{itemize}
propositionSuppose Assumptions (ref)--(ref), and (ref) hold. Then, Assumption (ref) is satisfied if either of the following two conditions holds: (a) Assumption (ref) and (ref) hold. (b) Assumption (ref) holds.

Proposition (ref) provides two distinct sets of primitive conditions under which Assumption (ref) holds. The first set, Condition (a), follows the approach of chiang2019causal. It requires the structural function to be smooth with respect to the entire vector of unobserved heterogeneity but does not impose monotonicity. The second set, Condition (b), utilizes an alternative set of assumptions similar to those in chernozhukov2015nonparametric and hoderlein2017corrigendum. It requires the existence of at least one continuously distributed scalar component within the unobserved heterogeneity vector, and that the structural function is smooth and monotonic with respect to this single scalar component, though not necessarily with respect to others.

An Alternative Weighted Interpretation

Our structural representation in Lemma (ref) establishes a connection between the LQTE and a sharp RKD analogue of the local average structural derivative (LASD) from hoderlein2007identification,hoderlein2009identification. Applying Lemma (ref) with $\phi = Q_\tau$ yields:

align*[align* omitted — 95 chars of source]

While this provides a direct interpretation as a structural marginal effect for a specific subpopulation $\{Y=y_\tau, X=x_0\}$, this interpretation involves conditioning on an event of Lebesgue measure zero. Handling such conditioning rigorously presents known theoretical difficulties (sasaki2015quantile).

Two approaches can address the issue of conditioning on a measure-zero event. The first approach, following chiang2019causal, involves replacing our Assumption (ref) with Assumption (ref). This leads to their weighted-average interpretation on the boundary $\partial\mathcal{V}(y_\tau,x_0)$, as given by Lemma (ref). However, integrating this approach within our unified framework would necessitate imposing the additional Assumption (ref) to ensure that Assumption (ref) still holds. We propose an alternative approach, showing that the measure-zero issue can be addressed by replacing Assumption (ref) with the following condition instead.

assumptionS\begin{itemize} • The set $\mathcal{A}_\varepsilon(\tau):=\{e\in\mathbb{R}^{d_\varepsilon}:f_{Y,X|\varepsilon}(y_\tau(x),x|e) > 0,\, \forall x\in I_{x_0}\}$ has a positive measure under $F_\varepsilon$ for all $\tau \in (0,1)$: $\oldint\nolimits_{\mathcal{A}_\varepsilon(\tau)} \,dF_\varepsilon(e) > 0$. • The conditional density \(f_{Y,X|\varepsilon}(\cdot,\cdot|e)\) is continuous on \(\mathcal{Y}_x \times I_{x_0}\) for all \(e \in I_\varepsilon\). \end{itemize}

Assumption (ref) imposes two conditions on the conditional density of the observables $(Y,X)$ given the unobserved heterogeneity $\varepsilon$. Part (i) requires this density to be strictly positive in a neighborhood of the kink point for a non-trivial subpopulation. Part (ii) requires this density to be continuous in the same neighborhood. Notably, Assumption (ref), in conjunction with Assumption (ref), provides sufficient conditions for Assumptions (ref) and (ref), which enable the derivation of the unified identification formula. The following proposition provides an alternative weighted causal interpretation for the QRKD estimand.

propositionSuppose Assumptions (ref)--(ref), (ref), (ref), (ref) and (ref) hold. Then, \begin{align*} \mathrm{QRKD}(\tau) = E\left[g_1(b_0,x_0,\varepsilon) \middle| Y=y_\tau,X=x_0\right] =\oldint\nolimits \omega_{y_\tau,x_0}(e) \cdot g_1(b_0,x_0,e)\,dF_\varepsilon(e) \end{align*} for all \(\tau \in (0,1)\), where the weight function is \(\omega_{y,x}(e):=f_{Y,X|\varepsilon}(y,x|e)/f_{Y,X}(y,x)\).

This proposition provides two results. The first equality shows that the QRKD estimand corresponds to the sharp RKD version of the local average structural derivative. The second equality addresses the issue of conditioning on the measure-zero event $\{Y=y_\tau,X=x_0\}$ by expressing the LASD as a weighted average of the structural derivative, $g_1$, over the entire population of unobserved heterogeneity. The weight function, $\omega_{y_\tau,x_0}(e)$, is well-defined under Assumptions (ref) and (ref) and represents the relative likelihood of observing an individual with type $\varepsilon=e$ having characteristics $\{Y=y_\tau, X=x_0\}$.

remark[The Role of Monotonicity] If the structural function $g(b,x,e)$ is monotonic in a scalar error term $e$, then for any given $(y,x)$, the boundary $\partial\mathcal{V}(y,x)$ consists of the unique point $h^{-1}(y,x)$. In this case, the conditional expectation simplifies to the structural derivative evaluated at this point: \begin{align*} E\left[g_1(b_0,x_0,\varepsilon) \middle| Y=y_\tau,X=x_0\right] &= E\left[g_1(b_0,x_0,\varepsilon)\middle|h(x_0,\varepsilon)=y_\tau, X=x_0\right] \\ &= g_1(b_0,x_0,h^{-1}(y_\tau,x_0)). \end{align*} This result is consistent with the special case in chiang2019causal, where the unobserved heterogeneity is one-dimensional ($M=1$) and the boundary $\partial\mathcal{V}(y_\tau,x_0)$ is a singleton.

The following table summarizes the causal interpretations of conventional RKD estimands as established in the prior literature and in this paper.

table[table omitted — 2,372 chars of source]

Estimation and Inference

Estimation Strategy

Our identification results in Section (ref) show that local treatment effects can be expressed either as direct Wald-type ratios or as smooth transformations thereof. This section develops a unified estimation and inference framework to handle both cases. To this end, we define two general classes of RKD estimands.

Type 1. (Local Wald Ratio Estimand): The first class of estimands takes the form of a local Wald ratio of derivatives:

align*[align* omitted — 407 chars of source]

Here, $m(\theta,x):=E[\varphi(Y,\theta)|X=x]$ is the conditional expectation of a known transformation $\varphi(\cdot,\theta)$ of the outcome, and $m^{(1)}(\theta,x)$ denotes its first-order derivative with respect to $x$. This foundational class subsumes many standard sharp RKD estimands. For instance, setting $\varphi(y,\theta)=\mathbbm{1}(y \leq \theta)$ yields the DRKD estimand ($\mathrm{RKD}_F = \mathrm{DRKD}$) for identifying the LDTE. Similarly, setting $\varphi(y,\theta)=y$ yields the mean RKD estimand ($\mathrm{RKD}_\mu = \mathrm{MRKD}$) for identifying the LATE.

Type 2. (Composite Estimand): The second class of estimands consists of composite functionals of the Type 1 estimands and the quantile RKD estimand:

align*[align* omitted — 150 chars of source]

Here, $\psi$ is a functional defined on a product of function spaces. This class is designed to handle LTEs that are identified via transformations of simpler effects. The local Lorenz treatment effect is a key example. By defining $\psi$ as

align[align omitted — 180 chars of source]

and choosing $g=\mathrm{MRKD}$ and $h=\mathrm{QRKD}$, the resulting estimand $\mathrm{RKD}_{\psi_L|\mu}$ identifies the LLTE, $\Delta_L$, where \(\theta' = \tau\) and \(\varTheta' = (0,1)\).

Under our assumptions, the function $x \mapsto E[\varphi(Y,\theta)|X=x]$ is continuous at the kink point $x=x_0$. This continuity holds for cases such as the conditional CDF $F_{Y|X}$. This continuity motivates a convenient estimation strategy that jointly estimates the function's level and its left- and right-hand derivatives within a single procedure. We therefore propose a $p$th-order local polynomial constrained regression estimator. For each $\theta$ in a compact set $\widebar{\varTheta}\subset \varTheta$, we solve the following minimization problem:

align[align omitted — 293 chars of source]

where \(K(\cdot)\) is a kernel function. The vector of regressors, $\bar{r}_p(u):=(1, u\delta_u^+, u\delta_u^-, \dots, u^p\delta_u^+, u^p\delta_u^-)^\top\in \mathbb{R}^{2p+1}$, is constructed to impose the continuity constraint, where \(\delta_u^+:= \mathbbm{1}(u\geq 0)\), and \(\delta_u^-:= \mathbbm{1}(u<0)\). It includes a single intercept but allows for different coefficients on the polynomial terms to the left ($u<0$) and right ($u\geq 0$) of the kink. Then, we obtain

align*[align* omitted — 245 chars of source]

The required function and its left- and right-hand derivatives can thus be obtained as $\hat{m}(\theta,x_0)=\iota_1^\top\hat{\alpha}(\theta)$, $\hat{m}^{(1)}(\theta,x_0^+) = \iota_{2}^\top\hat{\alpha}(\theta)$, and $\hat{m}^{(1)}(\theta,x_0^-) = \iota_{3}^\top\hat{\alpha}(\theta)$, where $\iota_j$ is the $j$th standard basis vector (e.g., \(\iota_2=(0,1,0,\ldots,0)^\top\)).

To estimate the quantile derivatives that form the numerator of the QRKD estimand, we follow chiang2019causal and use a local $p$th-order polynomial constrained quantile regression. In our notation, the estimator is:

align[align omitted — 272 chars of source]

where $\mathcal{T}\subset (0,1)$ is a closed interval; $\rho_\tau(u) = (\tau - \mathbbm{1}(u<0))u$ is the standard check function. The vector of regressors $\bar{r}_p(u)$ is the same as that used for the mean regression, thereby imposing the continuity of the conditional quantile function $x\mapsto Q_{Y|X}(\tau|x)$ at the kink. Let \(Q^{(\nu)}_{Y|X}(\tau|x):=\frac{\partial ^\nu}{\partial x^v}Q_{Y|X}(\tau|x)\) for an integer \(1 \leq \nu \leq p\). Analogous to the mean regression case, the estimated coefficient vector $\hat{\beta}(\tau)$ provides estimates of the conditional quantile function and its scaled derivatives at the kink:

align*[align* omitted — 282 chars of source]

The required quantile function and its left- and right-hand derivatives can thus be obtained as $\widehat{Q}_{Y|X}(\tau|x_0)=\iota_1^\top\hat{\beta}(\tau)$, $\widehat{Q}^{(1)}_{Y|X}(\tau|x_0^+) = \iota_2^\top\hat{\beta}(\tau)$, and $\widehat{Q}^{(1)}_{Y|X}(\tau|x_0^-) = \iota_3^\top\hat{\beta}(\tau)$.

remark[On the Constrained Regression Approach] Our constrained, one-step estimation approach presents several advantages over the common practice of fitting separate local polynomial regressions on each side of the kink (e.g., calonico2014robust,card2015inference). First, it is computationally efficient, particularly when estimating quantile effects over a large grid of quantiles. Second, it directly incorporates the known continuity of the underlying function (e.g., $m(\theta,x)$ or $Q_{Y|X}(\tau|x)$) at the kink point---the same continuity property used in the asymptotic theory. The concept of a constrained estimation approach has been noted in other contexts, for example by chen2020quantile for quantile treatment effects in a fuzzy RKD setting. It is important, however, to recognize the limitation of this approach. Our constrained estimators are designed specifically for regression kink designs, where the underlying conditional functions are continuous. They are not applicable to standard RDDs, as the primary identifying assumption in RDDs involves a discontinuity in these functions at the threshold (i.e., $m(\theta,x_0^+) \ne m(\theta,x_0^-)$).
remarkThe conditional quantile function is theoretically strictly increasing in $\tau$. This follows from its derivative, $\frac{\partial}{\partial \tau} Q_{Y|X}(\tau|x_0) = 1/f_{Y|X}(Q_{Y|X}(\tau|x_0)|x_0)$, which is positive under Assumption (ref). While the true quantile function is monotonic, its empirical counterpart, $\widehat{Q}_{Y|X}(\tau|x_0)$, may fail to exhibit monotonicity over the full range of $\tau$ due to sampling variation. To ensure a properly specified estimate, we apply the standard rearrangement procedure of chernozhukov2010quantile. This procedure transforms the potentially non-monotonic estimate into a monotonically increasing function that preserves the quantile properties.

We now detail the application of the general estimation framework developed above to the specific local treatment effects discussed previously. \setcounter{example}{0}

example[Average Effect, Continued] To estimate the LATE, $\Delta_\mu$, we apply our Type 1 estimator, the local polynomial constrained regression defined in ((ref)). The transformation function is the identity, $\varphi(Y,\theta) = Y$, and the parameter $\theta$ is degenerate. The conditional moment function simplifies to $m(x,\theta) = E[Y|X=x]=:\mu_{Y|X}(x)$. The estimated coefficient vector, $\hat{\alpha}$, thus provides estimates of the conditional mean and its derivatives at the kink point: \[ \hat{\alpha} = \left( \hat{\mu}_0, \frac{\hat{\mu}_{Y|X}^{(1)}(x_0^+)}{1!}, \frac{\hat{\mu}_{Y|X}^{(1)}(x_0^-)}{1!}, \dots , \frac{\hat{\mu}^{(p)}_{Y|X}(x_0^+)}{p!}, \frac{\hat{\mu}^{(p)}_{Y|X}(x_0^-)}{p!} \right)^\top . \] The required left- and right-hand derivatives can be obtained from $\hat\alpha$, yielding the estimator for the LATE: \begin{align} \widehat{\Delta}_\mu = \frac{\hat{\mu}_{Y|X}^{(1)}(x_0^+) - \hat{\mu}_{Y|X}^{(1)}(x_0^-)}{b'(x_0^+)-b'(x_0^-)} = \widehat{\mathrm{RKD}}_{\mu}.\notag \end{align}
example[Distributional and Quantile Effects, Continued] We now detail the estimation of the distributional and quantile effects. First, the LQTE, $\Delta_Q(\tau)$, is estimated directly using the local polynomial constrained quantile regression from ((ref)). As shown previously, the estimated coefficient vector $\hat{\beta}(\tau)$ yields the required left- and right-hand derivatives, which are then used to form the QRKD estimand: \begin{align} \widehat{\Delta}_{Q}(\tau) = \frac{\widehat{Q}_{Y|X}^{(1)}(\tau|x_0^+) - \widehat{Q}_{Y|X}^{(1)}(\tau|x_0^-)}{b'(x_0^+)-b'(x_0^-)} = \widehat{\mathrm{QRKD}}(\tau). \notag \end{align} Second, the LDTE, $\Delta_{Id}(y)$, is estimated using our Type 1 estimator from ((ref)), with the transformation function set to the indicator $\varphi(Y,y) = \mathbbm{1}(Y\leq y)$. Since the LDTE, $\Delta_{Id}(y)$, is identified pointwise for $y \in \mathcal{Y}_{x_0}$ by $\mathrm{RKD}_F(y)$, a common application involves evaluating this effect at specific quantiles of interest, $y_\tau = Q_{Y|X}(\tau|x_0)$. This yields the estimator for the LDTE: \begin{align} \widehat{\Delta}_{Id}(\hat{y}_\tau) = \frac{\widehat{F}_{Y|X}^{(1)}(\hat{y}_\tau|x_0^+) - \widehat{F}_{Y|X}^{(1)}(\hat{y}_\tau|x_0^-)}{b'(x_0^+)-b'(x_0^-)} = \widehat{\mathrm{RKD}}_{F}(\hat{y}_\tau). \notag \end{align} A feature of the constrained estimation approach is that the necessary components for this evaluation are obtained jointly. The local polynomial quantile regression used to estimate the LQTE also yields an estimate of the quantile level itself, $\hat{y}_\tau = \widehat{Q}_{Y|X}(\tau|x_0)$. The LDTE evaluated at the estimated $\tau$-th quantile is then estimated by substituting $\hat{y}_\tau$ into the DRKD estimator.
example[Lorenz Effect, Continued] The identification result for the LLTE, $\Delta_L(\tau)$, expresses it as a function of components related to the mean and quantile effects. Specifically, constructing an estimator for $\Delta_L(\tau)$ requires estimates of: the conditional mean ($\mu_0$), the mean RKD estimand ($\widehat{\mathrm{MRKD}}$), the conditional Lorenz curve ($L_{Y|X}(\tau|x_0)$), and the quantile RKD estimand ($\widehat{\mathrm{QRKD}}$). Each of these components is available from the estimation procedures described previously. The mean-related terms, $\hat{\mu}_0$ and $\widehat{\mathrm{MRKD}}$, are obtained from the local polynomial regression detailed in Example (ref). The quantile-related terms---namely $\widehat{\mathrm{QRKD}}(u)$ needed for the integral and $\widehat{Q}_{Y|X}(u|x_0)$ needed to construct the estimated Lorenz curve $\widehat{L}_{Y|X}(\tau|x_0)$---are obtained from the local polynomial quantile regression detailed in Example (ref). A plug-in estimator for the LLTE is therefore constructed by substituting these previously obtained estimates into the identification formula: \begin{align} \widehat{\Delta}_{L}(\tau) = \frac{1}{\hat{\mu}_0} \left(\oldint\nolimits_{0}^{\tau} \widehat{\mathrm{QRKD}}(u)\,du - \widehat{L}_{Y|X}(\tau|x_0) \cdot \widehat{\mathrm{RKD}}_{\mu} \right) = \widehat{\mathrm{RKD}}_{\psi_L |\mu}(\tau). \end{align} The integral in this expression can be computed using standard numerical methods.

Asymptotic Theory

This section establishes the asymptotic properties of our Type 1 and Type 2 estimators, $\widehat{\mathrm{RKD}}_m(\cdot)$ and $\widehat{\mathrm{RKD}}_{\psi|m}(\cdot)$. We begin by defining the necessary notation. Define the following kernel-dependent constant matrices and vectors: $\widebar{\varGamma}_p:= \oldint\nolimits_{\mathbb{R}} \bar{r}_p(u)\bar{r}_p(u)^\top K(u)\,du$, $\bar{\vartheta}_{p,q}^\pm := \oldint\nolimits_{\mathbb{R}_\pm}\bar{r}_p(u) u^{q} K(u)\,du$, and $\widebar{\varPsi}_p^\pm:=\oldint\nolimits_{\mathbb{R}_\pm} \bar{r}_p(u)\bar{r}_p(u)^\top K^2(u)\,du$. Let $\varepsilon^m(Y,X,\theta) = \varphi(Y,\theta) - m(\theta,X)$ be the error term from the conditional expectation. We define its conditional covariance function as $\sigma_{\varepsilon^m}(\theta_1,\theta_2|x):=E[\varepsilon^m(Y,X,\theta_1)\varepsilon^m(Y,X,\theta_2)|X=x]$.

assumption\begin{itemize} • (a) The data \(\{(Y_i,X_i)\}_{i=1}^n\) are i.i.d. copies of random vector \((Y,X)\) on the probability space \((\Omega,\mathcal{F},P)\). (b) The density $f_X(\cdot)$ is continuously differentiable on $I_{x_0}$ and satisfies \(0<f_X(x_0)<\infty\). • (a) \(m(\theta,x)\) is Lipschitz continuous at \(x=x_0\) and \(m^{(\nu)}(\theta,\cdot)\) is Lipschitz continuous on \( I_{x_0}\backslash\{x_0\}\) for each \(\theta\in\widebar{\varTheta}\) and \(\nu\in\{1,\ldots,p+1\}\). (b) The classes of real-valued functions \(\{y\mapsto \varphi(y,\theta):\theta \in \widebar{\varTheta}\} \) and \(\{x\mapsto m(\theta,x):\theta\in\widebar{\varTheta}\}\) are of Vapnik-Cervonenkis (VC) type with a common integrable envelope \(\chi:\mathcal{Y}\times\mathcal{X}\to\mathbb{R}\) satisfying \(\oldint\nolimits |\chi(y,x)|^{2+\eta}\,dF_{Y,X}(y,x) <\infty\) for some \(\eta >0\). (c) The conditional covariance \(\sigma_{\varepsilon^m}(\theta_1,\theta_2|\cdot) \in \mathcal{C}^1(I_{x_0}\backslash\{x_0\})\) and \(\sigma_{\varepsilon^m}(\theta_1,\theta_2|x_0^\pm) <\infty\) for each \(\theta_1, \theta_2 \in \widebar{\varTheta}\). • (a) The kernel function \(K:[-1,1]\to \mathbb{R}_+\) is bounded and continuous, and the class \(\{x\mapsto K((x-x_0)/h): h>0\}\) is of VC type. (b) \(\widebar{\varGamma}\) is positive definite. (c) \(\widebar{\varPsi}^+\), \(\widebar{\varPsi}^-\), and \(\oldint\nolimits_{\mathbb{R}_\pm} |\bar{r}_p(u) u^q K(u)|\,du\) are finite. • The bandwidth $h_{n,\theta}$ is specified as $h_{n,\theta} = \varsigma(\theta)h_n$, where $\varsigma:\varTheta \to (0,\infty)$ is a bounded, Lipschitz continuous function. The baseline bandwidth sequence $\{h_n\}$ is required to satisfy the following conditions as $n\to \infty$: (a) $h_n \to 0$ and $nh_n^3 \to \infty$; (b) $nh_n^{2p+3} \to 0$. \end{itemize}

Assumption (ref) collects the regularity conditions required to establish the uniform asymptotic expansion and weak convergence of our estimators. We briefly discuss each component. For Condition (i), Part (a) is the standard i.i.d. sampling assumption; Part (b) imposes smoothness on the density of the running variable, $f_X$. This is a standard condition used to rule out endogenous sorting around the kink point, and can be derived from the more primitive Assumption (ref). For Condition (ii), Part (a) imposes smoothness on the conditional moment function, $m(\theta,x)$. Continuity at the kink $x_0$ and differentiability away from the kink are the definitional feature of the RKD. Parts (b) and (c) are technical conditions on the complexity (i.e., VC-type) of the relevant function classes, and are analogous to Assumptions (ii)(a) and (ii)(c) in chiang2019robust. Condition (iii) imposes standard requirements on the kernel function, such as boundedness, which ensures a finite asymptotic variance. It is analogous to Assumption 1(iv) of chiang2019robust and notably rules out the use of the Gaussian kernel. Condition (iv) specifies the admissible rates for the bandwidth sequence $h_n$. Part (a) imposes a rate consistent with the RKD literature (card2015inference,chiang2019causal) and standard theory for local polynomial estimation of first-order derivatives. Part (b) is a standard undersmoothing requirement, which ensures that the bias term of the estimator is asymptotically negligible relative to the variance term.

lemma[Asymptotic Properties of \(\hat{m}^{(1)}(\cdot,x_0^\pm)\)] \begin{itemize} • (Uniform Bahadur Representation) Suppose Assumptions (ref)(i)--(iii), and (iv)(a) hold. Then, \begin{align*} &\begin{bmatrix} \sqrt{nh_{n,\theta}^3}\left(\hat{m}^{(1)}(\theta,x_0^+) - m^{(1)}(\theta,x_0^+) - h_{n,\theta}^p \frac{\iota_2^\top \widebar{\varGamma}_p^{-1} \left(m^{(p+1)}(\theta,x_0^+) \bar{\vartheta}_{p,p+1}^+ + m^{(p+1)}(\theta,x_0^-)\bar{\vartheta}_{p,p+1}^-\right)}{(p+1)!} \right) \\ \sqrt{nh_{n,\theta}^3}\left(\hat{m}^{(1)}(\theta,x_0^-) - m^{(1)}(\theta,x_0^-) - h_{n,\theta}^p \frac{\iota_3^\top \widebar{\varGamma}_p^{-1}\left(m^{(p+1)}(\theta,x_0^+) \bar{\vartheta}_{p,p+1}^+ + m^{(p+1)}(\theta,x_0^-)\bar{\vartheta}_{p,p+1}^-\right) }{(p+1)!} \right) \end{bmatrix}\\ & = \begin{bmatrix} \mathbb{Z}^m_n(\theta,2) \\ \mathbb{Z}^m_n(\theta,3) \end{bmatrix} + o_{P}(1) \end{align*} as \(n \to \infty\) uniformly in \(\theta \in \widebar{\varTheta} \), where \begin{align*} \mathbb{Z}^m_n(\theta, k) := \sum_{i=1}^{n} \frac{\iota_k \widebar{\varGamma}_p^{-1} \bar{r}_p\left(\frac{X_i-x_0}{h_{n,\theta}}\right) K\left(\frac{X_i-x_0}{h_{n,\theta} }\right)\varepsilon^m(Y_i,X_i,\theta)}{\sqrt{nh_{n,\theta}}f_X(x_0)}, \quad k \in \{2,3\}. \end{align*} • (Weak Convergence) Suppose Assumption (ref) holds. Then, \[ \begin{bmatrix} \mathbb{Z}^m_n(\cdot, 2) \\ \mathbb{Z}^m_n(\cdot, 3) \end{bmatrix} \leadsto \begin{bmatrix} \mathbb{Z}^m(\cdot, 2)\\ \mathbb{Z}^m(\cdot, 3) \end{bmatrix}, \] where \(\mathbb{Z}^m: \varOmega \to \ell^{\infty}(\widebar{\varTheta} \times \{2,3\}) \) is a tight zero-mean Gaussian process with covariance function \begin{align*} E\left[\mathbb{Z}^m(\theta_1, k_1) \mathbb{Z}^m(\theta_2, k_2)\right] = \frac{\iota_{k_1}^\top\widebar{\varGamma}_p^{-1}\widebar{\varXi}^m_p(\theta_1,\theta_2) \widebar{\varGamma}_p^{-1}\iota_{k_2}}{ f_X(x_0)} \end{align*} for all \(\theta_1, \theta_2 \in \widebar{\varTheta}\) and \(k_1, k_2 \in \{2,3\}\). The matrix \(\widebar{\varXi}^m_p\) in the covariance is given by \(\widebar{\varXi}^m_p(\theta_1,\theta_2) :=\widebar{\varPsi}_p^+(\theta_1,\theta_2)\sigma_{\varepsilon^m}(\theta_1,\theta_2|x_0^+) + \widebar{\varPsi}_p^-(\theta_1,\theta_2)\sigma_{\varepsilon^m}(\theta_1,\theta_2|x_0^-)\), where \begin{align*} \widebar{\varPsi}_p^\pm(\theta_1,\theta_2) &:=\frac{1}{\sqrt{\varsigma(\theta_1)\varsigma(\theta_2)}} \oldint\nolimits_{\mathbb{R}_\pm} \bar{r}_p\left(\frac{u}{\varsigma(\theta_1)}\right) \bar{r}_p\left(\frac{u}{\varsigma(\theta_2)}\right)^\top K\left(\frac{u}{\varsigma(\theta_1)}\right) K\left(\frac{u}{\varsigma(\theta_2)}\right) \,du. \end{align*} \end{itemize}

Part (i) of the lemma establishes the uniform Bahadur representation for our estimator and provides the explicit expression for the asymptotic bias. In contrast to standard local $p$th-order polynomial estimators (e.g., chiang2019robust, Lemma 1), the bias term exhibits two specific characteristics. First, it depends on the $(p+1)$th-order derivatives from both sides of the kink. Second, the associated constant matrices $\widebar{\varGamma}_p$ and vectors $\widebar{\vartheta}^\pm_{p,p+1}$ are constructed using the regressor vector $\bar{r}_p(u) \in \mathbb{R}^{2p+1}$ rather than the standard polynomial basis $r_p(u) \in \mathbb{R}^{p+1}$. Part (ii) of the lemma establishes the weak convergence of the normalized processes. This result requires the use of an undersmoothing bandwidth satisfying $nh_n^{2p+3}=o(1)$. This condition ensures that the higher-order bias term from Part (i) is asymptotically negligible. Consequently, the normalized process converges weakly to a zero-mean Gaussian process, free of asymptotic bias.

We now introduce the assumptions required to establish the asymptotic properties of the quantile-based and composite estimands, $\widehat{\mathrm{QRKD}}$ and $\widehat{\mathrm{RKD}}_{\psi|m}$.

assumption[chiang2019causal, Assumption 6] \begin{itemize} • Assumption (ref)(i). • (a) \(f_{Y|X}(Q_{Y|X}(\cdot|x_0)|x_0)\) is Lipschitz continuous on \(\mathcal{T}\). (b) There exists finite constants \(f_L>0\), \(f_U>0\), and \(\delta>0\) such that for all \(\tau\in\mathcal{T}\), \(|\eta|\leq\delta\) and \(x\in I_{x_0}\), \(f_{Y|X}(Q_{Y|X}(\tau|x)+\eta|x) \in [f_L,f_U]\). • (a) \(Q_{Y|X}(\cdot|x_0)\), \(\partial Q_{Y|X}(\cdot|x_0^\pm)/\partial\tau\) exist and are Lipschitz continuous on \(\mathcal{T}\). (b) \(Q_{Y|X}(\tau|\cdot)\) is continuous at \(x_0\); the function \((x,\tau)\mapsto \partial^v Q_{Y|X}(\tau|x)/\partial x^v\) exists and is Lipschitz continuous on \((x,\tau)\in (I_{x_0}\backslash\{x_0\})\times\mathcal{T}\) for all \(v \in \{0,1,\ldots,p+1\}\). • The kernel \(K\) is compactly supported, nonnegative, and satisfies \(K'(u) < \infty\), \(\oldint\nolimits K(u)\,du = 1\) and \(\oldint\nolimits uK(u)\,du = 0\). The matrix \(\widebar{\varGamma}_p\) is positive definite. • The bandwidths \(h_{n,\tau}\) satisfies \(h_{n,\tau} = c(\tau)h_n\), where \(nh_n^3 \to \infty\) and \(nh_n^{2p+3} \to 0\) as \(n\to \infty\) and \(c(\tau)\) is Lipschitz continuous satisfying \(0 < \underline{c} \leq c(\tau) \leq \bar{c} <\infty \) for all \(\tau\in \mathcal{T}\). \end{itemize}
assumption\begin{itemize} • For each \((g, h)\in \ell^{\infty}(\widebar{\varTheta}) \times \ell^{\infty}(\mathcal{T})\), \(\hat{\psi}(g,h)(\cdot) = \psi(g,h)(\cdot) + o_{P}(1)\) uniformly on \( \widebar{\varTheta}'\). • \(\psi\) is Hadamard differentiable at \((\mathrm{RKD}_m, \mathrm{QRKD})\) with derivative \(\psi'_{(\mathrm{RKD}_m, \mathrm{QRKD})}(\cdot,\cdot)\). \end{itemize}

Assumption (ref) is required to establish the asymptotic properties of $\widehat{\mathrm{QRKD}}$. This assumption is consistent with Assumption 6 of chiang2019causal, and a result that relies on it is restated for completeness as Lemma (ref) in Appendix (ref).

Assumption (ref) provides the high-level conditions required for the asymptotic analysis of our Type 2 composite estimators, $\widehat{\mathrm{RKD}}_{\psi|m}$. Part (i) requires the uniform consistency of any `nuisance' estimators that appear in the definition of the functional $\psi$. For example, when estimating the LLTE, the functional $\psi_L$ depends on the baseline mean $\mu_0$ and the conditional quantile function $Q_{Y|X}(\cdot|x_0)$. The consistent estimators for these nuisance components are conveniently provided by the same constrained regressions, ((ref)) and ((ref)), used for the main effects. Part (ii) imposes the standard Hadamard differentiability condition on the functional $\psi$. This condition enables the application of the functional Delta method to derive the asymptotic distribution of the composite estimator (e.g., kosorok2008introduction, Theorem 2.8).

theorem[Weak Convergence of RKD Estimators] \begin{itemize} • Suppose Assumptions (ref) and (ref) hold. Then, \begin{align*} &\sqrt{nh_{n,\theta}^3}\left(\widehat{\mathrm{RKD}}_m(\theta) - \mathrm{RKD}_m(\theta)\right) \leadsto \mathbb{G}^m(\theta):= \frac{\mathbb{Z}^m(\theta,2) - \mathbb{Z}^m(\theta,3) }{b'(x_0^+)-b'(x_0^-)} \end{align*} uniformly in \(\theta \in \widebar{\varTheta}\). The limiting process $\mathbb{G}^m(\cdot)$ is a zero-mean Gaussian process with covariance function \begin{align*} E\left[\mathbb{G}^m(\theta_1)\mathbb{G}^m(\theta_2)\right] = \frac{(\iota_{2}-\iota_{3})^\top\widebar{\varGamma}_p^{-1}\widebar{\varXi}^m_p(\theta_1,\theta_2) \widebar{\varGamma}_p^{-1}(\iota_{2}-\iota_{3})}{(b'(x_0^+)-b'(x_0^-))^2 f_X(x_0) }, \end{align*} where the terms \(\mathbb{Z}^m(\cdot,\cdot)\) and \(\widebar{\varXi}^m_p(\cdot,\cdot)\) are as defined in Lemma (ref). • Suppose Assumptions (ref), (ref), (ref), and (ref) hold, Then \begin{align*} \sqrt{nh_{n}^3}\left(\widehat{\mathrm{RKD}}_{\psi|m}(\theta') - \mathrm{RKD}_{\psi|m}(\theta') \right) \leadsto \mathbb{G}^{\psi|m}(\theta') :=\psi'_{(\mathrm{RKD}_m, \mathrm{QRKD})}\left(\frac{\mathbb{G}^m(\cdot)}{\sqrt{\varsigma^3(\cdot)}},\frac{\mathbb{G}^Q(\cdot)}{\sqrt{c^3(\cdot)}}\right)(\theta') \end{align*} uniformly in \(\theta'\in\widebar{\varTheta}'\), where $\mathbb{G}^{\psi|m}(\theta')$ is a zero-mean Gaussian process with covariance function \(E\left[\mathbb{G}^{\psi|m}(\theta'_1)\mathbb{G}^{\psi|m}(\theta'_2)\right]\), the form of which depends on \(\psi\). The limiting process \(\mathbb{G}^Q:\varOmega \to \ell^\infty(\mathcal{T})\) is the Gaussian process for the quantile estimator, as defined in Lemma (ref). \end{itemize}

Theorem (ref)(i) provides the weak convergence result for the Type 1 estimators. This result includes the distributional RKD estimator and the mean RKD estimator as special cases. While asymptotic theory for local Wald ratios exists (e.g., the general, higher-order case in chiang2019robust), the result here applies specifically to the first-order derivative and is derived for the constrained estimation approach used in this paper. Theorem (ref)(ii) establishes the weak convergence for the Type 2 (composite) estimators. One application is the local Lorenz effect estimator. The limiting process $\mathbb{G}^{\psi|m}$ for these estimators is a composite of the limiting processes $\mathbb{G}^m$ and $\mathbb{G}^Q$ for the Type 1 and quantile estimators. Because this limiting process is not a local Wald ratio, the inference framework of chiang2019robust, though general in other contexts, cannot be directly applied to these Type 2 estimators.

Multiplier Bootstrap and Pivotal Methods

This section develops a resampling method to approximate the asymptotic distributions derived in Theorem (ref) and to conduct uniform inference.

For the Type 1 estimator, $\widehat{\mathrm{RKD}}_m(\cdot)$, we employ a multiplier bootstrap to approximate the distribution of the limiting process $\mathbb{G}^m(\cdot)$. The bootstrap process is constructed based on the leading term of the uniform Bahadur representation established in Lemma (ref)(i). Let $\{\xi_i\}_{i=1}^n$ be a sequence of i.i.d. standard normal random variables, drawn independently of the original data. We define the estimated multiplier process (EMP) as:

align[align omitted — 356 chars of source]

for each $\theta\in\widebar{\varTheta}$. This process requires estimators for the unknown components in the asymptotic representation. The residuals are computed using the fitted values from the constrained regression, \(\hat{\varepsilon}^m(y,x,\theta):=\left(\varphi(y,\theta) - \bar{r}_p(x-x_0)\hat{\alpha}(\theta)\right)\mathbbm{1}\left(\left|x-x_0\right|\leq h_{n,\theta}\right)\), whose uniform consistency is established in Lemma (ref). The density \(f_X(x_0)\) is estimated using a standard kernel density estimator \(\hat{f}_X(x_0)=\frac{1}{nv_n}\sum_{i=1}^{n}K\left((X_i-x_0)/v_n\right)\).

Approximating the asymptotic distribution for the Type 2 estimator $\widehat{\mathrm{RKD}}_{\psi|m}$ requires to generate valid bootstrap analogues for both. The process $\mathbb{G}^m$ is simulated using the multiplier bootstrap $\widehat{\mathbb{G}}^m_{\xi}$ as described previously. For the quantile component $\mathbb{G}^Q$, we employ the pivotal method developed by chiang2019causal and qu2015nonparametric. Let $\{U_i\}_{i=1}^n$ be a sequence of i.i.d. Uniform$(0,1)$ random variables, drawn independently of the data. The estimated pivotal process for $\mathbb{G}^Q$ is then constructed as:

align*[align* omitted — 371 chars of source]

The bootstrap analogue for the composite limiting process $\mathbb{G}^{\psi|m}$ is then constructed by combining these two simulated processes using the estimated Hadamard derivative of the functional $\psi$:

align[align omitted — 289 chars of source]

The following high-level assumption is required to establish the uniform validity of our resampling procedures.

assumption[First Stage Estimation] \begin{itemize} • \(\hat{f}_X(x_0) = f_X(x_0) + o_{P}(1)\), and \(\hat{f}_{Y|X}(y|x_0) = f_{Y|X}(y|x_0) + o_{P}(1)\) uniformly in \(y \in \mathcal{Y}_{x_0}\). • For each \((g, h )\in \ell^{\infty}(\widebar{\varTheta}) \times \ell^{\infty}(\mathcal{T}) \), the estimated Hadamard derivative satisfies \[\hat{\psi}'_{(\mathrm{RKD}_m, \mathrm{QRKD})}(g,h)(\cdot) = \psi'_{(\mathrm{RKD}_m, \mathrm{QRKD})}(g,h)(\cdot) + o_{P}(1)\] uniformly on \(\widebar{\varTheta}'\) \end{itemize}

Assumption (ref) ensures the consistency of the nuisance components used in constructing $\widehat{\mathbb{G}}^m_{\xi}$ and $\widehat{\mathbb{G}}^{\psi|m}_{\xi,u}$. The role of Part (ii) can be illustrated using the Lorenz effect $\Delta_L$ as an example. For this effect, the functional $\psi_L$ is bilinear, and its Hadamard derivative, $[\psi_L]'$, depends on the nuisance parameters $\mu_0$ and the function $Q_{Y|X}(\cdot|x_0)$. In this context, Assumption (ref)(ii) is a high-level condition requiring uniform consistency of the estimators for these nuisance components, i.e., $\hat{\mu}_0 \overset{P}{\to} \mu_0$ and $\sup_{\tau \in \mathcal{T}} |\widehat{Q}_{Y|X}(\tau|x_0) - Q_{Y|X}(\tau|x_0)| \overset{P}{\to} 0$.

Let \(\overunderset{P}{M}{\leadsto}\) denote the conditional weak convergence, where the subscript $M$ indicates conditional expectation over the weights $M$ given the remaining data (kosorok2008introduction).

theorem[Conditional Weak Convergence] Suppose Assumptions (ref)--(ref) hold. Then: \begin{itemize} • \(\widehat{\mathbb{G}}^m_{\xi}(\theta) \overunderset{P}{\xi}{\leadsto} \mathbb{G}^m(\theta)\) uniformly in \(\theta \in \widebar{\varTheta}\). • \(\widehat{\mathbb{G}}^{\psi|m}_{\xi,u}(\theta') \overunderset{P}{\xi\times u}{\leadsto} \mathbb{G}^{\psi|m}(\theta')\) uniformly in \(\theta' \in \widebar{\varTheta}'\). \end{itemize}

Theorem (ref) establishes that the empirical distributions of the simulated processes, $\widehat{\mathbb{G}}^m_{\xi}(\theta)$ and $\widehat{\mathbb{G}}^{\psi|m}_{\xi,u}(\theta')$, validly approximate the asymptotic distributions of the normalized processes of $\widehat{\mathrm{RKD}}_m(\theta) $ and $\widehat{\mathrm{RKD}}_{\psi|m}(\theta')$, respectively. This result enables the construction of asymptotically valid uniform confidence bands and hypothesis tests in practice.

Hypothesis Tests and Uniform Confidence Bands

Building on the asymptotic results of Theorems (ref) and (ref), this section details the procedure for conducting uniform inference on the local treatment effects, including uniform hypothesis testing and the construction of uniform confidence bands. Our goal is to conduct inference on the functions $\theta \mapsto \Delta_m(\theta)$ and $\theta' \mapsto \Delta_{\psi|m}(\theta')$, which represent our Type 1 and Type 2 LTEs, respectively.

To test for the uniform significance of the treatment effect, we consider the null hypothesis that the LTE function is identically zero over the region of interest. For our Type 1 and Type 2 effects, respectively, the hypotheses are:

align*[align* omitted — 243 chars of source]

where \(\widebar{\varTheta} \subset \varTheta\) and \(\widebar{\varTheta}' \subset \varTheta'\) are compact spaces of interests. We test these hypotheses using Kolmogorov–Smirnov type test statistics, which measure the maximum deviation of the estimated process from zero:

align*[align* omitted — 326 chars of source]

To test for the homogeneity of the treatment effect, we consider the null hypothesis that the LTE function is constant, though not necessarily zero. The hypotheses are:

align*[align* omitted — 314 chars of source]

These hypotheses can be tested by measuring the deviation from the mean over the region of interest. The specific statistics are defined as:

align*[align* omitted — 597 chars of source]

The following corollary, which is a direct consequence of Theorem (ref), establishes the asymptotic distribution of our test statistics under their respective null hypotheses.

corollarySuppose Assumptions (ref)--(ref) hold. Then, \begin{itemize} • \(W_{m}^S(\widebar{\varTheta}) \leadsto \sup_{\theta \in \widebar{\varTheta}}\left|\mathbb{G}^m(\theta)\right|\) under the null hypothesis \(\mathscr{H}_{0,m}^S\); \(W_{m}^H(\widebar{\varTheta}') \leadsto \sup_{\theta \in \widebar{\varTheta} } \left|\mathbb{G}^m(\theta) - \frac{1}{|\widebar{\varTheta}|} \oldint\nolimits_{\widebar{\varTheta}} \mathbb{G}^m(\theta) \,d\theta \right|\) under the null hypothesis \(\mathscr{H}_{0,m}^H\). • \(W_{\psi|m}^S(\widebar{\varTheta}) \leadsto \sup_{\theta \in \widebar{\varTheta}}\left|\mathbb{G}^{\psi|m}(\theta)\right|\) under the null hypothesis \(\mathscr{H}_{0,\psi|m}^S\); \(W_{\psi|m}^H(\widebar{\varTheta}')\leadsto \sup_{\theta' \in \widebar{\varTheta}'} \left|\mathbb{G}^{\psi|m}(\theta') - \frac{1}{|\widebar{\varTheta}'|} \oldint\nolimits_{\widebar{\varTheta}'} \mathbb{G}^{\psi|m}(\theta') \,d\theta' \right|\) under the null hypothesis \(\mathscr{H}_{0,\psi|m}^H\). \end{itemize}

In practice, the asymptotic distributions of the test statistics in Corollary (ref) can be simulated based on the estimated limiting processes \(\widehat{G}^m_{\xi}\) and $\widehat{G}^{\psi|m}_{\xi,u}$, respectively.

Building on the bootstrap validity established in Theorem (ref), we now detail the construction of uniform confidence bands for our LTEs. The procedure for a generic LTE function, $\Delta(\cdot)$, follows a standard two-step process:

Step 1: Estimate the Critical Value. We simulate the distribution of the supremum norm of the relevant bootstrap process. For $\lambda\in (0,1)$, the estimated $(1-\lambda)$ critical value, $\hat{c}(1-\lambda)$, is then obtained as the $(1-\lambda)$ sample quantile of this simulated distribution.

Step 2: Construct the Band. The $100(1-\lambda)\%$ uniform confidence band for the function $\Delta(\cdot)$ is then formed by: $$ \left[ \widehat{\Delta}(\cdot) \pm \hat{c}(1-\lambda)\big/\sqrt{nh_{n,\cdot}^3}\right], $$ where $\widehat{\Delta}(\cdot)$ is the corresponding point estimator $\widehat{\mathrm{RKD}}_m(\theta)$ or $\widehat{\mathrm{RKD}}_{\psi|m}(\theta')$. Specifically, for the Type 1 LTE $\Delta_m$, we use the estimated bootstrap process $\widehat{\mathbb{G}}^m_{\xi}$ to compute $\hat{c}_m(1-\lambda)$, and for the Type 2 LTE $\Delta_{\psi|m}$, we use the estimated composite process $\widehat{\mathbb{G}}^{\psi|m}_{\xi,u}$ to compute $\hat{c}_{\psi|m}(1-\lambda)$.

This section concludes by detailing the implementation of our inference procedures for the specific LTEs discussed previously. Bandwidth selection procedures for each effect are provided in Appendix (ref). Let $B$ be the number of bootstrap replications. \setcounter{example}{0}

example[Average Effect, Continued] The LATE \(\Delta_\mu\) is estimated by $\widehat{\mathrm{RKD}}_\mu$. For each replication $b = 1, \dots, B$, we generate an i.i.d. sequence of \(N(0,1)\) random weights $\{\xi_i^b\}_{i=1}^n$, drawn independently of the data. Then, the distribution of \(\sqrt{nh_n^3}(\widehat{\mathrm{RKD}}_{\mu}-\Delta_\mu)\) is approximated by the empirical distribution of \(\{\widehat{\mathbb{G}}^{\mu}_{\xi^b}\}_{b=1}^B\), where \begin{align*} \widehat{\mathbb{G}}^{\mu}_{\xi^b} :=\sum_{i=1}^{n} \xi^b_i \frac{(\iota_2 - \iota_3)^\top \widebar{\varGamma}_p^{-1}\bar{r}_p\left(\frac{X_i-x_0}{h_n}\right) K\left(\frac{X_i-x_0}{h_n }\right) \hat{\varepsilon}^\mu(Y_i,X_i)}{(b'(x_0^+)-b'(x_0^-)) \sqrt{nh_n} \hat{f}_X(x_0)}. \end{align*} The random error is estimated by \(\hat{\varepsilon}^\mu(y,x)=\left(y- \bar{r}_p(x-x_0)\widehat{\alpha}\right)\mathbbm{1}\left(\left|x-x_0\right|\leq h_n\right)\). To test the significance hypothesis, $\Delta_\mu = 0$, we first compute the test statistic $W_\mu^S = |\sqrt{nh_n^3}\,\widehat{\mathrm{RKD}}_\mu|$. The $(1-\lambda)$ critical value, $\hat{c}_\mu(1-\lambda)$, is the $(1-\lambda)$ sample quantile of the simulated distribution $\{|\widehat{\mathbb{G}}^\mu_{\xi^b}|\}_{b=1}^B$. We reject the null if $W_\mu^S > \hat{c}_\mu(1-\lambda)$. The bootstrap $p$-value is calculated as $\hat{p}^S_\mu = \frac{1}{B}\sum_{b=1}^{B}\mathbbm{1}(|\widehat{\mathbb{G}}^\mu_{\xi^b}| > W_\mu^S)$. Finally, the uniform confidence band is constructed as $\left[\widehat{\mathrm{RKD}}_\mu \pm \hat{c}_\mu(1-\lambda)/\sqrt{nh_n^3}\right]$.
example[Distributional and Quantile Effects, Continued] \\ The LDTE $\Delta_{Id}(\cdot)$ is estimated by $\widehat{\mathrm{RKD}}_F(\cdot)$. A common application is to test properties of the LDTE $\Delta_{Id}$ over a range of quantiles, i.e., to conduct inference on the process $\{\Delta_{Id}(y_\tau)\}_{\tau \in \mathcal{T}}$. For each replication $b = 1, \dots, B$, generate an i.i.d. sequence of \(N(0,1)\) weights $\{\xi_i^b\}_{i=1}^n$. We construct the $b$th draw of the bootstrap process $\widehat{\mathbb{G}}^{Id}_{\xi^b}(\cdot)$ as \begin{align*} \widehat{\mathbb{G}}^{Id}_{\xi^b}(\theta):=\sum_{i=1}^{n} \xi^b_i \frac{(\iota_2 - \iota_3)^\top \widebar{\varGamma}_p^{-1}\bar{r}_p\left(\frac{X_i-x_0}{h_{n,\theta}}\right) K\left(\frac{X_i-x_0}{h_{n,\theta} }\right) \hat{\varepsilon}^F(Y_i,X_i,\theta)}{(b'(x_0^+)-b'(x_0^-)) \sqrt{nh_{n,\theta}} \hat{f}_X(x_0)}, \end{align*} where the residual is computed as $\hat{\varepsilon}^F(y,x,\theta)=\left(\mathbbm{1}(y\leq \theta) - \bar{r}_p(x-x_0)\widehat{\alpha}\right)\mathbbm{1}\left(\left|x-x_0\right|\leq h_n\right)$. For each replication $b$, evaluate the bootstrap process $\widehat{\mathbb{G}}^{Id}_{\xi^b}(\cdot)$ at the estimated quantile points $\{\hat{y}_\tau\}_{\tau \in \mathcal{T}}$ to obtain the simulated process $\{\widehat{\mathbb{G}}^{Id}_{\xi^b}(\hat{y}_\tau)\}_{\tau \in \mathcal{T}}$. To test the hypothesis of uniform treatment significance $\Delta_{Id}(y_\tau) = 0$ for all $\tau \in \mathcal{T}$, we first compute the test statistic \(W_{Id}^S:=\sup_{\tau \in \mathcal{T}}|\sqrt{nh_{n,y_\tau}^3}\, \widehat{\mathrm{RKD}}_F(\hat{y}_\tau)|\). Then, we use $\{\widehat{\mathbb{G}}^S_{Id,b}:=\sup_{\tau\in\mathcal{T}}|\widehat{\mathbb{G}}^{Id}_{\xi^b}(\hat{y}_\tau)|\}_{b=1}^B$ to approximate the distribution of \(W_{Id}^S\). The $(1-\lambda)$ critical value, $\hat{c}_{Id}(1-\lambda)$, is the $(1-\lambda)$ sample quantile of $\{\widehat{\mathbb{G}}^S_{Id,b}\}_{b=1}^B$. The bootstrap $p$-value is given by $\hat{p}^S_{Id}:=\frac{1}{B}\sum_{b=1}^{B}\mathbbm{1}(\widehat{\mathbb{G}}^S_{Id,b}>W_{Id}^S)$. To test the hypothesis of uniform treatment homogeneity $\Delta_{Id}(y_{\tau_1}) = \Delta_{Id}(y_{\tau_2})$ for all $\tau_1, \tau_2 \in \mathcal{T}$, we first compute the test statistic \[W^H_{Id}:=\sup_{\tau\in\mathcal{T}} \left| \sqrt{nh_{n,y_\tau}^3}\left( \widehat{\mathrm{RKD}}_F(\hat{y}_\tau) - \frac{1}{|\mathcal{T}|} \oldint\nolimits_{\mathcal{T}} \widehat{\mathrm{RKD}}_F(\hat{y}_\tau)\,d\tau \right) \right|. \] Then, we use $\{\widehat{\mathbb{G}}^H_{Id,b}:=\sup_{\tau\in\mathcal{T}}|\widehat{\mathbb{G}}^{Id}_{\xi^b}(\hat{y}_\tau) - \frac{1}{|\mathcal{T}|} \oldint\nolimits_{\mathcal{T}} \widehat{\mathbb{G}}^{Id}_{\xi^b}(\hat{y}_\tau)\,d\tau|\}_{b=1}^B$ to approximate the distribution of $W^H_{Id}$. The critical value and $p$-value are calculated analogously to the significance test. Finally, the uniform confidence band is constructed as $$\left[\widehat{\mathrm{RKD}}_F(\hat{y}_\tau) \pm \hat{c}_{Id}(1-\lambda)/\sqrt{nh_{n,y_\tau}^3}\right].$$ Inference for the LQTE, $\Delta_Q(\cdot)$, which is estimated by $\widehat{\mathrm{QRKD}}(\cdot)$, is conducted using the pivotal method detailed in Section (ref). This involves generating i.i.d. Uniform$(0,1)$ random variables to construct the pivotal process $\{\widehat{\mathbb{G}}^Q_{u^b}(\tau)\}_{\tau \in \mathcal{T}}$ for each replication $b$. The subsequent steps for constructing test statistics and obtaining critical values are analogous to those for the LDTE. For the sake of brevity, we omit the detailed formulas and refer readers to the comprehensive procedure outlined in Appendix C.2 of chiang2019causal, which our method follows directly.
example[Lorenz Effect, Continued] By Equation ((ref)), the LLTE \(\Delta_{L}(\tau)\) is estimated by \(\widehat{\mathrm{RKD}}_{\psi_L|\mu}(\tau)\). For each bootstrap iteration \(b\in\{1,\ldots,B\}\), we draw the i.i.d. sequence \(\{\xi^b_i\}_{i=1}^n\) from \(N(0,1)\), and the i.i.d. sequence \(\{u^b_i\}_{i=1}^n\) from Uniform$(0,1)$, respectively. Then, we approximate the distribution of \(\sqrt{n(h^L_{n,\tau})^3}(\widehat{\mathrm{RKD}}_{\psi_L|\mu}(\tau)-\Delta_{L}(\tau))\) by the empirical distribution of \(\{\widehat{\mathbb{G}}^L_{\xi^b,u^b}(\tau)\}_{b=1}^B\) for each \(\tau \in \mathcal{T}\), where \begin{align*} \widehat{\mathbb{G}}^L_{\xi^b,u^b}(\tau) := \frac{1}{\hat{\mu}_0} \left(\oldint\nolimits_{0}^{\tau} \widehat{\mathbb{G}}^Q_{u^b}(u)\,du - \widehat{L}_{Y|X}(\tau|x_0) \cdot \widehat{\mathbb{G}}^{\mu}_{\xi^b} \right). \end{align*} It is important to note the bandwidth choices used in the bootstrap procedures. For the mean effect bootstrap, $\widehat{\mathbb{G}}^{\mu}_{\xi^b}$, the single baseline bandwidth $h_n$ is employed. In contrast, for the quantile effect bootstrap, $\widehat{\mathbb{G}}^Q_{u}(\tau)$, the pointwise bandwidth $h_{n,\tau}$ is used. As established by Lemma (ref), using the pointwise bandwidth $h_{n,\tau}$ ensures that this bootstrap process correctly targets the desired limiting process $\mathbb{G}^Q(\tau)$, rather than the scaled version. To test the hypothesis of uniform treatment significance $\Delta_L(\tau) = 0$ for all \(\tau \in \mathcal{T}\), we first compute the test statistic $W_{L}^S:=\sup_{\tau \in \mathcal{T}}|\sqrt{n(h^L_{n,\tau})^3}\, \widehat{\mathrm{RKD}}_{\psi_L|\mu}(\tau)|$. Then, we use $\{\widehat{\mathbb{G}}^S_{L,b}:=\sup_{\tau\in\mathcal{T}}|\widehat{\mathbb{G}}^L_{\xi^b,u^b}(\tau)|\}_{b=1}^B$ to approximate the distribution of $W_{L}^S$. The \((1-\lambda)\)th critical value, $\hat{c}_{L}(1-\lambda)$, is \((1-\lambda)\) sample quantile of $\{\widehat{\mathbb{G}}^S_{L,b}\}_{b=1}^B$. The $p$-value is given by \(\hat{p}^S_{L}:=\frac{1}{B}\sum_{b=1}^{B}\mathbbm{1}(\widehat{\mathbb{G}}^S_{L,b}>W_{L}^S)\). To test the hypothesis of uniform treatment homogeneity $\Delta_{L}(\tau_1) = \Delta_{L}(\tau_2)$ for all $\tau_1, \tau_2 \in \mathcal{T}$, we first compute the test statistic the test statistic \[W^H_{L}:=\sup_{\tau\in\mathcal{T}} \left| \sqrt{n(h^L_{n,\tau})^3} \left(\widehat{\mathrm{RKD}}_{\psi_L|\mu}(\tau) - \frac{1}{|\mathcal{T}|} \oldint\nolimits_{\mathcal{T}} \widehat{\mathrm{RKD}}_{\psi_L|\mu}(\tau)\,d\tau \right) \right|. \] Then, we use $\{\widehat{\mathbb{G}}^H_{L,b}:=\sup_{\tau\in\mathcal{T}}|\widehat{\mathbb{G}}^L_{\xi^b,u^b}(\tau) - \frac{1}{|\mathcal{T}|} \oldint\nolimits_{\mathcal{T}} \widehat{\mathbb{G}}^L_{\xi^b,u^b}(\tau)\,d\tau|\}_{b=1}^B$ to approximate the distribution of $W^H_{L}$. The critical value is \((1-\lambda)\) sample quantile of $\{\widehat{\mathbb{G}}^H_{L,b}\}_{b=1}^B$. The $p$-value of the treatment homogeneity test is given by \(\hat{p}^H_{L}=\frac{1}{B}\sum_{b=1}^{B}\mathbbm{1}(\widehat{\mathbb{G}}^H_{L,b} > W^H_{L})\). Finally, the uniform confidence band is constructed as $$\left[\widehat{\mathrm{RKD}}_{\psi_L|\mu}(\tau) \pm \hat{c}_{L}(1-\lambda)/\sqrt{n(h^L_{n,\tau})^3}\right].$$

Numerical Illustrations

Empirical Application

In this section, we apply our framework to a classic dataset from the Continuous Wage and Benefit History (CWBH) to analyze the causal effects of unemployment insurance (UI) benefits on unemployment durations. The data cover unemployed individuals in Louisiana over two periods: September 1981--September 1982 and September 1982--December 1983. Our analysis builds upon the work of landais2015assessing, to which we refer readers for detailed descriptive statistics and institutional background. The dataset contains the following variables: (i) Running variable ($X$): The highest quarterly wage in an individual's base period. (ii) Treatment ($B$): The weekly UI benefit amount received. (iii) Outcome ($Y$): The duration of unemployment, measured in two ways: (a) the number of weeks benefits were claimed (UI Claimed), and (b) the number of weeks benefits were paid (UI Paid). This setting is a canonical example of a sharp regression kink design. The UI benefit amount $B$ is a deterministic, piecewise-linear function of the base wage $X$: \(b(x)=\min\{0.04\cdot x, b_{max}\}\). The maximum benefit level $b_{max}$, which defines the location of the kink point $x_0$, changed between the two periods. For the 1981--1982 period, $b_{max}$ was $\$ 4,575$, whereas for the 1982--1983 period, it increased to $\$ 5,125$.

Building on the theoretical framework, this analysis examines the local distributional (LDTE) and local Lorenz (LLTE) treatment effects of UI benefits on unemployment durations. This focus extends the existing literature, which has primarily examined the average effect (landais2015assessing) and the quantile effect (chiang2019causal). We estimate the LDTE, $\Delta_{Id}(y_\tau)$, and the LLTE, $\Delta_L(\tau)$, over a grid of quantiles $\tau$. The full estimated effect functions for both time periods are displayed graphically in Figures (ref) and (ref). For specific numerical results, Tables (ref) and (ref) present point estimates and standard errors for these effects at selected quantiles $\tau \in \{0.1, 0.2, \dots, 0.9\}$, along with $p$-values from the corresponding uniform hypothesis tests. All estimation and inference procedures follow the methods detailed in Section (ref). Consistent with chiang2019causal, we employ a tricube kernel function, $K(u) = \frac{70}{81}(1-|u|^3)^3\mathbbm{1}(|u|\leq 1)$, and set the local polynomial order to $p=2$ for bias-correction. Selection of the MSE-optimal bandwidths follows the procedures detailed in Algorithms (ref) and (ref) in Appendix (ref).

Panel A of the tables presents the results for the LDTE. The findings are summarized below. First, the estimated LDTE is negative across both time periods examined. This implies that a marginal increase in UI benefits reduces the cumulative probability of an unemployment spell ending within any given number of weeks for individuals at the kink wage. Second, the magnitude of the estimated effect varies with duration and changes between periods. In the 1981--1982 period, the effect's magnitude is smaller at shorter durations (e.g., the 10th percentile) and increases at longer durations. This pattern appears reversed in the 1982--1983 period, where the effect is largest in magnitude at short durations (e.g., the 20th percentile) and smaller at longer durations. Third, the LDTE is statistically significant at most quantile levels in both periods. Furthermore, formal tests for homogeneity are strongly rejected, confirming that the distributional effect is not constant. The results are also robust across the two different measures of unemployment duration (UI Claimed vs. UI Paid).

Panel B presents the results for the LLTE, which measures the effect of UI benefits on the inequality of the duration distribution. The estimated LLTE is positive at all quantiles considered in both periods. This indicates that a marginal increase in UI benefits increases the dispersion (inequality) of unemployment durations for individuals at the kink wage. This finding is consistent with the quantile effect patterns reported in chiang2019causal. Their observation---that benefit increases have a larger effect on longer unemployment spells---implies an increase in dispersion, which the LLTE directly measures. Similar to the LDTE, the estimated Lorenz effects are statistically significant in both periods. Tests also reject the null hypothesis of homogeneity. The findings remain consistent across the two measures of unemployment duration.

table[table omitted — 2,143 chars of source]
table[table omitted — 2,141 chars of source]
figure[figure omitted — 300 chars of source]
figure[figure omitted — 299 chars of source]

Simulation Experiment

This section evaluates the finite-sample performance of the proposed estimation and inference methods using Monte Carlo simulations. We generate i.i.d. samples $\{(Y_i,B_i,X_i)\}_{i=1}^n$ from the following data generating process:

align*[align* omitted — 446 chars of source]

The structure is similar to that in chiang2019causal but is extended to include conditional heteroskedasticity through the $(1+2B_i)\varepsilon_i$ term. The kink point is normalized to $x_0=0$. The distributional parameters are chosen to match those in calonico2014robust, with $\sigma_x = 0.1781742$, $\sigma_\varepsilon = 0.1295$, and correlation $\rho = 0.25$.

We evaluate the finite-sample performance of estimators for four local treatment effects: the average effect ($\Delta_\mu$), the distributional effect ($\Delta_{Id}$), the quantile effect ($\Delta_Q$), and the Lorenz effect ($\Delta_L$). The LATE $\Delta_\mu$ is estimated at a single point. The LQTE $\Delta_Q(\tau)$ and LLTE $\Delta_L(\tau)$ are evaluated over the quantile grid $\mathcal{T}=\{0.1, 0.2, \dots, 0.9\}$. The LDTE $\Delta_{Id}(y)$ is evaluated at points corresponding to the quantiles in $\mathcal{T}$, i.e., at $y = Q_{Y|X}(\tau|x_0)$ for $\tau \in \mathcal{T}$. Estimation and inference follow the procedures detailed in Section (ref). For the LLTE estimator, required numerical integrals are computed using a finer grid, $\mathcal{U}=\{0.01, \dots, 0.99\}$. Nuisance parameters (e.g., $\mu_0$, $Q_{Y|X}$) are estimated as part of the main procedures. We consider sample sizes \(n\in\{500,1000,2000,4000\}\) and set the number of Monte Carlo replications to 5,000 for each simulation design. In each replication, we employ both the multiplier bootstrap and the pivotal method, each based on 2,500 bootstrap draws, to construct the uniform confidence bands.

Figure (ref) plots the estimated effects across simulations against the true parameter values for different sample sizes. Table (ref) reports quantitative measures of finite-sample performance, including the absolute bias ratio ($|\widehat{\Delta}-\Delta|/|\Delta|$), root mean square error (RMSE), and empirical coverage probability for 95% uniform confidence bands. The results show a general improvement as the sample size $N$ increases. Specifically, the RMSE consistently decreases with $N$ for all estimators. The bias ratios also tend to decrease, although not always monotonically across all quantile points examined. This overall pattern indicates satisfactory finite-sample performance and suggests convergence of the estimators. Furthermore, the empirical coverage rates of the 95% uniform confidence bands approach the nominal 95% level as the sample size grows, supporting the validity of the resampling inference procedure.

table[table omitted — 3,128 chars of source]
figure[figure omitted — 375 chars of source]

Conclusion

This paper develops a unified framework for the identification, estimation, and uniform inference of a class of local treatment effects (LTEs) in the sharp regression kink design. The identification strategy applies to Hadamard differentiable functionals of the outcome distribution and utilizes a continuity assumption on the conditional density of the outcome variable at the kink point. This approach yields a unified identification formula and establishes a connection between the LTE and the local average structural derivative.

For estimation and inference, we consider two classes of local polynomial constrained regression estimators and develop their asymptotic theory, including a resampling method for uniform inference. The framework covers parameters such as effects on the mean, the distribution, and inequality measures. We apply the methods to re-examine the effect of unemployment insurance on unemployment durations, estimating the local distributional and Lorenz treatment effects. This analysis provides information on the impacts of the policy on the shape and dispersion of the duration distribution. An extension of this framework to the fuzzy regression kink design could be a direction for future research, potentially broadening its applicability.

\setcounter{equation}{0}