EconBase
← Back to paper

Identification and Estimation in Fuzzy Regression Discontinuity Designs with Covariates

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

156,462 characters · 25 sections · 25 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identification and Estimation in Fuzzy Regression Discontinuity Designs with Covariates

abstractWe study fuzzy regression discontinuity designs with covariates and characterize the weighted averages of conditional local average treatment effects (WLATEs) that are point identified. Any identified WLATE equals a Wald ratio of conditional reduced-form and first-stage discontinuities. We highlight the Compliance-Weighted LATE (CWLATE), which weights cells by squared first-stage discontinuities and maximizes first-stage strength. For discrete covariates, we provide simple estimators and robust bias-corrected inference. In simulations calibrated to common designs, CWLATE improves stability and reduces mean squared error relative to standard fuzzy RDD estimators when compliance varies. An application to Uruguayan cash transfers during pregnancy yields precise RDD-based effects on low birthweight.

Introduction

Regression Discontinuity Design (RDD) is a leading quasi-experimental design Imbens_Lemieux,Lee_Lemieux,Cattaneo_Escanciano,Cattaneo_Idrobo_Titiunik. In fuzzy RDDs, small first-stage discontinuities often render local Wald estimators imprecise and sensitive to tuning choices Feir_Lemieux_Marmer,Noack_Rothe. At the same time, conditional first-stage discontinuities typically vary with observed covariates, implying heterogeneous compliance. This motivates two questions: (i) which causal parameters are point identified in fuzzy RDDs with covariates, and (ii) among those, which are best supported by the design in the sense of leveraging the strongest first-stage information.

We answer both questions within a unified framework. We characterize a broad class of weighted local average treatment effects (WLATEs): weighted averages of covariate-specific LATEs at the cutoff with nonnegative weights summing to one. Our first result is an identification theorem: under conditions described below, a WLATE is point identified if and only if every covariate value with positive weight (i) has a nonzero conditional first-stage discontinuity at the cutoff, and (ii) appears with positive probability on both sides arbitrarily close to the cutoff. Any identified WLATE can be written as a Wald ratio in a transformed model where the “treatment” is the covariate-specific first-stage discontinuity and the “outcome” is the corresponding reduced-form discontinuity; equivalently, each WLATE is associated with an “instrument” given by a function of covariates, with different choices inducing different complier weights. This representation is a device for organizing what can be identified by cutoff discontinuities in fuzzy RDDs.

Within the identified WLATE class, we single out the Compliance-Weighted LATE (CWLATE). Choosing an “instrument” proportional to the conditional first-stage discontinuity implies weights proportional to the squared first-stage. Subject to the sign and support restrictions mandated by identification, this “instrument” has maximum correlation with the first-stage among admissible choices, making the CWLATE the member of the class that is most aligned with available first-stage information. In this sense, the CWLATE is the “easiest to estimate” in our class.\footnote{We borrow the “easy-to-estimate” terminology from goldsmith2024contamination, though here “easy” refers to strongest identification rather than to efficiency.}

We provide an economic interpretation of the identified WLATEs as the effect of a targeted policy where incentives are assigned by drawing a covariate value from a chosen distribution, and then offering the incentive to a random individual with that covariate value at the cutoff. Specifically, the “instrument” of a given WLATE is the ratio between the likelihood of that covariate value under the policy and the likelihood of that covariate value in the population. The “instrument” thus defines the relative tilt of the policy in favor of or against certain covariate values in comparison to the incidence of those values in the population. In that framework, the standard RDD estimand can be interpreted as the effect of a policy where incentives are offered randomly at the cutoff. In contrast, the CWLATE can be interpreted as the effect of a policy that targets compliers, by offering proportionally more incentives to groups that are more likely to have compliers.

Compliance-weighting has antecedents in the IV literature (e.g., joffe2003weighting,borusyak2020non,huntington2020instruments,coussens2021improving, CaetanoEscanciano21, hazard2022improving, abadie2023instrumental), where the predominant concern is efficiency in regular conditional moment models. Fuzzy RDDs, by contrast, target local and intrinsically nonregular parameters defined by discontinuities at a cutoff, so the usual efficiency bounds are not well-defined in the sense of chamberlain1992efficiency. Accordingly, we take an identification-first approach: we characterize the full class of point identified WLATEs, and show that the CWLATE is the member of this class least prone to identification issues.

Our paper also complements a growing literature on fuzzy RDDs with covariates that uses covariates predominantly to improve precision, while retaining the fuzzy-RDD LATE as the target estimand. Covariate-adjusted local regressions are advocated in Imbens_Lemieux and analyzed in depth by Calonico_Cattaneo_Farrell_Titiunik, who characterize conditions under which such adjustment identifies the unconditional LATE and yields efficiency gains. Frolich and Frolich_Huber provide nonparametric representations and estimators of the unconditional LATE based on covariate-specific effects, and noack2021flexible proposes flexible adjustment methods that can accommodate large covariate vectors. We share the concern that local Wald ratios can be highly sensitive to small first-stage discontinuities, but we shift emphasis to estimand selection: when conditional first-stage discontinuities vary with covariates, fuzzy RDDs point identify a class of cutoff-local effects, and the CWLATE concentrates on where the conditional first-stage is strongest. Our weak-monotonicity condition parallels related discussions in compliance-weighted IV interpretations in kolesa2013estimation,sloczynski2020should,huntington2020instruments.

We treat covariates as locally pre-determined features used solely to condition local discontinuities. Identification relies on four requirements near the cutoff: (a) local independence of unobservables and treatment effects from the running variable conditional on covariates; (b) weak monotonicity, i.e. for each covariate value either there are compliers or defiers near the cutoff, but not both; (c) overlap (each positively weighted covariate value occurs both just above and just below the cutoff); and (d) a nonzero conditional first-stage. Our conditions are equivalent to the standard assumptions for RDD, but conditional on covariates. Importantly, we do not require smoothness or balance in the marginal covariate distribution near the cutoff, nor do we require strong monotonicity (which completely rules out defiers). As such, the conditions used to establish the WLATE class and the identification of the CWLATE are generally weaker than the conditions in the previous literature on fuzzy RDDs with covariates.

For estimation, we focus on discrete (or coarsened) covariates. We propose simple plug-in estimators for the identified WLATE class, based on local-polynomial estimates of covariate-specific reduced-form and first-stage discontinuities, and develop robust bias-corrected (RBC) inference for the CWLATE, building on Calonico_Cattaneo_Titiunik; the Online Appendix provides the asymptotic theory and parallel results for more general WLATEs.

We investigate the finite sample performance of the CWLATE estimator in a Monte Carlo study, and show that, when compliance varies across covariates, the CWLATE estimator is markedly more stable and typically attains lower mean square error (MSE) than the standard RDD estimators (with or without covariate adjustment).

We also revisit amarante2016cash on cash transfers during pregnancy in Uruguay (eligibility via a predicted-income cutoff): conventional fuzzy RDD estimates are imprecise in preferred specifications, whereas CWLATE-based estimates point to sizable reductions in low birthweight among compliers, illustrating how aligning the estimand with where the design is strongest can “rescue” an otherwise weak-appearing fuzzy RDD in settings with heterogeneous compliance.

The rest of the paper proceeds as follows. Section (ref) formalizes the model, states the WLATE identification result, introduces the CWLATE, and discusses policy interpretations. Section (ref) describes estimation and RBC inference. Section (ref) reports Monte Carlo evidence. Section (ref) presents the empirical application. Section (ref) concludes. An Online Appendix contains additional results and proofs of the main results.

Identification in RDD with Covariates

Model Setup

We consider a fuzzy RDD with covariates. Let $X_i\in\{0,1\}$ denote treatment and $(Y_i(1),Y_i(0))$ the potential outcomes. The observed outcome is

\[ Y_i \;=\; Y_i(0) \;+\; \beta_i X_i, \] where $\beta_i = Y_i(1)-Y_i(0)$ may be heterogeneous and correlated with treatment. In addition to $(Y_i,X_i)$, we observe a continuously distributed running variable $Z_i$ with cutoff $z_0$, and a vector of covariates $W_i$ with support $\mathcal W$.

For each $z$, let $X_i(z)$ denote the potential treatment if $Z_i=z$. Define

\[ \Delta_i(e) \;=\; X_i(z_0+e)-X_i(z_0-e), \qquad \Delta_i \;=\; \lim_{e\downarrow0}\Delta_i(e), \] whenever the limit exists. Units with $\Delta_i=1$ are compliers, $\Delta_i=-1$ defiers, and $\Delta_i=0$ always- or never-takers. The symbol $\protect\mathpalette{\protect\independenT}{\perp}$ denotes independence.

assumption[RDD with Covariates] (i) $\mathbb{E}[Y_i(0)\mid Z_i=z,W_i]$ is continuous in $z$ at $z_0$ a.s.; (ii) there exists $\bar\epsilon>0$ such that, for all $z\in(z_0-\bar\epsilon,z_0+\bar\epsilon)$, $(\beta_i,X_i(z))\protect\mathpalette{\protect\independenT}{\perp} Z_i\mid W_i$; (iii) $\mathbb{E}[|\beta_i|]<\infty$; (iv) (weak monotonicity) for all $w\in\mathcal W$, either $\mathbb{P}(\Delta_i\ge0\mid W_i=w)=1$ or $\mathbb{P}(\Delta_i\le0\mid W_i=w)=1$.

Assumption (ref) is weaker than the assumptions in the RDD with covariates literature. Importantly, the distribution of $W_i|Z_i=z$ may be discontinuous at the cutoff. In fact, $W_i$ may cause $Y_i(0),$ $Y_i(1),$ $X_i$ and $Z_i$, but may not be caused by them near the cutoff. Additionally, we allow the existence of defiers, provided that there aren't also compliers with the same covariate value.

For any random variable $V_i$, define the conditional limits

equation[equation omitted — 382 chars of source]

and the associated conditional discontinuity

equation[equation omitted — 165 chars of source]

Covariate values observed on both sides arbitrarily close to the cutoff belong to

\[\mathcal W_{z_0}: =\{w\in\mathcal W:\; \forall\,\epsilon>0, \, \mathbb{P}\!\big(W_i=w \,\big|\, Z_i\in(z_0,\, z_0+\epsilon)\big),\mathbb{P}\!\big(W_i=w \,\big|\, Z_i\in(z_0-\epsilon,\, z_0)\big)>0\}.\]

Conditional LATEs and WLATEs

For $w\in\mathcal W$, define the conditional local average treatment effect (conditional LATE)

\[ \beta(w) = \mathbb{E}\!\big[\beta_i\mid W_i=w,\,Z_i=z_0,\,\Delta_i\neq0\big]. \] Let $\omega(W_i)\ge0$ a.s.\ with $\mathbb{E}[\omega(W_i)]=1$, and define the WLATE

\[ \beta_\omega \;=\; \mathbb{E}\!\big[\omega(W_i)\,\beta(W_i)\big]. \]

Henceforth, we assume all moments involved in WLATE definitions are finite.

theorem[Identification of Conditional LATEs and WLATEs] (a) Under Assumption (ref): for any $w\in\mathcal W$, $\delta_Y(w)=\beta(w)\,\delta_X(w)$. (b) The conditional LATE $\beta(w)$ is point identified iff (i) $\delta_X(w)\neq0$ and (ii) $w\in\mathcal W_{z_0}$, in which case $\beta(w)=\delta_Y(w)/\delta_X(w)$. (c) A WLATE $\beta_\omega=\mathbb{E}[\omega(W_i)\beta(W_i)]$ is point identified iff these two conditions hold for all $w$ with $\omega(w)>0$. (d) Moreover, whenever $\beta_\omega$ is identified there exists a measurable $b(W_i)$ with $b(W_i)\delta_X(W_i)\ge0$ a.s.\ and $\mathbb{E}[b(W_i)\delta_X(W_i)]>0$ such that \begin{equation} \beta_\omega=\frac{\mathbb{E}[b(W_i)\,\delta_Y(W_i)]}{\mathbb{E}[b(W_i)\,\delta_X(W_i)]}. \end{equation} The weights, then, can be written as $\omega(W_i)={b(W_i)\,\delta_X(W_i)}/{\mathbb{E}[b(W_i)\,\delta_X(W_i)]}$. Conversely, any ratio of the form (ref) corresponds to an identified WLATE with weights $\omega(W_i)$ as just described.

Equation (ref) resembles an IV-like Wald ratio. In this analogy, $\delta_X(W_i)$ plays the role of the “treatment” variable, $\delta_Y(W_i)$ plays the role of the “outcome” variable, and $b(W_i)$ plays the role of the “instrumental” variable. This is different from a typical IV setting, however, where the conditions are typically satisfied only for a specific IV. In our case, there is no “endogeneity” issue, because the cutoff conditions in Assumption (ref) already guarantee that $\delta_Y(w)=\beta(w)\delta_X(w)$. Thus, any measurable $b(W_i)$ that satisfies $b(W_i)\delta_X(W_i)\geq 0$ a.s. and $\mathbb{E}[b(W_i)\delta_X(W_i)]>0$ can play the role of the “IV” for an identified WLATE.

Note that the WLATEs are averages of treatment effects among compliers at the cutoff. The identification of the $\beta(w)$ relies on Assumption (ref), which is equivalent to the standard RDD conditions, but conditional on $W_i$. Interpretations of WLATEs as averages of treatment effects away from the cutoff would require additional assumptions for extrapolation that are beyond the scope of this paper (see e.g. angrist2015wanna).

Examples

We briefly show how common RDD targets fit into Theorem (ref). Full expressions for weights and “instruments” are given in Table (ref). Let $f_W$ denote the density of $W_i$.

example[Unconditional LATE] The target of the fuzzy RDD with covariates estimator in Calonico_Cattaneo_Farrell_Titiunik, $\beta_U$, is a WLATE under additional conditions provided in that paper (which include strong monotonicity and continuity of the covariate distribution across the cutoff). In our framework, $\beta_U$ can be written as (ref), with $b(W_i)$ proportional to the relative likelihood of $W_i$ at $z_0$.
example[Average Conditional LATE] The simple average $\beta_A=\mathbb{E}[\beta(W_i)]$ is a WLATE with $\omega(W_i)\equiv1$. By Theorem (ref), this requires $\delta_X(w)\neq0$ and $w\in\mathcal W_{z_0}$ for all $w$ in the support of $W_i$. In our framework, it corresponds to (ref) choosing a $b(W_i)$ that underweights high-compliance groups.
example[Counterfactual WLATE] Given a counterfactual covariate distribution $f^\ast$ with support $\mathcal W^\ast\subseteq\mathcal W_{z_0}$, the counterfactual target $\beta_C=\int\beta(w)f^\ast(w)\,dw$ is a WLATE with $\omega(w)=f^\ast(w)/f_W(w)$, identified when $\delta_X(w)\neq0$ for all $w\in\mathcal W^\ast$. In our framework, it corresponds to (ref) with a $b(W_i)$ that switches the measure from $f_W$ to $f^*$ and underweights high-compliance groups.
example[Maximal Average Social Welfare] Following the literature on Empirical Welfare Maximization (see e.g. kitagawa2018should), if a planner treats only cells with nonnegative $\beta(w)$ and strong monotonicity holds, the corresponding WLATE $\beta_S$ is identified when $\delta_X(w)>0$ for all weighted cells. In our framework, it corresponds to (ref) choosing a $b(W_i)$ that drops groups where $\delta_Y(W_i)<0$ and underweights high-compliance groups.
table[table omitted — 2,358 chars of source]

Example (ref) illustrates how large the WLATE class is. Note that one can choose weights such that the $\beta(w)$ are averaged over a distribution outside the cutoff, or even over an artificial distribution that has no relation to the data-generating process at all. All WLATEs must still be interpreted as an average effect at the cutoff $z_0$ because the $\beta(w)$ are local-to-the-cutoff treatment effects, but potentially the WLATE may use a different mix of those observations compared to the mix implied by the actual distribution near the cutoff.

Compliance-Weighted Identification

The identification from the representation in ((ref)) motivates us to maximize the correlation between the “instrument” $b(W_i)$ and the “treatment” $\delta_X(W_i)$. Specifically, consider measurable $b(\cdot)$ with $0<Var(b(W_i))<\infty$ satisfying $b(W_i)\delta_X(W_i)\ge0$ a.s. Let $\rho(b(W_i),\delta_X(W_i))$ denote the correlation between $b(W_i)$ and $\delta_X(W_i)$. We are interested in $b^{\ast}(\cdot),$ where

equation[equation omitted — 214 chars of source]
theorem[Compliance-Weighted “Instrument”] Let Assumption (ref), $\mathbb{P}(\delta_X(W_i)\neq0)>0,$ and $\mathbb{E}\!\big[\delta_X^2(W_i)\big]<\infty$ hold. Then $b^\ast$ solves (ref) for all admissible data generating processes if and only if $b^\ast(w)=c\,\delta_X(w)$ a.s. for some $c>0$.

Thus, in the transformed model of Theorem (ref), the admissible “instrument” that is maximally aligned with the first-stage is proportional to $\delta_X(W_i)$ itself. The associated WLATE equals

equation[equation omitted — 154 chars of source]

By Theorem (ref), the CWLATE is identified when (i) all weighted cells appear on both sides of the cutoff and (ii) $\mathbb{P}(\delta_X(W_i)\neq0)>0$.

While the RDD estimand $\beta_U$ (under strong monotonicity) weights conditional LATEs proportionally to the conditional first-stage at the cutoff, the compliance-weighted LATE $\beta_{CW}$ weights conditional LATEs proportionally to the square of the conditional first-stage in the whole distribution. The estimands differ in two ways. First, because the distribution of $W_i$ at the cutoff may be different from the distribution of $W_i$ in the population. Second, because $\beta_{CW}$ intensifies the tilting towards complier-heavy groups. Note that, if strong monotonicity holds and the first difference above is negligible, the ordering of the two estimands is informative of the direction of the selection on gains/costs of treatment. Specifically, $\beta_{CW}>\beta_U$ if, and only if, higher compliance rates are associated with higher treatment effects. Analogously, $\beta_{CW}<\beta_U$ if, and only if, higher compliance rates are associated with lower treatment effects, and $\beta_{CW}=\beta_U$ if, and only if, no association exists (see formal statement and proof in the end of Appendix (ref)).

WLATEs as Expected Policy Effects

WLATEs admit a policy interpretation under strong monotonicity. Consider a targeted policy where incentive packets are offered to individuals through a two-step process. First, a value of $w$ is drawn according to a policy-assigned probability function $p$. The policymaker can favor certain values $w$ over others by making $p(w)$ relatively higher. Once $w$ is drawn, the incentive is given randomly to someone with $W_i=w$ and $Z_i=z_0$.

The effect of this policy is the WLATE with “instrument” $b(W_i)=p(W_i)/f_{W}(W_i),$ where recall $f_W$ is the density of $W_i$ in the population. In this setting, $b(w)$ denotes the relative targeting “tilt” of the policy in comparison with a policy where no value of $w$ is targeted.

The converse is also true. A WLATE with “instrument” $b(W_i)$ corresponds to the policy where $w$ is drawn from the distribution $p(w)=b(w)f_W(w)/\mathbb{E}[b(W_i)]$. In particular, the CWLATE corresponds to a policy with probabilities given by $p(w)=\delta_X(w)f_W(w)/\mathbb{E}[\delta_X(W_i)],$ which “tilts” incentive-offers towards complier-heavy groups. Contrast this with the standard RDD estimand $\beta_U$, which corresponds to a policy with $p(w)=(f_{W|z_0}^+(w)+f_{W|z_0}^-(w))/2,$ where incentives are given according to the population likelihood of $w$ at the cutoff.

Details of policy interpretation of the WLATEs and formal results are in Online Appendix A.

Estimation of the CWLATE

We estimate the CWLATE in (ref) when $W_i$ has finite support (or is coarsened into finitely many cells). Let $\mathcal W_i=\{w_1,\ldots,w_m\}$, $\pi_j=\mathbb{P}(W_i=w_j)>0$, and $\{Y_i,X_i,Z_i,W_i\}_{i=1}^n$ be i.i.d. Define, for $V\in\{Y_i,X_i\}$,

\[ \mu_V^\pm(w)=\mathbb{E}[V\mid Z_i=z_0^\pm,W_i=w],\qquad \delta_V(w)=\mu_V^+(w)-\mu_V^-(w). \] The CWLATE equals

equation[equation omitted — 184 chars of source]

Relative to the standard fuzzy RDD estimand, CWLATE replaces unconditional discontinuities by an aggregation of cell-specific discontinuities, with weights proportional to squared first-stages.

\paragraph{Plug-in estimator.} Let $\hat\pi_j=n^{-1}\sum_{i=1}^n 1\{W_i=w_j\}$. Estimate $\delta_V(w_j)$ by local polynomials within each cell and set

equation[equation omitted — 206 chars of source]

\paragraph{Local linear implementation (stacked form).} For kernel $k$ and bandwidth $h_n$, define $r_1(z)=(1,z)'$, $\tilde W_i=(1\{W_i=w_1\},\ldots,1\{W_i=w_m\})'$, and $X_{i,1}=r_1(Z_i)\otimes\tilde W_i$, where $\otimes$ is the Kronecker product. Let $k_{h_n}^+(z)=h_n^{-1}k(z/h_n)1\{z\ge0\}$ and $k_{h_n}^-(z)=h_n^{-1}k(z/h_n)1\{z<0\}$. For $s\in\{+,-\}$,

equation[equation omitted — 193 chars of source]

where $e_0=[I_m,0]$ selects the $m$ intercepts. Then $\hat\delta_V(h_n)=\hat\mu_V^+(h_n)-\hat\mu_V^-(h_n)$, and $\hat\delta_V(w_j)$ is its $j$th component. Equivalently, one may run separate local linear regressions within each cell $W_i=w_j$.

\paragraph{Delta-method expansion.} With $\tau_Y,\tau_X$ as in (ref), a delta method yields

equation[equation omitted — 161 chars of source]

where $\hat{\mathcal R}=o_p\big((nh_n)^{-1/2}+h_n^2\big)$ under the regularity conditions in Appendix (ref). CWLATE is less sensitive to weak cells because $\tau_X=\sum_j \pi_j\delta_X^2(w_j)$ downweights small $|\delta_X(w_j)|$ automatically.

Bias-corrected inference for the CWLATE

With local-linear estimation of the cell discontinuities, $\hat\beta_{CW}$ has a leading smoothing bias of order $h_n^2$. Let $b_n$ denote a pilot bandwidth used to estimate this leading bias term via higher-order local polynomials (as done in in Calonico_Cattaneo_Titiunik). Define the debiased CWLATE estimator

equation[equation omitted — 113 chars of source]

where $\hat B_{CW,1,2}(h_n,b_n)$ is a consistent estimator of the leading bias constant (its explicit expression is given in Appendix (ref)).

Bias estimation affects the variance. Let $\mathbf V_{CW,1,2}^{bc}(h_n,b_n)$ denote the RBC asymptotic variance of $\hat\beta_{CW}^{\,bc}$ and $\hat{\mathbf V}_{CW,1,2}^{bc}(h_n,b_n)$ its consistent estimator; both are reported in Appendix (ref).

theorem[RBC inference for the CWLATE] Let the kernel, smoothness, and the RBC bandwidth conditions for $(h_n,b_n)$ stated in Appendix (ref) hold, then \[ \big(\mathbf V_{CW,1,2}^{bc}(h_n,b_n)\big)^{-1/2} \big(\hat\beta_{CW}^{\,bc}-\beta_{CW}\big) \ \xrightarrow{d}\ N(0,1), \] and $\hat{\mathbf V}_{CW,1,2}^{bc}(h_n,b_n)$, given in Appendix (ref), is consistent for $\mathbf V_{CW,1,2}^{bc}(h_n,b_n)$. Hence Wald confidence intervals based on standard normal critical values are asymptotically valid.

Monte Carlo Simulations

We compare the CWLATE estimator to (i) the standard fuzzy RDD and (ii) fuzzy RDD with covariates, using the Stata command RDROBUST (Calonico_Cattaneo_Titiunik).

Simulation Setup

Data are generated through the d.g.p.

align*[align* omitted — 190 chars of source]

where $(U_{X_i},U_{Y_i}) \sim \mathcal{N}((0,0)', (1, .5;.5 , 2))$. The variables $Z_i$ and $W_i$ are independent, with $Z \sim \mathcal{N}(0,2)$, and $W_i$ takes two values, $-1$ and $1$, with equal likelihood. We consider two different values of the parameter that influences the heterogeneity in first-stages across different values of $W_i$, $\alpha_{DW}\in\{0,1\}$, and two different values of the analogous parameter that influences the heterogeneity in treatment effects, $\beta_{XW}\in\{0,2\}$. Note that Assumption (ref) holds in this model. In fact, strong monotonicity holds for the parameter values we consider, so that the standard RDD estimators in the literature, which do not adapt to weak monotonicity, can be used as benchmark in this simulation.

For each $(\alpha_{DW},\beta_{XW})$ and $n\in\{300,\ldots,5000\}$ we run 10{,}000 replications and compute MSE for: the standard RDD $\hat\beta_U$, RDD with covariates $\hat\beta_U^{cov}$, and CWLATE $\hat\beta_{CW}$. We implement the two RDD estimators via RDROBUST with MSE‐optimal bandwidths (Calonico_Cattaneo_Titiunik,Calonico_Cattaneo_Farrell_Titiunik). To isolate weighting as the only difference, CWLATE uses the same bandwidth chosen for $\hat\beta_U^{cov}$ and restricts to observations within that bandwidth, thus differences in estimands cannot be due to differences in $f_W$ and $f_{W|z_0}$.

Figures (ref)--(ref) illustrate typical first‐stage and reduced‐form plots (bins of width $0.067$) for $n=5000$. The first-stage discontinuities are large, so this d.g.p. does not suffer from a weak-IV problem (note that the d.g.p. does not change with the sample size).

figure[figure omitted — 539 chars of source]
figure[figure omitted — 925 chars of source]
figure[figure omitted — 944 chars of source]

Note that this Monte Carlo was designed to favor the performance of standard RDD estimators for the following reasons: (a) we assumed $W_i\protect\mathpalette{\protect\independenT}{\perp} (Z_i,U_{X_i},U_{Y_i})$ and added no extra $Z_i\!\times\!W_i$ interactions; allowing them strengthens CWLATE further; (b) the baseline first-stage is relatively strong; weaker first-stages enlarge CWLATE's advantage, and may even lead to violations of strong monotonicity, which is required for the standard RDD estimators; (c) the bandwidth used for the CWLATE is the optimal one for $\hat{\beta}^{cov}_U$, and thus not necessarily the optimal one for the CWLATE itself.

Results

Table (ref) reports MSEs relative to the corresponding estimand (Unconditional LATE $\beta_U$ for $\hat\beta_U$ and $\hat\beta_U^{cov}$; Compliance-Weighted LATE $\beta_{CW}$ for $\hat\beta_{CW}$).

The main findings are as follows.

enumerate[(i)] • For small samples, the standard RDD estimators display weak‐identification patterns, while CWLATE is stable in all cases. • The first two panels show the edge case where there is no first-stage heterogeneity ($\alpha_{DW}=0$).\footnote{In this case, $\delta_X(W_i)$ is constant, and therefore $\rho(b(W_i),\delta_X(W_i))=0$ for all WLATEs, thus the CWLATE has no advantage in theory. Moreover, note that Theorem (ref) considers only “instruments” with a positive variance, which the CWLATE does not satisfy in this case.} In large samples, all estimators perform very similarly when there is no heterogeneity in treatment effects ($\beta_{XW}=0$), with the standard RDD estimators taking the lead when the heterogeneity in treatment effects increases ($\beta_{XW}=2$). • In the last two panels, the first-stages are heterogeneous ($\alpha_{DW}=1$). Here, Theorem (ref) establishes that the CWLATE has the strongest identification. Unsurprisingly, the CWLATE estimator outperforms the RDD estimators in all cases. In fact, although all estimators perform worse as the treatment effects heterogeneity increases, the comparative advantage of the CWLATE estimator remains (compare the third and fourth panels, and also the third and fourth panels in Table (ref) in Appendix (ref), where $\beta_{XW}=5$ and 10, respectively). This is not surprising, since the result that establishes that the CWLATE has the strongest identification (Theorem (ref)) does not depend on $\beta(w)$.
table[table omitted — 1,536 chars of source]

In Appendix (ref), we consider the scenario where the covariates that modulate the heterogeneity can take many values. We show that the performance of the CWLATE reduces only a little when the researcher uses a coarsened (binarized) version of the covariate, instead of the full covariate.

The results above suggest that standard RDD estimators tend to be more fragile to heterogeneity and to smaller samples than the CWLATE, which is to be expected given Theorem (ref). In Appendix (ref) we illustrate the issues of identification of the standard RDD estimand, $\beta_U,$ by considering the behavior of the three estimators when all of them target $\beta_U$. Here, the CWLATE is an inconsistent estimator of $\beta_U$. Yet, the first two panels in Table (ref) and Figure (ref) show that the variance and weak identification problems of the standard RDD are serious enough that an inconsistent estimator can perform better even when the inconsistency is substantial.

This does not mean that researchers should use the CWLATE as an estimator of $\beta_U$. Using an inconsistent estimator to target $\beta_U$ would result in wrongly-sized tests and wrong-coverage confidence intervals, as we show in the last two panels in Table (ref). Rather, our point is that substantive claims about the effects of treatment on compliers might be better established with different weighted averages that can be estimated with considerably more precision. This point is further illustrated in the next section.

Empirical Application

We consider the study of the effect of women receiving cash transfers while pregnant on the likelihood that their newborns have low birthweight (i.e., less than $2{,}500$ grams), as in amarante2016cash, henceforth AMMV. A monthly cash-transfer policy of around \$100 was implemented in Uruguay from April 2005 to December 2007. Eligibility to receive cash transfers depended on a continuously distributed predicted income score, which was a linear combination of several household characteristics obtained from an interview. After this score was computed, eligibility was determined based on whether it fell below a pre-determined cutoff.\footnote{Neither the interviewers nor the households were informed about the exact formula to compute the predicted score, nor the exact cutoff, mitigating concerns of manipulation around the cutoff.} AMMV restrict the sample to households with a pregnant woman and ask whether the treatment (defined as the household receiving at least one cash transfer during pregnancy) affected the likelihood that the child was born with low birthweight.

Figure (ref) shows the first-stage plot, where it is clear that families just below the cutoff are discontinuously more likely to have received a cash transfer while pregnant than those just above the cutoff.

figure[figure omitted — 344 chars of source]
figure[figure omitted — 541 chars of source]

Figure (ref) shows first-stages for different values of the covariates $W_i$, specifically for mothers aged 16–20 (left panel) and mothers aged 26–30 (right panel). The first-stages are clearly heterogeneous in age, with younger mothers being less likely to be compliers.

Table (ref) reports estimated effects of $X_i$ (receiving cash transfers while pregnant) on $Y_i$ (probability of low birthweight) for a range of bandwidths. Column 1 reports the standard fuzzy RDD Wald estimator based on unconditional discontinuities; Column 2 reports the conventional covariate-adjusted fuzzy RDD estimator (implemented via local regressions with covariates); Column 3 reports the CWLATE estimator, which aggregates covariate-cell reduced-form and first-stage discontinuities using compliance-weighting as in (ref)--(ref). To make the comparison as clean as possible, we align the estimators on the same local window: for each bandwidth choice, the CWLATE is computed using the same observations within that bandwidth, and differs from the covariate-adjusted Wald estimator only through the implied reweighting across covariate cells. As covariates $W_i$, we include interactions between the age of the mother, the education of the father, and the period when the initial visit to determine eligibility occurred.\footnote{Specifically, we consider interactions of the following variables: indicators of each age of the mother from age 16 to 46 (bottom coded and top coded for the two extreme ages), indicators for the education of the father (incomplete primary or less, incomplete secondary, incomplete tertiary, complete tertiary, or missing, which includes lack of presence of father), and indicators of the period when the official visit to the household to obtain the inputs of the income score happened (03/2005 or before, 04/2005--06/2005, 07/2005--09/2005, 10/2005--12/2005, 01/2006 or after). The findings were similar for different breakdowns of these variables.}

table[table omitted — 1,441 chars of source]

The two standard RDD estimates are very similar to those reported in AMMV (e.g. see AMMV's Figure A5), and tend to be insignificant at standard levels. However, because of the low precision of the RDD estimators in this application, AMMV's leading results are obtained using a different identification strategy.\footnote{Specifically, they implement a “localized difference-in-differences” approach with one of the differences being the regression discontinuity in the first column with a large bandwidth, equal to 0.10. For context, the optimal MSE-based bandwidths for both the standard RDD and RDD with covariates in this application are around 0.04.} With this approach, they find an effect of around -0.02 and a standard error of around 0.01.

In contrast to the standard RDD estimates, the CWLATE estimates in the last column tend to be significant at standard levels for most bandwidths.\footnote{To isolate weighting as the only difference, CWLATE uses the same bandwidth chosen for $\hat\beta_U^{cov}$ and restricts to observations within that bandwidth, thus differences in estimands cannot be due to differences in $f_W$ and $f_{W|z_0}$.} The main reason the CWLATE estimates tend to be more significant than the standard RDD estimates is that their standard error tends to be smaller. Nevertheless, because the CWLATE estimates are slightly more negative than the standard RDD estimators with the same bandwidth, there is also some evidence of selection on gains--values of $W_i$ with higher compliance tend to have better (more negative) treatment effects (see Section (ref)).\footnote{These results also allow us to indirectly assess the sign of $\beta_U$ from the sign of $\beta_{CW}$. Under strong monotonicity, if the sign of $\beta(W_i)$ does not change across values of $W_i$, then $\beta_U$ and $\beta_{CW}$ must have the same sign. Under this assumption, our results are evidence that $\beta_U$ is indeed negative, even if the estimates are not significant. This is because our estimates of $\beta_{CW}$ are negative and significant.} Therefore, the CWLATE estimator provides evidence based on a cross-sectional RDD identification strategy that the Compliance-Weighted LATE estimand is negative.

The results in Table (ref) showcase how the CWLATE estimator could be useful in practice. First, it provides a relatively precise estimate of a meaningful weighted average of treatment effects on compliers based on an RDD identification strategy. There is no ex ante reason to accept the Unconditional LATE as a satisfactory estimand and, at the same time, reject the CWLATE in this application. Both estimands consider only observations near the cutoff, and weight positively the exact same covariate groups of compliers. The only difference is that the Unconditional LATE weights by the proportion of compliers while the CWLATE weighs by the square of the proportion of compliers, thus tilting towards the higher compliance groups.

Second, the large standard errors of the standard RDD estimators may lead researchers to make “second best” decisions that effectively would take them farther away from the Unconditional LATE than the use of the CWLATE estimator would. For instance, researchers may use larger-than-optimal bandwidths, as done by AMMV. Note that the vertical differences in Table (ref) (across different bandwidths) are much more pronounced than the horizontal differences (across the different estimators, and different estimands in the case of the CWLATE), suggesting that the choice of bandwidth may be more consequential than the choice between these two estimands in this context. Researchers may go as far as to abandon the RDD identification strategy altogether in favor of an alternative that yields a more precise estimate, where the identification assumptions might be stronger and the estimand may not be the Unconditional LATE at all. For example, this is the case with the difference-in-differences with controls leading estimate chosen by AMMV, which often deviates from desirable target estimands and may not even identify a treatment effect (see caetano2022difference).

Overall, we are able to confirm with an RDD identification strategy the qualitative findings from amarante2016cash that the cash-transfers policy led to a reduction in the incidence of low birthweight among recipients. Our estimates are over twice as large as their estimates, suggesting that cash-transfers contributed to a reduction in the incidence of low-birthweight of the complier population from 10% to about 5–6%.

Conclusions

This paper studies identification and estimation in fuzzy RDDs with covariates when treatment compliance varies across observed subgroups. We characterize a broad class of WLATEs, defined as weighted averages of covariate-specific LATEs at the cutoff, and provide necessary and sufficient conditions for their point identification. Any identified WLATE admits an IV-type representation, where the specific WLATE is determined by the function of the covariates that plays the role of the “instrument.” This representation allows us to single out the Compliance-Weighted LATE (CWLATE) as the WLATE with the strongest “instrument” among all WLATEs. In comparison with the estimand of the standard RDD estimators, the CWLATE puts relatively more weight on groups with more compliers.

We provide a policy interpretation of the WLATEs. We also propose simple plug-in estimators in the case of discrete covariates, derive their asymptotic distribution, and extend robust bias-corrected inference methods to this setting, allowing researchers to implement CWLATE estimation with standard RDD tools and bandwidth selectors.

Monte Carlo simulations confirm the theoretical strength of the CWLATE target, and expose the practical issues with the standard RDD estimators. An application to Uruguay's cash transfer program for women during pregnancy illustrates these issues. While the standard RDD estimators yield inconclusive results, the CWLATE estimator shows that transfers substantially reduced the probability of low birthweight among compliers.

Overall, our results support a simple message for applied work with fuzzy RDDs and covariates: instead of defaulting to the unconditional LATE, researchers can transparently report WLATEs, and in particular the CWLATE, which are identified under standard RDD assumptions and explicitly aligned with where the design is strongest. This perspective preserves the credibility of the RDD while improving empirical content in applications where first-stage variation is unevenly distributed across observed covariates.

\singlespace

\setcounter{page}{1} \pagenumbering{arabic} \doublespacing \setcounter{section}{0} \setcounter{table}{0} \setcounter{figure}{0}

center[center omitted — 36 chars of source]

Policy Interpretation of the WLATEs

We interpret the WLATE estimands using a policy decision framework, and consider the interpretation of the CWLATE in that context.

Suppose that a policymaker targets the offer of a treatment incentive based on the value of covariates, $W_i.$ Each individual who will receive the incentive is selected in a two-step process where first a value $w$ is drawn from a distribution chosen by the policymaker (so that incentives are more or less likely to reach certain groups of the population according to the policymaker's wishes). The incentive is then given to a random individual among those with $W_i=w$ at the cutoff. We describe the policy precisely in the following assumption.

assumptionLet $W_i$ be continuously distributed in the population, with density function $f_W$.\footnote{This assumption is made for simplicity, as we can also allow $W_i$ to not be continuously distributed. For example, if $W_i$ is discrete with support $\mathcal{W}=\{w_1,\dots,w_m\},$ $p$ in Assumption (ref) may be defined as a set of probabilities that a given value of $W_i\in \mathcal{W}$ is drawn: $p_1,\dots,p_m,$ with $\sum_{j=1}^m p_j=1$. All quantities defined henceforth are analogous. For example, let $f_j=\mathbb{P}(W_i=w_j),$ $\bm{P_C}=\sum_{j=1}^m \delta_X(w_j)p_j=\sum_{j=1}^m \left(\frac{p_j}{f_j}\delta(w_j)\right) f_j$. Extension of results to the discrete case is straightforward.} The policymaker observes the value of $W_i$ for all individuals in the population. A treatment incentive is offered to some of these individuals. Each individual who receives the incentive is selected as follows.\begin{itemize} • Step 1: A value $w$ is drawn from a distribution $p,$ where $p(w)\geq0,$ $\int p(w)dw=1,$ $\int p(w)^2dw<\infty$ and, if $f_W(w)=0,$ then $p(w)=0$. • Step 2: The incentive is offered to an individual who is randomly drawn among those with $W_i=w$ and $Z_i=z_0$. \end{itemize} We further assume that \begin{enumerate}[(i)] • Nobody offered the incentive decreases the amount of treatment. • A sample of vectors $(Y_i,X_i,Z_i,W_i')'$ which is representative of the population is available, where identical incentives to those intended by the policy were offered to those with $Z_i\geq z_0$. \end{enumerate}

Let $A\in \mathcal{W}$ be a $p$-measurable set. The policy described above assures that roughly a proportion $\int_A p(w)dw$ of the incentives are allocated for individuals with $W_i\in A$. If one were to assign incentives randomly among the population, this likelihood would be $\int_A f_W(w)dw.$ Therefore, $p(w)/f_W(w)$ represents the relative “tilt” of the policy in targeting certain values of $w$ over others. Note that Step 1 guarantees that no incentives may be allocated outside the support of the distribution of $W_i$ (thus, for completeness, define $p(w)/f_W(w)=0$ whenever $f_W(w)=0$).

Since the incentive units assigned to those with $W_i=w$ and $Z_i=z_0$ are distributed randomly in that group, some of the incentives may be offered to individuals that will be treated anyway (always takers), and some of the incentives may be offered to individuals that will not take the treatment anyway (never takers). These individuals are not affected by the policy, since their treatment status will be the same whether or not they are offered the incentive. The policy only affects those who change their treatment status as a consequence of receiving the incentive. Assumption (ref)(ref) rules out the existence of defiers by assuming strong monotonicity. It is feasible, but more cumbersome, to devise a policy interpretation of the WLATEs under weak monotonicity.

Assumption (ref)(ref) makes it possible to tie the intended policy to the existing data. Precisely, while the data is generated from an RDD quasi-experiment, the policy would be applied in a population using the exact procedure described in Assumption (ref). Therefore, to identify the effects of the described policy, the data must be representative of the policy population, and the incentives considered by the policy must be identical to the incentives faced by the observations above the cutoff in the data.

If Assumption (ref) holds, then per incentive given, the probability it reaches a complier, denoted $\textbf{P}_\textbf{C},$ satisfies

equation*[equation* omitted — 136 chars of source]

The first equality follows because $p(w)$ is the probability that $w$ is selected, and $\delta_X(w)$ is the probability that a random individual with $W_i=w$ and $Z_i=z_0$ is a complier.

Since the incentive has no effect on always takers and never takers, the expected effect of each incentive offered, denoted $\textbf{APE}$ (for average policy effect per incentive unit) is

equation*[equation* omitted — 152 chars of source]

The expected effect per incentive unit on the compliers that are incentivized is the LAPE, where L stands for “local” to compliers. It is given by

equation*[equation* omitted — 207 chars of source]

The following result is an immediate consequence of Theorem (ref).

theoremIf Assumptions (ref) and (ref) hold, then the APE and the LAPE are identified if, and only if, $p(W_i)$ is identified and the conditions of Theorem (ref) hold a.s.\ when $p(W_i)\delta_X(W_i)>0$. The LAPE is a WLATE with weights \begin{equation*} \omega(W_i)=\frac{p(W_i)\delta_X(W_i)/f_W(W_i)}{E\left[p(W_i)\delta_X(W_i)/f_W(W_i)\right]}, \end{equation*} and “instrument” \begin{equation*} b(W_i)= c\cdot p(W_i)/f_W(W_i), \end{equation*} where $c>0$ is a constant. Conversely, a WLATE with “instrument” $W_i$ is a LAPE with $p(w)=b(w)f_W(w)/\mathbb{E}[b(W_i)]$.

This result informs us that every identifiable WLATE with “instrument” $b$ can be interpreted as the expected per-incentive unit effect of a policy of the type described in Assumption (ref) among those who are affected by it. For each incentive unit offered, the policy consists on drawing a value $w$ from the distribution function $p(w)=b(w)f_W(w)/E[b(W_i)],$ and giving the incentive to a random individual with $W_i=w$ and $Z_i=z_0$.

By this reasoning, the estimands discussed in the examples in Section (ref) can be interpreted as the expected effect among affected compliers per unit of incentive, where the incentives are distributed according to specific targeted policies. The corresponding targeting distribution $p$ is as follows.

example[Unconditional LATE, cont.] $\beta_U$ (Example (ref)) is the LAPE of a policy with \begin{equation*} p(w)=f_{W|z_{0}}(w). \end{equation*}

The Unconditional LATE estimand $\beta_U$ is well known to correspond to the average effect among compliers of a policy where incentives are distributed entirely randomly among those at the cutoff. This “random” policy can be understood equivalently as a “targeted” policy such as the one described in Assumption (ref) where the values of $w$ are drawn from the population distribution of $W_i$ at the cutoff.

example[Average LATE, cont.] $\beta_A$ (Example (ref)) is the LAPE of a policy with \begin{equation*} p(w)=\frac{f_W(w)}{\delta_X(w) E[\delta_X(W_i)^{-1}]}. \end{equation*}
example[Counterfactual Average LATE, cont.] $\beta_C$ (Example (ref)) is the LAPE of a policy with \begin{equation*} p(w)=\frac{f^*(w)}{\delta_X(w) E[\delta_X(W_i)^{-1}|W_i\in \mathcal{W}^*]}. \end{equation*}
example[Maximal Average Social Welfare Gain, cont.] $\beta_S$ (Example (ref)) is the LAPE of a policy with \begin{equation*} p(w)=\frac{f_W(w)}{\delta_X(W_i) E[\delta_X(W_i)^{-1}|\delta_Y(W_i)\geq 0]}\cdot\frac{1(\delta_Y(W_i)\geq0)}{\mathbb{P}(\delta_Y(W_i)\geq 0)}. \end{equation*}

The policies implied by the estimands $\beta_A,$ $\beta_C$ and $\beta_S$ tilt away from a random distribution in that they are less likely to give incentives to the groups with more compliers (i.e. when $\delta_X(W_i)$ is relatively large). $\beta_C$ also tilts in favor of $W_i$ values that are proportionally more likely in the counterfactual distribution than in the actual distribution. Note that $\beta_S$ only draws from groups with non-negative conditional LATEs.

example[Compliance-Weighted LATE] $\beta_{CW}$ (Section (ref)) is the LAPE of a policy with \begin{equation*} p(w)=\frac{\delta_X(w)f_W(w)}{ E[\delta_X(W_i)]}. \end{equation*}

The policy implied by the CWLATE estimand, $\beta_{CW},$ tilts from a random distribution in giving more incentives to groups with a larger proportion of compliers.

Policymakers might have a variety of different goals when distributing incentive units to individuals. For instance, a policymaker who wants to reduce income inequality might want to tilt an incentive policy towards values of $W_i$ that are associated with poor families. Moreover, the policy decision may depend on contrasting expected effects per incentive unit (APEs) or expected effects on affected compliers per incentive unit (LAPEs) of several policies at the same time, taking into account targeting parameters (e.g., the income cutoff to receive the treatment) and other criteria (e.g., the policymaker may want to take into account the policy cost in the decision). Theorem (ref) determines the range of policies for which the effect can be identified, and gives the formula of these effects.

Identification Proofs

proof[Proof of Theorem (ref)] (a) Using $Y_i = Y_i(0)+\beta_i X_i$, \begin{align*} \delta_Y(w) &= \mathbb{E}[Y_i\mid Z_i=z_0^+,W_i=w]-\mathbb{E}[Y_i\mid Z_i=z_0^-,W_i=w]\\ &= \underbrace{\mathbb{E}[Y_i(0)\mid Z_i=z_0^+,W_i=w]-\mathbb{E}[Y_i(0)\mid Z_i=z_0^-,W_i=w]}_{=0} \\[-0.2em] &\quad + \lim_{e\downarrow0}\Big\{ \mathbb{E}[\beta_i X_i\mid Z_i=z_0+e,W_i=w] -\mathbb{E}[\beta_i X_i\mid Z_i=z_0-e,W_i=w] \Big\}. \end{align*} By Assumption \ref{ass:rdd}(i) the $Y_i(0)$ term is continuous at $z_0$. For $e>0$, \[ \mathbb{E}[\beta_i X_i\mid Z_i=z_0+e,W_i=w] -\mathbb{E}[\beta_i X_i\mid Z_i=z_0-e,W_i=w] = \mathbb{E}[\beta_i \Delta_i(e)\mid Z_i=z_0,W_i=w], \] where $\Delta_i(e)=X_i(z_0+e)-X_i(z_0-e)$. Assumption (ref)(ii) and (iii), plus dominated convergence, imply $\lim_{e\downarrow0}\mathbb{E}[\beta_i \Delta_i(e)\mid Z_i=z_0,W_i=w] = \mathbb{E}[\beta_i\Delta_i\mid Z_i=z_0,W_i=w]$, where $\Delta_i=\lim_{e\downarrow0}\Delta_i(e)$. By weak monotonicity (iv), at each $w$ compliers and defiers do not coexist, so \[ \mathbb{E}[\beta_i\Delta_i\mid Z_i=z_0,W_i=w] = \mathbb{E}[\beta_i\mid Z_i=z_0,W_i=w,\Delta_i\neq0]\, \mathbb{E}[\Delta_i\mid Z_i=z_0,W_i=w]. \] By definition, $\beta(w)=\mathbb{E}[\beta_i\mid Z_i=z_0,W_i=w,\Delta_i\neq0]$ and $\delta_X(w)=\mathbb{E}[\Delta_i\mid Z_i=z_0,W_i=w]$, hence $\delta_Y(w)=\beta(w)\delta_X(w)$. (b) Sufficiency. If $\delta_X(w)\neq0$ and $w\in\mathcal W_{z_0}$, then $\mathbb{E}[Y_i\mid Z_i=z,W_i=w]$ and $\mathbb{E}[X_i\mid Z_i=z,W_i=w]$ are identified for $z\to z_0^\pm$, so $\delta_Y(w)$ and $\delta_X(w)$ are identified and $\beta(w)=\delta_Y(w)/\delta_X(w)$. Necessity. If $\delta_X(w)=0$, then for any $\phi$, $\tilde\beta(w)=\beta(w)+\phi$ satisfies $\tilde\beta(w)\delta_X(w)=\beta(w)\delta_X(w)=\delta_Y(w)$, so $\beta(w)$ is not unique. If $w\notin\mathcal W_{z_0}$, then $w$ has no support arbitrarily close to one side of the cutoff, so at least one of $\delta_Y(w),\delta_X(w)$ is not identified, and any $\beta(w)$ consistent with $\delta_Y(w)=\beta(w)\delta_X(w)$ is observationally equivalent. Hence conditions (i)–(ii) are necessary and sufficient. (c) Let $\omega\ge0$ a.s.\ with $\mathbb{E}[\omega]=1$ and $\beta_\omega=\mathbb{E}[\omega(W_i)\beta(W_i)]$. Sufficiency. If (i)–(ii) in (b) hold for all $w$ with $\omega(w)>0$, then each such $\beta(w)$ is identified, so $\beta_\omega$ is identified. \emph{Necessity.} Suppose there exists $w^\ast$ with $\omega(w^\ast)>0$ where either $\delta_X(w^\ast)=0$ or $w^\ast\notin\mathcal W_{z_0}$. By (b) we can change $\beta(w^\ast)$ without affecting observables, obtaining $\tilde\beta$ such that \[ \tilde\beta_\omega = \mathbb{E}[\omega(W_i)\tilde\beta(W_i)] = \beta_\omega + \phi\,\omega(w^\ast) \] for arbitrary $\phi$, so $\beta_\omega$ is not point identified. Thus the stated condition is also necessary. \textbf{(d)} \emph{From WLATE to IV form.} If $\beta_\omega$ is identified, then by (c) for all $w$ with $\omega(w)>0$, $\delta_X(w)\neq0$ and $w\in\mathcal W_{z_0}$. Define \[ b(w) = \begin{cases} c\,\dfrac{\omega(w)}{\delta_X(w)}, & \omega(w)>0,\\[0.3em] 0, & \omega(w)=0, \end{cases} \] with $c>0$ fixed. Then $b(w)\delta_X(w)=c\,\omega(w)\ge0$ and $\mathbb{E}[b(W_i)\delta_X(W_i)]=c>0$, so \[ \omega(W_i) = \frac{b(W_i)\delta_X(W_i)}{\mathbb{E}[b(W_i)\delta_X(W_i)]}. \] Using part (a), $\beta(W_i)\delta_X(W_i)=\delta_Y(W_i)$ wherever relevant, hence \[ \beta_\omega = \mathbb{E}[\omega(W_i)\beta(W_i)] = \frac{\mathbb{E}[b(W_i)\delta_Y(W_i)]}{\mathbb{E}[b(W_i)\delta_X(W_i)]}, \] which yields (ref). \emph{From IV form to WLATE.} Conversely, let $b(W_i)$ be identified with $b(W_i)\delta_X(W_i)\ge0$ a.s.\ and $\mathbb{E}[b(W_i)\delta_X(W_i)]>0$, and assume $\delta_Y(W_i),\delta_X(W_i)$ are identified where $b(W_i)\delta_X(W_i)>0$. Define \[ \omega(W_i) = \frac{b(W_i)\delta_X(W_i)}{\mathbb{E}[b(W_i)\delta_X(W_i)]}, \] so $\omega\ge0$ and $\mathbb{E}[\omega]=1$. By (a), \[ \frac{\mathbb{E}[b(W_i)\delta_Y(W_i)]}{\mathbb{E}[b(W_i)\delta_X(W_i)]} = \mathbb{E}[\omega(W_i)\beta(W_i)] = \beta_\omega. \] Thus any such ratio corresponds to an identified WLATE with weights $\omega$.
proof[Proof of Theorem (ref)] We restate the problem: \[ b^{\ast} = \arg\max_b\; |\rho(b(W_i),\delta_X(W_i))| \quad\text{s.t.}\quad b(W_i)\delta_X(W_i)\ge0\ \text{a.s.}, \] where the maximization is over measurable $b(\cdot)$ with $0<Var(b(W_i))<\infty$ and $\mathbb{P}(\delta_X(W_i)\neq0)>0$. Because $\rho$ is a correlation coefficient, we have $|\rho(b(W_i),\delta_X(W_i))|\le1$, with equality if and only if there exist constants $c_0\in\mathbb R$ and $c_1\neq0$ such that \[ b(W_i)=c_0 + c_1\delta_X(W_i) \quad\text{a.s.} \] Now impose the sign restriction $b(W_i)\delta_X(W_i)\ge0$ a.s.: \[ b(W_i)\delta_X(W_i) = c_0\delta_X(W_i) + c_1\delta_X(W_i)^2 \ge0\quad\text{a.s.} \] We interpret this restriction as a structural requirement: for any data generating process satisfying Assumption (ref) and $\mathbb{P}(\delta_X(W_i)\neq0)>0$, the resulting $b(\cdot)$ must satisfy $b(W_i)\delta_X(W_i)\ge0$ a.s.\ under that process. Under Assumption (ref)(iv), for each $w$ the sign of $\delta_X(w)$ is well-defined (weak monotonicity), but across $w$ both positive and negative values of $\delta_X(w)$ are admissible. Hence we must allow for designs in which $\delta_X(W_i)$ takes both positive and negative values. If $c_0\neq0$, then for some admissible design we can choose values of $\delta_X(W_i)$ with opposite signs and sufficiently small magnitude so that $c_0\delta_X(W_i)$ changes sign while $c_1\delta_X(W_i)^2$ is arbitrarily small, violating $b(W_i)\delta_X(W_i)\ge0$ with positive probability. Therefore, to satisfy the sign restriction for all admissible designs we must have $c_0=0$. Given $\mathbb{P}(\delta_X(W_i)\neq0)>0$, the inequality with $c_0=0$ implies $c_1>0$. Thus any admissible maximizer must be of the form \[ b^{\ast}(W_i)=c\,\delta_X(W_i),\qquad c>0. \] Conversely, for any $c>0$, the choice $b(W_i)=c\,\delta_X(W_i)$ satisfies $b(W_i)\delta_X(W_i)\ge0$ a.s.\ and yields $|\rho(b(W_i),\delta_X(W_i))|=1$, so it attains the maximum of the optimization problem. Hence $b^{\ast}$ solves (ref) if and only if $b^{\ast}(w)=c\,\delta_X(w)$ with $c>0$.

{\fontsize{12.5}{13.2}\selectfont Ordering of $\bm{\beta_U}$ and $\bm{\beta_{CW}}$

propositionSuppose that strong monotonicity holds, $f_{W|z_0}(w)=f_W(w)$ a.s., and that $\beta_U$ and $\beta_{CW}$ are identified. Then $\rho(\delta_X(W_i),\beta(W_i))$ and $\beta_{CW}-\beta_U$ have the same sign.
proofDefine a new probability measure $Q$ such that for any random variable $V$, the expectation under $Q$ is weighted by $\delta_X(W_i)$:$$E_Q[V] = \frac{E[\delta_X(W_i)V]}{E[\delta_X(W_i)]}$$ Under this measure, we can rewrite $\beta_U = E_Q[\beta(W_i)]$ and $\beta_{CW} = \frac{E_Q[\delta_X(W_i)\beta(W_i)]}{E_Q[\delta_X(W_i)]}$. By the definition of covariance under measure $Q$:$$E_Q[\delta_X(W_i)\beta(W_i)] - E_Q[\delta_X(W_i)]E_Q[\beta(W_i)] = \text{cov}_Q(\delta_X(W_i), \beta(W_i)).$$ Since $E_Q[\delta_X(W_i)] > 0$, we divide both sides by it, and obtain $$\beta_{CW}-\beta_U=\frac{\text{cov}_Q(\delta_X(W_i), \beta(W_i))}{E_Q[\delta_X(W_i)]}.$$ The result then follows because the sign of the left-hand side must be the same as the sign of $\rho(\delta_X(W_i),\beta(W_i))$.

Asymptotic Theory

Point Estimation and Asymptotic Theory for WLATEs

In this section, we develop point estimation and asymptotic theory for WLATEs. For estimation, we consider the practical situation where $W_i$ takes on a finite set of values. Let $\mathcal{W}=\{w_{1},...,w_{m}\}$ denote the support of $W,$ with $\pi_{j}=\mathbb{P}\left( W_i=w_{j}\right) >0.$ We assume we observe an independent and identically distributed (iid) random sample $\{Y_{i},X_{i},Z_{i},W_{i}\}_{i=1}^{n}$ of size $n.$ Let $\{Y_i,X_i,Z_i,W_i\}$ be a random vector with the same distribution as $\{Y_{i},X_{i},Z_{i} ,W_{i}\}.$ For a generic random vector $V,$ we use the notation

align*[align* omitted — 115 chars of source]

$\mu_{V}^{+}=(\mu_{V}^{+\prime}(w_{1}),...,\mu_{V}^{+\prime}(w_{m}))^{\prime },$ $\mu_{V}^{-}=(\mu_{V}^{-\prime}(w_{1}),...,\mu_{V}^{-\prime}(w_{m}))^{\prime}$ and $\delta_{V}=\mu_{V}^{+}-\mu_{V}^{-}.$ When $V$ is a function of $W,$ our notation is such that $\mu_{V}^{+}$ and $\mu_{V}^{-}$ correspond to the standard RDD quantities that do not condition on $W_i.$ For example, for $V=W_i,$ $ \mu_{V}^{+} =\mathbb{E}[W_i|Z_i=z_0^+],$ and $ \mu_{V}^{-} =\mathbb{E}[W_i|Z_i=z_0^-].$ We follow similar vector notation for functions of $w$. For example $\pi=(\pi_{1},...,\pi _{m})^{\prime}$. Without loss of generality, assume hereinafter that $z_{0}=0.$

Define $c=(c_{1},...,c_{m})^{\prime}$ with $c_{j}=\pi _{j}b(w_{j}),$ for $j=1,...,m.$ With this notation in place, the generic WLATE estimand is \[ \beta_{\omega}=\frac{\sum_{j=1}^{m}\pi_{j}b(w_{j})\delta_{Y}(w_{j})} {\sum_{j=1}^{m}\pi_{j}b(w_{j})\delta_{X}(w_{j})}=\frac{c^{\prime}\delta_{Y} }{c^{\prime}\delta_{X}}\equiv\frac{\tau_{Y}}{\tau_{X}}, \] where the dependence of the $\tau^{\prime}s$ on $c$ is dropped for simplicity. To account for the different examples mentioned above, we assume that $c$ is generated according to the following specification

equation[equation omitted — 58 chars of source]

for a known vector-valued function $g,$ and different choices of $V.$ When $g$ depends on $\mu_{V}^{+}$ and $\mu_{V}^{-}\ $only through the term $\mu_{V} ^{+}-\mu_{V}^{-},$ we write $c=g(\delta_{V},\pi).$

The proposed estimator of $\beta_{\omega}$ in this setting is \[ \hat{\beta}_{\omega}=\frac{\hat{\tau}_{Y}}{\hat{\tau}_{X}}, \] where $\hat{\tau}_{Y}=\hat{c}^{\prime}\hat{\delta}_{Y},$ $\hat{\tau}_{X} =\hat{c}^{\prime}\hat{\delta}_{X},$ $\hat{c}=g_{j}(\hat{\mu}_{V}^{+},\hat{\mu }_{V}^{-},\hat{\pi}),$ $\hat{\delta}_{V}=\hat{\mu}_{V}^{+}-\hat{\mu}_{V}^{-},$ $\hat{\pi}$ is the vector of sample probabilities, and $\left( \hat{\mu} _{V}^{+},\hat{\mu}_{V}^{-}\right) $ are local polynomial estimates. In the main text we present results for local-linear, but our asymptotic results in the Appendix allow for local-polynomial estimates of any order. To introduce the local-linear estimators, define for a generic random variable $V,$ kernel function $k$ and bandwidth $h_{n}$ the quantities

align[align omitted — 632 chars of source]

where $X_{i,1}=r_{1}(Z_{i})\otimes\tilde{W}_{i},$ $\tilde{W}_{i} =(W_{1i},...,W_{mi})^{\prime}$, $W_{ij}=1(W_{i}=w_{j})$, $r_{1}(z)=(1,z)^{\prime},$ $k_{h_{n}} ^{+}(z)=h_{n}^{-1}k(z/h_{n})1(z\geq0),$ $k_{h_{n}}^{-}(z)=h_{n}^{-1} k(z/h_{n})1(z<0),$ and $e_{0}$ is the conformable $m\times2m$ matrix $e_{0}=[I_{m},0]$ with $I_{m}$ the identity matrix of order $m$. Of course, the same estimator can be obtained by running separate local linear regressions over subsamples defined by covariates. The benefit of our approach in ((ref)) and ((ref)) is that simultaneous inference on $\beta(w)$ for different $w^{\prime}s$ can be constructed, based on the estimator $\hat{\beta}(w_{j})=\hat{\delta}_{Y}(w_{j})/\hat{\delta}_{X} (w_{j}),$ for $j=1,...,m,$ where $\hat{\delta}_{V}(w_{j})$ is the $j$-th component of $\hat{\delta}_{V}.$ Each of $\hat{\beta}(w_{j})$ is a standard local linear RDD estimator. However, we are particularly interested in the common situation where joint inference on $(\beta(w_{1}),...,\beta (w_{m}))^{\prime}$ is not reliable because $\delta_{X}(w_{j})$ for some $j$ is close to zero or sample sizes for some class $j$ are not large enough for inference to be reliable with nonparametric methods.

A Delta Method argument leads to the expansion

equation[equation omitted — 199 chars of source]

where the remainder term is given by

equation[equation omitted — 245 chars of source]

This delta method expansion is known to be sensitive to a small \textquotedblleft first-stage\textquotedblright\ $\tau_{X},$ which motivates our robust estimand. Indeed, a small $\tau_{X}$ makes the variance of the leading term in ((ref)) large, and introduces also a nonlinearity bias through the remainder term $\hat{\mathcal{R}}.$

In our setting where $\hat{\tau}_{V}$ involves two nonparametric estimators ($\hat{c}$ and $\hat{\delta}_{V}$), ((ref)) is not a linearization in terms of the more primitive objects $\hat{\mu}_{V}^{+},$ $\hat{\mu}_{V} ^{-}$ and $\hat{\delta}_{V}.$ Thus, a further linearization is required, which in turn may lead to further problems regarding close-to-zero denominators.

In this section, we establish the asymptotic theory for the proposed estimator $\hat{\beta}_{\omega}.$ We introduce some further notation and assumptions. Let $\varepsilon_{V_{i}}=V_{i}-\mathbb{E}[V_{i}|Z_{i},W_{i}]$ denote the regression errors for $V=Y_i$ and $V=X_i.$ In the proofs we use the generic notation, for a generic $V_{i},$ $j,g=1,...,m,$ \[ \mu_{jg}^{V}(z)=\mathbb{E}[W_{ij}W_{ig}V_{i}|Z_{i}=z]\text{ and }\sigma_{jg} ^{2,V}(z)=\mathbb{E}[W_{ij}W_{ig}V_{i}^{2}|Z_{i}=z] \] When $V$ is identically $1,$ we drop the reference to $V$ above.

We set $\mathcal{Y}_{n}=(Y_{1},...,Y_{n})^{\prime},$ $\mathcal{Z}_{n} =(Z_{1},...,Z_{n})^{\prime},$ $\mathcal{W}_{n}=(\tilde{W}_{1}^{\prime },...,\tilde{W}_{n}^{\prime}),$ $\mathcal{V}_{n}=(V_{1}^{\prime} ,...,V_{n}^{\prime})^{\prime},$ and $\mathcal{S}_{n}=\{\mathcal{Z} _{n},\mathcal{W}_{n}\}.$ Define the vector $\varepsilon_{V}=(\varepsilon _{V_{1}},...,\varepsilon_{V_{n}})^{\prime}$ and $\Sigma_{UV}=\mathbb{E}[\varepsilon _{U}\varepsilon_{V}^{\prime}|\mathcal{S}_{n}]=diag(\sigma_{UV}^{2} (Z_{1},\tilde{W}_{1}),\dots$ $\dots,\sigma_{UV}^{2}(Z_{n},\tilde{W}_{n})),$ with

align*[align* omitted — 312 chars of source]

With some abuse of notation, we denote $\mu_{W}(z)=\mathbb{E}[\tilde{W}_{i}|Z_{i}=z]$ and $\sigma_{W}^{2}(z)=\mathbb{E}[\tilde{W}_{i}\tilde{W}_{i}^{\prime}|Z_{i}=z].$

Consider the local right-hand side approximation

align*[align* omitted — 354 chars of source]

where $\otimes$ denotes Kronecker product, $\beta_{V,p}^{+}=(\mu_{V}^{+} {}^{\prime},\mu_{V}^{+(1)\prime},...,\mu_{V}^{+(p)\prime}/p!),$ and \[ \mu_{V}^{+(s)}\equiv\mu_{V}^{+(s)}(0)=\lim_{z\downarrow0}\frac{\partial^{s} \mu_{V}(z)}{\partial z^{s}}. \] Implicit in the notation above is that $\mu_{V}(z)=(\mu_{1V}(z),...,\mu _{mV}(z)),$ where \[ \mu_{jV}(z)=\mathbb{E}[V_{i}|Z_{i}=z,W_{i}=w_{j}]. \] We use the analogous notation for left hand side approximations and derivatives (with $-$ replacing $+)$.

We investigate the asymptotic properties of $\hat{\beta}_{\omega}$ under the following assumptions, which parallel those of HTV2001:

assumptionSuppose that with $z$ in a neighborhood of zero: \begin{enumerate} • The sample $\{\chi_{i}\}_{i=1}^{n}$ is an iid sample, where $\chi _{i}=(Y_{i},X_{i}^{\prime},W_{i}^{\prime},Z_{i})^{\prime}.$ • (i) the density of $Z_i,$ $f(z),$ is continuous, bounded and bounded away from zero; (ii) $\mathbb{E}[Y_{i}^{4}|Z_{i}=z,W_{i}=w]$ and $\sigma_{UV}^{2} (z,\tilde{w})$ are bounded. The matrices $\Gamma_{+,p}$ and $\Gamma_{-,p}$, defined below, are positive definite, and $\tau_{X}\neq0.$ • The kernel $k$ is continuous, symmetric and nonnegative-valued with compact support. • The functions $\mu_{jg}(z),$ $\sigma_{jg}^{2}(z),$ $\sigma_{UV}^{2}(z),$ $\mu_{jg}^{\varepsilon_{Y}\varepsilon_{X}}(z),$ $\sigma_{jg}^{2,\varepsilon _{Y}\varepsilon_{X}}(z)$ are uniformly bounded$,$ with well-defined and finite left and right limits to $z=0,$ for $j,g=1,...,m.$ • The bandwidth satisfies $nh_{n}^{2p+5}\rightarrow0$ and $nh_{n} \rightarrow\infty.$ • For $V_{i}=Y_{i}$ and $V_{i}=X_{i}$, for $z>0$ or $z<0$, and all $j:$ (i) $\mathbb{E}[V_{i}|Z_{i}=z,W_{i}=w_{j}]$ is $d$ times continuously differentiable, $d\geq p+2$; (ii) $Var[V_{i}|Z_{i}=z,W_{i}=w_{j}]$ are continuous in $z$ and bounded away from zero$.$$c=g(\mu_{V}^{+},\mu_{V}^{-},\pi),$ with $g$ continuously differentiable in its components at the true values $(\mu_{V}^{+},\mu_{V}^{-},\pi).$ When $V=W_{ij},$ for $z>0$ or $z<0,$ $\mathbb{E}[W_{ij}|Z_{i}=z]$ is $d$ times continuously differentiable, $d\geq p+2,$ and bounded away from zero and one. • The sequences $h_n$ and $b_n$ satisfy $n\min\{h_{n}^{5},b_{n}^{5}\}\max\{h_{n}^{2},b_{n}^{2} \}\rightarrow0$ and $n\min\{h_{n},b_{n}\}\rightarrow\infty.$ \end{enumerate}

To introduce the estimators, define for a generic random variable $V,$ degree $p,$ $0\leq v\leq p,$ and bandwidth $h_{n}$ the quantities

align*[align* omitted — 269 chars of source]
align*[align* omitted — 281 chars of source]

where $\hat{\mu}_{V,p}^{+}(h_{n})=\hat{\mu}_{V,p}^{+(0)}(h_{n}),$ $\hat{\mu }_{V,p}^{-}(h_{n})=\hat{\mu}_{V,p}^{-(0)}(h_{n}),$ $X_{i,p}=r_{p} (Z_{i})\otimes\tilde{W}_{i},$ $r_{p}(z)=(1,z,...,z^{p})^{\prime},$ $k_{h_{n} }^{+}(z)=h_{n}^{-1}k(z/h_{n})1(z\geq0),$ $k_{h_{n}}^{-}(z)=h_{n}^{-1} k(z/h_{n})1(z<0),$ and $e_{v}$ is a conformable $m\times m(1+p)$ matrix, that selects the $m$ elements corresponding to the $v$-th derivative. For example, $e_{0}=[I_{m},0...,0]$.

Define the matrices

align*[align* omitted — 739 chars of source]

When $U=V,$ $p=q$ or $h=b$ in the above expressions one of the terms is drop, so we denote, for example, $\Sigma_{U}=\Sigma_{UU}$ and $\Psi_{U,p} ^{+}(h)=\Psi_{UU,p,p}^{+}(h,h).$ Letting \[ H_{p}(h)=diag(1,...,1,h^{-1},...,h^{-1},...,h^{-p},...,h^{-p}), \] where each element is repeated $m$ times, it follows that

align*[align* omitted — 247 chars of source]

Define

align*[align* omitted — 225 chars of source]

and

align*[align* omitted — 447 chars of source]
lemmaUnder Assumption (ref)(1)-(6), \[ \Gamma_{+,p}(h_{n})\rightarrow_{p}\Gamma_{+,p}, \] where \[ \Gamma_{+,p}=f(0)\left( \Gamma_{p}^{+}\otimes\sigma_{W}^{2+}\right) . \] \begin{proof} Using that $X_{i,p}X_{i,p}^{\prime}=r_{p}(Z_{i}/h_{n})r_{p}^{\prime} (Z_{i}/h_{n})\otimes\tilde{W}_{i}\tilde{W}_{i}^{\prime}$ and by the change of variables $u=Z/h_{n},$ \begin{align*} \mathbb{E}[\Gamma_{+,p}(h_{n})] & =\mathbb{E}\left[ \frac{1}{n}\sum_{i=1}^{n}X_{i,p} X_{i,p}^{\prime}k_{ih_{n}}^{+}\right] \\ & =\mathbb{E}\left[ \left( r_{p}(Z_{i}/h_{n})r_{p}^{\prime}(Z_{i}/h_{n} )\otimes\sigma_{W}^{2}(Z_{i})\right) k_{ih_{n}}^{+}\right] \\ & =\int_{0}^{\infty}\left( r_{p}(u)r_{p}^{\prime}(u)\otimes\sigma_{W} ^{2}(uh_{n})\right) k(u)f(uh_{n})du\\ & \equiv\tilde{\Gamma}_{+,p}(h_{n}). \end{align*} By Assumptions (ref)(2) and (ref)(4) and Dominated Convergence theorem \[ \tilde{\Gamma}_{+,p}(h_{n})=\Gamma_{+,p}+o(1), \] where \[ \Gamma_{+,p}=f(0)\int_{0}^{\infty}\left( r_{p}(u)r_{p}^{\prime} (u)\otimes\sigma_{W}^{2+}\right) k(u)du. \] Let \[ \tau_{ljg}^{+}=\frac{1}{n}\sum_{i=1}^{n}\left( \frac{Z_{i}}{h_{n}}\right) ^{l}W_{ij}W_{ig}k_{ih_{n}}^{+},\qquad l=0,1,...,2p,j,g=1,...,m. \] Then, \begin{align*} Var(\tau_{ljg}^{+}) & \leq n^{-1}\mathbb{E}\left[ \left( \frac{Z_{i}}{h_{n} }\right) ^{2l}W_{ij}^{2}W_{ig}^{2}k_{ih_{n}}^{+2}\right] \\ & =\left( nh_{n}\right) ^{-1}\int_{0}^{\infty}u^{2l}k^{2}(u)\sigma_{jg} ^{2}(uh_{n})f(uh_{n})du\\ & =o(1), \end{align*} again by the Dominated Convergence theorem. Conclude by applying Markov's inequality. \end{proof}
lemmaUnder Assumption (ref)(1)-(6), \[ \Gamma_{-,p}(h_{n})\rightarrow_{p}\Gamma_{-,p}, \] where \[ \Gamma_{-,p}=f(0)\left( \Gamma_{p}^{-}\otimes\sigma_{W}^{2-}\right) . \] \begin{proof} As with $\Gamma_{+,p}(h_{n}),$ by the change of variables $u=Z/h_{n},$ \begin{align*} \mathbb{E}[\Gamma_{-,p}(h_{n})] & =\mathbb{E}\left[ \frac{1}{n}\sum_{i=1}^{n}X_{i,p} X_{i,p}^{\prime}k_{ih_{n}}^{-}\right] \\ & =\mathbb{E}\left[ \left( r_{p}(Z_{i}/h_{n})r_{p}^{\prime}(Z_{i}/h_{n} )\otimes\sigma_{W}^{2}(Z_{i})\right) k_{ih_{n}}^{-}\right] \\ & =\int_{-\infty}^{0}\left( r_{p}(u)r_{p}^{\prime}(u)\otimes\sigma_{W} ^{2}(uh_{n})\right) k(u)f(uh_{n})du\\ & =\int_{0}^{\infty}\left( r_{p}(-u)r_{p}^{\prime}(-u)\otimes\sigma_{W} ^{2}(-uh_{n})\right) k(-u)f(-uh_{n})du\\ & \equiv\tilde{\Gamma}_{-,p}(h_{n}). \end{align*} The rest of the proof follows as for $\Gamma_{+,p}(h_{n})$. \end{proof}
lemmaUnder Assumption (ref)(1)-(6), \[ \vartheta_{p,q}^{+}(h_{n})\rightarrow_{p}\vartheta_{p,q}^{+}, \] where \[ \vartheta_{p,q}^{+}=f(0)\left( \theta_{p,q}^{+}\otimes\mu_{W}^{+}\right) . \] \begin{proof} By the change of variables $u=Z/h_{n},$ \begin{align*} \mathbb{E}[\vartheta_{p,q}^{+}(h_{n})] & =\mathbb{E}\left[ \frac{1}{n}\sum_{i=1}^{n} X_{i,p}(Z_{i}/h_{n})^{q}k_{ih_{n}}^{+}\right] \\ & =\mathbb{E}\left[ \left( r_{p}(Z_{i}/h_{n})\otimes\mu_{W}(Z_{i})\right) (Z_{i}/h_{n})^{q}k_{ih_{n}}^{+}\right] \\ & =\int_{0}^{\infty}\left( r_{p}(u)\otimes\mu_{W}(uh_{n})\right) u^{q}k(u)f(uh_{n})du\\ & \equiv\tilde{\vartheta}_{p,q}^{+}(h_{n}). \end{align*} By Assumptions (ref)(2) and (ref)(4) and Dominated Convergence theorem \[ \tilde{\vartheta}_{p,q}^{+}(h_{n})=\vartheta_{p,q}^{+}+o(1), \] where \[ \vartheta_{p,q}^{+}=f(0)\int_{0}^{\infty}\left( r_{p}(u)\otimes\mu_{W} ^{+}\right) u^{q}k(u)du. \] The rest of the proof follows the same arguments as for $\Gamma_{+,p}(h_{n}).$ \end{proof}
lemmaUnder Assumption (ref)(1)-(6), \[ \vartheta_{p,q}^{-}(h_{n})\rightarrow_{p}\vartheta_{p,q}^{-}, \] where \[ \vartheta_{p,q}^{-}=f(0)\left( \theta_{p,q}^{-}\otimes\mu_{W}^{-}\right) . \] \begin{proof} By the change of variables $u=Z/h_{n},$ \begin{align*} \mathbb{E}[\vartheta_{p,q}^{-}(h_{n})] & =\mathbb{E}\left[ \frac{1}{n}\sum_{i=1}^{n} X_{i,p}(Z_{i}/h_{n})^{q}k_{ih_{n}}^{-}\right] \\ & =\mathbb{E}\left[ \left( r_{p}(Z_{i}/h_{n})\otimes\mu_{W}(Z_{i})\right) (Z_{i}/h_{n})^{q}k_{ih_{n}}^{-}\right] \\ & =\int_{0}^{\infty}\left( r_{p}(-u)\otimes\mu_{W}(-uh_{n})\right) (-u)^{q}k(-u)f(-uh_{n})du\\ & \equiv\tilde{\vartheta}_{p,q}^{-}(h_{n}). \end{align*} By Assumptions (ref)(2) and (ref)(4) and Dominated Convergence theorem \[ \tilde{\vartheta}_{p,q}^{-}(h_{n})=\vartheta_{p,q}^{-}+o(1), \] where \[ \vartheta_{p,q}^{-}=f(0)(-1)^{q}\int_{0}^{\infty}\left( r_{p}(-u)\otimes \mu_{W}^{-}\right) u^{q}k(u)du. \] The rest of the proof follows the same arguments as for $\Gamma_{+,p}(h_{n}).$ \end{proof}
lemmaUnder Assumption (ref)(1)-(6), for $U$ and $V$ equal $Y_i$ or $X_i$ \[ h_{n}\Psi_{UV,p}^{+}(h_{n})\rightarrow_{p}\Psi_{UV,p}^{+}, \] where \[ \Psi_{UV,p}^{+}=f(0)\int_{0}^{\infty}\left( r_{p}(u)r_{p}^{\prime} (u)\sigma_{UV}^{2+}\otimes\psi_{UV}^{+}\right) k^{2}(u)du. \] \begin{proof} Note \[ \mathbb{E}[\sigma_{UV}^{2}(Z_{i},\tilde{W}_{i})|Z_{i}=z]=\sigma_{UV}^{2}(z) \] and \[ \mathbb{E}[\tilde{W}_{i}\tilde{W}_{i}^{\prime}\sigma_{UV}^{2}(Z_{i},\tilde{W} _{i})|Z_{i}=z]=\psi_{UV}(z). \] By the change of variables $u=Z/h_{n},$ it follows that \begin{align*} \mathbb{E}[h_{n}\Psi_{UV,p}^{+}(h_{n})] & =\mathbb{E}[\frac{h_{n}}{n}\sum_{i=1}^{n} X_{i,p}X_{i,p}^{\prime}\sigma_{UV}^{2}(Z_{i},\tilde{W}_{i})k_{ih_{n}}^{2+}]\\ & =\int_{0}^{\infty}\left( r_{p}(u)r_{p}^{\prime}(u)\sigma_{UV}^{2} (uh_{n})\otimes\psi_{UV}(uh_{n})\right) k^{2}(u)f(uh_{n})du\\ & \equiv\tilde{\Psi}_{UV,p}^{+}(h_{n}). \end{align*} By Assumption (ref)(4) and Dominated Convergence theorem \[ \tilde{\Psi}_{UV,p}^{+}(h_{n})=\Psi_{UV,p}^{+}+o(1), \] Also, since $\sigma_{UV}^{2}(Z_{i},\tilde{W}_{i})$ is bounded, \begin{align*} h_{n}^{2}\mathbb{E}\left[ \left\vert \Psi_{UV,p}^{+}(h_{n})-\mathbb{E}[\Psi_{UV,p}^{+} (h_{n})]\right\vert ^{2}\right] & \leq Cn^{-1}h_{n}^{-1}\int_{0}^{\infty }\left\vert r_{p}(u)\right\vert ^{4}k^{4}(u)f(uh_{n})du\\ & =O\left( n^{-1}h_{n}^{-1}\right) . \end{align*} \end{proof}
lemmaUnder Assumption (ref)(1)-(6), \[ h_{n}\Psi_{UV,p}^{-}(h_{n})\rightarrow_{p}\Psi_{UV,p}^{-}, \] where \[ \Psi_{UV,p}^{-}=f(0)\int_{0}^{\infty}\left( r_{p}(u)r_{p}^{\prime} (u)\sigma_{UV}^{2-}\otimes\psi_{UV,p}^{-}\right) k^{2}(u)du. \] \begin{proof} The proof follows the same arguments as the previous one. \end{proof}
lemmaUnder Assumption (ref)(1)-(7), for $U$ and $V$ equal $Y_i$ or $X_i$ and $m_{n}=\min\{h_{n},b_{n}\}$ with $m_{n}\rightarrow0$ and $nm_{n}\rightarrow\infty,$ \[ \frac{h_{n}b_{n}}{m_{n}}\Psi_{UV,p,q}^{+}(h_{n},b_{n})=\tilde{\Psi}_{UV,p} ^{+}(h_{n},b_{n})+o_{P}(1), \] where \[ \tilde{\Psi}_{UV,p,q}^{+}(h_{n},b_{n})=\int_{0}^{\infty}\left( r_{p}\left( \frac{m_{n}u}{h_{n}}\right) r_{q}^{\prime}\left( \frac{m_{n}u}{b_{n} }\right) \sigma_{UV}^{2}(um_{n})\otimes\psi_{UV}(um_{n})\right) k\left( \frac{m_{n}u}{h_{n}}\right) k\left( \frac{m_{n}u}{b_{n}}\right) f(um_{n})du. \] \begin{proof} By the change of variables $u=Z/h_{n},$ it follows that \begin{align*} & \mathbb{E}\left[ \frac{h_{n}b_{n}}{m_{n}}\Psi_{UV,p,q}^{+}(h_{n},b_{n})\right] \\ & =\mathbb{E}\left[ \frac{h_{n}b_{n}}{m_{n}}\frac{1}{n}\sum_{i=1}^{n}X_{i,p} X_{i,q}^{\prime}\sigma_{UV}^{2}(Z_{i},\tilde{W}_{i})k_{ih_{n}}^{+}k_{ib_{n} }^{+}\right] \\ & =\int_{0}^{\infty}\left( r_{p}\left( \frac{m_{n}u}{h_{n}}\right) r_{q}^{\prime}\left( \frac{m_{n}u}{b_{n}}\right) \sigma_{UV}^{2} (um_{n})\otimes\psi_{UV}(um_{n})\right) k\left( \frac{m_{n}u}{h_{n}}\right) k\left( \frac{m_{n}u}{b_{n}}\right) f(um_{n})du\\ & \equiv\tilde{\Psi}_{UV,p,q}^{+}(h_{n},b_{n}). \end{align*} Also, since $\sigma_{UV}^{2}(Z_{i},\tilde{W}_{i})$ is bounded, \begin{align*} & \mathbb{E}\left[ \left\vert \frac{h_{n}b_{n}}{m_{n}}\Psi_{UV,p}^{+}(h_{n} )-\mathbb{E}[\frac{h_{n}b_{n}}{m_{n}}\Psi_{UV,p}^{+}(h_{n})]\right\vert ^{2}\right] \\ & \leq Cn^{-1}m_{n}^{-1}\int_{0}^{\infty}k^{2}\left( \frac{m_{n}u}{h_{n} }\right) k^{2}\left( \frac{m_{n}u}{b_{n}}\right) \left\vert r_{p}\left( \frac{m_{n}u}{h_{n}}\right) \right\vert ^{2}\left\vert r_{q}\left( \frac{m_{n}u}{b_{n}}\right) \right\vert ^{2}f(uh_{n})du\\ & =O\left( n^{-1}m_{n}^{-1}\right) . \end{align*} \end{proof}
lemmaUnder Assumption (ref)(1)-(7), for $U$ and $V$ equal $Y_i$ or $X_i$ and $m_{n}=\min\{h_{n},b_{n}\}$ with $m_{n}\rightarrow0$ and $nm_{n}\rightarrow\infty,$ \[ \frac{h_{n}b_{n}}{m_{n}}\Psi_{UV,p,q}^{-}(h_{n},b_{n})=\tilde{\Psi}_{UV,p} ^{-}(h_{n},b_{n})+o_{P}(1), \] where \[ \tilde{\Psi}_{UV,p,q}^{-}(h_{n},b_{n})=\int_{-\infty}^{0}\left( r_{p}\left( \frac{m_{n}u}{h_{n}}\right) r_{q}^{\prime}\left( \frac{m_{n}u}{b_{n} }\right) \sigma_{UV}^{2}(um_{n})\otimes\psi_{UV}(um_{n})\right) k\left( \frac{m_{n}u}{h_{n}}\right) k\left( \frac{m_{n}u}{b_{n}}\right) f(um_{n})du. \] \begin{proof} The proof is analogous to the previous one. \end{proof}

The following lemma gives the asymptotic bias for the local polynomial estimator of $\mu_{V}^{+(s)}.$ Define the bias terms

align[align omitted — 209 chars of source]

and

align[align omitted — 210 chars of source]
lemmaUnder Assumption (ref)(1)-(6), for $V=Y_i$ or $X_i,$ and $d\geq l+2$ \begin{align*} E[\hat{\mu}_{V,l}^{+(s)}(h_{n})|\mathcal{S}_{n}] & =s!e_{s}^{\prime} \beta_{V,l}^{+}+h_{n}^{1+l-s}\frac{\mu_{V}^{+(l+1)}}{(l+1)!}\mathcal{B} _{s,l,l+1}^{+}(h_{n})\\ & +h_{n}^{2+l-s}\frac{\mu_{V}^{+(l+2)}}{(l+2)!}\mathcal{B}_{s,l,l+2} ^{+}(h_{n})+o_{P}(h_{n}^{2+l-s}), \end{align*} and \begin{align*} E[\hat{\mu}_{V,l}^{-(s)}(h_{n})|\mathcal{S}_{n}] & =s!e_{s}^{\prime} \beta_{V,l}^{-}+h_{n}^{1+l-s}\frac{\mu_{V}^{-(l+1)}}{(l+1)!}\mathcal{B} _{s,l,l+1}^{-}(h_{n})\\ & +h_{n}^{2+l-s}\frac{\mu_{V}^{-(l+2)}}{(l+2)!}\mathcal{B}_{s,l,l+2} ^{-}(h_{n})+o_{P}(h_{n}^{2+l-s}). \end{align*} Hence, \begin{align*} E[\hat{\delta}_{V,l}(h_{n})|\mathcal{S}_{n}] & =\delta_{V}+h_{n} ^{l+1}B_{V,l,l+1}(h_{n})\\ & +h_{n}^{2+l}B_{V,l,l+2}(h_{n})+o_{P}(h_{n}^{2+l}), \end{align*} where \begin{equation} B_{V,l,m}=\frac{\mu_{V}^{+(m)}}{m!}\mathcal{B}_{l,m}^{+}(h_{n})-\frac{\mu _{V}^{-(m)}}{m!}\mathcal{B}_{l,m}^{-}(h_{n}), \end{equation} with $\mathcal{B}_{l,m}^{+}(h_{n})=\mathcal{B}_{0,l,m}^{+}(h_{n} )=e_{0}^{\prime}\Gamma_{+,l}^{-1}(h_{n})\vartheta_{l,m}^{+}(h_{n})$ and $\mathcal{B}_{l,m}^{-}(h_{n})=\mathcal{B}_{0,l,m}^{-}(h_{n})=e_{0}^{\prime }\Gamma_{-,l}^{-1}(h_{n})\vartheta_{l,m}^{-}(h_{n}).$ \begin{proof} A Taylor series expansion gives \begin{align*} \mathbb{E}[\hat{\beta}_{V,l}^{+}(h_{n})|\mathcal{S}_{n}] & =\beta_{V,l}^{+} +h_{n}^{1+l}H_{l}(h_{n})\Gamma_{+,l}^{-1}(h_{n})X_{l}^{\prime}(h_{n} )K^{+}(h_{n})S_{l+1}(h_{n})\frac{\mu_{V}^{+(l+1)}}{n(l+1)!}\\ & +h_{n}^{2+l}H_{l}(h_{n})\Gamma_{+,l}^{-1}(h_{n})X_{l}^{\prime}(h_{n} )K^{+}(h_{n})S_{l+2}(h_{n})\frac{\mu_{V}^{+(l+2)}}{n(l+2)!}+H_{l}(h_{n} )o_{P}(h_{n}^{2+l}),\\ & =\beta_{V,l}^{+}+h_{n}^{1+l}H_{l}(h_{n})\Gamma_{+,l}^{-1}(h_{n} )\vartheta_{l,l+1}^{+}(h_{n})\frac{\mu_{V}^{+(l+1)}}{(l+1)!}\\ & +h_{n}^{2+l}H_{l}(h_{n})\Gamma_{+,l}^{-1}(h_{n})\vartheta_{l,l+2}^{+} (h_{n})\frac{\mu_{V}^{+(l+2)}}{(l+2)!}+H_{l}(h_{n})o_{P}(h_{n}^{2+l}). \end{align*} The expression for $\mathbb{E}[\hat{\mu}_{V,l}^{+(s)}(h_{n})|\mathcal{S}_{n}]$ follows from previous lemmas and $e_{s}^{\prime}H_{l}(h_{n})=h_{n}^{-s}I_{m}.$ The same arguments apply to $\mathbb{E}[\hat{\beta}_{V,l}^{-}(h_{n})|\mathcal{S}_{n}].$ \end{proof}

The following lemma gives the asymptotic variance for the local polynomial estimator of $\mu_{V}^{+(s)}.$

lemmaUnder Assumption (ref)(1)-(6), for $V=Y\ $or $X,$ and $d\geq l+2$ \[ Var[\hat{\mu}_{V,l}^{+(s)}(h_{n})|\mathcal{S}_{n}]=\mathcal{V}_{V,l} ^{+(s)}(h_{n}), \] with \begin{align} \mathcal{V}_{V,l}^{+(s)}(h_{n}) & =\frac{1}{nh_{n}^{2s}}\left( s!\right) ^{2}e_{s}^{\prime}\Gamma_{+,l}^{-1}(h_{n})\Psi_{V,l}^{+}(h_{n})\Gamma _{+,l}^{-1}(h_{n})e_{s}\\ & =\frac{1}{nh_{n}^{1+2s}}\left( s!\right) ^{2}e_{s}^{\prime}\Gamma _{+,l}^{-1}\Psi_{V,l}^{+}\Gamma_{+,l}^{-1}e_{s}\left[ 1+o_{P}(1)\right] ,\nonumber \end{align} and \[ Var[\hat{\mu}_{V,l}^{-(s)}(h_{n})|\mathcal{S}_{n}]=\mathcal{V}_{V,l} ^{-(s)}(h_{n}), \] with \begin{align} \mathcal{V}_{V,l}^{-(s)}(h_{n}) & =\frac{1}{nh_{n}^{2s}}\left( s!\right) ^{2}e_{s}^{\prime}\Gamma_{-,l}^{-1}(h_{n})\Psi_{V,l}^{-}(h_{n})\Gamma _{-,l}^{-1}(h_{n})e_{s}\\ & =\frac{1}{nh_{n}^{1+2s}}\left( s!\right) ^{2}e_{s}^{\prime}\Gamma _{-,l}^{-1}\Psi_{V,l}^{-}\Gamma_{-,l}^{-1}e_{s}\left[ 1+o_{P}(1)\right] .\nonumber \end{align} Furthermore, \begin{equation} Var[\hat{\delta}_{V,l}(h_{n})|\mathcal{S}_{n}]=\mathbf{V}_{V,l}(h_{n} )\equiv\mathcal{V}_{V,l}^{+}(h_{n})+\mathcal{V}_{V,l}^{-}(h_{n}), \end{equation} with $\mathcal{V}_{V,l}^{+}(h_{n})\equiv\mathcal{V}_{V,l}^{+(0)}(h_{n})$ and $\mathcal{V}_{V,l}^{-}(h_{n})\equiv\mathcal{V}_{V,l}^{-(0)}(h_{n}).$ \begin{proof} By $e_{s}^{\prime}H_{l}(h_{n})=h_{n}^{-s}I_{m}$ and definition of $\Psi _{V,l}^{+}(h_{n})$ \[ Var[s!e_{s}^{\prime}\hat{\beta}_{V,p}^{+}(h_{n})|\mathcal{S}_{n}]=h_{n} ^{-2s}\left( s!\right) ^{2}e_{s}^{\prime}\Gamma_{+,l}^{-1}(h_{n})\Psi _{V,l}^{+}(h_{n})\Gamma_{+,l}^{-1}(h_{n})e_{s}. \] The result follows from previous lemmas. The same arguments apply to $\hat {\mu}_{V,l}^{-(s)}(h_{n}).$ \end{proof}
lemmaUnder Assumption (ref)(1)-(6), for $V=Y_i\ $or $X_i,$ $nh_{n}^{2p+5}\rightarrow0$ and $d\geq p+2$ \begin{equation} \left( \mathcal{V}_{V,p}^{+(s)}(h_{n})\right) ^{-1/2}\left( \hat{\mu} _{V,p}^{+(s)}(h_{n})-\mu_{V}^{+(s)}-h_{n}^{1+p-s}\frac{\mu_{V+}^{(p+1)} }{(p+1)!}\mathcal{B}_{s,p,p+1}^{+}(h_{n})\right) \rightarrow_{d}N\left( 0,I_{m}\right) , \end{equation} \[ \left( \mathcal{V}_{V,l}^{-(s)}(h_{n})\right) ^{-1/2}\left( \hat{\mu} _{V,p}^{-(s)}(h_{n})-\mu_{V}^{-(s)}-h_{n}^{1+p-s}\frac{\mu_{V-}^{(p+1)} }{(p+1)!}\mathcal{B}_{s,p,p+1}^{-}(h_{n})\right) \rightarrow_{d}N\left( 0,I_{m}\right) , \] and \[ \left( \mathbf{V}_{V,p}(h_{n})\right) ^{-1/2}\left( \hat{\delta} _{V,p}(h_{n})-\delta_{V,p}(h_{n})-h_{n}^{1+p}B_{p,p+1}^{+}(h_{n})\right) \rightarrow_{d}N\left( 0,I_{m}\right) . \] \begin{proof} Write the left hand side of ((ref)) as \begin{align*} \left( \mathcal{V}_{V,p}^{+(s)}(h_{n})\right) ^{-1/2}\left( s!e_{s} ^{\prime}\hat{\beta}_{V,p}^{+}(h_{n})-s!e_{s}^{\prime}\beta_{V}^{+} (h_{n})-h_{n}^{1+p-s}\frac{\mu_{V+}^{(p+1)}}{(p+1)!}\mathcal{B}_{s,p,p+1} ^{+}(h_{n})\right) & =\xi_{1n}+\xi_{2n},\\ & =\xi_{1n}+o_{P}(1), \end{align*} where \begin{align*} \xi_{1n} & =\left( \mathcal{V}_{V,p}^{+(s)}(h_{n})\right) ^{-1/2}\left( s!e_{s}^{\prime}\hat{\beta}_{V,p}^{+}(h_{n})-\mathbb{E}[s!e_{s}^{\prime}\hat{\beta }_{V,p}^{+}(h_{n})|\mathcal{S}_{n}]\right) \\ & =\left( \mathcal{V}_{V,p}^{+(s)}(h_{n})\right) ^{-1/2}s!e_{s}^{\prime }H_{p}(h_{n})\Gamma_{+,p}^{-1}(h_{n})X_{p}^{\prime}(h_{n})K^{+}(h_{n} )\varepsilon_{V}, \end{align*} \begin{align*} \xi_{2n} & =\left( \mathcal{V}_{V,p}^{+(s)}(h_{n})\right) ^{-1/2}\left( E[s!e_{s}^{\prime}\hat{\beta}_{V,p}^{+}(h_{n})|\mathcal{S}_{n}]-s!e_{s} ^{\prime}\beta_{V}^{+}(h_{n})-h_{n}^{1+p-s}\frac{\mu_{V+}^{(p+1)}} {(p+1)!}\mathcal{B}_{s,p,p+1}^{+}(h_{n})\right) \\ & =O_{P}\left( \sqrt{nh_{n}^{1+2s}}\right) O_{P}\left( h_{n} ^{2+p-s}\right) =O_{P}\left( \sqrt{nh_{n}^{5+2p}}\right) =o_{P}(1). \end{align*} Write \begin{align*} \xi_{1n} & =\sum_{i=1}^{n}\varpi_{n,i}\varepsilon_{V_{i}}+o_{P}(1)\\ & \equiv\tilde{\xi}_{1n}+o_{P}(1), \end{align*} where \[ \varpi_{n,i}=\left( n^{-1}h_{n}^{-1-2s}e_{l}^{\prime}\Gamma_{+,l}^{-1} \Psi_{V,l}^{+}\Gamma_{+,l}^{-1}e_{l}\right) ^{-1/2}e_{p}^{\prime}\Gamma _{+,p}^{-1}(h_{n})X_{i,p}k_{ih_{n}}^{+}/n. \] For any $\lambda\in\mathbb{R}^{m}$ with unit norm, $\{\lambda^{\prime} \varpi_{n,i}\varepsilon_{V_{i}}\}_{i=1}^{n}$ is a triangular array of independent random variables where $\mathbb{E}[\tilde{\xi}_{1n}]=0$ and $Var\left( \tilde{\xi}_{1n}\right) \rightarrow1.$ Thus, $\tilde{\xi}_{1n}\rightarrow _{d}N\left( 0,1\right) $ by Linderberg-Feller central limit theorem for triangular arrays because \begin{align*} \sum_{i=1}^{n}\mathbb{E}\left[ \left\vert \lambda^{\prime}\varpi_{n,i}\varepsilon _{V_{i}}\right\vert ^{4}\right] & \lesssim n^{2}h_{n}^{2+4s}h_{n}^{-4s} \sum_{i=1}^{n}\mathbb{E}\left[ \left\vert e_{s}^{\prime}\Gamma_{+,p}^{-1} (h_{n})X_{i,p}k_{ih_{n}}^{+}\right\vert ^{4}/n^{4}\right] \\ & =O(n^{-1}h_{n}^{-1})\\ & =o(1). \end{align*} The result for $\hat{\mu}_{V,p}^{-(s)}$ and $\hat{\delta}_{V,p}$ follows the same arguments. \end{proof}
lemmaUnder Assumption (ref)(1)-(7), for $d\geq p+2,$ $V=Y$ and $X,$ \begin{align*} \hat{c}-c & =O_{P}\left( \frac{1}{\sqrt{nh_{n}}}+h_{n}^{(p+1)}\right) ,\\ \hat{\delta}_{V}-\delta_{V} & =O_{P}\left( \frac{1}{\sqrt{nh_{n}}} +h_{n}^{(p+1)}\right) , \end{align*} \[ \hat{\mathcal{R}}=O_{P}\left( \frac{1}{nh_{n}}+h_{n}^{2(p+1)}\right) , \] \begin{proof} A Taylor expansion, $\hat{\pi}-\pi=O_{P}(n^{-1/2})$ and previous results yield \[ (\hat{c}-c)=G_{+}(\hat{\mu}_{V}^{+}-\mu_{V}^{+})+G_{-}\left( \hat{\mu} _{V}^{-}-\mu_{V}^{-}\right) +o_{P}(n^{-1/2}h_{n}^{-1/2}+h_{n}^{(p+1)}), \] where $G_{+}$ and $G_{-}$ are the derivatives of $g$ with respect to $\mu _{V}^{+}$ and $\mu_{V}^{-},$ respectively. Previous lemmas imply for $V=Y$ and $X_i$ \[ \left( \hat{\delta}_{V}-\delta_{V}\right) ^{2}=O_{P}\left( \frac{1}{nh_{n} }+h_{n}^{2(p+1)}\right) . \] Write \begin{equation} \hat{\tau}_{Y}-\tau_{Y}\equiv c^{\prime}\left( \hat{\delta}_{Y}-\delta _{Y}\right) +(\hat{c}-c)^{\prime}\delta_{Y}+(\hat{c}-c)^{\prime}\left( \hat{\delta}_{Y}-\delta_{Y}\right) \end{equation} Thus, by previous lemmas \[ \left( \hat{\tau}_{Y}-\tau_{Y}\right) ^{2}=O_{P}\left( \frac{1}{nh_{n} }+h_{n}^{2(p+1)}\right) . \] The same results hold for $\left( \hat{\tau}_{X}-\tau_{X}\right) ^{2}.$ Furthermore, \[ \left( \hat{\tau}_{Y}-\tau_{Y}\right) \left( \hat{\tau}_{X}-\tau _{X}\right) =O_{P}\left( \sqrt{\frac{1}{nh_{n}}+h_{n}^{2(p+1)}}\right) O_{P}\left( \sqrt{\frac{1}{nh_{n}}+h_{n}^{2(p+1)}}\right) . \] The results follows from the expression for $\hat{\mathcal{R}}$ in ((ref)). \end{proof}

Define the linear term in the expansion ((ref)) \[ \hat{l}_{\omega}=\frac{\hat{\tau}_{Y}-\tau_{Y}}{\tau_{X}}-\frac{\tau _{Y}\left( \hat{\tau}_{X}-\tau_{X}\right) }{\tau_{X}^{2}}, \] where each term $\hat{\tau}_{V}$ has been estimated by a $p-th$ order local polynomial.

lemmaUnder Assumption (ref)(1)-(7), for $d\geq p+2,$ \[ \mathbb{E}[\hat{l}_{\omega}|\mathcal{S}_{n}]=h_{n}^{p+1}\mathbf{B}_{\omega,p,p+1} (h_{n})+h_{n}^{p+2}\mathbf{B}_{\omega,p,p+2}(h_{n})+o_{P}(h_{n}^{2+p}), \] where \begin{equation} \mathbf{B}_{\omega,p,m}=\frac{B_{\tau_{Y},p,m}(h_{n})}{\tau_{X}}-\frac {\tau_{Y}B_{\tau_{X},p,m}(h_{n})}{\tau_{X}^{2}}, \end{equation} \[ B_{\tau_{Y},p,m}(h_{n})=c^{\prime}B_{Y,p,m}(h_{n})+B_{C,p,m}^{\prime} (h_{n})\delta_{Y}, \] \[ B_{\tau_{X},p,m}(h_{n})=c^{\prime}B_{X,p,m}(h_{n})+B_{C,p,m}^{\prime} (h_{n})\delta_{X}, \] with $B_{V,l,m}$ defined ((ref)) and \[ B_{C,p,m}(h_{n})=G_{+}\frac{\mu_{V}^{+(m)}}{m!}\mathcal{B}_{p,m}^{+} (h_{n})+G_{-}\frac{\mu_{V}^{-(m)}}{m!}\mathcal{B}_{p,m}^{-}(h_{n}). \] \begin{proof} By the delta method \[ \mathbb{E}[\hat{c}-c|\mathcal{S}_{n}]=h_{n}^{1+p}B_{C,p,p+1}(h_{n})+h_{n} ^{2+p}B_{C,p,p+2}(h_{n})+o_{P}(h_{n}^{2+p}). \] From ((ref)) and the bias expression for $\mathbb{E}[\hat{\delta}_{V,l} (h_{n})|\mathcal{S}_{n}]$ it holds \[ \mathbb{E}[\hat{\tau}_{Y}-\tau_{Y}|\mathcal{S}_{n}]=h_{n}^{1+p}B_{\tau_{Y},p,p+1} (h_{n})+h_{n}^{2+p}B_{\tau_{Y},p,p+2}(h_{n})+o_{P}(h_{n}^{2+p}), \] The same expression holds for $\mathbb{E}[\hat{\tau}_{X}-\tau_{X}|\mathcal{S}_{n}].$ The result follows from the expression for $\hat{l}_{\omega}.$ \end{proof}
lemmaUnder Assumption (ref)(1)-(7), for $d\geq p+2,$ \begin{equation} \mathbf{V}_{\omega,p}(h_{n})=Var[\hat{l}_{\omega}|\mathcal{S}_{n}]=\frac {1}{\tau_{X}^{2}}\mathbf{V}_{\tau,Y,p}(h_{n})-\frac{2\tau_{Y}}{\tau_{X}^{3} }\mathbf{C}_{\tau,p}(h_{n})+\frac{\tau_{Y}^{2}}{\tau_{X}^{4}}\mathbf{V} _{\tau,X,p}(h_{n}), \end{equation} where \begin{align*} \mathbf{V}_{\tau,Y,p}(h_{n}) & =c^{\prime}\mathbf{V}_{Y,p}(h_{n} )c+\delta_{Y}^{\prime}G_{+}\mathcal{V}_{V,p}^{+}(h_{n})G_{+}^{\prime} \delta_{Y}+\delta_{Y}^{\prime}G_{-}\mathcal{V}_{V,p}^{-}(h_{n})G_{-}^{\prime }\delta_{Y},\\ \mathbf{C}_{\tau,p}(h_{n}) & =c^{\prime}\mathbf{C}_{p}(h_{n})c+c^{\prime }\mathbf{C}_{YC,p}(h_{n})\delta_{X}+\delta_{Y}^{\prime}\mathbf{C}_{XC,p} (h_{n})c+\delta_{Y}^{\prime}\mathbf{V}_{C,p}(h_{n})\delta_{X}, \\ \mathbf{C}_{p}(h_{n}) & =\mathcal{C}_{YX,p}^{+}(h_{n})+\mathcal{C} _{YX,p}^{-}(h_{n}),\\ \mathbf{C}_{YC,p}(h_{n}) & =G_{+}\mathcal{C}_{YV,p}^{+}(h_{n})-G_{-} \mathcal{C}_{YV,p}^{-}(h_{n}),\\ \mathbf{C}_{XC,p}(h_{n}) & =G_{+}\mathcal{C}_{XV,p}^{+}(h_{n})-G_{-} \mathcal{C}_{XV,p}^{-}(h_{n}), \\ \mathbf{V}_{C,p}(h_{n}) & =G_{+}\mathcal{V}_{V,p}^{+}(h_{n})G_{+}^{\prime }+G_{-}\mathcal{V}_{V,p}^{-}(h_{n})G_{-}^{\prime},\\ \mathcal{C}_{UV,p}^{+}(h_{n}) & =n^{-1}e_{0}^{\prime}\Gamma_{+,p}^{-1} (h_{n})\Psi_{UV,p}^{+}(h_{n})\Gamma_{+,p}^{-1}(h_{n})e_{0}, \\ \mathcal{C}_{UV,p}^{-}(h_{n}) & =n^{-1}e_{0}^{\prime}\Gamma_{-,p}^{-1} (h_{n})\Psi_{UV,p}^{-}(h_{n})\Gamma_{-,p}^{-1}(h_{n})e_{0}. \end{align*} \begin{proof} From the expression of $\hat{l}_{\omega},$ \[ Var[\hat{l}_{\omega}|\mathcal{S}_{n}]=\frac{1}{\tau_{X}^{2}}Var[\hat{\tau} _{Y}-\tau_{Y}|\mathcal{S}_{n}]-\frac{2\tau_{Y}}{\tau_{X}^{3}}Cov[\hat{\tau }_{Y}-\tau_{Y},\hat{\tau}_{X}-\tau_{X}|\mathcal{S}_{n}]+\frac{\tau_{Y}^{2} }{\tau_{X}^{4}}Var[\hat{\tau}_{X}-\tau_{X}|\mathcal{S}_{n}]. \] Fromt the expansion of $\hat{\tau}_{V}-\tau_{V},$ \begin{align*} Var[\hat{\tau}_{Y}-\tau_{Y}|\mathcal{S}_{n}] & =c^{\prime}Var[\hat{\delta }_{Y}-\delta_{Y}|\mathcal{S}_{n}]c+\delta_{Y}^{\prime}G_{+}Var[\hat{\mu} _{Y}^{+}-\mu_{Y}^{+}|\mathcal{S}_{n}]G_{+}^{\prime}\delta_{Y}\\ & +\delta_{Y}^{\prime}G_{-}Var[\hat{\mu}_{Y}^{+}-\mu_{Y}^{+}|\mathcal{S} _{n}]G_{-}^{\prime}\delta_{Y},+o_{P}(1),\\ Cov[\hat{\tau}_{Y}-\tau_{Y},\hat{\tau}_{X}-\tau_{X}|\mathcal{S}_{n}] & =c^{\prime}Cov[\hat{\delta}_{Y}-\delta_{Y},\hat{\delta}_{X}-\delta _{X}|\mathcal{S}_{n}]c\\ & +c^{\prime}Cov[\hat{\delta}_{Y}-\delta_{Y},\hat{c}-c|\mathcal{S}_{n} ]\delta_{X}\\ & +\delta_{Y}^{\prime}Cov[\hat{c}-c,\hat{\delta}_{Y}-\delta_{Y} |\mathcal{S}_{n}]c\\ & +\delta_{Y}^{\prime}Var[\hat{c}-c|\mathcal{S}_{n}]\delta_{X}+o_{P}(1)\\ Var[\hat{\tau}_{X}-\tau_{X}|\mathcal{S}_{n}] & =c^{\prime}Var[\hat{\delta }_{X}-\delta_{X}|\mathcal{S}_{n}]c+\delta_{X}^{\prime}G_{+}Var[\hat{\mu} _{X}^{+}-\mu_{X}^{+}|\mathcal{S}_{n}]G_{+}^{\prime}\delta_{X}\\ & +\delta_{X}^{\prime}G_{-}Var[\hat{\mu}_{X}^{+}-\mu_{X}^{+}|\mathcal{S} _{n}]G_{-}^{\prime}\delta_{X},+o_{P}(1). \end{align*} Similarly to our previous results, it is shown that \begin{align*} Cov[\hat{\delta}_{Y}-\delta_{Y},\hat{\delta}_{X}-\delta_{X}|\mathcal{S}_{n}] & =Cov[\hat{\mu}_{Y}^{+}-\mu_{Y}^{+},\hat{\mu}_{X}^{+}-\mu_{X}^{+} |\mathcal{S}_{n}]+Cov[\hat{\mu}_{Y}^{-}-\mu_{Y}^{-},\hat{\mu}_{X}^{-}-\mu _{X}^{-}|\mathcal{S}_{n}]\\ & =\mathcal{C}_{YX,p}^{+}(h_{n})+\mathcal{C}_{YX,p}^{-}(h_{n}),\\ Cov[\hat{\delta}_{Y}-\delta_{Y},\hat{c}-c|\mathcal{S}_{n}] & =G_{+} \mathcal{C}_{YV,p}^{+}(h_{n})-G_{-}\mathcal{C}_{YV,p}^{-}(h_{n})+o_{P}(1),\\ Var[\hat{c}-c|\mathcal{S}_{n}] & =G_{+}\mathcal{V}_{V,p}^{+}(h_{n} )G_{+}^{\prime}+G_{-}\mathcal{V}_{V,p}^{-}(h_{n})G_{-}^{\prime}+o_{P}(1). \end{align*} \end{proof}
lemmaUnder Assumption (ref)(1)-(7), $nh_{n}^{2p+5} \rightarrow0$ and $d\geq p+2$ \[ \left( \mathbf{V}_{\omega,p}(h_{n})\right) ^{-1/2}\left( \hat{\beta }_{\omega}-\beta_{\omega}-h_{n}^{p+1}\mathbf{B}_{\omega,p,p+1}(h_{n})\right) \rightarrow_{d}N\left( 0,1\right) , \] \begin{proof} Write the left hand side of the last display as \begin{align*} \left( \mathbf{V}_{\omega,p}(h_{n})\right) ^{-1/2}\left( \hat{l}_{\omega }-h_{n}^{p+1}\mathbf{B}_{\omega,p,p+1}(h_{n})\right) +\left( \mathbf{V} _{\omega,p}(h_{n})\right) ^{-1/2}\hat{\mathcal{R}} & =\xi_{1n}+\xi_{2n},\\ & =\xi_{1n}+o_{P}(1), \end{align*} where \begin{align*} \xi_{2n} & =O_{P}\left( \frac{\sqrt{nh_{n}}}{nh_{n}}+\sqrt{nh_{n}} h_{n}^{2(p+1)}\right) \\ & =o_{P}(1) \end{align*} and \begin{align*} \xi_{1n} & =\left( \mathbf{V}_{\omega,p}(h_{n})\right) ^{-1/2}\left( \hat{l}_{\omega}-\mathbb{E}[\hat{l}_{\omega}|\mathcal{S}_{n}]\right) +\left( \mathbf{V}_{\omega,p}(h_{n})\right) ^{-1/2}\left( h_{n}^{p+2}\mathbf{B} _{\omega,p,p+2}(h_{n})+o_{P}(h_{n}^{2+p})\right) \\ & \rightarrow_{d}N\left( 0,1\right) , \end{align*} by Lemma (ref). \end{proof}
theoremLet Assumption (ref) hold. Suppose also that Assumption (ref)(1)-(7) in Section (ref) in the Appendix hold with $p=1$. Then \[ \left( \mathbf{V}_{\omega,1}(h_{n})\right) ^{-1/2}\left( \hat{\beta }_{\omega}-\beta_{\omega}-h_{n}^{2}\mathbf{B}_{\omega,1,2}(h_{n})\right) \rightarrow_{d}N\left( 0,1\right) , \] where $\mathbf{B}_{\omega,1,2}(h_{n})$ and $\mathbf{V}_{\omega,1}(h_{n})$ are given in (ref) and (ref), respectively, in Section (ref) in the Appendix.

\noindentProof of Theorem (ref): It follows from Lemma (ref) for $p=1.$ $\blacksquare$

When $g$ depends on $\delta_{X}$ the previous formulas simplify. In this case and with an abuse of notation, we define the linear term as a linear combination of $\hat{\delta}_{Y}-\delta_{Y}$ and $\hat{\delta}_{X}-\delta_{X}$ as \[ \hat{l}_{\omega}=c_{1}^{\prime}(\hat{\delta}_{Y}-\delta_{Y})+c_{2}^{\prime }\left( \hat{\delta}_{X}-\delta_{X}\right) , \] where $c_{1}=c/\tau_{X}$, $c_{2}^{\prime}=\left[ \delta_{Y}^{\prime} G_{\delta}-\beta_{w}(c^{\prime}+\delta_{X}^{\prime}G_{\delta})\right] /\tau_{X}$ and $G_{\delta}$ is the derivative of $g$ with respect to $\delta_{X}$.

lemmaUnder Assumption (ref)(1)-(7), for $d\geq p+2,$ \[ \mathbb{E}[\hat{l}_{\omega}|\mathcal{S}_{n}]=h_{n}^{p+1}\mathbf{B}_{\omega,p,p+1} (h_{n})+h_{n}^{p+2}\mathbf{B}_{\omega,p,p+2}(h_{n})+o_{P}(h_{n}^{2+p}), \] where now \begin{equation} \mathbf{B}_{\omega,p,m}=c_{1}^{\prime}B_{Y,p,m}(h_{n})+c_{2}^{\prime} B_{X,p,m}(h_{n}), \end{equation} and $B_{V,p,m}(h_{n})$ is defined in ((ref)). \begin{proof} It follows by a simple Taylor expansion and our previous results. \end{proof}

Recall the definitions for conditional covariances and variances of $(\hat{\mu}_{U}^{+}-\mu_{U}^{+})$ and $\left( \hat{\mu}_{V}^{+}-\mu_{V} ^{+}\right) $ (and similarly for $-$ parts)

align*[align* omitted — 418 chars of source]
lemmaUnder Assumption (ref)(1)-(7), for $d\geq p+2,$ \begin{equation} \mathbf{V}_{\omega,p}(h_{n})=Var[\hat{l}_{\omega}|\mathcal{S}_{n} ]=V_{+1}+V_{-1}, \end{equation} where \[ V_{+1}=c_{1}^{\prime}\mathcal{V}_{Y,p}^{+}(h_{n})c_{1}+2c_{1}^{\prime }\mathcal{C}_{YX,p}^{+}(h_{n})c_{2}+c_{2}^{\prime}\mathcal{V}_{X,p}^{+} (h_{n})c_{2}, \] and \[ V_{-1}=c_{1}^{\prime}\mathcal{V}_{Y,p}^{-}(h_{n})c_{1}+2c_{1}^{\prime }\mathcal{C}_{YX,p}^{-}(h_{n})c_{2}+c_{2}^{\prime}\mathcal{V}_{X,p}^{-} (h_{n})c_{2}. \] An equivalent representation is \[ \mathbf{V}_{\omega,p}(h_{n})=c_{1}^{\prime}\mathbf{V}_{Y,p}(h_{n})c_{1} +2c_{1}^{\prime}\mathbf{C}_{p}(h_{n})c_{2}+c_{2}^{\prime}\mathbf{V} _{X,p}(h_{n})c_{2}, \] where $\mathbf{V}_{V,p}(h_{n})$ is defined in ((ref)), and \[ \mathbf{C}_{p}(h_{n})=\mathcal{C}_{YX,p}^{+}(h_{n})+\mathcal{C}_{YX,p} ^{-}(h_{n}). \] \begin{proof} Write $\hat{l}_{\omega}=\hat{l}_{\omega}^{+}-\hat{l}_{\omega}^{+},$ where \begin{align*} \hat{l}_{\omega}^{+} & =c_{1}^{\prime}(\hat{\mu}_{Y}^{+}-\mu_{Y}^{+} )+c_{2}^{\prime}\left( \hat{\mu}_{X}^{+}-\mu_{X}^{+}\right) ,\\ \hat{l}_{\omega}^{-} & =c_{1}^{\prime}(\hat{\mu}_{Y}^{-}-\mu_{Y}^{-} )+c_{2}^{\prime}\left( \hat{\mu}_{X}^{-}-\mu_{X}^{-}\right) . \end{align*} The proof follows from previous results. \end{proof}

The following theory is for the bias-correction. Assume we use a $q$-th order polynomial for estimating the bias term (the $p+1$ derivative), $q\geq p+1$. Linearizing the bias-corrected estimator, we obtain

align*[align* omitted — 327 chars of source]

where

align*[align* omitted — 658 chars of source]

and $\mathcal{B}_{l,m}^{+}(h_{n})=e_{0}^{\prime}\Gamma_{+,l}^{-1} (h_{n})\vartheta_{l,m}^{+}(h_{n})$ and $\mathcal{B}_{l,m}^{-}(h_{n} )=e_{0}^{\prime}\Gamma_{-,l}^{-1}(h_{n})\vartheta_{l,m}^{-}(h_{n}).$

lemmaUnder Assumption (ref)(1)-(7), for $d\geq p+2,$ $n\min\{h_{n},b_{n}\}\rightarrow\infty,$ $h_{n}\rightarrow0,$ and $nh_{n}\rightarrow\infty,$ then \[ R_{n}=O_{P}\left( \frac{1}{nh_{n}}+h_{n}^{2(p+1)}\right) , \] \[ R_{n}^{bc}=O_{P}\left( \frac{h_{n}^{p+1}}{\sqrt{nh_{n}}}+h_{n}^{2(p+1)} \right) O_{P}\left( \frac{1}{\sqrt{nb_{n}^{3+2p}}}+1\right) . \] \begin{proof} Note \begin{equation} R_{n}=\frac{R_{\tau X}}{\tau_{X}}-\frac{\tau_{Y}R_{\tau Y}}{\tau_{X}^{2}} +\hat{\mathcal{R}}, \end{equation} where $R_{\tau X}$ and $R_{\tau Y}$ are implicitly defined from \[ \hat{\tau}_{Y}-\tau_{Y}\equiv c^{\prime}\left( \hat{\delta}_{Y}-\delta _{Y}\right) +\delta_{Y}^{\prime}G_{\delta}(\hat{\delta}_{X}-\delta _{X})+R_{\tau Y} \] and \[ \hat{\tau}_{X}-\tau_{X}\equiv\left( c^{\prime}+\delta_{X}^{\prime}G_{\delta }\right) (\hat{\delta}_{X}-\delta_{X})+R_{\tau X}. \] By continuous differentiability of $g$ and Lemmas (ref) and (ref) \[ R_{\tau V}=O_{P}\left( \frac{1}{nh_{n}}+h_{n}^{2(p+1)}\right) . \] The rate for $R_{n}$ then follows from ((ref)), the last display and Lemma (ref). On the other hand, by previous results \begin{align*} \left\vert R_{n}^{bc}\right\vert & \lesssim h_{n}^{p+1}\left\vert \hat {c}_{1}-c_{1}\right\vert \left\{ \left\vert e_{p+1}^{\prime}\hat{\beta} _{Y,q}^{+}(b_{n})\right\vert +\left\vert e_{p+1}^{\prime}\hat{\beta}_{Y,q} ^{-}(b_{n})\right\vert \right\} \\ & +h_{n}^{p+1}\left\vert \hat{c}_{2}-c_{2}\right\vert \left\{ \left\vert e_{p+1}^{\prime}\hat{\beta}_{X,q}^{+}(b_{n})\right\vert +\left\vert e_{p+1}^{\prime}\hat{\beta}_{X,q}^{-}(b_{n})\right\vert \right\} \\ & =O_{P}\left( \frac{h_{n}^{p+1}}{\sqrt{nh_{n}}}+h_{n}^{2(p+1)}\right) O_{P}\left( \frac{1}{\sqrt{nb_{n}^{3+2p}}}+1\right) , \end{align*} By\ Lemmas (ref), (ref) and (ref) . \end{proof}
lemmaUnder Assumption (ref)(1)-(7), for $d\geq p+2,$ $n\min\{h_{n},b_{n}\}\rightarrow\infty,$ $\max\{h_{n},b_{n}\}\rightarrow0,$ then \begin{align*} \mathbb{E}[\hat{l}_{\omega}^{bc}|\mathcal{S}_{n}] & =h_{n}^{p+2}\mathbf{B} _{\omega,p,p+2}(h_{n})[1+o_{P}(1)]\\ & +h_{n}^{p+1}b_{n}^{p-q}B_{\omega,p,q}^{bc}(h_{n},b_{n})[1+o_{P}(1)], \end{align*} where \[ B_{\omega,p,q}^{bc}(h_{n},b_{n})=c_{1}^{\prime}B_{Y,p,q}^{bc}(h_{n} ,b_{n})+c_{2}^{\prime}B_{X,p,q}^{bc}(h_{n},b_{n}) \] \[ B_{V,p,q}^{bc}(h_{n},b_{n})=\frac{\mu_{V}^{+(q+1)}}{(q+2)!}\mathcal{B} _{p+1,q,q+1}^{+}(b_{n})\frac{\mathcal{B}_{p,p+1}^{+}(h_{n})}{(p+1)!}-\frac {\mu_{V}^{-(q+1)}}{(q+1)!}\mathcal{B}_{p+1,q,q+1}^{-}(b_{n})\frac {\mathcal{B}_{p,p+1}^{-}(h_{n})}{(p+1)!}. \] \begin{proof} Note with \begin{align*} \mathbb{E}[\hat{l}_{\omega}^{bc}|\mathcal{S}_{n}] & =\mathbb{E}[\hat{l}_{\omega}-h_{n} ^{p+1}\mathbf{B}_{\omega,p,p+1}(h_{n})|\mathcal{S}_{n}]\\ & -h_{n}^{p+1}\mathbb{E}[\tilde{B}_{\omega,p,q}(h_{n},b_{n})-\mathbf{B}_{\omega ,p,p+1}(h_{n})|\mathcal{S}_{n}]\\ & \equiv B_{1}-h_{n}^{p+1}B_{2}. \end{align*} By\ Lemma (ref) \[ B_{1}=h_{n}^{p+2}\mathbf{B}_{\omega,p,p+2}(h_{n})[1+o_{P}(1)] \] and \begin{align*} B_{2} & =c_{1}^{\prime}\mathbb{E}[\hat{B}_{Y,p,q}(h_{n},b_{n})-B_{Y,p,p+1} (h_{n})|\mathcal{S}_{n}]+c_{2}^{\prime}\mathbb{E}[\hat{B}_{X,p,q}(h_{n},b_{n} )-B_{X,p,p+1}(h_{n})|\mathcal{S}_{n}]\\ & =c_{1}^{\prime}\mathbb{E}[e_{p+1}^{\prime}\hat{\beta}_{Y,q}^{+}(b_{n})-e_{p+1} ^{\prime}\beta_{Y}^{+}|\mathcal{S}_{n}]\mathcal{B}_{p,p+1}^{+}(h_{n} )-c_{1}^{\prime}\mathbb{E}[e_{p+1}^{\prime}\hat{\beta}_{Y,q}^{-}(b_{n})-e_{p+1} ^{\prime}\beta_{Y}^{-}|\mathcal{S}_{n}]\mathcal{B}_{p,p+1}^{-}(h_{n})\\ & +c_{2}^{\prime}\mathbb{E}[e_{p+1}^{\prime}\hat{\beta}_{X,q}^{+}(b_{n})-e_{p+1} ^{\prime}\beta_{X}^{+}|\mathcal{S}_{n}]\mathcal{B}_{p,p+1}^{+}(h_{n} )-c_{2}^{\prime}\mathbb{E}[e_{p+1}^{\prime}\hat{\beta}_{X,q}^{-}(b_{n})-e_{p+1} ^{\prime}\beta_{X}^{-}|\mathcal{S}_{n}]\mathcal{B}_{p,p+1}^{-}(h_{n})\\ & =c_{1}^{\prime}b_{n}^{p-q}\left[ \frac{\mu_{Y}^{+(q+1)}}{(q+1)!} \mathcal{B}_{p+1,q,q+1}^{+}(b_{n})\frac{\mathcal{B}_{p,p+1}^{+}(h_{n} )}{(p+1)!}-\frac{\mu_{Y}^{-(q+1)}}{(q+1)!}\mathcal{B}_{p+1,q,q+1}^{-} (b_{n})\frac{\mathcal{B}_{p,p+1}^{-}(h_{n})}{(p+1)!}\right] \\ & +c_{2}^{\prime}b_{n}^{p-q}\left[ \frac{\mu_{X}^{+(q+1)}}{(q+1)!} \mathcal{B}_{p+1,q,q+1}^{+}(b_{n})\frac{\mathcal{B}_{p,p+1}^{+}(h_{n} )}{(p+1)!}-\frac{\mu_{X}^{-(q+1)}}{(q+1)!}\mathcal{B}_{p+1,q,q+1}^{-} (b_{n})\frac{\mathcal{B}_{p,p+1}^{-}(h_{n})}{(p+1)!}\right] \\ & =b_{n}^{q-p}B_{\omega,p,p+2}^{bc}(h_{n},b_{n}). \end{align*} \end{proof}

Define

align*[align* omitted — 282 chars of source]

and

align*[align* omitted — 282 chars of source]

We drop one $U$ when $V=U$, e.g. $\mathcal{C}_{U,p,q}^{+}(h_{n},b_{n} )=\mathcal{C}_{UU,p,q}^{+}(h_{n},b_{n}).$

lemmaUnder Assumption (ref)(1)-(7), for $d\geq p+2,$ $n\min\{h_{n},b_{n}\}\rightarrow\infty,$ $\max\{h_{n},b_{n}\}\rightarrow0,$ then \begin{equation} \mathbf{V}_{\omega,p,q}^{bc}(h_{n},b_{n})\equiv Var[\hat{l}_{\omega} ^{bc}|\mathcal{S}_{n}]=\mathbf{V}_{\omega,p}(h_{n})+\mathbf{C}_{\omega ,p,q}^{bc}(h_{n},b_{n}), \end{equation} where \begin{equation} \mathbf{C}_{\omega,p,q}^{bc}(h_{n},b_{n})=2h_{n}^{p+1}\left( C_{+12} +C_{-12}\right) +h_{n}^{2p+2}\left( V_{+2}+V_{-2}\right) \end{equation} \[ C_{+12}=\left[ c_{1}^{\prime}\mathcal{C}_{Y,p,q}^{+}(h_{n},b_{n})c_{1} +c_{1}^{\prime}\mathcal{C}_{YX,p,q}^{+}(h_{n},b_{n})c_{2}+c_{1}^{\prime }\mathcal{C}_{XY,p,q}^{+}(h_{n},b_{n})c_{2}+c_{2}^{\prime}\mathcal{C} _{X,p,q}^{+}(h_{n},b_{n})c_{2}\right] \frac{\mathcal{B}_{p,p+1}^{+}(h_{n} )}{(p+1)!}, \] \[ C_{-12}=\left[ c_{1}^{\prime}\mathcal{C}_{Y,p,q}^{-}(h_{n},b_{n})c_{1} +c_{1}^{\prime}\mathcal{C}_{YX,p,q}^{-}(h_{n},b_{n})c_{2}+c_{1}^{\prime }\mathcal{C}_{XY,p,q}^{-}(h_{n},b_{n})c_{2}+c_{2}^{\prime}\mathcal{C} _{X,p,q}^{-}(h_{n},b_{n})c_{2}\right] \frac{\mathcal{B}_{p,p+1}^{-}(h_{n} )}{(p+1)!}, \] \[ V_{+2}=\left( c_{1}^{\prime}\mathcal{V}_{Y,p}^{+(q)}(b_{n})c_{1} +2c_{1}^{\prime}\mathcal{C}_{YX,p,q}^{+}(b_{n})c_{2}+c_{2}^{\prime} \mathcal{V}_{X,p}^{+(q)}(b_{n})c_{2}\right) \left( \frac{\mathcal{B} _{p,p+1}^{+}(h_{n})}{(p+1)!}\right) ^{2} \] and \[ V_{-2}=\left( c_{1}^{\prime}\mathcal{V}_{Y,p}^{-(q)}(b_{n})c_{1} +2c_{1}^{\prime}\mathcal{C}_{YX,p,q}^{-}(b_{n})c_{2}+c_{2}^{\prime} \mathcal{V}_{X,p}^{-(q)}(b_{n})c_{2}\right) \left( \frac{\mathcal{B} _{p,p+1}^{-}(h_{n})}{(p+1)!}\right) ^{2} \] \begin{proof} Note that $\hat{l}_{\omega}^{bc}=\hat{l}_{\omega}^{+bc}-\hat{l}_{\omega} ^{-bc},$ so that \[ Var[\hat{l}_{\omega}^{bc}|\mathcal{S}_{n}]=Var[\hat{l}_{\omega}^{+bc} |\mathcal{S}_{n}]+Var[\hat{l}_{\omega}^{-bc}|\mathcal{S}_{n}], \] where \begin{align*} \hat{l}_{\omega}^{+bc} & =\hat{l}_{\omega}^{+}-h_{n}^{p+1}\tilde{B} _{w,p,q}^{+}(h_{n},b_{n}),\\ \hat{l}_{\omega}^{-bc} & =-\hat{l}_{\omega}^{-}+h_{n}^{p+1}\tilde{B} _{w,p,q}^{-}(h_{n},b_{n}),\\ \hat{l}_{\omega}^{+} & =c_{1}^{\prime}(\hat{\mu}_{Y}^{+}-\mu_{Y}^{+} )+c_{2}^{\prime}\left( \hat{\mu}_{X}^{+}-\mu_{X}^{+}\right) ,\\ \hat{l}_{\omega}^{-} & =c_{1}^{\prime}(\hat{\mu}_{Y}^{-}-\mu_{Y}^{-} )+c_{2}^{\prime}\left( \hat{\mu}_{X}^{-}-\mu_{X}^{-}\right) ,\\ \tilde{B}_{w,p,q}^{\pm}(h_{n},b_{n}) & =c_{1}^{\prime}\hat{B}_{Y,p,q}^{\pm }(h_{n},b_{n})+c_{2}^{\prime}\hat{B}_{X,p,q}^{\pm}(h_{n},b_{n}),\\ \hat{B}_{V,p,q}^{\pm}(h_{n},b_{n}) & =\frac{\hat{\mu}_{V,q}^{\pm(p+1)} (b_{n})}{(p+1)!}\mathcal{B}_{p,p+1}^{\pm}(h_{n}). \end{align*} This decomposition implies \[ Var[\hat{l}_{\omega}^{+bc}|\mathcal{S}_{n}]=V_{+1}-2h_{n}^{p+1}C_{+12} +h_{n}^{2p+2}V_{+2}, \] where \begin{align*} C_{+12} & =Cov[\hat{l}_{\omega}^{+},\tilde{B}_{w,p,q}^{+}(h_{n} ,b_{n})|\mathcal{S}_{n}]\\ & =\left( c_{1}^{\prime}\mathcal{C}_{Y,p,q}^{+}c_{1}+c_{1}^{\prime }\mathcal{C}_{YX,p,q}^{+}c_{2}+c_{1}^{\prime}\mathcal{C}_{XY,p,q}^{+} c_{2}+c_{2}^{\prime}\mathcal{C}_{X,p,q}^{+}c_{2}\right) \frac{\mathcal{B} _{p,p+1}^{+}(h_{n})}{(p+1)!}, \end{align*} and \begin{align*} V_{+2} & =Var[\tilde{B}_{w,p,q}^{+}(h_{n},b_{n})|\mathcal{S}_{n}]\\ & =\left( c_{1}^{\prime}\mathcal{V}_{Y,p}^{-(q)}(b_{n})c_{1}+2c_{1}^{\prime }\mathcal{C}_{YX,p,q}^{+}(b_{n})c_{2}+c_{2}^{\prime}\mathcal{V}_{X,p} ^{-(q)}(b_{n})c_{2}\right) \left( \frac{\mathcal{B}_{p,p+1}^{+}(h_{n} )}{(p+1)!}\right) ^{2}. \end{align*} For the left hand side limits the proof is the same. \end{proof}
lemmaUnder Assumption (ref), \[ \left( \mathbf{V}_{\omega,p,q}^{bc}(h_{n},b_{n})\right) ^{-1/2}\left( \hat{\beta}_{\omega}^{bc}-\beta_{\omega}\right) \rightarrow_{d}N\left( 0,1\right) . \] \begin{proof} As in Calonico_Cattaneo_Titiunik it is shown that \begin{align*} \left( \mathbf{V}_{\omega,p,q}^{bc}(h_{n},b_{n})\right) ^{-1/2}\left( \hat{\beta}_{\omega}^{bc}-\beta_{\omega}\right) & =\xi_{1n}+\xi_{2n},\\ & =\xi_{1n}+o_{P}(1), \end{align*} where \[ \xi_{1n}=\left( \mathbf{V}_{\omega,p,q}^{bc}(h_{n},b_{n})\right) ^{-1/2} \hat{l}_{\omega}^{bc}. \] Proceeding as in Lemma (ref) it is shown that $\xi_{1n}\rightarrow _{d}N\left( 0,1\right) $ by Linderberg-Feller central limit theorem for triangular arrays. \end{proof}

\noindentProof of Theorem (ref): It follows from Lemma (ref) for $p=1,$ $q=2$ and $\omega=CW.$ $\blacksquare$

Consistent Standard Error Estimators

In the expressions for $\mathbf{V}_{\omega,p,q}^{bc}(h_{n},b_{n})\ $and $\mathbf{V}_{\omega,p}(h_{n})$ the only unknown quantities are the vectors $c_{1}$ and $c_{2},$ and the matrices $\Psi_{UV,p,q}^{+}(h_{n},b_{n})$ and $\Psi_{UV,p,q}^{-}(h_{n},b_{n}).$ The vectors $c_{1}$ and $c_{2}$ are estimated by $\hat{c}_{1}$ and $\hat{c}_{2},$ respectively. Natural estimators for \[ \Psi_{UV,p,q}^{+}(h_{n},b_{n})=\frac{1}{n}\sum_{i=1}^{n}X_{i,p}X_{i,q} ^{\prime}\sigma_{UV}^{2}(Z_{i},\tilde{W}_{i})k_{ih_{n}}^{+}k_{ib_{n}}^{+} \] and \[ \Psi_{UV,p,q}^{-}(h_{n},b_{n})=\frac{1}{n}\sum_{i=1}^{n}X_{i,p}X_{i,q} ^{\prime}\sigma_{UV}^{2}(Z_{i},\tilde{W}_{i})k_{ih_{n}}^{-}k_{ib_{n}}^{-} \] are \[ \hat{\Psi}_{UV,p,q}^{+}(h_{n},b_{n})=\frac{1}{n}\sum_{i=1}^{n}X_{i,p} X_{i,q}^{\prime}\hat{\varepsilon}_{U_{i}}^{+}\hat{\varepsilon}_{V_{i}} ^{+}k_{ih_{n}}^{+}k_{ib_{n}}^{+} \] and \[ \hat{\Psi}_{UV,p,q}^{-}(h_{n},b_{n})=\frac{1}{n}\sum_{i=1}^{n}X_{i,p} X_{i,q}^{\prime}\hat{\varepsilon}_{U_{i}}^{-}\hat{\varepsilon}_{V_{i}} ^{-}k_{ih_{n}}^{-}k_{ib_{n}}^{-} \] where $\hat{\varepsilon}_{V_{i}}^{+}=V_{i}-\tilde{W}_{i}^{\prime}\hat{\mu }_{V,1}^{+}(h_{n})$ and $\hat{\varepsilon}_{V_{i}}^{-}=V_{i}-\tilde{W} _{i}^{\prime}\hat{\mu}_{V,1}^{-}(h_{n}).$ Using this, a consistent estimator for $\mathbf{V}_{\omega,p}(h_{n})$ is \[ \mathbf{\hat{V}}_{\omega,p}(h_{n})=\hat{V}_{+1}+\hat{V}_{-1}, \] where \[ \hat{V}_{+1}=\hat{c}_{1}^{\prime}\mathcal{\hat{V}}_{Y,p}^{+}(h_{n})\hat{c} _{1}+2c_{1}^{\prime}\hat{C}_{YX,p}^{+}(h_{n})\hat{c}_{2}+\hat{c}_{2}^{\prime }\mathcal{\hat{V}}_{X,p}^{+}(h_{n})\hat{c}_{2}, \] and \[ \hat{V}_{-1}=\hat{c}_{1}^{\prime}\mathcal{\hat{V}}_{Y,p}^{-}(h_{n})\hat{c} _{1}+2\hat{c}_{1}^{\prime}\hat{C}_{YX,p}^{-}(h_{n})\hat{c}_{2}+\hat{c} _{2}^{\prime}\mathcal{\hat{V}}_{X,p}^{-}(h_{n})\hat{c}_{2}. \]

align*[align* omitted — 426 chars of source]

Likewise, we construct a consistent estimator for $\mathbf{V}_{\omega ,p,q}^{bc}(h_{n},b_{n})$ based on plug-in estimated residuals

equation[equation omitted — 159 chars of source]

where $\mathbf{\hat{C}}_{\omega,p,q}^{bc}$ replaces $c_{1}$ and $c_{2}$ by $\hat{c}_{1}$ and $\hat{c}_{2}$ and $\Psi_{UV,p,q}^{+}(h_{n},b_{n})$ and $\Psi_{UV,p,q}^{-}(h_{n},b_{n})$ by $\hat{\Psi}_{UV,p,q}^{+}(h_{n},b_{n})$ and $\hat{\Psi}_{UV,p,q}^{-}(h_{n},b_{n}),$ respectively.

lemmaUnder Assumption (ref), \[ \frac{\mathbf{\hat{V}}_{R,1,2}^{bc}(h_{n},b_{n})}{\mathbf{V}_{R,1,2} ^{bc}(h_{n},b_{n})}\rightarrow_{p}1. \] \begin{proof} We first show that \begin{equation} \hat{\Psi}_{UV,p,q}^{+}(h_{n},b_{n})=\Psi_{UV,p,q}^{+}(h_{n},b_{n} )+o_{P}\left( \frac{m_{n}}{h_{n}b_{n}}\right) , \end{equation} where $m_{n}=\min\{h_{n},b_{n}\}.$ Standard calculations show $\hat {\varepsilon}_{V_{i}}^{+}=\varepsilon_{V_{i}}^{+}-\tilde{W}_{i}^{\prime }\left( \hat{\mu}_{V,1}^{+}(h_{n})-\mu_{V,1}^{+}\right) $ and by previous results \begin{align*} \hat{\Psi}_{UV,p,q}^{+}(h_{n},b_{n}) & =\frac{1}{n}\sum_{i=1}^{n} X_{i,p}X_{i,q}^{\prime}\varepsilon_{U_{i}}^{+}\varepsilon_{V_{i}}^{+} k_{ih_{n}}^{+}k_{ib_{n}}^{+}\\ & -\left[ \frac{1}{n}\sum_{i=1}^{n}X_{i,p}X_{i,q}^{\prime}\varepsilon _{U_{i}}^{+}\tilde{W}_{i}^{\prime}k_{ih_{n}}^{+}k_{ib_{n}}^{+}\right] \left( \hat{\mu}_{V,1}^{+}(h_{n})-\mu_{V,1}^{+}\right) \\ & -\left[ \frac{1}{n}\sum_{i=1}^{n}X_{i,p}X_{i,q}^{\prime}\varepsilon _{V_{i}}^{+}\tilde{W}_{i}^{\prime}k_{ih_{n}}^{+}k_{ib_{n}}^{+}\right] \left( \hat{\mu}_{U,1}^{+}(h_{n})-\mu_{U,1}^{+}\right) \\ & +\left( \hat{\mu}_{U,1}^{+}(h_{n})-\mu_{U,1}^{+}\right) ^{\prime}\left[ \frac{1}{n}\sum_{i=1}^{n}X_{i,p}X_{i,q}^{\prime}\tilde{W}_{i}\tilde{W} _{i}^{\prime}k_{ih_{n}}^{+}k_{ib_{n}}^{+}\right] \left( \hat{\mu}_{V,1} ^{+}(h_{n})-\mu_{V,1}^{+}\right) \\ & =\frac{1}{n}\sum_{i=1}^{n}X_{i,p}X_{i,q}^{\prime}\varepsilon_{U_{i}} ^{+}\varepsilon_{V_{i}}^{+}k_{ih_{n}}^{+}k_{ib_{n}}^{+}+o_{P}\left( \frac{m_{n}}{h_{n}b_{n}}\right) \\ & \equiv\breve{\Psi}_{UV,p,q}^{+}(h_{n},b_{n})+o_{P}\left( \frac{m_{n} }{h_{n}b_{n}}\right) , \end{align*} By the change of variables $u=Z/h_{n},$ it follows that \begin{align*} & \mathbb{E}\left[ \frac{h_{n}b_{n}}{m_{n}}\breve{\Psi}_{UV,p,q}^{+}(h_{n} ,b_{n})\right] \\ & =\mathbb{E}\left[ \frac{h_{n}b_{n}}{m_{n}}\frac{1}{n}\sum_{i=1}^{n}X_{i,p} X_{i,q}^{\prime}\sigma_{UV}^{2}(Z_{i},\tilde{W}_{i})k_{ih_{n}}^{+}k_{ib_{n} }^{+}\right] \\ & =\tilde{\Psi}_{UV,p,q}^{+}(h_{n},b_{n}). \end{align*} Also, since $\sigma_{jg}^{2,\varepsilon_{U_{i}}\varepsilon_{V_{i}}}(z)$ is bounded, \begin{align*} & \mathbb{E}\left[ \left\vert \frac{h_{n}b_{n}}{m_{n}}\breve{\Psi}_{UV,p}^{+} (h_{n})-\mathbb{E}[\frac{h_{n}b_{n}}{m_{n}}\breve{\Psi}_{UV,p}^{+}(h_{n})]\right\vert ^{2}\right] \\ & \leq Cn^{-1}m_{n}^{-1}\int_{0}^{\infty}k^{2}\left( \frac{m_{n}u}{h_{n} }\right) k^{2}\left( \frac{m_{n}u}{b_{n}}\right) \left\vert r_{p}\left( \frac{m_{n}u}{h_{n}}\right) \right\vert ^{2}\left\vert r_{q}\left( \frac{m_{n}u}{b_{n}}\right) \right\vert ^{2}f(uh_{n})du\\ & =O\left( n^{-1}m_{n}^{-1}\right) . \end{align*} Hence with $m_{n}\rightarrow0$ and $nm_{n}\rightarrow\infty,$ \[ \frac{h_{n}b_{n}}{m_{n}}\breve{\Psi}_{UV,p,q}^{+}(h_{n},b_{n})=\tilde{\Psi }_{UV,p}^{+}(h_{n},b_{n})+o_{P}(1). \] Then, ((ref)) follows from Lemma (ref). Simple algebra shows the rest of the proof, using that \[ \hat{c}_{1}-c_{1}=O_{P}\left( \frac{1}{\sqrt{nh_{n}}}+h_{n}^{(p+1)}\right) \] and \[ \hat{c}_{2}-c_{2}=O_{P}\left( \frac{1}{\sqrt{nh_{n}}}+h_{n}^{(p+1)}\right) . \] \end{proof}

MSE-Optimal Bandwidth Selectors

We consider first the MSE optimal bandwidth selector for estimating the bias term, i.e. $b_{n}.$ Consider the general case where we use a $q-$th order polynomial for estimating the $p+1$ derivative, $q>p$. The leading term in the bias is

align*[align* omitted — 287 chars of source]

and $\mathcal{B}_{p,q}^{+}=e_{0}^{\prime}\Gamma_{+,p}^{-1}\vartheta_{p,q}^{+}$ and $\mathcal{B}_{p,q}^{-}=e_{0}^{\prime}\Gamma_{-,p}^{-1}\vartheta_{p,q} ^{-}.$ With a slightly abuse of notation denote the true bias

equation[equation omitted — 110 chars of source]

where \[ B_{V,p,q}=\frac{\mu_{V}^{+(p+1)}}{(p+1)!}\mathcal{B}_{p,q}^{+}-\frac{\mu _{V}^{-(p+1)}}{(p+1)!}\mathcal{B}_{p,q}^{-}. \]

lemmaUnder Assumption (ref), \begin{align*} \mathbf{MSE}(b_{n}) & \equiv \mathbb{E}[\left( \bar{B}_{w,p,q}(b_{n})-\mathbf{B} _{\omega,p,p+1}\right) ^{2}|\mathcal{S}_{n}]\\ & =b_{n}^{q-p}[B_{\omega,p,q}^{bc}+o_{P}(1)]\\ & +n^{-1}b_{n}^{-1-2(p+1)}[V_{\omega,p,q}^{bc}+o_{P}(1)], \end{align*} where \[ B_{\omega,p,q}^{bc}=c_{1}^{\prime}B_{\omega,p,q}^{Y,bc}+c_{2}^{\prime }B_{\omega,p,q}^{X,bc} \] \[ B_{\omega,p,q}^{V,bc}=\frac{\mu_{V}^{+(q+1)}}{(q+1)!}\frac{\mathcal{B} _{p+1,q,q+1}^{+}\mathcal{B}_{p,p+1}^{+}}{(p+1)!}-\frac{\mu_{V}^{-(q+1)} }{(q+1)!}\frac{\mathcal{B}_{p+1,q,q+1}^{-}\mathcal{B}_{p,p+1}^{-}}{(p+1)!} \] and \[ V_{\omega,p,q}^{bc}=V_{\omega,p,q}^{+bc}+V_{\omega,p,q}^{-bc} \] where \begin{align*} V_{\omega,p,q}^{+bc} & =\left( c_{1}^{\prime}\mathcal{V}_{Y,q}^{+(p+1)} c_{1}+2c_{1}^{\prime}\mathcal{C}_{YX,p,q}^{+}c_{2}+c_{2}^{\prime} \mathcal{V}_{X,q}^{+(p+1)}c_{2}\right) \left( \frac{\mathcal{B}_{p,p+1}^{+} }{(p+1)!}\right) ^{2},\\ \mathcal{V}_{V,q}^{+(p+1)} & =\left( (p+1)!\right) ^{2}e_{p+1}^{\prime }\Gamma_{+,q}^{-1}\Psi_{V,q}^{+}\Gamma_{+,q}^{-1}e_{p+1},\\ \mathcal{C}_{YX,p,q}^{+} & =\left( (p+1)!\right) ^{2}e_{p+1}^{\prime} \Gamma_{+,q}^{-1}\Psi_{YX,q}^{+}\Gamma_{+,q}^{-1}e_{p+1} \end{align*} and \begin{align*} V_{\omega,p,q}^{-bc} & =\left( c_{1}^{\prime}\mathcal{V}_{Y,q}^{-(p+1)} c_{1}+2c_{1}^{\prime}\mathcal{C}_{YX,p,q}^{-}c_{2}+c_{2}^{\prime} \mathcal{V}_{X,q}^{-(p+1)}c_{2}\right) \left( \frac{\mathcal{B}_{p,p+1}^{-} }{(p+1)!}\right) ^{2},\\ \mathcal{V}_{V,q}^{-(p+1)} & =\left( (p+1)!\right) ^{2}e_{p+1}^{\prime }\Gamma_{-,q}^{-1}\Psi_{V,q}^{-}\Gamma_{-,q}^{-1}e_{p+1},\\ \mathcal{C}_{YX,p,q}^{-} & =\left( (p+1)!\right) ^{2}e_{p+1}^{\prime} \Gamma_{-,q}^{-1}\Psi_{YX,q}^{-}\Gamma_{-,q}^{-1}e_{p+1}. \end{align*} The optimal MSE bandwidth has the form \begin{align*} b_{n}^{\ast} & =C_{MSEb,p}^{1/(2q+3)}n^{-1/(2q+3)}\\ C_{MSEb,p} & =\frac{(2p+3)V_{\omega,p,q}^{bc}}{2(q-p)\left( B_{\omega ,p,q}^{bc}\right) ^{2}}. \end{align*} \begin{proof} Note that by Lemma (ref) \begin{align*} \mathbb{E}[\big( \bar{B}_{w,p,q}(b_{n})-&\mathbf{B}_{\omega,p,p+1}\big) |\mathcal{S}_{n}] =c_{1}^{\prime}\mathbb{E}[\hat{B}_{Y,p,q}(b_{n})-B_{Y,p,p+1} |\mathcal{S}_{n}]+c_{2}^{\prime}\mathbb{E}[\hat{B}_{X,p,q}(b_{n})-B_{X,p,p+1} |\mathcal{S}_{n}]\\ & =c_{1}^{\prime}\mathbb{E}[e_{p+1}^{\prime}\hat{\beta}_{Y,q}^{+}(b_{n})-e_{p+1} ^{\prime}\beta_{Y}^{+}|\mathcal{S}_{n}]\mathcal{B}_{p,p+1}^{+}-c_{1}^{\prime }\mathbb{E}[e_{p+1}^{\prime}\hat{\beta}_{Y,q}^{-}(b_{n})-e_{p+1}^{\prime}\beta_{Y} ^{-}|\mathcal{S}_{n}]\mathcal{B}_{p,p+1}^{-}\\ & +c_{2}^{\prime}\mathbb{E}[e_{p+1}^{\prime}\hat{\beta}_{X,q}^{+}(b_{n})-e_{p+1} ^{\prime}\beta_{X}^{+}|\mathcal{S}_{n}]\mathcal{B}_{p,p+1}^{+}-c_{2}^{\prime }\mathbb{E}[e_{p+1}^{\prime}\hat{\beta}_{X,q}^{-}(b_{n})-e_{p+1}^{\prime}\beta_{X} ^{-}|\mathcal{S}_{n}]\mathcal{B}_{p,p+1}^{-}\\ & =c_{1}^{\prime}b_{n}^{p-q}\left[ \frac{\mu_{Y}^{+(q+1)}}{(q+1)!} \frac{\mathcal{B}_{p+1,q,q+1}^{+}\mathcal{B}_{p,p+1}^{+}}{(p+1)!}-\frac {\mu_{Y}^{-(q+1)}}{(q+1)!}\frac{\mathcal{B}_{p+1,q,q+1}^{-}\mathcal{B} _{p,p+1}^{-}}{(p+1)!}+o_{P}(1)\right] \\ & +c_{2}^{\prime}b_{n}^{p-q}\left[ \frac{\mu_{X}^{+(q+1)}}{(q+1)!} \frac{\mathcal{B}_{p+1,q,q+1}^{+}\mathcal{B}_{p,p+1}^{+}(h_{n})}{(p+1)!} -\frac{\mu_{X}^{-(q+1)}}{(q+1)!}\frac{\mathcal{B}_{p+1,q,q+1}^{-} \mathcal{B}_{p,p+1}^{-}(h_{n})}{(p+1)!}+o_{P}(1)\right] \\ & =b_{n}^{q-p}\left[ B_{\omega,p,q}^{bc}+o_{P}(1)\right] . \end{align*} As for the variance, write \begin{align*} \bar{B}_{w,p,q}(b_{n}) & =\bar{B}_{w,p,q}^{+}(b_{n})-\bar{B}_{w,p,q} ^{-}(b_{n}),\\ \bar{B}_{w,p,q}^{+}(b_{n}) & =c_{1}^{\prime}\hat{B}_{Y,p,q}^{+}(b_{n} )+c_{2}^{\prime}\hat{B}_{X,p,q}^{+}(b_{n}),\\ \bar{B}_{w,p,q}^{-}(b_{n}) & =c_{1}^{\prime}\hat{B}_{Y,p,q}^{-}(b_{n} )+c_{2}^{\prime}\hat{B}_{X,p,q}^{-}(b_{n}),\\ \bar{B}_{V,p,q}^{+}(b_{n}) & =\frac{\hat{\mu}_{V,q}^{+(p+1)}(b_{n})} {(p+1)!}\mathcal{B}_{p,p+1}^{+},\\ \hat{B}_{V,p,q}^{-}(b_{n}) & =\frac{\hat{\mu}_{V,q}^{-(p+1)}(b_{n})} {(p+1)!}\mathcal{B}_{p,p+1}^{-}. \end{align*} This decomposition and Lemma (ref) imply \[ Var[\bar{B}_{w,p,q}(b_{n})|\mathcal{S}_{n}]=n^{-1}b_{n}^{-1-2(p+1)}\left[ V_{\omega,p,q}^{+bc}+V_{\omega,p,q}^{-bc}+o_{P}(1)\right] , \] where \begin{align*} V_{\omega,p,q}^{+bc} & =\left( c_{1}^{\prime}\mathcal{V}_{Y,q}^{+(p+1)} c_{1}+2c_{1}^{\prime}\mathcal{C}_{YX,p,q}^{+}c_{2}+c_{2}^{\prime} \mathcal{V}_{X,q}^{+(p+1)}c_{2}\right) \left( \frac{\mathcal{B}_{p,p+1}^{+} }{(p+1)!}\right) ^{2},\\ \mathcal{V}_{V,q}^{+(p+1)} & =\left( (p+1)!\right) ^{2}e_{p+1}^{\prime }\Gamma_{+,q}^{-1}\Psi_{V,q}^{+}\Gamma_{+,q}^{-1}e_{p+1},\\ \mathcal{C}_{YX,p,q}^{+} & =\left( (p+1)!\right) ^{2}e_{p+1}^{\prime} \Gamma_{+,q}^{-1}\Psi_{YX,q}^{+}\Gamma_{+,q}^{-1}e_{p+1} \end{align*} and similarly \begin{align*} V_{\omega,p,q}^{-bc} & =\left( c_{1}^{\prime}\mathcal{V}_{Y,q}^{-(p+1)} c_{1}+2c_{1}^{\prime}\mathcal{C}_{YX,p,q}^{-}c_{2}+c_{2}^{\prime} \mathcal{V}_{X,q}^{-(p+1)}c_{2}\right) \left( \frac{\mathcal{B}_{p,p+1}^{-} }{(p+1)!}\right) ^{2},\\ \mathcal{V}_{V,q}^{-(p+1)} & =\left( (p+1)!\right) ^{2}e_{p+1}^{\prime }\Gamma_{-,q}^{-1}\Psi_{V,q}^{-}\Gamma_{-,q}^{-1}e_{p+1},\\ \mathcal{C}_{YX,p,q}^{-} & =\left( (p+1)!\right) ^{2}e_{p+1}^{\prime} \Gamma_{-,q}^{-1}\Psi_{YX,q}^{-}\Gamma_{-,q}^{-1}e_{p+1}. \end{align*} \end{proof}

The following MSE calculations for the linearized term in the WLATE estimators have been obtained before, and are summarized in the following result. Recall \[ \hat{l}_{\omega}=c_{1}^{\prime}(\hat{\delta}_{Y}-\delta_{Y})+c_{2}^{\prime }\left( \hat{\delta}_{X}-\delta_{X}\right) , \] where $c_{1}=c/\tau_{X}$, $c_{2}^{\prime}=\left[ \delta_{Y}^{\prime} G_{\delta}-\beta_{w}(c^{\prime}+\delta_{X}^{\prime}G_{\delta})\right] /\tau_{X}$ and $G_{\delta}$ is the derivative of $g$ with respect to $\delta_{X}$.

lemmaUnder Assumption (ref)(1)-(7), \begin{align*} \mathbf{MSE}(h_{n}) & \equiv \mathbb{E}[\hat{l}_{\omega}^{2}|\mathcal{S}_{n}]\\ & =h_{n}^{2(p+1)}[B_{\omega,p,p+1}+o_{P}(1)]\\ & +n^{-1}h_{n}^{-1}[V_{\omega,p}+o_{P}(1)], \end{align*} where $B_{\omega,p,p+1}$ is defined in ((ref)) and \begin{align*} V_{\omega,p} & =V_{+1}+V_{-1},\\ V_{\pm1} & =c_{1}^{\prime}\mathcal{V}_{Y,p}^{\pm}c_{1}+2c_{1}^{\prime }\mathcal{C}_{YX,p}^{\pm}c_{2}+c_{2}^{\prime}\mathcal{V}_{X,p}^{\pm}c_{2},\\ \mathcal{C}_{YX,p}^{\pm} & =e_{0}^{\prime}\Gamma_{\pm,p}^{-1}\Psi _{YX,p}^{\pm}\Gamma_{\pm,p}^{-1}e_{0},\\ \mathcal{V}_{V,p}^{+} & =\mathcal{C}_{VV,p}^{\pm}. \end{align*} The optimal MSE bandwidth has the form \begin{align*} h_{n}^{\ast} & =C_{MSEh,p}^{1/(2p+3)}n^{-1/(2p+3)}\\ C_{MSEh,p} & =\frac{V_{\omega,p}}{2(p+1)B_{\omega,p}^{2}}. \end{align*} \begin{proof} It follows from Lemmas (ref) and (ref). \end{proof}

Implementation of MSE-Optimal Bandwidth Selectors

In this section we describe the implementation of MSE-optimal bandwidths for local linear estimation $(p=1)$. We consider the following steps.

description• Choose a preliminary pilot bandwidth $c_{n}$ as \begin{align*} c_{n} & =C_{K}\times C_{sd}\times n^{-1/5},\\ C_{K} & =\left( \frac{8\sqrt{\pi}\int k^{2}(u)du}{3\left( \int u^{2}k(u)du\right) ^{2}}\right) ^{1/5}\\ C_{sd} & =\min\left( s,\frac{IQR}{1.349}\right) , \end{align*} where $s^{2}$ and $IQR$ are the sample variance and interquartile range of $\{Z_{i}\}_{i=1}^{n}.$ The constant $C_{K}$ is 1.843 when $K$ is the uniform kernel, and 2.576 when $K$ is the triangular kernel. • Choose an optimal bandwidth for bias estimation $b_{n}$ following Lemma (ref) with $q=p+1=2$: \begin{align*} \hat{b}_{n}^{\ast} & =\hat{C}_{MSEb,1}^{1/7}n^{-1/7}\\ \hat{C}_{MSEb,1} & =\frac{5\hat{V}_{\omega,1,2}^{bc}(c_{n})}{2\left( \hat {B}_{\omega,1,2}^{bc}(d_{n})\right) ^{2}}. \end{align*} The variance estimate $\hat{V}_{\omega,1,2}^{bc}(c_{n})$ is \[ \hat{V}_{\omega,1,2}^{bc}(c_{n})=\hat{V}_{\omega,1,2}^{+bc}(c_{n})+\hat {V}_{\omega,1,2}^{-bc}(c_{n}) \] where (with $p=1$ and $q=2)$ \begin{align*} \hat{V}_{\omega,p,q}^{\pm bc}(c_{n}) & =\left( \hat{c}_{1}^{\prime }\mathcal{V}_{Y,q}^{+(p+1)}(c_{n})\hat{c}_{1}+2\hat{c}_{1}^{\prime} \mathcal{C}_{YX,p,q}^{+}(c_{n})\hat{c}_{2}+\hat{c}_{2}^{\prime}\mathcal{V} _{X,q}^{+(p+1)}(c_{n})\hat{c}_{2}\right) \left( \frac{\mathcal{B} _{p,p+1}^{\pm}(c_{n})}{(p+1)!}\right) ^{2}\!,\\ \mathcal{V}_{V,q}^{\pm(p+1)}(c_{n}) & =\left( (p+1)!\right) ^{2} e_{p+1}^{\prime}\Gamma_{\pm,q}^{-1}(c_{n})\hat{\Psi}_{V,q}^{\pm}(c_{n} )\Gamma_{\pm,q}^{-1}(c_{n})e_{p+1},\\ \mathcal{C}_{YX,p,q}^{\pm}(c_{n}) & =\left( (p+1)!\right) ^{2} e_{p+1}^{\prime}\Gamma_{\pm,q}^{-1}(c_{n})\hat{\Psi}_{YX,q}^{\pm}(c_{n} )\Gamma_{\pm,q}^{-1}(c_{n})e_{p+1}. \end{align*} The bias term $\hat{B}_{\omega,1,2}^{bc}(d_{n})$ is obtained as (with $p=1$ and $q=2)$ \[ \hat{B}_{\omega,p,q}^{bc}(d_{n})=\hat{c}_{1}^{\prime}\hat{B}_{\omega ,p,q}^{Y,bc}(d_{n})+\hat{c}_{2}^{\prime}\hat{B}_{\omega,p,q}^{X,bc}(d_{n}) \] \[ \hat{B}_{\omega,p,q}^{V,bc}(d_{n})=\frac{\hat{\mu}_{V}^{+(q+1)}(d_{n} )}{(q+1)!}\frac{\mathcal{B}_{p+1,q,q+1}^{+}(c_{n})\mathcal{B}_{p,p+1} ^{+}(c_{n})}{(p+1)!}-\frac{\hat{\mu}_{V}^{-(q+1)}(d_{n})}{(q+1)!} \frac{\mathcal{B}_{p+1,q,q+1}^{-}(c_{n})\mathcal{B}_{p,p+1}^{-}(c_{n} )}{(p+1)!}, \] where $d_{n}$ is the optimal bandwidth for estimating the third derivatives $\mu_{V}^{\pm(q+1)}$ with a third order local polynomial. This optimal bandwidth is obtained for each variable $V$ and $\pm$ sides as (with $q=2)$ \begin{align*} d_{V,n}^{\pm} & =\hat{C}_{V,\pm,q+1}^{1/(2q+5)}n^{-1/(2q+5)}\\ \hat{C}_{V,\pm,q+1}^{1/(2q+5)} & =\frac{(2q+3)\hat{V}_{V,q+1}^{\pm (q+1)}(c_{n})}{2\left( \hat{B}_{q+1,q+1,q+2}^{\pm}(c_{n})\right) ^{2}}, \end{align*} where \[ \hat{V}_{V,q+1}^{\pm(q+1)}(c_{n})=e_{q+1}^{\prime}\Gamma_{+,q+1}^{-1} (c_{n})\hat{\Psi}_{V,q}^{\pm}(c_{n})\Gamma_{+,q+1}^{-1}(c_{n})e_{q+1} \] and \[ \hat{B}_{q+1,q+1,q+2}^{\pm}(c_{n})=\frac{\hat{\mu}_{V}^{+(q+2)}(c_{n} )}{(q+2)!}\mathcal{B}_{q+1,q+1,q+2}^{+}(c_{n}). \] • Choose optimal bandwidth $h_{n}$ for the estimator: use $c_{n}$ and $b_{n}$ to compute for $p=1$ \begin{align*} h_{n}^{\ast} & =\hat{C}_{MSEh,p}^{1/(2p+3)}n^{-1/(2p+3)}\\ \hat{C}_{MSEh,p} & =\frac{\hat{V}_{\omega,p}(c_{n})}{2(p+1)\hat{B}_{\omega ,p}^{2}(b_{n})}, \end{align*} where \begin{align*} \hat{V}_{\omega,p}(c_{n}) & =\hat{V}_{+1}(c_{n})+\hat{V}_{-1}(c_{n}),\\ \hat{V}_{\pm1}(c_{n}) & =\hat{c}_{1}^{\prime}\mathcal{V}_{Y,p}^{\pm} (c_{n})\hat{c}_{1}+2\hat{c}_{1}^{\prime}\mathcal{C}_{YX,p}^{\pm}(c_{n})\hat {c}_{2}+\hat{c}_{2}^{\prime}\mathcal{V}_{X,p}^{\pm}(c_{n})\hat{c}_{2},\\ \mathcal{C}_{YX,p}^{\pm}(c_{n}) & =e_{0}^{\prime}\Gamma_{\pm,p}^{-1} (c_{n})\hat{\Psi}_{YX,p}^{\pm}(c_{n})\Gamma_{\pm,p}^{-1}(c_{n})e_{0},\\ \mathcal{V}_{V,p}^{+}(c_{n}) & =\mathcal{C}_{VV,p}^{\pm}(c_{n}). \end{align*} and \[ \hat{B}_{\omega,p,q}(c_{n},b_{n})=\hat{c}_{1}^{\prime}\hat{B}_{Y,p,q} (c_{n},b_{n})+\hat{c}_{2}^{\prime}\hat{B}_{X,p,q}(c_{n},b_{n}), \] where \[ \hat{B}_{V,p,q}(c_{n},b_{n})=\frac{\hat{\mu}_{V,q}^{+(p+1)}(b_{n})} {(p+1)!}\mathcal{B}_{p,p+1}^{+}(c_{n})-\frac{\hat{\mu}_{V,q}^{-(p+1)}(b_{n} )}{(p+1)!}\mathcal{B}_{p,p+1}^{-}(c_{n}). \]

Additional Monte Carlo Simulations

In this appendix we consider a few other simulation scenarios.

More Intense Treatment Effect Heterogeneity

In Table (ref), we consider the case where $\beta_{XW}=5$ or $\beta_{XW}=10$ (instead of $\beta_{XW}=0$ or $\beta_{XW}=2$, as in Table (ref)). All other parameters from Section (ref) are kept the same. The increased heterogeneity in treatment effects lead to all estimators perform worse, with similar relative results between the CWLATE and the standard RDD estimators.

table[table omitted — 1,627 chars of source]

Coarsened $W_i$

Next, we consider what happens when more complex covariates are used to implement the CWLATE estimator. We use the same Monte Carlo set up discussed in Section (ref), but with $W_i\in \{-5/3,-4/3,$ $-1,-2/3,-1/3,1/3,2/3,1,4/3,$ $5/3\}$, in contrast to the binary covariate $W_i \in\{-1,1\}$ from Section (ref).

We consider two separate scenarios: (a) CWLATE is estimated with $\tilde{W}_i=W_i$, and (b) CWLATE is estimated with a coarsened binary version of the covariate ($\tilde{W}_i=1$ if $W_i>0,$ and $\tilde{W}_i=-1$ if $W_i<0$). As in the baseline case, the term representing the heterogeneity in the first-stages across different values of $W_i$ is $\alpha_{DW}(W_i)=W_i\cdot\alpha_{DW}$. Analogously, the term representing the heterogeneity in the treatment effects across different values of $W_i$ is $\beta_{XW}(W_i)=W_i\cdot\beta_{XW}$.

Table (ref) reports the MSE when CWLATE is estimated with the full covariate taking 10 different values. The results are similar to the results in Table (ref), with the small difference due to increased heterogeneity when either $\alpha_{DW}\neq0$ or $\beta_{XW}\neq0$, since now the range of $W_i$ increased from $[-1,1]$ to $[-5/3,5/3]$.

table[table omitted — 1,535 chars of source]

Table (ref) reports the MSE when CWLATE is estimated with the coarsened (binary) covariate. The only columns that are different from Table (ref) are the CWLATE ones, since $\hat{\beta}^{cov}_U$ is estimated with the full vector $W_i$ in both tables. The results in both simulations are similar, with a small advantage when using the full covariate vector $W_i,$ especially in smaller samples.

table[table omitted — 1,543 chars of source]

In our simulations, using a larger covariate set does not seem to be detrimental, and may in fact be beneficial (see also results in our application). At the same time, using a coarsened covariate in the CWLATE estimator had minimal impact on precision and seems to be a viable alternative.

Estimating $\beta_U$ with the CWLATE Estimator

In this section, we compare the CWLATE estimator to the standard RDD estimators when all target the unconditional LATE $\beta_U$. Note that, in the cases we examine, the CWLATE estimator is an inconsistent estimator of $\beta_U$ (since it is an unbiased estimator of $\beta_{CW}\neq \beta_U$). The results can be seen in Table (ref). We consider the case where $\alpha_{DW}=1$ and either $\beta_{XW}=2$ or $\beta_{XW}=10$. The differences between the estimands increase with $\beta_{XW},$ so the CWLATE is a more severely biased estimator of $\beta_U$ when $\beta_{XW}=10$. Nevertheless, the results show that not only is the MSE of the CWLATE estimator generally lower than that of the standard RDD estimators, but the comparatively better performance intensifies with the higher $\beta_{XW}$ value for moderate sample sizes.

table[table omitted — 1,939 chars of source]
figure[figure omitted — 673 chars of source]

The relative advantage of the CWLATE even when biased is due to the excessively large variance of the standard RDD estimators. The standard RDD estimators show signs of weak identification for the small samples, while the CWLATE's variance remains relatively low. Figure (ref) shows that, even for moderate samples, the variance of the standard RDD estimators is excessively high. A sample of 5,000 was necessary for the variance to be sufficiently low that the standard RDD with covariates could surpass the biased CWLATE for $\beta_{XW}=10.$ Note that for $\beta_{XW}=2,$ this level is not reached for any of the sample sizes considered.

These results do not mean that the CWLATE is a better estimator of $\beta_U$. Although technically it seems to be superior for point estimation (since it has a smaller MSE), it is not advisable to perform inference with a biased estimator because confidence interval coverages and test sizes can be wrong. In the right two panels of the table, we show that the coverage rate of the 95% confidence interval using the CWLATE estimator can be seriously off the mark.

Rather, the point of this exercise is to show that the issues of excessive variance should not be overlooked. The fact that an inconsistent estimator can perform better in such a wide range of scenarios is a cautionary tale against the standard RDD as an estimand. In many situations, such as in our application, the CWLATE can be a better target to substantiate causal claims.