EconBase
← Back to paper

Permutation Tests at Nonparametric Rates

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

284,758 characters · 11 sections · 67 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Permutation Tests at Nonparametric Rates

\ifblind \else \fi

abstractClassical two-sample permutation tests for equality of distributions have exact size in finite samples, but they fail to control size for testing equality of parameters that summarize each distribution. This paper proposes permutation tests for equality of parameters that are estimated at root-$n$ or slower rates. Our general framework applies to both parametric and nonparametric models, with two samples or one sample split into two subsamples. Our tests have correct size asymptotically while preserving exact size in finite samples when distributions are equal. They have no loss in local asymptotic power compared to tests that use asymptotic critical values. We propose confidence sets with correct coverage in large samples that also have exact coverage in finite samples if distributions are equal up to a transformation. We apply our theory to four commonly-used hypothesis tests of nonparametric functions evaluated at a point. Lastly, simulations show good finite sample properties, and two empirical examples illustrate our tests in practice.

Keywords: Randomization Inference, Hypothesis Tests, Confidence Sets.

Introduction

Applications of permutation tests have gained widespread popularity in empirical analyses in the social and natural sciences. Classical two-sample permutation tests appeal to applied researchers because they are easy to implement and have exact size in finite samples under the so-called “sharp null hypothesis”. The sharp null hypothesis states that the two population distributions are equal. However, researchers are often interested in testing equality of parameters that summarize the distributions. For example, one may want to test equality of average outcomes between treatment and control groups while nonparametrically controlling for age and income. Classical permutation tests fail to control size under such nulls, in both finite and large samples.\footnote{The lack of size control of the classical permutation test outside of the sharp null has been studied for a long time (e.g., romano1990). Theorem (ref) below confirms the lack of size control in our setting. }

This paper proposes robust two-sample permutation tests for equality of parameters that are estimated at root-$n$ or slower rates. The tests are robust in the sense that they control size asymptotically while preserving finite-sample exactness under the sharp null. Our general framework covers both parametric and nonparametric models, in cases with two samples from two populations or one sample from a union of two populations. In addition, the paper makes three further contributions. First, we derive the asymptotic permutation distribution in our general framework under both null and alternative hypotheses, which requires novel technical arguments. Second, we provide four examples of tests in widely-used nonparametric models, and we prove that they satisfy the conditions of our framework. Third, we construct robust confidence sets for differences between the two populations. The confidence sets are robust meaning that they have correct coverage asymptotically and exact coverage in finite samples if the sharp null holds under a class of transformations.

Our framework considers a summary parameter that can be consistently estimated using an asymptotically linear statistic. The influence function depends on the data, the population distribution, and the sample size. There may be two iid samples from two populations, or one iid sample from a union of two populations. In the case of one sample, there is a variable in the data that identifies the population of each observation, and the sample is split into two. The researcher applies the estimator to each of the two samples, computes the difference, and tests whether the two parameters are equal. The classical permutation test compares the estimated difference to critical values from the permutation distribution, that is, the distribution of estimates over all permutations of observations across the two samples.

We derive the asymptotic permutation distribution of the estimated difference, both with and without the null hypothesis, and find it to be generally different from its asymptotic sampling distribution. This leads the classical permutation test to have incorrect size, as shown in a variety of other settings (see related literature below). The derivation has two key technical features. First, we borrow the coupling approximation from chung2013 (henceforth CR) and apply it to our setting. The approximation consists of two steps: (a) couple the original sample with a new sample from a particular mixture of the two populations; and (b) prove that replacing the original sample with the new coupled sample renders no change in the asymptotic permutation distribution of the test statistic. We utilize the coupling approximation because it is more tractable to study the limiting behavior of permutation tests when both samples are drawn iid from the same mixture distribution, rather than different distributions. Our coupling approximation differs from that of CR in that our proof does not assume the null hypothesis. This allows us to study both size and power of the permutation test. Our proof requires a new argument to bound the variance of the approximation error (Section B.2.3 in the supplement). The second key technical feature is that our theory allows for random sample sizes. Sample sizes are random when the researcher splits one sample into two as a function of the data. Thus, the derivation of the limiting distribution must be valid conditional on any sequence of sample splits that occurs with probability one and requires novel arguments: (a) a conditional central limit theorem of weighted sums of triangular arrays, where the weights are random (Lemmas F.3 and F.4 in the supplement); and (b) an asymptotic linear representation for estimators that must hold uniformly over convex combinations of the two populations (Assumption (ref)).

Our proposed permutation test uses a studentized test statistic, which is the estimated difference of parameters divided by a consistent estimator of its standard deviation. We then show that both the asymptotic permutation and sampling distributions are standard normal. It follows that our permutation test has correct size in large samples, and its asymptotic power against local alternatives is identical to that of the test that relies on critical values from a standard normal. Finally, we construct a confidence set by inverting our test, which requires testing null hypotheses that are more general than simple equality of parameters. We propose ways to transform the data in order to test more general hypotheses and preserve finite sample exactness when populations are equal up to the data transformations.

Examples of applications of permutation tests abound. Table 3 in the supplement lists top-publications from the last decade in a variety of disciplines that use permutation tests. This broad applicability motivates our extension of the theory. We illustrate our framework using four nonparametric examples of hypothesis tests that are often used in empirical studies. The first and second examples test equality at a point of nonparametric conditional mean and quantile functions, respectively. The third and fourth examples test continuity at a point of nonparametric conditional mean or probability density functions (PDFs), respectively. We explain how to implement the permutation test in each case and give sufficient conditions to derive the limiting permutation distribution. Implementation requires a sign change and sample splitting in the third and fourth examples. We find that asymptotic size control requires studentization, except in the fourth example.

{\it Related Literature}

Randomization inference has recently received keen attention canay2017sym,shaikh2021. For the case of permutation inference, the insight of robustness through studentization has been proposed before in specific settings that differ from our general framework: neuhaus1993, janssen1997, janssen2005, neubert2007, chung2013, pauly2015, chung2016, and diciccio2017. In particular, the framework of janssen1997 applies to the problem of testing difference of means, while CR study the more general problem of difference of parameters that are estimable at root-$n$. It is important to emphasize that this paper is not a straightforward generalization of their work, and none of our applications fits CR's framework. For example, CR handles the difference of parametric quantiles for two independent samples; in contrast, our framework covers the cases of parametric or nonparametric quantiles, for two independent samples or one sample randomly split into two. Moreover, our verification of the coupling approximation differs from that of CR because we do not assume the null hypothesis. All these features make many of our proofs substantially different from theirs.

Studentization is not the only way to achieve robustness. Robustness is also obtained through prepivoting (chung2016multivariate; fogarty2021) and the Khmaladze transformation (chung2021). Prepivoting obtains robustness in examples where studentization does not help. For instance, consider a vector of differences in means where one is interested in the maximum absolute difference. The asymptotic distribution cannot be made pivotal by simply studentizing the maximum statistic. Likewise, the Khmaladze transformation is applied to empirical processes to make them asymptotically pivotal.

Previous works have also considered randomization tests for continuity of nonparametric models at a point. cattaneo2015 propose local randomization inference procedures for a sharp null hypothesis, while canay2016rd provide permutation tests for continuity of the whole distribution of an outcome variable conditional on a control variable at a point. In contrast, our permutation test applies to testing continuity of summary statistics of the conditional distribution such as mean, quantile, variance, etc. Our fourth example is related to bugni2018, who propose a sign-change test for continuity of PDFs at a point, where critical values come from maximizing a function of a binomial distribution. We show how the same null hypothesis fits into our framework and is simply testable using permutations. The last two papers use the insight that non-iid order statistics converge in distribution to iid variables, which is technically distinct from our coupling approximation. Finally, permutation-based confidence sets have previously been proposed only in specific settings. For example, the confidence sets of imbens2005 assume that treatment effects divided by treatment doses are constant across individuals, and that the distribution of treatment eligibility is known. To the best of our knowledge, ours is the first paper to provide valid two-sample permutation tests and confidence sets for scalar parameters estimated at nonparametric rates.

The rest of this paper is outlined as follows. Section (ref) presents the general framework, assumptions, and asymptotic distributions of the classical and robust permutation tests. Section (ref) studies how our theory applies to four nonparametric examples. Section (ref) explains how to invert permutation tests to build robust confidence sets. Section (ref) displays a simulation study that confirms our theory and illustrates good finite sample properties of the robust permutation test. Section (ref) illustrates the practical relevance of our procedures to test for continuity of conditional mean functions using real-world data. The supplement contains all proofs.

Theory

Consider two populations $P_1$ and $P_2$, and a real-valued parameter $\theta(P_k)$ summarizing distribution $P_k$, $k=1,2$. The null hypothesis is stated as $H_0 : \theta(P_1) = \theta(P_2)$. Note that our null hypothesis is a bigger set of distributions compared to the set of distributions in the so-called “sharp null hypothesis,” respectively, $\{ (P_1,P_2) : \theta(P_1) = \theta(P_2)\}$ vs. $\{ (P_1,P_2) : P_1=P_2 \}$. For each population $k$, there are $n_k$ iid observations $\boldsymbol{\mathrm{Z}}_{k,i} \in \mathbb{R}^q$ from distribution $P_k$, $i=1,\ldots, n_k$. Observations are independent across $k$, and the total number of observations is $n=n_1+n_2$. We define $\mathcal{P}$ to be the convex hull of $\{P_1, P_2 \}$, that is, $\mathcal{P} = \{P : P = \eta P_1 + (1-\eta) P_2, ~ \eta \in [0,1] \}$. Throughout this paper, random variables with subscript “$k$" indicate they have distribution $P_k$, e.g., $\boldsymbol{\mathrm{Z}}_k$. For any other distribution in $\mathcal{P}$, the random variable is denoted $\boldsymbol{\mathrm{V}}\in \mathbb{R}^q$. Operators such as $\mathbb{P}$ (probability), $\mathbb{E}$ (expectation), or $\mathbb{V}$ (variance) applied to $\boldsymbol{\mathrm{Z}}_k$ do not carry the subscript $P_k$, but operators applied to $\boldsymbol{\mathrm{V}}$ carry the subscript $P$, e.g., $\mathbb{E}[\boldsymbol{\mathrm{Z}}_k]$ vs. $\mathbb{E}_P[\boldsymbol{\mathrm{V}}]$. The parameter $\theta(P_k)$ is consistently estimated by $\widehat{\theta}_{k} = \theta_{n_k,n}(\boldsymbol{\mathrm{Z}}_{k,1}, \ldots, \boldsymbol{\mathrm{Z}}_{k,n_k})$, where the functions $\theta_{n_1,n}$ and $\theta_{n_2,n}$ satisfy the following assumption.

assumptionLet $\boldsymbol{\mathrm{V}}_1, \ldots, \boldsymbol{\mathrm{V}}_m$ be an iid sample from a distribution $P \in \mathcal{P}$. Let $m$ grow with $n$ such that $ m/n \to \gamma$, for some $\gamma \in (0,1)$. Use these observations to construct the estimator $\ha\theta = \theta_{m,n}(\boldsymbol{\mathrm{V}}_1, \ldots, \boldsymbol{\mathrm{V}}_m)$. Assume the following objects exist: a sequence of functions $\psi_n: \mathbb{R}^q \times \mathcal{P} \to \mathbb{R}$, a function $\xi: \mathcal{P} \to \mathbb{R}$, and a non-increasing sequence $ h_n$ such that $n h_n \to \infty$. Further assume \begin{align} & \forall \varepsilon>0 : \sup\limits_{P \in \mathcal{P} } \mathbb{P}_P \left\{ \left| \sqrt{m h_n } \left( \widehat{\theta} - \theta(P) \right) - \left( \frac{1}{\sqrt{m}} \sum_{i=1}^{m} \psi_n(\boldsymbol{\mathrm{V}}_{i}, {P} )\right) \right| >\varepsilon \right\} \to 0, \\ & \mathbb{E}_P[ \psi_n(\boldsymbol{\mathrm{V}}_{i}, {P} ) ] = 0 \forall P \in \mathcal{P}, \\ & \sup\limits_{P \in \mathcal{P} } \left| \mathbb{V}_P[ \psi_n(\boldsymbol{\mathrm{V}}_{i}, {P} ) ] - \xi^2(P) \right| \to 0, \\ & \sup\limits_{P \in \mathcal{P} } \mathbb{E} \left[ \psi_n^2(\boldsymbol{\mathrm{Z}}_k, {P} ) \right] < \infty , for k\in\{1,2\}, \\ & \exists \zeta>0 : n^{-\zeta/2} \sup\limits_{P \in \mathcal{P} } \mathbb{E}_P \left| \frac{\psi_n(\boldsymbol{\mathrm{V}}_{i}, {P} )}{ \mathbb{V}_P[ \psi_n(\boldsymbol{\mathrm{V}}_{i}, {P} ) ]^{1/2} } \right|^{2+\zeta} \to 0 , and \\ & \xi^2\left( \frac{m}{n} P_1 + \frac{n-m}{n} P_2 \right) \to \xi^2 \left( \gamma P_1 + (1-\gamma ) P_2 \right) . \end{align}

Assumption (ref) requires the estimator $\ha \theta_k$ to have an asymptotically linear expansion at rate $\sqrt{n_k h}$ (Eq. (ref)) with zero-mean influence function $\psi_n$ (Eq. (ref)) and asymptotic variance $\xi^2$ (Eq. (ref)). The expansion must hold for data drawn from any combination of $P_1$ and $P_2$ in $\mathcal{P}$ because the limiting permutation distribution of $\widehat{\theta}_1 - \widehat{\theta}_2$ behaves as if the data were drawn from a sequence of mixtures of $P_1$ and $P_2$ (Theorems (ref) and (ref) below). Assumption (ref) also imposes bounds on moments of $\psi_n$ (Eqs. (ref)--(ref)) to enable us to apply a central limit theorem and derive the limiting permutation distribution. This requires $\xi^2(P)$ to be smooth with respect to $P$ (Eq. (ref)) in order for the limit of the variance $\xi^2\left( \frac{m}{n} P_1 + \frac{n-m}{n} P_2 \right)$ to be well-defined.

Situations arise where the number of observations $n_k$ is random rather than deterministic. For example, suppose the researcher desires to compare the female and male subpopulations of a country but only has one iid sample with $n$ individuals from that country. The researcher splits the sample into two subsamples based on the gender of each observation and sample sizes are random. In order to accommodate both deterministic and random sample sizes, we consider a sampling scheme which is dictated by a vector of indicator variables $\underline{W}_n = (W_{1}, \ldots, W_{n})$, $W_{i} \in \{1,2\}$ for $i=1,\ldots,n$, where $\underline{W}_n$ has distribution $\boldsymbol{\mathrm{Q}}_n$. Conditional on $\underline{W}_n$, the sample $\underline{\boldsymbol{\mathrm{Z}}}_n = (\boldsymbol{\mathrm{Z}}_1, \ldots, \boldsymbol{\mathrm{Z}}_n)$ has $\boldsymbol{\mathrm{Z}}_i$ drawn from distribution $P_1$ if $W_{i}=1$ or from distribution $P_2$ if $W_{i}=2$, with observations independent across $i$. This accommodates the standard two-population sampling by making $\underline{W}_n$ non-random with $n_1$ entries equal to $1$, $n_2$ entries equal to $2$, and $n_1+n_2=n$. It also accommodates the example above of male and female subpopulations by making $W_{i}$ iid and $\mathbb{P}(W_{i}=1)$ equal to the probability of being female. Conditional on $\underline{W}_n$, there are $n_k$ iid observations $\boldsymbol{\mathrm{Z}}_{k,i} \in \mathbb{R}^q$ from distribution $P_k$, $k=1,2$, and observations are independent across $k$. As before, $\widehat{\theta}_{k} = \theta_{n_k,n}(\boldsymbol{\mathrm{Z}}_{k,1}, \ldots, \boldsymbol{\mathrm{Z}}_{k,n_k})$.

assumptionThere exists $\lambda \in (0,1)$ such that the sequence of distributions $\boldsymbol{\mathrm{Q}}_n$ satisfies ${n_1}/{n} \overset{p}{\to} \lambda $ as $n\to \infty$. Moreover, Assumption (ref) holds for all sequences of sample sizes $m$ such that $ m/n \to \lambda $ or $ m/n \to 1-\lambda $.

The test statistic $T_n$ is a function of the data $(\underline{W}_n, \underline{\boldsymbol{\mathrm{Z}}}_n)$ as follows:

gather[gather omitted — 164 chars of source]

where we omit the subscript $n$ from the sequence $h_n$ of Assumption (ref) to simplify notation.

The permutation test is constructed by permuting the order of observations in $\underline{\boldsymbol{\mathrm{Z}}}_n$, while keeping the indicator variables $\underline{W}_n$ unchanged, and recomputing the test statistic. A permutation is a one-to-one function $\pi: \{1,\ldots,n\} \to \{1, \ldots, n\}$, where $\pi(i) = j $ says that the $j$-th observation becomes the $i$-th observation once permutation $\pi$ is applied. Given permutation $\pi$, the permuted sample becomes $(\underline{W}_n,\underline{\boldsymbol{\mathrm{Z}}}_n^{\pi}) = (W_{1}, \ldots, W_{n},\boldsymbol{\mathrm{Z}}_{\pi(1)}, \ldots, \boldsymbol{\mathrm{Z}}_{\pi(n)})$, and the re-computed value of the test statistic is $T_n^{\pi} = T_n(\underline{W}_n, \underline{\boldsymbol{\mathrm{Z}}}_n^\pi)$. In other words, permutations swap individuals across the two samples to which they originally belonged according to $\underline{W}_n$, which remains fixed. The set $\boldsymbol{\mathrm{G}}_n$ is the set of all possible permutations $\pi$. The number of elements in $\boldsymbol{\mathrm{G}}_n$ is $n!$.

The two-sided permutation test with nominal level $\alpha \in (0,1)$ is constructed as follows. First, re-compute the test statistic $T_n(\underline{W}_n, \underline{\boldsymbol{\mathrm{Z}}}_n)$ for every $\pi \in \boldsymbol{\mathrm{G}}_n$. Rank the values of $T^{\pi}_n$ across $\pi$: $T^{(1)}_{n} \leq T^{(2)}_{n} \leq \ldots \leq T^{(n!)}_{n}$. Second, fix a nominal level $\alpha \in (0,1)$ and let $k^- = \lfloor n! \alpha/2 \rfloor$, that is, the largest integer less than or equal to $n! \alpha/2 $, and $k^+ = n! - k^- $. Third, compute the following quantities: (i) $M^+$, the number of values $T^{(j)}_{n}$, $j = 1, \ldots, n!$, that are strictly greater than $T^{(k^+)}_{n}$; (ii) $M^-$, the number of values $T^{(j)}_{n}$ that are strictly smaller than $T^{(k^-)}_{n}$; (iii) $M^0$, the number of values $T^{(j)}_{n}$ that are equal to either $T^{(k^+)}_{n}$ or $T^{(k^-)}_{n}$; and (iv) $a = \left( \alpha n! - M^+ - M^- \right)/M^0 $. Finally, the outcome of the test is based on the test function $\phi$:

align[align omitted — 363 chars of source]

For a given sample, if $\phi=1$, we reject the null hypothesis; if $\phi=a$, we randomly reject the null hypothesis with probability $a$; otherwise, if $\phi=0$, we fail to reject the null. A classic property of permutation tests is exact size in finite samples under the “sharp null,” that is, the null hypothesis stating $P_1 = P_2$.

lemmaFor any $n$, $\boldsymbol{\mathrm{Q}}_n$, $P_1$, and $P_2$, if $P_1=P_2$, then $\mathbb{E}[\phi (\underline{W}_n, \underline{\boldsymbol{\mathrm{Z}}}_n)] = \alpha$.
remarkRe-computing the test statistic for all $n!$ permutations is costly, even for small $n$. Lemma (ref) and remaining results in this section are unchanged if, instead of $\boldsymbol{\mathrm{G}}_n$, we use a random sample of $\boldsymbol{\mathrm{G}}_n$ with or without replacement (lehmann2005, page 636).
remarkThe randomized outcome in the case of ties is important for exact size in finite samples if $P_1=P_2$. However, it may be desirable to have a deterministic answer to a hypothesis test after observing a sample of data. An easy way to fix that is to set $\phi=0$ in the case of ties. This makes the test conservative, that is, the size becomes less than or equal to $\alpha$.

The set of distributions that satisfy the null hypothesis $\theta(P_1)=\theta(P_2)$ is in general larger than the set of distributions that satisfy the sharp null $P_1=P_2$. Thus, there is no finite sample size control in general. To investigate the asymptotic properties of the test in (ref), we derive the probability limit of the permutation distribution,

gather[gather omitted — 201 chars of source]

The hypothesis test (ref) utilizes critical values from $\widehat{R}_{T_n}$. The test has asymptotic size control if, under the null hypothesis, the probability limit of $\widehat{R}_{T_n}$ equals the cumulative distribution function (CDF) of the limiting distribution of $T_n$. In order to study both size and power, we derive these limiting distributions without imposing the null hypothesis in the following theorems.

theoremUnder Assumptions (ref)--(ref), the permutation distribution $\widehat{R}_{T_n}$ converges uniformly in probability to the CDF of a $N \left( 0, \tau^2 \right)$, i.e., $\sup_t \left| \widehat{R}_{T_n}(t) - \Phi\left( t / \tau \right) \right| \overset{p}{\to} 0,$ where $\tau^2 \doteq {\xi^2(\overline{P})} / {\lambda(1- \lambda)}$ and $\overline{P} \doteq \lambda P_1 + (1-\lambda) P_2$. Moreover, $T_n - \sqrt{nh}\left( \theta(P_1) - \theta(P_2) \right) \overset{d}{\to} N \left( 0, \sigma^2 \right),$ where $\sigma^2 \doteq { \xi^2(P_1)}/{\lambda} + { \xi^2(P_2) }/{(1-\lambda)}$.

The permutation distribution fails to control size asymptotically because the asymptotic variance of the permutation distribution $\tau^2$ generally differs from $\sigma^2$. To appreciate the implications of Theorem (ref), we present the parametric example below and four nonparametric examples in Section (ref) detailing cases where $\tau^2 = \sigma^2$ and $\tau^2 \neq \sigma^2$.

example\rm (Parametric Model) For $k=1,2$, $\boldsymbol{\mathrm{Z}}_k=(X_k,Y_k) \sim P_k$, $X_k \sim U[0,1]$, $\mathbb{E}{[Y_k|X_k]}=\theta(P_k) + \beta X_k$, and $\mathbb{V}{[Y_k|X_k]} = v_k$. There are two independent samples with $n_1/n \to \lambda$. The ordinary least squares estimator is $\widehat{\theta}_k = \frac{\mkern 1.5mu\overline{\mkern-1.5muX_k^2\mkern-1.5mu}\mkern 1.5mu\mkern 1.5mu\overline{\mkern-1.5muY_k\mkern-1.5mu}\mkern 1.5mu - \mkern 1.5mu\overline{\mkern-1.5muX_k\mkern-1.5mu}\mkern 1.5mu\mkern 1.5mu\overline{\mkern-1.5muX_kY_k\mkern-1.5mu}\mkern 1.5mu}{\mkern 1.5mu\overline{\mkern-1.5muX_k^2\mkern-1.5mu}\mkern 1.5mu - (\mkern 1.5mu\overline{\mkern-1.5muX_k\mkern-1.5mu}\mkern 1.5mu)^2},$ where $\mkern 1.5mu\overline{\mkern-1.5muX_k^s\mkern-1.5mu}\mkern 1.5mu = \frac{1}{n_k}\sum_{i=1}^{n_k} X_{k,i}^s$ for $s=1, 2$, and $\mkern 1.5mu\overline{\mkern-1.5muX_kY_k\mkern-1.5mu}\mkern 1.5mu = \frac{1}{n_k}\sum_{i=1}^{n_k} X_{k,i}Y_{k,i}$. For $\boldsymbol{\mathrm{V}} = (R, S) \sim P$, the setting satisfies Assumption $\ref{aspt:asy_linear}$ with $h=1$; $\psi_n(\boldsymbol{\mathrm{V}}, P)=\psi(\boldsymbol{\mathrm{V}}, P)=(S-\mathbb{E}_P{(S|R)})(4-6R)$; and $\xi^2(P) = \mathbb{E}_P\left[ (S-\mathbb{E}_P(S|R))^2 (4-6R)^2\right].$ Under the null hypothesis, the asymptotic variance of $T_n = \sqrt{n}(\widehat{\theta}_1 - \widehat{\theta}_2)$ and of the permutation distribution are, respectively, $\sigma^2 = 4 \left[ \frac{ v_1 }{ \lambda } + \frac{ v_2 }{1-\lambda } \right]$ and $\tau^2 = 4 \left[ \frac{ v_1 }{ 1- \lambda } + \frac{ v_2 }{\lambda } \right].$ Unless $\lambda=1/2$ or $v_1=v_2$, $\sigma^2$ and $\tau^2$ do not agree in general, distorting the rejection probability under the null for the permutation test. Section C in the supplement displays a graph that illustrates this distortion using simulated data.

Since $\tau^2$ and $\sigma^2$ are generally different, the test statistic $T_{n}$ must be transformed to become asymptotically pivotal. Thus, we divide $T_n$ by the square root of a consistent estimator for its asymptotic variance. For each population $k \in \{ 1, 2\}$, let $\hat \xi_{k}^2 = \xi_{n_k,n}^2(\boldsymbol{\mathrm{Z}}_{k,1}, \ldots, \boldsymbol{\mathrm{Z}}_{k,n_k})$ be a consistent estimator for $\xi^2(P_k)$ and assume the functions $\xi_{n_1,n}^2$ and $\xi_{n_2,n}^2$ satisfy the following assumption.

assumptionLet $\boldsymbol{\mathrm{V}}_1, \ldots, \boldsymbol{\mathrm{V}}_m$ be an iid sample from a distribution $P \in \mathcal{P}$. Use these observations to construct the estimator $\ha\xi^2 = \xi^2_{m,n}(\boldsymbol{\mathrm{V}}_1, \ldots, \boldsymbol{\mathrm{V}}_m)$. Assume that, for any sequence of sample sizes $m$ such that $ m/n \to \lambda $ or $ m/n \to 1-\lambda $, $\ha\xi^2 - \xi^2(P)$ converges in probability to $0$ uniformly over $P \in \mathcal{P}$.

Then, the studentized test statistic $S_n$ is

align[align omitted — 167 chars of source]

where $\ha \sigma_n$ is the square root of the consistent estimator for the asymptotic variance of $T_n$, that is, $\ha\sigma_n^2 = \frac{n}{n_1} \ha\xi_{1}^{2} + \frac{n}{n_2} \ha\xi_{2}^2.$

theoremLet $\widehat{R}_{S_n}$ be the permutation CDF defined in ((ref)) with $T_n$ replaced by $S_n$. Under Assumptions (ref)--(ref), $\widehat{R}_{S_n}$ converges uniformly in probability to the CDF of a $N \left( 0, 1 \right)$, i.e., $\sup_t \left| \widehat{R}_{S_n}(t) - \Phi\left( t \right) \right| \overset{p}{\to} 0.$ Moreover, $S_n - \frac{\sqrt{nh}\left( \theta(P_1) - \theta(P_2) \right)}{\ha \sigma_n} \overset{d}{\to} N \left( 0, 1 \right).$

\begin{example_contd}[(ref)]\rm If the test statistic $T_n$ is appropriately studentized by the usual ordinary least squares formula for standard errors, both permutation and sampling distributions of $S_n$ coincide asymptotically. Section C in the supplement simulates data for this example and presents these two distributions graphically. \end{example_contd}

Note that the standard deviation $\hat \sigma_n$ that divides $T_n$ must be consistent for $\sigma$, as opposed to $\tau$. However, when $\hat \sigma_n$ is evaluated using permuted samples, it converges in probability to $\tau$. Under the null hypothesis, both the permutation distribution and the test statistic $S_n$ are asymptotically standard normal. Therefore, our robust permutation test in (ref) with $T_n$ replaced by $S_n$ has asymptotic size equal to the nominal level $\alpha$, even if $P_1 \neq P_2$. In case $P_1 = P_2$, this test has exact size in finite samples. Since Theorems (ref) and (ref) are true regardless of whether the null hypothesis holds or not, we can now study the power properties of the permutation test.

corollaryLet $\phi(\underline{W}_n, \boldsymbol{\mathrm{Z}}_n)$ be the permutation test in (ref) with $T_n$ replaced by $S_n$, and suppose Assumptions (ref)--(ref) hold. If the null hypothesis holds, then $\mathbb{E}[\phi(\underline{W}_n, \boldsymbol{\mathrm{Z}}_n)] \to \alpha$; otherwise, $\mathbb{E}[\phi(\underline{W}_n, \boldsymbol{\mathrm{Z}}_n)] \to 1$. Moreover, assume $S_n$ has a limiting distribution under a sequence of local alternatives contiguous to the null. Then, the asymptotic power of the robust permutation test against local alternatives is the same as that of the test that uses critical values from the limiting null distribution of $S_n$.\footnote{Alternative resampling methods such as the bootstrap and subsampling also share the same asymptotic local power as the permutation test because they produce critical values that are consistent for standard Gaussian critical values under the null hypothesis.}$^,$\footnote{Suppose one uses an estimator $\widehat{\sigma}^2_n$ that assumes the null hypothesis is true; that is, $\widehat{\sigma}^2_n$ is consistent for $\sigma^2$ under the null hypothesis but has a different probability limit under the alternative. Regardless of whether the null is true, such an estimator applied to a random permutation of the data is generally consistent for $\tau^2$, and Corollary (ref) remains true. Consistency for $\tau^2$ comes from the fact that an estimator applied to a random permutation behaves as if it were applied to data from a mixture distribution, where the null is always true (Section B.3.2 in the supplement). }

Applications

In this section, we apply our theory to four different nonparametric problems: testing for equality of conditional mean and quantile functions evaluated at a point, and testing for continuity of conditional mean and PDF at a point. For a simple and intuitive presentation, we use the Nadaraya--Watson (NW) type of kernel estimators throughout, but proofs in this section generalize to other types of estimators, e.g., local-polynomial regression (LPR), sieves, etc. In particular, we demonstrate the generalization of our third application in Section (ref) to the LPR estimator in Section D.5 of the supplement. We obtain the usual rate restrictions on the bandwidth tuning parameter $h$, which allows researchers to choose among standard options available in the literature. We fit the four examples into the general framework of Section (ref), demonstrate the validity of Assumptions (ref) and (ref), and show that studentization is generally required for asymptotic size control, except for the PDF continuity test. The third and fourth examples (testing for continuity of conditional mean and PDF) illustrate the sample-splitting feature of our general framework.

Throughout this section, $K(u)$ denotes a symmetric, non-negative, bounded kernel density function with $\int_{-\infty}^{\infty} |u|^3 K(u) ~ du <\infty$. We denote $\kappa_{s,t} = \int_{-\infty}^{\infty} u^s K^t(u) ~ du$ and $\kappa_{s,t}^+ = \int_{0}^{\infty} u^s K^t(u) ~ du$.

Controlled Means

Researchers are often interested in comparing conditional mean functions between two different populations. For example, in randomized controlled trials, $P_1$ and $P_2$ are the populations of control and treatment individuals, respectively. Of interest is the average outcome $Y$ after controlling for an individual characteristic $X$. For instance, outcomes of a professional training program may differ between rich and poor individuals, and we would like to condition on income $X$.

There are two independent samples of bivariate variables: $n_1$ observations with $\boldsymbol{\mathrm{Z}}_{1,i}=(X_{1,i}, Y_{1,i})$ iid $P_1$ and $n_2$ observations with $\boldsymbol{\mathrm{Z}}_{2,i}=(X_{2,i}, Y_{2,i})$ iid $P_2$. The vector $\underline{W}_n$ is non-random with $n_1$ entries equal to $1$ and $n_2$ entries equal to $2$, where $n=n_1+n_2$. For a given interior point $x$, $\theta(P_k) = \mathbb{E}[Y_{k,i} | X_{k,i}=x]$, $k=1,2$, and the null hypothesis is $\theta(P_1) = \theta(P_2)$.

A common estimator for conditional mean functions is the NW kernel estimator, which computes a weighted average of $Y$ local to $X=x$ for each population $k=1,2$. For a bandwidth $h > 0$ and a kernel density function $K$, the NW estimator for $\theta(P_k)$ is written as

gather[gather omitted — 313 chars of source]

The superscript $b$ indicates that there is bias in the asymptotic distribution of $\sqrt{n h} (\widehat{\theta}_{k}^b - \theta(P_k))$ whenever the bandwidth choice converges to zero at the slowest possible rate, i.e., $h=O(n^{-1/5})$ (Proposition (ref)). This is the case of mean squared error (MSE) optimal bandwidths, and inference requires bias correction in this case. A conventional solution is to subtract a first-order bias term $h^2 B(P_k)$ from $\widehat{\theta}_{k}^b$, where $B(P_k)$ is nonparametrically estimated by $\widehat{B}_k = B_{n_k,n}(\boldsymbol{\mathrm{Z}}_{k,1}, \ldots, \boldsymbol{\mathrm{Z}}_{k,n_k})$. We give the analytical formulas for $B(P_k)$ and $\widehat{B}_k$ in Section D.1 (Equation D.6) of the supplemental appendix. Our permutation tests utilize the bias-corrected NW estimator $\widehat{\theta}_{k} = \theta_{n_k,n}(\boldsymbol{\mathrm{Z}}_{k,1}, \ldots, \boldsymbol{\mathrm{Z}}_{k,n_k}) \doteq \theta_{n_k,n}^b(\boldsymbol{\mathrm{Z}}_{k,1}, \ldots, \boldsymbol{\mathrm{Z}}_{k,n_k}) - h^2 B_{n_k,n}(\boldsymbol{\mathrm{Z}}_{k,1}, \ldots, \boldsymbol{\mathrm{Z}}_{k,n_k}) =\widehat{\theta}_{k}^b - h^2 \widehat{B}_k$. Note that no bias correction is needed if $h=o(n^{-1/5})$ because $\sqrt{n h} (\widehat{\theta}_{k}^b - \theta(P_k) - h^2 B(P_k)) = \sqrt{n h} (\widehat{\theta}_{k}^b - \theta(P_k)) +o(1)$.

An alternative solution to the bias issue consists of replacing the NW estimator $\widehat{\theta}_{k}^b$, which is the LPR estimator of order zero, with the LPR estimator of order two. For LPR estimation at an interior point $x$, if $\widehat{\theta}_{k}^b$ is LPR of order $\rho$, the asymptotic bias of $\sqrt{n h} (\widehat{\theta}_{k}^b - \theta(P_k))$ is $O\left( \sqrt{n h} h^{\rho+2} \right)$ if $\rho$ is even, or $O\left( \sqrt{n h} h^{\rho+1} \right)$ if $\rho$ is odd (Theorem 3.1 by fan1996). Thus, if an LPR of order $\rho$ has asymptotic bias, that bias vanishes if we increase the order of the polynomial by two if $\rho$ is even, or by one if $\rho$ is odd. We keep the NW estimator for a simple and intuitive presentation of our theory in this section, but our permutation tests also apply to LPR estimators. We demonstrate this in the context of our third application (Section (ref)) in the supplemental appendix (Section D.5). In the same spirit, our simulations in Section (ref) and empirical examples in Section (ref) utilize LPR of order one with MSE-optimal bandwidth and bias-correct it using LPR of order two. Other options of bias correction include higher-order kernels (liracine2007, Section 1.11) and the bootstrap (racine2001).

When the distribution of $(X_1,Y_1)$ equals that of $(X_2,Y_2)$, the permutation test in (ref) has exact size in finite samples. For other cases, we rely on asymptotic size control, which depends on Assumptions (ref) and (ref). Below, we describe regularity conditions such as continuous differentiability of conditional moments, and Proposition (ref) proves the asymptotic linear representation of the bias-corrected NW estimator.

assumptionAs $n\to \infty$, $ n_1/n \to \lambda \in (0,1) $, $h \to 0$, $n h \to \infty$, and $\sqrt{n h} h^2 \to c \in [0,\infty)$.
assumptionFor $k=1,2$, the distribution of $X_k$ has PDF $f_{X_k}(x_k)$ that is bounded, bounded away from zero, and three times differentiable with bounded derivatives.
assumptionFor $k=1,2$, $m_{Y_k|X_k}(x_k) \doteq \mathbb{E}[Y_k|X_k=x_k]$ is bounded, three times differentiable, and its derivatives are bounded; there exists $\zeta>0$ such that $\mathbb{E}[|Y_k |^{2+\zeta} | X_k]$ is almost surely bounded.
assumptionFor $k=1,2$, $v_{Y_k|X_k}(x_k) \doteq \mathbb{V}[Y_k|X_k=x_k]$ is bounded, differentiable, its derivative is bounded, and $v_{Y_k|X_k}(x) >0$.
assumptionLet $\boldsymbol{\mathrm{V}}_1, \ldots, \boldsymbol{\mathrm{V}}_m$ be an iid sample from a distribution $P \in \mathcal{P}$, where $m$ grows with $n$. Let $\widehat{B} = B_{m,n}(\boldsymbol{\mathrm{V}}_{1}, \ldots, \boldsymbol{\mathrm{V}}_{m})$ be a consistent estimator for the first-order bias term $B(P)$. Assume that $\left( \widehat{B} - B(P) \right) \overset{p}{\to} 0$ uniformly over $P \in \mathcal{P}$ for any sequence $m$ such that $ m/n \to \lambda$ or $ m/n \to 1-\lambda $.
propositionSuppose Assumptions (ref)--(ref) hold. Let $\boldsymbol{\mathrm{V}}_1=(R_1,S_1), \ldots, \boldsymbol{\mathrm{V}}_m=(R_m,S_m)$ be an iid sample from a distribution $P \in \mathcal{P}$, where $m$ grows with $n$, and consider the bias-corrected estimator $\widehat{\theta} = \theta_{m,n}(\boldsymbol{\mathrm{V}}_{1}, \ldots, \boldsymbol{\mathrm{V}}_{m})$ as described in the text. Then, Assumptions (ref) and (ref) hold for $\widehat{\theta}$ with $\psi_{n}(\boldsymbol{\mathrm{V}}, P) = K\left( \frac{ R - x }{h_n} \right) \left( \frac{S - m_{S|R}(R;P)}{\sqrt{h_n} f_{R}(x;P) } \right)$ and $\xi^2(P) = \frac{ v_{S|R}(x;P) \kappa_{0,2} }{ f_{R}(x;P) },$ where $m_{S|R}(x;P)$ is the conditional mean of $S$ given $R=x$, $v_{S|R}(x;P)$ is the conditional variance of $S$ given $R=x$, and $f_{R}(x;P)$ is the PDF of $R$ at $x$, all three assuming $\boldsymbol{\mathrm{V}}=(R,S) \sim P \in \mathcal{P}$.

The proof of Proposition (ref) adapts conventional arguments for nonparametric asymptotics (e.g., Theorem 2.2 by liracine2007) and is found in Section D.1 of the supplement. Under the null hypothesis, the asymptotic variance of $T_n$ and of the permutation distribution are, respectively, $\sigma^2 = \kappa_{0,2} \left( \frac{ v_{Y_1|X_1}(x) }{ \lambda f_{X_1}(x) } + \frac{ v_{Y_2|X_2}(x) }{ (1-\lambda) f_{X_2}(x) }\right)$ and $\tau^2 = \kappa_{0,2} \left( \frac{ f_{X_1}(x) v_{Y_1|X_1}(x) }{ (1-\lambda) f_{R}^2(x;\overline{P}) } + \frac{ f_{X_2}(x) v_{Y_2|X_2}(x) }{ \lambda f_{R}^2(x;\overline{P} ) } \right),$ where $f_{R}^2(x;\overline{P})$ is the square of the PDF of $R$ evaluated at $x$ under $\overline{P} = \lambda P_1 + (1-\lambda) P_2 \in \mathcal{P}$. These variances are generally different, except in special cases, e.g., when $f_{X_1}(x)=f_{X_2}(x)$ and $\lambda=1/2$ or when $f_{X_1}(x)=f_{X_2}(x)$ and $v_{Y_1|X_1}(x) = v_{Y_2|X_2}(x)$. Thus, in general, the researcher must use the studentized test statistic for the permutation test to have asymptotic size control.

Controlled Quantiles

In this subsection, we examine equality of conditional quantile functions for two populations. For example, a researcher may wish to compare not only averages (Section (ref)) but also other features of a conditional distribution between $P_1$ and $P_2$, such as the median, tails, interquartile range, etc. The goal is to test the difference of the $\chi$-th quantile of the outcomes $Y$ between the two populations, after controlling for a given value of the variable $X$. For instance, the immune response $Y$ of a certain treatment conditional on age $X$ may differ for individuals at the bottom, median, or top of the immunity distribution.

As in Section (ref), there are two independent samples, $\boldsymbol{\mathrm{Z}}_{1,i}=(X_{1,i}, Y_{1,i})$, $i=1, \ldots, n_1$, and $\boldsymbol{\mathrm{Z}}_{2,i}=(X_{2,i}, Y_{2,i})$, $i=1, \ldots, n_2$, and the vector $\underline{W}_n$ is non-random. For a given interior point $x$, the parameter of interest is the $\chi$-th conditional quantile, that is, $\theta(P_k) = \mathbb{Q}_{\chi}[Y_{k,i} | X_{k,i}=x] = \arg\min_a \mathbb{E}[\rho_{\chi}(Y_{k,i}-a)|X_{k,i}=x]~,$ where $\rho_{\chi}(u) = \left(\chi - \mathbb{I}(u < 0)\right)u$.

For a bandwidth $h > 0$ and a kernel density function $K$, a consistent estimator of the NW style is $\widehat{\theta}_{k}^b =\theta_{n_k,n}^b(\boldsymbol{\mathrm{Z}}_{k,1}, \ldots, \boldsymbol{\mathrm{Z}}_{k,n_k}) \doteq \mathop{\hbox{\rm argmin}}_a \sum\limits_{i=1}^{n_k} \rho_{\chi}(Y_{k,i}- a)K \left( \frac{ X_{k,i}-x }{ h } \right) \text{, for }k=1,2.$

The superscript $b$ indicates that there is bias in the asymptotic distribution of $\sqrt{n h} (\widehat{\theta}_{k}^b - \theta(P_k))$ whenever the bandwidth choice converges to zero at the slowest possible rate, i.e., $h=O(n^{-1/5})$ (Proposition (ref)). This is the case of mean squared error (MSE) optimal bandwidths, and inference requires bias correction in this case. A conventional solution is to subtract a first-order bias term $h^2 B(P_k)$ from $\widehat{\theta}_{k}^b$, where $B(P_k)$ is nonparametrically estimated by $\widehat{B}_k = B_{n_k,n}(\boldsymbol{\mathrm{Z}}_{k,1}, \ldots, \boldsymbol{\mathrm{Z}}_{k,n_k})$. We give the analytical formulas for $B(P_k)$ and $\widehat{B}_k$ in Section D.2 (Equation D.14) of the supplemental appendix. Our permutation tests utilize the bias-corrected estimator $\widehat{\theta}_{k} = \theta_{n_k,n}(\boldsymbol{\mathrm{Z}}_{k,1}, \ldots, \boldsymbol{\mathrm{Z}}_{k,n_k}) \doteq \theta_{n_k,n}^b(\boldsymbol{\mathrm{Z}}_{k,1}, \ldots, \boldsymbol{\mathrm{Z}}_{k,n_k}) - h^2 B_{n_k,n}(\boldsymbol{\mathrm{Z}}_{k,1}, \ldots, \boldsymbol{\mathrm{Z}}_{k,n_k}) =\widehat{\theta}_{k}^b - h^2 \widehat{B}_k$. Note that no bias correction is needed if $h=o(n^{-1/5})$ because $\sqrt{n h} (\widehat{\theta}_{k}^b - \theta(P_k) - h^2 B(P_k)) = \sqrt{n h} (\widehat{\theta}_{k}^b - \theta(P_k)) +o(1)$. An alternative solution to the bias issue consists of replacing $\widehat{\theta}_{k}^b$ with the LPR estimator, while keeping the MSE-optimal bandwidth choice for the NW estimator. See Sections (ref) and (ref) for details.

Proposition (ref) below demonstrates validity of Assumptions (ref)--(ref) for the NW quantile regression estimator. It relies on some of the same assumptions as Section (ref) plus a couple of assumptions on the distribution of $Y_k$ conditional on $X_k$, $k=1,2$.

assumptionp{(ref)'} For $k=1,2$, the distribution of $Y_k$ conditional on $X_k$ has PDF $f_{Y_k | X_k}(y_k | x_k)$ that is a bounded and differentiable function of $(x_k,y_k)$, has bounded partial derivatives with respect to $x_k$ and $y_k$, and is bounded away from zero over $y_k$ for $x_k=x$.
assumptionp{(ref)'} For $k=1,2$, the distribution of $Y_k$ conditional on $X_k$ has CDF $F_{Y_k | X_k}(y_k | x_k)$ that is three times partially differentiable with respect to $x_k$ with bounded partial derivatives.
propositionSuppose Assumptions (ref), (ref), (ref), (ref), and (ref) hold. Let $\boldsymbol{\mathrm{V}}_1=(R_1,S_1), \ldots, \boldsymbol{\mathrm{V}}_m=(R_m,S_m)$ be an iid sample from a distribution $P \in \mathcal{P}$, where $m$ grows with $n$, and consider the bias-corrected estimator $\widehat{\theta} = \theta_{m,n}(\boldsymbol{\mathrm{V}}_{1}, \ldots, \boldsymbol{\mathrm{V}}_{m})$ as described in the text. Then, $\widehat{\theta}$ satisfies Assumptions (ref) and (ref) with $\psi_{n}(\boldsymbol{\mathrm{V}}, P) = -K\left(\frac{R - x }{ h_n }\right) \frac{\left( \mathbb{I}\{S < \theta(P) \} - F_{S|R}(\theta(P)| R ;P) \right)} {f_{S|R}(\theta(P) | x ; P) f_R(x;P) \sqrt{h_n} } $ and $\xi^2(P) = \frac{\kappa_{0,2}\chi (1-\chi) }{f_{S|R}^2(\theta(P) | x ; P) f_R(x;P)}, $ where $f_{S|R}(s | r ; P)$ is the conditional PDF of $S$ given $R=r$, $f_R(r;P)$ is the PDF of $R$, and $F_{S|R}(s| r ;P) $ is the conditional CDF of $S$ given $R=r$, all three assuming $\boldsymbol{\mathrm{V}}=(R,S) \sim P \in \mathcal{P}$.

The proof of Proposition (ref) adapts arguments by pollard1991, chaudhuri1991, and fan1994 and is found in Section D.2 of the supplement. The asymptotic variance of $T_n$ and of the permutation distribution are, respectively, $\sigma^2 = \frac{\kappa_{0,2} \chi (1- \chi) }{ \lambda f_{Y_1 | X_1}^2(\theta(P_1)|x) f_{X_1}(x) } + \frac{ \kappa_{0,2} \chi (1- \chi) }{ (1-\lambda) f_{Y_2|X_2}^2(\theta(P_2)|x ) f_{X_2}(x ) }$ and $\tau^2 = \frac{ \kappa_{0,2} \chi (1- \chi) }{ \lambda (1-\lambda) f_{S|R}^2(\theta(\bar P) | x; \bar P) f_{R}(x; \bar P) }.$ These variances are generally different except in special cases. For example, the null hypothesis implies $\theta(P_1)=\theta(P_2) = \theta(\bar{P}) $. If $f_{Y_1|X_1}(\theta(P_1)|x) = f_{Y_2|X_2}(\theta(P_2)|x)$ and $f_{X_1}(x)=f_{X_2}(x)$, then $\sigma^2 = \tau^2$. Thus, in general, the researcher must use the studentized test statistic for the permutation test to have asymptotic size control.

Discontinuity of Conditional Mean

Numerous empirical studies in the social sciences have relied on estimation and inference on the size of a discontinuity in a conditional mean function at a certain point. In the so-called regression discontinuity design (RDD), an individual $i$ receives treatment if and only if a running variable $X_i$ is above a fixed threshold. If individuals do not know the threshold or do not have perfect manipulation over $X$, untreated individuals who barely missed the cutoff serve as a control group for treated individuals who barely made it across the cutoff. Assume the threshold for treatment is $0$ without loss of generality. The difference in side limits $\mathbb{E}[Y | X=0^+] - \mathbb{E}[Y | X=0^-]$ identifies the causal effect of treatment on an outcome variable $Y$. Thus, the null hypothesis of zero causal effect is equivalent to continuity of the conditional mean function $\mathbb{E}[Y | X=x]$ at $x=0$.

This idea first appeared in psychology thistlethwaite1960, made its way to economics hahn2001id, and has a growing number of applications in the social sciences. Examples include agarwal2016 in economics, valentine2017 in education, abou2020 in political science, and zoorob2020 in sociology. We focus on the conditional mean function in this subsection, but the whole argument goes through if one desires to use the conditional quantile function.

The researcher has a sample of $n$ iid observations $(X_i,Y_i)$ and the NW estimator for the discontinuity is

align*[align* omitted — 336 chars of source]

RDD fits in our two-population framework by splitting the sample based on $X$ being above or below the cutoff. Construct $W_{i} = 2 - \mathbb{I}\{X_i \geq 0\}$ so that $n_1 / n \overset{p}{\to} \lambda = \mathbb{P}[W_i=1]$. Re-order the sample such that the $n_1$ observations with $W_i=1$ come first, and the $n_2$ observations with $W_i=2$ come second. Define $\boldsymbol{\mathrm{Z}}_{1,i} = (X_{1,i}, Y_{1,i}) = (X_i, Y_i)$ for $i=1,\ldots, n_1$ and $\boldsymbol{\mathrm{Z}}_{2,i} = (X_{2,i}, Y_{2,i}) = (-X_{n_1+i}, Y_{n_1+i})$ for $i=1,\ldots, n_2$. We have $\underline{\boldsymbol{\mathrm{Z}}}_n = ( \boldsymbol{\mathrm{Z}}_{1,1}, \ldots, \boldsymbol{\mathrm{Z}}_{1,n_1}, \boldsymbol{\mathrm{Z}}_{2,1}, \ldots, \boldsymbol{\mathrm{Z}}_{2,n_2})$. Conditional on $\underline{W}_n$, the distribution of $Z_{1,i}$ is $P_1$, which equals the distribution of $(X,Y)$ conditional on $X \geq 0$. Likewise for $Z_{2,i}$, $P_2$ is the distribution of $(-X,Y)$ conditional on $X < 0$. Permutations re-order observations in $\underline{\boldsymbol{\mathrm{Z}}}_n$ but keep $\underline{W}_n$ unchanged. The RDD parameter becomes $\theta(P_1)-\theta(P_2)$, where $\theta(P_k) = \mathbb{E}[Y_k|X_k=0^+]$ for $k=1,2$.

The NW estimator for the RDD parameter is $\ha\theta_1^b - \ha\theta_2^b$, where $\ha \theta_k^b$ is defined in Equation (ref) for $k=1,2$ and evaluation point $x$ set to zero. The superscript $b$ indicates that there is bias in the asymptotic distribution of $\sqrt{n h} (\widehat{\theta}_{k}^b - \theta(P_k))$ whenever the bandwidth choice converges to zero at the slowest possible rate, i.e., $h=O(n^{-1/3})$ (Proposition (ref)). This is the case of MSE-optimal bandwidths, and inference requires bias correction in this case. A conventional solution is to subtract a first-order bias term $h B(P_k)$ from $\widehat{\theta}_{k}^b$, where $B(P_k)$ is nonparametrically estimated by $\widehat{B}_k = B_{n_k,n}(\boldsymbol{\mathrm{Z}}_{k,1}, \ldots, \boldsymbol{\mathrm{Z}}_{k,n_k})$. We give the analytical formulas for $B(P_k)$ and $\widehat{B}_k$ in Section D.3 (Equation D.21) of the supplemental appendix. Our permutation tests utilize the bias-corrected NW estimator $\widehat{\theta}_{k} = \theta_{n_k,n}(\boldsymbol{\mathrm{Z}}_{k,1}, \ldots, \boldsymbol{\mathrm{Z}}_{k,n_k}) \doteq \theta_{n_k,n}^b(\boldsymbol{\mathrm{Z}}_{k,1}, \ldots, \boldsymbol{\mathrm{Z}}_{k,n_k}) - h B_{n_k,n}(\boldsymbol{\mathrm{Z}}_{k,1}, \ldots, \boldsymbol{\mathrm{Z}}_{k,n_k}) =\widehat{\theta}_{k}^b - h \widehat{B}_k$. Note that no bias correction is needed if $h=o(n^{-1/3})$ because $\sqrt{n h} (\widehat{\theta}_{k}^b - \theta(P_k) - h B(P_k)) = \sqrt{n h} (\widehat{\theta}_{k}^b - \theta(P_k)) +o(1)$.

An alternative solution to the bias issue consists of replacing the NW estimator $\widehat{\theta}_{k}^b$, which is the LPR estimator of order zero, with the LPR estimator of order one, while keeping the MSE-optimal bandwidth choice for the NW estimator. For LPR estimation at a boundary point, if $\widehat{\theta}_{k}^b$ is LPR of order $\rho$, the asymptotic bias of $\sqrt{n h} (\widehat{\theta}_{k}^b - \theta(P_k))$ is $O\left( \sqrt{n h} h^{\rho} \right)$ (Theorem 3.2 by fan1996). Thus, if an LPR of order $\rho$ has asymptotic bias, that bias vanishes if we increase the order of the polynomial by one. We keep the NW estimator for a simple and intuitive presentation of our theory in this section, but we demonstrate that our permutation tests also apply to LPR estimators of any order in Section D.5 of the supplement. Section (ref) illustrates our procedures with two empirical examples of RDD using LPR estimators and MSE-optimal bandwidths, which is standard practice in applied literature. A third option for bias correction is the method proposed by armstrong2018.

In case the distribution of $(Y,X)$ equals that of $(Y,-X)$, then $P_1=P_2$ and the permutation test in (ref) has exact size in finite samples. Note that this X-symmetry restriction eliminates the impossibility problem in RDD tests (kamat2018 and bertanha2019) because there is no bias in estimation. For other cases, we rely on asymptotic size control, which depends on Assumptions (ref) and (ref). Proposition (ref) below builds on assumptions similar to those of Section (ref), which we rewrite below in terms of the originally sampled variables $(X,Y)$, that is, before the variables are transformed into $(X_k,Y_k)$, $k=1,2$.

assumptionp{(ref)$^*$} As $n\to \infty$, $h \to 0$, $n h \to \infty$, and $\sqrt{n h} h \to c \in [0,\infty)$.
assumptionp{(ref)$^*$} The distribution of $X$ has PDF $f_{X}(x)$ that is bounded, bounded away from zero, twice differentiable except at $x=0$, and has bounded derivatives.
assumptionp{(ref)$^*$} $\mathbb{E}[Y | X = x]$ is bounded, twice differentiable except at $x=0$, and has bounded derivatives. There exists $\zeta>0$ such that $\mathbb{E}[|Y|^{2+\zeta} | X]$ is almost surely bounded.
assumptionp{(ref)$^*$} $\mathbb{V}[Y | X = x]$ is bounded, differentiable except at $x=0$, has bounded derivative, $\mathbb{V}[Y|X=0^+]>0$, and $\mathbb{V}[Y|X=0^-]>0$, where $0^+$ and $0^-$ denote side limits as $x\to 0$.
propositionSuppose Assumptions (ref), (ref), (ref), (ref), and (ref) hold. Let $\boldsymbol{\mathrm{V}}_1=(R_1,S_1), \ldots, \boldsymbol{\mathrm{V}}_m=(R_m,S_m)$ be an iid sample from a distribution $P \in \mathcal{P}$, where $m$ grows with $n$, and consider the bias-corrected estimator $\widehat{\theta} = \theta_{m,n}(\boldsymbol{\mathrm{V}}_{1}, \ldots, \boldsymbol{\mathrm{V}}_{m})$ as described in the text. Then, $\widehat{\theta}$ satisfies Assumptions (ref) and (ref) with $\psi_{n}(\boldsymbol{\mathrm{V}}, P) = K\left( \frac{ R }{h_n} \right) \left( \frac{S - m_{S|R}(R;P)}{\sqrt{h_n} f_{R}(0^+;P)/2 } \right)$ and $\xi^2(P) = \frac{ v_{S|R}(0^+;P) \kappa_{0,2}^+ }{ f_{R}(0^+;P)/4 },$ where $m_{S|R}(r;P)$ is the conditional mean of $S$ given $R=r$, $v_{S|R}(r;P)$ is the conditional variance of $S$ given $R=r$, and $f_{R}(r;P)$ is the PDF of $R$ at $r$, all three assuming $\boldsymbol{\mathrm{V}}=(R,S) \sim P \in \mathcal{P}$.

The proof is in Section D.3 of the supplement. The conditions and proof of Proposition (ref) follow along the lines of the conditional mean case (Proposition (ref)). Unlike Section (ref), the evaluation point $x$ lies at the boundary of the support of $X_k$. As a result, $h$ has to converge faster to zero to bound the asymptotic bias of $T_n$ (i.e., $\sqrt{nh}h \to c$ in Assumption (ref) vs. $\sqrt{nh}h^2 \to c$ in Assumption (ref)). Proposition (ref) is extended to LPR estimators of any order $\rho$ in Section D.5 of the supplement. Common practice in RDD uses local-linear regression ($\rho=1$) with a MSE-optimal bandwidth, and the resulting asymptotic bias is often corrected by using a local-quadratic regression ($\rho=2$).

If agents do not manipulate $X$ to change their treatment status, which is a key assumption in RDD, then the PDF of $X$ should be continuous at the cutoff. This implies that $f_X(0^+) = f_{X_1}(0^+)\lambda = f_{X_2}(0^+)(1-\lambda) = f_X(0^-)=f_X(0)$. In this case, under the null hypothesis, the asymptotic variance of $T_n$ and of the permutation distribution are, respectively, $\sigma^2 = \frac{4 \kappa_{0,2}^+}{f_X(0)} ( v_{Y_1|X_1}(0^+) + v_{Y_2|X_2}(0^+) )$ and $\tau^2 = \frac{ \kappa_{0,2}^+ }{ \lambda (1-\lambda) f_X(0)} \left( v_{Y_1|X_1}(0^+) + v_{Y_2|X_2}(0^+) \right).$ These are generally different, except when $\lambda = 1/2$. Thus, in general, the researcher must use the studentized test statistic for the permutation test to have asymptotic size control.

Discontinuity of Density

In many settings, the distribution of a random variable may exhibit a discontinuity at a given point if a certain phenomenon of interest occurs. For example, estimating agents' responses to incentives is a central objective in the social sciences. A continuous distribution of agents who face a discontinuous schedule of incentives results in a distribution of responses with a discontinuity at a known point. For example, saez2010 looks for evidence of a mass point in the distribution of reported income at tax brackets as evidence of agents' responses to tax rates. caetano2015 proposes an exogeneity test in nonparametric regression models, where the distribution of the potentially endogenous regressor may have a mass point. Identification of causal effects with RDD depends heavily on continuity assumptions, and these imply that the PDF of the control variable is continuous at the cutoff.

Consider a scalar random variable $X$ with PDF $f$ that is continuous, except for point $x=0$. We want to test the null hypothesis of continuity of the PDF at $x=0$. For a sample with $n$ iid observations $X_i$, a kernel density estimator for the size of the discontinuity is

align*[align* omitted — 140 chars of source]

The problem fits in our two-population framework by randomly splitting the sample as follows. Make $n_1 \doteq \lfloor n/2 \rfloor $ and $n_2 \doteq n-n_1$. For observations $1 \leq i \leq n_1$, set $W_i=1$ and let $\boldsymbol{\mathrm{Z}}_{1,i} = X_{1,i} = X_i$; for observations $n_1 < i \leq n$, set $W_i=2$ and let $\boldsymbol{\mathrm{Z}}_{2,i-n_1} = X_{2,i-n_1} = -X_{ i}$. This implies that $n_1 / n \to 1/2.$ Permutations re-order observations in $\underline{\boldsymbol{\mathrm{Z}}}_n = (X_{1,1},\ldots, X_{1,n_1}, X_{2,1},\ldots, X_{2,n_2})$ but keep $\underline{W}_n$ unchanged. Conditional on $\underline{W}_n$, the distribution of $X_{1,i}$ is $P_1$, which equals the distribution of $X$. Likewise for $X_{2,i}$, $P_2$ is the distribution of $-X$.\footnote{We cannot split the sample based on $X$ being above or below $0$ as we do in Section (ref). If we split the sample based on $X$ and the distribution of $X$ is asymmetric, it becomes impossible to identify the side limit of $f$ at 0 using only data from either sample, as required by Assumption (ref).}

Let $\boldsymbol{\mathrm{V}}$ be a scalar variable $ R \sim P \in \mathcal{P}$. The parameter of interest is $\theta(P) = 1/2 (f_{R}(0^+;P) - f_{R}(0^-;P))$, where $f_{R}(0^+;P)$ and $f_{R}(0^-;P)$ are the side-limits at zero of the PDF of $R$ under the $P$ distribution. The discontinuity parameter is $\theta(P_1)-\theta(P_2)$, which equals $f_X(0^+)-f_X(0^-)$ in terms of the PDF of $X$. The kernel estimator for the density discontinuity becomes $\ha\theta_1^b - \ha\theta_2^b$, where $\ha \theta_k^b$ is defined as $\widehat{\theta}_{k}^b = \frac{1}{ n_k h} \sum\limits_{i=1}^{n_k} K \left( \frac{ X_{k,i} }{ h } \right) \left(\mathbb{I}\{X_{k,i} \ge 0\} - \mathbb{I}\{X_{k,i} < 0\} \right) \text{, for }k=1,2.$

The $b$ superscript denotes asymptotic bias in the limiting distribution of $\sqrt{n_k h} ( \widehat{\theta}_{k}^b - \theta(P_k) )$ whenever the bandwidth choice converges to zero at the slowest possible rate, i.e., $h=O(n^{-1/3})$ (Proposition (ref)). This is the case of MSE-optimal bandwidths, and inference requires bias correction in this case. A conventional solution is to subtract a first-order bias term $h B(P_k)$ from $\widehat{\theta}_{k}^b$, where $B(P_k)$ is nonparametrically estimated by $\widehat{B}_k = B_{n_k,n}(\boldsymbol{\mathrm{Z}}_{k,1}, \ldots, \boldsymbol{\mathrm{Z}}_{k,n_k})$. We give the analytical formulas for $B(P_k)$ and $\widehat{B}_k$ in Section D.4 (Equation D.24) of the supplemental appendix. Our permutation tests utilize the bias-corrected NW estimator $\widehat{\theta}_{k} = \theta_{n_k,n}(\boldsymbol{\mathrm{Z}}_{k,1}, \ldots, \boldsymbol{\mathrm{Z}}_{k,n_k}) \doteq \theta_{n_k,n}^b(\boldsymbol{\mathrm{Z}}_{k,1}, \ldots, \boldsymbol{\mathrm{Z}}_{k,n_k}) - h B_{n_k,n}(\boldsymbol{\mathrm{Z}}_{k,1}, \ldots, \boldsymbol{\mathrm{Z}}_{k,n_k}) =\widehat{\theta}_{k}^b - h \widehat{B}_k$. Note that no bias correction is needed if $h=o(n^{-1/3})$ because $\sqrt{n h} (\widehat{\theta}_{k}^b - \theta(P_k) - h B(P_k)) = \sqrt{n h} (\widehat{\theta}_{k}^b - \theta(P_k)) +o(1)$. Alternative solutions to the bias issue include using transformed-kernel estimators (marron1994) or LPR density estimators (cattaneo2020).

It is important to emphasize that if the sample size $n$ is even and split in half, then the test statistic $T_n=\sqrt{nh}\left(\ha \theta_1 - \ha \theta_2 \right)$ is invariant to the way the original sample is split. In case the distribution of $X$ is symmetric at $0$, then $P_1=P_2$ and the permutation test in (ref) has exact size in finite samples. For other cases, we rely on asymptotic size control, and thus need to verify Assumptions (ref) and (ref).

propositionSuppose Assumptions (ref), (ref), and (ref) hold. Let $\boldsymbol{\mathrm{V}}_1=(R_1,S_1), \ldots, \boldsymbol{\mathrm{V}}_m=(R_m,S_m)$ be an iid sample from a distribution $P \in \mathcal{P}$, where $m$ grows with $n$, and consider the bias-corrected estimator $\widehat{\theta} = \theta_{m,n}(\boldsymbol{\mathrm{V}}_{1}, \ldots, \boldsymbol{\mathrm{V}}_{m})$ as described in the text. Then, Assumptions (ref) and (ref) hold for $\widehat{\theta}$ with $\psi_{n}(\boldsymbol{\mathrm{V}}, P) = h_n^{-1/2} K\left( \frac{ R }{h_n} \right) \left(\mathbb{I}\{ R \ge 0\} - \mathbb{I}\{R < 0\} \right) - h_n^{-1/2} \mathbb{E}_P[ K\left( \frac{ R }{h} \right)(\mathbb{I}\{ R \ge 0\} - \mathbb{I}\{ R < 0\} )]$ and $\xi^2(P) = \kappa_{0,2}^+\left( f_{R}(0^+;P) + f_{R}(0^-;P) \right) $.

The proof of Proposition (ref) is in Section D.4 in the supplement. The asymptotic variance of $T_n$ and of the permutation distribution are, respectively,

align*[align* omitted — 208 chars of source]

These are the same regardless if the null hypothesis is true or not. Thus, unlike the previous examples, we do not need to studentize the test statistic for asymptotic validity of the permutation test.

Confidence Sets

This section constructs robust confidence sets for a discrepancy measure $\Psi (P_1,P_2)$ between the two populations by “inverting” the permutation test for the null hypothesis $\Psi (P_1,P_2) = \delta$, $\delta \in \mathbb{R}$. The discrepancy measure satisfies two requirements. First, there exists a unique $\delta_0 \in \mathbb{R}$ such that $\Psi (P_1,P_2)=\delta_0$ is equivalent to $\theta(P_1) = \theta(P_2)$. Second, for every $\delta \neq \delta_0$, there exists a known data transformation $\psi_{\delta}$ that applies to observations from the first sample such that the distribution $\ti P_1$ of $\psi_{\delta}(\boldsymbol{\mathrm{Z}}_{1,i})$ satisfies $\theta(\ti P_1)=\theta(P_2)$.

In terms of the examples of Sections (ref)--(ref), we set $\Psi(P_1,P_2) = \theta(P_1) - \theta(P_2)$, and it follows that $\delta_0=0$ and $\psi_{\delta}(\boldsymbol{\mathrm{Z}}_{1,i}) = (X_{1,i}, Y_{1,i} - \delta)$ satisfy the two requirements for the discrepancy measure. Note that the null hypothesis $\Psi(P_1,P_2) = \delta$ is equivalent to $\mathbb{E}[Y_1|X_1=x] - \mathbb{E}[Y_2|X_2=x] = \delta$ in the conditional mean case, to $\mathbb{Q}_{\chi}[Y_1|X_1=x] - \mathbb{Q}_{\chi}[Y_2|X_2=x] = \delta$ in the conditional quantile case, and to $\mathbb{E}[Y|X=0^+] - \mathbb{E}[Y|X=0^-] = \delta$ in the discontinuity of conditional mean case. For the discontinuity of PDF example of Section (ref), we make $\Psi(P_1,P_2) = f_{X_1}(0^+) / f_{X_2}(0^+)$, and we have that $\delta_0=1$ and $\psi_{\delta}(\boldsymbol{\mathrm{Z}}_{1,i}) = X_{1,i} \left( \delta \mathbb{I}\{X_{1,i} \geq 0 \} + (1/\delta) \mathbb{I}\{X_{1,i} < 0 \} \right)$ satisfy the two requirements. In this example, $\Psi(P_1,P_2) =\delta$ is equivalent to $f_{X}(0^+) / f_{X}(0^-) = \delta$.

Define $\phi_{\delta_0}(\underline{W}_n, \underline{\boldsymbol{\mathrm{Z}}}_n)$ to be the test described in Equation (ref) with the studentized test statistic $S_n$ of Equation (ref) replacing $T_n$. This test applies to the null hypothesis $\Psi(P_1,P_2)=\delta_0$. Next, for $\delta \neq \delta_0$, we first transform the data $\underline{\boldsymbol{\mathrm{Z}}}_n = (\boldsymbol{\mathrm{Z}}_{1,1}, \ldots, \boldsymbol{\mathrm{Z}}_{1,n_1}, \boldsymbol{\mathrm{Z}}_{2,1}, \ldots, \boldsymbol{\mathrm{Z}}_{2,n_2})$ to $\widetilde{ \underline{\boldsymbol{\mathrm{Z}}}}_n = (\widetilde{\boldsymbol{\mathrm{Z}}}_{1,1}, \ldots, \widetilde{\boldsymbol{\mathrm{Z}}}_{1,n_1}, \boldsymbol{\mathrm{Z}}_{2,1}, \ldots, \boldsymbol{\mathrm{Z}}_{2,n_2})$, where $\ti \boldsymbol{\mathrm{Z}}_{1,i} = \psi_{\delta}(\boldsymbol{\mathrm{Z}}_{1,i})$ for $i=1, \ldots, n_1$. The robust permutation test for the null hypothesis $\Psi(P_1,P_2)=\delta$ is defined as $\phi_{\delta}(\underline{W}_n, \underline{\boldsymbol{\mathrm{Z}}}_n) = \phi_{\delta_0}(\underline{W}_n, \widetilde{ \underline{\boldsymbol{\mathrm{Z}}}}_n)$.

Let $U$ be a uniform random variable in $[0,1]$ and independent of the data. The confidence set with $1-\alpha$ nominal coverage is $C(\underline{W}_n, \underline{\boldsymbol{\mathrm{Z}}}_n) \doteq \{\delta : U > \phi_\delta(\underline{W}_n, \underline{\boldsymbol{\mathrm{Z}}}_n) \}.$ The set almost-surely includes all values of $\delta$ for which the test fails to reject, and it excludes the ones the test rejects. For those values of $\delta$ for which the test outcome is randomized with rejection probability $a$, the inclusion in the confidence set occurs with probability $1-a$. The purpose of a randomized confidence set is to guarantee exact coverage whenever the test $\phi_\delta$ has exact size. A non-randomized confidence set is $\ti C(\underline{W}_n, \underline{\boldsymbol{\mathrm{Z}}}_n) \doteq \{\delta \in \mathbb{R}: \phi_\delta(\underline{W}_n, \underline{\boldsymbol{\mathrm{Z}}}_n) < 1 \}$, but its coverage is conservative, especially in small samples.

Lemma (ref) implies that $\phi_\delta(\underline{W}_n, \underline{\boldsymbol{\mathrm{Z}}}_n)$ has exact size $\alpha$ in finite samples for any $P_1$ and $P_2$ such that $\Psi(P_1,P_2) = \delta$ and $\widetilde{P}_1 = \psi_\delta P_1 = P_2$. This implies that the confidence set $C(\underline{W}_n, \underline{\boldsymbol{\mathrm{Z}}}_n)$ has exact coverage in finite samples if distributions are equal up to a transformation $\psi_{\delta}$. In the examples of Sections (ref)--(ref), exactness occurs when the distributions of $P_1$ and $P_2$ are such that there exists $\delta \in \mathbb{R}$ for which the distribution of $(Y_{1,i}-\delta,X_{1,i})$ equals that of $ (Y_{2,i} ,X_{2,i})$; in Section (ref), when there exists $\delta \in \mathbb{R}$ for which the distribution of $(Y-\delta,X)|X \geq 0$ equals that of $(Y , - X) | X<0$; and in Section (ref), when there exists $\delta \in (0,\infty)$ such that $\mathbb{I}\{x \geq 0\} f_X(x/\delta)/ \delta + \mathbb{I}\{x < 0\} f_X(x \delta) \delta$ is symmetric around $x=0$. If these restrictions do not apply, then the confidence set has correct coverage asymptotically.

corollaryConsider the discrepancy measure $\Psi$ and class of data transformations $\psi_\delta$ discussed above. For any $n$, $\boldsymbol{\mathrm{Q}}_n$, $P_1$, and $P_2$, if $\widetilde{P}_1 = \psi_\delta P_1 = P_2$ for $\delta = \Psi( P_1 , P_2 )$, then \\ $\mathbb{P} \left[ \Psi( P_1 , P_2 ) \in C(\underline{W}_n, \underline{\boldsymbol{\mathrm{Z}}}_n) \right] = 1 -\alpha.$ Assume instead that $\widetilde{P}_1 = \psi_\delta P_1 \neq P_2$ and Assumptions (ref)--(ref) hold. Then, as $n \to \infty$, $\mathbb{P} \left[ \Psi( P_1 , P_2 ) \in C(\underline{W}_n, \underline{\boldsymbol{\mathrm{Z}}}_n) \right] \to 1-\alpha.$\footnote{The data transformation $\psi_\delta$ is used to obtain finite sample exactness when distributions are equal up to a transformation $\psi_\delta$, but it is not necessary for correct asymptotic coverage. In fact, for the null hypothesis $\theta(P_1)-\theta(P_2) =\delta$, one may construct a permutation test $ \phi^*_\delta$ that compares the value of $S_n - \sqrt{nh}\delta/\widehat{\sigma}_n$ with the critical values from $\widehat{R}_{S_n}$. The test $ \phi^*_\delta$ has correct size asymptotically, and the confidence set constructed by inverting $ \phi^*_\delta$ has correct coverage in large samples. }

Monte Carlo Simulations

We conduct Monte Carlo simulations to compare the finite sample performance of our permutation test to the conventional $t$-test, that is, the test that rejects the null if $|S_n| > 1 - \Phi(1-\alpha/2)$. The goal is not to show that the permutation test dominates the $t$-test in all cases; instead, the goal is to verify the theoretical predictions of Section (ref) and explore DGP variations that illustrate pros and cons of our methods. The exercise confirms the theoretical findings of size control in large samples and in finite samples under the sharp null; it also shows similar power curves between permutation and $t$-tests in large samples. Moreover, we find several cases where the permutation test performs significantly better than the $t$-test, both in power and in size control outside of the sharp null. We also compare our permutation test to two other popular resampling procedures, namely, the bootstrap and subsample, and the permutation test compares favorably to them.

The basic setup of our DGP is as follows. For $k=1,2$, $X_k \sim U[0,1]$, $\varepsilon_k \sim N(0,\sigma_k^2)$, where $X_k$ is independent of $\varepsilon_k$, $Y_k = m_{y,k}(X_k) + \varepsilon_k$, and the variances are $(\sigma_1^2,\sigma_2^2) \in \{ (1,1), (1,5) \}$. The experiments simulate iid samples from the following designs:

Design 1: $m_{y,k}(x) = g_k(x)$, where $g_{1}(x) \doteq 5(x-0.2)(x-0.8) \mathbb{I}\{|x-0.5|>0.3\}$ and $g_{2}(x) \doteq -15 (x-0.2)(x-0.8) \mathbb{I}\{ |x-0.5|>0.3 \} $; the sample sizes are $(n_1,n_2) \in \{(100,1900), (250,4750), (500,9500), (40,1960), (100,4900), (200, 9800)\}$; and the null hypothesis is $H_0: \theta(P_1) = m_{y,1}(0.5) = m_{y,2}(0.5) = \theta(P_2)$;

Design 2: $m_{y,k}(x)=1+ g_k(x)$ for $g_k(x)$ of Design 1; $D_k = m_{d,k}(X_k) + \varepsilon_{d,k}$ for mutually independent $(X_k, \varepsilon_k, \varepsilon_{d,k})$ and $\varepsilon_{d,k} \sim N(0,1)$; $m_{d,k}(x) = \mu \in \{1,10\}$; the sample sizes are $(n_1,n_2) \in \{ (75,75), (150,150), (1000, 1000)\}$; and the null hypothesis is $H_0: \theta(P_1) = m_{y,1}(0.5)/m_{d,1}(0.5) = m_{y,2}(0.5)/m_{d,2}(0.5) = \theta(P_2)$.

Design 1 is an example of the controlled means case studied in Section (ref), and Design 2 is a variation of that case. Design 2 has controlled means of two different variables with the ratio of means being of interest. The asymptotic behavior of the ratio estimator can be obtained via the Delta method, and Assumptions (ref) and (ref) can be verified using arguments similar to those in Section (ref).

These designs represent practical situations in which the $t$-test is known to perform poorly in small samples. Design 1 corresponds to cases of sample imbalance, that is, cases where the sample sizes are very different. For example, a researcher has a much larger sample of women than men and is interested in comparing the average outcome from a professional training between men and women, after conditioning on a test score. Design 2 encompasses scenarios where the ratio of conditional mean functions is of interest, and the denominator may be small. An example is estimating the efficacy rate of a vaccine conditional on blood pressure, and comparing this rate between men and women. The efficacy rate in vaccine trials is the difference in proportions of infected individuals between treatment and control groups, divided by the proportion in the control group. Both designs have conditional mean functions for $Y$ given $X$ that are equal and flat for $|X-0.5| \leq 0.3$, but different and non-linear otherwise (Figure 4 in Section E of the supplement). This shape of conditional mean function allows us to experiment with scenarios with or without estimation bias, and inside or outside the sharp null, depending on the choice of bandwidth. Both designs fall under the null hypothesis; they fall under the sharp null if we further set $\sigma_1^2 = \sigma_2^2$ and restrict the sample to $|X-0.5| \leq 0.3$.

The test statistic $T_n$ is the difference of consistent estimators $\widehat{\theta}_1 - \widehat{\theta}_2$ multiplied by $\sqrt{nh}$ (Equation (ref)). The studentized statistic $S_n$ equals the difference of consistent estimators divided by the standard error of the difference (Equation (ref)). The conditional mean functions at point $0.5$ are consistently estimated by local linear regressions with triangular kernel and a bandwidth choice $h$ that shrinks to zero as $n$ increases. A practical choice for $h$ is the estimated MSE-optimal bandwidth for local-linear regression (LLR), which decreases at rate $n^{-1/5}$. In particular, we adapt the algorithm of imbens2012, denoted IK bandwidth, to our setting. This choice of bandwidth implies that $T_n$ and $S_n$ have asymptotic distributions not centered at zero. Thus, we employ local quadratic regressions, using the same kernel and bandwidth as before, to construct the test statistics and avoid the asymptotic bias.\footnote{For more details on bias correction see discussion in Section 3.1. Section D.5 of the supplement demonstrates validity of our permutation tests with the LPR estimator.} We use White's robust formula for local quadratic regressions to compute standard errors, where the the squared residuals are obtained by the nearest-neighbor matching estimator using three neighbors (abadie2006).

We consider 10,000 simulated samples and 1,000 random permutations for each variation of the two designs. We compare the null rejection probability of three different tests at 5% nominal size: the non-studentized permutation test (NSP), the studentized permutation test (SP), and the $t$-test ($t$). All tests use the same bandwidth choice, and we experiment with four possibilities: three fixed choices of $h$, $0.1,$ $0.3,$ and $0.5$, and the IK data-driven MSE-optimal bandwidth $\widehat{h}_{mse}$.

Table (ref) displays simulated rejection rates under the null hypothesis. DGPs with $\sigma_1^2=\sigma_2^2$ and $h=0.1$ or $0.3$ fall under the sharp null hypothesis. As the theory predicts, both NSP and SP control size in these cases, but $t$ fails to do so, most notably in Design 1 with $n_1=40$ and Design 2 with $\mu=1$. Models with $\sigma_1^2=\sigma_2^2$ and $h=0.5$ fall outside the sharp null, and the local quadratic estimators are biased. Bias makes the mean of $T_n$ diverge to infinity as the sample size increases, which explains the increasing size distortion of all tests. In these cases, all tests fail to control size, although the distortions are smaller for SP than for $t$.

The rows of Table (ref) with $\sigma_1^2 \neq \sigma_2^2$ fall outside of the sharp null. Cases with $\sigma_1^2 \neq \sigma_2^2$ and $h\leq 0.3$ violate the sharp null, but the distribution of $T_n$ does not diverge as in the case of $h=0.5$. Design 1 with $h=0.1$ or $0.3$ has a large size distortion of NSP and a small size distortion of SP that decreases with $n$, as predicted by our theory. The size distortion of SP is much smaller than that of $t$, especially for smaller samples. For Design 2 with $h=0.1$ or $0.3$, the size distortions of the permutation test are again smaller than $t$ and decrease with $n$. Finally, the IK MSE-optimal bandwidth $\widehat{h}_{mse}$ balances bias and variance, and does a good job keeping low size distortions of SP in all cases.

Since hall1990, many authors have proposed resampling procedures to test nonparametric hypotheses. It is natural to ask how some of these procedures compare to our robust permutation test. We implement the tests of Designs 1 and 2 using critical values for $S_n$ generated by the wild bootstrap of cao1991 and the subsample of politis1999subsampling. The last two columns of Table (ref) report the simulated rejection rates of studentized bootstrap (SB) and studentized subsample (SS) using the IK MSE-optimal bandwidth (Table 4 in Section E of the supplement reports the rates for fixed $h$). The performance of SB and SS are generally similar to that of the $t$-test, with SS having worse size control in some cases.\footnote{ It is worth noting that SB takes longer than SP to compute because it requires re-estimation of $\widehat{m}_k(X_{k,i})$ for two different bandwidths and multiple observations $i$ (cao1991, page 2227). We compare the computation time of all tests with two empirical illustrations in the next section (Table (ref)).}

sidewaystable[htbp] \caption{Simulated Rejection Rates - 5% Nominal Size} \begin{minipage}{18cm} \begin{tabular}{c c c c| c c c| c c c| c c c| c c c c c} \hline \hline \multicolumn{4}{c|}{Design 1} & \multicolumn{3}{c|}{$h=0.1$} & \multicolumn{3}{c|}{$h=0.3$} & \multicolumn{3}{c|}{$h=0.5$} & \multicolumn{5}{c}{$\widehat{h}_{mse}$} \\ $\sigma_1^2$ & $\sigma_2^2$ & $n_1$ & $n_2$ & NSP & SP & $t$ & NSP & SP & $t$ & NSP & SP & $t$ & NSP & SP & $t$ & SB & SS \\ \hline 1 & 1 & 100 & 1900 & 0.048 & 0.049 & 0.093 & 0.047 & 0.046 & 0.062 & 0.158 & 0.147 & 0.177 & 0.047 & 0.045 & 0.065 & 0.067 & 0.090 \\ 1 & 1 & 250 & 4750 & 0.048 & 0.047 & 0.065 & 0.049 & 0.048 & 0.055 & 0.327 & 0.321 & 0.340 & 0.048 & 0.046 & 0.056 & 0.057 & 0.056 \\ 1 & 1 & 500 & 9500 & 0.051 & 0.048 & 0.059 & 0.048 & 0.048 & 0.051 & 0.577 & 0.569 & 0.589 & 0.050 & 0.048 & 0.054 & 0.055 & 0.048 \\ \hline 1 & 1 & 40 & 1960 & 0.050 & 0.050 & 0.124 & 0.049 & 0.051 & 0.090 & 0.089 & 0.082 & 0.121 & 0.047 & 0.051 & 0.101 & 0.105 & 0.258 \\ 1 & 1 & 100 & 4900 & 0.051 & 0.052 & 0.099 & 0.050 & 0.048 & 0.065 & 0.166 & 0.153 & 0.181 & 0.049 & 0.050 & 0.073 & 0.078 & 0.097 \\ 1 & 1 & 200 & 9800 & 0.050 & 0.047 & 0.075 & 0.045 & 0.046 & 0.055 & 0.274 & 0.266 & 0.292 & 0.046 & 0.045 & 0.059 & 0.063 & 0.060 \\ \hline 5 & 1 & 100 & 1900 & 0.298 & 0.068 & 0.098 & 0.310 & 0.056 & 0.064 & 0.360 & 0.073 & 0.081 & 0.310 & 0.056 & 0.065 & 0.065 & 0.095 \\ 5 & 1 & 250 & 4750 & 0.314 & 0.059 & 0.069 & 0.314 & 0.053 & 0.056 & 0.426 & 0.104 & 0.112 & 0.316 & 0.053 & 0.055 & 0.052 & 0.056 \\ 5 & 1 & 500 & 9500 & 0.322 & 0.054 & 0.058 & 0.323 & 0.051 & 0.053 & 0.521 & 0.164 & 0.172 & 0.326 & 0.052 & 0.052 & 0.050 & 0.046 \\ \hline 5 & 1 & 40 & 1960 & 0.264 & 0.060 & 0.129 & 0.344 & 0.057 & 0.091 & 0.365 & 0.060 & 0.084 & 0.338 & 0.057 & 0.099 & 0.103 & 0.308 \\ 5 & 1 & 100 & 4900 & 0.335 & 0.058 & 0.100 & 0.351 & 0.053 & 0.066 & 0.392 & 0.072 & 0.085 & 0.347 & 0.051 & 0.069 & 0.069 & 0.096 \\ 5 & 1 & 200 & 9800 & 0.347 & 0.054 & 0.075 & 0.349 & 0.047 & 0.054 & 0.447 & 0.089 & 0.100 & 0.345 & 0.047 & 0.055 & 0.055 & 0.056 \\ \hline \hline \end{tabular} \end{minipage} \begin{minipage}{18cm} \begin{tabular}{c c c c| c c c| c c c| c c c| c c c c c} \hline \hline \multicolumn{4}{c|}{Design 2} & \multicolumn{3}{c|}{$h=0.1$} & \multicolumn{3}{c|}{$h=0.3$} & \multicolumn{3}{c|}{$h=0.5$} & \multicolumn{5}{c}{$\widehat{h}_{mse}$} \\ $\sigma_1^2$ & $\sigma_2^2$ & $\mu$ & $n_1=n_2$ & NSP & SP & $t$ & NSP & SP & $t$ & NSP & SP & $t$ & NSP & SP & $t$ & SB & SS \\ \hline 1 & 1 & 10 & 75 & 0.050 & 0.051 & 0.087 & 0.048 & 0.049 & 0.061 & 0.085 & 0.082 & 0.094 & 0.039 & 0.039 & 0.050 & 0.044 & 0.128 \\ 1 & 1 & 10 & 150 & 0.051 & 0.051 & 0.073 & 0.051 & 0.050 & 0.057 & 0.138 & 0.132 & 0.144 & 0.047 & 0.045 & 0.052 & 0.050 & 0.084 \\ 1 & 1 & 10 & 1000 & 0.047 & 0.049 & 0.053 & 0.050 & 0.049 & 0.052 & 0.605 & 0.606 & 0.616 & 0.048 & 0.047 & 0.051 & 0.049 & 0.055 \\ \hline 1 & 1 & 1 & 75 & 0.054 & 0.053 & 0.219 & 0.052 & 0.050 & 0.179 & 0.069 & 0.070 & 0.213 & 0.046 & 0.045 & 0.168 & 0.336 & 0.429 \\ 1 & 1 & 1 & 150 & 0.049 & 0.051 & 0.192 & 0.052 & 0.050 & 0.168 & 0.088 & 0.091 & 0.249 & 0.050 & 0.050 & 0.162 & 0.266 & 0.360 \\ 1 & 1 & 1 & 1000 & 0.047 & 0.045 & 0.167 & 0.050 & 0.049 & 0.161 & 0.333 & 0.339 & 0.582 & 0.048 & 0.048 & 0.163 & 0.184 & 0.273 \\ \hline 5 & 1 & 10 & 75 & 0.068 & 0.060 & 0.095 & 0.056 & 0.051 & 0.063 & 0.062 & 0.061 & 0.069 & 0.044 & 0.044 & 0.053 & 0.040 & 0.129 \\ 5 & 1 & 10 & 150 & 0.059 & 0.058 & 0.076 & 0.055 & 0.054 & 0.060 & 0.079 & 0.076 & 0.080 & 0.046 & 0.045 & 0.051 & 0.044 & 0.076 \\ 5 & 1 & 10 & 1000 & 0.050 & 0.048 & 0.053 & 0.047 & 0.048 & 0.050 & 0.253 & 0.254 & 0.260 & 0.048 & 0.049 & 0.051 & 0.049 & 0.053 \\ \hline 5 & 1 & 1 & 75 & 0.057 & 0.066 & 0.159 & 0.055 & 0.058 & 0.111 & 0.063 & 0.070 & 0.124 & 0.048 & 0.054 & 0.104 & 0.228 & 0.309 \\ 5 & 1 & 1 & 150 & 0.055 & 0.060 & 0.128 & 0.053 & 0.053 & 0.098 & 0.073 & 0.080 & 0.135 & 0.047 & 0.051 & 0.091 & 0.157 & 0.220 \\ 5 & 1 & 1 & 1000 & 0.050 & 0.049 & 0.092 & 0.046 & 0.048 & 0.087 & 0.201 & 0.208 & 0.298 & 0.047 & 0.048 & 0.087 & 0.096 & 0.145 \\ \hline \hline \end{tabular} \end{minipage} \caption*{ Notes: The table displays simulated rejection rates under the null hypothesis for the non-studentized permutation test (NSP), studentized permutation test (SP), and $t$-test ($t$). Rows correspond to variations of Designs 1 and 2, as explained in the text. The bandwidth choice $h$ equals one of three fixed values (0.1, 0.3, and 0.5) or the IK data-driven MSE-optimal bandwidth $\widehat{h}_{mse}.$ The last two columns display simulated rejection rates under the null hypothesis for the studentized bootstrap (SB) and studentized subsample (SS) tests, both using $\widehat{h}_{mse}.$ }
sidewaysfigure[htbp] \begin{minipage}{0.49\textwidth} \caption{Design 1 - Simulated Power Curves} \begin{center} \begin{minipage}{.5 \textwidth} (a) { $\sigma_1^2=5$, $\sigma_2^2 = 1$ \\ $n_1 / (n_1+n_2) = 0.05$ \\ SP(---) and $t$(- -) } \end{minipage} \begin{minipage}{.5 \textwidth} (b) { $\sigma_1^2=5$, $\sigma_2^2 = 1$ \\ $n_1 / (n_1+n_2) = 0.02$ \\SP(---) and $t$(- -) } \end{minipage} \end{center} \begin{center} \begin{minipage}{.5 \textwidth} (c) { $\sigma_1^2= 5$, $\sigma_2^2 = 1$ \\ $n_1 / (n_1+n_2) = 0.05$ \\ SP(---), SB(- -), and SS($\cdot\hspace{-.8mm}\cdot\hspace{-.8mm}\cdot$) } \end{minipage} \begin{minipage}{.5 \textwidth} (d) { $\sigma_1^2= 5$, $\sigma_2^2 = 1$ \\ $n_1 / (n_1+n_2) = 0.02$ \\ SP(---), SB(- -), and SS($\cdot\hspace{-.8mm}\cdot\hspace{-.8mm}\cdot$)} \end{minipage} \end{center} \caption*{ Notes: Panels (a)--(b) compare power of the studentized permutation test (SP) with the $t$-test ($t$). Panels (c)--(d) compare SP with the studentized bootstrap (SB) and studentized subsample (SS). Darker lines correspond to bigger sample sizes, i.e., Panels (a) and (c): black $(n_1,n_2) = (500,9500)$, dark gray $(n_1,n_2) = (250,4750)$, and light gray $(n_1,n_2) = (100,1900)$; Panels (b) and (d): black $(n_1,n_2) = (200,9800)$, dark gray $(n_1,n_2) = (100,4900)$, and light gray $(n_1,n_2) = (40,1960)$. The x-axis represents the difference $\theta(P_1)-\theta(P_2)$ and the y-axis, the simulated probability of rejection. Sizes of all tests are artificially adjusted such that the simulated rejection under the null is always equal to 5%. All estimates use the IK MSE-optimal bandwidth $\widehat{h}_{mse}$. } \end{minipage} \begin{minipage}{0.49\textwidth} \caption{Design 2 - Simulated Power Curves} \begin{center} \begin{minipage}{.5 \textwidth} (a) { $\sigma_1^2=5$, $ \sigma_2^2 = 1$ \\ $\mu=10$ \\ SP(---) and $t$(- -) } \end{minipage} \begin{minipage}{.5 \textwidth} (b) { $\sigma_1^2=5$, $\sigma_2^2 = 1$ \\ $\mu=1$ \\ SP(---) and $t$(- -) } \end{minipage} \end{center} \begin{center} \begin{minipage}{.5 \textwidth} (c) { $\sigma_1^2= 5$, $\sigma_2^2 = 1$ \\ $\mu=10$ \\ SP(---), SB(- -), and SS($\cdot\hspace{-.8mm}\cdot\hspace{-.8mm}\cdot$) } \end{minipage} \begin{minipage}{.5 \textwidth} (d) { $\sigma_1^2= 5$, $\sigma_2^2 = 1$ \\ $\mu=1$ \\ SP(---), SB(- -), and SS($\cdot\hspace{-.8mm}\cdot\hspace{-.8mm}\cdot$) } \end{minipage} \end{center} \caption*{ Notes: Panels (a)--(b) compare power of the studentized permutation test (SP) with the $t$-test ($t$). Panels (c)--(d) compare SP with the studentized bootstrap (SB) and studentized subsample (SS). Darker lines correspond to bigger sample sizes, i.e., black $(n_1,n_2) = (1000,1000)$, dark gray $(n_1,n_2) = (150,150)$, and light gray $(n_1,n_2) = (75,75)$. The x-axis represents the difference $\theta(P_1)-\theta(P_2)$ and the y-axis, the simulated probability of rejection. Sizes of all tests are artificially adjusted such that the simulated rejection under the null is always equal to 5%. All estimates use the IK MSE-optimal bandwidth $\widehat{h}_{mse}$. \\ } \end{minipage}

We shift the conditional mean function of $Y_2$ in both designs so that $\theta(P_1) \neq \theta(P_2)$. We then examine the power of SP and compare it to that of the $t$, SB, and SS tests. A direct comparison is difficult because not all tests control size in all cases. Thus, we artificially adjust the size of the tests to make sure they have a simulated rejection rate of 5% under the null hypothesis.\footnote{For the $t$-test, we obtain critical values from the simulated distribution under the null hypothesis, and keep those critical values to examine the simulated rejection rates under the alternative hypotheses. For the permutation, bootstrap, and subsampling tests, we numerically search for a nominal level $\alpha$ that gives us the simulated rejection rate of 5% under the null hypothesis. Once that artificial nominal level is found, we fix that nominal level and compute the simulated rejection probabilities under the various alternative hypotheses. Rejection may be randomized in case of ties (Eq. (ref)) in order for the numerical search to find a solution.} Figures (ref) and (ref) display the simulated power curves for Designs 1 and 2, respectively, for cases with $\sigma_1^2 > \sigma_2^2=1$. Cases with $\sigma_1^2 = \sigma_2^2=1$ are found in Figures 5--6 of Section E in the supplement. Panels (a)--(b) compare SP (solid line) to $t$ (dashed line), and Panels (c)--(d) compare SP (solid line) to SB (dashed line) and SS (dotted line). Each panel displays curves associated with three different sample sizes, with darker colors representing larger samples. The x-axis shows $\theta(P_1) - \theta(P_2)$, and the y-axis plots the simulated probability of rejection. All estimates in our power analysis use the IK MSE-optimal bandwidth $\widehat{h}_{mse}$.

In Design 1 (Figure (ref)), SP outperforms $t$, SB, and SS, particularly in cases with $n_1/(n_1+n_2)=0.02$; the dominance over SS is more pronounced for smaller samples. In Design 2 (Figure (ref)), there is no clear pattern of dominance between SP and $t$, however, SP dominates SB and SS in all cases. The discrepancies between SP and other tests converge to zero as $n$ increases, as predicted by our theory. Overall, we conclude that SP has size control superior to other tests without substantial costs, if any, in terms of power.

Empirical Examples

In this section, we revisit two classical examples in the RDD literature: lee2008 on US House elections and ludwig2007 on the Head Start (HS) funding program. We illustrate the performance of our permutation test in practice and compare it to that of the $t$, bootstrap, and subsample tests.

lee2008 studies the electoral advantage of incumbent parties, using data on US House of Representatives elections from 1946 to 1998. Since districts where a party's candidate narrowly won an election are comparable to districts where the party's candidate lost by a small margin, the difference in the electoral outcomes between these two groups in the subsequent election identifies the causal effect of party incumbency. lee2008 finds that an incumbent party has a significant causal advantage of a 0.08 vote share increase in the next election (Table 2 of lee2008).

ludwig2007 study the effects of HS on students' health and schooling. The HS program was established in 1965 to provide preschool, health, and social services to poor children, aged three to five, and their families. The program provided technical assistance to the 300 poorest counties in the US, based on the 1960 poverty rate. This created a discontinuity in program funding between the 300th and 301st poorest counties. Ludwig and Miller compare child mortality rates above and below this cutoff, and estimate that HS reduces mortality by 1.198 per 100,000 children (see Table 3 of ludwig2007, ages 5--9, HS-related causes, 1973--1983).

We test the null hypothesis of zero discontinuity using our robust permutation test on both datasets.\footnote{We downloaded the datasets from Michal Kolesar's repository, https://github.com/kolesarm.} We estimate the discontinuities with local-quadratic regressions and MSE-optimal bandwidths $\widehat{h}_{mse}$ for local-linear regressions, as explained in Section (ref). The number of observations is 6,558 for lee2008 and 3,103 for ludwig2007, respectively.

Table (ref) reports the $p$-values of our studentized permutation test (SP) and compares them to those of the $t$, studentized bootstrap (SB), and studentized subsample (SS) tests, as implemented in Section (ref). The party incumbency effect of lee2008 is strongly significant and robust across different tests for both choices of $\widehat{h}_{mse}$, that is, the $\widehat{h}_{mse}$ of calonico2014 in the first row, and that of imbens2012 in the second row. On the other hand, ludwig2007 find the effect of HS on child mortality to be marginally significant, and we find the significance level ranges between 1% to 12%, depending on the test and choice of $\widehat{h}_{mse}$. Finally, both examples demonstrate that our permutation test is feasible to compute in mere seconds, comparing favorably with the bootstrap in terms of computation time.

table[table omitted — 1,533 chars of source]

Conclusion

Classical two-sample permutation tests for the sharp null hypothesis of equal distributions are easy to implement and have exact size in finite samples. However, for testing equality of parameters that summarize distributions, classical permutation tests fail to control size. To fix this problem, we propose robust permutation tests based on studentized test statistics. Our framework is general enough to cover both parametric and nonparametric models with two samples or one sample split into two subsamples. We also propose confidence sets with correct asymptotic coverage that have exact coverage in finite samples if population distributions are the same up to a class of transformations. In a simulation study, our permutation test has good size control and power curves in finite samples, outperforming the conventional $t$-test and other resampling methods. Finally, we illustrate our permutation test with two empirical examples and show that its computation time is feasible, comparing favorably to the bootstrap.

\ifblind \else

Acknowledgements: We wish to thank Roger Koenker, Marcelo Moreira, Joe Romano, Azeem Shaikh, Xiaofeng Shao, the editor, associate editor, and two anonymous referees for many insightful comments. The paper benefited from feedback given by seminar participants at the U. of Sao Paulo (FEARP), U. of Michigan, Boston U., U. of Essex, IAAE-Rotterdam, Bristol-ESG, KER Int'l Conference, UIUC, and Penn State. Bertanha acknowledges financial support received while visiting the Kenneth C. Griffin Department of Economics at the University of Chicago, where part of this work was conducted. \fi

onehalfspace

\setcounter{page}{1}

center[center omitted — 143 chars of source]

This supplement is organized as follows. Section (ref) presents a table listing selected publications that use permutation tests. Section (ref) contains the proofs for the lemma, theorems, and corollary in Section (ref). Section (ref) graphically illustrates (non)validity of the (non)studentized permutation test using the parametric example of Section (ref). Section (ref) presents all the proofs for Section (ref). Section (ref) plots the conditional mean functions of the simulation designs of our Monte Carlo experiments and brings additional results. Lastly, Section (ref) contains auxiliary lemmas and their proofs.

appendix\begin{singlespace} \section{A List of Publications} \begin{table}[H] \caption{Selected Publications in Social and Natural Sciences} \begin{tabular}{l l} \hline \hline \multicolumn{2}{l}{Economics} \\ \hline \citeSM{alan2019} & Qtly. J. Economics\\ \citeSM{bick2018} & Ame. Econ. Rev. \\ \citeSM{bursztyn2019} & Rev. Econ. Stud. \\ \citeSM{cunningham2018} & Rev. Econ. Stud.\\ \citeSM{rao2019} & Ame. Econ. Rev. \\ \hline \multicolumn{2}{l}{Public, Environmental & Occupational Health} \\ \hline \citeSM{agha2015} & Int. J. Epidemiol.\\ \citeSM{arnup2016} & J. Clin. Epidemiol. \\ \citeSM{berk2020} & Ame. J. Public Health\\ \citeSM{burrows2017} & MMWR Morb Mortal Wkly Rep \\ \citeSM{huang2016} & Int. J. Epidemiol.\\ \hline \multicolumn{2}{l}{Medicine, General & Internal} \\ \hline \citeSM{berk2013} & JAMA\\ \citeSM{goldfine2013} & Lancet \\ \citeSM{hansen2016} & JAMA \\ \citeSM{rajagopalan2013} & N. Engl. J. Med.\\ \citeSM{ryan2016} & Lancet \\ \hline \multicolumn{2}{l}{Political Science} \\ \hline \citeSM{bechtel2015} & Ame. J. Pol. Sci.\\ \citeSM{lax2010} & J. Politics\\ \citeSM{ramos2019} & Comp. Polit. Stud. \\ \citeSM{schafer2020} & J. Politics\\ \citeSM{wood2020} & Ame. J. Pol. Sci.\\ \hline \multicolumn{2}{l}{Psychology, Multidisciplinary} \\ \hline \citeSM{bishara2012} & Psychol. Methods\\ \citeSM{hu2014} & Psychol. Methods\\ \citeSM{jorgensen2018} & Psychol. Methods \\ \citeSM{pietschnig2015} & Perspect Psychol Sci\\ \hline \hline \end{tabular} \caption*{ Notes: The table lists selected publications from top journals in various disciplines as categorized by the Journal Citation Report produced by the Web of Science. We ranked journals by the 2020 Journal Impact Factor and searched for publications in the last decade from a subset of top journals. } \end{table} \section{Proofs of Theorems and Corollaries} \subsection{Proof of Lemma (ref) - Exact Size in Finite Samples} Summing $\phi(\underline{W}_n, \underline{\boldsymbol{\mathrm{Z}}}_n^\pi)$ over $\pi \in \boldsymbol{\mathrm{G}}_n$ and taking the conditional expectation given $\underline{W}_n$ yields \begin{align} \alpha n! = \sum_{\pi \in \boldsymbol{\mathrm{G}}_n} \mathbb{E}[ \phi(W_n, \underline{\boldsymbol{\mathrm{Z}}}_n^\pi) | \underline{W}_n] = & n! \mathbb{E}[ \phi(\underline{W}_n, \underline{\boldsymbol{\mathrm{Z}}}_n) | \underline{W}_n], \end{align} which implies that $\mathbb{E}[ \phi(\underline{W}_n, \underline{\boldsymbol{\mathrm{Z}}}_n) ] = \mathbb{E} \left\{ \mathbb{E}[ \phi(\underline{W}_n, \underline{\boldsymbol{\mathrm{Z}}}_n) | \underline{W}_n] \right\} = \alpha$. $\square$ \subsection{Proof of Theorem (ref) - Asymptotic Distributions Without Studentization} \subsubsection{Asymptotic Distribution of Test Statistic} Consider a sequence of $\underline{W}_n$ that satisfies $ (n_1/n - \lambda) \to 0 $ for $\lambda \in (0,1)$. Condition on $\underline{W}_n$, \begin{align*} T_{n} - \sqrt{nh}\left( \theta(P_1) - \theta(P_2)\right) &= \sqrt{nh } \left[ \frac{ \sqrt{n_1 h} }{ \sqrt{n_1 h} } \left( \ha\theta_1 -\theta(P_1) \right) - \frac{ \sqrt{n_2 h} }{ \sqrt{n_2 h} } \left( \ha\theta_2 - \theta(P_2) \right) \right] \\ &= \frac{ \sqrt{n } }{ \sqrt{n_1 } } \left[ \frac{1}{\sqrt{n_1}} \sum_{i=1}^{n_1} \psi_{n}(\boldsymbol{\mathrm{Z}}_{1,i}, P_1) \right] - \frac{ \sqrt{n } }{ \sqrt{n_2 } } \left[ \frac{1}{\sqrt{n_2}} \sum_{i=1}^{n_2} \psi_{n}(\boldsymbol{\mathrm{Z}}_{2,i}, P_2) \right] \\ &+ \frac{ \sqrt{n } }{ \sqrt{n_1 } } o_{P_1}(1) - \frac{ \sqrt{n } }{ \sqrt{n_2 } } o_{P_2}(1), \end{align*} where $o_{P_k}(1)$ is a term that depends on $Z_{k,1}, \ldots, Z_{k,n_k}$ and converges in probability to zero as $n\to\infty$, $k=1,2$. First, \[ \frac{ \sqrt{n } }{ \sqrt{n_1 } } o_{P_1}(1) - \frac{ \sqrt{n } }{ \sqrt{n_2 } } o_{P_2}(1) \overset{p}{\to} 0~. \] Second, for each $k=1,2$, it suffices to show that the Lindeberg condition holds. Abbreviate $\psi_{n}(\boldsymbol{\mathrm{Z}}_{k,i},P_k)$ by $\psi_{n,k}$ and $\mathbb{V}[\psi_{n}(\boldsymbol{\mathrm{Z}}_{k,i},P_k)]$ by $\delta_{n,k}^2$. For every $\varepsilon>0$ and $\zeta$ of Assumption (ref)-(ref), note that \begin{align*} \left| \frac{ \psi_{n,k} }{\delta_{n,k} } \right|^{2+\zeta} & \geq \frac{ \psi_{n,k}^2 }{\delta_{n,k}^2 } \left( \varepsilon n_k \right)^{\zeta/2} \mathbb{I}\left\{ \frac{ \psi_{n,k}^2 }{\delta_{n,k}^2} > \varepsilon n_k \right\}, \end{align*} which implies that the Lindeberg condition holds and for $\xi^2(P_k) = \lim_{n\to\infty} \mathbb{V}[ \psi_{n}(\boldsymbol{\mathrm{Z}}_{k,i},P_k)]$, \begin{align*} \frac{1}{\sqrt{n_k}} \sum_{i=1}^{n_k} \psi_{n }(\boldsymbol{\mathrm{Z}}_{k,i},P_k) \overset{d}{\to} N\left(0 ; \xi^2(P_k) \right). \end{align*} Therefore, \begin{align*} & \frac{ \sqrt{n } }{ \sqrt{n_1 } } \left[ \frac{1}{\sqrt{n_1}} \sum_{i=1}^{n_1} \psi_{n }(\boldsymbol{\mathrm{Z}}_{1,i},P_1) \right] - \frac{ \sqrt{n } }{ \sqrt{n_2 } } \left[ \frac{1}{\sqrt{n_2}} \sum_{i=1}^{n_2} \psi_{n}(\boldsymbol{\mathrm{Z}}_{2,i},P_2) \right] \\ &\overset{d}{\to} N\left(0 ; \frac{ \xi^2(P_1)}{\lambda} + \frac{ \xi^2(P_2) }{ 1-\lambda }\right), \end{align*} which shows convergence in distribution conditional on $\underline{W}_n$. Lemma (ref) gives convergence in distribution unconditionally. \subsubsection{Asymptotic Linear Representation of Permuted Test Statistic} In this and the following subsections, we make the entire analysis conditional on a sequence of $\underline{W}_n$ that satisfies $ (n_1/n - \lambda) \to 0 $ for $\lambda \in (0,1)$. Without loss of generality, re-order observations in the sample such that $\underline{\boldsymbol{\mathrm{Z}}}_n=(\boldsymbol{\mathrm{Z}}_1, \ldots, \boldsymbol{\mathrm{Z}}_{n_1}, \boldsymbol{\mathrm{Z}}_{n_1+1}, \ldots, \boldsymbol{\mathrm{Z}}_{n})= (\boldsymbol{\mathrm{Z}}_{1,1}, \ldots, \boldsymbol{\mathrm{Z}}_{1,n_1}, \boldsymbol{\mathrm{Z}}_{2,1}, \ldots, \boldsymbol{\mathrm{Z}}_{1,n_2})$ and $\underline{W}_n=(W_{1,n}, \ldots, W_{n_1,n}, W_{n_1+1,n}, \ldots, W_{n,n})= (1, \ldots, 1, 2, \ldots, 2)$. Let $\pi$ be a random permutation that is uniformly distributed over $\boldsymbol{\mathrm{G}}_n$ and independent of the data. The goal of this subsection is to show that, for the test statistic $T_n$ defined in Equation (ref), \begin{align*} T_n(\underline{W}_n, \underline{\boldsymbol{\mathrm{Z}}}_n^{\pi}) = & \frac{ \sqrt{n} }{ \sqrt{n_1} } \left[ \frac{1}{\sqrt{n_1}} \sum_{i=1}^{n_1} \psi_{n}( \boldsymbol{\mathrm{Z}}_{\pi(i)}, \overline{P}_n ) \right] \\ - & \frac{ \sqrt{n} }{ \sqrt{n_2} } \left[ \frac{1}{\sqrt{n_2}} \sum_{i=n_1+1}^{n} \psi_{n}( \boldsymbol{\mathrm{Z}}_{\pi(i)}, \overline{P}_n ) \right] + o_p(1), \end{align*} where $\overline{P}_n = p_{1,n} P_1 + p_{2,n} P_2$, $p_{k,n} = n_k/n$, $k=1,2$, and $o_p(1)$ is a term that depends on the data and the random permutation, and it converges in probability to zero. For each $k=1,2$, let $\boldsymbol{\mathrm{V}}_{1,n}, \ldots, \boldsymbol{\mathrm{V}}_{n_k,n}$ be an \textit{iid} sample from the distribution $\overline{P}_n$. As $\overline{P}_n \in \mathcal{P}~~ \forall n$, the uniform asymptotic linear representation of Assumption (ref) guarantees that \begin{align} R_{k,n}(\boldsymbol{\mathrm{V}}_{1,n}, \ldots, \boldsymbol{\mathrm{V}}_{n_k,n}) \doteq & \sqrt{ n_k h} \left( \theta_{ n_k,n }(\boldsymbol{\mathrm{V}}_{1,n}, \ldots, \boldsymbol{\mathrm{V}}_{n_k,n}) - \theta(\overline{P}_n) \right) \\ & - \left( \frac{1}{\sqrt{n_k}} \sum_{i=1}^{n_k} \psi_{n} (\boldsymbol{\mathrm{V}}_{i,n}, \overline{P}_n) \right) \overset{p}{\to} 0. \end{align} Lemma 5.3 and Remark A.2 by \citeSM{chung2013} show that \begin{align*} R_{1,n}(\boldsymbol{\mathrm{Z}}_{\pi(1)}, \ldots, \boldsymbol{\mathrm{Z}}_{\pi(n_1)}) = & \sqrt{n_1h} \left( \theta_{n_1,n}(\boldsymbol{\mathrm{Z}}_{\pi(1)}, \ldots, \boldsymbol{\mathrm{Z}}_{\pi(n_1)}) - \theta(\overline{P}_n) \right) \\ & - \left( \frac{1}{\sqrt{n_1}} \sum_{i=1}^{n_1} \psi_{n} (\boldsymbol{\mathrm{Z}}_{\pi(i)}, \overline{P}_n) \right) \overset{p}{\to} 0, \\ R_{2,n}(\boldsymbol{\mathrm{Z}}_{\pi(n_1+1)}, \ldots, \boldsymbol{\mathrm{Z}}_{\pi(n)}) = & \sqrt{n_2 h } \left( \theta_{n_2,n}(\boldsymbol{\mathrm{Z}}_{\pi(n_1+1)}, \ldots, \boldsymbol{\mathrm{Z}}_{\pi(n)}) - \theta(\overline{P}_n) \right) \\ & - \left( \frac{1}{\sqrt{n_2}} \sum_{i=n_1+1}^{n} \psi_{n } (\boldsymbol{\mathrm{Z}}_{\pi(i)}, \overline{P}_n) \right) \overset{p}{\to} 0. \end{align*} Then, \begin{align*} T_n(\underline{W}_n, \underline{\boldsymbol{\mathrm{Z}}}_n^{\pi}) &= \frac{ \sqrt{n} }{ \sqrt{n_1} } \left[ \frac{1}{\sqrt{n_1}} \sum_{i=1}^{n_1} \psi_{n}( \boldsymbol{\mathrm{Z}}_{\pi(i)}, \overline{P}_n ) \right] - \frac{ \sqrt{n} }{ \sqrt{n_2} } \left[ \frac{1}{\sqrt{n_2}} \sum_{i=n_1+1}^{n} \psi_{n }( \boldsymbol{\mathrm{Z}}_{\pi(i)}, \overline{P}_n ) \right] \\ & + \frac{ \sqrt{n} }{ \sqrt{n_1} } R_{1,n}(\boldsymbol{\mathrm{Z}}_{\pi(1)}, \ldots, \boldsymbol{\mathrm{Z}}_{\pi(n_1)}) - \frac{ \sqrt{n} }{ \sqrt{n_2} } R_{2,n}(\boldsymbol{\mathrm{Z}}_{\pi(n_1+1)}, \ldots, \boldsymbol{\mathrm{Z}}_{\pi(n)}) \\ &= \frac{ \sqrt{n} }{ \sqrt{n_1} } \left[ \frac{1}{\sqrt{n_1}} \sum_{i=1}^{n_1} \psi_{n}( \boldsymbol{\mathrm{Z}}_{\pi(i)}, \overline{P}_n ) \right] - \frac{ \sqrt{n} }{ \sqrt{n_2} } \left[ \frac{1}{\sqrt{n_2}} \sum_{i=n_1+1}^{n} \psi_{n }( \boldsymbol{\mathrm{Z}}_{\pi(i)}, \overline{P}_n ) \right] \\ & + o_p(1). \end{align*} \subsubsection{Coupling Approximation} The goal of this section is to create a data set $\underline{\boldsymbol{\mathrm{Z}}}^*_n=( {\boldsymbol{\mathrm{Z}}}_1^*, \ldots, {\boldsymbol{\mathrm{Z}}}_n^*)$ that is \textit{iid} from $\overline{P}_n $ and a permutation $\pi_0$ such that, for a random permutation $\pi$, $T_n(\underline{W}_n,\underline{\boldsymbol{\mathrm{Z}}}^{*\pi\pi_0}_n)-T_n(\underline{W}_n,\underline{\boldsymbol{\mathrm{Z}}}^{\pi}_n) \overset{p}{\to} 0$. Given that $\underline{\boldsymbol{\mathrm{Z}}}_n=(\boldsymbol{\mathrm{Z}}_1, \ldots, \boldsymbol{\mathrm{Z}}_{n_1}, \boldsymbol{\mathrm{Z}}_{n_1+1}, \ldots, \boldsymbol{\mathrm{Z}}_{n})= (\boldsymbol{\mathrm{Z}}_{1,1}, \ldots, \boldsymbol{\mathrm{Z}}_{1,n_1}, \boldsymbol{\mathrm{Z}}_{2,1}, \ldots, \boldsymbol{\mathrm{Z}}_{1,n_2})$, the distribution of $\underline{\boldsymbol{\mathrm{Z}}}_n$ is $P_{1}^{n_1} \times P_{2}^{n_2}$, that is, the independent product of distributions $P_{1}$ ($n_1$ times) and $P_{2}$ (${n_2}$ times). In what follows, we construct a dataset $\underline{\boldsymbol{\mathrm{Z}}}^*_n=( {\boldsymbol{\mathrm{Z}}}_1^*, \ldots, {\boldsymbol{\mathrm{Z}}}_n^*)$ that is \textit{iid} from $\overline{P}_n$. For observation $i=1$, draw an index $k$ out of $\{1, 2\}$ at random with probabilities $(p_{1,n}, p_{2,n})$. Given the resulting index $k$, set ${\boldsymbol{\mathrm{Z}}}_1^*={\boldsymbol{\mathrm{Z}}}_{k,1}$. Move on to $i=2$ and draw an index $k'$ as before. If the index $k'$ drawn is the same as $k$, set ${\boldsymbol{\mathrm{Z}}}_2^* = {\boldsymbol{\mathrm{Z}}}_{k,2}$. Otherwise, if the index $k' \neq k$, then set ${\boldsymbol{\mathrm{Z}}}_2^* = {\boldsymbol{\mathrm{Z}}}_{k',1}$. Keep going until you reach a point where you may run out of observations from one of the two samples. For example, for observation $i$, if you randomly pick $k=1$, but you have already exhausted all $n_1$ observations from $P_{1}$, then randomly draw ${\boldsymbol{\mathrm{Z}}}_i^*$ from $P_{1}$. We end up with $\underline{\boldsymbol{\mathrm{Z}}}^*_n$ and $\underline{\boldsymbol{\mathrm{Z}}}_n$ having many of the same observations in common. Call $D$ the random number of observations that are different. Reorder observations in $\underline{\boldsymbol{\mathrm{Z}}}^*_n$ by a permutation $\pi_0$ so that $\boldsymbol{\mathrm{Z}}_{\pi_0(i)}^*$ equals ${\boldsymbol{\mathrm{Z}}}_i$ for most $i$, except for $D$ of them. The re-ordered sample $\underline{\boldsymbol{\mathrm{Z}}}^{* \pi_0}_n$ is constructed as follows. The sample $\underline{\boldsymbol{\mathrm{Z}}}_n$ has the $n_1$ observations from $P_{1}$ appear first, then the $n_2$ observations from $P_{2}$ appear second. Simply take all the observations from $P_{1}$ in $\underline{\boldsymbol{\mathrm{Z}}}^*_n$ and place them first, up to $n_1$. The observations from $P_{1}$ in $\underline{\boldsymbol{\mathrm{Z}}}^*_n$ that are equal to observations from $P_{1}$ in $\underline{\boldsymbol{\mathrm{Z}}}_n$ are placed first and in the same order as in $\underline{\boldsymbol{\mathrm{Z}}}_n$. If there are more than $n_1$ observations from $P_{1}$ in $\underline{\boldsymbol{\mathrm{Z}}}^*_n$, put the extra observations on the side. If there are less than $n_1$, leave the spots blank and start at spot $n_1+1$ with the observations from $P_{2}$ in $\underline{\boldsymbol{\mathrm{Z}}}^*_n$. Repeat the same procedure for those observations in $\underline{\boldsymbol{\mathrm{Z}}}^*_n$ that were drawn from $P_{2}$, and place them starting at spot $n_1+1$ up to spot $n_2$. There are remaining spots in the newly created sample $\underline{\boldsymbol{\mathrm{Z}}}^{* \pi_0}_n$, and these should be filled up with the observations from $\underline{\boldsymbol{\mathrm{Z}}}^*_n$ that were placed on the side, in any order. Note that the distribution of $\underline{\boldsymbol{\mathrm{Z}}}^{* \pi_0}_n$ is also \textit{iid} $\overline{P}_n$. We first need to show that the (random) number of observations $D$, that are different between $\underline{\boldsymbol{\mathrm{Z}}}^*_n$ and $\underline{\boldsymbol{\mathrm{Z}}}_n$, is “small” as a fraction of $n$. Let $n_1^*$ denote the number of observations in $\underline{\boldsymbol{\mathrm{Z}}}^*_n$ that are generated from $P_{1}$. Then, $n_1^*$ has the binomial distribution with $(n, p_{1,n})$ parameters. So the mean of $n_1^*$ is $n p_{1,n} = n_1$. We have that $D = | n_1^* - n_1|$. \[ \mathbb{E}[D] = \mathbb{E} | n_1^* - n_1 | = \mathbb{E} | n_1^* - n p_{1,n} | \] \[ \leq [\mathbb{E}\{(n_1^* - n p_{1,n})^2\}]^{1/2} = [n p_{1,n}(1-p_{1,n})]^{1/2}= O(n^{1/2} ), \] where the inequality follows from the Jensen's inequality and $p_{1,n}(1-p_{1,n}) \leq 1/4$. By a similar argument, $\mathbb{E}[D^2]=O(n)$ and $\mathbb{V}[D] = O(n)$. Let $\pi$ be a random permutation that is uniformly distributed over $\boldsymbol{\mathrm{G}}_n$ and independent of everything else. Define $\Delta_i = 1$ for $i \leq n_1$, and $\Delta_i = -\frac{n_1}{n_2 }$ for $i > n_1$. From before, we have \begin{align*} T_n(\underline{W}_n, \underline{\boldsymbol{\mathrm{Z}}}_n^{\pi}) &= \frac{ \sqrt{n} }{ \sqrt{n_1} } \left[ \frac{1}{\sqrt{n_1}} \sum_{i=1}^{n_1} \psi_{n}( \boldsymbol{\mathrm{Z}}_{\pi(i)}, \overline{P}_n ) \right] - \frac{ \sqrt{n} }{ \sqrt{n_2} } \left[ \frac{1}{\sqrt{n_2}} \sum_{i=n_1+1}^{n} \psi_{n }( \boldsymbol{\mathrm{Z}}_{\pi(i)}, \overline{P}_n ) \right] \\ & + o_p(1) \\ &\overset{d}{=} \underbrace{ \frac{ \sqrt{n} }{ n_1 } \sum_{i=1}^{n} \Delta_{\pi(i)} \psi_{n}( \boldsymbol{\mathrm{Z}}_{i}, \overline{P}_n ) }_{\doteq T_n^{\pi}} + o_p(1) = T_n^{\pi} + o_p(1), \end{align*} where $A \overset{d}{=} B$ means $A$ and $B$ have the same distribution. The test statistic that uses $\underline{\boldsymbol{\mathrm{Z}}}^{* \pi_0}_n$ in the place of $\underline{\boldsymbol{\mathrm{Z}}}_n$ as initial sample and then undergoes permutation $\pi$ is \[ T_n(\underline{W}_n, \underline{\boldsymbol{\mathrm{Z}}}_n^{*\pi\pi_0}) \overset{d}{=} \underbrace{ \frac{ \sqrt{n} }{ n_1 } \sum_{i=1}^{n} \Delta_{\pi(i)} \psi_{n}( \boldsymbol{\mathrm{Z}}^*_{\pi_0(i)}, \overline{P}_n ) }_{\doteq T_n^{* \pi \pi_0}} + o_p(1) = T_n^{* \pi \pi_0} + o_p(1). \] The rest of this subsection shows that $T_{n}^{* \pi \pi_0} - T_n^{\pi} \overset{p}{\to} 0$ as $n\to \infty$. First, the expected value. \begin{align*} \mathbb{E} \left[ T_{n}^{* \pi \pi_0} - T_n^{\pi} \right] & = \frac{\sqrt{n}}{n_1} \sum_{i=1}^{n} \mathbb{E}\left[ \Delta_{\pi(i)} \right] \mathbb{E}\left[ \left( \psi_{n}( \boldsymbol{\mathrm{Z}}^*_{\pi_0(i)}, \overline{P}_n ) - \psi_{n}( \boldsymbol{\mathrm{Z}}_{i}, \overline{P}_n ) \right) \right] \\ & = 0 \end{align*} because $\pi$ is independent of everything and $\mathbb{E}[\Delta_{\pi(i)}] = 0$. Second, the variance. \begin{align} \mathbb{V}\left( T_{n}^{* \pi \pi_0} - T_n^{\pi} \right) & = \mathbb{E}\left[ \mathbb{V}\left(\left. T_{n}^{* \pi \pi_0} - T_n^{\pi} \right| D, \pi, \pi_0 \right) \right] \\ & + \mathbb{V}\left[ \mathbb{E}\left( \left. T_{n}^{* \pi \pi_0} - T_n^{\pi} \right| D, \pi, \pi_0 \right) \right]. \end{align} \underline{Part (ref) :} The elements in $\underline{\boldsymbol{\mathrm{Z}}}^{* \pi_0}_n$ and $\underline{\boldsymbol{\mathrm{Z}}}_n$ are the same except for $D$ of them. This makes all the terms in the difference $T_{n}^{* \pi \pi_0} - T_n^{\pi}$ zero, except for at most $D$ of them. Conditioning on $D$, $\pi_0$, and $\pi$, the variance is \begin{gather} \mathbb{V}\left[\left. T_{n}^{* \pi \pi_0} - T_n^{\pi} \right| D, \pi, \pi_0 \right] = \frac{ n }{n_1^2} D \mathbb{V} \left[ \left. \Delta_{\pi(i)} \left( \psi_{n}( \boldsymbol{\mathrm{Z}}_{\pi_0(i)}^*, \overline{P}_n ) - \psi_{n}( \boldsymbol{\mathrm{Z}}_{i}, \overline{P}_n ) \right) \right| D, \pi, \pi_0 \right] \nonumber \\ \leq \frac{ n }{n_1^2} D \max\left\{\left( \frac{n_1}{n_2} \right)^2, 1 \right\} \left\{ \mathbb{V}[ \psi_{n}( \boldsymbol{\mathrm{Z}}_{1}, \overline{P}_n ) ] + \mathbb{V}[\psi_{n}( \boldsymbol{\mathrm{Z}}_{2}, \overline{P}_n ) ] \right\} \\ = \frac{n}{n_1^2 } D . O(1) \end{gather} because $n_1/n_2 = O(1)$ and $\mathbb{V}[\psi_{n}( \boldsymbol{\mathrm{Z}}_{k}, \overline{P}_n )] = O(1), ~ k=1,2,$ by Assumption (ref)-(ref). Taking the expectation, \begin{gather} \mathbb{E}\left[ \mathbb{V}\left(\left. T_{n}^{* \pi \pi_0} - T_n^{\pi} \right| D, \pi, \pi_0 \right) \right] \leq \frac{n}{n_1^2 } \mathbb{E}[ D ] O(1) =o(1). \end{gather} \underline{Part (ref) :} Equation (ref) is bounded by \begin{gather} \frac{nh}{\min \{n_1^2, n_2^2\}}[\theta(P_1)-\theta(P_2)+o(1)]^2 \mathbb{V}\{D\}, \end{gather} which converges to 0 under the null. However, when the null is not imposed, we need a new argument to bound the variance in (ref). To this end, let $S$ be the number of observations among those $D$ observations that have $\Delta_{\pi(i)}=1$. Conditioning on the random drawing of indices in the coupling construction (hence conditioning on $D$ and $\pi_0$), the distribution of $S$ is Hypergeometric with $D$ draws out of $n$ elements, among which $n_1$ have $\Delta_{\pi(i)}=1$. Then, \begin{align*} &\mathbb{E} \left[ \left. T_{n}^{* \pi \pi_0} - T_n^{\pi} \right| D, \pi, \pi_0 \right] \\ & = \frac{\sqrt{n}}{n_1} \left[ S \left(\frac{n}{n_2} \right) - \left(\frac{n_1}{n_2} \right) D \right] (-1)^{\mathbb{I}\{n_1^* \leq n_1 \} } \left[ \mathbb{E} \left[ \psi_{n}( \boldsymbol{\mathrm{Z}}_{1}, \overline{P}_n ) \right] - \mathbb{E} \left[ \psi_{n}( \boldsymbol{\mathrm{Z}}_{2}, \overline{P}_n ) \right] \right]. \end{align*} Take the variance conditional on $(D,\pi_0)$, \begin{align*} &\mathbb{V}\left[ \left. \mathbb{E} \left( \left. T_{n}^{* \pi \pi_0} - T_n^{\pi} \right| D, \pi, \pi_0 \right) \right| D, \pi_0 \right] \\ & = \frac{n }{n_1^2} \mathbb{V} \left[ \left. S \left(\frac{n}{n_2} \right) - \left(\frac{n_1}{n_2} \right) D \right| D, \pi_0 \right] \left\{ \mathbb{E} \left[ \psi_{n}( \boldsymbol{\mathrm{Z}}_{1}, \overline{P}_n ) \right] - \mathbb{E} \left[ \psi_{n}( \boldsymbol{\mathrm{Z}}_{2}, \overline{P}_n ) \right] \right\}^2 \\ & = \frac{n^3 }{n_1^2 n_2^2} \mathbb{V} \left[ \left. S \right| D, \pi_0 \right] O(1) \\ & = \frac{n^3 }{n_1^2 n_2^2} D \left( \frac{n_1}{n} \right) \left( \frac{n_2}{n} \right) \left( \frac{n-D}{n-1} \right) O(1) \\ & = \frac{n^2 }{n_1 n_2 (n-1) } \left[ D - D^2 \left( \frac{1}{n} \right) \right] O(1), \end{align*} where the $\left\{ \mathbb{E} \left[ \psi_{n}( \boldsymbol{\mathrm{Z}}_{1}, \overline{P}_n ) \right] - \mathbb{E} \left[ \psi_{n}( \boldsymbol{\mathrm{Z}}_{2}, \overline{P}_n ) \right] \right\}^2=O(1)$ because of Assumption (ref)-(ref). Take the expectation, \begin{align*} & \mathbb{E}\left\{ \mathbb{V}\left[ \left. \mathbb{E} \left( \left. T_{n}^{* \pi \pi_0} - T_n^{\pi} \right| D, \pi, \pi_0 \right) \right| D, \pi_0 \right] \right\} \\ & = \frac{n^2 }{n_1 n_2 (n-1) } \left[ \mathbb{E}(D) - \mathbb{E}(D^2) \left( \frac{1}{n} \right) \right]\left\{ \mathbb{E} \left[ \psi_{n}( \boldsymbol{\mathrm{Z}}_{1}, \overline{P}_n ) \right] - \mathbb{E} \left[ \psi_{n}( \boldsymbol{\mathrm{Z}}_{2}, \overline{P}_n ) \right] \right\}^2, \end{align*} where $\left\{ \mathbb{E} \left[ \psi_{n}( \boldsymbol{\mathrm{Z}}_{1}, \overline{P}_n ) \right] - \mathbb{E} \left[ \psi_{n}( \boldsymbol{\mathrm{Z}}_{2}, \overline{P}_n ) \right] \right\}^2=O(1)$ because of Assumption (ref)-(ref). Therefore, this bound converges to 0. Furthermore, we have that \\ $\mathbb{E}\left[ \left. \mathbb{E}\left( \left. T_{n}^{* \pi \pi_0} - T_n^{\pi} \right| D, \pi, \pi_0 \right) \right| D, \pi_0 \right]=0$, which makes\\ $\mathbb{V}\left\{ \mathbb{E}\left[ \left. \mathbb{E}\left( \left. T_{n}^{* \pi \pi_0} - T_n^{\pi} \right| D, \pi, \pi_0 \right) \right| D, \pi_0 \right] \right\} = 0$, and the law of total variance makes (ref) converge to $0$. \subsubsection{Hoeffding's CLT} Let $\pi$ and $\pi'$ be permutations that are mutually independent, uniformly distributed over $\boldsymbol{\mathrm{G}}_n$, and independent of everything else. The goal of this section is to show that \begin{align} ( T_{n}^{\pi }, T_{n}^{\pi'} ) \overset{d}{\to}(T,T'), \end{align} where $T,T'$ are independent normal random variables with the same distribution. To this end, we first show that \begin{align} ( T_{n}^{*\pi}, T_{n}^{*\pi'} ) = \left( \frac{\sqrt{n}}{n_1} \sum_{i=1}^n \Delta_{\pi(i)} \psi_{n}( \boldsymbol{\mathrm{Z}}_{i}^*, \overline{P}_n ), \frac{\sqrt{n}}{n_1} \sum_{i=1}^n \Delta_{\pi'(i)} \psi_{n}( \boldsymbol{\mathrm{Z}}_{i}^*, \overline{P}_n ) \right) \overset{d}{\to}(T,T'). \end{align} By the Cram\`er--Wold device, we need to verify that \begin{align} \sum_{i=1}^n \underbrace{\frac{\sqrt{n}}{n_1}(a\Delta_{\pi(i)} + b \Delta_{\pi'(i)})}_{\doteq C_{n,i}}\psi_{n}( \boldsymbol{\mathrm{Z}}_{i}^*, \overline{P}_n ) = \sum_{i=1}^n C_{n,i} \psi_{n}( \boldsymbol{\mathrm{Z}}_{i}^*, \overline{P}_n ) \end{align} is asymptotically normal for any choice of constants $a$ and $b$, where $a\neq 0$ or $b\neq 0$. Note that $C_{n,1}, \ldots, C_{n,n}$ is a sequence of random variables that are independent of $\psi_{n}( \boldsymbol{\mathrm{Z}}_{i}^*, \overline{P}_n )$, $i=1,\ldots, n$. Call $\delta_n^2 = \mathbb{V}[\psi_{n}( \boldsymbol{\mathrm{Z}}_{i}^*, \overline{P}_n )]$. In order to apply Lemma (ref) and conclude that \[\frac{\sum_{i=1}^n C_{n,i} \psi_{n}( \boldsymbol{\mathrm{Z}}_{i}^*, \overline{P}_n ) }{\delta_n \sqrt{\sum_{l=1}^n C_{n,l}^2}} \overset{d}{\to} N(0,1),\] we need to show that there exists $\zeta>0$ for which \begin{align} & \left( \frac{\max_{i=1,\ldots, n}C_{n,i}^2}{\sum_{l=1}^n C_{n,l}^2} \right)^{\zeta/2} \mathbb{E}\left| \frac{ \psi_{n}( \boldsymbol{\mathrm{Z}}_{i}^*, \overline{P}_n ) }{\delta_{n}} \right|^{2+\zeta} \overset{p}{\to} 0. \end{align} We verify (ref) in three steps. First, we show that $\max_{i=1,\ldots, n}C_{n,i}^2=O_p(n^{-1})$: \begin{align*} & C_{n,i}^2 = \frac{n}{n_1^2}\left(a^2\Delta_{\pi(i)} ^2 + 2ab\Delta_{\pi(i)} \Delta_{\pi'(i)} + b^2 \Delta_{\pi'(i)}^2\right) = \frac{n}{n_1^2} O_p(1) , \\ &\max_{i=1, \ldots, n} C_{n,i}^2 = \frac{n}{n_1^2}O_p(1) = O_p(n^{-1}). \end{align*} Second, we derive the probability limit of $\sum_{l=1}^n C_{n,l}^2$. Note that: \begin{align} & \mathbb{E}[\Delta_{\pi(i)}] = \mathbb{E}[\Delta_{\pi'(i)}] = 0, \nonumber\\ &\mathbb{V}[\Delta_{\pi(i)}] = \mathbb{V}[\Delta_{\pi'(i)}] = \frac{n_1}{n_2}, \nonumber \\ & \mathbb{C}[\Delta_{\pi(i)},\Delta_{\pi'(i)}] =\mathbb{E}[\Delta_{\pi(i)}\Delta_{\pi'(i)}] = 0,\nonumber \\ & \mathbb{E}\left[\sum_{i=1}^n C_{n,i}^2\right] \to \frac{1}{ \lambda (1- \lambda )} \left(a^2 + b^2 \right), \nonumber \\ & \mathbb{V}\left[ \sum_{i=1}^n \frac{n}{n_1^2}\left(a^2\Delta_{\pi(i)} ^2 + 2ab\Delta_{\pi(i)} \Delta_{\pi'(i)} + b^2 \Delta_{\pi(i)'}^2\right) \right] = o(1). \nonumber \end{align} Therefore, \[\sum_{i=1}^n C_{n,i}^2 \overset{p}{\to} \frac{1}{\lambda (1-\lambda)} \left(a^2 + b^2 \right), \] which is bounded away from zero. Third, combining the two previous steps \begin{align*} \left( \frac{\max_{i=1,\ldots, n}C_{n,i}^2}{\sum_{l=1}^n C_{n,l}^2} \right) ^{\zeta/2} = O_p \left( n^{-\zeta/2}\right) \end{align*} which combined with Assumption (ref)-(ref) yields (ref). Next, we derive the limiting distribution of $\sum_{i=1}^n C_{n,i} \psi_{n}( \boldsymbol{\mathrm{Z}}_{i}^*, \overline{P}_n )$, that is, without the standardization. Since we already know the probability limit of $\sum_{l=1}^n C_{n,l}^2$, we need the limit of $\delta^2_n$. Define $\overline{P} = \lambda P_1 + (1-\lambda) P_2$. \begin{align*} \delta^2_n & = \mathbb{V}[\psi_{n}( \boldsymbol{\mathrm{Z}}_{i}^*, \overline{P}_n )] - \xi^2(\overline{P}_n) + \xi^2(\overline{P}_n) - \xi^2(\overline{P}) + \xi^2(\overline{P}) \\ & \to \xi^2(\overline{P}), \end{align*} where we used Assumption (ref)-(ref) and Assumption (ref)-(ref). Therefore, \begin{align*} \sum_{i=1}^n C_{n,i} \psi_{n}( \boldsymbol{\mathrm{Z}}_{i}^*, \overline{P}_n ) & \overset{d}{\to} N\left( 0, \xi^2(\overline{P}) \left( \frac{1}{ \lambda (1-\lambda)} \left(a^2 + b^2 \right) \right) \right). \end{align*} By the Cram\`er--Wold device, we conclude that $( T_{n}^{*\pi}, T_{n}^{*\pi'} ) \overset{d}{\to} (T,T')$, where $(T,T')$ is bivariate normal with zero means, equal variances \begin{gather} \mathbb{V}[T]=\mathbb{V}[T'] = \left( \frac{\xi^2(\overline{P})}{\lambda(1- \lambda) } \right) \doteq \tau^2 \end{gather} and zero covariance. Thus, $T$ and $T'$ are independent. So far, we have shown that $( T_{n}^{*\pi}, T_{n}^{*\pi'} ) \overset{d}{\to} (T,T')$. Consider the permutation $\pi_0$ from Section (ref). The conditions on $(\pi, \pi')$ imply that the permutations $(\pi \pi_0, \pi' \pi_0)$ are also mutually independent, uniformly distributed over $\boldsymbol{\mathrm{G}}_n$, and independent of everything else. This implies that $( T_{n}^{*\pi\pi_0}, T_{n}^{*\pi'\pi_0} ) \overset{d}{\to} (T,T')$. Finally, \[ ( T_{n}^{\pi}, T_{n}^{\pi'} ) = (T_{n}^{\pi} - T_{n}^{*\pi\pi_0}, T_{n}^{\pi'} - T_{n}^{*\pi'\pi_0}) + ( T_{n}^{*\pi\pi_0}, T_{n}^{*\pi'\pi_0} ) \overset{d}{\to} (T,T'), \] where we use the coupling argument from Section (ref) to obtain $ T_n^{\pi} - T_n^{*\pi\pi_0} \overset{p}{\to} 0$ and $ T_n^{\pi'} - T_n^{*\pi'\pi_0} \overset{p}{\to} 0$, and the Slutsky theorem. \subsubsection{Unconditional Argument} In this subsection, we apply Lemma (ref) to show the conclusion of Section (ref) also holds unconditionally. The analysis in sections (ref) and (ref) are conditional on $\underline{W}^\infty$ that satisfies $\left| \frac{n_1}{n} - \lambda\right| \to 0$. Unconditionally, $n_1$ is a random variable. Assumption (ref) implies that $ (n_1/n - \lambda) \overset{p}{\to} 0 $ for $\lambda \in (0,1)$. By Lemma (ref), $ ( T_{n}^{\pi}, T_{n}^{\pi'} ) \overset{d}{\to} (T,T')$ unconditionally. By the Hoeffding's CLT (\citeSM{lehmann2005}, Theorem 15.2.3), \[ \widehat{R}_{T_n}(t) \overset{p}{\to} \Phi\left(\frac{t}{\tau}\right), \] where $\Phi$ is the CDF of the standard normal distribution, and $\tau^2$ is given by (ref). Uniform consistency follows from continuity and monotonicity of $\Phi$ (Lemma (ref)). \subsection{Proof of Theorem (ref) - Asymptotic Distributions With Studentization} \subsubsection{Asymptotic Distribution of Test Statistic} Consider a sequence of $\underline{W}_n$ that satisfies $ (n_1/n - \lambda) \to 0 $ for $\lambda \in (0,1)$. Condition on $\underline{W}_n$ for each $n$, \begin{align*} S_{n} - \sqrt{nh} \frac{\left( \theta(P_1) - \theta(P_2)\right)}{\ha \sigma_n} &= \frac{\sqrt{nh }}{\ha \sigma_n} \left[ \left( \ha\theta_1 -\theta(P_1) \right) - \left( \ha\theta_2 - \theta(P_2) \right) \right] \\ & \overset{d}{\to} N(0, 1) \end{align*} because ${\sigma }/{\ha \sigma_n} \overset{p}{\to} 1$ by Assumption (ref), $T_n - \sqrt{nh} \left( \theta(P_1) - \theta(P_2)\right) \overset{d}{\to} N(0,\sigma^2)$ by Theorem (ref), and the Slutsky theorem. The same is true unconditional on $\underline{W}_n$ by Lemma (ref). \subsubsection{Asymptotic Permutation Distribution} Again, consider a sequence of $\underline{W}_n$ that satisfies $ (n_1/n - \lambda) \to 0 $ for $\lambda \in (0,1)$. Condition on $\underline{W}_n$ for each $n$. Without loss of generality, re-order observations in the sample such that $\underline{\boldsymbol{\mathrm{Z}}}_n=(\boldsymbol{\mathrm{Z}}_1, \ldots, \boldsymbol{\mathrm{Z}}_{n_1}, \boldsymbol{\mathrm{Z}}_{n_1+1}, \ldots, \boldsymbol{\mathrm{Z}}_{n})= (\boldsymbol{\mathrm{Z}}_{1,1}, \ldots, \boldsymbol{\mathrm{Z}}_{1,n_1}, \boldsymbol{\mathrm{Z}}_{2,1}, \ldots, \boldsymbol{\mathrm{Z}}_{1,n_2})$ and $\underline{W}_n=(W_{1,n}, \ldots, W_{n_1,n}, W_{n_1+1,n}, \ldots, W_{n,n})= (1, \ldots, 1, 2, \ldots, 2)$. Let $\pi$ be a random permutation that is uniformly distributed over $\boldsymbol{\mathrm{G}}_n$ and independent of the data. For each $k=1,2$, let $\boldsymbol{\mathrm{V}}_{1,n}, \ldots, \boldsymbol{\mathrm{V}}_{n_k,n}$ be an \textit{iid} sample from the distribution $\overline{P}_n$. Assumption (ref) guarantees that \begin{align} \xi^2_{n_k,n}(\boldsymbol{\mathrm{V}}_{1,n}, \ldots, \boldsymbol{\mathrm{V}}_{n_k,n}) - \xi^2(\overline{P}_n) \overset{p}{\to} 0. \end{align} Lemma 5.3 and Remark A.2 by \citeSM{chung2013} show that \begin{align*} \xi^2_{n_1,n}(\boldsymbol{\mathrm{Z}}_{\pi(1)}, \ldots, \boldsymbol{\mathrm{Z}}_{\pi(n_1)}) - \xi^2(\overline{P}_n) & \overset{p}{\to} 0, \\ \xi^2_{n_2,n}(\boldsymbol{\mathrm{Z}}_{\pi(n_1+1)}, \ldots, \boldsymbol{\mathrm{Z}}_{\pi(n)}) - \xi^2(\overline{P}_n) & \overset{p}{\to} 0. \end{align*} Using Assumption (ref)-(ref), \begin{align*} & \xi^2_{n_1,n}(\boldsymbol{\mathrm{Z}}_{\pi(1)}, \ldots, \boldsymbol{\mathrm{Z}}_{\pi(n_1)}) \overset{p}{\to} \xi^2(\overline{P}), \\ & \xi^2_{n_2,n}(\boldsymbol{\mathrm{Z}}_{\pi(n_1+1)}, \ldots, \boldsymbol{\mathrm{Z}}_{\pi(n)}) \overset{p}{\to} \xi^2(\overline{P}), \\ & \ha\sigma_n^{2 \pi} \doteq \frac{n}{n_1} \xi^2_{n_1,n}(\boldsymbol{\mathrm{Z}}_{\pi(1)}, \ldots, \boldsymbol{\mathrm{Z}}_{\pi(n_1)}) + \frac{n}{n_2} \xi^2_{n_2,n}(\boldsymbol{\mathrm{Z}}_{\pi(n_1+1)}, \ldots, \boldsymbol{\mathrm{Z}}_{\pi(n)}) \overset{p}{\to} \frac{\xi^2(\overline{P})}{\lambda(1-\lambda)} = \tau^2. \end{align*} By Lemma (ref), $\ha\sigma_n^{2 \pi} \overset{p}{\to} \tau^2$ unconditional on $\underline{W}_n$. Section (ref) shows that $(T_n^\pi, T_n^{\pi'})\overset{d}{\to} (T,T')$, where $(T,T')$ are independent normals with variance $\tau^2$. Given these two facts together with Theorem 5.2 by \citeSM{chung2013}, the asymptotic permutation distribution of $T_n / \tau$ is the same as that of $S_n = T_n / \ha\sigma_n^{2}$. Therefore, $\ha R_{S_n}(t) \overset{p}{\to} \Phi(t)$. Uniform consistency follows from the monotonicity and continuity of $\Phi$ (Lemma (ref)). \subsection{Proof of Corollary (ref) - Asymptotic Size and Power} \begin{proof} For $a \in (0,1)$, define $r(a) = \inf\{t: \Phi (t) \geq a \} = \Phi^{-1}(a)$ and $\ha r_n( a ) = \inf\{t: \widehat{R}_{S_n}(t) \geq a\}$. Lemma 11.2.1 by \citeSM{lehmann2005} says that $\widehat{R}_{S_n} \overset{p}{\to} \Phi$ implies $\ha r_n(a) \overset{p}{\to} r(a)$. Rewrite the test $\phi$ as, \begin{math} \phi(\lbar W_n, \lbar \boldsymbol{\mathrm{Z}}_n) = \left\{\begin{array}{l l} 1 & \mbox{if} S_n > \ha r_n(1 - \alpha/2 ) \mbox{or} S_n < \ha r_n( \alpha/2 ) , \\ a & \mbox{if} S_n = \ha r_n(1 - \alpha/2 ) \mbox{or} S_n = \ha r_n( \alpha/2 ) , \\ 0 & \mbox{if} \ha r_n( \alpha/2 ) < S_n < \ha r_n(1 - \alpha/2 ). \end{array} \right. \end{math} First, assume the null hypothesis $\theta(P_1) - \theta(P_2)$ holds. For $n \to \infty$, $S_n \overset{d}{\to} S$, where $S$ is a standard normal. \begin{align*} & S_n - \ha r_n( a ) + r( a ) \overset{d}{\to} S, \\ &\mathbb{P} \left[ S_n < \ha r_n(a) \right] \to \mathbb{P} \left[ S < r(a) \right] = a, \\ &\mathbb{P} \left[ S_n > \ha r_n(a) \right] \to \mathbb{P} \left[ S > r(a) \right] = 1- a, \\ &\mathbb{P} \left[ S_n = \ha r_n(a) \right] \to \mathbb{P} \left[ S = r(a) \right] = 0. \end{align*} Then, the probability of rejection is \begin{align*} \mathbb{E}\left[ \phi(\lbar W_n, \lbar \boldsymbol{\mathrm{Z}}_n) \right] & \\ & = \mathbb{P}\left[ S_n (\lbar W_n, \lbar \boldsymbol{\mathrm{Z}}_n) > \ha r_n(1 - \alpha/2 ) \right] + \mathbb{P}\left[ S_n (\lbar W_n, \lbar \boldsymbol{\mathrm{Z}}_n) < \ha r_n( \alpha/2 ) \right] +o(1) \\ & \to 1- (1-\alpha/2) +\alpha/2 =\alpha. \end{align*} Second, assume that $\theta(P_1) - \theta(P_2) = \eta$ and $\eta>0$ without loss of generality. For $m$ fixed and $n\to \infty$, \begin{align*} & S_n - \sqrt{nh_n} \eta /\widehat{\sigma} + \sqrt{m h_m} \eta /\widehat{\sigma} - \ha r_n( a ) + r( a ) \overset{d}{\to} S + \sqrt{m h_m} \eta /{\sigma}. \end{align*} For $m$ fixed and $n$ larger than $m$, \begin{align*} \mathbb{P} \left[ S_n < \ha r_n(a) \right] &\leq \mathbb{P} \left[ S_n - \sqrt{nh_n} \eta /\widehat{\sigma} + \sqrt{m h_m} \eta /\widehat{\sigma} - \ha r_n(a) + r(a) < r(a) \right] , \\ \lim_{n\to\infty }\mathbb{P} \left[ S_n < \ha r_n(a) \right] & \leq \mathbb{P} \left[ S + \sqrt{m h_m} \eta /{\sigma} < r(a) \right] \\ &= \Phi\left(r(a)- \sqrt{m h_m} \eta /{\sigma} \right) , \\ \mathbb{P} \left[ S_n > \ha r_n(a) \right] &\geq \mathbb{P} \left[ S_n - \sqrt{nh_n} \eta /\widehat{\sigma} + \sqrt{m h_m} \eta /\widehat{\sigma} - \ha r_n(a) + r(a) > r(a) \right] , \\ \lim_{n\to\infty } \mathbb{P} \left[ S_n > \ha r_n(a) \right] &\geq \mathbb{P} \left[ S + \sqrt{m h_m} \eta /{\sigma} > r(a) \right] \\ &= 1- \Phi\left(r(a)- \sqrt{m h_m} \eta /{\sigma} \right). \end{align*} Take limits as $m\to\infty$ from both sides, \begin{align*} \lim_{n\to\infty }\mathbb{P} \left[ S_n < \ha r_n(a) \right] & \leq \lim_{m \to\infty } \Phi\left(r(a)- \sqrt{m h_m} \eta /{\sigma} \right) =0 , \\ \lim_{n\to\infty } \mathbb{P} \left[ S_n > \ha r_n(a) \right] & \geq \lim_{m\to\infty } 1- \Phi\left(r(a)- \sqrt{m h_m} \eta /{\sigma} \right) = 1. \end{align*} Then, \begin{align*} \mathbb{E}\left[ \phi(\lbar W_n, \lbar \boldsymbol{\mathrm{Z}}_n) \right] \to 1. \end{align*} Moreover, there is no loss in power in using permutation critical values. The asymptotic test rejects when $S_n > r(1-\alpha/2)$ or $S_n < r(\alpha/2)$, where $r(a) = \Phi^{-1}(a)$ is nonrandom. Suppose \[S_n (\lbar W_n, \lbar \boldsymbol{\mathrm{Z}}_n) \overset{d}{\to} \mathcal L_{\eta}\] for some $\mathcal L_{\eta}$ under a sequence of alternatives that are contiguous to some distribution satisfying the null hypothesis and $\theta(P_{1n})-\theta(P_{2n})=\eta/\sqrt{n h_n}$. Then the power of the test against local alternatives would tend to $1-\mathcal L_{\eta}(\Phi^{-1}(1-\alpha/2)) + \mathcal L_{\eta}(\Phi^{-1}(\alpha/2))$. For the permutation test, we have that $\hat r_{n}$ obtained from the permutation distribution satisfies $\hat r_n(a) \overset{p}{\to} \Phi^{-1}(a)$ under the null hypothesis. The same results follows under the sequence of contiguous alternatives, thus implying that the permutation test has the same limiting local power as the asymptotic test which uses nonrandom critical values. \end{proof} \subsection{Proof of Corollary (ref) - Confidence Set } \begin{proof} Fix $n$ and $\boldsymbol{\mathrm{Q}}_n$ arbitrary. Pick any pair $(P_1,P_2)$. Call $\delta = \Psi( P_1, P_2 )$. By assumption, $\psi_\delta P_1 = P_2$. Lemma (ref) says that $\mathbb{E}[\phi_{\delta}(\underline{W}_n,\underline{\boldsymbol{\mathrm{Z}}}_n)] = \mathbb{E}[\phi_{\delta_0}(\underline{W}_n,\widetilde{\underline{\boldsymbol{\mathrm{Z}}}}_n)] = \alpha$. Therefore, \begin{align*} \mathbb{P}\left[\Psi( P_1, P_2 ) \in C_n (\underline{W}_n,\underline{\boldsymbol{\mathrm{Z}}}_n)\right] &= \mathbb{P}\left[U > \phi_{\delta }(\underline{W}_n,\underline{\boldsymbol{\mathrm{Z}}}_n) \right] \\ &= \mathbb{E}\left[ \mathbb{P}\left( U > \phi_{\delta }(\underline{W}_n,\underline{\boldsymbol{\mathrm{Z}}}_n) | \underline{W}_n,\underline{\boldsymbol{\mathrm{Z}}}_n \right) \right] \\ & = 1- \mathbb{E}\left[ \phi_{\delta }(\underline{W}_n,\underline{\boldsymbol{\mathrm{Z}}}_n) \right] =1- \alpha. \end{align*} Now, suppose $\psi_\delta P_1 \neq P_2$ and Assumptions (ref)--(ref) hold. Corollary (ref) says that \\ $\mathbb{E}\left[ \phi_{\delta }(\underline{W}_n,\underline{\boldsymbol{\mathrm{Z}}}_n) \right] \to \alpha$. Take the limit as $n \to \infty$ on both sides of the equality above. It follows that the asymptotic coverage of $C_n (\underline{W}_n,\underline{\boldsymbol{\mathrm{Z}}}_n)$ is $1-\alpha$. \end{proof} \section{Example (ref) (Parametric Model) - Simulation} To illustrate the conclusions of Theorems (ref) and (ref), we consider the following simulation design: for $k=1,2$, $X_k \sim U[0,1], \varepsilon_k \sim N(0, v_k)$, where $X_k$ is independent of $\varepsilon_k$, and $Y_k=\theta(P_k) + (X_k-0.5) + \varepsilon_k$; the sample sizes are $n_1=20$ and $n_2 = 980$, and the variances are $v_1 = 5$ and $v_2 = 1$. We simulate 10,000 samples under the null hypothesis $H_0: \theta(P_1) = \theta(P_2) = 0$ and use 1,000 permutations for each sample. The left-hand side of Figure (ref) plots the permutation distribution and the sampling distribution when the test statistic is not studentized. As predicted by Theorem (ref), the variance of the permutation distribution is different from that of the sampling distribution because both $n_1 \neq n_2$ and $v_1 \neq v_2$. In our simulation, critical values obtained from the permutation distribution are smaller than critical values from the sampling distribution, leading to over-rejection. In contrast, the permutation distribution based on the studentized statistic is approximately equal to the sampling distribution as depicted on the right-hand side of Figure (ref). The permutation test now has a correct size (Theorem (ref)). \begin{figure}[H] \caption{Permutation Distribution vs. Sampling Distribution} \begin{center} \end{center} \end{figure} \section{Proof of Applications} This appendix gathers the proofs of Propositions (ref) -- (ref). Key steps in these proofs consist of demonstrating uniform convergence over $\mathcal{P}$. By uniformly over $\mathcal{P}$ we mean over $P$ distribution of $\boldsymbol{\mathrm{V}}$ and $P$ argument of functions, e.g., $f_R(x;P)$ or $m_{S|R}(x;P)$. For $A_n(P)$ random function of $P$, with distribution depending on $P$, we use $A_n(P) = o_{\mathcal{P}}(1)$ to denote $\sup_{P\in \mathcal{P}} \mathbb{P}_{P}[ |A_n(P)| > \varepsilon ] \to 0 $; we also use $A_n(P) = O_{\mathcal{P}}(1)$ to denote that, for every $\delta>0$, there exists $M_{\delta}<\infty$ such that $\sup_{P\in \mathcal{P}} \mathbb{P}_{P}[ |A_n(P)| > M_{\delta} ] <\delta $. The same notation is applied when $A_n(P)$ is a deterministic function of $P$. See Definition (ref) and Lemma (ref) in Appendix (ref). \subsection{Proof of Proposition (ref) - Controlled Means} The goal of this proof is to use the assumptions listed in Proposition (ref) to verify Assumptions (ref) and (ref). It builds on standard arguments from the literature on nonparametrics. See, for example, Theorem 2.2 by \citeSM{liracine2007}. Consider an \textit{iid} sample from $P\in\mathcal{P}$ with $m$ observations, $\boldsymbol{\mathrm{V}}_1=(R_1,S_1), \ldots, \boldsymbol{\mathrm{V}}_m=(R_m,S_m)$. The number $m$ grows with $n$ such that $m_n/n \to \gamma$, for some $\gamma \in (0,1)$. The parameter of interest is $\theta(P) = \mathbb{E}[S|R=x]$, and the NW estimator is \[ \widehat{\theta}^b = \theta_{m,n}^b(\boldsymbol{\mathrm{V}}_1 , \ldots, \boldsymbol{\mathrm{V}}_m) = \frac{\sum\limits_{i=1}^{m} K \left( \frac{ R_{i}-x }{ h } \right) S_{i} } {\sum\limits_{i=1}^{m} K \left( \frac{ R_{i}-x }{ h } \right) }. \] In this section we study the asymptotic representation of the bias-corrected estimator: $\widehat{\theta} = \widehat{\theta}^b - \theta(P) - h^2 \widehat{B} $. $\widehat{B}$ is a consistent estimator for the bias term $B(P)$ (Equation (ref) below). The assumptions in Proposition (ref) imply the following facts: \begin{enumerate} • \textit{As $m\to\infty$, $h \to 0$, $m h \to \infty$, $\sqrt{mh}h^2 = O(1)$, and $\sqrt{mh}h^3 = o(1)$}; • \textit{ $\int K(u) u ~ du =0$, $\int K^r(u) u^s g(u) ~ du < \infty$ for $1 \leq r <\infty$, $0 \leq s \leq 3$, and bounded function $g(u)$}; • \textit{The distribution of $R$ has PDF $f_{R}(r;P)$ that is three times differentiable with respect to (henceforth wrt) $r$: $\nabla_r f_{R}(r;P)$, $\nabla_{r^2} f_{R}(r;P)$, and $\nabla_{r^3} f_{R}(r;P)$; these derivatives are bounded as functions of $(r,P)$; $f_{R}(r;P)$ is bounded away from zero as a function of $(r,P)$;} To see this, note that $f_{R}(r;P)$ is a convex combination of $f_{X_1}(r)$ and $f_{X_2}(r)$, each bounded, with bounded derivatives, and bounded away from zero. • \textit{$m_{S|R}(r;P) = \mathbb{E}_P[S|R=r]$ has first, second, and third derivatives wrt $r$ denoted $\nabla_r m_{S|R}(r;P)$, $\nabla_{r^2} m_{S|R}(r;P)$, and $\nabla_{r^3} m_{S|R}(r;P)$, respectively; $m_{S|R}$, $\nabla_r m_{S|R}$, $\nabla_{r^2} m_{S|R}$, and $\nabla_{r^3} m_{S|R}$ are all bounded as functions of $(r,P)$;} To see this, take $P = \alpha P_1 + (1-\alpha) P_2$ and note that, \[ m_{S|R}(r;P) = \frac{\alpha f_{X_1}(r) }{ f_{R}(r;P) } m_{Y_1|X_1}(r) + \frac{(1-\alpha) f_{X_2}(r) }{ f_{R}(r;P) } m_{Y_2|X_2}(r). \] The expectations $m_{Y_k|X_k}(r)$ are bounded functions of $r$. The weights $\omega_1(r;P) \doteq \alpha \frac{f_{X_1}(r) }{ f_{R}(r;P) }$ and $\omega_2(r;P) \doteq (1-\alpha) \frac{f_{X_2}(r) }{ f_{R}(r;P) }$ are bounded functions of $(r,P)$ because they are positive and sum to $1$. The expectations $m_{Y_k|X_k}(r)$ and the weights $\omega_k(r;P)$ are three times differentiable wrt $r$. The derivatives of $m_{Y_k|X_k}(r)$ are bounded wrt $r$. The derivatives of the weights are bounded because the derivatives of the PDFs $f_{X_k}(r)$ are bounded plus the fact that $f_{R}(r;P)$ is bounded away from zero over $(r,P)$. • \textit{$v_{S|R}(r;P) = \mathbb{V}_P[S|R=r]$ has first derivative wrt $r$ denoted $\nabla_r v_{S|R}(r;P)$; $v_{S|R}$, $\nabla_r v_{S|R}$ are both bounded as functions of $(r,P)$; $v_{S|R}(x;P)$ is bounded away from zero as a function of $P$;} Again, $v_{S|R}(r;P) = m_{S^2|R}(r;P) - m_{S|R}^2(r;P)$, where \begin{align*} &m_{S^2|R}(r;P) = \omega_1(r;P) m_{Y_1^2|X_1}(r) + \omega_2(r;P) m_{Y_2^2|X_2}(r) \\ & =\omega_1(r;P) \left[ v_{Y_1|X_1}(r) + m_{Y_1|X_1}^2(r) \right] + \omega_2(r;P) \left[ v_{Y_2|X_2}(r) + m_{Y_2|X_2}^2(r) \right], \\ & m_{S|R}^2(r;P) = \left[ \omega_1(r;P) m_{Y_1|X_1}(r) + \omega_2(r;P) m_{Y_2|X_2}(r) \right]^2. \end{align*} A similar argument to Fact (ref) shows that $v_{S|R}$ and $\nabla_r v_{S|R}$ are bounded functions of $(r,P)$. Next, $v_{Y_k|X_k}(x)= m_{Y_k^2 |X_k}(x) - m_{Y_k|X_k}^2(x)$ is bounded away from zero, so that $m_{Y_k^2 |X_k}(x)$ is bounded away from $ m_{Y_k|X_k}^2(x)$. It follows that $m_{S^2|R}(x;P) = \omega_1(x;P) m_{Y_1^2|X_1}(x) + \omega_2(x;P) m_{Y_2^2|X_2}(x)$ is bounded away from $\omega_1(x;P) m_{Y_1|X_1}^2(x) + \omega_2(x;P) m_{Y_2|X_2}^2(x)$, which is greater than or equal to $m_{S|R}^2(x;P)$. Thus, $v_{S|R}(x;P)$ is bounded away from zero as a function of $P$. • \textit{Define $\eta(r;P) = \mathbb{E}_P[|S - m_{S | R }(r;P) |^{2+\zeta} | R=r]$. $\eta(r;P)$ is a bounded function of $(r,P)$.} By the $c_r$-inequality, \begin{align*} \eta(r;P) \leq & 2^{1+\zeta} \mathbb{E}_P[|S| ^{2+\zeta} | R=r] + 2^{1+\zeta} | m_{S | R }(r;P) |^{2+\zeta} \\ = & \omega_1(r;P) \mathbb{E}[|Y_1 |^{2+\zeta} | X_1=r] + \omega_2(r;P) \mathbb{E}[|Y_2 |^{2+\zeta} | X_2=r] \\ + & 2^{1+\zeta} | m_{S | R }(r;P) |^{2+\zeta} , \end{align*} which is a bounded function of $(r,P)$ because the weights $\omega_k(r;P)$ are bounded, $\mathbb{E}[|Y_k |^{2+\zeta} | X_k=r]$ are bounded, and $m_{S | R }(r;P)$ is bounded. \end{enumerate} We re-write $\sqrt{m h} \left( \ha\theta -\theta(P) \right)$ $= \sqrt{m h} \left( \ha\theta^b -h^2 \widehat{B} -\theta(P) \right)$ to find the asymptotic linear representation. \begin{align} \sqrt{m h} \left( \ha\theta -\theta(P) \right) &= \left( \frac{1}{\sqrt{m h} } \sum_{i=1}^{m} K\left( \frac{ R_{i} - x }{h} \right) \left( S_{i} - m_{S|R}( R_i;P) \right) \right) f_{R}(x;P)^{-1} \\ &+ \left( \frac{1}{\sqrt{m h} } \sum_{i=1}^{m} K\left( \frac{ R_{i} - x }{h} \right) \left( S_{i} - m_{S|R}( R_i;P) \right) \right) \notag \\ & \left[ \left( \frac{1}{ m h } \sum_{i=1}^{m} K\left( \frac{ R_{i} - x }{h} \right) \right)^{-1} - f_{R}(x;P)^{-1} \right] \\ &+ \left( \frac{1}{\sqrt{m h} } \sum_{i=1}^{m} K\left( \frac{ R_{i} - x }{h} \right) \left( m_{S|R}( R_i;P) - m_{S|R}( x;P) \right) \right) \notag \\* & \left( \frac{1}{ m h } \sum_{i=1}^{m} K\left( \frac{ R_{i} - x }{h} \right) \right)^{-1} - \sqrt{mh} h^2 \widehat{B}. \end{align} \begin{enumerate} • \underline{Assumption (ref) - (ref):} asymptotic expansion. Equation (ref) above gives the influence function $\psi_n$. \begin{align*} \frac{1}{\sqrt{m} } \sum_{i=1}^{m} \underbrace{ K\left( \frac{ R_{i} - x }{h} \right) \left( S_{i} - m_{S|R}( R_i;P) \right) h^{-1/2} f_{R}^{-1}(x;P) }_{\doteq \psi_{n}(\boldsymbol{\mathrm{V}}_{i}, P)} = \frac{1}{\sqrt{m} } \sum_{i=1}^{m} \psi_{n}(\boldsymbol{\mathrm{V}}_{i}, P). \end{align*} We need to show that Equations (ref) and (ref) converge in probability to zero uniformly over $\mathcal{P}$. \textit{Equation (ref):} is $o_{\mathcal{P}}(1)$. We show this in 3 steps. First, \begin{align*} \mathbb{V}_P\left( \frac{1}{\sqrt{m h} } \sum_{i=1}^{m} K\left( \frac{ R_{i} - x }{h} \right) \left( S_{i} - m_{S|R}(R_i;P) \right) \right) \\ = \int K^2(u) v_{S|R}(x+uh ; P) f_{R}(x+uh;P) du, \end{align*} which is bounded over $\mathcal{P}$ because of the kernel properties (Fact (ref) above), and $v_{S|R}(r;P)$ and $f_{R}(r;P)$ are bounded functions of $(r, P)$. Next, \begin{align*} \mathbb{E}_P\left( \frac{1}{\sqrt{m h} } \sum_{i=1}^{m} K\left( \frac{ R_{i} - x }{h} \right) \left( S_{i} - m_{S|R}(R_i;P) \right) \right) = 0. \end{align*} Use Lemma (ref), part 2, to conclude that \[ \frac{1}{\sqrt{m h} } \sum_{i=1}^{m} K\left( \frac{ R_{i} - x }{h} \right) \left( S_{i} - m_{S | R}(R_i;P) \right) =O_{\mathcal{P}}(1). \] Second, \begin{align*} \mathbb{E}_P\left[ \frac{1}{ m h } \sum_{i=1}^{m} K\left( \frac{ R_{i} - x }{h} \right) \right] - f_{R}(x;P) = \int K(u) \nabla_r f_{R}(x^*_{uh};P) uh du, \end{align*} where $x^*_{uh}$ is a point between $x$ and $x+uh$. The expression above converges to zero uniformly over $\mathcal{P}$ because of kernel properties (Fact (ref)), the derivative $\nabla_r f_{R}(r;P)$ is a bounded function of $(r,P)$, and $h \to 0 $ as $m\to\infty$. Next, the variance of the same term. \begin{align*} \mathbb{V}_P\left[ \frac{1}{ m h } \sum_{i=1}^{m} K\left( \frac{ R_{i} - x }{h} \right) \right] &= \frac{1}{ m h^2 } \mathbb{V}_P\left[ K\left( \frac{ R_{i} - x }{h} \right) \right] \\ &\leq \frac{1}{ m h^2 } \mathbb{E}_P\left[ K^2\left( \frac{ R_{i} - x }{h} \right) \right] \\ & = \frac{1}{ m h } \int K^2(u) f_{R}(x +uh;P ) du, \end{align*} which converges to zero uniformly over $\mathcal{P}$ because $mh\to\infty$, kernel properties (Fact (ref)), and $f_{R}(r;P)$ is a bounded function of $(r,P)$. Use Lemma (ref), parts 2 and 3 to arrive at: \begin{align} &\frac{1}{ m h } \sum_{i=1}^{m} K\left( \frac{ R_{i} - x }{h} \right) - f_{R}(x;P) = o_{\mathcal{P}}(1). \\ &\left( \frac{1}{ m h } \sum_{i=1}^{m} K\left( \frac{ R_{i} - x }{h} \right) \right)^{-1} - f_{R}^{-1} (x;P) = o_{\mathcal{P}}(1). \end{align} Third, combine steps 1 and 2 and use Lemma (ref) - part 1: \begin{align*} &\left( \frac{1}{\sqrt{m h} } \sum_{i=1}^{m} K\left( \frac{ R_{i} - x }{h} \right) \left( S_{i} - m_{S | R }(R_i;P) \right) \right) \\ &\left[ \left( \frac{1}{ m h } \sum_{i=1}^{m} K\left( \frac{ R_{i} - x }{h} \right) \right)^{-1} - \left( f_{R}(x;P) \right)^{-1} \right] \\ & = O_{\mathcal{P}}(1) o_{\mathcal{P}}(1) = o_{\mathcal{P}}(1). \end{align*} \textit{Equation (ref):} is $o_{\mathcal{P}}(1)$. We derive the probability limit of $\eqref{eq:control_mean:thetahat:bias} + \sqrt{m h} h^2 \widehat{B}$ in 3 steps. First, $\left( \frac{1}{ m h } \sum_{i=1}^{m} K\left( \frac{ R_{i} - x }{h} \right) \right)^{-1} - f_R^{-1}(x;P)= o_{\mathcal{P}}(1) $ by what was shown above (Equations (ref) and (ref)). Second, \begin{align*} & \mathbb{E}_P\left[ \frac{1}{\sqrt{m h} } \sum_{i=1}^{m} K\left( \frac{ R_{i} - x }{h} \right) \left( m_{S | R }(R_i;P) - m_{S | R }(x;P) \right) \right] \\ =& \sqrt{mh} \int K\left( u \right) [m_{S | R }(x+uh;P) - m_{S | R }(x;P)] f_{R}(x+uh;P) du \\ =& \sqrt{mh} \int K\left( u \right) \left[ \nabla_r m_{S | R }(x;P) uh + \nabla_{r^2} m_{S | R }(x;P) u^2 h^2/2 + \nabla_{r^3} m_{S | R }(x^{*}_{uh};P) u^3 h^3/6 \right] \\ & f_{R}(x + uh;P) du \\ =& \sqrt{mh} \int K\left( u \right) \nabla_r m_{S | R }(x;P) u h [f_{R}(x;P) + \nabla_r f_{R}(x;P) uh + \nabla_{r^2} f_{R}(x^{**}_{uh};P)u^2 h^2/2 ] du \\ &+ \sqrt{mh} \int K\left( u \right) \nabla_{r^2} m_{S | R }(x;P) u^2 (h^2/2) [f_{R}(x;P) + \nabla_r f_{R}(x^{***}_{uh} ; P) uh] du \\ &+ \sqrt{mh} \int K\left( u \right) \nabla_{r^3} m_{S | R }(x^{*}_{uh};P) u^3 (h^3/6) f_{R}(x + uh;P) du, \end{align*} where we use the existence of derivatives for the Taylor expansions and that $x^*_{uh},x^{**}_{uh},$ and $x^{***}_{uh}$ are points between $x+uh$ and $x$. Define $\kappa_{s,t} = \int u^s K^t(u) ~du$. The last equation above equals to \begin{align*} =& \sqrt{mh} h \nabla_r m_{S | R }(x;P) f_{R}(x;P) \int u K\left( u \right) du \\ &+ \sqrt{mh} h^2 \nabla_r m_{S | R }(x;P) \nabla_r f_{R}(x;P) \int u^2 K\left( u \right) du \\ & + \sqrt{mh} (h^3/2) \nabla_r m_{S | R }(x;P) \int u^3 K\left( u \right) \nabla_{r^2} f_{R}(x^{**}_{uh};P) du \\ &+ \sqrt{mh} (h^2/2) \nabla_{r^2} m_{S | R }(x;P) f_{R}(x;P) \int u^2 K\left( u \right) du \\ &+ \sqrt{mh} (h^3/2) \nabla_{r^2} m_{S | R }(x;P) \int u^3 K\left( u \right) \nabla_r f_{R}(x^{***}_{uh} ; P) du \\ &+ \sqrt{mh} (h^3/6) \int u^3 K\left( u \right) \nabla_{r^3} m_{S | R }(x^{*}_{uh};P) f_{R}(x + uh;P) du, \\ =& 0 \\ &+ \sqrt{mh} h^2 \nabla_r m_{S | R }(x;P) \nabla_r f_{R}(x;P) \kappa_{2,1} \\ & + O_{\mathcal{P}}\left( \sqrt{mh} h^3\right) \\ &+ \sqrt{mh} h^2 \nabla_{r^2} m_{S | R }(x;P) f_{R}(x;P) \frac{\kappa_{2,1}}{2} \\ &+ O_{\mathcal{P}}\left( \sqrt{mh} h^3\right) \\ &+ O_{\mathcal{P}}\left( \sqrt{mh} h^3\right) \\ =& \sqrt{mh} h^2 \underbrace{ \kappa_{2,1} \left[ \nabla_r m_{S | R }(x;P) \nabla_r f_{R}(x;P) + \nabla_{r^2} m_{S | R }(x;P) f_{R}(x;P) /2 \right]}_{\doteq B_0(P) } + o_{\mathcal{P}}\left( 1\right) \\ =& \sqrt{mh} h^2 B_0(P) + o_{\mathcal{P}}\left( 1\right), \end{align*} where we use the following: (i) $\int K(u) u ~ du =0$ and other kernel properties (Fact (ref)); (ii) $\nabla_{r^3} m_{S|R} (r;P)$, $f_{R}(r;P)$, $\nabla_r f_{R}(r;P)$, and $\nabla_{r^2} f_{R}(r;P)$ are bounded functions of $(r,P)$; (iii) $\sqrt{mh} h^2 =O(1)$ and $\sqrt{mh} h^3 =o(1)$. Next, the variance. \begin{align*} & \mathbb{V}_P\left[ \frac{1}{\sqrt{m h} } \sum_{i=1}^{m} K\left( \frac{ R_{i} - x }{h} \right) \left( m_{S | R }(R_i;P) - m_{S | R }(x;P) \right) \right] \\ =& \frac{1}{ h } \mathbb{V}_P\left[ K\left( \frac{ R_{i} - x }{h} \right) \left( m_{S | R }(R_i;P) - m_{S | R }(x;P) \right) \right] \\ \leq & \frac{1}{ h } \mathbb{E}_P \left[ K^2\left( \frac{ R_{i} - x }{h} \right) \left( m_{S | R }(R_i;P) - m_{S | R }(x;P) \right)^2 \right] \\ =& h^2 \int K^2\left( u \right) [\nabla_r m_{S | R }(x^*_{uh};P)]^2 u^2 f_{R}(x+uh;P) du \\ =& O_{\mathcal{P}} \left( h^2 \right) = o_{\mathcal{P}} \left( 1 \right), \end{align*} where we use that (i) $\nabla_r m_{S|R} (r;P)$ and $f_{R}(r;P)$ are bounded functions of $(r,P)$; (ii) $ h^2 \to 0$; and (iii) kernel properties (Fact (ref)). Apply Lemma (ref)- part 2 to get \[ \frac{1}{\sqrt{m h} } \sum_{i=1}^{m} K\left( \frac{ R_{i} - x }{h} \right) \left( m_{ S | R }(R_i;P) - m_{ S | R }(x;P) \right) = \sqrt{m h}h^2 B_0(P) + o_{\mathcal{P}}(1). \] Third, combine the first and second steps and apply Lemma (ref)- part 1 to arrive at \begin{align} & \left( \frac{1}{\sqrt{m h} } \sum_{i=1}^{m} K\left( \frac{ R_{i} - x }{h} \right) \left( m_{ S | R }(R_i;P) - m_{ S | R }(x;P) \right) \right) \left( \frac{1}{ m h } \sum_{i=1}^{m} K\left( \frac{ R_{i} - x }{h} \right) \right)^{-1} \notag \\ & = \sqrt{mh}h^2 \underbrace{ \kappa_{2,1} \left[ \frac{\nabla_r m_{S | R }(x;P) \nabla_r f_{R}(x;P)}{f_{R}(x;P)} + \frac{\nabla_{r^2} m_{S | R }(x;P)}{2} \right]}_{\doteq B(P)} +o_{\mathcal{P}}(1) \\ & = \sqrt{mh}h^2 B(P) + o_{\mathcal{P}}(1), \notag \end{align} which gives the bias term $B(P)$. A consistent estimator $\widehat{B}$ is, \[ \widehat{B} = B_{m,n}(\boldsymbol{\mathrm{V}}_1 , \ldots, \boldsymbol{\mathrm{V}}_m) \doteq \kappa_{2,1} \left[ \frac{\widehat{\nabla_r m_{S|R}}(x;P) \widehat{\nabla_r f_{R}}(x;P) } { \widehat{f_{R}}(x;P) } + \frac{ \widehat{\nabla_{r^2} m_{S|R}}(x;P) }{2} \right], \] that is, by replacing $f_{R}(x;P)$, $\nabla_r f_{R}(x;P)$, $\nabla_r m_{S|R}(x;P)$, and $\nabla_{r^2} m_{S|R}(x;P)$ in $B(P)$ by consistent nonparametric estimators readily available in the literature. The tuning parameters for these additional estimators and corresponding moment conditions may be set such that the bias-correction condition (Assumption (ref)) is met. Finally, (ref) equals to, \begin{align*} & \left( \frac{1}{\sqrt{m h} } \sum_{i=1}^{m} K\left( \frac{ R_{i} - x }{h} \right) \left( m_{ S | R }(R_i;P) - m_{ S | R }(x;P) \right) \right) \left( \frac{1}{ m h } \sum_{i=1}^{m} K\left( \frac{ R_{i} - x }{h} \right) \right)^{-1} \\ &- \sqrt{mh}h^2 \widehat{B} \\ & = \sqrt{mh}h^2 \left( B(P) - \widehat{B} \right) + o_{\mathcal{P}}(1) = o_{\mathcal{P}}(1), \end{align*} where we use $\sqrt{mh}h^2 = O(1)$ and Assumption (ref) that says $ B(P) - \widehat{B} = o_{\mathcal{P}}(1)$. • \underline{Assumption (ref) - (ref):} zero mean of influence function. $\mathbb{E}_P[ \psi_{n}(\boldsymbol{\mathrm{V}}_{i}, P) ]=0 ~~\forall P$ by construction. • \underline{Assumption (ref) - (ref):} variance of influence function. Define $\xi^2(P) = \kappa_{0,2} v_{ S | R} (x;P) /f_{R}(x;P)$, where $\kappa_{0,2} = \int_{-\infty}^{\infty}K^2(u) ~ du$. \begin{align*} &\mathbb{V}_P\left( \frac{1}{f_{R}(x;P) \sqrt{h} } K\left( \frac{ R_{i} - x }{h} \right) \left( S_{i} - m_{ S | R} (R_i;P) \right) \right) - \xi^2(P) \\ =&\frac{1}{f_{R}^2(x;P) } \int K^2(u) \left\{ \underbrace{v_{ S | R }(x+uh;P) f_{R}(x+uh;P)}_{\doteq g(uh;P)} - \underbrace{ v_{ S | R}(x;P) f_{R}(x;P) }_{\doteq g(x;P)} \right\} du \\ =&\frac{h}{f_{R}^2(x;P) } \int K^2(u) \nabla_r g(x^*_{uh};P ) u du = o_{\mathcal{P}}(1). \end{align*} where $\nabla_r g(r;P )$ denotes the derivative of $g(r;P)$ wrt $r$. The expression above converges to zero uniformly over $\mathcal{P}$ because $h\to0$, Fact (ref) on the kernel, the derivative $\nabla_r g(r;P )$ is a bounded function of $(r,P)$, and $f_{R}(x;P)$ is bounded away from zero as a function of $P$. Therefore, \[ \sup\limits_{P \in \mathcal{P} } \left| \mathbb{V}_P[ \psi_n(\boldsymbol{\mathrm{V}}_{i}, {P} ) ] - \xi^2(P) \right| \to 0. \]\underline{Assumption (ref) - (ref):} $ \sup\limits_{P \in \mathcal{P} } \mathbb{E} [\psi_n^2(\boldsymbol{\mathrm{Z}}_{k,i},P)] < \infty $ for $k=1,2$. We have, \begin{align*} & \psi_n(\boldsymbol{\mathrm{Z}}_k,P) = K\left( \frac{ X_{k} - x }{h} \right) \left( Y_{k} - m_{S|R}(X_k;P) \right) h^{-1/2} f_{R}^{-1}(x;P). \\ & \psi_n^2(\boldsymbol{\mathrm{Z}}_k,P) = \frac{K^2\left( \frac{ X_{k} - x }{h} \right)}{ h f_{R}^{2}(x;P)} \left[ \left(Y_{k} - m_{Y_k|X_k}(X_k) \right) + \left(m_{Y_k|X_k}(X_k) - m_{S|R}(X_k;P) \right) \right]^2 \\ & = \frac{K^2\left( \frac{ X_{k} - x }{h} \right)}{ h f_{R}^{2}(x;P) } \left[ \left( Y_{k} - m_{Y_k|X_k}(X_k) \right)^2 + \left(m_{Y_k|X_k}(X_k) - m_{S|R}(X_k;P) \right)^2 \right. \\ & \left. + 2 \left(Y_{k} - m_{Y_k|X_k}(X_k) \right) \left(m_{Y_k|X_k}(X_k) - m_{S|R}(X_k;P) \right) \right]. \end{align*} \begin{align*} &\mathbb{E}\left[ \psi_n^2(\boldsymbol{\mathrm{Z}}_k,P) | X_k \right] = \frac{K^2\left( \frac{ X_{k} - x }{h} \right)}{ h f_{R}^{2}(x;P) } \left[ v_{Y_k|X_k}(X_k) + \left(m_{Y_k|X_k}(X_k) - m_{S|R}(X_k;P) \right)^2 \right]. \\ &\mathbb{E}\left[ \psi_n^2(\boldsymbol{\mathrm{Z}}_k,P) \right] = \frac{1}{f_{R}^{2}(x;P)} \int K^2(u) \left[ v_{Y_k|X_k}(x+uh) \right. \\ & \left. + \left(m_{Y_k|X_k}(x+uh) - m_{S|R}(x+uh;P) \right)^2 \right] f_{X_k}(x+uh) du \\ & = O_{\mathcal{P}}(1), \end{align*} because $f_{R}(x;P)$ is bounded away from zero as a function of $P$, the conditional moment functions and $f_{X_k}$ inside the integral are bounded, and Fact (ref) on the kernel. • \underline{Assumption (ref) - (ref):} $(2+\zeta)$-th moment condition. We verify it in two steps. First, \begin{align*} \mathbb{V}_P \left( \psi_n(\boldsymbol{\mathrm{V}}_i,P) \right) = &\mathbb{V}_P\left( \frac{1}{f_{R}(x;P) \sqrt{h} } K\left( \frac{ R_{i} - x }{h} \right) \left( S_{i} - m_{S | R }(R_i;P) \right) \right) \\ =&\frac{1}{f_{R}^2(x;P) } \int K^2(u) \left\{ v_{S | R}(x+uh;P) f_{R}(x+uh;P) \right\} du \end{align*} is bounded away from zero uniformly over $\mathcal{P}$ and $n$ because (i) $f_{R}(x;P)$ is bounded as a function of $P$; and (ii) $v_{S | R}(r;P)$ and $f_{R}(r;P)$ are continuous functions of $r$ and bounded away from zero at $x=r$ and over $P$. Second, for $\zeta$ of the moment condition in Proposition (ref), call $\eta(r;P) = \mathbb{E}_P[|S_i - m_{S | R }(R_i;P) |^{2+\zeta} | R_i=r]$. \begin{align*} n^{-\zeta/2} \mathbb{E}_P \left| \psi_n(\boldsymbol{\mathrm{V}}_i,P) \right| ^{2+\zeta} & \\ & =\frac{(m/n)^{\zeta/2}}{ f_{R}^{2+\zeta}(x;P) (mh)^{\zeta/2}} \int |K(u)|^{2 +\zeta} \eta(x+uh;P) f_{R}(x+uh;P) du \\ & = o_{\mathcal{P}}(1), \end{align*} because (i) $mh \to \infty$, $m/n=O(1)$; (ii) $f_{R}(x;P) = O_{\mathcal{P}}(1) $ and $f_{R}^{-1}(x;P) = O_{\mathcal{P}}(1) $; (iii) $\eta(x;P) = O_{\mathcal{P}}(1)$; and (iv) Fact (ref) on the kernel. Combining steps 1 and 2, \begin{align*} n^{-\zeta/2} \sup\limits_{P \in \mathcal{P} } \mathbb{E}_P\left| \frac{ \psi_n(\boldsymbol{\mathrm{V}}_i,P) }{ \sqrt{\mathbb{V}_P\left( \psi_n(\boldsymbol{\mathrm{V}}_i,P) \right)} } \right|^{2 + \zeta} = o(1). \end{align*} • \underline{Assumption (ref) - (ref):} $\xi^2\left( \frac{m}{n} P_1 + \frac{n-m}{n} P_2 \right) \to \xi^2 \left( \gamma P_1 + (1-\gamma ) P_2 \right)$. Let $\overline{P}_n = \frac{m}{n} P_1 + \frac{n-m}{n} P_2$ and $\overline{P} = \gamma P_1 + (1-\gamma ) P_2$. We have that \\ $\xi^2(\overline{P}_n) = \kappa_{0,2} v_{S | R }(x; \overline{P}_n) /f_{R}(x;\bar{P}_n)$, so it suffices to show that $v_{ S | R }(x; \bar{P}_n ) \to v_{ S | R }(x; \bar{P})$ and $f_{R}(x;\bar{P}_n) \to f_{R}(x;\bar{P})$. First, convergence of the PDF, \begin{align*} f_{R}(x;\bar{P}_n) = \frac{m }{n } f_{X_1}(x) + \frac{n-m }{n } f_{X_2}(x) \to \gamma f_{X_1}(x) + (1-\gamma) f_{X_2}(x) = f_{R}(x;\bar{P}). \end{align*} Second, convergence of moments. For $g(x)=x$ or $g(x)=x^2$, \begin{align*} \mathbb{E}_{ \bar{P}_n }[g(S) | R =x ] = & \frac{m f_{X_1}(x) }{ n f_{R}(x;\bar{P}_n) } \mathbb{E}[g(Y_1)|X_1=x] + \frac{(n-m) f_{X_2}(x) }{n f_{R}(x;\bar{P}_n) } \mathbb{E}[g(Y_2) | X_2=x] \\ \to & \frac{\gamma f_{X_1}(x) }{ f_{R}(x;\bar{P}) } \mathbb{E}[g(Y_1)|X_1=x] + \frac{(1 -\gamma ) f_{X_2}(x) }{f_{R}(x;\bar{P}) } \mathbb{E}[g(Y_2) | X_2=x] \\ = & \mathbb{E}_{\bar{P}} [g( S ) | R =x ], \end{align*} which implies that $v_{ S | R }(x; \bar{P}_n ) \to v_{ S | R }(x; \bar{P})$. Therefore, $\xi^2(\overline{P}_n) \to \xi^2(\overline{P})$. • \underline{Assumption (ref) :} We have that $n_1$ is a deterministic sequence and $(n_1 /n - \lambda) \to 0$ by assumption. Assumption (ref) has already been verified above for any sequence $m_n$ such that $(m/n -\gamma) \to 0$ for arbitrary $\gamma \in (0,1)$. In particular it holds for $\gamma \in \{ \lambda, 1-\lambda\}.$ \end{enumerate} $\square$ \subsection{Proof of Proposition (ref) - Controlled Quantiles} The goal of this proof is to use the assumptions listed in Proposition (ref) to verify Assumptions (ref) and (ref). It adapts arguments from \citeSM{pollard1991}, \citeSM{chaudhuri1991}, and \citeSM{fan1994}. Consider an \textit{iid} sample from $P\in\mathcal{P}$ with $m$ observations, $\boldsymbol{\mathrm{V}}_1=(R_1,S_1), \ldots, \boldsymbol{\mathrm{V}}_m=(R_m,S_m)$. The number $m$ grows with $n$ such that $m/n \to \gamma$, for some $\gamma \in (0,1)$. The parameter of interest is $\theta(P) = \mathbb{Q}_{\chi}[S|R=x]$, $\chi \in (0,1)$, and the NW-style estimator is given by \[ \hat \theta^b = \theta_{m,n}^b (\boldsymbol{\mathrm{V}}_1 , \ldots, \boldsymbol{\mathrm{V}}_m) \doteq \arg\min_{\theta} \sum_{i=1}^{m} \rho_{\chi}(S_i - \theta) K\left(\frac{R_i-x}{h}\right)~, \] where $\rho_{\chi}(u) = (\chi - \mathbb{I}(u \leq 0))u$. This section studies the asymptotic behavior of $\widehat{\theta} = \widehat{\theta}^b - h \widehat{B} $, where the expression for $\widehat{B}$ is given below Equation (ref). Define $U = S - \theta(P)$. The assumptions in Proposition (ref) imply the following facts: \begin{enumerate} • \textit{As $m\to\infty$, $h \to 0$, $m h \to \infty$, $\sqrt{m h} h^2 =O(1)$, $\sqrt{m h} h^3 =o(1)$}; • \textit{ $\int K(u) u ~ du =0$, $\int K^r(u) u^s g(u) ~ du < \infty$ for $1 \leq r <\infty$, $0 \leq s \leq 3$, and bounded function $g(u)$}; • \textit{The distribution of $R$ has PDF $f_{R}(r;P)$ that is three times differentiable wrt $r$: $\nabla_r f_{R}(r;P)$, $\nabla_{r^2} f_{R}(r;P)$, and $\nabla_{r^3} f_{R}(r;P)$ respectively; $f_{R}(r;P)$ and these derivatives are bounded as functions of $(r,P)$; $f_{R}(r;P)$ is bounded away from zero as a function of $(r,P)$;} To see this, note that $f_{R}(r;P)$ is a convex combination of $f_{X_1}(r)$ and $f_{X_2}(r)$, each bounded, with bounded derivatives, and bounded away from zero. • \textit{The conditional distribution of $U$ given $R$ has PDF $f_{U|R}(u|r;P)$ that is a bounded function of $(u,r,P)$, $f_{U|R}(0|x;P)$ is bounded away from zero over $P$, $f_{U|R}(u|r;P)$ is differentiable as function of $(u,r)$ and has bounded partial derivatives;} To see this, take $P = \alpha P_1 + (1-\alpha) P_2$ and note that, \[ f_{U|R}(u|r;P) = \frac{\alpha f_{X_1}(r) }{ f_{R}(r;P) } f_{Y_1|X_1}(\theta(P) + u |r) + \frac{(1-\alpha) f_{X_2}(r) }{ f_{R}(r;P) } f_{Y_2|X_2}(\theta(P) + u |r). \] The PDFs $f_{Y_k|X_k}(y_k | x_k)$ are bounded functions of $(x_k,y_k)$. The weights $\omega_1(r;P) \doteq \alpha \frac{f_{X_1}(r) }{ f_{R}(r;P) }$ and $\omega_2(r;P) \doteq (1-\alpha) \frac{f_{X_2}(r) }{ f_{R}(r;P) }$ are bounded functions of $(r,P)$ because they are positive and sum to $1$. The PDFs $f_{Y_k|X_k}(y_k | x_k)$ are differentiable and so are the weights. The partial derivatives of $f_{Y_k|X_k}(y_k | x_k)$ are bounded. The derivatives of the weights wrt $r$ also are bounded because the derivatives of the PDFs $f_{X_k}(r)$ are bounded plus the fact that $f_{R}(r;P)$ is bounded away from zero over $(r,P)$. Finally, $f_{Y_k|X_k}(y_k | x)$ is bounded away from zero over $y_k$. • \textit{The conditional distribution of $U$ given $R$ has CDF $F_{U|R}(u|r;P)$ that is three times partially differentiable wrt $r$ and has partial derivatives $\nabla_r F_{U|R}(0|r;P)$, $\nabla_{r^2} F_{U|R}(0|r;P)$, and $\nabla_{r^3} F_{U|R}(0|r;P)$ that are bounded functions of $(r,P)$;} Again, \[ F_{U|R}(0|r;P) = \omega_1(r;P) F_{Y_1|X_1}(\theta(P) |r) + \omega_2(r;P) F_{Y_2|X_2}(\theta(P) |r). \] We have that and $F_{Y_k|X_k}(y_k |x_k)$ are three times partially differentiable wrt $x_k$. The weights $\omega_k(r;P)$ are three times differentiable wrt $r$ because the PDFs $f_{X_k}(x_k)$ are three times differentiable. The first three partial derivatives of $F_{Y_k|X_k}(y_k |x_k)$ wrt $x_k$ are bounded functions of $(x_k,y_k)$. The first three derivatives of $\omega_k(r;P)$ wrt $r$ are bounded functions of $(r,P)$ because the first three derivatives of $f_{X_k}(x_k)$ wrt $x_k$ are bounded and $f_{R}(r;P)$ is bounded away from zero over $(r,P)$. \end{enumerate} \begin{enumerate} • \underline{Assumption (ref) - (ref):} asymptotic expansion. \end{enumerate} Define $Z_m^b = \sqrt{mh}\left(\hat \theta^b - \theta(P)\right) $ and $Z_m = \sqrt{mh}\left(\hat \theta - \theta(P)\right) = \sqrt{mh}\left(\hat \theta^b - \theta(P) - h^2 \widehat{B} \right)$. We want to study the asymptotic behavior of $Z_m^b$ and $Z_m$, where $Z_m^b$ is the value that minimizes the objective function $L_m(z)$, i.e., \begin{align*} Z_m^b &= \arg\min_z \underbrace{\sum_{i=1}^m \rho_{\chi}\left(S_i - \theta(P) - \frac{z}{\sqrt{mh}}\right)K\left(\frac{R_i-x}{h}\right)}_{\doteq L_m(z)} \\ &= \arg\min_z L_m(z) - L_m(0), \end{align*} since $L_m(0)$ is not a function of $z$ and hence does not affect the argmin. Let $Q_m(z) = L_m(z) - L_m(0)$ and $U_i = S_i - \theta(P)$. We have, \[ Q_{m}(z) = \sum_{i=1}^{m}\left( \rho_{\chi}\left(U_{i} -\frac{1}{\sqrt{mh}}z\right) - \rho_{\chi}(U_{i})\right)K \left( \frac{R_i-x }{ h } \right) ,\] which is minimized by $Z_m^b = \sqrt{mh}\left(\hat \theta^b - \theta(P)\right)$. Notice that since $\rho_{\chi}(\cdot)$s are convex functions of z, so is $Q_m(z)$, which is a sum of convex functions. The derivation of the asymptotic linear representation is done in two parts. First, we approximate $Q_m(z)$ by a quadratic function $Q_m^*(z)$ whose minimizing value $z=\eta_m$ has an asymptotic linear representation plus a bias term. Second, we show that $Z_m^b = \sqrt{m h}(\widehat{\theta}^b - \theta(P) )$ converges to $\eta_m$ in probability and therefore they share the same asymptotic behavior. The bias-corrected version of $Z_m^b$ is $Z_m$, and $Z_m$ has an asymptotic linear representation with an influence function that satisfies Assumption (ref). \underline{Part I: approximating the objective function} Let $D_i = -\chi \mathbb{I}(U_i \geq 0) + (1-\chi)\mathbb{I}(U_i <0) = \mathbb{I}(U_i<0) - \chi $ and \[V_{i}(z) \doteq \rho_{\chi}\left(U_i-\frac{1}{\sqrt{mh}}z\right) -\rho_{\chi}(U_i)- \frac{1}{\sqrt{mh}}zD_i .\] We can rewrite $Q_m(z)$ in terms of $D_i$ and $V_i(z)$ by adding and subtracting the conditional expectation of $Q_m(z)$ as follows: \begin{align} Q_m(z) &= \mathbb{E}_P[Q_m(z)| \underline{\boldsymbol{\mathrm{R}}}_m ] \\ &+\frac{1}{\sqrt{mh}}\sum_{i=1}^m z\left(D_i-\mathbb{E}_P[D_i|R_i]\right) K \left( \frac{R_i-x }{ h } \right) \\ &+ \sum_{i=1}^m (V_{i}(z) - \mathbb{E}_P[V_{i}(z)|R_i])K \left( \frac{R_i-x }{ h } \right) , \end{align} where $\underline{\boldsymbol{\mathrm{R}}}_m$ is the vector $(R_1, \ldots, R_m)$. In what follows, we show that \begin{align*} &(ref) = \frac{1}{2} f_{U|R}(0| x ; P)f_R(x;P) z^2 \\ & + \sqrt{mh} h^2 z \kappa_{2,1} \left[ \nabla_{r} F_{U|R}(0|x;P) \nabla_r f_R(x ;P) +\frac{1}{2} \nabla_{r^2} F_{U|R}(0|x;P) f_R(x;P) \right] + o_{\mathcal P}(1). \\ &(ref) = o_{\mathcal P}(1), \end{align*} where $\kappa_{s,t} = \int u^s K^t(u) ~du$. Regarding $\eqref{eq:expanQ_1}$, define \[ M(t|r;P) \doteq \mathbb{E}_P\left[\rho_{\chi}(S - \theta(P) +t) |R=r \right] = \mathbb{E}_P\left[\rho_{\chi}(U +t) |R=r \right]. \] Notice that although the check function is not differentiable, the $M$ function is differentiable. \begin{align*} &\nabla_t M(t|r;P) = \chi(1-F_{U|R}(-t|r;P)) -(1-\chi) F_{U|R}(-t|r;P) = \chi - F_{U|R}(-t|r;P), \\ &\nabla_{t^2} M(t|r;P) = - \nabla_{t} \{ F_{U|R}(-t|r;P) \} = f_{U|R}(-t|r;P), \\ &\nabla_{t^3} M(t|r;P) = - \nabla_{u} f_{U|R}(-t|r;P), \\ &\nabla_{t r} M(t|r;P) = - \nabla_{r} F_{U|R}(-t|r;P), \\ &\nabla_{t r^2} M(t | r ; P ) =-\nabla_{r^2} F_{U|R}(-t|r;P), \\ &\nabla_{t r^3} M(t | r ; P ) =-\nabla_{r^3} F_{U|R}(-t|r;P), \end{align*} where we use the Leibniz rule, the existence of the derivatives $\nabla_u f_{U|R}(u|r;P)$, $\nabla_{r} F_{U|R}(-t|r;P)$, $\nabla_{r^2} F_{U|R}(-t|r;P)$, and $\nabla_{r^3} F_{U|R}(-t|r;P)$. We can write $\mathbb{E}[Q_m(z)|\underline{\boldsymbol{\mathrm{R}}}_m]$ in terms of $M$ as follows, \begin{align} &\mathbb{E}_P[Q_m(z)|\underline{\boldsymbol{\mathrm{R}}}_m] \nonumber\\* &= \sum_{i=1}^{m}\mathbb{E}_P \bigg[ \rho_{\chi}\left(U_i -\frac{1}{\sqrt{mh}}z\Big| R_i\right) - \rho_{\chi}\left(U_i \Big| R_i\right) \bigg] K \left( \frac{R_i-x }{ h } \right) \nonumber \\ &= \sum_{i=1}^m \left[ M\left( - \frac{z}{\sqrt{mh}}\Big|R_i ; P\right) - M\left( 0\big| R_i ; P\right) \right]K\left(\frac{R_i-x}{h}\right). \nonumber \end{align} Taylor expand $M$ as a function of $t$ around $0$, \begin{align} &\mathbb{E}_P[Q_m(z)|\underline{\boldsymbol{\mathrm{R}}}_m] \nonumber \\* &= \sum_{i=1}^m \left[ \nabla_{t} M\left( 0 \big| R_i ;P \right) \frac{-z}{\sqrt{mh}} + \frac{1}{2} \nabla_{t^2} M\left( 0 \big| R_i ; P \right) \frac{z^2}{mh} \right. \nonumber \\ & \left. - \frac{1}{6}\nabla_{t^3} M( q^* | R_i ; P) \frac{z^3}{(mh)^{3/2}} \right] K\left(\frac{R_i-x}{h}\right) \nonumber \\ & =\frac{-z}{\sqrt{mh}}\sum_{i=1}^m \nabla_{t} M\left(0 \big| R_i ; P \right) K\left(\frac{R_i-x}{h}\right) \\ & +\frac{1}{2}\frac{z^2}{mh}\sum_{i=1}^m f_{U|R}\left( 0 | R_i ; P \right) K\left(\frac{R_i-x}{h}\right) \\ & + \frac{z^3}{6\sqrt{mh}}\frac{1}{mh}\sum_{i=1}^m \nabla_u f_{U|R}\left( -q^*|R_i ; P \right) K\left(\frac{R_i-x}{h}\right) , \end{align} where $q^*$ is a point between $0$ and $- z/\sqrt{mh}$. The goal is to show that \begin{align*} & (ref)= \sqrt{mh} h^2 z \kappa_{2,1} \left[ \nabla_{r} F_{U|R}(0|x;P) \nabla_r f_R(x ;P) +\frac{1}{2} \nabla_{r^2} F_{U|R}(0|x;P) f_R(x;P) \right] + o_{\mathcal{P}}(1), \\ & (ref) = \frac{1}{2} f_{U|R}(0| x ; P)f_R(x;P) z^2 +o_{\mathcal P}(1), \\ & (ref) =o_{\mathcal{P}}(1). \end{align*} Expectation of ((ref)). We Taylor expand $\nabla_{t} M\left(0 \big| R_i ;P\right)$ as a function of $R_i$ around $R_i=x$ and use the fact that $\nabla_{t} M\left(0 \big| x ;P\right) = \chi - \mathbb{P}_P (S_i - \theta(P) \leq 0 | R_i = x) = 0 $: \begin{align*} &\mathbb{E}_P \left[\frac{-z}{\sqrt{mh}}\sum_{i=1}^m \nabla_{t} M\left(0 \big| R_i ;P\right) K\left(\frac{R_i-x}{h}\right)\right] \\ =& \mathbb{E}_P \left[\frac{-z}{\sqrt{mh}}\sum_{i=1}^m \underbrace{ \nabla_{t} M\left(0 \big| x ;P \right)}_{=0} K\left(\frac{R_i-x}{h}\right)\right] \\ &+ \mathbb{E}_P \left[\frac{-z}{\sqrt{mh}}\sum_{i=1}^m \underbrace{ \nabla_{tr} M\left(0 \big| x ;P \right)}_{= - \nabla_{r} F_{U|R}(0|x;P) } \left(R_i-x\right) K\left(\frac{R_i-x}{h}\right)\right] \\ &+ \mathbb{E}_P \left[\frac{-z}{\sqrt{mh}}\sum_{i=1}^m \frac{1}{2} \underbrace{ \nabla_{tr^2} M\left(0 \big| x ;P \right)}_{= - \nabla_{r^2} F_{U|R}(0|x;P)}\left(R_i-x\right)^2 K\left(\frac{R_i-x}{h}\right)\right] \\ &+ \mathbb{E}_P \left[\frac{-z}{\sqrt{mh}}\sum_{i=1}^m \frac{1}{6} \underbrace{ \nabla_{tr^3} M\left(0 \big| x^* ;P \right)}_{= - \nabla_{r^3} F_{U|R}(0|x^*;P)}\left(R_i-x\right)^3 K\left(\frac{R_i-x}{h}\right)\right] \\ =& z \sqrt{mh}\int \left[\nabla_{r} F_{U|R}(0|x;P) uh \right] K(u) f_R(x+uh;P)du \\ &+ z \sqrt{mh}\int \left[\nabla_{r^2} F_{U|R}(0|x;P) u^2 h^2 \right] K(u) f_R(x+uh;P)du \\ &+ z \sqrt{mh}\int \left[ \nabla_{r^3} F_{U|R}(0|x^*;P) u^3 h^3 \right] K(u) f_R(x+uh;P)du \\ =& z \nabla_{r} F_{U|R}(0|x;P) \sqrt{mh} h \int u K(u) [f_R(x;P) + \nabla_r f_R(x;P) uh + \frac{1}{2}\nabla_{r^2} f_R(x^{**};P) u^2 h^2] du \\ &+ \frac{z}{2} \nabla_{r^2} F_{U|R}(0|x;P) \sqrt{mh} h^2 \int u^2 K(u) [f_R(x;P) + \nabla_r f_R(x^{***};P) uh] du \\ &+ \frac{z}{6} \sqrt{mh} h^3 \int \nabla_{r^3} F_{U|R}(0|x^*;P) u^3 K(u) f_R(x+uh;P)du, \end{align*} where $x^*$ is a point between $R_i$ and $x$, $x^{**}$ and $x^{***}$ are points between $x+uh $ and $x$. The last equation above equals to \begin{align*} =& z \nabla_{r} F_{U|R}(0|x;P) f_R(x;P) \sqrt{mh} h \underbrace{\int u K(u) du}_{=0} \\ & + z \nabla_{r} F_{U|R}(0|x;P) \nabla_r f_R(x ;P) \sqrt{mh} h^2 \underbrace{\int u^2 K(u) du}_{= \kappa_{2,1}} \\ & + \frac{z}{2} \nabla_{r} F_{U|R}(0|x;P) \sqrt{mh} h^3 \int u^3 K(u) \nabla_{r^2} f_R(x^{**};P) du \\ &+ \frac{z}{2} \nabla_{r^2} F_{U|R}(0|x;P) f_R(x;P) \sqrt{mh} h^2 \underbrace{\int u^2 K(u) du}_{= \kappa_{2,1}} \\ &+ \frac{z}{2} \nabla_{r^2} F_{U|R}(0|x;P) \sqrt{mh} h^3 \int u^3 K(u) \nabla_r f_R(x^{***};P) du \\ &+ \frac{z}{6} \sqrt{mh} h^3 \int \left[ \nabla_{r^3} F_{U|R}(0|x^*;P) u^3 \right] K(u) f_R(x+uh;P)du \\ =& 0 \\ & + \sqrt{mh} h^2 z \kappa_{2,1} \nabla_{r} F_{U|R}(0|x;P) \nabla_r f_R(x ;P) \\ & + O_{\mathcal{P}} \left( \sqrt{mh} h^3 \right) \\ &+ \sqrt{mh} h^2 \frac{z \kappa_{2,1} }{2} \nabla_{r^2} F_{U|R}(0|x;P) f_R(x;P) \\ &+ O_{\mathcal{P}} \left( \sqrt{mh} h^3 \right) \\ &+ O_{\mathcal{P}} \left( \sqrt{mh} h^3 \right) \\ = & \sqrt{mh} h^2 z \kappa_{2,1} \left[ \nabla_{r} F_{U|R}(0|x;P) \nabla_r f_R(x ;P) +\frac{1}{2} \nabla_{r^2} F_{U|R}(0|x;P) f_R(x;P) \right] + o_{\mathcal{P}}(1), \end{align*} where we use the kernel properties, the facts that $\nabla_{r} F_{U|R}(0|x;P)$ and $\nabla_{r^2} F_{U|R}(0|x;P)$ are bounded functions of $P$, $\nabla_{r^3} F_{U|R}(0|r;P)$ is a bounded function of $(r,P)$, $f_R(r ; P )$ is a bounded function of $(r,P)$, $\nabla_r f_R(r ; P )$ and $\nabla_{r^2} f_R(r ; P )$ are bounded functions of $(r,P)$, and $\sqrt{mh}h^3 \to 0$. Variance of ((ref)). \begin{align*} &\mathbb{V}_P \left[\frac{-z}{\sqrt{mh}}\sum_{i=1}^m \nabla_{t} M\left(0 \big| R_i ;P\right) K\left(\frac{R_i-x}{h}\right)\right] \\ =& \frac{z^2}{h}\mathbb{V}_P \left[ \nabla_{t} M\left(0 \big| R_i ; P \right) K\left(\frac{R_i-x}{h}\right)\right] \\ \leq & \frac{z^2}{h} \mathbb{E}_P \left[ \left\{ \nabla_{t} M \left( 0 \big| R_i ; P \right) \right\}^2 K^2\left(\frac{R_i-x}{h}\right)\right] \\ = &z^2 h^2\int \left\{ \nabla_{r} F_{U|R}(0|x^*;P) \right\}^2 u^2f_{R}(x+uh;P)K^2\left(u\right)du = o_{\mathcal{P}}(1). \end{align*} Therefore, \begin{align*} & (ref)= \sqrt{mh} h^2 z \kappa_{2,1} \left[ \nabla_{r} F_{U|R}(0|x;P) \nabla_r f_R(x ;P) +\frac{1}{2} \nabla_{r^2} F_{U|R}(0|x;P) f_R(x;P) \right] + o_{\mathcal{P}}(1). \end{align*} It remains to show that the probability limits of (ref) and (ref) are zero. Expectation of (ref): \begin{align*} &\mathbb{E}_P \left[\frac{1}{2}\frac{z^2}{mh}\sum_{i=1}^m f_{U|R}\left( 0 |R_i ;P \right) K\left(\frac{R_i-x}{h}\right)\right] \\ =&\mathbb{E}_P\left[\frac{1}{2}\frac{z^2}{h}\left\{f_{U|R}\left(0|x;P \right) + \nabla_r f_{U|R}\left(0|x^* \right) (R_i-x) \right\} K\left(\frac{R_i-x}{h}\right)\right] \\ =&\frac{z^2}{2} \int \left\{f_{U|R}\left(0|x;P \right) + \nabla_r f_{U|R}\left(0|x^* \right) uh \right\} K\left(u\right) f_R(x+uh;P)du \\ =&f_{U|R}\left(0|x;P \right) \frac{z^2}{2} \int K\left(u\right) \left\{f_R(x;P) + \nabla_r f_R(x^{**} ; P ) uh \right\} du \\ &+ \frac{z^2}{2} h \int \nabla_r f_{U|R}\left(0|x^*;P \right) u K\left(u\right) f_R(x+uh;P)du \\ =&\frac{z^2}{2} f_{U|R}\left(0|x;P \right) f_R(x;P) + o_{\mathcal{P}}(1), \end{align*} where we use that $f_R(r ; P )$, $\nabla_r f_R(r ; P )$, $f_{U|R}\left(0|x;P \right)$, $\nabla_r f_{U|R}\left(0|r ;P\right)$ are bounded over $(r,P)$. Variance of (ref): \begin{align*} &\mathbb{V}_P\left[\frac{1}{2}\frac{z^2}{mh}\sum_{i=1}^m f_{U|R}\left( 0 | R_i ;P \right) K\left(\frac{R_i-x}{h}\right)\right] \\ & = \frac{z^4}{4}\frac{1}{mh^2} \mathbb{V}_P \left[f_{U|R}\left( 0 |R_i;P \right) K\left(\frac{R_i-x}{h}\right) \right] \\ & \leq \frac{z^4}{4}\frac{1}{mh^2}\mathbb{E}_P\left[f_{U|R}^2\left(0 |R_i ; P \right) K^2\left(\frac{R_i-x}{h}\right) \right] \\ &= \frac{z^4}{4}\frac{1}{mh}\int f_{U|R}^2\left( 0| x+uh ; P \right) K^2\left(u\right) f_{R}(x+uh;P)du = o_{\mathcal{P}}(1), \end{align*} because $f_{U|R}\left( 0 | r ; P \right)$ and $f_{R}(r;P)$ are bounded functions of $(r,P)$ and $mh \to \infty$. Therefore, we have that (ref) $=\frac{1}{2} f_{U|R}\left(0|x;P\right)f_R(x;P)z^2$ $+ o_{\mathcal{P}}(1)$. Moreover, $\eqref{eq:3rd} = o_{\mathcal{P}}(1)$ because $mh \to \infty$, $\frac{1}{mh}\sum_{i=1}^m K\left(\frac{R_i-x}{h}\right) = O_{\mathcal P}(1)$, and $\nabla_u f_{U|R}(u|r;P)$ is a bounded function of $(u,r,P)$. We are done in showing the probability limit of $\mathbb{E}_P[Q_m(z)| \underline{\boldsymbol{\mathrm{R}}}_m ]$, \begin{align*} & (ref) = \frac{1}{2} f_{U|R}\left(0|x;P\right)f_R(x;P)z^2 \\ & + \sqrt{mh} h^2 z \kappa_{2,1} \left[ \nabla_{r} F_{U|R}(0|x;P) \nabla_r f_R(x ;P) +\frac{1}{2} \nabla_{r^2} F_{U|R}(0|x;P) f_R(x;P) \right] + o_{\mathcal{P}}(1). \end{align*} It remains to show that $\eqref{eq:expanQ_3} = o_{\mathcal{P}}(1)$. We show that the expectation of (ref) is zero and its variance converges to zero. The expectation of (ref) is zero because it equals the expectation of \begin{align*} & \mathbb{E}_P \left[ \sum_{i=1}^n (V_{i}(z) - \mathbb{E}_P [V_{i}(z) | R_i ] ) K\left(\frac{R_i-x}{h}\right) \Big| \boldsymbol{\mathrm{R}}_m \right] =0. \end{align*} Variance of (ref): \begin{align*} &\mathbb{V}_P \left[ \sum_{i=1}^m (V_{i}(z) - \mathbb{E}_P[V_{i}(z)|R_i]) K\left(\frac{R_i-x}{h}\right) \right] \\ = & \mathbb{V}_P \left[\mathbb{E}_P \left[ \sum_{i=1}^m (V_{i}(z) - \mathbb{E}_P[V_{i}(z)|R_i]) K\left(\frac{R_i-x}{h} \right) \Big| \boldsymbol{\mathrm{R}}_m \right] \right] \\ & + \mathbb{E}_P \left[\mathbb{V}_P \left[ \sum_{i=1}^m (V_{i}(z) - \mathbb{E}_P [V_{i}(z) | R_i ] ) K\left(\frac{R_i-x}{h}\right)\Big| \boldsymbol{\mathrm{R}}_m \right] \right] \\ & \leq \sum_{i=1}^m \mathbb{E}_P \left[K^2\left(\frac{R_i-x}{h}\right)\mathbb{E}_P \left[ V_{i}(z)^2\Big|R_i\right] \right]\\ &= \sum_{i=1}^m \mathbb{E}_P \left[K^2\left(\frac{R_i-x}{h}\right)V_i(z)^2\right] \\ & \leq 4z^2 \int K^2(u)f_{R}(x+uh ;P) \left\{ F_{U|R}\left(\left|\frac{z}{\sqrt{mh}}\right| \bigg| x+uh;P \right) \right. \\ & \left. -F_{U|R}\left(-\left|\frac{z}{\sqrt{mh}}\right| \bigg| x+uh ;P\right)\right\} du \\ & = \frac{8 z^3}{\sqrt{mh}} \int K^2(u)f_{R}(x+uh ;P) \nabla_u F_{U|R}\left( u^* \bigg| x+uh; P \right) du \\ &=o_{\mathcal{P}}(1), \end{align*} where we use that $f_R$ and $\nabla_u F_{U|R} $ are bounded over $(u,r,P)$ and \begin{align*} |V_i(z)| &= \left| \left(\rho_{\chi}\left(U_i-z/\sqrt{mh}\right) - \rho_{\chi}(U_i) - D_iz/\sqrt{mh}\right)\right|\\ &\leq 2\left|z/\sqrt{mh}\right| \mathbb{I}\left(\left|U_i\right| \leq \left|\frac{z}{\sqrt{mh}}\right|\right). \end{align*} Consider (ref)--(ref) and the probability limits of (ref) and (ref) that we found. Define the following objects, \begin{align*} B_0(P) = & \kappa_{2,1} \left[ \nabla_{r} F_{U|R}(0|x;P) \nabla_r f_R(x ;P) +\frac{1}{2} \nabla_{r^2} F_{U|R}(0|x;P) f_R(x;P) \right], \\ Q_m^*(z) = & z^2 \frac{1}{2}f_{U|R}(0 | x ; P)f_R(x;P) \\ & + z \left[ \frac{1}{\sqrt{mh}}\sum_{i=1}^m \left(D_i-\mathbb{E}_P[D_i|R_i]\right) K \left( \frac{ R_i-x }{ h } \right) + \sqrt{mh} h^2 B_0(P) \right], \\ r_m(z) = & Q_m(z) - Q_m^*(z), \end{align*} so that $Q_m(z) = Q_m^*(z) + r_m(z)$ and $r_m(z) = o_{\mathcal{P}}(1)$ for fixed $z$. Rewrite $Q_m^*(z)$ as follows. \begin{align} Q_m^*(z) = & z^2 \frac{1}{2}f_{U|R}(0 | x ; P)f_R(x;P) \nonumber \\ & + z \underbrace{ \left[ \frac{1}{\sqrt{mh}}\sum_{i=1}^m \left(D_i-\mathbb{E}_P[D_i|R_i]\right) K \left( \frac{ R_i-x }{ h } \right) + \sqrt{mh} h^2 B_0(P) \right] }_{\doteq M_m} \nonumber \\ & = \frac{1}{2}f_{U|R}(0|x; P)f_R(x;P)z^2 +z M_m \nonumber \\ &= \frac{1}{2}f_{U|R}(0|x; P)f_R(x;P) \left(z + \underbrace{\frac{1}{f_{U|R}(0|x; P)f_R(x;P)} M_m}_{\doteq -\eta_m } \right)^2 \nonumber \\ & - \frac{1}{2f_{U|R}(0 | x ; P)f_{R}(x;P)}M_m^2 \nonumber \\ & = \frac{1}{2}f_{U|R}(0 | x ; P)f_R(x;P) \left(z -\eta_m \right)^2 - \frac{1}{2}f_{U|R}(0 | x ; P)f_{R}(x;P) \eta_m^2 , \end{align} which is minimized at \begin{align*} \eta_m & = -\frac{1}{f_{U|R}(0 | x ; P)f_R(x;P)} M_m \\ &= -\frac{1}{f_{U|R}(0 | x ; P) f_R(x;P)} \frac{1}{\sqrt{mh}} \sum_{i=1}^m \left(D_i-\mathbb{E}_P[D_i|R_i]\right) K \left( \frac{ R_i-x }{ h } \right) \\ &+ \sqrt{mh} h^2 \underbrace{\frac{-B_0(P) }{f_{U|R}(0 | x ; P) f_R(x;P)}}_{\doteq B(P)} \\ & = \frac{1}{\sqrt{mh}} \sum_{i=1}^m \left(\frac{-1}{f_{U|R}(0 | x ; P) f_R(x;P)}\right) \left(D_i-\mathbb{E}_P[D_i|R_i]\right) K \left( \frac{ R_i-x }{ h } \right) \\ &+ \sqrt{mh} h^2 B(P). \end{align*} The bias term $B(P)$ is \begin{align} B(P) = \frac{-\kappa_{2,1} }{f_{U|R}(0 | x ; P) f_R(x;P)} & \bigg[ \nabla_{r} F_{U|R}(0|x;P) \nabla_r f_R(x ;P) \notag \\ & +\frac{1}{2} \nabla_{r^2} F_{U|R}(0|x;P) f_R(x;P) \bigg] \end{align} and is consistently estimated by $\widehat{B}$, \begin{align*} & \widehat{B} = B_{m,n}(\boldsymbol{\mathrm{V}}_1 , \ldots, \boldsymbol{\mathrm{V}}_m) \\ & \doteq \frac{-\kappa_{2,1} }{\widehat{f_{U|R}}(0 | x ; P) \widehat{f_R}(x;P)} \left[ \widehat{\nabla_{r} F_{U|R}}(0|x;P) \widehat{\nabla_r f_R}(x ;P) +\frac{1}{2} \widehat{\nabla_{r^2} F_{U|R}}(0|x;P) \widehat{f_R}(x;P) \right], \end{align*} that is, by replacing $ f_R(x;P)$, $\nabla_r f_R(x ;P)$, $f_{U|R}(0 | x ; P)$, $\nabla_{r} F_{U|R}(0|x;P)$, and $\nabla_{r^2} F_{U|R}(0|x;P)$, in $B(P)$ by consistent nonparametric estimators readily available from the literature. The tuning parameters for these additional estimators and appropriate moment conditions may be set such that $\widehat{B} - B(P) = o_{\mathcal{P}}(1)$, as required by Assumption (ref). We have already shown above that $r_m(z) = o_{\mathcal{P}}(1)$ for fixed $z$. Now, we show that the convergence is also uniform over $z$ in a compact set $\mathcal{K} \subset \mathbb R$. To this end, consider \\ $\Lambda_m(z) \doteq Q_m(z) - z \frac{1}{\sqrt{mh}}\sum_{i=1}^m \left(D_i-\mathbb{E}_P[D_i|R_i]\right) K \left( \frac{ R_i-x }{ h } \right) - z \sqrt{mh}h^2 B_0(P)$, \\ $\Lambda(z) \doteq z^2 \frac{1}{2}f_{U|R}(0| r; P)f_R(x;P)$, and note that $\Lambda_m(z)$ is a convex function of $z$. Note also that $r_m(z) = Q_m(z) - Q_m^*(z) = \Lambda_m(z) - \Lambda(z)$. By the convexity lemma (\citeSM{pollard1991}, page 187), $\sup_{z \in \mathcal{K} } | \Lambda_m(z) - \Lambda(z) | = o_{\mathcal{P}}(1)$. We have that, for any compact subset $\mathcal{K} \subset \mathbb R$, $\sup_{z \in \mathcal{K}}\left|r_m(z)\right| = o_{\mathcal P}(1)$. \underline{Part II: $Z_m^b - \eta_m =o_{\mathcal P}(1)$.} We want to show that for each $\epsilon >0$, \[\inf_{P \in \mathcal{P} } \mathbb{P}_P \left(\left|Z_m^b - \eta_m\right| \leq \epsilon\right) \to 1.\] Consider the closed interval $B(m)$ with center $\eta_m$ and radius $\epsilon$. Since $\eta_m$ converges in distribution, it is stochastically bounded. The compact set $\mathcal{K}$ can be chosen to contain $B(m)$ with probability arbitrarily close to one, thereby implying $\sup_{z \in B(m) }\left| r_m(z) \right| = o_{\mathcal P}(1).$ For a value outside of the interval $B(m)$, suppose $z = \eta_m + \delta$ with $\delta > \epsilon$. For the boundary point $z^* = \eta_m+ \epsilon$, the convexity of $Q_m$ and ((ref)) imply \begin{align*} \frac{\epsilon}{\delta} Q_m(z) + &\left(1-\frac{\epsilon}{\delta}\right) Q_m(\eta_m) \geq Q_m\left(\frac{\epsilon}{\delta} z +\left(1-\frac{\epsilon}{\delta}\right) \eta_m\right) = Q_m(z^*) \\ &\geq \frac{f_{U|R}(0 | x ; P)f_R(x;P)}{2}(z^* - \eta_m)^2 - \frac{f_{U|R}(0| r ; P)f_R(x;P)}{2}\eta_m^2 - \sup_{z \in B(m)}\left|r_m(z)\right| \\ &\geq \frac{f_{U|R}(0 | x ; P)f_R(x;P)}{2}\epsilon^2 + Q_m(\eta_m) - 2\sup_{z \in B(m)}\left|r_m(z)\right|. \\ Q_m(z) & \geq Q_m(\eta_m) + \left( \frac{\delta}{\epsilon } \right) \left[ \frac{f_{U|R}(0 | x ; P)f_R(x;P)}{2}\epsilon^2 - 2 \sup_{z \in B(m)}\left|r_m(z)\right| \right]. \end{align*} An analogous argument holds for $z = \eta_m - \delta$. Define the event $A_m$ as \begin{align*} A_m : 2 \sup_{z \in B(m)}\left| r_m(z) \right| < \frac{f_{U|R}(0 | x ; P)f_R(x;P)}{4}\epsilon^2. \end{align*} We have that $\inf_{P \in \mathcal{P}} \mathbb{P}_P [A_m] \to 1$ because \\ $\sup_{P \in \mathcal{P}} \mathbb{P}_P \left[ \sup_{z \in B(m)}\left|r_m(z)\right| > f_{U|R}(0 | x ; P)f_R(x;P) \epsilon^2 / 4 \right] \to 0$. Conditional on $A_m$, \begin{align} Q_m(z) \geq Q_m(\eta_m) + \frac{f_{U|R}(0 | x ; P)f_R(x;P)}{4} \epsilon^2 \text{, for any $z : |z - \eta_m| > \epsilon$} \end{align} happens with probability one. The event in (ref) implies that $| Z_m^b - \eta_m | \leq \epsilon$ because $Z_m^b$ minimizes $Q_m(z)$ and thus $Q_m(Z_m^b) \leq Q_m(\eta_m)$. Finally, \begin{align*} &\mathbb{P}_P \left[ | Z_m^b - \eta_m | \leq \epsilon \right] \\ &\geq \mathbb{P}_P \left[ \inf\limits_{z : |z - \eta_m| > \epsilon} Q_m(z) \geq Q_m(\eta_m) + \frac{f_{U|R}(0 | x ; P)f_R(x;P)}{4} \epsilon^2 \right] \geq \mathbb{P}_P [ A_m ] \text{, so that } \\ & \inf_{P \in \mathcal{P}} \mathbb{P}_P \left[ | Z_m^b - \eta_m | \leq \epsilon \right] \geq \inf_{P \in \mathcal{P}} \mathbb{P}_P [ A_m ] \to 1. \end{align*} Therefore, we have the asymptotic linear representation of the estimator as follows: \begin{align*} & \sqrt{mh}\left(\hat \theta^b - \theta(P)\right) \\ & = \frac{1}{\sqrt{mh}} \sum_{i=1}^m \left(\frac{-1}{f_{U|R}(0 | x ; P) f_R(x;P)}\right) \left(D_i-\mathbb{E}_P[D_i|R_i]\right) K \left( \frac{ R_i-x }{ h } \right) \\ &+ \sqrt{mh} h^2 B(P) + o_{\mathcal{P}}(1). \end{align*} For $\widehat{B}$ such that $\widehat{B} = B(P) + o_{\mathcal{P}}(1)$, \begin{align*} & \sqrt{mh}\left(\hat \theta - \theta(P) \right) = \sqrt{mh}\left(\hat \theta^b - \theta(P) - h^2 \widehat{B} \right) = \sqrt{mh}\left(\hat \theta^b - \theta(P) - h^2 B(P) \right) + o_{\mathcal{P}}(1) \\ & = \frac{1}{\sqrt{m}}\sum_{i=1}^m \left( \frac{-1}{\sqrt{h} f_{U|R}(0 | x ; P)f_R(x;P)}\right) K\left(\frac{R_i-x}{h}\right)\left(D_i-\mathbb{E}_P[D_i |R_i ] \right) \\ & = \frac{1}{\sqrt{m}}\sum_{i=1}^m \underbrace{ \frac{-1}{f_{U|R}(0 | x ; P)f_R(x;P) \sqrt{h} } K\left(\frac{R_i-x}{h}\right) \left( \mathbb{I}\{S_i < \theta(P) \} - F_{S|R}(\theta(P)|R_i;P) \right) }_{\doteq \psi_n(\boldsymbol{\mathrm{V}}_i, P)} + o_{\mathcal P}(1) \\ & = \frac{1}{\sqrt{m}}\sum_{i=1}^m \psi_n(\boldsymbol{\mathrm{V}}_i, P) + o_{\mathcal P}(1), \end{align*} where we use that $D_i - \mathbb{E}_P[D_i | R_i] = \mathbb{I}\{S_i < \theta(P) \} - F_{S|R}(\theta(P) | R_i; P)$ and $\sqrt{mh}h^2 = o(1)$. \begin{enumerate} • \underline{Assumption (ref) - (ref):} zero mean of influence function. $\mathbb{E}_P[ \psi_{n}(\boldsymbol{\mathrm{V}}_{i}, P) ]=0 ~~\forall P$ by construction. • \underline{Assumption (ref) - (ref):} variance of influence function. Define \begin{align*} \xi^2(P) &\doteq \frac{1}{f_{U|R}^2(0 | x ; P) f_R(x;P)} \kappa_{0,2} \mathbb{V}_P [D_i | R_i =x] \\ & = \frac{1}{f_{U|R}^2(0 | x ; P) f_R(x;P)} \kappa_{0,2} F_{S|R}(\theta(P)| x ;P) (1-F_{S|R}(\theta(P)| x ;P)) \\ & = \frac{\chi (1-\chi) }{f_{U|R}^2(0 | x ; P) f_R(x;P)} \kappa_{0,2}. \end{align*} \begin{align*} &\mathbb{V}_P \left( \frac{-1}{f_{U|R}(0 | x ; P)f_R(x;P)\sqrt{h}}K\left(\frac{R_i-x}{h}\right)\left(D_i-\mathbb{E}_P[D_i|R_i]\right)\right) - \xi^2(P) \\ =&\frac{1}{f_{U|R}^2(0 | x ; P)f_{R}^2(x;P) } \int K^2(u) \bigg\{ \underbrace{ \mathbb{V}_P[D_i | R_i=x+uh] f_{R}(x+uh;P)}_{\doteq g(uh;P)} \\* & - \underbrace{ \mathbb{V}_P [D_i|R_i=x] f_R(x;P) }_{\doteq g(x;P)} \bigg\} du \\ =&\frac{h}{f_{U|R}^2(0 | x ; P)f_{R}^2(x;P) } \int K^2(u) \nabla_r g(x^*_{uh};P) u du = o_{\mathcal{P}}(1). \end{align*} where $x^*_{uh}$ is a point between $x+uh$ and $x$, and $\nabla_r g(r;P)$ denotes the derivative of $g $ wrt $r$. The expression above is $o_{\mathcal{P}}(1)$ because $h\to0$, the derivative $\nabla_r g(r;P)$ is a bounded function of $(r,P)$, and $f_R(x;P)$ and $f_{U|R}(0| x; P)$ are bounded away from zero over $P$. The derivative $\nabla_x g(x;P)$ is bounded because $f_R(r;P)$, $\nabla_r f_R(r;P)$, \\ $\mathbb{V}_P[D_i|R_i=r] = F_{U|R}(0 | r ;P) (1-F_{U|R}( 0 | r ;P))$, and \\ $\nabla_r \left\{F_{U|R}(0 | r ;P) (1-F_{U|R}( 0 | r ;P)) \right\} $ are bounded functions of $(r,P)$. Therefore, \[ \sup\limits_{P \in \mathcal{P} } \left| \mathbb{V}_P [ \psi_n(\boldsymbol{\mathrm{V}}_{i}, {P} ) ] - \xi^2(P) \right| \to 0. \]\underline{Assumption (ref) - (ref):} $ \sup_P \mathbb{E} [\psi_n^2(\boldsymbol{\mathrm{Z}}_k ,P)] < \infty .$ For $\boldsymbol{\mathrm{Z}}_k=(X_k,Y_k) \sim P_k, ~ k=1,2$, define $D_{k} = \mathbb{I}(Y_k - \theta(P_k) < 0) -\chi$, $m_{D_k|X_k}(x_k)=\mathbb{E}[ D_k | X_k=x_k ]$, and $v_{D_k|X_k}(x_k)=\mathbb{V}[ D_k | X_k=x_k ]$. For $\boldsymbol{\mathrm{V}}=(R,S) \sim P \in \mathcal{P}$, $D = \mathbb{I}(S - \theta(P) < 0) -\chi$ and $m_{D | R}(r;P)=\mathbb{E}_P[ D | R=r ]$. We have, \begin{align*} \psi_n(\boldsymbol{\mathrm{Z}}_k,P) & = \frac{-1}{f_{U|R}(0 | x ; P)f_{R}(x;P)\sqrt{h}}K\left(\frac{X_k-x}{h}\right)\left(D_k-m_{D | R}(X_k ; P )\right). \\ \mathbb{E}\left[ \psi_n^2(\boldsymbol{\mathrm{Z}}_k,P) \right] &= \frac{1}{ f_{U|R}^2(0 | x ; P) f_{R}^{2}(x;P)} \\ & \int K^2(u) \bigg[ v_{P_k}(x+uh) + \left(m_{D_k|X_k}(x+uh) - m_{D|R}(x+uh;P) \right)^2 \bigg] \\ & f_{X_k}(x+uh) du \\ & = O_{\mathcal{P}}(1), \end{align*} because $v_{D_k|X_k}(r)=F_{Y_k|X_k}(\theta(P_k)| r ) (1-F_{Y_k|X_k}(\theta(P_k)| r)) $,\\ $m_{D_k|X_k}(r)=F_{D_k|X_k}(\theta(P_k)| r ) - \chi$, $m_{D | R}(r;P)= F_{S|R}(\theta(P) | r) - \chi$, and $f_{X_k}(r)$ are bounded over $(r,P)$; and $f_R(x;P)$ and $f_{U|R}(0| x; P)$ are bounded away from zero over $P$. • \underline{Assumption (ref) - (ref):} $(2+\zeta)$-th moment condition. We verify it in two steps. First, $\mathbb{V}_P\left( \psi_n(\boldsymbol{\mathrm{Z}}_i,P) \right) - \xi^2(P) = o_{\mathcal{P}}(1)$ and $\xi^2(P)$ is bounded away from zero, uniformly over $\mathcal{P}$. Thus, $\mathbb{V}_P^{-1}\left( \psi_n(\boldsymbol{\mathrm{Z}}_i,P) \right) = O_{\mathcal{P}}(1)$. Second, for any $\zeta>0$, call $\eta(r;P) = \mathbb{E}_P[ | D_i - \mathbb{E}_P [ D_i | R_i ] |^{2+\zeta} | R_i=r]$ and note that $|\eta(r;P)| \leq 2.$ \begin{align*} &n^{-\zeta/2} \mathbb{E}_P \left| \psi_n(\boldsymbol{\mathrm{V}}_i,P) \right| ^{2+\zeta} \\ & =\frac{(m/n)^{\zeta/2}}{ f_{U|R}^{2+\zeta}(0 | x ; P) f_{R}^{2+\zeta}(x;P) (mh)^{\zeta/2}} \int |K(u)|^{2 +\zeta} \eta(x+uh;P) f_{R}(x+uh;P) du \\ & = o_{\mathcal{P}}(1). \end{align*} Combining steps 1 and 2, \begin{align*} n^{-\zeta/2} \sup\limits_{P \in \mathcal{P} } \mathbb{E}_P \left| \frac{ \psi_n(\boldsymbol{\mathrm{V}}_i,P) }{ \sqrt{\mathbb{V}_P\left( \psi_n(\boldsymbol{\mathrm{V}}_i,P) \right)} } \right|^{2 + \zeta} = o(1). \end{align*} • \underline{Assumption (ref) - (ref):} $\xi^2\left( \frac{m}{n} P_1 + \frac{n-m}{n} P_2 \right) \to \xi^2 \left( \gamma P_1 + (1-\gamma) P_2 \right)$. Let $\overline{P}_n = \frac{m}{n} P_1 + \frac{n-m}{n} P_2$ and $\overline{P} = \gamma P_1 + (1-\gamma) P_2$. Consider the expression for $\xi^2\left( P \right)$ given above. It suffices to show that $f_{R}(x;\bar{P}_n) \to f_{R}(x; \bar P)$ and $f_{U|R}(0|x; \bar{P_n}) \to f_{U|R }(0|x; \bar P)$. The first is straightforward (see Section (ref)). For the second, let $U_k = Y_k - \theta(P_k) $ and note that \begin{align*} f_{U|R}(0|x ; \bar P_n) &= \frac{m}{n}\frac{f_{X_1}(x)}{f_R(x;\bar P_n)} f_{U_1|X_1}(0|x) + \frac{n-m }{n } \frac{f_{X_2}(x)}{f_R(x;\bar P_n)}f_{U_2|X_2}(0|x) \\ &\to \gamma \frac{f_{X_1}(x )}{f_R(x;\bar P)}f_{U_1 | X_1 }(0|x ) + (1-\gamma)\frac{f_{X_2}(x )}{f_R(x;\bar P)} f_{U_2 | X_2 }(0|x) \\ &= f_{U|R}(0|x;\bar P). \end{align*} • \underline{Assumption (ref) :} $n_1 /n - \lambda \overset{p}{\to} 0$. In this case, $n_1$ is deterministic and $(n_1 /n - \lambda) \to 0$. Assumption (ref) has already been verified above for any sequence $m_n$ such that $(m/n -\gamma) \to 0$ for arbitrary $\gamma \in (0,1)$. In particular it holds for $\gamma \in \{ \lambda, 1-\lambda\}.$ \end{enumerate} $\square$ \subsection{Proof of Proposition (ref) - Discontinuity of Conditional Mean} The goal of this proof is to use the assumptions listed in Proposition (ref) to verify Assumptions (ref) and (ref). It follows the general lines of the proof of Proposition (ref), so the reader may refer to Section (ref) for the redundant details that we omit here. It builds on arguments from the literature on asymptotic approximations of the NW estimator at a boundary point. See, for example, Theorem 3.2 by \citeSM{fan1996} with $p=\nu=0$. For a generalization of this proof to the local polynomial regression estimator (LPR), see Section (ref) below. Consider an \textit{iid} sample from $P\in\mathcal{P}$ with $m$ observations, $\boldsymbol{\mathrm{V}}_1=(R_1,S_1), \ldots, \boldsymbol{\mathrm{V}}_m=(R_m,S_m)$, where the minimum value in the support of $R$ is 0. The number $m$ grows with $n$ such that $m/n \to \gamma$, for some $\gamma \in (0,1)$. The parameter of interest is $\theta(P) = \mathbb{E}[S|R=0^+]$, and the NW estimator is \[ \widehat{\theta}^b = \theta_{m,n}^b(\boldsymbol{\mathrm{V}}_1 , \ldots, \boldsymbol{\mathrm{V}}_m) \doteq \frac{\sum\limits_{i=1}^{m} K \left( \frac{ R_{i} }{ h } \right) S_{i} } {\sum\limits_{i=1}^{m} K \left( \frac{ R_{i} }{ h } \right) }. \] This section studies the asymptotic behavior of $\widehat{\theta} = \widehat{\theta}^b - h \widehat{B} $, where the expression for $\widehat{B}$ is given below Equation (ref). The assumptions in Proposition (ref) imply the following facts: \begin{enumerate} • \textit{As $m\to\infty$, $h \to 0$, $m h \to \infty$, $\sqrt{m h} h =O(1),$ and $\sqrt{m h} h^2 =o(1)$}; • \textit{The distribution of $R$ has PDF $f_{R}(r;P)$ that is twice differentiable wrt $r$ denoted $\nabla_r f_{R}(r;P)$ and $\nabla_{r^2} f_{R}(r;P)$; $f_{R}(r;P)$ and the derivatives are bounded functions of $(r,P)$; $f_{R}(r;P)$ is bounded away from zero as a function of $(r,P)$;} • \textit{$m_{S|R}(r;P) = \mathbb{E}_P[S|R=r]$ is twice differentiable wrt $r$ denoted $\nabla_r m_{S|R}(r;P)$ and $\nabla_{r^2} m_{S|R}(r;P)$; $m_{S|R}$ and the derivatives are bounded functions of $(r,P)$;} • \textit{$v_{S|R}(r;P) = \mathbb{V}_P[S|R=r]$ has first derivative wrt $r$ denoted $\nabla_r v_{S|R}(r;P)$; $v_{S|R}$, $\nabla_r v_{S|R}$ are both bounded as functions of $(r,P)$; $v_{S|R}(0^+;P)$ is bounded away from zero as a function of $P$;} • \textit{$\eta(r;P) \doteq \mathbb{E}_P[|S - m_{S | R }(R;P) |^{2+\zeta} | R=r]$ is a bounded function of $(r,P)$.} \end{enumerate} We re-write $\sqrt{m h} \left( \ha\theta - \theta(P) \right)$ $=\sqrt{m h} \left( \ha\theta^b - h \widehat{B} - \theta(P) \right)$ to find the asymptotic linear representation. \begin{align} \sqrt{m h} \left( \ha\theta -\theta(P) \right) &= \left( \frac{1}{\sqrt{m h} } \sum_{i=1}^{m} K\left( \frac{ R_{i} }{h} \right) \left( S_{i} - m_{S | R }(R_i;P) \right) \right)\left(f_{R}(0^+;P)/2\right)^{-1} \\ &+ \left( \frac{1}{\sqrt{m h} } \sum_{i=1}^{m} K\left( \frac{ R_{i} }{h} \right) \left( S_{i} - m_{S | R }(R_i;P) \right) \right) \notag \\ & \left[ \left( \frac{1}{ m h } \sum_{i=1}^{m} K\left( \frac{ R_{i} }{h} \right) \right)^{-1} - \left(f_{R}(0^+;P)/2\right)^{-1} \right] \\ &+ \left( \frac{1}{\sqrt{m h} } \sum_{i=1}^{m} K\left( \frac{ R_{i} }{h} \right) \left( m_{S | R }(R_i;P) - m_{S | R }(0^+;P) \right) \right) \notag \\ & \left( \frac{1}{ m h } \sum_{i=1}^{m} K\left( \frac{ R_{i} }{h} \right) \right)^{-1} -\sqrt{mh} h \widehat{B}. \end{align} \begin{enumerate} • \underline{Assumption (ref) - (ref):} asymptotic expansion. Equation (ref) above gives the influence function $\psi_n$. \begin{align*} \frac{1}{\sqrt{m} } \sum_{i=1}^{m} \underbrace{ K\left( \frac{ R_{i} }{h} \right) \left( S_{i} - m_{S | R }(R_i;P) \right) h^{-1/2} \left(f_{R}(0^+;P)/2\right)^{-1} }_{\doteq \psi_{n}(\boldsymbol{\mathrm{V}}_{i}, P)} = \frac{1}{\sqrt{m} } \sum_{i=1}^{m} \psi_{n}(\boldsymbol{\mathrm{V}}_{i}, P). \end{align*} We need to show that Equations (ref) and (ref) converge in probability to zero uniformly over $\mathcal{P}$. \textit{Equation (ref):} is $o_{\mathcal{P}}(1)$. We show this in 3 steps. First, \begin{align*} \mathbb{V}_P\left( \frac{1}{\sqrt{m h} } \sum_{i=1}^{m} K\left( \frac{ R_{i} }{h} \right) \left( S_{i} - m_{S | R }(R_i;P) \right) \right) \\ = \int_0^{\infty} K^2(u) v_{S|R}(uh;P) f_{R}(uh;P) du = O_{\mathcal{P}}(1). \end{align*} The expected value of the expression inside the variance above is zero, so we have that \[ \frac{1}{\sqrt{m h} } \sum_{i=1}^{m} K\left( \frac{ R_{i} }{h} \right) \left( S_{i} - m_{S | R }(R_i;P) \right) =O_{\mathcal{P}}(1). \] Second, \begin{align*} &\mathbb{E}_P\left[ \frac{1}{ m h } \sum_{i=1}^{m} K\left( \frac{ R_{i} }{h} \right) \right] - f_{R}(0^+;P)/2 = h \int_0^{\infty} K(u) \nabla_r f_{R}(x^*_{uh};P) u du = o_{\mathcal{P}}(1). \\ &\mathbb{V}_P\left[ \frac{1}{ m h } \sum_{i=1}^{m} K\left( \frac{ R_{i} }{h} \right) \right] \leq \frac{1}{ m h } \int K^2(u) f_{R}(uh;P ) du = o_{\mathcal{P}}(1). \end{align*} Therefore, \begin{align} &\frac{1}{ m h } \sum_{i=1}^{m} K\left( \frac{ R_{i} }{h} \right) - f_{R}(0^+;P)/2 = o_{\mathcal{P}}(1), \\ &\left( \frac{1}{ m h } \sum_{i=1}^{m} K\left( \frac{ R_{i} }{h} \right) \right)^{-1} - \left(f_{R}(0^+;P)/2\right)^{-1} = o_{\mathcal{P}}(1). \end{align} Third, combining steps 1 and 2 gives \begin{align*} &\left( \frac{1}{\sqrt{m h} } \sum_{i=1}^{m} K\left( \frac{ R_{i} }{h} \right) \left( S_{i} - m_{S | R }(R_i;P) \right) \right) \\ & \left[ \left( \frac{1}{ m h } \sum_{i=1}^{m} K\left( \frac{ R_{i} }{h} \right) \right)^{-1} - \left(f_{R}(0^+;P)/2\right)^{-1} \right] & = O_{\mathcal{P}}(1) o_{\mathcal{P}}(1) = o_{\mathcal{P}}(1). \end{align*} \textit{Equation (ref):} is $o_{\mathcal{P}}(1)$. We derive the probability limit of (ref) + $\sqrt{mh}h \widehat{B}$ in 3 steps.\\ First, $\left( \frac{1}{ m h } \sum_{i=1}^{m} K\left( \frac{ R_{i} }{h} \right) \right)^{-1} = 2 f_R^{-1}(0^+;P) + o_{\mathcal{P}}(1) $. \\ Second, \begin{align*} & \mathbb{E}_P\left[ \frac{1}{\sqrt{m h} } \sum_{i=1}^{m} K\left( \frac{ R_{i} }{h} \right) \left( m_{S | R }(R_i;P) - m_{S | R }(0^+;P) \right) \right] \\ =& \sqrt{mh} \int_0^{\infty} K\left( u \right) [m_{S | R }(uh;P) - m_{S | R }(0^+;P)] f_{R}(uh;P) du \\ =& \sqrt{mh} \int_0^{\infty} K\left( u \right) [ \nabla_r m_{S | R }(0^+;P)uh + \nabla_{r^2} m_{S | R }(x^*_{uh};P) u^2 h^2 ] f_{R}(uh;P) du \\ =& \sqrt{mh} \int_0^{\infty} K\left( u \right) [ \nabla_r m_{S | R }(0^+;P)uh ] [ f_{R}(0^+;P) + \nabla_r f_{R}(x^{**}_{uh};P) uh ] du \\ &+ \sqrt{mh} \int_0^{\infty} K\left( u \right) [ \nabla_{r^2} m_{S | R }(x^*_{uh};P) u^2 h^2 ] f_{R}(uh;P) du \\ =& \sqrt{mh} h \nabla_r m_{S | R }(0^+;P) f_{R}(0^+;P) \int_0^{\infty} u K\left( u \right) du \\ &+ \sqrt{mh} h^2 \nabla_r m_{S | R }(0^+;P) \int_0^{\infty} u^2 K\left( u \right) \nabla_r f_{R}(x^{**}_{uh};P) du \\ &+ \sqrt{mh} h^2 \int_0^{\infty} u^2 K\left( u \right) \nabla_{r^2} m_{S | R }(x^*_{uh};P) f_{R}(uh;P) du \\ =& \sqrt{mh} h \nabla_r m_{S | R }(0^+;P) f_{R}(0^+;P) \kappa_{1,2}^{+} \\ &+ O_{\mathcal{P}} \left( \sqrt{mh} h^2 \right) \\ &+ O_{\mathcal{P}} \left( \sqrt{mh} h^2 \right) \\ =& \sqrt{mh} h \underbrace{ \nabla_r m_{S | R }(0^+;P) f_{R}(0^+;P) \kappa_{1,2}^{+}}_{\doteq B_0(P) } + o_{\mathcal{P}} \left( 1 \right) \\ =& \sqrt{mh} h B_0(P) + o_{\mathcal{P}} \left( 1 \right), \end{align*} where $\kappa_{s,t}^{+}=\int_0^{\infty} u^s K^t(u) du$, and we use the existence and boundedness of derivatives, kernel moments, and the rate conditions $\sqrt{mh}h = O(1)$ and $\sqrt{mh}h^2 = o(1)$. \begin{align*} & \mathbb{V}_P\left[ \frac{1}{\sqrt{m h} } \sum_{i=1}^{m} K\left( \frac{ R_{i} }{h} \right) \left( m_{S | R }(R_i;P) - m_{S | R}(0^+;P) \right) \right] \\ =&h^2 \int_0^{\infty} K^2\left( u \right) [\nabla_r m_{S|R}(x^*_{uh};P)]^2 u^2 f_{R}(uh;P) du = h^2 O_{\mathcal{P}}(1) = o_{\mathcal{P}}(1). \end{align*} This gives, \begin{align*} & \frac{1}{\sqrt{m h} } \sum_{i=1}^{m} K\left( \frac{ R_{i} }{h} \right) \left( m_{S | R }(R_i ;P) - m_{S | R }( 0^+;P) \right) \\ & = \sqrt{mh} h B_0(P) + o_{\mathcal{P}}(1). \end{align*} Third, combine the first and second steps: \begin{align} & \left( \frac{1}{\sqrt{m h} } \sum_{i=1}^{m} K\left( \frac{ R_{i} }{h} \right) \left( m_{S | R }(R_i;P) - m_{S | R}(0^+;P) \right) \right) \left( \frac{1}{ m h } \sum_{i=1}^{m} K\left( \frac{ R_{i} }{h} \right) \right)^{-1} \notag \\ &= \sqrt{mh} h \underbrace{ 2 \kappa_{1,2}^{+} \nabla_r m_{S | R }(0^+;P)}_{\doteq B(P)} + o_{\mathcal{P}}(1) \\ &= \sqrt{mh} h B(P) + o_{\mathcal{P}}(1), \notag \end{align} which gives the bias term $B(P)$. That term is consistently estimated by $\widehat{B}$, \[ \widehat{B} = B_{m,n}(\boldsymbol{\mathrm{V}}_1 , \ldots, \boldsymbol{\mathrm{V}}_m) \doteq 2 \kappa_{1,1}^{+} \widehat{\nabla_r m_{S|R}}(x;P), \] that is, by replacing $\nabla_r m_{S|R}(x;P)$ in $B(P)$ by a consistent nonparametric estimator. The additional tuning parameters for such estimator and corresponding moment conditions may be set such that the bias-correction condition in Assumption (ref) is met. • \underline{Assumption (ref) - (ref):} zero mean of influence function. $\mathbb{E}_P[ \psi_{n}(\boldsymbol{\mathrm{V}}_{i}, P) ]=0 ~~\forall P$ by construction. • \underline{Assumption (ref) - (ref):} variance of influence function. Call $\xi^2(P) = 4 \kappa_{0,2}^{+} v_{S | R}(0^+;P) /f_{R}(0^+;P)$, where $\kappa_{0,2}^{+} = \int_{0}^{\infty}K^2(u) ~ du$. \begin{align*} &\mathbb{V}_P\left( \frac{2}{f_{R}(0^+;P) \sqrt{h} } K\left( \frac{ R_{i} }{h} \right) \left( S_{i} - m_{S | R}(R_i;P) \right) \right) - \xi^2(P) \\ =&\frac{4}{f_{R}^2(0^+;P) } \int_0^{\infty} K^2(u) \left\{ \underbrace{ v_{S | R}(uh;P) f_{R}(uh;P) }_{\doteq g(uh;P)} - \underbrace{ v_{S | R}(0^+;P) f_{R}(0^+;P) }_{\doteq g(0^+;P) } \right\} du \\ =&\frac{4h}{f_{R}^2(0^+;P) } \int_0^{\infty} K^2(u) \nabla_r g(x^*_{uh};P) u du = o_{\mathcal{P}}(1). \end{align*} Therefore, \[ \sup\limits_{P \in \mathcal{P} } \left| \mathbb{V}_P[ \psi_n(\boldsymbol{\mathrm{V}}_{i}, {P} ) ] - \xi^2(P) \right| \to 0. \]\underline{Assumption (ref) - (ref):} $ \sup_P \mathbb{E} [\psi_n^2(\boldsymbol{\mathrm{Z}}_{k,i},P)] < \infty .$ We have, \begin{align*} & \psi_n^2(\boldsymbol{\mathrm{Z}}_k,P) = \frac{4K^2\left( \frac{ X_{k} }{h} \right)}{ h f_{R}^{2}(0^+;P)} \left[ \left(Y_{k} - m_{Y_k|X_k}(X_k) \right) + \left(m_{Y_k|X_k}(X_k) - m_{S|R}(X_k;P) \right) \right]^2 , \\ &\mathbb{E}\left[ \psi_n^2(\boldsymbol{\mathrm{Z}}_k,P) \right] \\ &= \frac{4}{f_{R}^{2}(0^+;P)} \int_0^{\infty} K^2(u) \left[ v_{Y_k|X_k}(uh) + \left(m_{Y_k|X_k}(uh) - m_{S|R}(uh;P) \right)^2 \right] f_{X_k}(uh) du \\ &= O_{\mathcal{P}}(1). \end{align*} • \underline{Assumption (ref) - (ref):} $(2+\zeta)$-th moment condition. First, $\mathbb{V}_P^{-1}\left( \psi_n(\boldsymbol{\mathrm{V}}_i,P) \right) = O_{\mathcal{P}}(1). $ Second, $n^{-\zeta/2} \mathbb{E}_P \left| \psi_n(\boldsymbol{\mathrm{V}}_i,P) \right| ^{2+\zeta} = o_{\mathcal{P}}(1)$ because $\eta(r;P) = O_{\mathcal{P}}(1)$. Therefore, \begin{align*} n^{-\zeta/2} \sup\limits_{P \in \mathcal{P} } \mathbb{E}_P\left| \frac{ \psi_n(\boldsymbol{\mathrm{V}}_i,P) }{ \sqrt{\mathbb{V}_P\left( \psi_n(\boldsymbol{\mathrm{V}}_i,P) \right)} } \right|^{2 + \zeta} = o(1). \end{align*} • \underline{Assumption (ref) - (ref):} $\xi^2\left( \frac{m}{n} P_1 + \frac{n-m}{n} P_2 \right) \to \xi^2 \left( \gamma P_1 + (1-\gamma) P_2 \right)$. Let $\overline{P}_n = \frac{m}{n} P_1 + \frac{n-m}{n} P_2$ and $\overline{P} = \gamma P_1 + (1-\gamma ) P_2$. We have that \\ $\xi^2(\overline{P}_n) = \kappa_{0,2}^{+} v_{S | R }(0^+; \overline{P}_n) /f_{R}(0^+;\bar{P}_n)$, $v_{ S | R }(0^+; \bar{P}_n ) \to v_{ S | R }(0^+; \bar{P})$ and $f_{R}(0^+;\bar{P}_n) \to f_{R}(0^+;\bar{P})$ as in Section (ref). Therefore, $\xi^2(\overline{P}_n) \to \xi^2(\overline{P})$. • \underline{Assumption (ref) :} $ (n_1 /n - \lambda) \overset{p}{\to} 0$. In this setting, $n_1/n = \sum_i \mathbb{I}\{ X_i \geq 0 \} /n$ and $(n_1 /n - \lambda) \overset{p}{\to} 0$ holds. Assumption (ref) has already been verified above for any sequence $m_n$ such that $(m/n -\gamma) \to 0$ for arbitrary $\gamma \in (0,1)$. In particular it holds for $\gamma \in \{ \lambda, 1-\lambda\}.$ \end{enumerate} $\square$ \subsection{Proof of Proposition (ref) - Discontinuity of Density} The goal of this proof is to use the assumptions listed in Proposition (ref) to verify Assumptions (ref) and (ref). It follows the general lines of the proof of Proposition (ref), so the reader may refer to Section (ref) for the redundant details that we omit here. Additional references on the kernel density estimator at boundary points include \citeSM{marron1994} and Sections 1.9--1.10 of \citeSM{liracine2007}. Consider an \textit{iid} sample from $P\in\mathcal{P}$ with $m$ observations, $\boldsymbol{\mathrm{V}}_1=R_1, \ldots, \boldsymbol{\mathrm{V}}_m=R_m$, where $R$ is a scalar random variable and $0$ is an interior point of its compact support. The number $m$ grows with $n$ such that $m/n \to \gamma$, for some $\gamma \in (0,1)$. The parameter of interest is $\theta(P) = (f_R(0^+;P) - f_R(0^-;P))/2$, where $f_R(0^+;P)$ is the side limit as $x \downarrow 0$ of the PDF of $R$ under distribution $P$, and $f_R(0^-;P)$ is the negative side limit. The kernel density estimator for $\theta(P)$ is \[ \widehat{\theta}_{k}^b = \frac{1}{ n_k h} \sum\limits_{i=1}^{n_k} K \left( \frac{ X_{k,i} }{ h } \right) \left(\mathbb{I}\{X_{k,i} \ge 0\} - \mathbb{I}\{X_{k,i} < 0\} \right) \text{, for }k=1,2. \] This section investigates the asymptotic behavior of $\widehat{\theta} = \widehat{\theta}^b - h \widehat{B}$, where the expression for $\widehat{B}$ is given below Equation (ref). The assumptions in Proposition (ref) imply the following facts: \begin{enumerate} • \textit{As $m\to\infty$, $h \to 0$, $m h \to \infty$; $\sqrt{m h} h =O(1)$, and $\sqrt{m h} h^2 =o(1)$;} • \textit{The distribution of $R$ has PDF $f_{R}(r;P)$ that is twice differentiable wrt $r$ except at $r=0$; $f_{R}(r;P)$ and the derivatives $\nabla_r f_{R}(r;P)$ and $\nabla_{r^2} f_{R}(r;P)$ are bounded functions of $(r,P)$; $f_{R}(0^+;P)$ and $f_{R}(0^-;P)$ are bounded away from zero as functions of $P$;} \end{enumerate} We re-write $\sqrt{m h} \left( \ha\theta -\theta(P) \right)$ $= \sqrt{m h} \left( \ha\theta^b -h \widehat{B} -\theta(P) \right)$ to find the asymptotic linear representation. \begin{align} &\sqrt{m h} \left( \ha\theta -\theta(P) \right) \nonumber \\ &= \frac{1}{\sqrt{m h} } \sum_{i=1}^{m} \left\{ K\left( \frac{ R_i }{h} \right)\left(\mathbb{I}\{R_i \ge 0\} - \mathbb{I}\{R_i < 0\} \right) \right. \notag \\ & \left. - \mathbb{E}_P\left[ K\left( \frac{ R }{h} \right)\left(\mathbb{I}\{R \ge 0\} - \mathbb{I}\{R < 0\} \right)\right] \right\} \\ &+ \frac{1}{\sqrt{m h} } \sum_{i=1}^{m} \mathbb{E}_P\left[ K\left( \frac{ R }{h} \right)\left(\mathbb{I}\{R \ge 0\} - \mathbb{I}\{R < 0\} \right)\right] \notag \\ & - \frac{ \sqrt{mh} }{2} \left(f_R(0^+;P) - f_R(0^-;P) \right) - \sqrt{mh} h \widehat{B}. \end{align} \begin{enumerate} • \underline{Assumption (ref) - (ref):} asymptotic expansion. Equation (ref) above gives the influence function $\psi_n$, that is, $\eqref{eq:bunching:thetahat:clt}=m^{-1/2} \sum_{i=1}^m \psi_n(\boldsymbol{\mathrm{V}}_i,P)$, where \begin{align*} \psi_n(\boldsymbol{\mathrm{V}}_i,P) & \doteq \frac{1}{\sqrt{h} } \left\{ K\left( \frac{ R_i }{h} \right)\left(\mathbb{I}\{R_i \ge 0\} - \mathbb{I}\{R_i < 0\} \right) \right. \notag \\ & \left. - \mathbb{E}_P\left[ K\left( \frac{ R }{h} \right)\left(\mathbb{I}\{R \ge 0\} - \mathbb{I}\{R < 0\} \right)\right] \right\}. \end{align*} We need to show that Equation (ref) converges to zero uniformly over $\mathcal{P}$. \\ Define $\kappa_{s,t}^+ = \int_0^{\infty} u^s K^t(u) ~du$. First, consider the limit of (ref) $+ \sqrt{mh}h \widehat{B}$. \begin{align} &\frac{1}{\sqrt{m h} } \sum_{i=1}^{m} \mathbb{E}_P\left[ K\left( \frac{ R }{h} \right)\left(\mathbb{I}\{R \ge 0\} - \mathbb{I}\{R < 0\} \right)\right] - \frac{\sqrt{mh}}{2}\left(f_R(0^+;P)-f_R(0^-;P)\right) \notag \\ = & \sqrt{mh}\left( \int_0^{\infty} K(u) f_{R}(uh;P) du -\int_{-\infty}^0 K(u) f_{R}(uh;P) du \right. \notag \\ & \left. - \frac{1}{2}\left(f_R(0^+;P) - f_R(0^-;P)\right) \right) \notag \\ = &\sqrt{mh} \left( \int_0^{\infty} K(u) [ \nabla_r f_{R}(0^+;P)uh +\nabla_{r^2} f_{R}(x^*_{uh};P)u^2 h^2/2 ] du \right. \notag \\ & \left. - \int_{-\infty}^0 K(u) [ \nabla_r f_{R}(0^-;P) uh + \nabla_{r^2} f_{R}(x^{**}_{uh};P) u^2 h^2/2 ] du \right) \notag \\ = &\sqrt{mh} h \nabla_r f_{R}(0^+;P) \int_0^{\infty} u K(u) du \notag \\ &+\sqrt{mh} (h^2)/2 \int_0^{\infty} u^2 K(u) \nabla_{r^2} f_{R}(x^*_{uh};P) du \notag \\ & - \sqrt{mh} h \nabla_r f_{R}(0^-;P) \int_{-\infty}^0 u K(u) du \notag \\ & - \sqrt{mh} (h^2/2) \int_{-\infty}^0 u^2 K(u) \nabla_{r^2} f_{R}(x^{**}_{uh};P) du \notag \\ = &\sqrt{mh} h \nabla_r f_{R}(0^+;P) \kappa_{1,1}^{+} \notag \\ &+O_{\mathcal{P}} \left( \sqrt{mh} h^2 \right) \notag \\ & + \sqrt{mh} h \nabla_r f_{R}(0^-;P) \kappa_{1,1}^{+} \notag \\ & +O_{\mathcal{P}} \left( \sqrt{mh} h^2 \right) \notag \\ = & \sqrt{mh}h \underbrace{\kappa_{1,1}^{+} \left( \nabla_r f_{R}(0^+;P) + \nabla_r f_{R}(0^-;P) \right)}_{\doteq B(P)}+ o_{\mathcal{P}}(1) \\ = & \sqrt{mh}h B(P) + o_{\mathcal{P}}(1), \notag \end{align} where we used the existence and boundedness of derivatives, kernel moments, and the rate condition $\sqrt{mh}h^2=o(1)$. Equation (ref) defines the bias term $B(P)$. That term is consistently estimated by $\widehat{B}$, \[ \widehat{B} = B_{m,n}(\boldsymbol{\mathrm{V}}_1 , \ldots, \boldsymbol{\mathrm{V}}_m) \doteq \kappa_{1,1}^{+} \left[ \widehat{ \nabla_r f_{R}} (0^+;P) + \widehat{ \nabla_r f_{R}} (0^-;P) \right], \] that is, by replacing $\nabla_r f_{R}(0^+;P)$ and $\nabla_r f_{R}(0^-;P)$ in $B(P)$ by consistent nonparametric estimators. The tuning parameters for these additional estimators and corresponding moment conditions may be set such that the bias-correction condition in Assumption (ref) is met. Finally, the limit of (ref) is, \begin{align*} & \frac{1}{\sqrt{m h} } \sum_{i=1}^{m} \mathbb{E}_P\left[ K\left( \frac{ R }{h} \right)\left(\mathbb{I}\{R \ge 0\} - \mathbb{I}\{R < 0\} \right)\right] \notag \\ & - \frac{ \sqrt{mh} }{2} \left(f_R(0^+;P) - f_R(0^-;P) \right) - \sqrt{mh} h \widehat{B} \\ =& \sqrt{mh} h \left(B(P) - \widehat{B} \right) + o_{\mathcal{P}}(1) = o_{\mathcal{P}}(1), \end{align*} where we use $\widehat{B} - B(P) = o_{\mathcal{P}}(1)$ and $\sqrt{mh}h=O(1)$. • \underline{Assumption (ref) - (ref):} zero mean of influence function. $\mathbb{E}_P[ \psi_{n}(\boldsymbol{\mathrm{V}}_i, P) ]=0 ~~\forall P$ by construction. • \underline{Assumption (ref) - (ref):} variance of influence function. Call $\xi^2(P) = \kappa_{0,2}^{+} \left( f_{R}(0^+;P) + f_{R}(0^-;P) \right) $, where $\kappa_{0,2}^{+} = \int_{0}^{\infty}K^2(u) ~ du$. \begin{align*} &\mathbb{V}_P\left( \frac{1}{ \sqrt{h} } K\left( \frac{R_i }{h} \right)\left(\mathbb{I}\{R_i \ge 0\} - \mathbb{I}\{R_i < 0\} \right) \right) - \xi^2(P) \\ =&\frac{1}{h} \left[ \int_{-\infty}^{\infty} K^2(u)f_R(uh;P)h du \right. \\ & \left. - \left( \int_0^{\infty} K(u)f_R(uh;P)h du - \int_{-\infty}^0 K(u)f_R(uh;P)h du \right)^2 \right] - \xi^2(P) \\ =& \int_{0}^{\infty} K^2(u)[f_R(0^+;P) + \nabla_r f_{R}(x^*_{uh};P) uh ] du \\ &+ \int_{-\infty}^{0} K^2(u)[f_R(0^-;P) + \nabla_r f_{R}(x^{**}_{uh};P)uh ] du \\ & - h \left(\int_0^{\infty} K(u) f_R(uh;P) du - \int_{-\infty}^0 K(u)f_R(uh;P) du\right)^2 - \xi^2(P) \\ =&h \left[\int_{0}^{\infty} K^2(u) \nabla_r f_R(x^*_{uh};P) u du + \int_{-\infty}^{0} K^2(u) \nabla_r f_R(x^{**}_{uh};P) u du \right. \\ & \left. - \left(\int_0^{\infty} K(u) f_R(uh;P) du - \int_{-\infty}^0 K(u) f_R(uh;P) du\right)^2 \right] \\ =& h O_{\mathcal{P}}(1) = o_{\mathcal{P}}(1). \end{align*} Therefore, \[ \sup\limits_{P \in \mathcal{P} } \left| \mathbb{V}_P [ \psi_n(\boldsymbol{\mathrm{V}}_i, {P} ) ] - \xi^2(P) \right| \to 0. \]\underline{Assumption (ref) - (ref):} $ \sup_P \mathbb{E} [\psi_n^2(\boldsymbol{\mathrm{Z}}_k,P)] < \infty$ for $k=1,2.$ We have, \begin{align*} & \psi_n(\boldsymbol{\mathrm{Z}}_k,P) = h^{-1/2} K\left( \frac{X_{k} }{h} \right)\left(\mathbb{I}\{X_{k} \ge 0\} - \mathbb{I}\{X_{k} < 0\} \right) \\ & - h^{-1/2} \mathbb{E}_P \left[ K\left( \frac{ V }{h} \right)\left(\mathbb{I}\{ V \geq 0\} - \mathbb{I}\{V < 0\} \right)\right], \\ &\psi_n^2(\boldsymbol{\mathrm{Z}}_k,P) = \frac{1 }{h} K^2\left( \frac{X_{k} }{h} \right) \\ & + \frac{1}{h} \left\{ \mathbb{E}_P\left[ K\left( \frac{ V }{h} \right) \left(\mathbb{I}\{V \ge 0\} - \mathbb{I}\{V < 0\} \right)\right] \right\}^2 \\ & -\frac{2}{h} K\left( \frac{ X_k }{h} \right) \left(\mathbb{I}\{X_k \ge 0\} - \mathbb{I}\{X_k < 0\} \right) \mathbb{E}_P\left[ K\left( \frac{ V }{h} \right) \left(\mathbb{I}\{V \ge 0\} - \mathbb{I}\{V < 0\} \right)\right], \\ &\mathbb{E}\left[ \psi_n^2(\boldsymbol{\mathrm{Z}}_k,P) \right] = \int K^2\left( u \right) f_{X_k}(uh) du \\ & + h \left\{ \int_0^{\infty} K\left( u \right) f_{R}(uh;P) du - \int_{-\infty}^{0} K\left( u \right) f_{R}(uh;P) du \right\}^2 \\ & -2 \left\{ \int_0^{\infty} K\left( u \right) f_{X_k}(uh) du - \int_{-\infty}^{0} K\left( u \right) f_{X_k}(uh) du \right\} \\ & \times h \left\{ \int_0^{\infty} K\left( u \right) f_{R}(uh;P) du - \int_{-\infty}^{0} K\left( u \right) f_{R}(uh;P) du \right\}, \\ & \mathbb{E}\left[ \psi_n^2(\boldsymbol{\mathrm{Z}}_k,P) \right] = O_{\mathcal{P}}(1) + O_{\mathcal{P}}(h) + O_{\mathcal{P}}(h) = O_{\mathcal{P}}(1). \end{align*} • \underline{Assumption (ref) - (ref):} $(2+\zeta)$-th moment condition. We verify it in two steps. First, $\mathbb{V}_P\left( \psi_n(\boldsymbol{\mathrm{Z}}_i,P) \right) - \xi^2(P) = o_{\mathcal{P}}(1)$ and $\xi^2(P)$ is bounded away from zero, uniformly over $\mathcal{P}$. Thus, $\mathbb{V}_P^{-1}\left( \psi_n(\boldsymbol{\mathrm{Z}}_i,P) \right) = O_{\mathcal{P}}(1)$. Second, pick $\zeta>0$. From before, $\mathbb{E}_P\left[ K\left( \frac{ R }{h} \right)\left(\mathbb{I}\{R \ge 0\} - \mathbb{I}\{R< 0\} \right)\right] = O_{\mathcal{P}}(h)$. \begin{align*} & n^{-\zeta/2} \mathbb{E}_P \left| \psi_n(\boldsymbol{\mathrm{V}}_{i}, {P} ) \right|^{2+\zeta} \\ &= \frac{1}{(nh)^{\zeta/2} h } \mathbb{E}_P \left| K\left( \frac{ R_i }{h} \right)\left(\mathbb{I}\{V_i \ge 0\} - \mathbb{I}\{V_i < 0\} \right) - O_{\mathcal{P}}(h) \right| ^{2+\zeta} \\ &= \frac{1}{(nh)^{\zeta/2} } \int_0^{\infty} \left| K(u) - O_{\mathcal{P}}(h) \right|^{2+\zeta} f_R(uh;P) du \\ & + \frac{1}{(nh)^{\zeta/2} } \int_{-\infty}^{0} \left| - K(u) - O_{\mathcal{P}}(h) \right|^{2+\zeta} f_R(uh;P) du \\ &= \frac{1}{(nh)^{\zeta/2} } O_{\mathcal{P}}(1) = o_{\mathcal{P}}(1). \end{align*} Combining both steps gives, \begin{align*} n^{-\zeta/2} \sup\limits_{P \in \mathcal{P} } \mathbb{E}_P\left| \frac{ \psi_n(\boldsymbol{\mathrm{V}}_i,P) }{ \sqrt{\mathbb{V}_P\left( \psi_n(\boldsymbol{\mathrm{V}}_i,P) \right)} } \right|^{2 + \zeta} = o(1). \end{align*} • \underline{Assumption (ref) - (ref):} $\xi^2\left( \frac{m}{n} P_1 + \frac{n-m}{n} P_2 \right) \to \xi^2 \left( \gamma P_1 + (1-\gamma) P_2 \right)$. Let $\overline{P}_n = \frac{m}{n} P_1 + \frac{n-m}{n} P_2$ and $\overline{P} = \gamma P_1 + (1-\gamma) P_2$. Given that \\ $\xi^2(\overline{P}_n) = \kappa_{0,2}^{+} \left( f _{R}(0^+; \overline{P}_n ) + f _{R}(0^-; \overline{P}_n ) \right) ~ du$, the result follows because $m/n \to \gamma$, $f _{R}(0^+; \overline{P}_n ) \to f _{R}(0^+; \overline{P} ) $, and $f _{R}(0^-; \overline{P}_n ) \to f _{R}(0^-; \overline{P} ) $. • \underline{Assumption (ref) :} $ (n_1 /n - \lambda) \overset{p}{\to} 0$. In this setting, $n_1$ is a deterministic sequence, and we have that $n_1/n = \lfloor n/2 \rfloor / n \to 1/2$. Assumption (ref) has already been verified above for any sequence $m_n$ such that $(m/n -\gamma) \to 0$ for arbitrary $\gamma \in (0,1)$. In particular it holds for $\gamma = 1/2.$ \end{enumerate} $\square$ \subsection{Extension of Proposition (ref) - Local Polynomial Regression} Sections (ref) and (ref) gave sufficient conditions and demonstrated the asymptotic validity of our robust permutation test in the context of RDD using the NW estimator. This section generalizes that finding to the Local Polynomial Regression (LPR) estimator of order $\rho \in \mathbb{Z}_+$. The researcher has an \textit{iid} sample $(X_1,Y_1), \ldots, (X_n,Y_n)$ that is split into two samples $\boldsymbol{\mathrm{Z}}_{k,i} = (X_{k,i},Y_{k,i})$, $k=1,2$, $i =1, \ldots, n_k$, as explained in Section (ref), where the distribution of $\boldsymbol{\mathrm{Z}}_{k,i}$ conditional on $\underline{W}_n$ is $P_k$. We have that $P_1$ is the distribution of $(X,Y)|X \geq 0$, $P_2$ is the distribution of $(-X,Y)|X < 0$, and $\mathcal{P}$ is the set of all convex combinations of $P_1$ and $P_2$. The researcher selects a polynomial order $\rho \in \mathbb{Z}_+$ and a bandwidth $h>0$. The LPR estimator for $\theta(P_k)$ is defined as follows: \begin{align*} &(\widehat{a}, \widehat{\boldsymbol{\mathrm{b}}} ) = \mathop{\hbox{\rm argmin}}\limits_{(a, \boldsymbol{\mathrm{b}} )} \sum\limits_{i=1}^{n_k} K \left( \frac{X_{k,i} }{h } \right) \bigg[ Y_{k,i} - a - b_1 X_{k,i} - b_2 X_{k,i}^2 - \ldots -b_{\rho} X_{k,i}^{\rho} \bigg]^2, \\ &\widehat{\theta}^b_k = \theta_{n_k,n}^b(\boldsymbol{\mathrm{Z}}_{k,1} , \ldots, \boldsymbol{\mathrm{Z}}_{k,n_k}) = \widehat{a}, \end{align*} where ${a}$ is a scalar and ${\boldsymbol{\mathrm{b}}}$ is the vector $(b_1,\ldots,b_{\rho})$. Note that the LPR estimator becomes the NW estimator when we set $\rho=0$. The superscript $b$ in $\widehat{\theta}_{k}^b$ indicates there is bias in the asymptotic distribution of $\sqrt{n h} (\widehat{\theta}_{k}^b - \theta(P_k))$ whenever the bandwidth choice converges to zero at the slowest possible rate, i.e., $h=O(n^{-1/(2\rho +3)})$ (Proposition (ref) below). This is the case of MSE-optimal bandwidths and inference requires bias correction in that case. A conventional solution is to subtract a first-order bias term $h^{\rho+1} B(P_k)$ from $\widehat{\theta}_{k}^b$, where $B(P_k)$ is nonparametrically estimated by $\widehat{B}_k = B_{n_k,n}(\boldsymbol{\mathrm{Z}}_{k,1}, \ldots, \boldsymbol{\mathrm{Z}}_{k,n_k})$. We give the analytical formulas for $B(P_k)$ and $\widehat{B}_k$ in the proof of Proposition (ref) below (Equations (ref) and (ref)). Our permutation tests utilize the bias-corrected LPR estimator $\widehat{\theta}_{k} = \theta_{n_k,n}(\boldsymbol{\mathrm{Z}}_{k,1}, \ldots, \boldsymbol{\mathrm{Z}}_{k,n_k}) \doteq \theta_{n_k,n}^b(\boldsymbol{\mathrm{Z}}_{k,1}, \ldots, \boldsymbol{\mathrm{Z}}_{k,n_k}) - h^{\rho+1} B_{n_k,n}(\boldsymbol{\mathrm{Z}}_{k,1}, \ldots, \boldsymbol{\mathrm{Z}}_{k,n_k}) =\widehat{\theta}_{k}^b - h^{\rho+1} \widehat{B}_k$. Note that no bias correction is needed if $h=o(n^{-1/(2\rho +3)})$. \begin{proposition} Assume that: (i) as $n\to \infty$, $h \to 0$, $n h \to \infty$, and $\sqrt{n h} h^{\rho+1} \to c \in [0,\infty)$; (ii) $K$ is a kernel density function that is non-negative, bounded, symmetric, and $\int K(u) |u|^{4(\rho+1)}~du <\infty$; (iii) the distribution of $X$ has PDF $f_{X}$ that is bounded, bounded away from zero, $\rho+2$ times differentiable except at $x=0$, and has bounded derivatives; (iv) $\mathbb{E}[Y | X = x]$ is bounded, $\rho+2$ times differentiable except at $x=0$, and has bounded derivatives; (v) $\mathbb{V}[Y | X = x]$ is bounded, differentiable except at $x=0$, has bounded derivative, $\mathbb{V}[Y|X=0^+]>0$, and $\mathbb{V}[Y|X=0^-]>0$, where $0^+$ and $0^-$ denote side limits; and (vi) $\exists \zeta>0$ such that $\mathbb{E}[|Y|^{2+\zeta} | X]$ is almost surely bounded. Let $\boldsymbol{\mathrm{V}}_1=(R_1,S_1), \ldots, \boldsymbol{\mathrm{V}}_m=(R_m,S_m)$ be an \textit{iid} sample from a distribution $P \in \mathcal{P}$, where $m$ grows with $n$. Use the definitions above to construct the the LPR estimators: $\widehat{\theta}^b = \theta_{m,n}^b(\boldsymbol{\mathrm{V}}_{1}, \ldots, \boldsymbol{\mathrm{V}}_{m})$, $\widehat{B} = B_{m,n}(\boldsymbol{\mathrm{V}}_{1}, \ldots, \boldsymbol{\mathrm{V}}_{m})$, and $\widehat{\theta} = \theta_{m,n}(\boldsymbol{\mathrm{V}}_{1}, \ldots, \boldsymbol{\mathrm{V}}_{m})$. Finally, assume that (vii) $\widehat{B} - B(P) \overset{p}{\to} 0$ uniformly over $P \in \mathcal{P}$ for any sequence $m$ such that $ m/n \to \lambda $ or $ m/n \to 1-\lambda $. Then, the RDD setting with the LPR estimator defined above satisfies Assumptions (ref) and (ref) with \begin{align*} \psi_n(\boldsymbol{\mathrm{V}},P) & = \frac{1}{\sqrt{h} f_R(0^+;P) } K\left( \frac{ R }{h } \right) {\boldsymbol{\mathrm{e}}_1}' \boldsymbol{\mathrm{\Gamma}}^{-1} \widetilde{\boldsymbol{\mathrm{H}}}^{\rho} \left(S - m_{S|R}(R;P) \right), \\ \xi^2(P) = & {\boldsymbol{\mathrm{e}}_1}' \boldsymbol{\mathrm{\Gamma}}^{-1} \boldsymbol{\mathrm{\Delta}} \boldsymbol{\mathrm{\Gamma}}^{-1} {\boldsymbol{\mathrm{e}}_1} \frac{ v_{S|R}(0^+;P) }{f_R(0^+;P)}, \end{align*} where the following definitions are used, \begin{align*} & \boldsymbol{\mathrm{e}}_1 \text{ is the } (\rho_1+1 \times 1) \text{ column vector } \boldsymbol{\mathrm{e}}_1=[1 0 0 \cdots 0]', \\ & \widetilde{\boldsymbol{\mathrm{H}}}^{\rho} = \left[ 1 \left( \frac{R}{h} \right) \cdots \left( \frac{R }{h } \right)^{\rho } \right]', (\rho+1 \times 1) \text{ column vector, } \\ & \kappa_{s,t}^{+} = \int_{0}^{\infty} u^s K^t(u) du \text{ for integers s and t, } \\ & \boldsymbol{\mathrm{\Gamma}} = \left[ \begin{array}{ccc} \kappa_{0,1}^{+} & \ldots & \kappa_{\rho,1}^{+} \\ \vdots & \vdots & \vdots \\ \kappa_{\rho,1}^{+} & \ldots & \kappa_{2 \rho,1}^{+} \end{array} \right] \text{ and } \boldsymbol{\mathrm{\Delta}} = \left[ \begin{array}{ccc} \kappa_{0,2}^{+} & \ldots & \kappa_{\rho,2}^{+} \\ \vdots & \vdots & \vdots \\ \kappa_{\rho,2}^{+} & \ldots & \kappa_{2 \rho,2}^{+} \end{array} \right]. \end{align*} Moreover, $m_{S|R}(r;P)$ is the conditional mean of $S$ given $R=r$, $v_{S|R}(r;P)$ is the conditional variance of $S$ given $R=r$, and $f_{R}(r;P)$ is the PDF of $R$ at $r$, all three assuming $\boldsymbol{\mathrm{V}}=(R,S) \sim P \in \mathcal{P}$. \end{proposition} \textbf{\textit{Proof. }} The goal of this proof is to use the assumptions listed in the proposition to verify that the RDD setting with the LPR estimator satisfies Assumptions (ref) and (ref). It follows the general lines of the proof of Proposition (ref), so the reader may refer to Section (ref) for the redundant details that we omit here. The proof also builds on arguments from the literature on asymptotic approximations of the LPR estimator at a boundary point: porter2003 (Theorem 3(a)) and fan1996 (Theorem 3.2). Consider an \textit{iid} sample from $P\in\mathcal{P}$ with $m$ observations, $\boldsymbol{\mathrm{V}}_1=(R_1,S_1), \ldots, \boldsymbol{\mathrm{V}}_m=(R_m,S_m)$, where the minimum value in the support of $R$ is 0. The number $m$ grows with $n$ such that $m/n \to \gamma$, for some $\gamma \in (0,1)$. The assumptions listed in the proposition imply the following facts (see proof of Proposition (ref) in Section (ref) for details): \begin{enumerate} • \textit{As $m\to\infty$, $h \to 0$, $m h \to \infty$, $\sqrt{m h} h^{\rho+1} =O(1),$ and $\sqrt{m h} h^{\rho+2} =o(1)$}; • \textit{The distribution of $R$ has PDF $f_{R}(r;P)$ that is $\rho+2$ times differentiable wrt $r$ denoted $\nabla_r f_{R}(r;P)$, \ldots, $\nabla_{r^{\rho+2}} f_{R}(r;P)$; $f_{R}(r;P)$ and its derivatives are bounded functions of $(r,P)$; $f_{R}(r;P)$ is bounded away from zero as a function of $(r,P)$;} • \textit{$m_{S|R}(r;P) = \mathbb{E}_P[S|R=r]$ is $\rho+2$ times differentiable wrt $r$ denoted $\nabla_r m_{S|R}(r;P)$, \ldots, $\nabla_{r^{\rho+2}} m_{S|R}(r;P)$; $m_{S|R}$ and its derivatives are bounded functions of $(r,P)$;} • \textit{$v_{S|R}(r;P) = \mathbb{V}_P[S|R=r]$ has first derivative wrt $r$ denoted $\nabla_r v_{S|R}(r;P)$; $v_{S|R}$, $\nabla_r v_{S|R}$ are both bounded functions of $(r,P)$; $v_{S|R}(0^+;P)$ is bounded away from zero as a function of $P$;} • \textit{$\eta(r;P) \doteq \mathbb{E}_P[|S - m_{S | R }(R;P) |^{2+\zeta} | R=r]$ is a bounded function of $(r,P)$.} \end{enumerate} Use the data $\boldsymbol{\mathrm{V}}_1=(R_1,S_1), \ldots, \boldsymbol{\mathrm{V}}_m=(R_m,S_m)$ and the definitions above to construct the the LPR estimators: $\widehat{\theta}^b = \theta_{m,n}^b(\boldsymbol{\mathrm{V}}_{1}, \ldots, \boldsymbol{\mathrm{V}}_{m})$, $\widehat{B} = B_{m,n}(\boldsymbol{\mathrm{V}}_{1}, \ldots, \boldsymbol{\mathrm{V}}_{m})$, and $\widehat{\theta} = \theta_{m,n}(\boldsymbol{\mathrm{V}}_{1}, \ldots, \boldsymbol{\mathrm{V}}_{m})$. The proof studies the asymptotic behavior of $\widehat{\theta} = \widehat{\theta}^b - h \widehat{B} $. Consider the following definitions, \begin{align*} &\boldsymbol{\mathrm{e}}_1 \text{ is the } (\rho_1+1 \times 1) \text{ column vector } \boldsymbol{\mathrm{e}}_1=[1 0 0 \cdots 0]', \\ &\boldsymbol{\mathrm{H}}_i^{\rho} = \left[ 1 R_i \cdots R_i^{\rho} \right]', (\rho+1 \times 1) \text{ column vector, } \\ &\widetilde{\boldsymbol{\mathrm{H}}}_i^{\rho} = \left[ 1 \left( \frac{R_i}{h} \right) \cdots \left( \frac{R_i }{h } \right)^{\rho } \right]', (\rho+1 \times 1) \text{ column vector, } \\ &\boldsymbol{\mathrm{G}}_m = \left[ \frac{1}{m h } \sum_{i=1}^m K\left( \frac{R_i }{h } \right) \widetilde{\boldsymbol{\mathrm{H}}}_{i}^{\rho} \widetilde{\boldsymbol{\mathrm{H}}}_{i}^{' \rho} \right]^{-1}, (\rho+1 \times \rho+1) \text{ matrix, } \\ &\boldsymbol{\mathrm{\varphi}}_{\rho}(P) = \left[ m_{S|R}(0^+;P) \nabla_r m_{S|R}(0^+;P)/1! \cdots \nabla_{r^{\rho}} m_{S|R}(0^+;P)/\rho! \right]', \\ & (\rho+1 \times 1) \text{ column vector, } \\ & \varepsilon_i = S_i - m_{S|R}(R_i;P), \text{scalar, } \\ &\bar{\varepsilon}_i^{\rho} = S_i - \boldsymbol{\mathrm{H}}_i^{\rho'} \boldsymbol{\mathrm{\varphi}}_{\rho}(P) = \varepsilon_i + m_{S|R}(R_i;P) - \boldsymbol{\mathrm{H}}_i^{\rho'} \boldsymbol{\mathrm{\varphi}}_{\rho}(P), \text{scalar. } \end{align*} It is possible to rewrite $\sqrt{m h} \left( \widehat{\theta} - \theta(P) \right)$ as follows, \begin{align} & \sqrt{m h} \left( \widehat{\theta} - \theta(P) \right) = \sqrt{m h} \left( \widehat{\theta}^b - \theta(P) \right) - \sqrt{m h} h^{\rho+1} \widehat{B} \nonumber \\ & ={\boldsymbol{\mathrm{e}}_1}' \boldsymbol{\mathrm{G}}_m \left[ \frac{1}{\sqrt{m h } } \sum_{i=1}^m K\left( \frac{ R_i }{h } \right) \widetilde{\boldsymbol{\mathrm{H}}}_{i}^{\rho} \bar{\varepsilon}_i^{\rho} \right] - \sqrt{m h} h^{\rho+1} \widehat{B} \nonumber \\ & ={\boldsymbol{\mathrm{e}}_1}' \boldsymbol{\mathrm{G}}_m \left[ \frac{1}{\sqrt{m h } } \sum_{i=1}^m K\left( \frac{ R_i }{h } \right) \widetilde{\boldsymbol{\mathrm{H}}}_{i}^{\rho} \left( \bar{\varepsilon}_i^{\rho} - \mathbb{E}_P [\bar{\varepsilon}_i^{\rho} |R_i] \right) \right] \\ & + {\boldsymbol{\mathrm{e}}_1}' \boldsymbol{\mathrm{G}}_m \left\{ \frac{1}{\sqrt{m h}} \sum_{i=1}^m K\left( \frac{ R_i }{h } \right) \widetilde{\boldsymbol{\mathrm{H}}}_{i}^{\rho} \mathbb{E}_P [\bar{\varepsilon}_i^{\rho} |R_i] - \mathbb{E}_P \left[ \frac{1}{\sqrt{mh}} \sum_{i=1}^m K\left( \frac{ R_i }{h } \right) \widetilde{\boldsymbol{\mathrm{H}}}_{i}^{\rho} \bar{\varepsilon}_i^{\rho} \right] \right\} \\ &+ {\boldsymbol{\mathrm{e}}_1}' \boldsymbol{\mathrm{G}}_m \mathbb{E}_P \left[ \frac{1}{\sqrt{m h }} \sum_{i=1}^m K\left( \frac{ R_i }{h } \right) \widetilde{\boldsymbol{\mathrm{H}}}_{i}^{\rho} \bar{\varepsilon}_i^{\rho} \right] - \sqrt{m h} h^{\rho+1} \widehat{B}. \end{align} Equation (ref) gives the influence function representation and Equations (ref)--(ref) are $o_{\mathcal{P}}(1)$. The limit of the first term in (ref) characterizes the bias term $B(P)$ (Equation (ref) below). In what follows, we verify Assumptions (ref) and (ref) in 7 parts. \begin{enumerate} • \underline{Assumption (ref) - (ref):} asymptotic expansion. \underline{\textit{Equation (ref):}} Let $\kappa_{s,t}^{+} = \int_{0}^{\infty} u^s K^t(u) ~ du$ and define \begin{align*} & \boldsymbol{\mathrm{\Gamma}} = \left[ \begin{array}{ccc} \kappa_{0,1}^{+} & \ldots & \kappa_{\rho,1}^{+} \\ \vdots & \vdots & \vdots \\ \kappa_{\rho,1}^{+} & \ldots & \kappa_{2 \rho,1}^{+} \end{array} \right], \\ & \boldsymbol{\mathrm{G}}(P) = \frac{1}{f_R(0^+;P)} \boldsymbol{\mathrm{\Gamma}}^{-1}. \end{align*} Rewrite Equation (ref) as, \begin{align} & {\boldsymbol{\mathrm{e}}_1}' \boldsymbol{\mathrm{G}}_m \left[ \frac{1}{\sqrt{m h } } \sum_{i=1}^m K\left( \frac{ R_i }{h } \right) \widetilde{\boldsymbol{\mathrm{H}}}_{i}^{\rho} \varepsilon_i \right] \nonumber \\ =& {\boldsymbol{\mathrm{e}}_1}' \left[ \boldsymbol{\mathrm{G}}_m - \boldsymbol{\mathrm{G}}(P) \right] \left[ \frac{1}{\sqrt{m h } } \sum_{i=1}^m K\left( \frac{ R_i }{h } \right) \widetilde{\boldsymbol{\mathrm{H}}}_{i}^{\rho} \varepsilon_i \right] \\ &+ \left[ \frac{1}{\sqrt{m h } } \sum_{i=1}^m K\left( \frac{ R_i }{h } \right) {\boldsymbol{\mathrm{e}}_1}' \boldsymbol{\mathrm{G}}(P) \widetilde{\boldsymbol{\mathrm{H}}}_{i}^{\rho} \varepsilon_i \right]. \end{align} We show that Equation (ref) is $o_{\mathcal{P}}(1)$ by showing that $\boldsymbol{\mathrm{G}}_m - \boldsymbol{\mathrm{G}}(P) = o_{\mathcal{P}}(1)$ and \\ $\frac{1}{\sqrt{m h } } \sum_{i=1}^m K\left( \frac{ R_i }{h } \right) \widetilde{\boldsymbol{\mathrm{H}}}_{i} \varepsilon_i$ is $O_{\mathcal{P}}(1)$. We have that $\boldsymbol{\mathrm{G}}_m^{-1}$ is a sample average and $\boldsymbol{\mathrm{G}}_m^{-1} \overset{p}{\to} \boldsymbol{\mathrm{G}}^{-1}(P)$ (porter2003, top of page 45). The convergence is made uniform over $\mathcal{P}$ by the fact that $\nabla_r f_R(r;P)$ is a bounded function of $(r,P)$ and the first $2 (\rho+1)$ moments of $K(u)$ are bounded. The absolute value of the determinant of the matrix $\boldsymbol{\mathrm{G}}^{-1}(P)$ is bounded away from zero because $f_R(0^+;P)$ is bounded away from zero over $P \in \mathcal{P}$. Therefore, $\boldsymbol{\mathrm{G}}_m - \boldsymbol{\mathrm{G}}(P) = o_{\mathcal{P}}(1)$. We have that $\frac{1}{\sqrt{m h } } \sum_{i=1}^m K\left( \frac{ R_i }{h } \right) \widetilde{\boldsymbol{\mathrm{H}}}_{i} \varepsilon_i$ is $O_{\mathcal{P}}(1)$ because it has zero mean and variance $O_{\mathcal{P}}(1)$. The variance is bounded because the kernel has bounded moments, and $f_R(r;P)$ and $v_{S|R}(r;P)$ are bounded functions of $(r,P)$. Equation (ref) gives the influence function. \begin{align*} \psi_n(\boldsymbol{\mathrm{V}}_i,P) & = \frac{1}{\sqrt{h} } K\left( \frac{ R_i }{h } \right) {\boldsymbol{\mathrm{e}}_1}' \boldsymbol{\mathrm{G}}(P) \widetilde{\boldsymbol{\mathrm{H}}}_{i}^{\rho} \left(S_i - m_{S|R}(R_i;P) \right) \\ & = \frac{1}{\sqrt{h} f_R(0^+;P) } K\left( \frac{ R_i }{h } \right) {\boldsymbol{\mathrm{e}}_1}' \boldsymbol{\mathrm{\Gamma}}^{-1} \widetilde{\boldsymbol{\mathrm{H}}}_{i}^{\rho} \left(S_i - m_{S|R}(R_i;P) \right). \end{align*} \underline{\textit{Equation (ref):}} Equation (ref) is shown to be $o_{\mathcal{P}}(1)$: from above $\boldsymbol{\mathrm{G}}_m= O_{\mathcal{P}}(1)$ and we show that \[ \frac{1}{\sqrt{m h}} \sum_{i=1}^m K\left( \frac{ R_i }{h } \right) \widetilde{\boldsymbol{\mathrm{H}}}_{i}^{\rho} \mathbb{E}_P[\bar{\varepsilon}_i^{\rho} |R_i] - \mathbb{E}_P \left[ \frac{1}{\sqrt{mh}} \sum_{i=1}^m K\left( \frac{ R_i }{h } \right) \widetilde{\boldsymbol{\mathrm{H}}}_{i}^{\rho} \bar{\varepsilon}_i^{\rho} \right] = o_{\mathcal{P}}(1). \] To see this last equation, note that \[ \mathbb{E}_P \left\{ \frac{1}{\sqrt{m h}} \sum_{i=1}^m K\left( \frac{ R_i }{h } \right) \widetilde{\boldsymbol{\mathrm{H}}}_{i}^{\rho} \mathbb{E}_P [\bar{\varepsilon}_i^{\rho} |R_i] \right\} = \mathbb{E}_P \left[ \frac{1}{\sqrt{mh}} \sum_{i=1}^m K\left( \frac{ R_i }{h } \right) \widetilde{\boldsymbol{\mathrm{H}}}_{i}^{\rho} \bar{\varepsilon}_i^{\rho} \right]. \] It remains to show that \[ \mathbb{V}_P \left\{ \frac{1}{\sqrt{m h} } \sum_{i=1}^m K\left( \frac{ R_i }{h } \right) \left( \frac{R_i}{h}\right)^l \mathbb{E}_P [\bar{\varepsilon}_i^{\rho} |R_i] \right\} = o_{\mathcal{P}}(1), \] for $l=0, \ldots, \rho$. In fact, \begin{align*} &\mathbb{V}_P \left\{ \frac{1}{\sqrt{m h} } \sum_{i=1}^m K\left( \frac{ R_i }{h } \right) \left( \frac{R_i}{h}\right)^l \mathbb{E}_P [\bar{\varepsilon}_i^{\rho} |R_i] \right\} \\ &=\frac{1}{m h } \sum_{i=1}^m \mathbb{V}_P \left\{ K\left( \frac{ R_i }{h } \right) \left( \frac{R_i}{h} \right)^l \mathbb{E}_P [\bar{\varepsilon}_i^{\rho} |R_i] \right\} \\ &=\frac{1}{h } \mathbb{V}_P \left\{ K\left( \frac{ R_i }{h } \right) \left( \frac{R_i}{h} \right)^l \mathbb{E}_P [\bar{\varepsilon}_i^{\rho} |R_i] \right\} \\ &\leq \frac{1}{h } \mathbb{E}_P \left\{ K^2 \left( \frac{ R_i }{h } \right) \left( \frac{R_i}{h} \right)^{2l} \left( \mathbb{E}_P [\bar{\varepsilon}_i^{\rho} |R_i] \right)^2 \right\} \\ &= \frac{1}{h } \mathbb{E}_P \left\{ K^2 \left( \frac{ R_i }{h } \right) \left( \frac{R_i}{h} \right)^{2l} \left( \mathbb{E}_P [\bar{\varepsilon}_i^{\rho+1} |R_i] + R_i^{\rho+1} \nabla_{r^{\rho+1}} m_{S|R}(0^+;P)/(\rho+1)! \right)^2 \right\} \\ &\leq \frac{2}{h } \mathbb{E}_P \left\{ K^2 \left( \frac{ R_i }{h } \right) \left( \frac{R_i}{h} \right)^{2l} \left[ \left( \mathbb{E}_P [\bar{\varepsilon}_i^{\rho+1} |R_i] \right)^2 + \left( R_i^{\rho+1} \nabla_{r^{\rho+1}} m_{S|R}(0^+;P)/(\rho+1)! \right)^2 \right] \right\} \\ &= 2 \bigintssss_0^{\infty} K^2 \left( u \right) u^{2l} \left( \mathbb{E}_P [\bar{\varepsilon}_i^{\rho+1} |R_i=uh] \right)^2 f_R(uh;P) du \\ & + 2 \frac{ h^{2\rho+2} }{ [(\rho+1)!]^2 } \left[ \nabla_{r^{\rho+1}} m_{S|R}(0^+;P) \right]^2 \bigintssss_0^{\infty} K^2 \left( u \right) u ^{2l+2\rho+2} f_R(uh;P) du \\ &= 2 \bigintssss_0^{\infty} K^2 \left( u \right) u^{2l} \left[ \nabla_{r^{\rho+2}} m_{S|R}(r^*_u;P) (uh)^{\rho+2}/(\rho+2)! \right]^2 f_R(uh;P) du \\ & + 2 \frac{ h^{2\rho+2} }{ [(\rho+1)!]^2 } \left[ \nabla_{r^{\rho+1}} m_{S|R}(0^+;P) \right]^2 \bigintssss_0^{\infty} K^2 \left( u \right) u ^{2l+2\rho+2} f_R(uh;P) du \\ &= 2 \frac{ h^{2\rho+4} }{ [(\rho+2)!]^2 } \bigintssss_0^{\infty} K^2 \left( u \right) u^{2l + 2\rho + 4} \left[ \nabla_{r^{\rho+2}} m_{S|R}(r^*_u;P) \right]^2 f_R(uh;P) du \\ & + 2 \frac{ h^{2\rho+2} }{ [(\rho+1)!]^2 } \left[ \nabla_{r^{\rho+1}} m_{S|R}(0^+;P) \right]^2 \bigintssss_0^{\infty} K^2 \left( u \right) u ^{2l+2\rho+2} f_R(uh;P) du \\ & = O_{\mathcal{P}} \left( h^{2\rho+4} \right) + O_{\mathcal{P}} \left( h^{2\rho+2} \right) = O_{\mathcal{P}} \left( h^{2 (\rho+1) } \right) = o_{\mathcal{P}} \left( 1 \right), \end{align*} where $r^*_u$ lies between $uh$ and $0$, and we use the following facts: (i) the first $4\rho+4$ moments of $K$ are bounded, (ii) the functions $f_R(r;P)$, $\nabla_{r^{\rho+1}} m_{S|R}(0^+;P)$, and $\nabla_{r^{\rho+2}} m_{S|R}(r;P)$ are bounded functions of $(r,P)$. \underline{\textit{Equation (ref):}} is shown to be $o_{\mathcal{P}}(1)$. Define \begin{align} \boldsymbol{\mathrm{\gamma}}^* & = \left[ \begin{array}{c} \kappa_{\rho+1,1}^{+} \\ \vdots \\ \kappa_{2\rho+1,1}^{+} \end{array} \right], \\ B(P) & = \frac{1}{ (\rho+1)! } \nabla_{r^{\rho+1}} m_{S|R}(0^+;P) {\boldsymbol{\mathrm{e}}_1}' \boldsymbol{\mathrm{\Gamma}}^{-1} \boldsymbol{\mathrm{\gamma}}^*. \end{align} The term $B(P)$ is consistently estimated by \begin{align} & \widehat{B} = B_{m,n}(\boldsymbol{\mathrm{V}}_1 , \ldots, \boldsymbol{\mathrm{V}}_m) \doteq \frac{1}{ (\rho+1)! } \widehat{\nabla_{r^{\rho+1}} m_{S|R}}(0^+;P) {\boldsymbol{\mathrm{e}}_1}' \boldsymbol{\mathrm{\Gamma}}^{-1} \boldsymbol{\mathrm{\gamma}}^*, \end{align} that is, by replacing $\nabla_{r^{\rho+1}} m_{S|R}(0^+;P)$ in $B(P)$ by a consistent nonparametric estimator (condition (vii) of Proposition (ref)). First, we show that \begin{align*} & \frac{1}{h^{\rho+1}} \mathbb{E}_P \left[ \frac{1}{ m h } \sum_{i=1}^m K\left( \frac{ R_i }{h } \right) \widetilde{\boldsymbol{\mathrm{H}}}_{i}^{\rho} \bar{\varepsilon}_i^{\rho} \right] \\ &= \frac{f_R(0^+;P) }{ (\rho+1)! } \nabla_{r^{\rho+1}} m_{S|R}(0^+;P) \boldsymbol{\mathrm{\gamma}}^* + o_{\mathcal{P}}(1). \end{align*} It suffices to show that, for any $l=0, \ldots, \rho$, \begin{align*} & \frac{1}{h^{\rho+1}} \mathbb{E}_P \left[ \frac{1}{ m h } \sum_{i=1}^m K\left( \frac{ R_i }{h } \right) \left( \frac{ R_i }{h } \right)^l \bar{\varepsilon}_i^{\rho} \right] \\ &= \frac{f_R(0^+;P) }{ (\rho+1)! } \nabla_{r^{\rho+1}} m_{S|R}(0^+;P) \kappa_{\rho+1+l,1}^{+} + o_{\mathcal{P}}(1). \end{align*} To see that, \begin{align*} & \frac{1}{h^{\rho+1}} \mathbb{E}_P \left[ \frac{1}{ m h } \sum_{i=1}^m K\left( \frac{ R_i }{h } \right) \left( \frac{ R_i }{h } \right)^l \bar{\varepsilon}_i^{\rho} \right] \\ & = \frac{1}{h^{\rho+1}} \mathbb{E}_P \left\{ \frac{1}{ m h } \sum_{i=1}^m K\left( \frac{ R_i }{h } \right) \left( \frac{ R_i }{h } \right)^l \mathbb{E}_P\left[ \bar{\varepsilon}_i^{\rho} | R_i \right] \right\} \\ &= \frac{1}{h^{\rho+2} } \mathbb{E}_P \left\{ K \left( \frac{ R_i }{h } \right) \left( \frac{R_i}{h} \right)^{l} \left( \mathbb{E}_P [\bar{\varepsilon}_i^{\rho+1} |R_i] + R_i^{\rho+1} \nabla_{r^{\rho+1}} m_{S|R}(0^+;P)/(\rho+1)! \right) \right\} \\ &= \frac{1}{h^{\rho+1} } \bigintssss_0^{\infty} K \left( u \right) u^{l} \mathbb{E}_P [\bar{\varepsilon}_i^{\rho+1} |R_i=uh] f_R(uh;P) du \\ & + \frac{ \nabla_{r^{\rho+1}} m_{S|R}(0^+;P) }{ (\rho+1)! } \bigintssss_0^{\infty} K \left( u \right) u ^{l + \rho + 1} f_R(uh;P) du \\ &= \frac{1}{h^{\rho+1} } \bigintssss_0^{\infty} K \left( u \right) u^{l} \left[ \nabla_{r^{\rho+2}} m_{S|R}(r^*_u;P) (uh)^{\rho+2}/(\rho+2)! \right] f_R(uh;P) du \\ & + \frac{ \nabla_{r^{\rho+1}} m_{S|R}(0^+;P) }{ (\rho+1)! } \bigintssss_0^{\infty} K \left( u \right) u ^{l + \rho + 1} f_R(uh;P) du \\ &= \frac{h}{ (\rho+2)! } \bigintssss_0^{\infty} K \left( u \right) u^{l+\rho+2} \left[ \nabla_{r^{\rho+2}} m_{S|R}(r^*_u;P) \right] f_R(uh;P) du \\ & + \frac{ \nabla_{r^{\rho+1}} m_{S|R}(0^+;P) }{ (\rho+1)! } \bigintssss_0^{\infty} K \left( u \right) u ^{l + \rho + 1} f_R(uh;P) du \\ &= \frac{ \nabla_{r^{\rho+1}} m_{S|R}(0^+;P) }{ (\rho+1)! } \bigintssss_0^{\infty} K \left( u \right) u ^{l + \rho + 1} \left[ f_R(0^+;P) + \nabla_r f_R(r^*_u;P) uh \right] du \\ & +O_{\mathcal{P}} \left( h \right) \\ & = \frac{ \nabla_{r^{\rho+1}} m_{S|R}(0^+;P) f_R(0^+;P) }{ (\rho+1)! } \bigintssss_0^{\infty} K \left( u \right) u ^{l + \rho + 1} du \\ & - h \frac{ \nabla_{r^{\rho+1}} m_{S|R}(0^+;P) }{ (\rho+1)! } \bigintssss_0^{\infty} K \left( u \right) u ^{l + \rho + 2} \nabla_r f_R(r^*_u;P) du \\ & + O_{\mathcal{P}} \left( h \right) \\ & = \frac{ \nabla_{r^{\rho+1}} m_{S|R}(0^+;P) f_R(0^+;P) }{ (\rho+1)! } \bigintssss_0^{\infty} K \left( u \right) u ^{l + \rho + 1} du \\ & +O_{\mathcal{P}} \left( h \right) +O_{\mathcal{P}} \left( h \right) \\ & = \frac{ \nabla_{r^{\rho+1}} m_{S|R}(0^+;P) f_R(0^+;P) }{ (\rho+1)! } \kappa_{\rho+1+l,1}^{+} +o_{\mathcal{P}} \left( 1 \right), \end{align*} where the first $O_{\mathcal{P}} \left( h \right)$ term is obtained from the following facts: (i) the first $2 \rho+2$ moments of the kernel are bounded; (ii) $\nabla_{r^{\rho+2}} m_{S|R}(r;P)$ is bounded as a function of $(r,P)$; and (iii) $f_R(r;P)$ is bounded as a function of $(r,P)$. The Taylor expansion of the PDF $f_R(r;P)$ and the second $O_{\mathcal{P}} \left( h \right)$ term are obtained from the following facts: (i) the first $2 \rho+2$ moments of the kernel are bounded; (ii) $\nabla_r f_R(r;P)$ is a well-defined and bounded function of $(r,P)$; and (iii) $\nabla_{r^{\rho+1}} m_{S|R}(0^+;P)$ is a well-defined and bounded function of $(r,P)$. Rewrite Equation (ref) as, \begin{align*} & {\boldsymbol{\mathrm{e}}_1}' \boldsymbol{\mathrm{G}}_m \mathbb{E}_P\left[ \frac{1}{\sqrt{m h }} \sum_{i=1}^m K\left( \frac{ R_i }{h } \right) \widetilde{\boldsymbol{\mathrm{H}}}_{i}^{\rho} \bar{\varepsilon}_i^{\rho} \right] - \sqrt{m h} h^{\rho+1} \widehat{B} \\ &= \sqrt{m h } h^{\rho+1} \left\{ {\boldsymbol{\mathrm{e}}_1}' \boldsymbol{\mathrm{G}}_m \frac{1}{h^{\rho+1}} \mathbb{E}_P\left[ \frac{1}{ m h } \sum_{i=1}^m K\left( \frac{ R_i }{h } \right) \widetilde{\boldsymbol{\mathrm{H}}}_{i}^{\rho} \bar{\varepsilon}_i^{\rho} \right] - \widehat{B} \right\} \\ &= \sqrt{m h } h^{\rho+1} \left\{ {\boldsymbol{\mathrm{e}}_1}' \left[ \boldsymbol{\mathrm{G}}(P) + o_{\mathcal{P}} \left( 1 \right) \right] \left[ \frac{f_R(0^+;P) }{ (\rho+1)! } \nabla_{r^{\rho+1}} m_{S|R}(0^+;P) \boldsymbol{\mathrm{\gamma}}^* + o_{\mathcal{P}}(1) \right] - \widehat{B} \right\} \\ &= \sqrt{m h } h^{\rho+1} \left\{ B(P) + o_{\mathcal{P}}(1) - \widehat{B} \right\} = O(1) o_{\mathcal{P}}(1) = o_{\mathcal{P}}(1), \end{align*} where we use the facts that: $\boldsymbol{\mathrm{G}}_m = \boldsymbol{\mathrm{G}}(P) + o_{\mathcal{P}}(1)$, $\boldsymbol{\mathrm{G}}(P)=O_{\mathcal{P}}(1)$, \\ ${f_R(0^+;P) } \nabla_{r^{\rho+1}} m_{S|R}(0^+;P) \boldsymbol{\mathrm{\gamma}}^* /{ (\rho+1)! } = O_{\mathcal{P}}(1)$, $\widehat{B}= B(P) + o_{\mathcal{P}}(1)$, and $\sqrt{mh}h^{\rho+1} = O(1)$. • \underline{Assumption (ref) - (ref):} zero mean of influence function. $\mathbb{E}_P[ \psi_{n}(\boldsymbol{\mathrm{V}}_{i}, P) ]=0 ~~\forall P$ by construction. • \underline{Assumption (ref) - (ref):} variance of influence function. Recall that $\kappa_{s,t}^{+} = \int_{0}^{\infty} u^s K^t(u) ~ du$ and define \begin{align*} \boldsymbol{\mathrm{\Delta}} = & \left[ \begin{array}{ccc} \kappa_{0,2}^{+} & \ldots & \kappa_{\rho,2}^{+} \\ \vdots & \vdots & \vdots \\ \kappa_{\rho,2}^{+} & \ldots & \kappa_{2 \rho,2}^{+} \end{array} \right], \\ \xi^2(P) = & {\boldsymbol{\mathrm{e}}_1}' \boldsymbol{\mathrm{\Gamma}}^{-1} \boldsymbol{\mathrm{\Delta}} \boldsymbol{\mathrm{\Gamma}}^{-1} {\boldsymbol{\mathrm{e}}_1} \frac{ v_{S|R}(0^+;P) }{f_R(0^+;P)} \\ = & {\boldsymbol{\mathrm{e}}_1}' \boldsymbol{\mathrm{G}}(P) \boldsymbol{\mathrm{\Delta}} \boldsymbol{\mathrm{G}}(P) {\boldsymbol{\mathrm{e}}_1} v_{S|R}(0^+;P) f_R(0^+;P). \end{align*} We need to show that $ \mathbb{V}_P[ \psi_n(\boldsymbol{\mathrm{V}} , {P} ) ] - \xi^2(P) = o_{\mathcal{P}}(1).$ Note that, \begin{align*} & \mathbb{V}_P[ \psi_n(\boldsymbol{\mathrm{V}} , {P} ) ] = {\boldsymbol{\mathrm{e}}_1}' \boldsymbol{\mathrm{G}}(P) \left\{ \frac{1}{h} \mathbb{E}_P\left[ K^2\left( \frac{R}{h} \right) \widetilde{\boldsymbol{\mathrm{H}}}_{i}^{\rho} \widetilde{\boldsymbol{\mathrm{H}}}_{i}^{\rho'} v_{S|R}(R;P), \right] \right\} \boldsymbol{\mathrm{G}}(P) {\boldsymbol{\mathrm{e}}_1}. \end{align*} Thus, it suffices to show that \begin{align*} & \frac{1}{h} \mathbb{E}_P\left[ K^2\left( \frac{R}{h} \right) \widetilde{\boldsymbol{\mathrm{H}}}_{i}^{\rho} \widetilde{\boldsymbol{\mathrm{H}}}_{i}^{\rho'} v_{S|R}(R;P), \right] = \boldsymbol{\mathrm{\Delta}} v_{S|R}(0^+;P) f_R(0^+;P) + o_{\mathcal{P}}(1), \end{align*} or equivalently to show that \begin{align*} & \frac{1}{h} \mathbb{E}_P\left[ K^2\left( \frac{R}{h} \right) \left( \frac{R}{h} \right)^{j} \left( \frac{R}{h} \right)^{l} v_{S|R}(R;P), \right] = \kappa_{j+l,2}^{+} v_{S|R}(0^+;P) f_R(0^+;P) + o_{\mathcal{P}}(1), \end{align*} for every $j,l=0,\ldots, \rho$. In fact, \begin{align*} &\frac{1}{h} \mathbb{E}_P\left[ K^2\left( \frac{R}{h} \right) \left( \frac{R}{h} \right)^{j} \left( \frac{R}{h} \right)^{l} v_{S|R}(R;P) \right] - \kappa_{j+l,2}^{+} v_{S|R}(0^+;P) f_R(0^+;P) \\ =& \int_0^{\infty} K^2(u) u^{j+l} \left\{ \underbrace{ v_{S | R}(uh;P) f_{R}(uh;P) }_{\doteq g(uh;P) } - \underbrace{ v_{S | R}(0^+;P) f_{R}(0^+;P)}_{= g(0^+;P) } \right\} du \\ =& h \int_0^{\infty} K^2(u) u^{j+l} \nabla_r g(x^*_{uh};P) du = o_{\mathcal{P}}(1), \end{align*} where we use the fact that $g(r;P)$ is differentiable wrt $r$ with derivative $\nabla_r g(r;P)$ bounded over $(r,P)$, which is implied by bounded derivatives of $v_{S|R}(r;P)$ and $f_R(r;P)$ wrt $r$. • \underline{Assumption (ref) - (ref):} $ \sup_P \mathbb{E} [\psi_n^2(\boldsymbol{\mathrm{Z}}_{k,i},P)] < \infty .$ We have, \begin{align*} & \psi_n^2(\boldsymbol{\mathrm{Z}}_k,P) = \frac{1}{ h f_R^2(0^+;P) } K^2 \left( \frac{ X_k }{h } \right) \left( {\boldsymbol{\mathrm{e}}_1}' \boldsymbol{\mathrm{\Gamma}}^{-1} \widetilde{\boldsymbol{\mathrm{H}}}_{k}^{\rho} \right)^2 \left(Y_k - m_{S|R}(X_k;P) \right), \end{align*} where $\widetilde{\boldsymbol{\mathrm{H}}}_{k}^{\rho} = \left[ 1 ~~ \left( \frac{X_k}{h} \right) ~~ \cdots ~~ \left( \frac{X_k }{h } \right)^{\rho } \right]'$, which is a $ (\rho+1 \times 1)$ column vector. It follows that \begin{align*} & \psi_n^2(\boldsymbol{\mathrm{Z}}_k,P) = \frac{1}{ h f_R^2(0^+;P) } K^2 \left( \frac{ X_k }{h } \right) \left( {\boldsymbol{\mathrm{e}}_1}' \boldsymbol{\mathrm{\Gamma}}^{-1} \widetilde{\boldsymbol{\mathrm{H}}}_{k}^{\rho} \right)^2 \\* & \left[ \left(Y_{k} - m_{Y_k|X_k}(X_k) \right) + \left(m_{Y_k|X_k}(X_k) - m_{S|R}(X_k;P) \right) \right]^2 . \\ &\mathbb{E}\left[ \psi_n^2(\boldsymbol{\mathrm{Z}}_k,P) \right] = \frac{1}{ h f_R^2(0^+;P) } \mathbb{E}\left\{ K^2 \left( \frac{ X_k }{h } \right) \left( {\boldsymbol{\mathrm{e}}_1}' \boldsymbol{\mathrm{\Gamma}}^{-1} \widetilde{\boldsymbol{\mathrm{H}}}_{k}^{\rho} \right)^2 \right. \\ & \left. \left[ v_{Y_k|X_k}(X_k) + \left(m_{Y_k|X_k}(X_k) - m_{S|R}(X_k;P) \right)^2 \right] \right\} \\ &= \frac{1}{ f_R^2(0^+;P) } \int_0^{\infty} K^2(u) \left( {\boldsymbol{\mathrm{e}}_1}' \boldsymbol{\mathrm{\Gamma}}^{-1} \widetilde{\boldsymbol{\mathrm{H}}}^{\rho}(u) \right)^2 \\ & \left[ v_{Y_k|X_k}(uh) + \left(m_{Y_k|X_k}(uh) - m_{S|R}(uh;P) \right)^2 \right] f_{X_k}(uh) du \\ &= O_{\mathcal{P}}(1). \end{align*} where $\widetilde{\boldsymbol{\mathrm{H}}}^{\rho}(u) = \left[ 1 ~~ \left( u \right) ~~ \cdots ~~ u^{\rho } \right]'$, which is a $ (\rho+1 \times 1)$ column vector. • \underline{Assumption (ref) - (ref):} $(2+\zeta)$-th moment condition. First, $\mathbb{V}_P^{-1}\left( \psi_n(\boldsymbol{\mathrm{V}}_i,P) \right) = O_{\mathcal{P}}(1). $ Second, there exists $\zeta>0$ such that $n^{-\zeta/2} \mathbb{E}_P \left| \psi_n(\boldsymbol{\mathrm{V}}_i,P) \right| ^{2+\zeta} = o_{\mathcal{P}}(1)$ because $\eta(r;P) = O_{\mathcal{P}}(1)$ by assumption. Therefore, \begin{align*} n^{-\zeta/2} \sup\limits_{P \in \mathcal{P} } \mathbb{E}_P\left| \frac{ \psi_n(\boldsymbol{\mathrm{V}}_i,P) }{ \sqrt{\mathbb{V}_P\left( \psi_n(\boldsymbol{\mathrm{V}}_i,P) \right)} } \right|^{2 + \zeta} = o(1). \end{align*} • \underline{Assumption (ref) - (ref):} $\xi^2\left( \frac{m}{n} P_1 + \frac{n-m}{n} P_2 \right) \to \xi^2 \left( \gamma P_1 + (1-\gamma) P_2 \right)$. Let $\overline{P}_n = \frac{m}{n} P_1 + \frac{n-m}{n} P_2$ and $\overline{P} = \gamma P_1 + (1-\gamma ) P_2$. We have that $v_{ S | R }(0^+; \bar{P}_n ) \to v_{ S | R }(0^+; \bar{P})$ and $f_{R}(0^+;\bar{P}_n) \to f_{R}(0^+;\bar{P})$ as in Section (ref). This implies that $\xi^2(\overline{P}_n) \to \xi^2(\overline{P})$. • \underline{Assumption (ref) :} $ (n_1 /n - \lambda) \overset{p}{\to} 0$. In this setting, $n_1/n = \sum_i \mathbb{I}\{ X_i \geq 0 \} /n$ and $(n_1 /n - \lambda) \overset{p}{\to} 0$ holds. Assumption (ref) has already been verified above for any sequence $m_n$ such that $(m/n -\gamma) \to 0$ for arbitrary $\gamma \in (0,1)$. In particular it holds for $\gamma \in \{ \lambda, 1-\lambda\}.$ \end{enumerate} $\square$ \section{Monte Carlo Simulations} This section contains additional figures and tables regarding the Monte Carlo simulations in the main text (Section (ref)). \begin{figure}[H] \caption{Conditional Mean Functions} \begin{center} \end{center} \caption*{ Notes: Conditional mean functions of Design 1. The lines are equal and flat for $|X-0.5| \leq 0.3$. The conditional means of Design 2 equal those of Design 1 shifted upwards by $1$. } \end{figure} \begin{sidewaystable}[p] \caption{Simulated Rejection Rates - 5% Nominal Size} \begin{minipage}{18cm} \begin{tabular}{c c c c| c c c| c c c| c c c| c c c } \hline \hline \multicolumn{4}{c|}{\textbf{Design 1}} & \multicolumn{3}{c|}{$h=0.1$} & \multicolumn{3}{c|}{$h=0.3$} & \multicolumn{3}{c|}{$h=0.5$} & \multicolumn{3}{c}{$\widehat{h}_{mse}$} \\ $\sigma_1^2$ & $\sigma_2^2$ & $n_1$ & $n_2$ & SP & SB & SS & SP & SB & SS & SP & SB & SS & SP & SB & SS \\ \hline 1 & 1 & 100 & 1900 & 0.049 & 0.100 & 0.137 & 0.046 & 0.061 & 0.091 & 0.147 & 0.140 & 0.244 & 0.045 & 0.067 & 0.090 \\ 1 & 1 & 250 & 4750 & 0.047 & 0.067 & 0.062 & 0.048 & 0.054 & 0.062 & 0.321 & 0.277 & 0.376 & 0.046 & 0.057 & 0.056 \\ 1 & 1 & 500 & 9500 & 0.048 & 0.057 & 0.043 & 0.048 & 0.053 & 0.053 & 0.569 & 0.495 & 0.606 & 0.048 & 0.055 & 0.048 \\ \hline 1 & 1 & 40 & 1960 & 0.050 & 0.207 & 0.346 & 0.051 & 0.093 & 0.253 & 0.082 & 0.108 & 0.331 & 0.051 & 0.105 & 0.258 \\ 1 & 1 & 100 & 4900 & 0.052 & 0.104 & 0.146 & 0.048 & 0.064 & 0.097 & 0.153 & 0.148 & 0.250 & 0.050 & 0.078 & 0.097 \\ 1 & 1 & 200 & 9800 & 0.047 & 0.076 & 0.075 & 0.046 & 0.053 & 0.065 & 0.266 & 0.234 & 0.332 & 0.045 & 0.063 & 0.060 \\ \hline 5 & 1 & 100 & 1900 & 0.068 & 0.104 & 0.146 & 0.056 & 0.063 & 0.094 & 0.073 & 0.073 & 0.137 & 0.056 & 0.065 & 0.095 \\ 5 & 1 & 250 & 4750 & 0.059 & 0.071 & 0.056 & 0.053 & 0.053 & 0.060 & 0.104 & 0.097 & 0.134 & 0.053 & 0.052 & 0.056 \\ 5 & 1 & 500 & 9500 & 0.054 & 0.058 & 0.037 & 0.051 & 0.053 & 0.051 & 0.164 & 0.147 & 0.183 & 0.052 & 0.050 & 0.046 \\ \hline 5 & 1 & 40 & 1960 & 0.060 & 0.221 & 0.403 & 0.057 & 0.095 & 0.308 & 0.060 & 0.082 & 0.330 & 0.057 & 0.103 & 0.308 \\ 5 & 1 & 100 & 4900 & 0.058 & 0.106 & 0.147 & 0.053 & 0.063 & 0.097 & 0.072 & 0.078 & 0.139 & 0.051 & 0.069 & 0.096 \\ 5 & 1 & 200 & 9800 & 0.054 & 0.077 & 0.066 & 0.047 & 0.052 & 0.060 & 0.089 & 0.087 & 0.127 & 0.047 & 0.055 & 0.056 \\ \hline \hline \end{tabular} \end{minipage} \begin{minipage}{18cm} \begin{tabular}{c c c c| c c c| c c c| c c c| c c c } \hline \hline \multicolumn{4}{c|}{\textbf{Design 2}} & \multicolumn{3}{c|}{$h=0.1$} & \multicolumn{3}{c|}{$h=0.3$} & \multicolumn{3}{c|}{$h=0.5$} & \multicolumn{3}{c}{$\widehat{h}_{mse}$} \\ $\sigma_1^2$ & $\sigma_2^2$ & $\mu$ & $n_1=n_2$ & SP & SB & SS & SP & SB & SS & SP & SB & SS & SP & SB & SS \\ \hline 1 & 1 & 10 & 75 & 0.051 & 0.057 & 0.189 & 0.049 & 0.053 & 0.138 & 0.082 & 0.086 & 0.206 & 0.039 & 0.044 & 0.128 \\ 1 & 1 & 10 & 150 & 0.051 & 0.048 & 0.096 & 0.050 & 0.053 & 0.088 & 0.132 & 0.131 & 0.205 & 0.045 & 0.050 & 0.084 \\ 1 & 1 & 10 & 1000 & 0.049 & 0.052 & 0.043 & 0.049 & 0.051 & 0.057 & 0.606 & 0.560 & 0.636 & 0.047 & 0.049 & 0.055 \\ \hline 1 & 1 & 1 & 75 & 0.053 & 0.483 & 0.535 & 0.050 & 0.366 & 0.451 & 0.070 & 0.363 & 0.457 & 0.045 & 0.336 & 0.429 \\ 1 & 1 & 1 & 150 & 0.051 & 0.414 & 0.448 & 0.050 & 0.279 & 0.372 & 0.091 & 0.321 & 0.421 & 0.050 & 0.266 & 0.360 \\ 1 & 1 & 1 & 1000 & 0.045 & 0.217 & 0.323 & 0.049 & 0.179 & 0.257 & 0.339 & 0.555 & 0.645 & 0.048 & 0.184 & 0.273 \\ \hline 5 & 1 & 10 & 75 & 0.060 & 0.079 & 0.196 & 0.051 & 0.047 & 0.135 & 0.061 & 0.058 & 0.159 & 0.044 & 0.040 & 0.129 \\ 5 & 1 & 10 & 150 & 0.058 & 0.049 & 0.098 & 0.054 & 0.051 & 0.084 & 0.076 & 0.072 & 0.123 & 0.045 & 0.044 & 0.076 \\ 5 & 1 & 10 & 1000 & 0.048 & 0.048 & 0.037 & 0.048 & 0.047 & 0.054 & 0.254 & 0.222 & 0.277 & 0.049 & 0.049 & 0.053 \\ \hline 5 & 1 & 1 & 75 & 0.066 & 0.432 & 0.467 & 0.058 & 0.274 & 0.347 & 0.070 & 0.241 & 0.331 & 0.054 & 0.228 & 0.309 \\ 5 & 1 & 1 & 150 & 0.060 & 0.330 & 0.357 & 0.053 & 0.189 & 0.255 & 0.080 & 0.184 & 0.257 & 0.051 & 0.157 & 0.220 \\ 5 & 1 & 1 & 1000 & 0.049 & 0.129 & 0.189 & 0.048 & 0.094 & 0.146 & 0.208 & 0.264 & 0.330 & 0.048 & 0.096 & 0.145 \\ \hline \hline \end{tabular} \end{minipage} \caption*{ Notes: The table displays simulated rejection rates under the null hypothesis for the studentized permutation (SP), studentized bootstrap (SB), and studentized subsample (SS) tests. Rows correspond to variations of Designs 1 and 2, as explained in the text. The bandwidth choice $h$ equals one of three fixed values (0.1, 0.3, and 0.5) or the IK data-driven MSE-optimal bandwidth $\widehat{h}_{mse}.$ } \end{sidewaystable} \begin{sidewaysfigure}[p] \begin{minipage}{0.49\textwidth} \caption{Design 1 - Simulated Power Curves} \begin{center} \begin{minipage}{.5 \textwidth} (a) { $\sigma_1^2=\sigma_2^2 = 1$ \\ $n_1 / (n_1+n_2) = 0.05$ \\ SP(---) and $t$(- -)} \end{minipage} \begin{minipage}{.5 \textwidth} (b) { $\sigma_1^2=\sigma_2^2 = 1$ \\ $n_1 / (n_1+n_2) = 0.02$ \\ SP(---) and $t$(- -)} \end{minipage} \end{center} \begin{center} \begin{minipage}{.5 \textwidth} (c) { $\sigma_1^2= \sigma_2^2 = 1$ \\ $n_1 / (n_1+n_2) = 0.05$ \\ SP(---), SB(- -), and SS($\cdot\hspace{-.8mm}\cdot\hspace{-.8mm}\cdot$)} \end{minipage} \begin{minipage}{.5 \textwidth} (d) { $\sigma_1^2= \sigma_2^2 = 1$ \\ $n_1 / (n_1+n_2) = 0.02$ \\ SP(---), SB(- -), and SS($\cdot\hspace{-.8mm}\cdot\hspace{-.8mm}\cdot$)} \end{minipage} \end{center} \caption*{ Notes: Panels (a)--(b) compare power of the studentized permutation test (SP) with the $t$-test ($t$). Panels (c)--(d) compare SP with the studentized bootstrap (SB) and studentized subsample (SS). Darker lines correspond to bigger sample sizes, i.e., Panels (a) and (c): black $(n_1,n_2) = (500,9500)$, dark gray $(n_1,n_2) = (250,4750)$, and light gray $(n_1,n_2) = (100,1900)$; Panels (b) and (d): black $(n_1,n_2) = (200,9800)$, dark gray $(n_1,n_2) = (100,4900)$, and light gray $(n_1,n_2) = (40,1960)$. The x-axis represents the difference $\theta(P_1)-\theta(P_2)$ and the y-axis, the simulated probability of rejection. Sizes of all tests are artificially adjusted such that the simulated rejection under the null is always equal to 5%. All estimates use the IK MSE-optimal bandwidth $\widehat{h}_{mse}$. } \end{minipage} \begin{minipage}{0.49\textwidth} \caption{Design 2 - Simulated Power Curves} \begin{center} \begin{minipage}{.5 \textwidth} (a) { $\sigma_1^2= \sigma_2^2 = 1$ \\ $\mu=10$ \\ SP(---) and $t$(- -)} \end{minipage} \begin{minipage}{.5 \textwidth} (b) { $\sigma_1^2=\sigma_2^2 = 1$ \\ $\mu=1$ \\ SP(---) and $t$(- -)} \end{minipage} \end{center} \begin{center} \begin{minipage}{.5 \textwidth} (c) { $\sigma_1^2= \sigma_2^2 = 1$ \\ $\mu=10$ \\ SP(---), SB(- -), and SS($\cdot\hspace{-.8mm}\cdot\hspace{-.8mm}\cdot$)} \end{minipage} \begin{minipage}{.5 \textwidth} (d) { $\sigma_1^2= \sigma_2^2 = 1$\\ $\mu=1$ \\ SP(---), SB(- -), and SS($\cdot\hspace{-.8mm}\cdot\hspace{-.8mm}\cdot$)} \end{minipage} \end{center} \caption*{ Notes: Panels (a)--(b) compare power of the studentized permutation test (SP) with the $t$-test ($t$). Panels (c)--(d) compare SP with the studentized bootstrap (SB) and studentized subsample (SS). Darker lines correspond to bigger sample sizes, i.e., black $(n_1,n_2) = (1000,1000)$, dark gray $(n_1,n_2) = (150,150)$, and light gray $(n_1,n_2) = (75,75)$. The x-axis represents the difference $\theta(P_1)-\theta(P_2)$ and the y-axis, the simulated probability of rejection. Sizes of all tests are artificially adjusted such that the simulated rejection under the null is always equal to 5%. All estimates use the IK MSE-optimal bandwidth $\widehat{h}_{mse}$. \\ } \end{minipage} \end{sidewaysfigure} \section{Auxiliary Lemmas} \begin{definition} Consider a sequence of measurable functions $X_n : \Omega \times \mathcal{P} \to \mathbb{R}^q$ for a probability space $(\Omega, \mathcal{B}, P)$, where $P$ belongs to a set of distributions $\mathcal{P}$. In other words, for any fixed $P_0 \in \mathcal{P}$, $X_n(P_0)$ is a random variable defined on $(\Omega, \mathcal{B}, P)$ to the Euclidean space, whose probability distribution depends on $n$ and $P \in \mathcal{P}$. We say $X_n$ is uniformly bounded in probability over $\mathcal{P}$ by a deterministic sequence $\alpha_n$ if, for every $\delta>0$, there exists a deterministic constant $M_{\delta} \in (0,\infty)$ such that \[ \sup\limits_{P \in \mathcal{P} } \mathbb{P}_P \left[ \left\| X_n(P) \right\| > M_{\delta} \alpha_n \right] < \delta . \] We denote this as $X_n = O_{\mathcal{P}}( \alpha_n )$. Similarly, we say $X_n = o_{\mathcal{P}}( \alpha_n )$ if, for every $\varepsilon,\delta>0$ $\exists n_{\varepsilon,\delta}$ such that \[ \sup\limits_{P \in \mathcal{P} } \mathbb{P}_P \left[ \left\| X_n(P) \right\| > \varepsilon \alpha_n \right] < \delta ~~~ \forall n \geq n_{\varepsilon,\delta}. \] \end{definition} \begin{lemma} Consider a deterministic sequence $\alpha_n$ and a sequence of bivariate random functions $(X_n,Y_n): \Omega \times \mathcal{P} \to \mathbb{R}^2$ as in Definition (ref). \begin{enumerate} • If $X_n = O_{\mathcal{P}}(\alpha_n)$ and $Y_n = o_{\mathcal{P}}(1)$, then $X_n Y_n = o_{\mathcal{P}}(\alpha_n)$. • Assume $X_n(P)$ has expectation $\mathbb{E}_P [X_n(P)]$ and variance $\mathbb{V}_P [X_n(P)]$ that are bounded over $\mathcal{P}$. Then, \[ X_n = O_{\mathcal{P}}\left( \sup\limits_{P \in \mathcal{P} } \mathbb{E}^2_P [X_n(P) ] + \sup\limits_{P \in \mathcal{P} } \mathbb{V}_P [ X_n(P)] \right). \] • Suppose there exists a deterministic function $y:\mathcal{P} \to [\underline{y},\overline{y}]$ with $0<\underline{y} < \overline{y} <\infty$ such that $Y_n - y = o_{\mathcal{P}}(1) $. Then, \[ \frac{1}{Y_n} - \frac{1}{y} = o_{\mathcal{P}}(1). \] \end{enumerate} \end{lemma} \begin{proof} \textbf{Part 1:} Fix $\delta>0$ and $\varepsilon>0$. There exists $M_{\delta} \in (0,\infty)$ such that \[ \sup\limits_{P \in \mathcal{P} } \mathbb{P}_P \left[ \left| X_n(P) \right| > M_{\delta} \alpha_n \right] < \frac{\delta}{2} . \] There exists $\exists n_{\varepsilon,\delta}$ such that \[ \sup\limits_{P \in \mathcal{P} } \mathbb{P}_P \left[ \left| Y_n(P) \right| > \frac{\varepsilon}{M_{\delta}} \right] < \frac{\delta}{2} ~~~ \forall n \geq n_{\varepsilon,\delta}. \] Given that $\left| X_n(P) Y_n(P) \right| > \varepsilon \alpha_n$ implies $\left| X_n(P) \right| > M_{\delta} \alpha_n$ or $\left|Y_n(P) \right| > \varepsilon / M_{\delta}$ we have \begin{align*} &\sup\limits_{P \in \mathcal{P} } \mathbb{P}_P \left[ \left| X_n(P) Y_n(P) \right| > \varepsilon \alpha_n \right] \\ &\leq \sup\limits_{P \in \mathcal{P} } \mathbb{P}_P \left[ \left| X_n(P) \right| > M_{\delta} \alpha_n \right] + \sup\limits_{P \in \mathcal{P} } \mathbb{P}_P \left[ \left|Y_n(P) \right| > \varepsilon / M_{\delta} \right] < \delta \forall n \geq n_{\varepsilon,\delta}. \end{align*} \textbf{Part 2:} Call $A_n = \sup\limits_{P \in \mathcal{P} } \mathbb{E}^2_P [X_n(P)] + \sup\limits_{P \in \mathcal{P} } \mathbb{V}_P [ X_n(P)] $ and $B_n= \sup\limits_{P \in \mathcal{P} }\mathbb{E}_P\left( X_n^2(P) \right)$. Note that $A_n \geq B_n \geq \mathbb{E}_P\left( X_n^2(P) \right)$ for every $P \in \mathcal{P}$. For $\delta>0$ and $P\in \mathcal{P}$, \begin{align*} &\mathbb{P}_P\left[ |X_n(P)| > \delta^{-1/2} A_n \right] \\ &\leq \mathbb{P}_P\left[ |X_n(P)| > \delta^{-1/2} B_n \right] \leq \mathbb{P}_P\left[ |X_n(P)| > \delta^{-1/2} \mathbb{E}_P\left( X_n^2(P) \right) \right] \leq \delta, \end{align*} where the last inequality is the Markov inequality. The result follows by taking the supremum of both sides. \textbf{Part 3:} Pick $\eta>0$ such that $\underline{y}-\eta>0$. Define the set $A = \{y: \underline{y}-\eta \leq y \leq \overline{y} +\eta \}$. The function $g(y)=1/y$ is uniformly continuous over $A$. Thus, for every $\varepsilon>0$, there exists $\gamma_{\varepsilon}>0$ such that $|g(y') - g(y)|>\varepsilon \Rightarrow |y'-y|>\gamma_{\varepsilon}$ for any $y',y \in A$. This implies that, for fixed $P$, \begin{gather*} \mathbb{P}_P\left[ \left| 1/Y_n(P) - 1/y(P) \right| > \varepsilon, Y_n(P) \in A \right] \leq \mathbb{P}_P\left[ \left| Y_n(P) - y(P) \right| > \gamma_\varepsilon \right] \to 0. \end{gather*} This right-hand side is uniform over $\mathcal{P}$ because $Y_n - y = o_{\mathcal{P}}(1)$. Next, for fixed $P$, $|Y_n(P)-y(P)| < \eta \Rightarrow \underline{y} -\eta < Y_n(P) < \overline{y} +\eta \Rightarrow Y_n(P) \in A$. This implies that $\mathbb{P}_P [|Y_n(P)-y(P)| < \eta ] \leq \mathbb{P}_P[Y_n(P) \in A] $ and that $\mathbb{P}_P[Y_n(P) \in A ] \to 1 $ uniformly over $\mathcal{P}$ because $Y_n - y = o_{\mathcal{P}}(1)$ implies that $\mathbb{P}_P [|Y_n(P)-y(P)| < \eta ] \to 1$ uniformly over $\mathcal{P}$. Then, \begin{align*} &\mathbb{P}_P\left[ \left| 1/Y_n(P) - 1/y(P) \right| > \varepsilon \right] - \mathbb{P}_P\left[ \left| 1/Y_n(P) - 1/y(P) \right| > \varepsilon, Y_n(P) \in A \right] \\ &\leq 1 - \mathbb{P}_P\left[ Y_n(P) \in A \right] \to 0 \end{align*} uniformly over $\mathcal{P}$. \end{proof} \begin{lemma} Consider a sequence of random variables $X_n$ in the Euclidean space. $X_n \overset{p}{\to} X$ if, and only if, for every subsequence $X_{n_{k}}$ there exists a further subsequence $X_{n_{k_{j}}}$ such that $X_{n_{k_{j}}} \overset{as}{\to} X $. \end{lemma} \begin{proof} See the proof of Theorem 2.3.2 by \citeSM{durrett2019}. \end{proof} \begin{lemma} Consider a sequence of random variables $(Z_n,X_n)$, $n=1,2, \ldots$, and a random variable $Z$, all with domain on the measure space $(\Omega, \mathcal{B}, \mu)$. The image of $Z_n$ and $Z$ are in a Euclidean space, and the image of $X_n $ is $\mathcal{X}_n$. For a measurable function $F_n : \mathcal{X}_n \to \mathbb{R}$, assume $F_n(X_n) \overset{p}{\to} 0$. Moreover, suppose that for any nonrandom sequence $x_n \in \mathcal{X}_n$ such that $F_n(x_n) \to 0$, $Z_n$ conditional on $X_n=x_n$ converges in distribution to $Z$. Then, $Z_n \overset{d}{\to} Z$ unconditionally. \end{lemma} \begin{proof} For an arbitrary subsequence $n_k$ of the sequence $\left\{ F_n(X_n )\right\}_n$, Lemma (ref) says there is a further subsequence $n_{k_j}$ such that $ F_{n_{k_j}}(X_{n_{k_j}} ) \overset{as}{\to} 0$ as $j\to\infty$. Define $\mathbb{X}^{\infty}$ to be the space of all subsequences of the form $\left\{ x_{n_{k_{j}}} \right\}_{j=1}^{\infty}$, for values $x_{n_{k_j}} \in \mathcal{X}_{n_{k_j}}$ that satisfy $F_{n_{k_j}}(x_{n_{k_j}} ) \to 0$ as $j\to\infty$. We know that $\mathbb{P}\left[ \{ X_{n_{k_j}} \}^{\infty}_{j=1} \in \mathbb{X}^\infty \right]=1$. Let $G_{n}$ be the CDF of the distribution of $ Z_{n} $ conditional on ${X}_{n} = {x}_{n}$, where $x_n$ is an arbitrary sequence that satisfies $F_n(x_n) \to 0$. By assumption, $G_n \to G$ pointwise (for every continuity point of $G$), where $G$ is the CDF of $Z$. This implies that $G_{n_{k_j}} \to G$ for the subsequence $\{n_{k_j}\}_j$ from above. Next, \begin{align*} \mathbb{P}\left[ Z_{n_{k_j}} \leq z \right] =& \int_{\Omega} \mathbb{P}\left[\left. Z_{n_{k_j}} \leq z \right| {X}_{n_{k_j}} = {X}_{n_{k_j}}(\omega) \right] d \mu(\omega) \\ =& \int_{\Omega} G_{n_{k_j}}(z;\omega) \mathbb{I}\left\{ \omega : \{ X_{n_{k_j}} (\omega) \}^{\infty}_{j=1} \in \mathbb{X}^\infty \right\} d \mu(\omega), \end{align*} where $G_{n_{k_j}}(z;\omega)$ denotes the CDF of $ Z_{n_{k_j}} $ conditional on ${X}_{n_{k_j}} = {X}_{n_{k_j}}(\omega)$, and we used that $\mathbb{P}\left[\omega : \{ X_{n_{k_j}} (\omega) \}^{\infty}_{j=1} \in \mathbb{X}^\infty \right]=1$. For each $\omega$ such that the indicator is $1$, the conditional CDF $G_{n_{k_j}}(z;\omega)$ converges to $G(z)$ as $j \to \infty$. So, we have an integral of a measurable function of $\omega$ that changes with $j$, is bounded by $1$, and converges pointwise in $\omega$ to $G(z) \mathbb{I}\left\{ \omega : \{ X_{n_{k_j}} (\omega) \}^{\infty}_{j=1} \in \mathbb{X}^\infty \right\} $ as $j \to \infty$. By the dominated convergence theorem, the integral above converges to $G(z) \mathbb{P}\left[ \{ X_{n_{k_j}} \}^{\infty}_{j=1} \in \mathbb{X}^\infty \right] =G(z).$ This says that $Z_{n_{k_j}} \overset{d}{\to} Z$. Finally, call $H_{n} $ the unconditional CDF of $Z_{n}$. We have just shown that, for every subsequence $\{ n_k \}_k$ there exists a further subsequence $\{ n_{k_j} \}_j$ such that $H_{n_{k_j}}(z) \to G(z) $. This implies that $H_{n}(z) \to G(z) $. Therefore, $Z_n \overset{d}{\to} Z$. \end{proof} The next lemma is a generalization of Lemma 11.3.3 of \citeSM{lehmann2005}. \begin{lemma} For each $n$, let $Y_{n,1}, \ldots, Y_{n,n}$ be independently identically distributed with mean zero and finite variance $\sigma^2_n$. Let $C_{n,1}, \ldots, C_{n,n}$ be a sequence of random variables, independent of $Y_{n,1}, \ldots, Y_{n,n}$. Assume $\exists \zeta>0$ such that \begin{align} & \mathbb{E}\left[ | Y_{n,i}/\sigma_n |^{2 + \zeta} \right] \left( \frac{\max_{i=1}^n C_{n,i}^2 }{\sum_{l=1}^n C_{n,l}^2} \right)^{\zeta/2} \overset{p}{\to} 0 \text{ as }n \to \infty. \end{align} Then \begin{gather} \frac{\sum_{i=1}^n C_{n,i} Y_{i,n} }{ \sigma_n \sqrt{ \sum_{l=1}^n C_{n,l}^2 } }\overset{d}{\to} N(0,1). \end{gather} \end{lemma} \begin{proof} We use Lemma (ref) to prove this lemma. Call $\alpha_n = \mathbb{E}\left[ | Y_{n,i}/\sigma_n |^{2 + \zeta} \right]$. Mapping the lemma's notation to our case, we have \begin{gather*} Z_n = \frac{ \sum_{i=1}^{n} C_{n,i} Y_{n,i} } {\sigma_n \sqrt{ \sum_{l=1}^{n } C_{n,l}^2 } }, \\ X_n =\left(C_{n,1}, \ldots, C_{n,n} \right) , \\ \mathcal{X}_n = \mathbb{R}^n, \\ F_n(X_n) = \alpha_n \left( \frac{\max_{1 \leq i \leq n} C_{n,i}^2 }{ \sum_{j=1}^n C_{n,j}^2} \right)^{\zeta/2}. \end{gather*} Then, it suffices to derive the limiting distribution of the sequence of distributions of \\ $\sum_{i=1}^{n} C_{n,i} Y_{n,i} / \sqrt{ \sum_{l=1}^{n } C_{n,l}^2 }$ conditional on $\left(C_{n,1}, \ldots, C_{n,n} \right) = \left(c_{n,1}, \ldots, c_{n,n} \right) $, where the values $\left(c_{n,1}, \ldots, c_{n,n} \right)$ come from an arbitrary triangular array with infinitely many rows that satisfies $\alpha_n \left( {\max_{i=1}^n c_{n,i}^2 }/{ \sum_{j=1}^n c_{n,j}^2} \right)^{\zeta/2}$ $\to 0 $ as $n \to \infty$. For each $n$, we have a sum of a triangular array of random variables $C_{n,i} Y_{n, i}$ that are independent across $i=1, \ldots, n$ once we condition on $\left(C_{n,1}, \ldots, C_{n,n} \right)$. By assumption, $C_{n,1}, \ldots, C_{n,n}$ is independent of $Y_{n,1}, \ldots, Y_{n,n}$ for every $n$. Thus, \begin{gather} \mathbb{E} \left[ C_{n,i} Y_{n, i} | \left(C_{n,1}, \ldots, C_{n,n} \right) = \left(c_{n,1}, \ldots, c_{n,n} \right) \right] = 0, \\ \mathbb{V}\left[ C_{n,i} Y_{n, i} | \left(C_{n,1}, \ldots, C_{n,n} \right) = \left(c_{n,1}, \ldots, c_{n,n} \right) \right] = c_{n,i}^2 \sigma^2_{n} \end{gather} for $i=1, \ldots, n$. The sum of the variances is $s_{n}^2 = \sigma^2_{n} \sum_{i=1}^{n} c_{n,i}^2 $. For $n=1,2, \ldots$, the sequence of distributions of $ \sum_{i=1}^{n} C_{n ,i} Y_{n , i} / \sigma_{n } \sqrt{ \sum_{l=1}^{n} C_{n,l}^2 }$ conditional on $\left(C_{n,1}, \ldots, C_{n,n} \right) = \left(c_{n,1}, \ldots, c_{n,n} \right) $ is the same as the sequence of distributions of $Z_{n} = \sum_{i=1}^{n} c_{n,i} Y_{n, i}/ \sigma_{n} \sqrt{ \sum_{l=1}^{n} c_{n,l}^2 }$. We apply the Lindeberg CLT to derive the limiting distribution of $Z_{n}$. We need to verify the Lindeberg condition, that is, for any $\delta>0$ \begin{gather} \sum_{i=1}^{n } \frac{1}{s_{n }^2} \mathbb{E}\left[c_{{n },i}^2 Y_{n, i}^2 \mathbb{I}\{| c_{{n},i} Y_{n, i} | > \delta s_{n} \} \right] \to 0 \text{ as } n \to \infty \\ \Leftrightarrow \notag \\ \sum_{i=1}^{n } \frac{ c_{{n},i}^2 }{\sigma_{n }^2 \sum_{l=1}^{n} c_{{n},l}^2} \mathbb{E}\left[ Y_{n, i}^2 \mathbb{I}\left\{ c_{{n},i}^2 Y_{n, i}^2 > \delta^2 \sigma_{n}^2 \sum_{l=1}^{n} c_{{n},l}^2 \right\} \right] \to 0 \\ \Leftarrow \notag \\ \sum_{i=1}^{n} \frac{ c_{{n},i}^2 }{ \sum_{l=1}^{n} c_{{n},l}^2} \mathbb{E}\left[ \frac{ Y_{n, i}^2 }{\sigma_{n}^2} \mathbb{I}\left\{ \frac{ Y_{n, i}^2 }{\sigma_{n}^2} > \delta^2 \frac{ \sum_{l=1}^{n} c_{{n},l}^2}{ \max\limits_{1 \leq i \leq n} c_{{n},i}^2} \right\} \right] \to 0 \\ \Leftrightarrow \notag \\ \mathbb{E}\left[ \frac{ Y_{n, i}^2 }{\sigma_{n}^2} \mathbb{I}\left\{ \frac{ Y_{n, i}^2 }{\sigma_{n}^2} > \delta^2 \frac{ \sum_{l=1}^{n } c_{{n },l}^2}{ \max\limits_{1 \leq i \leq n } c_{{n },i}^2} \right\} \right] \to 0, \end{gather} where we used that $\mathbb{I}\left\{ c_{{n },i}^2 Y_{n , i}^2 > \delta^2 \sigma_{n }^2 \sum_{l=1}^{n } c_{{n },l}^2 \right\} \leq \mathbb{I}\left\{ \frac{ Y_{n , i}^2 }{\sigma_{n }^2} > \delta^2 \frac{ \sum_{l=1}^{n } c_{{n },l}^2}{ \max\limits_{1 \leq i \leq n } c_{{n },i}^2} \right\} $ and that $\mathbb{E}\left[ \frac{ Y_{n , i}^2 }{\sigma_{n }^2} \mathbb{I}\left\{ \frac{ Y_{n , i}^2 }{\sigma_{n }^2} > \varepsilon^2 \frac{ \sum_{l=1}^{n } c_{{n },l}^2}{ \max\limits_{1 \leq i \leq n } c_{{n },i}^2} \right\} \right]$ does not depend on $i$. Thus, it suffices to verify Equation (ref). Note that, \begin{align*} \left| \frac{ Y_{n, i} }{\sigma_{n} } \right|^{2+\zeta} & \geq \left| \frac{ Y_{n, i} }{\sigma_{n} } \right|^{2+\zeta} \mathbb{I}\left\{ \frac{ Y_{n, i}^2 }{\sigma_{n}^2} > \delta^2 \frac{ \sum_{l=1}^{n } c_{{n },l}^2}{ \max\limits_{1 \leq i \leq n } c_{{n },i}^2} \right\} \\ & = \frac{ Y_{n, i}^2 }{\sigma_{n}^2 } \left| \frac{ Y_{n, i} }{\sigma_{n} } \right|^{\zeta} \mathbb{I}\left\{ \frac{ Y_{n, i}^2 }{\sigma_{n}^2} > \delta^2 \frac{ \sum_{l=1}^{n } c_{{n },l}^2}{ \max\limits_{1 \leq i \leq n } c_{{n },i}^2} \right\} \\ & \geq \frac{ Y_{n, i}^2 }{\sigma_{n}^2 } \left( \delta^2 \frac{ \sum_{l=1}^{n } c_{{n },l}^2}{ \max\limits_{1 \leq i \leq n } c_{{n },i}^2} \right)^{\zeta/2} \mathbb{I}\left\{ \frac{ Y_{n, i}^2 }{\sigma_{n}^2} > \delta^2 \frac{ \sum_{l=1}^{n } c_{{n },l}^2}{ \max\limits_{1 \leq i \leq n } c_{{n },i}^2} \right\}. \end{align*} Re-arranging, \begin{align*} \left( \frac{ \max\limits_{1 \leq i \leq n } c_{{n },i}^2} { \delta^2 \sum_{l=1}^{n } c_{{n },l}^2} \right)^{\zeta/2} \mathbb{E} \left| \frac{ Y_{n, i} }{\sigma_{n} } \right|^{2+\zeta} & \geq \mathbb{E}\left[ \frac{ Y_{n, i}^2 }{\sigma_{n}^2 } \mathbb{I}\left\{ \frac{ Y_{n, i}^2 }{\sigma_{n}^2} > \delta^2 \frac{ \sum_{l=1}^{n } c_{{n },l}^2}{ \max\limits_{1 \leq i \leq n } c_{{n },i}^2} \right\} \right], \\ \delta^{-\zeta} \left( \frac{ \max\limits_{1 \leq i \leq n } c_{{n },i}^2} { \sum_{l=1}^{n } c_{{n },l}^2} \right)^{\zeta/2} \alpha_n & \geq \mathbb{E}\left[ \frac{ Y_{n, i}^2 }{\sigma_{n}^2 } \mathbb{I}\left\{ \frac{ Y_{n, i}^2 }{\sigma_{n}^2} > \delta^2 \frac{ \sum_{l=1}^{n } c_{{n },l}^2}{ \max\limits_{1 \leq i \leq n } c_{{n },i}^2} \right\} \right] \end{align*} and the left-hand side converges to zero, which implies the right-hand side converges to zero and (ref) holds. Therefore, \begin{gather*} \frac{ \sum_{i=1}^{n } c_{n ,i} Y_{n , i} }{ \sigma_{n } \sqrt{ \sum_{l=1}^{n } c_{n ,l}^2 } } \overset{d}{\to} N(0,1) \end{gather*} and Lemma (ref) implies that \begin{gather*} \frac{ \sum_{i=1}^{n } C_{n ,i} Y_{n , i} }{ \sigma_{n } \sqrt{ \sum_{l=1}^{n } C_{n ,l}^2 } } \overset{d}{\to} N(0,1). \end{gather*} \end{proof} \begin{lemma} Consider a sequence of random CDFs $\{ F_n \}_n$ that converges pointwise in probability to a continuous CDF $F$, that is, for every $x \in \mathbb{R}$, $F_n(x) \overset{p}{\to} F(x)$. Then, the convergence is also uniform, namely, \[ \sup_x | F_n(x) - F(x) | \overset{p}{\to} 0. \] \end{lemma} \begin{proof} Fix $\varepsilon>0$ and choose $m: \varepsilon>1/m$. Use continuity of $F$ to pick points $-\infty = x_0 < x_1 < \ldots < x_{m-1}<x_m=\infty$ such that $F(x_j)-F(x_{j-1})=1/m$, for $j=1, \ldots, m$. For any $x \in \mathbb{R}$, $\exists j$ such that $x_{j-1}\leq x \leq x_{j}$. Use monotonicity of $F$ and $F_n$ to show \begin{align*} F_n(x) - F(x) & \leq F_n(x_j) - F(x_{j-1}) = F_n(x_j) - F(x_{j}) + 1/m, \\ F_n(x) - F(x) & \geq F_n(x_{j-1}) - F(x_{j}) = F_n(x_{j-1}) - F(x_{j-1}) - 1/m, \\ |F_n(x) - F(x) | & \leq \max_j | F_n(x_{j}) - F(x_{j})| + 1/m. \end{align*} Therefore, \[ \mathbb{P}\left[ \sup_x |F_n(x) - F(x) | > \varepsilon\right] \leq \mathbb{P}\left[ \max_j |F_n(x_j) - F(x_j) | > \varepsilon - 1/m \right] \] \[ \leq \mathbb{P}\left[ \bigcup_j |F_n(x_j) - F(x_j) | > \varepsilon - 1/m \right] \] \[ \leq \sum_j \mathbb{P}\left[ |F_n(x_j) - F(x_j) | > \varepsilon - 1/m \right] \to 0 . \] \end{proof} \end{singlespace} \begin{onehalfspace} \bibliographystyleSM{econ} \bibliographySM{biblio} \end{onehalfspace}