EconBase
← Back to paper

Asymptotic Properties of Empirical Quantile-Based Estimators

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

46,439 characters · 0 sections · 27 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Asymptotic Properties of Empirical Quantile-Based Estimators

abstractWe consider inference for parameters of the form $\theta_0 = E[F_Y^{-1}\circ F_Z(X)]$ for some variables $X$, $Y$ and $Z$. Such parameters appear, in particular, in the “changes-in-changes” model of AtheyImbens2006. We first establish that $\widehat{\theta}$, a plug-in estimator of $\theta_0$, is root-$n$ consistent and asymptotically normal under weaker conditions than those previously available, allowing in particular for unbounded variables. Next, we propose a new estimator of the asymptotic variance of $\widehat{\theta}$ and show its consistency, also allowing for unbounded variables. Monte Carlo simulations suggest that the conditions for root-$n$ consistency and asymptotic normality are, in some sense, minimal. These simulations highlight that our variance estimator also leads to more accurate inference than some alternative approaches.

\noindentJEL Classification: C14, C21, C23.

\noindentKeywords: changes-in-changes, asymptotic inference, panel data.

\setcounter{page}{1} {0.5\baselineskip}

\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{Introduction}

Quantile-quantile transforms, namely objects of the kind $F^{-1}\circ G$ where $F$ and $G$ are two cumulative distribution functions, appear commonly in economics. In particular, they have been used to recover distributions of unobserved potential outcomes. Prominent examples include the “changes-in-changes” (CIC) causal inference model developed by AtheyImbens2006, and nonparametric instrumental variable quantile regression, see in particular vuong2017counterfactual and wuthrich2020comparison. In this setup, average treatment effects involve estimands of the form $\theta_0 = E[F_Y^{-1}\circ F_Z(X)]$ for some variables $X$, $Y$ and $Z$. The aim of this paper is to study inference for such parameters.

AtheyImbens2006 show that under suitable conditions, a plug-in estimator $\widehat{\theta}$ of $\theta_0$ is asymptotically normal, and establish the consistency of an estimator of its asymptotic variance. However, their results rely on strong assumptions. Specifically, they assume that the three variables (a) have bounded support, (b) each admit a continuously differentiable density, and (c) that these densities are bounded from above and below on their support. Such assumptions are overly restrictive for many variables of interest including, for instance, wages, prices or profits.

The goal of this paper is to obtain similar results under substantially weaker conditions. This is important for establishing that such methods remain applicable to key economic variables that may not satisfy Assumptions (a)--(c) above. To this end, we first establish asymptotic normality of $\widehat{\theta}$. The main difficulty is that standard tools are no longer applicable under these weaker conditions. When variables are bounded and their densities are bounded from below, the functional $(F_Y,F_X,F_Z)\mapsto \int F_Y^{-1}\circ F_Z dF_X$ is Hadamard differentiable CdCXdH2017. However, Hadamard differentiability fails otherwise, already because $F\mapsto \int gdF$ is not continuous with respect to the supremum norm when $g$ is unbounded. Similarly, we cannot directly exploit results on L-statistics ShorackWellner1986, as these correspond to the simpler case for which both $F_X$ and $F_Z$ are known. Instead, we rely on several results on weighted and unweighted empirical and quantile processes, see in particular Chapter 2, Section 7 and Chapter 11 in ShorackWellner1986 and csorgo1986. We also exploit auxiliary results, including (i) the fact that order statistics of uniforms and uniforms spacing follow beta distributions; (ii) known bounds for the mean absolute deviation of beta distributions.

Another contribution of this paper is to establish that in a sense that we make precise below, some of the conditions we impose for root-$n$ consistency and asymptotic normality are necessary as well. This implies fundamental constraints on the scope of methods relying on quantile-quantile transforms, at least if inference is based upon asymptotic normality. In the changes-in-changes model, for instance, this means that for the average treatment effect to be root-$n$ consistent and asymptotically normal, the distributions of the pre-treatment period outcome of the control and treatment groups must exhibit sufficiently similar tail behavior. These conditions are stronger than what is needed for root-$n$ consistency and asymptotic normality of quantile treatment effects. Intuitively, this is because unlike (non-extremal) quantile treatment effects, the average treatment effect depends on the tails of potential outcomes, whose corresponding quantiles are less precisely estimated.

Our second main contribution is to propose a new estimator of $\sigma^2$, the asymptotic variance of $\widehat{\theta}$. The plug-in estimator of AtheyImbens2006 includes a density term in its denominator. As a result, the consistency of this estimator becomes unclear when the density takes arbitrarily small values. A possible solution would be to trim the estimator, but this would introduce additional tuning parameters.

Instead, we consider an alternative estimator $\widehat{\sigma}^2$ based on a new expression of the asymptotic variance, still involving a density $f_U$ (that of $U:=F_Z(X)$) but without any denominator term. We consider a kernel density estimator of $f_U(u)$ with a varying bandwidth proportional to $u(1-u)$. Thus, the bandwidth shrinks as $u\to 0$ or $u\to 1$, a feature that is key to handling a possible explosion of $f_U(u)$ as $u\to 0$ or $u\to 1$. We show consistency of $\widehat{\sigma}^2$ under a slight strengthening of the conditions we impose for asymptotic normality. Notably, our proof does not require uniform or even pointwise consistency of our kernel-density estimator. Our estimator may be of interest for estimating functionals of probability densities, beyond the particular functional we consider here.

Finally, we investigate the finite-sample behavior of $\widehat{\theta}$ and inference based on asymptotic normality and $\widehat{\sigma}$ through Monte Carlo simulations. Our results suggest in particular that when our conditions for asymptotic normality hold, our inference method is already accurate with sample sizes around 100. Our estimator $\widehat{\sigma}$ also seems to perform better than that originally proposed by AtheyImbens2006. Finally, when our conditions for asymptotic normality are violated, the distribution of $\widehat{\theta}$ does not appear to be normal, and none of the inference methods we consider, including the bootstrap, performs well.

\paragraph{Related literature.} In their seminal work, AtheyImbens2006 derive the asymptotic normality of $\widehat{\theta}$ and propose a consistent variance estimator. As discussed above, our main contribution is to extend these results to allow for unbounded variables. We also delineate some restrictions that the variable distributions should satisfy for the asymptotic variance to exist.

sun2025debiased establish asymptotic normality of a debiased and semiparametrically efficient changes-in-changes estimator that flexibly accommodates continuous covariates. While they accommodate covariates, their result holds under a high-level condition (see their Assumption 4(a)). Our paper shows that establishing weak and low-level conditions for asymptotic normality is nontrivial, even in the absence of covariates.

In a concurrent and independent line of research, beare2026convergencedistributionppprocess establish the convergence in distribution in $L^1([0,1])$ of a process of the form $\widehat{F} \circ \widehat{G}^{-1}$ under similar conditions to those assumed in the present paper. However, we consider here the asymptotic normality of a quantity of a different kind, namely $\int_0^1 [\widehat{F}_Y^{-1}\circ \widehat{F}_Z] d\widehat{F}_X$, which does not seem to be a direct consequence of beare2026convergencedistributionppprocess, as it involves an additional source of randomness via $\widehat{F}_X$.

Our estimator of the asymptotic variance relies on nonparametric kernel density estimates with a varying bandwidth. Such estimators have been studied in mathematical statistics Jones1990, TerrellScott1992,chhor2025local, but we use this technique in a different context, where the focus is not the density itself but a functional of it. Our contribution is to show that such estimators lead to consistent estimation of the functional of interest under mild smoothness conditions, allowing in particular for the density to diverge at the boundaries. This is achieved by letting the bandwidth shrink appropriately near the boundary of the support.

\paragraph{Notation.} For any increasing function $F$ on the real line, we denote by $F^{-1}$ its left-continuous generalized inverse, $F^{-1}(q) = \inf\{x\in \mathbb R:F(x)\geq q\}$ for $q \in (0,1]$. In particular, for any real-valued random variable $W$ with cumulative distribution function (cdf) $F_W$, $F_W^{-1}$ is the corresponding quantile function. We denote by $\widehat{F}_W$ and $\widehat F_W^{-1}$ the corresponding empirical cdf and quantile function, obtained from a sample $(W_i)_{i=1,\dotsc, n}$. For any $(x,y) \in \mathbb R^2$, we denote by $x \land y$ and $x \lor y$ the minimum and maximum of $x$ and $y$, respectively. We let ${\rm B}(\cdot,\cdot)$ denote the beta function, i.e., for all $x,y>0$, ${\rm B}(x,y) = \int_0^1t^{x-1}(1-t)^{y-1}dt$. We let Beta$(\alpha,\beta)$ denote a random variable with the beta distribution with parameters $(\alpha,\beta)\in(0,\infty)^2$.

\paragraph{Organization.} Section (ref) provides the asymptotic normality result of $\widehat{\theta}$, the plug-in estimator of $\theta_0$, introduces our estimator $\widehat{\sigma}^2$ of the corresponding asymptotic variance and shows its consistency. Section (ref) studies the finite-sample behavior of $\widehat{\theta}$ and compares our estimator $\widehat{\sigma}^2$ with alternative ones. All the proofs are in the appendix, while the supplementary appendix gathers additional lemmas.

\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{Theory}

\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Asymptotic normality of the plug-in estimator}

As mentioned above, we seek to estimate $\theta_0=\int_0^1 F_Y^{-1}(u) dF_U(u)$, where $U$ is unobserved but satisfying $U=F_Z(X)$, whereas $X$ and $Z$ are observed; note that $\theta_0$ is well-defined under Assumption (ref) below (see Lemma (ref) in Appendix (ref)). We observe three samples, $(Y_i)_{i=1,\ldots,n_1}$, $(X_i)_{i=1,\ldots,n_2}$ and $(Z_i)_{i=1,\ldots,n_3}$. We consider the following plug-in estimator of $\theta_0$:

align*[align* omitted — 107 chars of source]

where $\widehat F_Y^{-1}$ is extended to $[0,1]$ by defining $\widehat F_Y^{-1}(0)=Y_{(1)}$.

We prove below that $\widehat{\theta}$ is asymptotically normal under the following conditions. Hereafter, we let $N := \min(n_1, n_2, n_3)$.

assumption[Sampling] \begin{enumerate}[label=(\roman*)] • $(Y_i)_{i=1,\ldots,n_1}$, $(X_i)_{i=1,\ldots,n_2}$ and $(Z_i)_{i=1,\ldots,n_3}$ are three samples of i.i.d. variables with respective cdfs $F_Y$, $F_X$ and $F_Z$. • $(Y_i)_{i=1,\ldots,n_1}$, $(X_i)_{i=1,\ldots,n_2}$, and $(Z_i)_{i=1,\ldots,n_3}$ are mutually independent. • For each $k \in \{1,2,3\}$, there exists $\lambda_k\in[0,1]$ such that $N/n_k\to \lambda_k$ as $N\to \infty$. \end{enumerate}
assumption[Smoothness] \begin{enumerate}[label=(\roman*)] • $F_Z$ is absolutely continuous with respect to the Lebesgue measure with density $f_Z$ supported on $[\underline{z},\overline z]$ with $-\infty\le \underline z<\bar z\le \infty$. • $F_Y$ is continuous and there exist $d_1,d_2 >0$ and $C_Y >0$ such that for all $t\in(0,1)$: \begin{equation} |F_Y^{-1}(t)| \le C_Y t^{-d_1} (1-t)^{-d_2}. \end{equation} • $F_U$ is absolutely continuous with respect to the Lebesgue measure, with continuous density $f_U$ supported on a subset of $[0,1]$. There exist $b_1,b_2>0$ and $C_U >0$ such that for all $u \in (0,1)$: \begin{equation} f_U(u) \le C_U u^{-b_1} (1-u)^{-b_2}. \end{equation} • $b_1+d_1<1/2$ and $b_2+d_2<1/2$. \end{enumerate}

While Assumption (ref)(ref) may be too stringent to cover panel-data versions of AtheyImbens2006's model, it is plausible in the context of independent repeated cross-sections; we discuss the panel case in Section (ref) below. Assumption (ref) imposes restrictions on the distributions of $X$, $Y$ and $Z$. First, their cdf must be continuous. Second, $F_Y$ and $f_U$ must satisfy tail restrictions. In particular, (ref) holds under the following moment condition on $Y$:

lemma[Lower-Level Conditions on $Y$] Assume $E[\vert Y \vert^p] < \infty$ for $p > 1$, then (ref) holds with $d_1=d_2=1/p$.

Given that $U=F_Z(X)$, and assuming that $F_X$ is differentiable, we have $f_U(u)=f_X(F_Z^{-1}(u))/f_Z(F_Z^{-1}(u))$. Hence, (ref) (together with the constraints on $(b_1,b_2)$ implied by Assumption (ref)(ref)) imposes that the tails of $X$ cannot be much heavier than those of $Z$. In the context of the changes-in-changes model, $X$ and $Z$ correspond to the pre-treatment period outcome of the control and treatment group, respectively. Thus, (ref) limits how different the distributions of the outcome in the two groups can be. To illustrate this, assume that $X \sim Z/c$ for some $c>0$ and $f_Z(z)\asymp K \exp(-L|z|^\alpha)$ for some $K, L,\alpha>0$ as $z\to-\infty$ (the same reasoning applies if we consider $z\to\infty$). Then, we show in Appendix (ref) that (ref) implies

equation[equation omitted — 70 chars of source]

Similarly, if the densities of $X$ and $Z$ have power-law tails $|x|^{-c\alpha-1}$ and $|x|^{-\alpha-1}$, respectively, for some $c, \alpha>0$, one can show that (ref) implies $c>1-b_1$.

Finally, Assumption (ref)(ref) implies a trade-off on the tails of $Y$ and $U$: the fatter the tails on $Y$, the lighter those on $U$ should be.

theoremIf Assumptions (ref) and (ref) hold and $\min(\lambda_1,\lambda_3)>0$, then, as $N\to\infty$, \begin{equation*} \sqrt{N}(\widehat{\theta}-\theta_0) \stackrel{d}{\longrightarrow} \mathcal{N}(0,\sigma^2), \end{equation*} where $\sigma^2 = \left[\lambda_1 + \lambda_3 \right]E[\eta^2] + \lambda_2 E[\varepsilon^2]$, with $\eta := - \int_0^{1} [\mathds{1}\{F_Y(Y)\leq t\} - t]f_U(t)\,dF_Y^{-1}(t)$ and $\varepsilon:= -\int_0^1 [\mathds{1}\{U\leq t\} - F_U(t)]\, dF_Y^{-1}(t)$.

The proof of Theorem (ref) is long and technical. The main difficulty lies in showing that various remainder terms are, indeed, negligible. To this end, we exploit several empirical process results, such as the convergence of the supremum of the weighted empirical quantile process csorgo1986. We also establish several results on quantile-quantile transforms that may be of independent interest, see in particular Lemma (ref). Our proof also relies on the fact that order statistics of uniform distributions and uniform spacings follow beta distributions, allowing us to leverage properties of such distributions. Finally, we handle the L-statistic term by relying on the characterization in hecker1976, as standard results on L-statistics, such as those in Chapter 19 of ShorackWellner1986, do not apply here.

The terms $\lambda_1 E[\eta^2]$, $\lambda_2E[\varepsilon^2]$ and $\lambda_3 E[\eta^2]$ in the asymptotic variance correspond to the contributions of the three samples. Specifically, $\lambda_1 E[\eta^2]$ corresponds to the contribution of the estimation of the cdf of $Y$. The term $\lambda_2 E[\varepsilon^2]$ is due to the fact that even if $F_Z$ were known, we would still estimate the cdf of $U$ using the sample $(F_Z(X_i))_{i=1,\ldots,n_2}$. The third term arises due to the estimation of $F_Z$. Perhaps surprisingly, it turns out that if $n_1=n_3$, so that $\lambda_1=\lambda_3$, this contribution is equal to that of the estimation of the cdf of $Y$.\footnote{Though this is not apparent in the expressions of AtheyImbens2006, some algebra show that their first and second variance terms $V^p$ and $V^q$ are in fact equal.} Even if the case $\lambda_3=0$ is not covered by the theorem, the proof of Theorem (ref) shows that if $U$ is observed (namely, if $F_Z$ is known), the estimator is still asymptotically normal with the same variance as above but with $\lambda_3$ set to 0.

To what extent is Assumption (ref) necessary for the result? We argue that, in some sense, Assumption (ref)(ref) is sharp. To see this, assume that $F_Y^{-1}$ is differentiable and

align[align omitted — 108 chars of source]

for some $\underline{C}>0$. This inequality implies that (ref) and (ref) are essentially sharp. The following proposition establishes that if this is the case, then Assumption (ref)(ref) is necessary for $E[\varepsilon^2+\eta^2]<\infty$ to hold. Hence, the restrictions on $X$, $Y$ and $Z$ mentioned above (and in particular that the distributions of $X$ and $Z$ must be sufficiently similar) are, to some extent, required to ensure $E[\varepsilon^2+\eta^2]<\infty$, and are not due to limitations in the proof of Theorem (ref).

propositionSuppose that $F_Y^{-1}$ is differentiable, (ref) holds, $\varepsilon$ and $\eta$ are well-defined and $E[\varepsilon^2+\eta^2]<\infty$. Then $b_k+d_k<1/2$ for $k=1,2$.

In a simpler setup than ours, mason1992necessary show the stronger result that under mild regularity conditions, L-statistics are root-$n$ consistent and asymptotically normal if and only if an integral similar to $E[\eta^2]$ is finite: see the condition $\sigma^2(0)<\infty$ in their Theorem 1.1. We could thus expect that here as well, $\widehat{\theta}$ is root-$n$ consistent and asymptotically normal if and only if $E[\varepsilon^2+\eta^2]<\infty$; our simulations below provide further support for this conjecture.

\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Consistent estimation of the asymptotic variance}

Recall that $\sigma^2=(\lambda_1+\lambda_3)E[\eta^2] + \lambda_2E[\varepsilon^2]$, with $\eta := - \int_0^{1} [\mathds{1}\{F_Y(Y)\leq t\} - t]f_U(t)\,dF_Y^{-1}(t)$ and $\varepsilon:= -\int_0^1 [\mathds{1}\{U\leq t\} - F_U(t)]\, dF_Y^{-1}(t)$. Note that

align*[align* omitted — 100 chars of source]

Then, let $\widehat U_i:=\widehat F_Z(X_i)$ and let $\widehat \varepsilon_i=\widehat \theta - \widehat F_Y^{-1}(\widehat U_i)$. We can simply estimate $E[\varepsilon^2]$ by the sample average of $(\widehat \varepsilon_i^2)_{i=1,...,n_2}$.

The estimation of $E[\eta^2]$ is more challenging. A natural idea would be to consider a plug-in estimator based on the definition of $\eta$. Let us assume, as we do in Assumption (ref) below, that $F_Y^{-1}$ is differentiable. Let also $P(y):=E[(U-\mathds{1}\left\{F_Y(y)\le U\right\})/f_Y(F_Y^{-1}(U)]$, so that $\eta=P(Y)$. Then, following AtheyImbens2006, we could estimate $E[\eta^2]$ by the sample average of $(\widehat{P}^2(Y_i))_{i=1,...,n_1}$, with

equation[equation omitted — 212 chars of source]

However, the inverse-density weighting appearing in (ref) makes it difficult to establish the consistency of this estimator, at least under the weak conditions we impose on the distributions of $Y$ and $U$. We circumvent this difficulty by employing another estimator, based on the following lemma.

lemmaSuppose that $F_Y$ is continuous and $\int_0^1 [t(1-t)]^{1/2} f_U(t)\,dF_Y^{-1}(t)<\infty$. Then, $\eta := - \int_0^{1} [\mathds{1}\{F_Y(Y)\leq t\} - t]f_U(t)\,dF_Y^{-1}(t)$ is well-defined almost surely and satisfies $E[\eta^2]<\infty$. Moreover, it holds that \begin{equation*} E[\eta^2]=\int_{\mathbb R^2} f_U(F_Y(y)) \; f_U(F_Y(y')) \;[F_Y(y)\wedge F_Y(y')] \;[\bar F_Y(y)\land \bar F_Y(y')]dydy'. \end{equation*} where $\bar F_Y=1-F_Y$.

Lemma (ref) shows that $E[\eta^2]$ can be expressed as an integral that does not involve any inverse-density weighting. We develop a plug-in estimator for $E[\eta^2]$ based on this integral. We consider sample-splitting, as it allows us to bound the variance of the estimator in our consistency proof, though the simulations below suggest that this is in fact unnecessary. To simplify notation, assume that $n_1$, $n_2$ and $n_3$ are multiples of $2$. Let $\widehat F_Z^{(1)}$, $\widehat F_Z^{(2)}$ denote two sample-splitting estimators of $F_Z$:

align*[align* omitted — 212 chars of source]

Let $\widehat F_Y^{(1)},\widehat F_Y^{(2)}$ denote two sample-splitting estimators of $F_Y$ defined analogously using the sample $(Y_i)_{i=1,\ldots,n_1}$. Let also $\widehat f_U^{(1)}$ and $\widehat f_U^{(2)}$ denote two sample-splitting kernel density estimators of $f_U$, namely for all $u\in(0,1)$,

align[align omitted — 352 chars of source]

where $h_{n_2,u}:=\varepsilon_{n_2}u(1-u)$ for some positive deterministic sequence $(\varepsilon_{n})_{n\ge 1}$ satisfying the following conditions:

assumption[Bandwidth conditions] For all $n\ge 1$, $\varepsilon_{n}\leq 1/2$ and as $n\to\infty$, $\varepsilon_{n}\to0$, $n\varepsilon_{n}\to\infty$, and $(\log(n)/\sqrt{n})^{1/2-b_j-d_j}=o(\varepsilon_{n})$ for $j\in\{1,2\}$.

We suggest choosing $\varepsilon_{n_2}:=1/\log(n_2)$, which satisfies these restrictions for any $b_1,b_2,d_1,d_2$ that verify the assumptions of Theorems (ref) below. For all $s,t\in[0,1]$, define $w(s,t)=(s\wedge t)(1-s\lor t) = (s \land t) (\bar s \land \bar t)$ where $\bar x = 1-x$ for any $x \in \mathbb R$. We let \[ \widehat E\left[\eta^2\right] := \int_{\mathbb R^2}\widehat f_U^{(1)}\!\left(\widehat F_Y^{(1)}(y)\right)\,\widehat f_U^{(2)}\!\left(\widehat F_Y^{(2)}(y')\right)\,w\!\left(\widehat F_Y^{(1)}(y), \widehat F_Y^{(2)}(y')\right)dydy'. \] Two remarks are in order. First, one could instead combine the subsamples of $U$ and $Y$ differently, replacing, for instance, $f_U^{(1)}\!\left(\widehat F_Y^{(1)}(y)\right)$ by $f_U^{(1)}\!\left(\widehat F_Y^{(2)}(y)\right)$, and then average the two estimators. Second, when a uniform kernel is used in $\widehat f_U^{(1)}$ and $\widehat f_U^{(2)}$, $E\left[\eta^2\right]$ is a double integral of a step function that vanishes outside the compact interval $[Y_{(1)},Y_{(n_1)}]$ and has jumps at the $Y_{(i)}$. Hence, its computation is straightforward.

Finally, given that $\sigma^2=(\lambda_1+\lambda_3)E[\eta^2] + \lambda_2E[\varepsilon^2]$, we estimate $\sigma^2$ by

align*[align* omitted — 320 chars of source]

We show in Theorem (ref) below that $\widehat \sigma^2$ is consistent under Assumptions (ref), (ref) and the following strengthening of Assumption (ref):

assumption[Smoothness] \begin{enumerate}[label=(\roman*)] • $F_Z$ is absolutely continuous with respect to the Lebesgue measure with density $f_Z$ supported on $[\underline{z},\overline z]$ with $-\infty\le \underline z<\bar z\le \infty$. • The support of $Y$ is $[\underline{y},\overline y]$ for some $-\infty\le \underline{y}<\bar y\le \infty$. Moreover, $F_Y^{-1}$ is differentiable on ($\underline{y},\overline{y})$ and there exists $c_Y>0$ such that for all $t\in(0,1)$: \begin{equation} \big(F_Y^{-1}\big)'(t) \leq c_Yt^{-(1+d_1)}(1-t)^{-(1+d_2)}. \end{equation} • The mapping $g:(0,1)^2\to\mathbb R_+$ defined as \[ g(s,t):=(s\wedge t)^{2b_1}(\bar s \land \bar t)^{2b_2}f_U(s)f_U(t), \quad \forall (s,t)\in(0,1)^2, \] is $\beta$-H\"older for some $\beta\in(0,1]$, i.e., there exists $c_U>0$ such that \[ \left\vertg(s,t)-g(s',t')\right\vert\leq c_U\left[\left\verts-s'\right\vert^\beta+\left\vertt-t'\right\vert^\beta\right] \quad \forall s,s',t,t'\in(0,1). \]$(b_1\lor b_2) +( d_1\lor d_2)<1/2$. \end{enumerate}

Condition (ref) is the same as in Assumption (ref), while Condition (ref) is a slight strenghthening of Assumption (ref)(ref). Condition (ref) is similar to, but stronger than, Assumption (ref)(ref). Condition (ref) is also a strengthening of Assumption (ref)(iii). To see this, let $\varphi(s):=s^{b_1} (1-s)^{b_2} f_U(s)$ and note that under Condition (ref), we have,

align*[align* omitted — 197 chars of source]

where the last inequality follows since $\varphi$ is nonnegative on $[0,1]$. Hence, $|\varphi(s) - \varphi(t)| \le (2c_U)^{1/2} |s-s'|^{\beta/2}$, which implies that $\varphi$ is $\beta/2$-Hölder, and thus bounded, on $[0,1]$. Hence, Assumption (ref)(ref) holds under Assumption (ref)(ref). We now state our second main theorem regarding the consistency of our asymptotic variance estimator.

theoremIf Assumptions (ref), (ref) and (ref) hold and $\min(\lambda_1,\lambda_2,\lambda_3)>0$, then, as $N\to\infty$, $$\widehat \sigma^2 \stackrel{P}{\longrightarrow} \sigma^2.$$

\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Panel data applications}

While Assumption (ref) may be reasonable in the repeated cross sections setting of AtheyImbens2006's model, it does not cover panel data applications where $Y$ and $Z$ ($Y_{00}$ and $Y_{01}$ in AtheyImbens2006) are observed on the same units and are thus possibly correlated. We adapt Assumption (ref) as follows. Hereafter, we let $N:=\min(n_1,n_2)$.

assumption[Panel data] \begin{enumerate}[label=(\roman*)] • $(Y_i,Z_i)_{i=1,\ldots,n_1}$ and $(X_i)_{i=1,\ldots,n_2}$ are two samples of i.i.d. variables with respective cdfs $F_{Y,Z}$ (with marginals $F_Y$ and $F_Z$) and $F_X$. • $(Y_i,Z_i)_{i=1,\ldots, n_1}$ and $(X_{i})_{i=1,\ldots, n_2}$ are mutually independent. • For each $k \in \{1,2\}$, there exists $\lambda_k\in[0,1]$, such that $N/n_k\to \lambda_k$ as $N\to \infty$. \end{enumerate}

To handle such cases, we introduce the following sample-splitting estimator of $\theta_0$: \[ \widetilde\theta:=\frac{1}{2}\left(\widehat\theta^{(1)}+\widehat\theta^{(2)}\right), \] where, assuming that $n_1$ and $n_2$ are multiples of $2$ to simplify notation,

align*[align* omitted — 270 chars of source]
theoremIf Assumptions (ref) and (ref) hold and $\min(\lambda_1,\lambda_2)>0$, then, as $N\to\infty$, we have \begin{equation*} \sqrt{N}(\widetilde{\theta}-\theta_0) \stackrel{d}{\longrightarrow} \mathcal{N}(0,\widetilde \sigma^2), \end{equation*} where $\widetilde \sigma^2$ is the quantity $\sigma^2$ defined in Theorem (ref) with $\lambda_3=\lambda_1$.

The proof follows directly from Theorem (ref): by sample splitting, $\widehat\theta^{(1)}$ and $\widehat\theta^{(2)}$ are independent and by Theorem (ref), as $N\to\infty$, we have \[ \sqrt{N}(\widehat{\theta}^{(j)}-\theta_0) \stackrel{d}{\longrightarrow} \mathcal{N}(0,2\widetilde \sigma^2), \quad j=1,2. \]

Similarly, a consistent estimator of the asymptotic variance $\widetilde\sigma^2$ can be obtained by considering a sample-splitting estimator of $\widetilde \sigma^2$ based on four splits of the sample $(Y_i,Z_i)_{i=1,\ldots,n_1}$ to ensure that $\widehat f_U^{(1)}$, $\widehat f_U^{(2)}$, $\widehat F_Y^{(1)}$ and $\widehat F_Y^{(2)}$, which appear in $\widehat E[\eta^2]$, are independent. To simplify notation, assume that $n_1$ is a multiple of $4$. Let

align*[align* omitted — 386 chars of source]

where, for $j\in\{1,2\}$, $\check \varepsilon_i^{(j)}=\widetilde \theta - {\widehat F_Y^{(j)}{}}^{-1}\left(\widehat F_Z^{(3-j)}(X_i)\right)$ and

align*[align* omitted — 729 chars of source]
theoremIf Assumptions (ref) and (ref) hold and $\min(\lambda_1,\lambda_2)>0$, then, as $N\to\infty$, $$\widehat{\widetilde{\sigma}}^2 \stackrel{P}{\longrightarrow} \widetilde \sigma^2.$$

Theorem (ref) follows directly by the independence induced by sample-splitting together with Theorem (ref).

\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{Monte Carlo simulations}

In this section, we investigate the finite sample properties of asymptotic confidence intervals based on Theorems (ref)--(ref). We consider a data generating process that provides a tight control on our assumptions. The random variables $Y_1,\ldots,Y_{n_1}$ are independently and identically distributed (i.i.d.) such that $Y_i=F_Y^{-1}(W_i)$ with $W_i\sim$Uniform$(0,1)$ and \[ F_Y^{-1}(t)=-t^{-d_1}+(1-t)^{-d_2}, \quad \forall t\in(0,1). \] We also assume that $Z_1,\ldots,Z_{n_3}$ are i.i.d. with distribution $\mathcal N(0,1)$, whose cdf is denoted as $\Phi$, and $X_1,\ldots,X_{n_2}$ are i.i.d. such that $X_i=\Phi^{-1}(V_i)$ with $V_i\sim{\rm Beta}(1-b_1,1-b_2)$. All the random variables are mutually independent, $U_i\sim$Beta$(1-b_1,1-b_2)$, and \[ \theta_0 = \frac{{\rm B}(1-b_1,1-b_2-d_2)-{\rm B}(1-b_1-d_1,1-b_2)}{{\rm B}(1-b_1,1-b_2)}. \] We consider $b_1=d_1=0$ and $$(b_2,d_2)\in\{(0.05,0.05),(0.20, 0.05), (0.05, 0.20), (0.20, 0.20), (0.30, 0.30), (0.40, 0.40)\}.$$ The sample size $N=n_1=n_2=n_3$ varies in $\{100, 500,1,000, 10,000\}$. The number of replications is $10,000$.

We first study the behavior of $\widehat{\theta}$ depending on $(b_2,d_2)$. Recall that by Theorem (ref) and since $b_1=d_1=0$, $\widehat{\theta}$ is root-$N$ consistent if $b_2+d_2<0.5$. On the other hand, our results do not cover the cases $b_2=d_2=0.3$ and $b_2=d_2=0.4$. Figure (ref) displays the log of the interquartile range of $\widehat{\theta}$, denoted as IQR$(\widehat{\theta}) := \text{quantile}_{\hat \theta}(0.75) - \text{quantile}_{\hat \theta}(0.25)$, as a function of $\log(N)$ for the different values $(b_2,d_2)$. We also plot straight lines with slope $-1/2$ starting from the initial point corresponding to $\log(N)=\log(100)$. Deviations from these straight lines indicate discrepancies from root-$N$ convergence. It appears that such deviations are moderate for $b_2+d_2< 0.5$, but are large otherwise, with slopes smaller than $-1/2$.

figure[figure omitted — 465 chars of source]

Next, we investigate in Figure (ref) how close the distribution of $\sqrt{N}(\widehat{\theta}-\theta_0)/\sigma$ is from a standard normal distribution. We consider both $(b_2,d_2)=(0.2,0.2)$ and $(b_2,d_2)=(0.3,0.3)$. Because $E[\eta^2+\varepsilon^2]=\infty$ here when $b_2+d_2>0.5$, we redefine, with a slight abuse of notation, $\sigma$ as IQR$(\sqrt{N}(\widehat{\theta}-\theta_0))/1.349$, so that this is object is well-defined even with $(b_2,d_2)=(0.3,0.3)$. Note also that if $b_N(\widehat{\theta}-\theta_0)\stackrel{d}{\longrightarrow} \mathcal{N}(0,\widetilde{\sigma}^2)$ for some diverging sequence $(b_N)_N$ and some $\widetilde{\sigma}^2>0$, $\sqrt{N}(\widehat{\theta}-\theta_0)/\sigma$ would still tend to a standard normal distribution.

Again, we observe a close match between the distribution of $\sqrt{N}(\widehat{\theta}-\theta_0)/\sigma$ and that of a standard normal distribution when $(b_2,d_2)=(0.2,0.2)$. When $(b_2,d_2)=(0.3,0.3)$, on the other hand, the discrepancy between the two distributions remains important even with $N=10,000$, the distribution of $\sqrt{N}(\widehat{\theta}-\theta_0)/\sigma$ being substantially left-skewed.

figure[figure omitted — 981 chars of source]

We now turn to inference. We consider six confidence intervals. The first five are based on asymptotic normality and different variance estimators, whereas the last relies on the bootstrap distribution. The first variance estimator (Split column) is $\widehat\sigma^2$, which is consistent for $\sigma^2$ under the assumptions of Theorem (ref). As suggested in Section (ref), we let $\varepsilon_{n_2}=1/\log(n_2)$. The second variance estimator (No Split column) is based on a variant of $\widehat \sigma^2$ without sample splitting. The third estimator (Unif column) is based on a variant of $\widehat \sigma^2$ without sample splitting and where $h_{n_2,u}:=\varepsilon_{n_2}$. The fourth variance estimator (AI column) is that of AtheyImbens2006. It estimates $E[\eta^2]$ by the sample average of $(\widehat{P}^2(Y_i))_{i=1,...,n_1}$, with $\widehat{P}$ defined in (ref). As in the simulations of AtheyImbens2006, the estimator of $f_Y$ appearing in $\widehat{P}$ is a kernel (Epanechnikov) estimator, with bandwidth equal $h_{n_1}=1.06n_1^{-1/5}/\widehat{\text{sd}}_Y$, where $\widehat{\text{sd}}_Y$ denotes the empirical standard deviation of $Y$. The fifth variance estimator (BSE) is based on the bootstrap (with $1,000$ random draws). The last column (BPC) reports the [0.025,0.975] percentile bootstrap confidence interval, based on 1,000 bootstrap samples.

Table (ref) reports the coverage rates and average lengths of the six confidence intervals. It shows that when $b_2+d_2<1/2$, all confidence intervals have coverage rates close to their nominal level as the sample size increases. The two confidence intervals whose coverage rate is closest to 95% are those based on the percentile bootstrap, and ours without sample-splitting. For $N\in\{500,1,000,10,000\}$, their coverage rates is always between 0.93 and 0.96. Even with $N=100$, their coverage is always greater than or equal to 0.89. This suggests that sample-splitting is not needed for consistency of the variance estimator, and that in fact it may slightly worsen its finite sample properties. The confidence interval based on asymptotic normality and bootstrap standard errors also performs well. Table (ref) also suggests that the AI estimator may be consistent, though the coverage of the corresponding confidence interval is systematically slightly below that of the other confidence intervals, except that based on our asymptotic variance estimator but using a constant bandwidth. The coverage of this latter confidence interval does not improve much with $N$ when $b_2=d_2=0.2$ and remains around 0.85. The results on this confidence interval underline the importance of allowing for a varying bandwidth when estimating the density of $U$.

The cases $b_2=d_2=0.3$ and $b_2=d_2=0.4$ are in line with Figures (ref) and (ref). In such cases, the coverage rates are well below 0.95, and it is unclear whether the coverage of one of the six confidence intervals converges to this level. Specifically, the last four confidence intervals, including those based on the bootstrap, do not display any improvement when $b_2=d_2=0.40$. The coverage of our confidence interval does improve, but still only reaches 0.63 for $N=10,000$. Coverage is better for $b_2=d_2=0.30$, but even in this case distortion remains significant for $N=10,000$. Again, this suggests that our results are sharp at least in terms of the conditions on $(b_2,d_2)$.

table[table omitted — 3,750 chars of source]