Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Inference on quantile processes with a finite number of clusters
\address{University of Michigan Ross School of Business, 701 Tappan Ave, Ann Arbor, MI 48109, USA. Tel.: +1 (734) 764-2355. Fax: +1 (734) 764-2769. E-mail: [email removed]}.}
.
I would like to thank two anonymous reviewers for helpful comments. All errors are my own.
}
abstractI introduce a generic method for inference on entire quantile and regression quantile processes in the presence of a finite number of large and arbitrarily heterogeneous clusters. The method asymptotically controls size by generating statistics that exhibit enough distributional symmetry such that randomization tests can be applied. The randomization test does not require ex-ante matching of clusters, is free of user-chosen parameters, and performs well at conventional significance levels with as few as five clusters. The method tests standard (non-sharp) hypotheses and can even be asymptotically similar in empirically relevant situations. The main focus of the paper is inference on quantile treatment effects but the method applies more broadly. Numerical and empirical examples are provided.
\vskip 1em
Keywords: cluster-robust inference, quantiles, treatment effects, randomization inference, difference in differences
\vskip .5em
JEL codes: C01, C21, C23
Introduction
Economic data often contain large clusters such as countries, regions, villages, or firms. Units within these clusters can be expected to influence one another or are influenced by the same political, environmental, sociological, or technical shocks. Several analytical and computer-intensive procedures such as the bootstrap are available to account for the presence of data clusters. These procedures generally achieve consistency by letting the number of clusters go to infinity. Numerical evidence by bertrandetal2004, mackinnonwebb2014, and others in the context of mean regression suggests that this type of asymptotic approximation often causes substantial size distortions when the number of clusters is small or the clusters are heterogenous. True null hypotheses are rejected far too often in both situations. hagemann2017 shows that this phenomenon is also present in quantile regression.
In this paper, I develop a generic method for inference on the entire quantile or regression quantile process in the presence of a finite number of large and arbitrarily heterogeneous clusters. The method, which I refer to as cluster-randomized Kolmogorov-Smirnov (CRK) test, asymptotically controls size by generating Kolmogorov-Smirnov statistics that exhibit enough distributional symmetry at the cluster level such that randomization tests fisher1935, canayetal2014 can be applied. The CRK test is not limited to the pure quantile regression setting and can be used in distributional difference-in-differences estimation callawayli2019 and related situations where quantile treatment effects are identified by between-cluster comparisons. The CRK test is free of user-chosen parameters, powerful against fixed and root-$n$ local alternatives, and performs well at conventional significance levels with as few as twelve clusters if parameters are identified between clusters. If parameters are identified within clusters, then even five clusters are sufficient for inference.
Quantile regression (QR), introduced by koenkerbassett1978, is an important empirical tool because it can quantify the effect of a set of covariates on the entire conditional outcome distribution. An issue with QR in the presence of clustering is that estimates normalized by their asymptotic covariance kernel have standard normal marginal limit distributions but are no longer pivotal for any choice of weight matrix hagemann2017. Cluster-robust tests about the QR coefficient function therefore have asymptotic distributions that cannot be tabulated for inference about
ranges of quantiles. Even if only individual quantiles are of interest, consistent covariance matrix estimation in large clusters is challenging. It requires knowledge of an explicit ordering of the dependence structure within each cluster combined with a kernel and bandwidth choice to give distant observations less weight. Because time has a natural order, this weighting is easily done for time-dependent data but ordering data within states or villages may be difficult or impossible. The common empirical strategy of simply assuming that the clusters are small and numerous enough to satisfy a central limit theorem circumvents these issues but can lead to substantial size distortions with as few as 20 clusters hagemann2017. This remains true if a cluster-robust version of the bootstrap is used. Distortions can be especially severe if clusters differ greatly in their size and dependence structure.
I show that the CRK test is robust to each of these concerns: It performs well even when the number of clusters is small, the dependence varies from cluster to cluster, and the cluster sizes are heterogenous. The reason for this robustness is that the CRK test does not rely on clustered covariance matrices to rescale the estimates. I instead use randomization inference to generate random critical values that automatically scale to the data. There are no kernels, bandwidths, or spatio-temporal orderings of the data to choose. The test achieves consistency with a finite number of large but heterogeneous clusters under interpretable high-level conditions. Despite being based on randomization inference, the CRK test can perform standard (non-sharp) inference on entire quantile or regression quantile processes. Randomization is performed with a fixed set of estimates and does not require repeated estimation to obtain its critical values.
The randomization method underlying the CRK test was first used in the cluster context by canayetal2014 as a way to perform inference on a finite-dimensional parameter with Student $t$ and Wald statistics in least squares regression. They do not consider inference on quantile functions or Kolmogorov-Smirnov statistics. Here, I considerably extend the scope of their method under explicit regularity conditions to allow for inference on the entire QR process and related objects. The proofs below are fundamentally different from those of canayetal2014\ to account for the infinite-dimensional setting and do not rely on the Skorokhod almost-sure representation theorem. A practical issue with their method is that they require treated clusters to be matched ex-ante with an equal number of control clusters. Each match corresponds to a separate test and two researchers working with the same data can reach different conclusions based on which matches they choose. If there is not an equal number of treated and control clusters, then some clusters have to be combined or dropped in an ad-hoc manner. The CRK test sidesteps these issues completely and explicitly merges all potential tests into a single, uniquely determined test decision using results of ruschendorf1982.
Cluster-robust inference in linear regression models has a long history; recent surveys include cameronmiller2014 and mackinnonnielsenwebb2022.
chenetal2003, wanghe2007, wang2009, parentesantossilva2013, and hagemann2017 provide bootstrap and analytical methods for cluster-robust inference in QR models. yoongalvao2020 discuss the situation where clusters arise from correlation of individual units over time. All of these papers require the number of clusters to go to infinity for consistency. The CRK test differs from these papers because it is based on randomization inference and is consistent with a finite number of clusters.
Several papers show that pointwise inference with a fixed number of clusters is possible under a variety of conditions. ibragimovmueller2010, ibragimovmueller2016 use special properties of the Student $t$ statistic to perform inference on scale mixtures of normal random variables. besteretal2014 use standard cluster-robust covariance matrix estimators but adjust critical values under homogeneity assumptions on the clusters. canayetal2018 show that certain cluster-robust versions of the wild bootstrap can be valid under strong homogeneity assumptions with a fixed number of clusters. hagemann2019b adjusts permutation inference for arbitrary heterogeneity at the cluster level but his bounds only apply to finite-dimensional objects. All of these methods can be used for inference at a single quantile but are not designed for simultaneous inference across ranges of quantiles. In contrast, the CRK test provides uniformly valid inference on the entire quantile process even if clusters are arbitrarily heterogeneous.
The remainder of the paper is organized as follows: Section (ref) establishes new results on randomization inference on Gaussian processes. Section (ref) uses these results to show consistency of the CRK test and gives specific examples where the test applies, including quantile difference-in-differences. Section (ref) illustrates the finite sample behavior of the test in Monte Carlo experiments and an empirical example using Project STAR data. The appendix contains proofs.
I use the following notation and definitions: $1\{\cdot\}$ is the indicator function, cardinality of a set $A$ is $|A|$, the smallest integer greater than or equal $a$ is $\lceil a \rceil$, and the largest integer smaller than or equal $a$ is $\lfloor a \rfloor$. The minimum of $a$ and $b$ is denoted by $a \wedge b$.
Limits are as $n\to\infty$ unless noted otherwise. Convergence in distribution under the parameter $\delta$ is denoted by $\mathchoice
{\raisebox{.0em}{ $\overset{\delta}{\leadsto}$ }}
{\raisebox{-.15em}{ $\overset{\raisebox{-.25em}{\scriptsize$\delta$}}{\leadsto}$ }}
{}
{}$. A stochastic process $\{\xi(t) : t\in\mathcal{T}\}$ indexed by a set $\mathcal{T}$ is a collection of random variables $\xi(t)\colon \Omega \to \mathbb{R}$ defined on the same probability space $(\Omega, \mathcal{F}, P)$. Such a process is Gaussian if and only if $(\xi(t_1), \dots, \xi(t_m))$ is multivariate normal for any finite collection of indices $t_1,\dots,t_m\in\mathcal{T}$.
Randomization inference on Gaussian processes
In this section I study the size of randomization tests when the data come from heterogeneous Gaussian processes. I then analyze asymptotic size when a limiting experiment is characterized by such processes. The next section applies these generic results to the quantile setting.
I first introduce some notation for randomization tests that I will use throughout the paper. Let $u\mapsto X_j(u)$, $1\leqslant j\leqslant q$, be independent mean-zero Gaussian processes indexed by $u\in\mathcal{U}$, where $\mathcal{U}$ is a compact subset of $(0,1)$.
Symmetry about zero implies that $(X_j(u_1), \dots, X_j(u_m))$ and $-(X_j(u_1), \dots, X_j(u_m))$ are identically distributed. Because this is true for every finite collection of indices $u_1,\dots,u_m\in\mathcal{U}$, $u\mapsto X_j(u)$ and $u\mapsto -X_j(u)$ have the same (finite-dimensional) distributions. Define $\mathcal{G} = \{1, -1\}^q$ as the $q$-dimensional product of $\{1, -1\}$ and, for $g = (g_1,\dots,g_q)\in \mathcal{G}$, define $g\mapsto gx$ as the direct product $gx = (g_1x_1,\dots, g_qx_q)$ of $g$ and $x\in\mathbb{R}^q$. Independence and symmetry together imply that $u\mapsto X(u) = (X_1, \dots, X_q)(u)$ and $u\mapsto gX(u)$ have the same distribution for every $g\in\mathcal{G}$ as long as $X$ has mean zero. The quantile and quantile-like processes discussed in the next section have this property under the null hypothesis. Deviations from the null cause non-zero means and therefore also asymmetry in $X$. The goal of this section is to develop a test of the null hypothesis of symmetry about zero,
equation[equation omitted — 119 chars of source]
To test this hypothesis, I use the Kolmogorov-Smirnov-type statistic
equation[equation omitted — 106 chars of source]
This statistic is large if symmetry is violated because the mean of the $X_j(u)$ is positive. I focus on one-sided tests to the right for simplicity but this is not restrictive. To test whether the mean is negative, simply use $-X$ instead of $X$ in the definition of $T$. These test statistics can be combined for two-sided tests. I explain this in detail at the end of Section (ref).
Randomization inference uses distributional invariance to generate null distributions and critical values. In the present case, $X$ is distributionally invariant to all transformations $g$ contained in $\mathcal{G}$ because $X$ is symmetric. Let $T^{(1)}(X, \mathcal{G} ) \leqslant $ $T^{(2)}(X, \mathcal{G} ) \leqslant \dots \leqslant T^{(|\mathcal{G}|)}(X, \mathcal{G} )$ be the $|\mathcal{G}| = 2^q$ ordered values of $T(g X)$ across $g\in\mathcal{G} $ and let
equation[equation omitted — 124 chars of source]
be the $1-\alpha$ quantile of these values. The randomization test function is then
equation[equation omitted — 112 chars of source]
If $\mathcal{U}$ is a finite set, distributional invariance under $H_0$ immediately implies ${\mathord \mathrm{E}} \varphi_{\alpha}(X, \mathcal{G} ) = {\mathord \mathrm{E}} \varphi_{\alpha}(gX, \mathcal{G} ).$ By an argument due to hoeffding1952, the test function must satisfy $|\mathcal{G} |\alpha \geqslant \sum_{g\in\mathcal{G} } \varphi_{\alpha}(gX, \mathcal{G} )$ and, after taking expectations on both sides, equality of the distributions yields $|\mathcal{G} |\alpha \geqslant {\mathord \mathrm{E}} \sum_{g\in\mathcal{G} } \varphi_{\alpha}(gX, \mathcal{G} ) = \sum_{g\in\mathcal{G} } {\mathord \mathrm{E}} \varphi_{\alpha}(gX, \mathcal{G} ) = |\mathcal{G} |{\mathord \mathrm{E}} \varphi_{\alpha}(X, \mathcal{G} )$. This implies ${\mathord \mathrm{E}} \varphi_{\alpha}(X, \mathcal{G} ) \leqslant \alpha$, which makes $T^{1-\alpha}(X,\mathcal{G} )$ an $\alpha$-level critical value.
If $\mathcal{U}$ is a not finite, this argument does not immediately go through because (ref) is a statement about possibly uncountably many $u \in \mathcal{U}$ but I have only established equivalence of the finite-dimensional distributions. However, as the following theorem shows, the conclusion that the test controls size holds nonetheless. The proof of the theorem extends hoeffding1952's proof to stochastic processes with smooth sample paths by showing that (ref) implies equality of the distributions of $(T(gX))_{g\in\mathcal{G}}$ and $(T(g\tilde{g}X))_{g\in\mathcal{G}}$ for every $\tilde{g}\in\mathcal{G}$. I prove that this is enough for hoeffding1952's argument to go through as long as at least one of the processes has positive variance at every $u$.
theoremLet $\{X_1(u)\colon u \in \mathcal{U}\}, \dots, \{X_q(u)\colon u\in \mathcal{U}\}$ be independent mean-zero Gaussian processes with continuous sample paths indexed by the compact set $\mathcal{U}\subset (0,1)$ and let $u\mapsto X(u) := (X_1,\dots, X_q)(u)$. If there is a $j\in\{1,\dots, q \}$ such that ${\mathord P}(X_j(u) = 0) = 0$ for all $u\in \mathcal{U}$, then ${\mathord \mathrm{E}} \varphi_{\alpha}(X, \mathcal{G} )\leqslant \alpha$.
remarks(i) If desired, the test decision can be randomized to construct an exact test. Take an independent variable $V$ with a uniform distribution on $[0,1]$ and the nonrandomized test function
\begin{equation}
\phi_{\alpha}(X, \mathcal{G}) = \begin{cases} 1 &if T(X) > T^{1-\alpha}(X,\mathcal{G} ),\\ a(X) & if T(X) = T^{1-\alpha}(X,\mathcal{G} ),\\ 0 &if T(X) < T^{1-\alpha}(X,\mathcal{G} ), \end{cases}
\end{equation}
where \[ a(X) = \frac{|\mathcal{G}|\alpha - |\{g\in\mathcal{G} : T(gX) > T^{1-\alpha}(X,\mathcal{G} )|}{|\{g\in\mathcal{G} : T(gX) = T^{1-\alpha}(X,\mathcal{G} )|}. \]
Using arguments of hoeffding1952, I show in the proof of Theorem (ref) that the randomized test indeed satisfies ${\mathord P} (\phi_{\alpha}(X, \mathcal{G}) \geqslant V) = \alpha$.
However, this type of test is uncommon in practice because rejecting the null if $\phi_{\alpha}(X, \mathcal{G})\geqslant V$ bases the test decision on a single draw from the uniform distribution. A researcher could therefore draw until a desired conclusion was reached.
(ii) Similar arguments arise in the context of conformal prediction vovketal2005 with exchangeable data. Such arguments do not apply here because $(T(gX))_{g\in\mathcal{G}}$ is generally not exchangeable.
\leavevmode\unskip\penalty9999 \hbox\nobreak
\quad\hbox{\ensuremath{\square}}
If $X$ is only an approximation in the sense that $X_n\leadsto X$ in $\ell^\infty(\mathcal{U})^q$, the space of bounded maps from $\mathcal{U}$ to $\mathbb{R}^q$, then the conclusions of the theorem still hold as long as the non-degeneracy conditions are strengthened. Here and in the following I tacitly assume that a process is indexed by a compact $\mathcal{U}\subset (0,1)$ and that $\ell^\infty(\mathcal{U})^q$ is equipped with the Borel $\sigma$-field induced by the uniform norm topology.
theoremIf $X_n \leadsto X = \{ (X_1,\dots, X_q)(u) \colon u\in\mathcal{U}\},$ where the $\{X_j(u) \colon u\in\mathcal{U}\}$ are independent mean-zero Gaussian processes with continuous sample paths that satisfy ${\mathord P}(X_j(u) = -X_j(u'))=0$ for all $u,u'\in \mathcal{U}$ and $1\leqslant j\leqslant q$, then ${\mathord \mathrm{E}} \varphi_{\alpha}(X_n, \mathcal{G} ) \to {\mathord \mathrm{E}} \varphi_{\alpha}(X , \mathcal{G} )$.
remarks(i) For the non-degeneracy assumption ${\mathord P}(X_j(u) = -X_j(u'))=0$ to fail, a Gaussian process with uniformly continuous sample paths has to traverse, with certainty, from $X_j(u)$ to $X_j(u') = -X_j(u)$ while maintaining a positive variance along the entire path. The process would have to have identical variances at time $u$ and $u'$ but be perfectly negatively correlated at those times, which is impossible for Brownian bridges and related processes that typically arise in a quantile context. Still, such Gaussian processes exist and have to be ruled out.
(ii) The main difficulty of the proof of Theorem (ref) is that the critical value $T^{1-\alpha}(X_n, \mathcal{G})$ does not settle down in the limit and is highly dependent on $T(X)$. The assumptions of Theorem (ref) rule out degeneracies in the limit process that could lead to ties in the order statistics of $\{T(gX) : g\in\mathcal{G} \}$. This would put probability mass on the boundary of the set $\{T(X) > T^{1-\alpha}(X, \mathcal{G})\}$ and prevent application of the portmanteau lemma. canayetal2014 use a delicate construction based on Skorokhod's representation theorem to account for the randomness in the limit. While these results could be extended from vectors to processes, I instead give a direct proof that I can also use to analyze the behavior of the test under both local and global alternatives when I discuss quantile processes in the next section.
(iii) Similar but less involved arguments show that if the supremum in the test statistic (ref) is replaced by an integral over $\mathcal{U}$, then Theorems (ref) and (ref) continue to hold. However, this implicitly changes (ref) to an hypothesis about the symmetry of $\int_\mathcal{U}X(u)du$. Other forms of the test statistic can also lead to valid tests, although the smoothness conditions described in parts (i) and (ii) of this remark may change.
\leavevmode\unskip\penalty9999 \hbox\nobreak
\quad\hbox{\ensuremath{\square}}
Inference on quantile processes with a finite number of clusters
This section gives high level conditions under which asymptotically valid inference on quantile processes and related objects can be performed even if the underlying data come from a fixed number of heterogeneous clusters.
Inference when parameters are identified within clusters
Suppose data from $q$ large clusters (e.g., counties, regions, schools, firms, or stretches of time) are available. Throughout the paper, the number of clusters $q$ remains fixed and does not grow with the number of observations $n$. Observations are independent across clusters but dependent within clusters. Data from each cluster $1\leqslant j\leqslant q$ separately identify a quantile or quantile-like scalar function $\delta : \mathcal{U}\to \mathbb{R}$.
The $\delta$ can be estimated by $\hat{\delta}_j$ using data from only cluster $j$ such that a total of $q$ separate estimates $(\hat{\delta}_1, \dots, \hat{\delta}_q) =: \hat{\delta}$ of $u \mapsto \delta(u)$ are available. The goal is to use randomization inference on a centered and scaled version of $\hat{\delta}$ to develop tests of the null hypothesis
equation[equation omitted — 98 chars of source]
for some known function $\delta_0 : \mathcal{U} \to \mathbb{R}$. The following two examples describe simple but empirically relevant situations that fit this framework.
example[Regression quantiles]
Suppose an outcome $Y_{i,j}$ of individual $i$ in cluster $j$ can be represented as $Y_{i,j} = X_{i,j}\delta(U_{i,j}) + Z_{i,j}'\beta_j(U_{i,j})$, where $u\mapsto X_{i,j}\delta(u) + Z_{i,j}'\beta_j(u)$ is strictly increasing in $u$ and $U_{i,j}$ is standard uniform conditional on covariates $(X_{i,j}, Z_{i,j})$. Here $X_{i,j}$ is the scalar covariate of interest and the $Z_{i,j}$ are additional controls. Monotonicity implies that the $u$-th conditional quantile of $Y_{i,j}$ is $X_{i,j}\delta(u) + Z_{i,j}'\beta_j(u)$ and linear QR as in koenkerbassett1978 can provide estimates $(\hat{\delta}_j,\hat{\beta}_j)$ of $(\delta, \beta_j)$ for each cluster. Testing (ref) with $\delta_0 \equiv 0$ tests whether $Y_{i,j}$ and $X_{i,j}$ are associated at any quantile after controlling for $Z_{i,j}$.
Several related models fit the framework of this example: (i) The $\beta_j$ can be constant across clusters. This does not impact the null hypothesis or the computation of the $\hat{\delta}_j$. (ii) The $\delta$ can vary by cluster in the QR model $Y_{i,j} = X_{i,j}\delta_j(U_{i,j}) + Z_{i,j}'\beta(U_{i,j})$ under the alternative. This has no impact on the computation of the ${\delta}_j$ and the null hypothesis simply becomes $H_0 \colon \delta_1 = \dots = \delta_q = \delta_0$. Identical $\delta_j$ are required only under the null hypothesis. (iii) If $\beta_j\equiv 0$ and $X_{i,j}\equiv 1$, then $u \mapsto \hat{\delta}(u)$ reduces to the $u$-th unconditional empirical quantile of $Y_{i,j}$. The null (ref) can then be used to test whether $\delta$ has a specific functional form, e.g., a standard normal quantile function.
\leavevmode\unskip\penalty9999 \hbox\nobreak
\quad\hbox{\ensuremath{\square}}
example[Quantile treatment effects] Consider predetermined pairs $\{(j, j+q) : 1\leqslant j\leqslant q \}$ of $2q$ groups. Suppose the first $q$ groups received treatment, indicated by $D_{j} = 1\{j\leqslant q\}$, and the remaining groups did not. Groups here could be manufacturing plants or villages. Treatment could be management consulting or introduction of a new technology. Denote treatment and control potential outcomes by $Y_{j}(1)\sim F_{Y(1)}$ and $Y_{j}(0)\sim F_{Y(0)}$, respectively. The observed outcome is $Y_{j} = D_{j}Y_{j}(1) + (1-D_{j})Y_{j}(0)$. For each group $j$, the experimenter observes identically distributed but potentially highly dependent copies $Y_{i,j}$ of $Y_{j}$ representing workers $i$ within group $j$. View each pair $(j, j+q)$ for $1\leqslant j\leqslant q$ as a cluster and define the quantile treatment effect (QTE) as\[ u\mapsto \delta(u) = F^{-1}_{Y(1)}(u) - F^{-1}_{Y(0)}(u). \] This QTE can be estimated as difference of the empirical quantiles \[ u\mapsto \hat{\delta}_j(u) = \hat{F}^{-1}_{Y_{j}}(u) - \hat{F}^{-1}_{Y_{j+q}}(u) \] or, alternatively, as the coefficient on $D_j$ in a QR of $Y_{i,j}$ on a constant and $D_{j}$ using data only from cluster $j$. The situation where $\delta$ varies with $j$ is again included in the analysis as long as the null hypotheses is $\delta_1 = \dots = \delta_q = \delta_0$. Estimation remains unchanged. I discuss the more complex scenario where the counterfactual $F_{Y(0)}$ has to be identified through difference-in-differences methods in Example (ref) ahead.
\leavevmode\unskip\penalty9999 \hbox\nobreak
\quad\hbox{\ensuremath{\square}}
The $\hat{\delta}$ is neither limited to the estimators discussed in the preceding two examples nor does it need to have a special functional form. However, I assume that it can be approximated by a Gaussian process as in Theorem (ref). Let $1_q$ be a $q$-vector of ones.
assumptionThe stochastic process $\{ \hat{\delta}(u) \colon u \in \mathcal{U}\}$ with $\hat{\delta}(u)\in\mathbb{R}^q$ satisfies
\begin{equation}
X_n := \{\sqrt{n}(\hat{\delta}-\delta 1_q)(u) : u\in \mathcal{U}\} \mathchoice
{\raisebox{.0em}{ $\overset{\delta}{\leadsto}$ }}
{\raisebox{-.15em}{ $\overset{\raisebox{-.25em}{\scriptsize$\delta$}}{\leadsto}$ }}
X = \{ (X_1,\dots, X_q)(u) \colon u\in\mathcal{U} \},
\end{equation}
where the components of $X$ are independent mean-zero Gaussian processes with continuous sample paths, ${\mathord P}(X_j(u) = -X_j(u'))=0$ for all $u,u'\in \mathcal{U}$ and $1\leqslant j\leqslant q$.
Examples of $X_n$ that can satisfy this assumption include unconditional quantile functions, coefficient functions in quantile regressions, quantile treatment effects, and other quantile-like objects. machkouriaetal2013 present invariance principles and moment bounds that can be used to establish the convergence condition (ref) under explicit weak dependence conditions.
I now connect the results from Section (ref) about heterogeneous Gaussian processes to tests about $\delta$ under Assumption (ref). The key property is that if $H_0$ in (ref) does not hold, then $\sqrt{n}(\hat{\delta}-\delta_0 1_q) = X_n + \sqrt{n}(\delta - \delta_0)1_q$. The $X_n$ converges to a symmetric process but $\sqrt{n}(\delta-\delta_0)(u)$ grows without bound for some $u$, which makes the distribution of $\sqrt{n}(\hat{\delta}-\delta_0 1_q)$ highly asymmetric. Testing for symmetry using randomization inference is therefore informative about the hypothesis that $\delta=\delta_0$. I refer to a test that uses $\hat{\delta} - \delta_0 1_q$ in place of $X$ in test function (ref) as the cluster-randomized Kolmorogov-Smirnov (CRK) test. From a practical perspective, the function $\delta_0$ is almost always $\delta_0 \equiv 0$. This tests the null of no effect at any quantile but more general hypotheses can be considered.
The test function $x\mapsto\varphi_{\alpha}(x, \mathcal{G} )$ is invariant to scaling of $x$ by positive constants. If $H_0\colon \delta = \delta_0$ is true, then the CRK test satisfies \[T(\hat{\delta}-\delta_0 1_q) > T^{1-\alpha}(\hat{\delta}-\delta_0 1_q,\mathcal{G} )\] if and only if $T(X_n) > T^{1-\alpha}(X_n,\mathcal{G} )$. That the CRK test is an asymptotic $\alpha$-level test
is then an immediate consequence of Theorems (ref) and (ref).
theorem[Size]
Suppose Assumption (ref) holds. If $H_0\colon \delta = \delta_0$ is true, then $\lim_{n\to\infty} {\mathord \mathrm{E}} \varphi_{\alpha}(\hat{\delta} - \delta_0 1_q, \mathcal{G} ) \leqslant \alpha.$
remarks(i) The canonical limit of quantile and regression quantile processes such as those in Examples (ref) and (ref) is a scaled version of a $q$-dimensional Brownian bridge. That process easily satisfies the non-standard condition ${\mathord P}(X_j(u) = -X_j(u'))=0$ imposed by Assumption (ref).
(ii) The inequality in the theorem becomes an equality if $(1-\alpha)2^q$ is an integer. In that case, the test in the limit experiment is “similar,” i.e., it has rejection probability exactly equal to $\alpha$ for all Gaussian processes that satisfy Assumption (ref). The CRK test can therefore be asymptotically similar in some situations. If desired, the test decision can be randomized to make the CRK test similar in the limit for all $\alpha$.
\leavevmode\unskip\penalty9999 \hbox\nobreak
\quad\hbox{\ensuremath{\square}}
To analyze the power of the CRK test, I consider fixed alternatives $\delta(u) = \delta_0(u) + \lambda(u)$ with a positive function $u\mapsto \lambda(u)$, and local alternatives $\delta(u) = \delta_0(u) + \lambda(u)/\sqrt{n}$ converging to the maintained null hypothesis $H_0\colon \delta = \delta_0$. In the local case, $\delta_0$ is fixed but $\delta$ now depends on $n$ and the convergence (ref) is under the sequence of functions $\delta = \delta_0 + \lambda/\sqrt{n}$. As the following results show, the CRK test has power against both types of alternatives.
theorem[Global and local power]
Suppose Assumption (ref) holds and $\alpha \geqslant 1/2^{q} $. If $H_1\colon \delta = \delta_0 + \lambda$ is true with $\lambda \colon \mathcal{U} \to [0,\infty)$ continuous and $\sup_{u\in\mathcal{U}} \lambda(u) > 0$, then $\lim_{n\to\infty}{\mathord \mathrm{E}} \varphi_{\alpha}(\hat{\delta}-\delta_0 1_q, \mathcal{G} ) = 1$. If $H_1\colon \delta = \delta_0 + \lambda/\sqrt{n}$ is true with $\sup_{u\in\mathcal{U}} \lambda(u) > {\mathord \mathrm{E}} \sup_{u\in\mathcal{U}} X_j(u)$, $1\leqslant j\leqslant q$, then \[ \lim_{n\to\infty }{\mathord \mathrm{E}} \varphi_{\alpha}(\hat{\delta}-\delta_0 1_q, \mathcal{G} )\geqslant \prod_{j=1}^q \Bigl( 1 - e^{-[\sup\lambda(u) - {\mathord \mathrm{E}} \sup X_j(u)]^2/2 \sup {\mathord \mathrm{E}} X^2_j(u)}\Bigr) > 0, \] where the suprema in the exponent are over $u\in\mathcal{U}$.
remarks(i) The lower bound used for the local power result comes from the Borell-Tsirelson-Ibragimov-Sudakov (Borell-TIS) inequality adlertaylor2007. For large $q$, the bound is relatively crude but for small $q$, the only crude part is the assumption that $\delta$ is moderately large when compared to $X$. This is reflected in the condition that $\sup_{u\in\mathcal{U}} \lambda(u) > {\mathord \mathrm{E}} \sup_{u\in\mathcal{U}} X_j(u)$ instead of $\sup_{u\in\mathcal{U}} \lambda(u) > 0$. The bound can be made arbitrarily close to $1$ by choosing $\sup_{u\in\mathcal{U}} \lambda(u)$ large enough.
(ii) If $(1-\alpha)|\mathcal{G} | > |\mathcal{G} |-1$, the power of the test is identically zero. In that case $T^{1-\alpha}(X,\mathcal{G}) = \max_{g\in\mathcal{G}}T(gX)$ and $T(X) > T^{1-\alpha}(X,\mathcal{G})$ becomes impossible because $T(X)$ is contained in $\{T(gX) : g\in\mathcal{G} \}$. I therefore I focus on the case $(1-\alpha)|\mathcal{G} |\leqslant |\mathcal{G} |-1$, which is equivalent to $\alpha \geqslant 1/2^q$.
(iii) The test also has power against alternatives where $\lambda$ varies with the cluster index $j$ and at least some of the $\lambda_j$ are large. However, a precise statement without additional conditions on the relative sizes of the $\lambda_j$ is involved. I do not pursue this here to prevent notational clutter.
\leavevmode\unskip\penalty9999 \hbox\nobreak
\quad\hbox{\ensuremath{\square}}
Inference when parameters are identified across clusters
In applications, the treatment effect is often not identified from within a cluster but by comparisons across two clusters. This is the case, for example, if treatment is assigned at random at the cluster level or if identification comes from comparing changes in one cluster to changes in another cluster in a quasi-experimental context. In this situation, each individual pairing of a treated cluster $j$ with a control cluster $k$ is generally informative about the treatment effect of interest $\delta$ and each $(j,k)$ pair gives rise to an estimate $\hat{\delta}_{j,k}$ of $\delta$ that could be used in a CRK-type test. The following example illustrates this for difference-in-differences estimation of quantile treatment effects.
example[Quantile difference in differences] Let $\Delta Y_t(0) = Y_t(0) - Y_{t-1}(0)$ be time differences of untreated outcomes. Periods $t \in \{0,-1 \}$ are pre-intervention periods and $t=1$ is the post-intervention period; $Y_1(1)$ is a treated potential outcome and $Y_t$ are observed outcomes.
Denote by $F_{Y\mid D=d}$ the distribution of a variable $Y$ conditional on the treatment indicator taking on the value $d\in\{0,1\}$. callawayli2019 show that the distribution $F_{Y_1(0)\mid D =1 }(y)$ of the untreated potential outcome of a treated observation at time $t=1$ can be identified as
\begin{equation}
{\mathord P}\Bigl( F^{-1}_{\Delta Y_1\mid D=0} \bigl(F_{\Delta Y_{0} \mid D = 1}(\Delta Y_{0})\bigr) + F^{-1}_{Y_{0}\mid D=0} \bigl(F_{Y_{-1} \mid D = 1}(Y_{-1})\bigr) \leqslant y\mid D = 1\Bigr)
\end{equation} as long as a distributional version of the standard parallel trends assumption and some additional stability and smoothness conditions hold. This identifies the quantile treatment on the treated (QTT) effect \[u\mapsto \delta(u) = F^{-1}_{Y_1(1)\mid D=1}(u) - F^{-1}_{Y_1(0)\mid D=1}(u),\] where $F^{-1}_{Y_1(1)\mid D=1}(u)$ can be estimated by the sample quantile $\hat{F}^{-1}_{Y_1\mid D=1}(u)$. To estimate the counterfactual quantile, callawayli2019
replace ${\mathord P}$ and every $F$ in (ref) with sample equivalents. This yields the estimated QTT
\begin{equation}
u\mapsto \hat{F}^{-1}_{Y_1\mid D=1}(u) - \hat{F}^{-1}_{Y_1(0)\mid D=1}(u).
\end{equation}
callawayli2019 show that $\sqrt{n}(\hat{F}^{-1}_{Y_1\mid D=1} - \hat{F}^{-1}_{Y_1(0)\mid D=1} - \delta)$ converges to a well-behaved Gaussian process under mild regularity conditions.
Suppose that data come from $q_1$ states that received treatment and $q_0$ states that did not. View a single state over time as a cluster. Then two clusters are enough to compute (ref): $\hat{F}^{-1}_{Y_1\mid D=1}$ can be computed from a treated cluster $j$ and $\hat{F}^{-1}_{Y_1(0)\mid D=1}$ can be computed from $j$ and an untreated cluster $k$. Denote by $\hat{\delta}_{j,k}$ the QTT estimated in this fashion using only data from clusters $j$ and $k$. Each $(j,k)$ pair provides a valid estimate of $\delta$ and each $\hat{\delta}_{j,k}$ could potentially be used in a CRK-type test of the null hypothesis $H_0\colon \delta = \delta_0$.
\leavevmode\unskip\penalty9999 \hbox\nobreak
\quad\hbox{\ensuremath{\square}}
I again assume that centered and scaled $\hat{\delta}_{j,k}$ converge in distribution to non-degenerate Gaussian processes with smooth sample paths as in Assumption (ref). I only adjust this condition for the fact that estimates are constructed from pairwise combination of clusters. Let $q_1$ be the number of treated clusters and let $q_0$ be the number of control clusters.
assumptionThe process $\{\sqrt{n}(\hat{\delta}_{j,k} - \delta)(u) : u\in \mathcal{U}\}$ converges, jointly in $j$ and $k$, in distribution to mean-zero Gaussian processes $X_{j,k}$ with continuous sample paths that satisfy ${\mathord P}(X_{j,k}(u) = -X_{j,k}(u'))=0$ for all $u,u'\in \mathcal{U}$, $1\leqslant j\leqslant q_1$, and $1\leqslant k\leqslant q_0$. If both $j\neq j'$ and $k\neq k'$, then $X_{j,k}$ and $X_{j',k'}$ are independent.
A na\"ive test of $H_0\colon \delta \equiv \delta_0$ would now take $X_{n,j,k} := \sqrt{n}(\hat{\delta}_{j,k} - \delta_{0})$ and generate randomization distributions from $\{X_{n,j,k} : 1\leqslant j\leqslant q_1, 1\leqslant k\leqslant q_0\}$ via sign changes. However, $X_{n,j,k}$ and $X_{n,j,k'}$ are dependent for any choice of $j,k,k'$ because $j$ is used twice. This remains true even in large samples and if the data from all $q_1+q_0$ groups are independent. Dependence causes problems because $(X_{n,j,k}, X_{n,j,k'})$ and $(X_{n,j,k}, -X_{n,j,k'})$ generally do not have the same joint distribution even when $n\to\infty$. Invariance under transformations with $g$ therefore fails. This issue can be avoided if one works with a subset of $\{X_{n,j,k} : 1\leqslant j\leqslant q_1, 1\leqslant k\leqslant q_0\}$ that uses each $j$ and $k$ only once. While this solves the dependence issue, it introduces another problem: each of the $q_1$ treatment groups now has to be paired with exactly one of the $q_0$ control groups. Unless these pairings are determined before the data are analyzed, two researchers working with the same data and methodology could arrive at different conclusions because they chose different pairings. To address this problem, I now develop a method that maintains invariance under sign changes but avoids any decisions on the part of the researcher.
I first introduce some notation. If $q_1\leqslant q_0$, there are $q_0 \times (q_0-1) \times \dots \times (q_0-q_1+1)$ ways of choosing $q_1$ ordered elements out of $(1,\dots, q_0)$. Identify each such choice with an $h$ and denote the collection of all $h$ by $\mathcal{H}$. The ordering within $\mathcal{H}$ will not affect the test decision. For each $h\in\mathcal{H}$, denote by
equation[equation omitted — 164 chars of source]
the vector that matches the subset of control groups associated with the label $h = (h(1),\dots, h(q_1))$ to the (unpermuted) treated groups. If there are more treated than control groups such that $q_1 > q_0$, permute treated groups instead and take $h$ as enumerating ways of choosing $q_0$ elements out of $(1,\dots, q_1)$ to define
equation[equation omitted — 159 chars of source]
By construction, the entries of $\hat{\delta}_{[h]}$ are independent of one another but $\hat{\delta}_{[h]}$ and $\hat{\delta}_{[h']}$ for $h,h'\in\mathcal{H}$ are potentially highly dependent.
To address the issue that there are multiple ways of combining clusters, I use an adjustment based on the randomization p-value
equation[equation omitted — 180 chars of source]
Testing with this $p$-value is equivalent to a test with a critical value because $T(X) > T^{1-\alpha}(X,\mathcal{G})$ if and only if $p(X, \mathcal{G}) \leqslant \alpha$. The multiple comparisons adjustment is based on an inequality of ruschendorf1982. It states that arbitrary, possibly dependent variables $U_h$ indexed by $h\in\mathcal{H}$ with the property that ${\mathord P}(U_h\leqslant u)\leqslant u$ for every $u\in [0,1]$ satisfy
equation[equation omitted — 174 chars of source]
This specific form of the inequality is given in vovkwang2020. Here the indexing set $\mathcal{H}$ is arbitrary and does not need to be related to permutations. The only condition is that $H = |\mathcal{H}| \geqslant 2$. The randomization $p$-value $p(\hat{\delta}_{[h]} - \delta_0 1_{q_1 \wedge q_0}, \mathcal{G})$ for testing whether the treatment effect of interest equals $\delta_0$ can be expected to behave like the $U_h$ in (ref) in a large enough sample. Combining $p$-values of the CRK test to reject the null if
equation[equation omitted — 134 chars of source]
does not exceed $\alpha$ should then asymptotically control size. The following theorem confirms that this is indeed true.
theorem[Size with combined p-values]
Suppose Assumption (ref) holds. If $\delta = \delta_0$, then \[\limsup_{n\to\infty} {\mathord P}\Biggl( \frac{2}{H}\sum_{h\in\mathcal{H}} p(\hat{\delta}_{[h]} - \delta_0 1_{q_1 \wedge q_0}, \mathcal{G}) \leqslant \alpha\Biggr)\leqslant \alpha .\]
remarks(i) The theorem can be improved slightly if $\alpha |\mathcal{G}| H/2$ is not an integer. In that case, the limit superior in the theorem is a proper limit that equals ${\mathord P}( (2/H)\sum_{h\in\mathcal{H}} p(X_{[h]},\mathcal{G}) \leqslant \alpha ),$ where $X_{[h]}$ is the weak limit of $\sqrt{n}(\hat{\delta}_{[h]} - \delta_0 1_{q_1 \wedge q_0})$.
This is because the sum in the preceding display can vary discontinuously at certain values. The limit inferior is ${\mathord P}( (2/H)\sum_{h\in\mathcal{H}} p(X_{[h]},\mathcal{G}) < \alpha ).$
(ii) Results of vovkwang2020 suggest that other ways of combining $p$-values such as $\exp(1)$ times the geometric mean of the $p$-value instead of a twice the average $p$-value are likely to be applicable here as well. However, the proof of the theorem given here relies crucially on the properties of the ruschendorf1982 inequality. In the Monte Carlo experiments in the next section, I do not find evidence that other ways of combining $p$-values lead to better results.
\leavevmode\unskip\penalty9999 \hbox\nobreak
\quad\hbox{\ensuremath{\square}}
The price paid for not matching treated and control clusters before the analysis is lower relative power. When $p$-values are averaged, ruschendorf1982's inequality essentially decreases $\alpha$ to $\alpha/2$ to control size. meng1993 shows that the constant $2$ cannot be improved. Still, as I establish below, the test has power against global and local alternatives if $\alpha > 1/2^{q_1 \wedge q_0-1}$, which is slightly stronger than what is needed in Theorem (ref). Compared to Theorem (ref), I also do not state an explicit bound for the local power analysis because applying the Borel-TIS inequality to the averaged $p$-values directly yields only relatively crude results. I instead show that if the alternatives $\lambda/\sqrt{n}$ converging to the null hypothesis are scaled up by a constant $c$, the test can detect these alternatives in the limit experiment with arbitrary accuracy if $c$ is large enough, that is, if first $n\to\infty$ and then $c\to\infty$.
theorem[Global and local power with combined p-values]
Suppose Assumption (ref) holds and $\alpha > 1/2^{q_1 \wedge q_0-1} $. If $H_1\colon \delta = \delta_0 + \lambda$ with $\lambda \colon \mathcal{U} \to [0,\infty)$ continuous and $\sup_{u\in\mathcal{U}}(u) > 0$, then $\lim_{n\to\infty} {\mathord P}( (2/H)\sum_{h\in\mathcal{H}} p(\hat{\delta}_{[h]} - \delta_0 1_{q_1 \wedge q_0}, \mathcal{G}) \leqslant \alpha) = 1$. If $H_1\colon \delta = \delta_0 + c \lambda/\sqrt{n}$, then \[ \lim_{c\to\infty} \liminf_{n\to\infty } {\mathord P}\Biggl( \frac{2}{H}\sum_{h\in\mathcal{H}} p(\hat{\delta}_{[h]} - \delta_0 1_{q_1 \wedge q_0}, \mathcal{G}) \leqslant \alpha\Biggr) = 1. \]
Implementation
I now turn to some practical aspects of the CRK test. I discuss (i) what to do if $\mathcal{G}$ is large, (ii) what to do if $\mathcal{H}$ is large, and (iii) how to implement the test with a step-by-step guide.
First, $\mathcal{G}$ can be prohibitively large if the number of clusters is large. If computing the entire randomization distribution is too costly, then $\mathcal{G}$ can be approximated by a random sample $\mathcal{G}_m$ consisting of $m$ draws from $\mathcal{G}$ with replacement. This is often referred to as “stochastic approximation.” The theorems presented in Sections (ref) and (ref) continue to hold if $\mathcal{G}_m$ is used in place of $\mathcal{G}$ as long as a limit superior or inferior as $m\to \infty$ is applied before $n\to\infty$. The order of limits is not restrictive because, in a given sample of size $n$, the number of draws can $m$ always be made as large as computationally feasible. Under stochastic approximation, the statement in Theorem (ref) becomes $\lim_{n\to\infty}\limsup_{m\to\infty} {\mathord \mathrm{E}} \varphi_{\alpha}(\hat{\delta} - \delta_0 1_q, \mathcal{G}_m ) \leqslant \alpha$, whereas statements about power use a limit inferior. Limit superior and inferior are needed here because of potential discontinuities but can be replaced by regular limits for most values of $\alpha$. Theorems (ref), (ref), and (ref) hold without additional conditions but the conditions of Theorem (ref) have to be strengthened marginally to avoid a discontinuity at $\alpha = 1/2^q$.
propositionSuppose $\mathcal{G}_m$ consists of $m$ iid draws from $\mathcal{G}$. If every instance of $\mathcal{G}$ is replaced by $\mathcal{G}_m$, then
\begin{enumerate}[ (i)]
• Theorem (ref) holds if $\lim_{n\to\infty}$ is replaced by $\lim_{n\to\infty}\limsup_{m\to\infty}$,
• Theorem (ref) holds if every $\lim_{n\to\infty}$ is replaced by $\lim_{n\to\infty}\liminf_{m\to\infty}$ and $\alpha > 1/2^q$,
• Theorem (ref) holds if $\limsup_{n\to\infty}$ is replaced by $\limsup_{n\to\infty}\limsup_{m\to\infty}$,
• Theorem (ref) holds if $\lim_{n\to\infty}$ is replaced by $\lim_{n\to\infty}\liminf_{m\to\infty}$ and $\liminf_{n\to\infty}$ is replaced by $\liminf_{n\to\infty}\liminf_{m\to\infty}$.
\end{enumerate}
If $\alpha \not\in \{ j/|G| : 1\leqslant j\leqslant |G| \}$, then $\liminf_{m\to\infty}$ and $\limsup_{m\to\infty}$ can be replaced by $\lim_{m\to\infty}$ in { (i)-(iv)}.
Second, the number of elements of $\mathcal{H}$ can similarly be large if the number of clusters is large or if there is a large discrepancy between the number of treated and the number of control clusters. In that case one can again work with a random subset $\mathcal{I}$ of $\mathcal{H}$. The crucial difference to the preceding result is that both Theorems (ref) and (ref) continue to hold even if $\mathcal{I}$ consists of only a finite number of random draws. In fact, the result goes through for any $\mathcal{I}$ as long as $\mathcal{I}$ is independent of the data.
propositionLet $\mathcal{I}$ with $|\mathcal{I}|\geqslant 2$ be a fixed or random subset of $\mathcal{H}$ independent of the data. Then Theorems (ref) and (ref) continue to hold if $\mathcal{H}$ is replaced by $\mathcal{I}$.
Finally, the following two algorithms outline and summarize how to apply the CRK test in practice. By Theorems (ref) and (ref), the procedures provide an asymptotically $\alpha$-level test in the presence of a finite number of large clusters that are arbitrarily heterogeneous. They are free of nuisance parameters and do not require any decisions on the part of the researcher. By Theorems (ref) and (ref), the tests are able to detect fixed and $1/\sqrt{n}$-local alternatives. The first algorithm describes the CRK test when the parameters are identified within clusters. The second algorithm describes the between-cluster case, which is needed for distributional difference in differences. The tests can be two-sided or one-sided in either direction.
algorithm[algorithm omitted — 1,477 chars of source]
algorithm[algorithm omitted — 1,549 chars of source]
In some contexts, Algorithm (ref) can be used even if the parameter of interest is identified by comparisons between treated and untreated clusters. For this to work, the researcher has to merge each treated cluster with an untreated cluster into a single cluster to recover within-cluster identification. If the number of treated clusters and control clusters is equal, then every treated cluster can be matched with a control cluster according to some rule. If the number of clusters is not equal, then two or more clusters can be merged to force an equal number of treated and control clusters. The merged clusters can then be reinterpreted as clusters and Algorithm (ref) can be applied to these new clusters. While this comes with a large number of decisions, it is a valid method for inference if these decisions are made before the data are analyzed. For example, when estimating quantile treatment effects, a pre-analysis plan can be put in place that prescribes how clusters that received treatment will be merged with clusters that did not receive treatment. This reduces the problem to the one described in Example (ref).
The next section investigates the finite sample performance of Algorithms (ref) and (ref) in several sitations.
Numerical results
This section presents several Monte Carlo experiments to investigate the small-sample properties of the CRK test in comparison to other methods of inference. I discuss significance tests on quantile regression coefficient functions (Example (ref)), inference in experiments when parameters are identified between clusters (Example (ref)), and estimation of QTEs in Project STAR (Example (ref)). I test one-sided hypotheses to the right but the results apply more broadly.
example[Regression quantiles, cont.] In this example, I adapt an experiment of hagemann2017 and use the data generating process (DGP)
\begin{align*}
Y_{i,j,k} = U_{i,j,k} + U_{i,j,k}Z_{i,j,k},
\end{align*}
where $U_{i,j,k} = \sqrt{\varrho}V_{j,k} + \sqrt{1-\varrho}W_{i,j,k}$ with $\varrho \in [0, 1)$; $V_{j,k}$ and $W_{i,j,k}$ are standard normal, independent of one another, and independent across indices. This ensures that the $U_{i,j,k}$ are standard normal and, for a given $j,k$, any pair $U_{i,j,k}$ and $U_{i',j,k}$ has correlation $\varrho$. The $Z_{i,j,k}$ satisfy $Z_{i,j,k} = X^2_{i,j,k}/3$ with $X_{i,j,k}$ standard normal independent of $U_{i,j,k}$ to ensure that the $U_{i,j,k} Z_{i,j,k}$ have mean zero and variance one. Both $X_{i,j,k}$ and $U_{i,j,k}$ are independent across $j$ and $k$, and $X_{i,j,k}$ is also independent across $i$. I discard information on $k$ after data generation and drop the $k$ subscripts in the following because they are not assumed to be known. This induces a dependence structure where each cluster $j = 1,\dots, q$ consists of several (unknown) neighborhoods $k = 1,\dots,K$ where observations are dependent if they come from the same $k$ but are independent otherwise. If $K\to\infty$ and the size of the neighborhoods is fixed or grows slowly with $K$, then this dependence structure is compatible with Assumptions (ref) and (ref) because it generates the weak dependence needed for central limit theory. In the experiments ahead, I set $K$ to either 10 or 20 and draw the size of each neighborhood from the uniform distribution on $\{5,6,\dots,15\}$.
The DGP in the preceding display corresponds to the QR model
\begin{align}
Q(u\mid X_{i,j}, Z_{i,j}) = \beta_0(u) + \beta_1(u)X_{i,j} + \beta_2(u)Z_{i,j}
\end{align}
with $\beta_1(u)\equiv 0$ and $\beta_0(u) = \Phi^{-1}(u) = \beta_2(u)$, where $\Phi$ is the standard normal distribution function.
For the CRK test, I estimated (ref) separately for each cluster, obtained $q$ estimates of $\beta_1$ and applied Algorithm (ref) with 1,000 new draws from $\mathcal{G}$ for each Monte Carlo replication.
To the best of my knowledge, there are no other methods of inference designed specifically for quantile functions or Kolmogorov-Smirnov statistics in data with few large clusters. I therefore compare the CRK test to inference with the wild gradient bootstrap hagemann2017, a cluster-robust version of the bootstrap that requires the number of clusters $q\to\infty$ for consistency. The wild gradient bootstrap is the default option for cluster-robust inference in the quantreg package in R. I use the package default settings with Mammen bootstrap weights and 200 bootstrap simulations.
Alternative analytical methods for cluster-robust inference in quantile regressions exist but can only perform pointwise inference because the QR process as $q\to\infty$ generally has an analytically intractable distribution. hagemann2017 shows that the wild gradient bootstrap can conduct uniform inference on quantile regression functions and that it outperforms other methods for pointwise inference in this context.
However, hagemann2017 notes that size distortions can occur when fewer than 20 clusters are present. I therefore focus on this situation in the following.
\begin{figure}
\resizebox{\textwidth}{!}{
\begin{tikzpicture}[x=1pt,y=1pt]
\definecolor{fillColor}{RGB}{255,255,255}
\path[use as bounding box,fill=fillColor,fill opacity=0.00] (0,0) rectangle (433.62,216.81);
\begin{scope}
\path[clip] ( 48.00, 42.00) rectangle (204.81,192.81);
\definecolor{drawColor}{RGB}{0,0,0}
\path[draw=drawColor,line width= 1.6pt,line join=round,line cap=round] ( 53.81,147.68) --
( 63.49,135.95) --
( 73.17,122.77) --
( 82.85,115.73) --
( 92.53,111.71) --
(102.21,109.47) --
(111.89,107.69) --
(121.57,102.10) --
(131.24,104.45) --
(140.92,101.65) --
(150.60, 99.53) --
(160.28, 97.19) --
(169.96, 93.39) --
(179.64, 99.42) --
(189.32, 95.06) --
(199.00, 94.62);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (433.62,216.81);
\definecolor{drawColor}{RGB}{0,0,0}
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 53.81, 42.00) -- (199.00, 42.00);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 53.81, 42.00) -- ( 53.81, 36.00);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (102.21, 42.00) -- (102.21, 36.00);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (150.60, 42.00) -- (150.60, 36.00);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (199.00, 42.00) -- (199.00, 36.00);
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 53.81, 20.40) {5};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (102.21, 20.40) {10};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (150.60, 20.40) {15};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (199.00, 20.40) {20};
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00, 47.59) -- ( 48.00,187.22);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00, 47.59) -- ( 42.00, 47.59);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00, 75.51) -- ( 42.00, 75.51);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00,103.44) -- ( 42.00,103.44);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00,131.37) -- ( 42.00,131.37);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00,159.30) -- ( 42.00,159.30);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00,187.22) -- ( 42.00,187.22);
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 33.60, 47.59) {0.00};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 33.60, 75.51) {0.05};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 33.60,103.44) {0.10};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 33.60,131.37) {0.15};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 33.60,159.30) {0.20};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 33.60,187.22) {0.25};
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00, 42.00) --
(204.81, 42.00) --
(204.81,192.81) --
( 48.00,192.81) --
cycle;
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (216.81,216.81);
\definecolor{drawColor}{RGB}{0,0,0}
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.0] at (126.41,200.68) {Bootstrap (size, $u \mapsto\beta_1(u)$)};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (126.41, 2.40) {Number of clusters};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 15.60,117.41) {Rejection frequency};
\end{scope}
\begin{scope}
\path[clip] ( 48.00, 42.00) rectangle (204.81,192.81);
\definecolor{drawColor}{RGB}{0,0,0}
\path[draw=drawColor,line width= 1.6pt,dash pattern=on 7pt off 3pt ,line join=round,line cap=round] ( 53.81,153.71) --
( 63.49,135.39) --
( 73.17,124.11) --
( 82.85,116.73) --
( 92.53,115.39) --
(102.21,112.71) --
(111.89,108.24) --
(121.57,101.21) --
(131.24, 97.19) --
(140.92,101.88) --
(150.60, 96.40) --
(160.28, 99.20) --
(169.96, 93.72) --
(179.64, 95.84) --
(189.32, 91.49) --
(199.00, 89.92);
\path[draw=drawColor,line width= 1.6pt,dash pattern=on 1pt off 3pt ,line join=round,line cap=round] ( 53.81,164.44) --
( 63.49,146.56) --
( 73.17,134.50) --
( 82.85,128.35) --
( 92.53,120.53) --
(102.21,118.97) --
(111.89,112.71) --
(121.57,105.56) --
(131.24,108.02) --
(140.92,104.78) --
(150.60,102.44) --
(160.28,100.76) --
(169.96, 98.19) --
(179.64, 99.20) --
(189.32, 95.84) --
(199.00, 94.17);
\path[draw=drawColor,line width= 0.4pt,dash pattern=on 4pt off 4pt ,line join=round,line cap=round] ( 48.00, 75.51) -- (204.81, 75.51);
\end{scope}
\begin{scope}
\path[clip] (264.81, 42.00) rectangle (421.62,192.81);
\definecolor{drawColor}{RGB}{0,0,0}
\path[draw=drawColor,line width= 1.6pt,line join=round,line cap=round] (270.62, 68.70) --
(280.30, 71.83) --
(289.98, 71.27) --
(299.66, 77.19) --
(309.34, 74.40) --
(319.02, 74.62) --
(328.70, 78.19) --
(338.38, 74.62) --
(348.05, 73.95) --
(357.73, 74.06) --
(367.41, 75.51) --
(377.09, 73.84) --
(386.77, 76.18) --
(396.45, 75.85) --
(406.13, 75.40) --
(415.81, 74.51);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (433.62,216.81);
\definecolor{drawColor}{RGB}{0,0,0}
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (270.62, 42.00) -- (415.81, 42.00);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (270.62, 42.00) -- (270.62, 36.00);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (319.02, 42.00) -- (319.02, 36.00);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (367.41, 42.00) -- (367.41, 36.00);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (415.81, 42.00) -- (415.81, 36.00);
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (270.62, 20.40) {5};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (319.02, 20.40) {10};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (367.41, 20.40) {15};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (415.81, 20.40) {20};
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81, 47.59) -- (264.81,187.22);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81, 47.59) -- (258.81, 47.59);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81, 75.51) -- (258.81, 75.51);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81,103.44) -- (258.81,103.44);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81,131.37) -- (258.81,131.37);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81,159.30) -- (258.81,159.30);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81,187.22) -- (258.81,187.22);
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (250.41, 47.59) {0.00};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (250.41, 75.51) {0.05};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (250.41,103.44) {0.10};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (250.41,131.37) {0.15};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (250.41,159.30) {0.20};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (250.41,187.22) {0.25};
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81, 42.00) --
(421.62, 42.00) --
(421.62,192.81) --
(264.81,192.81) --
cycle;
\end{scope}
\begin{scope}
\path[clip] (216.81, 0.00) rectangle (433.62,216.81);
\definecolor{drawColor}{RGB}{0,0,0}
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.0] at (343.21,200.68) {CRK test (size, $u \mapsto \beta_1(u)$)};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (343.21, 2.40) {Number of clusters};
\end{scope}
\begin{scope}
\path[clip] (264.81, 42.00) rectangle (421.62,192.81);
\definecolor{drawColor}{RGB}{0,0,0}
\path[draw=drawColor,line width= 1.6pt,dash pattern=on 7pt off 3pt ,line join=round,line cap=round] (270.62, 64.90) --
(280.30, 71.04) --
(289.98, 74.73) --
(299.66, 74.62) --
(309.34, 76.85) --
(319.02, 76.63) --
(328.70, 76.85) --
(338.38, 74.84) --
(348.05, 76.85) --
(357.73, 77.19) --
(367.41, 75.51) --
(377.09, 74.06) --
(386.77, 75.07) --
(396.45, 75.63) --
(406.13, 77.41) --
(415.81, 74.40);
\path[draw=drawColor,line width= 1.6pt,dash pattern=on 1pt off 3pt ,line join=round,line cap=round] (270.62, 65.57) --
(280.30, 68.92) --
(289.98, 71.94) --
(299.66, 77.08) --
(309.34, 76.85) --
(319.02, 73.61) --
(328.70, 74.40) --
(338.38, 76.30) --
(348.05, 73.61) --
(357.73, 76.41) --
(367.41, 77.75) --
(377.09, 76.74) --
(386.77, 75.96) --
(396.45, 76.52) --
(406.13, 73.06) --
(415.81, 76.18);
\path[draw=drawColor,line width= 0.4pt,dash pattern=on 4pt off 4pt ,line join=round,line cap=round] (264.81, 75.51) -- (421.62, 75.51);
\definecolor{fillColor}{RGB}{255,255,255}
\path[fill=fillColor] (325.68,192.81) rectangle (421.62,135.81);
\path[draw=drawColor,line width= 1.6pt,line join=round,line cap=round] (334.23,181.41) -- (351.33,181.41);
\path[draw=drawColor,line width= 1.6pt,dash pattern=on 7pt off 3pt ,line join=round,line cap=round] (334.23,170.01) -- (351.33,170.01);
\path[draw=drawColor,line width= 1.6pt,dash pattern=on 1pt off 3pt ,line join=round,line cap=round] (334.23,158.61) -- (351.33,158.61);
\path[draw=drawColor,line width= 0.4pt,dash pattern=on 4pt off 4pt ,line join=round,line cap=round] (334.23,147.21) -- (351.33,147.21);
\node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.85] at (359.88,178.14) {$K=10, \varrho = $.5};
\node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.85] at (359.88,166.74) {$K=20, \varrho = $.5};
\node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.85] at (359.88,155.34) {$K=10, \varrho = $.1};
\node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.85] at (359.88,143.94) {5% level};
\end{scope}
\end{tikzpicture}
}
\caption{Rejection frequencies in Example (ref) of a true null $H_0\colon \beta_1(u) = 0$ for all $u$ as a function of the number of clusters for the bootstrap (left) and the CRK test (right) with (i) $K=10$ neighborhoods per cluster with intra-neighborhood correlation $\varrho = .5$ (solid lines), (ii) $K=20$ with $\varrho = .5$ (long-dashed), and (iii) $K=10$ with $\varrho = .1$ (dotted). Short-dashed line equals nominal level $.05$.}
\end{figure}
Figure (ref) shows the rejection frequencies of a true null hypothesis $H_0\colon \beta_1(u) = 0$ for all $u$ as a function of the number of clusters $q\in \{5,6,\dots, 20\}$ for the wild gradient bootstrap (left) and the CRK test (right) at the 5% level (short-dashed line). The figure shows rejection frequencies in 5,000 Monte Carlo replications for each horizontal coordinate with (i) $K=10$ neighborhoods per cluster with intra-neighborhood correlation $\varrho = .5$ (solid lines), (ii) $K=20$ with $\varrho = .5$ (long-dashed), and (iii) $K=10$ with $\varrho = .1$ (dotted). Both methods were faced with the same data and I estimated $\beta_1$ at $u = .1, .2, \dots, .9$ for both methods. As can be seen, the wild gradient bootstrap over-rejected mildly with 20 clusters but over-rejected substantially for smaller numbers of clusters. It exceeded a 10% rejection rate if only 12 clusters were available. With 5 clusters, the wild gradient bootstrap falsely discovered an effect in up to 20.9% of all cases ($K=10, \varrho = .1$). In contrast, the CRK test rejected at or slightly below nominal level for all $q$ and all configurations of $K$ and $\varrho$.
I also experimented with a large number of alternative DGPs under the null. I considered (not shown) larger neighborhoods, different values of $\varrho$, different spatial dependence structures such as (spatial) autoregressive models, and different distributions for $X_{i,j,k}$. However, I found that these changes had little qualitative impact on the results described in the preceding paragraph or in hagemann2017. The wild gradient bootstrap generally performed very well but experienced size distortions with fewer than 20 clusters. The CRK test rejected at or slightly below nominal level in all situations I investigated.
I now turn to the behavior of the test under the alternative. I repeated the experiment but now tested the incorrect null hypothesis $H_0\colon \beta_2(u) = 0$ for all $u\in\mathcal{U}$. Figure (ref) shows the rejection frequencies of this null against the alternative $H_1\colon \beta_2(u) > 0$ for some $u\in\mathcal{U}$, where $\mathcal{U}$ was either $(0, 1)$ (black) or $(.5, 1)$ (grey). The null hypothesis is false in both situations but the case where $\mathcal{U} = (0, 1)$ is more challenging because $\beta_2(u) < 0$ for all $u < .5$ so that estimates below the median provide evidence in the direction away from the alternative. I again considered (i) $K=10$ neighborhoods per cluster with intra-neighborhood correlation $\varrho = .5$ (solid lines), (ii) $K=20$ with $\varrho = .5$ (long-dashed), and (iii) $K=10$ with $\varrho = .1$ (dotted). As could be expected, the bootstrap rejected a large fraction of null hypotheses mostly because it was unable to control the size of the test. However, it had high power when the number of clusters was above 20 and the size distortions disappeared (not shown). The CRK test had high power while maintaining size control even when the number of clusters was below 20. For example, at $q=12$ it detected a deviation from the null between 22.5% ($K=10, \varrho=.5, \mathcal{U} = (0,1)$) and 84.26% ($K=20, \varrho=.5, \mathcal{U} = (.5,1)$) of all cases. More generally, additional clusters, lower intra-cluster dependence, and additional neighborhoods per cluster increased the power of the CRK test.
\leavevmode\unskip\penalty9999 \hbox\nobreak
\quad\hbox{\ensuremath{\square}}
\begin{figure}
\resizebox{\textwidth}{!}{
\begin{tikzpicture}[x=1pt,y=1pt]
\definecolor{fillColor}{RGB}{255,255,255}
\path[use as bounding box,fill=fillColor,fill opacity=0.00] (0,0) rectangle (433.62,216.81);
\begin{scope}
\path[clip] ( 48.00, 42.00) rectangle (204.81,192.81);
\definecolor{drawColor}{RGB}{0,0,0}
\path[draw=drawColor,line width= 1.6pt,line join=round,line cap=round] ( 53.81,115.39) --
( 63.49,120.62) --
( 73.17,126.29) --
( 82.85,131.26) --
( 92.53,134.94) --
(102.21,143.07) --
(111.89,144.97) --
(121.57,152.26) --
(131.24,155.78) --
(140.92,159.58) --
(150.60,164.02) --
(160.28,166.53) --
(169.96,169.07) --
(179.64,171.56) --
(189.32,172.62) --
(199.00,175.36);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (433.62,216.81);
\definecolor{drawColor}{RGB}{0,0,0}
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 53.81, 42.00) -- (199.00, 42.00);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 53.81, 42.00) -- ( 53.81, 36.00);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (102.21, 42.00) -- (102.21, 36.00);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (150.60, 42.00) -- (150.60, 36.00);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (199.00, 42.00) -- (199.00, 36.00);
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 53.81, 20.40) {5};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (102.21, 20.40) {10};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (150.60, 20.40) {15};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (199.00, 20.40) {20};
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00, 47.59) -- ( 48.00,187.22);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00, 47.59) -- ( 42.00, 47.59);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00, 75.51) -- ( 42.00, 75.51);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00,103.44) -- ( 42.00,103.44);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00,131.37) -- ( 42.00,131.37);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00,159.30) -- ( 42.00,159.30);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00,187.22) -- ( 42.00,187.22);
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 33.60, 47.59) {0.0};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 33.60, 75.51) {0.2};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 33.60,103.44) {0.4};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 33.60,131.37) {0.6};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 33.60,159.30) {0.8};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 33.60,187.22) {1.0};
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00, 42.00) --
(204.81, 42.00) --
(204.81,192.81) --
( 48.00,192.81) --
cycle;
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (216.81,216.81);
\definecolor{drawColor}{RGB}{0,0,0}
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.0] at (126.41,200.68) {Bootstrap (power, $u \mapsto \beta_2(u)$)};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (126.41, 2.40) {Number of clusters};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 15.60,117.41) {Rejection frequency};
\end{scope}
\begin{scope}
\path[clip] ( 48.00, 42.00) rectangle (204.81,192.81);
\path[] (108.79,143.94) rectangle (204.73, 98.34);
\definecolor{drawColor}{RGB}{169,169,169}
\path[draw=drawColor,line width= 1.6pt,line join=round,line cap=round] (117.34,132.54) -- (134.44,132.54);
\path[draw=drawColor,line width= 1.6pt,dash pattern=on 7pt off 3pt ,line join=round,line cap=round] (117.34,121.14) -- (134.44,121.14);
\path[draw=drawColor,line width= 1.6pt,dash pattern=on 1pt off 3pt ,line join=round,line cap=round] (117.34,109.74) -- (134.44,109.74);
\definecolor{drawColor}{RGB}{255,255,255}
\node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.85] at (142.99,129.26) {$K=10, \varrho = $.5};
\node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.85] at (142.99,117.86) {$K=20, \varrho = $.5};
\node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.85] at (142.99,106.46) {$K=10, \varrho = $.1};
\path[] (108.87,145.91) rectangle (204.81, 88.91);
\definecolor{drawColor}{RGB}{0,0,0}
\path[draw=drawColor,line width= 1.6pt,line join=round,line cap=round] (117.42,134.50) -- (134.52,134.50);
\path[draw=drawColor,line width= 1.6pt,dash pattern=on 7pt off 3pt ,line join=round,line cap=round] (117.42,123.10) -- (134.52,123.10);
\path[draw=drawColor,line width= 1.6pt,dash pattern=on 1pt off 3pt ,line join=round,line cap=round] (117.42,111.70) -- (134.52,111.70);
\path[draw=drawColor,line width= 0.4pt,dash pattern=on 4pt off 4pt ,line join=round,line cap=round] (117.42,100.30) -- (134.52,100.30);
\node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.85] at (143.07,131.23) {$K=10, \varrho = $.5};
\node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.85] at (143.07,119.83) {$K=20, \varrho = $.5};
\node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.85] at (143.07,108.43) {$K=10, \varrho = $.1};
\node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.85] at (143.07, 97.03) {5% level};
\path[draw=drawColor,line width= 1.6pt,dash pattern=on 7pt off 3pt ,line join=round,line cap=round] ( 53.81,148.63) --
( 63.49,156.11) --
( 73.17,162.87) --
( 82.85,167.70) --
( 92.53,171.67) --
(102.21,176.22) --
(111.89,180.24) --
(121.57,181.53) --
(131.24,182.34) --
(140.92,184.04) --
(150.60,185.13) --
(160.28,185.44) --
(169.96,186.02) --
(179.64,186.50) --
(189.32,186.86) --
(199.00,186.97);
\path[draw=drawColor,line width= 1.6pt,dash pattern=on 1pt off 3pt ,line join=round,line cap=round] ( 53.81,157.26) --
( 63.49,163.77) --
( 73.17,169.21) --
( 82.85,172.53) --
( 92.53,175.36) --
(102.21,178.76) --
(111.89,180.86) --
(121.57,182.76) --
(131.24,184.07) --
(140.92,185.38) --
(150.60,185.10) --
(160.28,186.08) --
(169.96,186.22) --
(179.64,186.61) --
(189.32,186.69) --
(199.00,186.83);
\definecolor{drawColor}{RGB}{169,169,169}
\path[draw=drawColor,line width= 1.6pt,line join=round,line cap=round] ( 53.81,118.83) --
( 63.49,124.11) --
( 73.17,130.17) --
( 82.85,135.53) --
( 92.53,139.22) --
(102.21,147.29) --
(111.89,149.35) --
(121.57,156.00) --
(131.24,159.66) --
(140.92,163.49) --
(150.60,167.79) --
(160.28,169.57) --
(169.96,172.31) --
(179.64,174.80) --
(189.32,175.30) --
(199.00,177.90);
\path[draw=drawColor,line width= 1.6pt,dash pattern=on 7pt off 3pt ,line join=round,line cap=round] ( 53.81,151.92) --
( 63.49,159.69) --
( 73.17,166.08) --
( 82.85,170.47) --
( 92.53,174.04) --
(102.21,178.68) --
(111.89,181.86) --
(121.57,182.87) --
(131.24,183.87) --
(140.92,184.96) --
(150.60,185.80) --
(160.28,186.08) --
(169.96,186.36) --
(179.64,186.75) --
(189.32,187.06) --
(199.00,187.14);
\path[draw=drawColor,line width= 1.6pt,dash pattern=on 1pt off 3pt ,line join=round,line cap=round] ( 53.81,161.06) --
( 63.49,167.26) --
( 73.17,171.81) --
( 82.85,175.36) --
( 92.53,177.76) --
(102.21,180.55) --
(111.89,182.59) --
(121.57,184.15) --
(131.24,185.07) --
(140.92,185.74) --
(150.60,185.80) --
(160.28,186.36) --
(169.96,186.55) --
(179.64,186.86) --
(189.32,186.95) --
(199.00,187.14);
\definecolor{drawColor}{RGB}{0,0,0}
\path[draw=drawColor,line width= 0.4pt,dash pattern=on 4pt off 4pt ,line join=round,line cap=round] ( 48.00, 54.57) -- (204.81, 54.57);
\end{scope}
\begin{scope}
\path[clip] (264.81, 42.00) rectangle (421.62,192.81);
\definecolor{drawColor}{RGB}{0,0,0}
\path[draw=drawColor,line width= 1.6pt,line join=round,line cap=round] (270.62, 61.66) --
(280.30, 67.36) --
(289.98, 72.72) --
(299.66, 72.53) --
(309.34, 76.13) --
(319.02, 77.75) --
(328.70, 79.00) --
(338.38, 81.66) --
(348.05, 83.81) --
(357.73, 85.76) --
(367.41, 88.72) --
(377.09, 88.00) --
(386.77, 90.54) --
(396.45, 93.44) --
(406.13, 95.84) --
(415.81, 97.97);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (433.62,216.81);
\definecolor{drawColor}{RGB}{0,0,0}
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (270.62, 42.00) -- (415.81, 42.00);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (270.62, 42.00) -- (270.62, 36.00);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (319.02, 42.00) -- (319.02, 36.00);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (367.41, 42.00) -- (367.41, 36.00);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (415.81, 42.00) -- (415.81, 36.00);
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (270.62, 20.40) {5};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (319.02, 20.40) {10};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (367.41, 20.40) {15};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (415.81, 20.40) {20};
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81, 47.59) -- (264.81,187.22);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81, 47.59) -- (258.81, 47.59);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81, 75.51) -- (258.81, 75.51);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81,103.44) -- (258.81,103.44);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81,131.37) -- (258.81,131.37);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81,159.30) -- (258.81,159.30);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81,187.22) -- (258.81,187.22);
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (250.41, 47.59) {0.0};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (250.41, 75.51) {0.2};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (250.41,103.44) {0.4};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (250.41,131.37) {0.6};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (250.41,159.30) {0.8};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (250.41,187.22) {1.0};
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81, 42.00) --
(421.62, 42.00) --
(421.62,192.81) --
(264.81,192.81) --
cycle;
\end{scope}
\begin{scope}
\path[clip] (216.81, 0.00) rectangle (433.62,216.81);
\definecolor{drawColor}{RGB}{0,0,0}
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.0] at (343.21,200.68) {CRK test (power, $u \mapsto\beta_2(u)$)};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (343.21, 2.40) {Number of clusters};
\end{scope}
\begin{scope}
\path[clip] (264.81, 42.00) rectangle (421.62,192.81);
\definecolor{drawColor}{RGB}{0,0,0}
\path[draw=drawColor,line width= 1.6pt,dash pattern=on 7pt off 3pt ,line join=round,line cap=round] (270.62, 82.47) --
(280.30, 95.76) --
(289.98,107.46) --
(299.66,115.92) --
(309.34,124.08) --
(319.02,131.23) --
(328.70,138.04) --
(338.38,143.99) --
(348.05,148.63) --
(357.73,152.96) --
(367.41,157.45) --
(377.09,162.26) --
(386.77,165.19) --
(396.45,167.00) --
(406.13,171.42) --
(415.81,172.95);
\path[draw=drawColor,line width= 1.6pt,dash pattern=on 1pt off 3pt ,line join=round,line cap=round] (270.62, 82.75) --
(280.30, 98.61) --
(289.98,108.50) --
(299.66,116.20) --
(309.34,123.47) --
(319.02,129.11) --
(328.70,135.59) --
(338.38,140.92) --
(348.05,145.75) --
(357.73,149.13) --
(367.41,152.79) --
(377.09,155.72) --
(386.77,160.44) --
(396.45,162.76) --
(406.13,166.06) --
(415.81,168.62);
\definecolor{drawColor}{RGB}{169,169,169}
\path[draw=drawColor,line width= 1.6pt,line join=round,line cap=round] (270.62, 66.58) --
(280.30, 74.12) --
(289.98, 81.88) --
(299.66, 82.24) --
(309.34, 86.52) --
(319.02, 89.78) --
(328.70, 90.96) --
(338.38, 93.64) --
(348.05, 97.88) --
(357.73, 98.50) --
(367.41,101.68) --
(377.09,101.93) --
(386.77,104.03) --
(396.45,107.38) --
(406.13,110.12) --
(415.81,112.38);
\path[draw=drawColor,line width= 1.6pt,dash pattern=on 7pt off 3pt ,line join=round,line cap=round] (270.62, 97.49) --
(280.30,118.08) --
(289.98,130.95) --
(299.66,141.28) --
(309.34,149.91) --
(319.02,155.55) --
(328.70,161.31) --
(338.38,165.25) --
(348.05,169.43) --
(357.73,172.17) --
(367.41,174.74) --
(377.09,177.70) --
(386.77,178.85) --
(396.45,180.66) --
(406.13,181.58) --
(415.81,182.70);
\path[draw=drawColor,line width= 1.6pt,dash pattern=on 1pt off 3pt ,line join=round,line cap=round] (270.62,101.54) --
(280.30,123.60) --
(289.98,133.04) --
(299.66,141.90) --
(309.34,148.91) --
(319.02,152.59) --
(328.70,158.10) --
(338.38,162.28) --
(348.05,165.97) --
(357.73,168.29) --
(367.41,169.66) --
(377.09,171.33) --
(386.77,174.27) --
(396.45,175.58) --
(406.13,177.59) --
(415.81,179.29);
\definecolor{drawColor}{RGB}{0,0,0}
\path[draw=drawColor,line width= 0.4pt,dash pattern=on 4pt off 4pt ,line join=round,line cap=round] (264.81, 54.57) -- (421.62, 54.57);
\end{scope}
\end{tikzpicture}
}
\caption{Rejection frequencies in Example (ref) of false nulls $H_0\colon \beta_2(u) = 0$ for $u > .5$ (grey) and $H_0\colon \beta_2(u) = 0$ for all $u$ (black) as a function of the number of clusters for the bootstrap (left) and the CRK test (right) with (i) $K=10$ neighborhoods per cluster with intra-neighborhood correlation $\varrho = .5$ (solid lines), (ii) $K=20$ with $\varrho = .5$ (long-dashed), and (iii) $K=10$ with $\varrho = .1$ (dotted).
}
\end{figure}
example[Quantile treatment effects, cont.]
For this experiment, I reuse the setup of Example (ref) but replace the variable $X_{i,j,k}$ with a cluster-level treatment indicator $D_j$ that equals one if cluster $j$ received treatment and equals zero otherwise. I randomly assign $q_1 = \lfloor q/2 \rfloor$ clusters to treatment and $q_0 = \lceil q/2 \rceil$ to control. The coefficient of interest is $\delta$ in
\begin{align*}
Q(u\mid D_{j}) = \beta_0(u) + \delta(u)D_j + \beta_2(u)Z_{i,j}.
\end{align*}
I do not assume that pairings are predetermined and therefore use the adjusted $p$-values of the CRK test from Algorithm (ref). For each Monte Carlo replication, I drew a collection $\mathcal{I}$ with $|\mathcal{I}| = 50$ from $\mathcal{H}$ without replacement. The CRK test with unknown cluster parings requires $\alpha = .05 > 1/2^{q_1\wedge q_0-1}$ to have power, which is satisfied here as long as $q\geqslant 12$. I therefore restrict $q$ to be between 12 and 20. All other parameters of the experiment are exactly as in Example (ref).
\begin{figure}
\resizebox{\textwidth}{!}{
\begin{tikzpicture}[x=1pt,y=1pt]
\definecolor{fillColor}{RGB}{255,255,255}
\path[use as bounding box,fill=fillColor,fill opacity=0.00] (0,0) rectangle (433.62,216.81);
\begin{scope}
\path[clip] ( 48.00, 42.00) rectangle (204.81,192.81);
\definecolor{drawColor}{RGB}{0,0,0}
\path[draw=drawColor,line width= 1.6pt,line join=round,line cap=round] ( 53.81, 48.70) --
( 71.96, 48.37) --
( 90.11, 51.83) --
(108.26, 50.04) --
(126.41, 52.28) --
(144.55, 50.94) --
(162.70, 54.96) --
(180.85, 52.28) --
(199.00, 53.95);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (433.62,216.81);
\definecolor{drawColor}{RGB}{0,0,0}
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 53.81, 42.00) -- (199.00, 42.00);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 53.81, 42.00) -- ( 53.81, 36.00);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 90.11, 42.00) -- ( 90.11, 36.00);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (126.41, 42.00) -- (126.41, 36.00);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (162.70, 42.00) -- (162.70, 36.00);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (199.00, 42.00) -- (199.00, 36.00);
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 53.81, 20.40) {12};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 90.11, 20.40) {14};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (126.41, 20.40) {16};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (162.70, 20.40) {18};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (199.00, 20.40) {20};
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00, 47.59) -- ( 48.00,187.22);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00, 47.59) -- ( 42.00, 47.59);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00, 75.51) -- ( 42.00, 75.51);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00,103.44) -- ( 42.00,103.44);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00,131.37) -- ( 42.00,131.37);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00,159.30) -- ( 42.00,159.30);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00,187.22) -- ( 42.00,187.22);
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 33.60, 47.59) {0.00};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 33.60, 75.51) {0.05};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 33.60,103.44) {0.10};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 33.60,131.37) {0.15};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 33.60,159.30) {0.20};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 33.60,187.22) {0.25};
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] ( 48.00, 42.00) --
(204.81, 42.00) --
(204.81,192.81) --
( 48.00,192.81) --
cycle;
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (216.81,216.81);
\definecolor{drawColor}{RGB}{0,0,0}
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.0] at (126.41,200.68) {CRK test (size, $\delta(u)\equiv 0$)};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (126.41, 2.40) {Number of clusters};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at ( 15.60,117.41) {Rejection frequency};
\end{scope}
\begin{scope}
\path[clip] ( 48.00, 42.00) rectangle (204.81,192.81);
\path[] (108.80,143.66) rectangle (204.74, 98.06);
\definecolor{drawColor}{RGB}{169,169,169}
\path[draw=drawColor,line width= 1.6pt,line join=round,line cap=round] (117.35,132.26) -- (134.45,132.26);
\path[draw=drawColor,line width= 1.6pt,dash pattern=on 7pt off 3pt ,line join=round,line cap=round] (117.35,120.86) -- (134.45,120.86);
\path[draw=drawColor,line width= 1.6pt,dash pattern=on 1pt off 3pt ,line join=round,line cap=round] (117.35,109.46) -- (134.45,109.46);
\definecolor{drawColor}{RGB}{255,255,255}
\node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.85] at (143.00,128.99) {$K=10, \varrho = $.5};
\node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.85] at (143.00,117.59) {$K=20, \varrho = $.5};
\node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.85] at (143.00,106.19) {$K=10, \varrho = $.1};
\path[] (108.87,145.91) rectangle (204.81, 88.91);
\definecolor{drawColor}{RGB}{0,0,0}
\path[draw=drawColor,line width= 1.6pt,line join=round,line cap=round] (117.42,134.50) -- (134.52,134.50);
\path[draw=drawColor,line width= 1.6pt,dash pattern=on 7pt off 3pt ,line join=round,line cap=round] (117.42,123.10) -- (134.52,123.10);
\path[draw=drawColor,line width= 1.6pt,dash pattern=on 1pt off 3pt ,line join=round,line cap=round] (117.42,111.70) -- (134.52,111.70);
\path[draw=drawColor,line width= 0.4pt,dash pattern=on 4pt off 4pt ,line join=round,line cap=round] (117.42,100.30) -- (134.52,100.30);
\node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.85] at (143.07,131.23) {$K=10, \varrho = $.5};
\node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.85] at (143.07,119.83) {$K=20, \varrho = $.5};
\node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.85] at (143.07,108.43) {$K=10, \varrho = $.1};
\node[text=drawColor,anchor=base west,inner sep=0pt, outer sep=0pt, scale= 0.85] at (143.07, 97.03) {5% level};
\path[draw=drawColor,line width= 1.6pt,dash pattern=on 7pt off 3pt ,line join=round,line cap=round] ( 53.81, 50.15) --
( 71.96, 49.04) --
( 90.11, 51.50) --
(108.26, 50.71) --
(126.41, 53.39) --
(144.55, 52.39) --
(162.70, 53.62) --
(180.85, 52.84) --
(199.00, 54.96);
\path[draw=drawColor,line width= 1.6pt,dash pattern=on 1pt off 3pt ,line join=round,line cap=round] ( 53.81, 48.70) --
( 71.96, 48.26) --
( 90.11, 51.61) --
(108.26, 48.81) --
(126.41, 52.28) --
(144.55, 50.27) --
(162.70, 53.39) --
(180.85, 51.50) --
(199.00, 53.84);
\path[draw=drawColor,line width= 0.4pt,dash pattern=on 4pt off 4pt ,line join=round,line cap=round] ( 48.00, 75.51) -- (204.81, 75.51);
\end{scope}
\begin{scope}
\path[clip] (264.81, 42.00) rectangle (421.62,192.81);
\definecolor{drawColor}{RGB}{0,0,0}
\path[draw=drawColor,line width= 1.6pt,line join=round,line cap=round] (270.62, 83.53) --
(288.77, 81.91) --
(306.92,115.79) --
(325.07,114.70) --
(343.22,136.14) --
(361.36,134.66) --
(379.51,151.28) --
(397.66,149.35) --
(415.81,161.64);
\end{scope}
\begin{scope}
\path[clip] ( 0.00, 0.00) rectangle (433.62,216.81);
\definecolor{drawColor}{RGB}{0,0,0}
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (270.62, 42.00) -- (415.81, 42.00);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (270.62, 42.00) -- (270.62, 36.00);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (306.92, 42.00) -- (306.92, 36.00);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (343.22, 42.00) -- (343.22, 36.00);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (379.51, 42.00) -- (379.51, 36.00);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (415.81, 42.00) -- (415.81, 36.00);
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (270.62, 20.40) {12};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (306.92, 20.40) {14};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (343.21, 20.40) {16};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (379.51, 20.40) {18};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (415.81, 20.40) {20};
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81, 47.59) -- (264.81,187.22);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81, 47.59) -- (258.81, 47.59);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81, 75.51) -- (258.81, 75.51);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81,103.44) -- (258.81,103.44);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81,131.37) -- (258.81,131.37);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81,159.30) -- (258.81,159.30);
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81,187.22) -- (258.81,187.22);
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (250.41, 47.59) {0.0};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (250.41, 75.51) {0.2};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (250.41,103.44) {0.4};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (250.41,131.37) {0.6};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (250.41,159.30) {0.8};
\node[text=drawColor,rotate= 90.00,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (250.41,187.22) {1.0};
\path[draw=drawColor,line width= 0.4pt,line join=round,line cap=round] (264.81, 42.00) --
(421.62, 42.00) --
(421.62,192.81) --
(264.81,192.81) --
cycle;
\end{scope}
\begin{scope}
\path[clip] (216.81, 0.00) rectangle (433.62,216.81);
\definecolor{drawColor}{RGB}{0,0,0}
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.0] at (343.21,200.68) {CRK test (power, $\delta(u)\equiv .5$)};
\node[text=drawColor,anchor=base,inner sep=0pt, outer sep=0pt, scale= 1.00] at (343.21, 2.40) {Number of clusters};
\end{scope}
\begin{scope}
\path[clip] (264.81, 42.00) rectangle (421.62,192.81);
\definecolor{drawColor}{RGB}{0,0,0}
\path[draw=drawColor,line width= 1.6pt,dash pattern=on 7pt off 3pt ,line join=round,line cap=round] (270.62,132.37) --
(288.77,128.35) --
(306.92,165.92) --
(325.07,165.89) --
(343.22,178.59) --
(361.36,178.93) --
(379.51,183.23) --
(397.66,184.01) --
(415.81,185.55);
\path[draw=drawColor,line width= 1.6pt,dash pattern=on 1pt off 3pt ,line join=round,line cap=round] (270.62,153.01) --
(288.77,153.15) --
(306.92,179.99) --
(325.07,180.10) --
(343.22,185.13) --
(361.36,185.38) --
(379.51,186.53) --
(397.66,186.72) --
(415.81,186.92);
\definecolor{drawColor}{RGB}{169,169,169}
\path[draw=drawColor,line width= 1.6pt,line join=round,line cap=round] (270.62, 83.22) --
(288.77, 81.13) --
(306.92,114.39) --
(325.07,112.41) --
(343.22,134.19) --
(361.36,132.32) --
(379.51,147.18) --
(397.66,146.14) --
(415.81,158.88);
\path[draw=drawColor,line width= 1.6pt,dash pattern=on 7pt off 3pt ,line join=round,line cap=round] (270.62,131.20) --
(288.77,128.49) --
(306.92,162.96) --
(325.07,163.12) --
(343.22,176.44) --
(361.36,176.97) --
(379.51,182.59) --
(397.66,182.53) --
(415.81,184.93);
\path[draw=drawColor,line width= 1.6pt,dash pattern=on 1pt off 3pt ,line join=round,line cap=round] (270.62,147.90) --
(288.77,148.74) --
(306.92,177.37) --
(325.07,177.70) --
(343.22,184.12) --
(361.36,184.43) --
(379.51,185.86) --
(397.66,185.83) --
(415.81,186.61);
\definecolor{drawColor}{RGB}{0,0,0}
\path[draw=drawColor,line width= 0.4pt,dash pattern=on 4pt off 4pt ,line join=round,line cap=round] (264.81, 54.57) -- (421.62, 54.57);
\end{scope}
\end{tikzpicture}
}
\caption{Rejection frequencies in Example (ref) of a true null (left) $H_0\colon \delta(u) = 0$ for all $u$ and false nulls (right) $H_0\colon \delta(u) = 0$ for $u > .5$ (grey) and $H_0\colon \delta(u) = 0$ for all $u$ (black) as a function of the number of clusters for the CRK test when cluster pairings are not known with (i) $K=10$ neighborhoods per cluster with intra-neighborhood correlation $\varrho = .5$ (solid lines), (ii) $K=20$ with $\varrho = .5$ (long-dashed), and (iii) $K=10$ with $\varrho = .1$ (dotted).
}
\end{figure}
The left panel of Figure (ref) shows the rejection frequencies of a true null hypothesis $H_0\colon \delta(u) = 0$ for all $u$ in 5,000 Monte Carlo experiments per horizontal coordinate as $q$ increases. I again considered (i) $K=10$ neighborhoods per cluster with intra-neighborhood correlation $\varrho = .5$ (solid lines), (ii) $K=20$ with $\varrho = .5$ (long-dashed), and (iii) $K=10$ with $\varrho = .1$ (dotted). As can be seen, adjusting the CRK test for unknown cluster pairings results in a markedly more conservative test relative to an unadjusted test from Figure (ref). However, as the right panel of Figure (ref) shows, this did not translate into poor power under the alternative. When I repeated the experiment with $\delta(u)\equiv .5$, the CRK test with identification across clusters had no problem detecting that neither $H_0\colon \delta(u)$ for all $u\in (0,1)$ (black) nor $H_0\colon \delta(u)$ for $u > .5$ (grey) were true. Compared to Example (ref), the alternative where $\mathcal{U} = (0,1)$ rejects slightly more nulls because now every $u$ provides evidence against the null.
A noteworthy feature of the right panel of Figure (ref) is the “zig-zag” pattern in the rejection frequencies. The reason for this pattern is the treatment assignment mechanism. If $q=12$, then $q_1 = 6$ clusters receive treatment and $q_0 = 6$ do not. If $q=13$, then again $6 = \lfloor 13/2\rfloor$ clusters receive treatment but now $7 = \lceil 13/2 \rceil$ do not. Algorithm (ref) uses a large number of potential pairings of treatment to control for inference but effectively reduces the number of clusters to $\min\{q_1,q_0\}$. In this experiment, inference with $6+7$ clusters is therefore effectively the same as inference with $6+6$ clusters, which explains the similar performance of the test at $q$ and $q-1$ when $q$ is odd.
I also experimented with alternative methods for combining the $p$-values in Algorithm (ref). For this, I repeated the experiment in the right panel of Figure (ref) (not shown) but replaced the left-hand side of (ref) with either the standard Bonferroni correction or the geometric average of the $p$-values times $\exp(1)$. The latter method is due to mattner2012 and discussed in detail in vovkwang2020.
At $q=12$ with $K=20$ and $\varrho = .5$, Bonferroni rejected the false null $H_0\colon \delta(u) = 0$ for all $u$ in none of the 5,000 simulations, mattner2012's method rejected in 39.44% of all simulations, and Algorithm (ref) rejected in 60.76% of all simulations in the same data. At $q=20$, mattner2012's method and Algorithm (ref) rejected in about 99% of all cases. Bonferroni improved substantially and rejected in 96% of all cases. I conducted a large number of additional experiments but the results remained the same. The Bonferroni method had by far the lowest power. Relative to Algorithm (ref), mattner2012's method had significantly lower power when the number of clusters was small but both alternatives caught up as this number increased. Neither Bonferroni nor mattner2012's method improved the power of Algorithm (ref) in any of my experiments.
Finally, I compare Algorithm (ref) to the ideal situation where an equal number of treated and control clusters are paired through a pilot experiment or pre-analysis plan. For this, I repeated the experiment in Figure (ref) with a single, randomly chosen pairing. At $q_1 = q_0 =8$, $K=10$, and $\varrho = .5$, a test with pre-specified pairs rejected a false $H_0\colon \delta(u) = 0$ for all $u$ in 82.16% of all cases as opposed to the 63.40% achieved by Algorithm (ref). In the same experiment with $q_1 = q_0 =10$, the test with pre-specified pairs rejected 91.90% of all false nulls and Algorithm (ref) rejected 81.70% of all false nulls. However, there are 8! = 40,320 potential ways of paring $q_1 = 8$ treated clusters and $q_0=8$ control clusters. If $q_1=q_0=10$, there are 3,628,800 ways. Each separate set of pairs could be potentially selected for the test. If a researcher did not pre-specify pairs and instead searched over potential pairs to discover a significant result, the test would quickly lose size control. At $q_1 = q_0 =6$, $K=10$, and $\varrho = .5$, a correct null hypothesis was erroneously rejected in 8.14% of all cases when the lowest $p$-value among three potential pairings was used. If the researcher chooses the lowest $p$-value among ten potential matches, then a non-existent effect showed up as statistically significant in 13.96% of all cases in a test with 5% nominal level. Algorithm (ref) does not choose among these tests and completely avoids this loss of size control.
\leavevmode\unskip\penalty9999 \hbox\nobreak
\quad\hbox{\ensuremath{\square}}
example[Placebo interventions in Project STAR]
In this example, I revisit a challenging placebo exercise of hagemann2017 in data from the first year of the Tennessee Student/Teacher Achievement Ratio experiment, known as Project STAR. Details about the data can be found in \citetalias{wordetal1990} and graham2008. I only provide a brief summary.
In 1985, incoming kindergarten students in 79 project schools were randomly assigned to small classes (13-17 students) or regular-size classes (22-25 students) with or without a teacher's aide.
Each of the project schools was required to have at least one of each kindergarten class type. The outcome is standardized student performance on the Stanford Achievement Test (SAT) in mathematics and reading administered at the end of the school year. The raw test scores are standardized as in krueger1999. He finds across several mean regression models that students in small classes perform about five percentage points better on average than students in regular classrooms. (Assigning teachers aides had no effect uniformly across specifications and I do not consider such classes in the following.) jacksonpage2013 and hagemann2017 document similar effects in quantile regressions but show that the effects are smaller for students near the bottom and the very top of the conditional outcome distribution and larger near the center of the distribution. For example, in the model
\begin{align}
Q_{Y_{i,j}}(u\mid X_{i,j}) = \beta_0(u) + \delta(u) \mathit{small}_{i,j} + \beta_2(u)^T Z_{i,j}
\end{align}
where the treatment dummy $\mathit{small}$ indicates whether the student was assigned to a small class and $Z$ contains school dummies, the effect of being in a small class relative to a regular class varies between 2.78 percentage points at the 10th percentile to 7.23 percentage points at the 60th percentile. jacksonpage2013 hypothesize that this heterogeneity could be attributed to varying levels of student motivation to take advantage of increased individual attention from a teacher.
For the placebo experiment, I removed all small classes from the sample and only kept the 16 schools that had two regular-size classes without aide. In each of these 16 schools, I then randomly assigned one of the regular-size classes the treatment indicator $\mathit{small}=1$. This mimics the random assignment of class sizes within schools in the original sample, even though in this case no student actually attended a small class. I clustered at the classroom level and applied the CRK test as in Algorithm (ref) by running 16 separate quantile regressions, one for each school, on a constant and $\mathit{small}$ to get 16 separate estimates of $\delta$. The fixed effects as in (ref) are not needed here because the constant can vary freely by school in these quantile regressions. Algorithm (ref) applies here because each school in this experiment has only one class with $\mathit{small} = 1$ and one class with $\mathit{small} = 0$. (If multiple small classes per school were available, then Algorithm (ref) could be used instead.) For the wild gradient bootstrap, I reran the QR in (ref) in the placebo data and again clustered at the classroom level. For both methods, I tested at the 5% level the correct null hypothesis that $H_0\colon \delta(u) = 0$ jointly at $u \in \{.1, .2, \dots, .9 \}$ against the alternative that $\delta$ is positive.
The rejection frequencies in `size' column in Table (ref) show the outcome of repeating the placebo assignment 1,000 times. As can be seen, the CRK test provided a nearly exact test but the bootstrap over-rejected somewhat. The over-rejection for the bootstrap here was documented by hagemann2017 and
can be attributed to the very small number of clusters available in the placebo sample vis-\`a-vis the large number of clusters needed for the consistency of the bootstrap.
\begin{table}\caption{Rejection frequencies of $H_0\colon \delta(u) = 0$ for all $u$ in placebo interventions in Project STAR for the CRK test and the wild gradient bootstrap at 5% level}
{
\scalebox{1}{
\begin{tabular}{lp{0cm}cp{0cm}cccccc}
\hline
& & size & &\multicolumn{6}{c}{power} \\ \cline{3-3}\cline{5-10}
& &$\delta = 0$ & &$\delta = 2$ & $\delta = 3$ &$\delta = 4$ & $\delta = 5$ & $\delta = 6$ & $\delta = 7$\\ \hline
CRK test & &.043 & &.122 & .161 &.212 & .318 & .379 & .478 \\
Bootstrap & &.091 & &.233 & .316 &.428 & .580 & .691 & .814 \\ \hline
\end{tabular}
}}
\end{table}
I also investigated power by increasing the percentile scores of all students in the randomly drawn small classes of the placebo experiment by $\delta \in \{2,3,4,5,6 ,7\}$ percentage points. These increases are of the same or smaller magnitude as the estimated quantile treatment effects in the actual sample. Then I tested the incorrect hypothesis $H_0\colon \beta_1(u) = 0$ for all $u$ with the same experimental setup as before. The results are shown in `power' column of Table (ref). As can be seen, the CRK test was able to reliably detect effects for moderate deviations from the null hypothesis. The wild gradient bootstrap rejected more often, but this was likely caused by its tendency to over-reject in this data set.
\leavevmode\unskip\penalty9999 \hbox\nobreak
\quad\hbox{\ensuremath{\square}}
Conclusion
I introduce a generic method for inference on quantile and regression quantile processes in the presence of a finite number of large and arbitrarily heterogeneous clusters. The method asymptotically controls size by generating statistics that exhibit enough distributional symmetry such that randomization tests can be applied. This randomization test can even be asymptotically similar in empirically relevant situations. The test does not require ex-ante matching of clusters, is free of user-chosen parameters, and performs well at conventional significance levels with as few as five clusters. The main focus on the paper is inference on quantile treatment effects and quantile difference in differences but the method applies more broadly. Numerical examples and an empirical application are provided.