EconBase
← Back to paper

Some Impossibility Results for Inference With Cluster Dependence with Large Clusters

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

105,084 characters · 0 sections · 65 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
center[center omitted — 128 chars of source]
center[center omitted — 144 chars of source]

\fontsize{13}{14} \selectfont

abstract{ This paper focuses on a setting with observations having a cluster dependence structure and presents two main impossibility results. First, we show that when there is only one large cluster, i.e., the researcher does not have any knowledge on the dependence structure of the observations, it is not possible to consistently discriminate the mean. When within-cluster observations satisfy the uniform central limit theorem, we also show that a sufficient condition for consistent $\sqrt{n}$-discrimination of the mean is that we have at least two large clusters. This result shows some limitations for inference when we lack information on the dependence structure of observations. Our second result provides a necessary and sufficient condition for the cluster structure that the long run variance is consistently estimable. Our result implies that when there is at least one large cluster, the long run variance is not consistently estimable.} { Key words. Consistent Discrimination; Local Dependence; Unknown Dependence Structure; Consistent Estimation of Long-Run Variance; Cluster Dependence; Log Likelihood Process \ } { JEL Classification: C01, C12, C13}

\@startsection{section}{1} \z@{1.0\linespacing\@plus\linespacing}{.8\linespacing}{Introduction}

Statistical inference from data usually begins by imposing a form of a dependence structure on the data, by specifying which groups of observations exhibit strong within-group dependence. Various tools of asymptotic inference such as the law of large numbers and the central limit theorem are available for many typically imposed dependence structures. A standard case is the independence assumption or an assumption on time series dependence. However, it is well known that in the case of cross-sectional dependence, a researcher is often less confident about the correctness of the dependence structure used, despite its crucial role for inference.

A popular way to deal with this challenge is to use cluster dependence modeling, where the dependence structure among observations within each cluster is left unspecified, while independence is imposed between observations from different clusters. The inference procedures when there are many clusters are well known and can be analyzed using standard methods of asymptotic inference. However, less is known about the case where there are large clusters, and the dependence structure within such a cluster is unknown. Cameron/Gelbach/Miller:08:ReStat proposed a wild bootstrap procedure and showed by simulations that their tests perform well even when there are a small number of clusters. The robustness of this result was confirmed by MacKinnon/Webb:17:JAE even when the sizes of the clusters are highly heterogeneous. This large cluster issue has also drawn interest in the literature of difference-in-differences when there are only few treated clusters (see Conley/Taber:11:ReStat, Hagemann:19:JOE, and MacKinnon/Webb:20:JOE, and references therein). Djogbenou/MacKinnon/Nielsen:19:JOE studied inference on regression models with clustered errors. They provided conditions for the cluster sizes so that asymptotic and bootstrap inferences are asymptotically valid. They showed that their conditions exclude the presence of a large cluster.

There are several methods proposed to deal with the problem of inference with large clusters. Donald/Lang:07:ReStat and Bester/Conley/Hansen:11:JOE considered linear models and proposed inference where the asymptotic distribution of the long run variance estimator is fully known. This approach is related to the HAR (Heteroskedasticity-Autocorrelation Robust) inference of Kiefer/Vogelsang:2002:Eca and Sun:14:Eca in time series, which uses a normalization by an inconsistent long run variance estimator that has a stochastic limit.

Ibragimov/Muller:10:JBES proposed a $t$-test approach based on within-cluster estimators together with a $t$-distribution, where the degree of freedom in the $t$ distribution is determined by the number of clusters. They used the result of Bakirov/Szekely:05:ZNS and showed that their approach is asymptotically valid, even if the variances of the cluster specific estimators are different across the clusters. Ibragimov/Muller:16:ReStat extended these results to the problem of two-sample comparison and developed a testing procedure for the level of clustering.

Some studies adopted the approach of randomized testing to deal with cluster dependence with large clusters. Canay/Romano/Shaikh:17:Eca developed asymptotic inference procedures when the inference involves statistics whose limiting distribution satisfy symmetry properties. Hagemann:19:JOE proposed randomized tests for treatment effects when there are only a small number of clusters. Like Ibragimov/Muller:10:JBES, both proposals assumed large sample properties for within-cluster statistics. A recent work by Canay/Santos/Shaikh:21:ReStat use the analogue between wild bootstrap and randomized tests, and provided conditions under which the wild bootstrap for cluster-dependent regression models is asymptotically valid when there are only a small number of clusters.

Our paper focuses on observations with a cluster dependence structure and explores implications on statistical inference when there are large clusters. First, we show that when the sample consists of large clusters, the mean cannot be consistently discriminated if there is only one cluster, i.e., the researcher does not have any knowledge on the dependence structure of the data. Furthermore, when the observations form large clusters and within-cluster observations satisfy the uniform central limit theorem, a sufficient condition for the mean to be consistently discriminated at the rate of $\sqrt{n}$ is that the sample consists of at least two large clusters.

This impossibility result has a significant implication in a setting where the researcher does not know the dependence structure of observations. In such a case, consistent discrimination of the mean is not possible with uniform-in-$P$ asymptotic size control. Note that Song:16:arXiv proposed a randomized subsampling approach, and Leung:21:JAE provided a set of general conditions for the approach to produce asymptotically valid inference. Both focus on a setting where no knowledge on the dependence structure is required. Among other things, their results show that the mean is consistently discriminated. Our impossibility result on consistent discrimination considers a setting where there is no uniform upper bound of the long run variance in the null model, and this setting is excluded by part of their conditions. Hence, their results do not contradict our impossibility result.

Our second result is concerned with consistent estimation of long run variances. More specifically, suppose that $X_n = [X_{n,1},...,X_{n,n}]^\top$ is a given random vector of dimension $n$, where each observation $X_{n,i}$ has the same mean $\mu$. Let us define the long-run variance of $X_n$ as follows:\footnote{Note that when there is a common shock, say, $C_n$, such as cluster-specific fixed effects with few clusters, the analysis in this paper carries over to this case with $\sigma_{LR}^2$ replaced by the conditional variance given common shock $C_n$. Our impossibility results do not depend on whether there is a common shock of this form in the data or not. For simplicity, we consider a setting without such cluster-specific fixed effects.}

align[align omitted — 121 chars of source]

Recently, Hansen/Lee:19:JOE derived an asymptotic distribution theory for clustered data, including a law of large numbers and a central limit theorem. One of their results presents a condition for the cluster structure that is necessary and sufficient for the weak law of large numbers to hold for the sample average of the clustered observations. Our paper shows that the same condition is in fact necessary and sufficient for the consistent estimability of the long run variance of the clustered observations as well. Our condition for the cluster structure also implies that when there is at least one large cluster, i.e., the researcher does not know the dependence structure on a nonnegligible portion of the data, the long-run variance is not consistently estimable. It is not hard to show that the existing cluster-robust variance estimators are inconsistent when the cluster structure is severely misspecified. However, to the best of our knowledge, it has not been known whether there exists any consistent estimator of the long run variance when there is a lack of knowledge on the dependence structure on a nonnegligible portion of the data. Our result gives a negative answer to this question.

There has long been a strand of literature that studies impossibility of estimation and inference. (See, e.g., Bahadur/Savage:56:AMS, Dufour:97:Eca, Potscher:02:Eca, Bertanha/Moreira:20:JOE.) The impossibility of consistent estimation of a long run variance in this paper is related to Potscher:02:Eca who established a minimax risk lower bound for a general estimation problem. Among others, his result can be used to prove the impossibility of consistent estimation uniform in $P$ as shown in Corollary 3.2 there. However, we cannot apply this corollary in our setting, because our probability model is not indexed by a set of parameters fixed independently of the sample size, such as $\mathscr{H}$ in his paper. This stems from our setting where we have to deal with the joint distribution of the entire sample whose dependence structure varies in the model as $n$ changes. Bertanha/Moreira:20:JOE studied impossibility results of two types: indistinguishability of the null hypothesis from the alternative hypothesis and unbounded confidence sets. Their study of impossibility of the first type is related to impossibility of consistent discrimination of the mean in our paper. For this result, they assume that for each probability in the alternative hypothesis, there is a sequence of probabilities under the null hypothesis that weakly converge to this probability. Our setting does not satisfy this assumption in general. Hence, our result does not fall into their framework. Menzel:21:Eca recently developed and verified the validity of a bootstrap procedure in multi-way clustered observations with two or more dimensions. Part of his results shows that it is not possible to consistently estimate the distribution of the cluster dependent observations. Our results are not the special case of his results, because our impossibility result holds for models that exclude the counterexample that he used to prove the impossibility result. In particular, our cluster dependence accommodates within-cluster heterogeneity in terms of marginal distributions and dependence structures.

The rest of the paper is organized as follows. The next section studies the consistent discrimination of the mean. Section 3 is devoted to presenting the result of the impossibility of consistent estimation of the long run variance. In Section 4, we illustrate the implication of our results for the case of difference-in-difference models. In Section 5, we conclude. The mathematical proofs are found in the appendix.

\@startsection{section}{1} \z@{1.0\linespacing\@plus\linespacing}{.8\linespacing}{Cluster Dependence}

Let $X_n = [X_{n,1},...,X_{n,n}]^\top$ be a random vector with a joint distribution $P_n$ which belongs to the class of distributions $\mathcal{P}_n$. Throughout the paper, we assume that for each $P_n \in \mathcal{P}_n$,

align*[align* omitted — 87 chars of source]

and $\sigma_{LR}^2 < \infty$, where $\sigma_{LR}^2$ is defined in ((ref)). In many situations, the dependence structure is partially observed. Here we consider cluster dependence, where the dependence structure is entirely unknown within each cluster, and observations are independent between clusters. Let $N_{n,m}, m = 1,...,M_n$, be a partition of $N_n = \{1,...,n\}$ such that $|N_{n,m}| = n_m$ for each $m=1,...,M_n$, so that $\sum_{m=1}^{M_n} n_m = n$. Define $\mathcal{M}_n = \{N_{n,m}: m = 1,...,M_n\}$ and call it a \bi{cluster structure}. Throughout the paper we assume that $(X_{n,i})_{i \in N_{n,m}}$ are independent across $m$'s under all $P_n \in \mathcal{P}_n$, i.e., the joint distribution of $X_n$ has a cluster dependence structure. For future references, we define

align*[align* omitted — 188 chars of source]

so that $\overline X_{m,n}$ represents the within-cluster mean of $X_{n,i}$'s and $\sigma_{n,m}^2$ represents the within-cluster long-run variance of $X_{n,i}$'s.

Our impossibility results rely on the assumption that the probability model, $\mathcal{P}_n$, includes Gaussian experiments with what we call local-to-independence common shocks. For each cluster $m=1,...,M_n$, and for $\delta > 0$ and $\sigma^2>0$, we define

eqnarray*[eqnarray* omitted — 185 chars of source]

where $I_{n_m}$ denotes the $n_m$-dimensional identity matrix and $\mathbf{1}_{n_m}$ is the $n_m$-dimensional column vector of ones. Let $\Sigma_n(\sigma^2,\delta)$ be the $n \times n$ block diagonal matrix whose $m$-th block is given by $\Sigma_{n,m}(\sigma^2,\delta)$. Suppose that $\Sigma_n(\sigma^2,\delta)$ is positive definite. Then, for each $\mu_n \in \mathbf{R}^n$, we denote $\Phi(\mu_n,\Sigma_n(\sigma^2,\delta))$ the multivariate normal distribution with mean $\mu_n$ and covariance matrix $\Sigma_n(\sigma^2,\delta)$. Define

eqnarray*[eqnarray* omitted — 199 chars of source]

The set $\mathcal{P}_{n,\mathcal{N}}$ represents a set of Gaussian models, where each member is a multivariate normal distribution with a common mean and an equal covariance. We call each $\Phi(\mu_n,\Sigma_n(\sigma^2,\delta))$ the \bi{local-to-independence common shock (LTIC) Gaussian distribution} with parameters $\sigma^2$ and $\delta$. This Gaussian distribution represents the cross-sectional dependence structure of $X_{n,i}$'s generated as follows:

eqnarray*[eqnarray* omitted — 159 chars of source]

where $\mu_{n,i}$ is the $i$-th entry of $\mu_n$, $\varepsilon_i$'s are i.i.d.\ normal random variables with mean zero and variance $\sigma^2$, and $\eta_m$, $m=1,...,M_n$, are i.i.d.\ normal random variables with mean zero and variance $\sigma^2$, independent of $\varepsilon_i$'s. Each random variable $\eta_m$ represents a within-cluster “common shock”, and creates the within-cluster global dependence among $X_{n,i}$'s. The influence of this common shock on the random variable $X_{n,i}$ diminishes at the rate of $\sqrt{n_m}$.

\@startsection{section}{1} \z@{1.0\linespacing\@plus\linespacing}{.8\linespacing}{Consistent Discrimination of the Mean}

\@startsection{subsection}{2} \z@{.8\linespacing\@plus.7\linespacing}{.7\linespacing}{Consistent Discrimination of the Mean}

Let us explore the consistent discrimination of the mean under the general cluster dependence structure. We introduce the notion of consistent discrimination formally. Let $\mathcal{P}_n$ be a set of the distributions of $X_n \in \mathbf{R}^n$ such that $\mathbf{E}[X_{n,i}]$ is identical across $i$ for each $n \ge 1$. Let $\mathcal{P}_{n,0} = \{P_n \in \mathcal{P}_n: \mathbf{E}[X_{n,i}] = 0\}$, i.e., the set of probabilities under the null hypothesis of $\mathbf{E}[X_{n,i}] = 0$.

definitionThe mean of $X_{n,i}$, $n \ge 1$, is \bi{consistently discriminated at level $\alpha \in (0,1)$ in model $\mathcal{P}_n$}, if there is a sequence of (potentially randomized) tests $\{\varphi_n\}_{n \ge 1}$ such that \begin{align*} \limsup_{n \rightarrow \infty} \mathbf{E}[\varphi_n(X_n)] \le \alpha, \end{align*} along any sequence $P_{n,0} \in \mathcal{P}_{n,0}$, and \begin{align*} \liminf_{n \rightarrow \infty} \mathbf{E}[\varphi_n(X_n)] = 1, \end{align*} along any sequence $P_n \in \mathcal{P}_n$ such that $\liminf_{n \ge 1}\mathbf{E}[X_{n,i}] /\sigma_{LR} >0$.

The following theorem shows that when the sample consists of nonnegligible clusters, a necessary condition for the consistent discrimination of the mean is that there exist at least two clusters.

theoremSuppose that $\mathcal{P}_{n,\mathcal{N}} \subset \mathcal{P}_n$ for each $n \ge 1$. Suppose further that $\alpha \in (0,1/2)$, and $M_n = 1$ for each $n \ge 1$. Then, the mean cannot be consistently discriminated at level $\alpha$.

The theorem implies that when we do not know the local dependence structure of the random variables (i.e., $M_n = 1$), it is not possible to consistently discriminate the mean.

\@startsection{subsection}{2} \z@{.8\linespacing\@plus.7\linespacing}{.7\linespacing}{Consistent $\sqrt{n}$-Discrimination of the Mean}

We introduce the following notion of consistent $\sqrt{n}$-discrimination.

definitionThe mean of $X_{n,i}$, $n \ge 1$, is \bi{consistently $\sqrt{n}$-discriminated at level $\alpha \in (0,1)$ in model $\mathcal{P}_n$}, if there is a sequence of (potentially randomized) tests $\{\varphi_n\}_{n \ge 1}$ such that \begin{align*} \limsup_{n \rightarrow \infty} \mathbf{E}[\varphi_n(X_n)] \le \alpha, \end{align*} along any sequence $P_n \in \mathcal{P}_{n,0}$, and \begin{align*} \liminf_{n \rightarrow \infty} \mathbf{E}[\varphi_n(X_n)] = 1, \end{align*} along any sequence $P_n \in \mathcal{P}_n$ such that $\lim_{n \rightarrow \infty} \sqrt{n} \mathbf{E}[X_{n,i}]/\sigma_{LR} = \infty$.

We consider consistent discrimination against alternatives after normalizing by $\sigma_{LR}$ (which depends on $n$), so that when $\sigma_{LR}^2$ is larger, we focus on the alternative hypothesis that is farther away from the null hypothesis. Hence, if it is not possible to consistently $\sqrt{n}$-discriminate the mean, it is not necessarily due to the long run variance increasing to infinity fast.

The consistent $\sqrt{n}$-discrimination is often obtained when the parameter is in a finite dimensional space, and one knows the local dependence structure of the observations. To illustrate this point, suppose that $X_{n,1},...,X_{n,n}$'s are i.i.d. Then, often we have

align*[align* omitted — 124 chars of source]

where $\hat \sigma_n^2 \rightarrow_p \sigma^2 = \operatorname{Var}(X_{n,1}) >0$, and $\overline X_n = \frac{1}{n}\sum_{i =1}^n X_{n,i}$. Let us consider the usual $t$-test as follows:

align*[align* omitted — 118 chars of source]

Under the Pitman local alternatives such that $\mathbf{E}[X_{n,i}] = \overline \mu/\sqrt{n}$, $\overline \mu>0$, we have

align*[align* omitted — 140 chars of source]

where $\Phi$ denotes the CDF of $\mathcal{N}(0,1)$. The last term converges to 1 as $\overline \mu \rightarrow \infty$. Hence, the mean is consistently $\sqrt{n}$-discriminated. The discrimination results extend to the case with locally dependent observations where we know the local dependence structure and the long run variance is consistently estimable.

However, when we do not know the dependence structure, the consistent $\sqrt{n}$-discrimination of the mean is not guaranteed. We make this explicit in the following corollary which follows immediately from Theorem (ref).

corollarySuppose that $\mathcal{P}_{n,\mathcal{N}} \subset \mathcal{P}_n$. Suppose further that $\alpha \in (0,1/2)$, and $M_n = 1$. Then, the consistent $\sqrt{n}$-discrimination of the mean at level $\alpha$ is not possible.

On the other hand, if we have at least two large clusters and do not know the dependence structure within each cluster, we can consistently $\sqrt{n}$-discriminate the mean as long as the within-cluster sample means are asymptotically normal, as shown in the following theorem.

theoremSuppose that there exists a sub-partition $\mathcal{M}_n' \subset \mathcal{M}_n$ such that $|\mathcal{M}_n'| \ge 2$ for each $n \ge 1$, and \begin{align} \liminf_{n \rightarrow \infty} \min_{ N_{n,m} \in \mathcal{M}_n'} \frac{|N_{n,m}|}{n} >0. \end{align} Suppose further that the set $\mathcal{M}_n'$ satisfies that for each $P_n \in \mathcal{P}_n$ and for each $t \in \mathbf{R}$, \begin{align} \max_{1\le m \le M_n : N_{n,m} \in \mathcal{M}_n'}\left| P_n\left\{ \frac{\sqrt{n_m} (\overline X_{m,n} - \mathbf{E}[X_{n,i}])}{\sigma_{n,m}} \le t \right\} - \Phi(t) \right| \rightarrow 0, \end{align} as $n \rightarrow \infty$. Then, the mean is consistently $\sqrt{n}$-discriminated at level $\alpha \in (0,1)$.

For the theorem, we construct a $t$-test statistic as in Ibragimov/Muller:10:JBES and show that using the test, we can consistently $\sqrt{n}$-discriminate the mean, without knowing the dependence structure within the clusters.

The asymptotic normality condition ((ref)) is often satisfied if the within-cluster dependence is weak. As we show later, this does not mean that we can consistently estimate $\sigma_{n,m}$ for each cluster. (We will study this problem in the next section in detail.) Also, it is important to note that the within-cluster asymptotic normality is not enough to secure the consistent $\sqrt{n}$-discrimination of the mean, if there is only one cluster. In fact, the asymptotic normality condition alone does not exclude the possibility of $\mathcal{P}_{n,\mathcal{N}} \subset \mathcal{P}_n$, and in this case, Corollary (ref) shows that the mean is not consistently $\sqrt{n}$-discriminated.

As mentioned in the introduction, Song:16:arXiv and Leung:21:JAE considered the approach of randomized subsampling inference when one does not know the dependence structure at all. Hence, their situation corresponds to the setting with $M_n = 1$. Their procedure requires the following assumption:

align[align omitted — 100 chars of source]

as $n \rightarrow \infty$. If we know the upper bound of the long-run variance such that the upper bound does not change with $n$, it is not hard to see that we can consistently $\sqrt{n}$-discriminate the mean as long as the condition ((ref)) holds. Indeed, we can consider the test where we reject the null hypothesis of $\mathbf{E}[X_{n,i}] = 0$ against $\mathbf{E}[X_{n,i}] > 0$ if and only if

align*[align* omitted — 76 chars of source]

where $c$ is the known upper bound for the long-run variance. In our setting of hypothesis testing, however, the set of probabilities $\mathcal{P}_n$ does not have a finite upper bound for the long-run variance of the sample mean, reflecting the fact that the long-run variance is not known. Thus, the assumption ((ref)) does not hold uniformly over $P \in \mathcal{P}_n$ in our setting, and the results of Song:16:arXiv and Leung:21:JAE do not contradict the impossibility result of Theorem (ref).\footnote{This setting is analogous to that in the standard hypothesis testing with i.i.d.\ normal random variables with the unknown variance. In this standard setting, even with the unknown variance, the mean is typically consistently $\sqrt{n}$-discriminated, because the variance can be consistently estimated. However, as we will see later, in a setting with large clusters, the long-run variance is not consistently estimable.}

\@startsection{section}{1} \z@{1.0\linespacing\@plus\linespacing}{.8\linespacing}{Consistent Estimation of Variance}

Recently, Hansen/Lee:19:JOE showed that it is necessary and sufficient for the weak law of large numbers to hold for the sample average of the clustered observations with

align*[align* omitted — 108 chars of source]

In this section, we show that this condition is necessary and sufficient for consistent estimability of the long run variance. This implies that when there is a large cluster (i.e., which takes up an asymptotically nonnegligible fraction of observations), the long run variance is not consistently estimable. This is a consequence of lack of knowledge of the dependence structure within the large cluster. It means that the usual asymptotic inference based on the asymptotic normal approximation of statistics is generally not applicable in this situation.

\@startsection{subsection}{2} \z@{.8\linespacing\@plus.7\linespacing}{.7\linespacing}{Consistent Estimability}

Let us introduce the notion of consistent estimability of a parameter. Let $\mathcal{P}_n$ be the set of joint distributions of observed random variables, say, $\{X_1,...,X_n\}$. Given a parameter space $\Theta \subset \mathbf{R}^d$, we define our object of interest to be a map $\theta_n: \mathcal{P}_n \rightarrow \Theta$.

definitionFor any sequence of subsets $\mathcal{P}_n' \subset \mathcal{P}_n$, we say that $\theta_n$ is \bi{consistently estimable in} $\mathcal{P}_n'$, if there exists an estimator $\hat \theta$ such that along any sequence $P_n \in \mathcal{P}_n'$, \begin{align*} P_n\left\{\| \hat \theta - \theta_n(P_n) \| > \epsilon \right\} \rightarrow 0, \end{align*} as $n \rightarrow \infty$, for each $\epsilon>0$.

One can find a similar definition of consistent estimability in LeCam/Schwartz:60:AMS. They provide necessary and sufficient conditions for a parameter to be consistently estimable when the data are i.i.d. See also Section 1.4 of Ibragimov/Hasminskii:1981:StatEst and Section 6.2 of Pfanzagl:94:ParamStatTheory.

Our setting is somewhat nonstandard, requiring a different technique to prove impossibility of consistent estimation. It is usually assumed that the probability model is indexed by a certain set, i.e., $\mathcal{P}_n = \{P_{n,h}: h \in \mathcal{H}\}$, where each $P_{n,h}$ is a probability measure indexed by $h$ in some topological space $\mathcal{H}$ that is independent of the sample size $n$. One can then redefine the parameter $\psi_n(h) = \theta_n(P_{n,h})$, $h \in \mathcal{H}$, i.e., as a map on $\mathcal{H}$. As long as $\psi_n$ behaves “continuously” on $\mathcal{H}$, the parameter $\psi$ can be shown to be consistently estimable. (See, e.g., Theorem 4.1 of Ibragimov/Hasminskii:1981:StatEst and Theorem 6.2.11 of Pfanzagl:94:ParamStatTheory.) Then the impossibility of consistent estimation stems from the discontinuity of $\psi$ as a map on $\mathcal{H}$, which yields “non-identifiability” of the parameter Potscher:02:Eca.

However, we cannot apply this standard approach in our setting, because there is no natural space $\mathcal{H}$ that indexes $\mathcal{P}_n$ independently of $n$. The main reason is that we need to deal with a situation potentially with a large cluster with an unknown within-cluster dependence structure. This means that we need to require our probability model to accommodate a wide range of dependence structures for the entire sample. For example, suppose that there is only one large cluster, so that one does not know the dependence structure at all. This means, among other things, that our model needs to include various network dependence structures Kojevnikov/Marmer/Song:21:JOE for the joint distribution of the entire random vector $[X_1,...,X_n]$ whose dimension grows with the sample size $n$. One might consider parametrizing the probabilities in terms of the networks governing the dependence structure, but each network depends on the sample size $n$. To the best of our knowledge, there is no obvious way to topologize such a probability model and to define the continuity of the parameter $\theta_n$ on the probabilities, independently of sample size $n$.

Our approach relies on the following simple lemma that uses contiguity of probabilities at a primitive level. For any two sequences of probabilities $P_n$ and $P_n'$, we say that $P_n$ is \bi{contiguous with respect to} $P_n'$ if $P_n' (A_n) \rightarrow 0$ implies $P_n (A_n) \rightarrow 0$ for any sequence of Borel sets $A_n$, and write $P_n \triangleleft P_n'$. When $P_n \triangleleft P_n'$ and $P_n' \triangleleft P_n$, we say that $P_n$ and $P_n'$ are \bi{mutually contiguous}, and write $P_n \triangleleft \triangleright P_n'$. Contiguity between probabilities was introduced by LeCam:60:UCPS and is widely used, especially for deriving the limiting distribution of a test statistic under local alternatives. By tracing out the limiting distribution along a range of local alternatives, one obtains a limiting experiment which one can use to compute the asymptotic risk lower bound in statistical decision theory. (See, e.g., Chapter 6 of vanderVaart:98:AsympStat.)

The following lemma summarizes our scheme of proving the impossibility of consistent estimability of $\sigma_{LR}^2$.

lemmaSuppose that there exists a sequence $P_{n,0} \in \mathcal{P}_n$ such that $P_{n,1} \triangleleft P_{n,0}$ for every sequence $P_{n,1} \in \mathcal{P}_n$. Then, $\theta_n$ is consistently estimable in $\mathcal{P}_n$ if and only if $\lim_{n \rightarrow \infty} \left\| \theta_n(P_{n,1}) - \theta_n(P_{n,0}) \right\| = 0$ for any sequence $P_{n,1} \in \mathcal{P}_n$.

Later we use Lemma (ref) to prove the impossibility of consistent estimation of the long run variance, by selecting two Gaussian probabilities, $P_{n,0}$ and $P_{n,1}$, such that $P_{n,1} \triangleleft P_{n,0}$ and the values of the long run variance stay apart under $P_{n,0}$ and $P_{n,1}$ as $n \rightarrow \infty$. (See the discussion below Theorem (ref).)

The notion of consistent estimability in Definition (ref) coincides with consistent estimability uniform in $P$, i.e., the existence of an estimator $\hat \theta$ such that for each $\epsilon>0$, as $n \rightarrow \infty$,

align*[align* omitted — 129 chars of source]

(See Ibragimov/Hasminskii:1981:StatEst, p.31. See also Potscher:02:Eca for discussion on asymptotics uniform in $P$.) When $\mathcal{P}_{n} = \{P_{n,h}: h \in \mathcal{H}\}$ for some index set $\mathcal{H}$ which does not depend on the sample size $n$, uniform consistent estimability is stronger than pointwise consistent estimability which assumes the existence of an estimator $\hat \theta$ such that for each $h \in \mathcal{H}$ and for each $\epsilon>0$,

align*[align* omitted — 105 chars of source]

as $n \rightarrow \infty$. However, as explained above, in our setting, there is no space $\mathcal{H}$ that indexes $\mathcal{P}_n$ and is independent of $n$. Hence, there is no natural notion of pointwise consistency in $P$ in our set-up.

\@startsection{subsection}{2} \z@{.8\linespacing\@plus.7\linespacing}{.7\linespacing}{Consistent Estimability of Variance in Gaussian Experiments} \@startsection{subsubsection}{3} \z@{.5\linespacing\@plus.7\linespacing}{-.5em}{\normalfont}{A Necessary and Sufficient Condition for Consistent Estimability of Variance}

In our context, a major challenge is to show the contiguity condition (ii) of Lemma (ref). A standard argument proving contiguity utilizes the local asymptotic normality or local asymptotic mixed normality results for a log-likelihood process. However, these latter results often use an i.i.d.\ or time-series set-up where the researcher knows the dependence structure, and Hence, are not useful for our purpose here. For this reason, we focus on a Gaussian experiment, where we can explicitly compute the log-likelihood process in finite samples and investigate its asymptotic behavior as the dependence structure varies. In particular, we consider the following model for a fixed $\sigma^2>0$,

eqnarray*[eqnarray* omitted — 197 chars of source]

where $\mathbf{1}_n$ denotes the $n$-dimensional vector of ones. The set $\mathcal{P}_{n,\mathcal{N}}(\sigma^2)$ represents the set of LTIC Gaussian models, where each multivariate normal distributions with a common mean and the short run variance equal to $\sigma^2$.

For the impossibility result below, we require that the probability model does not exclude this Gaussian experiment.

theoremSuppose that $\mathcal{P}_{n,\mathcal{N}}(\sigma^2) \subset \mathcal{P}_n$ for each $n \ge 1$, for some $\sigma^2>0$ that is independent of $n$. Then, the long run variance $\sigma_{LR}^2$ is consistently estimable in $\mathcal{P}_n$ if and only if \begin{align} \lim_{n \rightarrow \infty} \sum_{m=1}^{M_n} \frac{n_m^2}{n^2} \rightarrow 0, \end{align} as $n \rightarrow \infty$.

The sufficiency part of the theorem is straightforward. To see this, suppose for simplicity that there is no singleton cluster in the data. If ((ref)) is satisfied, it means that the number of clusters $M_n$ grows to infinity as $n \rightarrow \infty$. Then, we consider the following estimator.

eqnarray*[eqnarray* omitted — 148 chars of source]

In fact, for the sufficiency part, we do not require that $\mathcal{P}_{n,\mathcal{N}}(\sigma^2) \subset \mathcal{P}_n$.

The nontrivial part of the theorem is to show that the condition ((ref)) is necessary for the consistent estimability of $\sigma_{LR}^2$ in $\mathcal{P}_{n,\mathcal{N}}$. Suppose that the condition of ((ref)) fails, which implies that one has at least one nonnegligible cluster. Then, we show that $\sigma_{LR}^2$ is not consistently estimable. For this, we employ Lemma (ref) after computing the log-likelihood process under cluster dependence. More specifically, suppose first that the observations consist of only a single large cluster. Then, we note that the model $\mathcal{P}_n$, due to the lack of knowledge on the dependence structure, does not exclude the LTIC Gaussian experiment: $\Phi(0,\Sigma_n(\sigma^2,\delta))$. Then we show that

align*[align* omitted — 92 chars of source]

whereas

align*[align* omitted — 154 chars of source]

as $n \rightarrow \infty$, for some nonzero constant $c$. Hence, by Lemma (ref), $\sigma_{LR}^2$ cannot be consistently estimated in any probability model that does not exclude the LTIC Gaussian experiment. It is not hard to extend the same arguments to a setting where there are potentially multiple large clusters.

Theorem (ref) then implies that if a nonnegligible portion of the sample belongs to non-singleton clusters, the long run variance $\sigma_{LR}^2$ is consistently estimable in $\mathcal{P}_{n,\mathcal{N}}$ if and only if the cluster structure consists of negligible clusters. We formalize this in the following corollary.

corollarySuppose that the conditions of Theorem (ref) hold, and $\liminf_{n \rightarrow \infty} n^*/n > 0$, where $n^* = \sum_{m=1: n_m \ge 2}^{M_n} n_m$. Then, the long run variance $\sigma_{LR}^2$ is consistently estimable in $\mathcal{P}_n$ if and only if the cluster structure $\mathcal{M}_n$ consists of negligible clusters, i.e., \begin{align} \lim_{n \rightarrow \infty} \max_{1 \le m \le M_n} \frac{n_m}{n} = 0. \end{align}

The condition $\liminf_{n \rightarrow \infty} n^*/n > 0$ requires that the fraction of random variables $X_{n,i}$ that do not belong to a singleton cluster is asymptotically nonnegligible. In this case, if the probability model in practice includes the Gaussian model $\mathcal{P}_{n,\mathcal{N}}$ as a subclass and there is at least one nonnegligible cluster, it is not possible to consistently estimate the long run variance. Certainly, this impossibility result carries over to a model where the long run variance $\sigma_{LR}^2$ is allowed to increase with the sample size $n$.

Hence, by combining Theorem (ref) with Theorem (ref), we find that when we have several large clusters, the long run variance is not consistently estimable because the sample contains large clusters, but the mean can still be consistently $\sqrt{n}$-discriminated.

\@startsection{subsubsection}{3} \z@{.5\linespacing\@plus.7\linespacing}{-.5em}{\normalfont}{Implications for Network Dependent Observations}

One might wonder whether the result extends to the case where the observations exhibit a dependence structure other than cluster dependence. Below we give a partial answer for the case of a dependency graph. Dependency graphs were introduced by Stein:86, and have been studied and used in statistics and econometrics. (See, e.g., Aronow&Samii:17, Song:18:ReStat, Leung:20:ReStat and Canen/Schwartz/Song:20:QE and references therein.)

A \bi{graph} (or network) is a pair $G_n = (N_n,E_n)$, where $N_n = \{1,...,n\}$ denotes the set of vertices and $E_n$ the set of edges, where we denote $N(i) = \{j: ij \in E_n\}$ to mean the neighborhood of vertex $i$. (Here, we consider only simple, undirected graphs, i.e., $ii \notin E_n$, for all $ i \in N_n$, and $ij \in E_n$ if and only if $ji \in E_n$.) We define

align*[align* omitted — 117 chars of source]

where $d_{mx}$ is called the \bi{maximum degree}, and $d_{av}$ the \bi{average degree} of the graph $G_n$. The maximum and average degrees are often used to capture the denseness of the graph. A subset of vertices in graph $G_n$ is called a \bi{clique} if any two distinct vertices are adjacent in $G_n$, and the number of vertices in the clique is called the size of the clique. The maximum clique size refers to the size of the clique that is largest in the graph $G_n$.

Recall that a graph $G_n = (N_n,E_n)$ on $N_n = \{1,...,n\}$ is called a \bi{dependency graph} for $X_n = (X_{n,i})_{i \in N_n}$, if for any subset $A \subset N_n$, $(X_{n,i})_{i \in A}$ and $(X_{n,i})_{i \in N_n \setminus \overline N_n(A)}$ is independent, where $\overline N_n(A) = \{j: ij \in E_n, \text{ for some } i \in A\} \cup \{i\}$. It is important to note that while the dependency graph imposes independence between $X_{n,i}$ and $X_{n,j}$ when they are not adjacent in the graph, it says nothing about dependence between them when they are adjacent. Thus, we allow in $\mathcal{P}_n$ any degree of dependence (including independence) between $X_{n,i}$ and $X_{n,j}$ whenever $i$ and $j$ are adjacent in $G_n$. As the dependency graph becomes denser, this reflects our limited knowledge on the dependence structure, similarly to large clusters in the cluster dependence case.

corollarySuppose that the conditions of Theorem (ref) hold and that for each $n \ge 1$, there exists a graph $G_n = (N_n,E_n)$ which has maximum degree $d_{mx}$, average degree $d_{av}$, maximum clique size $n_C$, and each distribution $P_n \in \mathcal{P}_n$ of $X_n$ has $G_n$ as a dependency graph. Then, the following holds. (i) If $\limsup_{n \rightarrow \infty} n_C/ n > 0$, the long run variance $\sigma_{LR}^2$ is not consistently estimable. (ii) If $\lim_{n \rightarrow \infty} d_{mx}^2 d_{av} /n = 0$, the long run variance $\sigma_{LR}^2$ is consistently estimable.

The impossibility result in (i) has an important implication in many models with network dependent observations. As in the case of a dependency graph, many models of network dependence do not specify the strength of dependence between observations that are adjacent in the network Kojevnikov/Marmer/Song:21:JOE. Weak dependence is usually imposed between observations that are far from each other in terms of the shortest path in the network. Hence, when there is a large clique in the network which constitutes a nonnegligible fraction of the entire sample in the limit as $n \rightarrow \infty$, Corollary (ref)(i) implies that the long run variance of the network dependent observations is not consistently estimable.

It is interesting to note that one cannot characterize a necessary and sufficient condition for the network solely in terms of its maximum degree. For example, if $X_n$ is a multivariate normal random vector such that each component has a bounded variance and has a star graph as a dependency graph, the long run variance is consistently estimable. To see this, let $X_n = [X_{n,1},...,X_{n,n}]^\top$ be a centered multivariate random vector which has a graph $G_n$ as a dependency graph. Let the graph $G_n$ be a star graph with the unit $1$ being its center.\footnote{The star graph as a dependency graph is different from an additive common shock model such as $X_{n,i} = C_n + \varepsilon_i$, where $C_n$ is a common shock, and $\varepsilon_i$'s are cross-sectionally independent idiosyncratic shock. In this case, the dependency graph is a complete graph, because every pair of random variables is correlated through the common shock. Hence, the center in the star graph as a dependency graph cannot be a source like a common shock. It is more plausible to imagine the center to be an aggregated outcome of independent sources. In this case, by simply eliminating the star, one obtains independent random variables.} In the context of multivariate normality, we can write

eqnarray*[eqnarray* omitted — 92 chars of source]

where the leading sum is the best linear projection, so that $\varepsilon$ is independent of $X_{n,i}$'s, $i \in N_n\setminus\{1\}$, which are independent from each other (due to the dependency graph being a star graph). Since the variance of $X_{n,1}$ is bounded, we should have $\sum_{i \in N_n \setminus \{1\}} \theta_i^2 < C$, for all $n \ge 1$, for some $C>0$. Note that we can identify

eqnarray*[eqnarray* omitted — 94 chars of source]

Now, we can write

align*[align* omitted — 155 chars of source]

The second term is written as

align*[align* omitted — 173 chars of source]

as $n \rightarrow \infty$, because the normalized sum in the parenthesis converges to zero in moments. Hence, we can simply take

eqnarray*[eqnarray* omitted — 73 chars of source]

to be an estimator of the long run variance. It is not hard to see that $\hat \sigma_{LR}^2$ is consistent for $\sigma_{LR}^2$. This example shows that one cannot express the condition for the consistent estimability solely in terms of the maximum degree of the dependency graph.

\@startsection{section}{1} \z@{1.0\linespacing\@plus\linespacing}{.8\linespacing}{Implications} \@startsection{subsection}{2} \z@{.8\linespacing\@plus.7\linespacing}{.7\linespacing}{A Linear Regression Model with Cluster-Dependent Errors} Let us consider the following regression model with cluster-dependent errors (see, e.g., Cameron/Gelbach/Miller:08:ReStat, Djogbenou/MacKinnon/Nielsen:19:JOE and Hansen/Lee:19:JOE and references therein):

align*[align* omitted — 31 chars of source]

where $y = [y_1',...,y_{M_n}']'$, $X = [X_1',...,X_{M_n}']'$ and $u = [u_1',...,u_{M_n}]'$, with $\mathbf{E}[u_m \mid X] = 0$ for each $m=1,...,M_n$, and each cluster $m$ has $n_m$ observations (so that $y_m$ and $u_m$ are $n_m$ dimensional column vectors, and $X_m$ is an $n_m \times k$ matrix.) We assume that $u_1,...,u_{M_n}$ are independent, but for each $m$, the dependence structure of $u_m$ is not known. We do not exclude the possibility that the error term follows a normal distribution.

Then, the OLS estimator of $\beta$ is given by

align*[align* omitted — 43 chars of source]

The sandwich form of the variance matrix of $\hat \beta$ is given by

align*[align* omitted — 121 chars of source]

Once we obtain a consistent estimator $\hat V$ of $V$, we can construct a standard error of the $j$-th entry of $\beta$, i.e., $\beta_j$, as $\hat \sigma_j^2 = [\hat V]_{jj}$, the $j$-th diagonal of $\hat V$. From the asymptotic normal inference applied to a $t$-statistic for $\beta_j$, we obtain the following confidence interval for $\beta_j$:

align[align omitted — 158 chars of source]

As for the consistent estimator $\hat V$, Djogbenou/MacKinnon/Nielsen:19:JOE considered the following estimator:

align*[align* omitted — 106 chars of source]

where $\hat u_m = y_m - X_m' \hat \beta$, and $d$ is a sequence such that $d \rightarrow 1$. They established the consistency of this estimator under a set of conditions, and showed that their conditions are not compatible with a setting in which one of the clusters is large, i.e., its size is proportional to the entire sample.

Our result implies that such an estimator $\hat V$ is not uniformly consistent when there is at least one large cluster. In fact, our result is much stronger than this. It shows that it is not possible to construct a uniformly consistent estimator of $V$ in such a case. Hence, in this case, we cannot construct a confidence interval of the form ((ref)) that is uniformly asymptotically valid. When a nonnegligible fraction of observations belong to a (non-singleton) cluster - which is the case with most cluster-dependence settings, the necessary and sufficient condition for the uniformly consistent estimability of $V$ is that each cluster is asymptotically negligible in the sense of ((ref)).

\@startsection{subsection}{2} \z@{.8\linespacing\@plus.7\linespacing}{.7\linespacing}{Difference-in-Differences with Spillovers} Let us explore the implications of the impossibility results in the context of a difference-in-differences approach to causal inference. (See Section 6.5 of Imbens/Wooldridge:09:JEL for an overview of this method. See also Roth/SantAnna/Bilinski/Poe:22:arXiv for an overview including recent advances in the literature.) Suppose that there are $n$ individuals who are subject to a treatment and the researcher observes their outcomes before and after the treatment. We let $Y_{i,t}(1)$ and $Y_{i,t}(0)$ denote the potential outcomes at time $t = 0,1$ for the treated state and the control state, respectively. As standard in the literature, we assume that in time 0, no individual is treated, and $Y_{i,0} = Y_{i,0}(0)$, which is observed. The observed outcome $Y_{i,1}$ at time $1$ is defined by

eqnarray*[eqnarray* omitted — 66 chars of source]

where $D_i$ is the indicator of treatment for $i$ that happens between times 0 and 1. Our parameter of interest is the average treatment effect on the treated:

eqnarray*[eqnarray* omitted — 89 chars of source]

Suppose that we have observations $\{(Y_{i,1},Y_{i,0},D_i)\}_{i=1}^n$, where $Y_{i,0}$ is the outcome for person $i$ at time 0. Furthermore, we assume that the researcher knows the probability $p = P\{D_i =1\}$. (The impossibility result we mention below carries over to the case where $p$ is not known.)

Let us introduce the standard parallel trend assumption used in the literature:

align*[align* omitted — 138 chars of source]

Under this assumption, we can identify

align*[align* omitted — 126 chars of source]

where $\Delta Y_{i} = Y_{i,1} - Y_{i,0}$. We can obtain a sample analogue estimator by

align*[align* omitted — 150 chars of source]

We consider settings where the observed outcomes are cross-sectionally dependent. Our interest is in constructing a confidence interval for $\text{ATT}$ that is uniformly asymptotically valid. Below we consider two situations, one with treatment spillover and the other with spillover of treatment effects. We explore implications of our impossibility results in these situations.

\@startsection{subsubsection}{3} \z@{.5\linespacing\@plus.7\linespacing}{-.5em}{\normalfont}{Treatment Spillover}

Suppose that there is a spillover of the treatments so that $D_i$'s are correlated across $i$, along some network among people. For example, one can think of a situation in a social program where two people $i$ and $j$ are neighbors and participating in the program by $i$ can induce the participation by $j$. Suppose that the researcher does not have information on the neighborhoods among the subjects. Then, this creates dependence among $Y_{i,1}$'s along a dependence structure that is unknown to the researcher. Then, our impossibility result shows that $\text{ATT}$ cannot be consistently discriminated.

In practice, the treatment assignment is often done at the cluster level, where the potential outcomes $Y_{i,1}(1)$ and $Y_{i,1}(0)$ may exhibit arbitrary dependence within each cluster. (See Section 5 of Roth/SantAnna/Bilinski/Poe:22:arXiv for examples and references studying such a setting.)

Suppose that we have at least two large clusters such that $(Y_{i,1}(1), Y_{i,0}(0), D_i)$ are independent across the clusters but arbitrarily correlated within each cluster. The researcher might attempt to test the null hypothesis of $\text{ATT} = 0$ by considering the usual $t$ statistic for testing the null hypothesis of $\text{ATT} = 0$ such that

align*[align* omitted — 88 chars of source]

where $\widehat{\sigma}^2$ is a consistent estimator of the variance of $\sqrt{n}(\widehat{\text{ATT}} - \text{ATT})$, and the critical values taken from the standard normal stable. Our impossibility result implies that it is not possible to consistently estimate the variance of $\sqrt{n}(\widehat{\text{ATT}} - \text{ATT})$, when there is at least one large cluster, and hence, such a $t$-test is not uniformly asymptotically valid. For the same reason, we cannot construct a confidence interval of the following familiar form:

align[align omitted — 180 chars of source]

such that the confidence interval is uniformly asymptotically valid. (See Section 5 of Roth/SantAnna/Bilinski/Poe:22:arXiv for various approaches.\footnote{To the best of our knowledge, there is no formal result that proposes a uniformly asymptotically valid confidence interval for $\text{ATT}$ in this setting. However, we expect that the bootstrap approach of Canay/Santos/Shaikh:21:ReStat can be used to construct a uniformly valid confidence interval under mild additional conditions.})

\@startsection{subsubsection}{3} \z@{.5\linespacing\@plus.7\linespacing}{-.5em}{\normalfont}{Spillover of Treatment Effects}

Suppose that the treatments $D_i$ themselves do not exhibit any spillover, but the cross-sectional dependence of $(Y_{i,1}(1), Y_{i,0}(0))$ arises due to the spillover of the treatment effects, for example, the treatment of a person $i$ influences the outcome of the person $j$ in the next period. Such a setting has been studied in the recent literature (see Aronow&Samii:17, Leung:20:ReStat, He/Song:22:WP and references therein.)

Suppose that the spillover of the treatment effects arises along some network among people, and yet the researcher does not have any information on the network. Then, our impossibility result implies that we cannot consistently discriminate ATT in such a situation. However, the researcher may observe a group structure where the spillover does not arise between groups, so that $(Y_{i,1}(1), Y_{i,0}(0), D_i)$ are independent across groups.

If each within-group sum of $(Y_{i,1}(1), Y_{i,0}(0), D_i)$ satisfies the central limit theorem, our result shows that the ATT can be consistently $\sqrt{n}$-discriminated. However, when there is at least one large group, there does not exist a consistent estimator of the variance of $\sqrt{n}(\widehat{\text{ATT}} - \text{ATT})$. Hence, similarly as before, we cannot construct a uniformly asymptotically valid $t$-test for the null hypothesis of $\text{ATT} = 0$ using the usual $t$ statistic and standard normal critical values, and cannot construct a uniformly asymptotically valid confidence interval of the form ((ref)) based on a normal approximation.

\@startsection{section}{1} \z@{1.0\linespacing\@plus\linespacing}{.8\linespacing}{Conclusion}

In this paper, we show two impossibility results on the inference on the mean when the dependence structure is not known. The first result is the impossibility of consistent estimation of the long run variance. The second result is the impossibility of the consistent $\sqrt{n}$-discrimination of the mean. We made an attempt to accommodate partial knowledge of the dependence structure through cluster dependence, and has obtained some necessary and sufficient conditions for the cluster structure for the impossibility results.

While cluster dependence is a popularly used specification of the cross-sectional dependence structure, it is not general enough to accommodate other forms of a dependence structure such as dependency graphs, Markov graphs, and network dependence. It would be interesting to investigate the implications of partial knowledge of a dependence structure for a more general setting. We leave this for future research.

\@startsection{section}{1} \z@{1.0\linespacing\@plus\linespacing}{.8\linespacing}{Appendix: Mathematical Proofs}

\@startsection{subsection}{2} \z@{.8\linespacing\@plus.7\linespacing}{.7\linespacing}{Preliminary Results}

For the proof of the main results, we first prove auxiliary lemmas. As a first step, we provide an explicit form of a log-likelihood process in Gaussian experiments in Lemma (ref). For this, we use the following auxiliary lemma.

lemmaLet $\Sigma_0 = U S U^\top$ be the spectral decomposition of an $n \times n$, symmetric positive definite matrix $\Sigma_0$ and let $\Sigma_1$ be an $n \times n$ matrix defined as \begin{eqnarray*} \Sigma_1 = \Sigma_0 + U A U^\top \end{eqnarray*} for some symmetric positive semidefinite matrix $A$. Let $B \Lambda B^\top$ be the spectral decomposition of $S^{-1/2} A S^{-1/2}$. Suppose that $|\lambda_i| < 1$ for all $i=1,...,n$, where $\lambda_i$ denote the $i$-th diagonal entry of $\Lambda$. Then the following results hold. (i) \begin{eqnarray} \log\left(|\Sigma_1|^{-1/2}\right) - \log\left(|\Sigma_0|^{-1/2}\right) = - \frac{1}{2}\sum_{i=1}^n \log\left( 1 + \lambda_i \right). \end{eqnarray} (ii) For any vectors $a,b \in \mathbf{R}^n$, \begin{eqnarray} a^\top(\Sigma_1^{-1} - \Sigma_0^{-1})b = -\sum_{i=1}^n \frac{\lambda_i}{1 + \lambda_i} \tilde a_i \tilde b_i, \end{eqnarray} where $\tilde a_i$ and $\tilde b_i$ are the $i$-th entries of $\tilde a$ and $\tilde b$, with \begin{eqnarray} \tilde a = B^\top S^{-1/2} U^\top a \quadand\quad \tilde b = B^\top S^{-1/2} U^\top b. \end{eqnarray}

Proof: Let $Q=U S^{1/2}$ such that $\Sigma_0 =Q B B^{\top}Q^{\top}$ and $\Sigma_{n,1}=Q B(I +\Lambda) B^{\top} Q^{\top}$. Thus,

align*[align* omitted — 97 chars of source]

and

align*[align* omitted — 166 chars of source]

where \[ \tilde{\Lambda} = \operatorname{diag}\left(\frac{\lambda_1}{1+\lambda_1},\ldots,\frac{\lambda_n}{1+\lambda_n}\right). \] $\blacksquare$

The following lemma provides an explicit form of a general log-likelihood process for Gaussian measures. Recall that $\Phi(\mu,\Sigma)$ denotes the multivariate normal distribution with mean vector $\mu$ and covariance matrix $\Sigma$.

lemmaLet $\Sigma_0$, $\Sigma_1, \Lambda, B, S$ and $U$ be the matrices in Lemma (ref). Then, for all $x, \mu_1, \mu_0 \in \mathbf{R}^n$, \begin{align*} \quad \quad \log \frac{d \Phi(\mu_1,\Sigma_1)}{d \Phi(\mu_0,\Sigma_0)}(x) = - \sum_{i=1}^n \log q_{i} + \frac{1}{2}\sum_{i=1}^n \frac{1}{q_i^2}\left(Z_i(x) (q_i + 1) - \tilde \mu_i \right)\left(Z_i(x) (q_i -1) + \tilde \mu_i \right), \end{align*} where $q_i = \sqrt{1 + \lambda_i}$, $\lambda_i$ is the $i$-th diagonal entry of $\Lambda$, $Z_i(x)$ is the $i$-th entry of $Z(x)$ and $\tilde \mu_i$ is the $i$-th entry of $\tilde \mu$ with \begin{eqnarray} Z(x) = B^\top S^{-1/2} U^\top(x - \mu_0) \quadand\quad \tilde \mu = B^\top S^{-1/2} U^\top (\mu_1 - \mu_0). \end{eqnarray}

Proof: We write

align[align omitted — 378 chars of source]

We apply Lemma (ref)(i) to the first term on the right hand side. As for the last term we let $\mu_{\Delta} = \mu_1 - \mu_0$, and $x_* = x - \mu_0$. Note that

align*[align* omitted — 268 chars of source]

Similarly, $\mu_{\Delta}^\top \Sigma_0^{-1} x_* = \tilde \mu^\top Z(x).$ We rewrite the last term in ((ref)) as

align*[align* omitted — 785 chars of source]

(by applying Lemma (ref)(ii)). By rearranging terms, we rewrite the last sum as

eqnarray[eqnarray omitted — 144 chars of source]

Combining this with an earlier result, we obtain the desired result. $\blacksquare$

lemmaLet $\Sigma_n = \sigma^2\left( (1 - \delta) I_n + \delta \mathbf{1}_n \mathbf{1}_n^\top \right)$, where $\delta$ is such that $n\delta \in (-1,1)$. Then, for any $\overline \mu \in \mathbf{R}$ and $x \in \mathbf{R}^n$, we have \begin{align*} \log \frac{d \Phi(\overline \mu \mathbf{1}_n,\Sigma_n)}{d \Phi(0,I_n)}(x) &= - \log \sqrt{\sigma^2(1 + (n-1) \delta)} - (n-1) \log \sqrt{\sigma^2(1 - \delta)}\\ & \quad + \frac{\sigma^2 (1 + (n-1) \delta) - 1}{2\sigma^2 (1 + (n-1) \delta)} Z_1^2(x) + \frac{\sigma^2 (1- \delta) - 1}{2 \sigma^2 (1 - \delta)} \sum_{k=2}^n Z_k^2(x)\\ & \quad + \frac{\sqrt{n}\overline \mu}{\sigma^2(1 + (n-1)\delta)} Z_1(x) - \frac{n \overline \mu^2}{2\sigma^2 (1 + (n-1)\delta)}, \end{align*} where $Z_k(x)$ is the $k$-th entry of $Z(x) \equiv B^\top x$, and $B = [b_1,...,b_n]$ is an $n \times n$ orthogonal matrix such that $b_1 = n^{-1/2} \mathbf{1}$, $b_k^\top \mathbf{1} = 0$ for all $k=2,...,n$, $b_k^\top b_\ell = 0$ for all $k \ne \ell = 2,...,n$.

Proof: We apply Lemma (ref) with $S = U = I_n$,

align*[align* omitted — 115 chars of source]

Note that the spectral decomposition of $A$ is given by $B \Lambda B^\top$, where $\Lambda$ is the diagonal matrix with the diagonal elements $\lambda_1,...,\lambda_n$ given as $\lambda_1 = (\sigma^2 - 1) + \sigma^2(n-1)\delta$, $\lambda_2 = ... = \lambda_n = - \sigma^2 \delta$, and the orthogonal matrix $B$ as given the lemma. The desired result follows from Lemma (ref). $\blacksquare$

Lemma (ref) yields the following result for the case with cluster dependence. From here on, we make the dimension of the matrices and vectors explicit. Let $I_{n_m}$ be the $n_m$-dimensional identity matrix and $\mathbf{1}_{n_m}$ denote the $n_m$-dimensional column vector of ones.

corollaryLet $\Sigma_n$ be the block diagonal matrix whose $m$-th block, $m = 1,...,M_n$, is given by \begin{eqnarray} \Sigma_{n,m} = \sigma^2 \left((1 - \delta_{n,m}) I_{n_m} + \delta_{n,m} \mathbf{1}_{n_m} \mathbf{1}_{n_m}^\top \right), \end{eqnarray} for some $\delta_{n,m} \in \mathbf{R}$ such that $n_m \delta_{n,m} \in (-1,1)$. Then, for any $\overline \mu_n \in \mathbf{R}$, and $x = [x_1,...,x_n]' \in \mathbf{R}^n$, \begin{align*} \log \frac{d \Phi(\overline \mu_n \mathbf{1}_n,\Sigma_n)}{d \Phi(0,I_n)}(x) &= - \sum_{m=1}^{M_n} \log \sqrt{\sigma^2(1 + (n_m-1) \delta_{n,m})} - \sum_{m=1}^{M_n} (n_m-1) \log \sqrt{\sigma^2(1 - \delta_{n,m})}\\ & \quad + \sum_{m=1}^{M_n} \frac{\sigma^2 (1 + (n_m-1) \delta_{n,m})-1}{2\sigma^2 (1 + (n_m-1) \delta_{n,m})}Z_{m,i_m}^2(x) \\ & \quad + \sum_{m=1}^{M_n} \frac{\sigma^2 (1-\delta_{n,m}) - 1}{2 \sigma^2(1 - \delta_{n,m})} \sum_{i \in N_{n,m} \setminus \{i_m\}} Z_{m,i}^2(x)\\ & \quad + \sum_{m=1}^{M_n} \frac{n\overline \mu_n/\sqrt{n_m}}{\sigma^2 (1 + (n_m-1)\delta_{n,m})} Z_{n,i_m}(x) - \frac{1}{2}\sum_{m=1}^{M_n} \frac{n^2\overline \mu_n^2/\sqrt{n_m}}{\sigma^2(1 + (n_m - 1)\delta_{n,m})}, \end{align*} where $Z_{m,i}(x)$, $i \in N_{n,m}$, are the entries of $Z_{m}(x) = B_m^\top x_{m,n}$, $B_m$ is an $n_m \times n_m$ orthogonal matrix, $i_m$ denotes the first index in $N_{n,m}$, and $x_{m,n} = [x_i]_{i \in N_{n,m}}$.
lemmaSuppose that $f: \mathbf{R} \rightarrow \mathbf{R}^{+}$ is a continuously differentiable function such that for some $\delta >0$, \begin{eqnarray*} \sup_{x \in [-\delta, \delta]} \left| \frac{d \log f(x)}{dx}\right| |\overline x| < 1, \end{eqnarray*} where $\overline x \in \mathbf{R}$ is such that \begin{eqnarray*} \sup_{x \in [-\delta,\delta]} f(x) = f(\overline x). \end{eqnarray*} Then, \begin{eqnarray*} f(\overline x) \le \left( 1 - \sup_{x \in [-\delta, \delta]}\left| \frac{d \log f(x)}{dx}\right| | \overline x | \right)^{-1} f(0). \end{eqnarray*}

Proof: Using the Mean Value Theorem,

align*[align* omitted — 264 chars of source]

where $x^*(x)$ is a point on the line segment between $0$ and $x$. Evaluating the inequality at $x=\overline x$ gives us the desired result. $\blacksquare$

Recall the defintion of $n^*$ in Corollary (ref):

eqnarray[eqnarray omitted — 71 chars of source]

The number $n - n^*$ represents the number of random variables, $X_{n,i}$, that are known to be mutually independent. Each variable outside this set belongs to a non-singleton cluster.

lemma$\lim_{n \rightarrow \infty} \sum_{m = 1}^{M_n} \left(n_m / n\right)^2 = 0$ if and only if (a) $\lim_{n \rightarrow \infty} n^*/n = 0$, or (b) $\lim_{n \rightarrow \infty} \sum_{m = 1: n_m \ge 2}^{M_n} \left(n_m / n^*\right)^2 = 0$.

Proof: For each $n \ge 1$, we have either

align*[align* omitted — 403 chars of source]

Since $(1/n)^2(n- n^*) = o(1)$, $\lim_{n \rightarrow \infty} \sum_{m = 1}^{M_n} \left(n_m / n\right)^2 = 0$ if and only if (a) or (b) holds. $\blacksquare$

lemmaSuppose that $n \ge 2$, and $\Sigma_n$ is a block diagonal matrix along a cluster structure $\mathcal{M}_n$, where the $m$-th block, denoted by $\Sigma_{n,m}$ is given by \begin{align*} \Sigma_{n,m} = \sigma^2 (1 -\delta_{n,m})I_{n_m} + \sigma^2 \delta_{n,m} \mathbf{1}_{n_m} \mathbf{1}_{n_m}^\top, \end{align*} where \begin{eqnarray} \delta_{n,m} = \frac{\overline \delta}{n^*}, \end{eqnarray} for some $\overline\delta \in [-a,a]$, with $0<a<1/2$, if $n_m \ge 2$, and $\sigma^2>0$ is independent of $n$. Then, the following holds for any random vector $X_n \in \mathbf{R}^n$ which follows $\Phi(0,\sigma^2 I_n)$. (i) $(d\Phi(0,\Sigma_n)/d\Phi(0,\sigma^2 I_n))(X_n)$ is uniformly integrable. (ii) $\log\left( (d\Phi(0,\Sigma_n)/d\Phi(0,\sigma^2 I_n))(X_n) \right)$ is uniformly tight.

Proof: For brevity, we focus on the case $\sigma^2 = 1$. By Corollary (ref),

align*[align* omitted — 78 chars of source]

where

align*[align* omitted — 352 chars of source]

and $i_m$ denotes the first $i$ in block $m$.

(i) Let us take small $\epsilon>0$ such that

align[align omitted — 72 chars of source]

We write (under $\Phi(0,I_n)$)

align*[align* omitted — 209 chars of source]

since $A_n$ and $R_n$ are independent. Let $t_m = (n_m-1)/n^*$, and write

align*[align* omitted — 466 chars of source]

Note that

align*[align* omitted — 231 chars of source]

because $t_m \le 1$ and $\overline \delta ( 1+ \epsilon) <1$ by ((ref)). This means that $f_n(\overline \delta)$ is increasing in $\overline \delta \in [-a,a]$ and achieves its maximum at $\overline \delta = a$. Hence,

align*[align* omitted — 355 chars of source]

because $\sum_{m=1}^{M_n} t_m^2 \le 1$ and we chose $\epsilon$ such that ((ref)) holds. By Lemma (ref), we have

align*[align* omitted — 138 chars of source]

The bound does not depend on $n$, and hence,

align*[align* omitted — 87 chars of source]

Now, we turn to $ \mathbf{E}\left[\exp((1 + \epsilon)R_n) \right]$. We can write

align[align omitted — 246 chars of source]

Using this expression, we rewrite

align*[align* omitted — 776 chars of source]

The last bound is a sequence converging to $\exp(a(1 + \epsilon)/2)$ as $n^* \rightarrow \infty$, and Hence, is a bounded sequence. Thus, we conclude that

align*[align* omitted — 87 chars of source]

This proves that

align*[align* omitted — 140 chars of source]

Hence, the proof of (i) is complete.

(ii) We rewrite

align[align omitted — 269 chars of source]

For any $x \in [0,1]$, we have

eqnarray*[eqnarray* omitted — 152 chars of source]

Hence,

align*[align* omitted — 279 chars of source]

It suffices to show the uniform tightness of the second sum in ((ref)). Under $\Phi(0,I_n)$, it has mean zero, and

align*[align* omitted — 427 chars of source]

because $\operatorname{Var}(Z_{m,i_m}^2) = 2$. Therefore, $A_n$ is uniformly tight.

As for $R_n$, we recall ((ref)), and can follow similar arguments to show that $R_n$ is uniformly tight as well. $\blacksquare$

\@startsection{subsection}{2} \z@{.8\linespacing\@plus.7\linespacing}{.7\linespacing}{Consistent Discrimination of the Mean}

We let

align[align omitted — 139 chars of source]

and for any $\overline \mu \in \mathbf{R}$, we write $\Phi(\overline \mu \mathbf{1}_n,\Sigma(\sigma^2,\delta))$ simply as $\Phi(\overline \mu,\sigma^2,\delta)$. Let us recall some basic notions of optimality of tests Lehmann/Romano:05:TSH. Given a model $\mathcal{P}_{n}$ which is partitioned as $\mathcal{P}_{n,0} \cup \mathcal{P}_{n,1}$, a test $\phi_n$ is said to be a \bi{UMP (uniformly most powerful) test} of $\mathcal{P}_{n,0}$ against $\mathcal{P}_{n,1}$ at level $\alpha \in (0,1)$, if under any $P_{n,0} \in \mathcal{P}_{n,0}$,

align*[align* omitted — 45 chars of source]

and for any alternative test $\phi_n'$ such that $\mathbf{E}[\phi_n'] \le \alpha$ under any $P_{n,0} \in \mathcal{P}_{n,0}$, we have

align*[align* omitted — 62 chars of source]

under any $P_{n,1} \in \mathcal{P}_{n,1}$.

A sequence of tests $\phi_n$ is said to be an \bi{AUMP (asymptotically uniformly most powerful) test} of $\mathcal{P}_{n,0}$ against $\mathcal{P}_{n,1}$ at level $\alpha \in (0,1)$, if under any sequence $P_{n,0} \in \mathcal{P}_{n,0}$,

align*[align* omitted — 76 chars of source]

and for any alternative test $\phi_n'$ such that $\limsup_{n \rightarrow \infty} \mathbf{E}[\phi_n'] \le \alpha$ under any sequence $P_{n,0} \in \mathcal{P}_{n,0}$, we have

align*[align* omitted — 93 chars of source]

under any sequence $P_{n,1} \in \mathcal{P}_{n,1}$.

lemmaSuppose that $A \subset \mathbf{R}$ is an open interval, and $\{\varphi_\alpha\}_{\alpha \in A}$ is a class of tests of $\mathcal{P}_{n,0}$ against $\mathcal{P}_{n,1}$, such that for each $\alpha \in A$, the test $\varphi_\alpha$ is UMP at level $\alpha$, and for any $\tilde \alpha = \alpha +o(1)$ as $n \rightarrow \infty$, \begin{align} \mathbf{E}[\varphi_{\alpha}] = \mathbf{E}[\varphi_{\tilde \alpha}] + o(1), \end{align} under any sequence $P_n \in \mathcal{P}_{n,0} \cup \mathcal{P}_{n,1}$. Then, $\varphi_\alpha$ is AUMP at level $\alpha$.

Proof: Choose any test $\tilde \varphi$ such that under any sequence $P_n \in \mathcal{P}_{n,0}$,

align*[align* omitted — 66 chars of source]

for some sequence $\epsilon_n \rightarrow 0$, as $n \rightarrow \infty$. Fix one such sequence $P_n\in \mathcal{P}_{n,0}$, together with the sequence $\epsilon_n$, and let $\tilde \alpha = \alpha + \epsilon_n$. Now, select a large enough $n$ such that $\tilde \alpha \in A$ and choose any $P_n' \in \mathcal{P}_{n,1}$. Then, since $\varphi_{\tilde \alpha}$ is UMP at level $\tilde \alpha$, we have

align*[align* omitted — 95 chars of source]

under $P_n'$. By ((ref)), we can see that $\varphi_{\alpha}$ is AUMP at level $\alpha$. $\blacksquare$

Proof of Theorem (ref): Suppose that $M_n = 1$. First, consider the case where $\mathcal{P}_{n} = \mathcal{P}_{n,\mathcal{N}}'$, with

align*[align* omitted — 133 chars of source]

Later, we generalize the result to the case where $\mathcal{P}_n$ contains the above probability model. Define $\mathcal{P}_{n,0} = \left\{\Phi(0,\sigma^2,\delta): (n-1) \delta \in (0,1), \sigma^2>0 \right\}$ and let $\mathcal{P}_{n,1} = \mathcal{P}_{n} \setminus \mathcal{P}_{n,0}$. In light of Lemma (ref), it suffices to construct a class of tests $\{\varphi_{\alpha}\}_{\alpha \in (0,1/2)}$ of $\mathcal{P}_{n,0}$ against $\mathcal{P}_{n,1}$ such that

(a) it satisfies ((ref)) in Lemma (ref),

(b) the test $\varphi_{\alpha}$ has power bounded by a constant below $1$ uniformly over $n$, and

(c) each test $\varphi_{\alpha}$ is a UMP test of $\mathcal{P}_{n,0}$ against $\mathcal{P}_{n,1}$.

Let us first construct such a test and show that (a)-(c) are satisfied. Define

align*[align* omitted — 69 chars of source]

For each $\alpha \in (0,1/2)$, let

align*[align* omitted — 179 chars of source]

for some $C_0>0$ and $\gamma_0 \in [0,1]$. Let $Z = (Z_{n,1} - \mathbf{E}[Z_{n,1}])/\sqrt{\operatorname{Var}(Z_{n,1})}$. Then the size control requires that under the null hypothesis,

align[align omitted — 120 chars of source]

Since $\alpha \in (0,1/2)$, we must have $C_0 = 1$ and $\gamma_0 = 2 \alpha$.

Let us first show that this test satisfies the condition (a). For any $\tilde \alpha$ such that $\tilde \alpha = \alpha + o(1)$, and under any sequence $P_n \in \mathcal{P}_{n}$,

align[align omitted — 193 chars of source]

Hence, the class of tests $\{\phi_{\alpha}(V_n)\}_{\alpha \in (0,1/2)}$ satisfies the condition ((ref)).

As for the condition (b), note that under any alternative hypothesis in $\mathcal{P}_{n,1}$, we have

align*[align* omitted — 159 chars of source]

Hence, the test does not have power exceeding $2\alpha < 1$.

Finally, we show that the condition (c) is satisfied. Let

align[align omitted — 442 chars of source]

Hence, $\mathcal{L}(X_n;\mu,\sigma^2,\delta_n)$ is the same as $\log\left(d \Phi(\mu \mathbf{1}_n,\Sigma(\sigma^2,\delta_n)/d \Phi(0,I_n)\right)(X_n)$ in Lemma (ref), except that the coefficient of $Z_{n,1}$ is $\sqrt{n}\mu$ and the last term is different. Define a probability measure $P_n(\mu,\delta_n)$ as follows: for any Borel $B$,

align*[align* omitted — 106 chars of source]

Similarly as before, we define

align*[align* omitted — 186 chars of source]

and let $\mathcal{\tilde P}_{n,1} = \mathcal{\tilde P}_{n} \setminus \mathcal{\tilde P}_{n,0}$. It is not hard to see that

align*[align* omitted — 130 chars of source]

Therefore, a UMP test of $\mathcal{\tilde P}_{n,0}$ against $\mathcal{\tilde P}_{n,1}$ is also a UMP test of $\mathcal{P}_{n,0}$ against $\mathcal{P}_{n,1}$. It suffices for condition (c) to show that the test $\varphi_\alpha$ is a UMP test of $\mathcal{\tilde P}_{n,0}$ against $\mathcal{\tilde P}_{n,1}$. From ((ref)), the sufficient statistics for $\mathcal{\tilde P}_{n}$ in the case of $M_n = 1$ are given by

align*[align* omitted — 71 chars of source]

where $Z_{n,k}$'s are as in Lemma (ref). For any $t \ge 0$,

align*[align* omitted — 367 chars of source]

under the null hypothesis. Hence, $V_n$ and $Z_{n,1}^2$ are independent under any probability in $\mathcal{\tilde P}_{n,0}$. Furthermore, under any probability in $\mathcal{\tilde P}_{n,0}$,

align*[align* omitted — 269 chars of source]

Note that $B_n^\top \mathbf{1}_n \mathbf{1}_n B_n$ is a matrix whose $(1,1)$-th entry is $b_1^\top \mathbf{1}_n \mathbf{1}_n^\top b_1 = 1$ and all the other entries are zeros. Hence, $Z_{n,k}$'s are independent across $k$'s under any probability in $\mathcal{\tilde P}_{n,0}$. Therefore, $V_n$ and $(Z_{n,1}^2,\sum_{k=2}^n Z_{n,k}^2)$ are independent under any probability in $\mathcal{\tilde P}_{n,0}$. By Theorem 5.1.1 of Lehmann/Romano:05:TSH, the randomized test $\varphi_\alpha(V_n)$ is an $\alpha$-level UMP test.

Next, consider the case where $\mathcal{P}_{n,\mathcal{N}}' \subset \mathcal{P}_n$. Take a sequence of tests $\varphi_n$ such that for any sequence of probabilities $P_n \in \mathcal{P}_{n,0}$, $\limsup_{n \rightarrow \infty} \mathbf{E}[\varphi_n(X_n)] \le \alpha$. Now, we take a sequence $P_{n,1} \in \mathcal{P}_{n,1} \cap \mathcal{P}_{n,\mathcal{N}}'$. For any $\alpha \in (0,1/2)$, the test $\phi_\alpha(V_n)$ is a UMP test at level $\alpha$ of the null hypothesis $\mathcal{P}_{n,0} \cap \mathcal{P}_{n,\mathcal{N}}'$ against $\mathcal{P}_{n,1} \cap \mathcal{P}_{n,\mathcal{N}}'$. Note that for any sequence of probabilities $P_n \in \mathcal{P}_{n,0} \cap \mathcal{P}_{n,\mathcal{N}}'$, $\limsup_{n \rightarrow \infty} \mathbf{E}[\varphi_n(X_n)] \le \alpha$. Hence, if we take $\epsilon>0$ such that $\alpha + \epsilon \in (0,1/2)$, there exists $n_0 \ge 1$ such that for all $n \ge n_0$, $\mathbf{E}[\varphi_n(X_n)] \le \alpha + \epsilon$. For all such $n$, under any $P_n \in \mathcal{P}_{n,1} \cap \mathcal{P}_{n,\mathcal{N}}'$, we have

align*[align* omitted — 114 chars of source]

Hence, we find that along any sequence $P_n \in \mathcal{P}_{n,1} \cap \mathcal{P}_{n,\mathcal{N}}'$, we have

align*[align* omitted — 103 chars of source]

Thus, the proof is complete. $\blacksquare$

The following lemma is used for the proof of Theorem (ref). For $w \in [0,1]$, define $\xi_1 = w Z_1$ and $\xi_2 = (1-w)Z_2$, where $Z_i \sim N(0,1)$, independent across $i=1,2$. Define

align*[align* omitted — 106 chars of source]

where $\bar \xi = (\xi_1 + \xi_2)/2$. Let us take $t \ge 0$, and define

align*[align* omitted — 40 chars of source]
lemma(i) For all $t \ge 1$ and all $w \in [0,1]$, $p(w;t) \ge p(1/2;t)$. (ii) For all $0\le t<1$ and all $w \in [0,1]$, $p(w;t) \le p(1/2;t)$.

Proof: First, we write

align*[align* omitted — 69 chars of source]

For (i) and (ii), since $[\xi_1,\xi_2]$ is symmetrically distributed around the origin, it suffices to show that $p(\,\cdot\,;t)$ is increasing on $[1/2,1]$ for all $t \ge 1$, and $p(\,\cdot\,;t)$ is decreasing on $[1/2,1]$ for all $0 \le t<1$. Let $A$ denote the event that $\xi_1 + \xi_2 <0$. Then, on the event $A$, for all $w \ge 0$ and all $t \ge 0$, we have $T(w) \le t$. Hence,

align*[align* omitted — 64 chars of source]

On the event $A^c$, $T(w) \le t$ if and only if

align*[align* omitted — 112 chars of source]

if and only if

align*[align* omitted — 98 chars of source]

where

align*[align* omitted — 83 chars of source]

Take $t \ge 1$. Let $B$ be the event $Z_1 Z_2 \le 0$. Certainly, if $Z_1 Z_2 \le 0$, $f(w; Z_1,Z_2) \ge 0$ for all $w \in [1/2,1]$. Hence

align*[align* omitted — 234 chars of source]

We show that $f(w;Z_1,Z_2)$ is increasing in $w$ on the event $A^c$. We take the derivative $f'(w;Z_1,Z_2)$ with respect to $w$:

align*[align* omitted — 92 chars of source]

The function is linear in $w$. First, we take $w =1$. Then, on the event $B^c$,

align*[align* omitted — 166 chars of source]

Second, we take $w=1/2$. Then,

align*[align* omitted — 56 chars of source]

Hence, for all $Z_1,Z_2$ such that $Z_1 Z_2 > 0$, $f(w;Z_1,Z_2)$ is increasing on $[1/2,1]$ for all $t \ge 1$. Therefore, whenever $t \ge 1$, $P\{T(w) \le t\}$ is increasing on $[1/2,1]$.

Take $0 \le t <1$. If $Z_1 Z_2 > 0$, then $f(w;Z_1,Z_2) < 0$. Hence

align*[align* omitted — 204 chars of source]

On the event $B$, when $w=1$,

align*[align* omitted — 73 chars of source]

and when $w=1/2$,

align[align omitted — 59 chars of source]

Hence, for all $Z_1,Z_2$ such that $Z_1 Z_2 \le 0$, $f(w;Z_1,Z_2)$ is decreasing on $[1/2,1]$ for all $0 \le t <1$. Therefore, whenever $0 \le t < 1$, $P\{T(w) \le t\}$ is decreasing on $[1/2,1]$. $\blacksquare$

Proof of Theorem (ref): Suppose that we have at least two nonnegligible clusters, i.e., $M_n \ge 2$ for all but finite number of $n$'s. Consider testing the null hypothesis of $\mathbf{E}[X_{n,i}] = 0$ against $\mathbf{E}[X_{n,i}] > 0$. Without loss of generality, we enumerate $\mathcal{M}_n' = \{N_1,...,N_{M_n'}\}$, $M_n' \ge 2$. Now, we construct a test that consistently $\sqrt{n}$-discriminates the mean. Define

align*[align* omitted — 72 chars of source]

and

align*[align* omitted — 156 chars of source]

We take

align*[align* omitted — 56 chars of source]

Let $c_{1 - \alpha}$ be the $1 - \alpha$ quantile of the $t$-distribution with degree of freedom $1$. Define

align*[align* omitted — 71 chars of source]

Note that $c_{0.75} = 1$.

We first show that this test controls the size of the test under $\alpha$ asymptotically under the null hypothesis. Define an infeasible test statistic

align*[align* omitted — 59 chars of source]

where

align*[align* omitted — 220 chars of source]

Here $\sigma_{n,m}^2$ is the variance of $\xi_m$ under the null hypothesis. Then $V_n''$ converges in distribution to the $t$-distribution with $1$ degree of freedom under the null hypothesis. By Lemma (ref), if $0< \alpha < 0.25$ so that $c_{1-\alpha} >1$,

align[align omitted — 137 chars of source]

and if $\alpha \ge 0.25$ so that $0 \le c_{1-\alpha} \le 1$,

align[align omitted — 148 chars of source]

as $n \rightarrow \infty$. Hence, the size of the test $\varphi_n$ is bounded by $\alpha$ asymptotically.

Suppose that we are under the local alternatives such that $\mathbf{E}[X_n] = \overline \mu_n \sigma_{LR}\mathbf{1}_n/\sqrt{n}$, for some sequence $\overline \mu_n \rightarrow \infty$. Define

align*[align* omitted — 108 chars of source]

Note that

align[align omitted — 424 chars of source]

because $n_m /n \le 1$. Note that

align[align omitted — 124 chars of source]

Since

align*[align* omitted — 145 chars of source]

there exist $\epsilon>0$ and $n_0>0$ such that for all $n \ge n_0$,

align*[align* omitted — 114 chars of source]

Therefore,

align*[align* omitted — 281 chars of source]

Since $(\xi_m - \mathbf{E}\xi_m)/\sigma_{n,m}$ converges in distribution to $N(0,1)$ under any sequence $P_n \in \mathcal{P}_n$ as $n \rightarrow \infty$ by ((ref)), we have $\bar U_n /\sigma_{LR} = O_P(1)$. Similarly, we can show that $T_n'/\sigma_{LR}^2 = O_P(1)$. Hence, the last probability in ((ref)) converges to one as $\overline \mu_n \rightarrow \infty$, proving that the mean is consistently $\sqrt{n}$-discriminated. $\blacksquare$

\@startsection{subsection}{2} \z@{.8\linespacing\@plus.7\linespacing}{.7\linespacing}{Impossibility of Consistent Estimation of Long Run Variance}

Proof of Lemma (ref): We first show sufficiency. Suppose that $\|\theta_{n}(P_{n,1}) - \theta_{n}(P_{n,0})\| = o(1)$ for any sequence $P_{n,1} \in \mathcal{P}_n$. Then we take $\hat \theta = \theta_n(P_{n,0})$, so that $\| \hat \theta - \theta_n(P_{n,1})\| \rightarrow 0$, along $P_{n,1} \in \mathcal{P}_n$, as $n \rightarrow \infty$. Hence, sufficiency follows.

Conversely, suppose that $\theta_n$ is consistently estimable in $\mathcal{P}_n$, so that there exists an estimator, say, $\tilde \theta$, such that $\tilde \theta - \theta_n(P_{n,1}) = o_P(1)$ along any $P_{n,1} \in \mathcal{P}_n$. Since $P_{n,0} \in \mathcal{P}_n$, this means that $\tilde \theta - \theta_n(P_{n,0}) = o_P(1)$ under $P_{n,0}$. Since $P_{n,1} \triangleleft P_{n,0}$, $\tilde \theta - \theta_n(P_{n,0}) = o_P(1)$ under any $P_{n,1} \in \mathcal{P}_n$. We choose any $P_{n,1} \in \mathcal{P}_n$ and write

eqnarray[eqnarray omitted — 144 chars of source]

The difference on the left hand side and the first difference on the right hand side are $o_P(1)$ under $P_{n,1}$. This implies that $\|\theta_n(P_{n,0}) - \theta_n(P_{n,1})\| = o(1)$. $\blacksquare$

Proof of Theorem (ref): Let us first show sufficiency. Suppose that either (a) or (b) in Lemma (ref) holds. Let us take

align*[align* omitted — 138 chars of source]

Since $\sigma_{LR}^2 \le c$ for all $n \ge 1$, we have

align*[align* omitted — 59 chars of source]

Hence

align*[align* omitted — 159 chars of source]

Note that

align[align omitted — 123 chars of source]

Choose $P_n \in \mathcal{P}_n$. Then, under $P_n$,

align[align omitted — 526 chars of source]

where $C'>0$ is a constant that does not depend on $n$. By ((ref)), the last term is $o(1)$, if either of the conditions (a) and (b) in Lemma (ref). Therefore, $\sigma_{LR}^2$ is consistently estimable in $\mathcal{P}_n$.

Now, let us show necessity. Suppose that both (a) and (b) in Lemma (ref) are violated. That is,

eqnarray[eqnarray omitted — 205 chars of source]

We fix $\sigma^2$ and show that $\sigma_{LR}^2$ is not consistently estimable in $\mathcal{P}_{n,\mathcal{N}}(\sigma^2)$. We choose $\delta_{n,m}>0$ for each cluster $m$ and $n$ such that ((ref)) holds for some $\overline\delta \in [-a,a]\setminus \{0\}$, $a< 1/2$, if $n_m \ge 2$. By ((ref)), there exists a subsequence $\{n_k\} \subset \{n\}$ such that

eqnarray[eqnarray omitted — 181 chars of source]

for some constants $\tilde c_1, \tilde c_2 \in (0,1]$. For simplicity, we fix this subsequence, and denote $n_k$ by $n$.

We let $\Sigma_n$ be the block diagonal $n \times n$ matrix whose $m$-th block is given by $\sigma^2((1-\delta_{n,m})I_{n_m} + \delta_{n,m} \textbf{1}_{n_m} \textbf{1}_{n_m}^\top)$. We show that $\Phi(0,\Sigma_n) \triangleleft \Phi(0,\sigma^2 I_n)$. First, we observe that by Lemma (ref)(ii), $\log (d \Phi(0,\Sigma_n)/d \Phi(0,\sigma^2 I_n))(X_n)$ is uniformly tight under $\Phi(0,\sigma^2 I_n)$. Furthermore, we find that by Prohorov's Theorem, there exists a subsequence $\{n_k\}$ of $\{n\}$ such that the sequence $\log (d\Phi(0,\Sigma_{n_k})/d \Phi(0,\sigma^2 I_{n_k}))(X_n)$ weakly converges. Let $W$ be a random variable whose distribution is identical to the weak limit. By the Continuous Mapping Theorem, we have

eqnarray*[eqnarray* omitted — 103 chars of source]

along the subsequence $\{n_k\}$. Note that $\mathbf{E}[(d \Phi(0,\Sigma_{n_k})/d\Phi(0,\sigma^2 I_{n_k}))(X_{n_k})] =1 $, where the expectation is under $\Phi(0,\sigma^2 I_{n_k})$. By Lemma (ref)(i), $(d \Phi(0,\Sigma_{n_k})/d\Phi(0,\sigma^2 I_{n_k}))(X_{n_k})$ is uniformly integrable under $\Phi(0,\sigma^2 I_{n_k})$. Hence, we find that $\mathbf{E} e^W =1$. By Le Cam's First Lemma vanderVaart:98:AsympStat, we conclude that $\Phi(0,\Sigma_{n_k}) \triangleleft \Phi(0,\sigma^2 I_{n_k})$.

On the other hand, note that the difference between the long-run variances under $\Phi(0,\Sigma_n)$ and under $\Phi(0,\sigma^2 I_n)$ is given by

align*[align* omitted — 341 chars of source]

where $\Delta_{n,m} = \delta_{n,m} \textbf{1}_{n_m} \textbf{1}_{n_m}^\top - \delta_{n,m} I_{n_m}$. We rewrite the last term as

align*[align* omitted — 383 chars of source]

as $n \rightarrow \infty$, where the last convergence is due to ((ref)). Certainly, $\sigma_{LR}^2$ is consistently estimable along $\Phi(0,\sigma^2 I_n)$. By Lemma (ref), we conclude that $\sigma_{LR}^2$ is not consistently estimable in $\mathcal{P}_{n,\mathcal{N}}(\sigma^2)$. Hence, it is not consistently estimable in $\mathcal{P}_{n}$ either. $\blacksquare$

Proof of Corollary (ref): Let us show sufficiency. First suppose that $\mathcal{M}_n$ consists of negligible clusters, so that $\max_{1 \le m \le M_n} n_m / n \rightarrow 0$, as $n \rightarrow \infty$. From ((ref)), this implies that

align*[align* omitted — 157 chars of source]

as $n \rightarrow \infty$. Therefore, $\sigma_{LR}^2$ is consistently estimable in $\mathcal{P}_{n}$ by Theorem (ref).

Conversely, suppose that for some $\epsilon>0$,

align*[align* omitted — 95 chars of source]

Then there exist subsequences $\{n_k\} \subset \{n\}$ and $\{n_{m(n_k)}\} \subset \{n_{m(n)}\}_{n \ge 1}$, $m(n) \in \{1,...,M_n\}$, such that

align*[align* omitted — 77 chars of source]

This implies that

align*[align* omitted — 110 chars of source]

Hence, $\sigma_{LR}^2$ is not consistently estimable in $\mathcal{P}_{n}$ by Theorem (ref). $\blacksquare$

Proof of Corollary (ref): (i) Let $N_\circ \subset \{1,...,n\}$ be the set of nodes in the clique with size $n_C$. Take $\mathcal{M}_n$ to be the cluster structure such that there is only one non-singleton cluster that is $N_\circ$. It suffices to show that $\sigma_{LR}^2$ is not consistently estimable in $\mathcal{P}_{n,\mathcal{N}}(\sigma^2)$ with the cluster structure $\mathcal{M}_n$. Note that

align*[align* omitted — 79 chars of source]

because we have only one non-singleton cluster in $\mathcal{M}_n$. By Theorem (ref), $\sigma_{LR}^2$ is not consistently estimable in $\mathcal{P}_{n,\mathcal{N}}(\sigma^2)$ with any fixed $\sigma^2>0$.

(ii) Let us define $\overline N(i) = \{j \in N_n: ij \in \overline E_n\}$, where $\overline E_n = E_n \cup \{ii: i \in N_n\}$, and consider

eqnarray*[eqnarray* omitted — 145 chars of source]

By rearranging terms, we can write

align[align omitted — 463 chars of source]

We can write the squared $L^2$ norm of the leading term on the right hand side as

align*[align* omitted — 202 chars of source]

If $\{i_1,j_1\}$ and $\{i_2,j_2\}$ are not adjacent in $G_n$, the covariance above is zero by the dependency graph assumption. The number of the terms in the above sum such that $\{i_1,j_1\}$ and $\{i_2,j_2\}$ are adjacent in $G_n$ is of the order $O(nd_{mx}^2 d_{av}) = o(n^2)$. The last rate comes from our assumption that $\lim_{n \rightarrow \infty} d_{mx}^2 d_{av} /n = 0$. Therefore, the leading term on the right hand side of ((ref)) is $o_P(1)$. Similarly, we can show that the remainder terms are $o_P(1)$. Hence, $\hat \sigma_{LR}^2$ is a consistent estimator of $\sigma_{LR}^2$. $\blacksquare$