Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
105,084 characters · 0 sections · 65 citation commands
\fontsize{13}{14} \selectfont
\@startsection{section}{1} \z@{1.0\linespacing\@plus\linespacing}{.8\linespacing}{Introduction}
Statistical inference from data usually begins by imposing a form of a dependence structure on the data, by specifying which groups of observations exhibit strong within-group dependence. Various tools of asymptotic inference such as the law of large numbers and the central limit theorem are available for many typically imposed dependence structures. A standard case is the independence assumption or an assumption on time series dependence. However, it is well known that in the case of cross-sectional dependence, a researcher is often less confident about the correctness of the dependence structure used, despite its crucial role for inference.
A popular way to deal with this challenge is to use cluster dependence modeling, where the dependence structure among observations within each cluster is left unspecified, while independence is imposed between observations from different clusters. The inference procedures when there are many clusters are well known and can be analyzed using standard methods of asymptotic inference. However, less is known about the case where there are large clusters, and the dependence structure within such a cluster is unknown. Cameron/Gelbach/Miller:08:ReStat proposed a wild bootstrap procedure and showed by simulations that their tests perform well even when there are a small number of clusters. The robustness of this result was confirmed by MacKinnon/Webb:17:JAE even when the sizes of the clusters are highly heterogeneous. This large cluster issue has also drawn interest in the literature of difference-in-differences when there are only few treated clusters (see Conley/Taber:11:ReStat, Hagemann:19:JOE, and MacKinnon/Webb:20:JOE, and references therein). Djogbenou/MacKinnon/Nielsen:19:JOE studied inference on regression models with clustered errors. They provided conditions for the cluster sizes so that asymptotic and bootstrap inferences are asymptotically valid. They showed that their conditions exclude the presence of a large cluster.
There are several methods proposed to deal with the problem of inference with large clusters. Donald/Lang:07:ReStat and Bester/Conley/Hansen:11:JOE considered linear models and proposed inference where the asymptotic distribution of the long run variance estimator is fully known. This approach is related to the HAR (Heteroskedasticity-Autocorrelation Robust) inference of Kiefer/Vogelsang:2002:Eca and Sun:14:Eca in time series, which uses a normalization by an inconsistent long run variance estimator that has a stochastic limit.
Ibragimov/Muller:10:JBES proposed a $t$-test approach based on within-cluster estimators together with a $t$-distribution, where the degree of freedom in the $t$ distribution is determined by the number of clusters. They used the result of Bakirov/Szekely:05:ZNS and showed that their approach is asymptotically valid, even if the variances of the cluster specific estimators are different across the clusters. Ibragimov/Muller:16:ReStat extended these results to the problem of two-sample comparison and developed a testing procedure for the level of clustering.
Some studies adopted the approach of randomized testing to deal with cluster dependence with large clusters. Canay/Romano/Shaikh:17:Eca developed asymptotic inference procedures when the inference involves statistics whose limiting distribution satisfy symmetry properties. Hagemann:19:JOE proposed randomized tests for treatment effects when there are only a small number of clusters. Like Ibragimov/Muller:10:JBES, both proposals assumed large sample properties for within-cluster statistics. A recent work by Canay/Santos/Shaikh:21:ReStat use the analogue between wild bootstrap and randomized tests, and provided conditions under which the wild bootstrap for cluster-dependent regression models is asymptotically valid when there are only a small number of clusters.
Our paper focuses on observations with a cluster dependence structure and explores implications on statistical inference when there are large clusters. First, we show that when the sample consists of large clusters, the mean cannot be consistently discriminated if there is only one cluster, i.e., the researcher does not have any knowledge on the dependence structure of the data. Furthermore, when the observations form large clusters and within-cluster observations satisfy the uniform central limit theorem, a sufficient condition for the mean to be consistently discriminated at the rate of $\sqrt{n}$ is that the sample consists of at least two large clusters.
This impossibility result has a significant implication in a setting where the researcher does not know the dependence structure of observations. In such a case, consistent discrimination of the mean is not possible with uniform-in-$P$ asymptotic size control. Note that Song:16:arXiv proposed a randomized subsampling approach, and Leung:21:JAE provided a set of general conditions for the approach to produce asymptotically valid inference. Both focus on a setting where no knowledge on the dependence structure is required. Among other things, their results show that the mean is consistently discriminated. Our impossibility result on consistent discrimination considers a setting where there is no uniform upper bound of the long run variance in the null model, and this setting is excluded by part of their conditions. Hence, their results do not contradict our impossibility result.
Our second result is concerned with consistent estimation of long run variances. More specifically, suppose that $X_n = [X_{n,1},...,X_{n,n}]^\top$ is a given random vector of dimension $n$, where each observation $X_{n,i}$ has the same mean $\mu$. Let us define the long-run variance of $X_n$ as follows:\footnote{Note that when there is a common shock, say, $C_n$, such as cluster-specific fixed effects with few clusters, the analysis in this paper carries over to this case with $\sigma_{LR}^2$ replaced by the conditional variance given common shock $C_n$. Our impossibility results do not depend on whether there is a common shock of this form in the data or not. For simplicity, we consider a setting without such cluster-specific fixed effects.}
Recently, Hansen/Lee:19:JOE derived an asymptotic distribution theory for clustered data, including a law of large numbers and a central limit theorem. One of their results presents a condition for the cluster structure that is necessary and sufficient for the weak law of large numbers to hold for the sample average of the clustered observations. Our paper shows that the same condition is in fact necessary and sufficient for the consistent estimability of the long run variance of the clustered observations as well. Our condition for the cluster structure also implies that when there is at least one large cluster, i.e., the researcher does not know the dependence structure on a nonnegligible portion of the data, the long-run variance is not consistently estimable. It is not hard to show that the existing cluster-robust variance estimators are inconsistent when the cluster structure is severely misspecified. However, to the best of our knowledge, it has not been known whether there exists any consistent estimator of the long run variance when there is a lack of knowledge on the dependence structure on a nonnegligible portion of the data. Our result gives a negative answer to this question.
There has long been a strand of literature that studies impossibility of estimation and inference. (See, e.g., Bahadur/Savage:56:AMS, Dufour:97:Eca, Potscher:02:Eca, Bertanha/Moreira:20:JOE.) The impossibility of consistent estimation of a long run variance in this paper is related to Potscher:02:Eca who established a minimax risk lower bound for a general estimation problem. Among others, his result can be used to prove the impossibility of consistent estimation uniform in $P$ as shown in Corollary 3.2 there. However, we cannot apply this corollary in our setting, because our probability model is not indexed by a set of parameters fixed independently of the sample size, such as $\mathscr{H}$ in his paper. This stems from our setting where we have to deal with the joint distribution of the entire sample whose dependence structure varies in the model as $n$ changes. Bertanha/Moreira:20:JOE studied impossibility results of two types: indistinguishability of the null hypothesis from the alternative hypothesis and unbounded confidence sets. Their study of impossibility of the first type is related to impossibility of consistent discrimination of the mean in our paper. For this result, they assume that for each probability in the alternative hypothesis, there is a sequence of probabilities under the null hypothesis that weakly converge to this probability. Our setting does not satisfy this assumption in general. Hence, our result does not fall into their framework. Menzel:21:Eca recently developed and verified the validity of a bootstrap procedure in multi-way clustered observations with two or more dimensions. Part of his results shows that it is not possible to consistently estimate the distribution of the cluster dependent observations. Our results are not the special case of his results, because our impossibility result holds for models that exclude the counterexample that he used to prove the impossibility result. In particular, our cluster dependence accommodates within-cluster heterogeneity in terms of marginal distributions and dependence structures.
The rest of the paper is organized as follows. The next section studies the consistent discrimination of the mean. Section 3 is devoted to presenting the result of the impossibility of consistent estimation of the long run variance. In Section 4, we illustrate the implication of our results for the case of difference-in-difference models. In Section 5, we conclude. The mathematical proofs are found in the appendix.
\@startsection{section}{1} \z@{1.0\linespacing\@plus\linespacing}{.8\linespacing}{Cluster Dependence}
Let $X_n = [X_{n,1},...,X_{n,n}]^\top$ be a random vector with a joint distribution $P_n$ which belongs to the class of distributions $\mathcal{P}_n$. Throughout the paper, we assume that for each $P_n \in \mathcal{P}_n$,
and $\sigma_{LR}^2 < \infty$, where $\sigma_{LR}^2$ is defined in ((ref)). In many situations, the dependence structure is partially observed. Here we consider cluster dependence, where the dependence structure is entirely unknown within each cluster, and observations are independent between clusters. Let $N_{n,m}, m = 1,...,M_n$, be a partition of $N_n = \{1,...,n\}$ such that $|N_{n,m}| = n_m$ for each $m=1,...,M_n$, so that $\sum_{m=1}^{M_n} n_m = n$. Define $\mathcal{M}_n = \{N_{n,m}: m = 1,...,M_n\}$ and call it a \bi{cluster structure}. Throughout the paper we assume that $(X_{n,i})_{i \in N_{n,m}}$ are independent across $m$'s under all $P_n \in \mathcal{P}_n$, i.e., the joint distribution of $X_n$ has a cluster dependence structure. For future references, we define
so that $\overline X_{m,n}$ represents the within-cluster mean of $X_{n,i}$'s and $\sigma_{n,m}^2$ represents the within-cluster long-run variance of $X_{n,i}$'s.
Our impossibility results rely on the assumption that the probability model, $\mathcal{P}_n$, includes Gaussian experiments with what we call local-to-independence common shocks. For each cluster $m=1,...,M_n$, and for $\delta > 0$ and $\sigma^2>0$, we define
where $I_{n_m}$ denotes the $n_m$-dimensional identity matrix and $\mathbf{1}_{n_m}$ is the $n_m$-dimensional column vector of ones. Let $\Sigma_n(\sigma^2,\delta)$ be the $n \times n$ block diagonal matrix whose $m$-th block is given by $\Sigma_{n,m}(\sigma^2,\delta)$. Suppose that $\Sigma_n(\sigma^2,\delta)$ is positive definite. Then, for each $\mu_n \in \mathbf{R}^n$, we denote $\Phi(\mu_n,\Sigma_n(\sigma^2,\delta))$ the multivariate normal distribution with mean $\mu_n$ and covariance matrix $\Sigma_n(\sigma^2,\delta)$. Define
The set $\mathcal{P}_{n,\mathcal{N}}$ represents a set of Gaussian models, where each member is a multivariate normal distribution with a common mean and an equal covariance. We call each $\Phi(\mu_n,\Sigma_n(\sigma^2,\delta))$ the \bi{local-to-independence common shock (LTIC) Gaussian distribution} with parameters $\sigma^2$ and $\delta$. This Gaussian distribution represents the cross-sectional dependence structure of $X_{n,i}$'s generated as follows:
where $\mu_{n,i}$ is the $i$-th entry of $\mu_n$, $\varepsilon_i$'s are i.i.d.\ normal random variables with mean zero and variance $\sigma^2$, and $\eta_m$, $m=1,...,M_n$, are i.i.d.\ normal random variables with mean zero and variance $\sigma^2$, independent of $\varepsilon_i$'s. Each random variable $\eta_m$ represents a within-cluster “common shock”, and creates the within-cluster global dependence among $X_{n,i}$'s. The influence of this common shock on the random variable $X_{n,i}$ diminishes at the rate of $\sqrt{n_m}$.
\@startsection{section}{1} \z@{1.0\linespacing\@plus\linespacing}{.8\linespacing}{Consistent Discrimination of the Mean}
\@startsection{subsection}{2} \z@{.8\linespacing\@plus.7\linespacing}{.7\linespacing}{Consistent Discrimination of the Mean}
Let us explore the consistent discrimination of the mean under the general cluster dependence structure. We introduce the notion of consistent discrimination formally. Let $\mathcal{P}_n$ be a set of the distributions of $X_n \in \mathbf{R}^n$ such that $\mathbf{E}[X_{n,i}]$ is identical across $i$ for each $n \ge 1$. Let $\mathcal{P}_{n,0} = \{P_n \in \mathcal{P}_n: \mathbf{E}[X_{n,i}] = 0\}$, i.e., the set of probabilities under the null hypothesis of $\mathbf{E}[X_{n,i}] = 0$.
The following theorem shows that when the sample consists of nonnegligible clusters, a necessary condition for the consistent discrimination of the mean is that there exist at least two clusters.
The theorem implies that when we do not know the local dependence structure of the random variables (i.e., $M_n = 1$), it is not possible to consistently discriminate the mean.
\@startsection{subsection}{2} \z@{.8\linespacing\@plus.7\linespacing}{.7\linespacing}{Consistent $\sqrt{n}$-Discrimination of the Mean}
We introduce the following notion of consistent $\sqrt{n}$-discrimination.
We consider consistent discrimination against alternatives after normalizing by $\sigma_{LR}$ (which depends on $n$), so that when $\sigma_{LR}^2$ is larger, we focus on the alternative hypothesis that is farther away from the null hypothesis. Hence, if it is not possible to consistently $\sqrt{n}$-discriminate the mean, it is not necessarily due to the long run variance increasing to infinity fast.
The consistent $\sqrt{n}$-discrimination is often obtained when the parameter is in a finite dimensional space, and one knows the local dependence structure of the observations. To illustrate this point, suppose that $X_{n,1},...,X_{n,n}$'s are i.i.d. Then, often we have
where $\hat \sigma_n^2 \rightarrow_p \sigma^2 = \operatorname{Var}(X_{n,1}) >0$, and $\overline X_n = \frac{1}{n}\sum_{i =1}^n X_{n,i}$. Let us consider the usual $t$-test as follows:
Under the Pitman local alternatives such that $\mathbf{E}[X_{n,i}] = \overline \mu/\sqrt{n}$, $\overline \mu>0$, we have
where $\Phi$ denotes the CDF of $\mathcal{N}(0,1)$. The last term converges to 1 as $\overline \mu \rightarrow \infty$. Hence, the mean is consistently $\sqrt{n}$-discriminated. The discrimination results extend to the case with locally dependent observations where we know the local dependence structure and the long run variance is consistently estimable.
However, when we do not know the dependence structure, the consistent $\sqrt{n}$-discrimination of the mean is not guaranteed. We make this explicit in the following corollary which follows immediately from Theorem (ref).
On the other hand, if we have at least two large clusters and do not know the dependence structure within each cluster, we can consistently $\sqrt{n}$-discriminate the mean as long as the within-cluster sample means are asymptotically normal, as shown in the following theorem.
For the theorem, we construct a $t$-test statistic as in Ibragimov/Muller:10:JBES and show that using the test, we can consistently $\sqrt{n}$-discriminate the mean, without knowing the dependence structure within the clusters.
The asymptotic normality condition ((ref)) is often satisfied if the within-cluster dependence is weak. As we show later, this does not mean that we can consistently estimate $\sigma_{n,m}$ for each cluster. (We will study this problem in the next section in detail.) Also, it is important to note that the within-cluster asymptotic normality is not enough to secure the consistent $\sqrt{n}$-discrimination of the mean, if there is only one cluster. In fact, the asymptotic normality condition alone does not exclude the possibility of $\mathcal{P}_{n,\mathcal{N}} \subset \mathcal{P}_n$, and in this case, Corollary (ref) shows that the mean is not consistently $\sqrt{n}$-discriminated.
As mentioned in the introduction, Song:16:arXiv and Leung:21:JAE considered the approach of randomized subsampling inference when one does not know the dependence structure at all. Hence, their situation corresponds to the setting with $M_n = 1$. Their procedure requires the following assumption:
as $n \rightarrow \infty$. If we know the upper bound of the long-run variance such that the upper bound does not change with $n$, it is not hard to see that we can consistently $\sqrt{n}$-discriminate the mean as long as the condition ((ref)) holds. Indeed, we can consider the test where we reject the null hypothesis of $\mathbf{E}[X_{n,i}] = 0$ against $\mathbf{E}[X_{n,i}] > 0$ if and only if
where $c$ is the known upper bound for the long-run variance. In our setting of hypothesis testing, however, the set of probabilities $\mathcal{P}_n$ does not have a finite upper bound for the long-run variance of the sample mean, reflecting the fact that the long-run variance is not known. Thus, the assumption ((ref)) does not hold uniformly over $P \in \mathcal{P}_n$ in our setting, and the results of Song:16:arXiv and Leung:21:JAE do not contradict the impossibility result of Theorem (ref).\footnote{This setting is analogous to that in the standard hypothesis testing with i.i.d.\ normal random variables with the unknown variance. In this standard setting, even with the unknown variance, the mean is typically consistently $\sqrt{n}$-discriminated, because the variance can be consistently estimated. However, as we will see later, in a setting with large clusters, the long-run variance is not consistently estimable.}
\@startsection{section}{1} \z@{1.0\linespacing\@plus\linespacing}{.8\linespacing}{Consistent Estimation of Variance}
Recently, Hansen/Lee:19:JOE showed that it is necessary and sufficient for the weak law of large numbers to hold for the sample average of the clustered observations with
In this section, we show that this condition is necessary and sufficient for consistent estimability of the long run variance. This implies that when there is a large cluster (i.e., which takes up an asymptotically nonnegligible fraction of observations), the long run variance is not consistently estimable. This is a consequence of lack of knowledge of the dependence structure within the large cluster. It means that the usual asymptotic inference based on the asymptotic normal approximation of statistics is generally not applicable in this situation.
\@startsection{subsection}{2} \z@{.8\linespacing\@plus.7\linespacing}{.7\linespacing}{Consistent Estimability}
Let us introduce the notion of consistent estimability of a parameter. Let $\mathcal{P}_n$ be the set of joint distributions of observed random variables, say, $\{X_1,...,X_n\}$. Given a parameter space $\Theta \subset \mathbf{R}^d$, we define our object of interest to be a map $\theta_n: \mathcal{P}_n \rightarrow \Theta$.
One can find a similar definition of consistent estimability in LeCam/Schwartz:60:AMS. They provide necessary and sufficient conditions for a parameter to be consistently estimable when the data are i.i.d. See also Section 1.4 of Ibragimov/Hasminskii:1981:StatEst and Section 6.2 of Pfanzagl:94:ParamStatTheory.
Our setting is somewhat nonstandard, requiring a different technique to prove impossibility of consistent estimation. It is usually assumed that the probability model is indexed by a certain set, i.e., $\mathcal{P}_n = \{P_{n,h}: h \in \mathcal{H}\}$, where each $P_{n,h}$ is a probability measure indexed by $h$ in some topological space $\mathcal{H}$ that is independent of the sample size $n$. One can then redefine the parameter $\psi_n(h) = \theta_n(P_{n,h})$, $h \in \mathcal{H}$, i.e., as a map on $\mathcal{H}$. As long as $\psi_n$ behaves “continuously” on $\mathcal{H}$, the parameter $\psi$ can be shown to be consistently estimable. (See, e.g., Theorem 4.1 of Ibragimov/Hasminskii:1981:StatEst and Theorem 6.2.11 of Pfanzagl:94:ParamStatTheory.) Then the impossibility of consistent estimation stems from the discontinuity of $\psi$ as a map on $\mathcal{H}$, which yields “non-identifiability” of the parameter Potscher:02:Eca.
However, we cannot apply this standard approach in our setting, because there is no natural space $\mathcal{H}$ that indexes $\mathcal{P}_n$ independently of $n$. The main reason is that we need to deal with a situation potentially with a large cluster with an unknown within-cluster dependence structure. This means that we need to require our probability model to accommodate a wide range of dependence structures for the entire sample. For example, suppose that there is only one large cluster, so that one does not know the dependence structure at all. This means, among other things, that our model needs to include various network dependence structures Kojevnikov/Marmer/Song:21:JOE for the joint distribution of the entire random vector $[X_1,...,X_n]$ whose dimension grows with the sample size $n$. One might consider parametrizing the probabilities in terms of the networks governing the dependence structure, but each network depends on the sample size $n$. To the best of our knowledge, there is no obvious way to topologize such a probability model and to define the continuity of the parameter $\theta_n$ on the probabilities, independently of sample size $n$.
Our approach relies on the following simple lemma that uses contiguity of probabilities at a primitive level. For any two sequences of probabilities $P_n$ and $P_n'$, we say that $P_n$ is \bi{contiguous with respect to} $P_n'$ if $P_n' (A_n) \rightarrow 0$ implies $P_n (A_n) \rightarrow 0$ for any sequence of Borel sets $A_n$, and write $P_n \triangleleft P_n'$. When $P_n \triangleleft P_n'$ and $P_n' \triangleleft P_n$, we say that $P_n$ and $P_n'$ are \bi{mutually contiguous}, and write $P_n \triangleleft \triangleright P_n'$. Contiguity between probabilities was introduced by LeCam:60:UCPS and is widely used, especially for deriving the limiting distribution of a test statistic under local alternatives. By tracing out the limiting distribution along a range of local alternatives, one obtains a limiting experiment which one can use to compute the asymptotic risk lower bound in statistical decision theory. (See, e.g., Chapter 6 of vanderVaart:98:AsympStat.)
The following lemma summarizes our scheme of proving the impossibility of consistent estimability of $\sigma_{LR}^2$.
Later we use Lemma (ref) to prove the impossibility of consistent estimation of the long run variance, by selecting two Gaussian probabilities, $P_{n,0}$ and $P_{n,1}$, such that $P_{n,1} \triangleleft P_{n,0}$ and the values of the long run variance stay apart under $P_{n,0}$ and $P_{n,1}$ as $n \rightarrow \infty$. (See the discussion below Theorem (ref).)
The notion of consistent estimability in Definition (ref) coincides with consistent estimability uniform in $P$, i.e., the existence of an estimator $\hat \theta$ such that for each $\epsilon>0$, as $n \rightarrow \infty$,
(See Ibragimov/Hasminskii:1981:StatEst, p.31. See also Potscher:02:Eca for discussion on asymptotics uniform in $P$.) When $\mathcal{P}_{n} = \{P_{n,h}: h \in \mathcal{H}\}$ for some index set $\mathcal{H}$ which does not depend on the sample size $n$, uniform consistent estimability is stronger than pointwise consistent estimability which assumes the existence of an estimator $\hat \theta$ such that for each $h \in \mathcal{H}$ and for each $\epsilon>0$,
as $n \rightarrow \infty$. However, as explained above, in our setting, there is no space $\mathcal{H}$ that indexes $\mathcal{P}_n$ and is independent of $n$. Hence, there is no natural notion of pointwise consistency in $P$ in our set-up.
\@startsection{subsection}{2} \z@{.8\linespacing\@plus.7\linespacing}{.7\linespacing}{Consistent Estimability of Variance in Gaussian Experiments} \@startsection{subsubsection}{3} \z@{.5\linespacing\@plus.7\linespacing}{-.5em}{\normalfont}{A Necessary and Sufficient Condition for Consistent Estimability of Variance}
In our context, a major challenge is to show the contiguity condition (ii) of Lemma (ref). A standard argument proving contiguity utilizes the local asymptotic normality or local asymptotic mixed normality results for a log-likelihood process. However, these latter results often use an i.i.d.\ or time-series set-up where the researcher knows the dependence structure, and Hence, are not useful for our purpose here. For this reason, we focus on a Gaussian experiment, where we can explicitly compute the log-likelihood process in finite samples and investigate its asymptotic behavior as the dependence structure varies. In particular, we consider the following model for a fixed $\sigma^2>0$,
where $\mathbf{1}_n$ denotes the $n$-dimensional vector of ones. The set $\mathcal{P}_{n,\mathcal{N}}(\sigma^2)$ represents the set of LTIC Gaussian models, where each multivariate normal distributions with a common mean and the short run variance equal to $\sigma^2$.
For the impossibility result below, we require that the probability model does not exclude this Gaussian experiment.
The sufficiency part of the theorem is straightforward. To see this, suppose for simplicity that there is no singleton cluster in the data. If ((ref)) is satisfied, it means that the number of clusters $M_n$ grows to infinity as $n \rightarrow \infty$. Then, we consider the following estimator.
In fact, for the sufficiency part, we do not require that $\mathcal{P}_{n,\mathcal{N}}(\sigma^2) \subset \mathcal{P}_n$.
The nontrivial part of the theorem is to show that the condition ((ref)) is necessary for the consistent estimability of $\sigma_{LR}^2$ in $\mathcal{P}_{n,\mathcal{N}}$. Suppose that the condition of ((ref)) fails, which implies that one has at least one nonnegligible cluster. Then, we show that $\sigma_{LR}^2$ is not consistently estimable. For this, we employ Lemma (ref) after computing the log-likelihood process under cluster dependence. More specifically, suppose first that the observations consist of only a single large cluster. Then, we note that the model $\mathcal{P}_n$, due to the lack of knowledge on the dependence structure, does not exclude the LTIC Gaussian experiment: $\Phi(0,\Sigma_n(\sigma^2,\delta))$. Then we show that
whereas
as $n \rightarrow \infty$, for some nonzero constant $c$. Hence, by Lemma (ref), $\sigma_{LR}^2$ cannot be consistently estimated in any probability model that does not exclude the LTIC Gaussian experiment. It is not hard to extend the same arguments to a setting where there are potentially multiple large clusters.
Theorem (ref) then implies that if a nonnegligible portion of the sample belongs to non-singleton clusters, the long run variance $\sigma_{LR}^2$ is consistently estimable in $\mathcal{P}_{n,\mathcal{N}}$ if and only if the cluster structure consists of negligible clusters. We formalize this in the following corollary.
The condition $\liminf_{n \rightarrow \infty} n^*/n > 0$ requires that the fraction of random variables $X_{n,i}$ that do not belong to a singleton cluster is asymptotically nonnegligible. In this case, if the probability model in practice includes the Gaussian model $\mathcal{P}_{n,\mathcal{N}}$ as a subclass and there is at least one nonnegligible cluster, it is not possible to consistently estimate the long run variance. Certainly, this impossibility result carries over to a model where the long run variance $\sigma_{LR}^2$ is allowed to increase with the sample size $n$.
Hence, by combining Theorem (ref) with Theorem (ref), we find that when we have several large clusters, the long run variance is not consistently estimable because the sample contains large clusters, but the mean can still be consistently $\sqrt{n}$-discriminated.
\@startsection{subsubsection}{3} \z@{.5\linespacing\@plus.7\linespacing}{-.5em}{\normalfont}{Implications for Network Dependent Observations}
One might wonder whether the result extends to the case where the observations exhibit a dependence structure other than cluster dependence. Below we give a partial answer for the case of a dependency graph. Dependency graphs were introduced by Stein:86, and have been studied and used in statistics and econometrics. (See, e.g., Aronow&Samii:17, Song:18:ReStat, Leung:20:ReStat and Canen/Schwartz/Song:20:QE and references therein.)
A \bi{graph} (or network) is a pair $G_n = (N_n,E_n)$, where $N_n = \{1,...,n\}$ denotes the set of vertices and $E_n$ the set of edges, where we denote $N(i) = \{j: ij \in E_n\}$ to mean the neighborhood of vertex $i$. (Here, we consider only simple, undirected graphs, i.e., $ii \notin E_n$, for all $ i \in N_n$, and $ij \in E_n$ if and only if $ji \in E_n$.) We define
where $d_{mx}$ is called the \bi{maximum degree}, and $d_{av}$ the \bi{average degree} of the graph $G_n$. The maximum and average degrees are often used to capture the denseness of the graph. A subset of vertices in graph $G_n$ is called a \bi{clique} if any two distinct vertices are adjacent in $G_n$, and the number of vertices in the clique is called the size of the clique. The maximum clique size refers to the size of the clique that is largest in the graph $G_n$.
Recall that a graph $G_n = (N_n,E_n)$ on $N_n = \{1,...,n\}$ is called a \bi{dependency graph} for $X_n = (X_{n,i})_{i \in N_n}$, if for any subset $A \subset N_n$, $(X_{n,i})_{i \in A}$ and $(X_{n,i})_{i \in N_n \setminus \overline N_n(A)}$ is independent, where $\overline N_n(A) = \{j: ij \in E_n, \text{ for some } i \in A\} \cup \{i\}$. It is important to note that while the dependency graph imposes independence between $X_{n,i}$ and $X_{n,j}$ when they are not adjacent in the graph, it says nothing about dependence between them when they are adjacent. Thus, we allow in $\mathcal{P}_n$ any degree of dependence (including independence) between $X_{n,i}$ and $X_{n,j}$ whenever $i$ and $j$ are adjacent in $G_n$. As the dependency graph becomes denser, this reflects our limited knowledge on the dependence structure, similarly to large clusters in the cluster dependence case.
The impossibility result in (i) has an important implication in many models with network dependent observations. As in the case of a dependency graph, many models of network dependence do not specify the strength of dependence between observations that are adjacent in the network Kojevnikov/Marmer/Song:21:JOE. Weak dependence is usually imposed between observations that are far from each other in terms of the shortest path in the network. Hence, when there is a large clique in the network which constitutes a nonnegligible fraction of the entire sample in the limit as $n \rightarrow \infty$, Corollary (ref)(i) implies that the long run variance of the network dependent observations is not consistently estimable.
It is interesting to note that one cannot characterize a necessary and sufficient condition for the network solely in terms of its maximum degree. For example, if $X_n$ is a multivariate normal random vector such that each component has a bounded variance and has a star graph as a dependency graph, the long run variance is consistently estimable. To see this, let $X_n = [X_{n,1},...,X_{n,n}]^\top$ be a centered multivariate random vector which has a graph $G_n$ as a dependency graph. Let the graph $G_n$ be a star graph with the unit $1$ being its center.\footnote{The star graph as a dependency graph is different from an additive common shock model such as $X_{n,i} = C_n + \varepsilon_i$, where $C_n$ is a common shock, and $\varepsilon_i$'s are cross-sectionally independent idiosyncratic shock. In this case, the dependency graph is a complete graph, because every pair of random variables is correlated through the common shock. Hence, the center in the star graph as a dependency graph cannot be a source like a common shock. It is more plausible to imagine the center to be an aggregated outcome of independent sources. In this case, by simply eliminating the star, one obtains independent random variables.} In the context of multivariate normality, we can write
where the leading sum is the best linear projection, so that $\varepsilon$ is independent of $X_{n,i}$'s, $i \in N_n\setminus\{1\}$, which are independent from each other (due to the dependency graph being a star graph). Since the variance of $X_{n,1}$ is bounded, we should have $\sum_{i \in N_n \setminus \{1\}} \theta_i^2 < C$, for all $n \ge 1$, for some $C>0$. Note that we can identify
Now, we can write
The second term is written as
as $n \rightarrow \infty$, because the normalized sum in the parenthesis converges to zero in moments. Hence, we can simply take
to be an estimator of the long run variance. It is not hard to see that $\hat \sigma_{LR}^2$ is consistent for $\sigma_{LR}^2$. This example shows that one cannot express the condition for the consistent estimability solely in terms of the maximum degree of the dependency graph.
\@startsection{section}{1} \z@{1.0\linespacing\@plus\linespacing}{.8\linespacing}{Implications} \@startsection{subsection}{2} \z@{.8\linespacing\@plus.7\linespacing}{.7\linespacing}{A Linear Regression Model with Cluster-Dependent Errors} Let us consider the following regression model with cluster-dependent errors (see, e.g., Cameron/Gelbach/Miller:08:ReStat, Djogbenou/MacKinnon/Nielsen:19:JOE and Hansen/Lee:19:JOE and references therein):
where $y = [y_1',...,y_{M_n}']'$, $X = [X_1',...,X_{M_n}']'$ and $u = [u_1',...,u_{M_n}]'$, with $\mathbf{E}[u_m \mid X] = 0$ for each $m=1,...,M_n$, and each cluster $m$ has $n_m$ observations (so that $y_m$ and $u_m$ are $n_m$ dimensional column vectors, and $X_m$ is an $n_m \times k$ matrix.) We assume that $u_1,...,u_{M_n}$ are independent, but for each $m$, the dependence structure of $u_m$ is not known. We do not exclude the possibility that the error term follows a normal distribution.
Then, the OLS estimator of $\beta$ is given by
The sandwich form of the variance matrix of $\hat \beta$ is given by
Once we obtain a consistent estimator $\hat V$ of $V$, we can construct a standard error of the $j$-th entry of $\beta$, i.e., $\beta_j$, as $\hat \sigma_j^2 = [\hat V]_{jj}$, the $j$-th diagonal of $\hat V$. From the asymptotic normal inference applied to a $t$-statistic for $\beta_j$, we obtain the following confidence interval for $\beta_j$:
As for the consistent estimator $\hat V$, Djogbenou/MacKinnon/Nielsen:19:JOE considered the following estimator:
where $\hat u_m = y_m - X_m' \hat \beta$, and $d$ is a sequence such that $d \rightarrow 1$. They established the consistency of this estimator under a set of conditions, and showed that their conditions are not compatible with a setting in which one of the clusters is large, i.e., its size is proportional to the entire sample.
Our result implies that such an estimator $\hat V$ is not uniformly consistent when there is at least one large cluster. In fact, our result is much stronger than this. It shows that it is not possible to construct a uniformly consistent estimator of $V$ in such a case. Hence, in this case, we cannot construct a confidence interval of the form ((ref)) that is uniformly asymptotically valid. When a nonnegligible fraction of observations belong to a (non-singleton) cluster - which is the case with most cluster-dependence settings, the necessary and sufficient condition for the uniformly consistent estimability of $V$ is that each cluster is asymptotically negligible in the sense of ((ref)).
\@startsection{subsection}{2} \z@{.8\linespacing\@plus.7\linespacing}{.7\linespacing}{Difference-in-Differences with Spillovers} Let us explore the implications of the impossibility results in the context of a difference-in-differences approach to causal inference. (See Section 6.5 of Imbens/Wooldridge:09:JEL for an overview of this method. See also Roth/SantAnna/Bilinski/Poe:22:arXiv for an overview including recent advances in the literature.) Suppose that there are $n$ individuals who are subject to a treatment and the researcher observes their outcomes before and after the treatment. We let $Y_{i,t}(1)$ and $Y_{i,t}(0)$ denote the potential outcomes at time $t = 0,1$ for the treated state and the control state, respectively. As standard in the literature, we assume that in time 0, no individual is treated, and $Y_{i,0} = Y_{i,0}(0)$, which is observed. The observed outcome $Y_{i,1}$ at time $1$ is defined by
where $D_i$ is the indicator of treatment for $i$ that happens between times 0 and 1. Our parameter of interest is the average treatment effect on the treated:
Suppose that we have observations $\{(Y_{i,1},Y_{i,0},D_i)\}_{i=1}^n$, where $Y_{i,0}$ is the outcome for person $i$ at time 0. Furthermore, we assume that the researcher knows the probability $p = P\{D_i =1\}$. (The impossibility result we mention below carries over to the case where $p$ is not known.)
Let us introduce the standard parallel trend assumption used in the literature:
Under this assumption, we can identify
where $\Delta Y_{i} = Y_{i,1} - Y_{i,0}$. We can obtain a sample analogue estimator by
We consider settings where the observed outcomes are cross-sectionally dependent. Our interest is in constructing a confidence interval for $\text{ATT}$ that is uniformly asymptotically valid. Below we consider two situations, one with treatment spillover and the other with spillover of treatment effects. We explore implications of our impossibility results in these situations.
\@startsection{subsubsection}{3} \z@{.5\linespacing\@plus.7\linespacing}{-.5em}{\normalfont}{Treatment Spillover}
Suppose that there is a spillover of the treatments so that $D_i$'s are correlated across $i$, along some network among people. For example, one can think of a situation in a social program where two people $i$ and $j$ are neighbors and participating in the program by $i$ can induce the participation by $j$. Suppose that the researcher does not have information on the neighborhoods among the subjects. Then, this creates dependence among $Y_{i,1}$'s along a dependence structure that is unknown to the researcher. Then, our impossibility result shows that $\text{ATT}$ cannot be consistently discriminated.
In practice, the treatment assignment is often done at the cluster level, where the potential outcomes $Y_{i,1}(1)$ and $Y_{i,1}(0)$ may exhibit arbitrary dependence within each cluster. (See Section 5 of Roth/SantAnna/Bilinski/Poe:22:arXiv for examples and references studying such a setting.)
Suppose that we have at least two large clusters such that $(Y_{i,1}(1), Y_{i,0}(0), D_i)$ are independent across the clusters but arbitrarily correlated within each cluster. The researcher might attempt to test the null hypothesis of $\text{ATT} = 0$ by considering the usual $t$ statistic for testing the null hypothesis of $\text{ATT} = 0$ such that
where $\widehat{\sigma}^2$ is a consistent estimator of the variance of $\sqrt{n}(\widehat{\text{ATT}} - \text{ATT})$, and the critical values taken from the standard normal stable. Our impossibility result implies that it is not possible to consistently estimate the variance of $\sqrt{n}(\widehat{\text{ATT}} - \text{ATT})$, when there is at least one large cluster, and hence, such a $t$-test is not uniformly asymptotically valid. For the same reason, we cannot construct a confidence interval of the following familiar form:
such that the confidence interval is uniformly asymptotically valid. (See Section 5 of Roth/SantAnna/Bilinski/Poe:22:arXiv for various approaches.\footnote{To the best of our knowledge, there is no formal result that proposes a uniformly asymptotically valid confidence interval for $\text{ATT}$ in this setting. However, we expect that the bootstrap approach of Canay/Santos/Shaikh:21:ReStat can be used to construct a uniformly valid confidence interval under mild additional conditions.})
\@startsection{subsubsection}{3} \z@{.5\linespacing\@plus.7\linespacing}{-.5em}{\normalfont}{Spillover of Treatment Effects}
Suppose that the treatments $D_i$ themselves do not exhibit any spillover, but the cross-sectional dependence of $(Y_{i,1}(1), Y_{i,0}(0))$ arises due to the spillover of the treatment effects, for example, the treatment of a person $i$ influences the outcome of the person $j$ in the next period. Such a setting has been studied in the recent literature (see Aronow&Samii:17, Leung:20:ReStat, He/Song:22:WP and references therein.)
Suppose that the spillover of the treatment effects arises along some network among people, and yet the researcher does not have any information on the network. Then, our impossibility result implies that we cannot consistently discriminate ATT in such a situation. However, the researcher may observe a group structure where the spillover does not arise between groups, so that $(Y_{i,1}(1), Y_{i,0}(0), D_i)$ are independent across groups.
If each within-group sum of $(Y_{i,1}(1), Y_{i,0}(0), D_i)$ satisfies the central limit theorem, our result shows that the ATT can be consistently $\sqrt{n}$-discriminated. However, when there is at least one large group, there does not exist a consistent estimator of the variance of $\sqrt{n}(\widehat{\text{ATT}} - \text{ATT})$. Hence, similarly as before, we cannot construct a uniformly asymptotically valid $t$-test for the null hypothesis of $\text{ATT} = 0$ using the usual $t$ statistic and standard normal critical values, and cannot construct a uniformly asymptotically valid confidence interval of the form ((ref)) based on a normal approximation.
\@startsection{section}{1} \z@{1.0\linespacing\@plus\linespacing}{.8\linespacing}{Conclusion}
In this paper, we show two impossibility results on the inference on the mean when the dependence structure is not known. The first result is the impossibility of consistent estimation of the long run variance. The second result is the impossibility of the consistent $\sqrt{n}$-discrimination of the mean. We made an attempt to accommodate partial knowledge of the dependence structure through cluster dependence, and has obtained some necessary and sufficient conditions for the cluster structure for the impossibility results.
While cluster dependence is a popularly used specification of the cross-sectional dependence structure, it is not general enough to accommodate other forms of a dependence structure such as dependency graphs, Markov graphs, and network dependence. It would be interesting to investigate the implications of partial knowledge of a dependence structure for a more general setting. We leave this for future research.
\@startsection{section}{1} \z@{1.0\linespacing\@plus\linespacing}{.8\linespacing}{Appendix: Mathematical Proofs}
\@startsection{subsection}{2} \z@{.8\linespacing\@plus.7\linespacing}{.7\linespacing}{Preliminary Results}
For the proof of the main results, we first prove auxiliary lemmas. As a first step, we provide an explicit form of a log-likelihood process in Gaussian experiments in Lemma (ref). For this, we use the following auxiliary lemma.
Proof: Let $Q=U S^{1/2}$ such that $\Sigma_0 =Q B B^{\top}Q^{\top}$ and $\Sigma_{n,1}=Q B(I +\Lambda) B^{\top} Q^{\top}$. Thus,
and
where \[ \tilde{\Lambda} = \operatorname{diag}\left(\frac{\lambda_1}{1+\lambda_1},\ldots,\frac{\lambda_n}{1+\lambda_n}\right). \] $\blacksquare$
The following lemma provides an explicit form of a general log-likelihood process for Gaussian measures. Recall that $\Phi(\mu,\Sigma)$ denotes the multivariate normal distribution with mean vector $\mu$ and covariance matrix $\Sigma$.
Proof: We write
We apply Lemma (ref)(i) to the first term on the right hand side. As for the last term we let $\mu_{\Delta} = \mu_1 - \mu_0$, and $x_* = x - \mu_0$. Note that
Similarly, $\mu_{\Delta}^\top \Sigma_0^{-1} x_* = \tilde \mu^\top Z(x).$ We rewrite the last term in ((ref)) as
(by applying Lemma (ref)(ii)). By rearranging terms, we rewrite the last sum as
Combining this with an earlier result, we obtain the desired result. $\blacksquare$
Proof: We apply Lemma (ref) with $S = U = I_n$,
Note that the spectral decomposition of $A$ is given by $B \Lambda B^\top$, where $\Lambda$ is the diagonal matrix with the diagonal elements $\lambda_1,...,\lambda_n$ given as $\lambda_1 = (\sigma^2 - 1) + \sigma^2(n-1)\delta$, $\lambda_2 = ... = \lambda_n = - \sigma^2 \delta$, and the orthogonal matrix $B$ as given the lemma. The desired result follows from Lemma (ref). $\blacksquare$
Lemma (ref) yields the following result for the case with cluster dependence. From here on, we make the dimension of the matrices and vectors explicit. Let $I_{n_m}$ be the $n_m$-dimensional identity matrix and $\mathbf{1}_{n_m}$ denote the $n_m$-dimensional column vector of ones.
Proof: Using the Mean Value Theorem,
where $x^*(x)$ is a point on the line segment between $0$ and $x$. Evaluating the inequality at $x=\overline x$ gives us the desired result. $\blacksquare$
Recall the defintion of $n^*$ in Corollary (ref):
The number $n - n^*$ represents the number of random variables, $X_{n,i}$, that are known to be mutually independent. Each variable outside this set belongs to a non-singleton cluster.
Proof: For each $n \ge 1$, we have either
Since $(1/n)^2(n- n^*) = o(1)$, $\lim_{n \rightarrow \infty} \sum_{m = 1}^{M_n} \left(n_m / n\right)^2 = 0$ if and only if (a) or (b) holds. $\blacksquare$
Proof: For brevity, we focus on the case $\sigma^2 = 1$. By Corollary (ref),
where
and $i_m$ denotes the first $i$ in block $m$.
(i) Let us take small $\epsilon>0$ such that
We write (under $\Phi(0,I_n)$)
since $A_n$ and $R_n$ are independent. Let $t_m = (n_m-1)/n^*$, and write
Note that
because $t_m \le 1$ and $\overline \delta ( 1+ \epsilon) <1$ by ((ref)). This means that $f_n(\overline \delta)$ is increasing in $\overline \delta \in [-a,a]$ and achieves its maximum at $\overline \delta = a$. Hence,
because $\sum_{m=1}^{M_n} t_m^2 \le 1$ and we chose $\epsilon$ such that ((ref)) holds. By Lemma (ref), we have
The bound does not depend on $n$, and hence,
Now, we turn to $ \mathbf{E}\left[\exp((1 + \epsilon)R_n) \right]$. We can write
Using this expression, we rewrite
The last bound is a sequence converging to $\exp(a(1 + \epsilon)/2)$ as $n^* \rightarrow \infty$, and Hence, is a bounded sequence. Thus, we conclude that
This proves that
Hence, the proof of (i) is complete.
(ii) We rewrite
For any $x \in [0,1]$, we have
Hence,
It suffices to show the uniform tightness of the second sum in ((ref)). Under $\Phi(0,I_n)$, it has mean zero, and
because $\operatorname{Var}(Z_{m,i_m}^2) = 2$. Therefore, $A_n$ is uniformly tight.
As for $R_n$, we recall ((ref)), and can follow similar arguments to show that $R_n$ is uniformly tight as well. $\blacksquare$
\@startsection{subsection}{2} \z@{.8\linespacing\@plus.7\linespacing}{.7\linespacing}{Consistent Discrimination of the Mean}
We let
and for any $\overline \mu \in \mathbf{R}$, we write $\Phi(\overline \mu \mathbf{1}_n,\Sigma(\sigma^2,\delta))$ simply as $\Phi(\overline \mu,\sigma^2,\delta)$. Let us recall some basic notions of optimality of tests Lehmann/Romano:05:TSH. Given a model $\mathcal{P}_{n}$ which is partitioned as $\mathcal{P}_{n,0} \cup \mathcal{P}_{n,1}$, a test $\phi_n$ is said to be a \bi{UMP (uniformly most powerful) test} of $\mathcal{P}_{n,0}$ against $\mathcal{P}_{n,1}$ at level $\alpha \in (0,1)$, if under any $P_{n,0} \in \mathcal{P}_{n,0}$,
and for any alternative test $\phi_n'$ such that $\mathbf{E}[\phi_n'] \le \alpha$ under any $P_{n,0} \in \mathcal{P}_{n,0}$, we have
under any $P_{n,1} \in \mathcal{P}_{n,1}$.
A sequence of tests $\phi_n$ is said to be an \bi{AUMP (asymptotically uniformly most powerful) test} of $\mathcal{P}_{n,0}$ against $\mathcal{P}_{n,1}$ at level $\alpha \in (0,1)$, if under any sequence $P_{n,0} \in \mathcal{P}_{n,0}$,
and for any alternative test $\phi_n'$ such that $\limsup_{n \rightarrow \infty} \mathbf{E}[\phi_n'] \le \alpha$ under any sequence $P_{n,0} \in \mathcal{P}_{n,0}$, we have
under any sequence $P_{n,1} \in \mathcal{P}_{n,1}$.
Proof: Choose any test $\tilde \varphi$ such that under any sequence $P_n \in \mathcal{P}_{n,0}$,
for some sequence $\epsilon_n \rightarrow 0$, as $n \rightarrow \infty$. Fix one such sequence $P_n\in \mathcal{P}_{n,0}$, together with the sequence $\epsilon_n$, and let $\tilde \alpha = \alpha + \epsilon_n$. Now, select a large enough $n$ such that $\tilde \alpha \in A$ and choose any $P_n' \in \mathcal{P}_{n,1}$. Then, since $\varphi_{\tilde \alpha}$ is UMP at level $\tilde \alpha$, we have
under $P_n'$. By ((ref)), we can see that $\varphi_{\alpha}$ is AUMP at level $\alpha$. $\blacksquare$
Proof of Theorem (ref): Suppose that $M_n = 1$. First, consider the case where $\mathcal{P}_{n} = \mathcal{P}_{n,\mathcal{N}}'$, with
Later, we generalize the result to the case where $\mathcal{P}_n$ contains the above probability model. Define $\mathcal{P}_{n,0} = \left\{\Phi(0,\sigma^2,\delta): (n-1) \delta \in (0,1), \sigma^2>0 \right\}$ and let $\mathcal{P}_{n,1} = \mathcal{P}_{n} \setminus \mathcal{P}_{n,0}$. In light of Lemma (ref), it suffices to construct a class of tests $\{\varphi_{\alpha}\}_{\alpha \in (0,1/2)}$ of $\mathcal{P}_{n,0}$ against $\mathcal{P}_{n,1}$ such that
(a) it satisfies ((ref)) in Lemma (ref),
(b) the test $\varphi_{\alpha}$ has power bounded by a constant below $1$ uniformly over $n$, and
(c) each test $\varphi_{\alpha}$ is a UMP test of $\mathcal{P}_{n,0}$ against $\mathcal{P}_{n,1}$.
Let us first construct such a test and show that (a)-(c) are satisfied. Define
For each $\alpha \in (0,1/2)$, let
for some $C_0>0$ and $\gamma_0 \in [0,1]$. Let $Z = (Z_{n,1} - \mathbf{E}[Z_{n,1}])/\sqrt{\operatorname{Var}(Z_{n,1})}$. Then the size control requires that under the null hypothesis,
Since $\alpha \in (0,1/2)$, we must have $C_0 = 1$ and $\gamma_0 = 2 \alpha$.
Let us first show that this test satisfies the condition (a). For any $\tilde \alpha$ such that $\tilde \alpha = \alpha + o(1)$, and under any sequence $P_n \in \mathcal{P}_{n}$,
Hence, the class of tests $\{\phi_{\alpha}(V_n)\}_{\alpha \in (0,1/2)}$ satisfies the condition ((ref)).
As for the condition (b), note that under any alternative hypothesis in $\mathcal{P}_{n,1}$, we have
Hence, the test does not have power exceeding $2\alpha < 1$.
Finally, we show that the condition (c) is satisfied. Let
Hence, $\mathcal{L}(X_n;\mu,\sigma^2,\delta_n)$ is the same as $\log\left(d \Phi(\mu \mathbf{1}_n,\Sigma(\sigma^2,\delta_n)/d \Phi(0,I_n)\right)(X_n)$ in Lemma (ref), except that the coefficient of $Z_{n,1}$ is $\sqrt{n}\mu$ and the last term is different. Define a probability measure $P_n(\mu,\delta_n)$ as follows: for any Borel $B$,
Similarly as before, we define
and let $\mathcal{\tilde P}_{n,1} = \mathcal{\tilde P}_{n} \setminus \mathcal{\tilde P}_{n,0}$. It is not hard to see that
Therefore, a UMP test of $\mathcal{\tilde P}_{n,0}$ against $\mathcal{\tilde P}_{n,1}$ is also a UMP test of $\mathcal{P}_{n,0}$ against $\mathcal{P}_{n,1}$. It suffices for condition (c) to show that the test $\varphi_\alpha$ is a UMP test of $\mathcal{\tilde P}_{n,0}$ against $\mathcal{\tilde P}_{n,1}$. From ((ref)), the sufficient statistics for $\mathcal{\tilde P}_{n}$ in the case of $M_n = 1$ are given by
where $Z_{n,k}$'s are as in Lemma (ref). For any $t \ge 0$,
under the null hypothesis. Hence, $V_n$ and $Z_{n,1}^2$ are independent under any probability in $\mathcal{\tilde P}_{n,0}$. Furthermore, under any probability in $\mathcal{\tilde P}_{n,0}$,
Note that $B_n^\top \mathbf{1}_n \mathbf{1}_n B_n$ is a matrix whose $(1,1)$-th entry is $b_1^\top \mathbf{1}_n \mathbf{1}_n^\top b_1 = 1$ and all the other entries are zeros. Hence, $Z_{n,k}$'s are independent across $k$'s under any probability in $\mathcal{\tilde P}_{n,0}$. Therefore, $V_n$ and $(Z_{n,1}^2,\sum_{k=2}^n Z_{n,k}^2)$ are independent under any probability in $\mathcal{\tilde P}_{n,0}$. By Theorem 5.1.1 of Lehmann/Romano:05:TSH, the randomized test $\varphi_\alpha(V_n)$ is an $\alpha$-level UMP test.
Next, consider the case where $\mathcal{P}_{n,\mathcal{N}}' \subset \mathcal{P}_n$. Take a sequence of tests $\varphi_n$ such that for any sequence of probabilities $P_n \in \mathcal{P}_{n,0}$, $\limsup_{n \rightarrow \infty} \mathbf{E}[\varphi_n(X_n)] \le \alpha$. Now, we take a sequence $P_{n,1} \in \mathcal{P}_{n,1} \cap \mathcal{P}_{n,\mathcal{N}}'$. For any $\alpha \in (0,1/2)$, the test $\phi_\alpha(V_n)$ is a UMP test at level $\alpha$ of the null hypothesis $\mathcal{P}_{n,0} \cap \mathcal{P}_{n,\mathcal{N}}'$ against $\mathcal{P}_{n,1} \cap \mathcal{P}_{n,\mathcal{N}}'$. Note that for any sequence of probabilities $P_n \in \mathcal{P}_{n,0} \cap \mathcal{P}_{n,\mathcal{N}}'$, $\limsup_{n \rightarrow \infty} \mathbf{E}[\varphi_n(X_n)] \le \alpha$. Hence, if we take $\epsilon>0$ such that $\alpha + \epsilon \in (0,1/2)$, there exists $n_0 \ge 1$ such that for all $n \ge n_0$, $\mathbf{E}[\varphi_n(X_n)] \le \alpha + \epsilon$. For all such $n$, under any $P_n \in \mathcal{P}_{n,1} \cap \mathcal{P}_{n,\mathcal{N}}'$, we have
Hence, we find that along any sequence $P_n \in \mathcal{P}_{n,1} \cap \mathcal{P}_{n,\mathcal{N}}'$, we have
Thus, the proof is complete. $\blacksquare$
The following lemma is used for the proof of Theorem (ref). For $w \in [0,1]$, define $\xi_1 = w Z_1$ and $\xi_2 = (1-w)Z_2$, where $Z_i \sim N(0,1)$, independent across $i=1,2$. Define
where $\bar \xi = (\xi_1 + \xi_2)/2$. Let us take $t \ge 0$, and define
Proof: First, we write
For (i) and (ii), since $[\xi_1,\xi_2]$ is symmetrically distributed around the origin, it suffices to show that $p(\,\cdot\,;t)$ is increasing on $[1/2,1]$ for all $t \ge 1$, and $p(\,\cdot\,;t)$ is decreasing on $[1/2,1]$ for all $0 \le t<1$. Let $A$ denote the event that $\xi_1 + \xi_2 <0$. Then, on the event $A$, for all $w \ge 0$ and all $t \ge 0$, we have $T(w) \le t$. Hence,
On the event $A^c$, $T(w) \le t$ if and only if
if and only if
where
Take $t \ge 1$. Let $B$ be the event $Z_1 Z_2 \le 0$. Certainly, if $Z_1 Z_2 \le 0$, $f(w; Z_1,Z_2) \ge 0$ for all $w \in [1/2,1]$. Hence
We show that $f(w;Z_1,Z_2)$ is increasing in $w$ on the event $A^c$. We take the derivative $f'(w;Z_1,Z_2)$ with respect to $w$:
The function is linear in $w$. First, we take $w =1$. Then, on the event $B^c$,
Second, we take $w=1/2$. Then,
Hence, for all $Z_1,Z_2$ such that $Z_1 Z_2 > 0$, $f(w;Z_1,Z_2)$ is increasing on $[1/2,1]$ for all $t \ge 1$. Therefore, whenever $t \ge 1$, $P\{T(w) \le t\}$ is increasing on $[1/2,1]$.
Take $0 \le t <1$. If $Z_1 Z_2 > 0$, then $f(w;Z_1,Z_2) < 0$. Hence
On the event $B$, when $w=1$,
and when $w=1/2$,
Hence, for all $Z_1,Z_2$ such that $Z_1 Z_2 \le 0$, $f(w;Z_1,Z_2)$ is decreasing on $[1/2,1]$ for all $0 \le t <1$. Therefore, whenever $0 \le t < 1$, $P\{T(w) \le t\}$ is decreasing on $[1/2,1]$. $\blacksquare$
Proof of Theorem (ref): Suppose that we have at least two nonnegligible clusters, i.e., $M_n \ge 2$ for all but finite number of $n$'s. Consider testing the null hypothesis of $\mathbf{E}[X_{n,i}] = 0$ against $\mathbf{E}[X_{n,i}] > 0$. Without loss of generality, we enumerate $\mathcal{M}_n' = \{N_1,...,N_{M_n'}\}$, $M_n' \ge 2$. Now, we construct a test that consistently $\sqrt{n}$-discriminates the mean. Define
and
We take
Let $c_{1 - \alpha}$ be the $1 - \alpha$ quantile of the $t$-distribution with degree of freedom $1$. Define
Note that $c_{0.75} = 1$.
We first show that this test controls the size of the test under $\alpha$ asymptotically under the null hypothesis. Define an infeasible test statistic
where
Here $\sigma_{n,m}^2$ is the variance of $\xi_m$ under the null hypothesis. Then $V_n''$ converges in distribution to the $t$-distribution with $1$ degree of freedom under the null hypothesis. By Lemma (ref), if $0< \alpha < 0.25$ so that $c_{1-\alpha} >1$,
and if $\alpha \ge 0.25$ so that $0 \le c_{1-\alpha} \le 1$,
as $n \rightarrow \infty$. Hence, the size of the test $\varphi_n$ is bounded by $\alpha$ asymptotically.
Suppose that we are under the local alternatives such that $\mathbf{E}[X_n] = \overline \mu_n \sigma_{LR}\mathbf{1}_n/\sqrt{n}$, for some sequence $\overline \mu_n \rightarrow \infty$. Define
Note that
because $n_m /n \le 1$. Note that
Since
there exist $\epsilon>0$ and $n_0>0$ such that for all $n \ge n_0$,
Therefore,
Since $(\xi_m - \mathbf{E}\xi_m)/\sigma_{n,m}$ converges in distribution to $N(0,1)$ under any sequence $P_n \in \mathcal{P}_n$ as $n \rightarrow \infty$ by ((ref)), we have $\bar U_n /\sigma_{LR} = O_P(1)$. Similarly, we can show that $T_n'/\sigma_{LR}^2 = O_P(1)$. Hence, the last probability in ((ref)) converges to one as $\overline \mu_n \rightarrow \infty$, proving that the mean is consistently $\sqrt{n}$-discriminated. $\blacksquare$
\@startsection{subsection}{2} \z@{.8\linespacing\@plus.7\linespacing}{.7\linespacing}{Impossibility of Consistent Estimation of Long Run Variance}
Proof of Lemma (ref): We first show sufficiency. Suppose that $\|\theta_{n}(P_{n,1}) - \theta_{n}(P_{n,0})\| = o(1)$ for any sequence $P_{n,1} \in \mathcal{P}_n$. Then we take $\hat \theta = \theta_n(P_{n,0})$, so that $\| \hat \theta - \theta_n(P_{n,1})\| \rightarrow 0$, along $P_{n,1} \in \mathcal{P}_n$, as $n \rightarrow \infty$. Hence, sufficiency follows.
Conversely, suppose that $\theta_n$ is consistently estimable in $\mathcal{P}_n$, so that there exists an estimator, say, $\tilde \theta$, such that $\tilde \theta - \theta_n(P_{n,1}) = o_P(1)$ along any $P_{n,1} \in \mathcal{P}_n$. Since $P_{n,0} \in \mathcal{P}_n$, this means that $\tilde \theta - \theta_n(P_{n,0}) = o_P(1)$ under $P_{n,0}$. Since $P_{n,1} \triangleleft P_{n,0}$, $\tilde \theta - \theta_n(P_{n,0}) = o_P(1)$ under any $P_{n,1} \in \mathcal{P}_n$. We choose any $P_{n,1} \in \mathcal{P}_n$ and write
The difference on the left hand side and the first difference on the right hand side are $o_P(1)$ under $P_{n,1}$. This implies that $\|\theta_n(P_{n,0}) - \theta_n(P_{n,1})\| = o(1)$. $\blacksquare$
Proof of Theorem (ref): Let us first show sufficiency. Suppose that either (a) or (b) in Lemma (ref) holds. Let us take
Since $\sigma_{LR}^2 \le c$ for all $n \ge 1$, we have
Hence
Note that
Choose $P_n \in \mathcal{P}_n$. Then, under $P_n$,
where $C'>0$ is a constant that does not depend on $n$. By ((ref)), the last term is $o(1)$, if either of the conditions (a) and (b) in Lemma (ref). Therefore, $\sigma_{LR}^2$ is consistently estimable in $\mathcal{P}_n$.
Now, let us show necessity. Suppose that both (a) and (b) in Lemma (ref) are violated. That is,
We fix $\sigma^2$ and show that $\sigma_{LR}^2$ is not consistently estimable in $\mathcal{P}_{n,\mathcal{N}}(\sigma^2)$. We choose $\delta_{n,m}>0$ for each cluster $m$ and $n$ such that ((ref)) holds for some $\overline\delta \in [-a,a]\setminus \{0\}$, $a< 1/2$, if $n_m \ge 2$. By ((ref)), there exists a subsequence $\{n_k\} \subset \{n\}$ such that
for some constants $\tilde c_1, \tilde c_2 \in (0,1]$. For simplicity, we fix this subsequence, and denote $n_k$ by $n$.
We let $\Sigma_n$ be the block diagonal $n \times n$ matrix whose $m$-th block is given by $\sigma^2((1-\delta_{n,m})I_{n_m} + \delta_{n,m} \textbf{1}_{n_m} \textbf{1}_{n_m}^\top)$. We show that $\Phi(0,\Sigma_n) \triangleleft \Phi(0,\sigma^2 I_n)$. First, we observe that by Lemma (ref)(ii), $\log (d \Phi(0,\Sigma_n)/d \Phi(0,\sigma^2 I_n))(X_n)$ is uniformly tight under $\Phi(0,\sigma^2 I_n)$. Furthermore, we find that by Prohorov's Theorem, there exists a subsequence $\{n_k\}$ of $\{n\}$ such that the sequence $\log (d\Phi(0,\Sigma_{n_k})/d \Phi(0,\sigma^2 I_{n_k}))(X_n)$ weakly converges. Let $W$ be a random variable whose distribution is identical to the weak limit. By the Continuous Mapping Theorem, we have
along the subsequence $\{n_k\}$. Note that $\mathbf{E}[(d \Phi(0,\Sigma_{n_k})/d\Phi(0,\sigma^2 I_{n_k}))(X_{n_k})] =1 $, where the expectation is under $\Phi(0,\sigma^2 I_{n_k})$. By Lemma (ref)(i), $(d \Phi(0,\Sigma_{n_k})/d\Phi(0,\sigma^2 I_{n_k}))(X_{n_k})$ is uniformly integrable under $\Phi(0,\sigma^2 I_{n_k})$. Hence, we find that $\mathbf{E} e^W =1$. By Le Cam's First Lemma vanderVaart:98:AsympStat, we conclude that $\Phi(0,\Sigma_{n_k}) \triangleleft \Phi(0,\sigma^2 I_{n_k})$.
On the other hand, note that the difference between the long-run variances under $\Phi(0,\Sigma_n)$ and under $\Phi(0,\sigma^2 I_n)$ is given by
where $\Delta_{n,m} = \delta_{n,m} \textbf{1}_{n_m} \textbf{1}_{n_m}^\top - \delta_{n,m} I_{n_m}$. We rewrite the last term as
as $n \rightarrow \infty$, where the last convergence is due to ((ref)). Certainly, $\sigma_{LR}^2$ is consistently estimable along $\Phi(0,\sigma^2 I_n)$. By Lemma (ref), we conclude that $\sigma_{LR}^2$ is not consistently estimable in $\mathcal{P}_{n,\mathcal{N}}(\sigma^2)$. Hence, it is not consistently estimable in $\mathcal{P}_{n}$ either. $\blacksquare$
Proof of Corollary (ref): Let us show sufficiency. First suppose that $\mathcal{M}_n$ consists of negligible clusters, so that $\max_{1 \le m \le M_n} n_m / n \rightarrow 0$, as $n \rightarrow \infty$. From ((ref)), this implies that
as $n \rightarrow \infty$. Therefore, $\sigma_{LR}^2$ is consistently estimable in $\mathcal{P}_{n}$ by Theorem (ref).
Conversely, suppose that for some $\epsilon>0$,
Then there exist subsequences $\{n_k\} \subset \{n\}$ and $\{n_{m(n_k)}\} \subset \{n_{m(n)}\}_{n \ge 1}$, $m(n) \in \{1,...,M_n\}$, such that
This implies that
Hence, $\sigma_{LR}^2$ is not consistently estimable in $\mathcal{P}_{n}$ by Theorem (ref). $\blacksquare$
Proof of Corollary (ref): (i) Let $N_\circ \subset \{1,...,n\}$ be the set of nodes in the clique with size $n_C$. Take $\mathcal{M}_n$ to be the cluster structure such that there is only one non-singleton cluster that is $N_\circ$. It suffices to show that $\sigma_{LR}^2$ is not consistently estimable in $\mathcal{P}_{n,\mathcal{N}}(\sigma^2)$ with the cluster structure $\mathcal{M}_n$. Note that
because we have only one non-singleton cluster in $\mathcal{M}_n$. By Theorem (ref), $\sigma_{LR}^2$ is not consistently estimable in $\mathcal{P}_{n,\mathcal{N}}(\sigma^2)$ with any fixed $\sigma^2>0$.
(ii) Let us define $\overline N(i) = \{j \in N_n: ij \in \overline E_n\}$, where $\overline E_n = E_n \cup \{ii: i \in N_n\}$, and consider
By rearranging terms, we can write
We can write the squared $L^2$ norm of the leading term on the right hand side as
If $\{i_1,j_1\}$ and $\{i_2,j_2\}$ are not adjacent in $G_n$, the covariance above is zero by the dependency graph assumption. The number of the terms in the above sum such that $\{i_1,j_1\}$ and $\{i_2,j_2\}$ are adjacent in $G_n$ is of the order $O(nd_{mx}^2 d_{av}) = o(n^2)$. The last rate comes from our assumption that $\lim_{n \rightarrow \infty} d_{mx}^2 d_{av} /n = 0$. Therefore, the leading term on the right hand side of ((ref)) is $o_P(1)$. Similarly, we can show that the remainder terms are $o_P(1)$. Hence, $\hat \sigma_{LR}^2$ is a consistent estimator of $\sigma_{LR}^2$. $\blacksquare$