Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
65,817 characters · 15 sections · 51 citation commands
Dependence-Robust Inference Using Resampled Statistics
\addcontentsline{toc}{part}{Main Paper}
\onehalfspacing
This paper builds on randomized subsampling tests due to song_ordering-free_2016 and proposes inference procedures for settings in which the dependence structure of the data is complex or unknown. This is useful for a variety of applications using, for example, network data, clustered data when cluster memberships are imperfectly observed, or spatial data with unknown locations. The proposed procedures compare a test statistic, constructed using a set of resampled observations, to a critical value constructed either using a normal approximation or by resampling. Computation is the same regardless of the dependence structure. We prove that our procedures are asymptotically valid under the weak requirement that the target parameter can be consistently estimated at the $\sqrt{n}$ rate, a condition satisfied by most forms of weakly dependent data for regular estimators. In this sense, inference using our resampled statistics is robust to quite general forms of weak dependence.
Consider the simple problem of inference on a scalar population mean $\mu_0$. Typically we assume that the sample mean $\bar{X}$ is asymptotically normal in the sense that
However, in a setting with complex or unknown forms of dependence, it may be difficult or unclear how to estimate $\bm{\Sigma}$ due to the covariance terms. A simple statistic we utilize for inference in this setting is
where $\bar{X}^*$ is the mean of $R_n$ draws with replacement from the data and $\hat{\bm{\Sigma}} = n^{-1} \sum_{i=1}^n (X_i-\bar{X})^2$ is the (naive) sample variance. Using the identity
where $\bm{X}$ is the data, we show that $[I] \stackrel{d}\longrightarrow \mathcal{N}(0,1)$ and the bias term $[II]$ is asymptotically negligible if $R_n$ is chosen to diverge at a sufficiently slow rate. We can therefore compare $\tilde T_M$ to a normal critical value to conduct inference on $\mu_0$ regardless of the underlying dependence structure. As we later discuss, larger values of $R_n$ generate higher power through a faster rate of convergence for $[I]$ but also a larger bias term $[II]$, which generates size distortion. Thus for practical implementation, we suggest a rule of thumb for $R_n$ that accounts for this trade-off.
song_ordering-free_2016 proposes tests for moment equalities when the data satisfies a particular form of weak dependence known as “local dependence” in which the dependence structure is characterized by a graph. The novelty of his procedure is that its implementation and asymptotic validity does not require knowledge of the graph. We study essentially the same test statistic but generalize his theoretical results, showing asymptotic validity under the substantially weaker requirement of estimability at a $\sqrt{n}$ rate, which significantly broadens the applicability of the method. Additionally, we propose new methods for testing moment inequalities.
Many resampling methods are available for spatial, temporal, and clustered data when the dependence structure is known cameron2008bootstrap,lahiri2013resampling,politis_subsampling_1999. These methods are used to construct critical values for a test statistic computed with the original dataset $\bm{X}$. For the critical values to be asymptotically valid, resampling is implemented in a way that mimics the dependence structure of the data, commonly by redrawing blocks of neighboring observations, but this requires knowledge of the dependence structure. In contrast, our procedures involve computing a resampled test statistic and critical values based on its limiting distribution conditional on the data. Hence, there is no need to mimic the actual dependence structure, which is why our procedures can be dependence-robust.
Of course, the broad applicability of our procedures comes at a cost. The main drawback is inefficiency due to the fact that the test statistics are essentially computed from a subsample of observations. In contrast, the test statistic under conventional subsampling, for example, utilizes the full sample. Thus, in settings where inference procedures exist, our resampled statistics suffer from slower rates of convergence, which yield tests with lower power and can exacerbate finite-sample concerns such as weak instruments. We interpret this as the cost of dependence robustness. Our objective is not to propose a procedure that is competitive with existing procedures but rather to provide a broadly applicable and robust inference procedure that can be useful when little is known about the dependence structure or when this structure is complex and no inference procedure is presently available.
We consider four applications. The first is regression with unknown forms of weak dependence. This setting is relevant when the dependent or independent variables are functions of a social network, as in network regressions chandrasekhar2016econometrics. Another special case is cluster dependence when the level of clustering is unknown and the number of clusters is small, settings in which conventional clustered standard errors can perform poorly cameron2015practitioner. The second application is estimating treatment spillovers on a partially observed network. The third is inference on network statistics, a challenging setting because different network formation models induce different dependence structures. The fourth application is testing for a power law distribution, a problem that has received a great deal of attention in economics, network science, biology, and physics barabasi1999emergence,gabaix2009power,newman2005power. Widely used methods in practice assume that the underlying data is i.i.d.\ clauset2009power,klaus2011statistical, which is often implausible in applications involving spatial, financial, or network data.
The outline of the paper is as follows. The next section introduces our inference procedures. We then discuss four applications in (ref), followed by an empirical illustration on testing for power law degree distributions in (ref). Next, (ref) states formal results on the asymptotic validity of our procedures. In (ref), we present simulation results from four different data-generating processes. Lastly, (ref) concludes.
We begin with a description of our proposed inference procedures. Throughout, let $\bm{X} = \{X_i\}_{i=1}^n \subseteq \mathbb{R}^m$ be a set of $n$ identically distributed random vectors with possibly dependent row elements. Denote the sample mean of $\bm{X}$ by $\bar{X}$. The goal is to conduct inference on some parameter $\mu_0 \in \mathbb{R}^m$. A simple example is the population mean $\mu_0 = {\bf E}[X_1]$, but we will also consider other parameters when discussing asymptotically linear estimators. Our main assumption will require $\bm{X}$ to be weakly dependent in the sense that $\bar{X}$ is $\sqrt{n}$-consistent for $\mu_0$.
We first consider testing the null hypothesis that $\mu_0 = \mu$ for some $\mu \in \mathbb{R}^m$ and constructing confidence regions for $\mu_0$. Let $R_n \geq 2$ be an integer and $\Pi$ the set of all bijections (permutation functions) on $\{1, \dots, n\}$. Let $\{\pi_r\}_{r=1}^{R_n}$ be a set of $R_n$ i.i.d.\ uniform draws from $\Pi$ and $\pi = (\pi_r)_{r=1}^{R_n}$. Define the sample variance matrix $\hat{\bm{\Sigma}} = n^{-1}\sum_{i=1}^n (X_i - \bar{X})(X_i - \bar{X})'$.
We focus on two test statistics. The first is the {\em mean-type statistic}, given by
That is, $\tilde T_M(\mu;\pi)$ is computed by drawing $R_n$ observations with replacement from $\{\hat{\bm{\Sigma}}^{-1/2}(X_i - \mu)\}_{i=1}^n$ and then taking the average and scaling up by $\sqrt{R_n}$. Note that we compute $\hat{\bm{\Sigma}}$ using the full sample.
The second test statistic is the {\em U-type statistic}, which is given by
and essentially follows song_ordering-free_2016. Unlike the mean-type statistic, here we resample pairs of observations with replacement and compute a quadratic form.
{\bf Inference Procedures.} We prove that if $\bar{X}$ is $\sqrt{n}$-consistent for $\mu_0$, then
(under regularity conditions), where $\bm{I}_m$ is the $m\times m$ identity matrix. CLTs for a wide range of notions of weak dependence, including mixing, near-epoch dependence, various forms of network dependence, etc., can be used to verify $\sqrt{n}$-consistency. There are some examples of dependent data that violate $\sqrt{n}$-consistency. One is cluster dependence with many clusters, large cluster sizes, and strongly dependent observations within clusters hansen2019asymptotic. We discuss in (ref) below how other rates of convergence can be accommodated by adjusting $R_n$.
Result (ref) enables us to construct critical values for testing. For example, to test the null that $\mu_0 = \mu$ against a two-sided alternative, we can use
where $z_{1-\alpha}$ and $q_{1-\alpha}$ are respectively the $(1-\alpha)$-quantiles of the standard normal distribution and chi-square distribution with $m$ degrees of freedom. Note that the U-type statistic is two-sided in nature, which is why in (ref) we simply compare $T_U(\mu; \pi)$, rather than its absolute value, to a normal quantile. To test one-sided alternatives with the U-type statistic, we can additionally exploit the sign of $\bar{X}$, as done in the application in (ref) and the moment inequality test below. For instance, we can choose to reject only if the test statistic exceeds its critical value and the sign of $\bar{X}$ is positive.
In the case of scalar data ($m=1$), we obtain the following simple confidence interval (CI) for $\mu_0$ using the mean-type statistic:
Alternatively, we can use the U-type statistic to obtain a CI by test inversion.
{\bf Why This Works.} To see the intuition behind (ref), consider the mean-type statistic. Define $W_{M,r} = \hat{\bm{\Sigma}}^{-1/2}(X_{\pi_r(1)} - \mu_0)$. Then
Some algebra shows that $[II] = (R_n/n)^{1/2} \hat{\bm{\Sigma}}^{-1/2} n^{-1/2} \sum_{i=1}^n (X_i-\mu_0)$. Since $R_n/n = o(1)$ and $n^{-1/2} \sum_{i=1}^n (X_i-\mu_0) = O_p(1)$ under the assumption of $\sqrt{n}$-consistency, we have $[II] = o_p(1)$, provided the sample variance converges to a positive-definite matrix. Since the random permutations are i.i.d.\ conditional on $\bm{X}$, $[I] \stackrel{d}\longrightarrow \mathcal{N}(\bm{0},\bm{I}_m)$. The proof for $T_U(\mu_0; \pi)$ follows a similar logic.
{\bf Choice of $R_n$ and Statistic.} When choosing $R_n$, we face the following trade-off. A larger value corresponds to using a larger number of observations to construct the test statistic, which translates to higher power through a faster rate of convergence for part $[I]$ of decomposition (ref). On the other hand, a smaller value ensures that the bias term $[II]$ in (ref) is negligible, which is important for size control. As is clear from the previous proof sketch, for the mean-type statistic, the bias and rate of convergence are respectively of order $(R_n/n)^{1/2}$ and $R_n^{-1/2}$. For the U-type statistic, they are instead $\sqrt{R_n}/n$ and $R_n^{-1/4}$, as discussed in (ref). For both statistics, we propose choosing $R_n$ to minimize the sum of these terms, yielding
for the mean- and U-type statistics, respectively. The choice of minimizing the sum of the two is only a heuristic but at least reflects an asymptotic trade-off between control of type I and type II errors. We only seek to provide a practical recommendation that accounts in some way for the trade-off and leave more sophisticated and data-dependent choices of $R_n$ to future work. We also note that simulation results in (ref) show that the tests perform similarly for a variety of choices of $R_n$ around (ref).
The rates of convergence of the mean- and U-type statistics when choosing $R_n$ according to (ref) are respectively $n^{-1/4}$ and $n^{-1/3}$. In general, the U-type statistic has better power properties, as shown theoretically in song_ordering-free_2016 in the context of locally dependent data. Our simulations confirm this for (ref) across a wider range of dependence structures, which leads us to recommend use of the U-type over the mean-type statistic for smaller samples. The main appeal of the mean-type statistic is the ease of constructing CIs, as in (ref).
{\bf Asymptotically Linear Estimators.} Suppose we observe identically distributed data $\bm{Z} = \{Z_i\}_{i=1}^n$ and are interested in a parameter $\beta_0 \in \mathbb{R}^d$. Suppose there exists a parameter $\theta_0$ and a function $\psi$ satisfying ${\bf E}[\psi(Z_1; \beta_0, \theta_0)] = 0$, and let $\hat\theta$ be an estimate of $\theta_0$. Consider an estimator $\hat\beta$ that is asymptotically linear in the sense that
For example, in the case of maximum likelihood, $\hat\theta$ is the sample Hessian, and $\psi$ is the score function times the Hessian. We can then apply our procedures to conduct inference on $\beta_0$ by defining
and $\mu_0 = {\bf E}[\psi(Z_1; \beta_0, \theta_0)] = 0$. Note that in this example, $\mu_0$ is not the population mean of the “data” $\bm{X}$. As discussed in (ref) below, under regularity conditions, our procedures are asymptotically valid if $(\hat\beta,\hat\theta)$ is $\sqrt{n}$-consistent for $(\beta_0,\theta_0)$.
We next consider testing the null $\mu_0 \leq 0$ for $\mu_0 = {\bf E}[X_1]$, where “$\leq$” denotes component-wise inequality. This is relevant, for example, for inference in strategic models of network formation sheng2014 and models of social interactions li2016partial. Let $T_{U,k}(\mu_k; \pi)$ be the U-type statistic applied to scalar data $\{X_{i,k}\}_{i=1}^n$, where $X_{i,k}$ is the $k$th component of $X_i$ and $\mu_k \in \mathbb{R}$. Also let $\bar{X}_k$ be the $k$th component of $\bar{X}$ and $\hat{\bm{\Sigma}}_{kk}$ the $k$th diagonal of $\hat{\bm{\Sigma}}$. We propose the test statistic
While complicated in appearance, this is easy to compute in practice. Define
and let $c_{1-\alpha}$ be the $(1-\alpha)$-quantile of the conditional-on-$\bm{X}$ distribution of $\tilde Q_n(\pi)$. Our proposed test is to reject if and only if $\phi_n = 1$ for
In practice, we can approximate $c_{1-\alpha}$ arbitrarily well by resampling $\pi$ $L$ times, computing $\tilde Q_n(\pi)$ for each draw, and then taking the appropriate sample quantile of this set of statistics. Formally, let $\tilde\pi_\ell = (\tilde\pi_{r\ell})_{r=1}^{R_n}$, where $\{\tilde\pi_{r\ell}\colon \ell \,{=}\, 1,\dots,L;\, r \,{=}\, 1,\dots,R_n\}$ are i.i.d.\ uniform draws from $\Pi$. The feasible critical value is
In (ref), we show that the proposed test uniformly exactly controls size. The intuition behind the test is as follows. Some algebra shows that
which is similar to $\tilde Q_n(\pi)$, except for the presence of the $\hat\lambda_k\mathbf{1}\{\bar{X}_k \geq 0\}$ term. The indicator serves to detect the sign of $\mu_{0k}$, the $k$th component of $\mu_0$, and thus give the test power. To see this, first note that, in the appendix, we show in (ref) and (ref) in the proof of (ref) that, for any $k$, $\hat\lambda_k \approx (R_n^{1/4}\mu_{0k})^2m^{-1/2}\hat{\bm{\Sigma}}^{-1}_{kk} + \mu_{0k}\cdot O_p(1)$, and $R_n^{1/4}\bar{X}_k \approx R_n^{1/4}\mu_{0k}$. We apply these two “facts” to the case $m=1$ for illustration.
First, under a fixed null, $\hat\lambda_k\mathbf{1}\{\bar{X}_k \geq 0\} = \hat\lambda_k\mathbf{1}\{R_n^{1/4}\bar{X}_k \geq 0\} \approx 0$ since either $R_n^{1/4}\mu_{0k} = 0$, in which case $\hat\lambda_k \approx 0$ by the first fact above, or $R_n^{1/4}\mu_{0k} < 0$, in which case the indicator is eventually zero by the second fact above. Consequently, $Q_n(\pi)$ and $\tilde Q_n(\pi)$ have the same asymptotic distributions, and the test controls size. Second, under the alternative, $R_n^{1/4}\bar{X}_k$ is instead eventually positive, so $\hat\lambda_k\mathbf{1}\{R_n^{1/4}\bar{X}_k \geq 0\} \approx \hat\lambda_k$, which is positive with high probability. Indeed, for a fixed alternative, $\hat\lambda_k$ diverges by the first fact above and therefore so does $Q_n(\pi)$. On the other hand, $\tilde Q_n(\pi)$ has a non-degenerate limit distribution, so the test is consistent.
Let $\bm{D}$ be an $n\times k$ matrix of covariates with $i$th row $D_i$, and
where $\{(D_i, \varepsilon_i)\}_{i=1}^n$ is identically distributed but possibly dependent. Our goal is inference on the $j$th component of $\beta_0$, denoted by $\beta_{0j}$. Common sources of dependence are clustering bertrand2004much, spatial autocorrelation barrios2012clustering,bester2009inference, and network autocorrelation acemoglu2015state. Often the precise form of dependence may be unknown, or there may be insufficient data to compute conventional standard errors, for example if we do not fully observe the clusters, the spatial locations of the observations, or the network. If the OLS estimator is $\sqrt{n}$-consistent, however, inference on $\beta_{0j}$ is possible using resampled statistics.
In the context of network applications, a common exercise is to regress an outcome on some measure of the network centrality of a node chandrasekhar2016econometrics. Such measures are inherently correlated across nodes even when links are i.i.d. Moreover, different models of network formation can lead to different expressions for the asymptotic variance. Because this model is an unknown nuisance parameter, it is useful to have an inference procedure that is valid for a large class of models.
To apply our procedure, we write the estimator in the form of a sample mean. Let $W_{ji}$ be the $ji$th component of the matrix $n (\bm{D}'\bm{D})^{-1} \bm{D}'$. Then
is the OLS estimator for $\beta_{0j}$. This fits into the setup of (ref) if we define $Z_i = (Y_i, D_i)$, $\hat\theta = n^{-1}\bm{D}'\bm{D}$, and $\psi(Z_i; \beta_0, \hat\theta) = W_{ji} Y_i - \beta_{0j}$. Thus, when computing our test statistics, we resample elements of $\bm{X}$ for $X_i = W_{ji}Y_i - \beta_{0j}$.
The main assumption required for the asymptotic validity of our procedure is $\sqrt{n}$-consistency of the estimators. For cluster dependence, this holds under conventional many-cluster asymptotics, where the number of observations in each cluster is small but the number of clusters is large. However, unlike with clustered standard errors, we need not know the right level of clustering or even observe cluster memberships. We can also allow the number of clusters to be small, possibly equal to one, so long as the data is weakly dependent within clusters in the sense of yielding $\sqrt{n}$-consistency. Within-cluster weak dependence is also required by canay2017randomization, ibragimov2010t, and ibragimov2016inference, who propose novel inference procedures for cluster dependence with a small number of clusters. Resampled statistics are advantageous because we can allow for only a single cluster and do not require knowledge of cluster memberships.
For spatial dependence, $\sqrt{n}$-consistency can be established using CLTs for mixing or near-epoch dependent data jenish2012spatial. CLTs for network statistics are mentioned in (ref).
Suppose we observe data from a randomized experiment on a single network, where for each node $i$, we observe an outcome $Y_i$, a binary treatment assignment $D_i$, the number of nodes connected to $i$ (“network neighbors”) $\gamma_i$, and the number of treated network neighbors $T_i$. Consider the following outcome model studied in leung2017treatment:
This departs from the conventional potential outcomes model by allowing $r(\cdot)$ to depend on $T_i$, which violates the stable unit treatment value assumption. The object of interest is the following measure of treatment/spillover effects:
where $t,t'\leq \gamma \in \mathbb{N}$. Variation across $d,d'$ identifies the direct causal effect of the treatment, while variation across $t,t'$ identifies a spillover effect, conditional on the number of neighbors. leung2017treatment provides conditions on the network and dependence structure of $\{\varepsilon_i\}_{i=1}^n$ under which the sample analog of (ref) is $\sqrt{n}$-consistent. For example, we can allow $\varepsilon_i$ and $\varepsilon_j$ to be correlated if $i$ and $j$ are connected. Therefore, the setup falls within the scope of our assumptions.
Suppose the econometrician obtains data $\{W_i\}_{i=1}^n$ for $W_i = (Y_i,D_i,T_i,\gamma_i)$ by snowball-sampling 1-neighborhoods. That is, she first obtains a random sample of units, from which she gathers $(Y_i,D_i)$, and then she obtains the network neighbors of those units and their treatment assignment, from which she gathers $(T_i,\gamma_i)$. This is a very common method of network sampling. However, standard error formulas provided by leung2017treatment may require knowledge of the path distances between observed units, which are usually only partially observed under this form of sampling.
Our proposed procedures can be used in this setting. Let $\mathbf{1}_i(d,t,\gamma) = \mathbf{1}\{D_i=d, T_i=t, \gamma_i=\gamma\}$. The analog estimator for (ref) is
This fits into the setup of (ref) by defining $Z_i = W_i$,
where $\beta_0$ is the hypothesized value of the true average treatment/spillover effect (ref).
Inference methods for network statistics are important for network regressions, as discussed in (ref), and strategic models of network formation sheng2014. They are also of inherent interest in the networks literature, where theoretical work is often motivated by “stylized facts” about the structure of real-world social networks barabasi2015,jackson2010. These facts are obtained by computing various summary statistics from network data. However, little attempt has been made to account for the sampling variation of these point estimates, perhaps due to the wide variety of network formation models, which induce different dependence structures. This motivates the use of resampled statistics, which can be used to conduct inference on network statistics without taking a stance on the network formation model.
We next consider two stylized facts that have arguably received the most attention in the literature: clustering and power law degree distributions. This subsection focuses on the former, while the latter is discussed in a more general context in (ref). For a set of $n$ nodes, let $\bm{A}$ be a binary symmetric adjacency matrix that represents a network. Its $ij$th entry $A_{ij}$ is thus an indicator for whether $i$ and $j$ are linked. Define the {\em individual clustering} for a node $i$ in network $\bm{A}$ as
with $Cl_i(\bm{A}) \equiv 0$ if $i$ has at most one link. The denominator counts the number of pairs linked to $i$, while the numerator counts the number of {\em linked} pairs linked to $i$. The {\em average clustering coefficient} of $\bm{A}$ is defined as $n^{-1}\sum_{i=1}^n Cl_i(\bm{A})$.
This statistic is a common measure of {\em transitivity} or {\em clustering}, the tendency for individuals with partners in common to associate. A well-known stylized fact in the network literature is that most social networks exhibit nontrivial clustering, where “nontrivial” is defined relative to the null model in which links are i.i.d.\ jackson2010. Under the null model, when $n$ is large, the average clustering coefficient is close to the probability of forming a link. Yet, the average clustering coefficient is typically larger than the empirical linking probability in practice, hence the stylized fact barabasi2015.
In order to formally assess whether average clustering is significantly different from the probability of link formation, we can use our tests (ref) with
Then $\bar{X}$ is the difference between the average clustering coefficient and the empirical linking probability. To verify $\sqrt{n}$-consistency, we can apply CLTs derived by, for example, bickel_method_2011 and leung2017normal.
Testing whether data follows a power law distribution is of wide empirical interest in economics, finance, network science, neuroscience, biology, and physics barabasi2015,gabaix2009power,klaus2011statistical,newman2005power. By “power law” we mean that the probability density or mass function of the data $f(x)$ is proportional to $x^{-\alpha}$ for some positive exponent $\alpha$. Many methods are available for estimating $\alpha$, for example maximum likelihood or regression estimators ibragimov2015heavy. Given an estimate of the power law exponent, it is of interest to test how well the data accords with or deviates from a power law. Standard methods assume that the underlying data is i.i.d.\ clauset2009power, but this is unrealistic for spatial, financial, and network data, motivating the use of resampled statistics.
In the networks literature, a well-known stylized fact is that real-world social networks have power law degree distributions barabasi1999emergence, where a node's degree is its number of connections. However, empirical evidence is commonly obtained by eyeballing log-log plots of degree distributions, rather than by using formal tests holme2018power. broido2019scale implement formal tests for power laws on a wide variety of network datasets, but their methods assume independent observations, despite the fact that network degrees are typically correlated.
The null hypothesis we next consider testing is motivated by klaus2011statistical, which is that the power law fits no better than some reference null distribution, for example exponential or log-normal. This is operationalized using a Vuong test of the null that the expected log-likelihood ratio is zero. The numerator of the likelihood ratio is the power law distribution with an estimated exponent, and the denominator is the estimated null distribution. Under general misspecification, the log-likelihood ratio is zero if both models poorly fit the data, and less (greater) than zero if the null distribution fits better (worse) pesaran1987global,vuong1989likelihood.
For i.i.d.\ data and non-nested hypotheses, we can test the null by comparing the absolute value of the normalized log-likelihood ratio with a normal critical value. If the former is smaller, the models are equally good. Otherwise, we reject in favor of the power law (null distribution) if the log-likelihood ratio is greater (less) than zero.
We modify this procedure to account for dependence using the U-type statistic as follows. For identically distributed data $\{Z_i\}_{i=1}^n$, let $\ell_{PL}(Z_i,\alpha)$ be the likelihood of observation $i$ under a power law and $\ell_0(Z_i,\gamma)$ the likelihood under the null distribution, which is parameterized by $\gamma$. Then the null hypothesis is
This fits into our setup (ref) by defining $\hat\theta = (\hat\alpha, \hat\gamma)$ (the estimates of $(\alpha,\gamma)$) and
We compute the U-type statistic using these $X_i$'s and compare it to a normal critical value. If the U-type statistic is smaller, then the models are equally good. Otherwise, we reject in favor of the power law (null distribution) if $\bar{X}$, the estimated log-likelihood ratio computed on the full dataset, is greater (less) than zero, as in vuong1989likelihood.
jackson2007meeting propose a model of network formation that generates a degree distribution parameterized by $r$, which interpolates between the exponential and power law distributions. Their model provides microfoundations for the different distributions. When $r \rightarrow \infty$, the network is formed primarily through random meetings, and the distribution is exponential. When $r \rightarrow 0$, the network is formed primarily through “network-based meetings,” as nodes are more likely to meet neighbors of nodes that were previously met. Since high-degree nodes are more likely to be met through network-based meetings, this corresponds to a “rich-get-richer” or “preferential-attachment” mechanism that generates a power law degree distribution.
The authors estimate $r$ using data on six distinct social networks and informally assess the extent to which the estimated distributions depart from a power law. See their paper for descriptions of the data. In this section, we use the same datasets to implement the test described in (ref), using the exponential distribution as the null. We set the lower support points of the exponential and power law distributions at one and estimate the parameters of the distribution using (pseudo) maximum likelihood.
(ref) displays the results of the tests. Row “Exp.” displays the estimated power law exponent, “LL” the normalized log-likelihood ratio, “$r$” the estimated value of $r$ from jackson2007meeting, and $n$ the sample size. “Naive” displays the conclusion of the conventional Vuong test that assumes the data is i.i.d., with the conclusion “P” denoting power law, “E” denoting exponential, and “N” meaning the null is not rejected. Finally, the bottom five rows display the conclusions of our test for different values of $R_n$ to assess the robustness of the conclusions. Recalling the definition of $R_n^U$ from (ref), the rows correspond to $R_n = \min\{R_n^U \cdot \epsilon, 100\text{k}\}$ for $\epsilon \in \{0.6,0.8,1,1.2,1.4\}$. Thus, the middle of the bottom rows is our suggested choice $R_n^U$. We truncate $R_n$ at 100k since this is large enough to draw a robust conclusion.
The results of all three methods are in agreement for the coauthor, prison, romance, and WWW networks. This is due to the large values of the normalized log-likelihood ratios, which make the conclusion rather obvious regardless of the test used. For the other datasets, however, the methods draw different conclusions.
For the ham radio network, jackson2007meeting estimate $r$ to be 5. As they note, this means network-based meetings are about eight times less common compared to the WWW network, so their degree distributions should be closer to the exponential than the power law distribution. However, our test finds insufficient evidence to reject the null for the ham radio network due to the very small sample size. The Vuong test rejects in favor of the exponential distribution, despite the normalized log-likelihood ratio being smaller (-1.75). This may be because the test assumes i.i.d.\ data, so the sample variance of the log-likelihoods may be underestimated.
jackson2007meeting estimate $r$ to be close to zero for the citation network, which favors the power law. In contrast, the Vuong test rejects in favor of the exponential distribution, while our test either concludes exponential or fails to reject, depending on the value of $R_n$. This is due to the negative normalized log-likelihood -3.84, which is large enough for the i.i.d.\ test to draw a clear conclusion but perhaps not quite large after adjusting for dependence, which explains the ambiguity in the conclusion of our test. Still, neither test concludes the data is consistent with a power law, unlike the estimate of jackson2007meeting.
This section considers a generalization of the setup in (ref) in which $\bm{X}$ is a triangular array, so $X_i$ and $\mu_0$ may implicitly depend on $n$. This is important to accommodate network applications since, for example, when the network is sparse, the linking probability decays to zero with $n$. All proofs can be found in the appendix.
For any vector $v$, let $\lVertv\rVert$ denote its sup norm. The next theorem shows that the U-type (mean-type) statistic is asymptotically normal (chi-square).
We next state formal results for test (ref). Let $\lambda_\text{min}(M)$ denote the smallest eigenvalue of a matrix $M$ and $\lVertM\rVert = \max_{i,j} \lvertM_{ij}\rvert$. Define $\mu_0(\mathbf{P}) = {\bf E}_\mathbf{P}[X_1]$, where ${\bf E}_\mathbf{P}[\cdot]$ denotes the expectation under the data-generating process (DGP) $\mathbf{P}$ (formally a probability measure). Let $\mu_{0k}(\mathbf{P})$ be the $k$th component of $\mu_0(\mathbf{P})$.
This shows that the test uniformly controls size. The theorem follows quite directly from the next lemma, which also provides results on the power of the test.
Part (a) shows the test asymptotically controls size under any sequence of null DGPs. Parts (b) and (c) describe the test's power, with (c) showing that the test has power against local alternatives that vanish no faster than rate $R_n^{-1/4}$. Since size control requires $\sqrt{R_n}/n \rightarrow 0$, the test does not have power against $\sqrt{n}$ local alternatives.
This section presents results from four simulation studies, each corresponding to one of the applications in (ref). We use asymptotic critical values to implement the tests and multiple values of $R_n$ to assess the robustness of the methods to the tuning parameter. The results are broadly summarized as follows. The size is largely close to the target level of 5 percent across all designs with weak dependence, more so for larger sample sizes. Power can be low in small samples, as expected from the convergence rates discussed in (ref). The U-type statistic has significantly better power properties than the mean-type, which leads us to recommend the former. Finally, results are similar across values of $R_n$.
{\bf Cluster Dependence.} Let $c$ index cities, $f$ index families, and $i$ index individuals. We generate outcomes according to the random effects model
where $\alpha_f \stackrel{iid}\sim \mathcal{N}(0,1)$ and $\varepsilon_{ifc} \stackrel{iid}\sim \mathcal{N}(0,1)$, the two mutually independent. The correct level of clustering is at the family level, and the true value of $\theta_0$ is one. Let $n_c$, $n_f$, and $n_i$ be the number of cities, families, and individuals, respectively, and $N = (n_c, n_f, n_i)$. Families have equal numbers of individuals and cities equal numbers of families.
We present results for resampled statistics and compare them to $t$-tests using clustered standard errors for each level of clustering. (ref) displays simulation results for the size and power of our tests, computed using 6000 simulations. The first two rows display rejection percentages for size (testing $H_0\colon \theta_0=1$) and power (testing $H_0\colon \theta_0=1.5$), respectively. To show the robustness of our test, we display results for five different values of $R_n$. For the U-type statistic, the columns correspond to $R_n = R_n^U \cdot \epsilon$ for $\epsilon \in \{0.6,0.8,1,1.2,1.4\}$ and $R_n^U$ defined in (ref). Thus, the middle of the five columns for both sample sizes corresponds to our suggested choice $R_n^U$. We do the same for the M-type statistic, except we use $R_n = R_n^M\cdot\epsilon$. Finally, (ref) displays analogous results for $t$-tests with clustered standard errors. The columns display the level of clustering with $c$ for city, $f$ for family, $i$ for individual.
The results show that the $t$-test overrejects when clustering at too coarse a level and the number of clusters is small (clustering at the city level). It also overrejects when clustering at too fine a level (clustering at the individual level) because this assumes more independence in the data than is warranted. In contrast, tests using resampled statistics properly control size. On the other hand, the $t$-test is clearly more powerful. U-type statistics show a significant power advantage over mean-type statistics. Our results are also similar across values of $R_n$.
These designs feature weak cluster dependence (required by all existing methods in the literature) since clustering is at the family level and the number of families is relatively large. To see how our methods break down under strong dependence, we next modify the design so that clustering is at the city level and the number of cities is small. For $N=(30,600,1200)$ (only 30 cities), the type I error of the mean-type statistic now ranges from 9.13 to 17.83 percent across the values of $R_n$, while that of the U-type statistic ranges from 25.70 to 33.58 percent. For $N=(10,600,1200)$, the error of the mean-type statistic ranges from 17.83 to 29.45 percent, and that of the U-type statistic ranges from 48.93 to 56.55 percent.
{\bf Network Statistics.} We generate a network according to a strategic model of network formation, following the simulation design of leung2017. There are $n$ nodes, and each node $i$ is endowed with a type $(X_i,Z_i)$, where $Z_i \stackrel{iid}\sim \text{Ber}(0.5)$ and $X_i \stackrel{iid}\sim U([0,1]^2)$, the two mutually independent. Let $\rho$ be the function such that $\rho(\delta)=0$ if $\delta \leq 1$ and equal to $\infty$ otherwise. Potential links satisfy
where $\lVert\cdot\rVert$ is the Euclidean norm on $\mathbb{R}^d$ and $\zeta_{ij} \stackrel{iid}\sim \mathcal{N}(0,\theta_4^2)$ is independent of types. We set $\theta = (-1, 0.25, 0.25, 1)$ and $r_n = (3.6/n)^{1/2}$ and use the selection mechanism in the design of leung2017; see his paper for details.
We are interested in two statistics that are functions of the network, the average clustering coefficient (defined in (ref)) and the average degree. leung2017normal prove $\sqrt{n}$-consistency of the sample statistics for their population analogs. Let $\theta_0$ be the expected value of the network statistic. Tables (ref) and (ref) in the appendix display rejection percentages for average clustering and degree for two nulls. The first is that $\theta_0$ equals its true value, which estimates the size. The second is that $\theta_0$ equals the true value plus the number indicated in the table, which estimates power. We use 6000 simulation draws each to simulate $\theta_0$ and rejection percentages.
For this design, our tests exhibit substantial size distortion in small samples ($n=100$), but the size tends toward the nominal level as $n$ grows. For the average degree, there is still some size distortion at larger samples ($n=1000$) for the U-type statistic, although this is less than 2 percentage points above the nominal level. The U-type statistic is significantly more powerful than the mean-type statistic, with rejection percentages sometimes more than twice as large.
{\bf Treatment Spillovers.} Consider the setup in (ref). We assign units to treatment with probability 0.3 and draw the network from the same model used for the network statistics above. The outcome model is $Y_i = \beta_1 + \beta_2 D_i + \beta_3 T_i + \beta_4 \gamma_i + \varepsilon_i$. For $\nu_j \stackrel{iid}\sim \mathcal{N}(0,1)$ independent of treatments, we set $\varepsilon_i = \nu_i + \sum_j A_{ij}\nu_j/\sum_j A_{ij}$, which represents exogenous peer effects in unobservables and generates network autocorrelation in the errors. We set $(\beta_1,\beta_2,\beta_3,\beta_4) = (1, 0.5, -1, 0.5)$.
We consider a linear regression estimator of $Y_i$ on $(1, D_i, T_i, \gamma_i)$ and test two hypotheses, $\beta_3=-1$ to estimate the size and $\beta_3=-1.8$ to estimate the power. (ref) in the appendix displays rejection percentages computed using 6000 simulations. As in the network statistics application, we display multiple values of $R_n$ corresponding to $R_n = R_n^U\cdot \epsilon$ for the U-type statistic and $R_n = R_n^M\cdot\epsilon$ for the mean-type statistics, for the same values of $\epsilon$ above. In this design, our tests control size well across all sample sizes, and the U-type statistic is substantially more powerful.
{\bf Power Laws.} We implement the test in (ref), which uses the U-type statistic. We set $R_n = R_n^U$. Following the notation in that section, we draw data $\{Z_i\}_{i=1}^n$ i.i.d.\ from either an $\text{Exp}(0.5)$ distribution or a power law distribution with exponent 2. The lower support point for both is set at 1. (ref) in the appendix reports rejection percentages from 6000 simulations under both alternatives (exponential and power law). Row “LL” displays the average normalized log-likelihood ratio, “Favor Exp” the percentage of simulations in which we reject in favor of the null distribution, and “Favor PL” the percentage of simulations in which we reject in favor of the power law. The power is around 55--60 percent for $n=100$ and 86--97 percent for $n=500$.
We develop tests for moment equalities and inequalities that are robust to general forms of weak dependence. The tests compare a resampled test statistic to an asymptotic critical value, in contrast to conventional resampling procedures, which compare a test statistic constructed using the original dataset to a resampled critical value. The validity of conventional procedures requires resampling in a way that mimics the dependence structure of the data, which in turn requires knowledge about the type of dependence. In contrast, the procedures we study are implemented the same way regardless of the dependence structure. We show our methods are asymptotically valid under the weak requirement that the target parameter can be estimated at a $\sqrt{n}$ rate. To illustrate the broad applicability of our procedure, we discuss four applications, including regression with unknown dependence, treatment effects with network interference, and testing network stylized facts.
\FloatBarrier \phantomsection \addcontentsline{toc}{section}{References}