EconBase
← Back to paper

Inference in Models of Discrete Choice with Social Interactions Using Network Data

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

122,337 characters · 26 sections · 124 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Inference in Models of Discrete Choice with Social Interactions Using Network Data

abstract{\sc Abstract.} This paper studies inference in models of discrete choice with social interactions when the data consists of a single large network. We provide theoretical justification for the use of spatial and network HAC variance estimators in applied work, the latter constructed by using network path distance in place of spatial distance. Toward this end, we prove new central limit theorems for network moments in a large class of social interactions models. The results are applicable to discrete games on networks and dynamic models where social interactions enter through lagged dependent variables. We illustrate our results in an empirical application and simulation study. {\sc JEL Codes}: C22, C31, C57 {\sc Keywords}: social networks, peer effects, empirical games, HAC estimator

\addcontentsline{toc}{part}{Main Paper}

Introduction

\onehalfspacing

Threshold models of social influence are the subject of a large theoretical literature in the social sciences ellison1993learning,granovetter1978threshold,jackson2010,morris2000contagion. These are models of social interactions with binary decisions, which have been used to study, for example, product adoption godinho2014peer, risky behavior bauman1996importance, health choices christakis2004social, protests gonzalez2017collective, and voting bond201261. The empirical content of these and related models has long been of interest in the econometric literature brock2001discrete,manski1993identification. Much of this work has focused on static models with cross-sectional data, where social interactions only operate within groups and the number of groups is large. This paper instead studies inference in static and dynamic models when the data consists of a single large network. The motivation for using network data is that it can be used to more accurately model social interactions, which is often heterogeneous by nature. For example, students in different classrooms may interact, and even within classrooms, they may only interact with a subset of their classmates.

The goal of this paper is to provide theoretically justified inference procedures for models of discrete choice with social interactions. We consider spatial and network heteroskedasticity and autocorrelation consistent (HAC) variance estimators, which have seen increasing use in applied work for inference in the single network setting. Our main theoretical results are new central limit theorems for static and dynamic models of social interactions and conditions under which the HAC estimators are consistent, or possibly conservative, for the variance. These results hold under an asymptotic approximation that sends the size of a single large network to infinity.

The motivation for large-network asymptotics is that network data commonly consists of observations on a single network, as opposed to many independent groups. This is theoretically challenging because social interactions inherently induce autocorrelation between different network subunits. Inference thus requires a large-sample theory under which the amount of “independent information” in a network grows with the network size. This is analogous to limit theory in time series that sends the number of time periods to infinity rather than the number of time series.

There are few inference procedures presently available in the single-network setting. In the medical literature, network autocorrelation is often ignored in practice, and i.i.d.\ standard errors are commonly used lee2019network. In economics, the strategy of clustering on subnetworks is quite common. By this we mean dividing a single network into many subnetworks using, say, geographic boundaries or a community detection algorithm blondel2008fast, and then clustering standard errors on the subnetworks. However, there is often a sizeable share of links that bridge clusters, which implies agents can interact across subnetworks. This renders implausible the assumption necessary for the validity of clustered standard errors, that subnetworks are independent. HAC estimators provide a viable alternative, as they do not rely on partitioning or independence of clusters. Instead, they account for correlation between each agent and the alters within a neighborhood of the agent, which can be thought of as defining agent-specific, overlapping clusters.

To provide a clearer sense of the practical import of our results, consider the study of conley2010learning on social learning among pineapple farmers in Ghana. Their main specification is a logistic model of the probability that a farmer's agricultural inputs change in response to above- or below-expectation returns from his network neighbors' crop yields. For the purposes of constructing clustered standard errors, it is unclear how to reasonably divide the network using geographic boundaries, as there are only three villages in their data. The authors instead use a spatial HAC estimator, which is valid when the data is spatially autocorrelated. We provide conditions for validity of the estimator when the data is potentially {\em network} autocorrelated, for example when the errors of network neighbors are dependent.

The spatial HAC estimator requires data on agents' spatial positions, but this is not always available. For instance, rather than physical space, the underlying space may correspond to a latent social space, whereby socially similar agents are more likely to form connections. Given network data, we can instead define an alternate notion of “distance,” namely {\em path distance}, which is the shortest number of links it takes to travel from one agent to another in the network, and use this in place of spatial distance to construct a network HAC estimator. This idea has been used in practice acemoglu2015state,eckles2016estimating. We provide conditions under which the estimator is valid.

Our results apply to static and dynamic models of social interactions. In dynamic models, an agent's action is a function of her network neighbors' lagged actions. Such models are useful when a short time series on a large network is available. We discuss how our results can be used for parametric and nonparametric inference in dynamic models. Static models instead allow actions to depend on neighbors' contemporaneous actions. They formally correspond to discrete games of complete information, which are useful when the data consists of a single network snapshot. In the existing literature, xl2015 propose a simulated method of moments estimator and li2016partial characterize the identified set using subnetwork moments. We discuss how our results can be used to conduct inference using their procedures.

We illustrate our results in a simulation study and empirical application. The simulation evidence shows that the normal approximation works well, and inference using the HAC estimators properly controls size, whereas naive i.i.d.\ standard errors can be highly anti-conservative. In the empirical application, we reanalyze data from conley2010learning and find that both spatial and network HAC estimators deliver similar standard errors.

Finally, we would like to highlight several technical contributions.

enumerate• To our knowledge, all existing results on HAC estimators require first-order stationarity conditions that essentially require moments to be centered at their conditional expectations. This holds in GMM-type settings. However, it is not satisfied in moment inequality models and often violated in nonparametric settings, for example nonparametric inference on the conditional choice probability. For the spatial HAC, we provide two positive results. The first is a general (albeit strong) sufficient condition under which first-order stationarity holds asymptotically for any network moments, whether centered or not. In this case, the spatial HAC is asymptotically conservative, which appears to be a new result. Second, our results justify use of the generalized spatial HAC proposed by leung2019normal, which is consistent regardless of whether stationarity holds. • To prove our CLTs, we first establish a general CLT under high-level “stabilization” conditions, which is a modification of a result due to leung2019normal. The difference is we consider an increasing domain setting to obtain consistency of the spatial HAC, whereas they study a quasi-infill setting.\footnote{Specifically, under our asymptotics, positions are scaled to be sufficiently spread out in space, which is the usual increasing domain asymptotics used to analyze spatial HAC estimators. Under the asymptotics of leung2019normal, positions are instead sampled from a bounded region, but preferences are scaled so that agents increasingly prefer nearby alters as the network size grows. } To derive primitive conditions for stabilization in our setting, we follow the methodology in leung2019normal, drawing on results in branching process theory. While the conditions we derive are qualitatively similar to those used in leung2019normal to establish CLTs for strategic models of network formation, our results are {\em not} special cases of theirs. They consider applications to network formation and network regressions, but their results do not apply to social interactions models. • Our results on the validity of the network HAC require an exogenous network in the static model. We provide intuition suggested by our proof strategy on why endogeneity is difficult to allow in general.

{\bf Related Literature.} To our knowledge, the only prior work on network HAC estimators is a recent paper by kojevnikov2019limit. kojevnikov2019bootstrap proposes novel bootstrap procedures for network data. Both papers assume the data satisfies a new conditional $\psi$-weak dependence notion modified from the time series literature but do not discuss applications to models of social interactions. We utilize a different notion of weak dependence proposed by leung2019normal and apply branching process results to verify weak dependence in our applications. See (ref) for further comparison of our approach and theirs.

There is a large literature on spatial autoregressive models, of which the widely used linear-in-means model of social interactions is a special case. For relevant results on spatial HAC estimators, see for example conley1999gmm, conley2007spatial, jenish2016, and kelejian2007hac. In the context of network formation, boucher2017my and leung2019normal show that spatial HAC estimators can be used for valid inference.

A growing literature in econometrics studies models with strategic interactions and many agents. menzel2015inference and shang2011two, among others, consider inference on discrete games of complete information. Unlike our setting, there is no network structure; instead, social interactions enter payoffs through a vector of aggregate statistics, such as the average action of all agents. li2016partial and xl2015 study settings with networked interactions and complete information. eraslan2017identification and xu2018social derive results for corresponding games of incomplete information. Large-network asymptotics are different in this setting because actions are i.i.d.\ conditional on observed characteristics and the network. kuersteiner2018dynamic develop large-sample theory for a dynamic version of the linear-in-means model where social interactions enter through lagged dependent variables. he2018measuring propose inference methods for nonparametric measures of diffusion in dynamic models with binary outcomes.

{\bf Outline.} The next two sections respectively introduce dynamic and static models of social interactions with binary outcomes. Then (ref) presents the HAC estimators and conditions for their asymptotic validity. We state formal conditions for central limit theorems in (ref) for the static model and (ref) for the dynamic model. Next, (ref) discusses results from a simulation study and empirical application using data from conley2010learning. Finally, (ref) concludes.

{\bf Notation.} If $f$ is a density function or random vector, let $\text{supp}(f)$ be its support. Given $n$ i.i.d.\ vectors $X_1, \dots, X_n$ and $H \subseteq \{1, \dots, n\}$, let $X_H$ be the submatrix $(X_i\colon i \in H)$ and $X_{-i} = (X_1, \dots, X_{i-1}, X_{i+1}, \dots, X_n)$. We use boldface letters to denote the entire collection $\bm{X} = (X_i)_{i=1}^n$ (as opposed to $X$ to avoid confusing this with a generic draw). For a symmetric $n\times n$ matrix $\bm{A}$, we let $A_{ij}$ denote the $ij$th entry, $A_i$ denote the $i$th column, and $A_{-i}$ denote $\bm{A}$ with the $i$th row and column deleted. Also for $H$ defined previously, let $A_H$ be the submatrix $(A_{ij}\colon i,j \in H)$ and $A_{H,i} = (A_{ij}\colon j \in H)$. For all of these submatrices and subvectors, rows and columns are ordered in the same way as the original matrices and vectors.

Throughout this paper we will only be concerned with undirected networks with no self-links. Accordingly, we represent a network $\bm{A}$ on $n$ agents as an $n\times n$ adjacency matrix with zeros on the diagonals. The $ij$th entry $A_{ij}$, which we call the {\em potential link} between $i$ and $j$, takes values in $\{0,1\}$. The {\em degree} of an agent $i$ in $\bm{A}$ is $\sum_j A_{ij}$. A {\em path} in $\bm{A}$ from agent $i$ to $j$ is a sequence of distinct agents starting with $i$ and ending with $j$ such that for each consecutive $k,k'$ in this sequence, $A_{kk'}=1$. The number of agents in this sequence minus one is the {\em length} of this path. The {\em path distance} between two agents is the length of the shortest path that connects them, assuming one exists; if one does not, then it is defined as $\infty$. The {\em $K$-neighborhood} of an agent $i$ in $\bm{A}$, denoted $\mathcal{N}_{\bm{A}}(i,K)$, is the set of all agents $j$ for which $\ell_{\bm{A}}(i,j) \leq K$. Note that this includes $i$. Finally the {\em component} of an agent $i$ in $\bm{A}$ is the set of all agents $j$ for which $\ell_{\bm{A}}(i,j) < \infty$.

Dynamic Model of Social Interactions

A large literature in microeconomics and computer science studies dynamic models of social influence ellison1993learning,kempe2003maximizing,montanari2010spread,morris2000contagion. In these models, a subset of agents are initially seeded as adopters (choosing action 1), and then agents myopically best respond in subsequent periods to the number of neighbors who are adopters. For example, if the share of period $t-1$ adopters in $i$'s neighborhood exceeds some threshold $\tau_i$, then agent $i$ might adopt in period $t$. This section introduces an econometric version of the model with observed and unobserved heterogeneity. Models of this sort are widely used in applied work in economics, marketing, and network science.\footnote{E.g.\ ameri2017structural, banerjee2013diffusion, christakis2007spread, iyengar2011opinion, katona2011network, and park2018social.} They are useful when the econometrician observes a short time series of a large network.

Model

The econometrician observes a set of agents $\mathcal{N}_n = \{1, \dots, n\}$ connected through a time-invariant network $\bm{A}$. Agents interact over a small number of time periods $t = 0, 1, \dots, T$ for $0<T<\infty$. Associate each agent $i$ and period $t$ with a binary outcome $Y_i^t$ and type $\tau_i^t$. We decompose $\tau_i^t = (X_i^t, \alpha_i, \varepsilon_i^t)$, where $X_i^t$ is a vector of observed covariates, $\alpha_i$ a fixed effect, and $\varepsilon_i^t$ an idiosyncratic error. We assume $X_i \equiv (X_i^0, \dots, X_i^T, \alpha_i)$ is i.i.d.\ across agents, and independent of $\varepsilon_i^t$ for any $t$, while the errors are i.i.d.\ across agents and time periods.

Outcomes are realized according to myopic best-response dynamics: for all agents $i$ and periods $t = 1, \dots, T$,

equation[equation omitted — 88 chars of source]

where $U(\cdot)$ is a real-valued function representing net utility. The first term $S_i^t$ is a finite-dimensional vector of statistics given by

equation*[equation* omitted — 91 chars of source]

for a vector-valued function $S(\cdot)$ with $Y^{<t} = (Y_j^s\colon j \in \mathcal{N}_n, s < t)$, $\tau_i^{\leq t} = (\tau_i^s\colon s \leq t)$, and $\tau_{-i}^{\leq t}$ the collection of $\tau_i^{\leq t}$'s excluding $i$ (see the definition of the $-i$ subscript in the introduction). Hence, $S_i^t$ can capture peer effects through lagged outcomes and types of other agents in the network. As a function of $\tau^{\leq t}_{-i}$, $S_i^t$ can also depend on the unobserved subvector of neighbors' types, which can generate network autocorrelation in unobservables.

exampleA typical payoff specification is the linear in parameters model \begin{equation*} U(S_i^t, \tau_i^t) = (X_i^t)'\beta_1 + \beta_2 \frac{\sum_j A_{ij} Y_j^{t-1}}{\sum_j A_{ij}} + \alpha_i + \nu_i^t. \end{equation*} Social interactions enter through the average outcome of neighbors in the previous period. The idiosyncratic term $\nu_i^t$ might simply equal $\varepsilon_i^t$, or it might be autocorrelated. For example, the errors may be jointly normal across agents with nonzero covariances for linked pairs. Alternatively, we could have \begin{equation*} \nu_i^t = \frac{\sum_j A_{ij} \varepsilon_j^t}{\sum_j A_{ij}} + \varepsilon_i^t, \end{equation*} where the first term captures exogenous peer effects in the errors and induces contemporaneous network autocorrelation between the $\nu_i^t$'s. In the latter case, $S_i^t$ is a two-dimensional vector consisting of this term and the term multiplying $\beta_2$.
exampleA large literature dating back to at least granovetter1978threshold studies threshold models of behavior, where for each agent $i$, $Y_i^t = 1$ if and only if the share of neighbors choosing action one in the previous period exceeds a threshold $\xi_i^t$ jackson2010,schelling2006micromotives. This is captured in our framework by setting \begin{equation*} U(S_i^t, \tau_i^t) = \frac{\sum_{j\neq i} A_{ij} Y_j^{t-1}}{\sum_{j\neq i} A_{ij}} - \xi_i^t. \end{equation*} If $\xi_i^t$ only depends on own type $\tau_i^t$, then in this model, $S_i^t$ is a scalar corresponding to the average lagged outcome on the right-hand side. In practice, heterogeneity in the threshold is often of interest, since this determines the extent of diffusion, so a typical exercise would be to parametrize it as a function of type and estimate the parameters.

Both of these examples restrict the dependence of $S_i^t$ on the network. We next impose this restriction more generally. Let $\mathcal{N}_{\bm{A}}^-(i,M) = \mathcal{N}_{\bm{A}}(i,M)\backslash\{i\}$, which is $i$'s $M$-neighborhood, excluding $i$ herself. Let $Y_H^{<t} = (Y_j^s\colon j \in H, s < t)$, and recall the notation for submatrices in the introduction.

assump[Local Interactions] There exists $M \in \mathbb{N}$ such that, for all $n\in\mathbb{N}$, $i \in \mathcal{N}_n$, and $t = 1, \dots, T$, \begin{equation*} S_i^t = S(Y_{\mathcal{N}(i)}^{<t}, \tau_i^{\leq t}, \tau_{\mathcal{N}(i)}^{\leq t}, A_{\mathcal{N}(i)}) \quadfor\quad \mathcal{N}(i) \equiv \mathcal{N}_{\bm{A}}^-(i,M). \end{equation*}

This states that $S(\cdot)$ is only a function of its arguments through the $M$-neighbors of $i$, which is clearly satisfied in the previous examples for $M=1$.

Model (ref) governs the evolution of the process from period 1 onward. It remains to specify a model for the initial condition. This model need not be known in practice to use our results, but we will need to impose some (nonparametric) restrictions for weak dependence in order to establish a CLT. If this process is observed shortly after its inception, it is reasonable to draw $Y_i^0$ from a single-agent discrete choice model, e.g.\ $Y_i^0 = \bm{1}\{U(\bm{0}, \tau_i^0) > 0\}$, where we zero out the regressors that are functions of lagged dependent variables because there is no previous time period. Alternatively, the initial condition might be viewed as the long-run outcome of a dynamic process. This can be reasonably approximated by a static model of social interactions, which is discussed in (ref). Since this nests the single-agent discrete choice model, we will assume a general static model for the initial condition, whose formal statement is postponed to (ref).

Network Moments

We next define the class of dynamic network moments, objects for which we seek to prove a CLT and construct variance estimators. Let $Y_i = (Y_i^t)_{t=0}^T$, $\tau_i$ be defined in the analogous way, and $Y_{-i}, \tau_{-i}$ be defined according to the notation in the introduction. We consider moments that are averages of {\em agent statistics} $\psi_i$, namely

equation*[equation* omitted — 136 chars of source]

where $\psi(\cdot)$ is $\mathbb{R}^m$-valued. The main technical contributions of this paper are twofold. In (ref), we provide conditions under which, as $n\rightarrow\infty$,

equation[equation omitted — 269 chars of source]

where $\bm{\Sigma}_n$ is the variance and $\bm{I}_m$ the $m\times m$ identity matrix. In (ref), we show that spatial and network HAC estimators are valid estimators for $\bm{\Sigma}_n$. The formal sequence of models along which we take these limits is defined in (ref).

We consider agent statistics $\psi(\cdot)$ satisfying the following {\em $K$-locality} condition. Recall the notation introduced prior to (ref).

assump[$K$-Locality] There exists $K \in \mathbb{N}$ such that, for any $n$ and $i\in\mathcal{N}_n$, \begin{equation*} \psi_i = \psi(Y_i, Y_{\mathcal{N}(i)}, A_{\mathcal{N}(i),i}, A_{\mathcal{N}(i)}, \tau_i, \tau_{\mathcal{N}(i)}) \quadfor\quad \mathcal{N}(i) \equiv \mathcal{N}_{\bm{A}}^-(i,K). \end{equation*}

This states that, for some positive constant $K$, the agent statistic of $i$ is only a function of its arguments through the agents in $i$'s $K$-neighborhood. We first walk through two basic illustrative examples and then discuss network moments useful for inference on social interactions.

exampleConsider $\psi_i = Y_i^t$ for some specified time period $t$. Then the network moment corresponds to the average outcome or empirical choice probability at time $t$. A related example is the average outcome over all observed time periods, where $\psi_i = T^{-1} \sum_{t=0}^T Y_i^t$. Both of these satisfy (ref) for $K=0$ because they are only functions of the first argument of $\psi(\cdot)$.
example(ref) also encompasses subnetwork moments, such as the number of dyads (linked pairs) such that both agents choose action 1 in period $t$. This corresponds to $\psi_i = \sum_j A_{ij}Y_i^tY_j^t$, which satisfies (ref) for $K=1$, since it is only a function of $i$'s neighbors.\footnote{Note that the network moment is then $n^{-1} \sum_{i,j} A_{ij}Y_i^tY_j^t$. We always scale by $n^{-1}$ because our assumptions will ensure a sparse network, which implies that $\sum_j A_{ij} = O_p(1)$ for any $i$. See (ref). } Another example is the number of intransitive triads (triplets with only two links) such that all agents choose action 0. This corresponds (up to scale) to $\psi_i = \sum_j \sum_k A_{ij}A_{jk}(1-A_{ik}) (1-Y_i^t)(1-Y_j^t)(1-Y_k^t)$, which satisfies (ref) for $K=2$, since $k$ may be a 2-neighbor of $i$.

We next provide examples of network moments useful for parametric and nonparametric inference on social interactions.

Method of Moments

In practice, a common exercise is to parametrize payoffs $U(\cdot)$ and the conditional distributions of $\alpha_i$ and $\varepsilon_i^t$ given $X_i^t$. For example, consider a two-period version of (ref) with $\alpha_i = 0$ for all $i$ and $\nu_i^t \sim \mathcal{N}(0,1)$. We can estimate the model using a probit regression of $Y_i^1$ on $(X_i^1, \sum_j A_{ij} Y_j^0 / \sum_j A_{ij})$. However, the usual probit likelihood may not be a true likelihood because the errors $\nu_i^t$ are potentially network autocorrelated. Nonetheless, we can treat this as a pseudo-likelihood, as in poirier1988probit.

To obtain a normal limit for the pseudo-MLE estimator, we need the average of the scores to obey a CLT, in addition to the usual regularity conditions. Here we take $\psi_i$ to be the score for agent $i$. This satisfies (ref) for $K=1$ because the regressors are a function of $i$'s 1-neighbors. The information matrix equality does not hold, since this is a pseudo-likelihood, but a consistent estimate of the asymptotic variance can obtained using the sandwich formula with the variance of the scores estimated using either our proposed spatial or network HAC.

Since the model is fully specified up to a vector of parameters $\theta$, we can also apply simulated method of moments using, for example, the moments introduced in Examples (ref) and (ref). This consists of computing their empirical analogs from data and matching them to simulated analogs. For inference, our results can be applied to construct GMM standard errors that account for network autocorrelation using the sandwich formula.

Nonparametric Inference

We next consider nonparametric inference on the average structural function (ASF), a parameter which provides a nonparametric measure of social interactions. Partition $(S_i^t, \tau_i^t) = (\tilde X_i^t, \tilde \alpha_i, \tilde\varepsilon_i^t)$, where $\tilde X_i^t$ contains all observed quantities, $\tilde\alpha_i$ all time-invariant unobservables, and $\tilde\varepsilon_i^t$ all time-varying unobservables that are independent across time. Thus we rewrite $U(S_i^t, \tau_i^t) = U(\tilde X_i^t, \tilde \alpha_i, \tilde\varepsilon_i^t)$. Using this representation, the ASF is given by

equation*[equation* omitted — 81 chars of source]

where $x$ is a vector of constants. Then if $\tilde X_i^t$ includes some function of the lagged outcome of neighbors, $\mu(x) - \mu(x')$ provides a nonparametric measure of peer effects.

The ASF is typically not point identified in this context chamberlain2010binary, so we use the partial identification approach of chernozhukov2013average. Let $\tilde X_i = (\tilde X_i^t)_{t=0}^T$, whose support is required to be discrete. Define

align*[align* omitted — 250 chars of source]

Note that $\mathcal{X}^t(x)$ is the set of values of $\tilde X_i$ for which the $t$th component first equals $x$ at time $t$. Then $\left\{ \mathcal{X}^\neq(x), \mathcal{X}^0(x), \dots, \mathcal{X}^T(x) \right\}$ partitions the support of $\tilde X_i$. Let

equation*[equation* omitted — 155 chars of source]

$\mu_\ell(x) = {\bf E}[\hat Y_i(x)]$, and $\mu_u(x) = \mu_\ell(x) + {\bf E}[P_i(x)]$. chernozhukov2013average show that

equation*[equation* omitted — 56 chars of source]

and that the bounds collapse to a point at a rate exponential in $T$.

To use these bounds in practice, we construct sample analog estimators

equation*[equation* omitted — 160 chars of source]

Given a joint CLT for these estimators, we can apply the method of woutersen2006simple asymptotic version of the GMS test andrews2010inference to construct confidence intervals for the ASF using a HAC estimate of the variance.

As a first step to apply our CLT, we need to verify that the analog estimators fall within the set of moments satisfying (ref). The agent statistic is $\psi_i = (\hat Y_i(x), P_i(x))$. By definition, $\hat Y_i(x)$ is a function of the outcome time series through $\{(Y_i^t, \tilde X_i^t)\}_{t=0}^T$. By (ref), $\tilde X_i^t$ (in particular the subvector corresponding to observed components of $S_i^t$) is a function of outcomes only through the 1-neighborhood of agent $i$. Hence, $\hat Y_i(x)$ satisfies (ref) for $K=1$. The same argument holds for $P_i(x)$.

Network Formation

This section introduces a (nonparametric) stochastic model of network formation. The model need not be known in practice, but it is used to prove our asymptotic results. Readers interested in inference methods can skip ahead to (ref).

We assume each agent $i$ is endowed with a {\em network type} $(\rho_i, \mu_i)$, which is a time-invariant subvector of $\tau_i^t$. Additionally, each agent pair $(i,j)$ is endowed with a time-invariant pair-specific random utility shock $\zeta_{ij}$, i.i.d.\ across pairs and independent of all other model primitives. For all $i,j \in \mathcal{N}_n$ with $i\neq j$, potential links in $\bm{A}$ satisfy

equation[equation omitted — 125 chars of source]

where $\lVert\cdot\rVert$ is a norm on $\mathbb{R}^d$ and $V(\cdot)$ a real-valued latent-index function, which we will later assume is eventually decreasing in its first argument ((ref)).\footnote{Since $A_{ij}$ is an undirected network, we assume $V(\lVert\rho_i-\rho_j\rVert, \mu_i, \mu_j, \zeta_{ij}) = V(\lVert\rho_i-\rho_j\rVert, \mu_j, \mu_i, \zeta_{ji})$.} In our applications, $V(\cdot)$ is an unknown function because the usual object of interest is some feature of $U(\cdot)$. Likewise, network types may be unobserved by the econometrician.

The second and third arguments of $V(\cdot)$ contain agent-specific characteristics that may influence link formation. If $V$ is monotonic in any of these elements, then this captures what graham2017econometric refers to as {\em degree heterogeneity}, where agents with more attractive characteristics $\mu_i$ are likely to have more connections (high degree). Alternatively, $V(\cdot)$ could depend on $\mu_i,\mu_j$ through $\bm{1}\{\mu_i = \mu_j\}$, which captures {\em homophily} in $\mu_i$. Homophily refers to the widely observed tendency for similar individuals to associate.

The first argument of $V(\cdot)$ requires homophily in {\em positions} $\rho_i$. This aspect of the model lends it a spatial dimension, which is essential for showing validity of the spatial HAC estimator. Under additional weak dependence conditions stated in (ref), spatial homophily implies that the correlation between $\psi_i$ and $\psi_j$ is smaller when $\lVert\rho_i-\rho_j\rVert$ is larger.

Note that the space on which agents are located need not correspond to physical space. They may be homophilous in other social dimensions, in which case we can define distance in terms of their social, rather than geographic, characteristics. In the case where positions are unobserved by the econometrician, the model is a nonlinear version of a latent space model commonly used in social network analysis hrh2002,breza_using_2017. In these models, agents are positioned on a latent “social space,” and connections form at a higher rate among socially similar agents. In this case, clearly a spatial HAC cannot be computed, since positions are unobserved, but our results show that the network HAC can be used instead.

remarkThe setup allows for correlation between $\mu_i$ and $\alpha_i$ (for example), in which case the network is endogenous. If agents are homophilous in $\mu_i$, then this captures {\em unobserved homophily}, a well-known hindrance to identifying social interactions shalizi2011homophily. For example, identification of peer effects in product adoption is confounded by peers with similar product preferences $\alpha_i$ forming connections at a higher rate.
remarkThe model does not allow for strategic interactions in link formation, meaning that $V(\cdot)$ is not a function of $\bm{A}$. Most papers in the econometric literature that address the problem of network endogeneity use similar models with no strategic interactions, viewing them heuristically as reduced-form approximations auerbach2018identification,hsieh2016social,johnsson2019estimation. leung2019normal prove a CLT for static models of network formation with strategic interactions. In principle, our CLTs can be generalized to their model, but we do not pursue this generalization because it mildly complicates the proofs in predictable ways without providing any new intuition.

Large-Network Asymptotics

This section formalizes the sequence of models along which we take limits in our asymptotic results. Recall that $\rho_i$, defined in the previous subsection, is a time-invariant subvector of $\tau_i^t$ for any $t$ and $\tau_i$ is the time series of $i$'s type $(\tau_i^t)_{t=0}^T$. Take the same time series but omit $\rho_i$ from each $\tau_i^t$, and call the result $Z_i$. Recall from the introduction that we define $\bm{\rho} = (\rho_i)_{i=1}^n$, and similarly define $\bm{Z}$ and $\bm{\zeta}$. Then the model is fully characterized by the tuple

equation[equation omitted — 85 chars of source]

where $U(\cdot)$ is the payoff function in (ref) and $V(\cdot)$ the latent index in the network formation model (ref). The new term $\lambda(\cdot)$ concerns the initial conditions model; being a static model with strategic interactions, initial outcomes will be determined by a selection mechanism $\lambda(\cdot)$, whose formal definition we postpone to (ref).

To establish the validity of spatial HAC estimators, we need positions $\rho_i$ to be sufficiently removed from each other asymptotically, which is the usual assumption of {\em increasing domain asymptotics} that is standard in spatial econometrics.

assump[Increasing Domain] Let $\tilde\rho_1, \tilde\rho_1, \dots$ be i.i.d., continuously distributed vectors in $\mathbb{R}^d$ with density $f$ bounded away from zero and infinity. Let $\omega_n = (n/\kappa)^{1/d}$ for some universal constant $\kappa>0$. Define $\rho_i = \omega_n\tilde\rho_i$ for all $i$. The observed outcome time series is realized according to the $n$th model of the sequence $\{(U, \lambda, V, \bm{\rho}, \bm{Z}, \bm{\zeta})\}_{n \in \mathbb{N}}$.

Under this sequence, we derive a CLT (ref) in (ref) and asymptotic properties of the HAC estimators in (ref).

remarkIn spatial econometrics, it is typically assumed that positions $\rho_i$ are non-random and separated by a universal minimum distance conley1999gmm,jenish_spatial_2012. (ref) is a slightly different model that follows the spatial graphs literature penrose2003 and is also sometimes used in spatial statistics lahiri2006resampling. Both assumptions have the same implication, that in the limit, any ball of fixed radius centered at an agent's position contains only an asymptotically finite number of other agents' positions. If $\omega_n$ diverges faster (slower) than the stated rate, then this ball will be asymptotically empty (save for the central agent); if it diverges slower, then it will contain an infinite number of agents in the limit.

Static Model of Social Interactions

This section introduces a static analog of the model studied in (ref), which is useful when the econometrician observes a snapshot of a large network. As before, we let $\mathcal{N}_n = \{1, \dots, n\}$ denote the set of agents, which are connected through an undirected network $\bm{A}$ realized according to the model in (ref). Each agent $i$ is endowed with a type $\tau_i = (X_i, \varepsilon_i)$, i.i.d.\ across agents, where $X_i$ is observed by the econometrician and $\varepsilon_i$ unobserved. Agents take a binary action, and payoffs may depend on the actions taken by others in the network. Agent $i$'s observed action satisfies

equation[equation omitted — 79 chars of source]

where the net payoff function $U(\cdot)$ depends on $i$'s type and a finite-dimensional vector of statistics

equation*[equation* omitted — 72 chars of source]

This captures strategic interactions through its dependence on $Y_{-i}$. Model (ref) states that observed actions are realized according to a pure-strategy Nash equilibrium; agents choose the action that maximizes payoffs given the actions of others in the network.

examplebramoulle2009identification study a network analog of the standard linear-in-means model manski1993identification, which is a model with continuous outcomes. Outcomes depend on the average outcome of peers (endogenous peer effects) and the average characteristics of peers (exogenous peer effects). The analogous specification for discrete choice is \begin{equation*} U(S_i, \tau_i) = S_i'\theta_1 + X_i'\theta_{-1} + \varepsilon_i, \quad S_i = \left( \frac{\sum_{j\neq i} A_{ij} Y_j}{\sum_{j\neq i} A_{ij}}, \frac{\sum_{j\neq i} A_{ij} X_j}{\sum_{j\neq i} A_{ij}} \right) \end{equation*} brock2001discrete,xu2018social. The first component of $S_i$ is the average action taken by neighbors. We can also consider type-weighted versions of the average action or nonlinear functions of $Y_{-i}$ and $A_i$ such as the minimum or maximum action hoxby2005taking.

Model (ref) is entirely analogous to the dynamic model (ref) but with contemporaneous rather than lagged dependent variables entering $S_i$. Hence, we can also consider analogs of Examples (ref) and (ref) in the static setting. In all of these examples, strategic interactions only operate through network neighbors of the ego $i$. We next impose this restriction more generally. Recall that $\mathcal{N}_{\bm{A}}^-(i,1)$ is the $1$-neighborhood of $i$, excluding $i$ herself.

assump[Local Interactions] For all $n$ and $i \in \mathcal{N}_n$, \begin{equation*} S_i = S(Y_{\mathcal{N}(i)}, \tau_i, \tau_{\mathcal{N}(i)}, A_i) \quadfor\quad \mathcal{N}(i) \equiv \mathcal{N}_{\bm{A}}^-(i,1). \end{equation*}

This states that $S_i$ is only a function of its arguments through agents connected to $i$.

Unlike the typical linear-in-means model for continuous outcomes, the nonlinearity of the discrete choice model typically gives rise to multiple equilibria brock2001discrete. This is obvious when $n=2$ and the two agents connected, since this is a $2 \times 2$ game of complete information. With multiple equilibria, the econometric model is incomplete in the sense that a reduced form does not (yet) exist tamer2003incomplete. To complete the model, we follow the empirical games literature and introduce a selection mechanism. Whether this is required to be known to the econometrician depends on the inference procedure, as discussed in the next subsection. Recall that $\bm{\tau} = (\tau_i)_{i=1}^n$. Let $\mathcal{E}(\bm{A},\bm{\tau}) \subseteq \{0,1\}^n$ be the set of Nash equilibria, i.e.\ the set of binary outcome vectors such that, for each $\bm{Y} \in \mathcal{E}(\bm{A},\bm{\tau})$, the $i$th component $Y_i$ satisfies (ref) for each $i$.

assump[Selection Mechanism] (a) A Nash equilibrium exists for any network size $n$, i.e.\ $|\mathcal{E}(\bm{A},\bm{\tau})| \geq 1$. (b) There exists a {\em selection mechanism} $\lambda(\bm{A},\bm{\tau})$ with range $\mathcal{E}(\bm{A},\bm{\tau})$ such that $\bm{Y} = \lambda(\bm{A},\bm{\tau})$ for any $n$.

Part (b) defines the selection mechanism as a reduced form mapping from the model primitives to the observed equilibrium outcome. Note that if $U(\cdot)$ does not vary in $S_i$, then $\lvert\mathcal{E}(\bm{A},\bm{\tau})\rvert = 1$, since this is simply a discrete-choice model with no strategic interactions, in which case $\lambda(\cdot)$ is trivial. In economic terms, $\lambda(\cdot)$ represents the process by which agents coordinate on a Nash equilibrium to play in the observed data. A simple example is a function that picks an element of $\mathcal{E}(\bm{A},\bm{\tau})$ uniformly at random. This is an econometric model used by bjorn1984simultaneous and soetevent2007discrete.\footnote{To see how this is formally represented in our notation, without loss of generality let the first component of $\varepsilon_i$, for each $i$, be a random variable $\gamma_i$ uniformly distributed on $[0,1]$ that is payoff-irrelevant. That is, it does not enter $U(\cdot)$. Partition the unit interval into $\lvert \mathcal{E}(\bm{A},\bm{\tau}) \rvert$ equally sized intervals and arbitrarily order the elements of $\mathcal{E}(\bm{A},\bm{\tau})$. Then let $\lambda(\cdot)$ be the function that selects the $k$th equilibrium if $\gamma_i$ (for any arbitrarily chosen $i$, say $i=1$) lies in the $k$th interval of the partition.} Another example is myopic best-response dynamics.

exampleAs discussed in (ref), the microeconomic literature on dynamic models of social influence predominantly considers models in which agents react myopically to the decisions of their peers in the previous period. Formally, fix $\bm{A},\bm{\tau}$ and an arbitrary $\bm{Y}^0 \in \{0,1\}^n$. Generate $Y^1 \in \{0,1\}^n$ by setting \begin{equation*} Y_i^1 = \bm{1}\left\{ U(S(Y_{-i}^0,\tau_i,\tau_{-i},A_i,A_{-i}), \tau_i) > 0 \right\} \end{equation*} for each $i \in \mathcal{N}_n$, and likewise generate $Y^2, Y^3, \dots$. This process is often referred to as {\em myopic best-response dynamics}. It is well known that if $\bm{Y}^0 = (1, \dots, 1)$, then under a game of strategic complements, this process converges to the “largest” Nash equilibrium $\bm{Y}^*$ in a precise sense jia2008happens. In that case, this process constitutes a mapping $\lambda(\cdot)$ from the primitives $(\bm{A},\bm{\tau})$ to a unique outcome $\bm{Y}^* \in \mathcal{E}(\bm{A},\bm{\tau})$.

Network Moments

Similar to (ref), we consider network moments that are averages of $\mathbb{R}^m$-valued {\em agent statistics} of the form

equation*[equation* omitted — 136 chars of source]

Our objective is to establish a CLT (ref) and validity of the HAC estimators under increasing domain asymptotics, as in the dynamic case ((ref)). We assume $\psi(\cdot)$ satisfies (ref), now under the new notation of the static setting. This condition simply states that $\psi_i$ is only a function of its arguments through $i$'s $K$-neighborhood in $\bm{A}$, which is satisfied for a wide variety of network moments useful for inference on social interactions. The remainder of this subsection provides illustrative examples.

exampleA trivial example of a network moment is $n^{-1}\sum_{i=1}^n Y_i$, which is the empirical choice probability and satisfies (ref) for $K=0$. A weighted version of the choice probability moment is used in the simulated method of moments estimator of xl2015 discussed in (ref). Another example is subnetwork moments \begin{equation*} \frac{1}{n} \sum_{i_1=1}^n \cdots \sum_{i_\ell=1}^n \bm{1}\{Y_H = y_H, A_H=a_H\}, \end{equation*} where $H = \{i_1, \dots, i_\ell\}$, $y_H \in \{0,1\}^\ell$, $a_H$ is a connected network on $H$, and $A_H$ is the subnetwork of $\bm{A}$ on $H$. Its expectation is proportional to the probability that agents in $H$ form subnetwork $a_H$ and choose outcomes $y_H$. As discussed below, this satisfies $K$-locality and can be used to construct moment inequalities for structural inference on $U(\cdot)$.

Method of Moments

The next two subsections discuss applications to parametric inference on $U(\cdot)$ that utilize moments falling within the scope of (ref). xl2015 consider a linear latent index model

equation*[equation* omitted — 135 chars of source]

where $\theta_1$ captures endogenous peer effects. They assume $\theta_1 \geq 0$, which implies that observed outcomes constitute a Nash equilibrium of a supermodular game. It is well known that the set of equilibria forms a complete lattice and therefore that there exists a “largest” equilibrium. For estimation, xl2015 assume the following.

assump(a) Realized outcomes correspond to the “largest” Nash equilibrium. (b) $(\bm{X},\bm{A}) \perp\!\!\!\perp \bm{\varepsilon}$. (c) The distribution of $\varepsilon_1$ given $X_1$ is known up to $\theta$.

Part (a) is equivalent to assuming that $\lambda(\cdot)$ is given as in (ref). By assuming a particular selection mechanism, this enables simulation of the model moments. Part (b) assumes an exogenous network.

The authors propose to estimate $\theta$ using simulated method of moments (SMM) based on the conditional choice probability, namely $n^{-1}\sum_i \psi_i$ for

equation*[equation* omitted — 106 chars of source]

Here, $\mathbf{P}_\theta$ refers to the probability under the model with structural parameters given by $\theta = (\theta_1, \theta_{-1})$, and $h(\cdot)$ is a vector-valued instrument function that converts to unconditional moments. In practice, the conditional probability can be simulated due to (ref). For inference, they propose to use the parametric bootstrap.

For the SMM estimator to have a normal limit, a key property to verify is a central limit theorem for the moments themselves (this is assumption (d) of the authors' Theorem B.1). xl2015 establish a CLT for the moments by invoking a CLT for near-epoch dependent data due to jenish_spatial_2012. Using our CLT instead requires significantly weaker restrictions on the network formation model.\footnote{In particular, we do not require their Assumption 7, which states that if two agents are linked, then their spatial distance must fall below some universal constant.} Additionally, our CLT applies to other network moments that can be used for estimation, such as subnetwork moments. Finally, our results enable the use of HAC variance estimators for inference. This consists of taking the usual GMM sandwich formula but replacing the sample variance of the moments in the middle of the sandwich with one of our HAC estimators.

For our CLT to apply to these moments, we need to verify $K$-locality ((ref)). This holds, for example, under the restriction

equation*[equation* omitted — 89 chars of source]

in which case (ref) is satisfied for $K=1$. The restriction states that the instruments only depend on the observed types and links involving neighbors of $i$. In their simulation study, xl2015 choose $h(X_i,X_{-i},A_i,A_{-i}) = (1, X_i, \sum_{j\neq i} A_{ij} X_j / \sum_{j\neq i} A_{ij})$, which satisfies this condition. This choice is likely motivated by the intuition in bramoulle2009identification that the average covariates of peers is an instrument for the endogenous peer effect.

Set Inference

Without imposing restrictions on equilibrium selection, $\theta$ is typically partially identified, and the identified set can be characterized in terms of moment inequalities beresteanu2011sharp,galichon2011set. For discrete games on networks, li2016partial propose moment inequalities that provide a conservative characterization of the identified set. These are based on subnetwork moments of the form $\mathbf{P}(Y_H = y_H \mid X_H, A_H)$ for $H\subseteq \mathcal{N}_n$, which gives the conditional joint distribution of outcomes of agents in $H$ given the subnetwork $A_H$.

The main identification result of li2016partial constructs a function $G(\cdot)$ such that

equation[equation omitted — 100 chars of source]

for any $y_H \in \{0,1\}^{\lvert H\rvert}$ when $\theta_0$ is the true parameter. The bound $G(\cdot)$ can be computed via simulation, while the conditional probability can be estimated nonparametrically, provided a law of large numbers holds.

While li2016partial focus on estimation, we next discuss how their bounds may be used for inference on $\theta_0$. Since their bounds come from conditional moment inequalities, we can apply a procedure due to andrews2013inference. To correct for autocorrelation, we need to use the asymptotic version of their test, consisting of steps 1 and 2 in their section 9, with a valid variance estimator in place of the sample variance. Asymptotic validity requires a CLT for the network moments, and our results can be applied for this purpose.

The procedure of andrews2013inference first converts to unconditional moment inequalities by multiplying moments with real-valued instrument functions $h(X_H, A_H)$, which can be done without loss of generality by using a large enough set of functions (see their section 3.3). Define

equation*[equation* omitted — 128 chars of source]

Then by (ref), we have $m(y_H; \theta) \leq 0$ when $\theta$ is the true parameter.

Consider the “sample analog” of these moments

equation*[equation* omitted — 167 chars of source]

where $H = \{i_1, \dots, i_\ell\}$; see (ref) for discussion of the $n^{-1}$ scaling. Its expectation is proportional to $m(y_H; \theta)$, where the constant of proportionality does not depend on $n$ or $\theta$. Since $A_H$ has finite support, we can consider instrument functions of the form

equation*[equation* omitted — 67 chars of source]

where $a_H$ is a network on $H$.

The sample moments can be rewritten as $n^{-1} \sum_i \psi_i$ for

equation*[equation* omitted — 144 chars of source]

where $H = \{i, i_2, \dots, i_\ell\}$. To apply our CLT, we need to verify (ref). If we use any subnetwork moment(s) in which $a_H$ is a connected network, then the assumption holds for $K=\ell$.

Variance Estimators

We postpone formal conditions for CLTs for the static and dynamic models to (ref) and (ref), respectively. This section is concerned with the properties of spatial and network HAC estimators for the variance $\bm{\Sigma}_n$ in (ref). The notation used is applicable to either the static or dynamic model. The first two subsections discuss the validity of these estimators under first-order stationarity conditions. To our knowledge, such conditions are always used to establish consistency of HAC estimators in time series and spatial econometrics. However, they are typically only satisfied in GMM-type settings and not, for example, in moment inequality models or many nonparametric models. The third subsection is concerned with inference when first-order stationarity fails.

Spatial HAC

The standard spatial HAC estimator is given by

equation*[equation* omitted — 149 chars of source]

where $h_n \in \mathbb{R}_+$ is the bandwidth, $K(\cdot)$ a real-valued kernel function with domain $\mathbb{R}^d$, and $\bar{\psi} = n^{-1}\sum_{i=1}^n \psi_i$, the vector of network moments of interest. In cases where ${\bf E}[\bar{\psi}]$ is a known constant, for example zero in GMM models, we can replace $\bar{\psi}$ in the HAC estimator with this constant. In practice, it is common to use a product kernel, which has the form

equation*[equation* omitted — 99 chars of source]

where $d$ is the dimension of $\rho_i$, $\rho_{ik}$ is the $k$th component of $\rho_i$, and $\tilde K$ is a real-valued kernel function with domain $\mathbb{R}$. A typical choice for $\tilde K(\cdot)$ is the Bartlett kernel $\tilde K(x) = (1 - \lvertx\rvert) \bm{1}\{\lvertx\rvert \leq 1\}$. For other examples, see e.g.\ andrews1991heteroskedasticity, p.\ 821.

In practice, it is best to show robustness of the standard errors across a reasonable range of bandwidths, as in conley2010learning. For example, one can recompute the HAC estimator for several bandwidth values in a neighborhood of some reference value. This value may be obtained from domain knowledge, say by examining the distribution of distances. Alternatively, one can calibrate $h_n$ by Monte Carlo simulation with a parametric submodel. Data-driven choice of $h_n$ is an important but difficult topic for future research and will not be addressed in this paper.

We next study the theoretical properties of $\hat{\bm{\Sigma}}_\rho$. To our knowledge, all existing results on the validity of HAC estimators, whether for time series, spatial processes, or network settings, require first-order stationarity andrews1991heteroskedasticity,conley1999gmm,kojevnikov2019limit. In our setting, this corresponds to the following.

assump[First-Order Stationarity] ${\bf E}[\psi_1] = {\bf E}[\psi_1 \mid \rho_1]$.\footnote{In spatial econometrics, the usual stationarity condition is that ${\bf E}[\psi_i \mid \bm{\rho}]$ does not depend on $i$. This is analogous to ours except conditional on the set of positions, since they are treated as fixed conley1999gmm,jenish2016. Our asymptotic results are unconditional, which is why the assumption is slightly different.}

This says that the absolute value of an agent's position is mean independent of her agent statistic. Its technical purpose is that the long-run covariance aggregates over covariances of agents $i$ and $j$ conditional on their positions. This covariance is a function of ${\bf E}[\psi_i \mid \rho_i]$, which needs to be consistently estimated. Under (ref), this is possible using the sample mean $\bar{\psi}$.

example(ref) typically holds in GMM-type settings. For example, consider the moments in (ref). We have \begin{equation} {\bf E}[\psi_i \mid \rho_i] = {\bf E}\big[Y_i h(X_i,X_{-i},A_i,A_{-i}) \mid \rho_i\big] - {\bf E}\big[ {\bf E}_\theta[Y_i \mid \bm{X},\bm{A}] h(X_i,X_{-i},A_i,A_{-i}) \mid \rho_i], \end{equation} where ${\bf E}_\theta[\cdot]$ is the expectation under the model with structural parameters given by $\theta$ and ${\bf E}[\cdot]$ is the expectation under the true model. Also, \begin{align*} {\bf E}\big[Y_i h(X_i,X_{-i},A_i,A_{-i}) \mid \rho_i\big] &= {\bf E}\big[ {\bf E}[Y_i h(X_i,X_{-i},A_i,A_{-i}) \mid \rho_i, \bm{X}, \bm{A}] \mid \rho_i\big] \\ &= {\bf E}\big[{\bf E}[Y_i \mid \rho_i, \bm{X}, \bm{A}] h(X_i,X_{-i},A_i,A_{-i}) \mid \rho_i\big] \\ &= {\bf E}\big[{\bf E}[Y_i \mid \bm{X}, \bm{A}] h(X_i,X_{-i},A_i,A_{-i}) \mid \rho_i\big], \end{align*} where the last line holds if positions only enter payoffs through the network $\bm{A}$; (ref)(b) suffices for this. If $\theta$ is the true parameter, then ${\bf E}[Y_i \mid \bm{X}, \bm{A}] = {\bf E}_\theta[Y_i \mid \bm{X}, \bm{A}]$, so (ref) $=\bm{0}$ and (ref) holds. A similar argument can be used to verify the assumption for the application in (ref).
assump[HAC Kernel] $K(0)=1$; $K(x) = 0$ for all $x \in \mathbb{R}^d$ such that $\lVertx\rVert > 1$; $\int \lvertK(x)\rvert \,\text{d}x < \infty$; $K$ is continuous at zero; and $K^* \equiv \sup_x \lvertK(x)\rvert < \infty$.

This assumption imposes standard restrictions on the kernel. Now, let $\lambda_\text{min}(\bm{M})$ denote the smallest eigenvalue of the matrix $\bm{M}$ and $\lVert\bm{M}\rVert$ the max norm of $\bm{M}$.

theoremSuppose $h_n = O(n^{1/(3d)})$ and $h_n \rightarrow \infty$. Under Assumptions (ref) and (ref) and the conditions required for a CLT (Theorems (ref) and (ref) in the static and dynamic cases, respectively), $\lVert\hat{\bm{\Sigma}}_\rho - \tilde{\bm{\Sigma}}_n\rVert \stackrel{p}\longrightarrow 0$ for some sequence of matrices $\{\tilde{\bm{\Sigma}}_n\}_{n\in\mathbb{N}}$ such that $\lambda_\text{min}(\bm{\Sigma}_n - \tilde{\bm{\Sigma}}_n) \geq 0$ for all $n$. If there exists a constant vector $c$ such that ${\bf E}[\psi_1]=c$ for any $n$, then $\bm{\Sigma}_n = \tilde{\bm{\Sigma}}_n$ for all $n$.
proofSee (ref) for the formal proof and a proof sketch.

The second conclusion of the theorem states that $\hat{\bm{\Sigma}}_\rho$ is consistent if ${\bf E}[\psi_1]$ does not vary with $n$. This is satisfied if $\psi(\cdot)$ is a moment function used for GMM estimation, since then ${\bf E}[\psi_1]=\bm{0}$. More generally, if the moment is centered at the right conditional expectation, usually both (ref) and ${\bf E}[\psi_1]=c$ hold. The applications in (ref) and (ref) are centered.

When moments are not centered, typically ${\bf E}[\psi_1]$ varies with $n$, and the first conclusion of (ref) shows that $\hat{\bm{\Sigma}}_\rho$ is asymptotically conservative in this case, which appears to be a new result. However, this also requires first-order stationarity to hold, which is often violated when moments are uncentered. In moment inequality models ((ref) and (ref)) and often in nonparametric models, moments are typically not centered, so the results here are inapplicable. We discuss alternatives in (ref).

Network HAC

Next we study the network HAC estimator for $\bm{\Sigma}_n$ used by acemoglu2015state and eckles2016estimating. This essentially consists of taking the usual spatial HAC estimator and replacing spatial distance with path distance. Recall from (ref) the definition of path distance $\ell_{\bm{A}}(i,j)$. Let $h_n \in \mathbb{R}_+$ be a bandwidth and $K\colon \mathbb{R} \rightarrow \mathbb{R}$ a kernel function. The network HAC estimator is given by

equation*[equation* omitted — 156 chars of source]

For consistency, the bandwidth $h_n$ will be required to grow essentially at a logarithmic rate. In practice, it is best to either show robustness of the standard errors across a reasonable range of bandwidths in some neighborhood of a reference value, as in conley1999gmm. This value may be obtained from domain knowledge, say by examining typical path lengths in the network. Another reference value is simply setting $h_n = \log n$, which is used in our simulation study. Finally, one can calibrate $h_n$ by Monte Carlo simulation with a parametric submodel.

The basic idea behind our proof for the validity of the network HAC estimator is as follows. In our model, agents are homophilous in positions, which implies that agents that are close in path distance should also be close in spatial distance. This suggests we can use the former to approximate the latter and expect $\hat{\bm{\Sigma}}_{\bm{A}} \approx \hat{\bm{\Sigma}}_\rho$.

remark[Related Literature] A recent paper by kojevnikov2019limit prove consistency of $\hat{\bm{\Sigma}}_{\bm{A}}$ for the variance of a class of moments satisfying a novel notion of network weak dependence they call “conditional $\psi$-weak dependence.” This is a modification of a concept in the time series literature using path distance in place of temporal distance. The condition is distinct from stabilization, the weak dependence conditions used in our CLTs, which do not condition on the network.\footnote{See (ref) for the formal stabilization conditions. These are taken from leung2019normal, which are, in turn, modifications of assumptions first proposed in the literature on geometric graphs penrose2003.} We view these as complementary contributions, since they consider applications for which stabilization cannot be used but do not consider applications to social interactions models. Additionally, kojevnikov2019limit impose general conditions on the network structure without assuming a specific model of network formation. We instead assume a particular (nonparametric) model of network formation, which enables us to derive lower-level conditions.
assump[First-Order Stationarity] Let $N$ be any random variable supported on the natural numbers, independent of all other primitives. Suppose the number of agents is given by $N$, and let $\psi_1^N$ denote agent 1's statistic (in either the static or dynamic model), $\bm{X}^N = (X_i)_{i=1}^N$, and $\bm{A}^N$ the network. Then ${\bf E}[\psi_1^N] = {\bf E}[\psi_1^N \mid \bm{X}^N, \bm{A}^N, N]$.

The simplest way to understand the assumption is to consider the case where $N = n$ a.s., which is the model originally specified in (ref) and (ref). Then this reduces to ${\bf E}[\psi_1] = {\bf E}[\psi_1 \mid \bm{X}, \bm{A}]$, which is analogous to (ref). As with the latter assumption, stationarity is satisfied in GMM-type settings or more generally when moments are centered at their conditional (on $\bm{X},\bm{A}$) expectations. This includes the applications in (ref) and (ref). Whether consistent variance estimators can be obtained when (ref) fails is an open question.

For technical reasons, we need stationarity to hold not only for models where $N=n$ but also for $N \sim \text{Poisson}(n)$, which is why the statement of (ref) is more complicated. This does not seem to rule out any applications of interest. Similar to the spatial HAC setting, the technical purpose of this assumption is that the long-run covariance aggregates over the conditional covariances of agents $i$ and $j$. Hence, it is a function of ${\bf E}[\psi_i \mid \bm{X}, \bm{A}]$, which needs to be consistently estimated. Under (ref), this is possible using the sample mean $\bar{\psi}$.

In order to approximate spatial with path distance, we require the following intermediate-level condition on the network formation model.

assump[Path Distance] (a) There exists $c>0$ such that \begin{equation*} \mathbf{P}(\ell_{\bm{A}}(i,j) > c \lVert\rho_i-\rho_j\rVert \mid \ell_{\bm{A}}(i,j) < \infty) = O(n^{-3/4}). \end{equation*} (b) For any $c > 0$, \begin{equation*} \lim_{\epsilon\rightarrow\infty} \limsup_{n\rightarrow\infty} \mathbf{P}(\ell_{\bm{A}}(i,j) > \epsilon \mid \ell_{\bm{A}}(i,j) < \infty, \lVert\rho_i-\rho_j\rVert) \bm{1}\{\lVert\rho_i-\rho_j\rVert < c\} = 0 a.s. \end{equation*}

This assumption concerns the relationship between spatial and path distance. Part (a) states that, for connected agents, their path distance is not asymptotically much larger than spatial distance with high probability. Part (b) requires the path distance of connected agents to be asymptotically bounded if their spatial distance is likewise bounded.

This seems to be a reasonable condition since agents are spatially homophilous, so agents that are spatially distant should be similarly distant in the network. However, verification of this condition is challenging because the relationship between spatial and path distance is a complicated graph-theoretic problem. To our knowledge, in random graph theory, the requisite results have only been developed for the random geometric graph model, which is a widely studied spatial graph model penrose2003. In this model, two agents connect if and only if their spatial distance falls below some fixed threshold. In (ref), we draw on recent theoretical results to verify (ref) for this model.

While no corresponding results are presently available for more general spatial graphs, it seems reasonable to conjecture that (ref) holds for the general class of models in (ref) under the regularity condition that the linking probability decays exponentially with distance, a condition we require anyway for a CLT in (ref). The basis of our conjecture is that proving a result under smooth exponential decay is usually expected once the corresponding result has been proven for a hard threshold decay, like the random geometric graph. For instance, compare limit theory for $M$-dependence, where neighbors beyond distance $M$ are uncorrelated with the ego, and $\alpha$-mixing, where the covariance instead may decay smoothly with distance.

assump[Network Exogeneity] For any $n\in\mathbb{N}$, $\bm{\varepsilon} \perp\!\!\!\perp \bm{A} \mid \bm{X}$.

This is another new condition not required by the spatial HAC. In the static case, it rules out network endogeneity, since $\bm{A}$ must be independent of unobservables given observables. Thus, unobserved homophily is not permitted.\footnote{The framework of kojevnikov2019limit also seems to rule out network endogeneity, since errors terms must be weakly dependent conditional on the network. Under endogeneity, the conditional dependence structure appears very difficult to characterize.} In the dynamic case, recall that $X_i$ collects the observed covariates and the fixed effects, while $\varepsilon_i$ collects the idiosyncratic errors. Then (ref) implies that the errors are independent of the network, but network endogeneity is still allowed through dependence between the fixed effect and the network.

theoremSuppose $h_n \rightarrow \infty$ and $\log h_n / \log n \rightarrow 0$. Under Assumptions (ref)--(ref) and the conditions required for a CLT (Theorems (ref) and (ref) in the static and dynamic cases, respectively), $\lVert\hat{\bm{\Sigma}}_{\bm{A}} - \tilde{\bm{\Sigma}}_n\rVert \stackrel{p}\longrightarrow 0$ for some sequence of matrices $\{\tilde{\bm{\Sigma}}_n\}_{n\in\mathbb{N}}$ such that $\lambda_\text{min}(\bm{\Sigma}_n - \tilde{\bm{\Sigma}}_n) \geq 0$ for all $n$. If there exists a constant vector $c$ such that ${\bf E}[\psi_1]=c$ for all $n$, then $\bm{\Sigma}_n = \tilde{\bm{\Sigma}}_n$ for all $n$.
proofSee (ref) for the formal proof and a proof sketch.

The conclusions are the same as those of (ref). When moments are centered, typically first-order stationarity and ${\bf E}[\psi_1]=c$ hold, in which case the result delivers consistency. On the other hand, if ${\bf E}[\psi_1]$ varies with $n$, then the estimator is asymptotically conservative.

remark[Bandwidth Rate] The bandwidth is required to grow at a sub-polynomial rate, unlike the spatial HAC. Actually, inspection of the proof shows that the result holds if $h_n$ diverges at a rate that is polynomial in $n$ with sufficiently small degree, but the degree depends on the dimension of $\rho_1$, which is unknown. Thus, a sub-polynomial rate is the price to pay for not observing positions.
remark[Network Endogeneity] (ref) is important to establish (ref). It serves to ensure that disconnected agents are uncorrelated, i.e.\ $\text{Cov}(\psi_i, \psi_j \mid \bm{X}, \bm{A}) = 0$ if $\ell_{\bm{A}}(i,j)=\infty$. If this covariance were non-zero, then the network HAC would be biased because it cannot account for the covariance between pairs of disconnected agents, since the kernel weight multiplying $(\psi_i-\bar{\psi}) (\psi_j-\bar{\psi})'$ in $\hat{\bm{\Sigma}}_{\bm{A}}$ would be zero. It may seem intuitive for this covariance to be zero, and it certainly is true when the network is exogenous. However, without exogeneity, the problem is that we are conditioning on the endogenous network $\bm{A}$, which can induce nonzero correlation between the agents.
remark[Positive Semidefiniteness] In our simulation study, $\hat{\bm{\Sigma}}_{\bm{A}}$ is always positive semidefinite, but in general, this is not guaranteed in finite sample. To ensure positive semidefiniteness, we can modify the estimator using a correction proposed by kojevnikov2019bootstrap. Let $\bm{Q}_n \bm{\Lambda}_n \bm{Q}_n'$ be the eigendecomposition of $\hat{\bm{\Sigma}}_{\bm{A}}$, $c$ a positive constant, and $\bm{I}_n$ the $n\times n$ identity matrix. Define \begin{equation*} \hat{\bm{\Sigma}}_{\bm{A}}^+ \equiv \bm{Q}_n \max\{\bm{\Lambda}_n, c\bm{I}_n\} \bm{Q}_n', \end{equation*} where “max” refers to the element-wise maximum. Then $\hat{\bm{\Sigma}}_{\bm{A}}^+$ is positive semidefinite because the smallest eigenvalue is at least $c>0$ by construction. Consistency requires $c \rightarrow 0$ as $n \rightarrow \infty$ kojevnikov2019bootstrap. In practice, we might choose, say, $c=0.1$.

Inference Without Stationarity

The theorems in the previous subsections justify the use of HAC estimators in GMM-type applications, such as (ref) and (ref), or more generally, settings where the moments are centered at the right conditional expectations. However, there are important settings where moments are uncentered, for example moment inequality models, which include the applications in (ref) and (ref). Another application in which moments may be uncentered is nonparametric inference, for example, on the conditional choice probability ${\bf E}[Y_1 \mid X_1]$, or other such moments.

We next discuss two ways in which nonstationarity can be addressed. First, we note that the generalized spatial HAC proposed by leung2019normal can be applied to our setting, an estimator that is consistent when both (ref) and ${\bf E}[\psi_1]=c$ fail to hold. Second, we derive new general (albeit strong) sufficient conditions under which arbitrary moments (whether centered or not) obey spatial first-order stationarity {\em asymptotically}. Combined with (ref), the latter result enables the use of the spatial HAC in a larger set of environments, with the caveat that it is conservative rather than consistent (since ${\bf E}[\psi_1]$ will typically vary with $n$). We are unable to prove corresponding results for the network HAC, since it appears difficult to derive general sufficient conditions for an asymptotic version of stationarity ((ref)) outside of GMM-type models.

{\bf Generalized Spatial HAC.} In settings where (ref) fails to hold, we can use the generalized spatial HAC estimator proposed in \S 6.1 of leung2019normal. This is given by

align[align omitted — 606 chars of source]

where $\tilde K(\cdot)$ is a kernel function and $b_n$ a bandwidth for the kernel estimator $\hat\theta(\rho_1)$ of ${\bf E}[\psi_1 \mid \rho_1]$. Thus, we avoid having to impose stationarity because we estimate ${\bf E}[\psi_1 \mid \rho_1]$ directly, nonparametrically. Consistency of the generalized spatial HAC requires $h_n^d = O(n^{1/4})$, $h_n\rightarrow\infty$, and $\hat\theta(p)$ to be uniformly consistent and converging faster than the $n^{-1/4}$ rate, as usual. Under these requirements and high-level stabilization conditions, leung2019normal prove consistency. Our CLTs verify these conditions and hence imply consistency for social interactions models.

remarkThis estimator differs slightly from leung2019normal due to the $n^{-1/d}$ scaling in $\hat\theta(\cdot)$ and the rate condition on $h_n$. The scaling addresses the problem that we want to estimate ${\bf E}[\psi_1 \mid \rho_1]$ by averaging over the agent statistics of agents with positions near $\rho_1$. However, positions are drifting apart as $n$ diverges, since $\rho_i = \omega_n \tilde\rho_i$. Fortunately, since $\omega_n$ drifts at rate $n^{1/d}$, we can simply scale down positions by this rate, converting $\rho_i$ to $\tilde\rho_i/\kappa^{1/d}$, the latter of which has bounded support and does not depend on $n$. The bandwidth $h_n$ has a different rate because, in this paper, we model positions as $\omega_n\tilde\rho_i$, whereas they model positions as simply $\tilde\rho_i$, which may be more reasonable when agents are closely located in space. This modification has no effect on the consistency proof, since the required high-level conditions (stated in (ref) and verified in our CLTs) are the same as those of leung2019normal except we change positions from $\{\tilde\rho_i\}_{i=1}^n$ to $\{\omega_n\tilde\rho_i\}_{i=1}^n$.

{\bf Asymptotic Stationarity.} We next provide general sufficient conditions applicable to {\em any} network moment under which it {\em asymptotically} satisfies first-order stationarity, i.e.\ $\lvert{\bf E}[\psi_1] - {\bf E}[\psi_1 \mid \rho_1=\omega_n p]\rvert = o(n^{-1/3})$ for any $p$.\footnote{The proof of (ref) uses exact stationarity ((ref)), but an inspection of the proof shows that asymptotic first-order stationarity suffices.} The conditions are strong, but they are useful for understanding what widely used stationarity conditions essentially require when moments are not centered. Recall the definition of $Z_i$ from (ref) in the static setting and (ref) in the dynamic setting.

assump(a) $Z_1 \perp\!\!\!\perp \rho_1$. (b) $f$ is the uniform distribution on some bounded subset of $\mathbb{R}^d$ for which the origin is an interior point.

Part (a) says that positions are independent of all other attributes, and (b) implies that agents are not more clustered in some regions of space than others.

theoremSuppose the CLT assumptions hold (Theorems (ref) and (ref) in the static and dynamic cases, respectively). Under (ref), there exists a sequence of constants $\{c_n\}_{n\in\mathbb{N}}$ such that $\lvert{\bf E}[\psi_1 \mid \rho_1=\omega_n p] - c_n\rvert = o(n^{-1/3})$ for any $p \in \text{supp}(f)$.
proofSee (ref).

Combined with (ref), this establishes that the spatial HAC is asymptotically conservative for any network moment under (ref).

The remainder of this subsection discusses (ref). Certainly it is a strong condition, but it clarifies what stationarity generally demands in spatial settings when moments are uncentered. The reason we need (b) is that, if some subset of $\text{supp}(f)$ had higher density, then agents positioned in that region would tend to have higher degrees because there are more nearby alters in that region, as well. Hence, degree would not be mean-independent of position, and we would expect other network statistics to similarly depend on an agent's position. We need (a) for the same reason; if an agent's position were correlated with some attribute of the agent, then agents with the right positions would have desirable realizations of the attribute and thus tend to have high degrees. In either case, the absolute value of an agent's position would influence her outcome, whereas, by definition, stationarity demands that only her relative position matters.

Thus, (ref) implies that {\em positions play no role in degree heterogeneity}. By this we mean the following. In general from specification (ref), variation in $\mu_i$ and $\rho_i$ can increase or decrease $i$'s degree, but (ref) rules out variation due to $\rho_i$. Thus, positions only generate homophily, not heterogeneity in degree. Of course, $\mu_i$ can still do either. Consequently, under (ref), the network can still be endogenous with respect to the outcome equation due to correlation between the error in the outcome equation and $\mu_i$.

CLT for the Static Model

In this section, we state formal conditions for a CLT for network moments defined in (ref). The first two subsections present and motivate our main conditions for weak dependence and the third states regularity conditions. The weak dependence conditions are analogous to those used by leung2019normal to prove a CLT for static models of network formation with strategic interactions. Their proof strategy is to verify high-level “stabilization” conditions for a normal approximation (their Theorem C.2) by constructing a dependency neighborhood for each observation, what they call the {\em relevant set}, and then using branching process techniques to bound the sizes of these neighborhoods. We use the same methodology for both the static and dynamic model, but the arguments need to be modified for the social interactions setting.

Additionally, rather than applying their Theorem C.2, we verify the high-level conditions of our (ref), which is similar to Theorem C.2 but makes two modifications. First, we use an increasing domain setup ((ref)), whereas leung2019normal use a quasi-infill setup (see (ref)). Second, Theorem C.2 provides a closed-form expression for the limit variance, whereas our theorem does not, which allows us to dispense with a high-level continuity condition.

Strategic Interactions

The first key condition required for a CLT is a restriction on the strength of strategic interactions. For motivation, consider the standard linear-in-means model as studied in bramoulle2009identification. The outcome equation is given by

equation*[equation* omitted — 115 chars of source]

where $\bm{A}$ may be row-normalized. For this model to have a reduced form,

equation[equation omitted — 80 chars of source]

is required, where $\lambda_\text{max}(\bm{A})$ is the largest eigenvalue of the adjacency matrix $\tilde{\bm{A}}$. This is evidently a restriction on the magnitude of peer effects. It ensures equilibrium existence and weak dependence.

We impose an assumption analogous to (ref). While it may appear superficially more complicated, this is unavoidable due to the nonlinearity of the model. First we require some definitions. Let

equation[equation omitted — 166 chars of source]

The significance of this indicator is that its expectation equals

equation[equation omitted — 124 chars of source]

(assuming measurability). This is our analog of $\beta$ in (ref). It measures the strength of strategic interactions, since it corresponds to the average causal effect on choosing action 1 of changing $S_i$ from its “lowest” to its “highest” possible value.

The next assumption imposes a restriction on (ref) in relation to the network topology. Recall that $i$'s position $\rho_i$ is a subvector of her type $\tau_i$. Let $Z_i$ be the subvector of $\tau_i$ that omits $\rho_i$ (so that $\tau_i = (\rho_i,Z_i)$), $d_z$ the dimension of $Z_i$, and $\Phi(\cdot \,|\, p)$ the conditional distribution of $Z_i$ given $\rho_i=p$. Define

equation[equation omitted — 136 chars of source]

For any $h\colon \mathbb{R}^d \times \mathbb{R}^{d_z} \rightarrow \mathbb{R}$, define the mixed norm

equation*[equation* omitted — 142 chars of source]

where $\Phi^*(\cdot)$ is a distribution given in the next assumption. Let $\bar{f} = \sup_{p\in\mathbb{R}^d} f(p)$.

assump[Strength of Interactions] There exists a distribution $\Phi^*(\cdot)$ on $\mathbb{R}^{d_z}$ such that, for any $p,p' \in \mathbb{R}^d$ and $z \in \mathbb{R}^{d_z}$, \begin{equation*} \int_{\mathbb{R}^{d_z}} \varphi(p,z;p',z') \,d\Phi(z' \,|\, p') \leq \int_{\mathbb{R}^{d_z}} \varphi(p,z;p',z') \,d\Phi^*(z'). \end{equation*} Furthermore, \begin{equation*} \lVerth\rVert_{\bm{m}} < 1 \quadfor\quad h(p,z) \equiv \kappa \bar{f}\, \int_{\mathbb{R}^d} \left( \int_{\mathbb{R}^{d_z}} \varphi(p,z;p',z')^2 \,d\Phi^*(z') \right)^{1/2} \,dp'. \end{equation*}

The first part of the assumption is a regularity condition on the conditional distribution of $Z_i$. The substantive requirement is $\lVerth\rVert_{\bm{m}} < 1$. To see how this assumption restricts (ref), note that, for any $n$,

equation[equation omitted — 247 chars of source]

using a change of variables and the fact that $n\omega_n^{-d} = \kappa$. Clearly $\lVerth\rVert_{\bm{m}}$ is an upper bound on the right-hand side, so a necessary condition for (ref) is that the left-hand side is strictly less than one. The interpretation of the left-hand side becomes clearer in the special case where the network is exogenous in the sense that $A_{ij} \perp\!\!\!\perp \mathcal{R}_j^c$. Then $\lVerth\rVert_{\bm{m}} < 1$ implies

equation[equation omitted — 158 chars of source]

for any $n$, since the left-hand side of (ref) equals the expectation of the left-hand side of (ref). This is transparently a restriction on the strength of strategic interactions, requiring it to be bounded below the inverse of the expected degree. The intuition is that, when mean degree is higher, an agent's action can impact more neighbors, so weak dependence requires weaker strategic interactions.

To see more clearly how (ref) compares with (ref), note that $\lambda_\text{max}(\bm{A}) \leq \max_i \sum_j A_{ij}$. Hence, a sufficient condition is

equation*[equation* omitted — 62 chars of source]

which has precisely the form of (ref). Whether $\bm{A}$ is row-normalized, our condition and (ref) have similar behavioral implications, as discussed in Remark A.1 of leung2019compute. Finally, it should be noted that both (ref) and (ref) are restrictions on the effect of $S_i$ on $Y_i$ only in a partial-equilibrium sense. The general equilibrium effect can be substantially larger, as seen in the simulation study of leung2016 in the context of network formation.

Conditions similar to (ref) have been utilized by xl2015 and xu2018social for social interactions models. leung2016 and menzel2015large use analogous conditions for weak dependence in network formation and matching models, respectively. The former paper discusses similarities between this condition and related assumptions for temporal and spatial autoregressive models.

Equilibrium Selection

Whereas in the linear-in-means model, (ref) is enough to guarantee equilibrium uniqueness, in our setting multiple equilibria are possible under (ref). By (ref), the selection mechanism $\lambda(\cdot)$ can be any function of $\bm{A}$ and $\bm{\tau}$, which can potentially induce strong dependence between the outcomes of agents distant in the network. For this reason, we require a restriction on $\lambda(\cdot)$ for weak dependence. We leave the formal statement of the assumption to (ref). Our goals here are (1) to motivate why a restriction is needed and (2) note that the assumption is satisfied under variants of myopic best-response dynamics, such as those used to define the dynamic model.

To illustrate why we need to restrict $\lambda(\cdot)$, let us take an extreme example. Suppose types are realized such that $\bm{A}$ has two components (disconnected subnetworks). By definition, agents in separate components are infinitely far apart in terms of path distance. Under (ref), the set of Nash equilibria for the full game is the Cartesian product of the set Nash equilibria on each component alone, since payoffs in one component do not depend on actions in the other component. In this respect, they appear to be two separate games. Nonetheless, it is easy to construct examples of $\lambda(\cdot)$ such that the realization of equilibria is correlated between components. For example, suppose types are realized such that there are two possible equilibria in each component. Consider an equilibrium selection mechanism that chooses one equilibrium over the other in each component, depending on the realization of agent 1's type $\tau_1$. Then the same underlying random vector determines equilibrium selection in both components, so their outcomes are correlated. In the cross-sectional setting, this problem is analogous to having a selection mechanism simultaneously determine equilibrium realizations across a set of distinct games, a scenario implicitly ruled out when researchers assume separate realizations of the same game are independent epstein2016robust.

In the previous example, all agents effectively coordinate on equilibria according to the realization of a common signal $\tau_1$. For weak dependence, we must rule out coordination between “distant” agents. In order to formalize what we mean by “distant” and state the required condition ((ref)), we need several new definitions that have to be motivated at length, which we defer to (ref). In brief, the basic idea of the condition is to only allow agents within certain subnetwork neighborhoods to coordinate on their equilibrium actions. Such a restriction holds under variants of {\em myopic best-response dynamics}. This is an interesting class of selection mechanisms to consider because they are widely used in theoretical work to model diffusion, as discussed in (ref), and they define the other of our main applications, the dynamic model.

To define these dynamics more formally, suppose an initial vector of outcomes $\bm{Y}^0$ is chosen randomly, and then agents iteratively best-respond according to (ref) until convergence. Under a game of strategic complements, the process necessarily converges to a Nash equilibrium milgrom1990rationalizability. Otherwise, convergence requires $\bm{Y}^0$ to lie along an improving path jackson_evolution_2002. In either case, this constitutes a selection mechanism because it maps $(\bm{A},\bm{\tau})$ to a unique Nash equilibrium.

Regularity Conditions

We next formalize the sequence of models along which we take limits. Similar to (ref), the static model is given by the tuple

equation[equation omitted — 86 chars of source]

where $U(\cdot)$ is the payoff function in (ref), $\lambda(\cdot)$ the equilibrium selection mechanism given in (ref), $V(\cdot)$ the latent index in the network formation model (ref), $\bm{\rho} = (\rho_i)_{i=1}^n$, $\bm{Z} = (Z_i)_{i=1}^n$ with $Z_i$ defined prior to (ref), and $\bm{\zeta}$ is the matrix of $\zeta_{ij}$'s. We can then directly import (ref) using this new notation.

Several regularity conditions are used to establish the CLT. The first imposes a restriction on the network formation model, requiring the linking probability to decay exponentially quickly in spatial distance. Exponential decay is useful for a CLT in the same way that mixing coefficients are usually required to tend to zero with distance at an exponential rate.

assump[Network Formation] \begin{enumerate}[(a)] • There exists $(\bar{\mu},\bar{\mu}')$ such that $V(\delta, \mu, \mu', \zeta) \leq V(\delta, \bar{\mu}, \bar{\mu}', \zeta)$ for any $\delta \in \mathbb{R}_+$, and $(\mu,\mu',\zeta) \in \text{supp}( (\mu_i,\mu_j,\zeta_{ij}) )$. • $\text{supp}(\zeta_{ij}) \subseteq \mathbb{R}$, and for any $\delta,\mu,\mu'$, $V(\delta,\mu,\mu',\cdot)$ is strictly monotonic. • Let $\tilde V^{-1}(r,\cdot)$ be the inverse of $V(r, \bar{\mu}, \bar{\mu}', \cdot)$. There exist $c_1,c_2>0$ such that \begin{equation*} \tilde{\Phi}_\zeta\left( \tilde V^{-1}(r, 0) \right) \leq c_1 e^{-c_2 r}, \end{equation*} where $\tilde{\Phi}_\zeta(\cdot)$ denotes the complementary CDF of $\zeta_{ij}$. \end{enumerate}

Parts (a) and (b) are common regularity conditions; (a) says the $\mu_i$'s have a bounded impact on the latent index $V(\cdot)$. Part (c) is the main restriction, which limits the nonlinearity of $V(\cdot)$ and the tail mass of the distribution of $\zeta_{ij}$. It holds, for example, if $V(\cdot)$ is linear in its arguments and $\zeta_{ij}$ has exponential tails. Note that $\mathbf{P}(A_{ij}=1 \mid \delta_{ij}=r) \leq \tilde{\Phi}_\zeta( \tilde V^{-1}(r, 0))$. Hence, part (c) implies that connection probabilities are exponentially decaying with distance. This is a typical requirement in latent space models breza_using_2017,hrh2002.

remark[Sparsity] Under the increasing domain asymptotics of (ref), the previous assumption implies that the observed network is {\em sparse} in the sense that the expected number of links involving any arbitrary agent is $O(1)$.\footnote{The expected number of links is ${\bf E}[\sum_j A_{ij}] \leq n {\bf E}[A_{ij}] \leq n{\bf E}[\tilde{\Phi}_\zeta( \tilde V^{-1}(\omega_n\lVert\tilde\rho_i-\tilde\rho_j\rVert, 0)] \leq \kappa\bar{f} C \int_{r\geq 0} c_1 e^{-c_2 r} \,\text{d}r = O(1)$ for some constant $C$ that appears from a change of variables to hyperspherical coordinates..} This is a commonly used definition of network sparsity that captures the stylized fact that, in most real world social networks, the number of connections involving the typical agent is small relative to the network size chandrasekhar2016.

For the next assumption, define

equation[equation omitted — 121 chars of source]
assump[Non-Degeneracy] Either $\varphi(p,z;p',z')=0$ for any $(p,z,p',z')$, or for $\Phi^*(\cdot)$ defined in (ref), \begin{equation*} \inf_{(p,z) \in \mathcal{T}} \int_{\mathbb{R}^d} \int_{\mathbb{R}^{d_z}} \varphi(p,z;p',z') \,d\Phi^*(z') \,dp' > 0. \end{equation*}

If $\varphi(p,z;p',z')=0$ everywhere, then this simply means there are no strategic interactions, which is the uninteresting case. To interpret the other case, note that when the network is exogenous, so that $\mathcal{R}_i^c \perp\!\!\!\perp \bm{A}$, then by (ref),

multline*[multline* omitted — 254 chars of source]

Hence, a sufficient condition is for the infimum of the right-hand side to be strictly positive. Note that if interactions exist, then necessarily ${\bf E}[\mathcal{R}_i^c] > 0$, so the sufficient condition strengthens this slightly by requiring it to hold conditional on type, for all types. Additionally, it requires a non-degenerate network in the sense that the limiting expected degree of any agent, conditional on their type, is positive.

For the last regularity condition, let $\bm{\rho}^m = (\rho_i)_{i=1}^{m+2}$, $\{H_n\}_{n\in\mathbb{N}}$ be a sequence of subsets of $\mathbb{R}^d$, and $\psi_i^m(H_n)$ be $i$'s agent statistic under a modified version of model (ref) where instead of $\bm{\rho}$ as our set of positions, we instead have $(\rho_i)_{i=1}^{m+2} \cap H_n$.\footnote{To be clear, the modified model is the same as the original model, except we draw types as follows. First, we generate the set of positions $(\rho_i)_{i=1}^{m+2} \cap H_n$. Then independently across each element $\rho_j$ in this set, we draw $Z_j$ from $\Phi(\cdot \,|\, \rho_j)$. (ref) in the appendix is the same as (ref) but with the formal data-generating process completely spelled out.} The original model is a special case where $m = n-2$ and $H_n = \mathbb{R}^d$ for all $n$.

assump[Bounded Moments] (a) There exists a finite constant $C$ such that ${\bf E}[\lvert\psi_1^m(H_n)\rvert^8 \mid X_1=x, X_2=y] < C$ for any $n$ sufficiently large, $m \in [n/2,3n/2]$, $x,y \in \mathbb{R}^d$, and sequence of subsets $\{H_n\}_{n\in\mathbb{N}}$ of $\mathbb{R}^d$. (b) There exists $c\geq 0$ such that $\lvert\psi_1^m(H_n)\rvert \leq c\, m^c$ a.s.\ for any $m,n\in\mathbb{N}$.

Part (b) says that agent statistics are uniformly bounded by a polynomial in the network size. This is easily satisfied by the moments in (ref) and (ref) if instrument functions are uniformly bounded. Part (a) essentially requires uniformly bounded 8th moments for agent statistics but is slightly stronger for technical reasons. Many of the moments in the applications are uniformly bounded, including (ref) and (ref), in which case part (a) is automatic. For subnetwork moments (used in (ref)), (ref) holds under (ref) by Lemma I.17 of leung2019normal.

Main Result

Before stating the main result, we need some additional notation. Let $\psi_i^{n+1}$ be $i$'s agent statistic in a model with $n+1$ agents. Consider counterfactually removing agent $n+1$ from the model, and let $\bm{A}^-$ and $\bm{\tau}^-$ be the network and type vector with agent $n+1$ removed. Let $\bm{Y}^- = \lambda(\bm{A}^-, \bm{\tau}^-)$ be the equilibrium outcome vector in the resulting model with the remaining $n$ agents. Let $\psi_i^{-(n+1)}$ be $i$'s agent statistic constructed from $(\bm{Y}^-, \bm{A}^-, \bm{\tau}^-)$. Define the {\em add-one cost} $\Xi_n = \psi_{n+1}^{n+1} + \sum_{i=1}^n (\psi_i^{n+1} - \psi_i^{-(n+1)})$, which is the total change in the sum of all agent statistics from adding a single agent $n+1$. Let $\bm{\Sigma}_n = \text{Var}(n^{-1/2} \sum_{i=1}^n \psi_i)$ and $\bm{I}_m$ be the $m\times m$ identity matrix. Recall that $\lambda_\text{min}(\bm{\Sigma}_n)$ denotes the smallest eigenvalue of $\bm{\Sigma}_n$. Finally, recall that agent statistics are vectors of dimension $m$.

theoremSuppose $c'\Xi_n$ is asymptotically non-degenerate for any $c \in \mathbb{R}^m\backslash\{\bm{0}\}$. Also suppose the model satisfies Assumptions (ref) and (ref), weak dependence conditions (Assumptions (ref) and (ref)), and regularity conditions (Assumptions (ref)--(ref)). If $\psi(\cdot)$ satisfies (ref), then as $n\rightarrow\infty$, \begin{equation*} \frac{1}{\sqrt{n}} \sum_{i=1}^n \bm{\Sigma}_n^{-1/2} \left( \psi_i - {\bf E}[\psi_i] \right) \stackrel{d}\longrightarrow \mathcal{N}(\bm{0}, \bm{I}_m) \quadand\quad \liminf_{n\rightarrow\infty} \lambda_min(\bm{\Sigma}_n) > 0 \end{equation*} under the sequence given in (ref).
proofSee (ref) for the formal proof and a proof sketch.

(ref) is discussed in (ref) and (ref). Non-degeneracy of the add-one cost essentially just requires a non-trivial choice of agent statistics. When agent $n+1$ is added to the model, the “direct effect” on the total statistic is that it increases by the value of her agent statistic $\psi_{n+1}^{n+1}$, which should be non-degenerate. However this also has an “indirect effect” on the statistics of other agents $i$ given by $\psi_i^{n+1} - \psi_i^{-(n+1)}$, for example due to strategic interactions. Non-degeneracy would fail, for instance, if the direct and indirect effects exactly cancel, which is unlikely to happen for non-trivial statistics.

CLT for the Dynamic Model

We first formalize the model for the initial conditions $\bm{Y}^0 = (Y_i^0)_{i=1}^n$. This is simply given by the static model in (ref). That is, agents choose their period-0 actions best-responding to the contemporaneous actions of her peers in that period. We view this as an approximation of the long-run outcome of the dynamic model (cf.\ (ref)). Define $\bm{\tau}^0 = (\tau_i^0)_{i=1}^n$. Let $U_0(\cdot)$ be the payoff function in period 0, which may potentially differ from $U(\cdot)$, and let $S_0(\cdot)$ be the analog of $S(\cdot)$ (statistics that capture strategic interactions). Let $\mathcal{E}(\bm{A},\bm{\tau}^0)$ be the set of outcomes $\bm{Y}^0$ such that each component $Y_i^0$ satisfies

equation[equation omitted — 158 chars of source]
assump[Initial Condition] (a) $\lvert\mathcal{E}(\bm{A},\bm{\tau}^0)\rvert \geq 1$. (b) There exists a function $\lambda(\bm{A},\bm{\tau}^0)$ with range $\mathcal{E}(\bm{A},\bm{\tau}^0)$ such that the realized initial condition satisfies $\bm{Y}^0 = \lambda(\bm{A},\bm{\tau}^0)$.

This is just a restatement of (ref) with period-0 types.

Define the {\em add-one cost} analogously to (ref): $\Xi_n = \psi_{n+1}^{n+1} + \sum_{i=1}^n (\psi_i^{n+1} - \psi_i^{-(n+1)})$, which is the total change in the sum of all agent statistics from removing a single agent $n+1$. Also recall that $\bm{\Sigma}_n = \text{Var}(n^{-1/2} \sum_{i=1}^n \psi_i)$ and $\bm{I}_m$ is the $m\times m$ identity matrix.

theoremSuppose $c'\Xi_n$ is asymptotically non-degenerate for any $c \in \mathbb{R}^m\backslash\{\bm{0}\}$. Also suppose \begin{itemize} • the best-response model satisfies (ref), • the initial conditions model defined in (ref) satisfies Assumptions (ref), (ref), (ref), and (ref), using $U_0(\cdot), S_0(\cdot)$ in place of $U(\cdot), S(\cdot)$, and • regularity conditions (Assumptions (ref) and (ref)) hold. \end{itemize} If $\psi(\cdot)$ satisfies (ref), then as $n\rightarrow\infty$, \begin{equation*} \frac{1}{\sqrt{n}} \sum_{i=1}^n \bm{\Sigma}_n^{-1/2} \left( \psi_i - {\bf E}[\psi_i] \right) \stackrel{d}\longrightarrow \mathcal{N}(\bm{0}, \bm{I}_m) \quadand\quad \liminf_{n\rightarrow\infty} \lambda_min(\bm{\Sigma}_n) > 0 \end{equation*} under the sequence given in (ref).
proofSee (ref).

The assumptions imposed on the initial conditions model are the same as those imposed on the static model for a CLT. Primitive sufficient conditions for (ref) are discussed following its statement in (ref). As in the static case, non-degeneracy of the add-one cost essentially just requires that the choice of $\psi(\cdot)$ is nontrivial; see the discussion following the statement of (ref).

Numerical Illustrations

This section illustrates the performance of the HAC estimators in an empirical application and simulation study.

Empirical Application

We revisit the study of conley2010learning on the diffusion of an agricultural technology through a network of farmers. They show that as farmers learn to use a new technology (pineapple production), they respond to information from their social contacts. In particular, they find that when the returns from a network neighbor's crop yields are above or below expectation, they adjust their agricultural inputs accordingly. Their main specifications are logistic models of the probability that a farmer's agricultural inputs change in response to the shares of network neighbors with above- or below-expectation yields in the previous period.

We replicate Table 4, column B of their paper. The outcome is an indicator for whether farmer $i$ adjusts her inputs in period $t$. There are four main regressors of interest comprising $S_i^t$, which are the share of good / bad news events from network neighbors who choose the same / different inputs as farmer $i$ last period. We include the same controls in their specification, including a full set of village and planting round dummies.

The authors use a spatial HAC estimator out of concern for potential spatial autocorrelation in the errors. Our theoretical results show that the estimator also accounts for forms of network autocorrelation. If spatial positions are unavailable, our theory shows that the network HAC can be used instead.

(ref) reports point estimates and spatial and network HAC standard errors for a range of bandwidths using the Bartlett kernel. The four main regressors of interest correspond to the first four rows in the table. The remaining rows are control variables. conley2010learning use a bandwidth of 1500 and say that the results are similar for 1000 and 2000. This is replicated in our table. The path distances used in the network HAC are obtained from the data on farmers' advice networks used to construct the main regressors. Across the three villages, average path distance ranges from about 0.5 to 2 and the maximum distance from 3 to 8, so we report bandwidth values in the range of 0 to 3.

From the table, we see that the results are not generally sensitive to bandwidth choice for positive values of the bandwidth. The table shows that spatial and network standard errors are similar, which is consistent with our theory. For comparison, we report a bandwidth of zero, which corresponds to i.i.d.\ standard errors. Depending on the coefficient, these standard errors can be conservative or anti-conservative relative to the HAC standard errors.

table[table omitted — 3,503 chars of source]

Monte Carlo

We simulate a dynamic probit model with two periods and period-1 payoffs given by

equation*[equation* omitted — 126 chars of source]

where $\nu_i^t \sim \mathcal{N}(0,1)$. To estimate the model, we use a standard probit regression. As discussed in (ref), to account for autocorrelation in the errors and regressors, we estimate the variance using the sandwich formula $H^{-1}SH^{-1}$, where $H$ is the Hessian and $S$ is a HAC estimate of the variance of the scores.

We are faced with the task of generating network autocorrelation in the errors while still maintaining the assumption of standard normal marginals, so that the probit can be used for estimation. Toward this end, we use the following modification of (ref):

equation*[equation* omitted — 123 chars of source]

where $\varepsilon_i^t \stackrel{iid}\sim \mathcal{N}(0,1)$ and $\omega_i$ is a constant. The first term is the average error of network neighbors, which captures exogenous peer effects in unobservables and generates autocorrelation. We choose the weight to ensure standard normal marginals for $\nu_i^t$: $\omega_i = (1 + (\sum_j A_{ij})^{-1})^{-1/2}$.

Let covariates follow an AR(1) model $X_i^1 = 0.5 X_i^0 + u_i^t$, where $u_i^t \stackrel{iid}\sim \mathcal{N}(0,1)$ and $X_i^0 \stackrel{iid}\sim \text{exp}(1)$. We set the true $\beta$ to $(0.5, -0.3, 1)$. Initial conditions $Y_i^0$ are given by the same outcome model as period 1, except we set $\beta_3$ to zero and replace period-1 types $(X_i^1,\varepsilon_i^1)$ with period-0 types $(X_i^0,\varepsilon_i^0)$. To generate the network, we use the random geometric graph model

equation*[equation* omitted — 95 chars of source]

Recall from (ref) that positions are given by $\rho_i = \omega_n \tilde\rho_i$. We draw $\tilde\rho_i \stackrel{iid}\sim \mathcal{U}([0,1]^2)$ and set the scaling factor $\omega_n$ to $(n / (5\pi) )^{1/2}$. This choice of $\omega_n$ ensures that the limiting expected degree is five, which generates a giant component penrose2003. Additionally, this model clearly induces geographic homophily, and as a consequence $\bm{A}$ typically has a high clustering coefficient. All of these are well-known to be realistic network features barabasi2015.

We compute standard errors four different ways. The first three use HAC estimators, all with the Bartlett kernel. For the standard spatial HAC, we set $h_n = n^{1/(3d)}$ (here $d=2$), and for the network HAC, $h_n = \log n$. The third uses the generalized spatial HAC from (ref), for which we set $h_n = n^{1/(4d)}$. For $\hat\theta(x)$ in the generalized spatial HAC, we use a kernel estimator and choose $b_n = (\log n/n)^{1/(2p+d)}$ for smoothness $p=d/2+1$, which achieves a $n^{-1/4}$ uniform convergence rate in the i.i.d.\ case. The fourth way of computing standard errors, which we term “oracle” standard errors, is by simulating the true variance of the probit estimator using 7500 simulation draws. This is used to assess the quality of the HAC estimators and our CLT.

We report results for $n = 500, 1\text{k}, 2\text{k}$. The point estimates are all extremely close to the truth on average for all sample sizes. (ref) displays standard errors, averaged across 15,000 simulations. Column “GS HAC” corresponds to the generalized spatial HAC. The table shows that all HAC standard errors are quite close on average to the target oracle standard errors. They can be slightly anti-conservative in smaller samples, which is a common issue with HAC estimators.

(ref) reports rejection percentages from two-sided $t$-tests at the 5 percent level of the null hypothesis that $\beta$ equals its true value, using the previous standard errors. The oracle standard errors control size well across all sample sizes, which illustrates the accuracy of the normal approximation. For the spatial and network HAC, rejection rates are fairly close to the nominal level, although there is a tendency toward overrejection. The network HAC happens to perform the best, while the generalized spatial HAC has the greatest tendency to overreject. The spatial HAC has better performance than the latter, at least when first-order stationarity holds. However, the story will be different when stationarity fails, as we will see below. While the network and generalized spatial HAC are not guaranteed to be positive definite in finite samples, they are positive definite in all of our simulation draws. Finally, the third column, labeled “Naive,” reports $t$-tests computed using i.i.d.\ standard errors, and we see that rejection rates can be many times higher than the nominal level.

table[table omitted — 1,036 chars of source]
table[table omitted — 1,246 chars of source]

The moments used in pseudo-maximum likelihood are centered, since they are conditionally mean zero. Hence, by Theorems (ref) and (ref), the HAC estimators are asymptotically exact, as confirmed by the probit results. We next report results for the weighted average period-1 outcome $n^{-1} \sum_i Y_i^1 X_i^1$, which is generated from the same data-generating process as the previous tables. This is clearly not centered and not first-order stationary. However, an asymptotic analog of (ref) does hold by (ref). Thus, our results predict that the spatial HAC is conservative while the generalized spatial HAC is exact. We have no predictions for the network HAC.

(ref) displays standard errors and rejection percentages for the level-5% two-sided $t$-test of the null that ${\bf E}[Y_i^1]$ equals its true value. The true value is computed from 7500 simulation draws, and the values displayed in the table are averages over 15,000 separate simulation draws. “Oracle” standard errors are obtained by simulating the variance of $n^{-1} \sum_i Y_i^1X_i^1$ from 7500 separate simulation draws.

From the table, we see that both the spatial and network HAC SEs are quite conservative, being about twice as large as the oracle SEs. In contrast, the generalized spatial HAC SEs produce rejection percentages close to the nominal level. The network HAC is also asymptotically conservative and comparable to the spatial HAC. This is interesting because network first-order stationarity ((ref)) does not apparently hold, so this is a case not covered by our theory.

table[table omitted — 952 chars of source]

Conclusion

The increasing availability of network data has enabled researchers to better study heterogeneous interactions and subject to empirical inquiry a large body of theoretical work on the impact of network topology on the diffusion of behaviors, products, and information jackson2017economic. The central challenge of drawing valid inference from network data is that it typically consists of autocorrelated observations from a single large network. This is an inherent problem in models with social interactions, since by definition, an agent's outcome is a function of her neighbors'. Consequently, standard procedures that assume i.i.d.\ data can produce highly misleading inference.

In the context of discrete choice models, we show that spatial and network HAC estimators can be used to correct for network autocorrelation. In an empirical application to conley2010learning, the spatial and network HAC estimators deliver similar results, which supports the theory. Our simulation study illustrates the potential for naive i.i.d.\ standard errors to be highly anti-conservative and demonstrates the accuracy of the normal approximation and HAC estimators.

Our main technical results are CLTs for static and dynamic models of social interactions and consistency results for the HAC estimators. In the existing literature, first-order stationarity conditions are always required for HAC estimators to approximate the variance. We also provide results for the spatial HAC in settings where stationarity fails to hold, which include models defined by moment inequalities and many nonparametric models. For the network HAC, its validity when first-order stationarity fails is an open question. The advantage of the network HAC is that it does not require additional information on spatial/social positions. On the other hand, for it to be valid in static models, we require the network to be exogenous (orthogonal to the errors in the outcome model), whereas no such assumption is needed for the spatial HAC. When exogeneity fails, the dependence structure of the data conditional on the network is difficult to characterize, and it is an open question whether the network HAC is valid in this setting.

\FloatBarrier \phantomsection \addcontentsline{toc}{section}{References}

\part{Supplemental Appendix}