Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
124,478 characters · 22 sections · 144 citation commands
Identifying Network Ties from Panel Data: Theory and an Application to Tax Competition
In many economic environments, behavior is shaped by social interactions between agents. In individual decision problems, social interactions have been key to understanding outcomes as diverse as educational test scores, the demand for financial assets, and technology adoption ( Sacerdote2001; Bursztynetal2014; ConleyUdry2010 ). In macroeconomics, the structure of firms' production and credit networks propagate shocks, or help firms to learn (Acemogluetal2012; Chaney2014). In political economy and public economics, ties between jurisdictions are key to understanding tax setting behavior (Tiebout1956; Shleifer1985; BesleyCase1994).
Underpinning all these bodies of research is some measurement of the underlying social ties between agents. However, information on social ties does not exist in most publicly available and widely used datasets. To overcome this limitation, studies of social interaction either postulate ties based on common observables or homophily, or elicit data on networks. However, it is increasingly recognized that postulated and elicited networks remain imperfect solutions to the fundamental problem of missing data on social ties, because of econometric concerns that arise with either method, or simply because of the cost of collecting network data. \footnote{ As detailed in DePaula2017, elicited networks are often self-reported and can introduce error to the outcome of interest. Network data can be censored if only a limited number of links can feasibly be reported. Incomplete survey coverage of nodes in a network may lead to biased aggregate network statistics. ChandrasekharLewis2016 show that even when nodes are randomly sampled from a network, partial sampling leads to non-classical measurement error and biased estimation. Collecting social network data is also a time- and resource-intensive process. In response to these concerns, a nascent strand of literature explores cost-effective alternatives to full elicitation to recover aggregate network statistics (Brezaetal2017).}
Two consequences are that (i) the classes of problems in which social interactions occur are understudied, because social networks data is missing or too costly to collect; and (ii) there is no way to validate social interactions analysis in contexts where ties are postulated. In this paper, we tackle this challenge by deriving sufficient conditions under which global identification of the entire structure of social networks\ is obtained, using only observational panel data that itself contains no information on network ties. Our identification results allow the study of social interactions without data on social networks, and the validation of structures of social interaction where social ties have hitherto been postulated. The recovered networks are economically meaningful to explain the effects under study, since they are entirely estimated from the data itself, and not driven by ex ante assumptions on how individuals interact.
A researcher is assumed to have panel data on individuals $i=1,...,N$ for instances $t=1,...,T$. An instance refers to a specific observation for $i$ and need not correspond to a time period (for example, if $i$ refers to a firm, $t$ could refer to market $t$). The outcome of interest for individual $i$ in instance $t$ is $y_{it}$ and is generated according to a canonical structural model of social interactions:\footnote{Blumeetal2015 present micro-foundations for this estimating equation based on non-cooperative games of incomplete information for individual choice problems.}
Outcome $y_{it}$ depends on the outcomes of other individuals to whom $i$ is socially tied, $y_{jt}$, and $x_{jt}$ includes characteristics of those individuals.\footnote{ In the case in which $t$ is considered to be a time period, $x_{it}$ may also include lagged values of $y_{it}$.} $W_{0,ij}$ measures how the outcome and characteristics of $j$ causally impact the outcome for $i$. The network is initially assumed to be fixed over time, and we later provide an extension of the method for time-varying networks. As outcomes for all individuals obey equations analogous to ((ref)), the system of equations can be written in matrix notation, where the structure of interactions is captured by the adjacency matrix, denoted by $W_{0}$. Our approach allows for unobserved heterogeneity across individuals $\alpha _{i}$ and common shocks to individuals $\alpha _{t}$. This framework encompasses\ a classic linear-in-means specification as in Manski1993. In his terminology, $ \rho _{0}$ and $\gamma _{0}$ capture endogenous and exogenous social effects, and $\alpha _{t}$ captures correlated effects. The distinction between endogenous and exogenous peer effects is critical, as only the former generates social multiplier effects. In line with the literature, we maintain that the same $W_{0}$ governs the structure of both endogenous and exogenous effects. We later discuss relaxing this assumption when more than one regressor is used.
Manski's seminal contribution set out the reflection problem of separately identifying endogenous, exogenous, and correlated effects in linear models. However, it has been somewhat overlooked that he also set out another challenge in the identification of the social network in the first place. \footnote{Manski1993 highlights difficulties (and potential restrictions) in identifying $\rho _{0},\beta _{0}$ and $\gamma _{0}$ when all individuals interact with each other, and when this is observed by the researcher. In ((ref)), this corresponds to $ W_{0,ij}=N^{-1} $, for $i,j=1,\dots ,N$. At the same time, he states (p. 536), \textquotedblleft I have presumed that researchers know how individuals form reference groups and that individuals correctly perceive the mean outcomes experienced by their supposed reference groups. There is substantial reason to question these assumptions (...) If researchers do not know how individuals form reference groups and perceive reference-group outcomes, then it is reasonable to ask whether observed behavior can be used to infer these unknowns (...) The conclusion to be drawn is that informed specification of reference groups is a necessary prelude to analysis of social effects.\textquotedblright} This is the problem we tackle, and thus, we expand the scope of identification beyond $\rho _{0}$, $\beta_{0}$, and $\gamma_{0}$. Our point of departure from much of the literature is therefore to presume $W_{0}$ is entirely unknown to the researcher. We derive sufficient conditions under which all the entries in $W_{0}$, and the endogenous and exogenous social effect parameters, $\rho _{0}$ and $\gamma _{0},$ are globally identified from \textquotedblleft reduced form\textquotedblright\ parameters. By identifying the social interactions matrix $W_{0}$, our results allow the recovery of aggregate network characteristics, such as the degree distribution and patterns of homophily, as well as node-level statistics such as the strength of social interactions between nodes, and the centrality of nodes. Such aggregate and node-level statistics often map back to underlying models of social interaction ( Ballesteretal2006; Jacksonetal2017; DePaula2017 ).
Our identification strategy is new and fundamentally different from those employed elsewhere in the literature and does not rely on requirements about network sparsity. However, it delivers sufficient conditions that are mild and relate to existing results on the identification of social effects parameters when $W_{0}$ is known (Bramoulleetal2009; DeGiorgietal2010; Blumeetal2015). The intuition for our identification result is simple: model (1) has $N^{2}$ reduced-form parameters, and there are $N(N-1)+3$ structural unknowns (as no unit affects itself, so $W_{0,ii}=0$). So there are more equations than unknowns if $ N\geq 2$, and we demonstrate those can be solved for the parameters of interest under the assumptions we invoke. Our identification result is also useful in other estimation contexts, such as when a researcher has partial knowledge of $W_{0}$,\footnote{ One such example is the nascent literature of Aggregate Relational Data (ARD) as in Brezaetal2017. Another possibility is that individuals are known to belong to subgroups, so $W_{0}$ is block diagonal. } or in navigating between priors on reduced-form and structural parameters in a Bayesian framework (see, e.g., gefanghalltavlas2023), thus avoiding issues the raised by KlineTamer2016.
Global identification is a necessary requirement for consistency of extremum estimators such as those based on the GMM (Hansen 1982; Newey and McFadden 1994). Our identification analysis provides primitives for this condition. To estimate the model, we employ the Adaptive Elastic Net GMM method ( CanerZhang2014), as this allows us to deal with a potentially high-dimensional parameter vector (in comparison to the time dimension in the data) including all the entries of the social interactions matrix $W_{0}$ , although other estimation protocols may also be entertained (e.g. using Bayesian methods or a priori information).\footnote{The Elastic Net was introduced by zou2005regularization in part to circumvent difficulties faced by alternative estimation protocols (e.g., LASSO) when the number of parameters, $p$, exceeds the number of observations, $n$ (where $p$ and $n$ follow the notation in that paper). Whereas the theoretical results on the large-sample properties of elastic net estimators usually have not exploited sparsity, several articles have demonstrated their performance in data scenarios where this occurs. In Section 3, we provide an informal discussion on the performance in our context. }
We showcase the method using Monte Carlo simulations based on stylized random network structures as well as real-world networks. In each case, we take a fixed network structure $W_{0}$ and simulate panel data as if the data generating process were given by ((ref)). We then apply the method to the simulated panel data to recover estimates of all elements in $W_{0}$, as well as the endogenous and exogenous social effect parameters ($\rho _{0}$, $\gamma _{0}$). The networks considered vary in size, complexity, and their aggregate and node-level features. In small samples, we find that the majority of links are identified even for $T=5$, and the proportion of true non-links (zeros in $W_{0}$) captured correctly as zeros is over $85$% even when $T=5$. Of course, there are important limitations to the use of the method in small-$T$ cases. Biases are expected and manifest themselves in two ways. First, weak links can be shrunk to zero, and the strength of strong edges can be overestimated. Second, the estimates of $\rho$ and $\gamma$ can suffer from small-sample bias, being analogous to well-known results for autoregressive time series models. Both properties rapidly improve with $T$. For instance, biases in the estimation of endogenous and exogenous effects parameters ($\hat{\rho},$ $\hat{\gamma}$) fall quickly with $T$ and are close to zero for large sample sizes. The endogenous and exogenous social effects are also correctly captured as $T$ increases. A fortiori, we estimate aggregate and node-level statistics of each network, demonstrating the accurate recovery of key players in networks, for example.
In the final part of our analysis, we apply the method to shed new light on a classic real-world social interactions problem:\ tax competition between US\ states. The literatures in political economy and public economics have long recognized the behavior of state governors might be influenced by decisions made in “neighboring” states. The typical empirical approach has been to postulate the relevant neighbors as being geographically contiguous states. Our approach allows us to infer the set of “economic” neighbors determining social interactions in tax setting behavior from panel data on outcomes and covariates alone. In this application, the panel data dimensions cover mainland US states, $N=48$, for the years 1962-2015, $T=53$.
The identified network structure of tax competition differs markedly from the assumption of competition between geographic neighbors. The identified economic network has fewer edges, and we identify non-adjacent states that influence tax setting behaviors. Differences in the structure of the identified economic and geography-based networks are reflected in\ the far lower clustering coefficient in the former ($.042$ versus $.419$). With the recovered social interactions matrix we establish, beyond geography, which covariates correlate to the existence of ties between states and so shed new light on hypotheses for social interactions in tax setting: factor mobility and yardstick competition (Tiebout1956; Shleifer1985; BesleyCase1994). The identified network highlights significant predictors of tax competition between states beyond distance:\ political homophily reduces the likelihood of a link, suggesting any yardstick competition driving social interactions occurs when voters compare their governor to those of the opposing party in other states. Tax haven states appear to be less influential in tax setting behaviors, easing concerns over a race-to-the-bottom in tax setting. Labor mobility between states does not robustly predict the existence of economic ties between states in tax setting behavior.
Given the relatively long study period in this application, at a final stage of analysis we extend our method to allow the strength of social interactions in tax competition ($\rho _{0}$, $\gamma _{0}$) and the structure of links in the economic network ($W_{0}$) to vary over time as we change the weight placed on observations from any given time period. We document the gradual increase in strength of social interactions over time, and the changing nature of the network of interactions. We utilize these findings to conduct counterfactual simulations of the general equilibrium propagation of tax shocks from a given state to all other mainland US\ states, and how these general equilibrium effects of the same policy shock vary as we place weight on observations later in our study period.
Our paper contributes to the literature on the identification of social interactions models. The first generation of papers studied the case where $ W_{0}$ is known, so only the endogenous and exogenous social effects parameters needed to be identified. It is now established that if the known $ W_{0}$ differs from the linear-in-means example where all units are linked with equal weights, $\rho _{0}$ and $\gamma _{0}$ can be identified ( Bramoulleetal2009; DeGiorgietal2010). Intuitively, identification in those cases can use peers-of-peers, are not necessarily connected to individual $i$ and can be used to leverage variation from exclusion restrictions in ((ref)), or can use groups of different sizes within which all individuals interact with each other (Lee2007). Bramoulleetal2009 show these conditions are met if $I,$ $W_{0}$, and $W_{0}^{2}$ are linearly independent, which is shown to hold generically by Blumeetal2015. However, as made precise in Section (ref), the linear algebraic arguments employed by Bramoulleetal2009 or Blumeetal2015 do not apply when $W_{0}$ is unobserved, and other arguments have to be used instead.\footnote{ Alternative identification approaches when $W_{0}$ is known focus on higher moments (variances and covariances across individuals) of outcomes ( DePaula2017) and rely on additional restrictions on higher moments of $\epsilon _{it}$. Note that ((ref)) is a spatial autoregressive model. In that literature, $W_{0}$ is also typically assumed to be known (Anselin2010).}
Blumeetal2015 investigate the case when $W_{0}$ is partially observed and show that if two individuals are known not to be directly connected, the parameters of interest in a model related to ((ref)) can be identified. Blume et al. (2011) take an alternative approach: suggesting a parameterization of $W_{0}$ according to a pre-specified distance between nodes. We do not impose such restrictions, but note that partial observability of $W_{0}$ or placing additional structure on $W_{0}$ is complementary to our approach, as it reduces the number of parameters in $W_{0}$ to be retrieved. BonaldiHortacsuKastl2015 and Manresa2016 estimate models like ((ref)) when $ W_{0} $ is not observed, but where $\rho _{0}$ is set to zero so there are no endogenous social effects. They use sparsity-inducing methods from the statistics literature, but the presence of $\rho _{0}$ in our case complicates identification because it introduces issues of simultaneity that we address.\footnote{Manresa2016 allows for unit-specific $\beta _{0} $ parameters. While in many applications those are taken to be homogeneous, we also discuss extensions on how heterogeneity in those parameters can be handled when $\rho _{0}\neq 0$ in Appendix B.}
Rose2015 also presents related identification results for linear models like ((ref)), assuming the sparsity of the neighborhood structure. Intuitively, given two observationally equivalent systems, sparsity guarantees the existence of pairs that are not connected in either. Since observationally equivalent systems are linked via the reduced-form coefficient matrix, this pair allows one to identify certain parameters in the model. Having identified those parameters, Rose2015 shows that one can proceed to identify other aspects of the structure (see also GautierRose2016 ). This is related to the ideas in Blume et al. (2015), who show identification results can be leveraged if individuals are known not to be connected. Our main identification results do $not$ rely on properties of sparse networks, and make use of plausible and intuitive conditions, whereas the auxiliary rank conditions necessary may be computationally complex to verify. \textcolor{black}{More recently, Lewbeletal2019 propose an estimation strategy for the parameters $\rho_0$, $\beta_0$, and $\gamma_0$ of model (ref) in the absence of network links if many different groups can be observed.} \textcolor{black}{Battaglinietal2019 estimate a structural model specifically for the case of unobserved social connections in the US Congress.}
Finally, in the statistics literature, LamSouza2019 study the penalized estimation of model (ref) when $W_{0}$ is not observed, assuming the model and social interactions are identified. The statistical literature on graphical models has investigated the estimation of neighborhoods defined by the covariance structure of the random variables at hand (MeinshausenBuhlmann2006). This corresponds to a model where $y_{t}=(I-\rho _{0}W_{0})^{-1}\epsilon _{t}$ is jointly normal (abstracting from covariates). On a graph with $N$ nodes corresponding to the variables in the model, an edge between two nodes (variables) $i$ and $j$ is absent when these two variables are conditionally independent given the other nodes. In the model above, the inverse covariance matrix is $(I-\rho _{0}W_{0})^{\top }\Sigma _{\epsilon }^{-1}(I-\rho _{0}W_{0})$, where $\Sigma _{\epsilon }$ is the variance covariance structure for $\epsilon _{t}$. The discovery of zero entries in this matrix is not equivalent to the identification of $W_{0}$ and involves $\Sigma _{\epsilon }$ (as do identification strategies using higher moments when $W_{0}$ is known).\footnote{ MeinshausenBuhlmann2006's and LamSouza2019's neighborhood estimates rely on (penalized) regressions of $y_{it}$ on $y_{1t},\dots ,y_{i-1,t},y_{i+1,t},\dots ,y_{N,t}$, which do not address the endogeneity in estimating $W_{0}$.}
We build on these papers by studying the problem where $W_{0}$ is potentially entirely unknown to the researcher. In so doing, we open up the study of social interactions to realms where social network data does not exist. In our case, we consider the definition of the network as the one that mediates, together with the variables $x_{it}$, the outcome process $ y_{it}$ according to Equation (ref). The identified network may be a combination of elicited types of social interactions -- such as friendship formation, lending and borrowing relations, links with relatives -- or different from elicited data, as long as the links are relevant in determining the outcomes. In our case, and in line with the literature, the network ties $W_{ij}$ are considered to be deterministic parameters or predetermined. Alternatively, the networks are assumed to be the outcome of a stochastic process, such as the latent space model ( hoff2002latent; Brezaetal2017) or Exponential Random Graphs models (HollandLeinhardt1981).
Our conclusions discuss how our approach can be modified, and assumptions weakened, to integrate partial knowledge of $W_{0}$. We discuss further applications and the steps required to simultaneously identify models of network formation and the structure of social interactions. The practical use of our proposed method has already been demonstrated in applications. For example, fetzeretal2020 study the impact on conflict of the transition of security responsibilities between international and Afghan forces. Our proposed method is used to control for violation of SUTVA-type hypotheses that might occur because of spillover and displacement effects of insurgent forces across districts. Since the pattern of displacement is unobserved -- and, in fact, insurgents have incentives to obfuscate their strategy -- the current method is applied to fully recover the network and bound the effects of the end of the military occupation on conflict. \footnote{Zhou2019 applies our identification results, focusing on unobserved networks with grouped heterogeneity, to suggest a nonlinear least squares procedure for estimation on a single network observation. }
We proceed as follows. Section 2 presents our core result:\ the sufficient conditions under which the social interactions matrix, endogenous and exogenous social effects are globally identified. Section 3 describes the high-dimensional techniques used for estimation based on the Adaptive Elastic Net GMM method and presents simulation results from stylized and real-world networks. Section 4 applies our methods to study tax competition between US\ states. Section 5 concludes. The Appendix provides proofs and further details on estimation and simulations.
Consider a researcher with panel data covering $i=1,\dots ,N$ individuals repeatedly observed over $t=1,\dots ,T$ instances. The number of individuals $N$ in the network is fixed but potentially large. The aim is to use this data to identify a social interactions model with no data on actual social ties. For expositional ease, we first consider identification in a simpler version of the canonical model in ((ref)), where we drop individual-specific ($\alpha _{i}$) and time-constant fixed effects ($\alpha _{t}$) and assume $x_{it}$ is a one-dimensional regressor for individual $i$ and instance $t$. We later extend the analysis to include individual-specific, time-constant fixed effects and allow for multidimensional covariates $x_{k,it}$, $k=1,\dots ,K$. We adopt the subscript \textquotedblleft 0\textquotedblright\ to denote parameters generating the data, and non-subscripted parameters are generic values in the parameter space:
As the outcomes for all individuals $i=1,\dots ,N$ obey equations analogous to ( (ref)), the system of equations can be more compactly written in matrix notation as:
The vector of outcomes $y_{t}=(y_{1t},\dots ,y_{Nt})^{\prime }$ assembles the individual outcomes in instance $t$; the vector $x_{t}$ does the same with individual characteristics. $y_{t}$, $x_{t}$, and $\epsilon _{t}$ have dimension $N\times 1$, the social interactions matrix $W_{0}$ is $N\times N$ , and $\rho _{0}$, $\beta _{0}$, and $\gamma _{0}$ are scalar parameters. We do not make any distributional assumptions on $\epsilon _{t}$ beyond $ \mathbb{E}(\epsilon _{t}|x_{t})=0$ (or $\mathbb{E}(\epsilon _{t}|z_{t})=0$ for an appropriate instrumental variable $z_{t}$ if $x_{t}$ is endogenous). We assume the network structure is predetermined and constant, and that the number of individuals $N$ is fixed and repeated. In reality, networks may evolve over time. We thus later expand the method for dynamic network cases. The network structure $W_{0}$ is a parameter to be identified and estimated.
The social interaction model (ref) has been widely studied ( Manski1993; Manresa2016; and Blumeetal2015, among many others), but it is also restrictive in at least two senses. First, we consider $W_{ij}$ to be fixed and predetermined, and not through models of strategic network formation (jackson1996strategic; dePaulaRichardsTamer2018) or of stochastic nature, as in the class of Exponential Random Graphs (HollandLeinhardt1981) or Latent Distance models (hoff2002latent; Brezaetal2017). If there is feedback between outcome determination and link formation, and especially if this involves unobservables, it would be important to model network formation more explicitly.
A regression of outcomes on covariates corresponds, then, to the reduced form for ((ref)),
with $\Pi _{0}=(I-\rho _{0}W_{0})^{-1}(\beta _{0}I+\gamma _{0}W_{0})$ and $ \nu _{t}\equiv (I-\rho _{0}W_{0})^{-1}\epsilon _{t}$. If $W_{0}$ is observed, Bramoulleetal2009 note that a structure $ (\rho ,\beta ,\gamma )$ that is observationally equivalent to $(\rho _{0},\beta _{0},\gamma _{0})$ is such that $(I-\rho _{0}W_{0})^{-1}(\beta _{0}I+\gamma _{0}W_{0})=(I-\rho W_{0})^{-1}(\beta I+\gamma W_{0})$. This can be written as a linear equation in $I,W_{0}$, and $W_{0}^{2}$, and identification is established if those matrices are linearly independent. If $W_{0}$ is not observed, the putative unobserved structure comprises $W_{0}$, and an observationally equivalent parameter vector will instead satisfy $ (I-\rho _{0}W_{0})^{-1}(\beta _{0}I+\gamma _{0}W_{0})=(I-\rho W)^{-1}(\beta I+\gamma W)$. Following the strategy in Bramoulleetal2009 would lead to an equation in $I,W,W_{0}$, and $WW_{0}$, so the insights obtained in that paper do not carry over to the case we study when $W_{0}$ is unknown.
We establish identification of the structural parameters of the model, including the social interactions matrix $W_{0}$, from the coefficients matrix $\Pi _{0}$. Without data on the network $W_{0}$, we treat it as an additional parameter in an otherwise standard model relating outcomes and covariates. Our identification strategy relies on how changes in covariates $ x_{it}$ reverberate through the system and impact $y_{it}$, as well as outcomes for other individuals. These are summarized by the entries of the coefficient matrix $\Pi _{0}$, which, in turn, encode information about $ W_{0}$ and $(\rho _{0},\beta _{0},\gamma _{0})$. A non-zero partial effect of $ x_{it}$ on $y_{jt}$ indicates the existence of direct or indirect links between $i$ and $j$. When $\rho _{0}=0$ (and $\Pi _{0}=\beta _{0}I+\gamma _{0}W_{0}$), only direct links produce such a correlation. When $\rho \neq 0$, both direct and indirect connections may generate a non-zero response, but distant connections will lead to a lower response. Our results formally determine sufficient conditions to precisely disentangle these forces.
We set out six assumptions underpinning our main identification results. Three of these are entirely standard. A fourth is a normalization required to separately identify ($\rho _{0}$, $\gamma _{0}$) from $W_{0}$, and the fifth is closely related to known results on the identification of ($\rho _{0}$, $\gamma _{0}$) when $W_{0}$ is known (Bramoulleetal2009). The sixth assumption pertains to the relation between the nature of repeated multiple observations of the outcome and covariates and restrictions on the stability of $W$. These Assumptions (A1-A6) deliver an identified set of up to two points.
Our first assumption explicitly states that no individuals affect themselves and is a standard condition in social interaction models:
Assumption (A1) rules out applications with self-influence. For example, Input-Output matrices typically feature $(W_{0})_{ii}>0$, as firms tend to source from other firms in the same industry. With Assumption (A1), we can omit elements on the diagonal of $W_{0}$ from the parameter space. We thus can denote a generic parameter vector as $\theta =\left( W_{12},\dots ,W_{N,N-1},\rho ,\gamma ,\beta \right) ^{\prime }\in \mathbb{R}^{m}$, where $ m=N\left( N-1\right) +3$, and $W_{ij}$ is the $(i,j)$-th element of $W$. Reduced-form parameters can be tied back to the structural model ((ref)) by letting $\Pi :\mathbb{R}^{m}\rightarrow \mathbb{R}^{N^{2}}$ define the relation between structural and reduced-form parameters:
where $\theta \in \mathbb{R}^{m}$, and $\Pi _{0}\equiv \Pi (\theta _{0})$.
As $\epsilon _{t}$ (and, consequently, $\nu _{t}$) is mean-independent from $ x_{t}$, $\mathbb{E}[\epsilon _{t}|x_{t}]=0$, the matrix $\Pi _{0}$ can be identified as the linear projection of $y_{t}$ on $x_{t}$. We do not impose additional distributional assumptions on the disturbance term, except for conditions that allow us to identify the reduced-form parameters in (ref). If $x_{t}$ is endogenous, i.e., $\mathbb{E}[\epsilon _{t}|x_{t}]\neq 0$, a vector of instrumental variables $z_{t}$ may still be used to identify $\Pi _{0}$. In either case, identification of $\Pi _{0}$ requires variation of the regressor across individuals $i$ and through instances $t$. In other words, either $\mathbb{E}[x_{t}x_{t}^{\prime }]$ (if exogeneity holds) or $\mathbb{E}[x_{t}z_{t}^{\prime }]$ (otherwise) is full-rank.
Our next assumption controls the propagation of shocks and guarantees that they die as they reverberate through the network. This provides adequate stability and is related to the concept of stationarity in network models. It implies the maximum eigenvalue norm of $\rho _{0}W_{0}$ is less than one and ensures $(I-\rho _{0}W_{0})$ is a non-singular matrix. As the variance of $y_{t}$ exists, the transformation $\Pi (\theta _{0})$ is well-defined, and the Neumann expansion $(I-\rho _{0}W_{0})^{-1}=\sum_{j=0}^{\infty }(\rho _{0}W_{0})^{j}$ is appropriate.
We next assume that network effects do not cancel out, another standard assumption. As we will show, this assumption rules out the pathological case in which endogenous and exogenous effects exactly cancel each other out:
The need for this assumption can be shown by expanding the expression for $\Pi (\theta _{0})$, which is possible by (A2):
If Assumption (A3) were violated, $\beta _{0}\rho _{0}+\gamma _{0}=0$ and $ \Pi _{0}=\beta _{0}I$, so the endogenous and exogenous effects would balance each other out, and network effects would be altogether eliminated in the reduced form. \footnote{ \textcolor{black}{One important case is when networks do not determine outcomes, which we interpret as $\rho_0=\gamma_0=0$ or with $W_0$ representing the empty network. From equation (ref), it is clear that if $\Pi(\theta_0)$ is not diagonal with constant entries, then it must be that $(\rho_0\beta_0+\gamma_0)\neq0$, which implies that $\rho_0\neq0$ or $\gamma_0\neq0$, and also that $W_0$ is non-empty. Taken together, this suggests that the observation that $\Pi(\theta_0)$ is not diagonal is sufficient to ensure that network effects are present and Assumption (A3) is not violated.}}
Identification of the social effects parameters $(\rho _{0},\gamma _{0})$ requires that at least one row of $W_{0}$ adds to a fixed and known number. Otherwise, $\rho _{0}$ and $\gamma _{0}$ cannot be separately identified from $W_{0}$. Clearly, no such condition would be required if $W_{0}$ were observed.
Letting $W_{y}\equiv \rho _{0}W_{0}$ and $W_{x}\equiv \gamma _{0}W_{0}$ denote the matrices that summarize the influence of peers' outcomes (the endogenous social effects) and characteristics on one's outcome (the exogenous social effects), respectively, the assumption above can be seen as a normalization. In this case, $\rho _{0}$ and $\gamma _{0}$ represent the row-sum for individual $i$ in $W_{y}$ and $W_{x}$, respectively.\footnote{ \textcolor{black}{Alternatively, one could normalize $\rho^*=1$ and rescale the network accordingly. In this case, $W^*=\rho_0W_0$ would be identified instead. Also, $W_x=\frac{\gamma_0}{\rho_0}W^*$ so $\gamma_0$ would be identified relative to $\rho_0$. $W_y$ and $W_x$ would be unchanged.}}
The fifth assumption allows for a specific kind of network asymmetry. We require the diagonal of $W_{0}^{2}$ not to be constant as one of our sufficient conditions for identification.
In unweighted networks, the diagonal of the square of the social interactions matrix captures the number of reciprocated links for each individual or, in the case of undirected networks, the popularity of those individuals. Assumption (A5) hence intuitively suggests differential popularity across individuals in the social network.
This assumption is related to the network asymmetry condition proposed elsewhere, such as in Bramoulleetal2009. They show that when $W_{0}$ is known, the structural model ((ref)) is identified if $I$, $ W_{0}$, and $W_{0}^{2}$ are linearly independent. Given the remaining assumptions, this condition is satisfied if (A5) is satisfied, but the converse is not true: one can construct examples in which $I$, $W_{0}$ , and $W_{0}^{2}$ are linearly independent when $W_{0}^{2}$ has a constant diagonal, so $\Pi _{0}$ does not pin down $\theta _{0}$. See Example 1 in Appendix A. The strengthening of this hypothesis is the formal price to pay for the social interactions matrix $W_{0}$ being unknown to the researcher.
Before proceeding to our formal results, we provide a very simple illustration to shed light on how the assumptions above come together to provide identification. Suppose the observed reduced-form matrix is,
and that, following (A4), the first row is normalized to one. From the third row and column of $\Pi _{0}$, we see there is no path of any length connecting the individual in row 3 to or from those in rows 1 or 2, since her outcome is not affected by their covariates and their outcomes are not affected by her covariates. In other words, individual 3, is isolated and $ (W_{0})_{13}=(W_{0})_{23}=(W_{0})_{31}=(W_{0})_{32}=0$. On the other hand, individuals 1 and 2 cannot be isolated, as their covariates are correlated with the other individual's outcome, reflecting (A5).\footnote{ If on the other hand, $(W_{0})_{ij}=0.5,i\neq j$ in violation of (A5), and all agents were connected, the model would not be identified.} Due to the row-sum normalization of the first row, $(W_{0})_{12}=1$. Using (A3), it can be seen that $W_{0}$ is symmetric if $\Pi _{0}$ is symmetric. We thus find that $(W_{0})_{21}=1$. This and (A1) map all elements of $W_{0}$, and thus,
As the third individual is isolated, she will only be affected by her exogenous $x_{i}$ and not by endogenous or exogenous peer effects. Hence, the $(3,3)$ element of $\Pi _{0}$ is equal to $\beta _{0}=\frac{182}{455}=.4$. \ To find $\rho _{0}$, note that $(I-\rho _{0}W_{0})\Pi _{0}=\beta _{0}I+\gamma _{0}W_{0}$. Hence, focusing on the (1,1) elements of the matrices above, we find that $\frac{275}{455}-\rho _{0}\frac{310}{455}=.4$, implying $\rho _{0}=.3$ (complying with (A2)). Finally, $\gamma _{0}$ is identified from entry $(1,2)$, giving $\gamma _{0}=\frac{310}{455}-.3\frac{ 275}{455}=.5$.
Our final assumption articulates the need for a constant network $W_0$ observed over multiple instances of $y_t$ and $x_t$:
Here, “instances” can refer to time but also to settings in which the same units are observed over multiple episodes. For example, if $i$ are firms, then $t$ can be segmented markets in which they operate. For simplicity, we refer to an instance as a time period. If $\Pi_0$ is known, the main identification result we articulate below will state that $W_0$, $\rho_0$, $ \beta_0$ and $\gamma_0$ are globally identified. However, in practice, $\Pi_0$ is rarely observed and thus all quantities need to be estimated. For this purpose, when $\Pi_0$ is not known, multiple observations of $y_t$ and $x_t$ with a constant $W_0$ are required to implement the estimator. We expand on estimation requirements in Section (ref).
Importantly, the main identification results (for a given $\Pi_0$) could, in principle, be applied for each time period $t$. That is, one can write a version of Equation (ref) as
where $W_{0t}$ is time-varying and, consequently, the reduced-form interaction matrix $\Pi_{0t}=(I-\rho_0 W_{0t})^{-1}(I\beta_0+W_{0t}\gamma_0)$ is also time-varying. If the reduced-form matrices were known, the identification results we develop below could be applied to the reduced-form element by element for each $\Pi_{0t}$. Again, one rarely observes $\Pi_{0t}$ , for all $t\in[1,T]$. This observation will motivate an extension of the method, presented in Section (ref), where Assumption (A6) is relaxed and $W_{0}$ is allowed to vary with $t$.
In an extension, we allow the network $W_{0t}$ to vary over time and introduce kernel weights. Akin to the nonparametric regression $ Y=f(X)+\epsilon $, $f$ is identified if $\mathbb{E}(\epsilon |X)=0$, and it is possible to estimate $f$ using neighboring observations if $f$ is sufficiently smooth or varies slowly. Similar considerations extend to varying-coefficient models and, in particular, time-varying coefficient models where local stability conditions as those discussed in Dahlhaus2012 are usually invoked (see also HastieTibshirani1993 and their Example (e)).
Under the assumptions above, we can begin to identify parameters related to the network. These results are then useful for our main identification theorems. Let $\lambda _{0j}$ denote an eigenvalue of $W_{0}$ with corresponding eigenvector $v_{0,j}$ for $j=1,\dots ,N$. Assumptions (A2) and (A3) allow us to identify the eigenvectors of $W_{0}$ directly from the reduced form. As $|\rho _{0}|<1$:
The infinite sum converges as $|\rho _{0}\lambda _{0,j}|<1$ by (A2). The equation above implies that $v_{0,j}$ is also an eigenvector of $\Pi _{0}$ with the associated eigenvalue $\lambda _{\Pi ,j}=\frac{\beta _{0}+\gamma _{0}\lambda _{0,j}}{1-\rho _{0}\lambda _{0,j}}$. The fact that eigenvectors of $W_{0}$ are also eigenvectors of $\Pi _{0}$ has a useful implication: eigencentralities may be identified from the reduced form, even when $W_{0}$ is not identified. As detailed in DePaula2017 and Jacksonetal2017, such eigencentralities often play an important role in empirical work as they allow a mapping back to underlying models of social interaction.\footnote{ To identify the eigencentralities, we identify the eigenvector that corresponds to the dominant eigenvalue. If $W_{0}$ is non-negative and irreducible, this is the (unique) eigenvector with strictly positive entries, by the Perron-Frobenius theorem for non-negative matrices (see HornJohnson2013, p. 534).}
Now let $\Theta \equiv \{\theta \in \mathbb{R}^{m}:$ A$\text{ssumptions (A1)-(A6) are satisfied}\}$ be the structural parameter space of interest. Our identification argument is structured as follows: a) we first establish local identification of the mapping $\Pi (\theta )$ using classical results on the rank of the gradient of Rothenberg1971 (Theorem 1); b) we then show that $\Pi (\theta )$ is proper (Corollary 1); and c) has a connected image (Lemma 2, in the Appendix); d) allowing us to state the cardinality of the pre-image $\Pi ^{-1}(\bar{\Pi})$ is constant for any $ \bar{\Pi}$ in the image of $\Pi (\cdot )$, and that the cardinality is at most 2 (Theorem 2). We then provide additional conditions to narrow the identified set to a singleton (Corollaries 2-4).
We now formally present our results. Our first theorem establishes local identification of the mapping. A parameter point $\theta _{0}$ is locally identifiable if there exists a neighborhood of $\theta _{0}$ containing no other $\theta $ which is observationally equivalent. Using classical results in Rothenberg1971, we show that our assumptions are sufficient to ensure that the Jacobian of $\Pi $ relative to $\theta $ is non-singular, which, in turn, suffices to establish local identification.
An immediate consequence of local identification is that the set $\{\theta \in \Theta :\Pi (\theta )=\Pi (\theta _{0})\}$ is discrete (i.e., its elements are isolated points). The following corollary establishes that $\Pi $ is a proper function, i.e. the inverse image $\Pi ^{-1}(K)$ of any compact set $K\subset \mathbb{R}^{N^{2}}$ is also compact (KrantzParks2013 , p.\ 124). Since it is discrete, the identified set must be finite.
Under additional assumptions, the identified set is at most a singleton in each of the partitioning sets $\Theta _{-}\equiv \Theta \cap \{\rho \beta +\gamma <0\}$ and $\Theta _{+}\equiv \Theta \cap \{\rho \beta +\gamma >0\}$.\footnote{ The global inversion results we use are related to, but different from, variations on a classic inversion result of Hadamard that has been used in the literature. In contrast, we employ results on the cardinality of the pre-image of a function, relying on less stringent assumptions. While the Hadamard result requires the image of the function to be simply-connected (Theorem 6.2.8 of KrantzParks2013), the results we rely on do not.}
Since $\Theta =\Theta _{-}\cup \Theta _{+}$, if the sign of $\rho _{0}\beta _{0}+\gamma _{0}$ is unknown, the identified set contains, at most, two elements. In the theorem that follows, we show global identification only for $\theta \in \Theta _{+}$, since arguments are mirrored for $\theta \in \Theta _{-}$.
Similar arguments apply if Theorem (ref) instead were to be restricted to $\theta\in\Theta_{-}$. The proof of the corollary below is immediate and therefore omitted.
We now turn our attention to the problem of identifying the sign of $\rho _{0}\beta _{0}+\gamma _{0}$ from the observation of $\Pi _{0}$. This would then allow us to establish global identification using Theorem (ref). It is apparent from ((ref)) that if $\rho _{0}>0$ and $(W_{0})_{ij}\geq 0$, for all $i,j=\{1,\dots ,N\}$, the off-diagonal elements of $\Pi _{0}$ identify the sign of $\rho _{0}\beta _{0}+\gamma _{0}$.
Real-world applications often suggest endogenous social interactions are positive ($\rho _{0}>0$), in which case global identification is fully established by Corollary (ref). On the other hand, if $\rho _{0}<0$ (e.g., if outcomes are strategic substitutes), $\rho _{0}^{k} $ in ((ref)) alternates signs with $k$, and the off-diagonal elements no longer carry the sign of $\rho _{0}\beta _{0}+\gamma _{0}$. Nonetheless, if $W_{0}$ is non-negative and irreducible (i.e., not permutable into a block-triangular matrix or, equivalently, a strongly connected social network), the model is also identifiable without further restrictions on $\rho _{0}$:
Corollary (ref) requires that $W_0$ be irreducible, i.e., that it is not permutable into a block upper-triangular matrix. In the context of directed graphs, this is similar to requiring that the matrix be strongly connected, that is, that any node can be reached from any other node. The corollary then rules out cases when the network is not connected, for example, if there are two disjoint groups (with no connection across groups), or a star network pointing from the center towards the edges. The corollary holds if there are at least two real eigenvalues, or if $\rho _{0}$ is appropriately bounded. Since $W_{0}$ is non-negative, it has at least one real eigenvalue by the Perron-Frobenius theorem. If $W_{0}$ is symmetric, for example, its eigenvalues are all real, and Corollary (ref) holds. It also holds if $(W_{0})_{ij}\leq 0$, as we can rewrite the model as $\rho W_{0}=-\rho |W_{0}|$, where $|W_{0}|$ is the matrix whose entries are the absolute values of the entries in $W_{0}$. However, Corollary (ref) rules out cases that mix positive $(W_0)_{ij}\geq 0$ and negative interactions $(W_0)_{ij}\leq 0$. In any case, the bound on $|\rho _{0}|$ is sufficient and holds in most (if not all) empirical estimates we are aware of obtained from either elicited or postulated networks, and in our application on tax competition.
We present three extensions of the method for individual fixed effects, common shocks, and time-varying $W$. Appendix B describes extensions for multivariate covariates and heterogeneous $\beta _{0}$.
We observe outcomes for $i=1,\dots ,N$ individuals repeatedly through $ t=1,\dots ,T$ instances. If $t$ corresponds to time, it is natural to think of there being unobserved heterogeneity across individuals, $\alpha _{i}$, to be accounted for when estimating $\Pi _{0}$. The structural model ((ref)) is then,
which can be written in matrix form as
where $\alpha ^{\ast }$ is the vector of fixed effects. Individual-specific and time-constant fixed effects can be eliminated using the standard subtraction of individual time averages. Defining $\bar{y} _{t}=T^{-1}\sum_{t=1}^{T}y_{t}$, $\bar{x}_{t}=T^{-1}\sum_{t=1}^{T}x_{t}$, and $\bar{\epsilon}_{t}=T^{-1}\sum_{t=1}^{T}\epsilon _{t}$,
if $W_{0}$ does not change with time. Identification from the reduced form follows from previous theorems, since $\Pi _{0}$ is unchanged when regressing $y_{t}-\bar{y}_{t}$ on $x_{t}-\bar{x}_{t}$.\footnote{ \textcolor{black}{As is the case in panel data, this would require strict exogeneity ($\mathbb{E}[\epsilon _{s}|x_{t}]=0$ for any $s$ and $t$) or predetermined errors ($\mathbb{E}[\epsilon _{s}|x_{t}]=0$ for $s \ge t$) so that the matrix $\Pi _{0}$ can be consistently estimated.}}
We next allow for unobserved common shocks to all individuals in the network in the same instance $t$. Such correlated effects\ $\alpha _{t}$ can confound the identification of social interactions. As we have not placed any distributional assumption on the covariance matrix of the disturbance term, our analysis readily incorporates correlated effects that are orthogonal to $x_{t}$. When this is not the case, one possibility is to model the correlated effects $\alpha _{t}$ explicitly. The model then is,
where $\alpha _{t}$ is a scalar capturing shocks in the network common to all individuals. Let $\Pi _{01}=\left( I-\rho _{0}W_{0}\right) ^{-1}$ and $ \Pi _{02}=\left( \beta _{0}I+\gamma _{0}W_{0}\right) $ such that $\Pi _{0}=\Pi _{01}\Pi _{02}$. The reduced-form model is
We propose a transformation to eliminate the correlated effects:\ exclude the individual-invariant $\alpha _{t}$, subtracting the mean of the variables in a given period (global differencing). For this purpose, define $ H=\frac{1}{n}\iota \iota ^{\prime }$. We note that in empirical and theoretical work, it is customary to strengthen Assumption (A4) and require that all rows of $W_{0}$ sum to one if no individual is isolated (see for example Blumeetal2015). This strengthened assumption is usually referred to as row-sum normalization, and is stated below:
This can be written compactly as $W_{0}\iota =\iota $. In this case, $W_{0}$ can be interpreted as the normalized adjacency matrix. Under row-sum normalization we have that,
because $\left( I-H\right) \left( I-\rho _{0}W_{0}\right) ^{-1}\alpha _{t}\iota =0$ if Assumption (A4$^\prime$) holds. It then follows that $\tilde{\Pi} _{0}=(I-H)\Pi _{0}$ is identified. The next proposition shows that, under row-sum normalization of $W_{0}$, $\Pi _{0}$ is identified from $\tilde{\Pi} _{0}$ (and, as a consequence, the previous results immediately apply).
Under row-sum normalization of $W_{0}$, a common group-level shock affects individuals homogeneously since $(I-\rho _{0}W_{0})^{-1}\alpha _{t}\iota =\alpha _{t}(I+\rho _{0}W_{0}+\rho _{0}^{2}W_{0}^{2}+\cdots )\iota =\frac{ \alpha _{t}}{1-\rho _{0}}\iota $, which is a vector with no variation across entries. Consequently, global differencing eliminates correlated effects and $\left( I-H\right) \left( I-\rho _{0}W_{0}\right) ^{-1}\alpha _{t}\iota =\left( I-\rho _{0}W_{0}\right) ^{-1}\alpha _{t}\left( I-H\right) \iota =0$. Absent row-sum normalization, global differencing does not ensure correlated effects are eliminated. To see this, note that $(I-\rho _{0}W_{0})^{-1}$ is no longer row-sum normalized and $\alpha _{t}(I-\rho _{0}W_{0})^{-1}\iota $ does not have constant entries.
The next proposition makes this point formally: that the stronger Assumption (A4$^\prime$) is necessary to eliminate group-level shocks by showing it is not possible to construct a data transformation that eliminates group effects in the absence of row-sum normalization.
It is useful to be able to test for row-sum normalization (A4$^\prime$) as it enables common shocks to be accounted for in the social interactions model. This is possible as
The last equality follows from the observation that, under row-normalization of $W_{0}$, $W_{0}^{k}\iota =W_{0}\iota =\iota $, $k>0$. This implies $\Pi _{0}$ has constant row-sums, which suggests row-sum normalization is testable. In the Appendix, we derive a Wald test statistic to do so.\footnote{ For ease of explanation, in the Appendix, we derive the test under the asymptotic distribution of the OLS estimator. The test generally holds with minor adjustments for estimators with known asymptotic distributions.}
We now relax Assumption (A6), which states that $W_0$ does not vary across the time periods $t=1,\dots,T$. The version of Equation (ref) with time-varying network is
with the reduced-form matrix $\Pi_{0t}=(I-\rho_0 W_{0t})^{-1}(I\beta_0+W_{0t}\gamma_0)$. We note that the identification results developed in Subsection (ref) can, in principle, be applied element by element to each $\Pi_{0t}$, leading to the identification of a time-varying $W_{0t}$ (and, potentially, of the parameters $\rho_0$, $ \beta_0 $ and $\gamma_0$).
In practice, implementing any estimation strategy with a time-varying $ \Pi_{0t}$ (or $W_{0t}$) is not feasible using only observation from the single time period $t$. We instead adopt a kernel-weighted version. Define period-specific weights $\omega _{t}$, and consider the transformed data $ \tilde{y}_{s}=\omega _{s}(t)y_{s}$ and $\tilde{x}_{s}=\omega_{s}(t)x_{s}$, $ s=1,\dots,T$. Evidently, uniform weights $\omega _{s}=1, s=1,\dots,T$ are equivalent to the strategy not considering time-varying networks, and assuming that the networks are fixed within those windows. Alternatively, one could estimate $W_{t}$ in time windows by setting $ \omega_{t}=1[\underbar{t}\leq t\leq \bar{t}]$, where $\underbar{t}$ and $ \bar{t}$ are the start and end of the time window for which $W_{t}$ is estimated. In this case, the minimum effective window length $\bar{t}- \underbar{t}$ can be computed as we discuss in Section 3.1. In the context of DSGE models with time-varying parameters, kapetanios2019time suggests a Gaussian kernel with positive weights throughout the entire sample. As in nonparametric regression with smooth kernel weights, it also assumes that the network evolves slowly over time. We further discuss this strategy in the estimation section, and it is implemented in the empirical application section below.
We now transition from our core identification results to their practical implementation. In practice, Ordinary Least Squares (OLS) can only be used to estimate $\theta $ if $T\gg N$, which is in practice unlikely to be met, as this is a high-dimensional problem. Our preferred approach makes use of penalized estimation techniques that can be used for any given $T$. More specifically, we make use of the Adaptive Elastic Net GMM (CanerZhang2014), which is based on the penalized GMM objective function. Given the identification results presented in Section (ref), the population moments used in forming the GMM objective function will be satisfied at the true parameter vector.
After setting out the estimation procedure, we showcase the method using Monte Carlo simulations based on stylized and real-world network structures. In each case, we take a fixed network structure $W_{0}$, and simulate panel data as if the data generating process were given by the model in ((ref)). We apply the method to the simulated panel data to recover estimates of all elements in $W_{0}$, as well as the endogenous and exogenous social effect parameters.
The parameter vector to be estimated is high-dimensional: $\theta =\left( W_{12},\dots ,W_{N,N-1},\rho ,\gamma ,\beta \right) ^{\prime }\in \mathbb{R} ^{m}$, where $m=N\left( N-1\right) +3$ and $W_{ij}$ is the $(i,j)$-th element of the $N\times N$ social interactions matrix $W_{0}$. \textcolor{black}{To be clear, in a network with $N$ individuals, there are $N(N-1)$ potential interactions because an individual could interact with everyone else but herself (which would violate Assumption A1). As a consequence, even with a modest $N$, there are many more parameters to estimate, and $m$ is large. For example, a network with $N=50$ implies more than 2,000 parameters to estimate. While we consider $N$ (and thus $m$) fixed, we still refer to $\theta$ as high-dimensional.} OLS estimation requires $m\ll NT(\Rightarrow N\ll T)$, so many more time periods than individuals: a requirement often met in finance data sets ( vanVliet2018) \textcolor{black}{or in other fields (see, e.g., Section 4.2 in rothenhausler2015)}. Instead, to estimate a large number of parameters with limited data, we utilize high-dimensional estimation methods, which are the focus of a rapidly growing literature.
Sparsity is a key assumption underlying many high-dimensional estimation techniques. In the context of social interactions, we say that $W_{0}$ is sparse if $\tilde{m}$, the number of non-zero elements of $W_{0}$, is such that $\tilde{m}\ll NT$. The notion of sparsity thus depends on the number of time periods. Sparsity corresponds to assuming that individuals influence or are influenced by a small number of others, relative to the overall size of the potential network and the time horizon in the data. As such, sparsity is typically not a binding constraint in social networks analysis. \footnote{ Common stylized networks are sparse, such as the star, lattice (each individual is a source of spillover only to one other individual), or interactions in pairs, triads or small groups (DeGiorgietal2010 ). Real-world economic networks are also sparse. The sparsity in AddHealth friendship network is around $98$%. Sparsity of the production networks in the US\ is above $99$% (Atalayetal2011).}
In the estimation of sparse models, the \textquotedblleft effective number of parameters\textquotedblright\ (or \textquotedblleft effective degrees of freedom\textquotedblright ) relates to the number of variables with non-zero estimated coefficients (TibshiraniTaylor2012). In the context of the current social network model, this is equivalent to $m$ parameters, where $m=dN(N-1)+K$ and $d$ is the network density defined as $\tilde{m}/(N(N-1))$. The Adaptive Elastic Net GMM estimator presented by CanerZhang2014 converges at a rate of $ \sqrt{NT/\tilde{m}}=\sqrt{NT/[dN(N-1)+K]}=O(\sqrt{T/(dN)})$ (see remark 7 in CanerZhang2014). Hence the quality of the large sample results relies on a comparison between $T$ and $dN$. In line with this, we thus require $NT\gg dN(N-1)+K$. For example, in the high-school network of Coleman1964 that is part of our simulation exercise, $N=70$ and $ d=0.076$. Assuming $K=3$, $N(N-1)\times d+3=370.1$.\footnote{ As pointed out by a referee, variation in $x$ will also matter for estimation precision. This is reflected in the asymptotic distribution for this estimator, shown later in this subsection.}
Finally, to reiterate, our identification results themselves do not depend on the sparsity of networks. In particular, Assumptions (A1)-(A6) do not impose restrictions on the number of links in $W_{0}$, or $ \tilde{m}$.\footnote{ If $N\rightarrow \infty $, Assumption (A2) would imply vanishing $ (W_{0})_{ij}$ entries. As highlighted previously, we consider $N$ to be fixed, in line with many practical applications. Furthermore, Assumption (A2) is used to represent inverse matrices as Neumann series in our identification results. What is necessary for this to hold is that a sub-multiplicative norm on $\rho W$ be less than one. Here we use a specific norm (i.e., the maximum row-sum norm), but other (induced) norms are also possible (i.e., the 2-norm or the 1-norm) (see HornJohnson2013, Chapter 5.6).} The identification results presented in Section (ref) apply more broadly and irrespective of the estimation procedure.
Our preferred approach estimates the interaction matrix in the reduced form while penalizing and imposing sparsity on the structural object $W_{0}$. \textcolor{black}{We impose sparsity and penalization in the structural-form matrix $W_{0}$ because this is a weaker requirement than imposing sparsity and penalization in the reduced-form matrix $\Pi _{0}$.}\footnote{ \textcolor{black}{Note that even if $W$ is sparse, $\Pi$ may not be sparse.} In Appendix C.1, we show that $[\Pi _{0}]_{ij}=0$ if, and only if, there are no paths between $i$ and $j$ in $W_{0}$, so the pair is not connected. So, sparsity in $\Pi _{0}$ is understood as $W_{0}$ being “sparsely connected”, which is a stronger assumption than sparsity in $W_{0}$.} To do so, we use the Adaptive Elastic Net GMM (CanerZhang2014), which is based on the penalized GMM objective function,
where $\theta =(W_{1,2},\dots ,W_{N,N-1}\,\rho ,\gamma ,\beta )^{\prime }$ with dimension $m=N(N-1)+3$, and $p_{1}$ and $p_{2}$ are the penalization terms. The term $g_{NT}\left( \theta \right) ^{\prime }M_{T}g_{NT}\left( \theta \right) $ is the unpenalized GMM objective function with moment conditions based on orthogonality between the structural disturbance term and the covariates: $g_{NT}\left( \theta \right) =\sum_{t=1}^{T}\left[ x_{1t}e_{t}(\theta )^{\prime }\;\cdots \;x_{Nt}e_{t}(\theta )^{\prime } \right] ^{\prime }$, $e_{t}(\theta )=y_{t}-\rho Wy_{t}-\beta x_{t}-\gamma Wx_{t}$. There are $q\equiv N^{2}$ moment conditions since $x_{it}$ is orthogonal to $e_{jt}$ for each $i,j=1,\dots ,N$. Hence, the GMM weight matrix $M_{T}$ is of dimension $N^{2}\times N^{2}$, symmetric, and positive definite. For simplicity, we use $M_{T}=I_{N^{2}\times N^{2}}$. Note that if $x_{t}$ is econometrically endogenous, one can also exploit moment conditions with respect to available instrumental variables.\footnote{ For expositional ease, we describe estimation in the context of the reduced-form model ((ref)), thereby abstaining from individual fixed or correlated effects. As the GMM estimator uses moments between the structural disturbance terms and covariates, this endogeneity is built into the estimation procedure.} Given the identification results presented in Section (ref), if $\theta \neq \theta _{0}$ and does not belong to the identified set, then $\Pi (\theta )\neq \Pi (\theta _{0})$. Consequently, the populational version of the GMM objective function is uniquely minimized at the true parameter vector $\theta _{0}$.
The penalization terms in Equation ((ref)) are what makes this different from a standard GMM problem. The first term, $p_{1}\sum_{{ i,j=1,i\neq j}}^{N}\left\vert W_{i,j}\right\vert $, penalizes the sum of the absolute values of $W_{ij}$, i.e., the sum of the strength of links, for all node-pairs. Depending on the choice of $p_{1}$, some $W_{i,j}$'s will be estimated as exact zeros. A larger share of parameters will be estimated as zeros if $p_{1}$ increases. The second term, $p_{2}\sum_{{i,j=1,i\neq j} }^{N}\left\vert W_{i,j}\right\vert ^{2}$, penalizes the sum of the square of the parameters. This term has been shown to provide better model-selection properties, especially when explanatory variables are correlated ( ZouZhang2009). The first-stage estimate is
where $(1+p_{2}/T)$ is a bias-correction term also used by CanerZhang2014.
Implementing the numerical optimization embedded in Equation (ref) is computationally challenging, as $m=N(N-1)+3$ may entail a large number of function arguments. We instead implement the following modification to use fast Least-Angle Regression (LARS) algorithms ( efron2004least). For any given $\rho $, $\beta$, and $\gamma $, the expression for $e_{t}(\theta )$ is linear in $W$:
where $\tilde{y}_{it}(\beta )\equiv y_{t}-x_{t}\beta $ and $\tilde{x} _{t}(\rho ,\gamma )\equiv \rho y_{t}+x_{t}\gamma $ and, following the strategy above, is instrumented with $x_{t}$. This motivates a two-step optimization routine:
where the expression in brackets has a computationally efficient solution through the LARS algorithm. The numerical optimization is then subsequently conducted over the parameter space of $(\rho ,\beta ,\gamma )$ only. We also impose row-sum normalization. Details of the implementation are expanded in Appendix Subsection C.2.
A second (adaptive) step provides improvements by re-weighting the penalization by the inverse of the first-step estimates (Zou2006):
where $\tilde{W}_{i,j}$ is the $\left( i,j\right) $-th element of the first-step estimate of $W$. We follow CanerZhang2014 and set $c =2.5$ . If $\tilde{W}_{i,j}<0.05$, we set $\tilde{W}_{i,j}=0.05$. This ensures that the second-stage estimates can be non-zero even if the first-stage estimates were zero or small. The computational improvement -- described above for the first-stage estimator -- is also applied in the adaptive stage.
As a third and final step, we fix the support of $\tilde{\theta}^{\ast }(p)$ , $\mathcal{S}=\{\rho ,\beta ,\gamma \}\cup \{W_{ij}:\tilde{W}_{ij}^{\ast }\neq 0\}$ and estimate the final parameters without penalization. This takes as arguments only the elements of $\tilde{\theta}^{\ast }(p)$ that were estimated as non-zero in the adaptive step. In essence, this step boils down to a standard GMM approach,
Importantly, CanerZhang2014 show that the third-step estimator is asymptotically normal, with a known and easy-to-compute distribution,
where $\hat{G}\equiv \hat{G}(\hat{\theta})=\nabla g_{NT}(\theta )$ and $ \Omega \equiv E[g_{NT}(\theta )g_{NT}(\theta )^{\prime }]$.\footnote{ This applies in the case of small $p_{2}$. In the case of large $p_{2}$, the asymptotic distribution is pre-multiplied by $K_{n}=\frac{I+p_{2}\left[ \hat{G}(\hat{\theta})^{\prime }\hat{\Omega}^{-1}\hat{G}(\hat{\theta})\right] ^{-1}}{1+p_{2}/{NT}}$. See Theorem 4 of CanerZhang2014.} This allows us to conduct hypothesis testing and inference on the $\rho $, $\beta $, $ \gamma $ and the non-zero elements of $W$.
We write $p=(p_{1},p_{1}^{\ast },p_{2})$ as the final set of penalization parameters. Conditional on $p$, the estimate of the procedure is $\hat{\theta }(p)$. As in CanerZhang2014, the penalization parameters $p$ are chosen by the BIC criterion. This balances model fit with the number of parameters included in the model.\footnote{\textcolor{black}{Following CanerZhang2014,} the choice of $p$, which we denote as $\hat{p}$, is the one that minimizes
where $A\Big(\hat{\theta}(p)\Big)$ counts the number of non-zero coefficients among $\{W_{1,2},\dots ,W_{N,N-1}\}$, and larger than a numerical tolerance, which we set at $10^{-5}$. See also zouhastietibshirani2007.}
In Appendix C.2, we provide further implementation details, including the choice of initial conditions. Of course, other estimation methods are available, and our identification results do not hinge on any particular estimator. Our aim is to demonstrate the practical feasibility of using the Adaptive Elastic Net estimator rather than claim it is the optimal estimator.\footnote{ See the alternative approaches of GautierTsybakov2014, Manresa2016, lamsouza2016, and GautierRose2016.}
We showcase the method using Monte Carlo simulations. We describe the simulation procedures, results, and robustness checks in more detail in Appendix D.1. Here, we just provide a brief overview to highlight how well the method works to recover social networks even in relatively short panels.
For each simulated network, we take a fixed network structure $W_{0}$ and simulate panel data as if the data generating process were given by ((ref)). We then apply the method to the simulated panel data to recover estimates of all elements in $W_{0}$, as well as the endogenous and exogenous social effect parameters ($\rho _{0}$, $\gamma _{0}$). Our result identifies entries in $W_{0}$ and so naturally recovers links of varying strength. It is long recognized that link strength might play an important role in social interactions (Granovetter1973). Data limitations often force researchers to postulate some ties to be weaker than others (say, based on interaction frequency). In contrast, our approach identifies the continuous strength of ties, $W_{0,ij}$, where $W_{0,ij}>0$ implies node $j$ influences node $i$.
The stylized networks we consider are a random network and a political party network in which two groups of nodes each cluster around a central node. The real-world networks we consider are the high-school friendship network in Coleman1964 from a small high school in Illinois, and one of the village networks elicited in banerjeeetal2013 from rural Karnataka, India.
Summary statistics for each network are presented in \textcolor{bluez}{Panel A} of \textcolor{bluez}{Table A1}. The four networks differ in their size, complexity, and the relative importance of strong and weak ties. For example, the Erd\"{o}s-Renyi network only has strong ties, while the political party network has twice as many strong as weak ties. For the real-world networks, the mean out-degree distributions are higher, so the majority of ties are weak, with the high school network having around $80$% of its edges being weak ties. All four networks are also sparse.
For the stylized networks, we assess the performance of the estimator for a fixed network size, $N=30$. We simulate the real-world networks using non-isolated nodes in each (so $N=70$ and $65$ respectively).\footnote{ Like Bramoulleetal2009, we exclude isolated nodes because they do not conform to row-sum normalization.}
We evaluate the procedure over varying panel lengths (starting from short panels with $T=5$), using various metrics. Given our core contribution is to identify the social interactions matrix, we first examine the proportion of true zero entries in $W_{0}$ estimated as zeros and the proportion of true non-zero entries estimated as non-zeros. A global perspective of the proximity between the true and estimated networks can be inferred from their average absolute distance between elements. This is the mean absolute deviation of $\hat{W}$ and $\hat{\Pi}$ relative to their true values, defined as $MAD(\hat{W})=\frac{1}{N(N-1)}\sum_{i,j,i\neq j}|\hat{W} _{ij}-W_{ij,0}|$ and $MAD(\hat{\Pi})=\frac{1}{N(N-1)}\sum_{i,j,i\neq j}|\hat{ \Pi}_{ij}-\Pi _{ij,0}|$. As these metrics are closer to zero, more of the elements in the true matrix are correctly estimated. Finally, we evaluate the procedure's performance using averaged estimates of the endogenous and exogenous social effect parameters, $\hat{\rho}$ and $\hat{\gamma}$. In keeping with the estimation strategy in our empirical application, we report unpenalised GMM.
\textcolor{bluez}{Figure A1} shows the simulation results as evaluated using the six metrics described above. \textcolor{bluez}{Panel A} shows that for each network, the proportion of zero entries in $W_{0}$ correctly estimated as zeros is above $95$% even when $T=5$. The proportion approaches $100$% as $T$ grows. Conversely, \textcolor{bluez}{Panel B} shows the proportion of strong non-zero entries estimated as non-zeros (defined as larger than 0.3) is also high for a small $T$. It is above $70$% from $T=5$ for the Erd\"{o}s-Renyi network, being at least $85$% across networks for $T=25$, and increasing as $T$ grows. As discussed above, the Adaptive Elastic Net estimator may only recover strong edges well, and not necessarily the weaker ones, due to the well-known issue with shrinkage estimators that they tend to shrink small parameters to zero. We return to this issue below.
\textcolor{bluez}{Panels C and D} show that for each simulated network, the mean absolute deviation between estimated and true networks for $\hat{W}$ and $\hat{\Pi}$ falls quickly with $T$ and is close to zero for large sample sizes. Finally, \textcolor{bluez}{Panels E and F} show that biases in the endogenous and exogenous social effects parameters, $\hat{\rho}$ and $\hat{ \gamma}$, also fall in $T$ (we do not report the bias in $\hat{\beta}$ since it is close to zero for all $T$). The fact that biases are not zero is as expected for a small $T$, being analogous to well-known results for autoregressive time series models. \textcolor{black}{\footnote{\textcolor{black}{The bias in spatial auto-regressive models with a small number of observations even when the network is observed is similarly documented by Smith2009, NeumanMizruchi2010, Wangetal2014, and others.}}}
\textcolor{bluez}{Figure A2} shows that, as $T$ increases, the procedure detects weaker links. The figure also shows that, with low sample sizes, weak edges are generally not detected. This pattern is consistent with the well-known fact that small parameters are likely shrunk to zero due to the penalization (BelloniChernozhukov2011b). The absence of weak edges also implies that the strength of strong edges may also be over-estimated, since rows are normalized to one. In \textcolor{bluez}{Panel A}, we show the distribution of the estimates of $\hat{W}_{ij}$, with $T=25$ and for the high-school network. We show the distribution for the five most common values of $W_{0,ij}$. We find that most edges weaker than $.5$ are not detected; edges with a strength of $.75$ are substantially more likely to be estimated as non-zeros. When they are detected as non-zeros, they are more likely to be over-estimated. When we estimate $W$ with $T=150$, \textcolor{bluez}{Panel B} shows that virtually all edges with strength greater than $.5$ are estimated as non-zeros, and most edges with strength $ .375$ are also detected. We further see are more continuous distribution of estimates of edge strength. Only edges smaller than $.25$ are not detected. \textcolor{bluez}{Panels C and D} show a similar conclusion for the village network.
\textcolor{bluez}{Figure A3} shows the simulated and actual networks under $ T=100$ time periods. The network size is set to $N=30$ in the two stylized networks, $N=70$ for the high school network, and $N=65$ for the village household network. In comparing the simulated and true networks, \textcolor{bluez}{Figure A3} distinguishes between kept edges, added edges, and removed edges. Kept edges are depicted in blue:\ these links are estimated as non-zero in at least $5$% of the iterations, and are also non-zero in the true network. Added edges are depicted in green:\ these links are estimated as non-zero in at least $5$% of the iterations but the edge is zero in the true network. Removed edges are depicted in red:\ these links are estimated as zero in at least $5$% of the iterations but are non-zero in the true network. \textcolor{bluez}{Figure A3} further distinguishes between strong and weak links:\ strong links are shown as solid edges ($W_{0,ij}>.3$), and weak links are shown as dashed edges.
\textcolor{bluez}{Panel A of Figure A3} compares the simulated and true Erd\"{o}s-Renyi networks. All links are recovered. For the political party network, Panel B shows that all strong edges are correctly estimated. However, around half the weak edges are recovered (blue dashed edges), with the others being missed (red dashed edges). As discussed above, this is not surprising given that shrinkage estimators force small non-zero parameters to zero. Hence, a larger $T$ is needed to achieve similar performance to the other simulated networks in terms of detecting weak links. For the more complex and larger real-world networks, \textcolor{bluez}{Panel C} shows that in the high-school network, the strong edges are all recovered. However, around half the weak edges are missing (red dashed edges), and there are a relatively small number of added edges (green edges):\ these amount to $87$ edges, or approximately $1.9 $% of the $4,534$ zero entries in the true high-school network. A similar pattern of results is seen in the village network in \textcolor{bluez}{Panel D}: the strong edges are all recovered, and here the majority of weak edges are also recovered.
\textcolor{bluez}{Panel B of Table A1} compares the network- and node-level statistics calculated from the recovered social interactions matrix $\hat{W}$ to those in Panel A from the true interactions matrix $W_{0}$. The random Erd\"{o}s-Renyi network is perfectly recovered. For the political party network, the number of recovered edges is slightly lower than in the true network ($41$ vs. $45$), and all edges are classified as strong. The mean of the in- and out-degree distributions are slightly lower in the recovered network, and all three nodes with the highest out-degree are correctly captured (nodes $1$ , $11$, and $28$), include both party leaders (individuals $1$ and $11$ ). We then move to discussing the performance in the two real-world networks. In the high-school network, 30% of all edges are correctly recovered, and they are all strong edges. As already noted in \textcolor{bluez}{Figure A2}, weak edges are not well estimated in the high-school network. This draws two main consequences. First, the average in- and out- degrees are smaller in the recovered network relative to the true network. Second, we over-estimate the number of strong edges ($61$ vs. $113$ ). This is a downside of row-sum normalization: because some weak edges get estimated as zeros, the non-zeros are over-estimated so that the row adds to one. We do, however, recover all three individuals with the highest out-degree. Finally, in the village network, half the edges are recovered. The same phenomena of underestimating weak and overestimating strong edges are again observed. We again recover the three households with the highest out-degree (nodes $16$, $35$, and $57$).
In the Appendix, we show the robustness of the simulation results to (i)\ varying network sizes and (ii)\ alternative parameter choices and richening the structure of shocks across nodes. We also demonstrate the gains from using the Adaptive Elastic Net GMM estimator over alternative estimators, such as the Adaptive Lasso and OLS.
We apply our results to shed new light on a classic social interactions problem:\ tax competition between US states ( Wilson1999). Defining competing “neighbors” remains the central empirical challenge in this literature. Theory provides some guidance on the issue through two mechanisms driving interactions across jurisdictions:\ factor mobility and yardstick competition.
On factor mobility, Tiebout1956 first argued that labor and capital can move in response to differential tax rates across jurisdictions. Factor mobility leads naturally to the postulated social interactions matrix being (i) geographic neighbors, given labor mobility, and (ii) jurisdictions with similar economic or demographic characteristics, given capital mobility ( Caseetal1989).
Yardstick competition is driven by voters making comparisons between states to learn about their own politician's quality (Shleifer1985). BesleyCase1995 formalize the idea in a model where voters use taxes set by governors in other states to infer their own governor's quality. Yardstick competition leads naturally to the postulated interactions matrix being “political neighbors”: states that voters make comparisons to.
In this application, the number of nodes and time periods is relatively low:\ the data covers mainland US states, $N=48$, for the years 1962-2015, $T=53$ . Our approach identifies the structure of social interactions among “economic neighbors”, denoted ${W}_{econ}$. We contrast this against a null that state taxes are influenced by geographic neighbors, $W_{geo}$, as shown in Panel A of \textcolor{bluez}{Figure 1A}. With $W_{econ}$ recovered, we can establish, beyond geography, what predicts the strength of ties between states and provide fresh insights on drivers of tax competition.
Before using the real data, we confirm the estimator's performance when the true network is $W_{geo}$ in simulated settings. In line with the findings of the previous section, Panel B of Figure 1A shows that (i) the procedure recovers strong edges frequently (more specifically, $89$% of the true strong edges are recovered) and (ii)\ performance deteriorates when recovering weak edges ($72$%). In all cases, the estimator does not add edges not in the true network. This suggests recovered economic links that deviate from geographic links may indeed carry signal, while weak links may not get detected. Finally, the estimator for $\rho $ and $\gamma $ may show some downward bias with the sample sizes in the application, consistent with the simulations in \textcolor{bluez}{Appendix Figure A1}.
We denote state tax liabilities for state $i$ in year $t$ as $\tau _{it}$, covering state taxes collected from real per capita income, sales, and corporate taxes. We extend the sample used by BesleyCase1995, that runs from 1962-1988 ($T=26$).\footnote{BesleyCase1995 test their political agency model using a two-equation set-up: (i) on gubernatorial re-election probabilities and (ii)\ on tax setting. Our application focuses on the latter because this represents a social interaction problem. They use two tax series: (i) TAXSIM data (from the NBER), which runs from 1977-1988 and (ii)\ state tax liabilities series constructed from data published annually in the Statistical Abstract of the US, that runs from 1962-1988. All their results are robust to either series. We extend the second series.} The outcome considered, $\Delta \tau _{it}$, is the change in tax liabilities between years $t$ and $(t-2)$ because it might take a governor more than a year to implement a tax program. Their model implies a standard social interactions specification for the tax setting behavior of state governors:
where $k=1,\dots ,K$ are the covariates for state $i$ in period $t$. Tax setting behavior is determined by (i)\ endogenous social effects arising through neighbors' tax changes ($\sum_{j=1}^{N}W_{0,ij}\Delta \tau _{jt});$ (ii) exogenous social effects arising through neighbors' characteristics ($ \sum_{j=1}^{N}W_{0,ij}x_{jkt})$; and (iii)\ state $i$'s characteristics ($ x_{ikt} $), including income per capita, the unemployment rate, and the proportions of young and elderly in the state's population. All specifications include state and time effects ($\alpha _{i},$ $\alpha _{t}$ ). Due to the inclusion of the time effects $\alpha _{t}$, we normalize the rows of $W_{econ}$ to one. \textcolor{bluez}{Table A6} presents descriptive statistics for the BesleyCase1995 sample and our extended sample.
Much of the earlier literature on tax competition has focused on endogenous social effects and ignored exogenous social effects by setting $\gamma =0$. Our identification result allows us to relax this restriction and estimate the full typology of social effects described by Manski1993. This is important because only endogenous social effects lead to social multipliers from tax competition, and they are crucial to identify as they can lead to a race-to-the-bottom or suboptimal public goods provision ( BrennanBuchanan1980; Wilson1986; OatesSchwab1988).
\textcolor{bluez}{Table 1} presents our preliminary findings and comparison to BesleyCase1995. Throughout this section, we refer to \lq\lq OLS estimates\rq\rq\ as the estimates of the main equation ((ref)) when $W_0$ is postulated as $W_{geo}$ or $W_{econ}$ and $\rho_0$, $\gamma_{0,k}$, and $\beta_{0,k}$ are estimated by OLS.\footnote{We postulate that $W$ is $W_{econ}$ obtained by running the procedure in Section 3, retrieving $\hat{W}$, and re-running model ((ref)) with $W=\hat{W}$. For such OLS estimates, we use robust standard errors and ignore the sampling uncertainty in the estimated $W_{econ}$.} \textcolor{bluez}{Column 1} shows those estimates where the postulated social interactions matrix is based on geographic neighbors, exogenous social effects are ignored and the panel includes all $48$ mainland states but runs only from 1962-1988 as in BesleyCase1995. Social interactions influence gubernatorial tax setting behavior:\ $\widehat{\rho }_{OLS}=.375$. \textcolor{bluez}{Column 2} shows this to be robust to instrumenting neighbors' tax changes using the instrument set proposed by BesleyCase1995:\ namely, instrumenting for $\Delta \tau _{jt}$ using geographic neighbors' lagged changes in per capita income and unemployment rates. These instruments are in the spirit of using exogenous social effects to instrument for neighbors' tax changes. $ \widehat{\rho }_{2SLS}$ is more than double the magnitude of $\widehat{\rho } _{OLS}$, suggesting tax setting behaviors across jurisdictions are strategic complements.
\textcolor{bluez}{Columns 3 and 4} replicate both specifications over the longer sample, confirming Besley and Case's (1995) \nocite{BesleyCase1995} finding on social interactions to be robust. $\widehat{\rho }_{2SLS}$ is again more than double $\widehat{\rho }_{OLS}$. The result in \textcolor{bluez}{Column 4} implies that for every dollar increase in the average tax rates among geographic neighbors, a state increases its own taxes by $64$ cents. This is similar to the headline estimate of BesleyCase1995.\footnote{ Nor is the magnitude very different from earlier work examining fiscal expenditure spillovers. For example, Caseetal1989 find that US state governments' levels of per-capita expenditure are significantly impacted by the expenditures of their neighbors, with a one-dollar increase in neighbors' expenditures leading to a seventy-cent increase in own-state expenditures.}
We now move beyond much of the earlier literature to first establish whether there are endogenous and exogenous social interactions in tax setting. We first focus on the endogenous and exogenous social interaction parameters, and in the next subsection, we detail the identified social interactions matrix, $\hat{W}_{econ}$. To do so, we need to modify slightly how we instrument for neighbors' tax changes:\ the instrument set proposed by BesleyCase1995 based on geographic neighbors' characteristics will generally be weaker when estimating the full specification in ((ref)) because the instruments are now directly controlled for in ((ref)). We use an Adaptive Elastic Net GMM approach, which instruments neighbors' tax changes with the characteristics of all other states. With the inclusion of endogenous and exogenous social effects, this represents our preferred approach.
\textcolor{bluez}{Columns 1 and 2 of Table 2} show OLS\ and GMM estimates for $\rho $ obtained from the Adaptive Elastic Net procedure, where we still set $\gamma =0$ but use our preferred instrument set:\ $\widehat{\rho } _{GMM}=.709>$ $\widehat{\rho }_{OLS}=.649$. \textcolor{bluez}{Columns 3 and 4} estimate the full model in ((ref)). Relative to when exogenous social effects are assumed away ($\gamma =0$), the OLS\ and GMM estimates of $\rho $ are smaller, but we continue to find robust evidence of endogenous social interactions in tax setting. The specification in \textcolor{bluez}{Column 4} represents our preferred one: $\widehat{\rho } _{GMM}=.452$ (with a standard error of $.132$). This value meets the requirements on $\rho $ in Corollaries 3 and 4 for global identification. \footnote{\textcolor{bluez}{Table A7} shows the full set of exogenous social effects (so \textcolor{bluez}{Columns 1 and 2} refer to the same specifications as \textcolor{bluez}{Columns 3 and 4 in Table 2}). Exogenous social effects operate through economic neighbors' unemployment rate, demographic characteristics, and their governor's age.}
\textcolor{bluez}{Figure 1B} shows how the structures of economic ($\hat{W} _{econ}$) and geographic networks ($W_{geo}$) differ, where connected edges imply that two states are linked in at least one direction (state $i$ causally impacts state taxes in $j$, and/or vice versa). This comparison makes clear whether all states geographically adjacent to $i$ matter for its tax setting behavior and whether there are non-adjacent states that influence its tax rate.
The left-hand panel of \textcolor{bluez}{Figure 1B}\ shows the network of geographic neighbors (whose edges are colored blue), onto which we superimpose edges not identified as links in $W_{econ}$;\ dropped edges are in red. The vast majority of geographically adjacent states are irrelevant for tax setting behavior. The right-hand panel of \textcolor{bluez}{Figure 1B}\ adds new edges identified in $\hat{W}_{econ}$ that are not part of $W_{geo}$;\ these added edges are in green and represent non-geographically adjacent states through which social interactions occur. For tax-setting behavior, economic distance is imperfectly measured if we simply assume interactions depend only on physical distance.
\textcolor{bluez}{Table 3} summarizes the comparison between $W_{geo}$ and $ \hat{W}_{econ}$. $W_{geo}$ has $214$ edges, while $\hat{W}_{econ}$ has only $ 49$. $\hat{W}_{econ}$ and $W_{geo}$ have $9$ edges in common. Hence, the vast majority of geographical neighbors ($205/214=96$%) are not relevant for tax setting. $\hat{W}_{econ}$ has $40$ edges that are absent in $W_{geo}$ , and the identified social interactions are more spatially dispersed than under the assumption of geographic networks. This is reflected in\ the far lower clustering coefficient in $\hat{W}_{econ}$ than in $W_{geo}$ ($.042$ versus $.419$).\footnote{ The clustering coefficient is the frequency of the number of fully connected triplets over the total number of triplets. }
Our estimation strategy identifies the continuous strength of links, $ W_{0,ij}$, where $W_{0,ij}>0$ is interpreted as state $j$ influencing outcomes in state $i$. This is useful because recent developments in tax competition theory, using insights from the social networks literature, suggest links need not be reciprocal (JanebaOsterleh2013).
\textcolor{bluez}{Table 3} reveals that only $12.2$% of edges in $\hat{W} _{econ}$ are reciprocal (all edges in $W_{geo}$ are reciprocal by definition). Hence, tax competition is both spatially disperse and asymmetric. In most cases where tax setting in state $i$ is influenced by taxes in state $j$, the opposite is not true.
Given common time shocks $\alpha _{t}$ in ((ref)), row-sum normalization is required and ensures $\sum_{j}W_{0,ij}=1$. Hence, for every state $i$, there will be at least one economic neighbor state $j^{\ast }$ that impacts it, so $W_{0,ij^{\ast }}>0$. This just reiterates that social interactions matter. On the other hand, our procedure imposes no restriction on the derived columns of $\hat{W}_{econ}$. It could be that a state does not affect any other state. To see this in more detail, the final rows of \textcolor{bluez}{Table 3} report the degree distribution across states, splitting for in-networks and out-networks. In $W_{geo}$, the in-degree is by construction equal to the out-degree, as all ties are reciprocal. The greater sparsity of the network of economic neighbors is reflected in the degree distribution being lower for $\hat{W}_{econ}$ than $W_{geo}$. In $ \hat{W}_{econ}$, the dispersion of in- and out-degree networks is very different (as measured by the standard deviation), being nearly nine times higher for the in-degree. Hence one reason for so few reciprocal ties being in the economic network is that out-degree network ties are rarely also in-degree ties.
This asymmetry in $\hat{W}_{econ}$ further suggests that some highly influential states drive tax setting behavior in other states. To see which states these are, \textcolor{bluez}{Figure 2} shows a histogram for the number of out-degree links from states. Twenty states have an out-degree of zero, so their tax rates have no direct impact on any other state's tax setting behavior. The most influential states in terms of the highest out-degree are Alabama (directly impacts tax setting behavior in five other states) and South Carolina, Pennsylvania, and Montana (which each directly impact tax setting behavior in four other states). Taking South Carolina as an example, the four states that it directly impacts include its geographic neighbor, Georgia, as well as non-geographic neighbors Missouri, Montana, and Virginia.
We use $\hat{W}_{econ}$ to shed light on the roles of factor mobility and yardstick competition in driving tax competition. To do so, we estimate the factors correlated with the existence of links between states $i$ and $j$ in $\hat{W}_{econ}$ relative to $\hat{W}_{geo}$. For state pairs with non-zero links in either $\hat{W}_{econ}$ or $\hat{W}_{geo}$, we define a dummy outcome $\hat{W}_{econ,ij}=1$ if a link between states $i$ and $j$ is estimated under $\hat{W}_{econ}$ and $\hat{W}_{econ,ij}=0$ if a link between states $i$ and $j$ exists under $\hat{W}_{geo}$ but not under $\hat{W} _{econ} $. We examine correlates of links using the following dyadic regression:
estimated using a linear probability model. The elements $X_{ij},$ $X_{i},$ and $X_{j}$ correspond to characteristics of the pair of states ($i,j$), state $i$, and state $j$, respectively. Covariates are time-averaged over the sample period, and robust standard errors are reported.
\textcolor{bluez}{Table 4} presents the dyadic regression results. \textcolor{bluez}{Column 1} controls only for the distance between states $i$ and $j$:$\ $this is highly predictive of an economic link between them. This reflects that the economic network of state $i$ often comprises states are in the same region, but not necessarily contiguous to state $i$. \textcolor{bluez}{Column 2} adds two $X_{ij}$ covariates to capture the economic and demographic homophily between states $i$ and $j$. GDP homophily is the absolute difference in the states' GDP per capita. Demographic homophily is the absolute difference in the share of young people (aged 5-17) plus the absolute difference of the share of elderly people (aged 65+) across the states. GDP\ homophily does not predict economic ties, whereas demographic homophily does.
\textcolor{bluez}{Columns 3-5} then sequentially add in several sets of controls. For labor mobility, we use net state-to-state migration data to control for the net migration flow of individuals from state $i$ to state $i$ (defined as the flow from $i$ to $j$ minus the flow from $j$ to $i$). \footnote{ We also experimented with alternative measures of labor migration, and the results were qualitatively the same. State-to-state migration data are based on year-to-year address changes reported on individual income tax returns filed with the IRS. The data cover filing years 1991 through 2015 and include the number of returns filed, which approximates the number of households that migrated, and the number of personal exemptions claimed, which approximates the number of individuals who migrated. The data are available at https://www.irs.gov/statistics/soi-tax-stats-migration-data (accessed September 2017).} We then add a political homophily variable between states. For any given year, this is set to one if a pair of states have governors of the same political party. As this is time-averaged over our sample, this element captures the share of the sample in which the states have governors of the same party. Lastly, we include whether state $j$ is considered a tax haven (and so might have disproportionate influence on other states). Based on Findleyetal2012, the following states are coded as tax havens: Nevada, Delaware, Montana, South Dakota, Wyoming, and New York.
\textcolor{bluez}{Column 5} shows that with this full set of controls, distance remains a robust predictor of the existence of economic links between states. However, the identified economic network highlights additional significant predictors of tax competition between states:\ political homophily reduces the likelihood of a link, suggesting any yardstick competition driving social interactions occurs when voters compare their governor to those of the opposing party in other states. Tax haven states appear to be especially less influential in the tax setting behaviors of other states. This mirrors what was observed in \textcolor{bluez}{Figure 2}, where some of the prominent tax havens -- Nevada, Delaware, and New York -- were all identified to have zero out-degree links. The relatively weak influence of tax haven states eases concerns over a race-to-the-bottom in tax setting behaviors.
\textcolor{bluez}{Column 6} controls for state $i$ and state $j$ fixed effects. This reinforces the idea that distance and political homophily correlate to the strength of influence states tax setting has on others (the tax haven dummy cannot be separately identified in this specification). Labor mobility between states does not robustly predict the existence of economic ties.
As in our identification result, our empirical approach has taken the network structure as fixed over the entire sample period. In the context of tax competition over our study period, this might be a strong assumption. We examine the issue in more detail by allowing the estimated $W_{econ}$ matrix to vary over time by changing the weight placed on observations from any given time period. More precisely, for any given time period $t$, we weight observations using a Gaussian kernel with its center varying period-by-period from $1962$ to $2015$. The variance of the kernel is set such that $75$% of the weight is given to the first half of the data (i.e., pre-$1988$) when the kernel is centered in $1962$. \textcolor{bluez}{Figure A4} shows the kernel employed as we vary its center: the solid kernel is centered in $1962$, the start of our sample -- when we place the most weight on observations from $1962$. The static case considered previously is akin to using a uniform kernel over all periods. This kernel weighting approach is outlined in Section (ref). We fully describe the algorithm in Appendix Section C.2.
We begin by considering time-varying estimates of the endogenous and exogenous social interaction parameters from the full model in ((ref)). The results for the endogenous social interactions parameter are shown in \textcolor{bluez}{Figure 3}, where the shaded areas are the $95$% confidence intervals of the period-by-period estimates. Panel A shows that OLS\ estimates of $\widehat{\rho }$ drift up over time, so the strength of endogenous interactions increases from around $.35$ in the late 1960s to $.50 $ by the 2010s. In all periods, we can reject the null that the endogenous social effect is zero. Recall the earlier static estimate was $ \widehat{\rho }_{OLS}=.375$.
Panel B shows the estimated endogenous social effect when we use GMM based on the characteristics of all other states as IVs. This also drifts up from around $.35$ in the late 1960s to $.50$ by the 2010s. In the majority of periods, we can reject the null of no endogenous social effect.\footnote{ Standard errors are estimated without imposing the restriction that parameters vary slowly over time, and fluctuations across periods reflect variations in the network across periods.}
\textcolor{bluez}{Figure 4} shows the evolution of $\hat{W}_{econ}$ over time as we center the kernel on different periods, following the same color-coding as in \textcolor{bluez}{Figure 1}. In all periods, geography-based edges play little role, and over time, the economic network becomes denser. This highlights\ not only that economic networks for tax competition always differ starkly from geography-based networks, but that the nature of economic networks relevant to tax competition has changed steadily over time.
\textcolor{bluez}{Figure 5} shows how the features of $\hat{W}_{econ}$ evolve as we place greater weight from early to later periods. For each statistic, we plot the period-by-period estimate when we center the kernel in any given period. The resulting smoothed estimates are then shown. To ease exposition, networks edges with $W_{0,ij}<1/47$ are removed. This cutoff is chosen as, in theory, states can only link at maximum with $47$ other states. \textcolor{bluez}{Panel A} shows the share of edges that are kept from the previous estimate (centered in the previous period). We see relatively high stability in $\hat{W}_{econ}$ with the smoothed estimate suggesting more than $60$% of edges always being kept from one estimate to the next, with this stability increasing from the late 1980s.
\textcolor{bluez}{Panel B} shows how the overlap between $\hat{W}_{econ}$ and $W_{geo}$ varies over time, as measured by the share of edges that are only present in $\hat{W}_{econ}$. There is little overlap between the two networks over the entire sample. The smoothed estimate suggests that at least $80$% of identified edges in $\hat{W}_{econ}$ are never in $W_{geo}$. The divergence between economic and geographic neighbors becomes starker from the mid-$1980$s onwards.
\textcolor{bluez}{Panels C and D} show how the clustering and reciprocity of links in $\hat{W}_{econ}$ vary as we shift the weight to later observations. Clustering of $\hat{W}_{econ}$ increases from the $1960$s through to the early $2000$s. Thereafter, social interactions in tax competition become sparser. We also observe a reversal in the extent to which social interactions are reciprocal, with reciprocity rising to a peak in the early 1980s -- when $20$% of ties were reciprocal -- and slowly falling thereafter.
Taken together, the results suggest the nature of tax competition between US\ states has changed over time through two mechanisms: (i) the strength of endogenous social interactions ($\widehat{\rho }$) has increased over time and (ii) the\ network of states interacted with ($\hat{W}_{econ}$) varies over time. This has important implications for policy evaluation:\ the same intervention might have different spillover effects if implemented at different moments in time due to the evolution of $\widehat{\rho }$ and $ \hat{W}_{econ}$. We consider this next using counterfactual simulations.
We use a counterfactual exercise to contrast how shocks to tax setting in a given state propagate under $\hat{W}_{econ}$, relative to what would have been predicted under $W_{geo}$. We do so for both static and dynamic estimates of $\hat{W}_{econ}$. We focus on South Carolina (SC), a state with one of the highest out-degree, as shown in \textcolor{bluez}{Figure 2}. We consider a scenario in which SC exogenously increases its taxes per capita by $10$%. We measure the differential change in equilibrium state taxes in state $j$ under the two network structures using the following statistic:
so that positive (negative) values imply equilibrium taxes being higher (lower) under $\hat{W}_{econ}$.\footnote{ For $W_{geo}$, we calculate the counterfactual at $\hat{\rho}_{GMM}=.452$, the endogenous effect parameter estimated in our preferred specification, \textcolor{bluez}{Column 4 of Table 2}.}
Starting with the static case, \textcolor{bluez}{Panel A of Figure 6} shows for each mainland US state the spillover effects through the economic network of tax competition. This highlights positive spillovers on tax rates in many states that are not geographic neighbors of SC. \textcolor{bluez}{Panel B} graphs $\Upsilon _{j}$ to make precise how spillovers derived from $\hat{W }_{econ}$ diverge from those predicted under $W_{geo}$. In $26$ states, $ \Upsilon _{j}$ is smaller than $.01$% because both networks predict negligible spillovers to those states. In the remaining $22$ mainland states, there is a wide discrepancy between the equilibrium state tax rates predicted under $\hat{W}_{econ}$ relative to $W_{geo}$: $\Upsilon _{j}$ varies from $-1$ to $4.03$. The long-run effect in SC itself is also higher under $\hat{W}_{econ}$ than under $W_{geo}$. The former states that given feedback effects, the long-run increase in tax rates in SC from a $10$% increase is $ 11.4$%, while the geographically based network implies a smaller equilibrium increase of $10.3$%.
As $\hat{W}_{econ}$ is spatially more dispersed than $W_{geo}$, the general equilibrium effects are different under the two network structures. \textcolor{bluez}{Table 5} summarizes the general equilibrium implications for tax inequality under $\hat{W}_{econ}$ and the $W_{geo}$ counterfactual. The average tax rate increase under $\hat{W}_{econ}$ is three times that estimated under $W_{geo}$. Moreover, the dispersion of tax rates across states increases under $\hat{W}_{econ}$ relative to $W_{geo}$. Finally, assuming interactions are based solely on geographic neighbors, we miss the fact that many states have relatively small tax increases.
We can repeat the exercise using the dynamically estimated economic network. Throughout, we calculate the general equilibrium effects of the same policy experiment: SC increasing its taxes per capita by $10$%. These general equilibrium effects vary over the sample period because the strength of social interactions in tax competition vary ($\widehat{\rho }_{GMM}$), as shown in \textcolor{bluez}{Figure 3}, and identified economic neighbors vary over time ($\hat{W}_{econ}$), as shown in \textcolor{bluez}{Figure 4}. The results are summarized in \textcolor{bluez}{Figure 7}. Placing weight on the early or later part of the sample generates similar changes in average tax rates and their variance in general equilibrium -- with both being lower than simulated under the static model. Placing more weight on the middle of the sample period generates higher changes in average tax rates and their variance in general equilibrium.
The differential general equilibrium impacts found as we place different weights across sample observations links to recent discussions on the external validity of internally valid causal impacts based on micro-evidence. While the earlier literature has emphasized the potential interaction of treatment effects with aggregate shocks (rosenzweig2020external) or how behavioral responses to social insurance policies vary over the business cycle (kroft2016should), our analysis provides another explanation for the changing impacts of policies where social interactions determine behavior:\ changes in the strength of social interactions and the network of economic interactions.
In a canonical social interactions model, we provide sufficient conditions under which the social interactions matrix, and endogenous and exogenous social effects are all globally identified, even absent information on social links. Our identification strategy is novel and may bear fruit in other areas. The method is immediately applicable to other classic social interactions problems, but where data on social links is either missing or partial. In fields such as macroeconomics, political economy, and trade, there are core areas of research where social interactions across jurisdictions/countries etc. drive key outcomes, panel data exist over many periods, and the number of nodes is relatively fixed. Moreover, while our discussion and application have focused on a continuous policy response (state taxes), our methods can also be applied to the extensive margin of policy adoption and diffusion. Such diffusion models might generate network interactions where some states influence the later adoption of economic and social policies in other jurisdictions. This issue is studied by dellavigna2022policy in the context of US\ state policies -- they examine the diffusion of over $700$ policies in the past $70$ years. Their work also suggests the nature of interactions across states has changed:\ while geographic proximity is a good predictor of policy diffusion, they also find that since $2000$, political alignment across states has become the strongest predictor of diffusion.
In finance, high-frequency panel data is readily available and relevant for the study of core research questions. For example, a long-standing question has been whether CEOs are subject to relative performance evaluation, and if so, what is the comparison set of firms/CEOs used (edmansgabaix2016 ). More generally, our method can be readily applied to a large class of economic questions around contagion, risk, and the fragility of economic and environmental systems. For example, since the financial crisis of 2008, it has become clear that linkages between actors such as firms or banks are complex and often hidden, yet because endogenous network interactions cause feedback loops and have multiplier effects, they can have enormous implications for the evolution of a financial crisis or the propagation of supply shocks in aggregate. Identifying such synchronicity is a critical first step to putting in place policies to reduce the fragility of economic systems ( vanVliet2018; elliott2022networks; goldstein2022synchronicity).
Advances in the availability of administrative data, data from social media, mobile technologies, and online economic transactions all offer new possibilities to identify social interactions with long panels or high-frequency data collection, where data on social ties will typically be missing.
Three further directions for future research are of note. First, under partial observability of $W_{0}$ (as in Blumeetal2015), the number of parameters in $W_{0}$ to be retrieved falls quickly. Our approach can then still be applied to complete knowledge of $W_{0}$, such as if Aggregate Relational Data is available, and this could be achieved with potentially weaker assumptions for identification, and in even shorter panels. To illustrate possibilities, \textcolor{bluez}{Figure A5} shows results from a final simulation exercise in which we assume the researcher starts with partial knowledge of $W_{0}$. We do so for Banerjee et al. (2013) \nocite{banerjeeetal2013} village family network, showing simulation results for scenarios in which the researcher knows the social ties of the three (five, ten) households with the highest out-degree. For comparison, we also show the earlier simulation results when $W_{0}$ is entirely unknown. This clearly illustrates that with partial knowledge of the social network, performance on all metrics improves rapidly for any given $T$.
Second, we have developed our approach in the context of the canonical linear social interactions model ((ref)). This builds on Manski1993 when $W_{0}$ is known to the researcher, and the reflection problem is the main challenge in identifying endogenous and exogenous social effects. However, the reflection problem is functional-form dependent and may not apply to many non-linear models (Blumeetal2011 , Blumeetal2015). An important topic for future research is to extend the analysis to non-linear social interaction settings. Relatedly, the canonical social interaction model assumes that the same $W_{0}$ governs the endogenous and exogenous channels. Despite the relaxation we propose in Section 2.3.3, we see this as a limitation of the current method, and future research is needed to allow for a fully flexible approach.
Finally, an important part of the social networks literature examines endogenous network formation (Jacksonetal2017; DePaula2017). Our analysis allows us to begin probing the issue in two ways. First, the kind of dyadic regression analysis in Section 4 on the correlates of entries in $W_{0,ij}$ suggests factors driving link formation and dissolution. Second, this leads naturally to a broad agenda going forward, to address the challenge of simultaneously identifying and estimating time varying models of network formation and social interaction, all in cases where data on social networks is not required.
\singlespacing