EconBase
← Back to paper

Identifying Network Ties from Panel Data: Theory and an Application to Tax Competition

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

124,471 characters

Identifying Network Ties from Panel Data: Theory and an Application to Tax Competition



\title{\textbf{Identifying Network Ties from Panel Data: Theory and an
Application to Tax Competition}\thanks{{\scriptsize We gratefully
acknowledge financial support from the ESRC through the Centre for the
Microeconomic Analysis of Public Policy (ES/T014334/1), the Centre for
Microdata Methods and Practice (RES-589-28-0001) and the Large Research
Grant ES/P008909/1 and from the ERC (SG338187). We thank Edo Airoldi, Luis
Alvarez, Michele Aquaro, Oriana Bandiera, Larry Blume, Yann Bramoull\'{e},
Stephane Bonhomme, Vasco Carvalho, Gary Chamberlain, Andrew Chesher,
Christian Dustmann, S\'{e}rgio Firpo, Jean-Pierre Florens, Eric Gautier, Giacomo
de Giorgi, Matthew Gentzkow, Stefan Hoderlein, Bo Honor\'{e}, Matt Jackson, Dale
Jorgensen, Christian Julliard, Maximilian Kasy, Miles Kimball, Thibaut
Lamadon, Simon Sokbae Lee, Arthur Lewbel, Tong Li, Xiadong Liu, Elena
Manresa, Charles Manski, Marcelo Medeiros, Angelo Mele, Francesca Molinari,
Pepe Montiel, Andrea Moro, Whitney Newey, Ariel Pakes, Eleonora Pattachini,
Michele Pelizzari, Martin Pesendorfer, Christiern Rose, Adam Rosen, Bernard
Salanie, Olivier Scaillet, Sebastien Siegloch, Pasquale Schiraldi, Tymon
Sloczynski, Kevin Song, John Sutton, Adam Szeidl, Thiago Tachibana, Elie
Tamer, and seminar and conference participants for valuable comments. We
also thank Tim Besley and Anne Case for comments and sharing data. Daniel
Barbosa provided outstanding research assistance. A previous version of
this paper was circulated as \textquotedblleft Recovering Social Networks
from Panel Data: Identification, Simulations and an
Application.\textquotedblright\ All errors remain our own. Codes are available on Zenodo (\texttt{https://zenodo.org/XXX}) and github (\texttt{https://github.com/YYY}) repositories.
}}}
\author{\'{A}ureo de Paula \and Imran Rasul \and Pedro CL Souza\thanks{
{\scriptsize de Paula: University College London, CeMMAP and IFS,
[email removed]; Rasul: University College London and IFS,
[email removed]; Souza: Queen Mary University, [email removed].}}.}
\date{October 2023}

\maketitle

\begin{abstract}
\noindent Social interactions determine many economic behaviors, but
information on social ties does not exist in most publicly available and
widely used datasets. We present results on the identification of social
networks from observational panel data that contains no information on
social ties between agents. In the context of a canonical social
interactions model, we provide sufficient conditions under which the social
interactions matrix, endogenous and exogenous social effect parameters are
globally identified if networks are constant over time. We also provide an
extension of the method for time-varying networks. We then describe how
high-dimensional estimation techniques can be used to estimate the
interactions model based on the Adaptive Elastic Net Generalized Method of
Moments. We employ the method to study tax competition across US states. The identified social interactions matrix implies that tax competition
differs markedly from the common assumption of competition between
geographically neighboring states, providing further insights into the
long-standing debate on the relative roles of factor mobility and yardstick
competition in driving tax setting behavior across states. Most broadly, our
identification and application show that the analysis of social interactions
can be extended to economic realms where no network data exists.\newline
\textit{JEL Classification: C31, D85, H71.}$\smallskip \smallskip \smallskip
\smallskip \smallskip \smallskip $
\end{abstract}

\section{Introduction}

In many economic environments, behavior is shaped by social interactions
between agents. In individual decision problems, social interactions have
been key to understanding outcomes as diverse as educational test scores,
the demand for financial assets, and technology adoption (
\citealp{Sacerdote2001}; \citealp{Bursztynetal2014}; \citealp{ConleyUdry2010}
). In macroeconomics, the structure of firms' production and credit networks
propagate shocks, or help firms to learn (\citealp{Acemogluetal2012};
\citealp{Chaney2014}). In political economy and public economics, ties between jurisdictions are
key to understanding tax setting behavior (\citealp{Tiebout1956};
\citealp{Shleifer1985}; \citealp{BesleyCase1994}).

Underpinning all these bodies of research is some measurement of the
underlying social ties between agents. However, information on social ties
does not exist in most publicly available and widely used datasets. To
overcome this limitation, studies of social interaction either \emph{
postulate} ties based on common observables or homophily, or \emph{elicit}
data on networks. However, it is increasingly recognized that postulated and
elicited networks remain imperfect solutions to the fundamental problem of
missing data on social ties, because of econometric concerns that arise with
either method, or simply because of the cost of collecting network data.
\footnote{
As detailed in \citet{DePaula2017}, elicited networks are often
self-reported and can introduce error to the outcome of interest. Network
data can be censored if only a limited number of links can feasibly be
reported. Incomplete survey coverage of nodes in a network may lead to
biased aggregate network statistics. \citet{ChandrasekharLewis2016} show
that even when nodes are randomly sampled from a network, partial sampling
leads to non-classical measurement error and biased estimation. Collecting
social network data is also a time- and resource-intensive process. In
response to these concerns, a nascent strand of literature explores
cost-effective alternatives to full elicitation to recover aggregate network
statistics (\citealp{Brezaetal2017}).}

Two consequences are that (i) the classes of problems in which social
interactions occur are understudied, because social networks data is missing
or too costly to collect; and (ii) there is no way to validate social
interactions analysis in contexts where ties are postulated. In this paper,
we tackle this challenge by deriving sufficient conditions under which
global identification of the \emph{entire structure} of social networks\ is
obtained, using only observational panel data that itself contains \emph{no}
information on network ties. Our identification results allow the study of
social interactions without data on social networks, and the validation of
structures of social interaction where social ties have hitherto been
postulated. The recovered networks are economically meaningful to explain
the effects under study, since they are entirely estimated from the data
itself, and not driven by \emph{ex ante} assumptions on how individuals
interact.

A researcher is assumed to have panel data on individuals $i=1,...,N$ for
instances $t=1,...,T$. An instance refers to a specific observation for $i$
and need not correspond to a time period (for example, if $i$ refers to a
firm, $t$ could refer to market $t$). The outcome of interest for individual
$i$ in instance $t$ is $y_{it}$ and is generated according to a canonical
structural model of social interactions:\footnote{\citet{Blumeetal2015}
present micro-foundations for this estimating equation based on
non-cooperative games of incomplete information for individual choice
problems.}
\begin{equation}
y_{it}=\rho _{0}\sum_{j=1}^{N}W_{0,ij}y_{jt}+\beta _{0}x_{it}+\gamma
_{0}\sum_{j=1}^{N}W_{0,ij}x_{jt}+\alpha _{i}+\alpha _{t}+\epsilon _{it}.
\label{model motivation}
\end{equation}
Outcome $y_{it}$ depends on the outcomes of other individuals to whom $i$ is
socially tied, $y_{jt}$, and $x_{jt}$ includes characteristics of those
individuals.\footnote{
In the case in which $t$ is considered to be a time period, $x_{it}$ may
also include lagged values of $y_{it}$.} $W_{0,ij}$ measures how the outcome
and characteristics of $j$ causally impact the outcome for $i$. The network
is initially assumed to be fixed over time, and we later provide an extension of
the method for time-varying networks. As outcomes for all individuals obey
equations analogous to (\ref{model motivation}), the system of equations can
be written in matrix notation, where the structure of interactions is
captured by the adjacency matrix, denoted by $W_{0}$. Our approach allows
for unobserved heterogeneity across individuals $\alpha _{i}$ and common
shocks to individuals $\alpha _{t}$. This framework encompasses\ a classic
linear-in-means specification as in \citet{Manski1993}. In his terminology, $
\rho _{0}$ and $\gamma _{0}$ capture endogenous and exogenous social
effects, and $\alpha _{t}$ captures correlated effects. The distinction
between endogenous and exogenous peer effects is critical, as only the
former generates social multiplier effects. In line with the literature, we
maintain that the same $W_{0}$ governs the structure of both endogenous and
exogenous effects. We later discuss relaxing this assumption when more than
one regressor is used.

Manski's seminal contribution set out the reflection problem of separately
identifying endogenous, exogenous, and correlated effects in linear models.
However, it has been somewhat overlooked that he also set out another
challenge in the identification of the social network in the first place.
\footnote{\citet{Manski1993} highlights difficulties (and potential
restrictions) in identifying $\rho _{0},\beta _{0}$ and $\gamma _{0}$ when
\emph{all} individuals interact with each other, and when this is observed by the
researcher. In (\ref{model motivation}), this corresponds to $
W_{0,ij}=N^{-1} $, for $i,j=1,\dots ,N$. At the same time, he states (p.
536), \textquotedblleft I have presumed that researchers know how
individuals form reference groups and that individuals correctly perceive
the mean outcomes experienced by their supposed reference groups. There is
substantial reason to question these assumptions (...) If researchers do not
know how individuals form reference groups and perceive reference-group
outcomes, then it is reasonable to ask whether observed behavior can be used
to infer these unknowns (...) The conclusion to be drawn is that informed
specification of reference groups is a necessary prelude to analysis of
social effects.\textquotedblright} This is the problem we tackle, and thus, we
expand the scope of identification beyond $\rho _{0}$, $\beta_{0}$, and $\gamma_{0}$. Our point of departure from much of the literature is
therefore to presume $W_{0}$ is \emph{entirely unknown} to the researcher. We
derive sufficient conditions under which all the entries in $W_{0}$, and the
endogenous and exogenous social effect parameters, $\rho _{0}$ and $\gamma
_{0},$ are globally identified from \textquotedblleft reduced
form\textquotedblright\ parameters. By identifying the social interactions
matrix $W_{0}$, our results allow the recovery of aggregate network
characteristics, such as the degree distribution and patterns of homophily,
as well as node-level statistics such as the strength of social interactions
between nodes,   and the centrality of nodes. Such aggregate and node-level
statistics often map back to underlying models of social interaction (
\citealp{Ballesteretal2006}; \citealp{Jacksonetal2017}; \citealp{DePaula2017}
).

Our identification strategy is new and fundamentally different from those
employed elsewhere in the literature and does not rely on requirements about
network sparsity. However, it delivers sufficient conditions that are mild
and relate to existing results on the identification of social effects
parameters when $W_{0}$ is known (\citealp{Bramoulleetal2009};
\citealp{DeGiorgietal2010}; \citealp{Blumeetal2015}). The intuition for our
identification result is simple: model (1) has $N^{2}$ reduced-form
parameters, and there are $N(N-1)+3$ structural unknowns (as no unit affects
itself, so $W_{0,ii}=0$). So there are more equations than unknowns if $
N\geq 2$, and we demonstrate those can be solved for the parameters of
interest under the assumptions we invoke. Our identification result is also
useful in other estimation contexts, such as when a researcher has partial
knowledge of $W_{0}$,\footnote{
One such example is the nascent literature of Aggregate Relational Data
(ARD) as in \citet{Brezaetal2017}. Another possibility is that individuals
are known to belong to subgroups, so $W_{0}$ is block diagonal.
} or in navigating between priors on reduced-form and structural parameters
in a Bayesian framework (see, e.g., \citealp{gefanghalltavlas2023}), thus avoiding issues the raised by
\citet{KlineTamer2016}.

Global identification is a necessary requirement for consistency of extremum
estimators such as those based on the GMM (Hansen 1982; Newey and McFadden
1994). Our identification analysis provides primitives for this condition.
To estimate the model, we employ the Adaptive Elastic Net GMM method (
\citealp{CanerZhang2014}), as this allows us to deal with a potentially
high-dimensional parameter vector (in comparison to the time dimension in
the data) including all the entries of the social interactions matrix $W_{0}$
, although other estimation protocols may also be entertained (e.g. using
Bayesian methods or \emph{a priori} information).\footnote{The Elastic Net was introduced by \citet{zou2005regularization} in part to circumvent difficulties faced by alternative estimation protocols (e.g., LASSO) when the number of parameters, $p$, exceeds the number of observations, $n$ (where $p$ and $n$ follow the notation in that paper).  Whereas the theoretical results on the large-sample properties of elastic net estimators usually have not exploited sparsity, several articles have demonstrated their performance in data scenarios where this occurs.
 In Section 3, we provide an informal discussion on the performance in our context. \label{footnote:elasticnet}}


We showcase the method using Monte Carlo simulations based on stylized
random network structures as well as real-world networks. In each case, we
take a fixed network structure $W_{0}$ and simulate panel data as if the
data generating process were given by (\ref{model motivation}). We then
apply the method to the simulated panel data to recover estimates of all
elements in $W_{0}$, as well as the endogenous and exogenous social effect
parameters ($\rho _{0}$, $\gamma _{0}$). The networks considered vary in
size, complexity, and their aggregate and node-level features. In small
samples, we find that the majority of links are identified even for $T=5$,
and the proportion of true non-links (zeros in $W_{0}$) captured correctly
as zeros is over $85$\% even when $T=5$. Of course, there are important
limitations to the use of the method in small-$T$ cases. Biases are expected
and manifest themselves in two ways. First, weak links can be shrunk to
zero, and the strength of strong edges can be overestimated. Second, the estimates of $\rho$ and $\gamma$ can
suffer from small-sample bias, being analogous to well-known results for
autoregressive time series models. Both properties rapidly improve with $T$.
For instance, biases in the estimation of endogenous and exogenous effects
parameters ($\hat{\rho},$ $\hat{\gamma}$) fall quickly with $T$ and are
close to zero for large sample sizes. The endogenous and exogenous social
effects are also correctly captured as $T$ increases. \emph{A fortiori}, we
estimate aggregate and node-level statistics of each network, demonstrating
the accurate recovery of key players in networks, for example.

In the final part of our analysis, we apply the method to shed new light on
a classic real-world social interactions problem:\ tax competition between
US\ states. The literatures in political economy and public economics have
long recognized the behavior of state governors might be influenced by
decisions made in ``neighboring'' states. The typical empirical approach has
been to postulate the relevant neighbors as being geographically contiguous
states. Our approach allows us to infer the set of ``economic'' neighbors
determining social interactions in tax setting behavior from panel data on
outcomes and covariates alone. In this application, the panel data
dimensions cover mainland US states, $N=48$, for the years 1962-2015, $T=53$.

The identified network structure of tax competition differs markedly from
the assumption of competition between geographic neighbors. The identified
economic network has fewer edges, and we identify non-adjacent states that
influence tax setting behaviors. Differences in the structure of the identified
economic and geography-based networks are reflected in\ the far lower
clustering coefficient in the former ($.042$ versus $.419$). With the
recovered social interactions matrix we establish, beyond geography, which
covariates correlate to the existence of ties between states and so shed new
light on hypotheses for social interactions in tax setting: factor mobility
and yardstick competition (\citealp{Tiebout1956}; \citealp{Shleifer1985};
\citealp{BesleyCase1994}). The identified network highlights significant
predictors of tax competition between states beyond distance:\ political
homophily \emph{reduces} the likelihood of a link, suggesting any yardstick
competition driving social interactions occurs when voters compare their
governor to those of the opposing party in other states. Tax haven states
appear to be less influential in tax setting behaviors, easing concerns
over a race-to-the-bottom in tax setting. Labor mobility between states does
not robustly predict the existence of economic ties between states in tax
setting behavior.

Given the relatively long study period in this application, at a final stage
of analysis we extend our method to allow the strength of social
interactions in tax competition ($\rho _{0}$, $\gamma _{0}$) and the
structure of links in the economic network ($W_{0}$) to vary over time as we
change the weight placed on observations from any given time period. We
document the gradual increase in strength of social interactions over time,
and the changing nature of the network of interactions. We utilize these
findings to conduct counterfactual simulations of the general equilibrium
propagation of tax shocks from a given state to all other mainland US\
states, and how these general equilibrium effects of the same policy shock
vary as we place weight on observations later in our study period.

Our paper contributes to the literature on the identification of social
interactions models. The first generation of papers studied the case where $
W_{0}$ is known, so only the endogenous and exogenous social effects
parameters needed to be identified. It is now established that if the known $
W_{0}$ differs from the linear-in-means example where all units are linked
with equal weights, $\rho _{0}$ and $\gamma _{0}$ can be identified (
\citealp{Bramoulleetal2009}; \citealp{DeGiorgietal2010}). Intuitively,
identification in those cases can use peers-of-peers, are not
necessarily connected to individual $i$ and can be used to leverage
variation from exclusion restrictions in (\ref{model motivation}), or can
use groups of different sizes within which all individuals interact with
each other (\citealp{Lee2007}). \citet{Bramoulleetal2009} show these
conditions are met if $I,$ $W_{0}$, and $W_{0}^{2}$ are linearly independent,
which is shown to hold generically by \citet{Blumeetal2015}. However, as
made precise in Section \ref{sec:Identification}, the linear algebraic
arguments employed by \citet{Bramoulleetal2009} or \citet{Blumeetal2015} do
not apply when $W_{0}$ is unobserved, and other arguments have to be used
instead.\footnote{
Alternative identification approaches when $W_{0}$ is known focus on higher
moments (variances and covariances across individuals) of outcomes (
\citealp{DePaula2017}) and rely on additional restrictions on higher
moments of $\epsilon _{it}$. Note that (\ref{model motivation}) is a spatial
autoregressive model. In that literature, $W_{0}$ is also typically assumed to be
known (\citealp{Anselin2010}).}

\citet{Blumeetal2015} investigate the case when $W_{0}$ is \emph{partially}
observed and show that if
two individuals are \emph{known} not to be directly connected, the
parameters of interest in a model related to (\ref{model motivation}) can be
identified. Blume \emph{et al.} (2011) take an alternative approach: suggesting a parameterization of $W_{0}$ according to a
pre-specified distance between nodes. We do not impose such restrictions,
but note that partial observability of $W_{0}$ or placing additional structure on $W_{0}$ is complementary to our approach, as it reduces the number of
parameters in $W_{0}$ to be retrieved. \citet{BonaldiHortacsuKastl2015} and
\citet{Manresa2016} estimate models like (\ref{model motivation}) when $
W_{0} $ is not observed, but where $\rho _{0}$ is set to zero so there are
no endogenous social effects. They use sparsity-inducing methods from the
statistics literature, but the presence of $\rho _{0}$ in our case
complicates identification because it introduces issues of simultaneity that
we address.\footnote{\citet{Manresa2016} allows for unit-specific $\beta
_{0} $ parameters. While in many applications those are taken to be
homogeneous, we also discuss extensions on how heterogeneity in those
parameters can be handled when $\rho _{0}\neq 0$ in Appendix B.}

\citet{Rose2015} also presents related identification results for linear
models like (\ref{model motivation}), assuming the sparsity of the neighborhood
structure. Intuitively,
given two observationally equivalent systems, sparsity guarantees the
existence of pairs that are not connected in either. Since observationally
equivalent systems are linked via the reduced-form coefficient matrix, this
pair allows one to identify certain parameters in the model. Having
identified those parameters, \citet{Rose2015} shows that one can proceed to
identify other aspects of the structure (see also \citealp{GautierRose2016}
). This is related to the ideas in Blume \emph{et al.} (2015),
who show identification results can be leveraged if individuals are \emph{
known} not to be connected. Our main identification results do $not$ rely on
properties of sparse networks, and make use of plausible and intuitive
conditions, whereas the auxiliary rank conditions necessary may be computationally complex to verify.
\textcolor{black}{More recently,
\citet{Lewbeletal2019} propose an estimation strategy for the parameters
$\rho_0$, $\beta_0$, and $\gamma_0$ of model \eqref{model motivation} in the
absence of network links if many different groups can be observed.}
\textcolor{black}{\citet{Battaglinietal2019} estimate a structural model
specifically for the case of unobserved social connections in the US
Congress.}

Finally, in the statistics literature, \citet{LamSouza2019} study the
penalized estimation of model \eqref{model motivation} when $W_{0}$ is not
observed, assuming the model and social interactions are identified. The
statistical literature on graphical models has investigated the estimation
of neighborhoods defined by the covariance structure of the random variables
at hand (\citealp{MeinshausenBuhlmann2006}). This corresponds to a model
where $y_{t}=(I-\rho _{0}W_{0})^{-1}\epsilon _{t}$ is jointly normal
(abstracting from covariates). On a graph with $N$ nodes corresponding to
the variables in the model, an edge between two nodes (variables) $i$ and $j$
is absent when these two variables are conditionally independent given the
other nodes. In the
model above, the inverse covariance matrix is $(I-\rho _{0}W_{0})^{\top
}\Sigma _{\epsilon }^{-1}(I-\rho _{0}W_{0})$, where $\Sigma _{\epsilon }$ is
the variance covariance structure for $\epsilon _{t}$. The discovery of zero
entries in this matrix is not equivalent to the identification of $W_{0}$
and involves $\Sigma _{\epsilon }$ (as do identification strategies using
higher moments when $W_{0}$ is known).\footnote{
\citet{MeinshausenBuhlmann2006}'s and \citet{LamSouza2019}'s neighborhood
estimates rely on (penalized) regressions of $y_{it}$ on $y_{1t},\dots
,y_{i-1,t},y_{i+1,t},\dots ,y_{N,t}$, which do not address the endogeneity
in estimating $W_{0}$.}

We build on these papers by studying the problem where $W_{0}$ is
potentially entirely unknown to the researcher. In so doing, we open up the
study of social interactions to realms where social network data does not
exist. In our case, we consider the definition of the network as the one
that mediates, together with the variables $x_{it}$, the outcome process $
y_{it}$ according to Equation \eqref{model motivation}. The identified
network may be a combination of elicited types of social interactions --
such as friendship formation, lending and borrowing relations, links with
relatives -- or different from elicited data, as long as the links are
relevant in determining the outcomes. In our case, and in line with the
literature, the network ties $W_{ij}$ are considered to be deterministic
parameters or predetermined. Alternatively, the networks are assumed to be
the outcome of a stochastic process, such as the latent space model (
\citealp{hoff2002latent}; \citealp{Brezaetal2017}) or Exponential Random
Graphs models (\citealp{HollandLeinhardt1981}).

Our conclusions discuss how our approach can be modified, and assumptions
weakened, to integrate partial knowledge of $W_{0}$. We discuss further
applications and the steps required to simultaneously identify models of
network formation and the structure of social interactions. The practical
use of our proposed method has already been demonstrated in applications.
For example, \citet{fetzeretal2020} study the impact on conflict of the
transition of security responsibilities between international and Afghan
forces. Our proposed method is used to control for violation of SUTVA-type
hypotheses that might occur because of spillover and displacement effects of
insurgent forces across districts. Since the pattern of displacement is
unobserved -- and, in fact, insurgents have incentives to obfuscate their
strategy -- the current method is applied to fully recover the network and
bound the effects of the end of the military occupation on conflict.
\footnote{\citet{Zhou2019} applies our identification results, focusing on
unobserved networks with grouped heterogeneity, to suggest a nonlinear least
squares procedure for estimation on a single network observation.
}

We proceed as follows. Section 2 presents our core result:\ the sufficient
conditions under which the social interactions matrix, endogenous and
exogenous social effects are globally identified. Section 3 describes the
high-dimensional techniques used for estimation based on the
Adaptive Elastic Net GMM method and presents simulation results from
stylized and real-world networks. Section 4 applies our methods to study tax
competition between US\ states. Section 5 concludes. The Appendix provides
proofs and further details on estimation and simulations.

\section{Identification\label{sec:Identification}}

\subsection{Setup\label{subsec:setup}}

Consider a researcher with panel data covering $i=1,\dots ,N$ individuals
repeatedly observed over $t=1,\dots ,T$ instances. The number of individuals
$N$ in the network is fixed but potentially large. The aim is to use this
data to identify a social interactions model with no data on actual social
ties. For expositional ease, we first consider identification in a simpler
version of the canonical model in (\ref{model motivation}), where we drop
individual-specific ($\alpha _{i}$) and time-constant fixed effects ($\alpha
_{t}$) and assume $x_{it}$ is a one-dimensional regressor for individual $i$
and instance $t$. We later extend the analysis to include
individual-specific, time-constant fixed effects and allow for
multidimensional covariates $x_{k,it}$, $k=1,\dots ,K$. We adopt the
subscript \textquotedblleft 0\textquotedblright\ to denote parameters
generating the data, and non-subscripted parameters are generic values in
the parameter space:
\begin{equation}
y_{it}=\rho _{0}\sum_{j=1}^{N}W_{0,ij}y_{jt}+\beta _{0}x_{it}+\gamma
_{0}\sum_{j=1}^{N}W_{0,ij}x_{jt}+\epsilon _{it}.  \label{eq:modelSF}
\end{equation}
As the outcomes for all individuals $i=1,\dots ,N$ obey equations analogous to (
\ref{eq:modelSF}), the system of equations can be more compactly written in
matrix notation as:
\begin{equation}
y_{t}=\rho _{0}W_{0}y_{t}+\beta _{0}x_{t}+\gamma _{0}W_{0}x_{t}+\epsilon
_{t}.  \label{eq:Model}
\end{equation}
The vector of outcomes $y_{t}=(y_{1t},\dots ,y_{Nt})^{\prime }$ assembles
the individual outcomes in instance $t$; the vector $x_{t}$ does the same
with individual characteristics. $y_{t}$, $x_{t}$, and $\epsilon _{t}$ have
dimension $N\times 1$, the social interactions matrix $W_{0}$ is $N\times N$
, and $\rho _{0}$, $\beta _{0}$, and $\gamma _{0}$ are scalar parameters. We
do not make any distributional assumptions on $\epsilon _{t}$ beyond $
\mathbb{E}(\epsilon _{t}|x_{t})=0$ (or $\mathbb{E}(\epsilon _{t}|z_{t})=0$
for an appropriate instrumental variable $z_{t}$ if $x_{t}$ is endogenous).
We assume the network structure is predetermined and constant, and that the
number of individuals $N$ is fixed and repeated. In reality, networks may
evolve over time. We thus later expand the method for dynamic network cases.
The network structure $W_{0}$ is a parameter to be identified and estimated.

The social interaction model \eqref{eq:Model} has been widely studied (
\citealp{Manski1993}; \citealp{Manresa2016}; and \citealp{Blumeetal2015}, among many
others), but it is also restrictive in at least two senses. First, we
consider $W_{ij}$ to be fixed and predetermined, and not through models
of strategic network formation (\citealp{jackson1996strategic};
\citealp{dePaulaRichardsTamer2018}) or of stochastic nature, as in the class
of Exponential Random Graphs (\citealp{HollandLeinhardt1981}) or Latent
Distance models (\citealp{hoff2002latent}; \citealp{Brezaetal2017}). If
there is feedback between outcome determination and link formation, and
especially if this involves unobservables, it would be important to model
network formation more explicitly.

A regression of outcomes on covariates corresponds, then, to the reduced
form for (\ref{eq:Model}),
\begin{equation}
y_{t}=\Pi _{0}x_{t}+\nu _{t},  \label{eq2:modelRF}
\end{equation}
with $\Pi _{0}=(I-\rho _{0}W_{0})^{-1}(\beta _{0}I+\gamma _{0}W_{0})$ and $
\nu _{t}\equiv (I-\rho _{0}W_{0})^{-1}\epsilon _{t}$.   If $W_{0}$ is observed, \citet{Bramoulleetal2009} note that a structure $
(\rho ,\beta ,\gamma )$ that is observationally equivalent to $(\rho
_{0},\beta _{0},\gamma _{0})$ is such that $(I-\rho _{0}W_{0})^{-1}(\beta
_{0}I+\gamma _{0}W_{0})=(I-\rho W_{0})^{-1}(\beta I+\gamma W_{0})$. This can
be written as a linear equation in $I,W_{0}$, and $W_{0}^{2}$, and
identification is established if those matrices are linearly independent. If
$W_{0}$ is not observed, the putative unobserved structure comprises $W_{0}$,
and an observationally equivalent parameter vector will instead satisfy $
(I-\rho _{0}W_{0})^{-1}(\beta _{0}I+\gamma _{0}W_{0})=(I-\rho W)^{-1}(\beta
I+\gamma W)$. Following the strategy in \citet{Bramoulleetal2009} would lead
to an equation in $I,W,W_{0}$, and $WW_{0}$, so the insights obtained in
that paper do \emph{not} carry over to the case we study when $W_{0}$
is unknown.

We establish identification of the structural parameters of the model,
including the social interactions matrix $W_{0}$, from the coefficients
matrix $\Pi _{0}$. Without data on the network $W_{0}$, we treat it as an
additional parameter in an otherwise standard model relating outcomes and
covariates. Our identification strategy relies on how changes in covariates $
x_{it}$ reverberate through the system and impact $y_{it}$, as well as
outcomes for other individuals. These are summarized by the entries of the
coefficient matrix $\Pi _{0}$, which, in turn, encode information about $
W_{0}$ and $(\rho _{0},\beta _{0},\gamma _{0})$. A non-zero partial effect of $
x_{it}$ on $y_{jt}$ indicates the existence of direct \emph{or} indirect
links between $i$ and $j$. When $\rho _{0}=0$ (and $\Pi _{0}=\beta
_{0}I+\gamma _{0}W_{0}$), only direct links  produce such a
correlation. When $\rho \neq 0$, both direct and indirect connections may
generate a non-zero response, but distant connections will lead to a lower
response. Our results formally determine sufficient conditions to precisely
disentangle these forces.

We set out six assumptions underpinning our main identification results.
Three of these are entirely standard. A fourth is a normalization required
to separately identify ($\rho _{0}$, $\gamma _{0}$) from $W_{0}$, and the
fifth is closely related to known results on the identification of ($\rho
_{0}$, $\gamma _{0}$) when $W_{0}$ is known (\citealp{Bramoulleetal2009}).
The sixth assumption pertains to the relation between the nature of repeated
multiple observations of the outcome and covariates and restrictions on the
stability of $W$. These Assumptions (A1-A6) deliver an identified set of up
to two points.

Our first assumption explicitly states that no individuals affect
themselves and is a standard condition in social interaction models:

\begin{itemize}
\item[(A1)] $(W_{0})_{ii}=0$, $i=1,\dots ,N$.
\end{itemize}

Assumption (A1) rules out applications with self-influence. For example,
Input-Output matrices typically feature $(W_{0})_{ii}>0$, as firms tend to
source from other firms in the same industry. With Assumption (A1), we can
omit elements on the diagonal of $W_{0}$ from the parameter space. We thus
can denote a generic parameter vector as $\theta =\left( W_{12},\dots
,W_{N,N-1},\rho ,\gamma ,\beta \right) ^{\prime }\in \mathbb{R}^{m}$, where $
m=N\left( N-1\right) +3$, and $W_{ij}$ is the $(i,j)$-th element of $W$.
Reduced-form parameters can be tied back to the structural model (\ref
{eq:Model}) by letting $\Pi :\mathbb{R}^{m}\rightarrow \mathbb{R}^{N^{2}}$
define the relation between structural and reduced-form parameters:
\begin{equation*}
\Pi (\theta )=\left( I-\rho W\right) ^{-1}\left( \beta I+\gamma W\right) ,
\end{equation*}
where $\theta \in \mathbb{R}^{m}$, and $\Pi _{0}\equiv \Pi (\theta _{0})$.

As $\epsilon _{t}$ (and, consequently, $\nu _{t}$) is mean-independent from $
x_{t}$, $\mathbb{E}[\epsilon _{t}|x_{t}]=0$, the matrix $\Pi _{0}$ can be
identified as the linear projection of $y_{t}$ on $x_{t}$. We do not impose
additional distributional assumptions on the disturbance term, except for
conditions that allow us to identify the reduced-form parameters in
\eqref{eq2:modelRF}. If $x_{t}$ is endogenous, i.e., $\mathbb{E}[\epsilon
_{t}|x_{t}]\neq 0$, a vector of instrumental variables $z_{t}$ may still be
used to identify $\Pi _{0}$. In either case, identification of $\Pi _{0}$
requires variation of the regressor across individuals $i$ and through
instances $t$. In other words, either $\mathbb{E}[x_{t}x_{t}^{\prime }]$ (if
exogeneity holds) or $\mathbb{E}[x_{t}z_{t}^{\prime }]$ (otherwise) is
full-rank.

Our next assumption controls the propagation of shocks and guarantees that they
die as they reverberate through the network. This provides adequate
stability and is related to the concept of stationarity in network models.
It implies the maximum eigenvalue norm of $\rho _{0}W_{0}$ is less than one
and ensures $(I-\rho _{0}W_{0})$ is a non-singular matrix. As the variance
of $y_{t}$ exists, the transformation $\Pi (\theta _{0})$ is well-defined,
and the Neumann expansion $(I-\rho _{0}W_{0})^{-1}=\sum_{j=0}^{\infty }(\rho
_{0}W_{0})^{j}$ is appropriate.

\begin{itemize}
\item[(A2)] \textcolor{black}{$\sum_{j=1}^N|\rho_0 (W_0)_{ij}| < 1$} for
every $i=1,\dots ,N$,
\textcolor{black}{$\|W_0\| < C$ for some positive $C
\in \mathbb{R}$} and $|\rho _{0}|<1$.
\end{itemize}

\noindent We next assume that network effects do not cancel out, another
standard assumption. As we will show, this assumption rules out the
pathological case in which endogenous and exogenous effects exactly cancel
each other out:

\begin{itemize}
\item[(A3)] $\beta _{0}\rho _{0}+\gamma _{0}\neq 0$.
\end{itemize}

\noindent The need for this assumption can be shown by expanding the
expression for $\Pi (\theta _{0})$, which is possible by (A2):
\begin{equation}
\Pi (\theta _{0})=\beta _{0}I+(\rho _{0}\beta _{0}+\gamma
_{0})\sum_{k=1}^{\infty }\rho _{0}^{k-1}W_{0}^{k}.  \label{eq:RFexpanded}
\end{equation}
If Assumption (A3) were violated, $\beta _{0}\rho _{0}+\gamma _{0}=0$ and $
\Pi _{0}=\beta _{0}I$, so the endogenous and exogenous effects would balance each
other out, and network effects would be altogether eliminated in the reduced form.
\footnote{
\textcolor{black}{One important case is when networks do not
determine outcomes, which we interpret as $\rho_0=\gamma_0=0$ or with $W_0$
representing the empty network. From equation \eqref{eq:RFexpanded}, it is
clear that if $\Pi(\theta_0)$ is \emph{not} diagonal with constant entries,
then it must be that $(\rho_0\beta_0+\gamma_0)\neq0$, which implies that
$\rho_0\neq0$ or $\gamma_0\neq0$, and also that $W_0$ is non-empty. Taken
together, this suggests that the observation that $\Pi(\theta_0)$ is not
diagonal is sufficient to ensure that network effects are present and
Assumption (A3) is not violated.}}

Identification of the social effects parameters $(\rho _{0},\gamma _{0})$
requires that at least one row of $W_{0}$ adds to a fixed and known number.
Otherwise, $\rho _{0}$ and $\gamma _{0}$ cannot be separately identified
from $W_{0}$. Clearly, no such condition would be required if $W_{0}$ were
observed.

\begin{itemize}
\item[(A4)] There is an $i$ such that $\sum_{j=1,\dots ,N}(W_{0})_{ij}=1$.
\end{itemize}

Letting $W_{y}\equiv \rho _{0}W_{0}$ and $W_{x}\equiv \gamma _{0}W_{0}$
denote the matrices that summarize the influence of peers' outcomes (the
endogenous social effects) and characteristics on one's outcome (the
exogenous social effects), respectively, the assumption above can be seen as
a normalization. In this case, $\rho _{0}$ and $\gamma _{0}$ represent the
row-sum for individual $i$ in $W_{y}$ and $W_{x}$, respectively.\footnote{
\textcolor{black}{Alternatively, one could normalize $\rho^*=1$
and rescale the network accordingly. In this case,
$W^*=\rho_0W_0$ would be identified instead. Also,
$W_x=\frac{\gamma_0}{\rho_0}W^*$ so $\gamma_0$ would be identified relative
to $\rho_0$. $W_y$ and $W_x$ would be unchanged.}}

The fifth assumption allows for a specific kind of network asymmetry. We
require the diagonal of $W_{0}^{2}$ not to be constant as one of our
sufficient conditions for identification.

\begin{itemize}
\item[(A5)] There exists $l,$ $k$ such that $(W_{0}^{2})_{ll}\neq
(W_{0}^{2})_{kk}$, i.e., the diagonal of $W_{0}^{2}$ is not proportional to $
\iota $, where $\iota$ is the $N\times1$ vector of ones.
\end{itemize}

In unweighted networks, the diagonal of the square of the social
interactions matrix captures the number of reciprocated links for each
individual or, in the case of undirected networks, the popularity of those
individuals. Assumption (A5) hence intuitively suggests differential
popularity across individuals in the social network.

This assumption is related to the network asymmetry condition proposed
elsewhere, such as in \citet{Bramoulleetal2009}. They show that when $W_{0}$
is known, the structural model (\ref{eq:modelSF}) is identified if $I$, $
W_{0}$, and $W_{0}^{2}$ are linearly independent. Given the remaining
assumptions, this condition is satisfied if (A5) is satisfied, but the
converse is \emph{not} true: one can construct examples in which $I$, $W_{0}$
, and $W_{0}^{2}$ are linearly independent when $W_{0}^{2}$ has a constant
diagonal, so $\Pi _{0}$ does not pin down $\theta _{0}$. See Example 1
in Appendix A. The strengthening of this hypothesis is the formal price to
pay for the social interactions matrix $W_{0}$ being unknown to the
researcher.

Before proceeding to our formal results, we provide a very simple
illustration to shed light on how the assumptions above come together to
provide identification. Suppose the observed reduced-form matrix is,
\begin{equation*}
\Pi _{0}=\frac{1}{455}\left[
\begin{array}{ccc}
275 & 310 & 0 \\
310 & 275 & 0 \\
0 & 0 & 182
\end{array}
\right] ,
\end{equation*}
and that, following (A4), the first row is normalized to one. From the third
row and column of $\Pi _{0}$, we see there is no path of any length
connecting the individual in row 3 to or from those in rows 1 or 2, since her
outcome is not affected by their covariates and their outcomes are not
affected by her covariates. In other words, individual 3, is isolated and $
(W_{0})_{13}=(W_{0})_{23}=(W_{0})_{31}=(W_{0})_{32}=0$. On the other hand,
individuals 1 and 2 cannot be isolated, as their covariates are correlated
with the other individual's outcome, reflecting (A5).\footnote{
If on the other hand, $(W_{0})_{ij}=0.5,i\neq j$ in violation of (A5), and
all agents were connected, the model would not be identified.} Due to the
row-sum normalization of the first row, $(W_{0})_{12}=1$. Using (A3), it can
be seen that $W_{0}$ is symmetric if $\Pi _{0}$ is symmetric. We thus find
that $(W_{0})_{21}=1$. This and (A1) map all elements of $W_{0}$, and thus,
\begin{equation*}
W_{0}=\left[
\begin{array}{ccc}
0 & 1 & 0 \\
1 & 0 & 0 \\
0 & 0 & 0
\end{array}
\right] .
\end{equation*}
As the third individual is isolated, she will only be affected by her
exogenous $x_{i}$ and not by endogenous or exogenous peer effects. Hence, the
$(3,3)$ element of $\Pi _{0}$ is equal to $\beta _{0}=\frac{182}{455}=.4$. \
To find $\rho _{0}$, note that $(I-\rho _{0}W_{0})\Pi _{0}=\beta
_{0}I+\gamma _{0}W_{0}$. Hence, focusing on the (1,1) elements of the
matrices above, we find that $\frac{275}{455}-\rho _{0}\frac{310}{455}=.4$,
implying $\rho _{0}=.3$ (complying with (A2)). Finally, $\gamma _{0}$ is
identified from entry $(1,2)$, giving $\gamma _{0}=\frac{310}{455}-.3\frac{
275}{455}=.5$.

Our final assumption articulates the need for a constant network $W_0$
observed over multiple instances of $y_t$ and $x_t$:

\begin{itemize}
\item[(A6)] $y_t$ and $x_t$ are observed for individuals $i=1,\dots,N$, and
instances $t=1,\dots,T$, and the network $W_0$ does not depend on $t$
\end{itemize}

Here, ``instances'' can refer to time but also to settings in which the same
units are observed over multiple episodes. For example, if $i$ are firms,
then $t$ can be segmented markets in which they operate. For simplicity, we
refer to an instance as a time period. If $\Pi_0$ is known, the main
identification result we articulate below will state that $W_0$, $\rho_0$, $
\beta_0$ and $\gamma_0$ are globally identified. However, in practice, $\Pi_0$
is rarely observed and thus all quantities need to be estimated. For this
purpose, when $\Pi_0$ is not known, multiple observations of $y_t$ and $x_t$
with a constant $W_0$ are required to implement the estimator. We expand on
estimation requirements in Section \ref{sec:estimation}.

Importantly, the main identification results (for a given $\Pi_0$) could, in
principle, be applied for each time period $t$. That is, one can write a
version of Equation \eqref{eq:Model} as
\begin{eqnarray*}
y_t&=&\rho_0W_{0t}y_t+\beta_0x_t+\gamma_0W_{0t}x_t+\epsilon_t,
\end{eqnarray*}
where $W_{0t}$ is time-varying and, consequently, the reduced-form
interaction matrix $\Pi_{0t}=(I-\rho_0 W_{0t})^{-1}(I\beta_0+W_{0t}\gamma_0)$
is also time-varying. If the reduced-form matrices were known, the
identification results we develop below could be applied to the reduced-form
element by element for each $\Pi_{0t}$. Again, one rarely observes $\Pi_{0t}$
, for all $t\in[1,T]$. This observation will motivate an extension of the
method, presented in Section \ref{subsec:timevaryingW}, where Assumption
(A6) is relaxed and $W_{0}$ is allowed to vary with $t$.

In an extension, we allow the network $W_{0t}$ to vary over time and
introduce kernel weights. Akin to the nonparametric regression $
Y=f(X)+\epsilon $, $f$ is identified if $\mathbb{E}(\epsilon |X)=0$, and it
is possible to estimate $f$ using neighboring observations if $f$ is
sufficiently smooth or varies slowly. Similar considerations extend to
varying-coefficient models and, in particular, time-varying coefficient
models where local stability conditions as those discussed in
\citet{Dahlhaus2012} are usually invoked (see also
\citet{HastieTibshirani1993} and their Example (e)).

\subsection{Main Identification Results\label{subsec:mainresults}}

Under the assumptions above, we can begin to identify parameters related to
the network. These results are then useful for our main identification
theorems. Let $\lambda _{0j}$ denote an eigenvalue of $W_{0}$ with
corresponding eigenvector $v_{0,j}$ for $j=1,\dots ,N$. Assumptions (A2) and
(A3) allow us to identify the eigenvectors of $W_{0}$ directly from the
reduced form. As $|\rho _{0}|<1$:
\begin{eqnarray}
\Pi _{0}v_{0,j} &=&\beta _{0}v_{0,j}+(\rho _{0}\beta _{0}+\gamma
_{0})\sum_{k=1}^{\infty }\rho _{0}^{k-1}W_{0}^{k}v_{0,j}  \notag \\
&=&\left[ \beta _{0}+(\rho _{0}\beta _{0}+\gamma _{0})\sum_{k=1}^{\infty
}\rho _{0}^{k-1}\lambda _{0,j}^{k}\right] v_{0,j}  \notag \\
&=&\frac{\beta _{0}+\gamma _{0}\lambda _{0,j}}{1-\rho _{0}\lambda _{0,j}}
v_{0,j}.  \label{eq:eigenvaluePiW}
\end{eqnarray}
The infinite sum converges as $|\rho _{0}\lambda _{0,j}|<1$ by (A2). The
equation above implies that $v_{0,j}$ is also an eigenvector of $\Pi _{0}$
with the associated eigenvalue $\lambda _{\Pi ,j}=\frac{\beta _{0}+\gamma
_{0}\lambda _{0,j}}{1-\rho _{0}\lambda _{0,j}}$. The fact that eigenvectors
of $W_{0}$ are also eigenvectors of $\Pi _{0}$ has a useful implication:
eigencentralities may be identified from the reduced form, even when $W_{0}$
is not identified. As detailed in \citet{DePaula2017} and
\citet{Jacksonetal2017}, such eigencentralities often play an important role
in empirical work as they allow a mapping back to underlying models of
social interaction.\footnote{
To identify the eigencentralities, we identify the eigenvector that
corresponds to the dominant eigenvalue. If $W_{0}$ is non-negative and
irreducible, this is the (unique) eigenvector with strictly positive
entries, by the Perron-Frobenius theorem for non-negative matrices (see
\citealp{HornJohnson2013}, p. 534).}

Now let $\Theta \equiv \{\theta \in \mathbb{R}^{m}:$ A$\text{ssumptions
(A1)-(A6) are satisfied}\}$ be the structural parameter space of interest.
Our identification argument is structured as follows: a) we first establish
local identification of the mapping $\Pi (\theta )$ using classical results
on the rank of the gradient of \citealp{Rothenberg1971} (Theorem 1); b) we
then show that $\Pi (\theta )$ is proper (Corollary 1); and c) has a
connected image (Lemma 2, in the Appendix); d) allowing us to state the
cardinality of the pre-image $\Pi ^{-1}(\bar{\Pi})$ is constant for any $
\bar{\Pi}$ in the image of $\Pi (\cdot )$, and that the cardinality is at
most 2 (Theorem 2). We then provide additional conditions to narrow the
identified set to a singleton (Corollaries 2-4).

We now formally present our results. Our first theorem establishes local
identification of the mapping. A parameter point $\theta _{0}$ is locally
identifiable if there exists a neighborhood of $\theta _{0}$ containing no
other $\theta $ which is observationally equivalent. Using classical results
in \citet{Rothenberg1971}, we show that our assumptions are sufficient to
ensure that the Jacobian of $\Pi $ relative to $\theta $ is non-singular,
which, in turn, suffices to establish local identification.

\begin{thm}
\label{thm:localidentification} Assume (A1)-(A6). $\theta_0\in\Theta$ is
locally identified.
\end{thm}

An immediate consequence of local identification is that the set $\{\theta
\in \Theta :\Pi (\theta )=\Pi (\theta _{0})\}$ is discrete (i.e., its
elements are isolated points). The following corollary establishes that $\Pi
$ is a proper function, i.e. the inverse image $\Pi ^{-1}(K)$ of any compact
set $K\subset \mathbb{R}^{N^{2}}$ is also compact (\citealp{KrantzParks2013}
, p.\ 124). Since it is discrete, the identified set must be finite.

\begin{cor}
\label{thm:finiteidentification} Assume (A1)-(A6). Then $\Pi(\cdot)$ is a
proper mapping. Moreover, the set $\{\theta:\Pi(\theta) = \Pi(\theta_0)\}$
has a finite number of elements.
\end{cor}

\noindent Under additional assumptions, the identified set is at most a
singleton in each of the partitioning sets $\Theta _{-}\equiv \Theta \cap
\{\rho \beta +\gamma <0\}$ and $\Theta _{+}\equiv \Theta \cap \{\rho \beta
+\gamma >0\}$.\footnote{
The global inversion results we use are related to, but different from,
variations on a classic inversion result of Hadamard that has been used in
the literature. In contrast, we employ results on the cardinality of the
pre-image of a function, relying on less stringent assumptions. While the
Hadamard result requires the image of the function to be simply-connected
(Theorem 6.2.8 of \citealp{KrantzParks2013}), the results we rely on do not.}

Since $\Theta =\Theta _{-}\cup \Theta _{+}$, if the sign of $\rho _{0}\beta
_{0}+\gamma _{0}$ is unknown, the identified set contains, at most, two
elements. In the theorem that follows, we show global identification only
for $\theta \in \Theta _{+}$, since arguments are mirrored for $\theta \in
\Theta _{-}$.

\begin{thm}
\label{thm:globalidentification} Assume (A1)-(A6). Then for every $
\theta\in\Theta_{+}$, we have $\Pi(\theta)=\Pi(\theta_0)\Rightarrow\theta=
\theta_0$. That is, $\theta_0$ is globally identified with respect to the
set $\Theta_{+}$.
\end{thm}

\noindent Similar arguments apply if Theorem \ref{thm:globalidentification}
instead were to be restricted to $\theta\in\Theta_{-}$. The proof of the
corollary below is immediate and therefore omitted.

\begin{cor}
\label{cor:globalidentification1} Assume (A1)-(A6). If $\rho _{0}\beta
_{0}+\gamma _{0}>0$, then the identified set contains at most one element,
and similarly if $\rho _{0}\beta _{0}+\gamma _{0}<0$. Hence, if the sign of $
\rho _{0}\beta _{0}+\gamma _{0}$ is unknown, the identified set contains, at
most, two elements.\footnote{
\textcolor{black}{Under some special
conditions, the mirror image of $\theta_0$ can be characterized from
equation \eqref{eq:RFexpanded}. If $-W_0$ satisfies Assumption (A4), we may
set $\rho^*=-\rho_0$, $\beta^*=\beta_0$, $\gamma^*=-\gamma_0$ and
$W^*=-W_0$. Then, $\rho_0\beta_0+\gamma_0=-(\rho^*\beta^*+\gamma^*)$. Also
note that $\sum_{k=1}^\infty \rho_0^{k-1}W_0^k=-\sum_{k=1}^\infty
(\rho^*)^{k-1}(W^*)^k$, so $(\rho _{0}\beta
_{0}+\gamma_{0})\sum_{k=1}^{\infty }\rho _{0}^{k-1}W_{0}^{k}=(\rho ^*\beta
^*+\gamma^*)\sum_{k=1}^{\infty }(\rho^*)^{k-1}(W^*)^{k}$. It follows that
$\Pi(\theta_0)=\Pi(\theta^*)$, where
$\theta^*=(\rho^*,\beta^*,\gamma^*,W^*)$.}}
\end{cor}

We now turn our attention to the problem of identifying the sign of $\rho
_{0}\beta _{0}+\gamma _{0}$ from the observation of $\Pi _{0}$. This would
then allow us to establish global identification using Theorem \ref
{thm:globalidentification}. It is apparent from (\ref{eq:RFexpanded}) that
if $\rho _{0}>0$ and $(W_{0})_{ij}\geq 0$, for all $i,j=\{1,\dots ,N\}$, the
off-diagonal elements of $\Pi _{0}$ identify the sign of $\rho _{0}\beta
_{0}+\gamma _{0}$.

\begin{cor}
\label{cor:globalidentification2} Assume (A1)-(A6). If $\rho_0 > 0$ and $
(W_0)_{ij} \ge 0$, the model is globally identified.
\end{cor}

Real-world applications often suggest endogenous social interactions are
positive ($\rho _{0}>0$), in which case global identification is fully
established by Corollary \ref{cor:globalidentification2}. On the other hand,
if $\rho _{0}<0$ (e.g., if outcomes are strategic substitutes), $\rho
_{0}^{k} $ in (\ref{eq:RFexpanded}) alternates signs with $k$, and the
off-diagonal elements no longer carry the sign of $\rho _{0}\beta
_{0}+\gamma _{0}$. Nonetheless, if $W_{0}$ is non-negative and irreducible
(i.e., not permutable into a block-triangular matrix or, equivalently, a
strongly connected social network), the model is also identifiable without
further restrictions on $\rho _{0}$:

\begin{cor}
\label{cor:globalidentification3} Assume (A1)-(A6), $(W_0)_{ij} \ge 0$ and $
W_0$ is irreducible. If $W_0$ has at least two real eigenvalues or $|\rho_0|
< \sqrt{2}/2$, then the model is globally identified.
\end{cor}

\noindent Corollary \ref{cor:globalidentification3} requires that $W_0$ be
irreducible, i.e., that it is not permutable into a block upper-triangular
matrix. In the context of directed graphs, this is similar to requiring that
the matrix be strongly connected, that is, that any node can be reached from
any other node. The corollary then rules out cases when the network is not
connected, for example, if there are two disjoint groups (with no connection
across groups), or a star network pointing from the center towards the
edges. The corollary holds if there are at least two real eigenvalues, or if
$\rho _{0}$ is appropriately bounded. Since $W_{0}$ is non-negative, it has
at least one real eigenvalue by the Perron-Frobenius theorem. If $W_{0}$ is
symmetric, for example, its eigenvalues are all real, and Corollary \ref
{cor:globalidentification3} holds. It also holds if $(W_{0})_{ij}\leq 0$, as
we can rewrite the model as $\rho W_{0}=-\rho |W_{0}|$, where $|W_{0}|$ is
the matrix whose entries are the absolute values of the entries in $W_{0}$.
However, Corollary \ref{cor:globalidentification3} rules out cases that mix
positive $(W_0)_{ij}\geq 0$ and negative interactions $(W_0)_{ij}\leq 0$. In
any case, the bound on $|\rho _{0}|$ is sufficient and holds in most (if not
all) empirical estimates we are aware of obtained from either elicited or
postulated networks, and in our application on tax competition.

\subsection{Extensions}

We present three extensions of the method for individual fixed effects,
common shocks, and time-varying $W$. Appendix B describes extensions for
multivariate covariates and heterogeneous $\beta _{0}$.

\subsubsection{Individual Fixed Effects\label{subsec:corrfixedeffects}}

We observe outcomes for $i=1,\dots ,N$ individuals repeatedly through $
t=1,\dots ,T$ instances. If $t$ corresponds to time, it is natural to think
of there being unobserved heterogeneity across individuals, $\alpha _{i}$,
to be accounted for when estimating $\Pi _{0}$. The structural model (\ref
{eq:modelSF}) is then,
\begin{equation*}
y_{it}=\rho _{0}\sum_{j=1}^{N}W_{0,ij}y_{jt}+\beta _{0}x_{it}+\gamma
_{0}\sum_{j=1}^{N}W_{0,ij}x_{jt}+\alpha _{i}+\epsilon _{it},
\end{equation*}
which can be written in matrix form as
\begin{equation*}
y_{t}=\rho _{0}W_{0}y_{t}+x_{t}\beta _{0}+W_{0}x_{t}\gamma _{0}+\alpha
^{\ast }+\epsilon _{t},
\end{equation*}
where $\alpha ^{\ast }$ is the vector of fixed effects. Individual-specific
and time-constant fixed effects can be eliminated using the standard
subtraction of individual time averages. Defining $\bar{y}
_{t}=T^{-1}\sum_{t=1}^{T}y_{t}$, $\bar{x}_{t}=T^{-1}\sum_{t=1}^{T}x_{t}$, and
$\bar{\epsilon}_{t}=T^{-1}\sum_{t=1}^{T}\epsilon _{t}$,
\begin{equation*}
y_{t}-\bar{y}_{t}=\rho _{0}W_{0}\left( y_{t}-\bar{y}_{t}\right) +\left(
x_{t}-\bar{x}_{t}\right) \beta _{0}+W_{0}\left( x_{t}-\bar{x}_{t}\right)
\gamma _{0}+\epsilon _{t}-\bar{\epsilon}_{t},
\end{equation*}
if $W_{0}$ does not change with time. Identification from the reduced form
follows from previous theorems, since $\Pi _{0}$ is unchanged when
regressing $y_{t}-\bar{y}_{t}$ on $x_{t}-\bar{x}_{t}$.\footnote{
\textcolor{black}{As is the case in panel data, this would require strict
exogeneity ($\mathbb{E}[\epsilon _{s}|x_{t}]=0$ for any $s$ and $t$) or
predetermined errors ($\mathbb{E}[\epsilon _{s}|x_{t}]=0$ for $s \ge t$) so
that the matrix $\Pi _{0}$ can be consistently estimated.}}

\subsubsection{Common Shocks\label{subsec:commonshocks}}

We next allow for unobserved common shocks to all individuals in the network
in the same instance $t$. Such correlated effects\emph{\ }$\alpha _{t}$ can
confound the identification of social interactions. As we have not placed
any distributional assumption on the covariance matrix of the disturbance
term, our analysis readily incorporates correlated effects that are
orthogonal to $x_{t}$. When this is not the case, one possibility is to
model the correlated effects $\alpha _{t}$ explicitly. The model then is,
\begin{equation*}
y_{t}=\rho _{0}W_{0}y_{t}+x_{t}\beta _{0}+\gamma _{0}W_{0}x_{t}+\alpha
_{t}\iota +\epsilon _{t},
\end{equation*}
where $\alpha _{t}$ is a scalar capturing shocks in the network common to
all individuals. Let $\Pi _{01}=\left( I-\rho _{0}W_{0}\right) ^{-1}$ and $
\Pi _{02}=\left( \beta _{0}I+\gamma _{0}W_{0}\right) $ such that $\Pi
_{0}=\Pi _{01}\Pi _{02}$. The reduced-form model is
\begin{equation*}
y_{t}=\Pi _{0}x_{t}+\alpha _{t}\Pi _{01}\iota +v_{t}.
\end{equation*}
We propose a transformation to eliminate the correlated effects:\ exclude
the individual-invariant $\alpha _{t}$, subtracting the mean of the
variables in a given period (global differencing). For this purpose, define $
H=\frac{1}{n}\iota \iota ^{\prime }$. We note that in empirical and
theoretical work, it is customary to strengthen Assumption (A4) and require
that \emph{all} rows of $W_{0}$ sum to one if no individual is isolated (see
for example \citealp{Blumeetal2015}). This strengthened assumption is
usually referred to as row-sum normalization, and is stated below:

\begin{itemize}
\item[(A4$^\prime$)] For all $i=1,...N$, we have that $\sum_{j=1,\dots
,N}(W_{0})_{ij}=1$.
\end{itemize}

This can be written compactly as $W_{0}\iota =\iota $. In this case, $W_{0}$
can be interpreted as the normalized adjacency matrix. Under row-sum
normalization we have that,
\begin{eqnarray*}
\left( I-H\right) y_{t} &=&\left( I-H\right) \left( I-\rho _{0}W_{0}\right)
^{-1}\left( \beta _{0}I+\gamma _{0}W_{0}\right) x_{t}+\left( I-H\right)
\left( I-\rho _{0}W_{0}\right) ^{-1}\epsilon _{t} \\
&=&\left( I-H\right) \Pi _{0}x_{t}+\left( I-H\right) v_{t},
\end{eqnarray*}
because $\left( I-H\right) \left( I-\rho _{0}W_{0}\right) ^{-1}\alpha
_{t}\iota =0$ if Assumption (A4$^\prime$) holds. It then follows that $\tilde{\Pi}
_{0}=(I-H)\Pi _{0}$ is identified. The next proposition shows that, under
row-sum normalization of $W_{0}$, $\Pi _{0}$ is identified from $\tilde{\Pi}
_{0}$ (and, as a consequence, the previous results immediately apply).

\begin{prop}
\label{prop:globalidentification1} If $W_0$ is non-negative, irreducible, and
row-sum normalized, $\Pi _{0}$ is identified from $\tilde{\Pi}_{0}$.
\end{prop}

Under row-sum normalization of $W_{0}$, a common group-level shock affects
individuals homogeneously since $(I-\rho _{0}W_{0})^{-1}\alpha _{t}\iota
=\alpha _{t}(I+\rho _{0}W_{0}+\rho _{0}^{2}W_{0}^{2}+\cdots )\iota =\frac{
\alpha _{t}}{1-\rho _{0}}\iota $, which is a vector with no variation across
entries. Consequently, global differencing eliminates correlated effects and
$\left( I-H\right) \left( I-\rho _{0}W_{0}\right) ^{-1}\alpha _{t}\iota
=\left( I-\rho _{0}W_{0}\right) ^{-1}\alpha _{t}\left( I-H\right) \iota =0$.
Absent row-sum normalization, global differencing does not ensure correlated
effects are eliminated. To see this, note that $(I-\rho _{0}W_{0})^{-1}$ is
no longer row-sum normalized and $\alpha _{t}(I-\rho _{0}W_{0})^{-1}\iota $
does not have constant entries.

The next proposition makes this point formally: that the stronger Assumption
(A4$^\prime$) is \emph{necessary} to eliminate group-level shocks by showing it is
not possible to construct a data transformation that eliminates group
effects in the absence of row-sum normalization.

\begin{prop}
\label{prop:globalidentification2} Define $r_{W_{0}}=(I-\rho
_{0}W_{0})^{-1}\iota $. If in space $\Theta =\{\theta \in \mathbb{R}^{m}:$
Assumptions (A1)-(A6) are satisfied$\}$, there are $N$ matrices $
W_{0}^{(1)},\dots ,W_{0}^{(N)}$ such that $[r_{W_{0}^{(1)}}\;\cdots
\;r_{W_{0}^{(N)}}]$ has rank $N$, then the only transformation such that $(I-
\tilde{H})(I-\rho _{0}W_{0})^{-1}\iota =0$ is $\tilde{H}=I$.
\end{prop}

It is useful to be able to test for row-sum normalization (A4$^\prime$) as it
enables common shocks to be accounted for in the social interactions model.
This is possible as
\begin{eqnarray}
\Pi _{0}\iota &=&\beta _{0}\iota +(\rho _{0}\beta _{0}+\gamma
_{0})\sum_{k=1}^{\infty }\rho _{0}^{k-1}W_{0}^{k}\iota  \notag \\
&=&\left[ \beta _{0}+(\rho _{0}\beta _{0}+\gamma _{0})\sum_{k=1}^{\infty
}\rho _{0}^{k-1}\right] \iota  \notag \\
&=&\frac{\beta _{0}+\gamma _{0}}{1-\rho _{0}}\iota .  \label{eq:PiRowSum}
\end{eqnarray}
The last equality follows from the observation that, under row-normalization
of $W_{0}$, $W_{0}^{k}\iota =W_{0}\iota =\iota $, $k>0$. This implies $\Pi
_{0}$ has constant row-sums, which suggests row-sum normalization is
testable. In the Appendix, we derive a Wald test statistic to do so.\footnote{
For ease of explanation, in the Appendix, we derive the test under the
asymptotic distribution of the OLS estimator. The test generally holds with
minor adjustments for estimators with known asymptotic distributions.}

\subsubsection{Time-varying $W$\label{subsec:timevaryingW}}

We now relax Assumption (A6), which states that $W_0$ does not vary across
the time periods $t=1,\dots,T$. The version of Equation \eqref{eq:Model}
with time-varying network is
\begin{eqnarray*}
y_t&=&\rho_0W_{0t}y_t+\beta_0x_t+\gamma_0W_{0t}x_t+\epsilon_t,
\end{eqnarray*}
with the reduced-form matrix $\Pi_{0t}=(I-\rho_0
W_{0t})^{-1}(I\beta_0+W_{0t}\gamma_0)$. We note that the identification results
developed in Subsection \ref{subsec:mainresults} can, in principle, be
applied element by element to each $\Pi_{0t}$, leading to the identification
of a time-varying $W_{0t}$ (and, potentially, of the parameters $\rho_0$, $
\beta_0 $ and $\gamma_0$).

In practice, implementing any estimation strategy with a time-varying $
\Pi_{0t}$ (or $W_{0t}$) is not feasible using only observation from the
single time period $t$. We instead adopt a kernel-weighted version. Define
period-specific weights $\omega _{t}$, and consider the transformed data $
\tilde{y}_{s}=\omega _{s}(t)y_{s}$ and $\tilde{x}_{s}=\omega_{s}(t)x_{s}$, $
s=1,\dots,T$. Evidently, uniform weights $\omega _{s}=1, s=1,\dots,T$ are
equivalent to the strategy not considering time-varying networks, and assuming that the networks are fixed within those windows.
Alternatively, one could estimate $W_{t}$ in time windows by setting $
\omega_{t}=1[\underbar{t}\leq t\leq \bar{t}]$, where $\underbar{t}$ and $
\bar{t}$ are the start and end of the time window for which $W_{t}$ is
estimated. In this case, the minimum effective window length $\bar{t}-
\underbar{t}$ can be computed as we discuss in Section 3.1. In the context
of DSGE models with time-varying parameters, \citet{kapetanios2019time}
suggests a Gaussian kernel with positive weights throughout the entire
sample. As in nonparametric regression with smooth kernel weights, it also
assumes that the network evolves slowly over time. We further discuss this
strategy in the estimation section, and it is implemented in the empirical
application section below.

\section{Implementation}

We now transition from our core identification results to their practical
implementation. In practice, Ordinary Least Squares (OLS) can only be used to
estimate $\theta $ if $T\gg N$, which is in practice unlikely to be met, as
this is a high-dimensional problem.   Our preferred approach makes use of penalized estimation techniques that
can be used for any given $T$. More specifically, we make use of the
Adaptive Elastic Net GMM (\citealp{CanerZhang2014}), which is based on the
penalized GMM objective function. Given the identification results presented
in Section \ref{sec:Identification}, the population moments used in forming
the GMM objective function will be satisfied at the true parameter vector.

After setting out the estimation procedure, we showcase the method using
Monte Carlo simulations based on stylized and real-world network structures.
In each case, we take a fixed network structure $W_{0}$, and simulate panel
data as if the data generating process were given by the model in (\ref
{model motivation}). We apply the method to the simulated panel data to
recover estimates of all elements in $W_{0}$, as well as the endogenous and exogenous
social effect parameters.

\subsection{Estimation\label{sec:estimation}}

The parameter vector to be estimated is high-dimensional: $\theta =\left(
W_{12},\dots ,W_{N,N-1},\rho ,\gamma ,\beta \right) ^{\prime }\in \mathbb{R}
^{m}$, where $m=N\left( N-1\right) +3$ and $W_{ij}$ is the $(i,j)$-th
element of the $N\times N$ social interactions matrix $W_{0}$.
\textcolor{black}{To be clear, in a network with $N$ individuals, there are
$N(N-1)$ potential interactions because an individual could interact with
everyone else but herself (which would violate Assumption A1). As a
consequence, even with a modest $N$, there are many more parameters to
estimate, and $m$ is large. For example, a network with $N=50$ implies more
than 2,000 parameters to estimate. While we consider $N$ (and thus
$m$) fixed, we still refer to $\theta$ as high-dimensional.} OLS estimation
requires $m\ll NT(\Rightarrow N\ll T)$, so many more time periods than
individuals: a requirement often met in finance data sets (
\citealp{vanVliet2018})
\textcolor{black}{or in other fields (see, e.g.,
Section 4.2 in \citealp{rothenhausler2015})}. Instead, to estimate a large
number of parameters with limited data, we utilize high-dimensional
estimation methods, which are the focus of a rapidly growing literature.

Sparsity is a key assumption underlying many high-dimensional estimation
techniques. In the context of social interactions, we say that $W_{0}$ is
sparse if $\tilde{m}$, the number of non-zero elements of $W_{0}$, is such
that $\tilde{m}\ll NT$. The notion of sparsity thus depends on the number of
time periods. Sparsity corresponds to assuming that individuals influence or
are influenced by a small number of others, relative to the overall size of
the potential network and the time horizon in the data. As such, sparsity is
typically \emph{not} a binding constraint in social networks analysis.
\footnote{
Common stylized networks are sparse, such as the star, lattice (each
individual is a source of spillover only to one other individual), or
interactions in pairs, triads or small groups (\citealp{DeGiorgietal2010}
). Real-world economic networks are also sparse. The sparsity in
\textit{AddHealth} friendship network is around $98$\%. Sparsity of the
production networks in the US\ is above $99$\% (\citealp{Atalayetal2011}).}

In the estimation of sparse models, the \textquotedblleft effective number
of parameters\textquotedblright\ (or \textquotedblleft effective degrees of
freedom\textquotedblright ) relates to the number of variables with non-zero
estimated coefficients (\citealp{TibshiraniTaylor2012}). In the context of
the current social network model, this is equivalent to $m$ parameters,
where $m=dN(N-1)+K$ and $d$ is the network density defined as $\tilde{m}/(N(N-1))$. The Adaptive Elastic Net
GMM estimator presented by \citet{CanerZhang2014} converges at a rate of $
\sqrt{NT/\tilde{m}}=\sqrt{NT/[dN(N-1)+K]}=O(\sqrt{T/(dN)})$ (see remark 7 in
\citealp{CanerZhang2014}). Hence the quality of the large sample results
relies on a comparison between $T$ and $dN$. In line with this, we thus
require $NT\gg dN(N-1)+K$. For example, in the high-school network of
\citet{Coleman1964} that is part of our simulation exercise, $N=70$ and $
d=0.076$. Assuming $K=3$, $N(N-1)\times d+3=370.1$.\footnote{
As pointed out by a referee, variation in $x$ will also matter for
estimation precision. This is reflected in the asymptotic distribution for
this estimator, shown later in this subsection.}

Finally, to reiterate, our identification results themselves do \emph{not}
depend on the sparsity of networks. In particular, Assumptions (A1)-(A6)
\emph{do not} impose restrictions on the number of links in $W_{0}$, or $
\tilde{m}$.\footnote{
If $N\rightarrow \infty $, Assumption (A2) would imply vanishing $
(W_{0})_{ij}$ entries. As highlighted previously, we consider $N$ to be
fixed, in line with many practical applications. Furthermore, Assumption
(A2) is used to represent inverse matrices as Neumann series in our
identification results. What is necessary for this to hold is that a
sub-multiplicative norm on $\rho W$ be less than one. Here we use a specific
norm (i.e., the maximum row-sum norm), but other (induced) norms are also
possible (i.e., the 2-norm or the 1-norm) (see \citealp{HornJohnson2013},
Chapter 5.6).} The identification results presented in Section \ref
{sec:Identification} apply more broadly and irrespective of the estimation
procedure.

Our preferred approach estimates the interaction matrix in the reduced form
while penalizing and imposing sparsity on the structural object $W_{0}$.
\textcolor{black}{We impose sparsity and penalization in the structural-form matrix $W_{0}$ because this is a weaker requirement than imposing sparsity and
penalization in the reduced-form matrix $\Pi _{0}$.}\footnote{
\textcolor{black}{Note that even if $W$ is sparse, $\Pi$ may not be sparse.}
In Appendix C.1, we show that $[\Pi _{0}]_{ij}=0$ if, and only if, there are no paths
between $i$ and $j$ in $W_{0}$, so the pair is not connected. So,
sparsity in $\Pi _{0}$ is understood as $W_{0}$ being ``sparsely connected'',
which is a stronger assumption than sparsity in $W_{0}$.} To do so, we use
the Adaptive Elastic Net GMM (\citealp{CanerZhang2014}), which is based on
the penalized GMM objective function,
\begin{equation}
G_{NT}(\theta ,p)\equiv g_{NT}\left( \theta \right) ^{\prime
}M_{T}g_{NT}\left( \theta \right) +p_{1}\sum_{\substack{ i,j=1  \\ i\neq j}}
^{N}\left\vert W_{i,j}\right\vert +p_{2}\sum_{\substack{ i,j=1  \\ i\neq j}}
^{N}\left\vert W_{i,j}\right\vert ^{2}  \label{eq:enetGMMs1}
\end{equation}
where $\theta =(W_{1,2},\dots ,W_{N,N-1}\,\rho ,\gamma ,\beta )^{\prime }$
with dimension $m=N(N-1)+3$, and $p_{1}$ and $p_{2}$ are the penalization
terms. The term $g_{NT}\left( \theta \right) ^{\prime }M_{T}g_{NT}\left(
\theta \right) $ is the unpenalized GMM objective function with moment
conditions based on orthogonality between the structural disturbance term
and the covariates: $g_{NT}\left( \theta \right) =\sum_{t=1}^{T}\left[
x_{1t}e_{t}(\theta )^{\prime }\;\cdots \;x_{Nt}e_{t}(\theta )^{\prime }
\right] ^{\prime }$, $e_{t}(\theta )=y_{t}-\rho Wy_{t}-\beta x_{t}-\gamma
Wx_{t}$. There are $q\equiv N^{2}$ moment conditions since $x_{it}$ is
orthogonal to $e_{jt}$ for each $i,j=1,\dots ,N$. Hence, the GMM weight
matrix $M_{T}$ is of dimension $N^{2}\times N^{2}$, symmetric, and positive
definite. For simplicity, we use $M_{T}=I_{N^{2}\times N^{2}}$. Note that if
$x_{t}$ is econometrically endogenous, one can also exploit moment
conditions with respect to available instrumental variables.\footnote{
For expositional ease, we describe estimation in the context of the reduced-form model (\ref{eq2:modelRF}), thereby abstaining from individual fixed or
correlated effects. As the GMM estimator uses moments between the structural
disturbance terms and covariates, this endogeneity is built into the
estimation procedure.} Given the identification results presented in Section
\ref{sec:Identification}, if $\theta \neq \theta _{0}$ and does not belong
to the identified set, then $\Pi (\theta )\neq \Pi (\theta _{0})$.
Consequently, the populational version of the GMM objective function is
uniquely minimized at the true parameter vector $\theta _{0}$.

The penalization terms in Equation (\ref{eq:enetGMMs1}) are what makes this
different from a standard GMM problem. The first term, $p_{1}\sum_{{
i,j=1,i\neq j}}^{N}\left\vert W_{i,j}\right\vert $, penalizes the sum of the
absolute values of $W_{ij}$, i.e., the sum of the strength of links, for all
node-pairs. Depending on the choice of $p_{1}$, some $W_{i,j}$'s will be
estimated as exact zeros. A larger share of parameters will be estimated as
zeros if $p_{1}$ increases. The second term, $p_{2}\sum_{{i,j=1,i\neq j}
}^{N}\left\vert W_{i,j}\right\vert ^{2}$, penalizes the sum of the square of
the parameters. This term has been shown to provide better model-selection
properties, especially when explanatory variables are correlated (
\citealp{ZouZhang2009}). The first-stage estimate is
\begin{equation}  \label{eq:1ststage}
\tilde{\theta}(p)=(1+p_{2}/T)\cdot \underset{\theta \in \mathbb{R}^{m}}{\arg
\min }\;\;G_{NT}(\theta ,p)
\end{equation}
where $(1+p_{2}/T)$ is a bias-correction term also used by
\citet{CanerZhang2014}.

Implementing the numerical optimization embedded in Equation
\eqref{eq:1ststage} is computationally challenging, as $m=N(N-1)+3$ may
entail a large number of function arguments. We instead implement the
following modification to use fast Least-Angle Regression (LARS) algorithms (
\citealp{efron2004least}). For any given $\rho $, $\beta$, and $\gamma $,
the expression for $e_{t}(\theta )$ is linear in $W$:
\begin{equation*}
e_{t}(\theta )=y_{t}-x_{t}\beta -W(\rho y_{t}+x_{t}\gamma )\;\;=\;\;\tilde{y}
_{it}(\beta )-W\tilde{x}_{t}(\rho ,\gamma )
\end{equation*}
where $\tilde{y}_{it}(\beta )\equiv y_{t}-x_{t}\beta $ and $\tilde{x}
_{t}(\rho ,\gamma )\equiv \rho y_{t}+x_{t}\gamma $ and, following the
strategy above, is instrumented with $x_{t}$. This motivates a two-step
optimization routine:
\begin{equation*}
\underset{\theta \in \Theta =\Theta _{1}\times \Theta _{2}}{\min }
G_{NT}(\theta ,p)=\underset{(\rho ,\beta ,\gamma )\in \Theta _{1}}{\min }\;\;
\left[ \underset{W_{ij}\in \Theta _{2}}{\min }G_{NT}(\theta ,p)\right] ,
\end{equation*}
where the expression in brackets has a computationally efficient solution
through the LARS algorithm. The numerical optimization is then subsequently
conducted over the parameter space of $(\rho ,\beta ,\gamma )$ only. We also
impose row-sum normalization. Details of the implementation are expanded in
Appendix Subsection C.2.

A second (adaptive) step provides improvements by re-weighting the
penalization by the inverse of the first-step estimates (\citealp{Zou2006}):
\begin{equation}
\tilde{\theta}^*(p)=\left( 1+p_{2}/T\right) \cdot \underset{\theta \in
\Theta}{\arg \min }\;\left\{ g_{NT}\left( \theta \right) ^{\prime
}M_{T}g_{NT}\left( \theta \right) +p_{1}^{\ast }\sum_{\substack{ \{i,j:
\tilde{W}_{ij}\neq 0,  \\ i,j=1,\dots ,N,  \\ i\neq j\}}}\frac{|W_{i,j}|}{|
\tilde{W}_{i,j}|^{c }}+p_{2}\sum_{\substack{ \{i,j:\tilde{W}_{ij}\neq 0,  \\
i,j=1,\dots ,N,  \\ i\neq j\}}}\left\vert W_{i,j}\right\vert ^{2}\right\} ,
\label{eq:enetGMMs2}
\end{equation}
where $\tilde{W}_{i,j}$ is the $\left( i,j\right) $-th element of the
first-step estimate of $W$. We follow \citet{CanerZhang2014} and set $c =2.5$
. If $\tilde{W}_{i,j}<0.05$, we set $\tilde{W}_{i,j}=0.05$. This ensures
that the second-stage estimates can be non-zero even if the first-stage estimates
were zero or small. The computational improvement -- described above for the
first-stage estimator -- is also applied in the adaptive stage.

As a third and final step, we fix the support of $\tilde{\theta}^{\ast }(p)$
, $\mathcal{S}=\{\rho ,\beta ,\gamma \}\cup \{W_{ij}:\tilde{W}_{ij}^{\ast
}\neq 0\}$ and estimate the final parameters without penalization. This
takes as arguments only the elements of $\tilde{\theta}^{\ast }(p)$ that
were estimated as non-zero in the adaptive step. In essence, this step boils
down to a standard GMM approach,
\begin{equation}
\hat{\theta}_{\mathcal{S}}(p)=\underset{\theta \in \mathcal{S}}{\arg \min }
\;\left\{ g_{NT}\left( \theta \right) ^{\prime }M_{T}g_{NT}\left( \theta
\right) \right\}.  \label{eq:enetGMMs3}
\end{equation}
Importantly, \citet{CanerZhang2014} show that the third-step estimator is
asymptotically normal, with a known and easy-to-compute distribution,
\begin{equation*}
\delta ^{\prime }\left[ \big(\hat{G}^{\prime }M_{T}\hat{G}\big)^{-1}\cdot
\big(\hat{G}^{\prime }M_{T}\Omega M_{T}\hat{G}\big)\cdot \big(\hat{G}
^{\prime }M_{T}\hat{G}\big)^{-1}\right] ^{1/2}\cdot \sqrt{NT/\tilde{m}}\cdot
(\hat{\theta}_{\mathcal{S}}-\theta _{0})\overset{d}{\longrightarrow }N(0,1),
\end{equation*}
where $\hat{G}\equiv \hat{G}(\hat{\theta})=\nabla g_{NT}(\theta )$ and $
\Omega \equiv E[g_{NT}(\theta )g_{NT}(\theta )^{\prime }]$.\footnote{
This applies in the case of small $p_{2}$. In the case of large $p_{2}$,
the asymptotic distribution is pre-multiplied by $K_{n}=\frac{I+p_{2}\left[
\hat{G}(\hat{\theta})^{\prime }\hat{\Omega}^{-1}\hat{G}(\hat{\theta})\right]
^{-1}}{1+p_{2}/{NT}}$. See Theorem 4 of \citet{CanerZhang2014}.} This allows
us to conduct hypothesis testing and inference on the $\rho $, $\beta $, $
\gamma $ and the non-zero elements of $W$.

We write $p=(p_{1},p_{1}^{\ast },p_{2})$ as the final set of penalization
parameters. Conditional on $p$, the estimate of the procedure is $\hat{\theta
}(p)$. As in \citet[p.
35]{CanerZhang2014}, the penalization parameters $p$ are chosen by the BIC
criterion. This balances model fit with the number of parameters included in
the model.\footnote{\textcolor{black}{Following \citealp{CanerZhang2014},}
the choice of $p$, which we denote as $\hat{p}$, is the one that minimizes
\begin{equation*}
\text{BIC}(p)=\log \left[ g_{NT}\Big(\hat{\theta}(p)\Big)^{\prime
}M_{T}g_{NT}\Big(\hat{\theta}(p)\Big)\right] +A\Big(\hat{\theta}(p)\Big)
\cdot \frac{\log T}{T}
\end{equation*}
where $A\Big(\hat{\theta}(p)\Big)$ counts the number of non-zero
coefficients among $\{W_{1,2},\dots ,W_{N,N-1}\}$, and larger than a
numerical tolerance, which we set at $10^{-5}$. See also
\citet{zouhastietibshirani2007}.}

In Appendix C.2,
we provide further implementation details, including the choice of initial
conditions. Of course, other estimation methods are available, and our
identification results do not hinge on any particular estimator. Our aim is
to demonstrate the practical feasibility of using the Adaptive Elastic Net
estimator rather than claim it is the optimal estimator.\footnote{
See the alternative approaches of \citet{GautierTsybakov2014},
\citet{Manresa2016}, \citet{lamsouza2016}, and \citet{GautierRose2016}.}

\subsection{Simulations}

\label{subsec:simulation}

We showcase the method using Monte Carlo simulations. We describe the
simulation procedures, results, and robustness checks in more detail in
Appendix D.1. Here, we just provide a brief overview to highlight how well
the method works to recover social networks even in relatively short panels.

For each simulated network, we take a fixed network structure $W_{0}$ and
simulate panel data as if the data generating process were given by (\ref
{model motivation}). We then apply the method to the simulated panel data to
recover estimates of all elements in $W_{0}$, as well as the endogenous and
exogenous social effect parameters ($\rho _{0}$, $\gamma _{0}$). Our result
identifies entries in $W_{0}$ and so naturally recovers links of varying
strength. It is long recognized that link strength might play an important
role in social interactions (\citealp{Granovetter1973}). Data limitations
often force researchers to postulate some ties to be weaker than others
(say, based on interaction frequency). In contrast, our approach identifies
the continuous strength of ties, $W_{0,ij}$, where $W_{0,ij}>0$ implies node
$j$ influences node $i$.

The stylized networks we consider are a random network and a political
party network in which two groups of nodes each cluster around a central
node. The real-world networks we consider are the high-school friendship
network in \citet{Coleman1964} from a small high school in Illinois, and one
of the village networks elicited in \citet{banerjeeetal2013} from rural
Karnataka, India.

Summary statistics for each network are presented in
\textcolor{bluez}{Panel
A} of \textcolor{bluez}{Table A1}.
The four networks differ in their size, complexity, and the relative
importance of strong and weak ties. For example, the Erd\"{o}s-Renyi network
only has strong ties, while the political party network has twice as many strong
as weak ties. For the real-world networks, the mean out-degree distributions
are higher, so the majority of ties are weak, with the high school network
having around $80$\% of its edges being weak ties. All four networks are
also sparse.

For the stylized networks, we assess the performance of the estimator for a
fixed network size, $N=30$. We simulate the real-world networks using
non-isolated nodes in each (so $N=70$ and $65$ respectively).\footnote{
Like \citet{Bramoulleetal2009}, we exclude isolated nodes because they do
not conform to row-sum normalization.}

We evaluate the procedure over varying panel lengths (starting from short
panels with $T=5$), using various metrics. Given our core contribution is to
identify the social interactions matrix, we first examine the proportion of
true zero entries in $W_{0}$ estimated as zeros and the proportion of true
non-zero entries estimated as non-zeros. A global perspective of the
proximity between the true and estimated networks can be inferred from their
average absolute distance between elements. This is the mean absolute
deviation of $\hat{W}$ and $\hat{\Pi}$ relative to their true values,
defined as $MAD(\hat{W})=\frac{1}{N(N-1)}\sum_{i,j,i\neq j}|\hat{W}
_{ij}-W_{ij,0}|$ and $MAD(\hat{\Pi})=\frac{1}{N(N-1)}\sum_{i,j,i\neq j}|\hat{
\Pi}_{ij}-\Pi _{ij,0}|$. As these metrics are closer to zero, more of the
elements in the true matrix are correctly estimated. Finally, we evaluate
the procedure's performance using averaged estimates of the endogenous
and exogenous social effect parameters, $\hat{\rho}$ and $\hat{\gamma}$. In
keeping with the estimation strategy in our empirical application, we report
unpenalised GMM.

\subsection{Results}

\textcolor{bluez}{Figure A1} shows the simulation results as evaluated using
the six metrics described above. \textcolor{bluez}{Panel A} shows that for
each network, the proportion of zero entries in $W_{0}$ correctly estimated
as zeros is above $95$\% even when $T=5$. The proportion approaches $100$\%
as $T$ grows. Conversely, \textcolor{bluez}{Panel B} shows the proportion of
strong non-zero entries estimated as non-zeros (defined as larger than 0.3)
is also high for a small $T$. It is above $70$\% from $T=5$ for the Erd\"{o}s-Renyi network, being at least $85$\% across networks for $T=25$, and
increasing as $T$ grows. As discussed above, the Adaptive Elastic Net
estimator may only recover strong edges well, and not necessarily the weaker
ones, due to the well-known issue with shrinkage estimators that they tend to
shrink small parameters to zero. We return to this issue below.

\textcolor{bluez}{Panels C and D} show that for each simulated network, the
mean absolute deviation between estimated and true networks for $\hat{W}$
and $\hat{\Pi}$ falls quickly with $T$ and is close to zero for large sample
sizes. Finally, \textcolor{bluez}{Panels E and F} show that biases in the
endogenous and exogenous social effects parameters, $\hat{\rho}$ and $\hat{
\gamma}$, also fall in $T$ (we do not report the bias in $\hat{\beta}$ since
it is close to zero for all $T$). The fact that biases are not zero is as
expected for a small $T$, being analogous to well-known results for
autoregressive time series models.
\textcolor{black}{\footnote{\textcolor{black}{The bias in spatial auto-regressive models with a small
number of observations \emph{even when the network is observed} is similarly
documented by
\citet{Smith2009}, \citet{NeumanMizruchi2010}, \citet{Wangetal2014}, and others.}}}

\textcolor{bluez}{Figure A2} shows that, as $T$ increases, the procedure detects weaker links. The figure also shows that, with low sample sizes,
weak edges are generally not detected. This pattern is consistent with the
well-known fact that small parameters are likely shrunk to zero due to the
penalization (\citealp{BelloniChernozhukov2011b}). The absence of weak edges
also implies that the strength of strong edges may also be over-estimated,
since rows are normalized to one. In \textcolor{bluez}{Panel A}, we show the
distribution of the estimates of $\hat{W}_{ij}$, with $T=25$ and for the
high-school network. We show the distribution for the five most common
values of $W_{0,ij}$. We find that most edges weaker than $.5$ are not
detected; edges with a strength of $.75$ are substantially more likely to be
estimated as non-zeros. When they are detected as non-zeros, they are more
likely to be over-estimated. When we estimate $W$ with $T=150$,
\textcolor{bluez}{Panel B} shows that virtually all edges with strength
greater than $.5$ are estimated as non-zeros, and most edges with strength $
.375$ are also detected. We further see are more continuous distribution of
estimates of edge strength. Only edges smaller than $.25$ are not detected.
\textcolor{bluez}{Panels C and D} show a similar conclusion for the village
network.

\textcolor{bluez}{Figure A3} shows the simulated and actual networks under $
T=100$ time periods. The network size is set to $N=30$ in the two stylized
networks, $N=70$ for the high school network, and $N=65$ for the village
household network. In comparing the simulated and true networks,
\textcolor{bluez}{Figure A3} distinguishes between kept edges, added edges,
and removed edges. Kept edges are depicted in blue:\ these links are
estimated as non-zero in at least $5$\% of the iterations, and are also
non-zero in the true network. Added edges are depicted in green:\ these
links are estimated as non-zero in at least $5$\% of the iterations but the
edge is zero in the true network. Removed edges are depicted in red:\ these
links are estimated as zero in at least $5$\% of the iterations but are
non-zero in the true network. \textcolor{bluez}{Figure A3} further
distinguishes between strong and weak links:\ strong links are shown as
solid edges ($W_{0,ij}>.3$), and weak links are shown as dashed edges.

\textcolor{bluez}{Panel A of Figure A3} compares the simulated and true Erd\"{o}s-Renyi networks. All links are recovered. For the political party network,
Panel B shows that all strong edges are correctly estimated. However, around
half the weak edges are recovered (blue dashed edges), with the others being
missed (red dashed edges). As discussed above, this is not surprising given
that shrinkage estimators force small non-zero parameters to zero. Hence,
a larger $T$ is needed to achieve similar performance to the other
simulated networks in terms of detecting weak links. For the more complex
and larger real-world networks, \textcolor{bluez}{Panel C} shows that in the
high-school network, the strong edges are all recovered. However, around half
the weak edges are missing (red dashed edges), and there are a relatively
small number of added edges (green edges):\ these amount to $87$ edges, or
approximately $1.9 $\% of the $4,534$ zero entries in the true high-school
network. A similar pattern of results is seen in the village network in
\textcolor{bluez}{Panel D}: the strong edges are all recovered, and here the
majority of weak edges are also recovered.

\textcolor{bluez}{Panel B of Table A1} compares the network- and node-level
statistics calculated from the recovered social interactions matrix $\hat{W}$
to those in Panel A from the true interactions matrix $W_{0}$. The random Erd\"{o}s-Renyi network is perfectly recovered. For the political party network,
the number of recovered edges is slightly lower than in the true network ($41$
vs. $45$), and all edges are classified as strong. The mean of the in- and
out-degree distributions are slightly lower in the recovered network, and
all three nodes with the highest out-degree are correctly captured (nodes $1$
, $11$, and $28$), include both party leaders (individuals $1$ and $11$
). We then move to discussing the performance in the two real-world
networks. In the high-school network, 30\% of all edges are correctly
recovered, and they are all strong edges. As already noted in
\textcolor{bluez}{Figure A2}, weak edges are not well estimated in the high-school network. This draws two main consequences. First, the average in- and
out- degrees are smaller in the recovered network relative to the true
network. Second, we over-estimate the number of strong edges ($61$ vs. $113$
). This is a downside of row-sum normalization: because some weak edges get
estimated as zeros, the non-zeros are over-estimated so that the row adds to
one. We do, however, recover all three individuals with the highest
out-degree. Finally, in the village network, half the edges are recovered.
The same phenomena of underestimating weak and overestimating strong
edges are again observed. We again recover the three households with the
highest out-degree (nodes $16$, $35$, and $57$).

In the Appendix, we show the robustness of the simulation results to (i)\
varying network sizes and (ii)\ alternative parameter choices and richening the
structure of shocks across nodes. We also demonstrate the gains from using
the Adaptive Elastic Net GMM estimator over alternative estimators, such as
the Adaptive Lasso and OLS.

\section{Application: Tax Competition between US States}

\label{sec:application} We apply our results to shed new light on a classic
social interactions problem:\ tax competition between US states (
\citealp{Wilson1999}). Defining competing ``neighbors'' remains the central
empirical challenge in this literature. Theory provides some guidance on the
issue through two mechanisms driving interactions across jurisdictions:\
factor mobility and yardstick competition.

On factor mobility, \citet{Tiebout1956} first argued that labor and capital
can move in response to differential tax rates across jurisdictions. Factor
mobility leads naturally to the postulated social interactions matrix being
(i) geographic neighbors, given labor mobility, and (ii) jurisdictions with
similar economic or demographic characteristics, given capital mobility (
\citealp{Caseetal1989}).

Yardstick competition is driven by voters making comparisons between states
to learn about their own politician's quality (\citealp{Shleifer1985}).
\citet{BesleyCase1995} formalize the idea in a model where voters use taxes
set by governors in other states to infer their own governor's quality.
Yardstick competition leads naturally to the postulated interactions matrix
being ``political neighbors'': states that voters make comparisons to.

In this application, the number of nodes and time periods is relatively
low:\ the data covers mainland US states, $N=48$, for the years 1962-2015, $T=53$
. Our approach identifies the structure of social interactions among
``economic neighbors'', denoted ${W}_{econ}$. We contrast this against a null
that state taxes are influenced by geographic neighbors, $W_{geo}$, as shown
in Panel A of \textcolor{bluez}{Figure 1A}. With $W_{econ}$ recovered, we
can establish, beyond geography, what predicts the strength of ties between
states and provide fresh insights on drivers of tax competition.

Before using the real data, we confirm the estimator's performance when the
true network is $W_{geo}$ in simulated settings. In line with the findings
of the previous section, Panel B of Figure 1A shows that (i) the procedure recovers strong edges frequently (more specifically, $89$\% of the true
strong edges are recovered) and (ii)\ performance deteriorates when recovering
weak edges ($72$\%). In all cases, the estimator does not add edges not in
the true network. This suggests recovered economic links that deviate from
geographic links may indeed carry signal, while weak links may not get
detected. Finally, the estimator for $\rho $ and $\gamma $ may show some
downward bias with the sample sizes in the application, consistent with the
simulations in \textcolor{bluez}{Appendix Figure A1}.

\subsection{Data and Empirical Specification}

We denote state tax liabilities for state $i$ in year $t$ as $\tau _{it}$,
covering state taxes collected from real per capita income, sales, and
corporate taxes. We extend the sample used by \citet{BesleyCase1995}, that
runs from 1962-1988 ($T=26$).\footnote{\citet{BesleyCase1995} test their
political agency model using a two-equation set-up: (i) on gubernatorial
re-election probabilities and (ii)\ on tax setting. Our application focuses
on the latter because this represents a social interaction problem. They use
two tax series: (i) TAXSIM data (from the NBER), which runs from 1977-1988 and
(ii)\ state tax liabilities series constructed from data published annually
in the Statistical Abstract of the US, that runs from 1962-1988. All their
results are robust to either series. We extend the second series.} The
outcome considered, $\Delta \tau _{it}$, is the change in tax liabilities
between years $t$ and $(t-2)$ because it might take a governor more than a
year to implement a tax program. Their model implies a standard social
interactions specification for the tax setting behavior of state governors:
\begin{equation}
\Delta \tau _{it}=\rho_0 \sum_{j=1}^{N}W_{0,ij}\Delta \tau
_{jt}+\sum_{k=1}^{K}\sum_{j=1}^{N}W_{0,ij}x_{jkt}\gamma
_{0,k}+\sum_{k=1}^{K}\beta _{0,k}x_{ikt}+\alpha _{i}+\alpha _{t}+\epsilon
_{it},  \label{BC Estimation}
\end{equation}
where $k=1,\dots ,K$ are the covariates for state $i$ in period $t$.
Tax setting behavior is determined by (i)\ endogenous social effects arising
through neighbors' tax changes ($\sum_{j=1}^{N}W_{0,ij}\Delta \tau _{jt});$
(ii) exogenous social effects arising through neighbors' characteristics ($
\sum_{j=1}^{N}W_{0,ij}x_{jkt})$; and (iii)\ state $i$'s characteristics ($
x_{ikt} $), including income per capita, the unemployment rate, and the
proportions of young and elderly in the state's population. All
specifications include state and time effects ($\alpha _{i},$ $\alpha _{t}$
). Due to the inclusion of the time effects $\alpha _{t}$, we normalize the
rows of $W_{econ}$ to one. \textcolor{bluez}{Table A6} presents descriptive
statistics for the \citet{BesleyCase1995} sample and our extended sample.

Much of the earlier literature on tax competition has focused on endogenous
social effects and ignored exogenous social effects by setting $\gamma =0$.
Our identification result allows us to relax this restriction and estimate
the full typology of social effects described by \citet{Manski1993}. This is
important because only endogenous social effects lead to social multipliers
from tax competition, and they are crucial to identify as they can lead to a
race-to-the-bottom or suboptimal public goods provision (
\citealp{BrennanBuchanan1980}; \citealp{Wilson1986};
\citealp{OatesSchwab1988}).

\subsection{Preliminary Findings}

\textcolor{bluez}{Table 1} presents our preliminary findings and comparison
to \citet{BesleyCase1995}.  Throughout this section, we refer to \lq\lq OLS estimates\rq\rq\ as the estimates of the main equation (\ref{BC Estimation}) when $W_0$ is postulated as $W_{geo}$ or $W_{econ}$ and $\rho_0$, $\gamma_{0,k}$, and $\beta_{0,k}$ are estimated by OLS.\footnote{We postulate that $W$ is $W_{econ}$ obtained by running the procedure in Section 3, retrieving $\hat{W}$, and re-running model (\ref{BC Estimation}) with $W=\hat{W}$. For such OLS estimates, we use robust standard errors and ignore the sampling uncertainty in the estimated $W_{econ}$.} \textcolor{bluez}{Column 1} shows those estimates where the postulated social interactions matrix is
based on geographic neighbors, exogenous social effects are ignored and the
panel includes all $48$ mainland states but runs only from 1962-1988 as in
\citet{BesleyCase1995}. Social interactions influence gubernatorial tax
setting behavior:\ $\widehat{\rho }_{OLS}=.375$. \textcolor{bluez}{Column 2}
shows this to be robust to instrumenting neighbors' tax changes using the
instrument set proposed by \citet{BesleyCase1995}:\ namely, instrumenting
for $\Delta \tau _{jt}$ using geographic neighbors' lagged changes in per capita income and unemployment rates. These instruments are in the spirit of
using exogenous social effects to instrument for neighbors' tax changes. $
\widehat{\rho }_{2SLS}$ is more than double the magnitude of $\widehat{\rho }
_{OLS}$, suggesting tax setting behaviors across jurisdictions are strategic
complements.

\textcolor{bluez}{Columns 3 and 4} replicate both specifications over the
longer sample, confirming Besley and Case's (1995) \nocite{BesleyCase1995}
finding on social interactions to be robust. $\widehat{\rho }_{2SLS}$ is
again more than double $\widehat{\rho }_{OLS}$. The result in
\textcolor{bluez}{Column 4} implies that for every dollar increase in the
average tax rates among geographic neighbors, a state increases its own
taxes by $64$ cents. This is similar to the headline estimate of
\citet{BesleyCase1995}.\footnote{
Nor is the magnitude very different from earlier work examining fiscal
expenditure spillovers. For example, \citet{Caseetal1989} find that US state
governments' levels of per-capita expenditure are significantly impacted by
the expenditures of their neighbors, with
a one-dollar increase in neighbors' expenditures leading to a seventy-cent increase in
own-state expenditures.}

\subsection{Endogenous and Exogenous Social Interactions}

We now move beyond much of the earlier literature to first establish whether
there are endogenous and exogenous social interactions in tax setting. We
first focus on the endogenous and exogenous social interaction parameters,
and in the next subsection, we detail the identified social interactions
matrix, $\hat{W}_{econ}$. To do so, we need to modify slightly how we
instrument for neighbors' tax changes:\ the instrument set proposed by
\citet{BesleyCase1995} based on geographic neighbors' characteristics will
generally be weaker when estimating the full specification in (\ref{BC
Estimation}) because the instruments are now directly controlled for in (\ref
{BC Estimation}). We use an Adaptive Elastic Net GMM approach, which
instruments neighbors' tax changes with the characteristics of all other
states. With the inclusion of endogenous and exogenous social effects, this
represents our preferred approach.

\textcolor{bluez}{Columns 1 and 2 of Table 2} show OLS\ and GMM estimates
for $\rho $ obtained from the Adaptive Elastic Net procedure, where we still
set $\gamma =0$ but use our preferred instrument set:\ $\widehat{\rho }
_{GMM}=.709>$ $\widehat{\rho }_{OLS}=.649$.
\textcolor{bluez}{Columns 3 and
4} estimate the full model in (\ref{BC Estimation}). Relative to when
exogenous social effects are assumed away ($\gamma =0$), the OLS\ and GMM
estimates of $\rho $ are smaller, but we continue to find robust evidence of
endogenous social interactions in tax setting. The specification in
\textcolor{bluez}{Column 4} represents our preferred one: $\widehat{\rho }
_{GMM}=.452$ (with a standard error of $.132$). This value meets the
requirements on $\rho $ in Corollaries 3 and 4 for global identification.
\footnote{\textcolor{bluez}{Table A7} shows the full set of exogenous social
effects (so \textcolor{bluez}{Columns 1 and 2} refer to the same
specifications as \textcolor{bluez}{Columns 3 and 4 in Table 2}). Exogenous
social effects operate through economic neighbors' unemployment rate,
demographic characteristics, and their governor's age.}

\subsection{Identified Social Interactions Matrix}

\textcolor{bluez}{Figure 1B} shows how the structures of economic ($\hat{W}
_{econ}$) and geographic networks ($W_{geo}$) differ, where connected edges
imply that two states are linked in at least one direction (state $i$
causally impacts state taxes in $j$, and/or \emph{vice versa}). This
comparison makes clear whether all states geographically adjacent to $i$
matter for its tax setting behavior and whether there are non-adjacent
states that influence its tax rate.

The left-hand panel of \textcolor{bluez}{Figure 1B}\ shows the network of
geographic neighbors (whose edges are colored blue), onto which we
superimpose edges \emph{not} identified as links in $W_{econ}$;\ dropped
edges are in red. The vast majority of geographically adjacent states are
irrelevant for tax setting behavior. The right-hand panel of
\textcolor{bluez}{Figure 1B}\ adds new edges identified in $\hat{W}_{econ}$
that are \emph{not} part of $W_{geo}$;\ these added edges are in green and
represent non-geographically adjacent states through which social
interactions occur. For tax-setting behavior, economic distance is
imperfectly measured if we simply assume interactions depend only on
physical distance.

\textcolor{bluez}{Table 3} summarizes the comparison between $W_{geo}$ and $
\hat{W}_{econ}$. $W_{geo}$ has $214$ edges, while $\hat{W}_{econ}$ has only $
49$. $\hat{W}_{econ}$ and $W_{geo}$ have $9$ edges in common. Hence, the
vast majority of geographical neighbors ($205/214=96$\%) are not relevant
for tax setting. $\hat{W}_{econ}$ has $40$ edges that are absent in $W_{geo}$
, and the identified social interactions are more spatially dispersed than
under the assumption of geographic networks. This is reflected in\ the far
lower clustering coefficient in $\hat{W}_{econ}$ than in $W_{geo}$ ($.042$
versus $.419$).\footnote{
The clustering coefficient is the frequency of the number of fully connected
triplets over the total number of triplets.
}

\subsection{Links and Reciprocity}

Our estimation strategy identifies the continuous strength of links, $
W_{0,ij}$, where $W_{0,ij}>0$ is interpreted as state $j$ influencing
outcomes in state $i$. This is useful because recent developments in tax
competition theory, using insights from the social networks literature,
suggest links need not be reciprocal (\citealp{JanebaOsterleh2013}).

\textcolor{bluez}{Table 3} reveals that only $12.2$\% of edges in $\hat{W}
_{econ}$ are reciprocal (all edges in $W_{geo}$ are reciprocal by
definition). Hence, tax competition is both spatially disperse and
asymmetric. In most cases where tax setting in state $i$ is influenced by
taxes in state $j$, the opposite is not true.

Given common time shocks $\alpha _{t}$ in (\ref{BC Estimation}), row-sum
normalization is required and ensures $\sum_{j}W_{0,ij}=1$. Hence, for every
state $i$, there will be at least one economic neighbor state $j^{\ast }$
that impacts it, so $W_{0,ij^{\ast }}>0$. This just reiterates that social
interactions matter. On the other hand, our procedure imposes no restriction
on the derived columns of $\hat{W}_{econ}$. It could be that a state does
not affect any other state. To see this in more detail, the final rows of
\textcolor{bluez}{Table 3} report the degree distribution across states,
splitting for in-networks and out-networks. In $W_{geo}$, the in-degree is
by construction equal to the out-degree, as all ties are reciprocal. The
greater sparsity of the network of economic neighbors is reflected in the
degree distribution being lower for $\hat{W}_{econ}$ than $W_{geo}$. In $
\hat{W}_{econ}$, the dispersion of in- and out-degree networks is very
different (as measured by the standard deviation), being nearly nine times
higher for the in-degree. Hence one reason for so few reciprocal ties
being in the economic network is that out-degree network ties are rarely
also in-degree ties.

This asymmetry in $\hat{W}_{econ}$ further suggests that some highly
influential states drive tax setting behavior in other states. To see which
states these are, \textcolor{bluez}{Figure 2} shows a histogram for the
number of out-degree links from states. Twenty states have an out-degree of
zero, so their tax rates have no direct impact on any other state's tax
setting behavior. The most influential states in terms of the highest
out-degree are Alabama (directly impacts tax setting behavior in five
other states) and South Carolina, Pennsylvania, and Montana (which each
directly impact tax setting behavior in four other states). Taking South
Carolina as an example, the four states that it directly impacts
include its geographic neighbor, Georgia, as well as non-geographic
neighbors Missouri, Montana, and Virginia.

\subsection{Factor Mobility or Yardstick Competition?}

We use $\hat{W}_{econ}$ to shed light on the roles of factor mobility and
yardstick competition in driving tax competition. To do so, we estimate the
factors correlated with the existence of links between states $i$ and $j$ in
$\hat{W}_{econ}$ relative to $\hat{W}_{geo}$. For state pairs with non-zero
links in either $\hat{W}_{econ}$ or $\hat{W}_{geo}$, we define a dummy
outcome $\hat{W}_{econ,ij}=1$ if a link between states $i$ and $j$ is
estimated under $\hat{W}_{econ}$ and $\hat{W}_{econ,ij}=0$ if a link between
states $i$ and $j$ exists under $\hat{W}_{geo}$ but not under $\hat{W}
_{econ} $. We examine correlates of links using the following dyadic
regression:
\begin{equation}
\hat{W}_{econ,ij}=\lambda _{0}+\lambda _{1}X_{ij}+\lambda _{2}X_{i}+\lambda
_{3}X_{j}+u_{ij},
\end{equation}
estimated using a linear probability model. The elements $X_{ij},$ $X_{i},$
and $X_{j}$ correspond to characteristics of the pair of states ($i,j$),
state $i$, and state $j$, respectively. Covariates are time-averaged over
the sample period, and robust standard errors are reported.

\textcolor{bluez}{Table 4} presents the dyadic regression results.
\textcolor{bluez}{Column 1} controls only for the distance between states $i$
and $j$:$\ $this is highly predictive of an economic link between them. This
reflects that the economic network of state $i$ often comprises states are
in the same region, but not necessarily contiguous to state $i$.
\textcolor{bluez}{Column 2} adds two $X_{ij}$ covariates to capture the
economic and demographic homophily between states $i$ and $j$. GDP homophily
is the absolute difference in the states' GDP per capita. Demographic
homophily is the absolute difference in the share of young people (aged
5-17) plus the absolute difference of the share of elderly people (aged 65+)
across the states. GDP\ homophily does not predict economic ties, whereas
demographic homophily does.

\textcolor{bluez}{Columns 3-5} then sequentially add in several sets of
controls. For labor mobility, we use net state-to-state migration data to
control for the net migration flow of individuals from state $i$ to state $i$
(defined as the flow from $i$ to $j$ minus the flow from $j$ to $i$).
\footnote{
We also experimented with alternative measures of labor migration, and
the results were qualitatively the same. State-to-state migration data are based
on year-to-year address changes reported on individual income tax returns
filed with the IRS. The data cover filing years 1991 through 2015 and
include the number of returns filed, which approximates the number of
households that migrated, and the number of personal exemptions claimed, which
approximates the number of individuals who migrated. The data are available
at https://www.irs.gov/statistics/soi-tax-stats-migration-data (accessed
September 2017).} We then add a political homophily variable between states.
For any given year, this is set to one if a pair of states have governors of
the same political party. As this is time-averaged over our sample, this
element captures the share of the sample in which the states have governors
of the same party. Lastly, we include whether state $j$ is considered a tax
haven (and so might have disproportionate influence on other states). Based
on \citet{Findleyetal2012}, the following states are coded as tax havens:
Nevada, Delaware, Montana, South Dakota, Wyoming, and New York.

\textcolor{bluez}{Column 5} shows that with this full set of controls,
distance remains a robust predictor of the existence of economic links
between states. However, the identified economic network highlights
additional significant predictors of tax competition between states:\
political homophily \emph{reduces} the likelihood of a link, suggesting any
yardstick competition driving social interactions occurs when voters compare
their governor to those of the opposing party in other states. Tax haven
states appear to be especially less influential in the tax setting behaviors
of other states. This mirrors what was observed in
\textcolor{bluez}{Figure
2}, where some of the prominent tax havens -- Nevada, Delaware, and New York
-- were all identified to have zero out-degree links. The relatively weak
influence of tax haven states eases concerns over a race-to-the-bottom in
tax setting behaviors.

\textcolor{bluez}{Column 6} controls for state $i$ and state $j$ fixed
effects. This reinforces the idea that distance and political homophily
correlate to the strength of influence states tax setting has on others (the
tax haven dummy cannot be separately identified in this specification).
Labor mobility between states does not robustly predict the existence of
economic ties.

\subsection{Dynamics}

As in our identification result, our empirical approach has taken the
network structure as fixed over the entire sample period. In the context of
tax competition over our study period, this might be a strong assumption. We
examine the issue in more detail by allowing the estimated $W_{econ}$ matrix
to vary over time by changing the weight placed on observations from any
given time period. More precisely, for any given time period $t$, we weight
observations using a Gaussian kernel with its center varying
period-by-period from $1962$ to $2015$. The variance of the kernel is set
such that $75$\% of the weight is given to the first half of the data (i.e.,
pre-$1988$) when the kernel is centered in $1962$.
\textcolor{bluez}{Figure
A4} shows the kernel employed as we vary its center: the solid kernel is
centered in $1962$, the start of our sample -- when we place the most weight
on observations from $1962$. The static case considered previously is akin
to using a uniform kernel over all periods. This kernel weighting approach
is outlined in Section \ref{subsec:timevaryingW}. We fully describe the
algorithm in Appendix Section C.2.

We begin by considering time-varying estimates of the endogenous and
exogenous social interaction parameters from the full model in (\ref{BC
Estimation}). The results for the endogenous social interactions parameter
are shown in \textcolor{bluez}{Figure 3}, where the shaded areas are the $95$\%
confidence intervals of the period-by-period estimates. Panel A shows that OLS\
estimates of $\widehat{\rho }$ drift up over time, so the strength of
endogenous interactions increases from around $.35$ in the late 1960s to $.50
$ by the 2010s. In all periods, we can reject the null that the
endogenous social effect is zero. Recall the earlier static estimate was $
\widehat{\rho }_{OLS}=.375$.

Panel B shows the estimated endogenous social effect when we use GMM based
on the characteristics of all other states as IVs. This also drifts up from
around $.35$ in the late 1960s to $.50$ by the 2010s. In the majority of
periods, we can reject the null of no endogenous social effect.\footnote{
Standard errors are estimated without imposing the restriction that
parameters vary slowly over time, and fluctuations across periods reflect
variations in the network across periods.}

\textcolor{bluez}{Figure 4} shows the evolution of $\hat{W}_{econ}$ over
time as we center the kernel on different periods, following the same color-coding as in \textcolor{bluez}{Figure 1}. In all periods, geography-based
edges play little role, and over time, the economic network becomes denser.
This highlights\ not only that economic networks for tax competition always
differ starkly from geography-based networks, but that the nature of
economic networks relevant to tax competition has changed steadily over
time.

\textcolor{bluez}{Figure 5} shows how the features of $\hat{W}_{econ}$ evolve as
we place greater weight from early to later periods. For each statistic, we
plot the period-by-period estimate when we center the kernel in any given
 period. The resulting smoothed estimates are then shown. To ease
exposition, networks edges with $W_{0,ij}<1/47$ are removed. This cutoff is
chosen as, in theory, states can only link at maximum with $47$ other
states. \textcolor{bluez}{Panel A} shows the share of edges that are kept
from the previous estimate (centered in the previous period). We see
relatively high stability in $\hat{W}_{econ}$ with the smoothed estimate
suggesting more than $60$\% of edges always being kept from one estimate to
the next, with this stability increasing from the late 1980s.

\textcolor{bluez}{Panel B} shows how the overlap between $\hat{W}_{econ}$
and $W_{geo}$ varies over time, as measured by the share of edges that are
only present in $\hat{W}_{econ}$. There is little overlap between the two
networks over the entire sample. The smoothed estimate suggests that at
least $80$\% of identified edges in $\hat{W}_{econ}$ are never in $W_{geo}$.
The divergence between economic and geographic neighbors becomes starker
from the mid-$1980$s onwards.

\textcolor{bluez}{Panels C and D} show how the clustering and reciprocity of
links in $\hat{W}_{econ}$ vary as we shift the weight to later observations.
Clustering of $\hat{W}_{econ}$ increases from the $1960$s through to the
early $2000$s. Thereafter, social interactions in tax competition become
sparser. We also observe a reversal in the extent to which social
interactions are reciprocal, with reciprocity rising to a peak in the early
1980s -- when $20$\% of ties were reciprocal -- and slowly falling
thereafter.

Taken together, the results suggest the nature of tax competition between
US\ states has changed over time through two mechanisms: (i) the strength of
endogenous social interactions ($\widehat{\rho }$) has increased over time and
(ii) the\ network of states interacted with ($\hat{W}_{econ}$) varies
over time. This has important implications for policy evaluation:\ the same
intervention might have different spillover effects if implemented at
different moments in time due to the evolution of $\widehat{\rho }$ and $
\hat{W}_{econ}$. We consider this next using counterfactual simulations.

\subsection{Counterfactuals}

We use a counterfactual exercise to contrast how shocks to tax setting in a
given state propagate under $\hat{W}_{econ}$, relative to what would have
been predicted under $W_{geo}$. We do so for both static and dynamic estimates
of $\hat{W}_{econ}$. We focus on South Carolina (SC), a state with one of the
highest out-degree, as shown in \textcolor{bluez}{Figure 2}. We consider a
scenario in which SC exogenously increases its taxes per capita by $10$\%.
We measure the differential change in equilibrium state taxes in state $j$
under the two network structures using the following statistic:
\begin{equation}
\Upsilon _{j}=\log (\Delta \tau _{jt}|\hat{W}_{econ})-\log (\Delta \tau
_{jt}|W_{geo}),  \label{Impulse}
\end{equation}
so that positive (negative) values imply equilibrium taxes being higher
(lower) under $\hat{W}_{econ}$.\footnote{
For $W_{geo}$, we calculate the counterfactual at $\hat{\rho}_{GMM}=.452$,
the endogenous effect parameter estimated in our preferred specification,
\textcolor{bluez}{Column 4 of Table 2}.}

Starting with the static case, \textcolor{bluez}{Panel A of Figure 6} shows
for each mainland US state the spillover effects through the economic
network of tax competition. This highlights positive spillovers on tax rates
in many states that are not geographic neighbors of SC.
\textcolor{bluez}{Panel
B} graphs $\Upsilon _{j}$ to make precise how spillovers derived from $\hat{W
}_{econ}$ diverge from those predicted under $W_{geo}$. In $26$ states, $
\Upsilon _{j}$ is smaller than $.01$\% because both networks predict
negligible spillovers to those states. In the remaining $22$ mainland
states, there is a wide discrepancy between the equilibrium state tax rates
predicted under $\hat{W}_{econ}$ relative to $W_{geo}$: $\Upsilon _{j}$
varies from $-1$ to $4.03$. The long-run effect in SC itself is also higher
under $\hat{W}_{econ}$ than under $W_{geo}$. The former states that given feedback
effects, the long-run increase in tax rates in SC from a $10$\% increase is $
11.4$\%, while the geographically based network implies a smaller equilibrium
increase of $10.3$\%.

As $\hat{W}_{econ}$ is spatially more dispersed than $W_{geo}$, the general
equilibrium effects are different under the two network structures.
\textcolor{bluez}{Table 5} summarizes the general equilibrium implications
for tax inequality under $\hat{W}_{econ}$ and the $W_{geo}$ counterfactual.
The average tax rate increase under $\hat{W}_{econ}$ is three times that
estimated under $W_{geo}$. Moreover, the dispersion of tax rates across
states increases under $\hat{W}_{econ}$ relative to $W_{geo}$. Finally,
assuming interactions are based solely on geographic neighbors, we miss the
fact that many states have relatively small tax increases.

We can repeat the exercise using the dynamically estimated economic network.
Throughout, we calculate the general equilibrium effects of the \emph{same}
policy experiment: SC increasing its taxes per capita by $10$\%. These
general equilibrium effects vary over the sample period because the strength
of social interactions in tax competition vary ($\widehat{\rho }_{GMM}$), as
shown in \textcolor{bluez}{Figure 3}, and identified economic neighbors vary
over time ($\hat{W}_{econ}$), as shown in \textcolor{bluez}{Figure 4}. The
results are summarized in \textcolor{bluez}{Figure 7}. Placing weight on the
early or later part of the sample generates similar changes in average tax
rates and their variance in general equilibrium -- with both being lower
than simulated under the static model. Placing more weight on the middle of
the sample period generates higher changes in average tax rates and their
variance in general equilibrium.

The differential general equilibrium impacts found as we place different
weights across sample observations links to recent discussions on the external
validity of internally valid causal impacts based on micro-evidence. While
the earlier literature has emphasized the potential interaction of treatment
effects with aggregate shocks (\citealp{rosenzweig2020external}) or how
behavioral responses to social insurance policies vary over the business
cycle (\citealp{kroft2016should}), our analysis provides another explanation
for the changing impacts of policies where social interactions determine
behavior:\ changes in the strength of social interactions and the network of
economic interactions.

\section{Discussion}

In a canonical social interactions model, we provide sufficient conditions
under which the social interactions matrix, and endogenous and exogenous social
effects are all globally identified, even absent information on social
links. Our identification strategy is novel and may bear fruit in other
areas. The method is immediately applicable to other classic social
interactions problems, but where data on social links is either missing or
partial. In fields such as macroeconomics, political economy, and trade,
there are core areas of research where social interactions across
jurisdictions/countries etc. drive key outcomes, panel data exist over many
periods, and the number of nodes is relatively fixed. Moreover, while our
discussion and application have focused on a continuous policy response
(state taxes), our methods can also be applied to the extensive margin of
policy adoption and diffusion. Such diffusion models might generate network
interactions where some states influence the later adoption of economic and
social policies in other jurisdictions. This issue is studied by
\citet{dellavigna2022policy} in the context of US\ state policies -- they
examine the diffusion of over $700$ policies in the past $70$ years. Their
work also suggests the nature of interactions across states has changed:\
while geographic proximity is a good predictor of policy diffusion, they
also find that since $2000$, political alignment across states has become the
strongest predictor of diffusion.

In finance, high-frequency panel data is readily available and relevant for
the study of core research questions. For example, a long-standing question
has been whether CEOs are subject to relative performance evaluation, and if
so, what is the comparison set of firms/CEOs used (\citealp{edmansgabaix2016}
).
More generally, our method can be readily applied to a large class of
economic questions around contagion, risk, and the fragility of economic and
environmental systems. For example, since the financial crisis of 2008, it
has become clear that linkages between actors such as firms or banks are
complex and often hidden, yet because endogenous network interactions cause
feedback loops and have multiplier effects, they can have enormous
implications for the evolution of a financial crisis or the propagation of supply
shocks in aggregate. Identifying such synchronicity is a critical first step
to putting in place policies to reduce the fragility of economic systems (
\citealp{vanVliet2018}; \citealp{elliott2022networks};
\citealp{goldstein2022synchronicity}).

Advances in the availability of administrative data, data from social media,
mobile technologies, and online economic transactions all offer new
possibilities to identify social interactions with long panels or high-frequency data collection, where data on social ties will typically be
missing.

Three further directions for future research are of note. First, under
partial observability of $W_{0}$ (as in \citealp{Blumeetal2015}), the number
of parameters in $W_{0}$ to be retrieved falls quickly. Our approach can
then still be applied to complete knowledge of $W_{0}$, such as if Aggregate
Relational Data is available, and this could be achieved with potentially
weaker assumptions for identification, and in even shorter panels. To
illustrate possibilities, \textcolor{bluez}{Figure A5} shows results from a
final simulation exercise in which we assume the researcher starts with
partial knowledge of $W_{0}$. We do so for Banerjee \emph{et al.} (2013)
\nocite{banerjeeetal2013} village family network, showing simulation results
for scenarios in which the researcher knows the social ties of the three
(five, ten) households with the highest out-degree. For comparison, we also
show the earlier simulation results when $W_{0}$ is entirely unknown. This
clearly illustrates that with partial knowledge of the social network,
performance on all metrics improves rapidly for any given $T$.

Second, we have developed our approach in the context of the canonical
linear social interactions model (\ref{model motivation}). This builds on
\citet{Manski1993} when $W_{0}$ is known to the researcher, and the
reflection problem is the main challenge in identifying endogenous and
exogenous social effects. However, the reflection problem is functional-form
dependent and may not apply to many non-linear models (\citealp{Blumeetal2011}
, \citealp{Blumeetal2015}). An important topic for future research is to
extend the analysis to non-linear social interaction settings. Relatedly,
the canonical social interaction model assumes that the same $W_{0}$ governs
the endogenous and exogenous channels. Despite the relaxation we propose in
Section 2.3.3, we see this as a limitation of the current method, and future
research is needed to allow for a fully flexible approach.

Finally, an important part of the social networks literature examines
endogenous network formation (\citealp{Jacksonetal2017};
\citealp{DePaula2017}). Our analysis allows us to begin probing the issue in
two ways. First, the kind of dyadic regression analysis in Section 4 on the
correlates of entries in $W_{0,ij}$ suggests factors driving link formation
and dissolution. Second, this leads naturally to a broad agenda going
forward, to address the challenge of simultaneously identifying and
estimating time varying models of network formation and social interaction,
all in cases where data on social networks is not required.

\singlespacing
\bibliographystyle{ectaCUSTOM}
\bibliography{bibTRIMMED}

\newpage