EconBase
← Back to paper

Non-Identifiability in Network Autoregressions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

137,295 characters · 18 sections · 63 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
center[center omitted — 290 chars of source]

\noindentAbstract

We study identifiability of the parameters in autoregressions defined on a network. Most identification conditions that are available for these models either rely on the network being observed repeatedly, are only sufficient, or require strong distributional assumptions. This paper derives conditions that apply even when the individuals composing the network are observed only once, are necessary and sufficient for identification, and require weak distributional assumptions. We find that the model parameters are generically, in the measure theoretic sense, identified even without repeated observations, and analyze the combinations of the interaction matrix and the regressor matrix causing identification failures. This is done both in the original model and after certain transformations in the sample space, the latter case being relevant, for example, in some fixed effects specifications.

\noindentKeywords: fixed effects, invariance, networks, quasi maximum likelihood estimation.\newline\noindentJEL Classification: C12, C21.

Introduction

In a wide range of empirical settings, data are available for an outcome variable and some covariates for each of the nodes of a network, as well as some measure of the pairwise interaction between the nodes. Autoregressive processes offer a simple way to study how the covariates affect the outcome variable, taking into account the network interaction. Models of this type can be traced back at least to whittle1954, and have since proved useful in many applications, across many scientific fields. In economics, and the social sciences more generally, they are currently particularly popular in the analysis of peer effects and social networks. The models are known as simultaneous autoregressions in the statistics literature Cressie1993, spatial autoregressions in the econometrics literature LeSagePace2009, are closely related to linear-in-means models Manski93, and have important connections to linear structural equation models Drton2011. To emphasize their wide applicability, we refer to them as network autoregressions.

This paper is concerned with identifiability of the parameters in a network autoregression. We employ a classical notion of identifiability, according to which parameters are identified if they are uniquely recovered from the distribution of the observables. Lack of identification has, of course, serious consequences for inference. For example, identification is necessary for the limiting objective function of an extremum estimator to be uniquely maximized at the true value of the parameter, which is a standard condition for consistency NeweyMcFadden94. In addition, inference is expected to be difficult near the cases in which identification fails. Given how fundamental the problem of identification is, and given that establishing consistency of some extremum estimator is sufficient for identification, it is not surprising that there is a vast amount of work that is relevant for the present study. We mention in particular two very influential papers: Lee2004 and Bramoulle2009. Lee2004, in a rigorous analysis of the asymptotic properties of the quasi maximum likelihood estimator based on the Gaussian distribution, provides conditions that are sufficient for consistency and hence for identifiability. Bramoulle2009 investigates identifiability by looking at the mapping from the reduced form parameters to the structural parameters, an approach that has become standard in the social network literature.

The present paper studies identifiability directly from the first moment, or the first two moments, of the outcome variable. Compared to the approach via reduced form parameters, identification from moments does not require the nodes of the network being observed over multiple instances (e.g., over time). We show that identification from the first moment is generally possible, and characterize the cases when it is impossible. One class of cases when identifiability from the first moment is impossible is particularly relevant in fixed effects models (for example, the classical linear-in-means model with group fixed effects belongs to this class of cases). In that class, non-identifiability from the first moment is linked to the impossibility of invariant inference; more precisely, the parameters cannot be identified from any statistic that is invariant with respect to a certain group of transformations under which the model itself is invariant. This type of non-identifiability occurs despite the fact that the parameters may be identifiable from the second moment of the outcome variable. Hence, identifiability from the second moment is of very questionable value in these case, because it can only lead to non-invariant inference.

mycommentwe aim to understand what combinations of the interaction matrix $W$ and the regressor matrix $X$ lead to a failure of identification. .......... The present paper departs from previous studies in two main ways. [[[this bit was in main text, move it to intro: Also, Condition (ref) will be analyzed in detail in Sections (ref) and (ref). In particular, it will become clear in Section (ref) that a failure of Condition (ref) implies a specific type of non-identifiability.]]]

Section (ref) sets out the framework. Section (ref) studies identifiability from the first and second moments of the outcome variable, and Section (ref) discusses identifiability after reduction by invariance. Section (ref) reports simulation evidence on the consequences of being close to non-identifiability. Section (ref) briefly concludes. The appendices contain additional material and all proofs. Throughout the paper the results are illustrated by means of several examples.

Notation. Matrices are denoted by capital letters, vectors and scalars by lowercase letters. We reserve bold letters for random quantities (scalars, vectors, or matrices), so, for example, $\boldsymbol{y}$ denotes a random vector and $y$ a realization of $\boldsymbol{y}$. Throughout the paper, $\iota_{n}$ denotes the $n\times1$ vector of ones, $\mathbb{R}^{n\times m}$ denotes the set of real $n\times m$ matrices, $\operatorname{col}(A)$ denotes the column space of a matrix $A$, $M_{A}\coloneqq I_{n}-A(A^{\prime} A)^{-1}A^{\prime}$ for a full column rank matrix $A$, $\mu_{\mathbb{R}^{n}}$ denotes the Lebesgue measure on $\mathbb{R}^{n}$, \textquotedblleft a.s.\textquotedblright\ stands for almost surely, with respect to $\mu_{\mathbb{R}^{n}}$, and $A\oplus B$ denotes the direct sum of the matrices $A$ and $B$ (that is, if $A$ is $n\times m$ and $B$ is $p\times q$, $A\oplus B$ is the $(n+p)\times(m+q)$ block diagonal matrix with $A$ as top diagonal block and $B$ as bottom diagonal block).

The model

The model of interest is the network autoregression

equation[equation omitted — 105 chars of source]

where $\boldsymbol{y}$ is the $n\times1$ vector of outcomes, $\lambda$ is a scalar parameter, $W$ is an interaction matrix, $X$ is an $n\times k$ matrix of regressors with full column rank and with $k\leq n-2$, $\beta\in \mathbb{R}^{k}$, $\sigma$ is a positive scale parameter, and $\boldsymbol{\varepsilon}$ is an unobservable $n\times1$ random vector with $\mathrm{E}(\boldsymbol{\varepsilon})=0$. The matrices $W$ and $X$ are assumed to be nonstochastic and known as, for instance, in Lee2004 .\footnote{At the cost of some additional complexity, one could, alternatively, take $W$ and/or $X$ to be stochastic, and condition on $W$ and/or $X$ under suitable exogeneity assumptions Bramoulle2009,Gupta2019. In that case the assumption $\mathrm{E}(\boldsymbol{\varepsilon})=0$ would be replaced by $\mathrm{E} (\boldsymbol{\varepsilon}\mid W,X)=0$. Allowing for endogeneity of $W$ and/or $X$ would instead require different methods; see Section (ref).} The entries of $W$ are supposed to reflect the pairwise interaction between the observational units; in particular, the $(i,j)$-th entry of $W$ is zero if unit $j$ is not deemed to be a neighbor of unit $i$. Some of the columns of $X$ may be spatial lags of some other columns (the spatial lag of a vector $x$ being the vector $Wx$). That is, in the terminology of social networks, we allow for \textquotedblleft contextual effects\textquotedblright\ or \textquotedblleft exogenous spillovers\textquotedblright. We assume that $\lambda$ is such that the model has a unique reduced form, or, in other words, that $\boldsymbol{y}$ is uniquely determined given $X$\ and $\boldsymbol{\varepsilon}$.\footnote{Identification analysis when a unique reduced form does not exist would require different tools; see, e.g., ChesherRosen2008.} This requires $S(\lambda )\coloneqq I_{n}-\lambda W$ to be nonsingular. We refer to the set $\Lambda_{\mathrm{u} }\coloneqq\{\lambda\in\mathbb{R}:\det(S(\lambda))\neq0\}$ as the unrestricted parameter space for $\lambda$. Note that the values of the (real) parameter $\lambda$ such that $\det(S(\lambda))=0$ are $\lambda=\omega^{-1}$ for any nonzero real eigenvalue $\omega$ of $W$, so $\Lambda_{\mathrm{u}}$ is the whole real line minus a number (less or equal to $n$) of isolated points.

mycommentIf there is no $X$, the model is referred to as a pure network autoregression.

When the index set of $\boldsymbol{y}$ has more than one dimension (e.g., individuals and time, or individuals and networks), it is often useful to include in the error term additive unobserved components relative to those dimensions. In that case, we take a fixed effects approach and treat the unobserved effects as parameters to be estimated. Accordingly, for inferential purposes, we incorporate the fixed effects into $\beta$ and the corresponding dummy variables into $X$. Two examples of fixed effects specifications that can be nested into the general model ((ref)) are given next.

mycomment(so the dimension of $\beta$ may be increasing with the sample size)

\begin{example2} (Panel data model) There are $N$ individuals, followed over $T$ time periods. Let $W_{t}$ be an $N\times N$ matrix describing the interaction between individuals at time $t$, and $\widetilde{X} $ an $NT\times\tilde{k}$ regressor matrix. A panel data version of the network autoregression ((ref)) is given by $\boldsymbol{y}_{it}=\lambda\sum _{ij}(W_{t})_{ij}\boldsymbol{y}_{jt}+\widetilde{x}_{it}^{\prime}\tilde{\beta }+\boldsymbol{u}_{it}$, for $i=1,\ldots,N$ and $t=1,\ldots,T$, where $(W_{t})_{ij}$ are the entries of $W_{t}$, and $\widetilde{x}_{it}^{\prime}$ are the $\tilde{k}\times1$ rows of $\widetilde{X}.$ The error $\boldsymbol{u} _{it}$ is decomposed into $\boldsymbol{c}_{i}+\sigma\boldsymbol{\varepsilon }_{it}$ (one-way model) or $\boldsymbol{c}_{i}+\boldsymbol{\alpha}_{t} +\sigma\boldsymbol{\varepsilon}_{it}$ (two-way model), where $\boldsymbol{c} _{i}$ and $\boldsymbol{\alpha}_{t}$ are, respectively, individual specific effects and time specific effects, and $\boldsymbol{\varepsilon}_{it}$ is an idiosyncratic error. Following a fixed effects approach (i.e., treating the random components $\boldsymbol{c}_{i}$ and $\boldsymbol{\alpha}_{t}$ as parameters to be estimated), the model can be written in the notation of equation ((ref)), with $W=\bigoplus_{t=1}^{T}W_{t}$, and, for the two-way model, $X=(\widetilde{X},\iota_{T}\otimes I_{N},I_{T}\otimes\iota_{N})$ and $\beta=(\tilde{\beta}^{\prime},c^{\prime},\alpha^{\prime})^{\prime}$, where $c$ and $\alpha$ are the vectors with entries $c_{i}$ and $\alpha_{t}$, respectively.\footnote{Obviously, for identification of $\beta$, one column of the matrix $(\iota_{T}\otimes I_{N},I_{T}\otimes\iota_{N})$ should be omitted from $X$, or some normalization should be imposed on the fixed effects, and no regressor should be constant over time or over individuals.} In most applications, $W_{t}$ is taken to be time invariant, say $W_{t}=W^{\ast}$ for all $t=1,\ldots,T$, so that $W=I_{T}\otimes W^{\ast}$.\qed

\end{example2}

\begin{example2} (Network fixed effects) There are $R$ networks, with network $r$ having $m_{r}$ individuals. The model is

equation[equation omitted — 215 chars of source]

where $W_{r}$ is the $m_{r}\times m_{r}$ interaction matrix of network $r$, $\widetilde{X}_{r}$ is the $m_{r}\times\tilde{k}$ regressor matrix of network $r$, $\tilde{\beta}$ is a $\tilde{k}\times1$ parameter, and $\boldsymbol{\alpha}_{r}$ is a network fixed effect. Stacking the equations in ((ref)) vertically, and following a fixed effects approach, the model can be written in the notation of equation ((ref)), with $\boldsymbol{y}=(\boldsymbol{y}_{1}^{\prime},\ldots,\boldsymbol{y}_{R} ^{\prime})^{\prime}$, $W=\bigoplus_{r=1}^{R}W_{r}$, $\beta=(\tilde{\beta }^{\prime},\alpha^{\prime})^{\prime}$, $\boldsymbol{\varepsilon} =(\boldsymbol{\varepsilon}_{1}^{\prime},\ldots,\boldsymbol{\varepsilon} _{R}^{\prime})^{\prime}$, and $X=(\widetilde{X},\bigoplus_{r=1}^{R} \iota_{m_{r}})$, where $\widetilde{X}\coloneqq(\widetilde{X}_{1}^{\prime },\ldots,\widetilde{X}_{R}^{\prime})^{\prime}$.\qed

\end{example2}

In the rest of the paper, unobserved effects are always treated as parameters. Two specific network autoregressions that will be used to illustrate our results are as follows.

\begin{example2} (Group Interaction model) A particular case of model ((ref)), which we refer to as the Group Interaction model, is when all members of a group interact homogeneously, that is, $W_{r}=\frac {1}{m_{r}-1}(\iota_{m_{r}}\iota_{m_{r}}^{\prime}-I_{m_{r}})\eqqcolon B_{m_{r} }$, for $r=1,\ldots,R$. Following Manski93, this specific structure has played a central role in the literature on peer effects. We say that the Group Interaction model is balanced if all group sizes $m_{r}$ are the same. In that case, letting $m$ denote the common group size, $W=I_{R}\otimes B_{m}$.\qed

\end{example2}

\begin{example2} (Complete Bipartite model) In a complete bipartite network the $n$ observational units are partitioned into two groups, of sizes $p$ and $q$ say, with all units within a group interacting with all in the other group, but with none in their own group wasserman94. In economics, such a structure arises commonly when modeling two-sided markets, before any specific matching between the two groups has taken place or when information about matchings is not available. The two groups could be, for instance, buyers and sellers, with each seller interacting with all buyers, and each buyer interacting with all sellers. For $p=1$ or $q=1$ this corresponds to the network known as a star. The adjacency matrix of a complete bipartite network is \[ A\coloneqq\mathopen\mathclose\bgroup\originalleft(

array[array omitted — 96 chars of source]

\aftergroup\egroup\originalright) . \] The associated row-normalized interaction matrix is\footnote{For an entrywise nonnegative matrix $B$ having all row-sums different from zero, the row-normalized version of $B$ is obtained by dividing each entry of $B$ by the corresponding row-sum, and is therefore a row-stochastic matrix.}

equation[equation omitted — 249 chars of source]

Alternatively, $A$ can be rescaled by its largest eigenvalue, yielding the symmetric interaction matrix

equation[equation omitted — 56 chars of source]

We refer to the network autoregressions with interaction matrix ((ref)) or ((ref)), as, respectively, the row-normalized Complete Bipartite model and the symmetric Complete Bipartite model. \qed

\end{example2}

Identifiability

This section explores identifiability of the parameters of a network autoregression in two cases. First, Section (ref) discusses what can be identified when no probabilistic assumptions beyond the maintained assumption $\mathrm{E}(\boldsymbol{\varepsilon})=0$ are imposed on the model. Then, Section (ref) considers adding assumptions on the second moment of $\boldsymbol{\varepsilon}$. Connections to the literature are discussed in Section (ref).

We employ a classical notion of global identifiability, according to which a parameter is said to be identified if it can be uniquely recovered from the distribution of the observables Koopmans1950, Rothenberg1971, Matzkin2007. The precise definitions we need are as follows. Consider a statistical model, defined as a family of distributions $\mathopen{}\mathclose\bgroup\originalleft\{ P_{\theta}:\theta\in\Theta\subseteq \mathbb{R}^{p}\aftergroup\egroup\originalright\} $ for some observable random vector on a certain sample space. A particular value $\theta_{1}\in\widetilde{\Theta} \subseteq\Theta$ of $\theta$ is said to be identified (from the distribution $P_{\theta}$) on a set $\widetilde{\Theta}$ if there is no other $\theta_{2}\in\widetilde{\Theta}$ such that $P_{\theta_{1}}=P_{\theta_{2}}$. If all values of $\theta$ in $\widetilde{\Theta}$ are identified on $\widetilde{\Theta}$, we say that the parameter $\theta$ is identified on $\widetilde{\Theta}$ (if the set $\widetilde{\Theta}$ is the whole parameter space $\Theta$, one often simply says that the model is identified). If all values of $\theta$ in $\widetilde{\Theta}$ except for those in a $\mu_{\mathbb{R}^{p}}$-null set are identified on $\widetilde{\Theta}$, we say that the parameter $\theta$ is generically identified on $\widetilde{\Theta}$. Identifiability can also be applied to functions of the parameter $\theta$, so that a function $f(\theta)$ is identified if it can be uniquely recovered from $P_{\theta}$. Formally, the function $f(\theta)$ is said to be identified on $\widetilde{\Theta}$ if $f(\theta_{1})\neq f(\theta_{2})$ implies $P_{\theta_{1}}\neq P_{\theta_{2}}$ for any $\theta _{1},\theta_{2}\in\widetilde{\Theta}$. The function $f(\theta)$ may extract a component of $\theta$. Note that a sufficient condition for a function $g(\theta)$ to be identified (on a set $\widetilde{\Theta}$) is that it can be recovered uniquely from a function $f(\theta)$ that is identified.\footnote{To see this, take arbitrary $\theta_{1},\theta_{2}$ (in $\widetilde{\Theta}$) such that $g(\theta_{1})\neq g(\theta_{2})$, and assume $g(\theta)$ can be recovered uniquely from $f(\theta)$, so that $f(\theta_{1})\neq f(\theta_{2} )$. Then $P_{\theta_{1}}\neq P_{\theta_{2}}$, because $f(\theta)$ is identified, which shows that $g(\theta)$ is identified.} Moments of $P_{\theta}$ (for example the mean, the variance matrix) are certainly identified functions of $\theta$ (because different moments imply different distributions), and so a sufficient condition for a function $g(\theta)$ to be identified (on a set $\widetilde{\Theta}$) is that it can be recovered uniquely from a moment of $P_{\theta}$; in this case we say that $g(\theta)$ is identified from a moment.

Note that identification from a moment is different from the concept, relevant in GMM estimation, of identification from a moment condition. Indeed, the concept of identification this paper refers to is detached from the choice of an estimator, unless one assumes that the distribution $P_{\theta}$ is known (up to $\theta$), in which case identification is equivalent to identification based on the (correctly specified) likelihood Rothenberg1971. Identification based on an extremum estimator is sufficient, but in general not necessary, for identification NeweyMcFadden94.

Identifiability from first moment

We start by studying generic identification, according to the definition just given, of $\lambda$ and $\beta$ when no distributional assumptions beyond $\mathrm{E}(\boldsymbol{\varepsilon})=0$ are imposed on the model.\footnote{Of course, the scale parameter $\sigma$, as well as any parameters that might be used to parametrize the variance of $\boldsymbol{\varepsilon}$, cannot be identified assuming only $\mathrm{E}(\boldsymbol{\varepsilon})=0$. Identification of $\sigma$ requires only very weak conditions (or a normalization) on the variance of $\boldsymbol{\varepsilon}$, whereas identifiability of any parameters affecting the variance of $\boldsymbol{\varepsilon}$ will depend on the particular parametric specification.}

propositionIn the network autoregression ((ref)), \begin{enumerate} • if $\mathrm{rank}(X,WX)>k$, the parameter $(\lambda,\beta)$ is generically identified on $\Lambda_{\mathrm{u}}\times\mathbb{R}^{k}$; • if $\mathrm{rank}(X,WX)=k$, no value of the parameter $(\lambda,\beta)$ is identified on $\Lambda_{\mathrm{u}}\times\mathbb{R}^{k}$. \end{enumerate}
mycommentNeedless to say, the crucial assumption here is that $\mathrm{E}(\boldsymbol{\varepsilon})$ does not depend on $\lambda$ and $\beta$. Also, recall that we are assuming that $W$ and $X$ are nonstochastic; otherwise, the assumption $\mathrm{E}(\boldsymbol{\varepsilon})=0$ should be replaced by $\mathrm{E}(\boldsymbol{\varepsilon}|W,X)=0$.
mycommentIn a pure model, $\mathrm{E}(\boldsymbol{y})$ is zero and hence cannot identify any parameters.
mycommentmaybe just state in the lemma conditions for param to be identified from first moment. then connect to identifiability afterwards, since identifiability of the moments may not be straightforward; see Goldsmith Imbens 2013 (for identif from second moment see de paula review, p 279). BUT\ WAIT, it is actually straightforward. If $\mathrm{rank}(X,WX)>k$ then $S^{-1}(\lambda)X\beta =S^{-1}(\tilde{\lambda})X\tilde{\beta}$ implies $\mathopen{}\mathclose\bgroup\originalleft( \lambda,\beta\aftergroup\egroup\originalright) =(\tilde{\lambda},\tilde{\beta})$ for any $\mathopen{}\mathclose\bgroup\originalleft( \lambda,\beta\aftergroup\egroup\originalright) ,(\tilde{\lambda},\tilde{\beta})\in\Lambda_{\mathrm{u}}\times\mathbb{R}^{k}$. That is, if $\mathrm{rank}(X,WX)>k$ $\mathopen{}\mathclose\bgroup\originalleft( \lambda,\beta\aftergroup\egroup\originalright) $ is identifiable

Proposition (ref) says that the parameters $\lambda$ and $\beta$ are generically identified (from the first moment of $\boldsymbol{y}$) if the matrices $X$ and $W$ are such that $\mathrm{rank}(X,WX)>k$. Conversely, if $\mathrm{rank}(X,WX)=k$, $\lambda$ and $\beta$ cannot be identified, and hence consistently estimated, without distributional assumptions beyond $\mathrm{E}(\boldsymbol{\varepsilon})=0$. A manifestation of this is that the 2SLS estimators of Kelejian98 and Lee2003, which are based on the specification of the first moment only of $\boldsymbol{y}$, are not defined if $\mathrm{rank}(X,WX)=k$.\footnote{The 2SLS estimator of Kelejian98 uses as instruments for the $k+1$ variables $(X,Wy)$ a subset of the columns of the linearly independent columns of $(X,WX,W^{2} X,\ldots)$. But when $\mathrm{rank}(X,WX)=k$, there are at most $k$ such linearly independent columns.} Note that if $\lambda$ is (generically) identified on $\Lambda_{\mathrm{u}}$ then it is also (generically) identified on any subset of $\Lambda_{\mathrm{u}}$.

The condition $\mathrm{rank}(X,WX)=k$ is trivially satisfied when $k=0$; in that case, $\mathrm{E}(\boldsymbol{y})$ is zero and therefore cannot identify any parameter. When $k>0$, the condition is typically very strong,\footnote{Indeed, for any given $W$ that is not a scalar multiple of the identity matrix, the set of (full rank) $n\times k$ matrices $X$ such that $\mathrm{rank}(X,WX)=k$ is a $\mu_{\mathbb{R}^{n\times k}}$-null set (with $\mu_{\mathbb{R}^{n\times k}}$ denoting the Lebesgue measure on the set of real $n\times k$ matrices). Accordingly, Proposition (ref) could be stated by saying that identification from the first moment of $\boldsymbol{y}$ is possible for generic parameter values $(\lambda,\beta)$ and for generic regressor matrices $X$.} but specific combinations of $W$ and $X$ such that $\mathrm{rank}(X,WX)=k$ may arise in some cases of interest, particularly in fixed effects models. An important class of cases when $\mathrm{rank}(X,WX)=k$ is characterized by a failure of the following condition.

conditionThere is no real eigenvalue $\omega$ of $W$ for which $M_{X}(\omega I_{n}-W)=0.$

Indeed, $M_{X}(\omega I_{n}-W)=0$ implies $M_{X}WX=0$, which is equivalent to $\mathrm{rank}(X,WX)=k.$ Thus, by Proposition (ref), any pair of matrices $X$ and $W$ violating Condition (ref) gives rise to a model in which $\lambda$ and $\beta$ cannot be identified from $\mathrm{E}(\boldsymbol{y})$. A condition equivalent to $M_{X}(\omega I_{n}-W)=0$ is $\operatorname{col}(\omega I_{n}-W)\subseteq\operatorname{col} (X)$. That is, a pair $(X,W)$ causes Condition (ref) to fail if and only if $\operatorname{col}(X)$ contains the subspace $\operatorname{col} (\omega I_{n}-W)$, for some real eigenvalue $\omega$ of $W$. Also, note that if, for a given $W$, Condition (ref) is violated for some $X=X_{1}$, then it is also violated for $X=(X_{1},X_{2})$, for any $X_{2}$ (such that $X$ is full rank). It is helpful to look at how failures of Condition (ref) can arise in the contexts of Examples (ref) and (ref).

\begin{example2} For the matrix $W=I_{R}\otimes B_{m}$ of a Balanced Group Interaction model (Example (ref)), the smallest eigenvalue is $\omega_{\min}=-\frac{1}{m-1}$, and $\operatorname{col}(\omega_{\min} I_{n}-W)=\operatorname{col}(I_{R}\otimes\iota_{m})$. Since $I_{R}\otimes \iota_{m}$ is the design matrix of the group fixed effects, it follows that Condition (ref) is violated if group fixed effects are included in a Balanced Group Interaction model.\qed

\end{example2}

\begin{example2} For both the row-normalized Complete Bipartite model and the symmetric Complete Bipartite model (Example (ref)), $\operatorname{col}(W)$ is spanned by the vectors $(\iota_{p}^{\prime} ,0_{q}^{\prime})^{\prime}$ and $(0_{p}^{\prime},\iota_{q}^{\prime})^{\prime}$. Hence, for both models, Condition (ref) is violated (for $\omega=0$) if $\operatorname{col}(X)$ contains $(\iota_{p}^{\prime},0_{q}^{\prime })^{\prime}$ and $(0_{p}^{\prime},\iota_{q}^{\prime})^{\prime}$. This is the case if the model contains an intercept for each of the two groups. \qed

\end{example2}

Examples (ref) and (ref) give cases in which Condition (ref) is violated, and hence $\mathrm{rank}(X,WX)=k$. Further examples in which Condition (ref) fails are given in Appendix (ref). We now turn our attention to examples in which $\mathrm{rank}(X,WX)=k$ even though Condition (ref) is satisfied.

\begin{example2} The condition $\mathrm{rank}(X,WX)=k$ is satisfied in the following instances of model ((ref)), when there are at least two groups ($R>1$):\footnote{In case (i), Condition (ref) is satisfied for generic matrices $\widetilde{X} _{1},\ldots,\widetilde{X}_{R}$ if the model is unbalanced, and is violated if the model is balanced (see Example (ref)). In cases (ii) and (iii), Condition (ref) is satisfied for any $\widetilde{X}$.}

enumerate• A Balanced Group Interaction model with contextual effects Liu2017. The model equation is $\boldsymbol{y}_{r}=\lambda B_{m}\boldsymbol{y}_{r}+\alpha\iota_{m}+\widetilde{X}_{r}\tilde{\beta} +B_{m}\widetilde{X}_{r}\delta+\sigma\boldsymbol{\varepsilon}_{r},$ for $r=1,\ldots,R,$ where $\alpha$ is an intercept and $\widetilde{X}_{r}$ is $m\times\tilde{k}$, with $0<\tilde{k}<R$. The matrix $X$ is given by $(\iota_{n},\widetilde{X},(I_{r}\otimes B_{m})\widetilde{X}),$ and has therefore $k=2\tilde{k}+1$ columns (recall that $\widetilde{X}$ is the $n\times\tilde{k}$ matrix obtained by stacking the matrices $\widetilde{X} _{1},\ldots,\widetilde{X}_{r}$ vertically).\footnote{The condition $\mathrm{rank} (X,WX)=k$ is satisfied whether the intercept is included in the model or not. Also, note that in this model, group fixed effects cannot be added, because $(\widetilde{X},(I_{r}\otimes B_{m})\widetilde{X},I_{R}\otimes\iota_{m})$ cannot have full column rank (this follows from $(m-1)^{-1}x+(I_{r}\otimes B_{m})x\in\operatorname{col}(I_{R}\otimes\iota_{m})$ for any $x\in \mathbb{R}^{n}$).} \begin{mycomment} if R=1 Condition (ref) violated, and $rank(\widetilde{X},W \widetilde{X})<2\tilde{k}$ if $\tilde{k}>1$ \end{mycomment} \begin{mycomment} ....(if R=1, col(X) is spanned by $\iota_{n}$ and something in the orth complement of $\iota_{n}$). ....(bramoulle sec 2.4.1.2) \end{mycomment} • The network fixed effects model ((ref)) with each $W_{r}$ being the symmetric or row-normalized adjacency matrix of a complete bipartite network, and with contextual effects. In this case, the matrix $X$ is given by $(\widetilde{X},W\widetilde{X},\bigoplus_{r=1}^{R}\iota_{m_{r}})$, and has therefore $k=R+2\tilde{k}$ columns, with $0\leq\tilde{k}<R$ .\qed \begin{mycomment} useful to rule out R=1 because in that case $rankX<k$ if $\tilde{k}>1$, and Condition (ref) violated \end{mycomment} \begin{mycomment} (.....bram discuss identif after removal FE in this model) (see complete-bipartite-CBG-assumption-1.m). \end{mycomment} • A Group Interaction model with group specific regression coefficients (denoted by $\tilde{\beta}_{r}$) and group fixed effects. The model equation is $\boldsymbol{y}_{r}=\lambda B_{m_{r}}\boldsymbol{y} _{r}+\widetilde{X}_{r}\tilde{\beta}_{r}+\alpha_{r}\iota_{m_{r}}+\sigma \boldsymbol{\varepsilon}_{r},$ for $r=1,\ldots,R,$ where the regressor matrix $\widetilde{X}_{r}$ is $m_{r}\times k_{r}$, with $0\leq k_{r}<m_{r}$. In this case, the matrix $X$ is given by $\bigoplus_{r=1}^{R}(\widetilde{X}_{r} ,\iota_{m_{r}})$, and has therefore $k=R+\sum_{r=1}^{R}k_{r}$ columns. \begin{mycomment} remember we need to make sure k<n (don't need to state this) \end{mycomment}\begin{mycomment} R=1 Condition (ref) violated \end{mycomment}

\end{example2}

\begin{example2} The condition $\mathrm{rank}(X,WX)=k$ is satisfied in the following network autoregressions with fixed effects and no regressors (i.e., the matrix $X$ contains only the dummies corresponding to the fixed effects):\footnote{In case (i) Condition (ref) is satisfied for any $W$. In cases (ii) and (iii) Condition (ref) is violated if and only if $W$ equals the interaction matrix of a Balanced Group Interaction model (i.e., if and only if $W=I_{T}\otimes B_{N}$ and $W=I_{R}\otimes B_{m}$ in cases (ii) and (iii), respectively).}

enumerate• The one-way model of Example (ref) with no regressors and time invariant interaction matrix, as, for instance, in Robinson2015. In this case, letting $W_{t}=W^{\ast}$ for each $t=1,\ldots,T$, we have $W=I_{T}\otimes W^{\ast}$, and $X$ contains only the individual fixed effects, i.e., $X=\iota_{T}\otimes I_{N}$, so that $k=N$. Since $WX=\iota_{T}\otimes W^{\ast}$, it follows that $\operatorname{rank}(X,WX)=\operatorname{rank} (\iota_{T}\otimes I_{N},\iota_{T}\otimes W^{\ast})=\operatorname{rank} (\iota_{T}\otimes(I_{N},W^{\ast}))=k.$ \begin{mycomment} here it is assumed $T>1$ but no need to say that \end{mycomment} \begin{mycomment} ......for the case of panel SAR with only individual fixed effects see individual_fixed_effects_identif_assumption_1.m \end{mycomment} \begin{mycomment} see individual_fixed_effects_identif_assumption_1 \end{mycomment} • The network fixed effects model of Example (ref) with no regressors and all matrices $W_{r}$'s being row-stochastic (a matrix is said to be row-stochastic if all its row sums are 1). In this case, $W=\bigoplus_{r=1}^{R}W_{r}$ and $X$ contains only the network fixed effects, i.e., $X=\bigoplus_{r=1}^{R}\iota_{m_{r}}$, so that $k=R$. Since each $W_{r}$ is row-stochastic, $\operatorname{rank}(X,WX)=\operatorname{rank} (\bigoplus_{r=1}^{R}\iota_{m_{r}},\bigoplus_{r=1}^{R}\iota_{m_{r}})=k.$ Note that, when $R=1$, this reduces to an intercept-only network autoregression ((ref)) with row-stochastic interaction matrix. • When $r$ is time, case (ii) also covers the case of a panel data model with time fixed effects. Hence, putting cases (i) and (ii) together, another example when $\mathrm{rank}(X,WX)=k$ is the two-way model of Example (ref) with no regressors (i.e., $X$ contains $k=N+T-1$ columns of $(\iota_{T}\otimes I_{N},I_{T}\otimes\iota_{N})$) and all matrices $W_{t}$'s being row-stochastic.\qed

\end{example2}

Examples (ref)--(ref) contain several cases in which $\mathrm{rank}(X,WX)=k$ and therefore $\lambda$ and $\beta$ cannot be identified from $\mathrm{E}(\boldsymbol{y})$ alone. The identification prospects in such cases are markedly different depending on whether Condition (ref) is satisfied (Examples (ref) and (ref)) or not (Examples (ref) and (ref)). In the former case, identification can be achieved, for example, by imposing higher moments restrictions (see Section (ref)). In the latter case, the identification problem is deeper, and a solution would require more drastic changes to the model (see Section (ref)).

mycommentfor nonexistence var 2SLS with 2 instruments see check_2SLS.m
mycommentSLM_MLE_OLS_2SLS_diff_distrib.m
mycommentWHAT HAPPENS CLOSE TO THE CASES WHERE THERE ARE PROBLEMS??? FOR EXAMPLE WHAT HAPPENS IF $M_{X}W$ IS CLOSE TO $0$ - see Xfixed-pert for CBG in plot-lik-score-SLM-with-adj.m

Identifiability from first two moments

Proposition (ref) gives a condition for $\lambda$ and $\beta$ to be generically identified when the model only specifies the first moment of $\boldsymbol{y}$. When that condition fails, identification may be achieved by imposing further restrictions on the model. The simplest of such restrictions is $\mathrm{var}(\boldsymbol{\varepsilon})=I_{n}$, in which case identification can occur via the first two moments of $\boldsymbol{y}$.

To see this, it is convenient to focus on a parameter space for $\lambda$ that is smaller than the unrestricted parameter space $\Lambda_{\mathrm{u} }\coloneqq\{\lambda\in\mathbb{R}:\det(S(\lambda))\neq0\}$. Indeed, the parameter space for $\lambda$ is usually taken to be a much smaller set than $\Lambda_{\mathrm{u}}$. Consider the case when $W$ has at least one (real) negative eigenvalue and at least one (real) positive eigenvalue. This is typically satisfied in both applications and theoretical studies. Denote the smallest real eigenvalue of $W$ by $\omega_{\min}$, and, without loss of generality, normalize the largest real eigenvalue to 1. The parameter space for $\lambda$ is often restricted to the largest interval containing the origin in which $S(\lambda)$ is nonsingular, that is, \[ \Lambda\coloneqq(\omega_{\min}^{-1},1), \] or a subset thereof (possibly independent of $n$) such as $(-1,1)$. Without such parameter space restrictions, the models are believed to be too erratic to be useful in practice, and $\lambda$ is difficult to interpret.

mycommentNote that the condition that $W$ has at least one negative eigenvalue and at least one positive eigenvalue rules out the case when $W$ is a scalar multiple of $I_{n}$, which trivially leads to non-identification of $\lambda,\beta$ from the first moment, or of $(\lambda,\beta,\sigma)$ from the first two moments.
propositionIn the network autoregression ((ref)) assume that $\mathrm{var}(\boldsymbol{\varepsilon})=I_{n}$ and that $W$ has at least one negative eigenvalue and at least one positive eigenvalue. The parameter $(\lambda,\beta,\sigma)$ is identified on $\Lambda\times\mathbb{R}^{k} \times(0,\infty)$.
mycommentseems to me for identifiability we can leave var matrix unknown. for estimation we would have to estimate it

In Proposition (ref), the parameters $\lambda$ and $\sigma$ are identified from $\mathrm{var}(\boldsymbol{y})$. Once $\lambda\ $is identified, $\beta$ can be identified from the first moment $\mathrm{E} (\boldsymbol{y})=(I_{n}-\lambda W)^{-1}X\beta$ (because $X$ has full rank). It should be noted that the restriction $\mathrm{var}(\boldsymbol{\varepsilon })=I_{n}$ is imposed only for simplicity, and one could certainly study identification under more general parametric structures for $\mathrm{var} (\boldsymbol{\varepsilon})$.

mycommentin first subm I had this, but removed following one comment by AE: is by no means necessary for identification from $\mathrm{var}(\boldsymbol{y})$. Indeed, one could assume some parametric structure for $\mathrm{var}(\boldsymbol{\varepsilon})$, say $\mathrm{var}(\boldsymbol{\varepsilon})=\Sigma(\eta)$, and study identifiability of the parameter $(\lambda,\sigma,\eta)$ from $\mathrm{var} (\boldsymbol{y})=\sigma^{2}(I_{n}-\lambda W)^{-1}\Sigma(\eta)(I_{n}-\lambda W^{\prime})^{-1}$, but we refrain from doing this here.
mycommentIf one can identify $\mathrm{var}(\boldsymbol{\varepsilon})$ independently of $\lambda$ and $\sigma^{2}$, then $\lambda$ and $\sigma^{2}$ can be identified by an obvious extension of Proposition (ref), in which $I_{n}$ is replaced by the value of $\mathrm{var}(\boldsymbol{\varepsilon})$. This is unrealistic! proof: proof of Proposition (ref) with $\mathrm{var}(\boldsymbol{\varepsilon} )=\Sigma.$ This proof is similar to the proof of Lemma 4.2 in Preinerstorfer2015. Assume that $\Sigma\coloneqq\mathrm{var} (\boldsymbol{\varepsilon})$ is positive definite and does not depend on $\lambda$ and $\sigma^{2}$. We establish that, if $\tilde{\sigma}^{2}S^{\prime} (\lambda)\Sigma^{-1}S(\lambda)=\sigma^{2}S^{\prime}(\tilde{\lambda} )\Sigma^{-1}S(\tilde{\lambda})$ for any two parameter values $(\lambda ,\sigma^{2}),(\tilde{\lambda},\tilde{\sigma}^{2})\in\Lambda\times(0,\infty)$, then $(\lambda,\sigma^{2})=(\tilde{\lambda},\tilde{\sigma}^{2})$. The maintained assumption that $W$ has at least one negative eigenvalue and at least one positive eigenvalue guarantees the existence of a nonzero vector $f\in\mathrm{null}(W-I_{n})$ and a nonzero vector $g\in\mathrm{null} (W-\omega_{\min}I_{n})$. Multiplying both sides of the equality $\tilde {\sigma}^{2}S^{\prime}(\lambda)\Sigma^{-1}S(\lambda)=\sigma^{2}S^{\prime }(\tilde{\lambda})\Sigma^{-1}S(\tilde{\lambda})$ by $f^{\prime}$ on the left and $f$ on the right gives $\tilde{\sigma}^{2}\mathopen{}\mathclose\bgroup\originalleft( 1-\lambda\aftergroup\egroup\originalright) ^{2}f^{\prime}\Sigma^{-1}f=\sigma^{2}(1-\tilde{\lambda})^{2}f^{\prime} \Sigma^{-1}f$. Since $1-\lambda>0$ for any $\lambda\in\Lambda$ , and $f^{\prime}\Sigma^{-1}f\neq0,$ the last equality is equivalent to $\tilde{\sigma}/\sigma=(1-\tilde{\lambda})/\mathopen{}\mathclose\bgroup\originalleft( 1-\lambda\aftergroup\egroup\originalright) $. Repeating with $g$ in place of $f$ gives $\tilde{\sigma}/\sigma=(1-\tilde {\lambda}\omega_{\min})/\mathopen{}\mathclose\bgroup\originalleft( 1-\lambda\omega_{\min}\aftergroup\egroup\originalright) $. Thus, we must have $\mathopen{}\mathclose\bgroup\originalleft( 1-\lambda\omega_{\min}\aftergroup\egroup\originalright) /\mathopen{}\mathclose\bgroup\originalleft( 1-\lambda\aftergroup\egroup\originalright) =(1-\tilde{\lambda}\omega_{\min})/(1-\tilde{\lambda})$. Since the function $\lambda\mapsto\mathopen{}\mathclose\bgroup\originalleft( 1-\lambda\omega_{\min}\aftergroup\egroup\originalright) /\mathopen{}\mathclose\bgroup\originalleft( 1-\lambda \aftergroup\egroup\originalright) $ is strictly increasing, we have $\lambda=\tilde{\lambda}$, and hence $\sigma^{2}=\tilde{\sigma}^{2}$.

At this point, it is worth considering the network (or spatial) error model

equation[equation omitted — 143 chars of source]

even though this specification is considerably less popular than the network autoregression ((ref)) in economic applications. The same set of assumptions as in the paragraph after equation ((ref)) will be maintained for model ((ref)). Proposition (ref) also applies to the network error model, because equations ((ref)) and ((ref)) imply the same variance structure for $\boldsymbol{y}$, and $\beta$ is trivially identified from the first moment in this model.

On the other hand, in the network error model ((ref)) the first moment of $\boldsymbol{y}$, $X\beta$, does not depend on $\lambda$, and hence cannot identify $\lambda$. Proposition (ref) can be interpreted as saying that $(\lambda,\beta)$ cannot be identified from $\mathrm{E} (\boldsymbol{y})$ in a network autoregression that \textquotedblleft behaves\textquotedblright\ like a network error model. This point is made precise by the following argument. If $\mathrm{rank}(X,WX)=k$, there exists a unique $k\times k$ matrix $A$ such that $WX=XA$, and hence $S^{-1} (\lambda)X=X(I_{k}-\lambda A)^{-1},$ for any $\lambda$ such that $S(\lambda)$ is invertible (it is easily seen that the eigenvalues of $A$ are eigenvalues of $W$, and therefore $I_{k}-\lambda A$ is invertible if $S(\lambda)$ is). It follows that, when $\mathrm{rank}(X,WX)=k$, the network autoregression $\boldsymbol{y}=S^{-1}(\lambda)X\beta+\sigma S^{-1}(\lambda )\boldsymbol{\varepsilon}$ can be written as $\boldsymbol{y}=X(I_{k}-\lambda A)^{-1}\beta+\sigma S^{-1}(\lambda)\boldsymbol{\varepsilon}$, which is a network error model with regression coefficients $(I_{k}-\lambda A)^{-1}\beta $.\footnote{According to Lemma (ref) in Appendix (ref), the model $\boldsymbol{y}=X(I_{k}-\lambda A)^{-1} \beta+\sigma S^{-1}(\lambda)\boldsymbol{\varepsilon}$ has the same profile quasi log-likelihood $l(\lambda,\sigma^{2})$ as model ((ref)), even though, clearly, the MLE of $\beta$ will be different in the two models.}

Connection with existing results

This section discusses connections between Propositions (ref) and (ref) and some related results available in the literature.

mycommentIt is useful to briefly compare Proposition (ref) with some related results available in the literature, obtained by two different approaches.

Identification from reduced form parameter

In the social network literature, identification of the structural parameters in model ((ref)) is typically established by verifying that those parameters can be uniquely recovered from the reduced form parameters Bramoulle2009,Blume2011,Kwok19. Such a strategy relies on the reduced form parameters being identified, which (in the case of a fixed $W$) would typically require individuals being observed repeatedly, over time or some other dimension.\footnote{Assuming, for simplicity, that we have only one covariate, $x$, the reduced form of the network autoregression $\boldsymbol{y}=\lambda W\boldsymbol{y}+\beta x+\sigma\boldsymbol{\varepsilon}$ is $\boldsymbol{y}=\Pi x+\sigma(I-\lambda W)^{-1}\boldsymbol{\varepsilon},$ where $\Pi\coloneqq\beta(I-\lambda W)^{-1}$ is an $n\times n$ matrix and therefore cannot, without further restrictions, be identified from the distribution of a single realization of the vector $\boldsymbol{y}$.} Hence, identification via reduced form parameters may not be appropriate in applications where a single observation of a network is available. Proposition (ref) can establish identifiability not only when repeated observations are available (in which case $W$ is block diagonal with identical blocks), but also when a single observation of the network is available. The following example considers a case when parameters can only be identified with repeated observations.

\begin{example2} Consider the row-normalized or symmetric Complete Bipartite model of Example (ref) with an intercept, a regressor, and a contextual effect term, so that $X=\mathopen{}\mathclose\bgroup\originalleft( \iota_{n},x,Wx\aftergroup\egroup\originalright) $ for some arbitrary vector $x\in\mathbb{R}^{n}$ (such that $X$ is full rank). This model is a particular case of equation (1) in Bramoulle2009. Because the matrices $I_{n}$, $W$, $W^{2}$ are linearly independent, Proposition 1 in Bramoulle2009 establishes that the parameters $\lambda$ and $\beta$ are identified from an i.i.d.\ sample of observations from the model, as long as $\beta_{2}\lambda+\beta_{3}\neq0$.\footnote{Parameter restrictions such as $\beta_{2}\lambda+\beta_{3}\neq0$ do not appear in Proposition (ref), due to the fact that generic identification is considered there.} However, in this model $\mathrm{rank} (X,WX)=k$, and therefore $\lambda$ and $\beta$ cannot be identified from a single observation of the model (without assumptions beyond $\mathrm{E} (\boldsymbol{\varepsilon})=0$), by Proposition (ref) .\footnote{In fact, Condition (ref) fails in this model (see Example (ref) in Appendix (ref) with a number of partitions equal to two), which implies $\mathrm{rank}(X,WX)=k$.} \qed

mycomment\footnote{We stress that non-identification here is for any $x$, not for generic $X$. As discussed above, Proposition (ref) implies identification for generic $X$.}

\end{example2}

To understand why, in Example (ref), identifiability requires repeated observations of the network it is helpful to apply Proposition (ref) to the case in which we observe the bipartite network $R\geq1$ times. Note that $R$ observations of a network autoregression with $X=\mathopen{}\mathclose\bgroup\originalleft( \iota_{n},x,Wx\aftergroup\egroup\originalright) $ correspond to a network autoregression with interaction matrix $W^{\ast}=I_{R}\otimes W$ and regressor matrix $X^{\ast}=\mathopen{}\mathclose\bgroup\originalleft( \iota_{nR},x^{\ast},W^{\ast}x^{\ast}\aftergroup\egroup\originalright) $ for some $x^{\ast}\in\mathbb{R}^{nR}$ (the regressor $x$ is not kept constant across the $R$ repetitions). When $W$ is the row-normalized or symmetric Complete Bipartite model, $\mathrm{rank}(X^{\ast},W^{\ast}X^{\ast})>k$ if and only if $R>1$. That is, Proposition (ref) establishes that identification is achieved if and only if $R>1$.

The applicability to the case of a single observation of a network is the most important difference between Proposition (ref) and Proposition 1 of Bramoulle2009. Note also that, contrary to Bramoulle2009's result, Proposition (ref) does not restrict attention to the case when $X$ contains contextual effects; our results can be used for that case, but also when no contextual effects are included, or only some contextual effects are included.

Asymptotic identification

The MLE that is typically used for a network autoregression is the one based on the likelihood that would obtain if $\boldsymbol{\varepsilon}$ were distributed as $\mathrm{N}(0,I_{n})$. Following common usage, we refer to this specific quasi MLE (QMLE) simply as the QMLE. Lee2004 studies asymptotic properties of the QMLE. The condition $\mathrm{rank}(X,WX)>k$ appearing in Proposition (ref) can be interpreted as a finite sample counterpart of Assumption 8 in Lee2004. Indeed, under the latter assumption (and other regularity assumptions) the limit of the Gaussian quasi-likelihood has a unique maximum at the true value of the parameters, which is a sufficient condition for identification; see NeweyMcFadden94 . Similarly, Conditions for $(\lambda,\sigma)$ to be identified from $\mathrm{var}(\boldsymbol{y})$ can be seen as finite sample counterparts of Assumption 9 in Lee2004.

mycommentFor the condition to be also necessary, however, correct specification of the likelihood is required [[[[in general???]]]])................

Identification from second moment

Proposition (ref) complements two results available in the literature that are concerned with identifiability from $\mathrm{var} (\boldsymbol{y})$ on a parameter space for $\lambda$ different from $\Lambda$. Firstly, Proposition (ref) extends Lemma 4.2 in Preinerstorfer2015, which establishes identification of $(\lambda ,\sigma)$ from $\mathrm{var}(\boldsymbol{y})$ on $(0,1)\times(0,\infty)$. Secondly, Lemma 4 in LeeYu2015 gives a sufficient condition for $(\lambda,\sigma)$ to be identified from $\mathrm{var}(\boldsymbol{y})$ on $\Lambda_{\mathrm{u}}\times(0,\infty)$, namely that the matrices $I_{n}$, $W+W^{\prime}$ and $W^{\prime}W$ are linearly independent. It is instructive to look at an example in which $(\lambda,\sigma)$ is identified from $\mathrm{var}(\boldsymbol{y})$ on $\Lambda_{\mathrm{u}}\times(0,\infty)$ but not on $\Lambda_{\mathrm{u}}\times(0,\infty)$, and therefore identification can be established by Proposition (ref) but not by Lemma 4 in LeeYu2015.

\begin{example2} Consider a balanced group interaction model (see Example (ref)) with $\mathrm{var}(\boldsymbol{\varepsilon})=I_{n}$. By Proposition (ref), $(\lambda,\sigma)$ is identified on $\Lambda\times(0,\infty)$ (and hence on any subset thereof). On the other hand, Lemma 4 in LeeYu2015 cannot establish identifiability, or non-identifiability, on $\Lambda\times(0,\infty)$ because the matrices $I_{n} $, $W+W^{\prime}$ and $W^{\prime}W$ are not linearly independent when $W=I_{R}\otimes B_{m}$. Indeed, when $W=I_{R}\otimes B_{m}$, $\sigma_{1} ^{2}\mathopen{}\mathclose\bgroup\originalleft( S^{\prime}(\lambda_{1})S(\lambda_{1})\aftergroup\egroup\originalright) ^{-1}=\sigma_{2} ^{2}\mathopen{}\mathclose\bgroup\originalleft( S^{\prime}(\lambda_{2})S(\lambda_{2})\aftergroup\egroup\originalright) ^{-1}$ if and only if $\sigma_{2}^{2}=m^{2}\sigma_{1}^{2}/(2\lambda_{1}+m-2)^{2}$ and $\lambda_{2}=-((m-2)\lambda_{1}+2(1-m))/(2\lambda_{1}+m-2)$, which shows that $(\lambda,\sigma)$ is not identifiable on $\Lambda_{\mathrm{u}}\times (0,\infty)$. Note however that $\lambda_{2}\notin\Lambda$ if $\lambda_{1} \in\Lambda$, which confirms that $(\lambda,\sigma)$ is identifiable on the smaller set $\Lambda\times(0,\infty)$.\qed

\end{example2}

Exploiting second moment restrictions to achieve identifiability has been considered in many areas of econometrics KomunjerNg11, and in particular in the peer effects literature; see, for instance, Graham2008, Theorem 3.2 in Davezies09, Rose2017, and Liu2017. The latter paper studies identifiability in the context of the model in Example (ref)(i) above.

mycomment(note that the model in the footnote in the example above is different from the one in case (b.iii) in Section (ref) because it has the intercept $\iota_{nR}$ instead of the group intercepts $I_{R}\otimes\iota_{n}$)
mycommentin the prev example ...allowing $X$ to vary across the $R$ observations
mycommentin the example 2SLS defined only with repeated obs
mycomment- so, are we saying that with $W^{\ast}=I_{R}\otimes W$ and $X^{\ast}=\mathopen{}\mathclose\bgroup\originalleft( \iota_{nR},x,Wx\aftergroup\egroup\originalright) $ we can always identify from first moment if $I_{n}$, $W$, $W^{2}$ are l.i. and $R>1$? yeah I\ think that's true: $\mathrm{rank} (X^{\ast},W^{\ast}X^{\ast})=\mathrm{rank}\mathopen{}\mathclose\bgroup\originalleft( \iota_{nR},x,W^{\ast }x,W^{\ast}\iota_{nR},W^{\ast}x,W^{\ast2}x\aftergroup\egroup\originalright) =[$with $W$ row-stoch$]\mathrm{rank}\mathopen{}\mathclose\bgroup\originalleft( \iota_{nR},x,W^{\ast}x,W^{\ast2}x\aftergroup\egroup\originalright) $ - A1 violated so MLE constant (=0) and OLS degenerate. - concerning the model with $R$ observations, if $\iota_{nR}$ is replaced by the FE $I_{R}\otimes\iota_{n}$ then $\mathrm{rank}(X^{\ast},W^{\ast}X^{\ast })=k$ for any $R$ so no identif from first moment even under repeated observations (this example discussed above, where we also say that Assumption 1 violated only if R=1)). For this model, non identification (after removal of FE)\ according to Bram, because $I,W^{\ast},W^{\ast2},W^{\ast3}$ are lin dep (in this case $\mathrm{rank}(X,WX)>k$)
mycommentFOR\ MYSELF: The presence of two distinct eigenvalues of $W$ is not sufficient for identifiability over the set of $\lambda\ $such that $S(\lambda)$ is nonsingular, because $\mathopen{}\mathclose\bgroup\originalleft\vert \mathopen{}\mathclose\bgroup\originalleft( 1-\lambda_{2}\aftergroup\egroup\originalright) /\mathopen{}\mathclose\bgroup\originalleft( 1-\lambda_{1}\aftergroup\egroup\originalright) \aftergroup\egroup\originalright\vert =\mathopen{}\mathclose\bgroup\originalleft\vert \mathopen{}\mathclose\bgroup\originalleft( 1-\lambda_{2} \omega_{\min}\aftergroup\egroup\originalright) /\mathopen{}\mathclose\bgroup\originalleft( 1-\lambda_{1}\omega_{\min}\aftergroup\egroup\originalright) \aftergroup\egroup\originalright\vert $ admits the sol $\lambda_{1}=\frac{\lambda_{2}\omega+\lambda_{2}-1} {2\lambda_{2}\omega-\omega}-1$ i.e. with 3 distinct eigenv we have If $\sigma_{1}^{2}(S^{\prime}(\lambda_{1})S_{\lambda_{1}})^{-1}=\sigma _{2}^{2}(S^{\prime}(\lambda_{2})S_{\lambda_{2}})^{-1}$ holds for $\lambda _{i}\ $such that $S_{\lambda_{i}}$ is nonsingular and $0<\sigma_{i}<\infty$ ($i=1,2$), then $\lambda_{1}=\lambda_{2}$ and $\sigma_{1}=\sigma_{2}.\bigskip$ $\sigma_{2}^{2}/\sigma_{1}^{2}=\mathopen{}\mathclose\bgroup\originalleft( 1-\lambda_{2}\aftergroup\egroup\originalright) ^{2}/\mathopen{}\mathclose\bgroup\originalleft( 1-\lambda_{1}\aftergroup\egroup\originalright) ^{2}$ $\sigma_{2}^{2}/\sigma_{1}^{2}=\mathopen{}\mathclose\bgroup\originalleft( 1-\lambda_{2}\omega_{\min}\aftergroup\egroup\originalright) /\mathopen{}\mathclose\bgroup\originalleft( 1-\lambda_{1}\omega_{\min}\aftergroup\egroup\originalright) $ $\sigma_{2}^{2}/\sigma_{1}^{2}=\mathopen{}\mathclose\bgroup\originalleft( 1-\lambda_{2}\omega\aftergroup\egroup\originalright) /\mathopen{}\mathclose\bgroup\originalleft( 1-\lambda_{1}\omega\aftergroup\egroup\originalright) $ 1) $\mathopen{}\mathclose\bgroup\originalleft( 1-\lambda_{2}\aftergroup\egroup\originalright) /\mathopen{}\mathclose\bgroup\originalleft( 1-\lambda_{1}\aftergroup\egroup\originalright) =\pm\mathopen{}\mathclose\bgroup\originalleft( 1-\lambda_{2}\omega_{\min}\aftergroup\egroup\originalright) /\mathopen{}\mathclose\bgroup\originalleft( 1-\lambda_{1}\omega_{\min}\aftergroup\egroup\originalright) $ 2) $\mathopen{}\mathclose\bgroup\originalleft( 1-\lambda_{2}\aftergroup\egroup\originalright) /\mathopen{}\mathclose\bgroup\originalleft( 1-\lambda_{1}\aftergroup\egroup\originalright) =\pm\mathopen{}\mathclose\bgroup\originalleft( 1-\lambda_{2}\omega\aftergroup\egroup\originalright) /\mathopen{}\mathclose\bgroup\originalleft( 1-\lambda_{1}\omega\aftergroup\egroup\originalright) $ 1) with + is impossible, but is possible with -, in which case we have $\lambda_{1}=\frac{\lambda_{2}\omega+\lambda_{2}-1}{2\lambda_{2}\omega-\omega }-1$ (see identification_3_eigenv.mw). We cannot have $\lambda_{1} =\frac{\lambda_{2}\omega+\lambda_{2}-1}{2\lambda_{2}\omega-\omega}-1$ and $\lambda_{1}=\frac{\lambda_{2}\omega_{\min}+\lambda_{2}-1}{2\lambda_{2} \omega_{\min}-\omega_{\min}}-1$ for $\omega\neq\omega_{\min}$
mycomment$ \begin{array} [c]{cc} 1 & -1\\ 0 & 1 \end{array} $, eigenvalues: $1$ $\frac{d}{d\rho}\frac{1+\rho^{2}b}{\mathopen{}\mathclose\bgroup\originalleft( 1-\rho\omega\aftergroup\egroup\originalright) ^{2} }=\allowbreak-\frac{2}{\mathopen{}\mathclose\bgroup\originalleft( \omega\rho-1\aftergroup\egroup\originalright) ^{3}}\mathopen{}\mathclose\bgroup\originalleft( \omega +b\rho\aftergroup\egroup\originalright) $ $\frac{d}{\mathop{}\!\mathrm{d}\lambda}\frac{1-\lambda\omega_{2}}{1-\lambda\omega_{1}} =\allowbreak\frac{1}{\mathopen{}\mathclose\bgroup\originalleft( \lambda\omega_{1}-1\aftergroup\egroup\originalright) ^{2}}\mathopen{}\mathclose\bgroup\originalleft( \omega_{1}-\omega_{2}\aftergroup\egroup\originalright) $ $1-\lambda\omega_{1}>0$, Solution is: $\mathopen{}\mathclose\bgroup\originalleft\{ \begin{array} [c]{ccc} \mathopen{}\mathclose\bgroup\originalleft( -\infty,\frac{1}{\omega_{1}}\aftergroup\egroup\originalright) & \text{if} & 0<\omega_{1} \wedge\frac{1}{\omega_{1}}\in \mathbb{R} \\ \mathopen{}\mathclose\bgroup\originalleft( \frac{1}{\omega_{1}},\infty\aftergroup\egroup\originalright) & \text{if} & \omega_{1} <0\wedge\frac{1}{\omega_{1}}\in \mathbb{R} \\ \emptyset & \text{if} & \frac{1}{\omega_{1}}\in \mathbb{R} \wedge\lnot\omega_{1}\in \mathbb{R} \\% \mathbb{R} & \text{if} & \omega_{1}=0\\ \mathopen{}\mathclose\bgroup\originalleft\{ 0\aftergroup\egroup\originalright\} & \text{if} & \omega_{1}\in \mathbb{C} \setminus \mathbb{R} \wedge\frac{1}{\omega_{1}}\in \mathbb{C} \setminus \mathbb{R} \end{array} \aftergroup\egroup\originalright. \allowbreak$

Invariance

So far, we have considered identifiability of a parameter $\theta$ from the distribution $P_{\theta}$ of an observable random vector $\boldsymbol{y}$. Sometimes, it may be appropriate to consider identifiability from some transformation of $\boldsymbol{y}$. Such a transformation might be dictated by the desire to eliminate some nuisance parameters, or, more generally, by invariance considerations ChamberlainMoreira2009. Suppose, for example, that $\theta$ is partitioned as $(\theta_{1}^{\prime },\theta_{2}^{\prime})^{\prime}$, where $\theta_{2}$ is not of direct interest. Particularly when the dimension of $\theta_{2}$ is large compared to the sample size, one may want to consider identification from a transformation of $\boldsymbol{y}$ whose distribution does not depend on $\theta_{2}$. In general, if $\theta_{1}$ is identified from the distribution of $\boldsymbol{y}$ it is also identified from the distribution of the transformation of $\boldsymbol{y}$. However, as we shall see in this section, this is not always the case. To analyze this point, we will need to discuss the full identifiability content of Condition (ref). We have seen in Section (ref) that, in a network autoregression, Condition (ref) is not necessary for identifiability from the distribution of $\boldsymbol{y}$; more precisely, it is necessary for identifiability from the first moment of $\boldsymbol{y}$, but not for identifiability from the second moment. Nonetheless, we will show that when Condition (ref) fails it is impossible to conduct inference that respects the symmetry properties of the model. More precisely, when Condition (ref) fails the model is invariant under certain transformations of the sample space, but identifiability is impossible from any function of the data that is invariant under those transformations. This result requires some general group invariance notions Lehmann2005, which are reviewed in Section (ref), and then applied to network autoregressions in Section (ref). Section (ref) contains the main identifiability result, and Section (ref) discusses some implications for likelihood inference.

General invariance notions

Let $\mathcal{G}$ be a group of one-to-one functions (transformations) from a space $\mathcal{S}$ into itself, with the group operation being the composition of functions. The group $\mathcal{G}$ induces a partition of $\mathcal{S}$ into equivalence classes, called orbits, with two elements of $\mathcal{S}$ being in the same orbit if there exists an element of $\mathcal{G}$ transforming one element into the other. The orbit of an element $x\in\mathcal{S}$ is therefore the set $\mathopen{}\mathclose\bgroup\originalleft\{ g(x):g\in\mathcal{G}\aftergroup\egroup\originalright\} $. A function on $\mathcal{S}$ is said to be invariant under $\mathcal{G}$ (or $\mathcal{G}$-invariant) if it is constant on the orbits of $\mathcal{G}$. A function on $\mathcal{S}$ is said to be a maximal invariant under $\mathcal{G}$ if it is invariant and takes different values on each orbit. A necessary and sufficient condition for a function on $\mathcal{S}$ to be invariant under $\mathcal{G}$ is that it depends on $x\in\mathcal{S}$ only through a maximal invariant under $\mathcal{G}$.

\begin{example2} (Scale invariance) Consider the group $\mathcal{G} =\{g_{\kappa}:\kappa>0\}$, where $g_{\kappa}$ is the function $y\mapsto\kappa y$ from $\mathbb{R}^{n}$ to $\mathbb{R}^{n}$. A maximal invariant under $\mathcal{G}$ is $y/\mathopen{}\mathclose\bgroup\originalleft\Vert y\aftergroup\egroup\originalright\Vert $, where $\mathopen{}\mathclose\bgroup\originalleft\Vert y\aftergroup\egroup\originalright\Vert \coloneqq\sqrt{y^{\prime}y}$, and hence a function on $\mathbb{R}^{n}$ is $\mathcal{G}$-invariant is and only if it depends on $y$ only through $y/\mathopen{}\mathclose\bgroup\originalleft\Vert y\aftergroup\egroup\originalright\Vert $.\footnote{Here we use the convention that $y/\mathopen{}\mathclose\bgroup\originalleft\Vert y\aftergroup\egroup\originalright\Vert =0$ if $y=0$. Invariance of $y/\mathopen{}\mathclose\bgroup\originalleft\Vert y\aftergroup\egroup\originalright\Vert $ holds because $y/\mathopen{}\mathclose\bgroup\originalleft\Vert y\aftergroup\egroup\originalright\Vert =\kappa y/\mathopen{}\mathclose\bgroup\originalleft\Vert \kappa y\aftergroup\egroup\originalright\Vert $ for any $\kappa>0$, maximality because, for any $y,\tilde{y}\in\mathbb{R}^{n}$, $y/\mathopen{}\mathclose\bgroup\originalleft\Vert y\aftergroup\egroup\originalright\Vert =\tilde {y}/\mathopen{}\mathclose\bgroup\originalleft\Vert \tilde{y}\aftergroup\egroup\originalright\Vert $ if and only if $\tilde{y}=ay$ for some $a>0$.}\qed

\end{example2}

The notion of group invariance can also be applied to a statistical model, defined to be a family of distributions on a certain sample space. In this case the set $\mathcal{S}$ upon which the group $\mathcal{G}$ acts is the sample space, and the functions $g$ in $\mathcal{G}$ are required to be measurable (with respect to the sample space $\sigma$-algebra), so that when $\boldsymbol{z}$ is a random variable with values in $\mathcal{S}$, $g(\boldsymbol{z})$ is too. The family of distributions $\mathopen{}\mathclose\bgroup\originalleft\{ P_{\theta }:\theta\in\Theta\aftergroup\egroup\originalright\} $ is said to be invariant under $\mathcal{G}$ if it is closed under the action of $\mathcal{G}$, that is, if every pair $g,\theta\in\mathcal{G}\times\Theta$ determines a unique element in $\Theta$ denoted by $\bar{g}(\theta)$, such that when $\boldsymbol{z}$ has distribution $P_{\theta}$, then $g(\boldsymbol{z})$ has distribution $P_{\bar{g}(\theta)}$. Note that the definition of invariance of a family of distributions requires $\theta$ to be identified. The functions $\theta\mapsto\bar{g}(\theta)$ form a group acting on the parameter space, denoted by $\mathcal{\bar{G}}$.

When a model for $\boldsymbol{y}$ is invariant under a group $\mathcal{G}$, it is natural to require that any inferential conclusion should be the same whether $y$ or $g(y)$ is observed, for any $g\in\mathcal{G}$. Accordingly, the data $y$ can be reduced to any $\mathcal{G}$-invariant function of $y$, with maximal reduction being achieved by reduction to a maximal invariant. This is what the so called principle of invariance advocates. For example, one would typically want to restrict attention to invariant loss functions (and, correspondingly, equivariant estimators) and, when the both the null and the alternative hypotheses are preserved under the group $\mathcal{G}$, to invariant tests. One fundamental results in the theory of invariance says that the distribution of any invariant statistic depends on $\theta$ only through a maximal invariant under $\mathcal{\bar{G}}$. Thus, in addition to a reduction in the sample space (from the dimension of $y$ to the dimension of the maximal invariant under $\mathcal{G}$), the principle of invariance generally also implies a reduction in the parameter space (from the dimension of $\theta$ to that of the maximal invariant under $\mathcal{\bar{G}}$).

\begin{example2} (A simple scale invariant model)\ Consider the statistical model defined by $\boldsymbol{y}=\sigma\boldsymbol{\varepsilon}$, where the distribution of $\boldsymbol{\varepsilon}$ depends on a parameter $\eta$, and does not depend on the parameter $\sigma>0$. Provided that $\eta$ is identified, the model is invariant under the group of scale transformations $\mathcal{G}=\{g_{\kappa}:\kappa>0\}$ (Example (ref)), because if $\boldsymbol{y}$ has distribution $P_{\sigma,\eta}$ $g(\boldsymbol{y})$ has distribution $P_{\bar{g}(\sigma,\eta)},$ with $\bar{g}(\theta)=(\kappa \sigma,\eta)$, for any $g\in\mathcal{G}$. The maximal invariant under $\mathcal{\bar{G}}$ is $\eta$, and indeed the distribution of the maximal invariant under $\mathcal{G}$, $\boldsymbol{y}/\mathopen{}\mathclose\bgroup\originalleft\Vert \boldsymbol{y} \aftergroup\egroup\originalright\Vert $, depends only on $\eta$.\qed

\end{example2}

Invariance of a network autoregression

We start our discussion of the invariance properties of a network autoregression by providing an invariance interpretation of Proposition (ref). For this we need to introduce a group that is often used for regression models kariya1980, and we need a definition of invariance of an expectation. For a given $n\times m$ full column rank matrix $Z$, define the group $\mathcal{G}_{Z}\coloneqq\{g_{\kappa ,\delta}:\kappa>0,\delta\in\mathbb{R}^{m}\}$, where $g_{\kappa,\delta}$ denotes the function $y\mapsto\kappa y+Z\delta$ (a one-to-one transformation of $\mathbb{R}^{n}$), and its subgroup $\mathcal{G}_{Z}^{1} \coloneqq\{g_{1,\delta}:\delta\in\mathbb{R}^{m}\}$. A maximal invariant under $\mathcal{G}_{Z}^{1}$ is $C_{Z}\boldsymbol{y}$, where $C_{Z}$ is an $\mathopen{}\mathclose\bgroup\originalleft( n-m\aftergroup\egroup\originalright) \times n$ matrix such that $C_{Z}C_{Z}^{\prime}=I_{n-m}$ and $C_{Z}^{\prime}C_{Z}=M_{Z}$, and a maximal invariant under $\mathcal{G}_{Z}$ is $v(\boldsymbol{y})\coloneqq C_{Z}\boldsymbol{y}/\mathopen{}\mathclose\bgroup\originalleft\Vert C_{Z} \boldsymbol{y}\aftergroup\egroup\originalright\Vert $ (with the convention that $v=0$ if $C_{Z} y=0$).\footnote{This can be established by direct verification of the definition of a maximal invariant. We provide the argument for $\mathcal{G}_{Z}$ (the argument for $\mathcal{G}_{Z}^{1}$ is similar). The statistic $v(y)$ is invariant because $v(\kappa y+Z\delta)=v(y)$ for any $\kappa>0$ and any $\delta\in\mathbb{R}^{m}$, since $C_{Z}Z=0$. It takes on different values on different orbits because, for any $y,\tilde{y} \in\mathbb{R}^{n}$, $v(y)=v(\tilde{y})$ if and only if $C_{Z}\tilde{y} =aC_{Z}y$ for some $a>0$, that is, $C_{Z}(\tilde{y}-ay)=0$, which is equivalent to $\tilde{y}=ay+Zb$ for some $b\in\mathbb{R}^{m}$.} We say that the expectation $\mathrm{E}_{\theta}(\boldsymbol{y})$ of a family of distributions $\mathopen{}\mathclose\bgroup\originalleft\{ P_{\theta}:\theta\in\Theta\aftergroup\egroup\originalright\} $ is $\mathcal{G} $-invariant if every pair $g,\theta\in\mathcal{G}\times\Theta$ determines a unique $\bar{g}(\theta)$ such that $\mathrm{E}_{\theta}(g(\boldsymbol{y} ))=\mathrm{E}_{\bar{g}(\theta)}(\boldsymbol{y})$.\footnote{Note that $\mathcal{G}$-invariance of $\mathrm{E}_{\theta}(\boldsymbol{y})$ is necessary but not sufficient for $\mathcal{G}$-invariance of $\mathopen{}\mathclose\bgroup\originalleft\{ P_{\theta} :\theta\in\Theta\aftergroup\egroup\originalright\} $ (the latter requires that every pair $g,\theta \in\mathcal{G}\times\Theta$ determines a unique $\bar{g}(\theta)$ such that $\mathrm{E}_{\theta}(\varphi(g(\boldsymbol{y})))=\mathrm{E}_{\bar{g}(\theta )}(\varphi(\boldsymbol{y}))$, for every measurable function $\varphi$).} Armed with these definitions, the non-identifiability result in Proposition (ref)(ii) can be seen as a consequence of the fact that, when $\mathrm{rank}(X,WX)=k$, the expectation $\mathrm{E}_{\lambda,\beta }(\boldsymbol{y})\coloneqq S^{-1}(\lambda)X\beta$ is $\mathcal{G}_{X}^{1} $-invariant.\footnote{If $\mathrm{rank}(X,WX)=k$, there exists a unique $k\times k$ matrix $A$ such that $WX=XA$, and hence $S^{-1}(\lambda )X=X(I_{k}-\lambda A)^{-1},$ for any $\lambda$ such that $S(\lambda)$ is invertible (note that $I_{k}-\lambda A$ is invertible if $S(\lambda)$ is, because the eigenvalues of $A$ must be eigenvalues of $W$). Hence $\mathrm{E}_{\lambda,\beta}(g_{1,\delta}(\boldsymbol{y}))=\mathrm{E}_{\bar {g}\mathopen{}\mathclose\bgroup\originalleft( \lambda,\beta\aftergroup\egroup\originalright) }(\boldsymbol{y})$, with $\bar{g}\mathopen{}\mathclose\bgroup\originalleft( \lambda,\beta\aftergroup\egroup\originalright) =(I_{k}-\lambda A)^{-1}\beta+\delta.$} This type of invariance implies that, when $\mathrm{rank}(X,WX)=k$, the expectation $\mathrm{E}_{\lambda,\beta}(\boldsymbol{y})$ can only identify a $k$-dimensional parameter, not the $(k+1)$-dimensional parameter $(\lambda,\beta)$. We will now show that, under the same rank condition, the full family of distributions underlying a network autoregression (not the mean only) is invariant under $\mathcal{G}_{X}^{1}$, in fact under $\mathcal{G} _{X}$. But before stating the result for the network autoregression, it is helpful to consider the network error model (ref).

assumptionThe distribution of $\boldsymbol{\varepsilon}$ does not depend on the parameters $\lambda$, $\beta$, and $\sigma$.

Let $P_{\theta}$ denote the distribution for $\boldsymbol{y}$ in the network error model, with $\theta\coloneqq(\lambda,\beta,\sigma,\eta)$, $\eta$ being a parameter indexing the distribution of $\boldsymbol{\varepsilon}$. Under Assumption (ref) and provided that $\theta$ is identified, the network error model is $\mathcal{G}_{X}$-invariant, because $g(\boldsymbol{y} )$ has distribution $P_{\bar{g}(\theta)},$ with $\bar{g}(\theta)=(\lambda ,\kappa\beta+\delta,\kappa^{2}\sigma^{2},\eta)$, for any $g\in\mathcal{G}_{X} $. Using the same parametrization, the result for network autoregression is as follows.

mycommentthe following was a footnote but actually not needed as I say it earlier that ran(X,WX) is trivially satisfied when X=0. Note that Lemma (ref) also holds for a pure network autoregression, by setting $X=0$.
mycommentwhen the model is semiparametric, check if I\ need to assume anything more
mycommentLemma (ref) also holds for a pure network autoregression, in which case $\operatorname{col}(X)$ is the trivial invariant subspace $\{0\}$ and $\mathcal{G}_{X}$ is the group $\mathcal{G}_{0}$ of scale transformations.
mycommentsee SEM_SLM_estimates.m
lemmaSuppose that, in the network autoregression ((ref)) Assumption (ref) holds and $\theta$ is identified. The model is $\mathcal{G}_{X}$-invariant if and only if $\mathrm{rank}(X,WX)=k$.

Of course, invariance under a certain group implies invariance under a subgroup of that group. A subgroup of $\mathcal{G}_{X}$ that will play an important role in Section (ref) is $\mathcal{G}_{X^{\ast}}^{1}$, for some $n\times k^{\ast}$ submatrix $X^{\ast}$ of $X$ ($k^{\ast}\leq k$). This is the group of transformations $y\mapsto y+X^{\ast}\delta$. For a network autoregression, it is clear from the proof of Lemma (ref) that a sufficient condition for the model to be $\mathcal{G}_{X^{\ast}}^{1} $-invariant (and also $\mathcal{G}_{X^{\ast}}$-invariant) is that $\mathrm{rank}(X^{\ast},WX^{\ast})=k^{\ast}$. Recall now that the principle of invariance says that if a model is invariant under a group $\mathcal{G}$ then the data should be reduced to $\mathcal{G}$-invariant functions of the data, i.e., to functions of the data that depend on $\boldsymbol{y}$ only through the maximal invariant under $\mathcal{G}$. In particular, imposition of invariance under $\mathcal{G}_{X^{\ast}}^{1}$ implies that $X^{\ast}$ is removed from the model, because the maximal invariant under $\mathcal{G} _{X^{\ast}}^{1}$ is $C_{X^{\ast}}\boldsymbol{y}$ and $C_{X^{\ast}}X^{\ast}=0$. In the following example, $X^{\ast}$ is a matrix of fixed effects.

mycomment$\boldsymbol{y}=S^{-1}(\lambda)X^{\ast}\beta^{\ast}+S^{-1}(\lambda)X^{\ast \ast}\beta^{\ast\ast}+\sigma S^{-1}(\lambda)\boldsymbol{\varepsilon}$ $g(\boldsymbol{y})=\kappa\boldsymbol{y+}X^{\ast}\delta$ $g(\boldsymbol{y})=\kappa S^{-1}(\lambda)X^{\ast}\beta^{\ast}+X^{\ast} \delta+\kappa S^{-1}(\lambda)X^{\ast\ast}\beta^{\ast\ast}+\kappa\sigma S^{-1}(\lambda)\boldsymbol{\varepsilon}$ inv if $\operatorname{col}(S^{-1}(\lambda)X^{\ast})=\operatorname{col} (X^{\ast})$ iff $\operatorname{col}(S(\lambda)X^{\ast})=\operatorname{col} (X^{\ast})$ iff $\operatorname{col}(WX^{\ast})\subseteq\operatorname{col} (X^{\ast})$ iff $\mathrm{rank}(X^{\ast},WX^{\ast})=k^{\ast}$ inv if $\operatorname{col}(S^{-1}(\lambda)X^{\ast\ast})=\operatorname{col} (X^{\ast})$ iff $\operatorname{col}(S(\lambda)X^{\ast})=\operatorname{col} (X^{\ast\ast})$ iff $\boldsymbol{y}=S^{-1}(\lambda)X_{1}\beta_{1}+S^{-1}(\lambda)X_{2}\beta _{2}+\sigma S^{-1}(\lambda)\boldsymbol{\varepsilon}$ invariant under $y\mapsto y+X_{1}\delta$ if (not iff prpbably?????? because I\ could also have $\operatorname{col}(S^{-1}(\lambda)X_{2})=\operatorname{col}(X_{1})$) $\operatorname{col}(S^{-1}(\lambda)X_{1})=\operatorname{col}(X_{1})$, or, which is the same, $\operatorname{col}(S(\lambda)X_{1})=\operatorname{col} (X_{1}).$ equivalent to $\mathrm{rank}(X_{1},WX_{1})=k_{1}$. $\mathrm{rank} (X^{\ast},WX^{\ast})=k^{\ast}$ is not necessary for $\mathcal{G}_{X^{\ast}} $-invariance when $k^{\ast}\leq k$ so for example $\boldsymbol{y}=\lambda W\boldsymbol{y}+X\beta+\sigma\boldsymbol{\varepsilon}$ is inv to $y\mapsto y+\iota\delta$ if W is row-stoch and X includes intercept

\begin{example2} Consider the network fixed effects model of Example (ref), and let $X_{\mathrm{FE}}:=\bigoplus_{r=1}^{R}\iota_{m_{r}} $, the $n\times R$ matrix of network fixed effects. Under Assumption (ref), and as long as each $W_{r}$ is row-stochastic, the model is $\mathcal{G}_{X_{\mathrm{FE}}}^{1}$-invariant, because $W_{r}\iota_{m_{r} }=\iota_{m_{r}}$ and therefore $\mathrm{rank}(X_{\mathrm{FE}},WX_{\mathrm{FE} })=R$. In this case, the principle of invariance suggests that data should be reduced to $\mathcal{G}_{X_{\mathrm{FE}}}^{1}$-invariant functions of $y$, that is, to functions that depend on $y$ only through the maximal invariant $C_{X_{\mathrm{FE}}}\boldsymbol{y}$ under $\mathcal{G}_{X_{\mathrm{FE}}}$. Since $C_{X_{\mathrm{FE}}}X_{\mathrm{FE}}=0$, reduction to $\mathcal{G} _{X_{\mathrm{FE}}}^{1}$-invariant statistics removes the fixed effects. We note that transformation by $C_{X_{\mathrm{FE}}}$ is equivalent to the transformation proposed in LeeLiuLin2010 to eliminate the network fixed effects (see Appendix (ref)). Other transformations used to remove fixed effects in this model may or may not satisfy the principle of invariance. For example, the transformation referred to as global differences in Bramoulle2009 does, whereas that referred to as local differences in the same paper does not. \end{example2}

The role of Condition (ref)

We are now in a position to discuss the implications of Condition (ref). Recall from Section (ref) that $\mathrm{rank} (X,WX)=k$ if Condition (ref) is violated. Thus, by Lemma (ref) and the principle of invariance, any failure of Condition (ref) is a case in which one would want to reduce the data to $\mathcal{G}_{X}$-invariant functions of the data. However, the imposition of $\mathcal{G}_{X}$-invariance causes an identifiability issue when Condition (ref) fails. To see this, observe that if Condition (ref) fails then $C_{X}S(\lambda)=(1-\lambda\omega)C_{X}$, and therefore premultiplying both sides of the network autoregression equation $S(\lambda)\boldsymbol{y}=X\beta+\sigma\boldsymbol{\varepsilon}$ by $C_{X}$ yields

equation[equation omitted — 111 chars of source]

Note that $\lambda$ and $\sigma$ appear together in the scale factor in front of $C_{X}\boldsymbol{\varepsilon}$. Thus, when Assumption (ref) is satisfied but Condition (ref) is not, $(\lambda,\beta,\sigma)$ cannot be separately identified from the distribution of $C_{X}\boldsymbol{y}$ and hence, since $C_{X}\boldsymbol{y}$ is a maximal invariant under $\mathcal{G}_{X}^{1}$, cannot be identified from the distribution of any $\mathcal{G}_{X}^{1}$-invariant statistic. Exactly the same conclusion obtains starting from the network error model $\boldsymbol{y}=X\beta+\sigma S^{-1}(\lambda)\boldsymbol{\varepsilon}$. The result is particularly perverse for the network autoregression: when Condition (ref) fails, and under Assumption (ref), the model is $\mathcal{G}_{X}^{1}$-invariant provided that its parameter $\theta$ is identifiable from the distribution of $\boldsymbol{y}$, and yet $\theta$ cannot be identified from any $\mathcal{G}_{X}^{1}$-invariant statistic.

It is possible to be more precise about the cause of this identification failure. Suppose Condition (ref) is violated for some eigenvalue $\omega$ of $W$, and let $\mathcal{\gamma}_{\omega}$ be the geometric multiplicity of $\omega$.\footnote{Note that, for fixed $W$ and $X$, the condition $M_{X}(\omega I_{n}-W)=0$ that leads to a violation of Condition (ref) can be satisfied at most by one eigenvalue $\omega$. This is because $M_{X}(\omega_{1}I_{n}-W)=M_{X}(\omega_{2}I_{n}-W)$ implies $\omega_{1}=\omega_{2}$. Also, note that $M_{X}(\omega I_{n}-W)=0$ implies that $\omega$ is real.} Recall from Section (ref) that a pair $(X,W)$ causes Condition (ref) to fail if and only if some of the columns of $X$ span the subspace $\operatorname{col}(\omega I_{n}-W)$. Observe that this requires $k\geq n-\mathcal{\gamma}_{\omega}$, because the dimension of $\operatorname{col}(\omega I_{n}-W)$ is $\mathrm{rank}(\omega I_{n}-W)=n-\mathrm{nullity}(\omega I_{n}-W)=n-\mathcal{\gamma}_{\omega}$. Let $X_{\omega}$ be the $n\times(n-\mathcal{\gamma}_{\omega})$ matrix containing the columns of $X$ that span $\operatorname{col}(\omega I_{n}-W)$, and reorder the columns of $X$ as in $X=(X_{\omega},X^{\ast})$, where $X^{\ast}$ is $n\times(k-(n-\mathcal{\gamma}_{\omega}))$, with $k-(n-\mathcal{\gamma }_{\omega})\geq0$. Generalizing the argument leading to equation ((ref)), if Condition (ref) fails then $C_{X_{\omega}} S(\lambda)=(1-\lambda\omega)C_{X_{\omega}}$, and therefore

equation[equation omitted — 192 chars of source]

where $\beta^{\ast}$ is the component of $\beta$ corresponding to $X^{\ast}$. This shows that, under Assumption (ref), $(\lambda,\beta ,\sigma)$ cannot be identified from the distribution of $C_{X_{\omega} }\boldsymbol{y}$ if Condition (ref) fails. That is, what really causes non-identification when Condition (ref) fails is the imposition of invariance with respect to the subgroup $\mathcal{G}_{X_{\omega }}^{1}$ of $\mathcal{G}_{X}$, and what we said above about $\mathcal{G} _{X}^{1}$-invariant statistics applies to the (larger) set of $\mathcal{G} _{X_{\omega}}^{1}$-invariant statistics. We formally state this result in the following theorem, and then provide an example.

theoSuppose that, in the network autoregression ((ref)) or in the network error model ((ref)), Assumption (ref) is satisfied. If Condition (ref) fails for some eigenvalue $\omega$ of $W$, then $(\lambda,\beta,\sigma)$ cannot be identified from the distribution of any $\mathcal{G}_{X_{\omega}}^{1}$-invariant statistic.

Theorem (ref) says that, for any $W$, there are matrices of regressors that make invariant inference impossible---these are the matrices leading to a violation of Condition (ref), that is, the matrices whose column space contains a subspace $\operatorname{col}(\omega I_{n}-W)$, for some eigenvalue $\omega$ of $W$. It is worth emphasizing that this result does not require any distributional assumption other than Assumption (ref). We provide an illustration of Theorem (ref) by revisiting a well-known identification failure in the context of the balanced group interaction model (see Example (ref)).

mycommentThe result about non-identifiability after reduction by $\mathcal{G}_{X}^{1}$-invariance is a direct consequence of the fact that the distribution of $Cy$ depends on $\lambda$ and $\sigma^{2}$ only through $\sigma/(1-\lambda\omega)$ (if the distribution of $\boldsymbol{\varepsilon}$ does not depend on $\lambda$ or $\sigma^{2}$)

\begin{example2} Consider a balanced group interaction model with group fixed effects. We have seen in Example (ref) that in this model Condition (ref) fails, because the columns of the fixed effects matrix $I_{R}\otimes\iota_{m}$ span $\operatorname{col}(\omega_{\min}I_{n} -W)$. That is, in the notation introduced just before equation ((ref) ), $X_{\omega_{\min}}=I_{R}\otimes\iota_{m}$. Theorem (ref) therefore implies that, under Assumption (ref), $(\lambda ,\beta,\sigma)$ cannot be identified from any statistic that is invariant under $\mathcal{G}_{I_{R}\otimes\iota_{m}}^{1}$, even though the model is invariant under that group (as shown in Example (ref)). Also, recall from Example (ref) that reducing the data to $\mathcal{G} _{I_{R}\otimes\iota_{m}}^{1}$-invariant functions of the data removes the group fixed effects. Thus, in this model, lack of identifiability from $\mathcal{G}_{I_{R}\otimes\iota_{m}}^{1}$-invariant statistics corresponds to the well-known identification failure that occurs upon removal of the group fixed effects Lee2007b.\qed

\end{example2}

Theorem (ref) is connected to the results obtained in Section (ref). Recall that the parameters of a network autoregression can be identified by suitable restrictions on the variance structure of $\boldsymbol{\varepsilon}$, regardless of whether Condition (ref) holds; for example, this is certainly the case if $\mathrm{var} (\boldsymbol{\varepsilon})=I_{n}$, by Proposition (ref). According to Theorem (ref), however, any result establishing identification from the distribution of $\boldsymbol{y}$ is pointless when Condition (ref) is not satisfied, because in that case the model is invariant under the group $\mathcal{G}_{X_{\omega}}^{1}$, and yet identification from $\mathcal{G}_{X_{\omega}}^{1}$-invariant functions of the data is impossible. We illustrate this point in the context of Example (ref).

\begin{example2} Consider the model in Example (ref). Due to the failure of Condition (ref), any result establishing identification from the distribution of $\boldsymbol{y}$ cannot help to achieve inference that respects the invariance properties of the model. This is so, for example, for Proposition 2 in dePaula2017, which establishes identification from the variance of $\boldsymbol{y}$ for the particular case $R=1$, when $\mathopen{}\mathclose\bgroup\originalleft\vert \lambda\aftergroup\egroup\originalright\vert <1$. While this identification result is correct, it should be noted that inference based on it cannot respect the invariance properties of the model. Indeed, the model is invariant under the group $\mathcal{G} _{\iota_{n}}^{1}$ of transformations $y\mapsto y+\alpha\iota_{n}$, $\alpha \in\mathbb{R}$, and yet, by Theorem (ref), identification is lost if, as advocated by the principle of invariance, we require inference to satisfy the same symmetries. So, for instance, any invariant test in this model can have only trivial power, and any equivariant estimator will be useless.\qed

mycommentprin of invariance advocates invariant tests for invariant testing problems
mycomment\begin{align*} y & =\mathopen\mathclose\bgroup\originalleft( I_{n}-\lambda B_{n}\aftergroup\egroup\originalright) ^{-1}(\beta_{0}\iota_{n}+X^{\ast }\beta^{\ast})+\sigma\mathopen\mathclose\bgroup\originalleft( I_{n}-\lambda B_{m}\aftergroup\egroup\originalright) ^{-1}\varepsilon\\ & =\frac{1}{1-\lambda}\beta_{0}\iota_{n}+\mathopen\mathclose\bgroup\originalleft( I_{n}-\lambda B_{m}\aftergroup\egroup\originalright) ^{-1}X^{\ast}\beta^{\ast}+\sigma\mathopen\mathclose\bgroup\originalleft( I_{n}-\lambda B_{m}\aftergroup\egroup\originalright) ^{-1}\varepsilon \end{align*} \begin{align*} g(y) & =\frac{1}{1-\lambda}\beta_{0}\iota_{n}+\alpha\iota_{n}+\mathopen\mathclose\bgroup\originalleft( I_{n}-\lambda B_{m}\aftergroup\egroup\originalright) ^{-1}X^{\ast}\beta^{\ast}+\sigma\mathopen\mathclose\bgroup\originalleft( I_{n}-\lambda B_{m}\aftergroup\egroup\originalright) ^{-1}\varepsilon\\ & =\frac{1}{1-\lambda}\mathopen\mathclose\bgroup\originalleft( \beta_{0}+(1-\lambda)\alpha\aftergroup\egroup\originalright) \iota _{n}+\mathopen\mathclose\bgroup\originalleft( I_{n}-\lambda B_{m}\aftergroup\egroup\originalright) ^{-1}X^{\ast}\beta^{\ast} +\sigma\mathopen\mathclose\bgroup\originalleft( I_{n}-\lambda B_{m}\aftergroup\egroup\originalright) ^{-1}\varepsilon \end{align*} induced group \[ \bar{g}\mathopen{}\mathclose\bgroup\originalleft( \lambda,\beta_{0},\beta^{\ast},\sigma\aftergroup\egroup\originalright) =\mathopen{}\mathclose\bgroup\originalleft( \lambda,\beta_{0}+(1-\lambda)\alpha,\beta^{\ast},\sigma\aftergroup\egroup\originalright) \] say what equivariance would be in this context An estimator $\hat{\theta}$ of a parameter $\theta$ based on data $y$ is said to be equivariant if $\hat{\theta}(y+\alpha\iota_{n})=\bar{g}(\hat{\theta}(y))$

\end{example2}

\begin{rem2} Theorem (ref) is related to some previous results in the literature. Martellosio2011 considers non-identifiability from $G_{X}$-invariant statistics in a network autoregression ((ref)) or spatial error model ((ref)) when $W$ is the matrix $B_{n}\coloneqq\frac{1}{n-1}(\iota_{n}\iota_{n}^{\prime}-I_{n})$. Preinerstorfer2015, p. 30, generalizes Martellosio2011's results to a regression model with correlated errors, which includes the particular case of a spatial error model ((ref)) with arbitrary $W$. Neither of these two papers, however, (i) discusses the relationship between non-identifiability from invariant statistics and identifiability from the first moment or from the first two moments; (ii) considers the set of $\mathcal{G}_{X_{\omega}}^{1}$-invariant statistics, which is larger than the set of $\mathcal{G}^{1}_{X}$-invariant statistics. \end{rem2}

mycommentmax inv under $\mathcal{G}_{X_{\omega}}^{1}$ $\mathcal{G}_{X}\coloneqq\{g_{\kappa,\delta}:\kappa>0,\delta\in\mathbb{R} ^{k}\}$, where $g_{\kappa,\delta}$ denotes the function $y\mapsto\kappa y+X\delta$ (a one-to-one transformation of $\mathbb{R}^{n}$), and its subgroup $\mathcal{G}_{X}^{1}\coloneqq\{g_{1,\delta}:\delta\in\mathbb{R}^{k}\}$. $r(y)=C_{X_{\omega}}y$ This can be established by direct verification of the definition of a maximal invariant. For $\mathcal{G}_{X_{\omega}},$ $r(y)$ is invariant because $r(y+X_{\omega}\delta)=v(y)$ for any $\delta\in\mathbb{R}^{...}$, since $C_{X_{\omega}}X_{\omega}=0$, and it is different on different orbits because, for any $y,\tilde{y}\in\mathbb{R}^{n}$, $r(y)=r(\tilde{y})$ if and only if $C_{X_{\omega}}(\tilde{y}-y)=0$, which is equivalent to $\tilde{y} =y+X_{\omega}b$ for some $b\in\mathbb{R}^{...}$, thus proving maximality.. statistic invariant under $\mathcal{G}_{X}^{1}$ (i.e. depends on $y$ only through $C_{X}y$) then it is invariant under $\mathcal{G}_{X_{\omega}}^{1}$ (i.e. depends on $y$ only through $C_{X_{\omega}}y$)
mycommentremoved as too cumbersome (but still interesting) remark: At the end of Section (ref) it was argued that Lemma (ref) cannot be useful for inference when Condition (ref) fails. It is also worth pointing out that, when Condition (ref) is satisfied, Proposition (ref) is particularly useful if $(\lambda,\beta)$ cannot be identified from the first moment. This is the case for a network autoregression with $\mathrm{rank}(X,WX)=k$ (or for a network error model). Examples of network autoregressions such that Condition (ref) is satisfied and $\mathrm{rank}(X,WX)=k$ are given in Section (ref) .
mycommentof course in (ref) the same holds on replacing $\mathcal{G}_{I_{R}\otimes\iota_{m}}^{1}$ with $\mathcal{G}_{X}^{1}\supseteq \mathcal{G}_{I_{R}\otimes\iota_{m}}^{1}$
mycommentSome technical remarks on the results of this section can be found in Section (ref).
mycommentQUESTION: How much of Propositions (ref) and (ref) can be extended to regression model with general $\Sigma(\lambda)$? \begin{proposition} For $y=X\beta+u,$ if $C\Sigma(\lambda)C^{\prime}$ is a scalar multiple of $I_{n-k}$ \begin{enumerate} • $l(\lambda)$ depends on y only through a term that does not depend on $\lambda.$ ???unbounded from above in a neighb of $a$$l_{\mathrm{a}}(\lambda)$ does not depend on $\lambda$ \end{enumerate} \end{proposition} \begin{pff} \begin{equation} l(\lambda)\coloneqq -\frac{n}{2}\log\mathopen\mathclose\bgroup\originalleft( y^{\prime}U(\lambda)y\aftergroup\egroup\originalright) -\frac{1} {2}\log\mathopen\mathclose\bgroup\originalleft( \det\mathopen\mathclose\bgroup\originalleft( \Sigma(\lambda)\aftergroup\egroup\originalright) \aftergroup\egroup\originalright) , \end{equation} [remember that King's lemma says $U(\lambda)=C^{\prime}(C\Sigma(\lambda )C^{\prime})^{-1}C]$ \begin{equation} l(\lambda)=-\frac{n}{2}\log\mathopen\mathclose\bgroup\originalleft( y^{\prime}C^{\prime}(C\Sigma(\lambda )C^{\prime})^{-1}Cy\aftergroup\egroup\originalright) -\frac{1}{2}\log\mathopen\mathclose\bgroup\originalleft( \det\mathopen\mathclose\bgroup\originalleft( \Sigma (\lambda)\aftergroup\egroup\originalright) \aftergroup\egroup\originalright) , \end{equation} if $C\Sigma(\lambda)C^{\prime}=f(\lambda)I_{n-k}$, \begin{align*} l(\lambda) & \coloneqq -\frac{n}{2}\log\mathopen\mathclose\bgroup\originalleft( f(\lambda)y^{\prime}M_{X}y\aftergroup\egroup\originalright) +\frac{1}{2}\log\mathopen\mathclose\bgroup\originalleft( \det\mathopen\mathclose\bgroup\originalleft( \Sigma^{-1}(\lambda)\aftergroup\egroup\originalright) \aftergroup\egroup\originalright) \\ & =-\frac{n}{2}\log\mathopen\mathclose\bgroup\originalleft( y^{\prime}M_{X}y\aftergroup\egroup\originalright) -\frac{n}{2}\log\mathopen\mathclose\bgroup\originalleft( f(\lambda)\aftergroup\egroup\originalright) +\frac{1}{2}\log\mathopen\mathclose\bgroup\originalleft( \det\mathopen\mathclose\bgroup\originalleft( \Sigma^{-1} (\lambda)\aftergroup\egroup\originalright) \aftergroup\egroup\originalright) \end{align*} $\log\mathopen{}\mathclose\bgroup\originalleft( \det\mathopen{}\mathclose\bgroup\originalleft( \Sigma^{-1}(\lambda)\aftergroup\egroup\originalright) \aftergroup\egroup\originalright) \rightarrow -\infty$ to establish lim$l(\lambda)$ we need to know $\lim f(\lambda)$ Let $\Sigma^{-1}(\lambda)=DD^{\prime}$. under some conditions (or always?) $f(\lambda)=\lambda_{1}^{2}(D)$-------check this Letting $\lambda_{i}$ denote the distinct eigenvalues ($s$ and $n_{i}$ are indep of $\lambda$ except for isolated points) \begin{align*} l(\lambda) & =-\frac{n}{2}\log\mathopen\mathclose\bgroup\originalleft( f(\lambda)y^{\prime}M_{X}y\aftergroup\egroup\originalright) +\log\mathopen\mathclose\bgroup\originalleft( \det\mathopen\mathclose\bgroup\originalleft( DD^{\prime}\aftergroup\egroup\originalright) \aftergroup\egroup\originalright) =-\frac{n}{2}\log\mathopen\mathclose\bgroup\originalleft( y^{\prime}M_{X}y\aftergroup\egroup\originalright) -\frac{n}{2}\log f(\lambda)+\log\mathopen\mathclose\bgroup\originalleft( \det\mathopen\mathclose\bgroup\originalleft( D\aftergroup\egroup\originalright) \aftergroup\egroup\originalright) \\ & =-\frac{n}{2}\log\mathopen\mathclose\bgroup\originalleft( y^{\prime}M_{X}y\aftergroup\egroup\originalright) -\frac{n}{2}\log\mathopen\mathclose\bgroup\originalleft( \lambda_{1}^{2}(D)\aftergroup\egroup\originalright) +\log\mathopen\mathclose\bgroup\originalleft( \prod_{i=1}^{s}\lambda_{i}^{n_{i} }(D)\aftergroup\egroup\originalright) \end{align*} careful as D not symmetric................ I would not be surprised if what we need here is Ass 4 in PP! from PP: If B is a symmetric and nonnegative de\ldots nite l l matrix, every l l matrix A that satis\ldots es AA' = B is called a square root of B; with B$^{1/2}$ we denote its unique symmetric and nonnegative de\ldots nite square root. Note thatevery square root of B is of the form B$^{1/2}$U for some orthogonal matrix U. $\Sigma^{-1}(\lambda)=DD^{\prime}=\Sigma^{-1/2}(\lambda)UU^{\prime} \Sigma^{-1/2}(\lambda)$............. ii) yes this certainly holds for general $\Sigma(\lambda)$ \end{pff}
mycommentSo, neither $l(\lambda)$ nor $l_{\mathrm{a} }(\lambda)$ can provide a suitable basis for inference when Condition (ref) is violated.
mycommentFurther properties of $l(\lambda)$ when Condition (ref) is violated are given by Lemma (ref) in the Supplement.
mycommentsee potscher prein particularly remark 2.3
mycommentOLD for lambda only:The profile log-likelihood functions of $\lambda$ in a SAR and in a network error model, given in equation ((ref)) and equation ((ref)) in the Supplement, respectively, are identical if and only if $M_{S(\lambda)X}=M_{X}$, for any $\lambda$ such that $S(\lambda)$ is invertible. This occurs if and only if $\operatorname{col}(S(\lambda )X)=\operatorname{col}(X)$, for any $\lambda$ such that $S(\lambda)$ is invertible, which, as noted above, is equivalent to the condition that $\operatorname{col}(X)$ is an invariant subspace of $W$.
mycommentNote that $(\lambda,\sigma)$ is identifiable from $\mathrm{var}(y|W,X)$, via Proposition (ref). but inference based on $y$ is imposs - see (ref) Provided that $\mathrm{var}(\boldsymbol{\varepsilon}|W,X)$ does not depend on $\lambda$ or $\sigma$, the parameter $(\lambda,\sigma)$ is identified on $\Lambda \times(0,\infty)$, by Proposition (ref). However, since Condition (ref) fails for this model, $(\lambda,\sigma)$ cannot be identified from $\mathrm{var}(Cy|W,X)$.

Likelihood

We now study the consequences of Theorem (ref) for likelihood inference. We consider the QMLE based on the original data, as introduced in Section (ref), but also the QMLE after transformation by $C_{X_{\omega}}$ and the QMLE after transformation by $C_{X}$. Transformation by $C_{X_{\omega}}$ is relevant, for example, when $X_{\omega}$ is a matrix of fixed effects and one wishes to remove the fixed effects prior to estimation Lee2007b,LeeLiuLin2010,LeeYu2010. Transformation by $C_{X}$, on the other hand, is relevant when the model is $\mathcal{G}_{X}^{1}$-invariant, which is certainly the case when Condition (ref) fails. Note that, when the model is $\mathcal{G}_{X}^{1}$-invariant, the QMLE after transformation by $C_{X}$ is equivalent to the so-called adjusted QMLE, which is obtained from the QMLE by centering the profile score for $(\lambda ,\sigma^{2})$ Yu2015. For estimation of $(\lambda,\sigma^{2})$, it is well known that the adjusted QMLE usually performs better than the QMLE when the dimension of $\beta$ is large with respect to the sample size $n$ (including in fixed effects models, in which case the dimension of $\beta$ is increasing with $n$).

We denote by $l(\lambda,\beta,\sigma^{2};y)$ the Gaussian quasi log-likelihood for $(\lambda,\beta,\sigma^{2})$ in a network autoregression or in a network error model, by $l(\lambda,\beta,\sigma^{2};Ay)$ the corresponding log-likelihood obtained after premultiplying the model by a matrix $A$, and by $l(\lambda;y)$ the profile likelihood for $\lambda$. The next result shows that, when Condition (ref) fails, likelihood estimation is fruitless, before or after transformation of the data.

mycommentcould alternatively express the prop in terms of $l(\lambda,\sigma^{2})$ and $l_{\mathrm{a}} (\lambda,\sigma^{2})$ (the latter corresponds to density max inv under $G_^1_X$)
mycommentunnique max of limiting quasi lik is nec and suff for identification from the quasi lik - see Lee and Yu 2015 and cf Remark (ref)
mycommentThe reverse implication $\operatorname{col}(X)$ is an invariant subspace of $W$ then Condition (ref) is violated does not hold. For example, consider the group interaction weights matrix of Example (ref) (UGI) and the fixed effects matrix $X=\bigoplus_{i=1}^{R}\iota_{m_{i}}$. In that case, $\operatorname{col}(X)$ is an invariant subspace of $W$ for any $R$ and for any $m_{1},\ldots,m_{R}$, but Condition (ref) is violated (with $\omega=\omega_{\min}$) only if the model is balanced (i.e. $m_{1}=\ldots=m_{r}$).
propositionConsider the network autoregression ((ref)) or the network error model ((ref)), with $\lambda$ such that $\det(S(\lambda))\neq0$, and $y\notin\operatorname{col}(X)$. If Condition (ref) is violated for some eigenvalue $\omega$ of $W$, then: \begin{enumerate} • the log-likelihood functions $l(\lambda,\beta,\sigma^{2} ;C_{X_{\omega}}y)$ and $l(\lambda,\beta,\sigma^{2};C_{X}y)$ do not depend on $(\lambda,\beta,\sigma^{2})$; • the profile score associated with the profile log-likelihood function $l(\lambda;y)$ does not depend on $y\ $and $X$. \end{enumerate}

Part (i) of Proposition (ref) says that the functions $l(\lambda,\beta,\sigma^{2};C_{X_{\omega}}y)$ and $l(\lambda ,\beta,\sigma^{2};C_{X}y)$ are constant. by Proposition (ref) this cannot be the case for the likelihood based on $y$, $l(\lambda,\beta,\sigma^{2};y)$. However, part (ii) of Proposition (ref) establishes that the first derivative of the profile likelihood $l(\lambda;y)$ does not depend on the data, and consequently that a QMLE based on $l(\lambda;y)$ cannot depend on the data if it exists.\footnote{This is similar to a result contained in Theorem 1 of KelejianPruchaYuzefovich2006. That result establishes that, in a network autoregression with $W=B_{n}$, the 2SLS estimator of $\lambda$ considered there does not depend on the data.} This results can be understood in terms of the invariance results in Section (ref): if Condition (ref) is violated for an eigenvalue $\omega$ of $W$, then $l(\lambda;y)$ is a $\mathcal{G}_{X_{\omega}}^{1}$-invariant loss function (see equation ((ref)) in the proof of the proposition), and as such it cannot produce a useful estimator.

mycommentREVISEEEE It is easily verified that the maximal invariant under the group induced by $\mathcal{G}_{X}$ on the parameter space is $\lambda$ (it would be $\mathopen{}\mathclose\bgroup\originalleft( \lambda,\eta\aftergroup\egroup\originalright) $ in the presence of a parameter $\eta$ in the distribution of $\boldsymbol{\varepsilon}$). This may seem to contradict one of the fundamental results on invariance, which is usually stated by saying that the distribution of an invariant statistic depends only on a maximal invariant induced on the parameter space Lehmann2005. The apparent contradiction is due to the non-identification caused by the violation of Condition (ref)
mycommentThis result is a consequence of the fact that the only part of the profile log-likelihood function $l(\lambda;y)$ that depends on the data $y\ $and $X$ does not depend on $\lambda$;$\ $see equation ((ref)) in the proof of the proposition.

\begin{example2} We have seen in Example (ref) that, for a network fixed effects model in which all interaction matrices $W_{r}$ are row-stochastic, transformation by $C_{X_{\mathrm{FE}}}$ is equivalent to the transformation proposed in LeeLiuLin2010. Thus the QMLE based on the Gaussian likelihood $l(\lambda,\beta,\sigma^{2};C_{X_{\omega}}y)$ is equivalent to the QMLE considered in LeeLiuLin2010. In the particular case of a balanced group interaction model (i.e., $W_{r}=B_{m_{r}}$, for each $r=1,\ldots,R$, and $m_{1}=m_{2}=\ldots= m_{R}$; see Example (ref), Condition (ref) fails, with $X_{\omega}=X_{\mathrm{FE}}$; see Example (ref). Thus, for the balanced group interaction model, Proposition (ref) establishes that the Gaussian quasi-likelihood based $C_{X_{\mathrm{FE}}}y$ and that based on $C_{X}y$ are constant functions of the parameters. \end{example2}

mycommentnote MLE is equiv under group under which model is inv. adj MLE is inv
mycomment............Consider the network autoregression ((ref)) with $W$ and $X$ such that Condition (ref) fails, and suppose Assumption (ref) holds. Recall that the model is $\mathcal{G}_{X}$-invariant by Lemma (ref), and that a maximal invariant under $\mathcal{G}_{X}$ is $v\coloneqq C_{X}y/\mathopen{}\mathclose\bgroup\originalleft\Vert C_{X}y\aftergroup\egroup\originalright\Vert $. Since $\sigma /(1-\lambda\omega)$ appears as a scale factor in (ref), it follows that the distribution of $v$, and hence of any $\mathcal{G}_{X}$-invariant statistic, is free of $\lambda$ if the distribution of $\boldsymbol{\varepsilon}$ does not depend on $\lambda$ (and of course is also free of $\sigma^{2}$ and $\beta$ under the full Assumption (ref)). It is easily verified that the maximal invariant induced on the parameter space is $\lambda$. This might seem to contradict one of the fundamental results on invariance, which is usually stated by saying that the distribution of an invariant statistic depends only on a maximal invariant induced on the parameter space Lehmann2005. The apparent contradiction is due to the non-identifiability caused by the violation of Condition (ref) .\textcolor{red}{same applies to SEM}
mycommentWe say that $\lambda$ is identifiable from the quasi profile likelihood $l(\lambda)$ if $l(\lambda_{1})=l(\lambda_{2})$ for almost all $y\in \mathbb{R}^{n}$. Of course, identifiability of $\lambda$ from $\mathrm{E} (y|W,X)$ (Proposition (ref)) or from $\mathrm{var}(y|W,X)$ (Lemma (ref)) implies identifiability from $l(\lambda)$.
mycommentfor the fact that the eigenvalues of $A$ are eigenvalues of $W$ see $\backslash$ citep[e.g.,][Theorem 3.9]\{StewartSun1990\}[[[[[[[[[[no need for reference this is straightforward]]]]]]]]]]
mycommentThis not need anymore as I moved assumption: Proposition (ref) does not require the assumption, made in Section (ref), that $W$ has at least one negative eigenvalue and at least one positive eigenvalue.

Simulations

Section (ref) contains several examples in which $\mathrm{rank} (X,WX)=k$, and therefore $\lambda$ and $\beta$ cannot be identified from the first moment of $\boldsymbol{y}$, that is, cannot be identified when the only assumption about $\boldsymbol{\varepsilon}$ is $\mathrm{E} (\boldsymbol{\varepsilon})=0$. As noted in that section, however, the condition $\mathrm{rank}(X,WX)=k$ is very strong in general. What might be more relevant in typical applications is that the condition is close, in some sense, to being satisfied. In that case, it would be natural to expect that identification from the first moment will be weak. This section analyses, by simulation, the consequences of near non-identification. We consider two Monte Carlo experiments, designed to study what happens close to, respectively, (i) a case when Condition 1 holds but $\mathrm{rank}(X,WX)=k$, (ii) a case when Condition 1 fails.

In both experiments, we draw $10,000$ replications from model ((ref)), with errors drawn from either a standard normal distribution or a gamma distribution with shape parameter 1 and scale parameter 1, demeaned by the population mean. Mean, variance, skewness, and kurtosis are 0, 1, 0, and 3 for the former distribution and 0, 1, 2, and 9 for the latter. The main objective of the simulations is to study the behavior of an estimator that uses only first moment information. We focus on the two-stage least squares estimator (2SLSE) with instruments for $(X,Wy)$ given by (the linearly independent columns of) $(X,WX,W^{2}X)$ Kelejian98. We study two implications of near non-identification: (i) accuracy of the estimator; (ii) adequacy of first-order asymptotic approximation to the distribution of the estimator. Accuracy of the estimator is measured by the median square error. The root median square error is reported rather than the more usual root mean square error because, in the setting we are considering, the variance of the 2SLS estimator does not exist Roberts95. Adequacy of the first-order asymptotic approximation is measured by the coverage of 95% Wald confidence intervals.\footnote{For a parameter $\phi$, the 95%\ Wald confidence interval is $\hat{\phi}\pm1.96\sqrt{\hat{v}}$, where $\hat{\phi}$ is an estimator of $\phi$ and $\hat{v}$ its estimated asymptotic variance. Expressions for the asymptotic variance of the 2SLS is standard, and that of the QMLE is given in Lee2004, Theorem 3.2.} As a benchmark, we consider the (quasi) maximum likelihood estimator based on the Gaussian likelihood, abbreviated by (Q)MLE; see Section (ref)). Contrary to the 2SLSE, the QMLE also uses second moment information.

First experiment

In the first experiment, $n$ is either $100$ or $1000$, and $W$ is a row-normalized 2-ahead 2-behind interaction matrix (before row-normalization, this is a matrix with all entries in the two diagonals above and the two diagonals below the main diagonal equal to one, and zero everywhere else), and the model has a single regressor equal to $\iota_{n}+bz$, where $b\in \mathbb{R}$ and $z\thicksim\mathrm{N}(0,I_{n})$, with $z$ being generated once, for each $n$, and then kept fixed across replications. If $b=0$, then $\mathrm{rank}(X,WX)=k=1$ (see Example (ref)(ii), with $R=1$), and therefore the parameters $\lambda$ and $\beta$ cannot be identified from the first moment. Thus, we expect any estimator of $\lambda$ and $\beta$ that relies entirely on the specification of the first moment of $\boldsymbol{y}$ to perform poorly if $b$ is close to $0$ (and to be undefined when $b=0$). The true values of $\lambda$, $\beta$, $\sigma$ are set to $\lambda_{0}=0$, $\beta_{0}=0.1,1$, and $\sigma_{0}=1$. Table (ref) displays the root median square error of the 2SLSE and (Q)MLE of $\lambda$ and $\beta$. Dots indicate nonexistence of the estimator. The 2SLSE exploits the (correct) specification of the first moment, whereas the QMLE also exploits the (correct, given that the data are generated under the assumption $\operatorname{var}(\boldsymbol{\varepsilon})=I_{n}$) specification of the second moment. Let us look at the case $\beta_{0}=1$ first. For both $\lambda$ and $\beta$, and for both the normal and the gamma distributions, the performance of the 2SLSE is satisfactory, compared to the (Q)MLE benchmark, when $b=1$, but deteriorates rapidly as $b$ gets smaller. Such a deterioration is due to both the bias and the dispersion of the 2SLSE growing large as $b$ decreases, for any $n$. When $b=0,$ the 2SLSE is not defined. On the contrary, due to the fact that it also exploits second moment information, the (Q)MLE is not much affected by the lack of identifiability from the first moment that occurs when $b=0$. Indeed, the root median square error of the (Q)MLE is considerably less sensitive to $b$, and the (Q)MLE does well even when $b=0$. Moving to the case $\beta_{0}=0.1$, it is natural to expect that identifiability from the fist moment may become more difficult when $\beta$ is near zero. Indeed, all values of $(\lambda,\beta)$ in the $\mu_{\mathbb{R} ^{2}}$-null set $\Lambda_{\mathrm{u}}\times\{0\}$ cannot be identified from the first moment --- see case (a) in the proof of Proposition 3.1. In particular, the Monte Carlo results show that the 2SLE of $\lambda$ performs much worse than the (Q)MLE, even when $b=1$.

Table (ref) confirms that estimation based on the first moment is impossible when $\mathrm{rank}(X,WX)=k$ and difficult in cases close to $\mathrm{rank}(X,WX)=k$. A second consequence of $X$ and $W$ being such that $\mathrm{rank}(X,WX)$ is close to being equal to $k$ is that typical asymptotic approximations may become unreliable. Table (ref) displays coverages of 95% two-sided Wald confidence intervals based on asymptotic normality, in the same setting at Table (ref). When $\beta_{0}=1$, the empirical coverages for the 2SLSE are close to the nominal one when $b=1$, but get further and further away as $b$ decreases. When $\beta_{0}=0.1,$ the empirical coverages for the 2SLSE are poor even when $b=1$. The (Q)MLE, on the other hand does well in terms of coverage even when $b=0$ and even when $\beta_{0}=0.1$, again due to the fact that, in these simulations, it does not rely on the specification of the first moment only, but also exploits the (correct in these simulations) specification of the second moment.

mycommentI haven;t used the sandwic yet for gamma but this should not make much difference
table[table omitted — 7,764 chars of source]

Second experiment

The second data generating process is the Group Interaction model of Example (ref) with fixed effects. The model equation is

equation[equation omitted — 211 chars of source]

The number of groups is $R=50,100,200,$ and for the group sizes we consider 6 cases, all with average group size equal to 10 (so that, corresponding to $R=50,100,200$, we have $n=500,1000,2000$), but with various degrees of unbalancedness. Specifically, the sequence of group sizes $m_{1},\ldots,m_{R}$ is periodic, with period 10 (i.e., $m_{i}=m_{i+10}$, for any $i=1,\ldots ,R-10$), and the first $10$ group sizes $m_{1},\ldots,m_{10}$ are given in Table (ref). For each $R$, the single regressor $\widetilde{x} =(\widetilde{x}_{1}^{\prime},\ldots,\widetilde{x}_{R}^{\prime})^{\prime}$ is drawn once from $\mathrm{N}(0,I_{n})$ and then kept fixed across replications, and $\boldsymbol{\alpha}=(\boldsymbol{\alpha}_{1},\ldots,\boldsymbol{\alpha }_{R})^{\prime}$ is drawn from $\mathrm{N}(0,I_{R})$ in each replication. The true values of $\lambda$, $\beta$, $\sigma$ are set to $\lambda_{0}=0$, $\tilde{\beta}_{0}=1$, and $\sigma_{0}=1$.

table[table omitted — 762 chars of source]

Estimation is performed after removal of the fixed effects by premultiplication by $C_{X_{\mathrm{FE}}}$ (see Section (ref) ).\footnote{Note that the 2SLSE of $(\lambda,\tilde{\beta} )$ after premultiplication by $C_{X_{\mathrm{FE}}}$ is the same as the 2SLSE prior to removal of the fixed effects.} When the model is balanced (case 6), both the 2SLSE and the (Q)MLE do not exist. The 2SLSE does not exist because $\mathrm{rank}(X,WX)=k=1$ and therefore the instrument matrix is singular. The (Q)MLE does not exist because Condition (ref) fails and therefore the likelihood based on $C_{X_{\mathrm{FE}}}\boldsymbol{y}$ is constant (see Example (ref)). Table (ref) shows that the root median square error of both 2SLS and (Q)ML estimators increases as the model becomes more balanced. Similarly to the first experiment, it is not surprising that the (Q)MLE performs better than the 2SLSE, given that $\operatorname{var} (\boldsymbol{\varepsilon})=I_{n}$ in the DGP. It is worth noting, however, that in the current experiment, the root median square error of the (Q)MLE, not only that of the 2SLSE, increases substantially close to the non-identifiability case. This is due to the two different types of non-identifiability studied in the two experiments.

Table (ref) shows that coverages of Wald confidence intervals can be very far from the nominal coverage when the model is close to being balanced. This is true for both the 2SLSE and (Q)MLE.

table[table omitted — 8,279 chars of source]

Conclusion

This paper has studied identification of an autoregression defined on a general network, under weak distributional assumptions and without requiring repeated observations of the network. In this context, identification is possible for generic parameter values and for generic regressor matrices, whatever the network. Nevertheless, important cases do exist when identification fails, either in the original sample space or after some transformation of the sample space (this could be, for instance, a transformation aimed at removing fixed effects). We have shown that, in particular, there are cases where it is impossible to conduct inference that respects the invariance properties of the model, despite the fact the parameters may be identifiable from the distribution on the original sample space.

For practical purposes, it may be useful to construct a measure of the distance from non-identifiability. This goes beyond the scope of the present paper, but, for example, one may want to have a measure of distance from the non-identifiability condition $\mathrm{rank}(X,WX)=k$ in Proposition (ref). One such measure would be the $\mathopen{}\mathclose\bgroup\originalleft( k+1\aftergroup\egroup\originalright) $-th largest singular value of $(X,WX)$, or some norm of the matrix $M_{X}WX$, possibly upon some normalization of $X$ and $W.$\footnote{To justify these two measures note that, since $\mathrm{rank}(X)=k,$ (i) $\mathrm{rank}(X,WX)=k$ if the $\mathopen{}\mathclose\bgroup\originalleft( k+1\aftergroup\egroup\originalright) $-th largest singular value of $(X,WX)$ is zero; (ii) $\mathrm{rank}(X,WX)=k$ is equivalent to $\operatorname{col}(WX)\subseteq \operatorname{col}(X)$, or, which is the same, to $M_{X}WX=0$.} Such measures should help model users to avoid not only the cases in which inference based on the first moment is impossible, but also cases close to these, in which inference is likely to be very challenging without additional distributional assumptions.

Finally, it is important to remark that the results in this paper have been derived under the assumption that the network is fully known and exogenous, which may be unrealistic in many applications. The study of identification when the network is (partially) unknown and/or endogenous remains a key challenge in the literature Blume2015,dePaula2020,Lewbel2019, and we hope that the results obtained in this paper can prove useful in that setting too.