EconBase
← Back to paper

Consistent specification testing under spatial dependence

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

117,383 characters · 17 sections · 114 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Consistent specification testing under spatial dependence

abstractWe propose a series-based nonparametric specification test for a regression function when data are spatially dependent, the `space' being of a general economic or social nature. Dependence can be parametric, parametric with increasing dimension, semiparametric or any combination thereof, thus covering a vast variety of settings. These include spatial error models of varying types and levels of complexity. Under a new smooth spatial dependence condition, our test statistic is asymptotically standard normal. To prove the latter property, we establish a central limit theorem for quadratic forms in linear processes in an increasing dimension setting. Finite sample performance is investigated in a simulation study, with a bootstrap method also justified and illustrated. Empirical examples illustrate the test with real-world data.

Keywords: Specification testing, nonparametric regression, spatial dependence, cross-sectional dependence

JEL Classification: C21, C55

\allowdisplaybreaks

Introduction

Models for spatial dependence have recently become the subject of vigorous research. This burgeoning interest has roots in the needs of practitioners who frequently have access to data sets featuring inter-connected cross-sectional units. Motivated by these practical concerns, we propose a specification test for a regression function in a general setup that covers a vast variety of commonly employed spatial dependence models and permits the complexity of dependence to increase with sample size. Our test is consistent, in the sense that a parametric specification is tested with asymptotically unit power against a nonparametric alternative. The `spatial' models that we study are not restricted in any way to be geographic in nature, indeed `space' can be a very general economic or social space. Our empirical examples feature conflict alliances and technology externalities as examples of `spatial dependence', for instance.

Specification testing is an important problem, and this is reflected in a huge literature studying consistent tests. Much of this is based on independent, and often also identically distributed, data. However data frequently exhibit dependence and consequently a branch of the literature has also examined specification tests under time series dependence. Our interest centers on dependence across a `space', which differs quite fundamentally from dependence in a time series context. Time series are naturally ordered and locations of the observations can be observed, or at least the process generating these locations may be modelled. It can be imagined that concepts from time series dependence be extended to settings where the data are observed on a geographic space and dependence can be treated as a decreasing function of distance between observations. Indeed much work has been done to extend notions of time series dependence in this type of setting, see e.g. Jenish2009,Jenish2012.

However, in a huge variety of economics and social science applications agents influence each other in ways that do not conform to such a setting. For example, farmers affect the demand of farmers in the same village but not in different villages, as in case1991spatial. Likewise, price competition among firms exhibits spatial features (pinkse2002), input-output relations lead to complementarities between sectors (Conley2003), co-author connections form among scientists (Oettl2012, Mohnen2017), R&D spillovers occur through technology and product market spaces (Bloom2013), networks form due to allegiances in conflicts (konig2017networks) and overlapping bank portfolios lead to correlated lending decisions (Gupta2018a). Such examples cannot be studied by simply extending results developed for time series and illustrate the growing need for suitable methods.

A very popular model for general spatial dependence is the spatial autoregressive (SAR) class, due to cliff1973spatial. The key feature of SAR models, and various generalizations such as SARMA (SAR moving average) and matrix exponential spatial specifications (MESS, due to LeSage2007), is the presence of one or more spatial weight matrices whose elements characterize the links between agents. As noted above, these links may form for a variety of reasons, so the `spatial' terminology represents a very general notion of space, such as social or economic space. Key papers on the estimation of SAR models and their variants include kelejian1998generalized and lee2004asymptotic, but research on various aspects of these is active and ongoing, see e.g. Robinson2015, Hillier2018, Hillier2018a, Kuersteiner2020, Han2021, Hahn2021.

Unlike work focusing on independent or time series data, a general drawback of spatially oriented research has been the lack of general unified theory. Typically, individual papers have studied specific special cases of various spatial specifications. A strand of the literature has introduced the notion of a cross-sectional linear-process to help address this problem, and we follow this approach. This representation can accommodate SAR models in the error term (so called spatial error models (SEM)) as a special case, as well as variants like SARMA and MESS, whence its generality is apparent. The linear-process structure shares some similarities with that familiar from the time series literature (see e.g. hannan1970multiple). Indeed, time series versions may be regarded as very special cases but, as stressed before, the features of spatial dependence must be taken into account in the general formulation. Such a representation was introduced by Robinson2011 and further examined in other situations by Robinson2012b (partially linear regression), Delgado2015 (non-nested correlation testing), Lee2016 (series estimation of nonparametric regression) and Hidalgo2017 (cross-sectionally dependent panels).

In this paper, we propose a test statistic similar to that of Hong1995, based on estimating the nonparametric specification via series approximations. Assuming an independent and identically distributed sample, their statistic is based on the sample covariance between the residual from the parametric model and the discrepancy between the parametric and nonparametric fitted values. Allowing additionally for spatial dependence through the form of a linear process as discussed above, our statistic is shown to be asymptotically standard normal, consistent and possessing nontrivial power against local alternatives of a certain type. To prove asymptotic normality, we present a new central limit theorem (CLT) for quadratic forms in linear processes in an increasing dimension setting that may be of independent interest. A CLT for quadratic forms under time series dependence in the context of series estimation can be found in Gao2000, and our result can be viewed as complementary to this. The setting of Su2017 is a very special case of our framework. There has been recent interest in specification testing for spatial models, see for example Sun2020 for a kernel-based model specification test and Lee2020 for a consistent omnibus test. We contribute to this literature by studying a linear process based increasing parameter dimension framework.

Our linear process framework permits spatial dependence to be parametric, parametric with increasing dimension, semiparametric or any combination thereof, thus covering a vast variety of settings. A class of models of great empirical interest are `higher-order' SAR models in the outcome variables, but with spatial dependence structure also in the errors. We initially present the familiar nonparametric regression to clarify the exposition, and then cover this class as the main model of interest. Our theory covers as special cases SAR, SMA, SARMA, MESS models for the error term. These specifications may be of any fixed spatial order, but our theory also covers the case where they are of increasing order.

Thus we permit a more complex model of spatial dependence as more data become available, which encourages a more flexible approach to modelling such dependence as stressed by Gupta2013,Gupta2018 in a higher-order SAR context, huber1973robust, portnoy1984asymptotic,portnoy1985asymptotic and Anatolyev2012 in a regression context and Koenker1999 for the generalized method of moments setting, amongst others. This literature focuses on a sequence of true models, rather than a sequence of models approximating an infinite true model. Our paper also takes the same approach. On the other hand, in the spatial setting, Gupta2018d considers increasing lag models as approximations to an infinite lag model with lattice data and also suggests criteria for choice of lag length.

Our framework is also extended to the situation where spatial dependence occurs through nonparametric functions of raw distances (these may be exogenous economic or social distances, say), as in pinkse2002. This allows for greater flexibility in modelling spatial weights as the practitioner only has to choose an exogenous economic distance measure and allow the data to determine the functional form. It also adds a degree of robustness to the theory by avoiding potential parametric misspecification. The case of geographical data is also covered, for example the important classes of Mat\'ern and Wendland (see e.g. Gneiting2002) covariance functions. Finally, we introduce a new notion of smooth spatial dependence that provides more primitive, and checkable, conditions for certain properties than extant ones in the literature.

To illustrate the performance of the test in finite samples, we present Monte Carlo simulations that exhibit satisfactory small sample properties. The test is demonstrated in three empirical examples, including two based on recently published work on social networks: Bloom2013 (R&D spillovers in innovation), konig2017networks (conflict alliances during the Congolese civil war). Another example studies cross-country spillovers in economic growth. Our test may or may not reject the null hypothesis of a linear regression in these examples, illustrating its ability to distinguish well between the null and alternative models.

The next section introduces our basic setup using a nonparametric regression with no SAR structure in responses. We treat this abstraction as a base case, and Section (ref) discusses estimation and defines the test statistic, while Section (ref) introduces assumptions and the key asymptotic results of the paper. Section (ref) examines the most commonly employed higher-order SAR models, while Section (ref) deals with nonparametric spatial error structures. Nonparametric specification tests are often criticized for poor finite sample performance when using the asymptotic critical values. In Section (ref) we present a bootstrap version of our testing procedure. Sections (ref) and (ref) contain a study of finite sample performance and the empirical examples respectively, while Section (ref) concludes. Proofs are contained in appendices, including a supplementary online appendix which also contains additional simulation results.

For the convenience of the reader, we collect some frequently used notation here. First, we introduce three notational conventions for any parameter $\nu$ for the rest of the paper: $\nu\in\mathbb{R}^{d_\nu}$, $\nu_0$ denotes the true value of $\nu$ and for any scalar, vector or matrix valued function $f(\nu)$, we denote $f\equiv f(\nu_0)$. Let $\overline{\varphi}(\cdot)$ (respectively $\underline{\varphi}(\cdot)$) denote the largest (respectively smallest) eigenvalue of a generic square nonnegative definite matrix argument. For a generic matrix $A$, denote $\left\Vert A\right\Vert=\left[\overline{\varphi}(A'A)\right]^{1/2}$, i.e. the spectral norm of $A$ which reduces to the Euclidean norm if $A$ is a vector. $\left\Vert A\right\Vert_R$ denotes the maximum absolute row sum norm of a generic matrix $A$ while $\left \Vert A \right \Vert _{F}=\left[tr(AA')\right]^{1/2}$, the Frobenius norm. Throughout the paper $|\cdot|$ is absolute value when applied to a scalar and determinant when applied to a matrix. Denote by $c$ ($C$) generic positive constants, independent of any quantities that tend to infinity, and arbitrarily small (big).

Setup

To illustrate our approach, we first consider the nonparametric regression

equation[equation omitted — 92 chars of source]

where $\theta _{0}(\cdot )$ is an unknown function and $x_{i}$ is a vector of strictly exogenous explanatory variables with support $\mathcal{X}\subset \mathbb{R}^k$. Spatial dependence is explicitly modeled via the error term $u_{i}$, which we assume is generated by:

equation[equation omitted — 90 chars of source]

where $\varepsilon _{s}$ are independent random variables, with zero mean and identical variance $\sigma_0^2$. Further conditions on the $\varepsilon_s$ will be assumed later. The linear process coefficients $b_{is}$ can depend on $n$, as may the covariates $x_i$. This is generally the case with spatial models and implies that asymptotic theory ought to be developed for triangular arrays. There are a number of reasons to permit dependence on sample size. The $b_{is}$ can depend on spatial weight matrices, which are usually normalized for both stability and identification purposes.

Such normalizations, e.g. row-standardization or division by spectral norm, may be $n$-dependent. Furthermore, $x_i$ often includes underlying covariates of `neighbors' defined by spatial weight matrices. For instance, for some $n\times 1$ covariate vector $z$ and exogenous spatial weight matrix $W\equiv W_n$, a component of $x_i$ can be $e_i'Wz$, where $e_i$ has unity in the $i$-th position and zeros elsewhere, which depends on $n$. Thus, subsequently, any spatial weight matrices will also be allowed to depend on $n$. Finally, treating triangular arrays permits re-labelling of quantities that is often required when dealing with spatial data, due to the lack of natural ordering, see e.g. Robinson2011. We suppress explicit reference to this $n$-dependence of various quantities for brevity, although mention will be made of this at times to remind the reader of this feature.

Now, assume the existence of a $d_\gamma\times 1$ vector $\gamma _{0}$ such that $ b_{is}=b_{is}(\gamma _{0})$, possibly with $d_\gamma\rightarrow \infty $ as $n\rightarrow \infty $, for all $i=1,\ldots ,n$ and $s\geq 1$. Let $u$ be the $n\times 1$ vector with typical element $u_{i}$, $\varepsilon $ be the infinite dimensional vector with typical element $\varepsilon _{s},$ and $B$ be an infinite dimensional matrix Cooke1950 with typical element $b_{is}.$ In matrix form,

equation[equation omitted — 205 chars of source]

We assume that $\gamma _{0}\in \Gamma $, where $\Gamma $ is a compact subset of $\mathbb{R}^{d_\gamma}$. With $d_\gamma$ diverging, ensuring $\Gamma $ has bounded volume requires some care, see Gupta2018. For a known function $f(\cdot)$, our aim is to test

equation[equation omitted — 170 chars of source]

against the global alternative $H_{1}:P\left[\theta _{0}\left( x_{i}\right)\neq f(x_{i},\alpha ) \right]>0,\text{ for all }\alpha \in \mathcal{A}$.

We now nest commonly used models for spatial dependence in ((ref)). Introduce a set of $n\times n$ spatial weight (equivalently network adjacency) matrices $W_j$, $j=1,\ldots,m_1+m_2$. Each $W_j$ can be thought of as representing dependence through a particular space. Now, consider models of the form $\Sigma(\gamma)=A^{-1}(\gamma)A'^{-1}(\gamma)$. For example, with $\xi$ denoting a vector of iid disturbances with variance $\sigma_0^2$, the model with SARMA$(m_1,m_2)$ errors is $ u=\sum_{j=1}^{m_1}\gamma_jW_ju+\sum_{j=m_1+1}^{m_1+m_2}\gamma_jW_j\xi+\xi$, with $A(\gamma)=\left(I_n+\sum_{j=m_1+1}^{m_1+m_2}\gamma_jW_j\right)^{-1}\left(I_n-\sum_{j=1}^{m_1}\gamma_jW_j\right)$, assuming conditions that guarantee the existence of the inverse. Such conditions can be found in the literature, see e.g. Lee2010 and Gupta2018. The SEM model is obtained by setting $m_2=0$ while the model with SMA errors has $m_1=0$. The model with MESS$(m)$ errors (LeSage2007, Debarsy2015) is $ u=\exp\left(\sum_{j=1}^{m} \gamma_jW_j\right)\xi,A(\gamma)=\exp\left(-\sum_{j=1}^{m} \gamma_jW_j\right). $

In some cases the space under consideration is geographic i.e. the data may be observed at irregular points in Euclidean space. Making the identification $u_i\equiv U\left(t_i\right)$, $t_i\in\mathbb{R}^d$ for some $d>1$, and assuming covariance stationarity, $U(t)$ is said to follow an isotropic model if, for some function $\delta$ on $\mathbb{R}$, the covariance at lag $s$ is $r(s)=\mathcal{E}\left[U(t)U(t+s)\right]=\delta(\Vert s\Vert)$. An important class of parametric isotropic models is that of Matern1986, which can be parameterized in several ways, see e.g. Stein1999. Denoting by $\Gamma_f$ the Gamma function and by $\mathcal{K}_{\gamma_1}$ the modified Bessel function of the second kind (Gradshteyn1994), take $ \delta(\left\Vert s\right\Vert,\gamma)=\left(2^{\gamma_1-1}\Gamma_f(\gamma_1)\right)^{-1}\left(\gamma_2^{-1}\sqrt{2\gamma_1}\left\Vert s\right\Vert\right)^{\gamma_1}\mathcal{K}_{\gamma_1}\left(\gamma_2^{-1}\sqrt{2\gamma_1}\left\Vert s\right\Vert\right), $ with $\gamma_1,\gamma_2>0$ and $d_\gamma=2$. With $d_\gamma=3$, another model takes $\delta(\left\Vert s\right\Vert,\gamma)=\gamma_1\exp\left(-\left\Vert s/\gamma_2\right\Vert^{\gamma_3}\right)$, see e.g. DeOliveira1997, Stein1999. Fuentes2007 considers this model with $\gamma_3=1$, as well as a specific parameterization of the Mat\`{e}rn covariance function.

Test statistic

We estimate $\theta _{0}(\cdot )$ via a series approximation. Certain technical conditions are needed to allow for $\mathcal{X}$ to have unbounded support. To this end, for a function $g(x)$ on $\mathcal{X}$, define a weighted sup-norm (see e.g. Chen2005, Chen2007, Lee2016) by $\left\Vert g\right\Vert_{w}=\sup_{x\in\mathcal{X}}\left\vert g(x)\right\vert\left(1+\left\Vert x\right\Vert^2\right)^{-w/2}, \text{ for some } w>0$. Assume that there exists a sequence of functions $\psi _{i}:=\psi \left( x_{i}\right) : \mathbb{R}^{k}\mapsto \mathbb{R}^{p}$, where $p\rightarrow \infty $ as $ n\rightarrow \infty $, and a $p\times 1$ vector of coefficients $\beta_0 $ such that

equation[equation omitted — 129 chars of source]

where $e(\cdot)$ satisfies: \setcounter{assumption}{0}

assumptionThere exists a constant $\mu>0$ such that $\left \Vert e \right \Vert_{w_x} =O\left( p^{-\mu}\right),$ as $p\rightarrow\infty$, where $w_x\geq 0$ is the largest value such that $\sup_{i=1,\ldots,n} \mathcal{E}\left\Vert x_i\right\Vert^{w_x}<\infty$, for all $n$.

By Lemma 1 in Appendix B of Lee2016, this assumption implies that

equation[equation omitted — 136 chars of source]

Due to the large number of assumptions in the paper, sometimes with changes reflecting only the various setups we consider, we prefix assumptions with R in this section and the next, to signify `regression'. In Section (ref) the prefix is SAR, for `spatial autoregression', while in Section (ref) we use NPN, for `nonparametric network'.

Let $ y=(y_{1},\ldots,y_{n})^{\prime },{\theta _{0}}=(\theta _{0}\left( x_{1}\right) ,\ldots,\theta _{0}\left( x_{n}\right) )^{\prime }, \Psi =(\psi _{1},\ldots,\psi _{n})^{\prime }$. We will estimate $\gamma_0$ using a quasi maximum likelihood estimator (QMLE) based on a Gaussian likelihood, although Gaussianity is nowhere assumed. For any admissible values $\beta $, $\sigma ^{2}$ and $\gamma $, the (multiplied by $2/n$) negative quasi log likelihood function based on using the approximation ((ref)) is

equation[equation omitted — 279 chars of source]

which is minimised with respect to $\beta $ and $\sigma ^{2}$ by

eqnarray[eqnarray omitted — 335 chars of source]

where $M(\gamma )=I_{n}-E(\gamma )\Psi \left( \Psi ^{\prime }\Sigma (\gamma )^{-1}\Psi \right) ^{-1}\Psi ^{\prime }E(\gamma )^{\prime }$ and $E(\gamma )$ is the $n\times n$ symmetric matrix such that $E(\gamma )E(\gamma )^{\prime }=\Sigma (\gamma )^{-1}$. The use of the approximate likelihood relies on the negligibility of $e(\cdot)$, which in turn permits the replacement of $\theta_0(\cdot)$ by $\psi'\beta_0$ with asymptotically negligible cost. Thus the concentrated likelihood function is

equation[equation omitted — 180 chars of source]

We define the QMLE of $\gamma _{0}$ as $\widehat{\gamma}=\text{arg min}_{\gamma \in \Gamma }\mathcal{L}(\gamma )$ and the QMLEs of $\beta_0 $ and $\sigma_0 ^{2}$ as $\widehat{\beta}=\bar{\beta}\left( \widehat{\gamma}\right) $ and $\widehat{\sigma} ^{2}=\bar{\sigma}^{2}\left( \widehat{\gamma}\right) $. At a given $x_1,\ldots,x_n$, the series estimate of $\theta _{0} $ is defined as

equation[equation omitted — 212 chars of source]

Let $\widehat{\alpha }_{n}\equiv \widehat{\alpha }$ denote an estimator consistent for $\alpha _{0}$ under $H_{0}$, for example the (nonlinear) least squares estimator. Note that $\widehat\alpha$ is consistent only under $H_0$, so we introduce a general probability limit of $\widehat{\alpha } $, as in Hong1995.

assumptionThere exists a deterministic sequence $\alpha_n^*\equiv\alpha^*$ such that $\widehat\alpha-\alpha^*=O_p\left(1/\sqrt{n}\right)$.

Examples of estimators that satisfy this assumption include (nonlinear) least squares, generalized method of moments estimators or adaptive efficient weighted least squares Stinchcombe1998.

Following Hong1995, define the regression error $u_{i}\equiv y_{i}-f(x_{i},\alpha ^{\ast })$ and the specification error $v_{i}\equiv \theta _{0}(x_{i})-f(x_{i},\alpha ^{\ast })$. Our test statistic is based on a scaled and centered version of $\widehat{m}_{n}=\widehat{\sigma }^{-2}\widehat{{v}}^{\prime }\Sigma \left( \widehat{\gamma }\right) ^{-1}\widehat{ {u}}/n=\widehat{\sigma } ^{-2}\left(\widehat{ {\theta }}- {f}\left(x,\widehat{\alpha }\right)\right)^{\prime }\Sigma \left( \widehat{\gamma }\right) ^{-1}\left(y- {f}\left(x, \widehat{\alpha }\right)\right)/n$, where $f(x,\alpha)=\left(f\left(x_1,\alpha\right),\ldots,f\left(x_n,\alpha\right)\right)'$. Precisely, it is defined as

equation[equation omitted — 87 chars of source]

The motivation for such a centering and scaling stems from the fact that, for fixed $p$, $n\widehat{m}_{n}$ has an asymptotic $\chi^2_p$ distribution. Such a distribution has mean $p$ and variance $2p$, and it is a well-known fact that $\left(\chi^2_p-p\right)/{\sqrt{2p}}\overset{d}\longrightarrow N(0,1),\text{ as }p\rightarrow \infty$. This motivates our use of ((ref)) and explains why we aspire to establish a standard normal distribution under the null hypothesis. Intuitively, the test statistic is based on the sample covariance between the residual from the parametric model and the discrepancy between the parametric and nonparametric fitted values, as in Hong1995.

Hong1995 also note that, due to the nonparametric nature of the problem, such a statistic vanishes faster than the parametric ($n^{\frac{1}{2}}$) rate, thus a $n^{\frac{1}{2}}$-normalization leads to degeneracy of the test. A proper normalization as in ((ref)) will yield a non-degenerate limiting distribution. As Hong1995 noted, our test is one-sided. This is because asymptotically negative values of our test statistic can occur only under the null, while under the alternative it tends to a positive, increasing number. Thus, we reject the null if our test statistic is on the right tail.

Asymptotic theory

Consistency of $\widehat{\protect \gamma }$

We first provide conditions under which our estimator $\widehat\gamma$ of $\gamma_0$ is consistent. Such a property is necessary for the results that follow. The following assumption is a rather standard type of asymptotic boundedness and full-rank condition on $\Sigma(\gamma)$.

assumption\[\varlimsup_{n\rightarrow\infty}\sup_{\gamma\in\Gamma}\bar\varphi\left(\Sigma(\gamma)\right)<\infty \text { and } \varliminf_{n\rightarrow\infty}\inf_{\gamma\in\Gamma}\underline\varphi\left(\Sigma(\gamma)\right)>0.\]
assumptionThe $u_i, i=1,\ldots,n,$ satisfy the representation ((ref)). The $\varepsilon _{s}$, $s\geq 1$, have zero mean, finite third and fourth moments $\mu _{3}$ and $\mu _{4}$ respectively and, denoting by $\sigma _{ij}(\gamma )$ the $(i,j)$-th element of $\Sigma (\gamma )$ and defining $ b_{is}^{\ast }={b_{is}}/{\sigma _{ii}^{\frac{1}{2}}},\;i=1,\ldots ,n,\;n\geq 1,s\geq 1, $ we have \begin{equation} \underset{n\rightarrow \infty }{\overline{\lim }}\sup_{i=1,\ldots ,n}\sum_{s=1}^{\infty }\left \vert b_{is}^{\ast }\right \vert +\sup_{s\geq 1} \underset{n\rightarrow \infty }{\overline{\lim }}\sum_{i=1}^{n}\left \vert b_{is}^{\ast }\right \vert <\infty . \end{equation}

By Assumption (ref), $\sigma_{ii}$ is bounded and bounded away from zero, so the normalization of the $b_{is}$ in Assumption (ref) is well defined. The summability conditions in ((ref)) are typical conditions on linear process coefficients that are needed to control dependence; for instance in the case of stationary time series $b^*_{is}=b^*_{i-s}$. The infinite linear process assumed in ((ref)) is further discussed by Robinson2011, who introduced it, and also by Delgado2015. These assumptions imply an increasing-domain asymptotic setup and preclude infill asymptotics.

Because we often need to consider the difference between values of the matrix-valued function $\Sigma(\cdot)$ at distinct points, it is useful to introduce an appropriate concept of `smoothness'. This concept has been employed before in economics, see e.g. Chen2007, and is defined below.

definitionLet $\left(X,\left\Vert\cdot\right\Vert_X\right)$ and $\left(Y,\left\Vert\cdot\right\Vert_Y\right)$ be Banach spaces, $\mathscr{L}(X,Y)$ be the Banach space of linear continuous maps from $X$ to $Y$ with norm $\left\Vert T\right\Vert_{\mathscr{L}(X,Y)}=\sup_{\left\Vert x\right\Vert_X\leq 1}\left\Vert T(x)\right\Vert_Y$ and $U$ be an open subset of $X$. A map $F:U\rightarrow Y$ is said to be Fr\'echet-differentiable at $u\in U$ if there exists $L\in\mathscr{L}(X,Y)$ such that \begin{equation} \lim_{\left\Vert h\right\Vert_X\rightarrow 0}\frac{F(u+h)-F(u)-L(h)}{\left\Vert h\right\Vert_X}=0. \end{equation} $L$ is called the Fr\'echet-derivative of $F$ at $u$. The map $F$ is said to be Fr\'echet-differentiable on $U$ if it is Fr\'echet-differentiable for all $u\in U$.

The above definition extends the notion of a derivative that is familiar from real analysis to the functional spaces and allows us to check high-level assumptions that past literature has imposed. To the best of our knowledge, this is the first use of such a concept in the literature on spatial/network models. Denote by $\mathcal{M}^{n\times n}$ the set of real, symmetric and positive semi-definite $n\times n$ matrices. Let $\Gamma^o$ be an open subset of $\Gamma$ and consider the Banach spaces $\left(\Gamma,\left\Vert\cdot\right\Vert_g\right)$ and $\left(\mathcal{M}^{n\times n},\left\Vert\cdot\right\Vert\right)$, where $\left\Vert \cdot\right\Vert_g$ is a generic $\ell_p$ norm, $p\geq 1$. The following assumption ensures that $\Sigma(\cdot)$ is a `smooth' function, in the sense of Fr\'echet-smoothness.

assumptionThe map $\Sigma:\Gamma^o\rightarrow \mathcal{M}^{n\times n}$ is Fr\'echet-differentiable on $\Gamma^o$ with Fr\'echet-derivative denoted $D\Sigma\in\mathscr{L}\left(\Gamma^o,\mathcal{M}^{n\times n}\right)$. Furthermore, the map $D\Sigma$ satisfies \begin{equation} \sup_{\gamma\in\Gamma^o}\left\Vert D\Sigma(\gamma)\right\Vert_{\mathscr{L}\left(\Gamma^o,\mathcal{M}^{n\times n}\right)}\leq C. \end{equation}

Assumption (ref) is a functional smoothness condition on spatial dependence. It has the advantage of being checkable for a variety of commonly employed models. For example, a first-order SEM has $\Sigma(\gamma)=A^{-1}(\gamma)A'^{-1}(\gamma)$ with $A=I_n-\gamma W$. Corollary (ref) in the supplementary appendix shows $\left(D\Sigma(\gamma)\right)\left(\gamma^\dag\right)=\gamma^\dag A^{-1}(\gamma)\left(G'(\gamma)+G(\gamma)\right)A'^{-1}(\gamma)$, at a given point $\gamma\in\Gamma^o$, where $G(\gamma)=WA^{-1}(\gamma)$. Then, taking

equation[equation omitted — 120 chars of source]

yields Assumption (ref). Condition ((ref)) limits the extent of spatial dependence and is very standard in the spatial literature; see e.g. lee2004asymptotic and numerous subsequent papers employing similar conditions.

Fr\'echet derivatives for higher-order SAR, SMA, SARMA and MESS error structures are computed in supplementary appendix (ref), in Lemmas (ref)-(ref) and Corollaries (ref)-(ref). Strictly speaking, Gateaux differentiability might suffice for the type of results that we target. We opt for Fr\'echet differentiability because this derivative map is linear and continuous or, equivalently, a bounded linear operator, a property that makes Assumption (ref) more reasonable.

The following proposition is very useful in `linearizing' perturbations in the $\Sigma(\cdot)$.

propositionIf Assumption (ref) holds, then for any $\gamma_1,\gamma_2\in\Gamma^o$, \begin{equation} \left\Vert\Sigma\left(\gamma_1\right)-\Sigma\left(\gamma_2\right)\right\Vert\leq C \left\Vert\gamma_1-\gamma_2\right\Vert. \end{equation}

To illustrate how the concept of Fr\'echet-differentiability allows us to check high-level assumptions extant in the literature, a consequence of Proposition (ref) is the following corollary, a version of which appears as an assumption in Delgado2015.

corollaryFor any $\gamma ^{* }\in \Gamma^o $ and any $\eta >0 $, \begin{equation} \underset{n\rightarrow \infty }{\overline{\lim }}\sup_{\gamma \in \left \{ \gamma :\left \Vert \gamma -\gamma ^{*}\right \Vert <\eta \right \} \cap \Gamma^o }\left \Vert \Sigma (\gamma )-\Sigma \left( \gamma ^{* }\right) \right \Vert <C\eta . \end{equation}

We now introduce regularity conditions needed to establish the consistency of $\hat\gamma$. Define \[ \sigma ^{2}\left( \gamma \right) =n^{-1}\sigma ^{2}tr\left( \Sigma (\gamma )^{-1}\Sigma \right) =n^{-1}\sigma ^{2}\left \Vert E(\gamma )E^{-1}\right \Vert _{F}^{2}, \] which is nonnegative by definition and bounded by Assumption (ref), red with the matrix $E(\gamma)$ defined after ((ref)).

assumption$c\leq \sigma ^{2}\left( \gamma \right)\leq C$ for all $\gamma \in \Gamma$.
assumption$\gamma _{0}\in \Gamma $ and, for any $\eta >0$, \begin{equation} \varliminf_{n\rightarrow \infty }\inf_{\gamma \in \overline{\mathcal{N}} ^{\gamma }(\eta )}\frac{n^{-1}tr\left( \Sigma (\gamma )^{-1}\Sigma \right) }{ \left \vert \Sigma (\gamma )^{-1}\Sigma \right \vert ^{1/n}}>1, \end{equation} where $\overline{\mathcal{N}}^{\gamma }(\eta )=\Gamma \setminus \mathcal{N} ^{\gamma }(\eta )$ and $\mathcal{N}^{\gamma }(\eta )=\left \{ \gamma :\left \Vert \gamma -\gamma _{0}\right \Vert <\eta \right \} \cap \Gamma $.
assumption$\left \{ \underline{\varphi }\left( n^{-1}\Psi ^{\prime}\Psi \right) \right \} ^{-1}+\overline{\varphi }\left( n^{-1}\Psi ^{\prime}\Psi \right)=O_{p}(1)$ .

Assumption (ref) is a boundedness condition originally considered in Gupta2018, while Assumptions (ref) and (ref) are identification conditions. Indeed, Assumption (ref) requires that $\Sigma(\gamma)$ be identifiable in a small neighborhood around $\gamma_0$. This is apparent on noticing that the ratio in ((ref)) is at least one by the inequality between arithmetic and geometric means, and equals one when $\Sigma(\gamma)=\Sigma$. Similar assumptions arise frequently in related literature, see e.g. lee2004asymptotic, Delgado2015. Assumption (ref) is a typical asymptotic boundedness and non-multicollinearity condition, see e.g. Newey1997 and much other literature on series estimation. Primitive conditions for this assumption to hold require the convergence (in matrix norm) of $n^{-1}\Psi'\Psi$ to its expectation, and this entails restrictions on the extent of spatial dependence in the $x_i$. A reference is Lee2016, wherein consider Assumption A.4 and the proof of Theorem 1. By Assumption (ref), (ref) implies $\sup_{\gamma\in\Gamma}\left\{\underline{\varphi}\left(n^{-1}\Psi'\Sigma(\gamma)^{-1}\Psi\right)\right\}^{-1}=O_p(1)$.

theorem\sloppy Under either $H_0$ or $H_1$, Assumptions (ref)-(ref) and $p^{-1}+\left(d_\gamma+p\right)/n\rightarrow 0$ as $n\rightarrow\infty$, $\left \Vert \left(\widehat{\gamma},\hat{\sigma}^2\right)-\left(\gamma_0,\sigma_0^2\right) \right \Vert \overset{p}{ \longrightarrow }0.$

Asymptotic properties of the test statistic

Write $\Sigma_j(\gamma)=\partial\Sigma(\gamma)/\partial\gamma_j$, $j=1,\ldots,d_\gamma$, the matrix differentiated element-wise. While Assumption (ref) guarantees that these partial derivatives exist, the next assumption imposes a uniform bound on their spectral norms.

assumption$\varlimsup_{n\rightarrow\infty}\sup_{j=1,\ldots,d_\gamma}\left\Vert \Sigma_j(\gamma)\right\Vert<C$.

We will later consider the sequence of local alternatives

equation[equation omitted — 152 chars of source]

where $h$ is square integrable on the support $\mathcal{X}$ of the $x_i$. Under the null $H_0$, we have $h(x_i)=0$, a.s..

assumption\sloppy For each $n\in\mathbb{N}$ and $i=1,\ldots,n$, the function $f:\mathcal{X}\times \mathcal{A}\rightarrow\mathbb{R}$ is such that $f\left(x_i,\alpha\right)$ is measurable for each $\alpha\in \mathcal{A}$, $f\left(x_i,\cdot\right)$ is a.s. continuous on $\mathcal{A}$, with $\sup_{\alpha\in \mathcal{A}}f^2\left(x_i,\alpha\right)\leq D_n\left(x_i\right)$, where $\sup_{n\in\mathbb{N}}D_n\left(x_i\right)$ is integrable and $\sup_{\alpha\in \mathcal{A}}\left\Vert\partial f\left(x_i,\alpha\right)/\partial\alpha\right\Vert^2\leq D_n\left(x_i\right)$, $\sup_{\alpha\in \mathcal{A}}\left\Vert\partial^2 f\left(x_i,\alpha\right)/\partial\alpha\partial\alpha'\right\Vert\leq D_n\left(x_i\right)$, all holding a.s..

Define the infinite-dimensional matrix $\mathscr{V}=B^{\prime }\Sigma ^{-1}\Psi \left( \Psi ^{\prime }\Sigma ^{-1}\Psi \right) ^{-1}\Psi ^{\prime }\Sigma ^{-1}B$, which is symmetric, idempotent and has rank $p$. We now show that our test statistic is approximated by a quadratic form in $\varepsilon$, weighted by $\mathscr{V}$.

theoremUnder Assumptions (ref)-(ref), $p^{-1}+p\left(p+d_\gamma^2\right)/n+\sqrt{n}/p^{\mu+1/4}\rightarrow 0$, as $n\rightarrow\infty$, and $H_0$, $ \mathscr{T}_n-{\left(\sigma_0^{-2}\varepsilon'\mathscr{V}\varepsilon-p\right)}/{\sqrt{2p}}=o_p(1). $
assumption$\underset{n\rightarrow \infty }{\overline{\lim }}\left \Vert \Sigma ^{-1}\right \Vert _{R}<\infty.$

Because $\left \Vert \Sigma ^{-1}\right \Vert\leq \left \Vert \Sigma ^{-1}\right \Vert _{R}$, this restriction on spatial dependence is somewhat stronger than a restriction on spectral norm but is typically imposed for central limit theorems in this type of setting, cf. lee2004asymptotic, Delgado2015, Gupta2018. The next assumption is needed in our proofs to check a Lyapunov condition. A typical approach would be assume moments of order $4+\epsilon$, for some $\epsilon>0$. Due to the linear process structure under consideration, taking $\epsilon=4$ makes the proof tractable, see for example Delgado2015.

assumptionThe $\varepsilon_s$, $s\geq 1$, have finite eighth moment.

The next assumption is strong if the basis functions $\psi_{ij}(\cdot)$ are polynomials, requiring all moments to exist in that case.

assumption$\mathcal{E}\left\vert\psi_{ij}\left(x\right)\right\vert<C$, $i=1,\ldots,n$ and $j=1,\ldots,p$.

The next theorem establishes the asymptotic normality of the approximating quadratic form introduced above.

theoremUnder Assumptions (ref), (ref), (ref), (ref)-(ref) and $p^{-1}+p^3/n\rightarrow 0$, as $n\rightarrow\infty$, $ {\left(\sigma _{0}^{-2}\varepsilon ^{\prime }\mathscr{V}\varepsilon -p\right)}/{\sqrt{2p}}\overset{d}{\longrightarrow}N(0,1). $

This is a new type of CLT, integrating both a linear process framework as well as an increasing dimension element. A linear-quadratic form in iid disturbances is treated by Kelejian2001, while a quadratic form in a linear process framework is treated by Delgado2015. However both results are established in a parametric framework, entailing no increasing dimension aspect of the type we face with $p\rightarrow\infty$.

Next, we summarize the properties of our test statistic in a theorem that records its asymptotic normality under the null, consistency and ability to detect local alternatives at $p^{1/4}/n^{1/2}$ rate. This rate has been found also by DeJong1994 and Gupta2018c. Introduce the quantity $\varkappa=\left({\sqrt{2}\sigma_0^2}\right)^{-1}\operatorname{plim}_{n\rightarrow\infty} {n^{-1}h'\Sigma^{-1}h}$, where $h=\left(h\left(x_1\right),\ldots,h\left(x_n\right)\right)'$ and $h\left(x_i\right)$ is from ((ref)).

theoremUnder the conditions of Theorems (ref) and (ref), (1) $\mathscr{T}_n\overset{d}{\rightarrow} N(0,1)$ under $H_0$, (2) $\mathscr{T}_n$ is a consistent test statistic, (3) $\mathscr{T}_n\overset{d}{\rightarrow}N \left(\varkappa,1\right)$ under local alternatives $H_{\ell}$.

Models with SAR structure in responses

We now introduce the SAR model

equation[equation omitted — 143 chars of source]

where $W_j$, $j=1,\ldots,d_{\lambda}$, are known spatial weight matrices with $i$-th rows denoted $w_{i,j}'$, as discussed earlier, and $\lambda_{0j}$ are unknown parameters measuring the strength of spatial dependence. We take $d_\lambda$ to be fixed for convenience of exposition. The error structure remains the same as in ((ref)). Here spatial dependence arises not only in errors but also responses. For example, this corresponds to a situation where agents in a network influence each other both in their observed and unobserved actions. Note that the error term $u_i$ can be generated by the same $W_j$, or different ones.

While the model in ((ref)) is new in the literature, some related ones are discussed here. Models such as ((ref)) but without dependence in the error structure are considered by Su2010 and Gupta2013,Gupta2018, but the former consider only $d_{\lambda}=1$ and the latter only parametric $\theta_0(\cdot)$. Linear $\theta_0(\cdot)$ and $d_{\lambda}>1$ are permitted by Lee2010, but the dependence structure in errors differs from what we allow in ((ref)). Using the same setup as Su2010 and independent disturbances, a specification test for the linearity of $\theta_0(\cdot)$ is proposed by Su2017. In comparison, our model is much more general and our test can handle more general parametric null hypotheses. We thank a referee for pointing out that ((ref)) is a particular case of Sun2016 when $u_i$ are iid and of Malikov2017 when $d_\lambda=1$.

Denoting $S(\lambda)=I_n-\sum_{j=1}^{d_{\lambda}} \lambda_j W_j$, the quasi likelihood function based on Gaussianity and conditional on covariates is

eqnarray[eqnarray omitted — 364 chars of source]

at any admissible point $\left(\beta',\phi',\sigma^2\right)'$ with $\phi=\left(\lambda',\gamma'\right)'$, for nonsingular $S(\lambda)$ and $\Sigma(\gamma)$. For given $\phi=\left(\lambda',\gamma'\right)'$, ((ref)) is minimised with respect to $\beta$ and $\sigma^2$ by

eqnarray[eqnarray omitted — 325 chars of source]

The QMLE of $\phi_0$ is $\widehat\phi=\operatorname*{arg\,min}_{\phi\in\Phi}\mathcal{L}\left(\phi\right)$, where

equation[equation omitted — 204 chars of source]

and $\Phi=\Lambda\times\Gamma$ is taken to be a compact subset of $\mathbb{R}^{d_{\lambda}+d\gamma}$. The QMLEs of $\beta_{0}$ and $\sigma_0^2$ are defined as $\bar{\beta}\left(\widehat{\phi}\right)\equiv\widehat{\beta}$ and $\bar{\sigma}^{2}\left(\widehat{\phi}\right)\equiv\widehat{\sigma}^2$ respectively. The following assumption controls spatial dependence and is discussed below equation ((ref)). \setcounter{assumption}{0}

assumption$\max_{j=1,\ldots,d_{\lambda}}\left\Vert W_j\right\Vert+\left\Vert S^{-1}\right\Vert<C$.

Writing $T(\lambda)=S(\lambda)S^{-1}$ and $\phi=\left(\lambda',\gamma'\right)'$, define the quantity \[\sigma ^{2}\left( \phi \right) =n^{-1}\sigma_0^2tr\left(T'(\lambda)\Sigma(\gamma)^{-1}T(\lambda)\Sigma\right)=n^{-1}\sigma_0^2\left\Vert E(\gamma)T(\lambda)E^{-1}\right\Vert_F^2, \] which is nonnegative by definition and bounded by Assumptions (ref) and (ref). The assumptions below directly extend Assumptions (ref) and (ref) to the present setup.

assumption$c\leq\sigma ^{2}\left( \phi \right)\leq C$, for all $\phi\in\Phi$.
assumption$\phi_0\in\Phi$ and, for any $\eta>0$, \begin{equation} \varliminf_{n\rightarrow\infty}\inf_{\phi\in\overline{\mathcal{N}}^{\phi}(\eta)}\frac{n^{-1}tr\left(T'(\lambda)\Sigma(\gamma)^{-1}T(\lambda)\Sigma\right)}{\left\vert T'(\lambda)\Sigma(\gamma)^{-1}T(\lambda)\Sigma\right\vert ^{1/n}}>1, \end{equation} where $\overline{\mathcal{N}}^{\phi}(\eta)=\Phi\setminus \mathcal{N}^{\phi}(\eta)$ and $\mathcal{N}^{\phi}(\eta)=\left\{\phi:\left\Vert\phi-\phi_0\right\Vert<\eta\right\}\cap\Phi$.

We now introduce an identification condition that is required in the setup of this section.

assumption$\beta_0\neq 0$ and for any $\eta>0$, \begin{equation} {P\left(\varliminf_{n\rightarrow\infty}\inf_{\left(\lambda',\gamma'\right)'\in\Lambda\times \overline{\mathcal{N}}^{\gamma}(\eta)}n^{-1}\beta_0'\Psi^{\prime }T^{\prime}(\lambda ) E(\gamma)'M\left( \gamma \right) E(\gamma) T(\lambda )\Psi\beta_0/\left\Vert \beta_0\right\Vert^2>0\right)=1.} \end{equation}

Upon performing minimization with respect to $\beta$, the event inside the probability in ((ref)) is equivalent to the event \[{\varliminf_{n\rightarrow\infty}\min_{\beta\in\mathbb{R}^p}\inf_{\left(\lambda',\gamma'\right)'\in\Lambda\times\overline{\mathcal{N}}^{\gamma}(\eta)}n^{-1}\left( \Psi\beta-T(\lambda)\Psi\beta_0\right)'\Sigma(\gamma)^{-1}\left( \Psi\beta-T(\lambda)\Psi\beta_0\right)/\left\Vert \beta_0\right\Vert^2>0,} \] which is analogous to the identification condition for the nonlinear regression model with a parametric linear factor in robinson1972non, weighted by the inverse of the error covariance matrix. This reduces the condition to a scalar form of a rank condition, making the identifying nature of the assumption transparent. A similar identifying assumption is used by Gupta2018.

theoremUnder either $H_0$ or $H_1$, Assumptions (ref)-(ref), (ref), (ref)-(ref) and \[p^{-1}+\left(d_\gamma+p\right)/n\rightarrow 0, \text{ as }n\rightarrow\infty,\] $\left\Vert\left(\widehat\phi,\widehat{\sigma}^2\right)-\left(\phi_0,{\sigma_0}^2\right)\right\Vert\overset{p}\longrightarrow 0$ as $n\rightarrow\infty$.

The test statistic $\mathscr{T}_n$ can be constructed as before but with the null residuals redefined to incorporate the spatially lagged terms, i.e. $\hat u=S(\hat\lambda)y-f(x,\hat\alpha)$. Then we have the following theorem.

theoremUnder Assumptions (ref)-(ref), (ref)-(ref), (ref)-(ref), \[p^{-1}+p\left(p+d_\gamma^2\right)/n+\sqrt{n}/p^{\mu+1/4}+d^2_\gamma/p\rightarrow 0, \text{ as }n\rightarrow\infty,\] and $H_0$, $ \mathscr{T}_n-{\left(\sigma_0^{-2}\varepsilon'\mathscr{V}\varepsilon-p\right)}/{\sqrt{2p}}=o_p(1). $
theoremUnder the conditions of Theorems (ref), (ref) and (ref), (1) $\mathscr{T}_n\overset{d}{\rightarrow} N(0,1)$ under $H_0$, (2) $\mathscr{T}_n$ is a consistent test statistic, (3) $\mathscr{T}_n\overset{d}{\rightarrow}N \left(\varkappa,1\right)$ under local alternatives $H_{\ell}$.

Nonparametric spatial weights

In this section we are motivated by settings where spatial dependence occurs through nonparametric functions of raw distances (this may be geographic, social, economic, or any other type of distance), as is the case in pinkse2002, for example. In their kind of setup, $d_{ij}$ is a raw distance between units $i$ and $j$ and the corresponding element of the spatial weight matrix is given by $w_{ij}=\zeta_0\left(d_{ij}\right)$, where $\zeta_0(\cdot)$ is an unknown nonparametric function. pinkse2002 use such a setup in a SAR model like ((ref)), but with a linear regression function. In contrast, in keeping with the focus of this paper we instead model dependence in the errors in this manner. Our formulation is rather general, covering, for example, a specification like $w_{ij}=f\left(\gamma_0,\zeta_0\left(d_{ij}\right)\right)$, with $f(\cdot)$ a known function, $\gamma_0$ an unknown parameter of possibly increasing dimension, and $\zeta_0(\cdot)$ an unknown nonparametric function. For the sake of simplicity, we do not permit the $x_i$ in this section to be generated by such nonparametric weight matrices although they can be generated from other, known weight matrices.

Let $\Xi$ be a compact space of functions, on which we will specify more conditions later. For notational simplicity we abstract away from the SAR dependence in the responses. Thus we consider ((ref)), but with

equation[equation omitted — 134 chars of source]

where $\zeta_0(\cdot)=\left(\zeta_{01}(\cdot),\ldots,\zeta_{0d_\zeta}(\cdot)\right)'$ is a fixed-dimensional vector of real-valued nonparametric functions with $\zeta_{0\ell}\in\Xi$ for each $\ell=1,\ldots,d_\zeta$, and ${z}_i$ a fixed-dimensional vector of data, independent of the $\varepsilon_s$, $s\geq 1$, with support $\mathcal{Z}$. One can also take $z_i$ to be a fixed distance measure. We base our estimation on approximating each $\zeta_{0\ell}({z_i})$, $\ell=1,\ldots,d_\zeta$, with the series representation $\delta_{0\ell}'\varphi_\ell({z_i})$, where $\varphi_\ell\left({z_i}\right)\equiv\varphi_{\ell}$ is an $r_\ell\times 1$ ($r_\ell\rightarrow\infty$ as $n\rightarrow\infty$) vector of basis functions with typical function $\varphi_{\ell k}$, $k=1,\ldots,r_\ell$. The set of linear combinations $\delta_\ell'\varphi_\ell({z_i})$ forms the sequence of sieve spaces $\Phi_{r_\ell}\subset \Xi$ as $r_\ell\rightarrow \infty$, for any $\ell=1,\ldots,d_\zeta$, and

equation[equation omitted — 114 chars of source]

with the following restriction on the function space $\Xi$: \setcounter{assumption}{0}

assumptionFor some scalars $\kappa_\ell>0$, $\left \Vert \nu_\ell \right \Vert_{w_z} =O\left(r_\ell^{-\kappa_\ell}\right),$ as $r_\ell\rightarrow\infty$, $\ell=1,\ldots,d_\zeta$, where $w_z\geq 0$ is the largest value such that $\sup_{z\in\mathcal{Z}} \mathcal{E}\left\Vert z\right\Vert^{w_z}<\infty$

Just as Assumption (ref) implied ((ref)), by Lemma 1 of Lee2016, we obtain

equation[equation omitted — 164 chars of source]

Thus we now have an infinite-dimensional nuisance parameter $\zeta_0(\cdot)$ and increasing-dimensional nuisance parameter $\gamma$. Writing $\sum_{\ell=1}^{d_\zeta}r_\ell=r$ and $\tau=(\gamma',\delta'_1,\ldots,\delta'_{d_\zeta})'$, which has increasing dimension $d_\tau=d_\gamma+r $, define $ \varsigma(r)=\sup_{z\in\mathcal{Z}; \ell=1,\ldots,d_\zeta}\left\Vert\varphi_{\ell}\right\Vert.$ Write $\Sigma(\tau)$ for the covariance matrix of the $n\times 1$ vector of $u_i$ in ((ref)), with $\delta_{\ell}'\varphi_\ell$ replacing each admissible function $\zeta_{\ell}(\cdot)$. This is analogous to the definition of $\Sigma(\gamma)$ in earlier sections, and indeed after conditioning on $z$ it can be treated in a similar way because $d_\gamma\rightarrow\infty$ was already permitted. For example, suppose that $u=(I_n-W)^{-1}\varepsilon$, where $\left\Vert W\right\Vert<1$ and the elements satisfy $w_{ij}=\zeta_0\left(d_{ij}\right)$, $i,j=1,\ldots,n$, for some fixed distances $d_{ij}$ and unknown function $\zeta_0(\cdot)$, see e.g. Pinkse1999. Approximating $\zeta_0(z)= \tau_0'\varphi(z)+\nu$, for some $r\times 1$ basis function vector $\varphi(z)$ and approximation error $\nu$, we define $W(\tau)$ as the $n\times n$ matrix with elements $w_{ij}(\tau)=\tau_0'\varphi\left(d_{ij}\right)$, and set $\Sigma(\tau)=\text{var}\left((I_n-W(\tau))^{-1}\varepsilon\right)=\sigma_0^2(I_n-W(\tau))^{-1}(I_n-W'(\tau))^{-1}$.

For any admissible values $\beta $, $\sigma ^{2}$ and $\tau $, the redefined (multiplied by $2/n$) negative quasi log likelihood function based on using the approximations ((ref)) and ((ref)) is

equation[equation omitted — 278 chars of source]

which is minimised with respect to $\beta $ and $\sigma ^{2}$ by

eqnarray[eqnarray omitted — 331 chars of source]

where $M(\tau )=I_{n}-E(\tau )\Psi \left( \Psi ^{\prime }\Sigma (\tau )^{-1}\Psi \right) ^{-1}\Psi ^{\prime }E(\tau )^{\prime }$ and $E(\tau )$ is the $n\times n$ symmetric matrix such that $E(\tau )E(\tau )^{\prime }=\Sigma (\tau )^{-1}$. Thus the concentrated likelihood function is

equation[equation omitted — 179 chars of source]

Again, for compact $\Gamma$ and sieve coefficient space $\Delta$, the QMLE of $\tau _{0}$ is $\widehat{\tau}=\text{arg min}_{\tau \in \Gamma\times\Delta }\mathcal{L}(\tau )$ and the QMLEs of $\beta $ and $\sigma ^{2}$ are $\widehat{\beta}=\bar{\beta}\left( \widehat{\tau}\right) $ and $\widehat{\sigma} ^{2}=\bar{\sigma}^{2}\left( \widehat{\tau}\right) $. The series estimate of $\theta _{0} $ is defined as in ((ref)). Define also the product Banach space $\mathcal{T}=\Gamma\times \Xi^{d_\zeta}$ with norm $\left\Vert \left(\gamma',\zeta'\right)'\right\Vert_{\mathcal{T}_w}=\left\Vert \gamma\right\Vert+\sum_{\ell=1}^{d_\zeta}\left\Vert \zeta_\ell\right\Vert_{w}$, and consider the map $\Sigma:\mathcal{T}^{o}\rightarrow \mathcal{M}^{n\times n}$, where $\mathcal{T}^{o}$ is an open subset of $\mathcal{T}$.

assumptionThe map $\Sigma:\mathcal{T}^o\rightarrow \mathcal{M}^{n\times n}$ is Fr\'echet-differentiable on $\mathcal{T}^o$ with Fr\'echet-derivative denoted $D\Sigma\in\mathscr{L}\left(\mathcal{T}^o,\mathcal{M}^{n\times n}\right)$. Furthermore, conditional on ${z}$, the map $D\Sigma$ satisfies \begin{equation} \sup_{t\in\mathcal{T}^o}\left\Vert D\Sigma(t)\right\Vert_{\mathscr{L}\left(\mathcal{T}^o,\mathcal{M}^{n\times n}\right)}\leq C, \end{equation} on its domain $\mathcal{T}^o$.

This assumption can be checked in a similar way to how we checked Assumption (ref), where a diverging dimension for the argument was already permitted.

propositionIf Assumption (ref) holds, then for any $t_1,t_2\in\mathcal{T}^o$, conditional on $z$, \begin{equation} \left\Vert\Sigma\left(t_1\right)-\Sigma\left(t_2\right)\right\Vert\leq C \varsigma(r)\left\Vert t_1-t_2\right\Vert. \end{equation}
corollaryFor any $t ^{* }\in \mathcal{T}^o $ and any $\eta >0 $, conditional on $z$, \begin{equation} \underset{n\rightarrow \infty }{\overline{\lim }}\sup_{t \in \left \{ t :\left \Vert t -t ^{*}\right \Vert <\eta \right \} \cap \mathcal{T}^o }\left \Vert \Sigma (t )-\Sigma \left( t ^{* }\right) \right \Vert <C\varsigma(r)\eta . \end{equation}
assumption$c\leq \sigma ^{2}\left( \tau \right)\leq C$ for $\tau \in \Gamma\times \Delta$, conditional on $z$.

Denote $\Sigma\left(\tau_0\right)=\Sigma_0$. Note that this is not the true covariance matrix, which is $\Sigma\equiv \Sigma\left(\gamma_0,\zeta_0\right)$.

assumption$\tau _{0}\in \Gamma\times\Delta $ and, for any $\eta >0$, conditional on $z$, \begin{equation} \varliminf_{n\rightarrow \infty }\inf_{\tau \in \overline{\mathcal{N}} ^{\tau }(\eta )}\frac{n^{-1}tr\left( \Sigma (\tau )^{-1}\Sigma_0 \right) }{ \left \vert \Sigma (\tau )^{-1}\Sigma_0 \right \vert ^{1/n}}>1, \end{equation} where $\overline{\mathcal{N}}^{\tau }(\eta )=(\Gamma\times\Delta) \setminus \mathcal{N} ^{\tau }(\eta )$ and $\mathcal{N}^{\tau }(\eta )=\left \{ \tau :\left \Vert \tau -\tau _{0}\right \Vert <\eta \right \} \cap (\Gamma\times \Delta) $.
remarkExpressing the identification condition in Assumption (ref) in terms of $\tau$ implies that identification is guaranteed via the sieve spaces $\Phi_{r_\ell}$, $\ell=1,\ldots,d_\zeta$. This approach is common in the sieve estimation literature, see e.g. Chen2007, p. 5589, Condition 3.1.
theoremUnder either $H_0$ or $H_1$, Assumptions (ref)-(ref) (with (ref) and (ref) holding for $t\in\mathcal{T}$ rather than $\gamma\in\Gamma$), (ref), (ref)-(ref) and $p^{-1}+\left(\min_{\ell=1,\ldots,d_\zeta}r_\ell\right)^{-1}+\left(d_\gamma+p+\max_{\ell=1,\ldots,d_\zeta}r_\ell\right)/n\rightarrow 0$ as $n\rightarrow\infty$, $\left \Vert \left(\widehat{\tau},\hat\sigma^2\right)-\left(\tau _{0},\sigma^2_0\right)\right \Vert \overset{p}{ \longrightarrow }0. $
theoremUnder the conditions of Theorems (ref) and (ref), but with $\tau$ and $\mathcal{T}$ replacing $\gamma$ and $\Gamma$ in assumptions prefixed with R and $p\rightarrow\infty$, \[\left(\min_{\ell=1,\ldots,d_\zeta}r_\ell\right)^{-1}+ \frac{p^2}{n}+\frac{\sqrt{n}}{p^{\mu+1/4}} +p^{1/2}\varsigma (r) \left(\frac{ d_{\gamma }+\displaystyle\max_{\ell=1,\ldots,d_\zeta}r_\ell}{\sqrt{n}}+\sqrt{\sum_{\ell=1}^{d_\zeta}r_\ell^{-2\kappa_\ell}}\right )\rightarrow 0, \] as $n\rightarrow\infty$, and $H_0$, $ \mathscr{T}_n-\left({\sigma_0^{-2}\varepsilon'\mathscr{V}\varepsilon-p}\right)/{\sqrt{2p}}=o_p(1). $
theoremLet the conditions of Theorems (ref) and (ref) hold, but with $\tau$ and $\mathcal{T}$ replacing $\gamma$ and $\Gamma$ in assumptions prefixed with R. Then (1) $\mathscr{T}_n\overset{d}{\rightarrow} N(0,1)$ under $H_0$, (2) $\mathscr{T}_n$ is a consistent test statistic, (3) $\mathscr{T}_n\overset{d}{\rightarrow}N \left(\varkappa,1\right)$ under local alternatives $H_{\ell}$.

Fixed-regressor residual-based bootstrap test

The performance of nonparametric tests based on asymptotic distributions often leaves something to be desired in finite samples. An alternative approach is to use the bootstrap approximation. In this section, we propose a bootstrap version of our test, focusing on the setting of Section (ref). In our simulations and empirical studies, we consider test statistics based on both $\widehat{m}_{n}=\widehat{\sigma }^{-2}\widehat{v}^{\prime }\Sigma \left( \widehat{\gamma }\right) ^{-1}\widehat{u}/n$ and $\widetilde{m}_{n}=\widehat{\sigma }^{-2}(\widehat{u}^{\prime }\Sigma \left( \widehat{\gamma }\right) ^{-1}\widehat{u}-\widehat{\eta }^{\prime }\Sigma \left( \widehat{\gamma }\right) ^{-1}\widehat{\eta })/n$, where $\widehat{\eta }=S(\hat\lambda)y-\widehat{\theta }$, i.e., the residual from nonparametric estimation, $\hat u=S(\hat\lambda)y-f(x,\hat \alpha)$, and $\hat v=\hat\theta-f(x,\hat \alpha)$. Analogous to the definition of $\mathscr{T}_n$, define the statistic $ \mathscr{T}_{n}^{a}={\left(n\widetilde{m}_{n}-p\right)}/{\sqrt{2p}}.$ In the case of no spatial autoregressive term, and under the power series, $\mathscr{T}_{n}^{a}$ and $\mathscr{T}_{n}$ are numerically identical, as was observed by Hong1995. However, in the SARSE setting a difference arises due to the spatial structure in the response $y$. We show that $\mathscr{T}_{n}^{a}-\mathscr{T}_{n}=o_{p}(1)$ under the null or local alternatives in Theorem (ref) in the online supplementary appendix.

The bootstrap versions of the test statistics $\mathscr{T}_n$ and $\mathscr{T}_n^{a}$ are

eqnarray*[eqnarray* omitted — 594 chars of source]

respectively, where $\widehat{{u}}^{\ast }$ is the bootstrap residual vector under the null, $\widehat{\eta }^*$ is the bootstrap residual vector under the alternative, $ \widehat{{v}}^{\ast }=\widehat{\theta }^{\ast }(x)-f(x,\widehat{ \alpha }^{\ast })$, and $\left(\widehat{ \gamma }^{\ast },\lambda^*,\widehat{\sigma }^{\ast 2},\widehat{\theta }^{\ast },\widehat{\alpha }^{\ast }\right)$ is the estimator using the bootstrap sample. We elaborate on the bootstrap statistics using the SARARMA($m_{1}$,$ m_{2},m_{3}$) model as an example:

equation*[equation* omitted — 172 chars of source]

Following Jin2015, we first deduct the empirical mean of the residual vector from

equation*[equation* omitted — 271 chars of source]

to obtain $\widetilde{{\xi }}=(I_{n}-\frac{1}{n}l_{n}l_{n}^{\prime }) \widehat{{\xi }}$. Next, we sample randomly with replacement $n$ times from elements of $\widetilde{{\xi }}$ to obtain a vector of $ \mathbf{\xi }^{\ast }.$ After this, we generate the bootstrap sample $ y^{\ast }$ by treating $\widehat{{f}}=f(x,\widehat{\alpha })$, $\hat\lambda$ and $\widehat{\gamma }$ as the true parameter:

equation*[equation* omitted — 294 chars of source]

We estimate the model based on the bootstrap sample $y^{\ast }$ using QMLE to obtain the estimator $\widehat{{\theta }}^{\ast }=\psi ^{\prime }\widehat{\beta }^{\ast },$ $\widehat{\lambda }^{\ast }$, and $ \widehat{\gamma }^{\ast }$ under the alternative hypothesis and $\widehat{ \alpha }^{\ast }$ under the null hypothesis of $\theta (x)=f(x,\alpha _{0}).$ Then, $\widehat{{\eta}}^{\ast }=y^{\ast }-\sum_{k=1}^{m_{1}} \widehat{\lambda }_{k}^{\ast }W_{1k}y^{\ast }-\widehat{{\theta }}^{\ast }$, $\widehat{{u }}^{\ast }=y^{\ast }-\sum_{k=1}^{m_{1}} \widehat{\lambda }_{k}^{\ast }W_{1k}y^{\ast }-f(x,\widehat{\alpha }^{\ast }).$

This procedure is repeated $B$ times to obtain the sequence $\left \{ \mathscr{T}_{nj}^{\ast }\right \} _{j=1}^{B}$. We reject the null when $p^{\ast }=B^{-1}\sum_{j=1}^{B}\mathbf{1(}\mathscr{T}_n<\mathscr{T}_{nj}^{\ast })$ is smaller than the given level of significance. An identical procedure holds for the test based on $\mathscr{T}_n^{a\ast }.$ The asymptotic validity of the bootstrap method can be shown as in Theorem 4 of Su2017 and Lemma 2 in Jin2015, and detailed analysis can be found in the supplementary appendix, see proof of Theorem TS.1.

Finite sample performance

Parametric error spatial structure

\sloppy Taking $n=60,100,200$, we choose two specifications to generate $y$ from the SARARMA($m_{1}$,$m_{2},m_{3}$) models:

eqnarray*[eqnarray* omitted — 213 chars of source]

where $\xi$ is $N(0,I_n)$. The DGP of $\theta (x)$ is

equation*[equation* omitted — 108 chars of source]

where $x_{i}^{\prime }\alpha =1+x_{1i}+x_{2i}$, with $ x_{1i}=(z_{i}+z_{1i})/2 $, $x_{2i}=(z_{i}+z_{2i})/2$. We choose two settings: compactly supported regressors where $z_{i},z_{1i}$ and $z_{2i}$ are i.i.d., $U[0,2\pi ]$ and unboundedly supported regressors where $z_{i}, z_{1i}$ and $z_{2i}$ are i.i.d. $N(0,1).$ We report the compact support setting in the main text, while the results for unbounded support are reported in the online supplement.

\sloppy We use three series bases for our experiments: power (polynomial) series of third and fourth order ($p=10,p=15$), trigonometric series $trig_1=\left(1, \sin\left(x_1\right), \sin\left(x_1/2\right), \sin\left(x_2\right), \sin\left(x_2/2\right),\cos\left(x_1\right),\cos\left(x_1/2\right),\cos\left(x_2\right),\cos\left(x_2/2\right)\right)'$ and $trig_2=\left(trig_1',\sin\left(x_1^2\right),\cos\left(x_1^2\right),\sin\left(x_2^2\right),\cos\left(x_2^2\right)\right)'$, and the B-spline bases of fourth and seventh order ($p=9,p=14$), We also set $\gamma _{2}=0.3$, $\lambda _{1}=0.3$ and $\gamma _{3}=0.4$; the value $c=0,3,6$ indicates the null hypothesis and the local alternatives. The spatial weight matrices are generated using LeSage's code make_neighborsw from http://www.spatial-econometrics.com/, where the row-normalized sparse matrices are generated by choosing a specific number of the closest locations from randomly generated coordinates and we set the number of neighbors to be $n/20$. We employ 100 bootstrap replications in each of 500 Monte Carlo replications except for the SARARMA(1,0,1) design with $n=200$, where we set 50 bootstrap replications in view of the computation time. We report the rejection frequencies of tests based on bootstrap critical values in the main text, while tests based on asymptotic critical values are reported in the online supplement.

Tables (ref)-(ref) report the empirical rejection frequencies using the bootstrap test statistics $\mathscr{T}_n^\ast$ (Tables (ref), (ref)) and $\mathscr{T}_n^{a \ast}$ (Tables (ref), (ref)), when nominal levels are given by 1%, 5% and 10%. To see how the choice of $p$ and the basis functions affect small sample outcomes, we report two sets of results for each basis function family: the first row for each value of $c$ is from the smaller $p$ ($p=9$ or $10$), while the second row is from the larger $p$ ($p=14$ or $15$). We summarize some important findings. First, we see that for most DGPs, our bootstrap test is closer to the nominal level than the asymptotic test (reported in the online supplement) although the sizes of both types of tests improve generally as the sample size increases. Second, both bootstrap and asymptotic tests are powerful in detecting any deviations from linearity in the local alternatives. The patterns are similar across all cases: the bootstrap generally affords better size control, albeit not always.

All three types of bases give qualitatively similar results, but we note that $\mathscr{T}_n^\ast=\mathscr{T}_n^{ \ast a }$ when using polynomial series under the SARARMA(0,1,0) model, as observed in Hong1995. When using trigonometric and B-spline series, tests based on these two statistics give slightly different rejection rates. However, under the SARARMA(1,0,1) model, all series give quantitatively different results, as illustrated in Tables (ref) and (ref). When using B-spline bases, $p=14$ does not perform well compared to $p=9$. In the other cases, both choices of $p$ work well.

Nonparametric error spatial structure

Now we examine finite sample performance in the setting of Section (ref). The three DGPs of $\theta (x)$\ are the same as the parametric setting but we generate the $n\times n$ matrix $W^*$ as $w^*_{ij}=\Phi (-d_{ij}) I(c_{ij}<0.05)$ if $ i\neq j$, and $w^*_{ii}=0$, where $\Phi (\cdot )$ is the standard normal cdf, $d_{ij}\sim$iid $U[-3,3]$, and $c_{ij}\sim$iid $U[0,1]$. From this construction, we ensure that $W^*$ is sparse with no more than $5\%$ elements being nonzero. Then, $y$ is generated from $ y=\theta (x)+u,\text{ }u=Wu+\xi , $ where $\xi\sim N(0,I_n)$ and $W=W^*/{1.2\overline{\varphi }\left(W^*\right)}$, ensuring the existence of $(I-W)^{-1}$. In estimation, we know the distance $d_{ij}$ and the indicator $I(c_{ij}<0.05)$, but we do not know the functional form of $w_{ij}$, so we approximate elements in $W$ by $ \widehat{w}_{ij}=\sum_{l=0}^{r}a_{l}d_{ij}^{l}I(c_{ij}<0.05)\text{ if } i\neq j\text{; }\widehat{w}_{ii}=0.$

Table (ref) reports the rejection rates using 500 Monte Carlo simulation at the 5% asymptotic level 1.645 using polynomial bases with $r=2,3,4, 5$ and $p=10,15, 20$. We take $n=150, 300, 500, 600, 700$, larger sample sizes than earlier because two nonparametric functions must be estimated in this spatial setting. The two largest bandwidths ($r=5, p=20$) are only employed for the largest sample size $n=700$. We observe a clear pattern of rejection rates approaching the theoretical level as sample size increases. Power improves as $c$ increases for all designs and is non-trivial in all cases even for $c=3$. Sizes are acceptable for $n=500$, particularly when $p=15$. Size performance improves further as $n=600$, indicating asymptotic stability. Note that with two diverging bandwidths ($p$ and $r$), we expect sizes to improve in a diagonal pattern going from top left corner to bottom right corner in Table (ref). This is indeed the case. For $n=700$, we observe that the pairs $(r,p)=(5,15),(5,20)$ deliver acceptable sizes.

Empirical applications

{ In this section, we illustrate the specification test presented in previous sections using several empirical examples. }

Conflict alliances

This example is based on a study of how a network of military alliances and enmities affects the intensity of a conflict, conducted by konig2017networks. They stress that understanding the role of informal networks of military alliances and enmities is important not only for predicting outcomes, but also for designing and implementing policies to contain or put an end to violence. konig2017networks obtain a closed-form characterization of the Nash equilibrium and perform an empirical analysis using data on the Second Congo War, a conflict that involves many groups in a complex network of informal alliances and rivalries.

To study the fighting effort of each group the authors use a panel data model with individual fixed effects, where key regressors include total fighting effort of allies and enemies. They further correct the potential spatial correlation in the error term by using a spatial heteroskedasticity and autocorrelation robust standard error. We use their data and the main structure of the specification and build a cross-sectional SAR(2) model with two weight matrices, $W^A$ ($W^A_{ij}=1$ if group $i$ and $j$ are allies, and $W^A_{ij}=0$ otherwise) and $W^E$ ($W^E_{ij}=1$ if group $i$ and $j$ are enemies, and $W^E_{ij}=0$ otherwise):

equation*[equation* omitted — 96 chars of source]

where $y$ is a vector of fighting efforts of each group and $X$ includes the current rainfall, rainfall from the last year, and their squares.\footnote{We follow the analysis in the original paper and do not row normalize. This is because the economic content of the weight matrices is defined by total fights of allies or enemies.} To consider the spatial correlation in the error term, we consider both the Error SARMA(1,0) and Error SARMA(0,1) structures. For these, we employ a spatial weight matrix $W^d$, based on the inverse distance between group locations and set to be 0 after 150 km, following konig2017networks. The idea is that geographical spatial correlation dies out as groups become further apart. We also report results using a nonparametric estimator of the spatial weights, as described in Section (ref) and studied in simulations in Section (ref). For the nonparametric estimator we take $r=2$.

In the original dataset, there are 80 groups, but groups 62 and 63 have the same variables and the same locations, so we drop one group and end up with a sample of 79 groups. We use data from 1998 as an example and further use the pooled data from all years as a robustness check. $H_{0}$ stands for restricted model where the linear functional form of the regression is imposed, while $H_{1}$ stands for the unrestricted model where we use basis functions comprising of power series with $p=10$. In all our specifications, the test statistics are negative, so we cannot reject the null hypothesis that the model is correctly specified. As Table (ref) indicates, this failure to reject the null persists when we use pooled data from 13 years, yielding 1027 observations. Thus we conclude that a linear specification is not inappropriate for this setting. One possible reason is that the original regression, though linear, has already included the squared terms of the rainfall as regressors. This finding is robust to using the bootstrap tests of Section (ref), which generally yield smaller p-values but unchanged conclusions.

Innovation spillovers

This example is based on the study of the impact of R&D on growth from Bloom2013. They develop a general framework incorporating two types of spillovers: a positive effect from technology (knowledge) spillovers and a negative `business stealing' effect from product market rivals. They implement this model using panel data on U.S. firms.

We consider the Productivity Equation in Bloom2013:

equation[equation omitted — 134 chars of source]

where $y$ is a vector of sales, $R\&D$ is a vector of R&D stocks, and regressors in $X$ include the log of capital ($Capital$), log of labor ($Labor$), $R\&D$, a dummy for missing values in $R\&D$, a price index, and two spillover terms constructed as the log of $W_{SIC} R\&D$ ($Spsic$) and the log of $W_{TEC} R\&D$ ($Sptec$), where $W_{SIC}$ measures the product market proximity and $W_{TEC}$ measures the technological proximity. Specifically, they define \[W_{SIC,ij}={S_{i}S_{j}^{\prime }}/{(S_{i}S_{i}^{\prime })^{1/2}(S_{j}S_{j}^{\prime })^{1/2}},W_{TEC,ij}={ T_{i}T_{j}^{\prime }}/{(T_{i}T_{i}^{\prime })^{1/2}(T_{j}T_{j}^{\prime })^{1/2}},\] where $S_{i}=(S_{i1},S_{i2},\ldots,S_{i597})'$, with $S_{ik}$ being the share of patents of firm $i$ in the four digit industry $k$ and $T_{i}=(T_{i1},T_{i2},\ldots,T_{i426})'$, with $T_{i\tau }$ being the share of patents of firm $i$ in technology class $\tau $. Focusing on a cross-sectional analysis, we use observations from the year 2000 and obtain a sample size of 577. Both weight matrices are row normalized.

The column FE of Table (ref) is from Table 5 of Bloom2013 based on their panel fixed effects estimation and we use it as a baseline for comparison. This table reports results for SARARMA(0,1,0) models using $W_{SIC}$ and $W_{TEC}$ separately. We use both $W_{SIC}$ and $W_{TEC}$ simultaneously in SARARMA(0,2,0), SARARMA(0,2,0), and Error MESS(2) models, reported in Table (ref). In all of these specifications, the test statistics are larger than 1.645, so we reject the null hypothesis of the linear specification. This rejection also persists with the bootstrap tests, albeit the p-values go up compared to the asymptotic ones. However, we can say even more as our estimation also sheds light on spatial effects in the disturbances in ((ref)). As before $H_{0}$ imposes linear functional form of the regressors, while $H_{1}$ uses the nonparametric series estimate employing power series with $p=10$. Regardless of the specification of the regression function, the disturbances suggest a strong spatial effect as the coefficients on $W_{TEC}$ and $W_{SIC}$ are large in magnitude.

Economic growth

{ The final example is based on the study of economic growth rate in Ertur2007. Knowledge accumulated in one area might depend on knowledge accumulated in other areas, especially in its neighborhoods, implying the possible existence of spatial spillover effects. These questions are of interest to both economists as well as regional scientists. For example, Autant-Bernard2011 examine spatial spillovers associated with research expenditures for French regions, while Ho2013 examine the international spillover of economic growth through bilateral trade amongst OECD countries, Cuaresma and Feldkircher (2013) study spatially correlated growth spillovers in the income convergence process of Europe, and Evans2014 study the spatial dynamics of growth and convergence in Korean regional incomes. }

{ In this section, we want to test the linear SAR model specification in Ertur2007. Their dataset covers a sample of 91 countries over the period 1960-1995, originally from Heston2002, obtained from the Penn World Tables (PWT version 6.1). The variables in use include per worker income in 1960 ($y60$) and 1995 ($ y95 $), average rate of growth between 1960 and 1995 $(gy)$, average investment rate of this period ($s$), and average rate of growth of working-age population ($n_{p}$).

{ Ertur2007 consider the model

equation[equation omitted — 89 chars of source]

where the dependent variable is log real income per worker $\ln (y95)$, elements of the explanatory variable $X=(x_1',x_2')$ include log investment rate $ \ln (s)=x_1$ and log physical capital effective rate of depreciation\ $ \ln (n_{p}+0.05)=x_2$, with corresponding subscripted coefficients $\beta_1,\beta_2,\theta_1,\theta_2$. A restricted regression based on the joint constraints $\beta _{1}=-\beta _{2}$\ and $\theta _{1}=-\theta _{2}$ (these constraints are implied by economic theory) is also considered in Ertur2007. The model ((ref)) has regressors $(X,WX)$ and iid errors, so the test derived in Section (ref) can be directly applied here. Denoting by $d_{ij}$ the great-circle distance between the capital cities of countries $i$ and $j$, one construction of $W$ takes $w_{ij}=d_{ij}^{-2}$ while the other takes $w_{ij}=e^{-2d_{ij}}$, following Ertur2007.

Table (ref) presents the estimation and testing results based on using linear and quadratic power series basis functions with $p=10$ and a sample size of $n=91$. We impose additive structure in our estimation to at least alleviate the curse of dimensionality, always a concern in nonparametric estimation. We also use only linear and quadratic basis functions to reduce the number of terms for series estimation.

We cannot reject linearity of the regression function for the unrestricted model. On the other hand, linearity is rejected for the restricted model, which is the preferred specification of Ertur2007, with $w_{ij}=e^{-2d_{ij}}$. Thus, not only can we conclude that the specification of the model is under suspicion we can also infer this is due to constraints from economic theory. The findings are supported by the bootstrap tests of Section (ref).

Conclusion

This paper justifies a specification test for the regression function in a model where data are spatially dependent. The test is based on a nonparametric series approximation and is consistent. The paper also allows for some robustness in error spatial dependence by permitting this to be a nonparametric function of an underlying economic distance. On the other hand, our Section (ref) imposes correct specification of the spatial weight matrices $W_j$ in the SAR context, while Sun2020 allows these to be nonparametric functions as well. Thus our work acts as a complement to existing results in the literature and future work might combine both aspects.

{

table[table omitted — 4,660 chars of source]
table[table omitted — 4,564 chars of source]
table[table omitted — 4,592 chars of source]
table[table omitted — 4,582 chars of source]

}

table[table omitted — 2,016 chars of source]
table[table omitted — 3,115 chars of source]

{

table[table omitted — 2,115 chars of source]
table[table omitted — 3,860 chars of source]

}

table[table omitted — 1,718 chars of source]
center[center omitted — 48 chars of source]