Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
89,685 characters · 17 sections · 64 citation commands
Statistical Inference for the Logarithmic Spatial Heteroskedasticity Model with Exogenous Variables $^1$Department of Statistics & Actuarial Science, University of Hong Kong, Hong Kong $^2$School of Mathematics, Jilin University, Changchun 130012, China $$Corresponding author. E-mail: [email removed]
The data with spatial dependence across spatial units are often observed in urban, real estate, regional, public, agricultural, and environmental economics (see the surveys in LeSage2008 and Anselin2010). To capture the spatial dependence in mean, the spatial autoregressive (SAR) model introduced by Cliff1981 is widely used, and its generalizations and statistical inference methods are investigated by Lee2004, YDL, Kelejian2010, SJ, and Sun2018 to name just a few. However, the SAR model is inadequate to deal with the spatial dependence in variance, which could co-exist with the spatial dependence in mean under many circumstances; see, for example, the empirical elaborations in Sato2017 and Otto2018.
To study the spatial dependence in variance or spatial heteroscedasticity, Sato2017 propose the logarithmic spatial autoregressive conditional heteroscedasticity (log-SARCH) model defined as
where $Y_n=(y_{1,n},...,y_{n,n})'$ is an $n$-dimensional vector of dependent variables with $y_{i,n}$ being the observation in the spatial unit $i$ and $n$ being the total number of spatial units, ${Y_n^{2}}= (y_{1,n}^2,\dots,y_{n,n}^2)'$, $H_n=(h_{1,n},...,h_{n,n})'$ is an $n$-dimensional vector of non-negative variables, $\text{diag}(A)$ is an $n\times n$ diagonal matrix with entries of $A$ on the diagonal, $\log(A)$ is an entry-wise vector-valued function on $A$, $1_n$ is an $n$-dimensional vector of ones, $M_n=[m_{ij,n}]$ is an $n\times n$ spatial weight matrix, $V_n=(v_{1,n},...,v_{n,n})'$ is an $n$-dimensional vector of independent and identically distributed (i.i.d.) errors with mean zero, and $\alpha$ and $\rho$ are two unknown parameters. Under the log-SARCH model, $y_{i,n}$ has the spatial heteroscedasticity, that is, it has different variances across $i$. With a slight modification, the log-SARCH model is extended to the logarithmic spatial generalized ARCH model in Sato2021.
Although the log-SARCH model is able to capture the spatial heteroscedasticity, it has three major shortcomings. First, the log-SARCH model does not include any exogenous variables, which have significant effects to explain the spatial dependence in mean, as demonstrated by many existing works on SAR models. Second, the log-SARCH model adopts the spatial weight matrix $M_n$ with a nonzero entry $m_{ij,n}$ to capture the direct influence of $y_{j,n}$ to the variance of $y_{i,n}$ for $j\not=i$. However, this influence has the local character in nature, and it overlooks a global character that the variance of $y_{i,n}$ could also be affected by $y_{j,n}$ in an indirect way through some intermediate transmissions within the spatial units (Anselin2003). Third, the log-SARCH model is used without a systematic statistical inference procedure. So far only Sato2017 propose a two-step maximum likelihood (2SML) estimator for this model, based on the spatial auto-regressive moving average (SARMA) transformation. Their methodology is limited in practice, since the 2SML estimator is deemed to be sub-efficient due to the SARMA transformation, and no valid statistical inference tools (e.g., the classical tests on parameters and model diagnostic checking) is provided for the log-SARCH model.
This paper contributes to the literature in two aspects. First, we propose the generalized logarithmic spatial heteroscedasticity model with exogenous variables (denoted by the log-SHE model) to study the spatial dependence in variance. This log-SHE model nests the log-SARCH model as a special case. It not only accounts for the exogenous variables as in the SAR model, but also enables the handling of global character in the variance by employing the structures of spatial matrix exponential (SME) and spatial moving average (SMA) (LeSage2007; Fingleton2008). Moreover, we study the spatial near-epoch dependence (NED) property of the log-SHE model, and the NED property is crucial to provide the asymptotic theory of inference for the log-SHE model; see the related discussions in Jenish2009,Jenish2012, Xu2015a,Xu2015b, Liu2019, and Jin2020.
Second, we provide a systematic statistical inference procedure for the log-SHE model. Specifically, the maximum likelihood (ML) estimator is proposed by assuming that the errors $\{v_{i,n}\}_{i=1}^n$ is a sequence of i.i.d. $N(0,1)$ random variables. Since the ML estimator is often inconsistent for the non-normal errors, the generalized method of moments (GMM) estimator is constructed, and the corresponding optimal GMM (OGMM) estimator is further raised to improve the estimation efficiency. Under regular conditions, the consistency and asymptotic normality of ML, GMM, and OGMM estimators are established. Furthermore, based on the OGMM estimator, the Wald, Lagrange multiplier (LM), and likelihood ratio (LR) type D tests are proposed to detect the constraints of model parameters, and the overidentification test is given to check the model adequacy. Under regular conditions, all proposed tests are shown to have the chi-square limiting null distributions, and they thus are easy-to-implement in practice. Finally, the importance of our entire methodologies is illustrated by simulation studies and one real example on the house selling price in the U.S.
Besides the log-SHE model, there has another strand of literature to study the spatial dependence in variance by the following SARCH model (Otto2018):
where all notations are inherited from model ((ref)) except that $\alpha\geq0$ and $\rho\geq 0$. Compared with the log-SHE model, the SARCH model and its extensions in Otto2019b, Merk2021, and Otto2019a have two major drawbacks. First, the SARCH-type models need the bounded model errors to guarantee their existence, and this condition rules out the commonly used error distributions. Second, the SARCH-type models need complex constrains on parameters for the positivity of $h_{i,n}$ when the exogenous variables (taking either positive or negative values) are included, and it seems infeasible to account for those constraints in the computation of model estimation.
The remaining paper is organized as follows. Section (ref) introduces the specification of the log-SHE model, and studies its related NED properties. Section (ref) proposes the ML, GMM, and OGMM estimators, and establishes their asymptotic properties. Section (ref) constructs the Wald, LM, and D tests to detect the parameter constraints and the overidentification test to examine the model adequacy. Simulations are given in Section (ref). A real application is provided in Section (ref). Concluding remarks are offered in Section (ref). Technical lemmas and proofs are deferred into appendices.
Throughout the paper, $\mathcal{R}$ is the one-dimensional Euclidean space. For a matrix $A=[a_{ij}]\in\mathcal{R}^{p\times q}$, $A'$ is its transpose, ${\rm tr}(A)$ is its trace when $p=q$, $|A|=\|A\|_{F}$ is its Frobenius norm, and $\|A\|_{\infty}$ is its $L_{\infty}$-norm. For a random variable $\xi\in\mathcal{R}$, $||\xi||_p=({\rm E}|\xi|^p)^{1/p}$ is its $L_p$-norm for $0<p<\infty$. A sequence of matrices $\{A_{i,n}\}_{i=1}^n$ is uniformly $L_{\infty}$-bounded if $\sup_{i,n}||A_{i,n}||_{\infty}<\infty$, and a sequence of random variables $\{\xi_{i,n}\}_{i=1}^n$ is uniformly $L_p$-bounded if $\sup_{i,n} ||\xi_{i,n}||_p<\infty$. Moreover, $I_n$ denotes the $n\times n$-dimensional identity matrix, $O(1)$ denotes a generic bounded constant, $o_p(1) (O_p(1))$ denotes a sequence of random vectors converging to zero (bounded) in probability, “$\stackrel{\text{p}}{\longrightarrow}$” denotes the convergence in probability, and “$\stackrel{\text{d}}{\longrightarrow}$” denotes the convergence in distribution. All limits are taken as $n\to\infty$, unless stated otherwise.
In this section, we first give the specification of the log-SHE model and then investigate its related NED properties.
Let $\mathcal{A}_n(\rho)$ be a polynomial function of the spatial weight matrix $M_n$ having the form
where all coefficients $a_{l}(\rho)\in\mathcal{R}$ are deterministic functions of $\rho$ satisfying $\mathcal{A}_n(0)=I_n$. Using $\mathcal{F}_n(\rho)=I_n-\mathcal{A}_n(\rho)$, our log-SHE model is defined as
where $Z_n$ is an $n\times K$-dimensional matrix of variables (including exogenous variables) with rows $Z_{i,n}'=(z_{i1,n},...,z_{iK,n})$, $\gamma$ is its corresponding $K$-dimensional vector of parameters, and other notations are inherited from model ((ref)). Based on model ((ref)), we have
From ((ref))--((ref)), we know that the distribution of $H_n$ is invariant when $V_n$ is replaced by $-V_n$. Therefore, if $v_{i,n}$ is symmetric about zero (e.g., $v_{i,n}\sim N(0, 1)$), it follows that ${\rm E}(y_{i,n})=0$ and ${\rm Var}(y_{i,n})={\rm E}(h_{i,n}v_{i,n}^2)$ for all $i$. This indicates that the role of $h_{i,n}$ in the log-SHE model is mainly to depict the variance of $y_{i,n}$.
Needless to say, the structure of $h_{i,n}$ depends on both $Z_n\gamma$ and $\mathcal{A}_n(\rho)$ (or $\mathcal{F}_n(\rho)$). The term $Z_n\gamma$ determines how the exogenous variables affect $h_{i,n}$. For example, we can follow the spatial Durbin model (see Anselin1988) to take
where $\gamma=(\alpha,\beta',\beta_m')'$, and $M_nX_n$ is the spatial lag of exogenous variables $X_n$. The polynomial function $\mathcal{A}_n(\rho)$ (or $\mathcal{F}_n(\rho)$) controls how the dependent variables over the space affect $h_{i,n}$. For example, we can consider the following three special cases of $\mathcal{A}_n(\rho)$:
where the corresponding log-SHE models are named as the log-SARHE, log-SMAHE, and log-SMEHE models, respectively. In the SAR-type case, there is a direct impact from $y_{j,n}$ to the variance of $y_{i,n}$ (via $h_{i,n}$) when $M_n$ has a nonzero entry $m_{ij,n}$. This kind of impact is local, since $y_{j,n}$ only affects the variances of its immediate neighbors based on the specification of $M_n$. In the SMA- and SME-type cases, the impact from $y_{j,n}$ to the variance of $y_{i,n}$ is allowed to have a global character, meaning that $y_{j,n}$ can have an indirect influence on the variances of its non-immediate neighbors through the powers of $M_n$; see LeSage2007 and Fingleton2008 for more discussions on this aspect.
When $Z_n\gamma$ satisfies the spatial Durbin specification in ((ref)) with $(\beta',\beta_m')'=0$ and $\mathcal{A}_n(\rho)$ has the SAR-type form, our log-SHE model reduces to the log-SARCH model in ((ref)). Note that the log-SHE model (including the log-SARCH model) essentially studies the unconditional variance of $y_{i,n}$. Hence, the wording “conditional heteroscedasticity” used by the log-SARCH model seems inappropriate, and we instead use the wording “heteroscedasticity” to form the name of log-SHE model in this paper.
In order to study the NED properties of the log-SHE model, some notation and definitions are used throughout the paper. We assume that all spatial units are located in a region $D_n\subset D\subset\mathcal{R}^r$. For convenience, we endow $\mathcal{R}^r$ with the metric $d(i,j)=\max_{1\leq l\leq r}|i_l-j_l|$, and the corresponding norm $|i|=\max_{1\leq l\leq r}|i_{l}|$, where $i_l$ denotes the $l$-th component of $i$. Furthermore, the cardinality of a finite subset $U$ is defined as $|U|_{card}$.
We first make the following regular assumption for the increasing domain asymptotics; see Jenish2009,Jenish2012.
Next, we introduce the definition of $L_p$-NED in Jenish2009.
For non-linear spatial models (e.g., the log-SHE model in ((ref))), how to investigate the NED properties is key to studying their asymptotic properties, which commonly rely on the law of large numbers (LLN) and central limit theory (CLT) for the spatial NED process in Jenish2009,Jenish2012. This is different from linear spatial models (e.g., the SAR model in Cliff1981), which can simply use the LLN and CLT for linear-quadratic forms in Kelejian2001 to derive their asymptotic properties.
Let $\theta=(\rho,\gamma')'\in \Theta$ be the parameter vector in model ((ref)), and $\theta_0=(\rho_0,\gamma_0')'\in \Theta$ be its true value, where $\Theta=\Theta_{\rho}\times \Theta_{\gamma}\in \mathcal{R}\times \mathcal{R}^{K}$ is the parameter space, $\gamma=(\gamma_{1},...,\gamma_{K})'$, and $\gamma_0=(\gamma_{0,1},...,\gamma_{0,K})'$. By ((ref)), we can write
provided that $\mathcal{A}_n(\rho)$ is non-singular, where $\bar{\mathcal{A}}_n(\rho)$ is the inverse of $\mathcal{A}_n(\rho)$, and $\dot{\mathcal{A}}_n(\rho)$, $\ddot{\mathcal{A}}_n(\rho)$ and $\dddot{\mathcal{A}}_n(\rho)$ are the first, second, and third derivatives of $\mathcal{A}_n(\rho)$ with respect to $\rho$, respectively. Denote
To establish the NED properties of the log-SHE model, we need the following additional assumptions:
Assumption (ref)(i) gives common conditions about spatial units and the weight matrix $M_n$. Assumption (ref)(ii) accommodates the theoretical results for spatial NED processes. It allows units far from each other to have spatial dependence, but the strength of dependence declines with the distance. Assumption (ref)(iii) bounds the influence to one unit from other units by $|\rho|$. Under Assumptions (ref)--(ref) and Lemma A.1(iii) in Jenish2009, it follows that $M_n$ is uniformly bounded in both row and column sums (see the definition in Lee2004). All of conditions in Assumption (ref) can describe most spatial dependence in practice, and they hold for the often used Rook and $k$--nearest neighbors weight matrices.
Assumption (ref)(i) gives a sufficient condition for the existence of the log-SHE model. Let $\zeta=\sup_{n,\rho}||\rho M_n||_{\infty}$. For log-SARHE and log-SMAHE models, Assumption (ref)(i) holds when $\zeta<1$, under which $I_n-\rho M_n$ is non-singular by Corollary 5.6.16 of Horn2012. For the log-SMEHE model, Assumption (ref)(i) holds when $\zeta<\infty$, under which $\big(\text{Exp}(\rho M_n)\big)^{-1}=\text{Exp}(-\rho M_n)$. Assumption (ref)(ii) is useful to establish the $L_p$-bounded properties in Proposition (ref) below. By using the sub-multiplicative property of the matrix norm, it is not hard to see that the above conditions on $\zeta$ are also sufficient for Assumption (ref)(ii) under log-SARHE, log-SMAHE, and log-SMEHE models. Clearly, if $||M_n||_{\infty}=1$ under Assumption (ref)(i), we have $\zeta=\rho_*$, where $\rho_*=\sup_{\rho}|\rho|$; in this case, Assumption (ref)(i) holds when
Assumption (ref) is necessary to obtain the NED properties in Proposition (ref) below. A sufficient condition for Assumption (ref) is ${\rm limsup}_{l\rightarrow\infty}|a_{l+1,\kappa}|/|a_{l,\kappa}|<1$, provided that $||M_n||_{\infty}=1$. Therefore, when Assumption (ref)(i) holds, the condition for $\rho_*$ in ((ref)) suffices the validity of Assumption (ref) under log-SARHE, log-SMAHE, and log-SMEHE models. Assumption (ref) is basic and satisfied by most variables $Z_{i,n}$ and errors $v_{i,n}$.
Assumption (ref) poses some uniform moment conditions, and it is used to prove the $L_2$-bounded and $L_{2}$-NED properties in Propositions (ref)--(ref) below. According to Proposition 3.2 in Francq2013, the condition $\sup_{i,n} \|\exp\left(6 c_a|\log(v_{i,n}^2)|\right)\|<\infty$ holds when (i) $v^2_{i,n}$ is uniformly $L_{6 c_a}$-boundness; and (ii) $v_{i,n}$ admits a density $f$ around 0 with $f(v^{-1})=O(|v|^{\iota})$ for $\iota < 1$ and $|v|\rightarrow\infty$. Clearly, the two sufficient conditions above can be checked easily.
Denote
where $f_{ij,n}(\rho)$, $a_{ij,n}(\rho)$, $\bar{a}_{ij,n}(\rho)$, $\dot{a}_{ij,n}(\rho)$, $\ddot{a}_{ij,n}(\rho)$, and $\dddot{a}_{ij,n}(\rho)$ are defined in ((ref)). We are now ready to show some uniformly $L_p$-bounded and $L_2$-NED properties for the log-SHE model, and these results are key to using the LLN and CLT for the spatial NED process (see Lemma (ref) below) in our technical proofs.
In this section, we study the asymptotics of the ML, GMM, and OGMM estimators for the log-SHE model, and discuss some issues of the existing 2SML estimator for the model.
In order to propose the ML estimator for the log-SHE model, we follow the convention to assume normal distributed errors.
Under Assumption (ref), the likelihood function $L_n(\theta)$ of $Y_n$ can be given by
where $\varphi(\cdot)$ is the density function of $N(0,1)$ distribution, $A_n^{\dagger}(\theta)=\big[\dfrac{\partial v_{i,n}(\theta)}{\partial y_{j,n}}\big]_{i,j=1,\dots,n}$ is an $n\times n$ matrix, and $v_{i,n}(\theta)$ is defined in ((ref)). Using ((ref)), we have $$A_n^{\dagger}(\theta)=\text{diag}\Big(\sqrt{h_{1,n}(\theta)},...,\sqrt{h_{n,n}(\theta)}\Big) \big(I_n-\text{diag}(Y_n)\mathcal{F}_n(\rho)\text{diag}(Y_n)^{-1}\big),$$ where $h_{i,n}(\theta)$ is defined in ((ref)). Since $\det(I_n-\text{diag}(Y_n)\mathcal{F}_n(\rho)\text{diag}(Y_n)^{-1})=\det(I_n-\mathcal{F}_n(\rho))=\det(\mathcal{A}_n(\rho))$, the log-likelihood function can be written as
For the first term on the right-hand side of ((ref)), we need to compute $f_{ij,n}(\rho)$ (involved in $h_{i,n}(\theta)$) via $\mathcal{F}_{n}(\rho)$. Under log-SARHE and log-SMAHE models, $\mathcal{F}_{n}(\rho)=\rho M_n$ and $I_n-(I_n-\rho M_n)^{-1}$, respectively, which both can be directly computed. Under the log-SMEHE model, $\mathcal{F}_{n}(\rho)=I_n-\text{Exp}(\rho M_n)$, however, $\text{Exp}(\rho M_n)$ is an infinite series of matrices and can not be directly computed. To deal with this issue, we follow LeSage2007 to approximate $\text{Exp}(\rho M_n)$ by $\sum_{i=0}^{p}\frac{\rho^i}{i!}M_n^i$ for some integer $p>0$ (say, i.e., $p=10$) in our computation.
For the second term on the right-hand side of ((ref)), we make the following assumption:
Assumption (ref) above ensures that $\log|\det\left(\mathcal{A}_n(\rho)\right)|=\log\det\left(\mathcal{A}_n(\rho)\right)$, so our log-likelihood function becomes
Since $\log |\det\left(\mathcal{A}_n(\rho)\right)|$ is not differentiable but $\log\det\left(\mathcal{A}_n(\rho)\right)$ is, this assumption not only simplifies our computation but also brings us the convenience to establish the asymptotic results. Note that $\det(I_n-\rho M_n)=\prod_{i=1}^n(1-\rho \omega_{i,n})$, where $\omega_{i,n}$ are eigenvalues of $M_n$. Thus, for log-SARHE and log-SMAHE models, a sufficient condition for Assumption (ref) is
and the computation of $\log\det\left(\mathcal{A}_n(\rho)\right)$ relies on the result $$\log\det\left({I}_n-\rho M_n\right)=\sum_{i=1}^n\log(1-\rho\omega_{i,n}).$$ Particularly, when Assumption (ref)(i) holds, $M_n$ is row-standardized so that the upper bound $(\max_{i,n}\omega_{i,n})^{-1}$ in ((ref)) becomes one; in this case, conditions ((ref)) and ((ref)) hold if
For the log-SMEHE model, Assumption (ref) holds automatically under Assumption (ref)(i). This is because if $\text{tr}(M_n)=0$ as implied by Assumption (ref)(i), we have $$\det\left(\mathcal{A}_n(\rho)\right)=\det[\text{Exp}(\rho M_n)]=\exp[\text{tr}(\rho M_n)]=1.$$ As a result, the term $\log\det\left(\mathcal{A}_n(\rho)\right)$ in ((ref)) is absent for the log-SMEHE model, provided that Assumption (ref)(i) holds.
Based on the log-likelihood function in ((ref)), our ML estimator $\hat\theta_{ML}$ is defined as
In order to obtain the asymptotics of $\hat\theta_{ML}$, we denote
see their explicit formulas in ((ref))--((ref)) below. The following assumptions are needed to show the consistency of $\hat\theta_{ML}$.
Assumption (ref) is standard in the spatial NED literature (see, e.g., Xu2015a,Xu2015b). Assumption (ref) is regular for most model estimation theories. Assumption (ref) is similar to that in Xu2015a,Xu2015b, Debarsy2015, and Jin2018, and it is a sufficient condition for the unique global identification of $\theta_0$, that is, $\lim\inf\big({\rm E}(\log L_n(\theta_0))-{\rm E}(\log L_n(\theta_1))\big)>0$ for any $\theta_1\neq\theta_0$.
Now, we are ready to study the consistency of $\hat\theta_{ML}$.
From the proof of Theorem (ref), we have
where $\theta^*\in\Theta$ lies between $\hat{\theta}_{ML}$ and $\theta_0$, and
Hence, the consistency of $\hat\theta_{ML}$ follows from the condition that $\frac{1}{n} {\rm E}\big(\frac{\partial \log L_n(\theta_0)}{\partial\theta}\big)=o(1)$, which is equivalent to the conditions
according to the result in ((ref)) below. When $v_{i,n}\sim N(0,1)$, we have ${\rm E}\left((v_{i,n}^2-1)\log(v_{i,n}^2)\right)=2$ and ${\rm E}\left(v^2_{i,n}-1\right)=0$, so the conditions in ((ref)) hold, leading to the consistency of $\hat\theta_{ML}$ in Theorem (ref). However, when $v_{i,n}$ is not $N(0,1)$ distributed, the conditions in ((ref)) generally do not hold, except for some special cases. For example, for the log-SMEHE model, we have $$\text{tr}\big(\dot{\mathcal{A}}_n(\rho) \bar{\mathcal{A}}_n(\rho) \big)=\text{tr}\Big(\text{Exp}(-\rho M_n)\frac{\partial \text{Exp}(\rho M_n)}{\partial \rho}\Big)=\text{tr}(M_n)=0$$ under Assumption (ref)(i), so the conditions in ((ref)) hold provided that ${\rm E}\left(v^2_{i,n}\right)=1$. Therefore, $\hat\theta_{ML}$ is always consistent for the log-SMEHE model as long as ${\rm E}\left(v^2_{i,n}\right)=1$, however, it is highly possible inconsistent for other models when $v_{i,n}$ is not $N(0,1)$ distributed (see the numerical evidence in Section (ref) below).
To establish the asymptotic normality of $\hat\theta_{ML}$, we need two more additional assumptions:
Assumption (ref) is regular to ensure the existence of the asymptotic variance-covariance matrix of $\hat\theta_{ML}$. Assumption (ref) accommodates the conditions on the CLT for spatial NED processes (Jenish2012). With two assumptions above and others, we have the following result:
To establish the results in Theorems (ref)--(ref), the key is to prove that all terms involved in $\log L_n(\theta)$, $\frac{\partial \log L_n(\theta)}{\partial\theta}$, and $\frac{\partial^2 \log L_n(\theta) }{\partial \theta\partial \theta'}$ are uniformly $L_2$-NED and $L_p$-bounded, so that the LLN and CLT for spatial NED processes can be used.
Since the ML estimator is not always consistent, it motivates us to propose the GMM estimator for the log-SHE model in this subsection. The GMM estimation has been considered for the spatial mean model (see, e.g., Kelejian1999, Lee2007, Lin2010, SU, LeeYu and many others), however, it has not been studied for the spatial variance model.
To facilitate our GMM estimator, we assume that ${\rm E}(v_{i,n}^2)=1$ and ${\rm E}[(v_{i,n}^2-1)(v_{j,n}^2-1)]=0$ for $i\not=j$, and then consider two types of moment conditions. As ${\rm E}[(v_{i,n}^2-1)(v_{j,n}^2-1)]=0$ for $i\not=j$, the first type of moment condition is
where $P_n\in \mathcal{P}_n$ is an $n\times n$ deterministic matrix, $\mathcal{P}_n$ is the class of $n\times n$ deterministic matrices with zero trace, and
with $V_n^2(\theta)=\exp\{\mathcal{A}_n(\rho)\log(Y_n^2)-Z_n\gamma\}$ and ${v}^{*2}_{i,n}(\theta)=v^2_{i,n}(\theta)-1$. As ${\rm E}(v_{i,n}^2)=1$, the second type of moment condition is
where $Q_n$ is an $n\times K_{q}$ instrumental variable matrix constructed by $Z_n$ and $M_n$, and it is independent of $V_n^*(\theta_0)$ by Assumption (ref). Typically, we can take
With selected matrices $P_{1,n},...,P_{K_p,n}\in \mathcal{P}_n$ and $Q_n$, we accommodate the moment conditions in ((ref))--((ref)) to construct a $(K_p+K_q)$-dimensional estimating vector
Since ${\rm E}(R_n(\theta_0))=0$, we propose the GMM estimator as
where $\Xi$ is a $(K_p+K_q)\times (K_p+K_q)$ positive definite deterministic weighting matrix.
To study the asymptotics of $\hat{\theta}_{GMM}$, we first denote
see their explicit formulas in ((ref))--((ref)). Next, we make the following assumptions:
Assumption (ref)(i) is core to propose $\hat\theta_{GMM}$. Assumption (ref)(ii) lists some technical conditions on $v_{i,n}$. Assumption (ref) is made for the unique global identification of $\theta_0$ in the GMM framework; see the similar assumptions in Jenish2012 and Qu2015. Assumption (ref) poses some technical conditions on matrices $P_{s,n}$ and $Q_n$. It is not hard to see that the choices of $P_n$ in ((ref)) satisfy Assumption (ref)(i) if Assumption (ref)(ii)-(iii) hold, and that of $Q_n$ in ((ref)) satisfies Assumption (ref)(ii) if $\{z_{ik,n}\}_{i=1}^n$ is uniformly $L_{6+2\delta}$-bounded. Assumption (ref) is basic for the GMM estimation (Lee2007).
The following theorems establish the consistency and asymptotic normality of $\hat\theta_{GMM}$.
To compute $\hat\theta_{GMM}$ in ((ref)), one handy way is to set $\Xi=I_{K_p+K_q}$. However, the resulting $\hat\theta_{GMM}$ is suboptimal. Following standard arguments in the GMM framework, we know that the optimal choice of $\Xi$ is $\Omega_{R,n}^{-1}$, which can minimize the asymptotic covariance matrix $\Sigma_{GMM}$. Since $\Omega_{R,n}$ is unobserved, we need to estimate it by its sample version $\tilde\Omega_{R,n}$ based on an initial consistent estimator of $\theta_0$ (e.g., $\hat\theta_{GMM}$ with $\Xi=I_{K_p+K_q}$). Then, using $\tilde\Omega_{R,n}$, we are able to propose the following optimal GMM (OGMM) estimator:
The following theorem gives the asymptotic normality of $\hat\theta_{OGMM}$.
Note that $\hat\theta_{OGMM}$ not only achieves the best efficiency among all GMM estimators, but also leads to many useful tests in the next section.
Besides the ML, GMM, and OGMM estimators, one may also apply the idea of Sato2017,Sato2021 to propose the 2SML estimator for the log-SHE model. To illustrate the idea of 2SML estimation, we consider the log-SHE model with $Z_n\gamma$ satisfying ((ref)), which can be equivalently transformed into a spatial linear structure:
where $c_e={\rm E}(\log(v_{i,n}^2))+\alpha$, $\widetilde{V}_n=\log(V_n^2)-{\rm E}(\log(V_n^2))$ with ${\rm E}(\widetilde{V}_n)=0$, and ${\rm Var}(\widetilde{V}_n)=\widetilde\sigma^21_n$. Following Sato2017,Sato2021, we propose the 2SML estimator $\hat\theta_{2SML}=(\hat\alpha,\hat\beta',\hat\beta_m')'$ as follows:
The estimation in Step 1 is based on the spatial linear structure in ((ref)), and that in Step 2 is based on the optimization of log-likelihood function in ((ref)) with respect to $\alpha$. From the viewpoint of efficiency, the 2SML estimator is less efficient than our ML estimator due to the transformation in ((ref)), although both of them require the $N(0, 1)$ distributed $v_{i,n}$ (see our numerical evidence in Section (ref) below). From the viewpoint of theory, it seems inconvenient to derive the asymptotic normality of $\hat\alpha$ or $\hat\theta_{2SML}$ in the two-step estimation fashion, making the statistical inference infeasible for practitioners.
In this section, we propose the Wald, LM, and LR-type D tests to examine parameter constraints in the log-SHE model, and construct an overidentification test to detect the model adequacy.
Checking model parameter constraints is an important step in real applications. For the log-SHE model, we aim to examine the following null hypothesis on the parameter constraint:
where $\mathbb{G}(\cdot)$ is a $c_g$-dimensional constraint function, and $G=\frac{\partial \mathbb{G}(\theta_0)}{\partial \theta}$ has the rank $c_g$. Note that $\mathbb{G}(\cdot)$ can be either a linear or a nonlinear function of $\theta_0$. For example, we can test the spatial dependence in variance by setting $\mathbb{G}(\theta_0)=J_1'\theta_0$ with $J_1=(1,0,...,0)'\in\mathcal{R}^{(K+1)\times 1}$ and $c_g=1$; and we can also test the significance of $\gamma_0$ by setting $\mathbb{G}(\theta_0)=J_2'\theta_0$ with $J_2=(0, I_{K})'\in\mathcal{R}^{(K+1)\times K}$ and $c_g=K$.
To facilitate our tests for $H_0$, we first give some notations. Let $\hat{\theta}^c_{OGMM}$ be the constrained OGMM estimator under $H_0$, and $\hat G^c$, $\hat\Sigma^c_{R,n}$, and $\hat\Omega^c_{R,n}$ be the sample versions of $G$, $\Sigma_{R,n}$, and $\Omega_{R,n}$ based on $\hat{\theta}^c_{OGMM}$, respectively. Similarly, let $\hat G$, $\hat\Sigma_{R,n}$, and $\hat\Omega_{R,n}$ denote the sample versions of $G$, $\Sigma_{R,n}$, and $\Omega_{R,n}$, respectively, based on the unconstrained OGMM estimator $\hat{\theta}_{OGMM}$. Using these notations, our Wald, LM, and D tests are defined as
respectively, where $R_n(\theta)$ and $D_{n}(\theta)$ are defined as in ((ref)) and ((ref)), respectively.
Next, we show the equivalence of the aforementioned three tests and establish their limiting null distributions.
Based on Theorem (ref), we set the rejection regions of Wald, LM, and D tests at the significance level $\tau\in(0,1)$ as $\{\hat \xi_{Wald}> \upchi^2_{c_g,\tau}\}$, $\{\hat \xi_{LM}> \upchi^2_{c_g,\tau}\}$, and $\{\hat \xi_{D}> \upchi^2_{c_g,\tau}\}$, respectively, where $\upchi^2_{c_g,\tau}$ is the $\tau(\times 100)$th upper percentile of $\upchi^2(c_g)$.
Note that our tests are formed based on the OGMM estimator. If other GMM estimators are used, the corresponding Wald and LM tests have quite complex formulas, and the corresponding D test no longer has the limiting chi-square distribution; see the related illustrations in Newey1987.
Model diagnostic checking is a common step in real data analysis, however, it has not been investigated for spatial variance models. Below, we aim to provide an overidentification test to check the adequacy of the log-SHE model.
Note that if the log-SHE model is correctly specified, the moment conditions in ((ref))--((ref)) hold. Based on this observation, we can check the validity of moment conditions in ((ref))--((ref)) to examine the adequacy of the log-SHE model. This testing idea is the same as that of the overidentification test in Sargan1958 and Hansen1982; see also Lee2007, Sun2012, Dovonon2017, and Jin2019 for more studies in this aspect.
Following the above arguments, our overidentification test is defined as
where $R_n(\theta)$ is defined in ((ref)). Clearly, the test $\hat J$ is motivated by examining whether ${\rm E}(R_n(\theta_0))=0$, and a large value of $\hat J$ conveys the rejection evidence for the model adequacy. Under certain conditions, the limiting distribution of $\hat J$ is given below:
Based on the preceding theorem, we set the rejection region of the overidentification test at the significance level $\tau\in(0,1)$ as $\{\hat J> \upchi^2_{K_p+K_q-(K+1),\tau}\}$.
In this section, we assess the finite-sample performance of the proposed estimators and tests for the log-SHE model.
In this subsection, we examine the finite-sample performance of the ML estimator $\hat\theta_{ML}$ in ((ref)), GMM estimator $\hat{\theta}_{GMM}$ in ((ref)), and OGMM estimator $\hat\theta_{OGMM}$ in ((ref)). As a comparison, the 2SML estimator $\hat{\theta}_{2SML}$ in ((ref))--((ref)) is also considered.
We generate 1000 replications of sample size $n=200$ and $500$ from the following log-SHE model with the true value $\theta_0=(\rho_0,\gamma_0')'$ and exogenous variables $Z_n=(1_n,X_{n},M_n X_{n})$:
where $\rho_0=0.3$, $\gamma_0=(\alpha_0, \beta_0, \beta_{0,m})'=(1, 3, 3)'$, $X_n\in\mathcal{R}^{n\times 1}$ has its each entry generated from $N(0, 1)$, $M_n$ is the $5$-nearest neighbors matrix, and $v_{i,n}$ follows $N(0, 1)$, $MN(2,1)$, and $U(-\sqrt{3}, \sqrt{3})$ satsifying ${\rm E}(v_{i,n}^2)=1$. Here, $MN(a, b)$ denotes a mixed normal (MN) distribution, the density of which is a mixture of two normal densities of $N(a/c,b/c^2)$ and $N(-a/c,b/c^2)$ with the equal weighting probabilities, where $c=\sqrt{a^2+b}$. For each replication, we compute all of considered estimators, where $\hat{\theta}_{GMM}$ and $\hat\theta_{OGMM}$ are calculated by choosing $P_{\kappa,n} = M_n^{\kappa} - (\text{tr}(M_n^{\kappa})/n)I_n$ for ${\kappa}=1,2,3,4$ and $Q_n=(1_n,X_{n},M_n X_{n},M_n^2X_{n})$.
Based on 1000 replications, Tables (ref), (ref), and (ref) report the averaged bias and root mean squared error (RMSE) of all considered estimators, when $\mathcal{F}_n(\rho_0)$ is SAR-type, SMA-type, and SME-type, respectively. From these three tables, we can have the following findings:
Overall, when $v_{i,n}$ is normal, $\hat\theta_{ML}$ performs better than its competitors regardless of the structure of $\mathcal{F}_n(\rho_0)$. However, when $v_{i,n}$ is non-normal, $\hat\theta_{ML}$ is recommended only for the SME-type $\mathcal{F}_n(\rho_0)$, and $\hat\theta_{OGMM}$ is preferred for other types of $\mathcal{F}_n(\rho_0)$.
In this subsection, we first assess the finite-sample performance of the Wald ($\hat \xi_{Wald}$), LM ($\hat \xi_{LM}$), and D ($\hat \xi_{D}$) tests in ((ref)). We generate 1000 replications of sample size $n=200$, $500$, and $1000$ from the following log-SARHE model:
where the settings of $\alpha_0$, $\beta_0$, $\beta_{0,m}$, $X_n$, $M_n$, and $v_{i,n}$ are the same as those in ((ref)), and the value of $\rho_0$ varies from $-0.6$ to $0.6$. For each replication, we use the Wald, LM, and D tests to detect the null hypothesis of $\rho_0=0$ (i.e., $\mathbb{G}(\theta_0)=J_1'\theta_0$ with $J_1=(1,0,0,0)'$ and $c_g=1$ in ((ref))). Based on 1000 replications, Table (ref) reports the sizes of the Wald, LM and D tests at the significance level $\tau=1\%$, $5\%$, and $10\%$. From this table, we find that (i) the sizes of the D test are close to their nominal levels for small $n$, whereas those of Wald and LM tests are less accurate in this case; (ii) when the value of $n$ becomes larger, the sizes of all three tests become more accurate.
Since the sizes of Wald and LM are not accurate for small $n$, we consider the size-adjusted power of Wald, LM, and D tests for the purpose of comparison. Based on 1000 replications, Table (ref) reports the size-adjusted power of all three tests across different non-zero values of $\rho_0$ at the significance level $\tau=5\%$. From this table, we find that all three tests have a similar and satisfactory power performance, although the Wald test seems to be less powerful than the LM and D tests for smaller values of $\rho_0$ when $n=200$.
Next, we assess the finite-sample performance of the overidentification test $\hat J$ in ((ref)). We choose our null model as the log-SARHE model in ((ref)) with $\rho_0=0.3$, and use the following two alternative models to study the power of $\hat J$:
where $\rho_0^*=0.6$, $M_n^*$ is the $2$-nearest neighbors matrix, and the settings of $\alpha_0$, $\beta_0$, $\beta_{0,m}$, $X_n$, $M_n$, and $v_{i,n}$ are the same as those in the null model. For each null or alternative model, we generate 1000 replications of sample size $n=200$, $500$, and $1000$, and apply $\hat J$ to check whether the null model can fit each replication adequately. Table (ref) shows the sizes and power of $\hat J$ at the significance level $\tau=1\%$, $5\%$, and $10\%$, based on 1000 replications. From this table, we find that $\hat J$ has accurate sizes in general, and its power to detect two alternative models is satisfactory.
Overall, our simulation results imply that (i) the D test is better than the Wald and LM tests, since it has much more accurate sizes even for a small sample size; and (ii) the overidentification test is useful to detect the inadequate models. It is worthwhile to mention that we have also examined the finite-sample performance of all proposed tests for the SMA-type and SME-type models. Since the related findings are similar, they are not reported here for saving the space but are available upon request.
In this section, we study how the house characteristics affect the house selling price by revising the data in Bille2017. The data we consider contain the selling price of single-family home and the corresponding values of ten different house characteristics (see Table (ref) for their descriptions) in some adjacent regions of Lucas counties (Ohio, USA) during years 1993--1998. Figure (ref) plots the percentile graph of the selling price within 1993--1998. From this figure, we can find a clear spatial structure in the selling price. Bille2017 analyze the spatial mean structure of the selling price based on all observations during 1993--1998. To investigate more useful variance and temporal information, we study both the spatial mean structure and the spatial variance structure of the selling price below in each year, which has around 1000 observations in total.
For each year, we let $\bar{Y}_n$ denote the $n$-dimensional vector of logged selling price, where $n$ is the sample size, and the log transformation applied to the selling price (after adding one) is to decrease the skewness. Moreover, we let $X_n$ denote the $n\times 10$-dimensional matrix constructed by the ten exogenous variables, where the same log transformation (as for the variable “Price”) is applied to the exogenous variables “TLA”, “Frontage”, “Depth”, “Garagesqft”, and “Lotsize”. Based on $\bar{Y}_n$ and $X_n$, we first apply the following spatial Durbin model (SDM) to study the spatial mean structure of $\bar{Y}_n$:
where $\bar\alpha$ is the intercept, $\bar\beta$ and $\bar\beta_m$ are two coefficients, $\bar\rho$ is the spatial parameter, $\varepsilon_n$ is the vector of model errors, and $M_n$ is $10$-nearest neighbors matrix chosen by Bayesian information criterion (BIC). For model ((ref)), we apply the backward elimination procedure to select the significant parameters based on the ML estimation. That is, we start from model ((ref)) with all parameters, and then remove the most insignificant variable in a stepwise way, until all remaining variables are significant. Under this procedure, we obtain the estimation results of the fitted model ((ref)) in Table (ref). From this table, we find that four exogenous variables “TLA”, “Garagespft”, “Lotsize”, and “Age” are significant in most years, and the selling price tends to have the strongest spatial dependence in mean in 1996, since the estimates of $\bar\rho$ have an increasing trend from 1993 to 1996 but then a decreasing trend afterward.
Next, we let $Y_n$ denote the residuals of fitted model ((ref)), and study the spatial variance structure of $\bar{Y}_n$ by using the log-SHE model to fit $Y_n$. Table (ref) reports the values of BIC for different choices of log-SHE models, where each log-SHE model fitted by the backward elimination procedure as for model ((ref)) has the specification
with $M_n$ being $10$-nearest neighbors matrix. From Table (ref), we find that (i) the inclusion of exogenous variables can significantly decrease the value of BIC for each model; and (ii) the log-SARHE model (with $\mathcal{F}_{n}(\rho)=\rho M_n$ in ((ref))) has the smallest value of BIC in all years. Therefore, we only report the estimation results for the log-SARHE model in Table (ref), where all fitted log-SARHE models are adequate with the p-values of the overidentification test larger than 0.196, and their parameters are jointly significant with the p-values of Wald, LM, and D tests smaller than 0.001. From Table (ref), we find four significant exogenous variables “TLA”, “Garagespft”, “Baths”, and “Age” for the variance of the selling price in most years, and in view of the estimates of $\rho$, we also detect that the selling price has the strongest spatial dependence in variance in 1996 as observed for its spatial dependence in mean above. It is interesting to see that both spatial mean and variance models have three common significant variables “TLA”, “Garagespft”, and “Age”, whereas the variable “Lotsize” (or “Baths”) only plays an important role in the spatial mean (or variance) model.
To further explore the impact of significant exogenous variables, we consider their average total effect (ATE), average direct effect (ADE), and average indirect effect (AIE) on the mean and variance of selling price. Following LeSage2008, the ATE, ADE, and AIE of $k$-th exogenous variable on ${\rm E}(\bar{Y}_n)$ are defined as
respectively, where $\bar{Y}_n=(\bar{y}_{1,n},...,\bar{y}_{n,n})'$. However, the above way to compute three effects in ((ref)) can not be directly extended for the variance of selling price, since ${\rm E}\big(\frac{\partial y^2_{i,n}}{\partial x_{jk,n}}\big)$ does not have a closed form based on the log-SARHE model. To remedy this deficiency, we use $\frac{\partial y^2_{i,n}}{\partial x_{jk,n}}$ rather than ${\rm E}\big(\frac{\partial y^2_{i,n}}{\partial x_{jk,n}}\big)$ to form the following ATE, ADE, and AIE of $k$-th exogenous variable on ${\rm E}(Y_n^2)$:
Note that the three effects in ((ref)) are close to those based on ${\rm E}\big(\frac{\partial y^2_{i,n}}{\partial x_{jk,n}}\big)$. Table (ref) shows the effects of all significant exogenous variables on the mean and variance of selling price across six years. From Table (ref), we have the following findings:
Overall, our empirical study finds that not only the mean but also the variance of selling price have a prominent spatial structure. For the real estate investors, they should target a house with large values of “TLA” and “Garagespft” while a small value of “Age”, since this kind of house most likely has a high and stable selling price in the market. The advantage can become more substantial, if the houses in the neighborhood of the targeted house also have the same conditions of “TLA”, “Garagespft”, and “Age”.
In this paper, we first propose a new log-SHE model to study the spatial dependence in variance. The log-SHE model has a general form to include the exogenous variables and allow for the structures of SAR, SME, SMA, and many others in the variance. Next, we investigate the spatial NED property of the log-SHE model. This is the first attempt to study the probabilistic structure of the spatial variance model in the literature. More importantly, we provide a complete statistical inference procedure for the log-SHE model, whereas the existing spatial variance models lack a sound statistical inference procedure. Using the spatial NED tool kit, the asymptotics of our proposed estimators and tests are established. Simulations show that our proposed estimators and tests have a good finite-sample performance. Interesting findings on the spatial variance structure of the house selling price are detected by the log-SHE model in the application.
With slight modifications, our proposed methodology and its related technical treatment can be extended to study the following high-order and generalized log-SHE models: $Y_n=\text{diag}({H}_n)^{1/2}V_n$ with
where $\mathcal{F}_{1,n}(\rho_1)$ and $\mathcal{F}_{2,n}(\rho_2)$ are defined in a similar way as $\mathcal{F}_{n}(\rho)$ in ((ref)). A more challenging extension is to study the spatio-temporal variance model when both cross-section and time dimensions are large, although some studies for the fixed cross-section dimension have been given in Holleland2020 and Zhou2020. To fulfill this goal, some new technical developments on the LLN and CLT under spatio-temporal NED conditions are needed. We leave this interesting topic for future study.
F. Zhu's work is supported in part by the National Natural Science Foundation of China (No. 12271206) and Natural Science Foundation of Jilin Province (No. 20210101143JC). K. Zhu's work is supported in part by the General Research Fund, Research Grants Council of Hong Kong (Nos. 17304421 and 17302622).