Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
208,099 characters · 15 sections · 0 citation commands
Inference without smoothing for large panels with cross-sectional and temporal dependence
\address{Economics Department\\ London School of Economics\\ Houghton Street \\ London WC2A 2AE\\ UK} \email{[email removed]} \address{Economics Department\\ London School of Economics\\ Houghton Street \\ London WC2A 2AE\\ UK} \email{[email removed]}
Nowadays we often encounter panel data sets where both the number of individuals, $n$, and the time dimension, $T$, are large or increase without limit. Phillips and Moon $\left( 1999\right) $\ and Pesaran and Yamagata $ \left( 2008\right) $ provide some theoretical results for the parameter estimators in large panel data models, that is where both $n$\ and $T$\ tend to infinity. These works were done under the assumption of no dependence among the cross-sectional units. Yet, it is well recognized that the latter assumption is not very realistic, and there has been a surge of work on how to provide valid inferences when this type of dependence is present. The issues are closely related to Zellner's $\left( 1962\right) $ $SURE$ (Seemingly Unrelated Regression) model, be it that here both dimensions are allowed to increase without limit.
Once one accepts the possibility that the errors of the model may exhibit cross-sectional and/or temporal dependence, a key component to make valid inferences is the consistent estimation of the asymptotic covariance matrix of the estimators. For that purpose, we might proceed by explicitly assuming some specific dependence structure on the error term. In our context, this route appears to be quite cumbersome mainly for two reasons. First, it is quite difficult to specify an appropriate model in the presence of cross-sectional dependence as there are ample generic models capable to justify such a dependence. Some examples are the Spatial Autoregressive ($ SAR $) model of Cliff and Ord $\left( 1973\right) $, which has its origins in Whittle $\left( 1954\right) $, Andrews' $\left( 2005\right) $\ proposal to capture common shocks (e.g., macroeconomic, technological, legal/institutional) across observations and Pesaran's $\left( 2006\right) $ factor model. Second, in many settings it may be quite unrealistic to assume that the temporal dependence is the same for all individuals, so finding a correct specification may be infeasible as $n$ increases with no limit. Inferential properties based on parameter estimates that use a specific (wrong)\ structure, moreover, may be worse than the least squares estimates\ ($LSE$).\ The latter observation was first documented in Engle $\left( 1974\right) $\ and latter examined in Nicholls and Pagan $\left( 1977\right) $, who illustrated the adverse consequences of imposing incorrect temporal dependence assumptions on inferences, say when the practitioner assumes an $ AR\left( 1\right) $\ model instead of the true underlying $AR\left( 2\right) $\ specification.
As the task of finding an appropriate model for the dependence can be very daunting, one of our main aims in this paper is then to provide inferences in panel data not only when the error term (potentially) exhibits both temporal and cross-sectional dependence, but more importantly doing so without relying on any parametric functional form for such a dependence. Under these circumstances, one standard methodology is based on the HAC estimator, whose implementation requires the choice of one (or more) bandwidth parameter(s).\footnote{ In a time series regression model context several proposals, both in the time and frequency domain, have been employed and bootstrap applications commonly approximate the long term covariance by using a long AR polynomial (sieve method). Other methods include the use of orthogonal polynomials, see, e.g., Sun (2013) and Phillips (2005), instead of the use of Fourier sequences. All of them have in common that they require the choice of a bandwidth parameter and/or base function. Lazarus et al. (2018) provide an interesting simulation study.} While this approach is often invoked and used in the context of time series regression models, application of spatial HAC estimators is less common. The use of HAC\ estimators in spatial econometrics was advocated by Conley $\left( 1999\right) $\ and Kelejian and Prucha $\left( 2007\right) $ studied its use in Cliff-Ord type spatial models. Recently, a HAC\ estimator accounting for both temporal and spatial correlation been considered by Kim and Sun $\left( 2013\right) $. The implementation requires not only the selection of a bandwidth parameter but, more importantly, an associated measure of distance between the cross-sectional units. This explicitly assumes that there is some type of ordering among the individuals or cross-sectional units which, in contrast with\ the time dimension, is not unambiguous. Even if one accepts the existence of such an ordering, it is likely that various economics and/or geographical distance measures, each requiring their own bandwidth, may be required to encapsulate the order. For instance, simply relying on the geographic \textquotedblleft as the crow flies\textquotedblright\ distance measure for ordering is questionable as one cannot expect that two cross-sectional units located in the Rockies would behave the same as if they were in the Midwest. Clearly, a distance measure which captures the topography and other economic measures may be required.\ In addition, if we recognize that the temporal dependence may not be the same for all individuals, even the selection of a bandwidth parameter to account for the temporal dependence may become infeasible. Any cross-validation algorithm used to determine the bandwidth parameter for temporal dependence may then need to be performed for each individual.
To deal with the potential caveats of HAC estimators, we shall propose a cluster based estimator which is able to take into account both types of dependence and permits the temporal dependence to vary across individuals, see Condition $C1$ and its discussion, extending the work of Arellano $ \left( 1987\right) $\ and Driscoll and Kraay $\left( 1998\right) $ in a substantial way. While Driscoll and Kraay (1998) employed a cluster type of estimator to account for the cross-sectional dependence, they relied on the HAC methodology to accommodate the temporal dependence subjecting it to the drawback mentioned before.\ We avoid the use of the HAC\ methodology altogether. In addition, we provide a new CLT that accounts for an unknown and general temporal spatial dependence structure that permits strong spatial dependence. Our approach allows for more general dependence structure than permitted by Kim and Sun $\left( 2013\right) $ and Driscoll and Kraay $\left( 1998\right) $. Our new results can therefore be regarded as providing primitive conditions that guarantee Kim and Sun's and Driscoll and Kraay's assumption of the existence of a suitable CLT.
Our approach is based on the observation that the spectral representation of the fixed effect panel data model $\left( \text{\ref{model_1}}\right) $ is such that the errors become approximately temporally uncorrelated whilst heteroskedastic.\ It is this observation that enables us to conduct inference without any smoothing. To provide finite sample improvements for inference based on our cluster estimator, we present and examine bootstrap schemes which also do not require the choice of any bandwidth parameter , contrary to the sieve or moving block bootstrap (henceforth denoted MBB). Two bootstrap algorithms are presented, one where we assume homogeneous temporal dependence, which we shall denote as the na\"{\i} ve bootstrap, and a second one, denoted the wild bootstrap, where we allow for heterogeneous temporal dependence. Our bootstrap schemes can be viewed as wild bootstraps in the frequency domain which are shown to have good finite sample properties.
We compare our proposal to other methods that also do not require any ordering of the cross-sectional units. In particular, we consider Driscoll and Kraay's HAC\ estimator and the fixed-b asymptotic framework advocated by Vogelsang $\left( 2012\right) .$ We also consider the MBB bootstrap applied to the vector containing all the individual observations at each point in time as proposed by Gon\c{c}alves $\left( 2011\right) $.
While our estimator does permit more general spatio temporal dependence and does not require any smoothing parameters, in line with Robinson $\left( 1989\right) $, the approach examined in Section 2 precludes the presence of conditional heteroskedasticity.\ In Section 4, we examine how we can relax this by introducing a multiplicative error structure, $v_{pt}=\sigma _{1}(w_{p})\sigma _{2}\left( \varrho _{t}\right) u_{pt},$\ where $w_{p}$\ and $\varrho _{t}$\ can be functions of the fixed effects and/or variables which are correlated with the included regressors. It is worth noting that we do not need to observe these variables, as is the case when $w_{p},$ say, is the fixed effect. That is, we can allow for \textquotedblleft groupwise\textquotedblright\ heteroskedasticity and applications in development economics are commonplace, see Deaton (1996) and Greene (2018). Of particular interest, here, is the realisation that our cluster based inference is robust to the presence of heteroskedasticity that is only cross-sectional in nature (i.e., where $\sigma _{2}\left( \varrho _{t}\right) $\ is constant). In the presence of a non-constant $\sigma _{2}\left( \varrho _{t}\right) $, we propose a simple way to robustify our cluster based inference. Whereas more general forms of heteroskedasticity, where $v_{pt}=\sigma (x_{pt})u_{pt}$, can be permitted, its implementation would require the use of nonparametric methods which would require the selection of a bandwidth parameter to estimate the heteroskedasticity function.\ We shall indicate how we should proceed if this were the case.\ Finally, a benefit of our estimator is that it permits the temporal dependence to vary across individuals, which is more realistic. It is important to point out that MBB would not be valid in these settings as it depends on some type of temporal homogeneity or even stationarity.
The remainder of the paper is organized as follows. In the next section we discuss the regularity conditions for our model and describe the main results. In Section 3 we introduce our bandwidth parameter free bootstrap schemes and we demonstrate their validity. Section 4 discusses a generalisation of our model that permits (conditional) heteroskedasticity. Section 5 presents a Monte Carlo simulation experiment to shed some light on the finite sample performance of our cluster estimator and its comparison to others and we illustrate the finite sample benefits of our bootstrap schemes. In Section 6 we summarize. The proofs of our main results are given in Appendix A, which employs a series of lemmas given in Appendix B.
We shall begin by considering the panel data model
where $\beta $ is a $k\times 1$ vector of unknown parameters, $x_{pt}$ is a $ k\times 1$ vector of covariates, $\alpha _{t}$ and $\eta _{p}$ represent respectively the time and individual fixed effects and $\left\{ u_{pt}\right\} _{t\in \mathbb{Z}}$, $p\in \mathbb{N}^{+}$, are sequences of zero mean errors with heterogeneous variance $E\left( u_{pt}^{2}\right) =\sigma _{p}^{2}$, $p\in N^{+}$. We allow for general (unknown) temporal and cross-sectional dependence structures of the sequence $\left\{ u_{pt}\right\} _{t\in \mathbb{Z}}$, $p\in N^{+}$, detailed in Condition $C1$ and $\left\{ x_{pt}\right\} _{t\in \mathbb{Z}}$, $p\in N^{+},$ detailed in Condition $C2.$ Further details are provided in our discussion of these conditions below. For simplicity, we shall assume that the sequences $ \left\{ x_{pt}\right\} _{t\in \mathbb{Z}}$, $p\in N^{+}$, are mutually independent of the error term $\left\{ u_{pt}\right\} _{t\in \mathbb{Z}}$, $ p\in N^{+}$, whilst allowing for dependence of the covariates with the fixed effects $\eta _{p}$\ and/or $\alpha _{t}$.\footnote{ In fact, all that is needed is that the first conditonal moment of the error is zero and the second conditonal moment is equal to the unconditonal one.} In Section 4, we shall relax this condition allowing for heteroskedasticity.\ A straightforward extension that allows for lagged endogenous variables $ \left\{ y_{p,t-\ell }\right\} _{\ell =1}^{k_{1}},$\ as in Hidalgo and Schafgans $\left( 2017\right) $, requires the use of the instrumental variable estimator, where $\left\{ x_{p,t-\ell }\right\} _{\ell =1}^{k_{1}}$ \ provide natural instruments for $\left\{ y_{p,t-\ell }\right\} _{\ell =1}^{k_{1}}$. We have avoided this generalization as it would detract from the main contribution of the paper and it will only add some extra technicalities and/or considerations which are well known and understood when $n=1$.
Our first aim in the paper is to perform inference on the slope parameters $ \beta $ in the presence of a very general and unknown spatio-temporal dependence structure. To that end, we first need to extend a Central Limit Theorem\ provided in Phillips and Moon $\left( 1999\right) $, see also Hahn and Kuersteiner $\left( 2002\right) $. The reason for this is that in their work the sequences of random variables, say $\left\{ \psi _{pt}\right\} _{t\in \mathbb{Z}}$, $p\in \mathbb{N}^{+}$, are assumed to be independent, that is $\left\{ \psi _{pt}\right\} _{t\in \mathbb{Z}}$ and $ \left\{ \psi _{qt}\right\} _{t\in \mathbb{Z}}$ are mutually independent for any $p\neq q$, which is ruled out in our context as we permit cross-sectional dependence. Moreover, as we shall allow for \textquotedblleft strong-dependence\textquotedblright\ in our error and regressor sequences,\ we cannot use results and arguments based on any type of \textquotedblleft strong-mixing\textquotedblright\ conditions, so that results in Jenish and Prucha $\left( 2009,2012\right) $ cannot be invoked in our framework either. Our theorem also extends the results provided in Hidalgo and Schafgans $\left( 2017\right) $ by allowing the errors $u_{pt}$ to exhibit temporal dependence. A second aim of the paper is to extend the work of Driscoll and Kraay $\left( 1998\right) $ by examining, in the presence of individual and temporal fixed effects, a cluster estimator of the asymptotic covariance of the estimator of the slope parameters that does not require the ordering of the observations (in the cross-sectional dimension) or the selection of a bandwidth parameter.
The fixed effect model and the estimator for the slope parameters we consider is well known. Denoting for any generic sequence $\left\{ \varsigma _{pt}\right\} _{t=1}^{T}$,\ $p=1,...,n$, the transformation
the estimator of $\beta $ is obtained by performing least squares on the transformed model (where the individual and time effects are removed)
so that $\widehat{\beta }$ is defined as
It is obvious that we can take $Ex_{pt}=0$\ as $\widetilde{x}_{pt}$\ is invariant to additive constants, say $\mu _{t}$\ or $\nu _{p}$, to $x_{pt}$.
In this paper, we shall focus on an equivalent frequency domain formulation of $\left( \text{\ref{model_1}}\right) $ and $\left( \text{\ref{model_1R}} \right) $. It is the application of the Discrete Fourier Transform (DFT) to our model, as will become clear shortly, that plays an important role in describing and motivating the cluster estimator of the asymptotic covariance matrix of $\widehat{\beta },$ or equivalently $\tilde{\beta}$\ given in $ \left( \text{\ref{beta_fef}}\right) $\ below, and the bootstrap schemes described in Section 3.
For this purpose, we denote the DFT for generic sequences $\left\{ \varsigma _{pt}\right\} _{t=1}^{T}$, $p\geq 1$, by
and $\mathcal{J}_{\varsigma ,p}\left( \lambda _{j}\right) =\mathcal{J} _{\varsigma ,p}\left( -\lambda _{T-j}\right) $, $j=\widetilde{T}+1,..,T.$ We can then rewrite $\left( \text{\ref{model_1R}}\right) $ as
Given that our sequences $\left\{ \widetilde{\varsigma }_{pt}\right\} _{t=1}^{T},$\ $p\geq 1,$ are centered around their sample means, we can leave out the frequency $\lambda _{j}$ for $j=T$ (and 0) as $J_{\widetilde{ \varsigma },p}\left( 0\right) =\frac{1}{T^{1/2}}\sum_{t=1}^{T}\widetilde{ \varsigma }_{pt}=0$. The interesting property of $\mathcal{J}_{\widetilde{u} ,p}\left( \lambda _{j}\right) ,$ $j=1,...,T-1,$ that allows us to formulate our new cluster estimator that accounts for both types of dependence, is that it is serially uncorrelated over the Fourier frequencies for large $T,$ \ whilst possibly heteroskedastic. Based on the frequency domain formulation of our model $\left( \text{\ref{model_4freq}}\right) $, we can also compute our estimator of $\beta $ as
We introduce our regularity conditions next. To that end, and in what follows, we denote for any generic sequence $\left\{ v_{pt}\right\} _{t\in \mathbb{Z}}$, $p\in N,$
We now comment on our conditions. Conditions $C1$ and $C2$ indicate that $ \left\{ u_{pt}\right\} _{t\in \mathbb{Z}}$\ and\ $\left\{ x_{pt}\right\} _{t\in \mathbb{Z}}$, $p\in \mathbb{N}^{+}$, are linear processes that permit the usual $SAR$ (or more generally $SARMA$) model. Indeed, by definition of the $SAR$ model, with $W$\ a spatial weight matrix we have
so that $u_{p}=\sum_{q=0}^{n}\psi _{q}\left( p\right) \varepsilon _{q}$, which implies that the $SAR$\ model satisfies Condition $C1$. Unlike the SAR model, Condition $C1$ does permit the sequence $\sum_{p=1}^{n}\left\vert a_{\ell }\left( p\right) \right\vert $ to grow with $n$. One can allow the weights $a_{\ell }\left( p\right) $ to depend on the sample size \textquotedblleft $n$\textquotedblright\ as is often done in $SAR$ models with weight matrices $W$ row-normalized, but it does not add anything significant. Our conditions, therefore, appear to be weaker than those typically assumed when cross-sectional dependence is allowed while being similar to those of Lee and Robinson\ $\left( 2013\right) $. As the sequences may exhibit long memory spatial dependence, the condition of strong mixing for the spatial dependence in Jenish and Prucha $\left( 2012\right) $ is ruled out. This appears to be the case as Ibragimov and Rozanov $\left( 1978\right) $ showed; if the sequence $\left\{ \gamma _{u,pq}\left( j\right) \right\} _{j\in \mathbb{Z}}$ is not summable, the process $\left\{ u_{pt}\right\} _{t\in \mathbb{Z}},p\in \mathbb{N}^{+}$, cannot be strong-mixing. The long memory dependence also rules out that the process is Near Epoch Dependent with size greater than $1/2$, which appears to be a necessary condition for standard asymptotic CLT results.
Conditions $C1$ and $C2$ do rule out long\ memory temporal dependence on the sequences $\left\{ x_{pt}\right\} _{t\in \mathbb{Z}}$ and $ \left\{ u_{pt}\right\} _{t\in \mathbb{Z}}$ for each $p$. Even though there are several results available allowing their temporal dependence to exhibit long memory, see Robinson and Hidalgo $\left( 1997\right) $ or Hidalgo $ \left( 2003\right) $, we have decided to assume the temporal dependence of the regressors and errors to be weakly dependent to simplify the arguments. It is worth pointing out that our Conditions $C1\left( \mathbf{i}\right) $ and $C2\left( \mathbf{i}\right) $ can be relaxed to some extend to allow some type of mixing condition such as $L^{4}-$Near Epoch dependence with size greater than or equal to $2$. The latter condition is often invoked when we allow the errors to have a nonlinear type of dependence structure or if $\left( \text{\ref{model_1}}\right) $ were replaced by a nonlinear panel data model
In fact, we expect the conclusions of our results to hold under such a mixing condition\ as it has been shown in numerous situations. Conditions $C1$ and $C2$\ do permit heterogeneity in its second moments as $E\left( \xi _{pt}^{2}\mid \mathcal{V}_{p,t-1}\right) =\sigma _{\xi ,p}^{2}$\ and\ $Cov\left( \chi _{pt}\mid \Upsilon _{p,t-1}\right) =\Sigma _{\chi ,p}$. This follows from our conditions because $E\left( \xi _{pt}^{2}\mid \mathcal{V}_{p,t-1}\right) =\sum_{\ell =1}^{\infty }\left\vert a_{\ell }\left( p\right) \right\vert ^{2}$ clearly depends on $p$. Furthermore, we allow for some trending behaviour of the sequences $\left\{ x_{pt}\right\} _{t\in \mathbb{Z}}$, $p\in \mathbb{N}^{+}$ , as we allow the mean of $x_{pt}$ to depend on time.
An important consequence of Conditions $C1$ and $C2$ is that they guarantee that the covariance structure of the sequences $\left\{ u_{pt}\right\} _{t\in \mathbb{Z}}$ and $\left\{ x_{pt}\right\} _{t\in \mathbb{Z}}$, $p\in \mathbb{N}^{+}$, is multiplicative. For instance, Condition $C1$ implies that, for all $p,q\in \mathbb{N}^{+}$,
Following the spatio-temporal literature, see Cressie and Huang $\left( 1999\right) $, we can denote this covariance structure as separable. Of course, there are nonseparable covariance structures, see Gneiting $\left( 2002\right) $ and tests for separability are available, see Fuentes $ \left( 2006\right) $ or Matsuda and Yajima $\left( 2004\right) $.\ Notice that in the absence of cross-sectional dependence, $E\left( \xi _{p1}\xi _{q1}\right) =\sigma _{\xi p}^{2}\mathbf{1}\left( p=q\right) $ and $E\left( u_{pt}u_{qs}\right) =\sigma _{\xi p}^{2}\gamma _{u;pp}\left( t-s\right) \mathbf{1}\left( p=q\right) $. Here, and in what follows, $\mathbf{1}\left( A\right) $ denotes the indicator function.
Condition $C3$ assumes that the sequences $\left\{ x_{pt}\right\} _{t\in \mathbb{Z}}$ and $\left\{ u_{pt}\right\} _{t\in \mathbb{Z}}$, $p\in \mathbb{N}^{+}$ are independent, although we envisage that it can be relaxed to require only conditional independence in first and second moments. To simplify the arguments somewhat, we have preferred to keep the condition as it stands. Even though we allow long memory spatial dependence of the individual sequences, the absolute summability requirement in $\left( \text{ \ref{c3}}\right) $\ limits the combined cross-sectional dependence, that is the dependence of the sequence $\left\{ z_{pt}=u_{pt}x_{pt}\right\} _{t\in \mathbb{Z}}$, $p\in N^{+}$, is \textquotedblleft weakly spatially dependent\textquotedblright , see also Hidalgo and Schafgans $\left( 2017\right) $.\ We have adopted the convention that $\gamma _{u;pp}\left( t-s\right) =E\left( u_{pt}u_{ps}\right) /\varphi _{u}\left( p,p\right) .$ Importantly, as we assume that the errors and regressors are uncorrelated, the spectral density matrix of the sequences $\left\{ z_{pt}=:u_{pt}x_{pt}\right\} _{t\in \mathbb{Z}}$, $p\in \mathbb{N}^{+}$ is given by the convolution of the spectral density matrix of $\left\{ x_{pt}\right\} _{t\in \mathbb{Z}}$ and spectral density function of $\left\{ u_{pt}\right\} _{t\in \mathbb{Z}}$, that is
where Conditions $C1$ and $C2$ imply that $f_{p}\left( \lambda \right) $ is twice continuous differentiable. By Fuller's $\left( 1996\right) $ Theorem 3.4.1, or Corollary 3.4.1.2, the Fourier coefficients of $f_{p}\left( \lambda \right) $ are given by $\gamma _{p}\left( j\right) =\gamma _{x,p}\left( j\right) \gamma _{u,p}\left( j\right) $, \ $p\in \mathbb{N}^{+}$ , so that
With the convention that $\gamma _{u,pq}\left( 0\right) =\gamma _{x,pq}\left( 0\right) =1$, $Cov\left( z_{pt},z_{qt}\right) =\varphi \left( p,q\right) =:\varphi _{u}\left( p,q\right) \varphi _{x}\left( p,q\right) $ as defined in Condition $C3$.
Conditions $C1-C3,$ therefore, imply that the \textquotedblleft average \textquotedblright\ long-run variance of the sequences $\left\{ z_{pt}=:u_{pt}x_{pt}\right\} _{t\in \mathbb{Z}}$, $p\in \mathbb{N}^{+},$ is given by\
Observe that standard algebra yields that
or, using its spectral domain formulation,
Finally, we denote
where $\Sigma _{x}>0$ was defined in Condition $C2$.
The following theorem presents our result establishing the CLT for our slope parameter estimates in the presence of general temporal and cross-sectional dependence.
With \textsl{V} defined in $\left( \text{\ref{v_2}}\right) $, Theorem (ref) indicates that to make inferences on $\beta $, we need to provide a consistent estimator of $\Phi $. A first glance at $\left( \text{\ref{v_1}} \right) $ or $\left( \text{\ref{v_1freq}}\right) $ suggests that this might be complicated or computationally burdensome due to the general spatio-temporal dependence structure of the data. As we pointed out in the introduction, the standard approach to deal with such dependence, that is a HAC type of estimator, has various potential drawbacks in the presence of cross-sectional dependence. While choosing a bandwidth parameter associated with the cross-sectional dependence requires or induces an artificial and/or nontrivial ordering, the presence of individual heterogeneous temporal dependence (as assumed in Conditions $C1$\ and $C2)$\ would even render any cross validation method used to choose the temporal bandwidth parameter intractable.
While Kim and Sun's $\left( 2013\right) $\ approach is subject to both these criticisms, Driscoll and Kraay $\left( 1989\right) $ avoid the need to specify an ordering of individuals by introducing an HAC estimator of cross-sectional averages, so that one can consider their estimator as a hybrid between an HAC and a cluster one: they employ the HAC methodology to deal with the temporal dependence whereas they employ a cluster type of estimator to account for the cross-sectional dependence. We advocate to use an approach that does not require any ordering and/or selection of a bandwidth parameter and also permits a more general spatio-temporal dependence than allowed by either Driscoll and Kraay $\left( 1989\right) $ or Kim and Sun $\left( 2013\right) $ and permits the cross-sectional dependence to be "long-memory" which latter work ruled out. Moreover our approach permits the temporal dependence to be heterogeneous across individuals, which is more realistic.
Our approach can be regarded as a natural extension of the earlier work by Robinson $\left( 1998\right) $\ on inference without smoothing in a time series regression model context. In his case, abstracting from cross-sectional dependence
Applying his estimator to our model, would yield the estimator
where $\widehat{\gamma }_{x,p}\left( j\right) $ and $\widehat{\gamma } _{u,p}\left( j\right) $ are respectively the standard sample moment estimators of $\gamma _{x,p}\left( j\right) $ and $\gamma _{u,p}\left( j\right) $ and $I_{u,p}\left( \lambda \right) =T^{-1}\left( \sum_{t=1}^{T}u_{pt}e^{it\lambda }\right) \left( \sum_{t=1}^{T}u_{pt}e^{-it\lambda }\right) ^{\prime }$ with $I_{x,p}\left( \lambda \right) $ similarly defined. When cross-sectional dependence is allowed, the latter arguments suggest that $\left( \text{\ref{rob_1}}\right) $ is not a consistent (cluster) estimator of $\Phi $. The reason for this (see also the proof of Proposition (ref) below) is that
as expected since the first moment of $\left( \text{\ref{rob_1}}\right) $ does not capture the cross-sectional dependence. The purpose of the next section is therefore to provide a consistent \textquotedblleft cluster\textquotedblright\ estimator of $\Phi $ that accounts for the presence of cross-sectional dependence.
$\left. {}\right. $
We shall present a simple cluster estimator of $\Phi $ using the \textquotedblleft frequency\textquotedblright\ domain methodology. Obviously, there is a time domain analogue, which we shall briefly describe at the end of the section. Our cluster estimator appears to be the first one which permits both time and cross-sectional dependence and gives a formal justification of its statistical properties. Our estimator therefore becomes an extension of previous cluster estimators in the literature such as that in Arellano $\left( 1987\right) $\ (where only temporal dependence is present) or Bester, Conley and Hansen $\left( 2011\right) $ (where only cross-sectional dependence is present).
Our main motivation to propose a cluster estimator using the frequency domain methodology comes from the well known observation that for all $j\neq k$, $J_{u,p}\left( \lambda _{j}\right) $ and $J_{u,q}\left( \lambda _{k}\right) $ can be considered as being uncorrelated although possibly heteroskedastic. This observation was employed in the landmark paper by Hannan $\left( 1963\right) $ on adaptive estimation in a time series regression model. The fact that we may therefore consider $\mathcal{J}_{ \widetilde{x},p}\left( \lambda _{j}\right) \mathcal{J}_{u,p}\left( -\lambda _{j}\right) $ as a sequence of uncorrelated and heteroskedastic random variables in $j$, although not in $p$, suggests that, in a spirit similar to White's $\left( 1980\right) $ estimator, we may estimate $\Phi $ by
Based on the DFT formulation, we denote the estimator of $\Sigma _{x}$ by
The following proposition establishes the consistency of our cluster estimator for the \textquotedblleft average\textquotedblright\ long-run variance of the sequences $\left\{ z_{pt}=:u_{pt}x_{pt}\right\} _{t\in \mathbb{Z}}$, $p\in N^{+}.$
Denoting \textsl{\^{V}}$=:\widetilde{\Sigma }_{x}^{-1}\breve{\Phi}\widetilde{ \Sigma }_{x}^{-1}$, we now obtain the following corollary.
We now describe the time domain analogue estimator of $\Phi $. For that purpose, using $\sum_{t=1}^{T}e^{it\lambda _{\ell }}=0$ if $1\leq \ell \leq T-1$, we have after standard algebra that
where due to $\left( \ref{cov_u}\right) $
and $\widehat{u}_{pt}=\widetilde{y}_{pt}-\widetilde{\beta }^{\prime } \widetilde{x}_{pt}$, \ $p=1,...,n;$ $t=1,...,T$.
Our motivation to introduce bootstrap schemes emanates from findings in our Monte-Carlo experiment, which suggest that the asymptotic distribution of $ \left( nT\right) ^{1/2}$\textsl{\^{V}}$^{-1/2}\left( \widetilde{\beta } -\beta \right) $ does not appear to provide a good approximation of its finite sample distribution. In such situations, the use of the bootstrap has been advocated as it has been shown to improve the finite sample performance. The general spatio-temporal dependence inherent in our model suggests that a valid bootstrap mechanism may not to be easy to implement since one of the basic requirements for its validity is that it has to preserve the covariance structure of the data/model. Drawing analogies from the time series literature, one might be tempted to use the block bootstrap\ ($BB$) principle as it is not clear how the sieve bootstrap can be implemented under cross-sectional dependence in the absence of a clear ordering of the data. Applying a $BB$\ in both dimensions, however, would also be sensitive to the particular ordering chosen by the practitioner and be subject to the absence of weak stationarity, where the dependence structure of say $\left( x_{p_{1},t},...,x_{p_{1}+m,t}\right) ^{\prime }$\ and $\left( x_{p_{1}+1,t},...,x_{p_{1}+m+1,t}\right) ^{\prime }$\ need not be identical.
Avoiding the need to establish a particular ordering of the cross sectional units, Gon\c{c}alves $\left( 2011\right) $\ proposes to apply a moving block bootstrap ($MBB$) to the vector containing all individual observations for each\ $t$, that is it only applies a $BB$ in the time dimension.\ The $MBB$, however, does require the choice of the block size and is known to be sensitive to its choice in finite samples. In the absence of temporal dependence, the block size equals one, and the approach is similar to Hidalgo and Schafgans (2017).\
Here we propose valid bootstrap schemes with the interesting feature that they are computationally simple (there is no need to estimate, either by parametric or nonparametric methods, the time and/or cross-sectional dependence of the error term) and do not require the choice of any bandwidth parameter for its implementation, thereby avoiding any level of arbitrariness.
Both bootstrap schemes considered are in the frequency domain. We recall the DFT for generic sequences $\left\{ \varsigma _{pt}\right\} _{t=1}^{T}$, $ p\geq 1$, as $\mathcal{J}_{\varsigma ,p}\left( \lambda _{j}\right) =\frac{1}{ T^{1/2}}\sum_{t=1}^{T}\varsigma _{pt}e^{-it\lambda _{j}}$, \ \ $j=1,..., \widetilde{T}=\left[ T/2\right] $, \ $\lambda _{j}=\frac{2\pi j}{T}$, and define its periodogram
The first scheme, labelled the na\"{\i}ve bootstrap, imposes Condition $C4$\ under which the time dependence is assumed to be homogeneous across individuals. We relax this assumption in the second scheme, labelled the wild bootstrap, in line with Conditions $C1$ and $C2$.
For homogenous temporal dependence, therefore, we impose
Denoting $\sigma _{u}^{2}\left( p\right) =Eu_{pt}^{2}$ and $f_{u,p}\left( \lambda \right) $ the spectral density function of the sequence $\left\{ u_{pt}\right\} _{t\in \mathbb{Z}}$, for any $p=1,2,...$, Condition $C4$ ensures that
That is, the spectral density function normalized by the variance does not depend on $p$. This enables us to use the average periodogram of standardized residuals in constructing the valid bootstrap model. What really matters here is that the "correlation" structure is the same.
The na\"{\i}ve bootstrap, involves resampling from the residuals of the model \footnote{ See also Remark (ref) below.} and involves the following simple 3 STEPS:
The key feature of this na\"{\i}ve bootstrap, is that there is no need to choose any bandwidth parameter for its implementation. Under Condition $C4$ ,\ uniformly in $j=1,...,T-1$, we have that
and
The last displayed expressions suggest that, under Condition $C4$, we can consider $\left( \frac{1}{n}\sum_{q=1}^{n}\mathcal{I}_{\check{u} ,q}\left( \lambda _{j}\right) \right) ^{1/2}\mathcal{J}_{u^{\ast },p}\left( \lambda _{j}\right) $\ as some type of wild bootstrap in the frequency domain because under homogeneous time dependence
The following theorem is used to establish the validity of our na\"{\i}ve bootstrap scheme.
With the bootstrap cluster estimator of the asymptotic covariance, given by,
the next proposition establishes the consistency of the bootstrap cluster estimator.
The previous results can be extended to incorporate the more realistic situation where the temporal dynamics might differ by individual, as allowed by Conditions $C1$ and $C2.$ This bootstrap, labelled the wild bootstrap, merges ideas from Hidalgo $\left( 2003\right) $ and Chan and Ogden $\left( 2009\right) $. As the DFT residuals are heterogeneous whilst independent over the Fourier frequencies, it applies the wild-type bootstrap approach to the increasing dimensional vector $\left\{ \mathcal{J}_{\hat{u},p}\left( \lambda _{j}\right) \right\} _{p=1}^{n}$. It requires a modification of the above bootstrap, which primarily involves replacing STEP 2. For completeness we provide all steps:
The validity of the wild bootstrap scheme follows from the following proposition.
We conclude with the stating the validity of the standardised bootstrap statistic
In this section, we extend our model to permit general forms of heteroskedasticity. Specifically, we begin by considering
where
The error $\left\{ u_{pt}\right\} _{t\in \mathbb{Z}}$, $p\in \mathbb{N}^{+}$ satisfies the same regularity conditions given in Condition $C1$, exhibiting general spatial and temporal dependence. The sequences $\left\{ w_{p}\right\} _{p\in \mathbb{N}}$\ and $\left\{ \varrho _{t}\right\} _{t\in \mathbb{Z}}$, which can even be functions of the fixed effects, are not required to be mutually independent of the regressors\ $\left\{ \acute{x}_{pt}\right\} _{t\in \mathbb{Z}}$, $p\in N^{+}$. Without loss of generality we will normalize $\sigma _{u,p}^{2}=1$ in Condition $C1$; $\sigma _{u,p}^{2}$ is not separately identified from $ \sigma _{1}^{2}(w_{p})$ and $\sigma _{2}^{2}\left( \varrho _{t}\right) $. The error $\left\{ u_{pt}\right\} _{t\in \mathbb{Z}}$, $p\in \mathbb{N}^{+}$ is assumed to be independent of the regressors $\left\{ \acute{x} _{pt}\right\} _{t\in \mathbb{Z}}$, $p\in N^{+}$, $\left\{ w_{p}\right\} _{p\in \mathbb{N}^{+}}$ and $\left\{ \varrho _{t}\right\} _{t\in \mathbb{Z}}$ , see also footnote (ref). Here the errors $v_{pt}$\ permit conditional heteroskedasticity. This is an extension of the so-called groupwise heteroskedasticity, where observations belonging to different groups have distinct variances, see for instance Greene $\left( 2018\right) $. This type of heteroskedasticity is not uncommon in applications such as in development economics, where it has been suggested that observations within a villages or strata would have the same (conditional) variance while differences over villages or strata exist (Deaton, $1996$), that is the variance depends on some specific village variable(s).
Before we modify our Condition $C3$ to ensure we can permit this generalisation, it is useful to introduce some notation. We shall denote $ \left\{ \ddot{x}_{pt}\right\} _{t\in \mathbb{Z}},p\in \mathbb{N}^{+}$, the sequence that applies the usual transformation to remove the fixed effects, see ((ref)), to the sequence $\left\{ \acute{x}_{pt}\right\} _{t\in \mathbb{Z}},p\in \mathbb{N}^{+},$ such that
Observe that as it happens with $\tilde{x}_{pt},$ we can take $E\left( \ddot{ x}_{pt}\right) =0.$ Our new Condition $C3^{\prime }$ is given next.
The requirement given in ((ref)) limits the combined cross-sectional dependence in $v_{pt}$ and $\ddot{x}_{pt}$ ($\acute{x}_{pt}$) needed to ensure the existence of a consistent estimator of the \textquotedblleft average\textquotedblright\ long-run variance of the sequences $ \left\{ z_{pt}=:v_{pt}\ddot{x}_{pt}\right\} _{t\in \mathbb{Z}}$, $p\in \mathbb{N}^{+}$ in this framework. This is an obvious extension of our previous Condition $C3$, since our expression for $\Phi $ under our generalisation becomes
We shall now give some examples. We can allow
for given $h_{1},$ $h_{2}$\ and $h$, where $x_{pt}$ satisfies Condition $C2$ . It is clear that under the additive structure in $\mathbf{(i)}$, the transformed variables that account for the fixed effects, recall ((ref) ), satisfy
which renders this, potentially, the most straightforward setting. In this case we have
where \H{x}$_{pt}=x_{pt}\sigma _{1}\left( w_{p}\right) \sigma _{2}\left( \varrho _{t}\right) $. The behaviour of the second moments of \H{x}$_{pt}$ are essentially those of $x_{pt}$ because
With the multiplicative structure in $\left( \mathbf{ii}\right) $, it is basically the same since
where now \H{x}$_{pt}=h\left( w_{p};\varrho _{t}\right) \sigma _{1}\left( w_{p}\right) \sigma _{2}\left( \varrho _{t}\right) x_{pt}$,and $\left\vert Cov\left( \text{\H{x}}_{pt},\text{\H{x}}_{qs}\right) \right\vert \leq K\left\vert Cov\left( x_{pt},x_{qs}\right) \right\vert $ using Markov inequality.\ The same caveats mentioned in the last remark apply in this case.
We now turn to the consistent estimator of the \textquotedblleft average\textquotedblright\ long-run variance of the sequences $\left\{ z_{pt}=:v_{pt}\ddot{x}_{pt}\right\} _{t\in \mathbb{Z}}$, $p\in \mathbb{N} ^{+} $ in this framework (recognizing that we have established the necessary regularity conditions for its existence). Following a rescaling of our regressors,
for given $\sigma _{1}\left( w_{p}\right) \sigma _{2}\left( \varrho _{t}\right) ,$ our estimator for $\Phi ,$ see also $\left( \ref{v_estimate} \right) ,$ becomes
where $\hat{u}_{pt}:=\hat{v}_{pt}/\left( \hat{\sigma}_{1}\left( w_{p}\right) \hat{\sigma}_{2}(\varrho _{t}\right) )$ and $\widetilde{\dot{x}}_{pt}=\hat{ \sigma}_{1}\left( w_{p}\right) \hat{\sigma}_{2}(\varrho _{t})\widetilde{ \acute{x}}_{pt}.$ Implementation of this estimator only requires a consistent estimator of $\sigma _{2}^{2}\left( \varrho _{t}\right) $ (up to unknown scale of proportionality), and a natural estimator we can use is
The estimator for $\sigma _{1}^{2}\left( w_{p}\right) $, $\hat{\sigma} _{1}^{2}\left( w_{p}\right) $, indeed cancels out when considering the product $\mathcal{J}_{\widetilde{\dot{x}},p}\left( \lambda _{j}\right) \mathcal{J}_{\hat{u},p}\left( -\lambda _{j}\right) $, as
Moreover, this result shows that when $\sigma _{2}\left( \varrho _{t}\right) $\ is a constant, our results in Section 2 and 3 continue to hold true. That is, our estimators in previous section are robust to groupwise heteroskedasticity in the cross-sectional unit, a result supported by our Monte-Carlo simulations in Table 4 in the next section.
The intuition of the validity of this estimator comes form the standard observation that
so that
and
from the above arguments. Of course the details can be lengthy, but otherwise has been done in other contexts many times.
Our bootstrap algorithms also require some obvious and minimal change. The only adjustment to the wild bootstrap algorithm relates to the use of the robust estimator of $\Phi $ provided in ((ref)). For the na \"{\i}ve bootstrap, a straightforward modification involves the following steps
We now discuss the scenario where the conditional moment of the error term depends on the regressors $\acute{x}_{pt}=:x_{pt}$\ themselves, i.e.,
As mentioned in the introduction, this would require us to estimate the conditional expectation, $\sigma ^{2}\left( \acute{x}_{pt}\right) ,$\ nonparametrically. Several methods are available such as the Kernel regression method or sieve estimation. As this approach would require the selection of a bandwidth parameter which we set out to avoid in this paper, we do not consider this in detail although we outline how to proceed. Regardless of the approach used, we anticipate that the estimator would be pretty accurate as the number of observations in large panel data will normally be huge. For instance in a typical data set, with $T=20$\ and $ n=1000,$\ we can use $20,000$\ observations to estimate the nonparametric function. The estimator for $\Phi ,$\ see also $\left( \ref {v_estimate_robust}\right) ,$\ becomes
where $\hat{u}_{pt}:=\hat{v}_{pt}/\hat{\sigma}\left( \acute{x}_{pt}\right) $ \ and $\widetilde{\dot{x}}_{pt}=\hat{\sigma}\left( \acute{x}_{pt}\right) \widetilde{x}_{pt}.$\ For the associated na\"{\i}ve bootstrap procedure, we can proceed as above where $\widehat{\sigma _{1}(w_{p})\sigma _{2}(\varrho _{t})}$\ is replaced by $\hat{\sigma}\left( \acute{x}_{pt}\right) .$
In this section, we discuss the finite sample performance of our cluster-based inference procedure in the presence of cross-sectional and temporal dependence of unknown form. We contrast this performance with the HAC-based inference procedure proposed by Driscoll and Kraay (1989), which unlike ours, requires the choice of smoothing parameters that may be arbitrary and erroneous. We also provide evidence of the potential finite sample improvements of our frequency domain bootstrap schemes and implement the MBB time domain bootstrap to the vector containing all individual observations for each $t$. Our frequency domain approaches have the benefit that they do not rely on the choice of any smoothing parameter or require an ordering of cross-sectional units, which, as we argued before, may be arbitrary and erroneous. Another benefit of our estimator we address in our simulations is the fact that our estimator permits heterogeneity in the temporal dependence. In our simulations, we also consider a multiplicative error structure that permits groupwise heteroskedasticity and we reveal the robustness of our estimator to this setting.
In our Monte-Carlo experiments, we first consider the following data generating process
The time fixed effects $\alpha _{t}$ and individual fixed effects $\eta _{p}$ are drawn independently $(\alpha _{t}\sim IIDN(1,1)$ and $\eta _{p}\sim IIDN(1,1)$) and are held fixed across replications and without loss of generality $\beta $ is set equal to zero. The independently drawn errors and regressors are postulated to exhibit a variety of scenarios for the temporal and cross-sectional dependence that are assumed to be the same for simplicity.
To evaluate the performance of our proposed cluster estimator, we analyze the empirical size and power for testing the significance of our parameter, \ $H_{0}:\beta =0$ against $H_{A}:\beta \neq 0$, at the nominal 5% level for various pairs of $n$ and $T$ using 5,000 simulations. In addition to presenting the rejection rates based on the asymptotic distribution of the Wald statistic $nT\hat{\beta}_{FE}^{\prime }\hat{V}^{-1} \hat{\beta}_{FE},$\ with $\widehat{V}=:\widetilde{\Sigma }_{x}^{-1}\breve{ \Phi}\widetilde{\Sigma }_{x}^{-1}$\ where $\breve{\Phi}$\ is defined in ((ref)) (or equivalently the asymptotic t-test as $\beta $ is scalar), we present rejection rates based on the empirical distribution of the bootstrapped test statistic
where $\hat{\beta}_{FE}^{\ast }$ and $\hat{V}^{\ast }$\ are the bootstrapped estimators of $\beta $ and $V$ defined in Section 3. As inference based on the asymptotic distribution might not provide a good approximation to the finite sample one, this allows us to assess the finite sample improvements our bootstrap schemes may yield.
We compare the finite sample performance of our cluster-based inference procedure to the HAC based inference procedure and select the bandwidth parameter, denoted $m_{T}$, using the parametric AR(1) plug-in method suggested in Andrews $\left( 1991\right) .$\footnote{$m_{T}$ is chosen to be upward rounded integers.}\ This lag window is designed to minimize (approximately) the mean square of the standard error.\footnote{ Kiefer and Vogelsang (2002) discuss the use of HAC estimators with bandwidth equal to the sample size $(b=1)$. This bandwidth free approach does come at the cost of power relative to Andrews' popular data driven optimal bandwidth selection, see also\ Vogelsang (2012).} For the HAC based inference we provide rejection rates of the Wald statistic $nT\hat{\beta}_{FE}^{\prime } \hat{V}_{m_{T}}^{-1}\hat{\beta}_{FE}$ based on the asymptotic critical values (asy) and the critical values based on the fixed-b asymptotics (fixb) of Kiefer and Vogelsang (2005) as this is shown to lead to more reliable inference, see also Vogelsang $(2012)$. With $\widehat{V}_{m_{T}}=: \widetilde{\Sigma }_{x}^{-1}\hat{\Phi}_{m_{T}}\widetilde{\Sigma }_{x}^{-1},$ $\hat{\Phi}_{m_{T}}$ is defined as
where $\hat{z}_{pt}=\widetilde{x}_{pt}\widehat{u}_{pt}$ and $K(h)=\left( 1-\left\vert h\right\vert \right) \mathbf{1}\left( \left\vert h\right\vert \leq 1\right) $ is the Bartlett kernel. The fixed-b asymptotic distribution is non-standard and our critical values are obtained by simulation.\footnote{ Let $W_{q}(r)$ denote a $q$ dimensional vector of independent standard Wiener processes and define $\tilde{W}_{q}(r)=W_{q}\left( r\right) -rW_{q}\left( 1\right) .$ The limiting distribution of the t-test is $ W_{1}(1)/\sqrt{C_{1}}$ with $C_{q}=\frac{2}{b}\int_{0}^{1}\tilde{W}_{q}(r) \tilde{W}_{q}(r)^{\prime }dr$ $-\frac{1}{b}\int_{0}^{1-b}\left[ \tilde{W} _{q}(r+b)\tilde{W}_{q}(r)^{\prime }+\tilde{W}_{q}(r)\tilde{W} _{q}(r+b)^{\prime }\right] dr$ given the use of the Bartlet kernel, where $ b\in (0,1]$ with $m_{T}=bT$ (see Theorem 4, Vogelsang, 2012); the limiting distribution of the Wald test is $W_{q}(1)^{\prime }C_{q}^{-1}W_{q}(1).$ We obtain the critical values using 500,000 simulations.}
We also provide critical values for HAC based inference that rely on the pairs moving block bootstrap proposed by Gon\c{c}alves\ $\left( 2011\right) . $ She obtained bootstrapped samples $z_{it}^{\ast }=(y_{it}^{\ast },x_{it}^{\ast \prime })^{\prime }$ by arranging $k$ resampled blocks of $ \ell $ observations from the set of $T-\ell +1$\ overlapping blocks $\left\{ B_{1,\ell },..,B_{T-\ell +1,\ell }\right\} $\ with $B_{t,\ell }=\left\{ z_{t,n},z_{t+1,n},..,z_{t+\ell -1,n}\right\} $\ and $z_{t,n}=\left( z_{1t},..,z_{nt}\right) ^{\prime }$ in sequence (for notational simplicity $ T=k\ell $). When $\ell =1$\ this corresponds to the standard iid bootstrap on $\left\{ z_{t,n}\right\} _{t=1}^{T}$. The MMB based critical value are based on the standardized test statistic $T\left( \hat{\beta}_{FE}^{\ast }- \hat{\beta}_{FE}\right) ^{\prime }\allowbreak \left[ \QTR{sl}{\hat{V}}_{\ell }^{\ast }\right] ^{-1}\allowbreak \left( \hat{\beta}_{FE}^{\ast }-\hat{\beta} _{FE}\right) .$ Here $\QTR{sl}{\hat{V}}_{\ell }^{\ast }=\left( \widetilde{ \Sigma }_{x}^{\ast }\right) ^{-1}\breve{\Phi}_{\ell }^{\ast }\left( \widetilde{\Sigma }_{x}^{\ast }\right) ^{-1}$, $\widetilde{\Sigma } _{x}^{\ast }=\frac{1}{nT}\tsum\nolimits_{p=1}^{n}\tsum\nolimits_{t=1}^{T} \widetilde{x}_{pt}^{\ast }\widetilde{x}_{pt}^{\ast }$ and
where $\hat{s}_{nt}^{\ast }=\tsum\nolimits_{p=1}^{n}\widetilde{x}_{pt}^{\ast }\left( \tilde{y}_{pt}^{\ast }-\tilde{x}_{pt}^{\ast \prime }\hat{\beta} _{FE}^{\ast }\right) $ with $\widetilde{y}_{pt}^{\ast }=y_{pt}^{\ast }- \overline{y}_{p\cdot }^{\ast }-\overline{y}_{\cdot t}^{\ast }+\overline{ \overline{y}}_{\cdot \cdot }^{\ast }$\ and $\widetilde{x}_{pt}^{\ast }=x_{pt}^{\ast }-\overline{x}_{p\cdot }^{\ast }-\overline{x}_{\cdot t}^{\ast }+\overline{\overline{x}}_{\cdot \cdot }^{\ast }$ (see also G\"{o}tze and K \"{u}nsch, 1996). The block size used is given by the integer part of the automatic bandwidth chosen by the Andrews (1991) as proposed by Gon\c{c} alves\ $\left( 2011\right) .$
In the first set of simulations, we assume that the time dependence is homogenous among individuals $p=1,..,n$. In particular, we assume that the error and regressors are mutually independent, homoscedastic, first order auto regressive random variables with $\rho =0.7$\ or $\rho =0.9.$ The error term, therefore, takes the form
where $\eta _{pt}$ characterizes the spatial dependence inherent in the error.\footnote{ We generated the spatial data with $49+T$\ periods and take the last $T$\ periods as our sample using $0$\ as the starting value.} We consider both a weak and strong cross-sectional dependence scenarios for $u_{pt}$ $(\eta _{pt})$. To describe the cross-sectional dependence, we follow Lee and Robinson $\left( 2013\right) $ and draw random locations for individual units along a line, denoted $s=\left( s_{1},...s_{n}\right) ^{\prime }$ with $s_{p}\sim IIDU[0,n]$ for $p=1,..,n$. Using the linear time dependence representation, $\eta _{pt}=\sigma _{p}\left( \sum_{\ell =1}^{\infty }c_{\ell }\left( p\right) e_{\ell t}\right) $ with $e_{\ell t}\sim IIDN(0,1), $ we set $c_{\ell }(p)=(1+|s_{\ell }-s_{p}|_{+})^{-10}$ to permit weak dependence; $\sigma _{p}$ is such that $Var(\eta _{pt})=1$. For the strong spatial dependence setting, we use $c_{\ell }(p)=(1+|s_{\ell }-s_{p}|_{+})^{-0.7}$ instead, see also Hidalgo and Schafgans (2017). The same discussion holds for the independently drawn, strictly exogenous regressor $x_{pt}$, where, to allow for some time heterogeneity, we may, without loss of generality add $\mu _{t},$ which is independently drawn $ \left( \mu _{t}\sim IIDN(1,1)\right) .$
In Table (ref), we report the empirical size for testing the significance of $\beta $ at the 5% level of significance based on our cluster estimator of the variance of $\tilde{\beta}$ in the columns labelled $HS$ (Cluster). In addition to presenting the rejection rates based on the asymptotic critical values (asy), we report the empirical size based on the na\"{\i}ve bootstrap (nb), and the wild bootstrap (nb). The empirical size based on the HAC based inference procedure proposed by Driscoll and Kraay are reported in the columns labelled $DK$ (HAC). For the HAC based inference, we provide rejection rates based on the asymptotic critical values (asy), the critical values based on the fixed-b asymptotics (fixb) of Kiefer and Vogelsang (2005), and Gon\c{c}alves'\ $\left( 2011\right) $ MBB (mbb). We used the parametric AR(1) plug-in method suggested by Andrews $\left( 1991\right) $ to determine the window lag $ m_{T} $ and the block length $\ell $.
The results from Table (ref) reveal that our cluster based inference performs remarkably well even in the presence of strong cross sectional dependence for moderately large panels. As before, the rejection rates based on the asymptotic critical values tend to be closer to the nominal rejection rates as $n$ and $T$ increase. The finite sample performance using these asymptotic critical values does suffer, in particular, from $T$ being small, more so when the temporal dependence is stronger.\ This suggests that the cluster variance's finite sample performance, in particular, appears to require larger $T$, in order for us to be able to rely on the asymptotic critical value. Nevertheless, finite sample improvements in inference can be made using either frequency domain bootstrap schemes as rejection rates based on them are typically closer to the nominal rejection rates, with the differences typically smaller as sample sizes increase. Given that we assume the temporal dynamics to be the same for all individuals in this simulation, both bootstrap schemes are valid. The na \"{\i}ve bootstrap approach tends to perform better in the sense of providing a size closer to the nominal rejection rate.
Our cluster based inference, using the na\"{\i}ve bootstrap for small panels, suggests large improvements in size relative to HAC based inference. While the use of fixed-b asymptotic critical values for HAC based inference does indeed improve its performance, in accordance with Vogelsang, $\left( 2012\right) $, the gains in improvement in size achieved by our cluster based estimator remain significant and are larger when the temporal or spatial dependence is stronger. Our cluster based inference, however, does not necessarily perform superior to the HAC based inference that use the critical values based on Gon\c{c}alves' pairwise MBB. Her approach indeed performs very well in this setting where the temporal dependence is homogenous across individuals. As we will see in Table (ref) relaxing this assumption, which is more realistic, does reveal a marked improvement of our cluster based performance over HAC based inference using the MBB. But even in the homogenous setting, it should be noted that the MBB approach is sensitive to the chosen blocksize, and its selection here was appropriate given the imposed AR(1) temporal dependence (which is unknown in practical applications). Contrary to the MBB we do not need to choose a block size.
In Table (ref), we present the empirical power of our test for the significance of the slope when $\beta =0.1$ for a selection of $(n,T)$ pairs and compare the performance of our cluster-based inference procedure to the HAC\ based inference procedure proposed by Driscoll and Kraay as before.
The results from Table (ref) show that our cluster based inference has good power to reject $H_{0}:\beta =0$ when $\beta =0.1$ in both time dependence scenarios, even for small panels, in particular when the spatial dependence is not strong. For the reported sample sizes, the cluster based inference using the na\"{\i}ve bootstrap only showed limited size distortions. As expected its power approaches one as the sample size, and therefore the precision of our estimator, increases. This improved power performance comes about faster when the cross sectional and/or temporal dependence is lower and improved power performance appears stronger with increases in $T$ relative to $n.$ The power for our cluster based inference, using the na\"{\i}ve bootstrap, compares well with that of power of HAC based inference. Where the size-distortions for HAC\ based inference are smallest, any apparent power loss of cluster based inference disappears. Both cluster based inference and HAC\ based inference have a comparable loss of power when both spatial and temporal dependence is large.
In our second set of simulations, we allow individual heterogeneity in the time dependence of the error and the strictly exogenous regressor. The error term $u_{pt}$ is generated using various heterogeneous ARMA processes
with $L\ $denoting the lag operator, such that, e.g., $Lu_{pt}=u_{p,t-1},$ $ \rho _{1,p}$ and $\theta _{1,p}$ are individual specific AR and MA coefficients, and $(\rho _{2},\rho _{3})$ and $\left( \theta _{2},\theta _{3}\right) $ are additional non-varying higher order AR and MA coefficients. As before $\eta _{pt}$ characterizes the spatial dependence. A similar description holds for the independently drawn, strictly exogenous regressor, $x_{it},$ which is assumed to have the same spatial temporal dependence as the error for simplicity. We allow the variance of $u_{pt}$ to vary across individuals $p=1,...,n$.
We consider four heterogeneous specifications: Mixed AR(1), Mixed AR(1)/MA(1), Mixed AR(3), and Mixed AR(3)/MA(3). The individual specific parameters $\rho _{1,p}$ and $\theta _{1,p}$, where non-zero, reflect equidistant points on $[0.5,0.9].$ The full details of these heterogeneous specifications are provided at the bottom of Table (ref).
In Table (ref), we report the empirical size for testing the significance of $\beta $ in the presence of heterogenous time dependence for panels where $n=100$ and $T=$ $64,$ $128,$ and $256.$ As before, we consider both weak and strong spatial dependence scenarios. For HAC based inference we used the parametric AR(1) plug-in method suggested by Andrews $\left( 1991\right) $ again, to determine the window lag $m_{T}$ and the block length $\ell .$ A common approach, which does not recognize the temporal heterogeneity nor the higher order (autoregressive) nature of the temporal dependence under consideration.
The results in Table (ref) show that our cluster estimator of the variance is robust to the presence of individual specific time dependence. The rejection rates based on the asymptotic critical values in the heterogeneous AR(1) time dependence setting, with $\left\{ \rho _{1,p}\right\} _{p=1}^{n}$ in the range $[0.5,0.9],$ are comparable to the rejection rates in the homogenous AR(1) setting with $\rho =0.7.$\footnote{ Associated simulations considering the power to reject $H_{0}:\beta =0$ when $\beta =0.1$ in the presence of heterogenous temporal dependence show comparable results as in the homogenous time dependence setting, see also Hidalgo and Schafgans (2018).} As in the homogeneous time dependence setting, the rejection rates based on the asymptotic critical values approach the nominal rejection rate of 5% as the sample size increases. The rejection rates based on both frequency-based bootstrap schemes show that finite sample improvements in inference can be made. The improvements achieved when applying the wild bootstrap, proven to be valid in the heterogeneous time dependence scenario, are more modest than those suggested by the na\"{\i}ve bootstrap, which assumes homogeneous time dependence. Our cluster based inference reveals a similar pattern when we permit higher order heterogeneous autoregressive/moving average temporal dependence with the na\"{\i}ve bootstrap performing remarkably well again, suggesting that the na\"{\i}ve bootstrap may be robust to violations of the homogeneous time dependence such as those considered in these simulations. Whereas the wild bootstrap does perform less well then expected, in particular in the presence of strong spatial dependence, the discrepancy between the rejection rates based on the two bootstrap schemes does appear to be smaller than in the homogeneous time dependence scenario.
Importantly, our cluster based inference suggests large improvements over HAC\ based inference in these heterogeneous time dependence settings, whether we use the asymptotic, the fixed-b asymptotic critical values or base its rejection rates on the MBB. The inferior HAC based inference may be explained by the inappropriate use of a single smoothing parameter in these heterogeneous settings, as is common practice, in addition to the fact that the parametric AR(1)\ plug-in method does not account for other, and possible higher order (autoregressive) processes, than AR(1). Our cluster based inference benefits from not requiring the choice of any smoothing parameter, and is therefore not subject to this deterioration in size. Aside from the ease of implementation, the robustness of our approach to the presence of individual specific time dependence is a particularly attractive feature of our cluster robust inference.
Finally, we consider simulations that make use of our "modified" cluster based inference that permits general forms of heteroskedasticity. Here the data generating process is given by
We consider both an additive and multiplicative specification for $\acute{x} _{pt}$, in particular
Here $u_{pt}$ and $x_{pt}$ are drawn independently with weak temporal and weak cross-sectional dependence; $w_{p}$ and $\varrho _{t}$ are additional regressors where $w_{p}$ exhibits strong spatial dependence and $\varrho _{t} $ follows an AR(1) with coefficient equal to $0.7$. Without loss of generality $\beta =0$ again. Due to the presence of the multiplicative error $\sigma _{1}\left( w_{p}\right) \sigma _{2}\left( \varrho _{t}\right) u_{pt}:=v_{t}$, this setting does permit (conditional) heteroskedasticity with $Var\left( v_{pt}|x_{pt},w_{p}\right) =\sigma _{1}^{2}\left( w_{p}\right) \sigma _{2}^{2}\left( \varrho _{t}\right) $ after a normalization of the variance of $u_{pt}$ to one, for simplicity. In particular, we consider
The severity of heteroskedasticity, which we can measure using the coefficient of variation of $\sigma _{1}^{2}\left( w_{p}\right) \sigma _{2}^{2}\left( \varrho _{t}\right) ,$ increases with the values of $\delta _{1}$ and $\delta _{2}.$ The coefficient of variation is defined as the ratio of the standard deviation of $\sigma _{1}^{2}\left( w_{p}\right) \sigma _{2}^{2}\left( \varrho _{t}\right) $ to its mean. The average coefficient of variation of $\sigma _{1}^{2}\left( w_{p}\right) \sigma _{2}^{2}\left( \varrho _{t}\right) $ over our simulations with $\delta _{2}=0.5$ ranges from 42% ($\delta _{1}=0.5)$ to 260% ($\delta _{1}=2.0).$ The constant $\sigma $ in chosen in such a way that the expected variability of $\sigma _{1}^{2}\left( w_{p}\right) \sigma _{2}^{2}\left( \varrho _{t}\right) ,$ equals one for comparability across simulations.
In Table (ref), we report the empirical size for testing the significance of $\beta $ in the presence of (conditional) heteroskedasticity. The average coefficient of variation for each specification across the simulations is given in the first column. We provide two sets of simulations for our cluster based inference: first we apply the original cluster based inference, which is robust to the presence of heteroskedasticity that is only cross-sectional in nature, followed by the heteroskedasticity robust cluster based inference. In the top panel, we report the results based on the additive specification of the regressor. The multiplicative specification of the regressors is in the bottom panel. As before, we will compare the empirical size of our (robust) cluster based inference with the HAC based inference, in particular those using the MMB based critical values. As we impose an AR(1) temporal dependence, we have ensured that the use of the parametric AR(1) plug-in method suggested by Andrews $\left( 1991\right) $ to determine the window lag $m_{T}$ and the block length $\ell $ required for this approach is suitable.
The results in Table (ref) show that under the additive formulation of the regressor, the performance of the cluster based inference and robust cluster based inference (which accounts for a non-constant $\sigma _{2}\left( \rho _{t}\right) )$, are quite similar. The robust cluster based inference is required for our first three formulations, where $\sigma (w_{p},\varrho _{t})=\sigma _{1}(w_{p})\sigma _{2}(\varrho _{t}),$ whereas the final formulation permits the original cluster based inference. Compared to the results in Table (ref) (case with weak spatial and temporal dependence), the rejection rates in the presence of (conditional) heteroskedasticity are only slightly larger (in part explained by the need to use estimates for $\sigma _{2}\left( \rho _{t}\right) )$. Since $ \widetilde{\acute{x}}_{pt}=\widetilde{x}_{pt}$ under the additive formulation of the regressor, a comparison with the results in Table (ref) is more straightforward here than under the multiplicative formulation. Rejection rates that rely on the na\"{\i}ve bootstrap compare favourably with that of the HAC based inference that uses the MBB and there does not appear a serious deterioration in the performance of the (robust) cluster based inference when the severity of heteroskedasticity increases, either via $\delta _{1}$ or $\delta _{2},$ in this setting.
Under the multiplicative formulation of the regressor, the performance of the robust cluster based inference is clearly superior in our first three specifications where robust cluster based inference is required. The cluster based inference that does not account for (conditional) heteroskedasticity that is not purely cross-sectional in nature (i.e., in the presence of non-constant $\sigma _{2}\left( \rho _{t}\right) ),$ deteriorates quite quickly with $\delta _{2}$ (parameter reflecting the severity of temporal heteroskedasticity). The rejection rates that use the robust estimator of the long-run variance, (ref), are much closer to the nominal 5% rejection rates, whether we use the asymptotic critical values or the bootstrap algorithms. In fact, rejection rates that rely on the na \"{\i}ve bootstrap compare again quite well with the HAC based inference that use the MBB, which reveals the robustness of our estimator to this type of (conditional) heteroskedasticity. This is a welcome result, given that our estimator is simple to apply and does not require the choice of any smoothing parameter.
In this paper we extend the literature on inference in panel data models in the presence of both temporal and cross-sectional dependence\ of unknown form. While a standard methodology, based on the $HAC$ estimator, is often invoked and used in the context of time series regression models, in the presence of cross-sectional dependence its implementation has only recently been considered, see Kim and Sun $\left( 2013\right) $, Driscoll and Kraay $ \left( 1998\right) $ or Vogelsang $\left( 2012\right) $. To deal with various potential caveats of the $HAC$\ estimator, we propose a cluster based estimator which is able to take into account both types of dependence and allows the temporal dependence to be heterogeneous across individuals, extending the work of Arellano $\left( 1987\right) $\ and Driscoll and Kraay $\left( 1998\right) $\ in a substantial way. We provide a new CLT that accounts for an unknown and general temporal spatial dependence structure that permits strong spatial dependence. We thereby provide primitive conditions that guarantee Kim and Sun's $\left( 2013\right) ,$ Driscoll and Kraay's $\left( 1998\right) $ and Gon\c{c}alves' $\left( 2011\right) $ assumption of the existence of a suitable CLT.
Our approach is based on the insightful observation that the spectral representation of the fixed effect panel data model is such that the errors become approximately temporally uncorrelated and heteroskedastic allowing the use of a cluster estimator of the long run variance in the frequency domain. As the cluster estimator may not be reliable in small samples, and therefore may not provide a good approximation to make accurate inferences, we present and examine bootstrap schemes in the frequency domain that are also bandwidth parameter free.
Our simulation results reveal that our cluster estimator performs quite well even in the presence of strong spatial dependence. For large panels, inference based on our cluster estimator is properly sized even in the presence of heterogeneous time dependence unlike Driscoll and Kraay's HAC based inference of cross sectional averages that ignores such heterogeneity. Our bootstrap schemes provide small sample improvements, where inference that use the na\"{\i}ve bootstrap, in particular, is well sized, and reveal large improvements in size relative to HAC based inference when fixed-b asymptotic critical values are used. Improvements over MBB based inference are more limited, except in the presence of heterogeneous time dependence. We have shown the robustness of our cluster based inference to the presence of \textquotedblleft groupwise\textquotedblright\ heteroskedasticity. To enable us to adapt to the presence of \textquotedblleft groupwise\textquotedblright\ heteroskedasticity that is not purely cross-sectional in nature, a simple robust cluster based inference procedure was proposed that also does not require the selection of any smoothing parameter.
\setcounter{equation}{0} \setcounter{lemma}{0} \setcounter{remark}{0}
We first introduce some notation. For a generic function $h$, we shall abbreviate $h\left( \lambda _{j}\right) $ by $h\left( j\right) $ and for generic sequences $\left\{ \psi _{pt}\right\} _{t=1}^{T}$, $p=1,...,n$,
Using expression $\left( 10.3.12\right) $ of Brockwell and Davis $\left( 1991\right) $, we also have the useful relation
where $\mathcal{B}_{u,p}\left( j\right) =:\mathcal{B}_{u,p}\left( e^{i\lambda _{j}}\right) $, $\mathcal{B}_{x,p}\left( j\right) =:\mathcal{B} _{x,p}\left( e^{i\lambda _{j}}\right) $ and
Finally, we shall make use of the well know result
$\left. {}\right. $
For completeness, we provide the proof using the time domain estimator, $ \hat{\beta},$\ and the frequency domain estimator,\ $\tilde{\beta}$ . \newline We begin with $\hat{\beta}$. Without loss of generality assume that $x_{pt}$ is scalar. Using $\left( \ref{not}\right) $ and standard arguments, we obtain
Because the second and third terms on the right of $\left( \ref{the_3} \right) $ are handled similarly, we shall only look at the second. Now
The latter displayed expression holds true because Conditions $C1$ and $C2$ imply that
whereas Condition $C3$, see also Remark (ref), implies that \footnote{ For two nonnegative sequences $\left\{ \alpha _{p}\right\} $ and $\left\{ \beta _{p}\right\} $, $\sum \alpha _{p}\beta _{p}<C$ implies that $\sum \alpha _{p}\sum \beta _{p}=o\left( n\right) $ if $\sum \left( \alpha _{p}+\beta _{p}\right) =o\left( n\right) $.}
so that
Proceeding similarly with $\sum_{t=1}^{T}\sum_{p=1}^{n}\overline{x}_{p\cdot }u_{pt}$ and $\overline{x}_{\cdot \cdot }\sum_{t=1}^{T}\sum_{p=1}^{n}u_{pt}$ , we can conclude using $\left( \ref{the_3}\right) $ that
by Lemma (ref). From here it is standard to conclude that $\left( nT\right) ^{1/2}\left( \widehat{\beta }-\beta \right) \rightarrow _{d} \mathcal{N}\left( 0,\Sigma ^{-1}\Phi \Sigma ^{-1}\right) $.
We now show that $\left( nT\right) ^{1/2}\left( \widetilde{\beta }-\beta \right) \rightarrow _{d}\mathcal{N}\left( 0,\Sigma ^{-1}\Phi \Sigma ^{-1}\right) $. Proceeding similarly as we did above, we shall examine
The first term of $\left( \ref{the_4}\right) $ converges in distribution to $ \mathcal{N}\left( 0,\Phi \right) $ by Lemma (ref). So, to complete the proof it suffices to show that the last two terms of $\left( \ref{the_4} \right) $ are $o_{p}\left( 1\right) $. We examine the second term only, with the third term being handled similarly. By standard algebra and $\left( \ref {bartlett}\right) $, this term is
We examine the second term of $\left( \ref{a_5}\right) $ first. Using $ \left( \ref{dft_cov}\right) $, we have that its second moment is bounded by
by Lemma (ref) and $\left( \ref{the_5}\right) $. Likewise the third and fourth terms of $\left( \ref{a_5}\right) $ are $o_{p}\left( T^{-1/2}\right) $ . So to complete the proof we need to examine the first term of $\left( \ref {a_5}\right) $, whose second moment is, by $\left( \ref{the_5}\right) $ and using $\left( \sup_{p,q}\left\vert f_{x,pq}\left( j\right) \right\vert +\sup_{p,q}\left\vert f_{u,pq}\left( j\right) \right\vert \right) \leq C$, bounded by
This concludes the proof of the theorem. $\square $
$\left. {}\right. $
We begin with part $\left( \mathbf{a}\right) $. We need to show that, for any $k_{1},k_{2}=1,...,k$,
To simplify the notation we shall assume that $k=1$. Now, after observing that
we have that $\breve{\Phi}=:\breve{\Phi}_{1,1}$ is
The third term of $\left( \ref{Prop_5}\right) $ is $O_{p}\left( T^{-1}\right) $ by Lemma (ref) and $\widetilde{\beta }-\beta =O_{p}\left( \left( nT\right) ^{-1/2}\right) $. The second term of $\left( \ref{Prop_5}\right) $ is also $o_{p}\left( 1\right) $ by Cauchy-Schwarz's inequality if we show that the first term converges in probability to $\Phi $ . Since
this result holds true if we show that
and
First we examine $\left( \ref{prop_55}\right) $. We begin with the first term on the left of $\left( \ref{prop_55}\right) $, whose first moment is
using Lemma (ref), after we observe that the factor in brackets is $ n^{1/2}\mathcal{J}_{\overline{x},\cdot }\left( j\right) \mathcal{J}_{ \overline{u},\cdot }\left( -j\right) $. Using $\left( \ref{pp}\right) ,$ we conclude that the last displayed expression is $o\left( 1\right) $. Next, we observe that Lemma (ref) implies, for instance, that
The variance of the first term on the left of $\left( \ref{prop_55}\right) ,$ therefore, is bounded by
using Condition $C3$ and $\left( \ref{the_5}\right) $. Hence the first term on the left of $\left( \ref{prop_55}\right) $ is $o_{p}\left( 1\right) $. The same conclusion holds true for the second term of $\left( \ref{prop_55} \right) $.
To complete the proof of part $\left( \mathbf{a}\right) $, it remains to show $\left( \ref{prop_5}\right) $. Using $\left( \ref{bartlett}\right) $, we have that $\left( \ref{prop_5}\right) $ holds true if the following expressions $\left( \ref{prop_51}\right) -\left( \ref{prop_53}\right) $ are $ o_{p}\left( 1\right) $;
We begin by showing that $\left( \ref{prop_51}\right) $ is $o_{p}\left( 1\right) $. First, the expectation of $\left( \ref{prop_51}\right) $ is
because, by continuous differentiability of $f_{x,pq}\left( -\lambda \right) f_{u,pq}\left( \lambda \right) $, we have that
Next, because $\left( \ref{dft_cov}\right) $ implies that
standard algebra yields that the second moment of $\left( \ref{prop_51} \right) $ is $o\left( 1\right) $, when recognizing
and
since $\sum_{p_{1}=1}^{n}\varphi _{x}\left( p_{1},p_{2}\right) \varphi _{u}\left( p_{1},p_{2}\right) =O\left( 1\right) $ implies $\varphi _{x}\left( p_{1},p_{2}\right) =O\left( p_{1}^{-\alpha }\right) $ and $ \varphi _{u}\left( p_{1},p_{2}\right) =O\left( p_{1}^{-\beta }\right) $ with $\alpha +\beta >1$.
Next consider $\left( \ref{prop_52}\right) $. Because $\sup_{p}\left\vert \mathcal{B}_{x,p}\left( -j\right) \mathcal{B}_{u,p}\left( j\right) \right\vert <C$, the second moment of $\left( \ref{prop_52}\right) $ is bounded by
From here, proceeding as with $\left( \ref{prop_51}\right) $ but using Lemmas (ref) and (ref) as needed, we easily conclude that $\left( \ref{prop_52}\right) =o_{p}\left( 1\right) $ by Markov's inequality, since for instance
The proof of part $\left( \mathbf{a}\right) $ now concludes since $\left( \ref{prop_53}\right) =o_{p}\left( 1\right) $ by standard algebra and Lemmas (ref) and (ref).
Part $\left( \mathbf{b}\right) $. Because the continuous differentiability of $f_{x,p}\left( \lambda \right) $, we have that $T^{-1} \sum_{j=1}^{T}f_{x,p}\left( j\right) \rightarrow \int_{0}^{2\pi }f_{x,p}\left( \lambda \right) d\lambda =:\Sigma _{x,p}$, see Brillinger $ \left( 1981,\text{p. 15}\right) $, so we can conclude by Lemma (ref) and $\left( \ref{dif_dft}\right) $, that to finish the proof, it suffices to show that
are both $o_{p}\left( 1\right) $. However this is the case proceeding similarly as with the proof of $\left( \ref{prop_55}\right) $, so it is omitted. $\square $
$\left. {}\right. $
Because Lemma (ref) implies that $\left( nT\right) ^{-1}\sum_{p=1}^{n}\sum_{j=1}^{T-1}\mathcal{I}_{\widetilde{x},p}\left( j\right) \overset{P}{\rightarrow }\Sigma _{x}$ and abbreviating $\widehat{f} _{u}\left( j\right) =\frac{1}{n}\sum_{q=1}^{n}\mathcal{I}_{\widehat{u} ,q}\left( j\right) $, it suffices to show
We begin with part $\left( \mathbf{ii}\right) $. The left hand side of $ \left( \ref{The2_2}\right) $ is
The second (bootstrap) moment of the second term of $\left( \ref{The2_4} \right) $ is
using
By Lemma (ref) and $\left( \ref{bartlett}\right) $,
Hence it easily follows that the expected value of equation $\left( \ref {The2_5}\right) $ is $o\left( 1\right) $ and consequently the second term of $\left( \ref{The2_4}\right) $ is $o_{p^{\ast }}\left( 1\right) $, after we observe that $\left( \ref{The2_5}\right) $ is a nonnegative expression.
Turning to the first term of $\left( \ref{The2_4}\right) ,$ let us denote
Standard algebra yields that the first term of $\left( \ref{The2_4}\right) $ is
where to simplify the notation we assume that $\varphi _{x}\left( p,p\right) =\varphi _{u}\left( p,p\right) =1$ for all $p=1,...,n$ and $\phi \left( r\right) $ is the $r$th Fourier coefficient of $\mathcal{G}\left( j\right) $ . Hence the right hand side of $\left( \ref{Lem3_3*}\right) $ can now be written as
Because $\phi \left( r\right) =O\left( r^{-2}\right) $ by Conditions $C1$ and $C2,$ given the independence of the sequences of random variables $ n^{-1/2}\sum_{p=1}^{n}\chi _{pt}u_{p,t+\ell }^{\ast }$ and \ $ n^{-1/2}\sum_{p=1}^{n}\chi _{p,t+\ell }u_{pt}^{\ast }$ in $t$, to complete the proof of part $\left( \mathbf{ii}\right) $, it suffices to to show that
The second bootstrap moment of $\Lambda _{t,n}^{\ast }$ is
by standard algebra and Theorem (ref). Now, Conditions $C1$ and $C2$ imply that
Moreover, because $E\left( u_{p_{1},t+\ell }u_{q_{1},t+\ell }u_{p_{2},s+\ell }u_{q_{2},s+\ell }\right) =E\left( u_{p_{1}t}u_{q_{1}t}u_{p_{2}s}u_{q_{2}s}\right) $
because $E\left( u_{ps}u_{qr}\right) =\varphi _{u}\left( p,q\right) \gamma _{u,pq}\left( r-s\right) $, $\sum_{r,s=1}^{T}\left\vert \gamma _{u,pq}\left( r-s\right) \right\vert =O\left( T\right) $ and $\left( \ref{A}\right) $. This shows that the second moment converges to the square of the first moment, and hence $E^{\ast }\left\vert \Lambda _{t,n}^{\ast }\right\vert ^{2}-\frac{T-\ell }{T}\frac{1}{n}\sum_{p,q=1}^{n}\varphi \left( p,q\right) =o_{p}\left( 1\right) $.
Thus, it remains to show the Lindeberg's condition to complete the proof of part $\left( \mathbf{ii}\right) $. To that end, it suffices to show that
The left hand side of the last displayed expression is
which completes the proof of part $\left( \mathbf{ii}\right) $.
Next we prove part $\left( \mathbf{i}\right) $. The left side of $\left( \ref {The2_1}\right) $ is
We shall only show explicitly that the first term of $\left( \ref{a_27} \right) $ is $o_{p^{\ast }}\left( 1\right) $, the second term following similarly if not easier proceeding as with the second term of $\left( \ref {The2_4}\right) $ and Lemma (ref). Now by $\left( \ref{e}\right) $, the first term of $\left( \ref{a_27}\right) $ has second bootstrap moment given by
Because the last displayed expression is a nonnegative expression, to show that it is $o_{p}\left( 1\right) $, it suffices to show that its first moment converges to zero. To that end, we first observe that
using standard arguments and Theorem 1 under Condition\ $C4$. On the other hand, proceeding similarly as in Proposition (ref), we obtain easily that
and thus the proof of part $\left( \mathbf{i}\right) $, and thereby the theorem, is completed if
But the left hand side of the last displayed expression is
by Condition $C3,$ which completes the proof of the theorem. $\square $
$\left. {}\right. $
As with the proof of Proposition (ref), we shall assume that $k=1$. Now, after observing that
we have that $\breve{\Phi}^{\ast }$\ equals the sum of the following expressions $\left( \ref{prop2_1}\right) -\left( \ref{prop2_3} \right) $;
That $\left( \ref{prop2_3}\right) $ is $o_{p^{\ast }}\left( 1\right) $ follows straightforwardly by Theorem (ref) and Lemma (ref) and $\left( \ref{prop2_2}\right) $ is $o_{p^{\ast }}\left( 1\right) $ by Cauchy-Schwarz's inequality if we show that $\left( \ref{prop2_1}\right) $ is $o_{p^{\ast }}\left( 1\right) $. To that end, using $\left( \ref{dif_dft} \right) $ and $\left( \ref{e}\right) $, we have
Because $\widehat{\sigma }_{u,pq}=\varphi _{u}\left( p,q\right) \left( 1+o_{p}\left( 1\right) \right) $ and $\breve{\Phi}-\Phi =o_{p}\left( 1\right) $ by Proposition (ref), proceeding as in the proof of Theorem (ref) part $\left( \mathbf{i}\right) $, it suffices to examine the behaviour of
$\left( \ref{Prop2_4}\right) $ is $o_{p}\left( 1\right) $ as we now show. As it is a nonnegative sequence, it suffices to show that its first mean converges to zero. Using $\left( \ref{bartlett}\right) $ and then Lemmas (ref) and (ref), we have that its first moment is proportional to
by $\left( \ref{pp}\right) $. Because the first moment of $\left( \ref {Prop2_5}\right) $ is $o\left( 1\right) $, it then remains to show that the (bootstrap) variance of $\left( \ref{prop2_1}\right) ,$ with $\mathcal{J}_{ \widetilde{x},p}\left( j\right) $ replaced by $\mathcal{J}_{x,p}\left( j\right) ,$ converges to zero. Using $\left( \ref{e}\right) $, the (bootstrap) variance is
with Lemma (ref) guaranteeing
From here we proceed as before after noticing that $\widehat{\sigma } _{u,p_{1}p_{2}}=\varphi _{u}\left( p_{1},p_{2}\right) \left( 1+o_{p}\left( 1\right) \right) $. This completes the proof of the proposition. $\square $
$\left. {}\right. $
As with the proof of Theorem (ref), it suffices to show that
Because $\eta _{j}$ are normally distributed it suffices to show
This is the case as we now show. The left hand side of the last displayed expression is
as $\widehat{u}_{pt}-u_{pt}=\left( \widetilde{\beta }-\beta \right) x_{pt}$ and $\widetilde{\beta }-\beta =O_{p}\left( T^{-1/2}n^{-1/2}\right) $. Using $ \left( \ref{dif_dft}\right) $ and proceeding as in the proof of part $\left( \mathbf{a}\right) $ of Proposition (ref), we now have that the right hand side is
The first term converges in probability to $\Phi $, whereas the second term follows by Cauchy-Schwarz's inequality if the third term is also $ o_{p}\left( 1\right) $. But that term is $o_{p}\left( 1\right) $ proceeding as in the proof of part $\left( \mathbf{a}\right) $ of Proposition (ref) using Lemma (ref). Again observe that the expression is nonnegative. This concludes the proof. $\square \bigskip $
\setcounter{equation}{0} \setcounter{lemma}{0} \setcounter{remark}{0}
First denoting $\Upsilon _{\ell ,p}\left( j\right) =\left\{ \sum_{t=1-\ell }^{T-\ell }-\sum_{t=1}^{T}\right\} \xi _{pt}e^{-it\lambda _{j}}$ and $\Psi _{\ell ,p}\left( j\right) =\left\{ \sum_{t=1-\ell }^{T-\ell }-\sum_{t=1}^{T}\right\} \chi _{pt}e^{-it\lambda _{j}}$, we have that $ \mathrm{Y}_{u,p}\left( j\right) $ and $\mathrm{Y}_{x,p}\left( j\right) $ given in $\left( \ref{y_p}\right) $ can be decomposed as
where
The next lemma extends a Central Limit Theorem in Phillips and Moon $\left( 1999\right) $ when their independence condition fails.