Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
228,238 characters · 37 sections · 109 citation commands
Heterogeneous Autoregressions in Short $T$ Panel Data Models
\thispagestyle{empty}
\setcounter{page}{1}
\doublespacing
The importance of cross-sectional heterogeneity in panel regressions is becoming increasingly recognized in the literature. When the time dimension of the panel, $T$, is short, significant advances have been made in the case of random coefficient models with strictly exogenous regressors, for example, Chamberlain1992, Wooldridge2005, and GrahamPowell2012. A trimmed version of the mean group estimator proposed by PesaranSmith1995 can also be applied to ultra short $T$ panels when the regressors are strictly exogenous. See PesaranYang2024. In contrast, there are only a few papers that consider the estimation of heterogeneous dynamic panels when the time dimension is short.
There are some limitations to applying existing estimation methods to such heterogeneous short $T$ dynamic panels. The generalized method of moments (GMM) estimators applied after first differencing by AndersonHsiao1981,AndersonHsiao1982, ArellanoBond1991, BlundellBond1998, and ChudikPesaran2021, allow for intercept heterogeneity but not for possible heterogeneity in the autoregressive (AR) coefficients, and as shown in this paper, can lead to biased estimates and distorted inference. GuKoenker2017 and Liu2023 consider the estimation of panel AR(1) models with exogenous regressors using Bayesian techniques. While they assume random coefficients on strictly exogenous regressors, they still impose homogeneity on the AR coefficients. The mean group estimator and the hierarchical Bayesian estimator proposed by HsiaoEtal1999 allow for heterogeneity but require that $T$ is reasonably large relative to the cross section dimension, $n$.
For moderate values of $T$, analytical, Bootstrap and Jackknife bias correction approaches have also been proposed to deal with the small sample bias of the mean group and other related estimators. See, for example, PesaranZhao1999, OkuiYanagi2019 and OkuiYanagi2020. Even with bias corrections, $n$ cannot be too large compared with $T$, since a valid inference based on the asymptotic distribution often requires $ nT^{-c}\rightarrow 0,$ for some constant $c>2$. In short, none of the above approaches are appropriate and can lead to seriously biased estimates and distorted inference when $T$ is small and fixed as $n\rightarrow \infty $. Nonetheless, heterogeneity in dynamics can play an important role in many empirical studies using short $T$ panel data models. Examples include earnings dynamics studied by MeghirPistaferri2004, unemployment dynamics by BrowningCarro2014, and firm's growth by Liu2023
This paper considers a relatively simple panel AR(1) model, but allows for both individual fixed effects and heterogeneous AR coefficients, $\phi _{i}$ , where some of the individual processes, $\{y_{it}\},$ could have unit roots, $\phi _{i}=1$. We eliminate the fixed effects by first differencing, $ \Delta y_{it}=y_{it}-y_{i,t-1}$, and establish conditions under which the mean and variance of $\phi _{i}$ can be identified from the autocovariances of $\Delta y_{it}$, averaged over $i$. We show that existing GMM estimators of $E(\phi _{i})=\mu _{\phi }$ are asymptotically biased, and derive analytical expressions for their bias in simple cases. We then propose estimators for the moments of $\phi _{i}$, in particular, $E(\phi _{i})$ and $E(\phi _{i}^{2})$, using cross-sectional averages of the autocorrelation coefficients of the first differences. In terms of the estimation approach, the most relevant paper to ours is by Robinson1978, who considered a random coefficient AR(1) model without fixed effects. Assuming the \textquotedblleft usual" stationary conditions, he proposed identifying the moments of $\phi _{i}$ as functions of autocovariances of $y_{it}$.
In particular, we propose two new estimators for the moments $\theta _{s}=E(\phi _{i}^{s})$ for $s=1,2,...,T-3$. A relatively simple estimator based on autocorrelations of first differences, denoted by FDAC, and a generalized method of moments (GMM) estimator based on autocovariances of first differences, which we denote by HetroGMM. We also consider estimation of $Var(\phi _{i})=\sigma _{\phi }^{2}=\theta _{2}-\theta _{1}^{2}$, when the true value of $\sigma _{\phi }^{2}$ is not too close to zero. We do not make any assumptions about the fixed effects and allow them to have arbitrary correlations with $\phi _{i}$, but require the underlying AR(1) processes to be stationary after first differencing and assume $\phi _{i}$ and the error variances are independently distributed. It is possible to extend our analysis to higher-order panel AR processes and dynamic panels with exogenous regressors. However, these important extensions are outside the scope of the present paper.
We compare FDAC and HetroGMM estimators to a kernel-weighted likelihood estimator proposed by MavroeidisEtal2015, MSW. Assuming independently distributed Gaussian errors with cross-sectional heteroskedasticity, MSW show that the unknown distribution of heterogeneous coefficients can be identified, provided the linear operator that maps the unknown distribution to the joint distribution of data is complete (or \textquotedblleft invertible"). They provide an estimation algorithm for the parametric version of their estimator assuming the heterogeneous coefficients, including the intercepts and $\phi _{i}$, follow a multivariate normal distribution. The estimation algorithm becomes computationally very demanding if the parametric assumption about the distribution of $\phi _{i}$ is relaxed.
We investigate small sample properties of FDAC and HetroGMM estimators using Monte Carlo (MC) experiments. The simulations show that the relatively simple FDAC estimator performs better than the HetroGMM estimator uniformly across different sample sizes, and is robust to non-Gaussian errors and conditional error heteroskedasticity.
We also compare the small sample properties of the FDAC estimator of $\mu _{\phi }$ with several GMM estimators proposed in the literature for homogenous AR panels, including the popular ArellanoBond1991, AB, and BlundellBond1998, BB, estimators. We refer to these as HomoGMM estimators, to be distinguished from the HetroGMM estimator proposed in this paper. The simulation results confirm the neglected heterogeneity bias of the HomoGMM estimators, and show that the FDAC estimator of $\mu _{\phi }$ performs well for all values of $T=4,6,10$ and $n=100,1,000$ and $5,000$, so long as the underlying processes are stationary after first differencing. This is true for bias, root mean square errors, and size. Both FDAC and HetroGMM estimators are robust to the presence of unit roots and non-Gaussian errors, but can be subject to bias and size distortions if the distribution of the initial values, $y_{i0}$, significantly depart from stationarity. Similar comparative outcomes are also obtained when estimating $\sigma _{\phi }^{2}$, except that much larger sample sizes ($n$ and/or $T$) are required for reliable estimation and inference. In addition, it is important that the true value of $\sigma _{\phi }^{2}$ is not too close to the boundary value of $0$. When $n$ and $T$ are not sufficiently large, estimates of $\sigma _{\phi }^{2}$ obtained using the plugging estimator, $ \hat{\sigma}_{\phi }^{2}=\hat{\theta}_{2}-\hat{\theta}_{1}^{2}$, can be negative. This occurs with a high frequency when $n=100$, and $T=5$. The occurrence of negative estimates declines rapidly when $T=10$ and $n\geq 1,000$.
Using Monte Carlo experiments we also provide a limited comparison of the MSW and FDAC estimators of $\mu_{\phi}$, and find that in general, the MSW estimator does not have satisfactory small-sample performance under the data generating process in the paper. As the MSW estimator depends on the assumed Gaussian distribution of $\phi _{i}$, it can be severely biased with uniformly and categorically distributed $\phi _{i}$ that we consider in our MC experiments.
Finally, we provide an empirical application using five and ten yearly samples from the Panel Study of Income Dynamics (PSID) dataset over the 1976--1995 period to estimate the persistence of real earnings. To this end, we extend the basic panel AR(1) model to allow for linear trends. Following the empirical literature we report estimates for three educational categories (high school dropouts, high school graduates and college graduates) and all three categories combined. We find comparable estimates for the linear trend coefficients across sub-periods and educational categories, around 2 per cent per annum. The FDAC estimates of mean persistence $(\mu _{\phi })$ for the sub-periods 1991--1995 and 1986--1995 fall in the range of $0.570-0.734$, and tend to rise with the level of educational attainment, with college graduates showing the highest degree of persistence. No such patterns are observed for other estimates, which are around $0.3,0.9$ and $0.41$ for the AB, BB and MSW estimators, respectively. The FDAC estimates of $\sigma _{\phi }^{2}$ for all three categories combined are statistically significant and are given by $0.100$ $(0.042)$ and $0.129$ $(0.023)$ for the sub-periods 1991--1995 ($n=1,366$) and 1986--1995 $(n=1,139)$, respectively, providing further evidence of heterogeneity in real earnings persistence.
The rest of the paper is set out as follows. Section (ref) sets out the model and assumptions. Section (ref) derives the autocovariances of the first differences, $\Delta y_{it}=y_{it}-y_{i,t-1}$, and establishes conditions under which they are stationary. Section (ref) shows that the HomoGMM estimators are biased in the heterogeneous panel AR(1) model. Section (ref) establishes conditions under which the moments of $ \phi _{i}$ can be identified from the autocorrelation functions of first differences. Section (ref) proposes FDAC and HetroGMM estimators of the moments of $\phi _{i}$. The respective asymptotic distributions are also derived. Section (ref) evaluates the performance of FDAC, HetroGMM, HomoGMM, and MSW estimators by Monte Carlo simulations. Section (ref) presents the empirical application, and Section (ref) concludes. Some of the mathematical derivations, Monte Carlo evidence and additional empirical results are provided in an online supplement.
We consider the following first-order autoregressive panel data model
where the fixed effects, $\alpha _{i}$, are restricted, $\alpha _{i}=\mu _{i}\left( 1-\phi _{i}\right) $. This restriction is necessary for $y_{it}$ to have a fixed mean irrespective of whether $\phi _{i}=1$ or $\left\vert \phi _{i}\right\vert <1$. If $\alpha _{i}$ is unrestricted, a linear trend is introduced in $y_{it}$ when $\phi _{i}=1$. The restriction on $\alpha _{i} $ is not binding when $\left\vert \phi _{i}\right\vert <1$. We impose the restriction since we will be considering a mixture of processes with and without unit roots. With $\alpha _{i}=\mu _{i}(1-\phi _{i})$, ((ref)) can be written equivalently as
Suppose that $y_{it}$ is generated starting at time $t=-M_{i}\leq 0$ with the initial value, $y_{i,-M_{i}}$. We assume observations on all the $n$ units are available over the periods $t=1,2,3,...,T$, yielding a total of $ nT $ observations $\left\{ y_{i1},y_{i2},...,y_{iT}\text{, } i=1,2,...,n\right\} $. The parameters of interest are first and higher order moments of $\phi _{i}$, which we denote by $\theta _{s}=E(\phi _{i}^{s})$, $ s=1,2,...,T-2$. The key feature of our analysis is to allow for a high degree of parameter heterogeneity when $T$ is short as $n\rightarrow \infty $ . We allow $\phi _{i}$ to take any values in the non-explosive interval $ [-1+\epsilon ,1]$ for some $\epsilon > 0$, which includes the unit root case, $\phi _{i}=1$ for some of the units, but rules out a negative unit root, namely it is required that $\inf_{i}(1+\phi _{i})>0$. We are able to accommodate distributions of $\phi _{i}$ with a non-zero mass on $\phi _{i}=1 $, by basing our estimation of $\theta _{s}$ on autocorrelations of first differences, $\Delta y_{it}=y_{it}-y_{i,t-1}$, rather than the autocovariances of $y_{it}$ considered by Robinson1978. As examples, we consider a uniform distribution of $\phi _{i}$ defined over the interval $ (-1,1-\epsilon ]$ with $\epsilon \geq 0$, and a categorical distribution where $ \phi _{i}$ takes two values, $\phi _{H}$ (high) and $\phi _{L}$ (low), with probabilities $(1-\pi )$ and $\pi $, respectively. The unit root case arises when $\epsilon =0$ (for the uniform distribution), and $\phi _{H}=1$ with $ 0<\phi _{L}<1$ (for the categorical distribution). Our analysis does not allow for a negative unit root, namely when $\phi _{i}=-1$.
The key identification assumption is the stationarity of the first differences. First differencing of ((ref)) eliminates the fixed effects, $\alpha _{i}=\mu _{i}(1-\phi _{i}),$ but does not remove the effects of initial values, $y_{i,-M_{i}}$, on the first differences when $T$ is small. Under slope heterogeneity, the effects of initial values on first differences do not vanish for processes whose $\phi _{i}$ falls in the stable region, $-1<\phi _{i}<1$, unless they are all initialized at a distant past, namely only if $M_{i}\rightarrow \infty $, otherwise the realized values $y_{it}$ and/or their first differences $\left\{ \Delta y_{it}\text{, for }t=2,3,...,T\right\} $ will depend on $y_{i,-M_{i}}-\mu _{i}$. Including the observations $y_{i0}$ amongst the realizations does not resolve the problem, since we move one period backward and the distribution of $y_{i,-1}$ must still be specified and so on. For processes with unit roots, $\phi _{i}=1$, we have $\Delta y_{it}=u_{it}$, and initialization will not be an issue, at least not for the unit-root AR(1) process.
To accommodate the possible mixture of stationary and unit-root processes and achieve identification of the moments of $\phi _{i}$, we make the following assumptions regarding unit-specific parameters, $\boldsymbol{\psi } _{i}=(\mu _{i},\phi _{i},\sigma _{i}^{2})^{\prime }$, the error terms, $ u_{it}$, and the initial value deviations, $y_{i,-M_{i}}-\mu _{i}$ for $ i=1,2,...,n$, where $\sigma _{i}^{2}=Var(u_{it})$.
Assumption (ref) imposes minimal restrictions on $\mu _{i}$ or on the fixed effects $\alpha _{i}$ for $\left\vert \phi _{i}\right\vert <c<1$. But as noted earlier, to ensure that $y_{it}$ is not subject to a drift, as it is standard in the unit root literature, $\alpha _{i}$ is set to $0$ when $ \phi _{i}=1$. Assumptions (ref) and (ref) are standard in the literature on short $T$ dynamic panels. They allow for cross-sectional as well as conditional time series heteroskedasticity, such as GARCH effects, but rule out unconditional time series heteroskedasticity. Denoting the available information at time $t-1$ by $\mathcal{I}_{i,t-1}$, $ E(u_{it}^{2}\left\vert \mathcal{I}_{i,t-1}\right.)$ could be time-varying, so long as $E(u_{it}^{2})=\sigma _{i}^{2}$ as required by Assumption (ref).
Assumptions (ref) and (ref) ensure that $\Delta y_{it}$ is covariance stationary if $M_{i}\rightarrow \infty $, without requiring $ y_{it}$ to be stationary for all $n$ units in the panel.
Before setting out our approach to the identification of $\theta _{s}=E(\phi _{i}^{s}),$ we need to derive expressions for the autocovariances of $\Delta y_{it}$. Given the available data and after first differencing ((ref)), we have
Also setting $t=1$ and using ((ref)) we obtain
Iterating ((ref)) forward from $t=2$ and using the above expression for $\Delta y_{i1}$, we obtain
It is clear that in general, $\Delta y_{it}$ depends on $y_{i0}-\mu _{i}$, and Assumption (ref) is required if we are to eliminate the impact of initial values on the autocovariances of $\Delta y_{it}$. Iterating equation ((ref)) forward from $y_{i,-M_{i}}$ to $t=0$ we have
Substituting $y_{i0}-\mu _{i}$ from ((ref)) in ((ref)) now yields
where $R_{i}\left( y_{i,-M_{i}}\right) =-\phi _{i}^{M_{i}-1}(1-\phi _{i})\left( y_{i,-M_{i}}-\mu _{i}\right) $. For a fixed $T$, the remainder term, $R_{i}$, does not vanish unless $M_{i}\rightarrow \infty $. Note that under Assumption (ref) $\sup_{i}\left\vert y_{i,-M_{i}}-\mu _{i}\right\vert <C$, and $\left\vert R_{i}\left( y_{i,-M_{i}}\right) \right\vert \leq \left\vert \phi _{i}\right\vert ^{M_{i}-1}\left\vert 1-\phi _{i}\right\vert \left\vert y_{i,-M_{i}}-\mu _{i}\right\vert \leq C\left\vert \phi _{i}\right\vert ^{M_{i}-1}\left\vert 1-\phi _{i}\right\vert ,$ and $ \left\vert R_{i}\left( y_{i,-M_{i}}\right) \right\vert \rightarrow 0$, for all $i$ (irrespective of whether $\phi _{i}=1$ or $\left\vert \phi _{i}\right\vert <1$), if and only if $M_{i}\rightarrow \infty $. Under this condition
and the available first differences, $\Delta y_{it}$ for $t=2,3,...,T,$ do not depend on $y_{i0}$, and can be used to derive expressions for $\gamma _{\Delta }(h)=E\left( \Delta y_{it}\Delta y_{i,t-h}\right) $\ for $ h=0,1,...,T-2$. But first, we need to establish that these autocovariances do exist, particularly given that we allow for some $y_{it}$ processes to have unit roots. This requirement is easily established when the distribution of $\phi _{i}$ is categorical. In this case we have
where $0<\pi \leq 1$. By application of Minkowski's inequality to ((ref) ) we have (for $p\geq 1$)
where $\left\Vert \Delta y_{it}\right\Vert _{p}=E\left( \left\vert \Delta y_{it}\right\vert ^{p}\right) ^{1/p}$. By Assumption (ref) $ sup_{i,t}\left\Vert u_{it}\right\Vert _{4}<C$, and for units with $ \left\vert \phi _{i}\right\vert <c<1$, we have $\left\Vert \phi _{i}^{\ell }\right\Vert _{p}=\left\vert \phi _{i}\right\vert ^{\ell }<c^{\ell }$. Hence, conditional on $\left\vert \phi _{i}\right\vert <c<1$, we have $ \left\Vert \Delta y_{it}\right\Vert _{4}\leq \frac{2C}{1-c}<\infty $. Also by Cauchy-Schwarz inequality $\left\vert E\left( \Delta y_{it}\Delta y_{i,t-h}\right) \right\vert \leq \left[ E\left( \Delta y_{it}\right) ^{2}E\left( \Delta y_{i,t-h}\right) ^{2}\right] ^{1/2}$, and $ \sup_{i}\left\vert E\left( \Delta y_{it}\Delta y_{i,t-h}\left\vert \left\vert \phi _{i}\right\vert <c<1\right. \right) \right\vert <\infty $. In the unit root case
and overall $\left\vert \gamma _{\Delta }(h)\right\vert <\infty $, for $ h\leq T-2$. Existence of $\gamma _{\Delta }(h)$ when $\phi _{i}$ is distributed uniformly over the closed interval $[0,1]$ involves some algebra and is established in Section (ref) of the online supplement.
General expressions for the mean, variance and autocovariances of the first differences (covering unit root processes) are given in the following lemma and will be used in our subsequent analysis.
A proof is provided in Section (ref) of the online supplement.
Our identification and estimation strategy is based on matching sample estimates of autocorrelations of first differences (denoted as $\rho _{h}$) with first and higher order moments of $\phi _{i}$. But before providing the details of our proposed estimators, we first show that the HomoGMM estimators of $E(\phi _{i})$ that neglect heterogeneity of $\phi _{i}$ over $ i$ are biased even as $n\rightarrow \infty $, for any fixed $T$, and inferences based on them could be misleading. It is recognized that neglecting heterogeneity in dynamic panels can lead to biased estimates, but to the best of our knowledge, there is no formal analysis of the extent of the bias for short $T$ panels. In the case of heterogeneous dynamic panels when both $n$ and $T$ are large, PesaranSmith1995 provide expressions for asymptotic bias of fixed effects estimators.
Under homogeneity where $\phi _{i}=\phi $ for all $i$, $\phi $ can be consistently estimated by the method of moments after eliminating $\alpha _{i}$, for example by first differencing. We begin our analysis by showing the HomoGMM estimators are biased when $\phi _{i}$ are heterogeneous. The extent of the bias depends on the degree of heterogeneity. To simplify the exposition, without loss of generality, we consider the case where $T=4$, the minimum value required for identification of $\mu _{\phi }=E(\phi _{i})$ under heterogeneity established in Section (ref). For the Anderson-Hsiao (AH) estimator, $\hat{\phi}_{AH}=\left( \sum_{i=1}^{n}\Delta y_{i4}\Delta y_{i2}\right) /\left( \sum_{i=1}^{n}\Delta y_{i3}\Delta y_{i2}\right) $, and using ((ref)) for $t=4$ we have
Since $E\left( \Delta u_{i4}\Delta y_{i2}\right) =0,$ then under Assumptions (ref) to (ref) and assuming $M_{i}\rightarrow \infty $ for units with $|\phi _{i}|<1$, we have (as $n\rightarrow \infty $)
where $E\left( \Delta y_{i3}\Delta y_{i2}\right) $ and $E\left( \phi _{i}\Delta y_{i3}\Delta y_{i2}\right) $ are given by ((ref)) and ((ref)), respectively. Using these results
In the homogeneous case ($\phi _{i}=\phi $), we have $\hat{\phi} _{AH}\rightarrow _{p}\mu _{\phi }=\phi $, as expected. Under heterogeneity, $ \hat{\phi}_{AH}$ is clearly not a consistent estimator of $E(\phi _{i})$. The extent of the asymptotic bias of the AH estimator depends on the distribution of $\phi _{i}$. The expression for the neglected heterogeneity bias of the AH estimator is summarized in the following proposition.
A proof is provided in Section (ref) of the online supplement.
The asymptotic biases of the AB and BB estimators under heterogeneous slopes are derived in Section (ref) of the online supplement. The magnitude of the asymptotic bias of AH, AB and BB estimators depends on the distribution of $\phi _{i}$. For example, suppose that $\phi _{i}$ are random draws from a uniform distribution centered at $E(\phi _{i})=\mu _{\phi }>0$, with $\phi _{i}=\mu _{\phi }+v_{i}$, where $v_{i}$ $\sim IIDU[-a,a],$ $a>0$.\footnote{ To ensure that $\left\vert \phi _{i}\right\vert \leq 1$ we also require that $a\leq 1-\mu _{\phi }$.} Then
where $\delta =a/(1+\mu _{\phi })\leq \left( 1-\mu _{\phi }\right) /(1+\mu _{\phi })<1$. It is easily seen that $\hat{\phi}_{AH}-\mu _{\phi }\rightarrow 0$ with $a\rightarrow 0$. The magnitudes of the asymptotic bias of the AH estimator for $\mu _{\phi }\in \{0.4,0.5\}$ and $a=0.5$ are around $-0.186$ and $-0.204$, respectively, which are very close to the corresponding simulated bias in Tables (ref) and (ref) in the online supplement.
In the case where $\phi _{i}$ follows a categorical distribution, $\phi _{i}=\phi _{L}$ $\left( 0<\phi _{L}<1\right) $ with probability $\pi $ and $ \phi _{i}=\phi _{H}>\phi _{L}$ with probability $1-\pi $, we have
As to be expected the asymptotic bias is negative, and its magnitude depends on the degree of dispersion of $\phi _{i}$ which is given by $Var(\phi _{i})=\sigma _{\phi }^{2}=\pi (1-\pi )(\phi _{H}-\phi _{L})^{2}$. The unit root case arises for the units with $\phi _{H}=1$.
Asymptotic bias, even if small, can lead to substantial size distortions when $n$ is sufficiently large. See sub-section (ref) for Monte Carlo evidence on the bias and size distortions of AH and other HomoGMM estimators.
In this section, we formally establish conditions necessary for identification of $E(\phi _{i}^{s})$ without making any specific distributional assumptions on $\phi _{i}$. Suppose Assumptions (ref) to (ref) hold. We consider the minimum number of periods needed to consistently estimate $E(\phi _{i}^{s}),$ for $s=1,2,...,S$. Denote the $ h^{th}$-order autocorrelation coefficients of $\Delta y_{it}$ as $\rho _{h} $ given by
for $h=1,2,...$, with $\left\vert \rho _{h}\right\vert \leq 1$. Since by assumption $\phi _{i}$ and $\sigma _{i}^{2}$ are independently distributed (see part (b) of Assumption (ref)), then using the results in Lemma (ref) we have
for $h=1,2,...$, with $\left\vert \rho _{h}\right\vert \leq 1$.
Suppose that $\rho _{h}$ can be consistently estimated by the moment estimators of $E\left( \Delta y_{it}\Delta y_{i,t-h}\right) $ and $E\left[ \left( \Delta y_{it}\right) ^{2}\right] $. Then the identification condition of $E(\phi _{i}^{s})$ can be derived by the system of equations in ((ref)). For $h=1$, $2E\left( \frac{1}{1+\phi _{i}}\right) \rho _{1}=-E\left( \frac{1-\phi _{i}}{1+\phi _{i}}\right) =1-2E\left( \frac{1}{ 1+\phi _{i}}\right) $, which can be equivalently written as $2E\left( \frac{1 }{1+\phi _{i}}\right) =\frac{1}{1+\rho _{1}}$. Using this result and noting that for $h=2$, $2E\left( \frac{1}{1+\phi _{i}}\right) \rho _{2} =-2+E\left( \phi _{i}\right) +2E\left( \frac{1}{1+\phi _{i}}\right) $, we have
Similarly, for $h=3$ we have $2E\left( \frac{1}{1+\phi _{i}}\right) \rho _{3}=-E\left( 2\phi _{i}-2-\phi _{i}^{2}+\frac{2}{1+\phi _{i}}\right) $, which yields
For $h=4$, $2E\left( \frac{1}{1+\phi _{i}}\right) \rho _{4}=-E\left( 2\phi _{i}^{2}-\phi _{i}^{3}-2\phi _{i}+2-\frac{2}{1+\phi _{i}}\right)$, and upon using the results of the lower-order moments we obtain
Higher-order moments of $\phi _{i}$ can be obtained similarly. To identify the $s^{th}$ order moment of $\phi _{i}$ requires consistent estimation of $ \rho _{h}$ for $h=1,2,...,s+1$. In general, we must have $T\geq s+3$, as $ n\rightarrow \infty $ to identify $E\left( \phi _{i}^{s}\right) $.
We now turn our attention to consistent estimation of the moments of $\phi _{i}$, namely $\theta _{s}=E\left( \phi _{i}^{s}\right) $, for $s=1,2,$ and $ 3$. We consider a simple moment estimator which we refer to as the first differenced autocorrelation (FDAC) estimator, and a GMM-type estimator that we refer to as HetroGMM to be distinguished from the GMM estimators proposed in the literature for estimation of the homogeneous AR coefficient assuming $ \phi _{i}=\phi $ for all $i$.
The FDAC estimator uses the sample analogs of autocorrelations of the first differences, $\rho _{h}$ given by ((ref)), in equations ((ref)), ((ref)) and ((ref)) to obtain consistent estimators of $\theta _{s}=E(\phi _{i}^{s})$ for $s=1,2$ and 3, respectively. Specifically, using $ \left\{ \Delta y_{it}\text{, }t=2,3,...,T;i=1,2,...,n\right\} $, $\rho _{h}$ can be consistently estimated by
Then plugging these estimators in ((ref))--((ref)) we have the following FDAC estimators
and
These estimators can also be viewed as moment estimators that place equal weights on the cross-section averages, $n^{-1}\sum_{i=1}^{n}\Delta y_{it}\Delta y_{i,t-h}$, for different $t$. This makes sense since under our assumptions for each $t$, $\Delta y_{it}\Delta y_{i,t-h}$ are cross-sectionally independent with finite second-order moments, and by the law of large numbers $n^{-1}\sum_{i=1}^{n}\Delta y_{it}\Delta y_{i,t-h}\rightarrow _{p}E\left( \Delta y_{it}\Delta y_{i,t-h}\right) $, and hence $\hat{\rho}_{h,nT}\rightarrow _{p}E\left( \Delta y_{it}\Delta y_{i,t-h}\right) /E\left( \Delta y_{it}\right) ^{2}=\rho _{h}$ as $n \rightarrow \infty$. Using this result and noting that $1+\hat{\rho} _{1,nT}\rightarrow _{p}1+\rho _{1}>0$, it then readily follows that $\hat{ \theta}_{1,FDAC}\rightarrow \frac{1+2\rho _{1}+\rho _{2}}{1+\rho _{1}} =\theta _{1}=E(\phi _{i})$. Similarly, $\hat{\theta}_{s,FDAC}\rightarrow _{p}E(\phi _{i}^{s})$, for $s=2$ and $3$. \ Since $\Delta y_{it}\Delta y_{i,t-h}$ for $h=1,2,...,T-2$ have second order moments, it also follows that the convergence of $\hat{\theta}_{s,FDAC}$ to $\theta _{s}$ is in the mean squared error sense which is stronger than convergence in probability.
The FDAC estimator is a plug-in type estimator and needs not be efficient. An alternative and arguably more efficient approach would be to base the estimation of $\theta _{s}$ directly on the sample moments of $E\left( \Delta y_{it}\Delta y_{i,t-h}\right) $ and then use standard results from the GMM literature to obtain asymptotically optimum weighted moment conditions rather than equally weighted moments which might not be efficient. In practice, the differences between the two approaches could depend on the degree of heterogeneity and the sampling uncertainty associated with the GMM weights. The relative performance of FDAC and heterogeneous GMM estimators of $\theta _{s}$ will be investigated by Monte Carlo simulations.
Given ((ref)), the moment condition ((ref)) can be written equivalently as
which yields $T-3$ moment conditions for $t=4,5,...,T$, requiring that $ T\geq 4$. These moment conditions can be written more compactly as
where $M_{nt}(\theta _{1,0})=n^{-1}\sum_{i=1}^{n}m_{it}(\theta _{1,0})$, $ m_{it}(\theta _{1})=\theta _{1}h_{it}-g_{it}$,
To optimally combine the moment conditions in ((ref)) set
Then $\mathbf{M}_{nT}(\theta _{1})=\left( m_{n,4}(\theta _{1}),m_{n,5}(\theta _{1}),...,m_{n,T}(\theta _{1})\right) ^{\prime }= \mathbf{G}_{nT}-\mathbf{H}_{nT}\theta _{1}$, where $\mathbf{G}_{nT}=\frac{1}{ n}\sum_{i=1}^{n}\mathbf{g}_{iT}$ and $\mathbf{H}_{nT}=n^{-1}\sum_{i=1}^{n} \mathbf{h}_{iT}$. Using ((ref)), it readily follows that $E\left[ \mathbf{M}_{nT}(\theta _{1,0})\right] =\mathbf{0}$. The HetroGMM estimator of $\theta _{1}$ is given by
where $\mathbf{A}_{nT}$ is a $\left( T-3\right) \times (T-3)$ positive definite stochastic weight matrix, and for any $T\geq 4$, it tends to a non-stochastic positive definite matrix \thinspace $\mathbf{A}_{T}$ as $ n\rightarrow \infty $. The most efficient HetroGMM estimator is given by
where $\mathbf{A}_{T}^{\ast }=\mathbf{S}_{T}^{-1}(\theta _{1})$ is the optimal weight matrix with
Given ((ref)), $E\left( \mathbf{g}_{iT}-\theta _{1,0}\mathbf{h} _{iT}\right) =\boldsymbol{0}$, and $\mathbf{g}_{iT}-\theta _{1,0}\mathbf{h} _{iT}$ are cross-sectionally independent, then
It is difficult to derive an analytical expression for $\mathbf{S} _{T}(\theta _{1,0})$, but for a given value of $\theta _{1}$, $\mathbf{S} _{T}(\theta _{1})$ can be consistently estimated by its sample mean given by
A standard two-step GMM estimator of $\theta _{1}$ can now be obtained using $\hat{\theta}_{1,FDAC}$ given by ((ref)) as an initial estimate to consistently estimate the optimal weight matrix, $\mathbf{S}_{T}^{-1}(\theta _{1,0})$, in the first step. Substituting $\hat{\theta}_{1,FDAC}$ into ((ref)) yields the following two-step HetroGMM estimator
where
It is also possible to obtain an iterated version of the above, where $\hat{ \theta}_{1,HetroGMM}$ is used to obtain a new estimate of $\mathbf{\hat{S}} _{T}\left( \theta _{1}\right) $, namely $\mathbf{\hat{S}}_{T}(\hat{\theta} _{1,HetroGMM})$, and so on. But there seems little gain in doing so since $ \hat{\theta}_{1,HetroGMM}$ is asymptotically efficient.
The above results are summarized in the following theorem.
Our use of the FDAC estimator as an initial estimator for the two-step GMM estimator is based on the observations that FDAC exploits the stationarity properties of moments in the first differences and is based on more information as compared to the first step GMM estimator. As an example, consider the exact identified case when $T=4$. Then (see ((ref)))
as compared to $\hat{\theta}_{1,FDAC}$ given by ((ref)) which can be written equivalently
Both estimators converge to $\theta _{1,0}$ at the rate of $\sqrt{n}$, but $ \hat{\theta}_{1,FDAC}$ exploits the stationary properties of the $\left( \Delta y_{it}\right) ^{2}$ and $\Delta y_{it}\Delta y_{i,t-1}$ more effectively. Specifically, $n^{-1}\sum_{i=1}^{n}\left( \Delta y_{i4}\right) ^{2}$ and $(1/3)\sum_{t=2}^{4}\left[ n^{-1}\sum_{i=1}^{n}\left( \Delta y_{it}\right) ^{2}\right] $ converge to the same limit, but the latter makes use of $\left( \Delta y_{i2}\right) ^{2}$ and $\left( \Delta y_{i3}\right) ^{2}$ obervations as well as $\left( \Delta y_{i4}\right) ^{2}$. Similarly, $ 2n^{-1}\sum_{i=1}^{n}\Delta y_{i4}\Delta y_{i3}$ and $\sum_{t=3}^{4}\left[ n^{-1}\sum_{i=1}^{n}\Delta y_{it}\Delta y_{i,t-1}\right] $ converge to the same limit, but the latter makes use of $\Delta y_{i3}\Delta y_{i2}$ in addition to $\Delta y_{i4}\Delta y_{i3}$.
Similarly, the HetroGMM estimator of $\theta _{2}=E\left( \phi _{i}^{2}\right) $ can be obtained based on the equation below for $ t=5,6,...,T$,
Let
with $h_{2,it}=( \Delta y_{it}) ^{2}+\Delta y_{it}\Delta y_{i,t-1}$ and $g_{2,it}=\left( \Delta y_{it}\right) ^{2}+2\Delta y_{it}\Delta y_{i,t-1}+2\Delta y_{it}\Delta y_{i,t-2}+\Delta y_{it}\Delta y_{i,t-3}$. Denote $\mathbf{G}_{2,nT}=n^{-1}\sum_{i=1}^{n} \mathbf{g}_{2,iT}$, and $\mathbf{H}_{2,nT}=n^{-1}\sum_{i=1}^{n}\mathbf{h} _{2,iT}$, where $\mathbf{G}_{2,nT}$ and $\mathbf{H}_{2,nT}$ are $\left( T-4\right) \times 1$ vectors (with $T>4$). Then, the two-step HetroGMM estimator of the second moment can be derived as
where the initial estimator can be the FDAC estimator of $\theta _{2}$ given by equation ((ref)), and $\mathbf{\hat{S}}_{2,T}\left( \theta _{2}\right) =\frac{1}{n}\sum_{i=1}^{n}\left( \mathbf{g}_{2,iT\ }-\theta _{2} \mathbf{h}_{2,iT}\right) \left( \mathbf{g}_{2,iT\ }-\theta _{2}\mathbf{h} _{2,iT}\right) ^{\prime }$. Finally, the asymptotic distribution of $\hat{ \theta}_{2,HetroGMM}$ is given by
where $\theta _{2,0}$ is the true value of $\theta _{2}$, and $V_{\theta _{2}}$ can be consistently estimated by
Consider now the estimation of $\sigma _{\phi }^{2}=Var(\phi _{i})$, and recall that in terms of $\boldsymbol{\theta }=(\theta _{1},\theta _{2})^{\prime }$ we have $\sigma _{\phi }^{2}=\theta _{2}-\theta _{1}^{2}$. Therefore, a plug-in estimator of $\sigma _{\phi }^{2}$ is given by
which is an asymptotically valid estimator of $\sigma _{\phi }^{2}$ if $\hat{ \theta}_{2}-\left( \hat{\theta}_{1}\right) ^{2}>0$. This condition will be met for $n$ sufficiently large, noting that $\boldsymbol{\hat{\theta}} =\left( \hat{\theta}_{1},\hat{\theta}_{2}\right) ^{\prime }$ is a consistent estimator of $\boldsymbol{\theta }_{0}=(\theta _{1,0},\theta _{2,0})^{\prime }$. The asymptotic distribution of $\boldsymbol{\hat{\theta}}=(\hat{\theta} _{1},\hat{\theta}_{2})^{\prime }$, $\sqrt{n}(\boldsymbol{\hat{\theta}}- \boldsymbol{\theta }_{0})\rightarrow _{d}N\left( 0,\mathbf{V}_{\boldsymbol{ \theta }}\right) $ is derived in Section (ref) of the online supplement. Then using the Delta method it follows that $\sqrt{n}\left( \hat{ \sigma}_{\phi }^{2}-\sigma _{\phi ,0}^{2}\right) \rightarrow _{d}N\left( 0,V_{\sigma ^{2}}\right) $, where $\sigma _{\phi ,0}^{2}=\theta _{2,0}-\theta _{10}^{2}$ denotes the true value of $\sigma _{\phi }^{2}$, and $V_{\sigma ^{2}}=\left( -2\theta _{1,0},1\right) \mathbf{V}_{\boldsymbol{ \theta }}\left( -2\theta _{1,0},1\right) ^{\prime }$. $V_{\sigma ^{2}}$ can be consistently estimated by $\hat{V}_{\sigma }=\left( -2\hat{\theta} _{1},1\right) \mathbf{\hat{V}}_{\boldsymbol{\theta }}\left( -2\hat{\theta} _{1},1\right) ^{\prime }$, where $\mathbf{\hat{V}}_{\boldsymbol{\theta }}$ $ \ $is a consistent estimator of $\mathbf{V}_{\boldsymbol{\theta }}$ given by ((ref)) in the online supplement. However, it is important to bear in mind that the asymptotic distribution of $\hat{\sigma}_{\phi }^{2}$ is valid only in the locality of the true value of $\sigma _{\phi }^{2}$, and only if this true value is sufficiently away from the boundary value of $0$. In practice, we recommend using the plug-in estimator of $\sigma _{\phi }^{2} $ only when $n$ is large, in excess of $1,000,$ judging by the Monte Carlo evidence to be discussed below.
For each $i=1,2,...,n$, the process $\{y_{it}\}$ is generated starting at time $t=-M_{i}+1$, with the initial value $y_{i,-M_{i}}$ using
We experiment with two distributions to generate $\phi _{i}\in (-1,1]$: (a) uniform and (b) categorial. Under the former we set $\phi _{i}=\mu _{\phi }+v_{i}$, with $v_{i}\thicksim IIDU[-a,a]$. To distinguish between cases when $\left\vert \phi _{i}\right\vert <1$ for all $i$ and when $\phi_{i} \in [-1+\epsilon , 1]$ for some $\epsilon>0$ with $\phi _{i}=1$ for some $i$, we fix $a=0.5$ and consider the values of $\mu _{\phi }=0.4$ and $0.5$, with $E(\phi _{i})=\mu _{\phi }$ and $\sigma _{\phi }^{2}=a^{2}/3=0.083$. Under case (b), we generate $\phi _{i}=\phi _{H}$ (high) and $\phi _{i}=\phi _{L}$ (low) with probabilities $1-\pi $ and $\pi $, respectively. Two sets of parameter values for $(\phi _{H},\phi _{L},\pi )$ are considered: $(0.8,0.5,0.85)$ with $|\phi _{i}|<1$ for all $i$, and $(1,0.5,0.95)$ with $\phi_{i} \in (-1, 1]$ for all $i$. Then $\mu _{\phi }=E(\phi _{i})=\phi _{L}\pi +\phi _{H}(1-\pi )=0.545$ and $0.525$, and $\sigma _{\phi }^{2}=\left[ \phi _{L}^{2}\pi +\phi _{H}^{2}(1-\pi )\right] -\mu _{\phi }^{2}=0.011$ and $0.012$, respectively. The individual-specific means of $\{y_{it}\}$ are generated as $\mu _{i}=\phi _{i}+\eta _{i}$ with $\eta _{i}\thicksim IIDN(0,1)$, allowing for a non-zero correlation between $\mu _{i}$ and $\phi _{i}$.
We consider two choices when generating $\varepsilon_{it}$: Gaussian $ \varepsilon _{it}\sim IIDN(0,1)$, and non-Gaussian $\varepsilon _{it}=\left( e_{it}-2\right) /2$, with $e_{it}\sim IID\chi _{2}^{2}$, where $\chi _{2}^{2} $ is a chi-squared variate with two degrees of freedom, for all $i$ and $t$. $\left\{ h_{it}\right\} $ is generated as a $GARCH(1,1)$ process, namely $h_{it}^{2}=\sigma _{i}^{2}(1-\psi _{0}-\psi _{1})+\psi _{0}h_{i,t-1}^{2}+\psi _{1}(h_{i,t-1}\varepsilon _{i,t-1})^{2}$, with $ \sigma _{i}^{2}\sim IID\left( 0.5+0.5z_{i}^{2}\right) $ and $z_{i}\sim IIDN(0,1)$. We set $\psi _{0}=0.6$ and $\psi _{1}=0.2$, with the initial values $h_{i,-M_{i}}=\sigma _{i}$.\footnote{ Our approach also allows the coefficients of the $GARCH(1,1)$ model to be heterogeneous across $i$, so long as they are drawn from the same common distribution. But to keep the MC design simple, we are only reporting for the case where $\psi _{0}$ and $\psi _{1}$ are homogeneous.} The case where errors are conditionally homoskedastic over time is obtained as a special case by setting $\psi _{0}=\psi _{1}=0$.
We generate the initial values of $\{y_{it}\}$ as $(y_{i,-M_{i}}-\mu _{i})\sim IIDN(b,\kappa \sigma _{i}^{2})$ with $b=1$ and $\kappa =2$ for all $i$. Again the choice of $M_{i}$ is set depending on whether $\left\vert \phi _{i}\right\vert <1$ or $\phi _{i}=1$. For the former case, we set $ M_{i}=100$, which applies to all the units when $\phi _{i}$ is uniformly distributed as it is not known which $\phi _{i}=1$, and units with $\phi _{i}<1$ in the case of categorical-distributed $\phi _{i}$. For draws with $ \phi _{i}=1$ in the categorical distribution, we set $M_{i}=1$ such that $ y_{it}$ for $t=1,2,...,T$ has finite moments as $T$ is fixed in our design.
To check the robustness of the results to non-stationary initialization for $ |\phi _{i}|<1$, when the processes start from a finite date in the past, we conduct two sets of experiments, one set with $M_{i}=1$, and another set with $M_{i}=3$ for all $i$.
The estimation of the moments of $\phi _{i}$, $\mu _{\phi } = E(\phi _{i})$ and $\sigma _{\phi }^{2} = Var(\phi_{i})$, are based on $\{y_{it}^{(r)},$ for $i=1,2,...,n;t=1,2,...,T\},$ where $r$ denotes the $r^{th}$ replication of DGP in ((ref)). We carry out $2,000$ replications for the experiments that compare the small sample performances of FDAC, HetroGMM, and a number of estimators proposed in the literature for the homogeneous slope case (denoted by HomoGMM), specifically, the estimators proposed by AndersonHsiao1981,AndersonHsiao1982 (AH), ArellanoBond1991 (AB), BlundellBond1998 (BB), and the augmented Anderson-Hsiao (AAH) estimator proposed by ChudikPesaran2021, as well as the FDLS estimator due to HanPhillips2010.\footnote{ We have downloaded the codes of the AH, AB, BB, and AAH estimators from the supplementary materials of ChudikPesaran2021 using the link: \url{https://www.econ.cam.ac.uk/people-files/emeritus/mhp1/fp21/CP_AAH_paper_July_2021_codes_and_data.zip} . We are grateful to Alexander Chudik for making the codes publicly available.} For experiments that compare our proposed estimator with the MSW estimator in MavroeidisEtal2015, we use $1,000$ replications as it takes a substantial amount of time to compute the MSW estimator.\footnote{ We have downloaded the codes of the MSW estimator used in empirical applications from the supplementary materials of MavroeidisEtal2015 using the link: \url{https://drive.google.com/file/d/1hdRFpcWo3r88YV_5Kc40ur-siCYGSBDN/view?usp=sharing} . We are grateful to Yuya Sasaki for also sharing the codes of the MSW estimator used in their Monte Carlo experiments by private correspondence.} To save space, the tables summarize the results of the MC experiments are all included in the online supplement.
Bias, root mean square errors (RMSE), and size of tests of FDAC and HetroGMM estimators of $\mu _{\phi }=E(\phi _{i})$ with uniformly distributed $\phi _{i}$ are summarized in Table (ref) in the online supplement. The results with categorically distributed $\phi _{i}$ are shown in Table (ref) in the online supplement. These tables provide results for the sample size combinations $T=4,5,6,10$ and $n=100,1000,5000$, in the case of Gaussian errors without GARCH effects. The parameters of distributions are chosen to distinguish between cases where $|\phi _{i}|<1$ and $ \phi _{i} \in (-1, 1]$, with the related results displayed in the left and right panels of the tables, respectively.
In line with our theoretical results, both FDAC and HetroGMM estimators offer reliable estimates for $\mu _{\phi }$ in the case of heterogeneous short $T$ panels under both uniform and categorical distributions. The categorical distribution yields marginally lower RMSEs, which is largely due to the fact that $\sigma _{\phi }^{2}$ is much smaller under the categorical distribution around $0.012$, as compared to $0.083$ under the uniform distribution. More importantly, the magnitudes of bias, RMSE, and size are very similar irrespective of whether $\left\vert \phi _{i}\right\vert <1$ or $\phi _{i} \in (-1,1]$. This result holds even if a fixed proportion of units have unit roots, as is the case with the categorical distribution where $ \phi _{i}=1$ in the case of 5 per cent of all units in the sample. The empirical power functions for FDAC and HetroGMM estimators of $\mu _{\phi }$ are displayed in Figure (ref) of the online supplement for the uniformly distributed AR coefficients with $\phi _{i} \in (-1,1]$ in the baseline case (Gaussian errors and no GARCH effects). The power functions for the other experiments are very similar and can be obtained from the authors upon request.
Compared with the HetroGMM estimator, the FDAC estimator has uniformly smaller biases across all sample size combinations, lower RMSE, and greater power for $T=4,5,6$, and $n=100,1000,$ and $5000$. The differences between the two estimators of $\mu _{\phi }$ become negligible only when $T=10$. In the light of our discussion in sub-section (ref), this could be because the FDAC estimator uses averages of the individual sample moments both over time and across all units given the stationary properties of the autocovariances of the first differences, and thus it is not subject to the many moment problem that could adversely impact the HetroGMM estimator. Consequently, tests based on the FDAC estimator are not adversely affected as $T$ is increased with $n$ small, and its size is mostly around the nominal size of five per cent. However, tests based on the HetroGMM estimator tend to over-reject slightly as $T$ is increased when $n$ is relatively small ($n=100$). For example, for the uniform distribution with $\phi _{i} \in (-1,1]$ and $n=100$, the size of the tests of $\mu _{\phi }=0.5$ based on the HetroGMM estimator rises from $5.7$ to $10.5$ per cent when $T$ is increased from $4$ to $10$. These findings are in line with the results obtained in the literature when GMM is applied to homogeneous dynamic panels. \footnote{ For GMM estimators with many moment conditions, some of the moment conditions can be weak. The small-sample bias associated with the weak moments will result in substantial size distortions, which become more severe with greater weights on the weak moments. See also Section 6 of ChudikPesaran2021.}
As can be seen from the empirical power functions in Figure (ref), the tests based on FDAC and HetroGMM estimators can not reject $\mu _{\phi }=1$ with 100 per cent certainty due to the small sample sizes with $n=100$. But as $n$ and $T$ increase, the empirical power functions become steeper, illustrating an enhanced ability to discern deviations from the null hypothesis.
As discussed in sub-section (ref), the FDAC and HetroGMM estimators of $\sigma _{\phi }^{2}$ are consistent so long as the true value of $\sigma _{\phi }^{2}$, namely $\theta _{2,0}-\theta _{0,1}^{2}$ is not too close to the boundary value of zero. Also to avoid negative estimates of the plug-in estimator of $\sigma _{\phi }^{2}$ given by ((ref)) we need $n$ to be sufficiently large. Table (ref) in the online supplement summarizes the number of replications, out of $2,000$, with negative or close to zero estimates (defined as estimates below $0.0001$) for the baseline experiments and sample size combinations $n=100,1000,2500,5000$ and $T=5,6,10$. The frequencies of the HetroGMM estimator are noticeably higher than those of the FDAC estimator for small $T$ and $n$. When $n=100$, a sizeable proportion of the estimates of $\sigma _{\phi }^{2}$ are negative, suggesting that $n=100$ is not sufficiently large for the asymptotic properties to hold. However, as to be expected, the number of negative estimates declines rapidly as $n$ and $T$ are increased. Accordingly, we only focus on samples with $n\geq 1000$, and report the bias and RMSE of the estimates of $\sigma _{\phi }^{2}$ for sample size combinations $ n=1000,2500,5000$ and $T=5,6,10$. The results for the positive estimates are summarized in Table (ref) of the online supplement for uniformly distributed $\phi _{i}$. For these sample size combinations, we only encounter very few negative estimates and none when $n=5000$ and $T\geq 6$. \footnote{ We did not consider estimating $\sigma _{\phi }^{2}$ under the categorical distributions of $\phi _{i}$ since the associated true values of $\sigma _{\phi }^{2}$ are too close to zero.}
Overall, both FDAC and HetroGMM estimators of $\sigma _{\phi }^{2}$ perform well when $T=10$ or $n$ is large, with comparable performances whether $ |\phi |<1$ or $\phi _{i} \in (-1,1]$. However, the FDAC estimator performs much better for smaller values of $T$ and $n$, as can be seen from the larger bias and RMSE of the HetroGMM estimator.
The empirical power functions for FDAC and HetroGMM estimators of $\sigma _{\phi }^{2}$ are shown in Figure (ref) of the online supplement for the uniformly distributed AR coefficients with $\phi _{i} \in (-1,1]$ in the baseline case (Gaussian errors and no GARCH effects). The empirical power functions are flat around the true value of $\sigma _{\phi }^{2}$ for $T=5$ and $n$ small. When $T=5$, large values of $n$ are required to achieve reasonable power in the locality of the null hypothesis. The power improves rapidly as $T$ and $n$ are increased and, in line with the earlier results, the FDAC estimator performs better than the HetroGMM estimator.
The FDAC estimators seem to be reasonably robust to departures from Gaussian errors and the presence of GARCH effects. Table (ref) of the online supplement provides results for the four combinations of error distributions, Gaussian and non-Gaussian, without and with GARCH effects for estimation of $\mu _{\phi }$. This table reports the results for the uniformly distributed AR coefficients with $\phi _{i} \in (-1,1]$ and $\mu _{\phi }=0.5$. We obtain similar results when we generate $\phi _{i}$ following a categorical distribution. The RMSE and size distortions of the FDAC estimator increase only slightly as we move from Gaussian to non-Gaussian errors and as we allow for GARCH effects. In contrast, the HetroGMM estimator is much more adversely affected by departures from Gaussian errors. Its bias and RMSE are much higher, with large size distortions, particularly with small $n$ ($n=100$). The performances of both estimators are adversely affected when non-Gaussian errors are combined with GARCH effects. Estimation of $\sigma _{\phi }^{2}$ is similarly adversely affected when we allow for non-Gaussian errors as well as GARCH effects. The related simulation results are summarized in Table (ref) of the online supplement.
Overall, the FDAC estimator outperforms the HetroGMM estimator and seems to be reasonably robust to non-Gaussian errors and GARCH effects. It is also simple to compute. In what follows we focus on the estimation of $\mu _{\phi }$ and compare the FDAC estimator with the HomoGMM estimators as well as the MSW estimator that allows for slope heterogeneity.
Tables (ref) and (ref) in the online supplement summarize the results comparing the FDAC estimator with FDLS, AH, AAH, AB and BB estimators, where $\phi _{i}$ is uniformly distributed, $\phi _{i}=\mu _{\phi }+v_{i}$ and $v_{i}\sim IIDU[-a,a]$ with $a=0.5$ and $\mu _{\phi }=0.4$ ($|\phi _{i}|<1$) and $\mu _{\phi }=0.5$ ($\phi _{i} \in (-1,1]$ ). We use the sample size combinations, $T=4,6,10$, and $n=100,1000,5000,$ in the baseline case where the errors are Gaussian without GARCH effects. The simulation results with the other error processes are available upon request.
In line with our theoretical derivations, the HomoGMM estimators that neglect heterogeneity are severely biased and show large size distortions, whilst the bias of the FDAC estimator is close to zero and its size is around the five per cent nominal level, irrespective of whether $|\phi _{i}|<1$ or $\phi _{i} \in (-1,1]$. Also, with increases in $n$ and/or $T$, the biases of the HomoGMM estimators do not shrink to zero and, as a result, the size distortions of the HomoGMM estimators become even more pronounced. The simulation results also confirm the magnitude of the asymptotic bias of the AH estimator given by ((ref)) in Section (ref), and those of AB and BB estimators provided in Section (ref) of the online supplement.
Since it is not known if the heterogeneity bias is serious, it is natural to ask if the FDAC estimator continues to perform equally well under homogeneity ($\phi _{i}=\mu _{\phi }=0.5$ for all $i$), and if its performance under homogeneity is comparable to those of HomoGMM estimators of $\phi $. Accordingly, we also compute bias, RMSE, and size of the FDAC and HomoGMM estimators under slope homogeneity $(a=0)$ with $\mu _{\phi }=0.5 $. The results for Gaussian errors without GARCH effects are summarized in Table (ref) of the online supplement. As can be seen, the FDAC estimator continues to perform well even under slope homogeneity. Its bias is close to zero and shows only a small degree of size distortions when $n=100$. In terms of assumptions, the FDAC estimator is closest to the FDLS estimator under homogeneity. Figure (ref) in the online supplement compares the empirical power functions of FDAC and FDLS estimators. Compared to the FDLS estimator, the FDAC estimator makes use of higher order autocorrelation of first differences that are not needed for identification of $\mu _{\phi }$ under homogeneity. As a result, the FDLS estimator is marginally more powerful than the FDAC for small $T=4$, while the opposite is the case for $T=10$.
When comparing the FDAC and the other HomoGMM estimators (such as AAH, BB, or AB) one needs to be cautious however, since these estimators do allow for the distribution of $y_{i0}$ to depart from the steady state distribution of $\{y_{it}\}$. With this in mind, we note that the FDAC estimator performs well when compared to AH, AAH and AB estimators, although it is marginally less efficient when compared to the BB estimator. Also, the FDAC estimator has less size distortion and better power performance compared to all HomoGMM estimators as $T$ is increased. In short, these results demonstrate the FDAC estimator is reliable and has desirable small-sample performance even in homogeneous panels with stationary outcome processes.
Figure (ref) in the online supplement shows the empirical power functions for the FDAC estimator under homogeneity with $ \phi_{i} = \mu_{\phi} = 0.5$ for all $i$ and heterogeneity with uniformly distributed $\phi _{i} \in (-1,1]$ $(\mu_{\phi} = 0.5)$, in the cases of Gaussian and non-Gaussian errors without GARCH effects. The empirical power functions for the FDAC estimator in the cases of Gaussian errors without and with GARCH effects are displayed in Figure (ref) of the online supplement. The power functions become steeper as $n$ and $T$ increase. In general, the power of the FDAC estimator is similar under heterogeneous and homogeneous $\phi _{i}$. Consistent with the previous findings, with non-Gaussian errors and/or GARCH effects, particularly for small $n=100$, the power functions become noticeably flatter, and the size distortions become more pronounced.
This section compares the small-sample performance of the FDAC estimator with the MSW estimator by MavroeidisEtal2015. Table (ref) in the online supplement reports bias, RMSE, and size of the FDAC and MSW estimators for $\mu _{\phi } $ for $T=4,6,10$, and $n=100,1000$, with uniformly distributed $\phi _{i}$ and Gaussian errors without GARCH effects. The left and right panels of the table report results for $\mu _{\phi }=0.4$ ($|\phi _{i}|<1$) and $\mu _{\phi }=0.5$ ($\phi _{i} \in (-1,1]$), respectively. The performance of the FDAC estimator is in line with the ones already discussed and as noted earlier is not affected by whether some $\phi _{i}=1$ or not. In contrast, the MSW estimator performs rather poorly in the presence of a high degree of heterogeneity in $\phi _{i}$ and shows large biases and substantial size distortions across the examined sample sizes. In the case of $\phi _{i} \in (-1,1]$, the MSW estimator shows greater bias, RMSE, and size distortions.
Since the first differences of $y_{it}$ do not depend on the initial values when $\phi _{i}=1$, non-stationary initialization matters only if $|\phi _{i}|<1$. In this case, using ((ref)) it is clear initial values matter only when $M_{i}$ is small. Therefore, to investigate the robustness of the FDAC estimator to different initializations we consider relatively small values of $M_{i}=1$ and $3$ for all $i$, compared with the baseline case where we set $M_{i}=100$ for all units with $\left\vert \phi _{i}\right\vert <1$. The initial values are generated as $(y_{i,-M_{i}}-\mu _{i})\sim IIDN(b,\kappa \sigma _{i}^{2})$ with $b=1$ and $\kappa =2$, compared to their steady state values of $b=0$ and $\kappa_{i} =1/(1-\phi _{i}^{2})$, respectively. When $\phi _{i}$ are generated from a categorical distribution we set $M_{i}=1$ for all units with $\phi _{H}=1$.
We consider both uniformly and categorically distributed $\phi _{i}\,$. The results for the uniformly distributed $\phi _{i}$ under the three initializations $M_{i}\in \{100,3,1\}$ are summarized in Table (ref) of the online supplement. Similar results when $\phi _{i}$ follow the categorical distribution are given in Table (ref) of the online supplement. It is clear that the FDAC estimator is adversely affected when $M_{i}=1$ and displays bias and substantial size distortions. As to be expected, the magnitude of the bias is not affected by $n$ but falls sharply with $T$. As a result when $M_{i}=1$ we observe substantial size distortions when $n$ is large. Comparing the upper and lower panels, as the first differences of a unit root process are not affected by the initial values, having some $\phi _{i}$ being close to one mitigates the negative impact of non-stationary initializations on the FDAC estimator. These impacts are more pronounced for categorically distributed $\phi _{i}$ where the variances of $\phi _{i}$ are smaller, as shown in Table (ref) versus Table (ref). More importantly, as to be expected, the bias and size distortion of the FDAC estimator disappear as $M_{i}$ is increased. When moving from $M_{i}=1$ to $M_{i}=3$, the bias and size distortion shrink fast, with only a slight size distortion observed when $M_{i}=3$.
We also consider the relative performance of the FDAC and HomoGMM estimators under different initialization scenarios, for both cases of homogeneous and heterogeneous panels. Results for the homogenous case when $\phi _{i}=\mu _{\phi }=0.5$ are summarized in Table (ref), and results for the heterogenous case are provided in Tables (ref) and (ref) for cases where $\mu _{\phi }=0.4$ ($|\phi _{i}|<1$) and $\mu _{\phi }=0.5$ ($\phi _{i} \in (-1.1]$), respectively. In the homogeneous case, when $M_{i}=1$, the FDAC, FDLS, and BB estimators all show sizeable bias and size distortions that do not vanish as $n$ increases. Also, as to be expected, under homogeneity, the AH, AAH, and AB estimators are robust to non-stationary initialization and have similar performances across different values of $M_{i}$. In the case of heterogeneous panels, the performance of the FDAC estimator is as discussed above. For the HomoGMM estimators, the magnitude of neglected heterogeneity bias is smaller with less serious size distortions when $M_{i}=1$ or $3$, as compared to $M_{i}=100$ (which approximately corresponds to the stationary case). The AH estimator seems to be an exception. Nonetheless, the HomoGMM estimators exhibit substantial size distortions across most of the considered sample sizes, leading to incorrect inference.
In short, for moderate values of $M_{i}$ (in the case of our experiments when $M_{i}>3)$, the performance of the FDAC estimator is satisfactory even when $y_{i0} - \mu_{i}$ are not drawn from the steady distribution of the underlying processes, $\{y_{it}-\mu _{i}\}$. Comparisons of the FDAC and HomoGMM estimators also highlight the trade-off that exists between the \textquotedblleft non-stationary initialization" bias of the FDAC estimator and the neglected heterogenous bias of the HomoGMM estimators. It remains a challenge to simultaneously deal with heterogeneity of $\phi _{i}$ and the non-stationarity of the initial values.
Estimating earnings equations is crucial for answering some of the most important economic
questions.\footnote{ See p. 58 in Guvenen2009 for a brief summary of several economic inquiries hinging on the estimation of earnings functions.} Variance of earnings has been modeled and decomposed to measure income uncertainties in LillardWeiss1979, MaCurdy1982, CarrollSamwick1997, MeghirPistaferri2004, AltonjiEtal2013 and to quantify earnings mobility in LillardWillis1978 and GewekeKeane2000. The covariance structures between earnings and other households' characteristics, for example, work hours, consumptions and savings, have been studied by AbowdCard1989, HubbardEtal1995, Guvenen2007, and AlanEtal2018.
Among these studies, a homogeneous AR or ARMA process is often used as a component when modeling innovations in earnings processes. Based on the Restricted Income Profiles model that assumes homogenous linear trends proposed in MaCurdy1982, MaCurdy1982 and HubbardEtal1995 obtained close to unit root estimates for the AR(1) coefficient, ranging from 0.946 to 0.998.\footnote{ See Table 5 on p. 111 in MaCurdy1982 using an ARMA(1,1) process. See Table 2 on p. 380 in HubbardEtal1995 based on an AR(1) process.} Following this literature, a unit root assumption was imposed in CarrollSamwick1997 and MeghirPistaferri2004. On the other hand, using the Heterogeneous Income Profiles, by assuming unit-specific linear trends, LillardWeiss1979 obtained estimates of the AR(1) coefficient (assumed to be homogeneous) ranging from 0.153 to 0.860 for a sample with PhD degrees. Guvenen2009 obtained estimates ranging from 0.809 to 0.899 using PSID data.\footnote{ See Tables 2, 4, 6 and 7 in LillardWeiss1979, Table 1 on p. 64 in Guvenen2009, and the abstract of GuKoenker2017.}
There are also a number of studies that allow for heterogeneity in the AR(1) coefficients. Prominent examples are BrowningEtal2010, AlanEtal2018, BrowningEtal2010, and GuKoenker2017. These studies are typically based on panels with a moderate time dimension and make parametric assumptions regarding the distribution of the AR(1) coefficients; often using a Bayesian framework.\footnote{ See pp. 227--232 in BrowningEjrnaes2013 for a comprehensive survey of heterogeneity in parameters of earnings functions.} The application of the FDAC estimator to earnings equation allows for heterogeneity in the AR(1) coefficients without making any strong parametric assumptions, even when $T$ is as small as $5$. Also because of first differencing prior to estimation, the FDAC estimator is robust to unobserved individual-specific characteristics and is not subject to misspecification bias that could arise when log real wages are filtered for individual-specific characteristics before investigating the dynamics of the earnings process.
We consider estimating the earnings equation with fixed effects, heterogeneous autoregressive coefficients, without imposing any restrictions on the joint distributions of $\alpha _{i}$, $\phi _{i}$, and $y_{i0}$. However, to accommodate growth in real earnings we extend our baseline model in ((ref)) to allow for linear trends:
where $y_{it}=log(earnings_{it}/p_{t}),$ $earnings_{it}$ is the reported earnings of individual $i$ in year $t$, $p_{t}$ is a general price, and $ g_{i}$ is the growth rate of real earnings for individual $i$. ((ref)) can be written equivalently as
with $\tilde{y}_{it}(g_{i})=y_{it}-g_{i}t$ and $b_{i}=\alpha _{i}-g_{i}\phi _{i}$. For $\vert \phi_{i}\vert < 1$, the steady state distribution of $ y_{it}$ can now be derived using
When $T$ is sufficiently large, individual-specific growth rates, $g_{i}$, can be estimated $\sqrt{T}$-consistently by running individual least squares regressions of $y_{it}$ on an intercept and a linear trend, and then using the residuals from these regressions to estimate the moments of $\phi _{i}$. This approach requires $n$ and $T$ to be both large. In the case of the present empirical application where $T$ is short ($5$ or $10)$, we provide estimates of the moments of $\phi _{i}$ assuming that $g_{i}=g$ for individuals within a given group, but allow $g$ to differ across groups, classified by the educational attainment levels. $\sqrt{n}$-consistent estimators of $g$ can be obtained either from the pooled regression of $ y_{it}$ on fixed effects and a common linear trend, namely
with $\bar{y}_{\circ t}=n^{-1}\sum_{i=1}^{n}y_{it}$ and $\bar{y}_{\circ \circ }=T^{-1}\sum_{t=1}^{T}\bar{y}_{\circ t}$, or after first differencing of ((ref)) by
For small $T$ there is little to choose between these two estimators, and they are identical when $T=2$. Given either of the above estimators, generically denoted by $\hat{g}$, $\tilde{y}_{it}(\hat{g})=y_{it}-\hat{g} t$ can now be used to estimate the moments of $\phi _{i}$ using the FDAC or MSW procedures.\footnote{ Consistent estimation of $E(\phi _{i})$ in the presence of heterogeneity in both $\phi _{i}$ and $g_{i}$ requires moderate to large values of $T$. The approach used in the empirical literature whereby $y_{it}$ are first de-meaned and de-trended for each $i$ prior to the estimation of $E(\phi _{i})$ is subject to Nickell1981 bias in the case of short $T$ panels, even if $E(\phi _{i})=\phi $.}
In addition to the FDAC estimates, we also present estimates based on three estimation methods assuming homogeneous slope coefficients, namely AAH, AB, and BB estimators proposed by ChudikPesaran2021, ArellanoBond1991, and BlundellBond1998, and the MSW estimator of MavroeidisEtal2015. Following MeghirPistaferri2004, individuals in each time series sample are divided into three education categories, where \textquotedblleft HSD" refers to high school dropouts with less than 12 years of education, \textquotedblleft HSG" refers to high school graduates with at least 12 but less than 16 years of education, and \textquotedblleft CLG" refers to college graduates with at least 16 years of education.\footnote{ The sample for all individuals in both $5$ and $10$ yearly samples covered $ 3,113$ individuals with consecutive observations of nine years or more, and 36,325 individual-year observations.} To allow for possible time variations in the estimates of mean earnings persistence we provide estimates for $five$ and $ten$ yearly non-overlapping sub-periods. The five yearly samples are 1976--1980, 1981--1985, 1986--1990 and 1991--1995. The ten yearly samples are 1976--1985, 1981--1990 and 1991--1995. For each sub-period, we provide estimates for all categories combined, as well as separate estimates for the three educational sub-categories.\footnote{ From 1997 PSID data are updated every two years. We confine our analysis to the years 1976 to 1995 to construct panels with $5$ and $10$ consecutive years.} To save space, the results for the last five and ten yearly samples are given in the paper. The estimates for the earlier sub-periods are provided in the online supplement.
Table (ref) gives the estimates of mean earnings persistence, $ \mu_{\phi} = E(\phi _{i}),$ and the common linear trend coefficient, $g$, for the sub-periods 1991--1995 ($T=5$) and 1986--1995 ($T=10$). The estimates of $g$ are on average around $2$ per cent per annum with some modest variations across the sub-samples and educational categories. The HomoGMM estimates (AAH, AB and BB) differ a great deal, both over sub-periods and across educational categories. The AAH estimates are all around 0.50 and show little variations across the two sub-periods and the educational categories. The AB estimates tend to be quite low and are not statistically significant for two of the educational categories in the shorter sub-period $(T=5)$. In contrast, the BB estimates are much larger and in many instances are close to unity. For example, for the sub-period 1986--1995 $(T=10)$, the BB estimates of earnings persistence for the three educational categories HSD, HSG and CLG are 0.923 (0.003), 0.914 (0.003) and 0.992 (0.004), respectively, with standard errors in brackets.
We also find sizeable differences in the estimates of mean earnings persistence when we consider the FDAC and MSW estimators. The MSW estimates are all around 0.45 and do not vary with the level of educational attainment. In contrast, the FDAC estimates are somewhat larger (lie in the range of 0.570--0.734) and rise with the level of educational attainment. This pattern can be seen in both sub-periods. For example, for the longer sub-period (1986-1995), the mean persistence for HSD, HSG and CLG categories are estimated to be 0.580 (0.071), 0.611 (0.028) and 0.735 (0.040), respectively. Similar results are obtained for the other sub-periods. See Tables (ref) and (ref) of the supplement. Interestingly, the higher earnings persistence of the college graduate category is a prominent feature of the FDAC estimates for all sub-periods. This result is also in line with a number of theoretical arguments in the literature in terms of higher mobility of college graduates and their relative job stability, for example, CarrollSamwick1997 and CarneiroEtal2023.
Although we have not developed a formal statistical test of the heterogeneity $\phi _{i}$, the estimates of $\sigma _{\phi }^{2}$ provide a good indication of the degree of within-group heterogeneity. Estimates of $ \sigma _{\phi }^{2}$ based on MSW and FDAC procedures for the various sub-periods are given in Tables (ref)--(ref) of the online supplement. The FDAC estimates are much larger than the MSW estimates. For example, for the sub-period 1986--1995 the MSW estimates of $ \sigma _{\phi }^{2}$ are all around 0.011 with standard errors in the range of 0.005--0.011, whilst the FDAC estimates of $\sigma _{\phi }^{2}$ for the same sub-period are 0.122 (0.06), 0.12 (0.031) and 0.141 (0.036) for the three educational categories of HSD, HSG and CLG, respectively. The degree of within-group heterogeneity also seems to vary over time. For example, for the shorter sub-period (1991-1995), the FDAC estimates of $\sigma _{\phi }^{2}$ are generally smaller with larger standard errors for the two categories of HSG and CLG.
This paper considers the estimation of heterogeneous panel AR(1) models with short $T$, as $n\rightarrow \infty $. It allows for individual fixed effects and proposes estimating the moments of the AR(1) coefficients, $E(\phi _{i}^{s})$, for $s=1,2,...,S$, using the autocorrelation functions of first differences. It is shown that the standard GMM estimators proposed in the literature for short $T$ homogeneous panels are inconsistent in the presence of slope heterogeneity. Analytical expressions for the bias are derived and shown to be very close to estimates obtained from stochastic simulations.
We propose two moment based estimators. A simple estimator based on autocorrelations of first differences, denoted by FDAC, and a GMM estimator based on autocovariances of first differences denoted by HetroGMM. Both estimators allow for some of the cross section units to have unit roots.
The small sample properties of the proposed estimators are investigated using Monte Carlo experiments. It is shown that the FDAC estimators of $\mu _{\phi }$ and $\sigma _{\phi }^{2}$ perform much better than the corresponding HetroGMM estimator. We also find that quite large samples might be required for reliable estimation of $\sigma _{\phi }^{2}$, assuming that the true value of $\sigma _{\phi }^{2}$ is not too close to zero.
The simulation results also show that the FDAC estimator of $\mu _{\phi }$ is robust to different distributions of autoregressive coefficients and error processes. Further, we find that the FDAC estimator performs well even under homogeneous AR(1) coefficients. The magnitudes of bias and RMSE of the FDAC estimator are comparable to the HomoGMM estimators, and the size of the tests based on the FDAC estimator is mostly around the 5 per cent nominal level. But when initializations of the outcome processes deviate from their associated steady state distributions, the FDAC estimator could suffer from bias and size distortions. There is a trade-off between heterogeneity bias and the bias due to the non-stationary initializations.
The utility of the FDAC estimators of $\mu _{\phi }$ and $\sigma _{\phi }^{2} $ is illustrated by an empirical application using the 1976--1995 PSID data to estimate heterogeneous AR(1) panels in log real earnings with a common linear trend. We provide estimates of $\mu _{\phi }$ and $\sigma _{\phi }^{2} $ over a number of $5$ and $10$ yearly sub-periods, with three educational groups. The estimates of $\mu _{\phi }$ differ systematically across the education groups, with the mean persistence of real earnings rising with the level of educational attainments (high school dropouts, high school graduates, and college graduates). The estimates of $\sigma _{\phi }^{2}$ differ across periods and levels of educational attainment but do not display any particular patterns.
It is important to acknowledge that the scope of the present paper is limited, with a number of remaining challenges: (a) allowing for individual-specific time-varying covariates, and (b) simultaneously dealing with heterogeneity and non-stationary initializations. It is not clear that such extensions will be possible without relaxing the assumption that $T$ is short and fixed, as $n\rightarrow \infty $. But these are clearly important topics for future research.
{ \setstretch{1.26} }
\makeatletter \setcounter{page}{1}\setcounter{table}{0} \setcounter{section}{0} \setcounter{figure}{0} \setcounter{footnote}{0} \setcounter{theorem}{0} \setcounter{proposition}{0} \setcounter{assumption}{0} \setcounter{lemma}{0} \setcounter{remark}{0} \makeatother
This online supplement is organized as follows. Section (ref) provides a proof of Lemma (ref) under the stationarity of the first differences, $\Delta y_{it} = y_{it} - y_{i,t-1}$. Section (ref) further illustrates the convergence property with uniformly distributed autoregressive coefficients, $\phi _{i}$. Section (ref) derives expressions for the analytical bias of the AB and BB estimators under heterogeneity of $\phi _{i}$ when $T=4$. Section (ref) derives the asymptotic covariance matrix for the HetroGMM estimator of the first two moments and its consistent estimator. Section (ref) provides formulae for empirical power functions of the tests based on our proposed estimators in the Monte Carlo simulations. Section (ref) provides additional Monte Carlo evidence. Section (ref) describes the sample (1976--1995) of the Panel Study of Income Dynamics (PSID) data used in the empirical application and provides estimation results for a number of sub-periods in addition to the ones reported in the main paper.
We first establish conditions under which first differences, $\Delta y_{it},$ are covariance stationary for any given $t$ and $i$. Consider the result ( (ref)) in the main paper which we reproduce here for convenience:
for $t=2,3,...,T,$ where $R_{i}\left( y_{i,-M_{i}}\right) =-\phi _{i}^{M_{i}-1}(1-\phi _{i})\left( y_{i,-M_{i}}-\mu _{i}\right) $. Assuming $ \phi _{i}$ and $u_{it}$ are independently distributed and since the initial values, $y_{i,-M_{i}}-\mu _{i},$ are given, we have
Also since $\phi _{i}\in (-1,1]$, then $E\left\vert \phi _{i}^{s}(1-\phi _{i})\right\vert \leq c^{s},$ for some $c<1$, and we have
Hence, $E\left\vert y_{it}\right\vert $ exists for all values of $\phi _{i}\in (-1,1]$ and is given by
It is clear that, since $t=1,2,...,T$ and $T$ is finite, then $E\left( \Delta y_{it} \right) $ varies with $t$ and in general depends on the initial values, $y_{i,-M_{i}}$. $E\left( \Delta y_{it}\left\vert y_{i,-M_{i}}-\mu _{i}\right. \right) $ is time-invariant if and only if $M_{i}\rightarrow \infty $, and hence unconditionally we have $E(\Delta y_{it})=0$, for all $i$ and $ t $, if $M_{i}\rightarrow \infty $.
By Cauchy-Schwarz inequality $\left\vert \gamma _{\Delta }(h)\right\vert =\left\vert E\left( \Delta y_{it}\Delta y_{i,t-h}\right) \right\vert \leq \left[ E\left( \Delta y_{it}\right) ^{2}E\left( \Delta y_{i,t-h}\right) ^{2} \right] ^{\frac{1}{2}}$, thus for the existence of autocovariances of $ \Delta y_{it}$, it is sufficient to show that $E\left( \Delta y_{it}\right) ^{2}<\infty $. Using ((ref)) in the main paper, it readily follows that
and given the independence of $\sigma _{i}^{2}$ and $\phi _{i}$ (see Assumption (ref) in the main paper) we have
We now show that $\sum_{s=0}^{\infty }E\left[ (1-\phi _{i})^{2}\phi _{i}^{2s} \right] $ is convergent for any probability distributions of $\phi _{i}$ defined over the interval $(-1,+1]$. Note that for any finite $M$
where $1+\phi _{i}>\epsilon >0$, and $-1 < \phi_{i} \leq 1$.
Hence,
But since $\phi_{i} \in (-1, 1]$, $E\left\vert \phi _{i}^{\ell }\right\vert \leq 1$ for any $\ell =0,1,...,$ and it follows that $\sum_{s=0}^{M}E\left[ (1-\phi _{i})^{2}\phi _{i}^{2s}\right] $ $\leq 4/\epsilon $ for any finite $M$ and as $M\rightarrow \infty $. Therefore, it follows that $\left\vert \gamma _{\Delta }(h)\right\vert <C$, as required.
Having established the existence of $\gamma _{\Delta }(h)$, using ((ref)) and recalling that under Assumption (ref) in the main paper $\phi _{i}$ and $\sigma _{i}^{2}$ are independently distributed we have
Similarly, to derive $\gamma _{\Delta }(h)=E\left( \Delta y_{it}\Delta y_{i,t-h}\right) $ we first note that $M_{i}\rightarrow \infty $, then using ((ref)) we have
and for $h=1,2,...,$
First, we consider the second term, and note that
Also
Hence
and since $\phi _{i}$ and $\sigma _{i}^{2}$ are independently distributed we have
As before, for all $\phi _{i}\in (-1,1],$ we have $E\left\vert (1-\phi _{i})^{2}\phi _{i}^{h+s}\right\vert \leq E\left\vert (1-\phi _{i})^{2}\left\vert \phi _{i}\right\vert ^{h+s}\right\vert \leq c^{h+s}$, where $c<1$, and the series is convergent, and we have
or
as required. Similarly
The results of Lemma (ref) are now established noting that $E\left( \sigma _{i}^{2}\right) =\sigma ^{2}$.
It is also instructive to consider the important case where $\phi _{i}$ is uniformly distributed. First suppose that $\phi _{i}\thicksim Uniform(0,a] $ for $0<a\leq 1$, then \thinspace $E\left( \phi _{i}^{\ell }\right) =\frac{ a^{\ell }}{\ell +1}$, and
Hence
When $a<1$, all the three individual sums in the above expression are bounded by $C/(1-a^{2})$. However, this does not follow when $a=1$, and the series $\sum_{s=0}^{\infty }\frac{1}{2s+1}$, $\sum_{s=0}^{\infty }\frac{2}{ 2s+2}$, and $\sum_{s=0}^{\infty }\frac{1}{2s+3}$, diverge individually. Hence, to investigate the convergence property of $\sum_{s=0}^{\infty }E \left[ (1-\phi _{i})^{2}\phi _{i}^{2s}\right] $ when $a=1$, we need to consider all the terms together. For $a=1$,
and after some algebra we have
Similarly, for $\phi _{i}\thicksim Unifrom(-1+\epsilon ,0]$ we have $E\left( \phi _{i}^{\ell }\right) =-\frac{(-1)^{\ell }\left( 1-\epsilon \right) ^{\ell }}{\ell +1}$, and we have
and $\sum_{s=0}^{\infty }E\left[ (1-\phi _{i})^{2}\phi _{i}^{2s}\right] $ is convergent for $\epsilon >0$, and diverges if $\epsilon =0$. The latter case is ruled out under Assumption (ref) in the main paper, which establishes the necessity of ruling out the boundary value of $\phi _{i}=-1$.
Result ((ref)) follows directly from ((ref)), after subtracting $ E\left( \phi _{i}\right) $ from\ both sides. Also, since $\phi _{i}\in \lbrack -1+\epsilon ,1]$, for some $\epsilon >0$, then $1+E(\phi _{i})>0$, and $E\left( \frac{1-\phi _{i}}{1+\phi _{i}}\right) >0$. Since $1/(1+\phi _{i})$ is a convex function of $\phi _{i}$ on $[-1+\epsilon ,1]$, then by Jensen inequality $E\left( \frac{1}{1+\phi _{i}}\right) \geq \frac{1}{ 1+E(\phi _{i})}$, and it follows that $plim_{n\rightarrow \infty }\hat{\phi} _{AH}\leq E(\phi _{i})=\mu _{\phi }$. Since $1+\mu _{\phi }=1+E(\phi _{i})>0$ , the asymptotic bias is zero if and only if $\frac{1}{1+\mu _{\phi }} =E\left( \frac{1}{1+\phi _{i}}\right) $, and due to the convexity of $ 1/(1+\phi _{i})$, this condition is met only if $\phi _{i}=\mu _{\phi }$ for all $i$.
The AB estimator proposed by ArellanoBond1991 is based on the following moment conditions:\footnote{ See equation (8) on p. 5 in ChudikPesaran2021.}
which can also be written as $E[y_{is}(\Delta y_{it}-\phi _{i}\Delta y_{i,t-1})]=0,$ with $(T-1)(T-2)/2$ moment conditions in total. When $T=4$, under homogeneity of $\phi_{i}$, the AB moment conditions are given by $ E[y_{i1}(\Delta y_{i3}-\phi \Delta y_{i2})]=0$, $E[y_{i1}(\Delta y_{i4}-\phi \Delta y_{i3})]=0$, and $E[y_{i2}(\Delta y_{i4}-\phi \Delta y_{i3})]=0$. With a fixed weight matrix $\mathbf{W}_{AB}$, the AB estimator can be written as
where $\boldsymbol{\bar{z}}_{na}=n^{-1}\left( \sum_{i=1}^{n}y_{i1}\Delta y_{i2},\sum_{i=1}^{n}y_{i1}\Delta y_{i3},\sum_{i=1}^{n}y_{i2}\Delta y_{i3}\right) ^{\prime },$ and
$\boldsymbol{\bar{z}}_{nb}=n^{-1}\left( \sum_{i=1}^{n}y_{i1}\Delta y_{i3},\sum_{i=1}^{n}y_{i1}\Delta y_{i4},\sum_{i=1}^{n}y_{i2}\Delta y_{i4}\right) ^{\prime }.$
Using ((ref)) in the main paper
and assuming that $\mu _{i}$ is distributed independently of $\{u_{it}\}$ (as assumed under AB) then using ((ref)) in the main paper and ((ref)) we have
As $M_{i}\rightarrow \infty $ for $|\phi _{i}|<1$ (with finite $M_{i}$ for $ \phi _{i}=1$),
Given ((ref)), if $Pr(\phi _{i}=1)=0$ and $\phi _{i}\in (-1,1]$, we have
Since $\phi _{i}$ is distributed independently of $\sigma _{i}^{2}$, for uniformly distributed $\phi _{i}=\mu _{\phi }+v_{i}$ with $v_{i}$ $\thicksim IIDU(-a,a)$, $a>0$ and $\phi _{i} \in (-1,1]$,
where $\boldsymbol{z}_{a}=-\sigma ^{2}(c_{\phi },1-c_{\phi },c_{\phi })^{\prime }$ and $\boldsymbol{z}_{a}=-\sigma ^{2}(1-c_{\phi },\mu _{\phi }-1+c_{\phi },1-c_{\phi })^{\prime }$ with $\sigma ^{2}=E(\sigma _{i}^{2})$ and $c_{\phi }=E\left( \frac{1}{1+\phi _{i}}\right) =\frac{1}{2a}\ln \left( \frac{1+\mu _{\phi }+a}{1+\mu _{\phi }-a}\right) $.
In addition to ((ref)), consider the following moment conditions, used in the system GMM estimator proposed by BlundellBond1998, and note that under homogeneity we have for $i=1,2,...,n$, and $t=3,4,...,T$,\footnote{ See equation (9) on p. 5 in ChudikPesaran2021.}
For $T=4$, with a given weight matrix $\boldsymbol{W}_{BB}$, the BB estimator based on the moment conditions in ((ref)) and ((ref)) is given by
where
Using ((ref)) in the main paper and ((ref)), similarly, we can derive the following equations as $M_{i}\rightarrow \infty $ for $|\phi _{i}|<1$,
and for $\phi _{i}=1$ and finite $M_{i}$, $E(\Delta y_{i,t-1}y_{it})=E(\Delta y_{i,t-1}y_{i,t-1})=\sigma _{i}^{2}$. In the case of $Pr(\phi _{i}=1)=0$ with $\phi _{i}\in (-1,1]$, it follows that
Since $\phi _{i}$ is distributed independently of $\sigma _{i}^{2}$, for uniformly distributed $\phi _{i}=\mu _{\phi }+v_{i}$ with $v_{i}$ $\thicksim IIDU(-a,a)$, $a>0$ and $\phi _{i}\in (-1,1]$,
where $\boldsymbol{z}_{c}=\sigma ^{2}(-c_{\phi },-1+c_{\phi },-c_{\phi },c_{\phi },c_{\phi })^{\prime }$ and $\boldsymbol{z}_{d}=\sigma ^{2}(c_{\phi }-1,1-\mu _{\phi }-c_{\phi },-1+c_{\phi },1-c_{\phi },1-c_{\phi })^{\prime },$ with $E(\sigma _{i}^{2})=\sigma ^{2}$, and $c_{\phi }=E\left( \frac{1}{1+\phi _{i}}\right) =\frac{1}{2a}\ln \left( \frac{1+\mu _{\phi }+a}{ 1+\mu _{\phi }-a}\right) $.
To approximate the values of the asymptotic bias of AB and BB estimators corresponding to our Monte Carlo experiments, we replace $\boldsymbol{W} _{AB} $ and $\boldsymbol{W}_{BB}$ by the simulated weight matrices\footnote{ The simulated weight matrices are calculated as the average of the weight matrices used in calculating the two-step AB and BB estimators across 2,000 replications.} with $a=0.5$, $\mu _{\phi }\in \{0.4,0.5\}$, and Gaussian errors without GARCH effects for $T=4$, and $n=5,000$. In this case, the biases of AB and BB estimators are around -0.055 and -0.045 for $\mu _{\phi }=0.4$, and -0.062 and -0.044 for $\mu _{\phi }=0.5$, respectively. These results are close to the simulated bias of these estimators reported in Tables (ref) ($\mu _{\phi }=0.4$) and (ref) ($ \mu _{\phi }=0.5$) for $T=4$ and $n=5,000$.
Suppose that Assumptions (ref)--(ref) in the main paper hold, $ T\geq 5$, and $M_{i}\rightarrow \infty $. The asymptotic distribution of $ \boldsymbol{\hat{\theta}}_{HetroGMM}=(\hat{\theta}_{1,HetroGMM},\hat{\theta} _{2,HetroGMM})^{\prime }$ is given by
with $\boldsymbol{\theta }_{0}=(\theta _{1,0},\theta _{2,0})^{\prime }$, and
where $\mathbf{H}_{\boldsymbol{\theta },nT}=\frac{1}{n}\sum_{i=1}^{n}\mathbf{ H}_{\boldsymbol{\theta },iT}$,
$\mathbf{S}_{\boldsymbol{\theta },T}(\boldsymbol{\theta }_{0})=\frac{1}{n} \sum_{i=1}^{n}(\mathbf{g}_{\boldsymbol{\theta },iT}-\mathbf{H}_{\boldsymbol{ \theta },iT}\boldsymbol{\theta }_{0})(\mathbf{g}_{\boldsymbol{\theta },iT}- \mathbf{H}_{\theta ,iT}\boldsymbol{\theta }_{0})^{\prime }$, $\mathbf{g}_{ \boldsymbol{\boldsymbol{\theta }},iT}=(\mathbf{g}_{iT}^{\prime },\mathbf{g} _{2,iT}^{\prime })^{\prime }$, and $\mathbf{h}_{iT}$, $\mathbf{h}_{2,iT}$, $ \mathbf{g}_{iT}$ and $\mathbf{g}_{2,iT}$ are given by ((ref)), ((ref)), ((ref)) and ((ref)) in the main paper, respectively. $ \mathbf{V}_{\boldsymbol{\theta }}$ can be consistently estimated by
with
The test statistics for $\mu _{\phi } =E(\phi _{i})$ and $\sigma _{\phi }^{2} = Var(\phi _{i})$ are given by
respectively, where FDAC and HetroGMM estimators of $\hat{\mu}_{\phi }=\hat{ \theta}_{1}$ are given by ((ref)) and ((ref)) in the main paper, respectively. $\hat{\sigma}_{\phi }^{2}$ is computed as the plug-in estimator given by ((ref)) in the main paper. In the Monte Carlo experiments, the empirical power functions (EPF) are computed as the simulated rejection frequencies for replications $r=1,2,...,R$:
and
Tables (ref) and (ref) summarize bias, RMSE, and size of FDAC and HetroGMM estimators of $\mu _{\phi }=E(\phi _{i})$ with uniformly and categorically distributed $\phi _{i}$, respectively, in the case of Gaussian errors without GARCH effects for the sample size combinations $n=100,1000,5000$ and $T=4,5,6,10$. The empirical power functions of FDAC and HetroGMM estimators of $\mu _{\phi }$ with uniformly distributed $\phi_{i} \in [-1+\epsilon, 1]$ for some $\epsilon>0$ are shown in Figure (ref).
Table (ref) reports the frequency where FDAC and HetroGMM estimates of $\sigma _{\phi }^{2}$ are either negative or very close to zero, using the threshold $\left( \hat{\sigma}_{\phi }^{2}\right) ^{(r)}<0.0001$, for replication $r=1,2,...,2000$, respectively, with uniformly distributed $\phi _{i}$ and Gaussian errors without GARCH effects for $n=100,1000,2500,5000$ and $T=5,6,10$. Table (ref) summarizes simulated outcomes with positive estimates of $\sigma_{\phi }^{2}=Var(\phi _{i})$ with uniformly distributed $\phi _{i}$ in the case of Gaussian errors without GARCH effects for the sample size combinations $n=100,1000,5000$ and $T=5,6,10$. The empirical power functions of FDAC and HetroGMM estimators of $\sigma _{\phi }^{2}$ (for simulated outcomes of positive estimates) are shown in Figure (ref) with $n=1000,2500,5000$ and $T=5,6,10$.
For the four combinations of error distributions, Gaussian and non-Gaussian, without and with GARCH effects, Tables (ref) and (ref) summarize simulation results of the estimation of $\mu _{\phi }$ and $ \sigma_{\phi}^{2}$ (for simulated outcomes of positive estimates), respectively, for uniformly distributed $\phi_{i} \in [-1+\epsilon, 1]$ for some $\epsilon>0$ with $\mu_{\phi} = 0.5$. Table (ref) reports the frequency where estimates of $ \sigma_{\phi}^{2}$ are not positive.
Tables (ref)--(ref) report bias, RMSE, and size of the FDAC, FDLS, AH, AAH, AB, and BB estimators with $\phi _{i}=\mu _{\phi }+v_{i}$, $v_{i}\sim IIDU(-a,a)$, $\mu _{\phi } \in \{0.4, 0.5\}$, $a = 0.5$, and Gaussian errors without GARCH effects. Table (ref) summarizes simulation results of FDAC and the above HomoGMM estimators with homogeneous $\phi_{i} = \mu_{\phi} = 0.5$ and Gassuain errors without GARCH effects.
Figure (ref) compares the empirical power functions of FDAC and FDLS estimators under homogeneity of $\phi_{i}$ for $T=4,10$, and $n=5,000$. Figures (ref) and (ref) plot the empirical power functions of the FDAC estimator in homogeneous ($\phi_{i} = \mu_{\phi} = 0.5$ for all $i$) and heterogeneous panel AR(1) panels, where the heterogeneous AR(1) coefficients are generated by the above uniform distribution with $\phi_{i} \in (-1,1]$ and $\mu_{\phi} = 0.5$, under different error processes for $T=4, 10$ and $n=100, 1000, 5000$.
Table (ref) summarizes bias, RMSE, and size of FDAC and MSW estimators of $\mu _{\phi }=E(\phi _{i})$ with uniformly distributed $\phi _{i}$ in the case of Gaussian errors without GARCH effects for the sample size combinations $n=100,1000$ and $T=4,6,10$.
Tables (ref) and (ref) summarize the bias, RMSE, and size of the FDAC estimator of $E(\phi _{i})$ with uniformly and categorically distributed $\phi_{i}$, respectively, under different initializations $M_{i}=100, 3, 1$ for all $i$, (except a case of categorically distributed $\phi_{i}$ where $M_{i}=100$ for units with $ \phi_{i}=\phi_{L} = 0.5$ and $M_{i}=1$ for units with $\phi_{i} = \phi_{H}=1$ ). Table (ref) reports the bias, RMSE, and sizes of the FDAC, FDLS, AH, AAH, AB, and BB estimators in homogeneous panels for $ M_{i}=100,3,1 $ for all $i$. The simulation results for heterogeneous panels with uniformly distributed $\phi_{i}$ are shown in Table (ref) for $\mu_{\phi}=0.4$, and Table (ref) for $\mu_{\phi}=0.5$. Table (ref) summarizes results of FDAC and MSW estimators in both homogeneous and heterogeneous panels for different initializations with $M_{i}=100,1$ for all $i$. .
Table (ref) shows the distribution of cross-sectional observation numbers by year based on the sample selection criterion in MeghirPistaferri2004. For different sub-periods, Tables (ref) and (ref) report the estimates of mean persistence of log real earnings in a panel AR(1) model with a common linear trend, and Tables (ref)--(ref) report the estimates of $ \sigma_{\phi}^{2} $ of the heterogeneous persistence parameters, $\phi_{i}$.
{ References }
{ Anderson, T. W., and Hsiao, C. (1981). \href{https://doi.org/10.2307/2287517} {Estimation of dynamic models with error components}. Journal of the American Statistical Association 76, 598-606. }
{ Anderson, T. W., and Hsiao, C. (1982). \href{https://doi.org/10.1016/0304-4076(82)90095-1} {Formulation and estimation of dynamic models using panel data}. Journal of Econometrics 18, 47-82. }
{ Arellano, M. and S. Bond (1991). \href{https://doi.org/10.2307/2297968} {Some tests of specification for panel data: Monte Carlo evidence and an application to employment equations}. The Review of Economic Studies 58, 277--297. }
{ Blundell, R. and S. Bond (1998). \href{https://doi.org/10.1016/S0304-4076(98)00009-8} {Initial conditions and moment restrictions in dynamic panel data models}. Journal of Econometrics, 87, 115--143. }
{ Chudik, A. and M. H. Pesaran (2021). \href{https://www.tandfonline.com/doi/abs/10.1080/07474938.2021.1971388?casa_token=2BwaEs6bQLoAAAAA:nYnfFDroaEJJN6-jZ4cm664YF0L-vppanHVwgPzw65e0Be9uIdHIywoaM4t36lsLXIDy1paeBHl-} {An augmented Anderson-Hsiao estimator for dynamic short-$T$ panels.} Econometric Reviews 1--32. }
{ Han, C. and P. C. Phillips (2010). \href{https://doi.org/10.1017/S026646660909063X} {GMM estimation for dynamic panels with fixed effects and strong instruments at unity}. Econometric Theory 26, 119--151. }
{ Mavroeidis, S., Y. Sasaki, and I. Welch (2015). \href{https://www.sciencedirect.com/science/article/pii/S0304407615001517?casa_token=HPGr9y-Oe7cAAAAA:u5R7ML9kmRODINHPmblJlnUV3z9-NQa4-xYGU9BBxm0dMTVqiwpp4Vh197-wAa-GWYSkgxvgjQ} {Estimation of heterogeneous autoregressive parameters with short panel data. } Journal of Econometrics 188, 219--235. }
{ Meghir, C. and L. Pistaferri (2004). \href{https://onlinelibrary.wiley.com/doi/abs/10.1111/j.1468-0262.2004.00476.x?casa_token=45Vh36Scuo8AAAAA:16mrkums6ZB4tQCcGjD8X1ETJCNKBOn-Fr1om-Kt4oywqd7zbvZshPd8FXk3Q4A0JTWSWoDkiIiLVb8} {Income variance dynamics and heterogeneity}. Econometrica 72, 1--32. }