Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
72,217 characters · 12 sections · 57 citation commands
Forward Orthogonal Deviations GMM and the Absence of Large Sample Bias
\onehalfspacing
In an influential paper, Alvarez2003 examined the asymptotic properties of a generalized method of moments (GMM) estimator that relied on forward orthogonal deviations (FOD) to remove fixed effects (FOD GMM). They showed that an FOD GMM estimator of the autoregression parameter in a first-order autoregressive (AR(1)) panel data model has a bias term in its asymptotic distribution if the number of time periods ($T$) increases too quickly relative to the number of cross-sectional units ($n$)---that is, if $T/n$ does not converge to zero. Though the model considered by Alvarez2003 was a special case, the paper's finding had an important implication. This is because if an estimator, say $\widehat{\delta}$, of parameter $\delta$ has a bias term in the asymptotic distribution of $\sqrt{nT}(\widehat{\delta} - \delta)$, then even though the estimator may be consistent, large sample confidence intervals and test statistics based on the estimator will be inaccurate.
Alvarez and Arellano's conclusion that the FOD GMM estimator they studied has asymptotic bias if $T/n$ does not converge to zero depends on using all available instrumental variables. If $T$ is not small, the number of instrumental variables can be large if all of them are used, and it is well-known that a GMM estimator that exploits many instrumental variables can be biased. Consequently, in practice researchers often resort to using fewer instrumental variables than all that are available. But left unanswered is the question: how many instrumental variables can be used while still avoiding bias when $T$ is not small compared to $n$? This paper addresses that question.
\sloppy Like this paper, Anderson2011, Bekker1994, Bun2006, Chao2005, Hansen2008, Hayakawa2019, and Koenker1999, among others, study how the properties of estimators are affected by the number of moment restrictions that are exploited. However, the papers most closely related to this paper are Alvarez2003 and Hsiao2017. As already noted, Alvarez2003 considered estimation of the AR(1) panel data model:
Assuming all available moment restrictions are exploited, Alvarez and Arelleno found that an FOD GMM estimator, say $\widehat{\beta}$, of the autoregression parameter, $\beta$, has a bias term in the asymptotic distribution of $\sqrt{nT}(\widehat{\beta} - \beta)$ if $T/n \rightarrow c >0$, as $n,T \rightarrow \infty$, but it does not have a bias term if $T/n \rightarrow 0$. Hsiao2017, on the other hand, investigated the effect of using fewer than all available moment restrictions. In particular, Hsiao2017 analyzed FOD GMM estimation of the model in ((ref)), but, instead of focusing on only using all available instrumental variables, Hsiao and Zhou also considered estimation based on a single instrumental variable per period. Upon doing so, they obtained results indicating that the FOD GMM estimators they considered---when based on a single instrumental variable per period---have no bias in their asymptotic distributions, as $n,T \rightarrow \infty$, regardless of what happens to $T/n$.\footnote{On the other hand, Hsiao2017 also argued that GMM based on first differences (FD GMM) and a single instrumental variable per period is asymptotically biased if $T/n \rightarrow c > 0$.} The papers by Alvarez2003 and Hsiao2017, therefore, indicate that whether or not there is a bias term in the asymptotic distribution of an FOD GMM estimator depends not just on what happens to $T/n$, as $n, T \rightarrow \infty$, but also on how many per-period instruments are used.
\fussy Using more general conditions than previously considered, this paper shows that whether or not an FOD GMM estimator is asymptotically biased can be summarized in terms of the largest number of instrumental variables used in a period. Specifically, I show that the FOD GMM estimator has no asymptotic bias regardless of what happens to $T/n$, if the maximum number of instruments used in a period increases with $T$ at a rate slower than $T^{1/2}$ increases. This conclusion is specific to using the FOD transformation. It does not necessarily carry over to other transformations that may be used to remove fixed effects.
Moreover, the conclusion is robust in the sense that the large sample bias result holds regardless of how $n$ and $T$ increase. Specifically, the conclusion is based on taking joint limits, which are limits obtained by letting $n$ and $T$ increase simultaneously.
A second less robust result is also provided: an asymptotic distribution result is provided that is based on taking limits sequentially. With a sequential limit, one index---$n$ or $T$---is taken to infinity, and then the other goes to infinity. Taking limits sequentially is often a more tractable tactic for obtaining results than joint limits, and sequential limit results may require weaker conditions than joint limit results. These advantages may explain why limits are sometimes taken sequentially in the literature. Hsiao2017, for example, used sequential limits to derive asymptotic distribution results for their FOD GMM estimators of the autoregression parameter in the model in ((ref)). A few other papers that exploit sequential limits are Kapetanios2008 and Hsiao2015.
However, sequential limits are not guaranteed to yield the same results one gets from joint limit analysis Moon1999, Moon2000. Indeed, I show that this is the case when all available instrumental variables are used. On the other hand, when fewer than all available instruments are used, Monte Carlo evidence indicates that the normal approximation obtained by taking limits sequentially appears to work well for constructing confidence intervals when $T$---in addition to $n$---is not small, provided the FOD transformation is used to remove fixed effects.
The regression model studied in this paper is
In this regression, $\boldsymbol{x}_{i,t}' := (x_{i,t,1},\ldots, x_{i,t,K})$ and $\boldsymbol{\beta}' := (\beta_1, \ldots, \beta_K)$ are vectors of regressors and parameters. Some or all of the $x_{i,t,k}$s may be lagged values of $ y_{i,t} $. The term $ \eta_i $ is an unobserved individual-specific or fixed effect, and $ v_{i,t} $ is an error term.
In order to estimate $\boldsymbol{\beta}$, the first step is to remove the fixed effect $\eta_i$ by transforming the dependent and explanatory variables. This paper studies transforming the variables using forward orthogonal deviations---the FOD transformation. The FOD transformation subtracts from each variable its within-group average over future periods. For example, the FOD transformed explanatory variables for the $i$th individual in the $t$th period are $ \ddot{\boldsymbol{x}}_{i,t} := c_{t} \left( \boldsymbol{x}_{i,t} - \overline{\boldsymbol{x}}_{i,t} \right) $, where $ \overline{\boldsymbol{x}}_{i,t} := \left( 1/(T-t)\right)\sum_{s=1}^{T-t} \boldsymbol{x}_{i,t+s} $ and $c_t^2 := (T-t)/(T-t+1) $. Similarly, the transformed value of the ($i,t$)th observation on the dependent variable is $ \ddot{y}_{i,t} := c_{t} \left( y_{i,t} - \overline{y}_{i,t} \right) $, where $ \overline{y}_{i,t} := \left( 1/(T-t)\right)\sum_{s=1}^{T-t} y_{i,t+s} $. The constant $c_t $ ensures the transformed errors (the $ \ddot{v}_{i,t} $s) are conditionally homoskedastic and uncorrelated if the original errors (the $ v_{i,t} $s) are conditionally homoskedastic and uncorrelated Arellano2003.
Now let $ \boldsymbol{z}_{i,t} $ denote a $ q_{t} \times 1 $ vector of instrumental variables for the $ i $th individual in the $ t $th period ($t=1,\ldots,T-1$). Also, set $ \boldsymbol{Z}_t' := \left(\boldsymbol{z}_{1,t}, \ldots, \boldsymbol{z}_{n,t} \right) $, $\ddot{\boldsymbol{X}}_t' := (\ddot{\boldsymbol{x}}_{1,t}, \ldots, \ddot{\boldsymbol{x}}_{n,t})$, $\ddot{\boldsymbol{y}}_t' := (\ddot{y}_{1,t}, \ldots, \ddot{y}_{n,t})$, and finally $ \boldsymbol{P}_t := \boldsymbol{Z}_t\left( \boldsymbol{Z}_t'\boldsymbol{Z}_t\right)^{-1}\boldsymbol{Z}_t' $. Then, the FOD GMM estimator can be written as
Arellano2003.
\sloppy As an alternative to Eq. ((ref)), the FOD GMM estimator can also be expressed as a two-stage least squares (TSLS) estimator after removing fixed effects with the FOD transformation. To see this, let $\boldsymbol{Z}_{d,i}$ be the block-diagonal matrix given by
Also, define $\dot{\boldsymbol{X}}_i' := (\ddot{\boldsymbol{x}}_{i,1}, \ldots, \ddot{\boldsymbol{x}}_{i,T-1})$ and $\dot{\boldsymbol{y}}_i' := (\ddot{y}_{i,1}, \ldots, \ddot{y}_{i,T-1})$. Then, note that $\sum_{t=1}^{T-1}\ddot{\boldsymbol{X}}_t'\boldsymbol{P}_t \ddot{\boldsymbol{X}}_t = \sum_{i=1}^n \dot{\boldsymbol{X}}_i'\boldsymbol{Z}_{d,i}(\sum_{i=1}^n\boldsymbol{Z}_{d,i}'\boldsymbol{Z}_{d,i})^{-1}\sum_{i=1}^n\boldsymbol{Z}_{d,i}'\dot{\boldsymbol{X}}_i $. Similarly, $\sum_{t=1}^{T-1}\ddot{\boldsymbol{X}}_t'\boldsymbol{P}_t \ddot{\boldsymbol{y}}_t = \sum_{i=1}^n \dot{\boldsymbol{X}}_i'\boldsymbol{Z}_{d,i}(\sum_{i=1}^n\boldsymbol{Z}_{d,i}'\boldsymbol{Z}_{d,i})^{-1}\sum_{i=1}^n\boldsymbol{Z}_{d,i}'\dot{\boldsymbol{y}}_i $. From this and Eq. ((ref)), we see that the FOD GMM estimator can be expressed as a TSLS estimator after applying the FOD transformation to the dependent and explanatory variables and upon using block-diagonal instrument matrices:
\fussy It is well-known that TSLS is efficient GMM when the errors are conditionally homoskedastic and uncorrelated. Moreover, as already noted, if the $v_{i,t}$s are conditionally homoskedastic and uncorrelated, then the transformed errors (the $\ddot{v}_{i,t}$s) are conditionally homoskedastic and uncorrelated. Hence, if the $v_{i,t}$s are conditionally homoskedastic and uncorrelated, the FOD GMM estimator is the asymptotically efficient GMM estimator given the moment restrictions $E(\boldsymbol{z}_{i,t}\ddot{v}_{i,t}) = \boldsymbol{0}$ ($t=1,\ldots,T-1$). This is a total of $\sum_{t=1}^{T-1}q_t$ moment restrictions, which can be a large number when $T$ is large, especially if $ q_{T}^{\ast} := \max_{1 \le t \le T-1}q_{t}$---i.e., the maximum per-period number of instrumental variables---increases with $T$. Therefore, although the FOD GMM estimator is efficient when the errors are conditionally homoskedastic and uncorrelated, we might anticipate it to be biased when $T$ is not small.
\sloppy Alvarez2003 provide conditions under which the distribution of an FOD GMM estimator, $\widehat{\beta}$, of the autoregression parameter, $\beta$, in the model in ((ref)) has an asymptotic bias term if $T/n \rightarrow c > 0$, as $n,T \rightarrow \infty$. In particular, note that
where $\theta_{n,T} := E(b_{n,T})/a_{n,T} $, $a_{n,T} := (1/(nT))\sum_{t=1}^{T-1}\ddot{\boldsymbol{y}}_{t-1}'\boldsymbol{P}_t \ddot{\boldsymbol{y}}_{t-1}$, $ b_{n,T} := (1/\sqrt{nT})\sum_{t=1}^{T-1}\ddot{\boldsymbol{y}}_{t-1}'\boldsymbol{P}_t \ddot{\boldsymbol{v}}_t $, $\ddot{\boldsymbol{y}}_{t-1} := c_t(\boldsymbol{y}_{t-1} - \overline{\boldsymbol{y}}_{t-1})$, $\boldsymbol{y}_{t-1}' := (y_{1,t-1},\ldots,y_{n,t-1})$, and $\overline{\boldsymbol{y}}_{t-1} := \left( 1/(T-t)\right)\sum_{s=1}^{T-t} \boldsymbol{y}_{t-1+s}$. If $a_{n,T}$ converges in probability to a fixed non-zero limit and $b_{n,T} - E(b_{n,T})$ has a limit distribution, then $\sqrt{nT}(\widehat{\beta} - \beta)$ has an asymptotic distribution that is centered at zero only if $\theta_{n,T} \overset{p}\rightarrow 0$, as $n,T \rightarrow \infty$. Alvarez2003 provide conditions that imply
with $\sigma^2 := \text{var}(v_{i,t})$ Alvarez2003. Using these results, they showed that if $\lim (T/n) = c$, with $ 0 \le c < \infty$, then
Alvarez2003. Therefore, for the AR(1) panel data model, $\sqrt{nT}( \widehat{\beta} - \beta)$ has an asymptotic bias term of $\theta := - \sqrt{c}(1 + \beta)$, which is zero only when $c = 0$.
\fussy On the other hand, if $T/n \rightarrow c > 0$, then $\theta_{n,T} \overset{p}\rightarrow \theta \ne 0$, and
The fact that $\sqrt{nT}( \widehat{\beta} - \beta)$ has a bias term in its asymptotic distribution does not imply $\widehat{\beta}$ is inconsistent. In fact, it is a consistent estimator of $\beta$ Alvarez2003.\footnote{This conclusion follows from ((ref)). This is because the limit distribution in ((ref)) implies that, for large $nT$, $\widehat{\beta}$ has an approximate normal distribution that is centered at $\beta +\theta/\sqrt{nT}$ with a variance of $(1-\beta)/(nT)$. Hence, the limit distribution of $\widehat{\beta}$, as $n,T \rightarrow \infty$, is degenerate at $\beta$, which implies $\widehat{\beta}$ converges in probability to $\beta$.} However, because the asymptotic distribution of $\sqrt{nT}( \widehat{\beta} - \beta)$ is not centered at zero, large sample confidence intervals and test statistics will be misleading. Specifically, the actual coverage of a confidence interval and the size of a test will differ from what they would be if the distribution of $\sqrt{nT}( \widehat{\beta} - \beta)$ were centered at zero.
The case considered by Alvarez and Arellano is instructive for two reasons. First, it illustrates that $\theta_{n,T} \overset{p}\rightarrow 0$, as $n,T \rightarrow \infty$, is required for the asymptotic distribution of $\sqrt{nT}( \widehat{\beta} - \beta)$ to be centered at zero. The second reason the example is instructive is less obvious and is the point of the first result (Theorem (ref)) provided here. The fact that $\theta_{n,T} \rightarrow \theta \ne 0$ when $c > 0$ is due to the number of instrumental variables that are used. Alvarez2003 assumed all available moment restrictions are exploited, in which case the maximum number of instrumental variables used in a period increases at the rate $T$ increases, and the total number of moment restrictions increases at the rate $T^2$ increases. It has long been known that using many moment restrictions has deleterious effects on the bias of a GMM estimator. However, Theorem (ref) sheds light on how many moment restrictions can be used without leading to large sample bias.
\sloppy Theorem (ref) provides conditions under which a generalization of $\theta_{n,T}$ converges in probability to a vector of zeros. Specifically, let $\boldsymbol{A}_{n,T} := (1/(nT)) \sum_{t=1}^{T-1}\ddot{\boldsymbol{X}}_t'\boldsymbol{P}_t \ddot{\boldsymbol{X}}_t $ and $\boldsymbol{b}_{n,T} := (1/\sqrt{nT}) \sum_{t=1}^{T-1}\ddot{\boldsymbol{X}}_t'\boldsymbol{P}_t \ddot{\boldsymbol{v}}_t$, with $ \ddot{\boldsymbol{v}}_{t} := c_{t} \left( \boldsymbol{v}_{t} - \overline{\boldsymbol{v}}_{t} \right) $, $\boldsymbol{v}_t' := \left(v_{1,t}, \ldots, v_{n,t} \right) $, and $ \overline{\boldsymbol{v}}_t := \left( 1/(T-t)\right)\sum_{s=1}^{T-t} \boldsymbol{v}_{t+s} $. Also, set $\boldsymbol{\theta}_{n,T} := \boldsymbol{A}_{n,T}^{-1}E(\boldsymbol{b}_{n,T})$. Then, analogous to ((ref)), we have
For every $n$ and $T$, the distribution of $\boldsymbol{b}_{n,T} - E(\boldsymbol{b}_{n,T})$ is centered at a vector of zeros, $\boldsymbol{0}$. It follows that if $\boldsymbol{A}_{n,T} \overset{p}\rightarrow \boldsymbol{A} > 0$,\footnote{The notation $\boldsymbol{B} >0$, when $\boldsymbol{B} $ is a matrix, means $\boldsymbol{B} $ is a positive definite matrix.} the distribution of $\sqrt{nT}\left(\widehat{\boldsymbol{\beta}} - \boldsymbol{\beta} \right) - \boldsymbol{\theta}_{n,T}$ becomes centered at $\boldsymbol{0}$ as the sample size grows. Therefore, whether or not the distribution of $\sqrt{nT}(\widehat{\boldsymbol{\beta}} - \boldsymbol{\beta} )$ is centered at $\boldsymbol{0}$ depends on whether or not $\boldsymbol{\theta}_{n,T} \overset{p}\rightarrow \boldsymbol{0}$. If the distribution of $\sqrt{nT}(\widehat{\boldsymbol{\beta}} - \boldsymbol{\beta} )$ is not centered at $\boldsymbol{0}$ in large samples, then the estimator $\widehat{\boldsymbol{\beta}}$ will henceforth be described as having asymptotic bias. Theorem (ref) shows that whether or not an FOD GMM estimator has asymptotic bias depends not just on the relative sizes of $n$ and $T$ but also on the number of instrumental variables used per period.
\fussy A few more definitions are needed in order to state Theorem (ref). Let $ \boldsymbol{w}_{i,t} $ be a column vector consisting of all of the distinct entries in $ \left\lbrace \left( \boldsymbol{x}_{i,s}', \boldsymbol{z}_{i,s}'\right) ; \; s = 1,\ldots, t \right\rbrace $. That is, $ \boldsymbol{w}_{i,t} $ contains all of the distinct values of the explanatory variables and instrumental variables for the $i$th individual from the first period up to the $ t $th period. Also, set $ \boldsymbol{u}_{i}' := (\boldsymbol{w}_{i,T}',v_{i,1},\ldots, v_{i,T},\eta_i ) $, and let $\gamma_{k,t,s} := \text{cov}\left( v_{1,t} , x_{1,t+s,k} |\boldsymbol{w}_{1,t} \right)$.
In addition to these definitions, Theorem (ref) relies on several conditions:
These conditions are weak enough to be satisfied by many dynamic regressions. For example, the estimation problems studied by Alvarez2003 and Hsiao2017, among others Bun2006,Okui2009, are special cases of the estimation problem examined here.
The first four assumptions are relatively straightforward. Nevertheless, it is worthwhile to review them to clarify what is---and what is not---required. For example, Assumption A1 does not require that the errors be conditionally homoskedastic or time series homoskedastic. Moreover, the i.i.d. assumption across $i$ can possibly be relaxed, but unfortunately relaxing the assumption would appear to require more technical detail.
The second assumption implies $n$ must increase with $T$ if the largest number of instrumental variables per period increases with $T$. Specifically, it implies $n \ge q_T^{\ast}$, and, therefore, if $q_T^{\ast}$ increases with $T$, then $n$ must likewise increase at least as fast as $q_T^{\ast}$ increases. On the other hand, Assumption A2 does not rule out cases where $n$ is fixed. For example, if at most $q$ instruments are used each period and $q$ does not increase with $T$, then $n$ can be fixed while $T$ increases provided $n \ge q$.
The third assumption implies the error in period $t$ ($v_{i,t}$) is uncorrelated with the current and past explanatory variables. The assumption also implies the current error is uncorrelated with the instruments used in period $t$. However, it imposes no additional restrictions on the instruments. The entries in $\boldsymbol{z}_{i,t}$ can consist of, but need not be restricted to, current and past explanatory variables. Moreover, they can be in levels, differenced, or transformed in some other way. Regardless of what instrumental variables are used or how they are constructed, Assumption A3 allows for dynamic panel data regressions with additional predetermined explanatory variables, provided the $v_{i,t}$s are uncorrelated across time.
\sloppy As for the fourth assumption, it implicitly entails some familiar conditions which are typically taken for granted. To see this, first note that Assumption A2 ensures $\left[(1/n)\boldsymbol{Z}_t'\boldsymbol{Z}_t\right]^{-1}$ is defined (wp 1). This, in turn, implies $\boldsymbol{\Psi}_{n,t} := (1/n)\ddot{\boldsymbol{X}}_t'\boldsymbol{Z}_t\left[(1/n)\boldsymbol{Z}_t'\boldsymbol{Z}_t\right]^{-1}(1/n)\boldsymbol{Z}_t'\ddot{\boldsymbol{X}}_t$ and $\boldsymbol{A}_{n,T} = (1/T)\sum_{t=1}^{T-1}\boldsymbol{\Psi}_{n,t}$ are defined. But in order for $\boldsymbol{\Psi}_{n,t}$ to be positive definite, the rank of $(1/n)\boldsymbol{Z}_t'\ddot{\boldsymbol{X}}_t$ must be $K$. That, in turn, requires the well-known necessary condition that $q_t \ge K$. In other words, it requires that we use at least as many instrumental variables each period as explanatory variables. Moreover, if the $\boldsymbol{\Psi}_{n,t}$s are positive definite, then the average of these positive definite matrices---that is, $\boldsymbol{A}_{n,T} $---is positive definite.\footnote{Not all of the $\boldsymbol{\Psi}_{n,t}$s need be positive definite in order for $\boldsymbol{A}_{n,T} $ to be positive definite.}. All of this is typically taken for granted, because if $\boldsymbol{A}_{n,T} $ is not positive definite, then $\widehat{\boldsymbol{\beta}}$ is not defined. Therefore, it is not unreasonable to assume the average of the $\boldsymbol{\Psi}_{n,t}$s ($\boldsymbol{A}_{n,T}$) converges (in probability), and that the limit ($\boldsymbol{A}$) is---like $\boldsymbol{A}_{n,T}$---positive definite. An example in which the condition is satisfied is provided by Alvarez2003.
\fussy The fifth assumption is a characterization of weak dependence. It is likely unfamiliar, but, like the other conditions, it does not appear to impose significant limitations on the types of panel data models to which the conclusion of Theorem 1 applies. Assumption A5 characterizes the linear association between the error in period $t$ and the $ k $th regressor $ s $ periods in the future relative to that period, conditional on the available information at time $t$. Because the conditional covariance between $ v_{1,t} $ and $ x_{1,t+s,k} $ is not restricted to be zero, Assumption A5 allows for predetermined regressors. Moreover, if the error term and future values of the explanatory variables are suitably weakly dependent, the bound on the sum of covariances in A5 will be satisfied.
This fact is illustrated by a stationary $K$th order autoregressive (AR($K$)) panel data model; see Lemma (ref). Proofs are provided in the appendix.
The AR(K) panel data model is just one example of the models that satisfy Assumption A5. In the Supplemental Appendix Phillips2024, I show that A5 is satisfied by a variety of models. These models include---but are not limited to---panel vector autoregressions and other panel data models with predetermined regressors. The panel data models considered by Alvarez2003, Bun2006, Hsiao2017, and Okui2009, among others, are special cases of models that satisfy A5.
Theorem (ref) can now be stated.
Theorem (ref) shows that the absence of asymptotic bias depends on the relative rate of increase in $n$ and $T$ and on the largest number of instrumental variables used per period ($q_T^{\ast}$). The maximum per-period number of instruments may depend on $T$. A bound on how fast $q_T^{\ast}$ increases with $T$ is parameterized in the theorem by $\alpha$.
The case $\alpha = 1$ has been widely studied in the literature. If all available instrumental variables are used, and past lags of explanatory variables are viable instruments, then the maximum number of instruments used in a period increases with $T$ at the rate $T$ increases. This corresponds to $\alpha = 1$. Then $T^{2\alpha}(\ln T)^2/(nT) = T(\ln T)^2/n$, which clearly does not go to zero for all sequences of $n$ and $T$. Instead, it only goes to zero for sequences of $n$ and $T$ for which either $T$ does not increase or it increases slowly enough relative to $n$ that $T(\ln T)^2/n \rightarrow 0$. The set of sequences for which $T(\ln T)^2/n \rightarrow 0$ is smaller than what Alvarez2003 found, for they found that there was no bias term if $T/n \rightarrow 0$. However, Alvarez2003 established their result for an AR(1) panel data model with i.i.d. homoskedastic errors. Theorem (ref), on the other hand, provides sufficient conditions for how fast $T$ can increase relative to $n$ for more general panel data regressions and under weaker assumptions about the errors.
Moreover, although $\boldsymbol{\theta}_{n,T} \overset{p}\rightarrow \boldsymbol{0}$ is only guaranteed for some sequences of $n$ and $T$ when $\alpha = 1$, we get $T^{2\alpha}(\ln T)^2/(nT) \rightarrow 0$ for all sequences of $n$ and $T$ if $\alpha < 1/2$. Hence, Theorem (ref) shows that $\sqrt{nT}(\widehat{\boldsymbol{\beta}} - \boldsymbol{\beta})$ has no asymptotic bias regardless of what happens to $T/n$, provided the maximum number of instrumental variables used in a period increases with $T$ more slowly than $T^{1/2}$ increases.
An important special case that satisfies the restriction $\alpha < 1/2$ is when the number of instrumental variables used per period never exceeds some fixed number, say $q$. Then, $\alpha = 0$. Thus, if a researcher uses a fixed number of instruments each period, the FOD GMM estimator has no asymptotic bias regardless of how $n$ and $T$ increase. In fact, although it is unusual for $n$ to be small while $T$ is large, Theorem (ref) nevertheless tells us there will be little bias in the distribution of $\sqrt{nT}(\widehat{\boldsymbol{\beta}} - \boldsymbol{\beta})$ even when $n$ is relatively small provided $T$ is large and $n \ge q$. This conclusion is similar to what is true for the fixed effects estimator when the explanatory variables are all strictly exogenous. That estimator, like the FOD GMM estimator, relies on eliminating fixed effects by subtracting from each variable a within cross-sectional average.
The absence of bias in the distribution of $\sqrt{nT}(\widehat{\boldsymbol{\beta}} - \boldsymbol{\beta})$ tells us only that. It does not tell us what distribution approximates the distribution of $\sqrt{nT}(\widehat{\boldsymbol{\beta}} - \boldsymbol{\beta})$ when $n$ and $T$ are both large. To address that question, Hsiao2017 studied a choice of $q_t$ for which $\alpha = 0$ in some detail. Specifically, they obtained asymptotic distribution results for FOD GMM estimators of the autoregression parameter in the AR(1) panel data model using a single instrumental variable per period. In doing so, they used sequential limits---specifically, they let $n \rightarrow \infty$, and then $T \rightarrow \infty$.\footnote{An alternative sequential-limit approach is to take limits in the order $T \rightarrow \infty$, then $n \rightarrow \infty$ (see, e.g., Phillips and Moon, 1999; Phillips and Moon, 2000; Kapetanios, 2008).}
Phillips and Moon remark that “Sequential limit theory is easy to derive and generally leads to quick results for a variety of model configurations” Moon1999. Moreover, sequential limit results can sometimes be obtained under weaker assumptions than those needed for joint limit results Moon1999, Moon2000. These advantages may explain why sequential limits have been used to analyze estimator properties when both $n$ and $T$ are large Kapetanios2008, Hsiao2015, Hsiao2017.
However, taking limits sequentially has a drawback: sequential limit results are not always robust. Specifically, a sequential limit is not, in general, guaranteed to be equal to a joint limit Moon2000.\footnote{Moon1999 provide conditions that, if satisfied, allows sequential limit results to be strengthened to joint limit results. However, their paper focuses on taking sequential limits with $T \rightarrow \infty$, and then $n \rightarrow \infty$. Moreover, even for limits are taken in this order, it can be difficult to verify the conditions that imply a sequential limit equals a joint limit.} Indeed, in the present application, letting $n \rightarrow \infty$ first, and then $T \rightarrow \infty$, always leads to the conclusion that there is no asymptotic bias in the distribution of $\sqrt{nT}(\widehat{\boldsymbol{\beta}} - \boldsymbol{\beta})$ regardless of $q_T^{\ast}$. This is illustrated by Theorem (ref)
Before stating Theorem (ref), some additional assumptions and definitions are needed. Specifically, assume
Assumption A6 implies the entries in the matrix $\boldsymbol{C}_t := E(\boldsymbol{z}_{1,t} \ddot{\boldsymbol{x}}_{1,t}^{\prime})$ are finite. Similarly, the entries in the matrices $\boldsymbol{Q}_t := E(\boldsymbol{z}_{1,t} \boldsymbol{z}_{1,t}^{\prime})$, and $E(\ddot{v}_{1,s}\ddot{v}_{1,t}\boldsymbol{z}_{1,s}\boldsymbol{z}_{1,t}')$ are finite. Moreover, given Assumption A2 implies $\boldsymbol{Q}_t$ is positive definite, if Assumptions A6 and A2 are both satisfied, we can define $\boldsymbol{\Sigma}_{t,t+s} := \boldsymbol{\Pi}_t^{\prime}E(\ddot{v}_{1,t}\ddot{v}_{1,t+s}\boldsymbol{z}_{1,t}\boldsymbol{z}_{1,t+s}')\boldsymbol{\Pi}_{t+s}$, where $\boldsymbol{\Pi}_t := \boldsymbol{Q}_t^{-1} \boldsymbol{C}_t$. Also, let $\boldsymbol{\Omega}_{T} = (1/T)\sum_{t=1}^{T-1}\boldsymbol{\Sigma}_{t,t} + (1/T)\sum_{t=1}^{T-2}\sum_{s=1}^{T-1-t}\left( \boldsymbol{\Sigma}_{t,t+s} + \boldsymbol{\Sigma}_{t+s,t}\right)$. Then the last set of assumptions for Theorem (ref) are
\fussy
The notation “$ ( n, T \rightarrow \infty)_{\text{seq}} $” means $ n \rightarrow \infty $, then $ T \rightarrow \infty $.
Theorem (ref) illustrates the problem with taking limits sequentially in the present context. Specifically, Theorem (ref) says the asymptotic distribution of $\sqrt{nT}(\widehat{\boldsymbol{\beta}} - \boldsymbol{\beta})$ is centered at $\boldsymbol{0}$, and it does so without referring to how many instrumental variables are used. However, as has already been noted, Alvarez2003 showed that if all instrumental variables are used, the FOD GMM estimator they considered has a nonzero bias term in its asymptotic distribution if $T/n \rightarrow c >0$, as $n, T \rightarrow \infty$. Alvarez and Arellano examined estimation of the AR(1) panel data model, which is a special case of the models covered by Theorem (ref). Therefore, Theorem (ref) contradicts the result provided in Alvarez2003.
The contradiction is not the fault of Alvarez and Arellano. Instead, it arises from how limits are executed. Taking a limit with $n \rightarrow \infty$ first, implicitly holds $T$ fixed initially, and GMM estimators do not generally have asymptotic bias, as $n \rightarrow \infty$, in the fixed $T$ case. Hence, the bias term is gone by the time the second step---which consists of letting $T \rightarrow \infty$---is reached.
On the other hand, under the conditions of Theorem (ref), there is also no bias term regardless of how $n$ and $T$ increase, provided the maximum number of per-period instruments, $q_T^{\ast}$, is selected so that $\alpha < 1/2$. In the next section, the accuracy of the normal approximation, for large $n$ and $T$ and $\alpha = 0$, is investigated with Monte Carlo experiments.
Phillips2020 provides Monte Carlo evidence that supports the claim that the distribution of $\sqrt{nT}(\widehat{\boldsymbol{\beta}} - \boldsymbol{\beta})$ is centered at $\boldsymbol{0}$ when, for example, $\alpha = 0$ and $T$ is not quite small compared to $n$. That paper reports Monte Carlo evidence illustrating that when the per-period number of instrumental variables is fixed, FOD GMM generally has less bias and is more efficient than one-step first difference GMM (FD GMM). In those experiments, the FOD GMM estimator was also less biased and usually more efficient than a two-step FD GMM in the presence of heteroskedastic errors, provided $T$ is not particularly small (e.g., $T=40$, $n=200$).\footnote{See Arellano1991 for a description of one-step and two-step FD GMM.}
This section, on the other hand, reports on Monte Carlo experiments that were used to investigate the reliability of confidence intervals based on the FOD GMM estimator. When $q_T^{\ast}$ was chosen so that $\alpha = 0$, the coverage of the confidence intervals based on the FOD GMM estimator turned out to be consistent with the conclusion of Theorem (ref). Specifically, for these experiments, the coverage of the FOD GMM confidence intervals, which relied on the normal approximation in Theorem (ref), was accurate for large $T$. In addition to confidence intervals, this section compares the estimator's precision, as measured by root mean squared error (RMSE), to the precision of the efficient GMM estimator.
Monte Carlo samples were generated using a sampling scheme similar to one of the sampling schemes described in Bun2006.\footnote{Bun2006 considered two sampling schemes, and sampling schemes similar to both were initially considered. However, the results were qualitatively similar across the two schemes. Therefore, for the sake of brevity, the results for only one of them are reported here.} In particular, the dependent variable in the regression model was generated according to
with $ \beta_1 \in \left\lbrace 0.25,0.75\right\rbrace $ and $ \beta_2 = 1 - \beta_1 $. The start-up value for $ y_{i,-50} $ was zero. Moreover, the error components $ v_{i,t} $ and $ \eta_i $ were generated independently of one another as i.i.d. standard normal variates. As for the explanatory variable $ x_{i,t} $, it was generated as
with $ \rho \in \left\lbrace 0.50, 0.95 \right\rbrace $, $ \kappa_1 \in \left\lbrace-1,0,+1 \right\rbrace $, and $ \phi_1 \in \left\lbrace-1,0,+1 \right\rbrace $. Moreover, the $ \varepsilon_{i,t} $s were generated as $ \varepsilon_{i,t} \overset{i.i.d.}{\sim} U(-\sqrt{12}/2,\,\sqrt{12}/2)$ independently of the $ v_{i,t} $s and $ \eta_i $s.
I considered a total of 36 experimental designs, where a design is a combination of parameter values. Table (ref) lists Designs 1 through 18. For these designs, I set $ \beta_1 = 0.25 $. Designs 19 through 36 were identical to Designs 1 through 18, except that $ \beta_1 $ was set to $0.75$ for Designs 19 through 36. The latter designs will henceforth be described as the weak instruments designs, for it is well-known that lagged values of the dependent variable become weaker instruments as $\beta_1$ approaches one Blundell1998.
\sloppy The effect of how each series was initialized was eliminated by dropping the first 50 values of each generated time series. As a result, estimation was based on the values $(x_{i,0},y_{i,0}),(x_{i,1},y_{i,1}),\ldots, (x_{i,T},y_{i,T}) $ ($ i = 1,\ldots,n $). Moreover, I set $n = 200$. The number of time periods, $T$, was set equal to 20, then 40, and finally 100 in order to examine the effect of increasing $T/n$ on confidence intervals and estimator precision. Finally, for each sample size and parameter design, 5,000 samples were generated. Estimates were calculated for each of these 5,000 samples.\footnote{The data were generated using GAUSS. Moreover, all calculations were performed with GAUSS.}
After a sample was generated, three estimates of $\boldsymbol{\beta} = (\beta_1, \beta_2)'$ were calculated: an FOD GMM estimate, an FD GMM estimate, and an efficient GMM estimate. For the FD and FOD GMM estimates, I set $ \boldsymbol{z}_{i,1}^{\prime} = (y_{i,0},x_{i,0},x_{i,1}) $ and $ \boldsymbol{z}_{i,t}^{\prime} = (y_{i,t-2},y_{i,t-1}, x_{i,t-2},x_{i,t-1},x_{i,t}) $ ($ t=2,\ldots,T-1 $, $ i=1,\ldots,n $). Given these instrumental variables, FOD GMM estimates were calculated using the formula in ((ref)).
The FD GMM estimates, on the other hand, were calculated using the formula
In this formula, $ \tilde{\boldsymbol{y}}_i' := (y_{i,2} - y_{i,1},\ldots,y_{i,T} - y_{i,T-1}) $; $ \tilde{\boldsymbol{X}}_i $ stacks $ \tilde{\boldsymbol{x}}_{i,t+1}' := \boldsymbol{x}_{i,t+1}' - \boldsymbol{x}_{i,t}' $, with $\boldsymbol{x}_{i,t}' = (y_{i,t-1},x_{i,t})$ ($ t=1,\ldots,T-1 $); and $\boldsymbol{Z}_{d,i}$ is defined in ((ref)). Also, $ \boldsymbol{G} $ is a $ (T-1) \times (T-1) $ matrix with twos running down the main diagonal, minus ones just above and below the main diagonal, and zeros everywhere else Arellano1991.
Finally, efficient GMM estimates were also calculated. The efficient GMM estimator exploited all available instrumental variables. For this estimator, I set $ \boldsymbol{z}_{i,1}^{\prime} = (y_{i,0},x_{i,0},x_{i,1}) $ and $ \boldsymbol{z}_{i,t}^{\prime} = (\boldsymbol{z}_{i,t-1}',y_{i,t-1},x_{i,t}) $ ($ t=2,\ldots,T-1 $, $ i=1,\ldots,n $). To calculate efficient estimates, either the formula in ((ref)) or the formula in ((ref)) can be used, because when all available instrumental variables are used, the two formulas give numerically identical results Phillips2019. Although ((ref)) and ((ref)) give the same estimate in this case, the formula in ((ref)) was used to calculate estimates because it is more efficient computationally Phillips2020. Moreover, efficient estimates were not calculated for $T=100$, because the requirement that $n$ be no smaller than $q_t$ was not satisfied---and hence $(\boldsymbol{Z}_t'\boldsymbol{Z}_t)^{-1}$ was not defined---for all $t$ when $T$ is this large relative to $n$ and all available instruments are used. Finally, because FD GMM and FOD GMM are the same when all available instruments are used, the efficient GMM estimator will henceforth be referred to as the FD/FOD GMM estimator.
\fussy In order to construct confidence intervals, standard errors were also calculated. Their calculation was simplified by the fact that the experimental $v_{i,t}$s were conditionally homoskedastic and uncorrelated. In this case, the variance-covariance matrix of $\sqrt{nT} (\widehat{\boldsymbol{\beta}}-\boldsymbol{\beta})$ simplifies to $\sigma^2 \boldsymbol{A}^{-1}$. Therefore, standard errors for the FOD GMM estimates were calculated by taking the square roots of the diagonal entries of $ \widehat{\sigma}^2\left( \sum_{t=1}^{T-1} \ddot{\boldsymbol{X}}_t'\boldsymbol{P}_t \ddot{\boldsymbol{X}}_t \right) ^{-1} $, where $ \widehat{\sigma}^2 :=\sum_{t=1}^{T-1}(\ddot{\boldsymbol{y}}_{t} - \ddot{\boldsymbol{X}}_{t}\widehat{\boldsymbol{\beta}})'(\ddot{\boldsymbol{y}}_{t} - \ddot{\boldsymbol{X}}_{t}\widehat{\boldsymbol{\beta}})/[n(T-1)] $. Moreover, when all available instruments were used, standard errors for FD/FOD estimates were calculated similarly. On the other hand, the standard errors for the FD estimator, $ \widetilde{\boldsymbol{\beta}} $, were calculated using the square roots of the diagonal entries in $ \widetilde{\sigma}^2 \left[\sum_{i=1}^n \tilde{\boldsymbol{X}}_i'\boldsymbol{Z}_{d,i}\left( \sum_{i=1}^n\boldsymbol{Z}_{d,i}'\boldsymbol{G}\boldsymbol{Z}_{d,i}\right)^{-1} \sum_{i=1}^n \boldsymbol{Z}_{d,i}'\tilde{\boldsymbol{X}}_i \right]^{-1} $, where $ \widetilde{\sigma}^2 := \sum_{i=1}^n( \tilde{\boldsymbol{y}}_{i} - \tilde{\boldsymbol{X}}_{i}\widetilde{\boldsymbol{\beta}})'( \tilde{\boldsymbol{y}}_{i} - \tilde{\boldsymbol{X}}_{i}\widetilde{\boldsymbol{\beta}}) /[2n(T-1)] $.
Tables (ref) through (ref) report the coverage of 95 percent confidence intervals for $\beta_1$ and $\beta_2$ using the FOD, FD, and FD/FOD GMM estimators, their respective standard errors, and the normal approximation---specifically, each interval extends from 1.96 standard errors below a regression parameter estimate to 1.96 standard errors above it. Tables (ref) and (ref) provide the coverage of the confidence intervals for $\beta_1$ and $\beta_2$, respectively, for Designs 1 through 18, and Tables (ref) and (ref) provide the coverage of the confidence intervals for $\beta_1$ and $\beta_2$ for Designs 19 through 36.
Coverage estimates for the FD/FOD confidence intervals are only provided for $T=20$. This is the best case for the FD/FOD GMM estimator because the estimator was developed for situations where $T$ is small compared to $n$. Nevertheless, even in this case the coverage of the FD/FOD confidence intervals is often unreliable. For example, consider the FD/FOD GMM intervals for $\beta_1$ given weak instruments (Table (ref)). For Designs 19 through 36, the coverage of the FD/FOD GMM intervals for $\beta_1$ ranged from a low of 51.8 percent (Design 27) to a high of 82.7 percent (Design 28). Unsurprisingly, increasing $T$ to 40 made for even worse coverage. This result is consistent with Alvarez and Arellano's finding of non-zero asymptotic bias in the distribution of the FOD GMM estimator of the AR(1) panel data model when $T$ is not small compared to $n$ and all available moment restrictions are exploited.
Using fewer than all available instrumental variables improved the reliability of the confidence intervals, as expected, but by how much depended on how the data were transformed to remove fixed effects. Usually, the coverage of the FOD confidence intervals approximated 95 percent at least as well, if not better, than did the FD confidence intervals. Moreover, confidence intervals based on the FOD GMM estimator were no less reliable for $T=100$ than for $T=20$. Indeed, the reliability of the FOD confidence intervals appears to improve as $T$ increases relative to $n$; see, for example, the FOD confidence intervals for $\beta_1$ and Designs 19 through 36 (Table (ref)). On the other hand, the FD confidence intervals often under estimated 95 percent, and their reliability did not improve as $T$ was increased relative to $n$; see, for example, Table (ref).\footnote{The finding that the coverage of the FD confidence intervals did not improve as $T$ was increased is consistent with analytical results Hsiao2017 provide for FD GMM estimation of the AR(1) panel data model.}
Moreover, using critical values from the standard normal distribution, 90 percent and 50 percent confidence intervals were also calculated. For the sake of brevity, the coverage of these intervals is omitted and is only summarized here: their reliability was similar to that of the 95 percent confidence intervals. For example, for large $n$ and $T$, the coverage of the 90 and 50 percent FOD confidence intervals approximated 90 and 50 percent, respectively. Apparently for these experiments, the FOD GMM estimator has little bias and the normal approximation works well when $n$ and $T$ are both large and $q=0$. On the other hand, the coverage of the 90 and 50 percent FD confidence intervals often under estimated 90 and 50 percent, especially for the weak instrument designs (Designs 19 through 36).
Although confidence intervals based on the FOD GMM estimator have superior coverage accuracy, the estimator sacrifices estimation efficiency to achieve this superior accuracy. In the experiments considered here, the FD and FOD GMM estimators use at most five instrumental variables each period. The FD/FOD GMM estimator, on the other hand, exploits all available moment restrictions. Therefore, the FD and FOD GMM estimators are inefficient relative to the FD/FOD GMM estimator. However, for $\alpha < 1/2$, the FOD GMM estimator does not have asymptotic bias, and, therefore, it can improve on the FD and FD/FOD GMM estimators in terms of precision, as measured by RMSE. The Monte Carlo experiments were also used to investigate this possibility.
\pgfplotsset{ myplotstyle/.style={ legend style={draw=none, font=}, legend cell align=left, legend pos=north east, ylabel style={align=center}, xlabel style={align=center}, scaled ticks=false, every axis plot/.append style={thick}, }, }
Figures (ref) through (ref) are line plots of relative precision estimates for the 36 designs. A relative precision estimate is an RMSE ratio. Specifically, an FD relative precision estimate is an FD RMSE divided by the corresponding FD/FOD RMSE, and a FOD relative precision estimate is an FOD RMSE divided by the FD/FOD RMSE. Figures (ref) through (ref) plot the FD and FOD GMM relative precision estimates for the 36 designs. Red line plots graph the FD relative precisions, whereas blue line plots graph the FOD relative precisions. Figures (ref) and (ref) provide relative precisions for $\beta_1$ for $T=20$ and $T=40$, whereas Figures (ref) and (ref) give the same for $\beta_2$. For all estimates, $n= 200$.
The plots make conclusions about relative precision obvious visually. If an estimator's relative precision plot is above one, the estimator is less precise than the FD/FOD GMM estimator in terms of RMSE. If the estimator's plot is below one, the estimator is more precise than the FD/FOD estimator.
The plots illustrate the effect of the absence, or presence, of bias on precision. For example, although the FOD GMM estimator of $\beta_1$ has a larger variance than the corresponding FD/FOD GMM estimator, the bias of the FOD estimator is less---in fact, much less, especially as $T$ increases relative to $n$. Hence, the FOD estimator of $\beta_1$ is usually more precise than the FD/FOD estimator. On the other hand, the FD/FOD GMM estimator of $\beta_2$ generally had smaller RMSE than the FOD GMM estimator of $\beta_2$ for $T=20$. The FOD estimator of $\beta_2$, however, became more competitive relative to the FD/FOD estimator in terms of RMSE for $T=40$. Again, finite sample bias, or its absence, provides the explanation. The finite sample bias of the FD/FOD GMM estimator was less pronounced when estimating $\beta_2$ than it was when $\beta_1$ was estimated, especially for $T=20$.
On the other hand, the FD GMM estimators of $\beta_1$ and $\beta_2$ have larger variances than the FD/FOD GMM estimators, but, unlike the FOD estimators, they do not benefit from having much less bias than the FD/FOD estimators. Consequently, the FD estimators of $\beta_1$ and $\beta_2$ are less precise than the FD/FOD estimators of $\beta_1$ and $\beta_2$.
When using panel data to estimate a dynamic regression, researchers typically remove fixed effects by transforming the data. Many transformations are available, but---due to historical precedence---differencing the data became a widely adopted approach. However, when not all of the available instrumental variables are used, how the data are transformed matters, and differencing may not be the best transformation.
The results provided in this paper show that transforming the data using forward orthogonal deviations produces a GMM estimator, $\widehat{\boldsymbol{\beta}}$, such that whether or not the distribution of $\sqrt{nT}(\widehat{\boldsymbol{\beta}} - \boldsymbol{\beta})$ is centered at a vector of zeros, as $n,T \rightarrow \infty$, depends on how fast the largest number of instruments used in a period increases with $T$. For the absence of large sample bias, this number must increase more slowly than $T^{1/2}$ increases. This observation is important because the reliability of large sample confidence intervals and test statistics based on an FOD GMM estimator depends on the distribution of $\sqrt{nT}(\widehat{\boldsymbol{\beta}} - \boldsymbol{\beta})$ being centered at a vector of zeros.