EconBase
← Back to paper

A Comparison of First-Difference and Forward Orthogonal Deviations GMM

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

65,689 characters · 9 sections · 29 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Comparison of First-Difference and Forward Orthogonal Deviations GMM

abstractThis paper provides a necessary and sufficient instruments condition assuring two-step generalized method of moments (GMM) based on the forward orthogonal deviations transformation is numerically equivalent to two-step GMM based on the first-difference transformation. The condition also tells us when system GMM, based on differencing, can be computed using forward orthogonal deviations. Additionally, it tells us when forward orthogonal deviations and differencing do not lead to the same GMM estimator. When estimators based on these two transformations differ, Monte Carlo simulations indicate that estimators based on forward orthogonal deviations have better finite sample properties than estimators based on differencing.

Keywords: system GMM; first-difference GMM; Arellano-Bond GMM; forward orthogonal demeaning; forward orthogonal deviations

Introduction

A popular method for removing time-invariant effects from panel data is to first difference the data. Arellano1991, for example, proposed one-step and two-step first-difference generalized method of moments --- henceforth FD-GMM --- estimators. Later Bover1995 and Blundell1998 showed how equations in differences and equations in levels could be estimated as a system with generalized method of moments --- system GMM. Since that time, FD-GMM and system GMM have become dominant estimation approaches in the literature on panel data estimation.

The dominance of the first-difference transformation may, in part, be attributed to invariance results in Schmidt1992 and Bover1995. These results say that, under suitable restrictions, two different transformations can lead to the same generalized method of moments (GMM) estimator. But if two different transformations lead to the same estimator, why bother with another transformation that gives the same result as the first-difference transformation? Hence, differencing may be all we need.

It turns out, however, significant computational advantages may be possible using another transformation (see, e.g., Arellano and Bover, 1995; Phillips, 2019b). Moreover, it is not always the case that different transformations lead to the same GMM estimator. In particular, invariance to transformation depends on the instruments. Schmidt1992 and Bover1995 focused on efficient estimation, and, in doing so, identified using all available instruments as sufficient for invariance to transformation conclusions. But using all available instruments is not necessary for an invariance to transformation result Phillips2019a. Phillips2019a, on the other hand, provided a sufficient {\em and necessary} instruments condition that assures a GMM estimator can be calculated using two-stage least squares (2SLS) after filtering the data.

But the result in Phillips2019a does not cover GMM when optimal weighting is used in the presence of conditional heteroskedasticity. This paper examines that case. It shows that the condition on the instruments identified in Phillips2019a is necessary and sufficient for two-step FD-GMM to be equivalent to two-step GMM based on the forward orthogonal deviations transformation --- henceforth, two-step FOD-GMM. In fact, the result provided in this paper applies more generally than to the first difference and forward orthogonal deviations transformations. All we need assume about the transformation is that the transformation matrix $ \boldsymbol{K} $ that sweeps out the time-invariant effects is such that $\boldsymbol{K}\boldsymbol{K}^{\prime}$ is a positive definite matrix. Moreover, I show that if the instruments condition is met, and only if the instruments condition is met, then the well-known system GMM estimator (Arellano and Bover, 1995; Blundell and Bond, 1998) can be calculated using the forward orthogonal deviations transformation rather than the first-difference transformation.

The necessity of the instruments condition tells us that, if a choice for instruments does not satisfy the condition, two different transformations of the data cannot lead to the same GMM estimator. For example, experience has taught researchers that first-differencing and forward orthogonal deviations do not lead to the same GMM estimator when only recent lags are used as instrumental variables. The reason this is true is because the instruments condition is not satisfied Phillips2019a.

But when different transformations must lead to different GMM estimators, as when only recent lags are used as instruments, the relevant question then becomes which transformation leads to the better GMM estimator? This question has received some attention in the literature; see Hay2009, Hsiao2017, and Phillips2019b. Hsiao2017 compared the asymptotic properties of method of moments estimators based on differencing the data versus using forward orthogonal deviations. That paper also provides some Monte Carlo evidence on the finite sample behavior of method of moments estimators based on differencing and on forward orthogonal deviations. Hay2009 examined the finite sample behavior of one-step GMM based on the forward orthogonal deviations transformation --- one-step FOD-GMM --- and one-step FD-GMM. He found that one-step FOD-GMM compared favorably to one-step FD-GMM. And Phillips2019b found that one-step FOD-GMM also outperformed two-step FD-GMM when the length of the time-series ($T$) is not small.

In this paper I compare the finite sample properties of two-step FOD-GMM to two-step FD-GMM with Monte Carlo experiments. I also investigate the finite sample properties of a system GMM estimator that exploits the forward orthogonal deviations transformation and compare its sampling behavior to that of the usual system GMM estimator, which relies on differencing. I find that the estimators based on forward orthogonal deviations dominate their counterparts based on differencing. They generally have smaller absolute bias and their standard deviations are almost always smaller.

The next section provides numerical equivalence results for GMM based on different transformations. Section (ref) provides the Monte Carlo evidence, and Section (ref) concludes. Proofs are relegated to Section (ref).

Numerically equivalent transformations

When panel data are used, the data are often transformed in order to remove time-invariant effects. Specifically, consider the model

equation[equation omitted — 174 chars of source]

where $\boldsymbol{X}_{i}$ is a matrix of observations on explanatory variables, $\boldsymbol{v}_{i}$ is a vector of errors that vary with time and individual, $\boldsymbol{\iota}$ is a vector of ones, and $\eta_i$ is an unobserved time-invariant effect. The time-invariant effect $\eta_i$ can be removed by premultiplying through ((ref)) by a transformation matrix $\boldsymbol{K}$ that satisfies $\boldsymbol{K}\boldsymbol{\iota}=\boldsymbol{0}$.

Moreover, if $\boldsymbol{K}\boldsymbol{K}^{\prime}$ is a positive definite matrix, there exists another transformation that yields exactly the same estimator if, and only if, any instrument used in period $s$ can be constructed from a linear combination of instruments used for period $t$, for every $t \geq s$. This condition is satisfied in the well-known case where the instruments consist of lagged predetermined variables and all available instruments are used (see, e.g., Arellano, 2003, p. 153; Phillips, 2019a). But other instrument choices also satisfy the instruments condition. The condition is satisfied whenever an instrument used in an earlier period is included, somehow, in the list of instruments used in a later period. That possibility allows for many different choices of instruments that satisfy the instruments condition; see Phillips2019a for examples.

The numerically equivalent transformation is provided in Theorem 1.

theoremLet $\boldsymbol{z}_{it}$ be a $k_{t} \times 1$ vector of instruments $(t=1,\ldots ,R)$. Also, let \begin{equation} \boldsymbol{Z}_{i}=\left( \begin{array}{cccc} \boldsymbol{z}_{i1}^{\prime } & \boldsymbol{0} & \cdots & \boldsymbol{0} \\ \boldsymbol{0} & \boldsymbol{z}_{i2}^{\prime } & \cdots & \boldsymbol{0} \\ \vdots & \vdots & \ddots & \vdots \\ \boldsymbol{0} & \boldsymbol{0} & \cdots & \boldsymbol{z}_{iR}^{\prime } \end{array} \right) . \end{equation} Moreover, let $\boldsymbol{K}$ be such that $\boldsymbol{K}\boldsymbol{\iota}=\boldsymbol{0}$ and $\boldsymbol{K}\boldsymbol{K}^{\prime}$ is positive definite. Let $\widehat{\boldsymbol{\beta}}$ be an initial estimator of $\boldsymbol{\beta}$ and set $\boldsymbol{e}_i=\boldsymbol{y}_i-\boldsymbol{X}_i\widehat{\boldsymbol{\beta}}$ $(i=1,\ldots,N)$. Furthermore, set $\boldsymbol{F}=\boldsymbol{U}\boldsymbol{K}$, where $\boldsymbol{U}$ is the upper-triangular Cholesky factorization of $(\boldsymbol{K}\boldsymbol{K}^{\prime})^{-1}$. Next, let $\boldsymbol{\tilde{y}}_i=\boldsymbol{K}\boldsymbol{y}_i$, $\boldsymbol{\tilde{X}}_i=\boldsymbol{K}\boldsymbol{X}_i$, and $\boldsymbol{\tilde{e}}_i=\boldsymbol{K}\boldsymbol{e}_i$ ($i=1,\ldots,N$). Also, set $\boldsymbol{\ddot{y}}_i=\boldsymbol{F}\boldsymbol{y}_i$, $\boldsymbol{\ddot{X}}_i=\boldsymbol{F}\boldsymbol{X}_i$, and $\boldsymbol{\ddot{e}}_i=\boldsymbol{F}\boldsymbol{e}_i$ ($i=1,\ldots,N$). Finally, define \begin{eqnarray} \widehat{\boldsymbol{\beta}}_K&=&\left[\sum_i\boldsymbol{\tilde{X}}_i^{\prime}\boldsymbol{Z}_i\left(\sum_i\boldsymbol{Z}_i^{\prime}\boldsymbol{\tilde{e}}_i\boldsymbol{\tilde{e}}_i^{\prime}\boldsymbol{Z}_i\right)^{-1}\sum_i\boldsymbol{Z}_i^{\prime}\boldsymbol{\tilde{X}}_i\right]^{-1} \nonumber \\ & & \times \sum_i\boldsymbol{\tilde{X}}_i^{\prime}\boldsymbol{Z}_i\left(\sum_i\boldsymbol{Z}_i^{\prime}\boldsymbol{\tilde{e}}_i\boldsymbol{\tilde{e}}_i^{\prime}\boldsymbol{Z}_i\right)^{-1}\sum_i\boldsymbol{Z}_i^{\prime}\boldsymbol{\tilde{y}}_i \end{eqnarray} and \begin{eqnarray} \widehat{\boldsymbol{\beta}}_F&=&\left[\sum_i\boldsymbol{\ddot{X}}_i^{\prime}\boldsymbol{Z}_i\left(\sum_i\boldsymbol{Z}_i^{\prime}\boldsymbol{\ddot{e}}_i\boldsymbol{\ddot{e}}_i^{\prime}\boldsymbol{Z}_i\right)^{-1}\sum_i\boldsymbol{Z}_i^{\prime}\boldsymbol{\ddot{X}}_i\right]^{-1} \nonumber \\ & & \times \sum_i\boldsymbol{\ddot{X}}_i^{\prime}\boldsymbol{Z}_i\left(\sum_i\boldsymbol{Z}_i^{\prime}\boldsymbol{\ddot{e}}_i\boldsymbol{\ddot{e}}_i^{\prime}\boldsymbol{Z}_i\right)^{-1}\sum_i\boldsymbol{Z}_i^{\prime}\boldsymbol{\ddot{y}}_i. \end{eqnarray} Then $\widehat{\boldsymbol{\beta}}_F=\widehat{\boldsymbol{\beta}}_K$ if, and only if, every entry in $\boldsymbol{z}_{is}$ is a linear combination of entries in $\boldsymbol{z}_{it}$ $(s = 1,\ldots, t$, $ t=1,\ldots ,R)$.

See Section (ref) for a proof.

An important special case of Theorem (ref) is a first-differenced panel data model. In this case, $\boldsymbol{K}=\boldsymbol{D}$, where

equation*[equation* omitted — 233 chars of source]

Given $\boldsymbol{K}=\boldsymbol{D}$, the appropriate $\boldsymbol{F}$ is the forward orthogonal deviations transformation matrix given by

eqnarray[eqnarray omitted — 605 chars of source]

(see Arellano, 2003, p. 17). It is well-known that if the errors in $ \boldsymbol{v}_i $ are conditionally homoskedastic and uncorrelated, then the forward orthogonal deviation errors in $ \boldsymbol{\ddot{v}}_i = \boldsymbol{F}\boldsymbol{v}_i $ are conditionally homoskedastic and uncorrelated. However, even if the entries in $ \boldsymbol{v}_i $ are not conditionally homoskedastic and uncorrelated, according to Theorem (ref), FD-GMM, based on its optimal weighting matrix, is equivalent to FOD-GMM, based on its optimal weighting matrix, if, and only if, every instrument used in period $s$ can be constructed as a linear combination of instruments used in period $t$, for $t \geq s$.

In addition to being able to compute FD-GMM estimates in an alternative manner, the system GMM estimator has an alternative representation, provided the instruments condition is met. As is well-known, the usual system GMM estimator uses both differenced and levels data. However, system GMM estimates can alternatively be calculated using levels data and forward orthogonal deviations if, and only if, the instruments condition in Theorem (ref) is satisfied.

To establish this claim, consider the model

equation*[equation* omitted — 164 chars of source]

Under suitable conditions, $\boldsymbol{\beta}=(\delta,\boldsymbol{\alpha}^{\prime})^{\prime}$ can be estimated with the system GMM estimator studied by Bover1995 and Blundell1998.

In order to write that estimator, let $\boldsymbol{y}_i = (y_{i1},\ldots,y_{iT})^{\prime}$, and let $\boldsymbol{X}_i$ denote a $T \times K$ matrix with $(y_{i,t-1},\boldsymbol{x}_{it}^{\prime})$ in its $t$th row ($t=1,\ldots, T$). Next set $\boldsymbol{y}_i^+ = (\boldsymbol{y}_i^{\prime},\boldsymbol{y}_i^{\prime})^{\prime}$ and $\boldsymbol{X}_i^+ = (\boldsymbol{X}_i^{\prime},\,\boldsymbol{X}_i^{\prime})^{\prime} $. The usual system GMM estimator relies on differencing the observations in the first $T$ rows in $\boldsymbol{y}_i^+$ and $\boldsymbol{X}_i^+$. Specifically, it uses the transformed data $\boldsymbol{\tilde{y}}_i^+=\boldsymbol{K}^+\boldsymbol{y}_i^+$ and $\boldsymbol{\tilde{X}}_i^+=\boldsymbol{K}^+\boldsymbol{X}_i^+$ ($i=1,\ldots,N$), where

equation[equation omitted — 163 chars of source]

For the instrument matrix, let $\boldsymbol{Z}_{1i}$ and $\boldsymbol{Z}_{2i}$ be block-diagonal instrument matrices, where $\boldsymbol{Z}_{1i}$ has $1\times k_t$ instrument vector $\boldsymbol{z}_{it}^{\prime}$ in its $t$th diagonal block ($t=1,\ldots,T-1$) and $\boldsymbol{Z}_{2i}$ has $\boldsymbol{x}_{i1}^{\prime}-\boldsymbol{x}_{i0}^{\prime}$ in its first diagonal block and $({y}_{i,t-1}-y_{i,t-2},\boldsymbol{x}_{it}^{\prime}-\boldsymbol{x}_{i,t-1}^{\prime})$ in diagonal blocks $t=2,\ldots,T$. Next, set

equation*[equation* omitted — 165 chars of source]

Given this transformation matrix and the preceding notation, the system GMM estimator can be expressed as

eqnarray[eqnarray omitted — 610 chars of source]

where $\boldsymbol{\tilde{e}}_i^+=\boldsymbol{\tilde{y}}_i^+ - \boldsymbol{\tilde{X}}_i^+\widehat{\boldsymbol{\beta}}$ ($i=1,\ldots,N$) and $\widehat{\boldsymbol{\beta}}$ is an initial estimator of $\boldsymbol{\beta}$.

Alternatively, if the instruments condition is satisfied, the same system GMM estimator can be constructed using forward orthogonal deviations rather than first differences. In this case, the transformation matrix is

equation[equation omitted — 163 chars of source]

where $\boldsymbol{F}$ is the forward orthogonal deviations transformation matrix given by Eq. ((ref)). Now let $\boldsymbol{\ddot{y}}_i^+=\boldsymbol{F}^+\boldsymbol{y}_i^+$ and $\boldsymbol{\ddot{X}}_i^+=\boldsymbol{F}^+\boldsymbol{X}_i^+$ ($i=1,\ldots,N$). Then define

eqnarray[eqnarray omitted — 602 chars of source]

where $\boldsymbol{\ddot{e}}_i^+=\boldsymbol{\ddot{y}}_i^+ - \boldsymbol{\ddot{X}}_i^+\widehat{\boldsymbol{\beta}}$ ($i=1,\ldots,N$).

We can now state Theorem (ref).

theoremSuppose $\widehat{\boldsymbol{\beta}}_{K^+}$ and $\widehat{\boldsymbol{\beta}}_{F^+}$ use the same initial estimator $\widehat{\boldsymbol{\beta}}$. Then $\widehat{\boldsymbol{\beta}}_{F^+}=\widehat{\boldsymbol{\beta}}_{K^+}$ if, and only if, every entry in $\boldsymbol{z}_{is}$ is a linear combination of entries in $\boldsymbol{z}_{it}$ $(s = 1,\ldots, t$, $ t=1,\ldots ,T-1)$.

The proof is provided in Section (ref).

Theorem (ref) applies to an important case. Specifically, it tells us when system GMM based on first differences is equivalent to system GMM based on forward orthogonal deviations. But the result holds more generally. In particular, we can replace $\boldsymbol{D}$ in the definition of $\boldsymbol{K}^+$ with another transformation matrix $\boldsymbol{K}$, provided $\boldsymbol{K}\boldsymbol{K}^{\prime}$ is positive definite and provided $\boldsymbol{F}$ in $\boldsymbol{F}^+$ is given by $\boldsymbol{F}=\boldsymbol{U}\boldsymbol{K}$, where $\boldsymbol{U}$ is the upper-triangular Cholesky factorization of $(\boldsymbol{K}\boldsymbol{K}^{\prime})^{-1}$.

When only recent lags are used as instruments

Theorem (ref) not only tells us when two different transformations lead to the same GMM estimator, it also tells us when they do not. For example, it tells us we cannot use two different transformations to get the same GMM estimator for a popular choice for instruments --- specifically, when only recent lags of predetermined variables are used as instruments. This is because, if only recent lags are used as instrumental variables, a lagged predetermined variable that is used as an instrument in an earlier period will not be used as an instrument in some later period, and consequently we cannot construct the instrument used in the earlier period from the instruments used in the later period. In other words, the instruments condition is violated. Similarly, Theorem (ref) tells us that when only recent lags of predetermined variables are used as instruments, the system GMM estimator based on first differences --- the FD-SYS estimator --- is not the same as the system GMM estimator that exploits forward orthogonal deviations --- the FOD-SYS estimator.

But these observations raise some questions. When two-step FOD-GMM is not the same as two-step FD-GMM, yet both rely on the same choice of instruments and both are based on their respective optimal weighting matrices, which estimator is the better choice? Also, when FD-SYS and FOD-SYS estimators are not the same, which system estimator should we use?

This section addresses these questions with Monte Carlo experiments.

Monte Carlo simulations

The experiments conducted for this paper are similar to those used in Phillips2019b. This allows the reader to compare the finite sample behavior of the estimators examined here to that of the estimators studied in Phillips2019b.

For all of the Monte Carlo simulations, data were generated according to the model

equation*[equation* omitted — 131 chars of source]

Moreover, the $x_{it}$s were generated as predetermined variables:

equation*[equation* omitted — 153 chars of source]

The start-up values $y_{i,-50}$ and $x_{i,-50}$ were set as $y_{i,-50}=0$ and $x_{i,-50}=5+10\xi _{i,-50}$. Moreover, start-up observations were discarded. In particular, for each sample, estimation was based on the $T+1$ observations $\left( x_{i0},y_{i0}\right) ,\ldots ,\left( x_{iT},y_{iT}\right) $ ($i=1,\ldots ,N$), with $N$ always set to 200 and $T$ set to either 10 or 30. For each sample size and combination of parameters, 10,000 independent samples were drawn.

In order to generate a sample, the parameters had to be specified and pseudo random numbers were generated. For the parameters, I set $\alpha=0.5$, and $\delta$ was either 0.5 or 0.9, while $\rho$ was either 0.3 or 0.8. As for the pseudo random variates, the $\xi _{it}$s were generated as independent uniform random variates with mean zero and variance one. Moreover, the individual-specific effects --- the $\eta _{i}$s --- were generated independently of the $\xi _{it}$s and $v_{it}$s as $\eta _{i}=\sigma _{\eta }\zeta _{i}$ ($i=1,\ldots,N$), with $ \zeta _{i}$ a standard normal random variable. Two values for the standard deviation $\sigma _{\eta }$ were considered: one or four.

Moreover, two models were used to generate the $v_{it}$s: a conditionally heteroskedastic errors model and a time-series heteroskedastic errors model. For conditionally heteroskedastic errors, I set $v_{it}=x_{it} \epsilon _{it}$, with $\epsilon _{it}$ a standard normal random variable, which was generated independently of $ x_{it}$, $x_{i,t-s}$, and $\epsilon _{i,t-s}$ for $s\geq 1$. For time-series heteroskedastic errors, $\lambda_{t}$s ($t=1, \ldots, T$) were first generated as uniform random variates with mean zero and variance one. Then I set $v_{it} = \lambda_{t} \epsilon_{it}$.

Results

To get some sense of magnitudes, Table (ref) provides bias, standard deviation, and root mean squared error estimates for the two-step FD-GMM estimator.\footnote{All computations were performed using GAUSS.} The results in Table (ref) are for an estimator that uses only recent lags of the predetermined explanatory variables as instruments. In particular, for $ \boldsymbol{z}_{it} $, I set $ \boldsymbol{z}_{i1}^{\prime} = (y_{i0},x_{i0},x_{i1}) $ and $ \boldsymbol{z}_{it}^{\prime} = (y_{i,t-2},y_{i,t-1}, x_{i,t-2},x_{i,t-1},x_{it}) $ ($ t=2,\ldots,T-1 $). To compute two-step FD-GMM estimates I used the formula in ((ref)) with $\boldsymbol{K} = \boldsymbol{D}$. These estimates require that one-step estimates first be calculated, and for those I used the formula in ((ref)) with the weighting matrix $ \left(\sum_i\boldsymbol{Z}_i^{\prime}\boldsymbol{\tilde{e}}_i\boldsymbol{\tilde{e}}_i^{\prime}\boldsymbol{Z}_i\right)^{-1} $ replaced by $\left(\sum_i\boldsymbol{Z}_i^{\prime}\boldsymbol{D}\boldsymbol{D}^{\prime}\boldsymbol{Z}_i\right)^{-1}$.

Table (ref) shows how the two-step FOD-GMM estimator compares to the two-step FD-GMM estimator when only recent lags of predetermined explanatory variables are used as instruments.\footnote{The two-step FOD-GMM estimator is based on the same instruments as the FD-GMM estimator. The two-step FOD-GMM estimator is given by ((ref)) with $ \boldsymbol{F} $ given by the transformation matrix in ((ref)). Like two-step FD-GMM estimates, two-step FOD-GMM estimates require one-step estimates first be calculated. For the two-step FOD-GMM estimates, the one-step estimates were one-step FOD-GMM estimates.} The table gives the percent reduction in absolute bias, standard deviation, and root mean squared error from using the FOD-GMM estimator rather than the FD-GMM estimator. Specifically, the estimates in Table (ref) were calculated as

equation*[equation* omitted — 40 chars of source]

where $ FD $ and $ FOD $ stand for the absolute bias, standard deviation, or root mean squared error of the two-step FD-GMM estimator and the two-step FOD-GMM estimator, respectively.

It is clear from the data in Table (ref) that for the vast majority of sample designs the FOD-GMM estimator has smaller absolute bias than the FD-GMM estimator. There are only eight cases for which the FD-GMM estimator has the smaller absolute bias --- i.e., cases for which the percent reduction in bias is negative. And, for most of those cases, the bias of both estimators is small relative to their standard deviations. This is obvious from the fact that even though the FOD-GMM estimator has larger bias in these cases, for most of these cases its percentage reduction in root mean squared error is similar to its percentage reduction in standard deviation, which indicates that the standard deviations contribute more to the root mean squared errors than the biases of the two estimators.

The data in Table (ref) also reveal that the two-step FOD-GMM estimator is almost always the more efficient estimator. The percent reduction in standard deviation is almost always positive. There is only one case for which it is not positive. And, for that case, the FOD-GMM estimator loses little efficiency relative to the FD-GMM estimator. This fact and the fact that the FOD-GMM estimator usually has smaller bias than the FD-GMM estimator implies that, when precision is measured in terms of root mean squared error, the FOD-GMM estimator is always the more precise estimator.

Table (ref) provides bias, standard deviation, and root mean squared error estimates for the FD-SYS estimator. This estimator is given by Eq. ((ref)) with $ \boldsymbol{K}^+ $ given by Eq. ((ref)). The results in Table (ref) are limited to $T = 10$ because for $T=30$ there are so many moment restrictions that $ \sum_i\boldsymbol{Z}_i^{+ \prime}\boldsymbol{\ddot{e}}_i^+\boldsymbol{\ddot{e}}_i^{+ \prime}\boldsymbol{Z}_i^+ $ is singular, and, therefore, the optimal weighting matrix cannot be computed.

Table (ref) reports the percent reduction in absolute bias, standard deviation, and root mean squared error provided by the FOD-SYS estimator. The FOD-SYS estimator is given by Eq. ((ref)) with $ \boldsymbol{F}^+ $ given by Eq. ((ref)). The message of Table (ref) is similar to that of Table (ref). The FOD-SYS estimator typically has smaller absolute bias, it is always more efficient, and it also always has the smaller root mean squared error.

table[table omitted — 8,785 chars of source]
table[table omitted — 8,527 chars of source]
table[table omitted — 4,828 chars of source]
table[table omitted — 4,721 chars of source]

Summary

This paper showed that the necessary and sufficient instruments condition provided in Phillips2019a applies to two-step GMM estimation based on heteroskedasticity-robust weighting matrices. If the condition is satisfied, a two-step GMM estimator, based on an optimal weighting matrix, can be calculated using another transformation and the optimal weighting matrix corresponding to the alternative transformation. The paper also showed when the system GMM estimator studied by Bover1995 and Blundell1998 can be computed using forward orthogonal deviations rather than first differencing.

Because the instruments condition is not just sufficient but also necessary, it tells us when GMM estimators are not invariant to transformation. One situation for which invariance to transformation is not possible is when only recent lags of predetermined variables are used as instrumental variables.

Monte Carlo experiments were used to examine two important cases: two-step FD-GMM estimation versus two-step FOD-GMM estimation and FD-SYS estimation versus FOD-SYS estimation. When these GMM estimators exploited only recent lags of predetermined variables as instruments, the estimators based on forward orthogonal deviations were generally less biased and almost always more efficient than their counterparts based on first differencing.

Proofs

The proof of Theorem (ref) relies on a corollary to Theorem 1 in Phillips2019a. That corollary is stated as Lemma (ref) here.

lemmaLet $\boldsymbol{U}$ be the upper-triangular Cholesky factorization of $(\boldsymbol{K}\boldsymbol{K}^{\prime})^{-1}$, and let $\boldsymbol{Z}_i$ be defined as in ((ref)). Then there is a nonsingular matrix $\boldsymbol{C}$ satisfying $\boldsymbol{C}\boldsymbol{Z}_i^{\prime} = \boldsymbol{Z}_i^{\prime}\boldsymbol{U}$ if, and only if, every entry in $\boldsymbol{z}_{is}$ is a linear combination of the entries in $\boldsymbol{z}_{it}$ ($s =1,\ldots, t$, $t=1,\ldots,R$).

Proof of Lemma (ref): The conclusion of the lemma follows from Theorem 1 in Phillips2019a. To apply that theorem, let $\boldsymbol{\Phi} = \boldsymbol{K}\boldsymbol{\Omega}\boldsymbol{K}^{\prime}$, where $\boldsymbol{\Omega}$ is a positive definite matrix. Theorem 1 in Phillips2019a says that for the upper-triangular Cholesky factorization of $\boldsymbol{\Phi}^{-1}$, say $\boldsymbol{U}^{\ast}$, there is a nonsingular matrix $\boldsymbol{C}$ satisfying $\boldsymbol{C}\boldsymbol{Z}_i^{\prime} = \boldsymbol{Z}_i^{\prime}\boldsymbol{U}^{\ast}$ if, and only if, every entry in $\boldsymbol{z}_{is}$ is a linear combination of the entries in $\boldsymbol{z}_{it}$ ($s = 1,\ldots, t$, $t=1,\ldots,R$). Set $\boldsymbol{\Omega}=\boldsymbol{I}$. Then $ \boldsymbol{U}^{\ast} = \boldsymbol{U} $, and the conclusion of Lemma (ref) follows. \\

Proof of Theorem (ref)

Under the conditions of the theorem, we have by Lemma (ref) that there is a nonsingular matrix $\boldsymbol{C}$ such that $\boldsymbol{C}\boldsymbol{Z}_i^{\prime}=\boldsymbol{Z}_i^{\prime}\boldsymbol{U}$ if, and only if, every entry in $\boldsymbol{z}_{is}$ is a linear combination of entries in $\boldsymbol{z}_{it}$ $(s = 1, \ldots, t$, $ t=1,\ldots ,R)$. This fact and $\boldsymbol{F}=\boldsymbol{U}\boldsymbol{K}$ imply

eqnarray[eqnarray omitted — 1,237 chars of source]

if, and only if, every entry in $\boldsymbol{z}_{is}$ is a linear combination of entries in $\boldsymbol{z}_{it}$ $(s = 1, \ldots, t$, $ t=1,\ldots ,R)$. By similar reasoning, we get

equation[equation omitted — 559 chars of source]

if, and only if, every entry in $\boldsymbol{z}_{is}$ is a linear combination of entries in $\boldsymbol{z}_{it}$ $(s = 1, \ldots, t$, $ t=1,\ldots ,R)$.

Proof of Theorem (ref)

By Lemma (ref), there is a nonsingular matrix $\boldsymbol{C}$ satisfying $\boldsymbol{C}\boldsymbol{Z}_{1i}^{\prime}=\boldsymbol{Z}_{1i}^{\prime}\boldsymbol{U}$ if, and only if, every entry in $\boldsymbol{z}_{is}$ is a linear combination of entries in $\boldsymbol{z}_{it}$ $(s = 1, \ldots, t$, $ t=1,\ldots ,T-1)$. Let

equation*[equation* omitted — 163 chars of source]

and

equation*[equation* omitted — 164 chars of source]

Then $\boldsymbol{C}^+$ is nonsingular and $\boldsymbol{C}^+\boldsymbol{Z}_i^{+ \prime}=\boldsymbol{Z}_i^{+ \prime}\boldsymbol{U}^+$ if, and only if, every entry in $\boldsymbol{z}_{is}$ is a linear combination of entries in $\boldsymbol{z}_{it}$ $(s = 1,\ldots, t$, $ t=1,\ldots ,T-1)$. This fact and $\boldsymbol{F}^+=\boldsymbol{U}^+\boldsymbol{K}^+$ gives

equation*[equation* omitted — 560 chars of source]

and

equation*[equation* omitted — 560 chars of source]

by derivations similar to those establishing Eq.s ((ref)) and ((ref)).

thebibliography\bibitem[\citeauthoryear{Arellano}{Arellano}{2003}]{Arellano2003} \newblock Arellano, M. (2003). \newblock {\em Panel Data Econometrics}. \newblock Oxford University Press, Oxford. \bibitem[\citeauthoryear{Arellano and Bond}{Arellano and Bond}{1991}]{Arellano1991} \newblock Arellano, M., & Bond, S. (1991). \newblock Some tests of specification for panel data: Monte Carlo evidence and an application to employment equations. \newblock {\em The Review of Economic Studies} {\em 58}, 277--297. \bibitem[\citeauthoryear{Arellano and Bover}{Arellano and Bover}{1995}]{Bover1995} \newblock Arellano, M., & Bover, O. (1995). \newblock Another look at the instrumental variable estimation of error-components models. \newblock {\em Journal of Econometrics} {\em 68}, 29--51. \bibitem[\citeauthoryear{Blundell and Bond}{Blundell and Bond}{1998}]{Blundell1998} \newblock Blundell, R. & Bond, S. (1998). \newblock Initial conditions and moment restrictions in dynamic panel data models. \newblock {\em Journal of Econometrics} {\em 87}, 115--143. \bibitem[\citeauthoryear{Hayakawa}{Hayakawa}{2009}]{Hay2009} \newblock Hayakawa, K. (2009). \newblock First difference or forward orthogonal deviation- which transformation should be used in dynamic panel data models?: A simulation study. \newblock {\em Economics Bulletin} {\em 29}, 2008--2017. \bibitem[\citeauthoryear{Hsiao and Zhuo}{Hsiao and Zhou}{2017}]{Hsiao2017} \newblock Hsiao, C. & Zhou, Q. (2017). \newblock First difference or forward demeaning: Implications for the method of moments estimators. \newblock {\em Econometric Reviews} {\em 36}, 883--897. \bibitem[\citeauthoryear{Phillips}{Phillips}{2019a}]{Phillips2019a} \newblock Phillips, R. F. (2019a). \newblock A numerical equivalence result for generalized method of moments. \newblock {\em Economics Letters} {\em 179}, 13--15. \bibitem[\citeauthoryear{Phillips}{Phillips}{2019b}]{Phillips2019b} \newblock Phillips, R. F. (2019b). \newblock Quantifying the advantages of forward orthogonal deviations for long time series. \newblock {\em Computational Economics\/}. https://doi.org/10.1007/s10614-019-09907-w. \bibitem[\citeauthoryear{Schmidt, Ahn, and Wyhowski}{Schmidt et al.}{1992}]{Schmidt1992} \newblock Schmidt, P., Ahn, S. C., & Wyhowski, D. (1992). \newblock Comment. \newblock {\em Journal of Business & Economic Statistics} {\em 10}, 10--14.