EconBase
← Back to paper

A Doubly Corrected Robust Variance Estimator for Linear GMM

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

96,665 characters · 14 sections · 0 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Doubly Corrected Robust Variance Estimator for Linear GMM

array[array omitted — 677 chars of source]

$}

abstractWe propose a new finite sample corrected variance estimator for the linear generalized method of moments (GMM) including the one-step, two-step, and iterated estimators. Our formula additionally corrects for the over-identification bias in variance estimation on top of the commonly used finite sample correction of Windmeijer (2005) which corrects for the bias from estimating the efficient weight matrix, so is doubly corrected. An important feature of the proposed double correction is that it automatically provides robustness to misspecification of the moment condition. In contrast, the conventional variance estimator and the Windmeijer correction are inconsistent under misspecification. That is, the proposed double correction formula provides a convenient way to obtain improved inference under correct specification and robustness against misspecification at the same time.

Introduction

The generalized method of moments (GMM) estimators (Hansen, 1982) are widely used in economics. Among the class of GMM estimators, the efficient GMM has the smallest asymptotic variance which can be obtained via a two-step procedure. However, researchers have found that the standard error of the two-step efficient GMM is often severely downward biased. To solve this problem, Windmeijer (2005) proposed a finite sample bias-corrected standard error formula for the two-step linear GMM. Specifically, his formula corrects for the bias arising from using the efficient weight matrix being evaluated at an estimate, rather than the true value. The correction formula (the Windmeijer correction, hereinafter) has been routinely used in practice.\footnote{More than 5,200 citations according to Google Scholar on May 26, 2020.}

However, the Windmeijer correction does not take into account the over-identification bias, which is another important source of bias in the GMM standard error. The over-identification bias arises from the fact that the over-identified sample moment condition is nonzero in general while it converges in probability to zero under correct specification.

We propose a new finite sample correction which takes into account the over-identification bias for the variance of the linear one-step, two-step, and iterated GMM estimators. For one-step GMM such as the two-stage least squares (2SLS) estimators, the proposed finite sample correction is new as the Windmeijer correction does not cover the one-step GMM. For two-step and iterated GMM, the proposed correction improves upon the Windmeijer correction by additionally correcting for the over-identification bias. Thus, we doubly correct the finite sample bias of the linear GMM variance estimator.

The order of our double correction terms equals the order of the sample moment condition. Under correct specification or local misspecification (where the moment condition is modeled as a drifting sequence within a $n^{-1/2}$-neighborhood), these terms are $O_{p}(n^{-1/2})$ so that the double correction is a finite-sample correction for the variance. We provide a stochastic expansion of the GMM estimators under local misspecification in the appendix which shows that the double correction estimates the (co)variance of higher-order terms which increase with the over-identification bias.

Under (global) misspecification, however, the double correction terms no longer degenerate to zero asymptotically because the stochastic order of the sample moment condition becomes $O_{p}(1)$. The conventional variance estimator and the Windmeijer correction omit these $O_{p}(1)$ terms under misspecification. This implies that the conventional variance estimator and the Windmeijer correction are inconsistent, while our doubly corrected variance estimator is consistent regardless of whether the moment condition model is (locally or globally) misspecified or not.

Since the doubly corrected variance estimators are robust to misspecification, it is not surprising that the formulas coincide with the misspecification-robust variance estimator in Lee (2014) for the one-step and two-step GMM and Hansen and Lee (2019) for the iterated GMM. Indeed the simulation results reported in those papers show that the misspecification-robust variance estimator often performs better than the conventional sandwich variance estimator under correct specification. This paper provides an answer to this seemingly puzzling result by taking an alternative path to obtain the misspecification-robust variance estimator formula. Our approach provides new insight into the misspecification-robust formula as a finite-sample correction closely related to the well-known Windmeijer (2005) correction. We show that the misspecification-robust variance estimators of Lee (2014) and Hansen and Lee (2019) provide the same order of finite-sample correction with the Windmeijer (2005) correction under correct specification and local misspecification. To the best of our knowledge, this paper is the first to show the equivalence between the finite-sample corrected variance formula and the robust variance formula in misspecified GMM.

From a practical point of view, this implies that accurate inference under correct specification and robust inference under misspecification can be achieved simultaneously, without knowing whether the moment condition is correctly specified or not. Moreover, it can be easily implemented to obtain more accurate $t$ tests and confidence intervals (smaller errors in the size and the coverage) by bootstrapping the $t$ statistic studentized by the doubly corrected variance estimator. Lee (2014) shows that this bootstrap procedure is robust to misspecification and does not require an ad hoc correction in the bootstrap sample called recentering.

The finite sample correction of the proposed formula and the Windmeijer formula work for linear models. For nonlinear models, the order of the remainder term is the same as the correction terms, so that the corrections do not necessarily provide improvements under correct specification.

Robust inference with possibly misspecified moment condition models has gained considerable attention in the literature. For linear instrumental variable (IV) models, Maasoumi and Phillips (1982) investigate the limiting distribution of inconsistent IV estimators. Guggenberger (2012) studies the behavior of the weak instrument robust tests under local misspecification. Kang (2018) derives higher-order expansions of IV estimators allowing for local violation of the instrument validity condition. Lee (2018) shows that the moment condition is misspecified under treatment effect heterogeneity and proposes a robust variance estimator for 2SLS.

For general moment condition models, Hall and Inoue (2003) derive the asymptotic distribution of GMM under misspecification. Schennach (2007) proposes an alternative GEL-type estimator robust to global misspecification. Ai and Chen (2007) investigate the asymptotic properties of the sieve minimum distance estimator under misspecified conditional moment restrictions model. Otsu (2011) analyses moderate deviation behaviors of GMM. Kitamura, Otsu, and Evdokimov (2013) propose an estimator that achieves optimal minimax robust properties under local misspecification. Lee (2014, 2016) propose a robust nonparametric bootstrap procedure for GMM and GEL estimators. Hansen and Lee (2019) provide robust inference theory for the iterated GMM. Rotemberg (1983) and Andrews (2019) characterize the estimands of the linear GMM under misspecification. Andrews, Gentzkow, and Shapiro (2017) propose to measure the effect of model misspecification on the sensitivity of parameter estimates for the minimum distance estimators. Bonhomme and Weidner (2018) and Armstrong and Koles\'{a}r (2019) consider minimax and GMM inference under possible misspecification, respectively.

Finite sample properties of GMM estimators, including the iterated and the continuously updating (CU) GMM are investigated by Hansen, Heaton, and Yaron (1996). Bond and Windmeijer (2005) provide simulation evidence on the finite sample performance of the asymptotic and bootstrap tests based on GMM estimators. Hwang (2020) develops fixed-cluster asymptotics and finite-sample corrected variance formula for cross-sectionally dependent data. Hwang and Sun (2018) employ fixed-smoothing asymptotics to provide a more accurate comparison between the one-step and two-step GMM procedures for time-series observations.

Our doubly corrected robust variance estimators are generally different than the many instruments and many weak instruments robust variance estimators for IV, GMM, and GEL estimators proposed by Bekker (1994), Han and Phillips (2006), Newey and Windmeijer (2009), and Evdokimov and Koles\'{a}r (2018). Since our double correction formula does not use the many (weak) instruments asymptotics, it is not robust under such sequences.

The remainder of the paper is organized as follows. Section (ref) reviews the Windmeijer correction. Section (ref) proposes the doubly corrected variance estimator. Section (ref) shows that the doubly corrected variance estimator is misspecification-robust. Section (ref) discusses the iterated GMM and the CU GMM. Section (ref) derives the double correction formula for cross-sectional IV and the difference GMM. Finally, Section (ref) provides extensive simulation results comparing the double correction and other variance estimators. All the proofs are collected in Appendix A. In Appendix B, we derive the stochastic expansion of the one-step and two-step GMM estimators under local misspecification and show that the double correction estimates the (co)variance of some higher-order terms.

Finite Sample Correction of Windmeijer (2005)

Suppose that we observe a sequence of i.i.d. random vectors $X_{i}\in \mathbb{R}^{d_{x}}$ for $i=1,...,n$. Let $g(X_{i},\theta)$ be a $q\times1$ moment function where $\theta$ is a $k\times1$ parameter vector. We assume $q>k$ so that the model is over-identified and $g(X_{i},\theta)$ is linear in parameter. When the model is just-identified ($q=k$) the correction terms are zero and the analysis becomes trivial. The moment condition model is correctly specified if

equation[equation omitted — 61 chars of source]

for a unique $\theta_{0}$. Assume $E[\|g(X_{i},\theta_{0})\|^{2}]<\infty$ so that

equation[equation omitted — 130 chars of source]

Thus, the sample moment condition converges in probability to zero at the rate of $n^{-1/2}$ under correct specification (ref). This will be used in determining the order of higher-order terms in Sections 2 and 3.

The one-step GMM estimator is defined as

equation[equation omitted — 117 chars of source]

where $W_{n}$ is a $q\times q$ positive definite weight matrix which takes the form of $n^{-1}\sum_{i=1}^{n}W(X_{i})$ and $W(X_{i})$ does not depend on any unknown parameter. Common choices of $W(X_{i})$ are the identity matrix and $Z_{i}Z_{i}^{\prime}$ where $Z_{i}$ is the instrument vector in IV regressions. Let $W=EW_{n}$, a positive definite matrix of constants.

The two-step efficient GMM estimator using $\hat{\theta}_{1}$ as a preliminary (initial) estimator is defined as

equation[equation omitted — 164 chars of source]

where

equation*[equation* omitted — 101 chars of source]

Define $\Omega =\Omega(\theta_{0})$ where $\Omega(\theta) = E\Omega_{n}(\theta)$. Since $\Omega_{n}(\hat{\theta}_{1})$ is consistent for the asymptotic variance of the moment function the two-step GMM is efficient.

We also define an infeasible two-step GMM estimator $\tilde{\theta}_{2}$ using $[\Omega_{n}(\theta_{0})]^{-1}$ as the weight matrix:

equation[equation omitted — 128 chars of source]

Investigating the limiting behavior of $\sqrt{n}(\tilde{\theta}_{2}-\theta_{0})$ will help us understand the higher-order behavior of the feasible two-step estimator $\sqrt{n}(\hat{\theta}_{2}-\theta_{0})$.

Let $G(X_{i})=\partial g(X_{i},\theta)/\partial\theta^{\prime}$. Note that it does not depend on $\theta$ due to linearity. Define $G_{n}=n^{-1}\sum_{i=1}^{n}G(X_{i})$ and $G=EG_{n}$, which is assumed full column-rank. By the first-order Taylor expansion, the first-order condition (FOC) of the (feasible) two-step GMM can be written as

align[align omitted — 241 chars of source]

and so we have

equation[equation omitted — 250 chars of source]

Using a similar expansion$,$ we can get

equation[equation omitted — 242 chars of source]

for the infeasible two-step GMM and

equation[equation omitted — 174 chars of source]

for the one-step GMM.

Asymptotically (ref) and (ref) have the same limiting distribution so that using $\Omega_{n}(\hat{\theta}_{1})$ instead of $\Omega_{n}(\theta_{0})$ does not affect the first-order asymptotic analysis. However, by expanding $\Omega_{n}(\hat{\theta}_{1})$ around $\theta_{0}$ and using (ref), Windmeijer (2005) shows that the extra finite sample variations caused by higher-order terms can be estimated and the accuracy of the variance estimator can be improved for linear moment condition models.

To see this, we use the first-order Taylor expansion of $\Omega_{n}(\hat{\theta}_{1})$ in the right-hand side (RHS) of (ref) around $\theta_{0}$:

align[align omitted — 426 chars of source]

where

align*[align* omitted — 500 chars of source]

are $k\times k$ matrices and $R_{n}$ is the remainder term. Since $g_{n}(\theta_{0})=O_{p}(n^{-1/2})$ both $F_{1n}$ and $F_{2n}$ are $O_{p}(n^{-1/2})$. Thus, the second term in the RHS of (ref) is of order $O_{p}(n^{-1/2})$ assuming that $\sqrt{n} (\hat{\theta}_{1}-\theta_{0}) = O_{p}(1)$. The remainder term $R_{n}$ is of order $O_{p}(n^{-1})$ because of the linearity of the moment function provided that the higher moments of $g(X_{i},\theta_{0})$ and $G(X_{i})$ exist (the formal justification is given in the proof of Theorem 1). Thus, by taking into account for the variation caused by the $O_{p}(n^{-1/2})$ term, the finite sample variance of $\sqrt{n}(\hat{\theta}_{2}-\theta_{0})$ can be more accurately approximated. Note that the expansion (ref) only holds for linear moment condition models.

The Windmeijer correction of the variance of $\sqrt{n}(\hat{\theta}_{2}-\theta_{0})$ is obtained by

equation[equation omitted — 287 chars of source]

where $D[.,j]$ denotes the $j$th column of $D$, $\theta_{[j]}$ denotes the $j$th element of $\theta$, and

align*[align* omitted — 952 chars of source]

Since the estimate of $F_{1n}$ equals to zero because of the FOC, $0=G_{n}^{\prime}[\Omega_{n}(\hat{\theta}_{1})]^{-1}g_{n}(\hat{\theta}_{2})$, it does not appear in the variance estimator formula. The standard error is obtained by taking the diagonal elements of $\sqrt{\widehat{V}_{w}(\hat{\theta}_{2})/n}$.

Double Correction

The Windmeijer correction accounts for the extra variability due to using an estimated parameter in the weight matrix. This correction is effective because $\widehat{D}_{n}\neq0$, which is due to $g_{n}(\hat{\theta}_{2})\neq0$ in finite sample. In fact, $g_{n}(\theta)\neq0$ for all $\theta$ almost surely if at least one of the moments is continuously distributed, which (trivially) implies $g_{n}(\theta_{0})\neq0$. We call this the over-identification bias, which is non-zero for any $n$ in general for over-identified models.

We show that the over-identification bias causes additional finite sample variability in (ref). These additional terms are not considered in the Windmeijer correction (ref). We propose alternative variance estimators that fully incorporate the additional variations induced by the over-identification bias. These variance estimators will replace $\widetilde{V}(\hat{\theta}_{2})$ and $\widetilde{V}(\hat{\theta}_{1})$ in (ref) without affecting the order of finite sample corrections, leading to our doubly corrected variance estimator.

Assume that

align[align omitted — 201 chars of source]

which hold under appropriate regularity conditions. Since $G'\Omega^{-1}g=0$ by the population FOC and

equation[equation omitted — 142 chars of source]

we can write

align[align omitted — 388 chars of source]

Using (ref), the expansion of the infeasible two-step GMM (ref) can be written as

align[align omitted — 412 chars of source]

Similarly,

align[align omitted — 366 chars of source]

which simplifies to

equation[equation omitted — 191 chars of source]

when $W_{n}=I$. From the above expansions we learn the followings. First, we need to consider the extra variations from $\sqrt{n}(G_{n}-G)$ and $\sqrt{n}(\Omega_{n}(\theta)-\Omega)$ (or $\sqrt{n}(W_{n}-W)$ for the one-step GMM) to account for the over-identification bias. Second, the order of the remainder term of the original expansion (ref) is not changed.

Using the expansions (ref)-(ref) an (ref)-(ref), the expansion of the two-step GMM can be written as

align[align omitted — 755 chars of source]

Note that (ref) is the sum of (ref), (ref), and $R_{n}$ in (ref). The first two terms are $O_{p}(n^{-1})$ under (ref)-(ref).

In finite sample, $g_{n}(\theta_{0})\neq0$ because $g_{n}(\theta)\neq0$ for all $\theta$ almost surely, and this causes extra variations through the terms in (ref) and (ref). Similar to the Windmeijer correction, by taking into account for these (asymptotically negligible) terms in estimating the variance we can make more accurate inference.

Since $D_{n}=O_{p}(n^{-1/2})$, the terms in (ref) multiplied by $D_{n}$ are $O_{p}(n^{-1})$, which is the same order as the remainder term. Thus, considering those terms in (ref) does not necessarily provide finite sample corrections. However, including these terms are critical to getting robustness to misspecification, which is shown in Section (ref).

The expansion for the one-step GMM is (ref)-(ref) and those terms in (ref) are $O_{p}(n^{-1/2})$. Thus, considering the finite sample variation caused by these terms provides a more accurate variance estimator formula. This correction for the one-step GMM is not considered in Windmeijer (2005) and is new.

The doubly corrected variance estimator of $\sqrt{n}(\hat{\theta}_{2}-\theta_{0})$ is

equation[equation omitted — 313 chars of source]

where

align[align omitted — 746 chars of source]

and

align[align omitted — 460 chars of source]

When $\Xi_{n}(\phi)=\Xi(X_{i},\phi)=I$, the last term of $m(\theta,\Xi_{n}(\phi))$ drops. Note that $G(X_{i})$ and $\Xi(X_{i},\phi)$ in $m_{i}(\theta,\Xi_{n}(\phi))$ are not centered because the FOCs hold evaluated at $(\hat{\theta}_{2},\Omega_{n}(\hat{\theta}_{1}))$ and $(\hat{\theta}_{1},W_{n})$, respectively.

The doubly corrected variance estimator for the two-step GMM, $\widehat{V}_{dc}(\hat{\theta}_{2})$, provides the same order of finite sample correction as the Windmeijer correction, $\widehat{V}_{w}(\hat{\theta}_{2})$. The standard error is obtained by taking the diagonal elements of $\sqrt{\widehat{V}_{dc}(\hat{\theta}_{2})/n}$.

The doubly corrected variance estimator for the one-step GMM, $\widehat{V}_{dc}(\hat{\theta}_{1})$, accounts for the variations up to the order of $O_{p}(n^{-1/2})$ in the expansion (ref)-(ref). This correction is not considered in Windmeijer (2005). The standard error is obtained by taking the diagonal elements of $\sqrt{\widehat{V}_{dc}(\hat{\theta}_{1})/n}$.

Robustness to Misspecification

The variance estimators considered so far, the doubly corrected, the Windmeijer corrected, and the conventional, are consistent for the asymptotic variance of $\sqrt{n}(\hat{\theta}_{2}-\theta_{0})$ under correct specification, $E[g(X_{i},\theta_{0})]=0$. In words, correct specification means that an over-identified model exactly holds at a unique parameter value $\theta_{0}$, but this may be too restrictive in reality. Indeed, the sample moment condition does not hold for any finite sample size $n$ almost surely if the model is over-identified, i.e., $g_{n}(\hat{\theta})\neq0$, provided that at least one of the moments is continuously distributed. Thus, it is reasonable to view the assumed moment condition model as the best-approximating model and to allow for possible misspecification.

Under (global) misspecification, which is defined as

equation[equation omitted — 119 chars of source]

where $\delta(\theta)$ is a vector of constants and $\Theta$ is the parameter space, the GMM estimator is consistent for the pseudo-true value, which is defined as the unique minimizer of the population GMM criterion given the weight matrix (Hall and Inoue, 2003). In addition, the asymptotic variance has more terms that are assumed away under correct specification. Thus, the conventional variance estimators are no longer consistent under misspecification. Lee (2014) proposes variance estimators for the one-step and two-step GMM under misspecification. Hansen and Lee (2019) propose a similar robust variance estimator for the iterated GMM. These variance estimators are shown to be consistent regardless of misspecification and they are referred to as the misspecification-robust variance estimator, hereinafter.

Since the misspecification-robust variance estimators contain additional terms that are not present in the conventional variance estimator, it has been generally conjectured less accurate than the conventional variance estimator under correct specification. We show that this conjecture is not true by showing that the doubly corrected variance estimator $\widehat{V}_{dc}(\hat{\theta}_{2})$ is the misspecification-robust variance estimator.

The robustness of $\widehat{V}_{dc}(\hat{\theta}_{2})$ holds for the following reasons. Recall that the formulas for $\widehat{V}_{dc}(\hat{\theta}_{2})$ and $\widehat{V}_{w}(\hat{\theta}_{2})$ are given by

align*[align* omitted — 546 chars of source]

The correction term $\widehat{D}_{n}$ corrects for the bias in the variance due to using the weight matrix $[\Omega_{n}(\hat{\theta}_{1})]^{-1}$ rather than $[\Omega_{n}(\theta_{1})]^{-1}$. Since both $\widehat{V}_{dc}(\hat{\theta}_{2})$ and $\widehat{V}_{w}(\hat{\theta}_{2})$ have $\widehat{D}_{n}$, this bias is corrected in both variance estimators. What is not accounted for in the Windmeijer corrected variance estimator is the additional variations in the sample Jacobian $G_{n}$ and the sample weight matrices $W_{n}$ and $\Omega_{n}(\theta_{1})$. These variations are asymptotically negligible under correct specification but become the first-order under misspecification. Our doubly corrected variance estimator $\widehat{V}_{dc}(\hat{\theta}_{2})$ accounts for these variations.

To formally show that $\widehat{V}_{dc}(\hat{\theta}_{2})$ is consistent for the asymptotic variance under misspecification, we introduce some definitions. Define the one-step and two-step GMM (pseudo-) true values as

align[align omitted — 253 chars of source]

In general $\theta_{1}\neq \theta_{2}$ but $\theta_{1}=\theta_{2}=\theta_{0}$ under correct specification. Write $g_{j}= E[g(X_{i},\theta_{j})]$ and $\Omega_{j} = \Omega(\theta_{j})$ for $j=1,2$. (Global) misspecification implies that $g_{n}(\theta_{j})=O_{p}(1)$ for $j=1,2$.

The expansion of the GMM estimators under misspecification is quite similar to those under correct specification, except that we need to allow for different pseudo-true values for the one-step and two-step GMM and the moment condition evaluated at the pseudo-true value is not equal to zero.

Consider the FOC of the two-step GMM (ref). By expanding $g_{n}(\hat{\theta}_{2})$ around $\theta_{2}$ and $\Omega_{n}(\hat{\theta}_{1})$ around $\theta_{1}$, we can write

align[align omitted — 394 chars of source]

where $\tilde{\theta}_{2}^{*}$ is defined as

equation[equation omitted — 160 chars of source]

and

align*[align* omitted — 513 chars of source]

and $R_{n}^{*}$ is the remainder term of order $O_{p}(n^{-1/2}\|g_{n}(\theta_{2})\|)$ (this and the order of other terms are formally justified in the proof of Theorem 1). Since $D_{n}^{*} = O_{p}(\|g_{n}(\theta_{2})\|)$, the order of finite sample correction depends on the degree of misspecification, from being $O_{p}(n^{-1/2})$ under correct specification to $O_{p}(1)$ under (global) misspecification. Note that both $\sqrt{n}(\tilde{\theta}_{2}^{*}-\theta_{2})$ and $D_{n}^{*} \sqrt{n}(\hat{\theta}_{1}-\theta_{1})$ in (ref) are $O_{p}(1)$ under misspecification and this will alter the first-order asymptotic variance.

Using the population FOC $G'\Omega_{j}^{-1}g_{j}=0$ for $j=1,2$, the FOC of the infeasible two-step GMM (ref) can be expanded as

align[align omitted — 473 chars of source]

The FOC of the one-step GMM can be expanded similarly:

align[align omitted — 338 chars of source]

The expansions (ref) and (ref) are misspecification-robust versions of (ref) and (ref), allowing for different probability limits of the one-step and two-step GMM estimators and taking into account for the misspecification (over-identification) bias. Under correct specification, $g_{1}=g_{2}=0$ and (ref) and (ref) coincide with (ref) and (ref).

Now we list assumptions for the main result.

assumption\ \begin{enumerate} • $\theta_{j}$ is unique and is in the interior of the parameter space $\Theta$ for $j=1,2$$X_{1},\cdots,X_{n}$ are i.i.d. • $W$ and $\Omega(\theta_{1})$ are nonsingular • $G$ is full column rank • $E[\|g(X_{i},\theta_{1})\|^{4}]<\infty$, $E[\|g(X_{i},\theta_{2})\|^{2}]<\infty$, $E[\|G(X_{i})\|^{4}]<\infty$, $E[\|W(X_{i})\|^{2}]<\infty$ \end{enumerate}

Assumption 1 is mild regularity conditions for consistency and asymptotic normality of the one-step and two-step linear GMM estimators allowing for global misspecification. For misspecified models, Hall and Inoue (2003) provide a list of conditions for one-step and two-step GMM in the time series context. Hansen and Lee (2019) provide a list of conditions for one-step and iterated GMM under the i.n.i.d. and clustered sampling.

The following theorem shows that the doubly corrected variance estimators of the one-step and two-step linear GMM are consistent for the asymptotic variance matrices under misspecification. The proof is given in the Appendix A.

theoremSuppose that Assumption 1 holds. As $n\rightarrow\infty$, for $j=1,2$, \begin{equation*} \sqrt{n}(\hat{\theta}_{j}-\theta_{j})\xrightarrow{d}N(0,V_{j}) \end{equation*} and \begin{equation*} \widehat{V}_{dc}(\hat{\theta}_{j})\xrightarrow{p}V_{j}. \end{equation*}

Theorem (ref) holds regardless of whether the model is correctly specified or not. Thus, $\widehat{V}_{dc}(\hat{\theta}_{j})$ for $j=1,2$, provides a finite sample correction under correct specification and it remains consistent under misspecification. In contrast, the Windmeijer corrected variance estimator, $\widehat{V}_{w}(\hat{\theta}_{2})$, is not first-order consistent under misspecification. Assuming the linearity of moment condition, the proofs of Lemma 2 and Theorem 1 in Appendix A can be used to rigorously justify the stochastic orders of the higher-order terms mentioned in the previous sections.

For nonlinear models, $\widehat{V}_{dc}(\hat{\theta}_{j})$ formula can be adjusted by considering the second derivative of the moment function, which coincides with the misspecification-robust formula of Lee (2014) and Hansen and Lee (2019). The same consistency result with Theorem 1 are shown in those papers, but stronger assumptions on the moment/Jacobian processes and the compact parameter space are required to use the uniform law of large numbers.

Theorem (ref) also implies that the nonparametric i.i.d. bootstrap $t$ test and confidence intervals (CIs) based on the GMM $t$ statistic studentized with the doubly corrected standard error automatically achieve higher-order refinements over the asymptotic $t$ test and CIs regardless of misspecification (Lee, 2014). In contrast, those bootstrap $t$ test and CIs based on the GMM $t$ statistic studentized with the conventional or the Windmeijer standard error require an additional recentering procedure in resampling to correct for the over-identification bias to achieve higher-order refinements, see Hall and Horowitz (1996) and Andrews (2002). Furthermore, the conventional nonparametric i.i.d. bootstrap procedure for GMM is not valid under misspecification. Thus, our doubly corrected variance estimator formula provides a very convenient way to get more accurate but also robust bootstrap tests and CIs.

\noindentRemark 1 (Stochastic expansion under local misspecification) From the way it is constructed, we can directly see that the doubly corrected variance formulas provides a finite-sample variance correction with the same argument as Windmeijer (2005). It is expected to work the best when the sample moment condition is large. Since a nonzero sample moment condition can also be due to a locally misspecified moment condition, one can seek an additional justification of the double correction by deriving the stochastic expansions of the GMM estimators under such a sequence. Appendix B provides formal stochastic expansions of the one-step and two-step GMM estimators by allowing the population moment condition evaluated at the true value is $E[g(X_{in},\theta_{0})]=\delta/\sqrt{n}$ for some $\delta\neq0$. In Theorems (ref) and (ref) in Appendix B, we discuss how the first-order terms $D_{n}^{*}$, (ref), and (ref) in the double correction are related to the higher-order terms under local misspecification. Interestingly, our analysis reveals that the double correction effectively estimates the (co)variances of higher-order terms that depend on $\delta$ and thus is fully robust to additional variations due to local misspecification. In contrast, the Windmeijer corrected variance estimator only partially considers the higher-order terms that depend on $\delta$, making it only partly robust to local misspecification. Finally, our stochastic expansions show that both the doubly corrected and the Windmeijer corrected variance estimators are not higher-order variance estimators which would estimate additional terms up to $O(n^{-1})$. For the general treatment of the stochastic expansion, see Rothenberg (1984). Newey and Smith (2004) derive the stochastic expansion of GMM and the generalized empirical likelihood (GEL) estimators under correct specification.

\noindentRemark 2 (Weight matrix) The main results hold if we replace $\Omega_{n}(\theta)$ with the centered weight matrix

equation[equation omitted — 184 chars of source]

with some specifics need to be modified accordingly. Specifically, the two-step GMM pseduo-true value $\theta_{2}$ defined in (ref) is now defined with $\Omega^{c}_{n}(\theta_{1})$. In addition, the derivative of the centered weight matrix is

equation*[equation* omitted — 219 chars of source]

so that

equation*[equation* omitted — 134 chars of source]

The centered weight matrix is consistent for the asymptotic variance matrix of the moment equation under misspecification. Hansen (2020) recommends using the centered weight matrix for this reason. Hall (2000) shows that the GMM over-identification test statistic with a centered heteroskedasticity-and-autocorrelation-consistent (HAC) weight matrix leads to more powerful tests in the time series setting.

Iterated GMM and Continuously Updating GMM

Both the Windmeijer and our double correction correct for the extra variation due to the weight matrix being evaluated at an estimate rather than the true value. A natural question is whether similar finite sample corrections can be obtained for other GMM estimators, namely the iterated GMM of B. Hansen and Lee (2019) and the continuously-updating (CU) GMM of L. Hansen, Heaton, and Yaron (1996). We show that the answer is yes for the iterated GMM and the double correction formula is the same as the misspecification-robust formula. For the CU GMM, the answer is negative.

Assume correct specification. The iterated GMM estimator is obtained by iterating the two-step efficient GMM estimator until convergence. By iteration the dependence of the final estimator on the previous step estimators disappears. The FOC is given by

equation[equation omitted — 74 chars of source]

where $\hat{\theta}$ is the iterated GMM. Assume that $g_{n}(\theta_{0})=O_{p}(n^{-1/2})$ and $\hat{\theta} -\theta_{0} = O_{p}(n^{-1/2})$ whose sufficient conditions are provided in Hansen and Lee (2019). By applying the first-order Taylor expansion around $\theta_{0}$ to $g_{n}(\hat{\theta})$ and $\Omega_{n}(\hat{\theta})$ sequentially

align[align omitted — 241 chars of source]

and thus

align[align omitted — 232 chars of source]

Windmeijer (2000) proposes a finite sample corrected variance estimator based on the expansion (ref). We proceed one additional step. By further expanding to take into account for the over-identification bias, we have

align[align omitted — 420 chars of source]

Since the remainder term is $O_{p}(n^{-1})$, by estimating the variance of the terms in (ref)-(ref) up to $O_{p}(n^{-1/2})$ we can get the same order of finite sample correction with the doubly corrected two-step GMM variance estimator.

The doubly corrected variance estimator for the iterated GMM is

align[align omitted — 267 chars of source]

where $\Sigma_{n}(\hat{\theta},\Omega_{n}(\hat{\theta}))$ is defined in (ref) and $\widehat{D}_{n}$ is evaluated at $\hat{\theta}$. Not surprisingly, this formula is identical to the misspecification-robust variance estimator for the iterated GMM of Hansen and Lee (2019). The finite sample corrected formula suggested by Windmeijer (2000) is

equation[equation omitted — 166 chars of source]

On the other hand, a similar finite sample correction may not be obtained for the CU GMM. Windmeijer (2005) showed that if the derivative of the moment function is a function of the parameter, then the proposed formula would not necessarily give finite sample corrections. The same argument applied to CU GMM. Let $\hat{\theta}$ be the CU GMM estimator. For simplicity, let $k=1$ so that $\theta$ is scalar. The FOC is

equation[equation omitted — 253 chars of source]

This shows that even when the moment function is linear, the effective Jacobian term in the FOC still depends on the parameter. Thus, the misspecification-robust variance formula for CU GMM does not necessarily provide a finite sample correction under correct specification. Since GEL estimators have similar non-linear FOC even with linear moment functions, we expect similar conclusions.

Examples

Cross-sectional IV

Consider the linear IV model $y_{i} = X_{i}^{\prime} \theta + e_{i}$ with the moment conditions $E[Z_{i} e_{i}]=0$. The two-stage least squares (2SLS) estimator is given by

equation[equation omitted — 129 chars of source]

where $Y = [y_{1}, \cdots, y_{n}]^{\prime}, X = [X_{1}, \cdots, X_{n}]^{\prime}$, and $Z = [Z_{1}, \cdots, Z_{n}]^{\prime}$ are $n\times 1$, $n\times k$, and $n\times q$ data matrices. Using the 2SLS as the preliminary estimator, the two-step efficient GMM estimator is given by

equation[equation omitted — 143 chars of source]

where

align*[align* omitted — 159 chars of source]

Also define $\hat{e}_{2i} = y_{i} - X_{i}^{\prime} \hat{\theta}_{2}$ and the $n\times 1$ residual vector $\hat{e}_{j} = Y-X\hat{\theta}_{j}$ for $j=1,2$.

The doubly corrected variance estimators of the 2SLS and two-step GMM are

align[align omitted — 544 chars of source]

where

align*[align* omitted — 1,382 chars of source]

It is worth observing that the doubly corrected variance estimator $\widehat{V}_{dc}(\hat{\theta}_{2})$ reduces to the Windmeijer corrected one $\widehat{V}_{w}(\hat{\theta}_{2})$ if (i) the last two terms in $\hat{m}_{2i}$ and $\hat{m}_{1i}$ are ignored and (ii) $\hat{e}_{1i}$ replaces $\hat{e}_{2i}$ in $\hat{m}_{2i}$. By (i) and (ii), the variance estimators $\widehat{V}(\hat{\theta}_{2})$ and $\widehat{V}_{dc}(\hat{\theta}_{1})$ reduce to conventional ones $\widetilde{V}(\hat{\theta}_{2})$ and $\widetilde{V}(\hat{\theta}_{1})$, and $\widehat{C}(\hat{\theta}_{1},\hat{\theta}_{2})$ becomes $\widetilde{V}(\hat{\theta}_{2})$. In general, however, $ \widehat{V}_{dc}(\hat{\theta}_{2})\neq \widehat{V}_{w}(\hat{\theta}_{2})$ because $Z^{\prime}\hat{e}_{j} \neq0$ for $j=1,2$, so the last two terms of $\hat{m}_{ji}$ are non-zero. Furthermore, it is critical (and reasonable) to use $\hat{e}_{2i}$ in $\hat{m}_{2i}$ to get robustness under misspecification.

The iterated GMM estimator is obtained as follows. Let $\hat{\theta}_{0}$ be any initial value. The $s$-step GMM estimator for $s\geq1$ is given by

equation[equation omitted — 148 chars of source]

where

equation*[equation* omitted — 126 chars of source]

We iterate the $s$-step GMM estimator until convergence given a preset tolerance $\epsilon$, i.e. $\|\hat{\theta}_{s}-\hat{\theta}_{s-1}\|<\epsilon$ to obtain the iterated GMM estimator $\hat{\theta}$. The residuals are $\hat{e}_{i} = y_{i} - X_{i}^{\prime} \hat{\theta}$. Also let $\hat{e} = Y-X\hat{\theta}$ be the $n\times1$ residual vector.

The doubly corrected variance estimator is

align[align omitted — 710 chars of source]

In comparison, the Windmeijer corrected and the conventional variance estimators are

align[align omitted — 263 chars of source]

A Panel Data Model

Consider a panel data model with a scalar regressor

equation[equation omitted — 76 chars of source]

for $i = 1, ..., N$ and $t = 1, ... ,T$ where $\eta_{i}$ is the unobserved individual effects, the unknown parameter of interest is $\beta$, and the single regressor $x_{it}$ is predetermined with respect to $v_{it}$ (possibly including lags of the dependent variable), i.e., $E(x_{it}v_{is})=0$ for all $s\geq t$. After first-differencing,

equation*[equation* omitted — 89 chars of source]

the standard approach to estimate $\beta$ is the first differenced GMM (Arellano and Bond (1991) estimator) with the moment conditions $E(Z_{i}^{\prime} \Delta v_{i}) = 0$ where $Z_{i}$ is the $(T-1)$ by $T(T-1)/2$ instrument matrix

equation*[equation* omitted — 67 chars of source]

with all possible lagged instruments $z_{it}=(x_{i1},\cdots,x_{it-1})^{\prime}$ for $2\leq t\leq T$ and $\Delta v_{i} = (\Delta v_{i2},\cdots, \Delta v_{iT})^{\prime}$. The total number of observations is $n=N(T-1)$.

Our doubly corrected variance estimator can be used for the model (ref) with additional strictly exogenous, predetermined, or endogenous variables as well as the system GMM estimator (Arellano and Bover (1995) and Blundell and Bond (1998)) by stacking and modifying additional moment conditions into the instrument sets $Z_{i}$. If the panel is unbalanced the instrument matrix can be constructed as described in Arellano and Bond (1991).

Using the initial weight matrix $\widehat{W} = n^{-1}\sum_{i=1}^{N}Z_{i}^{\prime} H Z_{i}$, where $H$ is a matrix with $2$'s on the main diagonal, $-1$'s on the first off-diagonals and zero elsewhere, the one-step GMM estimator is given by

equation*[equation* omitted — 154 chars of source]

where $Z = (Z_{1}^{\prime}, ..., Z_{N}^{\prime})^{\prime}$ is the instrument matrix, $\Delta Y = (\Delta y_{1}^{\prime}, ..., \Delta y_{N}^{\prime})^{\prime}$, $\Delta X = (\Delta x_{1}^{\prime}, ..., \Delta x_{N}^{\prime})^{\prime}$, $\Delta y_{i} = (\Delta y_{i2}, ..., \Delta y_{iT})'$, and $\Delta x_{i} = (\Delta x_{i2}, ..., \Delta x_{iT})'$. Note that scaling the weight matrix does not affect the estimator. The doubly corrected variance estimator of $\hat{\beta}_{1}$ is given by

align*[align* omitted — 563 chars of source]

where $\Delta \hat{v}_{1i} = \Delta y_{i} - \Delta x_{i} \hat{\beta}_{1}$ and $\Delta \hat{v}_{1} = (\Delta \hat{v}_{11}^{\prime}, ..., \Delta \hat{v}_{1N}^{\prime})^{\prime}$. The doubly corrected standard error is obtained by taking the diagonal elements of $\sqrt{\widehat{V}_{dc}(\hat{\beta}_{1})/n}$. In comparison, the conventional variance estimator is given by

equation*[equation* omitted — 285 chars of source]

where

equation[equation omitted — 135 chars of source]

Next, consider the two-step efficient GMM estimator

equation*[equation* omitted — 174 chars of source]

Let $\Delta \hat{v}_{2i} = \Delta y_{i} - \Delta x_{i} \hat{\beta}_{2}$ and $\Delta \hat{v}_{2} = (\Delta \hat{v}_{21}^{\prime}, ..., \Delta \hat{v}_{2N}^{\prime})^{\prime}$. The doubly corrected variance estimator of $\hat{\beta}_{2}$ is given by

equation*[equation* omitted — 294 chars of source]

where

align*[align* omitted — 1,378 chars of source]

The doubly corrected standard error is obtained by taking the diagonal elements of $\sqrt{\widehat{V}_{dc}(\hat{\beta}_{2})/n}$. Note that the Windmeijer corrected variance estimator is

equation*[equation* omitted — 255 chars of source]

where

equation[equation omitted — 136 chars of source]

Finally, the iterated GMM estimator is given as follows. Let $\hat{\beta}_{0}$ be any initial value. The $s$-step GMM estimator for $s\geq1$ is given by

equation[equation omitted — 175 chars of source]

where

equation*[equation* omitted — 176 chars of source]

We iterate the $s$-step GMM estimator until convergence given a preset tolerance $\epsilon$, i.e. $\|\hat{\beta}_{s}-\hat{\beta}_{s-1}\|<\epsilon$ to obtain the iterated GMM estimator $\hat{\beta}$. The residuals are $\Delta \hat{v}_{i} = \Delta y_{i} - \Delta x_{i}\hat{\beta}$. Also let $\Delta \hat{v} = (\Delta \hat{v}_{1}^{\prime}, ..., \Delta \hat{v}_{N}^{\prime})^{\prime}$ be the $n\times1$ residual vector.

The doubly corrected variance estimator for the iterated GMM is given by

align*[align* omitted — 901 chars of source]

and the doubly corrected standard error is obtained by taking the diagonal elements of $\sqrt{\widehat{V}_{dc}(\hat{\beta})/n}$.

In comparison, the Windmeijer corrected and the conventional variance estimators are

align[align omitted — 289 chars of source]

Simulation

We investigate the finite sample performance of the doubly corrected standard errors proposed in this paper and provide a thorough comparison with the conventional and the Windmeijer corrected ones under correct specification and misspecification. We consider three different setups: (i) a cross-sectional linear IV model with potentially invalid instruments; (ii) a linear dynamic panel model with a random coefficient; (iii) a linear dynamic panel model with possibly misspecified lag specifications. The number of Monte Carlo simulation is 100,000.

In an unreported simulation, we also investigate the performance of the estimators with the centered weight matrix (ref). Since the results are similar and there is no obvious pattern of better performance of the point and variance estimators based on the centered weight matrix compared with those based on the uncentered one (reported) they are not reported.

Cross-sectional IV

We use the following simulation design which is a simple linear instrumental variable regression with a single endogenous regressor. The model to be estimated is

align[align omitted — 116 chars of source]

where $x_{i}$ and $\beta_{0}$ are scalar and $z_{i}=(z_{1i},z_{2i} ,z_{3i},z_{4i})^{\prime}$ is a vector of instrumental variables. We estimate $\beta_{0}$ by 2SLS (one-step), two-step, and iterated GMM, and calculate the conventional, the Windmeijer corrected, and the doubly corrected standard errors. Our data-generating process (DGP) is

align[align omitted — 561 chars of source]

We set $\beta_{0}=1$, vary $\alpha_{0}$ from 0 to 1 in steps of 0.2, and set the first-stage coefficient $\pi_{0}$ so that the first-stage $R^{2}=0.2$. We set the number of observations as $n=50, 100, 500$.

The parameter $\alpha_{0}$ is the extent that the exclusion condition is locally violated. At $\alpha_{0}=0$, the model is correctly specified. For $\alpha_{0}\neq0$, we find $E (z_{i}e_{i}) = (\alpha_{0}, -\alpha_{0}, \alpha_{0}, -\alpha_{0})^{\prime}/\sqrt{n} \neq 0$, so the moment condition ((ref)) fails to hold in finite samples, but it holds asymptotically.

Means and standard deviations of one-step (2SLS), two-step, and iterated GMM estimators are computed in Table (ref). For all GMM estimators, we report means of the conventional standard errors (se $\hat{\beta}$), the Windmeijer corrected standard errors (se$_{w}$ $\hat{\beta}$), and the doubly corrected standard errors (se$_{dc}$ $\hat{\beta}$).

Table (ref) shows that our doubly corrected standard errors remain accurate regardless of misspecification, including the correct specification case ($\alpha_{0} = 0$); the means of corrected standard errors are very close to the standard deviations for all values of $\alpha_{0}$, especially for the two-step and the iterated GMM. Simulation evidence reassures our theory that the doubly corrected standard errors not only take into account variation in the estimation of the weight matrix but also extra variation due to the non-zero sample moments in the over-identified model even under correct specification. Furthermore, our doubly corrected standard errors are the only valid one under misspecification.

The conventional standard error for the one-step GMM (2SLS) estimator is downward biased under correct specification ($\alpha_{0}=0$) and this bias increases with $\alpha_{0}$. As is well known, the conventional standard error for the two-step is severely downward biased when $\alpha_{0}=0$, and this bias also increases with $\alpha_{0}$. The Windmeijer corrected standard error works well under correct specification, but does not fully account for additional variations when $\alpha_{0}$ is non-zero. The result is similar for the iterated GMM.

It is worth noting that the two-step and iterated GMM point estimates are sensitive to the local violation parameter $\alpha_0$. Interestingly, the point estimate becomes similar to the true value as the degree of misspecification $\alpha_{0}$ increases. This is because, in our DGP, the local misspecification bias (which depends on $\alpha_{0}$) and the higher-order asymptotic bias (which does not depend on $\alpha_{0}$) have opposite signs so that it happens to cancel out each other as $\alpha_{0}$ increases. In contrast, the one-step GMM point estimate varies little across $\alpha_{0}$ because the local misspecification bias is zero.\footnote{The local misspecification bias and the higher-order asymptotic bias can be calculated using the formula in Theorems 2 and 3 in Appendix B.} This is specific to this DGP and cannot be generalized. The size of bias in the point estimate decreases as the sample size gets larger.

table[table omitted — 3,254 chars of source]

Linear Dynamic Panel Model

Random Coefficient

We next explore the finite sample performance of the doubly corrected standard error in the presence of heterogeneous effects (random coefficient) in dynamic panel model. We consider the AR(1) dynamic panel model of Blundell and Bond (1998). For $i=1,...,N$ and $t=1,...,T$,

eqnarray[eqnarray omitted — 94 chars of source]

where $\eta_{i}$ is an unobserved individual-specific effect and $\nu_{it}$ is an error term. The parameter of interest $\rho_{0}$ is estimated by the difference GMM based on a set of moment conditions:

equation[equation omitted — 144 chars of source]

The moment conditions are derived from taking differences of (ref), and uses the lagged values of $y_{it}$ as instruments. The number of moment conditions is $(T-1)(T-2)/2$.

The moment conditions are correctly specified if there is a unique parameter that satisfies (ref). A sufficient condition for this to hold is that the model (ref) coincides with the true DGP, but this is unlikely to be true. A reasonable deviation from the assumed model (ref) is heterogeneity in $\rho_{0}$ across $i$. We assume the following DGP. For $i=1,...,N$ and $t=1,...T$,

eqnarray[eqnarray omitted — 292 chars of source]

where $\Phi(z)$ is the standard normal cdf. At $\alpha_{0}=0$, the model is correctly specified and $\rho_{i}=\rho_{0}=0.5$. For $\alpha_{0}\neq0$, the effective moment condition model can be written as

align*[align* omitted — 204 chars of source]

where $\gamma_{j}$ is the $j$th autocovariance. The last equation becomes zero at $\rho=E[\rho_{i}]$ if $\rho_{i}$ is independent of the $\{y_{it}\}$ process. If this is the case, then the moment condition model is correctly specified and the estimand is $E[\rho_{i}]$. Otherwise in general, the moment condition model fails to hold at a single unique parameter value because each of the moment condition imposes a restriction \[\rho = \frac{E[\rho_{i}y_{i,t-s}\Delta y_{i,t-1}]}{\gamma_{s-1}-\gamma_{s-2}}\] but there is no reason that this should hold at a unique $\rho$ for $s=2,3,...,t-1$. In the DGP, $\eta_{i}$ and $\rho_{i}$ are dependent through $\alpha_{0}$ and a larger $\alpha_{0}$ leads to larger heterogeneity. We vary $\alpha_{0}$ from 0 to 0.3 in steps of 0.05. The pseudo-true value would depend on the instrument set and the value of $\alpha_{0}$ under global misspecification. However, by varying $\alpha_{0}$ by a small amount we try to capture local behavior of the standard errors when the pseudo-true value is close to the true value. The sample sizes are $N=100, 500$ and $T=4, 6$.

We report the simulation results in Tables (ref) and (ref), which are qualitatively similar to the IV setup. Tables (ref) and (ref) show that the doubly corrected standard errors approximate the standard deviation of the GMM estimators well regardless of misspecification. For the two-step and iterated GMM estimators, the doubly corrected standard errors are as accurate as the Windmeijer correction for small values of $\alpha_{0}$ (including correct specification $\alpha_{0} = 0$) but dominate the other in terms of accuracy for larger values of $\alpha_{0}$. The doubly corrected standard error for the one-step GMM is slightly upward biased for small values of $\alpha_{0}$, but this bias decreases with a larger sample size $N =500$.

table[table omitted — 2,447 chars of source]
table[table omitted — 2,447 chars of source]

Misspecified Lag Length

We use the baseline linear panel model of Windmeijer (2005) allowing for possible lag length misspecification. The model is

equation[equation omitted — 87 chars of source]

for $i = 1, ..., N$ and $t = 1, ..., T$. The unknown parameter of interest is $\beta_{0}$, and the regressor $x_{it}$ is predetermined with respect to $v_{it}$, i.e., $E(x_{it}v_{it+s})=0$ for $s=0,...,T-t.$ We use the first differenced GMM estimator and the number of moment conditions is $T(T-1)/2$ as in Section (ref).

The DGP is

align[align omitted — 406 chars of source]

We generate initial $50$ time periods with $\tau_{t}=0.5$ for $t=-49,\ldots,0$ and $x_{i, -49}\sim N(\eta_{i}/0.5,1/0.75)$ same as Windmeijer (2005). The parameter $\alpha_{0}$ in (ref) governs the degree of misspecification. When $\alpha_{0}=0$, the model (ref) is correctly specified which reduces to that of Windmeijer (2005).

The model (ref) is misspecified for $\alpha_{0}\neq0$. We discuss the pseudo-true value and its interpretation. Since $0=E[x_{is}\Delta v_{it}]$, the effective moment condition can be written as

equation[equation omitted — 163 chars of source]

for all $1\leq s<t$ and $2\leq t\leq T$.

First, consider $T=2$ where there is only one moment condition so that the model is just-identified. By setting $s=1$ and $t=2$, the unique solution to (ref) is $\beta^{*}=\beta_{0}-\alpha_{0}$. Since (ref) equals to zero at $\beta^{*}$, the moment condition is not misspecified although the lag is misspecified. Since the model is just-identified $\beta^{*}=\beta_{0}-\alpha_{0}$ is considered as the true value but this is not equal to $\beta_{0}$ unless $\alpha_{0}=0$. An implication is that model misspecification and pseudo-true values are not only pertinent to over-identified models.

For $T\geq3$, the model is overidentified and each of the moment condition holds at

equation[equation omitted — 136 chars of source]

Note that $\beta^{*}_{t,s}$ can vary across $t$ and $s$ for a nonzero $\alpha_{0}$. For example, $\beta^{*}_{3,1} = \beta_{0} + 2\alpha_{0}$ and $\beta^{*}_{3,2} = \beta_{0}-\alpha_{0}$. The pseudo-true value $\beta^{*}$ is a weighted average of $\beta_{t,s}^{*}$'s given $T$ and the weight matrix. One can interpret that $\beta^{*}$ represents the causal effects of past and present values of $x_{it}$'s on $y_{it}$ in the true DGP.

Tables (ref) and (ref) report estimation results for $\beta_{0}=1,$ $N=100, 500$ and $T=4,6$. The degree of misspecification $\alpha_{0}$ is varied across $\{ 0, 0.05, 0.1, 0.2, 0.3 \}$. The first column $(\alpha_{0} =0)$ in Table (ref) replicates Monte Carlo studies in Windmeijer (2005, Table 1). Tables (ref) and (ref) also report the rejection rates of the $J$ test with the nominal size of 5%. Although the rejection rate increases as the degree of misspecification increases and the sample size increase, the $J$ test tends to not correctly detect the violation of the overidentifying moment restrictions in small samples.

The implication of the results in Tables (ref) and (ref) are largely unchanged as in two previous simulation experiments; doubly corrected standard errors approximate the standard deviations well regardless of model misspecification. In this simulation experiment, the Windmeijer correction works best under correct specification but becomes downward biased as $\alpha_{0}$ increases. Note that deviation from the correct specification makes the bias of the conventional standard error and the Windmeijer corrected standard error larger, and this bias does not disappear with a larger sample size of $N=500$.

table[table omitted — 2,785 chars of source]
table[table omitted — 2,780 chars of source]

Size

This section investigates the small sample performance of the $t$ tests with the proposed standard errors. We report the size of the $t$ tests under the correct specification for the simulation setups considered in previous sections.

Based on the one-step, two-step, and iterated GMM estimators, Table (ref) evaluates the size of the $t$ tests for nominal size 5% using various standard errors. In the column labeled $t$, we report the size of the test based on conventional heteroskedasticity-robust standard errors. $t_w$ is based on the Windmeijer corrected standard errors while $t_{dc}$ is based on the doubly corrected standard errors in this paper. Further, we report the size of the bootstrap $t$ test using the doubly corrected standard errors ($t_{dc}$-$bs$) which is equivalent to the misspecification-robust (MR) bootstrap of Lee (2014). To calculate bootstrap critical values, 1000 additional bootstrap replications are performed per Monte Carlo replication. Results for larger sample sizes ($n=500$ in Table (ref), Table (ref) and (ref)) are not reported here for brevity as the results are similar across different methods.

For Setups 1 and 2, tests using the conventional standard errors ($t$) are severely oversized. Using the Windmeijer corrected standard errors ($t_w$), and the doubly corrected standard errors ($t_{dc}$) for the two-step and iterated estimator improve the size and perform similarly, although both are moderately oversized. The size distortion decreases when we increase the sample size. Using the MR bootstrap with the doubly corrected standard errors ($t_{dc}$-$bs$) improves the size of the test dramatically. The excellent size property of the MR bootstrap $t$ test is theoretically justified by its asymptotic refinements, which is formally shown by Lee (2014). Perhaps surprisingly, using the doubly corrected standard errors for the $t$ test and bootstrap $t$ test improves size properties considerably for the one-step estimator where the Windmeijer correction is not available.

For Setup 3, $t$ based on the one-step estimator and $t_{w}$ with the two-step estimator have good size properties, and the same results can be found in Windmeijer (2005, Fig 1.). $t_{dc}$ has similar size properties to $t_{w}$, but is slightly undersized for the one-step estimator and slightly oversized for the two-step and iterated estimators. The MR bootstrap test is slightly undersized. Windmeijer (2005) and Bond and Windmeijer (2005) report similar results for the nonparametric bootstrap test of Hall and Horowitz (1996) based on the two-step estimator and explain that the performance of the bootstrap deteriorates with an increasing number of moment conditions. Since $t_{dc}$ becomes oversized with the number of moment conditions, this suggests that some components of the doubly corrected variance estimator may be sensitive to the number of moment conditions. This issue deserves further future investigation.

table[table omitted — 1,670 chars of source]