EconBase
← Back to paper

Asymptotic Refinements of a Misspecification-Robust Bootstrap for Generalized Method of Moments Estimators

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

80,933 characters · 14 sections · 0 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Asymptotic Refinements of a Misspecification-Robust Bootstrap for Generalized Method of Moments Estimators

abstractI propose a nonparametric iid bootstrap that achieves asymptotic refinements for $t$ tests and confidence intervals based on GMM estimators even when the model is misspecified. In addition, my bootstrap does not require recentering the moment function, which has been considered as critical for GMM. Regardless of model misspecification, the proposed bootstrap achieves the same sharp magnitude of refinements as the conventional bootstrap methods which establish asymptotic refinements by recentering in the absence of misspecification. The key idea is to link the misspecified bootstrap moment condition to the large sample theory of GMM under misspecification of Hall and Inoue (2003). Two examples are provided: Combining data sets and invalid instrumental variables.\\ Keywords: nonparametric iid bootstrap, asymptotic refinement, Edgeworth expansion, generalized method of moments, model misspecification. \\ JEL Classification: C14, C15, C31, C33

Introduction

This paper proposes a novel bootstrap procedure for the generalized method of moments (GMM) estimators of Hansen (1982). It extends the existing literature by establishing the same asymptotic refinements for $t$ tests and confidence intervals (CI's) (i) without recentering the bootstrap moment function, and (ii) without assuming correct model specification. In contrast, the conventional bootstrap achieves the refinements only if recentering is done and the assumed moment condition is correctly specified. Thus, the contribution of this paper may look too good to be true at first glance, but it becomes apparent once we realize that those two eliminations are in fact closely related, because recentering makes the bootstrap non-robust to misspecification.

Bootstrapping has been considered as an alternative to the first-order GMM asymptotic theory, which has been known to provide poor approximations of finite sample distributions of test statistics especially when the model is highly non-linear or the number of moments is large, e.g., Blundell and Bond (1998), Bond and Windmeijer (2005), Hansen, Heaton, and Yaron (1996), Kocherlakota (1990), and Tauchen (1986).\footnote{The 1996 special issue of the Journal of Business & Economic Statistics deals with this problem in various contexts.} Hahn (1996) proves the first-order validity of the bootstrap distribution of GMM estimators. Hall and Horowitz (1996) show asymptotic refinements of the bootstrap for $t$ tests and the $J$ test (henceforth the Hall-Horowitz bootstrap). Andrews (2002) proposes a computationally attractive $k$-step bootstrap procedure based on the Hall-Horowitz bootstrap. Inoue and Shintani (2006) extend the Hall-Horowitz bootstrap by allowing correlation of moment functions beyond finitely many lags. Brown and Newey (2002) suggest an alternative bootstrap procedure using the empirical likelihood (EL) probability (henceforth the Brown-Newey bootstrap).

In the existing bootstrap methods for GMM estimators, recentering is critical. Horowitz (2001) explains why recentering is important when applying the bootstrap to overidentified moment condition models, where the dimension of a moment function is greater than that of a parameter. In such models, the sample mean of the moment function evaluated at the estimator is not necessarily equal to zero, though it converges almost surely to zero if the model is correctly specified. In principle, the bootstrap considers the sample and the estimator as if they were the population and the true parameter, respectively. This implies that the bootstrap version of the moment condition, that the sample mean of the moment function evaluated at the estimator should equal to zero, does not hold when the model is overidentified. Recentering makes the bootstrap version of the moment condition hold. The Hall-Horowitz bootstrap analytically recenters the bootstrap moment function with respect to the sample moment condition. The Brown-Newey bootstrap recenters the bootstrap moment condition by employing the EL probability in resampling the bootstrap sample. Thus, both the Hall-Horowitz bootstrap and the Brown-Newey bootstrap can be referred as the recentered bootstrap.

A naive bootstrap is to apply the standard bootstrap procedure as is done for just-identified models, without any additional correction, such as recentering. However, it turns out that this naive bootstrap fails to achieve asymptotic refinements for $t$ tests and CI's, and jeopardizes first-order validity of the $J$ test. Hall and Horowitz (1996) and Brown and Newey (2002) explain that the bootstrap and sample versions of test statistics would have different asymptotic distributions without recentering, because of the violation of the moment condition in the sample.

Although they address that the failure of the naive bootstrap is due to the misspecification in the sample, they do not further investigate the conditional asymptotic distribution of the bootstrap GMM estimator under misspecification. Instead, they eliminate the misspecification problem by recentering. In contrast, I observe that the conditional asymptotic covariance matrix of the bootstrap GMM estimator under misspecification is different from the standard one. The conditional asymptotic covariance matrix is consistently estimable by using the result of Hall and Inoue (2003), and I construct the $t$ statistic of which distribution is asymptotically standard normal even under misspecification.

Hall and Inoue (2003) show that the asymptotic distributions of GMM estimators under misspecification are different from those of the standard GMM theory.\footnote{Hall and Inoue (2003) does not deal with bootstrapping, however.} In particular, the asymptotic covariance matrix has additional non-zero terms in the presence of misspecification. Hall and Inoue's formulas for the asymptotic covariance matrix encompass the case of correct specification as a special case. The variance estimator using their formula is denoted by the Hall-Inoue variance estimator, hereinafter. Imbens (1997) also describes the asymptotic covariance matrices of GMM estimators robust to misspecification by using a just-identified formulation of overidentified GMM. However, his description is general, rather than being specific to the misspecification problem defined in this paper.

I propose a bootstrap procedure that uses the Hall-Inoue variance estimators in constructing the sample and the bootstrap $t$ statistics. It ensures that the bootstrap $t$ statistic satisfies the asymptotic pivotal condition without recentering. Moreover, the sample $t$ statistic is also asymptotically pivotal regardless of misspecification in the population. In other words, my bootstrap applies to the robust $t$ statistic which is studentized with the Hall-Inoue variance estimator. Therefore, it works without assuming correct model specification in the population, and is referred to as the misspecification-robust (MR) bootstrap. In contrast, the conventional first-order asymptotics as well as the recentered bootstrap would not work under misspecification, because the conventional $t$ statistic is not asymptotically pivotal anymore.

The MR bootstrap achieves asymptotic refinements, a reduction in the error of test rejection probability and CI coverage probability by a factor of $n^{-1}$ for symmetric two-sided $t$ tests and symmetric percentile-$t$ CI's, over the asymptotic counterparts. The magnitude of the error is $O(n^{-2})$, which is sharp. This is the same magnitude of error shown in Andrews (2002), that uses the Hall-Horowitz bootstrap for independent and identically distributed (iid) data with slightly stronger assumptions than those of Hall and Horowitz (1996).

I note that the MR bootstrap is not for the $J$ test. To get the bootstrap distribution of the $J$ statistic, the bootstrap should be implemented under the null hypothesis that the model is correctly specified. The recentered bootstrap imposes the null hypothesis of the $J$ test because it eliminates the misspecification in the bootstrap world by recentering. In contrast, the MR bootstrap does not eliminate the misspecification and thus, it does not mimic the distribution of the $J$ statistic under the null. Since the conventional asymptotic and bootstrap $t$ tests and CI's are valid only in the absence of misspecification, it is important to conduct the $J$ test and report the result that the model is not rejected. However, even a significant $J$ statistic would not invalidate the estimation results if possible misspecification of the model is assumed and the validity of $t$ tests and CI's is established under such assumption, as is done in this paper.

Three papers in the literature are in a similar vein in terms of bootstrap methods under misspecification. Corradi and Swanson (2006) show the first-order validity of the block bootstrap for conditional distribution tests under dynamic misspecification. Kline and Santos (2012) examine the higher-order properties of the wild bootstrap in a linear regression model when the mean independent assumption of the error term is misspecified. In particular, a referee suggested to clarify the marginal contribution of this paper with respect to the work of Gon\c{c}alves and White (2004) which proves the first-order validity of the bootstrap for $t$ tests based on the quasi-maximum likelihood (QML) estimators studentized with the misspecification-robust variance estimator of White (1982).

First, the QML estimator is a special case of the GMM estimator when one uses the first-order condition of the QML as the moment condition. This also puts an additional restriction that the model is just-identified. Therefore, this paper covers a broader class of models than Gon\c calves and White (2004). For example, the proposed bootstrap applies to the two-stage least squares (2SLS) estimator. In addition, the definition of misspecified moment condition model should be distinguished from that of misspecified likelihood function. The former arises only when the model is overidentified, which implies that the first-order condition of the QML forms a correctly specified moment condition even if the likelihood function is misspecified. Thus, the misspecification-robust QML variance estimator corresponds to the conventional GMM variance estimator under correct specification, rather than the Hall-Inoue variance estimator.\footnote{Hall and Inoue (2003) explain their marginal contribution over Gallant and White (1988), White (1996), and Maasoumi and Phillips (1982) in this regard.}

Second, Gon\c calves and White (2004) neither provide a guidance whether to recenter or not, nor explain the relationship between recentering and misspecification. One of the contributions of Hall and Horowitz (1996) is that bootstrapping for GMM is non-standard so that one should recenter the moment function to achieve asymptotic refinements. I argue that recentering can be detrimental and is not even needed if we use the Hall-Inoue variance estimator. The key idea is to link the misspecified moment condition in the bootstrap world to the large sample theory of GMM under misspecification of Hall and Inoue (2003).

The remainder of the paper is organized as follows. Section (ref) discusses theoretical and empirical implications of misspecified models and explains the advantage of using the MR bootstrap $t$ tests and CI's. Section (ref) outlines the main result. Section (ref) defines the estimators and test statistics. Section (ref) defines the nonparametric iid MR bootstrap for iid data. Section (ref) states the assumptions and establishes asymptotic refinements of the MR bootstrap. Section (ref) presents Monte Carlo simulation results. Section (ref) concludes the paper. Lemmas and proofs are gathered in the Appendix.

Why We Care About Misspecification

Empirical studies in the economics literature often report a significant $J$ statistic along with GMM estimates, standard errors, and CI's. Such examples include Imbens and Lancaster (1994), Jondeau, Le Bihan, and Galles (2004), Parker and Julliard (2005), and Ag\"{u}ero and Marks (2008), among others. Significant $J$ statistics are also quite common in the instrumental variables literature using the 2SLS estimator, which is a special case of the GMM estimator.

A significant $J$ statistic means that the test rejects the null hypothesis of correct model specification. For 2SLS estimators, this implies that at least one of the instruments is invalid. The problem is that, even if models are likely to be misspecified, inferences are made using the asymptotic theory for correctly specified models and the estimates are interpreted with economic implications. Various authors justify this by noting that the $J$ test over-rejects the correct null in small samples.

On the other hand, comparing and evaluating the relative fit of competing models have been an important research topic. Vuong (1989), Rivers and Vuong (2002), and Kitamura (2003) suggest various tests of the null hypothesis that test whether two possibly misspecified models provide equivalent approximation to the true model in terms of the Kullback-Leibler information criteria (KLIC). Recent studies such as Chen, Hong, and Shum (2007), Marmer and Otsu (2012), and Shi (2013) generalize and modify the test in broader settings. Hall and Pelletier (2011) show that the limiting distribution of the Rivers-Vuong test statistic may not be consistently estimable unless both models are misspecified. In this framework, therefore, all competing models are misspecified and the test selects a less misspecified model. For applications of the Rivers-Vuong test, see French and Jones (2004), Gowrisankaran and Rysman (2009), and Bonnet and Dubois (2010).

Either for the empirical studies that report a significant $J$ statistic, or for a model selected by the Rivers-Vuong test, inferences about the parameters should take into account a possible misspecification in the model. Otherwise, such inferences would be misleading.

example[Example: Combining Micro and Macro Data]\ Imbens and Lancaster (1994) suggest an econometric procedure that uses nearly exact information on the marginal distribution of economic variables to improve accuracy of estimation. As an application, the authors estimate the following probit model for employment: For an individual $i$, \begin{eqnarray} P(L_{i}=1|Age_{i},Edu_{i})&=&\Phi(X_{i}'\theta)\\ \nonumber &=&\Phi(\theta_{0}+\theta_{1}\cdot Edu_{i} + \theta_{2}\cdot(Age_{i}-35) + \theta_{3}\cdot (Age_{i}-35)^{2}), \end{eqnarray} with $X_{i}=(1,Edu_{i},Age_{i}-35,(Age_{i}-35)^{2})'$ and $\Phi(\cdot)$ is the standard normal cdf. $L_{i}$ is labor market status ($L_{i}=1$ when employed), $Edu_{i}$ is education level in five categories, and $Age_{i}$ is age in years. The sample is a micro data set on Dutch labor market histories and the number of observations is 347. Typically, the probit model is estimated by the ML estimator. The first row of Table (ref) presents the ML point estimates and the standard errors. None of the coefficients are statistically significant except for that of the intercept. To reduce the standard errors of the estimators, the authors use additional information on the population from the national statistics. By using the statistical yearbooks for the Netherlands which contain 2.355 million observations, they calculated the probability of being employed given the age category (denoted by $p_{k}$ where the index for the age category $k=1,2,3,4,5$) and the probability of being in a particular age category (denoted by $q_{k}$). These probabilities are considered as the true population parameters. The authors suggest to use GMM estimators with the moment function that utilizes the information from the aggregate statistic. The second row of Table (ref) reports the two-step efficient GMM point estimates and the standard errors. Now the coefficient $\theta_{3}$ is statistically significant at 1% level and the authors argue “...Age is not ancillary anymore and knowledge about its marginal distribution is informative about $\theta$.” Although they could successfully improve the accuracy of the estimators by combining two data sets, their argument has a potential problem. The last column of Table (ref) reports the $J$ test statistic and its $p$-value. Since the $p$-value is 4.4%, the model is marginally rejected at 5% level. The problem is that, if the model is truly misspecified, the reported GMM standard errors are inconsistent because the conventional standard errors are only consistent under correct specification. Then the authors' argument about the coefficient estimates may be flawed. This problem could be avoided if the standard errors which are consistent even under misspecification were used. The formulas for the misspecification-robust standard errors for the GMM estimators are available in Section 4.\footnote{Since the original data sets used in Imbens and Lancaster (1994) are not available, I could not calculate the robust standard errors. Instead, I provide simulation result with a simple hypothetical model that utilizes additional population information in estimation in Section 7.1.}

When the model is misspecified, $Eg(X_{i},\theta)\neq 0$ for all $\theta$, where $\theta$ is a parameter of interest, $X_{i}$ is a random vector, $g(X_{i},\theta)$ is a known moment function, and $E[\cdot]$ denotes mathematical expectation. Let $\hat{\theta}$ be the GMM estimator and $\Omega^{-1}$ be a positive definite matrix, which is the probability limit of a weight matrix. According to Hall and Inoue (2003), (i) the probability limit of $\hat{\theta}$ is the pseudo-true value that depends on $\Omega^{-1}$ such that

equation[equation omitted — 135 chars of source]

and (ii) the asymptotic distribution of the GMM estimator is

equation[equation omitted — 110 chars of source]

where $\Sigma_{MR}$ is the asymptotic covariance matrix under misspecification that is different from $\Sigma_{C}$, the asymptotic covariance matrix under correct specification. If the model is correctly specified, then $\theta_{0}(\Omega^{-1})$ and $\Sigma_{MR}$ simplify to $\theta_{0}$ and $\Sigma_{C}$, respectively.

The pseudo-true value can be interpreted as the best approximation to the true value, if any, given the weight matrix. The dependence of the pseudo-true value on the weight matrix may make the interpretation of the estimand unclear. Nevertheless, the literature on estimation under misspecification considers the pseudo-true value as a valid estimand, see Sawa (1978), White (1982), and Schennach (2007) for more discussions. Other pseudo-true values that minimize the generalized empirical likelihood (GEL) without using a weight matrix, have better interpretations but comparing different pseudo-true values is beyond the scope of this paper.

Although we cannot fix a potential bias in the pseudo-true value in general, we can report the standard error of the GMM estimator as honest as possible. (ref) implies that the conventional $t$ tests and CI's are invalid under misspecification, because the conventional standard errors are based on the estimate of $\Sigma_{C}$. Misspecification-robust standard errors are calculated using the Hall-Inoue variance estimator of $\Sigma_{MR}$. By using the robust standard errors, the resulting asymptotic $t$ tests and CI's are robust to misspecification. The MR bootstrap $t$ tests and CI's improve upon these MR asymptotic $t$ tests and CI's in terms of the magnitude of errors in test rejection probability and CI coverage probability. A summary on the advantage of the MR bootstrap over the existing asymptotic and bootstrap $t$ tests and CI's is given in Table (ref).

One may consider local misspecification to model a slight misspecification which may not be detected by the $J$ test. A recent development on this topic includes the works of Bravo (2010), Berkowitz, Caner, and Fang (2008, 2012), DiTraglia (2012), Guggenberger (2012), Guggenberger and Kumar (2012), Hall (2005), and Otsu (2011). Local misspecification enables us to make a better interpretation of the pseudo-true value. To see this, let a triangular array $\{X_{n,i}\}_{i\leq n}$ be iid over $i$ for fixed $n$, where $n$ is the sample size. The moment condition is locally misspecified if

equation[equation omitted — 75 chars of source]

where $\theta_{0}$ is a true parameter and $\delta$ is an unknown vector of constants. Since the GMM estimator $\hat{\theta}$ is not $\sqrt{n}$-consistent for $\theta_{0}$ in this setting, the MR bootstrap CI as well as the conventional CI's does not give asymptotically correct coverage for $\theta_{0}$.

Outline of the Results

In this section, I outline the MR bootstrap. The idea of the MR bootstrap procedure can be best understood in the same framework with Hall and Horowitz (1996) and Brown and Newey (2002), as is described below.

Suppose that the random sample is $\chi_{n}=\{X_{i}:i\leq n\}$ from a probability distribution $P$. Let $F$ be the corresponding cumulative distribution function (cdf). The empirical distribution function (edf) is denoted by $F_{n}$. The GMM estimator, $\hat{\theta}$, minimizes a sample criterion function, $J_{n}(\theta)$. Suppose that $\theta$ is a scalar for notational brevity. Let $\hat{\Sigma}$ be a consistent estimator of the asymptotic variance of $\sqrt{n}(\hat{\theta}-plim(\hat{\theta}))$.

I also define the bootstrap sample. Let $\chi_{n_{b}}^{*}=\{X_{i}^{*}:i\leq n_{b}\}$ be a sample of random vectors from the empirical distribution $P^{*}$ conditional on $\chi_{n}$ with the edf $F_{n}$. In this section, I distinguish $n$ and $n_{b}$, which helps to understand the concept of the conditional asymptotic distribution.\footnote{$n_{b}$ is the resample size and should be distinguished from the number of bootstrap replication (or resampling), often denoted by $B$. See Bickel and Freedman (1981) for further discussion.} I set $n=n_{b}$ from the following section. Define $J_{n_{b}}^{*}(\theta)$ and $\hat{\Sigma}^{*}$ like $J_{n}(\theta)$ and $\hat{\Sigma}$ are defined, but with $\chi_{n_{b}}^{*}$ in place of $\chi_{n}$. The bootstrap GMM estimator $\hat{\theta}^{*}$ minimizes $J_{n_{b}}^{*}(\theta)$.

Consider a symmetric two-sided test of the null hypothesis $H_{0}:\theta=\theta_{0}$ with level $\alpha$. The $t$ statistic under $H_{0}$ is $T(\chi_{n})=(\hat{\theta}-\theta_{0})/\sqrt{\hat{\Sigma}/n}$, a functional of $\chi_{n}$. One rejects the null hypothesis if $|T(\chi_{n})|>z$ for a critical value $z$. I also consider a $100(1-\alpha)\%$ CI for $\theta_{0}$, $[\hat{\theta}\pm z \sqrt{\hat{\Sigma}/n}]$. For the asymptotic test or the asymptotic CI, set $z=z_{\alpha/2}$, where $z_{\alpha/2}$ is the $1-\alpha/2$ quantile of a standard normal distribution. For the bootstrap test or the symmetric percentile-$t$ interval, set $z=z^{*}_{|T|,\alpha}$, where $z^{*}_{|T|,\alpha}$ is the $1-\alpha$ quantile of the distribution of $|T(\chi^{*}_{n_{b}})|\equiv |\hat{\theta}^{*}-\hat{\theta}|/\sqrt{\hat{\Sigma}^{*}/n_{b}}$.

Let $H_{n}(z,F)=P(T(\chi_{n})\leq z|F)$ and $H_{n_{b}}^{*}(z,F_{n})=P(T(\chi_{n_{b}}^{*})\leq z|F_{n})$. According to Hall (1992), under regularity conditions, $H_{n}(z,F)$ and $H_{n_{b}}^{*}(z,F_{n})$ allow Edgeworth expansion of the form

eqnarray[eqnarray omitted — 252 chars of source]

uniformly over $z$, where $q_{1}(z,F)$ is an even function of $z$ for each $F$, $q_{2}(z,F)$ is an odd function of $z$ for each $F$, $q_{2}(z,F_{n})\rightarrow q_{2}(z,F)$ almost surely as $n\rightarrow\infty$ uniformly over $z$, $H_{\infty}(z,F)=\lim_{n\rightarrow\infty}H_{n}(z,F)$ and $H_{\infty}^{*}(z,F_{n})=\lim_{n_{b}\rightarrow\infty}H_{n_{b}}^{*}(z,F_{n})$. If $T(\cdot)$ is asymptotically pivotal, then $H_{\infty}(z,F)=H_{\infty}^{*}(z,F_{n})=\Phi(z)$ where $\Phi$ is the standard normal cdf, because $H_{\infty}(z,F)$ and $H_{\infty}^{*}(z,F_{n})$ do not depend on the underlying cdf.

Using (ref) and the fact that $q_{1}$ is even, it can be shown that under $H_{0}$,

equation[equation omitted — 129 chars of source]

where $CI=[\hat{\theta}\pm z_{\alpha/2}\sqrt{\hat{\Sigma}/n}]$. In other words, the error in the rejection probability and coverage probability of the asymptotic two-sided $t$ test and CI is $O(n^{-1})$.

For the bootstrap $t$ test and CI, subtract (ref) from (ref), use the fact that $q_{1}$ is even, and set $n_{b}=n$ to show, under $H_{0}$,

equation[equation omitted — 139 chars of source]

where $CI^{*}=[\hat{\theta} \pm z^{*}_{|T|,\alpha}\sqrt{\hat{\Sigma}/n}]$. The elimination of the leading terms in (ref) and (ref) is the source of asymptotic refinements of bootstrapping the asymptotically pivotal statistics (Beran, 1988; Hall, 1992).

First, suppose that the model is correctly specified, $Eg(X_{i},\theta_{0})=0$ for unique $\theta_{0}$, where $E[\cdot]$ is the expectation with respect to the cdf F. The conventional $t$ statistic $T_{C}(\chi_{n})=(\hat{\theta}-\theta_{0})/\sqrt{\hat{\Sigma}_{C}/n}$, where $\hat{\Sigma}_{C}$ is the standard GMM variance estimator, is asymptotically pivotal. However, a naive bootstrap $t$ statistic without recentering,\footnote{A naive bootstrap for GMM is constructing $\hat{\theta}^{*}$ and $\hat{\Sigma}^{*}$ in the same way we construct $\hat{\theta}$ and $\hat{\Sigma}$, using the bootstrap sample $\chi_{n_{b}}^{*}$ in place of $\chi_{n}$.} $T_{C}(\chi_{n_{b}}^{*})=(\hat{\theta}^{*}-\hat{\theta})/\sqrt{\hat{\Sigma}_{C}^{*}/n_{b}}$, is not asymptotically pivotal because the moment condition under $F_{n}$ is misspecified, $E_{F_{n}}g(X_{i}^{*},\hat{\theta})=n^{-1}\sum_{i=1}^{n}g(X_{i},\hat{\theta})\neq 0$ almost surely when the model is overidentified, where $E_{F_{n}}[\cdot]$ is the expectation with respect to $F_{n}$. If the moment condition is misspecified, the conventional GMM variance estimator is no longer consistent. Note that the bootstrap moment condition is evaluated at $\hat{\theta}$, where $\hat{\theta}$ is considered as the true value given $F_{n}$.

The recentered bootstrap makes the bootstrap moment condition hold so that the recentered bootstrap $t$ statistic is asymptotically pivotal. For instance, the Hall-Horowitz bootstrap uses a recentered moment function $g^{*}(X_{i}^{*},\theta)=g(X_{i}^{*},\theta)-n^{-1}\sum_{i=1}^{n}g(X_{i},\hat{\theta})$ so that $E_{F_{n}}g^{*}(X_{i}^{*},\hat{\theta})=0$ almost surely. The Brown-Newey bootstrap uses the EL distribution function $\hat{F}_{EL}(z)=n^{-1}\sum_{i=1}^{n}\hat{p}_{i}\mathbf{1}(X_{i}\leq z)$ in resampling, where $\hat{p}_{i}$ is the EL probability and $\mathbf{1}(\cdot)$ is an indicator function, instead of using $F_{n}$, so that $E_{\hat{F}_{EL}}g(X_{i}^{*},\hat{\theta})=0$ almost surely, where $E_{\hat{F}_{EL}}[\cdot]$ is the expectation with respect to $\hat{F}_{EL}$.

The MR bootstrap uses the original non-recentered moment function in implementing the bootstrap and resamples according to the edf $F_{n}$. This is similar to the naive bootstrap. The distinction is that the MR bootstrap uses the Hall-Inoue variance estimator in constructing the sample and the bootstrap versions of the $t$ statistic instead of using the conventional GMM variance estimator. The sample $t$ statistic is $T_{MR}(\chi_{n})=(\hat{\theta}-\theta_{0})/\sqrt{\hat{\Sigma}_{MR}/n}$, where $\hat{\Sigma}_{MR}$ is a consistent estimator of $\Sigma_{MR}$, the asymptotic variance of the GMM estimator regardless of misspecification. $T_{MR}(\chi_{n})$ is asymptotically pivotal.

The MR bootstrap $t$ statistic is $T_{MR}(\chi_{n_{b}}^{*})=(\hat{\theta}^{*}-\hat{\theta})/\sqrt{\hat{\Sigma}_{MR}^{*}/n_{b}}$, where $\hat{\Sigma}_{MR}^{*}$ uses the same formula as $\hat{\Sigma}_{MR}$ with $\chi_{n_{b}}^{*}$ in place of $\chi_{n}$. $\hat{\Sigma}_{MR}^{*}$ is consistent for the conditional asymptotic variance of the bootstrap GMM estimator, $\Sigma_{MR|F_{n}}$, almost surely, even if the bootstrap moment condition is not satisfied. As a result, $T_{MR}(\chi_{n_{b}}^{*})$ is asymptotically pivotal. Therefore, the MR bootstrap achieves asymptotic refinements without recentering under correct specification.

Now suppose that the model is misspecified in the population, $Eg(X_{i},\theta)\neq 0$ for all $\theta$. The advantage of the MR bootstrap is that neither the sample $t$ statistic nor the bootstrap $t$ statistic requires the assumption of correct model. Since $T_{MR}(\chi_{n})$ and $T_{MR}(\chi_{n_{b}}^{*})$ are constructed by using the Hall-Inoue variance estimator, they are asymptotically pivotal regardless of model misspecification. Thus, the ability of achieving asymptotic refinements of the MR bootstrap is not affected.

The conclusion changes dramatically for the recentered bootstrap, however. First of all, the conventional $t$ statistic $T_{C}(\chi_{n})$ is no longer asymptotically pivotal and this invalidates the use of the asymptotic $t$ test and CI. Moreover, the recentered bootstrap $t$ test and CI are not first-order valid because (i) they use the inconsistent conventional standard error and, (ii) they impose a wrong moment condition by recentering.\footnote{The conditional and unconditional distributions of the recentered bootstrap $t$ statistic is described in Supplementary Appendix available at the author's webpage.}

Let $z^{*}_{|T_{MR}|,\alpha}$ be the $1-\alpha$ quantile of the distribution of $|T_{MR}(\chi^{*}_{n_{b}})|$ and let $CI_{MR}^{*}=[\hat{\theta} \pm z^{*}_{|T_{MR}|,\alpha}\sqrt{\hat{\Sigma}_{MR}/n}]$. Using the MR bootstrap without assuming the correct model, I show that, under $H_{0}$,

equation[equation omitted — 155 chars of source]

This rate is sharp. The further reduction in the error from $o(n^{-1})$ of (ref) to $O(n^{-2})$ of (ref) is based on the argument given in Hall (1988). Andrews (2002) shows the same sharp bound using the Hall-Horowitz bootstrap and assuming the correct model.

Estimators and Test Statistics

Given an $L_{g}\times 1$ vector of moment conditions $g(X_{i},\theta)$, where $\theta$ is $L_{\theta}\times 1$, and $L_{g}\geq L_{\theta}$, define a correctly specified and a misspecified model as follows: The model is correctly specified if there exists a unique value $\theta_{0}$ in $\Theta\subset\mathbb{R}^{L_{\theta}}$ such that $Eg(X_{i},\theta_{0})=0$, and the model is misspecified if there exists no $\theta$ in $\Theta\subset\mathbb{R}^{L_{\theta}}$ such that $Eg(X_{i},\theta)=0$. That is, $Eg(X_{i},\theta)=g(\theta)$ where $g:\Theta\rightarrow\mathbb{R}^{L_{g}}$ such that $\|g(\theta)\|>0$ for all $\theta\in\Theta$, if the model is misspecified. Assume that the model is possibly misspecified.

The (pseudo-)true parameter $\theta_{0}$ minimizes the population criterion function,

equation[equation omitted — 84 chars of source]

where $\Omega^{-1}$ is the probability limit of a weight matrix. Since the model is possibly misspecified, the moment condition and the population criterion may not equal to zero for any $\theta\in\Theta$. In this case, the minimizer of the population criterion depends on $\Omega^{-1}$ and is denoted by $\theta_{0}(\Omega^{-1})$. We call $\theta_{0}(\Omega^{-1})$ the pseudo-true value. The dependence vanishes when the model is correctly specified.

Consider two forms of GMM estimator. The first one is a one-step GMM estimator using the identity matrix $I_{L_{g}}$ as a weight matrix, which is the common usage. The second one is a two-step GMM estimator using a weight matrix constructed from the one-step GMM estimator. Under correct specifications, the common choice of the weight matrix is an asymptotically optimal one. However, the optimality is not established under misspecification because the asymptotic covariance matrix of the two-step GMM estimator cannot be simplified to the efficient one under correct specification.

The one-step GMM estimator, $\hat{\theta}_{(1)}$, solves

equation[equation omitted — 174 chars of source]

The two-step GMM estimator, $\hat{\theta}_{(2)}$ solves

equation[equation omitted — 220 chars of source]

where\footnote{One may consider an $L_{g}\times L_{g}$ nonrandom positive-definite symmetric matrix for the one-step GMM estimator or the uncentered weight matrix, $W_{n}(\theta)=(n^{-1}\sum_{i=1}^{n}g(X_{i},\theta)g(X_{i},\theta)')^{-1}$, for the two-step GMM estimator. This does not affect the main result of the paper, though the resulting pseudo-true values are different. In practice, however, the uncentered weight matrix may not behave well under misspecification, because the elements of the uncentered weight matrix include bias terms of the moment function. See Hall (2000) for more discussion on the issue.}

equation[equation omitted — 134 chars of source]

and $g_{n}(\theta) = n^{-1}\sum_{i=1}^{n}g(X_{i},\theta)$. Suppress the dependence of $W_{n}$ on $\theta$ and write $W_{n}\equiv W_{n}(\hat{\theta}_{(1)})$. Under regularity conditions, the GMM estimators are consistent: $\hat{\theta}_{(1)}$ converges to a pseudo-true value $\theta_{0}(I)\equiv\theta_{0(1)}$, and $\hat{\theta}_{(2)}$ converges to a pseudo-true value $\theta_{0}(W)\equiv\theta_{0(2)}$. Under misspecification, $\theta_{0(1)}\neq\theta_{0(2)}$ in general. The probability limit of the weight matrix $W_{n}$ is $W = \left\{E[(g(X_{i},\theta_{0(1)})-g_{0(1)})(g(X_{i},\theta_{0(1)})-g_{0(1)})']\right\}^{-1}$, where $g_{0(j)}=Eg(X_{i},\theta_{0(j)})$ for $j=1,2$.

To further simplify notation, let $G(X_{i},\theta)=(\partial/\partial\theta')g(X_{i},\theta)$,

equation[equation omitted — 163 chars of source]

for $j=1,2$, and $L_{\theta}\times L_{\theta}$ matrices $H_{0(1)}=G_{0(1)}' G_{0(1)}+(g_{0(1)}'\otimes I_{L_{\theta}})G_{0(1)}^{(2)}$ and $H_{0(2)}=G_{0(2)}'W G_{0(2)}+(g_{0(2)}'W\otimes I_{L_{\theta}})G_{0(2)}^{(2)}$. Let

equation[equation omitted — 184 chars of source]

$G_{n(j)}=G_{n}(\hat{\theta}_{(j)})$ for $j=1,2$, and $H_{n(1)}=G_{n(1)}' G_{n(1)}+(g_{n(1)}'\otimes I_{L_{\theta}})G_{n(1)}^{(2)}$ and $H_{n(2)}=G_{n(2)}'W_{n} G_{n(2)}+(g_{n(2)}'W_{n}\otimes I_{L_{\theta}})G_{n(2)}^{(2)}$. Let $\Omega_{1}$ and $\Omega_{2}$ denote positive-definite matrices such that

equation[equation omitted — 275 chars of source]

and

equation[equation omitted — 305 chars of source]

To obtain the MR asymptotic covariance matrix for the GMM estimator, I use Theorems 1 and 2 of Hall and Inoue (2003):

equation[equation omitted — 93 chars of source]

where $\Sigma_{MR(j)}=H_{0(j)}^{-1}V_{j}H_{0(j)}^{-1'}$, for $j=1,2,$

eqnarray[eqnarray omitted — 504 chars of source]

Under correct specifications, $\Sigma_{MR(1)}$ and $\Sigma_{MR(2)}$ reduce to the standard asymptotic covariance matrices of the GMM estimators, $\Sigma_{C(1)}$ and $\Sigma_{C(2)}$ respectively, where

equation[equation omitted — 152 chars of source]

$G_{0}=EG(X_{i},\theta_{0})$, $\Omega_{C} = E[g(X_{i},\theta_{0})g(X_{i},\theta_{0})']$, and $\theta_{0}$ satisfies $Eg(X_{i},\theta_{0})=0$.

A consistent estimator of $\Sigma_{MR(j)}$ is $\hat{\Sigma}_{MR(j)} = H_{n(j)}^{-1}V_{n(j)}H_{n(j)}^{-1'}$ for $j=1,2,$ where

eqnarray[eqnarray omitted — 524 chars of source]

and $\Omega_{n(j)}$ is a consistent estimator of $\Omega_{j}$, with the population moments replaced by the sample moments. In particular,

eqnarray[eqnarray omitted — 1,512 chars of source]

where\footnote{Note that $W_{n}-W=-W(W_{n}^{-1}-W^{-1})W_{n}$.}

equation[equation omitted — 182 chars of source]

The diagonal elements of the covariance estimator $\hat{\Sigma}_{MR(j)}$ for $j=1,2$ are the Hall-Inoue variance estimators. In practice, the estimation of the MR covariance matrices does not involve much complication. What we need to calculate additionally is the second derivative of the moment function.

Let $\theta_{k}$, $\theta_{0(j),k}$, and $\hat{\theta}_{(j),k}$ denote the $k$th elements of $\theta$, $\theta_{0(j)}$, and $\hat{\theta}_{(j)}$ respectively. Let $(\hat{\Sigma}_{MR(j)})_{kk}$ denote the $(k,k)$th element of $\hat{\Sigma}_{MR(j)}$. The $t$ statistic for testing the null hypothesis $H_{0}:\theta_{k}=\theta_{0(j),k}$ is

equation[equation omitted — 110 chars of source]

where $j=1$ for the one-step GMM estimator and $j=2$ for the two-step GMM estimator. $T_{MR(j)}$ is robust to misspecification because it is asymptotically standard normal under $H_{0}$, without assuming the correct model. $T_{MR(j)}$ is different from the conventional $t$ statistic, because $\hat{\Sigma}_{C(j)}\neq\hat{\Sigma}_{MR(j)}$ in general even under correct specification, for $j=1,2$.\footnote{Applied researchers may be interested in the choice between $T_{MR(1)}$ and $T_{MR(2)}$. However, it is hard to compare them because (i) $\hat{\theta}_{(1)}$ and $\hat{\theta}_{(2)}$ have different probability limits, and (ii) efficiency gain of the two-step GMM does not hold anymore under misspecification. Nevertheless, comparing $T_{C(j)}$ and $T_{MR(j)}$ would be helpful in practice, where $T_{C(j)}$ is the conventional $t$ statistic studentized with $\hat{\Sigma}_{C(j)}$ for $j=1,2$. For example, one might want to use $T_{C(1)}$ instead of $T_{C(2)}$ to avoid a potential finite sample bias in the two-step GMM. In this case, it is recommended to calculate $T_{MR(1)}$ and compare it with $T_{C(1)}$. In general, $T_{MR(1)}$ is a better choice than $T_{C(1)}$ because it is robust to misspecification while it is not necessarily less powerful than $T_{C(1)}$ (see Section 7). A similar argument applies to $T_{C(2)}$ and $T_{MR(2)}$.} Note that $\hat{\Sigma}_{C(j)}$ is a consistent estimator for $\Sigma_{C(j)}$, the asymptotic covariance matrix under correct specification for $j=1,2$.

The MR bootstrap described in the next section achieves asymptotic refinements over the MR asymptotic $t$ test and CI, rather than the conventional non-robust ones. Define the MR asymptotic $t$ test and CI as follows. The symmetric two-sided $t$ test with asymptotic significance level $\alpha$ rejects $H_{0}$ if $|T_{MR(j)}|>z_{\alpha/2}$, where $z_{\alpha/2}$ is the $1-\alpha/2$ quantile of the standard normal distribution. The corresponding CI for $\theta_{0(j),k}$ with asymptotic confidence level $100(1-\alpha)\%$ is $CI_{MR(j)}=[\hat{\theta}_{(j),k}\pm z_{\alpha/2}\sqrt{(\hat{\Sigma}_{MR(j)})_{kk}/n}]$, $j=1,2$. The error in the rejection probability of the $t$ test with $z_{\alpha/2}$ and coverage probability of $CI_{MR(j)}$ is $O(n^{-1})$: Under $H_{0}$, $P\left(|T_{MR(j)}|>z_{\alpha/2}\right)=\alpha+O(n^{-1}) \mbox{ and } P\left(\theta_{0(j),k}\in CI_{MR(j)}\right)=1-\alpha+O(n^{-1}),$ for $j=1,2$.

The Misspecification-Robust Bootstrap

The nonparametric iid bootstrap is implemented by sampling $X_{1}^{*},\cdots,X_{n}^{*}$ randomly with replacement from the sample $X_{1},\cdots,X_{n}$.

The bootstrap one-step GMM estimator, $\hat{\theta}_{(1)}^{*}$ solves:

equation[equation omitted — 173 chars of source]

and the bootstrap two-step GMM estimator $\hat{\theta}_{(2)}^{*}$ solves

equation[equation omitted — 231 chars of source]

where

equation[equation omitted — 154 chars of source]

and $g_{n}^{*}(\theta) = n^{-1}\sum_{i=1}^{n}g(X_{i}^{*},\theta)$. Suppress the dependence of $W_{n}^{*}$ on $\theta$ and write $W_{n}^{*}\equiv W_{n}^{*}(\hat{\theta}_{(1)}^{*})$. To further simplify notation, let

equation[equation omitted — 261 chars of source]

$G_{n(j)}^{*}=G_{n}^{*}(\hat{\theta}^{*}_{(j)})$ for $j=1,2$, and $H_{n(1)}^{*}=G_{n(1)}^{*'} G_{n(1)}^{*}+(g_{n(1)}^{*'}\otimes I_{L_{\theta}})G_{n(1)}^{(2)*}$ and $H_{n(2)}^{*}=G_{n(2)}^{*'}W_{n}^{*}G_{n(2)}^{*}+(g_{n(2)}^{*'}W_{n}^{*}\otimes I_{L_{\theta}})G_{n(2)}^{(2)*}$.

The bootstrap version of the robust covariance matrix estimator $\hat{\Sigma}_{MR(j)}$ is $\hat{\Sigma}_{MR(j)}^{*} = H_{n(j)}^{*-1}V_{n(j)}^{*}H_{n(j)}^{*-1'}$ for $j=1,2,$ where

eqnarray[eqnarray omitted — 552 chars of source]

and $\Omega_{n(j)}^{*}$ is constructed by replacing the sample moments in $\Omega_{n(j)}$ with the bootstrap sample moments. In particular,

eqnarray[eqnarray omitted — 1,656 chars of source]

where

equation[equation omitted — 227 chars of source]

The MR bootstrap $t$ statistic is

equation[equation omitted — 127 chars of source]

for $j=1,2$. Let $z^{*}_{|T_{MR(j)}|,\alpha}$ denote the $1-\alpha$ quantile of $|T_{MR(j)}^{*}|$, $j=1,2$. Following Andrews (2002), we define $z^{*}_{|T_{MR(j)}|,\alpha}$ to be a value that minimizes $|P^{*}(|T_{MR(j)}^{*}|\leq z)-(1-\alpha)|$ over $z\in \mathbf{R}$, since the distribution of $|T_{MR(j)}^{*}|$ is discrete. The symmetric two-sided bootstrap $t$ test of $H_{0}:\theta_{k}=\theta_{0(j),k}$ versus $H_{1}:\theta_{k}\neq \theta_{0(j),k}$ rejects if $|T_{MR(j)}|>z^{*}_{|T_{MR(j)}|,\alpha}$, $j=1,2$, and this test is of asymptotic significance level $\alpha$. The $100(1-\alpha)\%$ symmetric percentile-$t$ interval for $\theta_{0(j),k}$ is, for $j=1,2$,

equation[equation omitted — 131 chars of source]

The MR bootstrap $t$ statistic differs from the recentered bootstrap $t$ statistic. First, unlike the Hall-Horowitz bootstrap, the MR bootstrap GMM estimator is calculated from the original moment function with the bootstrap sample. Second, the Hall-Inoue variance estimator is used to construct the bootstrap $t$ statistic. In the recentered bootstrap, the conventional variance estimator of Hansen (1982) is used.

Main Result

Assumptions

The assumptions are analogous to those of Hall and Horowitz (1996) and Andrews (2002). The main difference is that I do not assume correct model specification. If the model is misspecified, then the probability limits of the one-step and the two-step GMM estimators are different. Thus, we need to distinguish $\theta_{0(1)}$ from $\theta_{0(2)}$, the probability limit of $\hat{\theta}_{(1)}$ and $\hat{\theta}_{(2)}$, respectively. The assumptions are modified to hold for both pseudo-true values. If the model happens to be correctly specified, then the pseudo-true values become identical.

Let $f(X_{i},\theta)$ denote the vector containing the unique components of $g(X_{i},\theta)$ and $g(X_{i},\theta)g(X_{i},\theta)'$, and their derivatives through order $d_{1}\geq 6$ with respect to $\theta$. Let $(\partial^{m}/\partial\theta^{m})g(X_{i},\theta)$ and $(\partial^{m}/\partial\theta^{m})f(X_{i},\theta)$ denote the vectors of partial derivatives with respect to $\theta$ of order $m$ of $g(X_{i},\theta)$ and $f(X_{i},\theta)$, respectively.

assumption$X_{i},i=1,2,...$ are iid.
assumption\ (a) $\Theta$ is compact and $\theta_{0(1)}$ and $\theta_{0(2)}$ are interior points of $\Theta$.\\ (b) $\hat{\theta}_{(1)}$ and $\hat{\theta}_{(2)}$ minimize $J_{n}(\theta,I_{L_{g}})$ and $J_{n}(\theta,W_{n})$ over $\theta\in\Theta$, respectively; $\theta_{0(1)}$ and $\theta_{0(2)}$ are the pseudo-true values that uniquely minimize $J(\theta,I_{L_{g}})$ and $J(\theta,W)$ over $\theta\in\Theta$, respectively; for some function $C_{g}(x)$, $\|g(x,\theta_{1})-g(x,\theta_{2})\|<C_{g}(x)\|\theta_{1}-\theta_{2}\|$ for all $x$ in the support of $X_{1}$ and all $\theta_{1},\theta_{2}\in\Theta$; and $EC_{g}^{q_{1}}(X_{1})<\infty$ and $E\|g(X_{1},\theta)\|^{q_{1}}<\infty$ for all $\theta\in\Theta$ for all $0<q_{1}<\infty$.
assumptionThe followings hold for $j=1,2$.\\ (a) $\Omega_{j}$ is positive definite.\\ (b) $H_{0(j)}$ is nonsingular and $G_{0(j)}$ is full rank $L_{\theta}$.\\ (c) $g(x,\theta)$ is $d=d_{1}+d_{2}$ times differentiable with respect to $\theta$ on $N_{0(j)}$, where $N_{0(j)}$ is some neighborhood of $\theta_{0(j)}$, for all $x$ in the support of $X_{1}$, where $d_{1}\geq6$ and $d_{2}\geq5$.\\ (d) There is a function $C_{\partial f}(X_{1})$ such that $\|(\partial^{m}/\partial\theta^{m})f(X_{1},\theta)-(\partial^{m}/\partial\theta^{m})f(X_{1},\theta_{0(j)})\|\leq C_{\partial f}(X_{1})\|\theta-\theta_{0(j)}\|$ for all $\theta\in N_{0(j)}$ for all $m=0,...,d_{2}$.\\ (e) $EC^{q_{2}}_{\partial f}(X_{1})<\infty$ and $E\|(\partial^{m}/\partial\theta^{m})f(X_{1},\theta_{0(j)})\|^{q_{2}}\leq C_{f}<\infty$ for all $m=0,...,d_{2}$ for some constant $C_{f}$ (that may depend on $q_{2}$) and all $0<q_{2}<\infty$.\\ (f) $f(X_{1},\theta_{0(j)})$ is once differentiable with respect to $X_{1}$ with uniformly continuous first derivative.
assumptionFor $t\in\mathbf{R}^{dim(f)}$ and $j=1,2$, $\limsup_{\|t\|\rightarrow\infty}\left|E\left(\exp(it'f(X_{1},\theta_{0(j)}))\right)\right|<1,$ where $i=\sqrt{-1}$.

Assumption (ref) says that we restrict our attention to iid sample. Hall and Horowitz (1996) and Andrews (2002) deal with dependent data. I focus on iid sample and nonparametric iid bootstrap to emphasize the role of the Hall-Inoue variance estimator in implementing the MR bootstrap without recentering and to avoid the complications arising when constructing blocks to deal with dependent data. For example, the Hall-Horowitz bootstrap needs an additional correction factor as well as recentering for dependent data. The correction factor would also be needed in implementing the MR bootstrap for dependent data. I do not investigate this issue further in this paper.

Assumptions (ref)-(ref) are similar to Assumptions 2-3 of Andrews (2002), except that I eliminate the correct model assumption. In particular, I relax Assumption 2 of Hall and Horowitz (1996) and Assumption 2(b)(i) of Andrews (2002). The moment conditions in Assumptions (ref)-(ref) are not primitive, but they lead to simpler results as in Andrews (2002). Assumption (ref) is the standard Cram\'{e}r condition for iid sample, that is needed to get Edgeworth expansions.

Asymptotic Refinements of the Misspecification-Robust Bootstrap

Theorem (ref) shows that the MR bootstrap symmetric two-sided $t$ test has rejection probability that is correct up to $O(n^{-2})$, and the same magnitude of convergence holds for the MR bootstrap symmetric percentile-$t$ interval. This result extends the results of Theorem 3 of Hall and Horowitz (1996) and Theorem 2(c) of Andrews (2002), because their results hold only under correctly specified models. In other words, the following Theorem establishes that the MR bootstrap achieves the same magnitude of asymptotic refinements with the existing bootstrap procedures, without assuming the correct model and without recentering.

theoremSuppose Assumptions (ref)-(ref) hold. Under $H_{0}:\theta_{k}=\theta_{0(j),k}$, for $j=1,2,$ $$P(|T_{MR(j)}|>z^{*}_{|T_{MR(j)}|,\alpha})=\alpha+O(n^{-2})\hspace{5mm}\mbox{ or }\hspace{5mm}P(\theta_{0(j),k}\in CI_{MR(j)}^{*})=1-\alpha+O(n^{-2}),$$ where $z^{*}_{|T_{MR(j)}|,\alpha}$ is the $1-\alpha$ quantile of the distribution of $|T_{MR(j)}^{*}|$.

Since $P\left(|T_{MR(j)}|>z_{\alpha/2}\right)=\alpha+O(n^{-1})$, the bootstrap critical value has a reduction in the error of rejection probability by a factor of $n^{-1}$ for symmetric two-sided $t$ tests. The symmetric percentile-$t$ interval is formulated by the symmetric two-sided $t$ test, and the CI also has a reduction in the error of coverage probability by a factor of $n^{-1}$.

We note that neither asymptotic refinements nor first-order validity for the $J$ test are established in Theorem (ref). The MR bootstrap is implemented with a misspecified moment condition in the sample, $E^{*}g(X_{i}^{*},\hat{\theta})\neq 0$, where $E^{*}$ is the expectation over the bootstrap sample. Thus, the distribution of the MR bootstrap $J$ statistic does not consistently approximate that of the sample $J$ statistic under the null hypothesis, which is $Eg(X_{i},\theta_{0})=0$.

The proof of the Theorem proceeds by showing that the misspecification-robust $t$ statistic studentized with the Hall-Inoue variance estimator can be approximated by a smooth function of sample moments. Once we establish that the approximation is close enough, we can use the result of Edgeworth expansions for a smooth function in Hall (1992). The proof extensively follows those of Hall and Horowitz (1996) and Andrews (2002). The differences are that I allow for distinct probability limits of the one-step and the two-step GMM estimators, and that no special bootstrap version of the test statistic is needed for the MR bootstrap. Indeed, the recentering creates more complication than it seems even under correct specification, because $\hat{\theta}_{(1)}\neq\hat{\theta}_{(2)}$ in general, which in turn implies that there are two (pseudo-)true values in the bootstrap world. This issue is not explicitly explained in Hall and Horowitz (1996) and Andrews (2002). In contrast, I explicitly distinguish the pseudo-true values in the bootstrap world as well as in the population, which makes the proof given in this paper more straightforward than theirs.

Monte Carlo Experiments

In this section, I compare the actual finite sample coverage probabilities of the asymptotic and bootstrap CI's under correct specification and misspecification.

The conventional asymptotic CI with coverage probability $100(1-\alpha)\%$ is

equation[equation omitted — 90 chars of source]

where $z_{\alpha/2}$ is the $1-\alpha/2$th quantile of the standard normal distribution. The MR asymptotic CI using the Hall-Inoue variance estimator with coverage probability $100(1-\alpha)\%$ is

equation[equation omitted — 92 chars of source]

The only difference between $CI_{MR}$ and $CI_{C}$ is the choice of the variance estimator. Under correct model specification, both the asymptotic CI's have coverage probability $100(1-\alpha)\%$ asymptotically and the error in the coverage probability is $O(n^{-1})$. Under misspecification, $CI_{MR}$ still provides asymptotically correct coverage, but $CI_{C}$ does not because $\hat{\Sigma}_{C}$ is inconsistent.

The Hall-Horowitz and the Brown-Newey bootstrap CI's with coverage probability $100(1-\alpha)\%$ are given by

eqnarray[eqnarray omitted — 205 chars of source]

where $z_{|T_{HH}|,\alpha}^{*}$ and $z_{|T_{BN}|,\alpha}^{*}$ are the $1-\alpha$th quantiles of the bootstrap distribution of the absolute value of the $t$ statistic based on the Hall-Horowitz bootstrap and the Brown-Newey bootstrap, respectively. Both the recentered bootstrap CI's are expected to perform better than $CI_{C}$ under correct specification. However, similar to $CI_{C}$, they do not provide asymptotically correct coverage under misspecification.

The MR bootstrap CI with coverage probability $100(1-\alpha)\%$ is:

equation[equation omitted — 107 chars of source]

where $z_{|T_{MR}|,\alpha}^{*}$ is the $1-\alpha$th quantile of the MR bootstrap distribution of the absolute value of the $t$ statistic. $CI_{MR}^{*}$ is expected to perform better than $CI_{MR}$ regardless of misspecification by Theorem (ref).

Example 1: Combining Data Sets

Suppose that we observe $X_{i}=(Y_{i},Z_{i})'\in\mathbb{R}^{2}$, $i=1,...n$, and we have an econometric model based on $Z_{i}$ with a moment function $g_{1}(Z_{i},\theta)$, where $\theta$ is a parameter of interest. Also, suppose that we know the mean (or other population information) of $Y_{i}$. If $Y_{i}$ and $Z_{i}$ are correlated, we can exploit the known information on $EY_{i}$ to get more accurate estimates of $\theta$. This situation is common in survey sampling: A sample survey consists of a random sample from some population and aggregate statistics from the same population. Imbens and Lancaster (1994) and Hellerstein and Imbens (1999) show how to efficiently combine data sets and make an inference. For more examples, see Imbens (2002) and Section 3.10 of Owen (2001).

Let $g_{1}(Z_{i},\theta)=Z_{i}-\theta$, so that the parameter of interest is the mean of $Z_{i}$. Without the knowledge on $EY_{i}$, the natural estimator is the sample mean of $Z_{i}$. If an additional information, $EY_{i}=0$, is available, then we form the moment function as

equation[equation omitted — 203 chars of source]

Since the number of moment restrictions ($L_{g}=2$) is greater than that of the parameter ($L_{\theta}=1$), the model is overidentified and we can use GMM estimators to estimate $\theta$. If the assumed mean of $Y$ is not true, i.e., $EY_{i}\neq0$, then the model is misspecified because there is no $\theta$ that satisfies $Eg(X_{i},\theta)=0$.

The one-step GMM estimator solving (ref) is given by $\hat{\theta}_{(1)}=\bar{Z}\equiv n^{-1}\sum_{i=1}^{n}Z_{i}$. The two-step GMM estimator solving (ref) and the pseudo-true value are given by

equation[equation omitted — 183 chars of source]

where $\widehat{Var}(Y_{i})=n^{-1}\sum_{i=1}^{n}(Y_{i}-\bar{Y})^{2}$ and $\widehat{Cov}(Y_{i},Z_{i})=n^{-1}\sum_{i=1}^{n}(Y_{i}-\bar{Y})(Z_{i}-\bar{Z})$. Note that the pseudo-true value reduces to $\theta_{0(2)}=EZ_{i}$ when $EY_{i}=0$, i.e., the model is correctly specified.

The conventional asymptotic variance of $\hat{\theta}_{(2)}$ is $\Sigma_{C(2)}=(G_{0}'\Omega_{C}^{-1}G_{0})^{-1}$. The MR asymptotic variance of $\hat{\theta}_{(2)}$ is $\Sigma_{MR(2)}$, where the formula for $\Sigma_{MR(2)}$ is given in the previous section. Note that $\Sigma_{C(2)}$ is a special case of $\Sigma_{MR(2)}$ imposing no misspecification. The following example makes this case clear. Consider a simple data generating process (DGP)

equation[equation omitted — 541 chars of source]

where $0<\rho<1$ is a correlation between $Y_{i}$ and $Z_{i}$, and $(Y_{i},Z_{i})'$ is iid. The assumed mean of $Y_{i}$, zero, may not equal to the true value, $\delta$. Therefore, $\delta$ measures a degree of misspecification. As $\delta$ deviates farther from zero, the degree of misspecification becomes larger. The pseudo-true value is $\theta_{0(2)}=-\rho\delta$, and the asymptotic variances $\Sigma_{C(2)}$ and $\Sigma_{MR(2)}$ are\footnote{See Supplementary Appendix for details about the calculation.}

equation[equation omitted — 95 chars of source]

If the model is correctly specified, then using the additional information reduces the variance of the estimator by $\rho^{2}$, because the asymptotic variance of the sample mean $\bar{Z}$ is $Var(Z_{i})=1$. However, this reduction may not occur when the additional information is misspecified, and furthermore, the conventional variance estimator is inconsistent for the true asymptotic variance, $\Sigma_{MR(2)}$. In contrast, the Hall-Inoue variance estimator is consistent for the true asymptotic variance regardless of misspecification.

To better compare the coverage probabilities of the CI's, I modify the DGP (ref):

equation[equation omitted — 601 chars of source]

where $\sigma$ is a shape parameter.\footnote{Unreported simulation results based on the DGP (ref) are similar to the reported one, although the size distortion of the asymptotic CI's are less severe.} In this case, $Z_{i}$ has a shifted log-normal distribution, and the mean and the variance are 0 and $(e^{\sigma^{2}}-1)e^{\sigma^{2}}$, respectively. Estimating the mean of $Z_{i}$ is a common problem in economics, as many economic data are well approximated by log-normal distributions. The information on the mean of $Y_{i}$ is assumed to be relatively accurate, but may not be exact, which is the source of misspecification.

Table (ref) shows the coverage probabilities of 90% and 95% CI's based on the two-step GMM estimator, $\hat{\theta}_{(2)}$, when $\rho=0.5$ and $\sigma=1.5$ in (ref). The number of Monte Carlo repetition (r) is 5,000, and the number of bootstrap replication (B) is 1,000. $J$ ($J^{*}$) at 5% denotes the actual rejection probabilities of the asymptotic and the Hall-Horowitz bootstrap $J$ test at 5% level.

For a correctly specified model ($\delta=0$), the bootstrap CI's show better performance than the asymptotic CI's for $n=50$, $200$, and $1,000$. One might suspect that $CI_{MR}^{*}$ and $CI_{MR}$ may not work well compared to the conventional CI's under correct specification ($\delta=0$). Interestingly, $CI_{MR}^{*}$ works as good as $CI_{HH}^{*}$ and $CI_{BN}^{*}$, and $CI_{MR}$ works as good as $CI_{C}$ under correct specification. This implies that the two variance estimators $\hat{\Sigma}_{MR}$ and $\hat{\Sigma}_{C}$ do not differ much, but the difference is enough to achieve asymptotic refinements of the bootstrap without recentering. Since $\hat{\Sigma}_{MR}$ involves estimation of the fourth moment of the moment function $g(X_{i},\theta)$, rather than the second moment, $\hat{\Sigma}_{MR}$ may not work well if we consider more complicated nonlinear models and DGP's. Their relative performance under correct specification deserves more research.

For misspecified models ($\delta=-0.3,-0.6,0.6$), only $CI_{MR}^{*}$ and $CI_{MR}$ have asymptotically correct coverage. $CI_{MR}^{*}$ performs better than $CI_{MR}$ regardless of misspecification, which supports asymptotic refinements robust to misspecification. In contrast, the conventional asymptotic and bootstrap CI's are first-order invalid. Their coverage is either significantly lower (when $\delta=-0.6$) or significantly higher (when $\delta=0.6$) than the nominal coverage.\footnote{Under misspecification, the estimation of the empirical likelihood probabilities for the Brown-Newey bootstrap did not work well. For example, convergence failure occurred about 30% of the Monte Carlo repetition when $\delta=0.6$ and $n=1,000$. If this happens, $CI_{BN}^{*}$ has a length zero, which trivially does not cover the pseudo-true value.} In particular, the result when $\delta=0.6$ implies that the conventional CI's may be neither asymptotically correct nor shorter in finite sample under misspecification. Figure (ref) shows the coverage probabilities of the CI's when $n=200$ for different values of $\delta$, and also supports the findings above.

Example 2: Invalid Instrumental Variables

Suppose that there is endogeneity in the linear model $y_{i}=x_{i}\beta_{0} + \varepsilon_{i}$, where $y_{i},x_{i}\in\mathbb{R}$ and $Ex_{i}\varepsilon_{i}\neq 0$, so that the OLS estimator is inconsistent for $\beta_{0}$. Suppose that we have two instruments, $z_{1i}$ and $z_{2i}$. We can estimate $\beta_{0}$ using both instruments by GMM. The moment function is

equation[equation omitted — 230 chars of source]

where $X_{i}=(y_{i},x_{i},z_{1i},z_{2i})'$. This moment function is correctly specified when both instruments are valid, i.e., $Ez_{1i}\varepsilon_{i}=Ez_{2i}\varepsilon_{i}=0$. In practice, a commonly used weight matrix is $W_{n}=(n^{-1}\sum_{i=1}^{n}\mathbf{z}_{i}\mathbf{z}_{i}')^{-1}$, where $\mathbf{z}_{i}=(z_{1i},z_{2i})'$. With this choice of the weight matrix, the one-step GMM estimator $\hat{\beta}_{(1)}$ is equivalent to the 2SLS estimator. If at least one of the instruments is invalid, then only the Hall-Inoue variance estimator $\hat{\Sigma}_{MR}$ is consistent for the true asymptotic variance of $\hat{\beta}_{(1)}$. Neither the conventional GMM variance estimator nor the 2SLS variance estimator is consistent.\footnote{Maasoumi and Phillips (1982) points out that the calculation of the asymptotic variance of overidentified and misspecified IV estimator is very complicated. Their asymptotic variance is a special case of Hall and Inoue (2003).}

Let the DGP be

eqnarray[eqnarray omitted — 1,634 chars of source]

where $(z_{1i},z_{2i}^{0})'$, $(\varepsilon_{i}^{0},u_{i}^{0})'$ are iid. The error terms are log-normally distributed with the mean zero. This DGP satisfies $Ex_{i}\varepsilon_{i}\neq0$, $Ez_{1i}\varepsilon_{i}=0$, and $Ez_{2i}\varepsilon_{i}=\delta$, where $\delta$ measures a degree of misspecification. Therefore, the instrument $z_{1i}$ is valid, while $z_{2i}$ may not. Let $\beta_{0}=0$ for simplicity. The probability limit of $\hat{\beta}_{(1)}$ is

equation[equation omitted — 333 chars of source]

where $\rho_{\varepsilon u}=E\varepsilon_{i}u_{i}$. The pseudo-true value $\beta_{0(1)}$ depends on $\delta$, $\rho_{\varepsilon u}$, $\gamma_{1}$ and $\gamma_{2}$. Thus, it is different from $\beta_{0}=0$ in general. However, larger misspecification does not necessarily imply larger potential bias in the pseudo-true value. To see this, let

equation[equation omitted — 98 chars of source]

Then $\beta_{0(1)}=\beta_{0}=0$ regardless of the value of $\delta$, $\rho_{\varepsilon u}$, and $\gamma_{1}$. Therefore, we can consistently estimate the structural parameter even with invalid instrument in this special case. Moreover, this particular choice of $\gamma_{2}$ can be considered as a strong but potentially invalid instrument. Let $\gamma_{1}=0.25$ so that the first instrument $z_{1i}$ is relatively weak.\footnote{The strength of instruments depends on the magnitude of the reduced form coefficient as well as the number of instruments, e.g., Hahn and Hausman (2002, 2005) and Guggenberger (2008). Since the weak instruments problem is not the main issue of this paper, I do not further investigate it.} When $\delta=0$, then $\gamma_{2}=0$ so that $z_{2i}$ has no explanatory power. However, the instrument becomes stronger as $\delta$ deviates from zero given $\rho_{\varepsilon u}$ is not zero. We can significantly improve the finite sample coverage probability of CI's by using this instrument. Monte Carlo simulation results support this thought experiment.

Table (ref) shows the coverage probabilities of 90% and 95% CI's based on the one-step GMM estimator, $\hat{\beta}_{(1)}$ with $\gamma_{1}=0.25$ and $\gamma_{2}$ in (ref). First, consider the case when $\delta=0$ so that both the instruments are valid but the second one has no explanatory power. The bootstrap CI's provide more accurate coverage than the asymptotic CI's when the model is correctly specified, but the bootstrap does not solve the problem of using a relatively weak instrument, see Hall and Horowitz (1996) for more discussions. Interestingly, the MR CI's show better performance than the conventional CI's when $n=50$ and $n=200$. This finding further supports the use of the MR CI's in practice, especially when one suspects an over-rejection of the $J$ test. There is a noticeable size distortion in the reported $J$ tests. The Hall-Horowitz bootstrap $J$ test shows smaller size distortion than the asymptotic one. Note that the MR bootstrap is not for the $J$ test, because it does not impose the correct specification of the model in implementing the bootstrap.

Now consider the misspecified cases, $\delta=0.25$ and $\delta=0.5$. By using the invalid but relatively strong instrument, the coverage of the MR CI's improves overall. $CI_{MR}^{*}$ performs better than $CI_{MR}$ regardless of misspecification, and there is a significant improvement even when $n=1,000$. In contrast, the conventional CI's are first-order invalid. The $J$ tests seem less powerful to reject the null hypothesis compared to Example 1 (Table (ref)). Furthermore, the Hall-Horowitz bootstrap $J$ test are less powerful than the asymptotic $J$ test.

Figure (ref) shows the coverage probabilities of the CI's over different degrees of misspecification. It reinforces the previous finding: (i) The ability of achieving asymptotic refinements of the bootstrap CI's is clearly demonstrated at $\delta=0$, and $CI_{MR}^{*}$ maintains the ability regardless of misspecification, and (ii) the MR CI's may perform even better than the conventional CI's under correct specification.

Power

Asymptotic refinements of the bootstrap focus on the size, not the power of $t$ tests. Nevertheless, one may wonder the power property of the asymptotic and bootstrap $t$ tests. The null hypothesis is $H_{0}: \theta=\theta_{0(j)}$ for $j=1,2$. Similar to the CI's, we consider five types of two-sided symmetric $t$ tests. We have two $t$ statistics, $T_{C(j)}$ and $T_{MR(j)}$:

equation[equation omitted — 205 chars of source]

where $\hat{\Sigma}_{C(j)}$ and $\hat{\Sigma}_{MR(j)}$ are the conventional variance estimator and the Hall-Inoue variance estimator, respectievly. Let the asymptotic significance level be $\alpha$. The conventional asymptotic $t$ test rejects the null if $|T_{C(j)}|>z_{\alpha/2}$, and is denoted by $t_{C}$. The MR asymptotic $t$ test rejects the null if $|T_{MR(j)}|>z_{\alpha/2}$, and is denoted by $t_{MR}$. The Hall-Horowitz and the Brown-Newey bootstrap $t$ tests reject the null if $|T_{C(j)}|>z_{|T_{HH}|,\alpha}^{*}$ and $|T_{C(j)}|>z_{|T_{BN}|,\alpha}^{*}$, and are denoted by $t_{HH}^{*}$ and $t_{BN}^{*}$, respectively. Finally, the MR bootstrap $t$ test rejects the null if $|T_{MR(j)}|>z_{|T_{MR}|,\alpha}$, and is denoted by $t_{MR}^{*}$.

Figures (ref) and (ref) show the power curves of the $t$ statistics in Examples 1 and 2. Since the $t$ tests show large size distortion as we saw in the previous section, I use the 10% size-corrected critical values for the asymptotic and bootstrap $t$ tests. The number of Monte Carlo repetition (r) is 1,000 and the number of bootstrap replication (B) is 1,000. For each generated sample, the $t$ statistics are evaluated at various values of $\theta$ around the null and the rejection frequency of the $t$ tests is computed using the size-corrected critical values.

The conclusion is mixed. We find from the figures that under correct specification, (i) the asymptotic $t$ tests show better power properties than the bootstrap $t$ tests ($t_{MR}$ over $t_{MR}^{*}$; $t_{C}$ over $t_{HH}^{*}$ and $t_{BN}^{*}$), but (ii) it is difficult to rank between the asymptotic $t$ tests ($t_{MR}$ and $t_{C}$), and among the bootstrap $t$ tests ($t_{MR}^{*}$, $t_{HH}^{*}$, and $t_{BN}^{*}$). Under misspecification, the conventional asymptotic and bootstrap $t$ tests are inconsistent. The power of the MR asymptotic and bootstrap $t$ tests are not necessarily weaker than the ones using standard $t$ statistic (Figure (ref) Panels 2 and 3). In addition, the MR bootstrap $t$ test can be more powerful than the MR asymptotic $t$ test (Figure 4 Panel 3).

Conclusion

Bootstrap critical values allow more accurate inferences and CI's than the asymptotic critical values. To get the bootstrap refinements for GMM estimators, an ad hoc procedure called recentering has been considered as critical in the existing literature. In addition, the conventional bootstrap methods are not robust to unknown model misspecification. In contrast, the proposed MR bootstrap achieves the same rate of asymptotic refinements without recentering, and without assuming correct specification of the model. The key idea is to link the misspecified moment condition in the bootstrap world to the large sample theory of GMM under misspecification of Hall and Inoue (2003).

Possible extensions of this paper would be (i) to see whether the MR bootstrap still works conditional on the event that the $J$ test fails to reject the null as this is likely to happen in practice, and (ii) to apply the MR bootstrap to the GEL estimators.

Acknowledgment

I am very grateful to Bruce Hansen and Jack Porter for their guidance and helpful comments. I also thank Ken West, Xiaoxia Shi, Don Andrews, Ping Yu, and James Morley, as well as seminar participants at Auckland, Iowa, Sogang, Sungkyunkwan, Sydney, UNSW, Wisconsin-Madison, and Yale for their discussions and suggestions. An earlier version of this paper was presented at the 2011 NASM, 2011 AMES, and 2011 Midwest Econometrics Group. Finally, I thank the co-editor, the associate editor, and three referees for their comments and suggestions that greatly improved the presentation of the paper.