EconBase
← Back to paper

Efficient GMM and Weighting Matrix under Misspecification

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

122,072 characters · 16 sections · 0 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Efficient GMM and Weighting Matrix under Misspecification

abstractThis paper develops efficient GMM estimation when the moment conditions are misspecified. We observe that the influence function of the standard GMM estimator under misspecification depends on both the original moment conditions and their Jacobian, motivating a new class of estimators based on augmented moment conditions with recentering. The standard GMM estimator is a special case within this class, and generally suboptimal. By optimally weighting the augmented system, we obtain a misspecification-efficient (ME) estimator with the smallest asymptotic variance for the same GMM pseudo-true value. In linear models, the asymptotic variance of ME estimator reduces to the textbook efficient-GMM variance formula $(G'W^{*}G)^{-1}$, where $W^{*}$ is the inverse of the variance of residualized moments after projection on the Jacobian $G$. We consider a feasible double-recentered bootstrap estimator, which can be considered as a misspecification-robust and efficient version of Hall and Horowitz (1996) recentered bootstrap GMM estimator, and also consider a split-sample ME estimator. Finally, we establish uniform local asymptotic minimax bounds over a class of weighting matrices. We illustrate the proposed methods in simulation and empirical examples.\\ Keywords: Generalized method of moments, misspecification, pseudo-true value, semiparametric efficiency, augmented moment conditions, bootstrap, instrumental variables JEL classification: C12, C13, C14, C26.

Introduction

The generalized method of moments (GMM) offers a robust framework for estimation and inference, and is widely used in applied econometrics. When the model is over-identified, where the number of available moment conditions exceeds the number of parameters, the GMM estimator depends on the choice of weight matrix. Under the assumption of the correct model specification when there exists a unique true parameter that satisfies the population moment conditions exactly equal to zero, the efficient weighting matrix is the inverse of the variance-covariance matrix of the moments. This choice achieves the smallest asymptotic variance among standard GMM estimators (Hansen (1982), Chamberlain (1987)), and researchers often use this efficient weight in practice, for example, an inverse of the variance of the instruments in the instrumental variables (IV) setup, or a Heteroskedasticity and Autocorrelation Consistent (HAC) matrix estimator.

However, economic models are inherently structural approximations, and the overidentifying restriction test is often rejected due to various forms of misspecification; candidate instruments may be invalid, lag structures or functional forms are misspecified or treatment effect heterogeneity makes no parameter satisfies all moment equations exactly. The presence of misspecification has significant implications for the properties of estimators, potentially affecting asymptotic distribution, and consequently, the validity of standard inference procedures. Recognizing this challenge, a growing literature studies the inference for the pseudo-true value, defined as the population minimizer of the objective functions, including Hall and Inoue (2003), M\"{u}ller (2013), Hansen and Lee (2021), and Andrews and Kwon (2024) among others.\footnote{Defining the pseudo-true value as minimizers of population criterion has a long tradition in econometrics, including White (1982), Hall and Inoue (2003), M\"{u}ller (2013), Hansen and Lee (2021), etc. In the asset pricing literature, such as Kan and Robotti (2009) and Gospodinov, Kan, and Robotti (2014), the GMM pseudo-true value can be defined as a minimizer of the distance metric of Hansen and Jagannathan (1997).} Under misspecification, the asymptotic variance of the GMM estimator has different forms than the standard GMM variance estimator, and ignoring this can lead to inconsistent standard error estimation and asymptotically invalid confidence intervals for the pseudo-true values, which was also highlighted in Andrews, Chen and Techhio (2025).

More importantly, misspecification fundamentally changes the role of the weighting matrix, and has different implications than the correct specification setup. The choice of the weighting matrix no longer merely affects the statistical efficiency of the estimator, but it also affects what is being estimated; different weighting matrices corresponds to different target parameters.\footnote{We are well aware that pseudo-true values, in general, may be possibly different from quantities of economic interest (M\"{u}ller (2013), Andrews, Barnhard, and Carlson (2024)). The main focus of this paper is not on the identification and the causal interpretation of the widely used GMM/TSLS estimators, but on the valid and improved inference for the same GMM/TSLS estimand under moment misspecification. We refer the readers to other influential papers (and references therein) such as Koles\'{a}r (2013), Mogstad et al. (2021), Blandhol et al. (2025), and Sloczy\'{n}ski (2024) for advances in this direction in the linear IV model under treatment effect heterogeneity. See also Andrews, Barahona, et al. (2025) for the potentially misspecified general structural model.} This dual role of the weighting matrix under misspecification raises a number of questions that the standard GMM literature does not directly address. Is the “standard" optimal weight, defined under correct specification, still efficient under misspecification? Does any other class of estimators exist with a smaller asymptotic variance than the standard GMM under misspecification for the same target? How should we think about asymptotic efficiency across different weighting matrices when each weighting matrix targets a different estimand? This paper addresses these challenging questions and provides theoretical and some empirical guidance for researchers using overidentified GMM models.

First of all, by focusing on the fixed weighting matrix $W$ and the associated pseudo-true parameter $\theta_W$, we consider the class of M-estimator where the original moment functions and the Jacobian moments are “stacked" to form a vector of augmented moment conditions with recentering. This is based on the novel observation that the “standard" GMM estimators under misspecification can be expressed as an influence function of the augmented moments. The proposed M-estimator is a general class of estimators that include the “standard" misspecified-GMM as a special case. We then define the “efficient estimator" within this class, which we label as a “misspecification-efficient" (ME) estimator, that has the smallest asymptotic variance that is consistent for the same GMM pseudo-true value. By applying the optimal weighting derived from the inverse of the variance matrix of the augmented moments, the proposed ME estimator captures the correlation between the moments and the Jacobian and extracts additional identifying information from the moment misspecification; while the standard misspecified-GMM is suboptimal with particular linear combination of the available moments dictated by the usual GMM first-order condition.

The ME estimator is not directly feasible because its construction requires the population recentering of the original and Jacobian moments. However, our asymptotic result of the oracle-ME estimator is constructive in the sense that the variance of ME estimator can be easily estimated (without knowing these unknown quantities) and reported as a measure of efficiency frontier; showing researchers what could be gained by imposing an additional information from the misspecification and identification strength. Furthermore, our results provide new insights into the conventional (correct-specification) optimal weighting matrix and variance of the efficient GMM. Under the linear model, the asymptotic variance of the ME estimator simplifies to $ V_{ME}(W) = (G^{\prime} \Sigma_{11, 2}^{-1}G)^{-1}$, where $G$ is Jacobian, $\Sigma_{11, 2} = \Sigma_{11} - \Sigma_{12} \Sigma_{22}^{-1} \Sigma_{21} $, and $\Sigma = (

smallmatrix\Sigma_{11} & \Sigma_{12} \\ \Sigma_{21}&\Sigma_{22}

)$ is the variance of the augmented moments. The weighting matrix $W^{*} = \Sigma_{11, 2}^{-1}$ is the inverse of the variance of residualized moment after projection on the Jacobian, while the conventional (correct-specification) efficient weighting matrix $W_0 = \Sigma_{11}^{-1}$ is the inverse of the variance of the original moments.

This variance expression in linear models has two notable features. First, $V_{ME}(W)$ is invariant to $W$; under misspecification with linear moments, different $W$ targets different $\theta_W$, but does not affect how precisely it can be estimated. Second, it coincides with the textbook “efficient GMM" formula $(G'W^{*}G)^{-1}$ with $ W^{*} = \Sigma_{11, 2}^{-1}$, and our result provide a different angle to the well-known downward bias of the efficient-GMM standard errors. In practice, researchers have found that the standard error of the two-step efficient GMM is often severely downward biased, motivating the finite-sample correction of Windmeijer (2005). We show that, with the choice of $W^* = \Sigma_{11,2}^{-1}$ in the linear GMM, the “conventional" efficient GMM formula will be valid for the oracle ME estimator, but it is too “small" for the standard GMM estimator under misspecification, even in large samples. This further echoes the use of the misspecification-robust standard errors (SE), instead of the conventional formula (Andrews, Chen and Techhio (2025)).\footnote{Hwang, Kang, and Lee (2022) further show that, in linear models, the misspecification-robust SE also works as a finite-sample corrections even under correct specification.} We also recommend to report the efficiency frontier $V_{ME}(W) = (G^{\prime} \Sigma_{11, 2}^{-1}G)^{-1}$ alongside with the misspecification-robust standard errors as a transparent efficiency measure, without need to commit to a specific $W$ and specific target parameter $\theta_W$; perhaps it was reported anyway with an incorrect labelling of “conventional-correct specification" SE of the two-step efficient GMM when $ \Sigma_{11, 2}\approx\Sigma_{11}$. In general nonlinear model, however, above arguments fail to hold, and $V_{ME}(W)$ varies with $W$ and no longer have a textbook formula.

We propose three feasible procedures for the oracle-ME estimator. The first is a bootstrap ME-GMM estimator based on the GMM criterion function that recenters both the original moments and the Jacobian at their sample analogs evaluated at a preliminary GMM estimate, with the efficient choice $W^{*} = \Sigma_{11, 2}^{-1}$. We show that the asymptotic distribution of the bootstrap ME-GMM estimator is equal to those of the oracle ME estimator, which allows us to approximate the sampling variation of the oracle-ME estimator and the efficiency bounds considered in this paper. The proposed bootstrap estimator can be considered as a misspecification-robust and misspecification-efficient version of Hall and Horowitz (1996) recentered bootstrap GMM estimator, which is only valid under correct specification. We also consider a closely related “double recentered" (DR) GMM bootstrap that jointly perturbs the moment conditions and the Jacobian similarly to the FOC of the misspecified GMM, and show that the bootstrap distribution mimics the limiting distribution of the standard GMM estimators regardless of model specification. Under correct specification the additional Jacobian recentering correction is higher-order and the DR bootstrap nests the original Hall and Horowitz (1996) procedure as a special case; under misspecification it correctly approximates the Jacobian variation term that the Hall and Horowitz (1996) bootstrap omits. Finally, we also develop a repeated sample-splitting ME estimator, in the spirit of Angrist and Krueger (1995), that uses a held-out half of the sample to estimate the recentering quantities. In the linear model with $W = \Sigma_{11,2}^{-1}$, the moment recentering term drops out by the first-order condition, so only the Jacobian recentering term needs to be estimated.

The recentering has been considered important when applying the bootstrap in the over-identified GMM model to achieve higher-order improvements under correct specification (Hall and Horowitz (1996) and Brown and Newey (2002)). In the linear IV example, the proposed bootstrap estimators can be seen as a linear combinations of the HH recentered GMM/TSLS estimator and the Jacobian estimator, where the recentering of the Jacobian estimator ensure that the bootstrap expected Jacobian matrix evaluated the estimator is zero. Lee (2014) argues that recentering in the standard nonparametric GMM bootstrap can be detrimental and is not even needed if we use the analytic misspecification-robust variance estimator. In our paper, however, we clarify this argument that additional considerations of recentered Jacobian can also achieve asymptotic validity under misspecification. This “double recentering" has been also considered in weak-identification literature to accurately mimic the behaviour of the original Jacobian under weak identification (e.g., Dovonon and Goncalves (2017), Lee and Liao (2018)), yet under correct specification.

In this paper, we view the choice of $W$ as a different candidate models/designs, and we do not evaluate the performances of different weighting matrix $W$ and do not aim to choose the optimal $W$ (and thus $\theta_W$) under certain statistical criteria. The weighting matrix $W$ and the candidate class of $\mathcal{W}$ can be possibly chosen by the researchers a priori, but in general, it may be difficult for researchers view one candidate pseudo-true value better than another, even if it is based on statistically well-defined criteria. For example, we cannot tell that a more precisely estimated LATE (local average treatment effect) for one subpopulation is better than a noisier LATE for different subpopulation; they may answer different causal questions.

In this paper, we develop uniform semiparametric efficiency bounds over $W \in \mathcal{W}$ with the augmented system, which exhausts all additional information from the direction of misspecification and Jacobian, taking into account the correlation between them. We want to highlight that the notion of semiparametric efficiency would be defined differently under (global) misspecification. The pseudo-true value $\theta_W$ changes with $W$ under misspecification, but the semiparametric efficiency framework typically assumes a true-parameter and the efficiency bound is computed at the well-defined single target parameter. We first focus on fixed $W$ and consider the semiparametric efficiency bounds for $\theta_W$, and consider uniform bounds over $W\in \mathcal{W}$.

Our uniform local asymptotic minimax bounds will be not only useful for a researcher who chooses $W$ based on target a specific economic parameter and consider only $\theta_W$, but also useful for a researcher who is completely agnostic to the choice of $W \in \mathcal{W}$. Many applied researchers may just pick one $W$ (e.g., TSLS), however, applied researcher who runs GMM, already considers choosing $W$ over a class of different weighting matrices $\mathcal{W}$ routinely; one-step $W = I$ or $(Z'Z)^{-1}$, two-step, and iterated until convergence. Under misspecification each iteration changes the pseudo-true value. So when results differ across one-step, two-step, and iterated GMM, it can be a direct evidence of sensitivity to $W$.\footnote{Another example is Kan and Robotti (2009) and Gospodinov, Kan, and Robotti (2014) who explicitly study inference under misspecification using Hansen-Jagannathan (1997)'s distance measure, and $W$ denotes different asset portfolios in that context.}

We consider Monte Carlo simulations and three empirical applications to illustrate the proposed methods. The simulation results confirm that the oracle ME is meaningfully tighter, and the DR percentile CI delivers valid coverage, with similar or slightly shorter lengths than the misspecification-robust CI. Repeated sample splitting estimator also performs well, with slightly over coverage and larger lengths as expected. In both of Card (1995) and Angrist and Krueger (1991) returns to schooling examples, the ME efficiency frontier is roughly 10%-20% smaller than the conventional and misspecification-robust standard errors when using efficient weighting matrix $W^*= \Sigma_{11,2}^{-1}$ or $W_0 = \Sigma_{11}^{-1}$. From extensive simulations, and all of the three applications, we found that the ME efficiency bound is most informative when (i) $W$ is efficient weight (either $W_0$ or $W^*$) and far from $W=I$ or other naive choices; (ii) identification strength is moderate, so that Jacobian sampling variation is non-negligible. The identity matrix ($W=I$) is a poor default in our simulations and empirical examples, as the estimates are very different from using efficient weightings and the gap between conventional/robust SE with the oracle ME bound is large (20-40% in Angrist and Krueger (1991) example). Under $W^*$ with moderate identification strength, the gap is typically below 10%, and the misspecification-robust SE is close to the efficiency frontier. Overall, we find that the ME efficiency bounds can be useful, and the proposed DR bootstrap and sample-split procedures provide reasonable alternatives to the standard GMM with analytic misspecification-robust standard errors.

Our paper complements a growing literature on sensitivity analysis and robust inference under misspecification, such as Andrew, Gentzkow and Shapiro (2017), Armstrong and Koles\'{a}r (2021), Bonhomme and Weidner (2022), and Christensen and Connault (2023), among many others. More recent work focuses on the relationship between the pseudo-true value and the true value under misspecification, e.g., Andrews, Barnhard, and Carlson (2024), Andrews, Barahona et al. (2025), and also see Andrews, Chen, and Techhio (2025) for a summary of recent developments in the literature with some constructive recommendations for practitioners. While these papers focus on bias, sensitivity, or optimal inference, this paper analyse efficient estimation in the GMM setup under misspecification, and we provide a feasible efficiency measure that researchers can report alongside robust-standard errors.

There are some limitations to our paper. First, we do not consider two-step and the iterated GMM estimators, and we focus only on the one-step GMM estimator. The textbook motivation of going beyond the one-step GMM estimator is to achieve efficiency, and existing papers (e.g., Hall and Inoue (2003), Hwang, Kang and Lee (2022), and Hansen and Lee (2021)) uses the “standard" optimal weighting-matrix $W_0$, which we show in this paper suboptimal under misspecification. It is not clear why we use the “standard" optimal weighting matrix in each iteration under misspecification. It would be interesting to see the behavior of the two-step/iterated GMM using other weighting matrix schemes, for example, $W^{*} = \Sigma_{11,2}^{-1}(\theta)$ considered in our paper.

Second, our results are constructed under a standard full rank assumption of the Jacobian, and we do not consider the weak-identification settings, where the behavior of the GMM estimators may be nonstandard (e.g., Stock and Wright (2000), Andrews and Cheng (2012), Dovonon and Renault (2013)). While, in this paper, using a Jacobian of the moment condition comes from the influence function of the GMM under misspecification, there are some papers utilize a Jacobian estimation in the weak identification or the lack of first-order local identification in GMM literature. Kleibergen (2005) uses Jacobian estimator that is asymptotically uncorrelated with the original sample moments to develop test statistic that is valid without assuming identification. Lee and Liao (2018) uses zero Jacobian moment conditions, which holds when the rank condition fails to hold, as extra moment conditions in addition to the original moment conditions. Extension of our approach that allow for weak identification would be important, but more challenging. We leave these for future research.

The remainder of the paper is organized as follows. Section (ref) introduces misspecified GMM setup. Section (ref) develops the augmented moment framework and the ME estimator, and derives its asymptotic distribution. Section (ref) analyses the feasible bootstrap estimators - ME-GMM and the double recentered (DR) bootstrap procedure. Section (ref) considers split-sample ME estimator as an alternative. Section (ref) provides the uniform semiparametric efficiency bounds over classes of weighting matrices. Section (ref) reports the Monte Carlo evidence and Section (ref) shows illustrative empirical examples for the proposed methods. Section (ref) concludes. Appendix A includes all proofs of the results in the main paper. Appendix B considers worst-case variance bounds for sensitivity analysis over misspecification and identification strength in the spirit of Conley, Hansen, and Rossi (2012). Appendix C reports additional simulations. Appendix D contains the dynamic-panel application.

Notation

$||A ||$ denotes the spectral norm. $A \geq B$ denotes that the $A-B$ is positive semi-definite for real symmetric matrix A and B. Let $o_p(\cdot)$ and $O_p(\cdot)$ denote the usual stochastic order symbols, and $\overset{p}{\longrightarrow}$ and $ \overset{d}{\longrightarrow}$ denote convergence in probability, convergence in distribution, respectively. Superscript $\phantom{}^{*}$ denotes a probability or moment computed under the bootstrap distribution conditional on the original data set. Let $\xrightarrow{p^{*}}$ and $\xrightarrow{d^{*}}$ denote the convergence in probability, in probability, and the convergence in distribution in probability, respectively. In addition, we write $\xi_{n}=o_{p^{*}}(1)$ if $\xi_{n}\xrightarrow{p^{*}}0$ and $\xi_{n}=O_{p^{*}}(1)$ if $\xi_{n}$ is bounded in probability, in probability.

We introduce further notations on a neighborhood of the probability distribution and the loss function $l(\cdot)$. Let $\Pi$ be the set of all probability measures $F$ defined on the Borel sets in $\mathbb{R}^m$ with the support of a probability measure $\mathcal{X}$. Given an $F \in \Pi$, the basic neighborhoods of $F$ are sets of the form: \[ \left\{Q \in \Pi : \left|\int f_{j}dQ - \int f_{j}dF\right| < \epsilon_{j}, j=1,\dots,q\right\} \] where $\epsilon_{j} > 0$, $q$ is some positive integer, and $f_j: \mathcal{X} \rightarrow \mathbb{R}$ are measurable functions such that $\int |f_{j}| dF < \infty$. An arbitrary neighborhood of $F$ is formed by taking unions of sets of this form. We next define a class of loss function $l \in \mathcal{L}$, if $l(\cdot) \in \mathcal{L}$ satisfies the following conditions, if for all $u, v \in \mathbb{R}$: (i) $l(u)=l(|u|)$, (ii) $|u| \le |v|$ implies $l(u) \le l(v)$, (iii) $\int_{-\infty}^{\infty}l(u)\exp\left(-\frac{1}{2}\lambda u^{2}\right)du < \infty$ for $\lambda > 0$, (iv) $l(0)=0$. Many standard loss function satisfy these conditions, for example, quadratic $l(u) = u^2$, absolute function $l(u) = |u|$, and polynomial loss $l(u) = |u|^p$ for $ p > 0$.

The Setup

Suppose that we have data on observations $\{ X_i\}_{i=1}^{n}$ randomly drawn from the unknown probability distribution $F_0$, with a data vector $X$. Let $g(X, \theta)$ be an $m\times 1$ vector of moment functions and $\theta$ be a $p\times 1$ parameter vector. The model is based on the moment conditions, and a popular estimator of $\theta$ in this setup is the generalized method of moments (GMM) estimator of Hansen (1982),

align[align omitted — 134 chars of source]

where $g_n(\theta) = n^{-1}\sum_{i=1}^{n} g (X_i, \theta)$ and $W$ is the weighting matrix. Under the regularity conditions, the GMM estimator converges to the pseudo-true value $ \theta_{W} $, which is defined as a minimizer of the population objective function

align[align omitted — 117 chars of source]

where $g(\theta) = E[g_n(\theta)]$. Assuming global identification, $ \theta_{W} $ satisfies the following just-identified equation from the FOC of (ref);

align[align omitted — 72 chars of source]

where $G(\theta) = E [\frac{\partial g(X_i, \theta)}{\partial\theta^{\prime}}]$ is a population Jacobian. When the model is correctly specified, i.e., there exists unique $\theta_0$ satisfies the original moment equations

align[align omitted — 78 chars of source]

the pseudo-true value $\theta_W= \theta_0$ for all $W$, and it does not depend on $W$. Even if the moment condition is (globally) misspecified, i.e., $E[g_n (\theta)] \neq0,~\forall\theta \in \Theta$, the FOC condition (ref) still holds. However, the GMM pseudo-true values generally depend on the choice of $W$ under misspecification when the model is overidentified ($m>p$), as it satisfies the linear combinations of non-zero population moments $g(\theta)$ as in (ref), and that the varying $W$ changes the parameter being estimated.

In this paper, our object of interest is on providing valid inference procedures for the pseudo-true value $\theta_W$. It is difficult to give general economic interpretation to the pseudo-true parameter values, and instead, researchers typically analyze and interpret them case-by-case. For example, in the linear IV model, moment misspecification can arise due to the treatment effect heterogeneity, and it is common to focus attention on the TSLS estimand, which corresponds to $\theta_W$ with $W = E[Z_i Z_i^{\prime}]^{-1}$ and the instrumental variables $Z_i$. In this setup, $\theta_W$ can be interpreted as a weighted LATE under appropriate assumptions, see Angrist and Imbens (1995).\\

The novelty of this paper is to explicitly find the influence function of the “standard" GMM estimators under misspecification as a set of augmented moment conditions which include (i) an original moment restriction and (ii) a Jacobian of the moment function, (iii) with recentering. This result, while implicit in the form of the asymptotic variance of the misspecified-GMM in the literature, further suggests a new class of estimators based on the augmented moments. We show in the following sections that the proposed class of estimators include the “standard" GMM as a special case, and we can consider more efficient estimator within the class that has the smallest asymptotic variance that are consistent for the same pseudo-true value $ \theta_{W}$.

To see this, we first observe that the GMM estimator has the following expression under misspecification with the standard regularity conditions (see the proof of Theorem (ref) for details),

equation[equation omitted — 264 chars of source]

where

eqnarray[eqnarray omitted — 698 chars of source]

Based on (ref), we can derive an asymptotic distribution of the GMM estimators under misspecification,

align[align omitted — 145 chars of source]

where \[ V_{GMM}(W, \theta) = (A (W , \theta)\Gamma (\theta ))^{-1} A (W , \theta) \Sigma (\theta ) A (W , \theta)^{\prime} ( ( A (W , \theta)\Gamma (\theta))^{\prime})^{-1}, \] and $\Sigma (\theta) =E [ \psi_{i}(X_i, \theta) \psi_{i}(X_i, \theta)^{\prime}] -E [ \psi_{i}(X_i, \theta)] E [ \psi_{i}(X_i, \theta)]^{\prime} $. The asymptotic distribution results of the misspecified-GMM in (ref) are well-known in the literature; see Imbens (1997), Hall and Inoue (2003), Hansen and Lee (2021) for derivations of same result under misspecification for one-step/two-step GMM, and the iterated GMM estimator.

There exist a few papers in the literature, which consider over-identified GMM models as a just-identified GMM estimator with an augmented parameter vectors. One can generally interpret the overidentified GMM models as a just-identified GMM (Newey and McFadden (1994)), and Imbens (1997) finds the connection with the misspecification-robust variance from the just-identified GMM estimator with an augmented parameter vector. However, they did not consider the explicit asymptotic variance form of the misspecified GMM estimator and the choice of weighting matrix, which are the main focus of this paper.

Efficient Estimator under Misspecification

By inspection of the influence function in (ref), we observe that the misspecified-GMM estimator is related to the M-estimator where the original moment functions and the Jacobian moments are “stacked" to form a vector of augmented moment conditions. We consider the following recentered moment function with augmented moments, $\widetilde{\psi}(X_i, \theta):=\widetilde{\psi}(X_i, \theta; W) = \psi(X_i, \theta) - E[\psi(X_i, \theta_{W})]$, where $\psi(X_i, \theta)$ is defined in (ref), and $\theta_W$ is the GMM pseudo-true value in (ref) for fixed $W$. By construction, the population augmented moments $E[\widetilde{\psi}(X_i, \theta)] $ are zero when $\theta$ is evaluated at $\theta_{W}$.

We can then transform this to the just-identified moment $\Lambda E[\widetilde{\psi}(X_i, \theta_{W})] = 0$ using the transform matrix $\Lambda = [\Lambda_1 \quad \Lambda_2]$, where $\Lambda_1 \neq 0, \Lambda_2 \neq 0$ are $p \times m$, and $p\times mp$ matrix, respectively. Then, we consider the following general class of M-estimator (or Z-estimator), $\widetilde{\theta} (\Lambda)$, to denote an M-estimator based on the sample analogs of augmented moment conditions

equation[equation omitted — 107 chars of source]

where $\widetilde{\psi}_n (\theta) = \frac{1}{n} \sum_{i=1}^{n} \widetilde{\psi}(X_i, \theta)$. Alternative choices of $\Lambda$ are associated with alternative estimators that are consistent for $\theta_W$. This estimator is not feasible without estimation of $E[\psi(X_i, \theta_{W})]$, but we postpone the discussion to focus first on the efficient choice of $\Lambda$ and the asymptotic distribution of the oracle estimator.

Asymptotic Distribution of the M-Estimator with Augmented Moments

The following assumptions are used to show the asymptotic distribution of $\widetilde{\theta} (\Lambda)$.

assumption\begin{enumerate} • The observations $ \{ X_{i} \}_{i=1}^{n}$ are independent and identically distributed (i.i.d.). • For $W >0$, $\theta_{W} = \operatorname*{arg\,min}_{\theta} g(\theta)^{\prime} W g(\theta)$ is unique, and $\theta_{W}$ is in the interior of the compact parameter space $\Theta \subset \mathcal{R}^{p}$, where $g(\theta) \neq 0$ for all $\theta \in \Theta$. • $g(X, \theta)$ and $G(X, \theta)$ are continuous at each $\theta \in \Theta$ with probability one, and $E[\sup_{\theta} || g(X_i, \theta)||] < \infty, E[\sup_{\theta} || G(X_i, \theta)||] < \infty, E[\sup_{\theta} || F(X_i, \theta)||] < \infty$. • $g(X, \theta)$ is twice continuously differentiable in $\theta$, and $E[||g(X, \theta_W) ||^{2}] <\infty$, $E[||G(X, \theta_W) ||^{2}] <\infty$. • $\Lambda E[\widetilde{\psi}(X_i, \theta)] \neq 0 $ for $\theta \neq \theta_W$. $\Lambda \Gamma (\theta_{W})$ exists and is nonsingular. \end{enumerate}

Assumption (ref).3 ensures uniform convergence of the $\widetilde{\psi}_n (\theta)$ and the population objective function $E[\widetilde{\psi}_n (\theta)]$ to show consistency. Assumption (ref).4 is for the asymptotic normality of the normalized sum of $\widetilde{\psi}_n (\theta_{W})$. The first part of Assumption (ref).5 is an identification condition. Necessary and sufficient condition for (local) identification is that $\Lambda \Gamma(\theta_W) = \Lambda_1 G(\theta_W) + \Lambda_2 F(\theta_W)$ has a full column rank. This is also related to the second-order identification condition because even if the Jacobian $G(\theta_W)$ is zero (rank deficient), the model can still be identified by nonzero $\Lambda_2 F(\theta_W)$. More primitive conditions on the identification assumption can be considered. For example, when the moment functions are linear in $\theta$, this reduces to the condition that $\Lambda \Gamma (\theta_W) = \Lambda_1 G$ has a full-column rank, where $G = E[G(X_i)]$ does not depend on $\theta$. When, $\Lambda_1$ takes the form of $\Lambda_1 = G'W$ with $W>0$, it reduces to the standard GMM rank condition, which is implied by the global identification of $\theta_W$ in Assumption (ref).2. \\

theoremUnder Assumption (ref), \begin{align} \sqrt{n}(\widetilde{\theta} (\Lambda) - \theta_{W} ) \overset{d}{\longrightarrow} N(0, V_{M}(\Lambda, \theta_{W})). \end{align} where \[ V_{M} (\Lambda, \theta_W) = (\Lambda \Gamma (\theta_W))^{-1} \Lambda \Sigma (\theta_{W} ) \Lambda ^{\prime} ( (\Lambda \Gamma (\theta_W))^{\prime})^{-1} \] and $\Sigma (\theta) =E [ \psi_{i}(X_i, \theta) \psi_{i}(X_i, \theta)^{\prime}] -E [ \psi_{i}(X_i, \theta)] E [ \psi_{i}(X_i, \theta)]^{\prime} $.\\

Theorem (ref) shows that $\widetilde{\theta} (\Lambda)$ is consistent for the same GMM pseudo-true value $\theta_{W}$, and the augmented-M estimator is a sufficiently general class of estimator to include the “standard" misspecified-GMM as a special case. When $\Lambda = A(W, \theta_W) = [G(\theta_W)^{\prime} W \quad g(\theta_W)^{\prime} W \otimes I_p]$ as in (ref)-(ref), $\widetilde{\theta} (\Lambda)$ has the same influence function representation with the original misspecified-GMM estimator $\widehat{\theta}_{GMM} (W)$ and has the same asymptotic distribution as in (ref).

Efficient Choice of $\Lambda$ and the Efficient Estimator under Misspecification

More importantly, Theorem (ref) justifies the use of our new class of estimator, through the following corollaries, as we can consider “efficient" choice of $\Lambda$ to minimize the asymptotic variance $V_{M} (\Lambda, \theta_{W})$. By using the standard theory of optimal estimating equations (Godambe (1960)) or semi-parametric efficiency of GMM, the optimal choice of $\Lambda $ takes the following form;\footnote{When $\Sigma(\theta_W)$ is singular, we can use the generalized inverse $\Sigma^{-}$ and all the arguments below will still be valid. $\Sigma$ is singular when some components of $vec(G(X_i, \theta))$ has an overlap with the $g(X_i, \theta)$, which occurs for example, when $g(X_i, \theta) = ( X_i-\theta , (X_i-\theta)^2 - 1)^{\prime}$ as $G(X_i, \theta) = (-1, -2 (X_i-\theta))^{\prime}$. In this case, the Jacobian carries no information beyond the original moments, but $(\Gamma'\Sigma^{-}\Gamma)$ may still be nonsingular.}

eqnarray[eqnarray omitted — 292 chars of source]

With this $\Lambda^{*}_{W} $, we use the optimal linear combinations of the augmented moments, while the specific choice $\Lambda = A(W, \theta_W)$ based on the FOC of the GMM objective, is inefficient. When the estimator $\widetilde{\theta} (\Lambda)$ is constructed with $\Lambda^{*}_{W}$ (or any consistent estimator of $\Lambda^{*}_{W}$), we label it as the misspecification-efficient (ME) estimator, $\widehat{\theta}_{ME} (W) = \widetilde{\theta} (\Lambda^{*}_{W})$.\\

corollarySuppose Assumption (ref) holds for $\Lambda =\Lambda^{*}_{W} = \Gamma(\theta_W)^{\prime} \Sigma (\theta_W)^{-1}$. Then, we have \begin{equation} \sqrt{n}(\widehat{\theta}_{ME} (W) - \theta_{W} ) \overset{d}{\longrightarrow} N(0, V_{ME}(W)), \end{equation} where $V_{ME}(W) = (\Gamma(\theta_W)^{\prime} \Sigma (\theta_W)^{-1}\Gamma(\theta_W))^{-1}$, and the following holds for any $\Lambda$ that satisfies Assumption (ref), \[ V_{M} (\Lambda, \theta_W) \geq V_{ME}(W). \] Further, the asymptotic variance of the ME estimator has the following form; \begin{equation} V_{ME}(W) = (G(\theta_W)^{\prime} \Sigma_{11}^{-1} G(\theta_W) +F_G(\theta_W)^{\prime} \Sigma_{22, 1}^{-1} F_G(\theta_W))^{-1}, \end{equation} where $F_G(\theta_{W}) = F(\theta_{W}) - \Sigma_{21} \Sigma_{11}^{-1} G(\theta_{W})$, $\Sigma_{22,1} = (\Sigma_{22} - \Sigma_{21}\Sigma_{11}^{-1} \Sigma_{12})$, and $\Sigma (\theta_W) = \begin{bmatrix} \Sigma_{11} & \Sigma_{12} \\ \Sigma_{21}&\Sigma_{22}\end{bmatrix}$.\\

Corollary (ref) shows that the misspecification-efficient estimator has the smallest asymptotic variance in the class of M-estimator that we consider with the augmented moment conditions $\widetilde{\psi}(X_i, \theta) $. Since misspecified-GMM estimator is the augmented-M estimator $\widetilde{\theta} (\Lambda)$ with $\Lambda = A(W, \theta_W) = [G(\theta_W)^{\prime} W \quad g(\theta_W)^{\prime} W \otimes I_p]$, Corollary (ref) implies that

equation[equation omitted — 117 chars of source]

Corollary (ref) also provides some insights on the “conventional" asymptotic variance of the efficient GMM under correct-specification. The ME variance in (ref) decomposes as the correct-specification efficient GMM variance contribution ($G(\theta_W)^{\prime} \Sigma_{11}^{-1} G(\theta_W)$ evaluated at $\theta_W$ here instead of $\theta_0$) and the curvature term contribution from the second derivative of the moments $F$ after residualization $(F_G(\theta_W)^{\prime} \Sigma_{22, 1}^{-1} F_G(\theta_W))$, and we can deduce that

equation[equation omitted — 116 chars of source]

Thus, using the extra moment conditions from Jacobian does not worsen the asymptotic efficiency of estimators even under correct specification.\footnote{Under correct specification, $E[g (X_i, \theta_0)] = 0$, straightforward calculation shows that the asymptotic variance of the “standard" GMM, $V(W, \theta_0)$ in (ref) reduces to the classical one $ V (W, \theta_0) = (G(\theta_0)'WG(\theta_0))^{-1} G(\theta_0)'W \Sigma_{11}(\theta_0) W G(\theta_0) (G(\theta_0)'WG(\theta_0))^{-1},$ where $\Sigma_{11} (\theta)= E[g(X_i, \theta) g(X_i, \theta)'] - E[g(X_i, \theta)] E[g(X_i, \theta)]^{\prime}$ is the $m \times m$ upper-left submatrix of $\Sigma(\theta)$. Then, the “standard" optimal-weight matrix is $W_0 = \Sigma_{11}(\theta_0)^{-1} $ and the asymptotic variance of $\widehat{\theta}_{GMM}( \Sigma_{11} (\theta_0)^{-1})$ achieves smallest asymptotic variance in the class of (correctly specified) GMM estimator in the sense that $V (W, \theta_0) \geq V( \Sigma_{11}(\theta_0)^{-1}, \theta_0) = (G(\theta_0)' \Sigma_{11}(\theta_0)^{-1} G(\theta_0))^{-1}$ for any non-singular matrix $W>0$. } The ME estimator provides no efficiency gain when $F_G(\theta_{W}) = F(\theta_{W}) - \Sigma_{21} \Sigma_{11}^{-1} G(\theta_{W}) = 0$. In the linear IV model, where $F(\theta_W) = 0$ and $G$ is constant, this condition holds when $\Sigma_{12} = 0$, i.e., the covariance between the moment function $g(X_i, \theta)$ and the Jacobian $G(X_i)$ is zero. However, $\Sigma_{12}$ (and thus $F_G(\theta_{W})$) is generally nonzero both under correct specification and misspecification, implying that the ME estimator provides strictly positive efficiency gains in most empirically relevant settings.\footnote{To see this, consider the simple linear IV model where $g(X_i, \theta) =Z_i (Y_i - X_i\theta) $, and $\theta$ is a scalar. Define $ \varepsilon_i = Y_i - X_i \theta_W, \sigma_{\varepsilon X} = E[\varepsilon_i X_i |Z_i]$, $\sigma_{\varepsilon Z} =E[Z_i \varepsilon_i ]$, and $X_i = \pi^{\prime} Z_i + v_i, E[Z_i v_i] = 0$. Then, $\Sigma_{12} = E[Z_i Z_i^{\prime} X_i \varepsilon_i ] - \sigma_{\varepsilon Z} E[X_i Z_i^{\prime} ] = E[\sigma_{\varepsilon X} Z_i Z_i^{\prime}] - \sigma_{\varepsilon Z} E[X_i Z_i^{\prime} ] $. Under correct specification ($\sigma_{\varepsilon Z} = 0)$, $\Sigma_{12} =E[\sigma_{\varepsilon X} Z_i Z_i^{\prime}]$ is generally nonzero whenever there is endogeneity, $\sigma_{\varepsilon X} \neq 0$. Under misspecification ($\sigma_{\varepsilon Z} \neq 0)$, $\Sigma_{12}$ is generally nonzero unless there exists specific structure of the higher moments of $Z_i$ so that $E[\sigma_{\varepsilon X} (Z_i) Z_i Z_i^{\prime}] =\sigma_{\varepsilon Z}\pi^{\prime} E[Z_i Z_i^{\prime}]$.} Intuitively, variations of the Jacobian affect the distribution of the GMM estimator under misspecification, and thus improved efficiency can be achieved by imposing efficient weighting of the augmented moments accounting for the correlations between $g(X_i, \theta)$ and $G(X_i, \theta)$. \\

Next Corollary provides some further insights on the optimal choice of $\Lambda^{*}_{W}$ and the asymptotic variance of the ME estimator in the linear model. This will facilitate the discussions on the efficient weighting matrix under misspecification, compared with the“conventional" optimal weighting matrix that was defined under correct specification. When the model is linear, i.e., $G(X_i) = \frac{\partial g(X_i, \theta)}{\partial\theta^{\prime}}$ does not depend on $\theta$, and $F(X_i, \theta) = \frac{\partial vec [G(X_i, \theta)^{\prime}]}{\partial\theta^{\prime}}= 0$, thus we can further simplify the form of $\Lambda^{*}_{W}$ and the asymptotic variance of ME estimator.\\

corollaryIf $g(X, \theta)$ is linear in $\theta$, we have \begin{equation} \Lambda^{*}_{W} = [G^{\prime} \Sigma_{11,2}^{-1} \quad - G^{\prime} \Sigma_{11, 2}^{-1} \Sigma_{12} \Sigma_{22}^{-1}] \end{equation} where $G = E[ G(X_i)], \Sigma_{11, 2} = (\Sigma_{11} - \Sigma_{12} \Sigma_{22}^{-1} \Sigma_{21})$, $\Sigma (\theta_W) =\begin{bmatrix} \Sigma_{11} & \Sigma_{12} \\ \Sigma_{21}&\Sigma_{22}\end{bmatrix} $. The asymptotic variance of the ME estimator is \begin{equation} V_{ME} (W)= (G^{\prime} \Sigma_{11, 2}^{-1}G)^{-1},\\ \end{equation} and $\Sigma_{11,2}$ is invariant to $\theta_W$.

Under the linear model, Corollary (ref) shows that the asymptotic variance of the ME estimator is $(G^{\prime} \Sigma_{11, 2}^{-1}G)^{-1}$, where $\Sigma_{11, 2} = \Sigma_{11} - \Sigma_{12} \Sigma_{22}^{-1} \Sigma_{21} $. We refer $W^* = \Sigma_{11,2}^{-1}$ as the misspecification-efficient weighting matrix, in analogy with the conventional (correct-specification) efficient weighting matrix $W_0 = \Sigma_{11}^{-1}$. While, $W_0$ is the inverse of the variance of the original moments $g(X_i,\theta)$, $W^* = \Sigma_{11,2}^{-1}(\theta)$ is the inverse of the variance of $g(X_i, \theta) - \Sigma_{12}\Sigma_{22}^{-1} vec(G(X_i)')$, which is the residualized moment after the projection onto the Jacobian. Even under correct specification, the ME bound $(G^{\prime} \Sigma_{11, 2}^{-1}G)^{-1}$ is strictly smaller than the GMM estimator with the conventional efficient weighting matrix $W_0$ when $\Sigma_{12} \neq 0$ as we discussed.

Corollary (ref) provides new insight into the conventional (correct-specification) efficient GMM variance formula. Corollary (ref) shows that the conventional efficient GMM variance formula $(G^{\prime} WG)^{-1}$ achieves the efficiency bounds of the oracle-ME estimator, when $W^* = \Sigma_{11,2}^{-1}(\theta_W)$. The caveat is that the GMM estimator itself with the choice $W^*$ does not achieve the oracle ME bound as shown in Corollary (ref), $V_{GMM}(W^*, \theta_W) \ge (G'\Sigma_{11,2}^{-1}G)^{-1} $ under misspecification. It follows that conventional formula for efficient GMM is too “small", even in large samples, for the standard GMM estimator, as it will be valid for the ME estimator. Researchers have found that the standard errors of the efficient GMM are often severely downward biased.\footnote{Many researchers thus routinely used the finite-sample correction of Windmeijer (2005). Hwang et al. (2022) further shows, under the linear model, that misspecification-robust GMM variance will not only be valid under misspecification, but also work as a finite-sample correction even if the model is correctly specified.} Our results show that reporting the conventional efficient GMM variance formula in the linear GMM will still be useful as an efficiency frontier - showing researchers what could be gained by imposing moments from the Jacobian $G(X_i, \theta)$; but not to be used for an inference directly with the efficient GMM itself under misspecification. When there is no efficiency gain $\Sigma_{12}=0$, $W^* = W_0 = \Sigma_{11}^{-1}$, using the misspecification-efficient weighting matrix reduces to the conventional case.

We can easily report the oracle ME bound, $V_{ME} (W)=(G^{\prime} \Sigma_{11, 2}^{-1}G)^{-1}$ as an efficiency frontier alongside conventional/misspecification-robust standard errors with no significant cost. We can consistently estimate the $V_{ME} (W)= (G^{\prime} \Sigma_{11, 2}^{-1}G)^{-1}$ regardless of the model specification based on the sample estimator \[ \widehat\Sigma_{11,2} = \frac{1}{n}\sum_i (r_i(\widehat{\theta}) - \bar r_n(\widehat{\theta}))(r_i(\widehat{\theta}) - \bar r_n(\widehat{\theta}))^{\prime} \] with $r_i(\theta) = g_i(\theta) - \widehat\Sigma_{12}\widehat\Sigma_{22}^{-1}vec(G_i(\theta)')$ and the preliminary GMM estimator $\widehat{\theta} = \widehat{\theta}_{GMM} (W)$.\footnote{Under misspecification, we must use this centered covariance estimator because the uncentered covariance estimator $ \frac{1}{n}\sum_i r_i(\widehat{\theta}) r_i(\widehat{\theta})^{\prime}$ is not consistent for $\Sigma_{11,2}$.}

Furthermore, Corollary (ref) shows that under the linear model, $V_{ME} = (G'\Sigma_{11, 2}^{-1}G)^{-1}$ is same for all different pseudo-true values $\theta_W$, due to the invariance property of $\Sigma_{11, 2}$. This implies that the misspecification affects what is being estimated, but not how precisely it can be estimated. This is because the linear GMM model has a “location-like" structure in the sense that $g(X_i,\theta) = g(X_i, 0) + G(X_i)\theta$, so shifting $\theta$ moves the moment vector along the Jacobian direction. The invariance property holds because $\Sigma_{11, 2}$ is the variance of the residual from the population regression of $g(X_i,\theta)$ on $vec(G(X_i)')$, so all the $\theta$ dependent parts are projected out. Practically, this implies that $V_{ME} $ in the linear model can be consistently estimated from any preliminary GMM without committing to a particular weight matrix or pseudo-true value.

In general nonlinear settings, above arguments fail to hold as the Jacobian $G(X_i, \theta)$, and $\Sigma_{11,2}(\theta)$ changes with $\theta$. However, the distribution of the oracle-ME estimator $\widehat{\theta}_{ME} (W) $ and the oracle bounds $V_{ME}(W)$ can be approximated by the bootstrap methods (will be discussed in the next Section (ref)). This bootstrap method can be viewed as a misspecification-robust and efficient version of the Hall and Horowitz (1996) recentering method.

The joint covariance matrix of the moment and Jacobian vectors, $\Sigma (\theta) =E [ \psi_{i}(X_i, \theta) \psi_{i}(X_i, \theta)^{\prime}] -E [ \psi_{i}(X_i, \theta)] E [ \psi_{i}(X_i, \theta)]^{\prime}$ plays a central role in this paper, and the similar construction of the orthogonalized moments has been used in the weak-identification literature, though from a different perspective than ours. Kleibergen (2005) constructs a recentered Jacobian that is asymptotically uncorrelated with the sample moments, by projecting the Jacobian on the moment vector, and it was used to construct weak-identification robust test-statistics. Kleibergen and Zhan (2025) extend this construction to accommodate misspecification, proposing a double robust Lagrange multiplier test for the pseudo-true value of the continuously updating estimator (CUE), with $\Sigma_{11}(\theta)^{-1}$, which is the efficient choice under correct specification. \\

Example: Linear IV. Consider a linear instrumental variable regression $Y_i = X_i^{\prime} \theta + \varepsilon_i $ where $Y_i$ is an outcome variable, $X_i$ is a potentially endogenous regressor, and $Z_i$ is a $m\times 1$ vector of instrumental variables. Let $Y, X, Z$ are $n \times 1, n\times p$, and $n\times m$ data matrices with row $i$ equal to $Y_i^{\prime}, X_i^{\prime}, Z_i^{\prime}$, respectively. The moment function is $ g(\cdot, \theta) = Z_i (Y_i - X_i^{\prime} \theta)$, and we consider moment misspecification $E[Z_i (Y_i - X_i^{\prime} \theta)] \neq 0 $. With $m\times m$ matrix $W>0$, the standard GMM estimator is \[ \widehat{\theta}_{GMM}(W) = (X'Z WZ'X)^{-1} X'Z W Z'Y \] with the GMM pseudo-true value, \[ \theta_W = (E[X'Z] W E[Z'X])^{-1} (E[X'Z] W E[Z'Y]). \] When $W = E[Z'Z]^{-1}$, $\theta_W$ is the TSLS estimand. Based on (ref) and using the linearity of the moment condition, and the optimal $\Lambda^{*}_{W}$ in Corollary (ref), ME has a closed-form as follows

multline[multline omitted — 311 chars of source]

where $\Sigma_{ 11, 2} = (\Sigma_{11} - \Sigma_{12} \Sigma_{22}^{-1} \Sigma_{ 21}), \Sigma (\theta_W) = \big[

smallmatrix\Sigma_{11} & \Sigma_{12} \\ \Sigma_{ 21}&\Sigma_{22}

\big]$, and $\Sigma (\theta)$, i.e., $E [ \big(

smallmatrixZ_i Y_i-Z_i X_i'\theta) \\ vec(X_iZ_i^{\prime})

\big) \big(

smallmatrixZ_i Y_i-Z_i X_i'\theta) \\ vec(X_iZ_i^{\prime})

\big)^{\prime}] -E \big(

smallmatrixZ_i Y_i-Z_i X_i'\theta) \\ vec(X_iZ_i^{\prime})

\big) E \big(

smallmatrixZ_i Y_i-Z_i X_i'\theta) \\ vec(X_iZ_i^{\prime})

\big)^{\prime} $. Sample analogs of $ \Sigma_n (\widehat{\theta}_W)^{-1}$ can also be used here. In the linear IV example, the ME estimator, (ref) can be considered as efficient linear combinations of the recentered GMM/TSLS estimator and the Jacobian moment conditions.

Assuming that the data-generating process is $Y_i = \theta_0 X_i + Z_{2i} +\varepsilon_i, X_i = Z_{1i} + 2Z_{2i}+ v_i$, $Z_i = (Z_{1i}, Z_{2i})^{\prime} \overset{i.i.d.}{\sim} N(0, I_2), (\varepsilon_i, v_i)~N(0, [1, 0.5; 0.5, 1])$ with $\theta_0 = 0$. When we consider the weight matrix $W =(

smallmatrix1 &\rho\\ \rho &1

)$ for $\rho \in (-1, 1)$, the GMM pseudo-true value is $\theta_W(\rho) = \frac{\rho + 2}{5 + 4\rho}$, which is a strictly decreasing function of $\rho \in (-1, 1)$, and there is one-to-one mapping from $\rho$ to $\theta_W(\rho)$. By using the formula above, the asymptotic variance of the GMM estimator as a function of $\rho$ is: \[ V_{GMM}(\rho) = \frac{2(56\rho^4 + 131\rho^3 + 210\rho^2 + 232\rho + 100)}{(5 + 4\rho)^4}. \] The variance of the ME estimator can be calculated as $V_{ME} = \frac{969}{6116} \approx 0.1584$ (do not depend on $\rho$), and it can be verified that $V_{GMM}(\rho) > V_{ME}$ for all $\rho$. At $\rho = 0$ (TSLS), the variance simplifies to $V_{GMM}(0) = 0.32$, and the efficiency gain is over 50\% against TSLS. The efficiency gain ($1- \frac{V_{ME}}{V_{GMM}(\rho)}$) grows as $\rho$ decreases; for example, it is 72\% and 88\% for $\rho = -0.5$ and $-0.9$, respectively.\footnote{Although we do not report the results here for brevity, we also considered a general nonlinear instrumental variables (NLIV) model, where the moment function is $g(X_i, \theta) = Z_i (Y_i - exp(X_i'\theta))$. This model has been considered in the count data models with endogenous regressor (e.g., Windmeijer and Santos Silva (1997)). Using numerical methods, we verified that $V_{GMM} > V_{ME}$, and $V_{ME}$ varies with $W$ in the nonlinear model. }

remark[$s$-step and the iterated GMM estimators under misspecification] In this paper, we only consider the one-step GMM estimator in the usual sense with the fixed $W$. It is well-known that the GMM pseudo-true value changes with each step of the iteration under misspecification, and the asymptotic distribution of the conventional $s$-step GMM depends on that of the previous-step estimator at which the efficient weight matrix is evaluated, making the asymptotic analysis complicated (see, Hall and Inoue (2003), Hwang, Kang and Lee (2022), and Hansen and Lee (2021) for an asymptotic distribution of two-step and the iterated GMM estimator under misspecification). However, using the “standard" efficient weighting-matrix $W_0$ is ambiguous under misspecification and it is not clear why we want to stick with this choice in each iteration, which we show in this paper is suboptimal. The choice of $W^* = \Sigma_{11,2}^{-1}(\theta)$ deserves more attention, at least in the linear GMM setup. The FOC condition for the efficient GMM estimator based on $W^* = \Sigma_{11,2}^{-1}(\theta_W)$ with the one-step GMM and the initial weighting matrix $W$ is $G'\Sigma_{11,2}^{-1}(\theta_W) E[g(X_i, \theta_{W^*})] = 0$. In general, the FOC condition $G'\Sigma_{11,2}^{-1}(\theta_0) E[g(X_i, \theta_0)] =0 $ holds for the iterated GMM pseudo-true values $\theta_0$ with the $W = \Sigma_{11,2}^{-1}(\theta)$.

Misspecification Efficient Bootstrap GMM Estimators

We first consider a related class of estimators, misspecification-efficient (ME) GMM based on the augmented moment conditions. We then propose the bootstrap GMM estimator, which can be considered as a misspecification-robust and efficient version of Hall and Horowitz (1996) recentered bootstrap GMM estimator.

The ME-GMM estimator, $\widehat{\theta}_{ME-GMM} (\Delta)$, with the fixed weighting matrix $(mp+m) \times (mp+m)$ matrix $\Delta >0$ solves

equation[equation omitted — 254 chars of source]

where,

equation*[equation* omitted — 188 chars of source]

We can easily show that the ME-GMM with the efficient weighting matrix $\Delta = E[\widetilde{\psi}(X_i, \theta_W)\\ \widetilde{\psi}(X_i, \theta_W)^{\prime}]^{-1} = \Sigma(\theta_W)^{-1}$ is asymptotically equivalent to the ME estimator, $\widehat{\theta}_{ME} (W)$. The consistency and the asymptotic normality of $\widehat{\theta}_{ME-GMM} (W)$ immediately follow by the same Assumptions (Assumption (ref).2-(ref).4).

We first note that Lee and Liao (2018) consider the same GMM problems under correct specification $(E[g(X_i, \theta_W)] = 0)$ and singular Jacobian $(E[G(X_i, \theta_{W})] =0) $. However, the ME estimator, $\widehat{\theta}_{ME} (W)$, and the ME-GMM estimator in (ref) are not fully feasible without estimation of the population moments $E[g(X_i, \theta_{W})]$ and $E[vec (G(X_i, \theta_{W})^{\prime})]$, which was used for recentering. If we use the sample counterparts of the re-centered moments, e.g., replacing the sample analogs of $E[g(X_i, \theta_W)]$ and $E[vec(Z'X)]$, then the estimator degenerates to the original estimator $\widehat{\theta}_W$, so it does not correctly account for variations in the limiting distributions. We consider the sample-split estimators in Section (ref) as potential alternatives.

Here, we consider the bootstrap GMM estimator which mimics (ref) with the efficient weighting matrix $\Delta = \Sigma(\theta_W)^{-1}$. Let $\{X_i^*\}_{i=1}^n$ as the bootstrap samples, independently sampled from the original sample $\{X_i\}_{i=1}^n$ with replacement. With re-centering, the bootstrap ME-GMM estimator solves

equation[equation omitted — 278 chars of source]

where,

equation*[equation* omitted — 210 chars of source]

, \[ g_n(\theta) = n^{-1}\sum_{i=1}^{n} g (X_i, \theta), \quad G_n(\theta) = n^{-1}\sum_{i=1}^{n} G (X_i, \theta), \]

equation[equation omitted — 231 chars of source]

The recentering by $g_n(\widehat{\theta}_W), vec (G_n(\widehat{\theta}_W))$ ensures $E^*[g(X_i^*,\widehat{\theta}_W) - g_n(\widehat{\theta}_W)] = 0$, and the bootstrap moments $\widetilde{\psi}^{*}(X_i^{*}, \theta)$ has mean zero at $\widehat{\theta}_W$ conditional on the data. $\Delta^{*}$ is symmetric positive definite weighting matrix that is a bootstrap sample version of the efficient weighting matrix $\Delta = \Sigma(\theta_W)^{-1}$, where $\Sigma (\theta) =E [ \psi_{i}(X_i, \theta) \psi_{i}(X_i, \theta)^{\prime}] -E [ \psi_{i}(X_i, \theta)] E [ \psi_{i}(X_i, \theta)]^{\prime}$.

The proposed bootstrap estimator utilizes the re-centering of the Jacobian estimator to ensure that the bootstrap expected Jacobian matrix evaluated at the estimator is zero, which is essential to correctly mimic the asymptotic distribution under the misspecification. This “double re-centering" idea has been also considered in weak-identification literature to accurately reflect the behavior of the original Jacobian (e.g., Dovonon and Goncalves (2017)), yet under correct specification.

The standard nonparametric bootstrap without recentering provides first-order valid inference under correct specification (Hahn (1996)), however, the standard bootstrap does not achieve an asymptotic refinement in overidentified models. Again, this is because the bootstrap sample mean of the moment function evaluated at the estimator, $E^{*}[g(X_i^*, \theta_0)] = n^{-1} Z^{\prime} \widehat{e}$, is not necessarily equal to zero under correct specification. To achieve refinement, Hall and Horowitz (1996) proposed a recentered bootstrap GMM estimator;

equation[equation omitted — 256 chars of source]

where $W_0^*$ is the bootstrap version of efficient weighting matrix under correct specification, $W_0 = \Sigma_{11} (\theta)^{-1} $, and $\widehat{\theta}_W = \widehat{\theta}_{GMM}(W)$ is the GMM estimator. $\widehat{\theta}_{HH}^{*} $ can be considered as a special case of the proposed bootstrap ME-GMM estimator $\widehat{\theta}_{ME}^{*}$ by ignoring the Jacobian augmented parts and only recentering the original sample moments, i.e., $\Delta^{*}_{11} =W_0^*, \Delta_{12}^{*} = 0$, and $\Delta_{22}^{*} = 0$.

The conventional GMM and the HH bootstrap estimator would have different asymptotic distributions under misspecification because Hall-Horowitz recentered bootstrap always creates a correctly-specified bootstrap world, regardless of whether the actual model is misspecified or not. Hall-Horowitz bootstrap eliminates the additional source of Jacobian sampling variability that arises under misspecification, while the proposed ME bootstrap can consider these variations, and furthermore it considers more efficient combinations of the sampling uncertainty for original and Jacobian moments. \\

Example: Linear IV (continued). Based on the FOC from (ref) and using the linearity of the moment condition, ME-GMM has a closed-form

multline[multline omitted — 256 chars of source]

where $\Delta =

bmatrix[bmatrix omitted — 70 chars of source]

$. When $\Delta_{11} = (Z'Z)^{-1} $ and $\Delta_{12} = 0$, ME-GMM estimator is related to a recentered TSLS estimator,

equation[equation omitted — 179 chars of source]

The choice of $\Delta_{11} = (Z'Z)^{-1}, \Delta_{12} = 0$ for ME-GMM is not efficient under misspecification, although its choice was motivated by the efficiency of GMM under correct specification and homoskedasticity.

With the bootstrap samples $ \{ Y_i^{*}, X_i^{*}, Z_i^{*} \}_{i=1}^{n} $, independently sampled from the original sample $\{ Y_i, X_i, Z_i\}_{i=1}^{n} $, Hall and Horowitz (1996) version of the recentered bootstrap TSLS estimator is \[ \widehat{\theta}_{HH}^{*} = (X^{* \prime}Z^{*} (Z^{* \prime}Z^{*})^{-1} Z^{* \prime}X^{*})^{-1} X^{* \prime}Z^{* } (Z^{* \prime}Z^{*})^{-1} (Z^{* \prime}Y^{* } - Z^{\prime} \widehat{e}) \] where $\widehat{e} = Y - X^{\prime} \widehat{\theta}_{TSLS}$. This is essentially related to a version of recentered bootstrap estimator of (ref). Our proposed bootstrap GMM estimator in this setup is

multline[multline omitted — 422 chars of source]

where $\widehat{e} = Y - X^{\prime} \widehat{\theta}_{W}$, $\widehat{\Sigma}_{11, 2} = (\widehat{\Sigma}_{11} - \widehat{\Sigma}_{12} \widehat{\Sigma}_{22}^{-1} \widehat{\Sigma}_{21}), \widehat{\Sigma} (\widehat{\theta}_{W}) =

bmatrix[bmatrix omitted — 106 chars of source]

$, $\widehat{\Sigma} (\theta) = \frac{1}{n} \sum_{i=1}^{n} \psi(X_i, \theta) \psi(X_i, \theta)^{\prime} - \sum_{i=1}^{n} \psi(X_i, \theta) \sum_{i=1}^{n} \psi(X_i, \theta)^{\prime}$, and $\psi(X_i, \theta) = ((Z_i Y_i-Z_i X_i'\theta)', vec(X_iZ_i^{\prime})')'$. In the linear IV setup, $\widehat{\theta}_{ME}^{*}$ is not just a linear combination of the $\widehat{\theta}_{HH}^{*}$ and the recentered Jacobian moments, but an efficient linear combination with $\widehat{\Sigma}_{11, 2}$.\\

We show below that the asymptotic distribution of the bootstrap estimator $\widehat{\theta}_{ME}^{*}(W)$ is identical to that of the original ME estimator $\widehat{\theta}_{ME}(W) $. We consider the following assumptions for the bootstrap validity. Let bootstrap sample mean of moments and Jacobian as $ g_n^*(\theta) = n^{-1}\sum_{i=1}^{n} g (X_i^*, \theta), G_n^*(\theta) = n^{-1}\sum_{i=1}^{n} G (X_i^*, \theta)$.

assumption\begin{enumerate} • $\widehat{\theta}_{ME}^{*} (W) \overset{p^{*}}{\rightarrow} \theta_W.$$\sup_{\theta \in \Theta} \| g_n^*(\theta) - g_n(\theta)\| \overset{p^{*}}{\rightarrow} 0$, and $\sup_{\theta \in \Theta} \| G_n^*(\theta) - G_n(\theta)\| \overset{p^{*}}{\rightarrow} 0$. • $\sqrt{n}\big(g_n^*(\widehat{\theta}_W) - g_n(\widehat{\theta}_W),\; \mathrm{vec}(G_n^*(\widehat{\theta}_W) - G_n(\widehat{\theta}_W)\big)' \overset{d^{*}}{\rightarrow} N(0,\Sigma(\theta_W))$, where $ \widehat{\theta}_W = \widehat{\theta}_{GMM}(W)$, and $\Sigma(\theta_W)$ is the same covariance matrix as in Theorem (ref). \end{enumerate}

These conditions are standard assumptions under i.i.d. sampling, e.g., Gin\'{e} and Zinn (1990), and Hahn (1996). Assumption (ref).1 on the bootstrap consistency is a high-level assumptions, which can be verified from more primitive conditions using the standard bootstrap consistency of the extremum estimators, combined with the uniform convergence conditions in Assumption (ref).2. Assumption (ref).3 is implied by Assumptions (ref).1-(ref).4 under i.i.d. setup. \\

theoremSuppose that Assumption (ref).1-(ref).4 hold. Further, we assume that $\theta = \theta_W$ is a unique solution to $E[\widetilde{\psi}_i(X_i, \theta)] = 0 $, and $\Delta$ is positive definite. Then, we have \begin{align*} \sqrt{n}(\widehat{\theta}_{ME-GMM} (\Delta)- \theta_{W} ) \overset{d}{\longrightarrow} N(0,(\Gamma(\theta_W)^{\prime} \Delta \Gamma (\theta_W))^{-1} \Gamma(\theta_W)^{\prime} \Delta \Sigma (\theta_{W} ) \Delta \Gamma(\theta_W)( \Gamma(\theta_W)^{\prime} \Delta \Gamma (\theta_W))^{-1}). \end{align*} When $\Delta = \Sigma(\theta_W)^{-1}$, the asymptotic variance reduces to $(\Gamma(\theta_W)^{\prime} \Sigma (\theta_W)^{-1}\Gamma(\theta_W))^{-1}$, provided it is non-singular. In addition, suppose that Assumption (ref) holds, then \begin{align*} \sqrt{n}(\widehat{\theta}_{ME}^*(W) - \widehat{\theta}_{GMM} (W) ) \overset{d^*}{\longrightarrow} N(0,(\Gamma(\theta_W)^{\prime} \Sigma (\theta_W)^{-1}\Gamma(\theta_W))^{-1}).\\ \end{align*}

Theorem (ref) shows that the asymptotic distribution of the bootstrap estimator $\sqrt{n}(\widehat{\theta}_{ME}^{*}(W) - \widehat{\theta}_{GMM}(W))$ is equal to those of $\sqrt{n}(\widehat{\theta}_{ME} (W)- \theta_{W} )$. Theorem (ref) allows us to approximate the sampling variation of the oracle-ME estimator $\widehat{\theta}_{ME} (W)$, and the theorem justifies the use of the bootstrap to obtain the minimax bounds that will be established in Section (ref) (Theorem (ref)). Although it is necessary to have stronger conditions such as uniformly square integrable condition to guarantee convergence in moments from the convergence in distribution results, we can still use the trimmed bootstrap estimator of variance as the consistent estimator of the asymptotic variance $(\Gamma(\theta_W)^{\prime} \Sigma (\theta_W)^{-1}\Gamma(\theta_W))^{-1}$.

Double-Recentered Hall and Horowitz (1996) Bootstrap Estimator

We note, however, that Theorem (ref) does not guarantee the validity of the percentile bootstrap CI based on $\widehat{\theta}_{ME}^{*}(W) $ because it is centered at $\widehat{\theta}_{GMM}(W)$, which has a different asymptotic distribution. Using the double-recentering idea from the previous section to approximate the sampling variation of the oracle ME estimator, we propose a double-recentered (DR) that jointly perturbs the moment conditions and the Jacobian. The DR bootstrap correctly approximates the asymptotic distribution of the GMM estimator regardless of specification status, nesting the Hall and Horowitz (1996) bootstrap as a special case under correct specification.

Lee (2014) argues that recentering in the standard nonparametric GMM bootstrap can be detrimental and is not even needed if we use the analytic misspecification-robust variance estimator. In our paper, however, we clarify this argument that additional considerations of re-centered Jacobian can achieve asymptotic validity under misspecification. This is also the case for the standard GMM bootstrap estimator without recentering, while Hall and Horowitz (1996) do not achieve asymptotic validity because it only recenters the original moment functions.\footnote{Since, the main focus of the paper is on the efficiency results for $\theta_W$ and minimax bounds, we only considered the GMM estimator with fixed $W$. Thus, we do not need to consider variations from the weight matrix, which was typically considered for the construction of the Hessian and misspecification-robust variance constructions in the literature. Although it is beyond the scope of this paper, the idea of the “double-recentering" in this paper can be easily extended to the “triple-recentering" with the estimated $W_n$ (one-step) or the $W_n(\theta)$ (two-step) case by augmenting the weighting matrix into the $\psi(\cdot)$. }

We also note that DR bootstrap estimator is different from the bootstrap M-estimator based on the just-identified moment functions with an augmented parameter as in Imbens (1997). One can show that the bootstrap M-estimator based on the just-identified augmented moment functions as in Imbens (1997) corresponds to the standard GMM bootstrap estimator without recentering.

We first observe that the GMM-estimator satisfies the estimating equations (ref) for the augmented moment conditions with the particular choice of $\Lambda = A(W, \theta_W) = [G(\theta_W)^{\prime} W \quad g(\theta_W)^{\prime} W \otimes I_p]$. We then define the misspecification-robust recentered bootstrap estimator $\widehat{\theta}_{DR}^{*}$ as the solution to the bootstrap analog of the augmented estimating equation:

equation[equation omitted — 113 chars of source]

where

equation*[equation* omitted — 262 chars of source]

and \[ A_n(W, \widehat{\theta}_{W})= \big[G_n(\widehat{\theta}_{W})'W_n\;,\;\; g_n(\widehat{\theta}_{W})'W \otimes I_p\big] \] is the sample analog of the matrix $A(W, \theta_W)= [G(\theta_W)'W,\; g(\theta_W)'W \otimes I_p]$. While the HH bootstrap only recenters the original moments $g_n^*$ at $g_n(\widehat{\theta}_{W})$, the DR bootstrap recenters the full augmented vector $\psi_n$ at its sample counterpart $\psi_n^*$, including Jacobian estimation variability, appropriately weighted by $\Lambda_n = A_n(W, \widehat{\theta}_{W})$ to mimic the distribution of $\widehat{\theta}_{W}$.

For the implementation, we first, compute the original sample GMM estimator $\widehat{\theta}_{W}$, and calculate $A_n(W, \widehat{\theta}_W)$, which requires $g_n(\widehat{\theta}_{W})$ and $G_n(\widehat{\theta}_{W})$. Then, for each bootstrap replication $b = 1, \ldots, B$:

enumerate• draw $\{X_i^*\}_{i=1}^n$ with replacement; • with $A_n(W, \widehat{\theta}_W)$, solve ((ref)) for $\widehat{\theta}_{DR}^{*}$ via Newton--Raphson initialized at $\widehat{\theta}_{W}$; for the linear model, $\widehat{\theta}_{DR}^{*}$ has a closed-form solution (see below linear IV example); • Confidence intervals are constructed as standard percentile or percentile-$t$ intervals based on the $t^* = \frac{\widehat{\theta}_{DR}^{*} -\widehat{\theta}_{W}}{s.e.(\widehat{\theta}_{DR}^{*} )}$ from the $B$ bootstrap draws.

Example: Linear IV (continued). Since the sample and population Jacobian does not depend on $\theta$, $\widehat{\theta}_{DR}^{*}$ admits a closed-form solution by direct algebra from (ref):

eqnarray[eqnarray omitted — 279 chars of source]

So, DR bootstrap estimator $\widehat{\theta}_{DR}^{*}$ is approximately equal to $\widehat{\theta}_{HH}^{*} + \tilde{\Lambda}_2 \cdot \mathrm{vec}(Z^{*\prime}X^* - Z'X)$ with some weights $ \tilde{\Lambda}_2 $. The last term is a correction term for the Jacobian variability. Under correct specification, $g_n(\widehat{\theta}_{DR}) = n^{-1}Z'\hat{e} = O_p(n^{-1/2})$, so the last terms are finite-sample corrections.\\

corollarySuppose that Assumption (ref).1-(ref).5, and Assumption (ref) hold. Then, \begin{align*} \sqrt{n}(\widehat{\theta}_{DR}^* - \widehat{\theta}_{GMM} (W) ) \overset{d^*}{\longrightarrow} N(0,V_{GMM}(W, \theta_W)), \end{align*} where $V_{GMM}(W, \theta) $ is defined in (ref).

We recommend to consider $\widehat{\theta}_{DR}^* $ and the percentile confidence interval for the valid inference for $\theta_W$. In addition, separately report the $\widehat{V}_{ME}(W) $ and $\sup_{W \in \mathcal{W}} \widehat{V}_{ME}(W)$ based on the general double bootstrap ME estimator $\widehat{\theta}_{ME}^*$ or using analytical ME estimates with the standard GMM point estimate $\widehat{\theta}_{GMM}(W)$ as evidence of the attainable efficiency frontier by showing researchers what could be gained if a feasible ME point estimate were available with known misspecification and identification. This provides honest bounds that acknowledge the researcher doesn't know which $W$(and hence $\theta_W$) is “right."

Split-Sample ME Estimator

The bootstrap provides a consistent estimate of the misspecification-efficient variance, but the oracle ME point estimate that actually achieves that variance is computable only if “degree of misspecification" $(\gamma_1(W) = E[g(X_i, \theta_W)])$ and the “identification strength" $(\gamma_2(W) = E[G(X_i, \theta_W)])$ parameters are known, that was used for recentering in the ME estimator. While these components can be consistently estimated from the full sample, the population moments and Jacobian cannot be replaced by its sample analog, which suffers from a degeneracy in our setup by collapsing ME estimator back to the standard GMM estimator that invalidates the ME variance reduction.\footnote{In Appendix B, we consider the valid inference methods for $\theta_W$ without assuming $\gamma (W) = (E[g(X_i, \theta_W)]', vec(E[G(X_i,\theta_W)^{\prime}]))'$ is known using the worst-case bounds approach similar to Conley et al. (2012).}

We consider a sample-splitting ME estimator which uses one-half of a sample to estimate $\gamma(W) = (\gamma_1(W)', vec(\gamma_2(W))')'$. Then estimated parameters $\widehat{\gamma}(W)$ are then used to construct the final ME estimates using the other half of the samples. This construction parallels the split-sample IV estimator of Angrist and Krueger (1995), who uses the first-half of the sample to estimate the population-first stage $E[Z'Z]^{-1}E[Z'X]$ and use them in the IV estimates in the other half sample.

Specifically, we consider the following step to construct the sample-split estimator for $\widehat{\theta}_{ME} (W)$;

enumerate• Randomly divide the samples into $K=2$ independent subsample roughly equal size $n/2$, and let $\{Y_1,X_1\}, \{Y_2, X_2\}$ as data matrices for each subsample. • Use all observations with $\{ Y_2,X_2\}$ to estimate the $\gamma(W) =(\gamma_1(W)', vec(\gamma_2(W))')'$ with the GMM estimator $\widehat{\theta}_{W} $ and let these leave-fold-out estimators as $\widehat{\gamma}_{-1} (W)$. • Solve the single-split ME estimator $\widehat{\theta}^{SS}(W)$ using observations $\{X_1,Y_1\}$ with the $\widehat{\gamma}_{-1}(W)$, and calculate the standard error using “efficient" misspecification-robust variance formula $\widehat{V}_{ME}(W)$. • Repeat the above procedure, $s = 1, ..., S$, and set $\widehat{\theta}^{RSS}(W) = \frac{1}{S} \sum_{s=1}^{S} \widehat{\theta}_{s}^{SS}(W)$, $SD (\widehat{\theta}^{RSS}(W)) = \frac{1}{S-1} \sum_{s=1}^{S} (\widehat{\theta}_{s}^{SS}(W) - \widehat{\theta}^{RSS}(W))^2$. We also consider the variance estimators considered in Chernozhukov et al. (2018) based on median, i.e., $ median \{ \widehat{V}_{ME,s}(W) + ((\widehat{\theta}_{s}^{SS}(W) - \widehat{\theta}_{median}^{SS}(W))(\widehat{\theta}_{s}^{SS}(W) - \widehat{\theta}_{median}^{SS}(W)))' \}_{s=1}^{S}$. • In the linear model, when $W^* = \Sigma_{11,2}^{-1}$, then estimation of $\gamma_1(W)$ is not needed as the recentering part will be eliminated by the FOC.

Since $n_1 = n/2$ in the split-sample ME estimator, the asymptotic variance is at least twice the asymptotic variance of the oracle-ME. Repeated single sample-splitting procedure can be done to reduce the variance from the single sample. Furthermore, when we construct ME estimator based on $W = \Sigma_{11,2}^{-1}$, we don't need recentering part $\gamma_1(W)$ due to the FOC with $W^* = \Sigma_{11,2}^{-1}$. This helps reduce computational costs and estimation noises of the sample-splitting estimate. We do not consider the general symmetric sample-splitting estimator (i.e., swap the role of each subsample and then average) or cross-fitting estimator because the symmetry of averaging again has similar degeneracy problems.

Semiparametric Efficiency Under Misspecification

Based on the results from the previous sections, we derive the uniform (over $W\in \mathcal{W})$ asymptotic minimax bounds of the misspecified moment condition models. We can generally view misspecified-GMM model as a semiparametric model such that we consider a family of probability distributions of the observed data $X$, indexed by the finite dimensional parameter of interest $\theta$ for some weighting matrix $W$. In practice, researchers can choose different weighting matrix $W \in \mathcal{W}$ based on the application they consider, where $\mathcal{W}$ denotes a class of $m \times m$ positive definite matrices.

However, the notion of semiparametric efficiency and optimality need to be defined differently under (global) misspecification. As noted earlier, the fundamental difficulty here is that the pseudo-true value $\theta_W$ changes with $W \in \mathcal{W}$ in our misspecified-moment condition models. The semiparametric efficiency framework typically assumes the existence of a well-defined “true" parameter, and the efficiency bound is computed at this target parameter. For example, GMM estimator is semiparametrically efficient under correct specification (Chamberlain (1987)) for $\theta = \theta_0$ under the moment conditions $E[g (X_i, \theta)] = 0$ with an “efficient" weighting matrix $W = \Sigma_{11} (\theta_0)^{-1}$, the inverse of the variance matrix of the original moments.

Our uniform minimax bounds will be not only useful for a researcher who chooses $W$ based on target a specific economic parameter $\theta_W$, but also useful for a researcher who is completely agnostic to the choice of $W \in \mathcal{W}$. Many applied researchers may just pick one $W$ (e.g., TSLS); however, an applied researcher who runs GMM already considers choosing $W$ over a class of different weighting matrices $\mathcal{W}$ routinely; one-step $W = I$ or $(Z'Z)^{-1}$, two-step, and iterated until convergence. Under misspecification each iteration changes the pseudo-true value. So when results differ across one-step, two-step, and iterated GMM, that is direct evidence of sensitivity to $W$.

We may compare two or more different weighting matrices, and may want to choose a “better" $W$ and pseudo-true $\theta_W$ than the others, based on the pseudo-distance measure such as $J$-statistics, or based on the goodness of fit measure or the information criteria (e.g., Rivers and Vuong (2002), Marmer and Otsu (2012)). We can also consider the optimal weights $W^{*} = \operatorname*{arg\,min}_{W} V (W, \theta_{W})$, that minimizes the asymptotic variance of misspecified-GMM in (ref) or $W^{*} = \operatorname*{arg\,min}_{W} V_M (\Lambda^*, \theta_{W})$ minimizes the asymptotic variance of misspecification-efficient GMM in (ref). However, as noted in Hall and Pelletier (2011), economic theory dictates an appropriate choice of the weighting matrix in some cases and the researcher may choose $W$ based on which $\theta_W$ they care about substantively, and in the absence of these economic considerations, the choice of $W$ and the relative ranking of over $\mathcal{W}$ can be arbitrary.

In this paper, we treat $\mathcal{W}$ as fixed, but in practice $\mathcal{W}$ can be chosen by the researchers. For example, researchers may focus only on the identity matrix and diagonal matrices due to computational costs, or may put some restrictions to have particular causal effects interpretations.\\

Example: Linear IV (continued). Suppose we focus on a scalar $\theta$ ($p=1$), instruments $Z$ consists of $m$-single instruments $Z_{1}, \cdots Z_{m}$. Then, the GMM pseudo-true value can be written as a linear combination of the single-instrument IV estimands $\theta_m =E[Z_m^{\prime} Y]/E[Z_m^{\prime} X] $ \[ \theta_W = \frac{E[X'Z] W E[Z'Y]}{E[X'Z] W E[Z'X]} = \sum_{j=1}^{m} w_j \theta_{j} \] where the weight $w_j$ for the $j$-th instrument is given by: \[ w_j = \frac{\pi' W e_j e_j' \pi}{\pi' W \pi} =\frac{( \sum_{k=1}^{m} W_{jk} \pi_k) \pi_j}{{\pi' W \pi}} \] with the $j$-th unit vector $e_j$ and $\pi = E[Z^{\prime}X]$. While the weights always sum to 1, individual weights $w_j$ are not constrained to be positive. A weight $w_j$ can be negative if $\pi_j$ (the strength of $j$-th instruments $Z_j$) and $ \sum_{k=1}^{m} W_{kj} \pi_k$ (the weighted strength of full instrument vector $Z$) have opposite signs. When $W = E[Z'Z]^{-1}$, $\theta_W$ is the TSLS estimand and Angrist and Imbens (1995), Koles\'{a}r (2013) characterizes $\theta_W$ as a weighted average of LATE under monotonicity conditions (see also Andrews (2017)). However, the TSLS weights are not guaranteed to be positive.\footnote{For the same DGP we consider in the linear IV example in Section (ref), the pseudo-true value can be decomposed into $\theta_W = w_1\theta_1 + w_2 \theta_2$, where $w_1 = \frac{1 + 2\rho}{5 + 4\rho}, w_2 = \frac{2(\rho + 2)}{5 + 4\rho} $, $\theta_1=0, \theta_2 = 1/2$. For the TSLS ($\rho = 0$), the weights are $w_1 = \frac{1}{5}$ and $w_2 = \frac{4}{5}$, which are positive. Notably, both weights are positive for $\rho > -1/2$. However, for $\rho < -1/2$, the weight $w_1$ becomes negative.}

In general, researchers can restrict the class of weighting matrix $\mathcal{W}^{+}$ with positive weights $w_j$ as follows; \[ \mathcal{W}^{+} = \{ W \in \mathcal{W}: \pi^{\prime} W e_j e_j' \pi \geq 0 \textnormal{ for all } j=1, \cdots, m, \pi = E[ZX] \}. \] The class $\mathcal{W}^{+}$ consists of all matrices $W$ that preserve the “sign" of the instrument strength, i.e., such that $\text{sign}\left( (W \pi)_j \right) = \text{sign}(\pi_j)$. Note that any diagonal positive definite matrix (e.g., identity, or $\text{diag}((E[ZZ'])^{-1})$) is always in $\mathcal{W}^{+}$. The TSLS weighting matrix $W = (E[ZZ'])^{-1}$ is in this class only under strict conditions if we have covariates (e.g., Blandhol et al. (2025), and Sloczy\'{n}ski (2024)). Restricting to the class $\mathcal{W}^{+}$ allows us to interpret $\theta_W$ as a weighted average of the single-instrument estimands. Under the treatment effect heterogeneity, this condition guarantees that $\theta_W$ lies strictly between the minimum and maximum of the single-instrument LATE estimands $\theta_j$. Instead of considering the “right" causal target $\theta_{W^*}$, regardless of which convex combination that researchers care about with multiple instruments, we provide here uniform (over $W\in \mathcal{W}^{+})$ asymptotic minimax bounds of the class of causal LATE estimands.

Uniform Asymptotic Minimax Bounds

To formally state our results, we first define the class of distributions and the parameter spaces we consider. We also provide more specific definitions of correct/misspecified models.

Suppose that we observe an i.i.d. sample $\{ X_i \}_{i=1}^{n}$ drawn from the true (unknown) probability distribution $F_0$. Let $\mathcal{F}$ be the space of all probability distributions and let $\mathcal{F}^{C}$ as follows;

eqnarray[eqnarray omitted — 638 chars of source]

i.e., $\mathcal{F}^{C}$ is the set of probability distributions $F$, such that for some $\theta \in \Theta$, $( F, \theta)$ satisfies the regularity and the moment conditions in $\mathcal{F}^{C} $. Correct specification assumes $(F_0, \theta_0) \in \mathcal{F}^{C}$ such that the true probability distribution $F_0$ with unique parameter value $\theta_0$ satisfies conditions above. Chamberlain (1987) shows that $\widehat{\theta}_{GMM}( \Sigma_{11} (\theta_0)^{-1})$ achieves semiparametric efficiency bounds in the sense of H\'{a}jek's (1972) local asymptotic minimaxity for the neighborhood of $(F_0, \theta_0) \in \mathcal{F}^{C}$.

In this paper, we consider the misspecified moment condition models and the pseudo-true values associated with the choice of weighting matrix $W$. We define that the moment function $g(\cdot)$ is said to be misspecified if \[ (F_0, \theta) \notin \mathcal{F}^{C} \textnormal{ for all } \theta \in \Theta. \]

Then, for any $W \in \mathcal{W}$, where $\mathcal{W}$ is a set of positive definite matrices, we can define the pseudo-true values as a minimizer of goodness of fit measure of moment condition $g(\cdot)$ to the true data distributions $F_0$ using the weighted-norm associated with the weighting matrix $W$;\footnote{One can consider other pseudo-distance measure such as Kullback-Leibler (KL) divergence. The KL divergence from $F$ to $F_0$ is $d(F, F_0)= \int \log (dF_0/dF) dF_0$ if $F_0$ is absolutely continuous with respect to $F$, and $d(F, F_0) = \infty$, otherwise. Then, the pseudo distance from $\mathcal{F}^{C}$ to $F_0$ is defined by $d(\mathcal{F}^{C}, F_0) = \inf_{F \in \mathcal{F}^{C}} d(F, F_0) $. Vuong (1989), Kitamura (2000), Shi (2015) use this generic choice for model selection tests in moment moment equality/inequality models.} \[ \theta_{W} = \operatorname*{arg\,min}_{\theta} E_{F_0} g(X_i, \theta)'WE_{F_0} g(X_i, \theta). \]

To consider the uniform minimax bounds, we define the extended class of models $\mathcal{G}$ as the set of tuples, $\mathcal{G} = \{ (W, (F, \theta)) : W \in \mathcal{W}, (F, \theta) \in \mathcal{F}^{W} \}$, where $\mathcal{F}^{W}$ is the space of probability distributions that is consistent with $\theta_W$, i.e., the set of probability distributions $F$, such that for some $\theta \in \Theta$, $(F, \theta)$ satisfies the following regularity conditions;

eqnarray[eqnarray omitted — 965 chars of source]

Under the same identification assumption (Assumption (ref)), we have $(F_0, \theta_W) \in \mathcal{F}^{W}$. Condition $(i)$ requires that the original function $g(X, \theta)$ is twice continuously differentiable. Condition $(iv)$ requires that $\Gamma(\theta) = [G(\theta); F(\theta)]$ has rank $p$. This is a weaker condition than standard GMM identification (condition $(iv)$ in $\mathcal{F}^C$) because it allows identification to come from the curvature ($F(\theta)$) in addition to the first-order derivative ($G(\theta)$). $\mathcal{F}^{W} $ is defined through a recentering that depends on the pseudo-true value $\theta_W$ (and thus implicitly on $F_0$), and this is innocuous for the local asymptotic minimax analysis; it concerns the minimax risk over neighborhoods of $(F_0, \theta_W)$, and within such neighborhood, the recentering is fixed. \\

The next Theorem shows the semiparametric efficiency bounds in the sense of H\'{a}jek's (1972) local asymptotic minimaxity for the neighborhood of $(W, (F_0, \theta_W)) \in \mathcal{G} = \{ (W, (F, \theta)) : W \in \mathcal{W}, (F, \theta) \in \mathcal{F}^{W} \}$. \\

theoremSuppose that $(F_{0},\theta_W)$ satisfies Condition in $\mathcal{F}^W$ for all $W\in \mathcal{W}$. Let $\Delta^{W}$ be any neighborhood of $(F_{0},\theta_W)$, and let $\Gamma^{W}$ be the subset of $\Delta^W$ such that Condition in $\mathcal{F}^W$ is satisfied for all $(F,\theta) \in \Gamma^W$. Let $\mathcal{G}^{\Gamma} = \{ (W, (F, \theta)) : W \in \mathcal{W}, (F, \theta) \in \Gamma^W \}$. Let $\theta_{1, W}$ be the first component of $\theta_W$ and let $T_{n}(X_{1},\dots,X_{n})$ be any (measurable) estimator for $\theta_{1,W}$. Then for any loss function $l(\cdot) \in \mathcal{L}$, we have \[ \liminf_{n\rightarrow\infty} \sup_{(W, (F,\theta))\in \mathcal{G}^{\Gamma}} E_{F}\left\{l\left[\sqrt{n}(T_{n}-\theta_{1,W})\right]\right\} \ge \ \int_{-\infty}^{\infty}l(\sigma u)d\Phi(u), \] where $\Phi(u)$ is a cumulative distribution function of the standard normal distribution, $\sigma^2 = \sup_W \sigma_{W}^2 <\infty$, $\sigma_{W}^2$ is the (1,1) element of $V_{ME}(W) = (\Gamma(\theta_W)^{\prime} \Sigma (\theta_W)^{-1}\Gamma(\theta_W))^{-1}$.\\

Theorem (ref) states that the asymptotic variance of any semiparametric estimator under the model class $\mathcal{G} = \{ (W, (F, \theta)) : W \in \mathcal{W}, (F, \theta) \in \mathcal{F}^{W} \}$, is no smaller than the bounds with the least favorable weighting matrix $W$ that maximises risk over the class $\mathcal{W}$. This bound can be achieved by the worst-case variance of the ME estimator $\widehat{\theta}_{ME} (W), W \in \mathcal{W}$, which can be feasibly approximated by the analytic formula or the bootstrap method. Since $V_{ME}$ doesn't depend on $W$ in the linear IV model, the ME estimator eliminates the question of which $W$ is most efficient among $W \in \mathcal{W}$, and the uniformity can be obtained with no cost.

One can straightforwardly show that the misspecified-GMM achieves different semiparametric efficiency bounds based on $V_{GMM}(W, \theta_{W})$ by replacing moment condition ($ii$) in $\mathcal{F}^C$ with the (local) first order condition $E_F[ \partial g(X, \theta)/\partial \theta]^{\prime}W E_F [g(X, \theta)] = 0$ (Mukhin (2019)) or using the augmented parameter space approach as in Imbens (1997) with similar regularity (smoothness) conditions. However, as shown in Corollary (ref), the standard misspecified-GMM Estimator, which uses the linear combination $\Lambda = [G(\theta_W)^{\prime} W \quad g(\theta_W)^{\prime} W \otimes I_p]$ dictated by the FOC of the GMM objective, is inefficient relative to the bound we consider here, because it ignores the potential information contained in the variation of the Jacobian. By using the optimal linear combinations of the augmented moments $\Lambda^* = \Gamma(\theta_W)^{\prime} \Sigma (\theta_W)^{-1}$, the moment vector $\widetilde{\psi}(X, \theta)$ provides an over-identified system for $\theta$, exploiting more information from the degree of misspecification and the identification strength. \\

Monte Carlo Simulation

We compare the finite-sample performance of the standard GMM and the proposed misspecification-efficient estimator in a linear IV model. The data generating process is

equation[equation omitted — 89 chars of source]

with $Z_i = (Z_{1i},Z_{2i})' \sim_{iid} N(0,I_2)$ and $(\varepsilon_i, v_i) \sim N(0,[1, 0.5; 0.5, 1])$. We set $\theta=1$, and vary the “degree of misspecification" through the direct effect $\gamma=(0,\delta)'$ with $\delta \in \{0,0.5,1,2\}$; $\delta=0$ is the correctly specified benchmark, $\delta > 0$ introduces an exclusion violation through $Z_2$. The sample sizes are $n \in \{200,500, 1000\}$. We set the first-stage coefficient $\Pi=(\pi,2\pi)'$ so that the scaled concentration parameter ($\mu^2/m$) is 50 or 10, corresponding to moderately strong and weak instrument setups for $n=200$. The pseudo-true value $\theta_W$ is $W$-dependent and moves away from $\theta=1$ as $\delta$ changes. We consider $W = I, \widehat\Sigma_{11,2}^{-1}$, and we report the moderately strong instrument case here, and the weak instrument design results, together with an additional results under the conventional optimal weighting $W = \widehat\Sigma_{11}^{-1}$ are reported in Appendix C. For each estimator, we report the standard deviation (SD), empirical $95\%$ coverage (Cov), and median CI length (Len) over $2,000$ Monte Carlo replications, using $B=1,000$ bootstrap draws and $S=100$ sample-split replications. SD is normalized, relative to the SD of standard GMM with the correctly-specified benchmark $(n=200, \delta=0)$. We do not report the bias here as it is nearly zero for all estimators targeting the same $\theta_W$.

Tables (ref) and (ref) report results for the one-step $(W = I)$, and the two-step GMM $(W = \widehat\Sigma_{11,2}^{-1})$, respectively. For each $(n,\delta)$, we report the following estimators and standard errors (SE) with Wald CI: (i) standard GMM estimator with the conventional SE; (ii) GMM with the misspecification-robust SE; (iii) the oracle ME estimator with efficient $V_{ME}(W)$; (iv) Sample-splitting ME estimator, the median repeated sample-split estimates with the median standard deviation. We also considered the following bootstrap estimators; (v) HH (Hall and Horowitz, 1996) percentile bootstrap CI; and (vi) the DR (double-recentered) percentile bootstrap CI proposed in Section (ref).

As the degree of misspecification ($\delta$) increases, we observe several consistent patterns across DGPs; 1) the SD of GMM increases, and 2) the coverage of GMM with conventional SE undercovers, while the misspecification-robust SE restores coverage by correctly capturing the SD of GMM; 3) the oracle ME dominates in SD and CI length across all designs, while retaining coverage (coverage 94-95% across all $\{ n, \delta \}$ in Tables (ref) and (ref)), consistent with Corollary (ref) and (ref). Note that the SD of standard GMM grows substantially with $\delta$, whereas the SD of oracle ME grows only modestly (e.g., SD of GMM $0.43 \rightarrow 1.38$, SD of ME $0.40 \rightarrow 0.79$, when $\delta=0 \rightarrow 2$, $n=1000$ in Table (ref)); the oracle ME estimator take into account increases in variance through optimal Jacobian recentering rather than letting the asymptotic variance inflate. The efficiency gain of the oracle ME over GMM are $25.8\%$, $54.3\%$, $67.4\%$ for $W=I$, and $23.8\%$, $50.1\%$, $63.8\%$ for $W=\widehat\Sigma_{11,2}^{-1}$, when $\delta=0.5, 1, 2$, respectively. We also observe that the efficient weighting choice $W=\widehat\Sigma_{11,2}^{-1}$ reduces the SD of GMM.

Furthermore, among feasible procedures, coverage of the sample-splitting ME and DR is very close to the nominal level in most cases, while the HH bootstrap exhibits the coverage distortions. While the sample-splitting ME provides similar coverage/lengths, the DR CI lengths shows meaningfully smaller CI compare to the CI with analytic GMM misspecification-robust SE (10-17% smaller in Table (ref) for $\delta \neq 0$). The misspecification-robust SE adds a Jacobian uncertainty term that inflates SE with the moment violations, and this SE can be large under mild or severe misspecification, which leads to over-coverage and larger CI lengths in Table (ref).

The simulation results confirms the efficiency of the oracle ME estimator, and overall suggest that the DR bootstrap and sample-split procedures provides valid procedures, and reasonable alternatives to the standard GMM with analytic misspecification-robust standard errors.

table[table omitted — 3,770 chars of source]
table[table omitted — 2,775 chars of source]

Empirical Examples

We illustrate our proposed methods with two main empirical applications: returns to schooling example of Card (1995) and Angrist and Krueger (1991). We also consider the dynamic panel data model of income and democracy in Acemoglu et al. (2008) in Appendix D.

Card (1995)

We replicate the well-known college-proximity IV analysis by Card (1995) who uses the National Longitudinal Survey of Young Men for 1976. The sample size is $n=3010$. The outcome is log wage, the regressors include education and some demographic/geographic variables. The instruments include two binary indicators for living near a four-year public and private college, respectively.

Table (ref) compares several estimators across four specifications, and three weighting choices: $W=I$, $W=(Z'Z)^{-1}$ (2SLS), and $W=\widehat\Sigma_{11,2}^{-1}$ (efficient).\footnote{In both of applications we consider here, we also considered conventional optimal weight $W=\widehat\Sigma_{11}^{-1}$ for each specification. Results are almost identical and thus we do not report in the paper.} The first column (1) corresponds to the baseline specification with experience, experience squared, an indicator for Black, indicators for residence in an SMSA in 1976, and residence in the South in 1976. Column (3) treat experience and its square as endogenous, adding age and age squared to the instrument sets. These specifications correspond to Table 3 in Card (1995). In Columns (2) and (4), we additionally control for more covariates including cubic and quartic experience, and all pairwise interactions between baseline covariates, motivated by the “rich covariates" condition of Blandhol et al. (2025), required for the 2SLS estimand to admit a weakly causal estimand.

Table (ref) shows that GMM point estimates are very similar across specifications and $W$ (0.164-0.175 in four columns with three $W$). The conventional and misspecification-robust SEs are very similar in this example (within 1.5%). The Hansen $J$ test does not reject in any specifications (p-values $\ge 0.35$), however, misspecification can still exist even when $J$ test reject. The $J$-test is not informative to decide whether ME bound is informative, and the moment-Jacobian covariance $\Sigma_{12}$ is more relevant in this direction. Although, the ME efficiency bound is essentially similar across specifications (0.0346-0.0366), it is 15-20% smaller than the conventional/misspecification-robust SE. The ME-GMM bootstrap SD tracks the bound closely, and the DR percentile CI is comparable to the Wald CI with misspecification-robust SEs across specifications.

Similar to Mogstad and Torgovitsky (2024), who consider similar specifications in Card applications, we find that a Ramsey (1969) RESET test rejects the null of rich covariates in specifications (1) and (2). However, when treating experience as endogenous, we do not reject the RESET test (p-value = 0.88) in Columns in (3) and (4) at 1% significance level, which indicates that the first-stage nonlinearity is captured by the dummy interactions terms. Since the point estimates are similar across all specification, the same caveat emphasized by Mogstad and Torgovitsky (2024) still applies: two estimates can be very similar even when one corresponds to a non-negatively weighted causal estimand and the other does not.

table[table omitted — 3,567 chars of source]

Angrist and Krueger (1991)

In this section, we consider Angrist and Krueger (1991) quarter-of-birth instruments example. We use the US Census 1980 extract of men born 1930-1939 ($n=329, 509$). The outcome is the log wage, and endogenous variable is education with exogenous controls (year-of-birth, race, SMSA, marital status, region, state). The instruments are quarter-of-birth (QOB) with the interactions. In Table (ref), Column (1) uses the baseline QOB dummies $(m=3)$; (2) uses QOB$\times$YOB interactions ($m=30$), (3) adds QOB$\times$State interactions ($m=180$) which may exhibits many and weak instruments issues as noted in Bound et al. (1995). These specifications correspond to Angrist and Krueger (1991, Table V Columns 2 and 4 and Table VII Column 8). For each case, we consider the basic controls and the flexible specification that adds all interactions between controls.

The results in Table (ref) parallel Table (ref) in the Card (1995) example. The misspecification-robust SE only marginally exceeds the conventional SE, and the gap is larger for many instrument specifications. The ME-GMM bootstrap SD tracks the ME efficiency bound closely, in line with the simulation evidence. The ME efficiency bound is significantly smaller than the conventional/misspecification-robust SE when $W=I$. When we use 2SLS ($(Z'Z)^{-1}$) and efficient weights ($\widehat\Sigma_{11,2}^{-1}$), the ME efficiency bounds are almost similar to the conventional SE, but still 10-20% smaller than the misspecification-robust SE as predicted by our theory.

Adding interaction terms barely changes the GMM point estimates. However, GMM point estimate under $W=I$ is systematically higher ($0.091$--$0.113$) than under 2SLS or efficient weighting ($0.081$--$0.099$). Furthermore, DR bootstrap CI is close to the Wald CIs under 2SLS and efficient weighting, but it is noticeably wider when $W = I$. In all specifications, we reject a Ramsey (1969) RESET test. ME-GMM inference still applies to the pseudo-true value $\theta_W$, but the rich covariates condition of Blandhol et al. (2025) fails, so causal interpretations of all of 2SLS/GMM estimates in Table (ref) may be fragile in this application.

table[table omitted — 4,163 chars of source]

Conclusion

This paper develops the theory of efficient GMM estimation under moment misspecification. The key insight is that the influence function of the standard misspecified-GMM estimator depends both on the original moment conditions and their Jacobian. By optimally weighting this augmented system, we propose a misspecification-efficient estimator that achieves a smaller asymptotic variance than standard GMM for the same pseudo-true value. We also establish semiparametric efficiency bounds uniform over a class of weighting matrices, and propose a bootstrap implementation based on double recentering of both moment and Jacobian conditions.

Several directions remain for future work. First, extending the theory to dependent data with HAC-type weighting matrices would significantly broaden the applicability of the ME estimator to time-series settings. Second, the relationship between the ME estimator and generalized empirical likelihood (GEL) estimators under misspecification deserves further investigation. Under misspecification, these estimators have different pseudo-true values and different asymptotic distributions, so generally difficult to compare in a unified framework. We note however, that the continuously updating estimator (CUE) implicitly updates the weighting matrix as a function of $\theta$, which may share connections with the augmented-moment approach. Third, extending the framework to accommodate weak identification, where the Jacobian $G(\theta_W)$ is near-singular, would be important given the connection to the literature (e.g., Lee and Liao (2018) and Kleibergen (2005)) that already exploits Jacobian information. Finally, developing practical guidance on the choice of the weighting matrix class $\mathcal{W}$ would provide applied researchers with concrete recommendations for implementation.