Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
88,995 characters · 14 sections · 2 citation commands
Sensitivity, Informativeness, and Misspecification in GMM Estimation
\noindentKeywords: Generalized method of moments; misspecification; sensitivity analysis; influence function; informativeness; structural estimation.\\ \noindentJEL classification: C13, C26, C36, C52.
\thispagestyle{empty}
\setcounter{page}{1}
Assessing sensitivity to model specification is central to evaluating the robustness of empirical conclusions. In reduced-form settings, researchers have access to a rich toolkit: for example, the omitted variable bias formula in linear regression provides a transparent benchmark for assessing the impact of confounding variables, and formal sensitivity diagnostics have been developed for a range of reduced-form estimators \citep*{rosenbaum1983assessing, imbens2003sensitivity, oster2019unobservable, cinelli2020making}. Comparable tools for structural estimators are far less standardized. The relationship between parameters and identifying moment conditions is implicit and often nonlinear, and misspecification is prevalent in applied work.
Empirical researchers increasingly rely on sensitivity diagnostics to assess how GMM estimates depend on particular moment conditions. In many applications, however, overidentifying restrictions are strongly rejected, and the maintained moment conditions fail to hold exactly at any parameter value. In such settings, the GMM estimator does not converge to a structurally “true” parameter satisfying the moment restrictions; it converges instead to a pseudo-true value defined as the minimizer of a misspecified population objective function. Local sensitivity can no longer be interpreted as a perturbation around a correctly specified benchmark, and the relevant object is the mapping from the moment conditions to the estimator's misspecified probability limit. This paper develops sensitivity and informativeness measures defined relative to the pseudo-true value and shows that this change of benchmark gives the diagnostics different economic content and interpretation.
A growing literature has begun to formalize sensitivity analysis for structural estimators using large-sample approximations. \citet*{gentzkow2014measuring} introduce the concepts of asymptotic sensitivity and asymptotic sufficiency, which characterize how parameter estimates covary with auxiliary sample statistics. \citet*{andrews2017measuring} (hereafter AGS) formalize sensitivity for GMM through local perturbations of the moment conditions, mapped linearly into asymptotic bias by a closed-form sensitivity matrix. These approaches evaluate sensitivity at a benchmark where the moment conditions are correctly specified. The natural reference point under misspecification is the pseudo-true value, which retains economic meaning in several settings of empirical interest. For instance, in overidentified IV with heterogeneous treatment effects, the 2SLS estimator converges to a weighted average of the local average treatment effects \citep*{angrist1996identification, lee2018consistent}.
Recent work argues that applied researchers should default to misspecification-robust inference for overidentified GMM \citep*{andrews2025purpose}. We take this recommendation as our starting point and ask what additional diagnostic information the misspecification-robust influence function provides. We propose a misspecification-robust sensitivity (MRS) matrix $\Lambda$ that characterizes how local deviations in the moment conditions affect the first-order behavior of GMM estimators around their pseudo-true probability limits. To construct MRS, we derive influence function representations for one-step, two-step, and iterated GMM, explicitly accounting for misspecification and for variation arising from estimated weight matrices; continuously updating GMM is treated in Appendix (ref). Relative to AGS, MRS contains additional terms operating through the Jacobian of the moment conditions and the estimated weight matrix; these terms are closely related to components of the misspecification-robust asymptotic variance \citep*[e.g.,][]{imbens1997one, hall2003large, hansen2021inference, hwang2022doubly}.
The misspecification-robust influence function also yields an informativeness measure $\Delta_k$, defined as the population $R^2$ from the linear projection of the estimator's influence function on the moments. \citet*{gentzkow2014measuring} show that the analogous quantity, called sufficiency in their framework and informativeness in \citet*{andrews2020informativeness}, equals one under correct specification; this equivalence breaks down under misspecification. We reinterpret $\Delta_k$ as a measure of structural efficiency: the proportion of the estimator's asymptotic variance attributable to sampling variation in the moments, as opposed to variation arising from the Jacobian or the estimated weight matrix. Unlike classical GMM efficiency, which concerns the choice of optimal weights to minimize variance, structural efficiency quantifies a different source of variance loss. This loss exists even at the efficient weight matrix when the model is misspecified. $\Delta_k$ is complementary to the Hansen $J$-test \citep*{hansen1982large}. The $J$-test asks whether the moment restrictions are jointly violated; $\Delta_k$ asks how much of the estimator's variance is left unexplained by the moments after such a violation. The income-and-democracy application of \citet*{acemoglu2008income} (AJRY) in Section (ref) makes the distinction concrete. At the iterated fixed point the income coefficient has informativeness $\widehat\Delta_\gamma=0.77$: only $77\%$ of its asymptotic variance is explained by the moments, the rest a structural-efficiency loss the $J$-test does not quantify. $\widehat\Delta_\gamma$ is invariant to the centering of the weight matrix, whereas the $J$-test verdict flips with it, from failing to reject without centering ($p=0.42$) to rejecting with centering ($p=0.007$).
We make three contributions. First, we define a projection-based sensitivity measure for GMM estimators evaluated at pseudo-true values that nests the AGS sensitivity matrix under correct specification. Second, we derive misspecification-robust influence functions for one-step, two-step, and iterated GMM, organize them as a channel decomposition into moment, Jacobian, weight-matrix, and first-step components, and use this decomposition to characterize when informativeness falls strictly below one. Third, we apply the framework to three canonical structural settings: the automobile demand model of \citet*{berry1995automobile} (BLP), the consumption insurance model of \citet*{blundell2008consumption} (BPP), and the AJRY income-and-democracy regressions. Accounting for misspecification reorders sensitivity rankings in BLP; in BPP it lowers $\widehat\Delta$ below one for every parameter under the optimal weight but leaves it near one under the diagonally weighted scheme actually used in BPP; and in AJRY it reveals structural-efficiency loss at the iterated fixed point that is invariant to the weight-matrix centering on which the conclusion of the Hansen $J$-test depends. The diagnostics are computable directly from the same inputs as misspecification-robust standard errors.
As a by-product, the framework rationalizes a widespread practice in minimum distance estimation: since the efficient weight matrix contributes a misspecification channel that simpler weight matrices largely avoid (Proposition (ref)), the common use of diagonally weighted or fixed-weight estimators trades efficiency for informativeness. This trade-off is an asymptotic counterpart to the small-sample bias argument of \citet*{altonji1996small}.
Our paper is also related to a recent literature on local sensitivity in structural models, including \citet*{jorgensen2023sensitivity, iskrev2019expect} on calibrated parameters, \citet*{christensen2023counterfactual} on counterfactual conclusions, \citet*{armstrong2021sensitivity} on moment selection, and \citet*{bonhomme2022minimizing} on estimators designed to minimize the impact of misspecification. Our contribution is complementary in focusing on sensitivity to moment conditions evaluated at the estimator's misspecified probability limit.
The remainder of the paper is organized as follows. Section (ref) defines misspecification-robust sensitivity and informativeness in terms of the conditional expectation of the estimator given the moments, and expresses both measures through influence functions. Section (ref) derives the misspecification-robust influence functions for one-step, two-step, and iterated GMM (Section (ref)). It then characterizes the influence function of optimal minimum distance, in which the optimal weight's misspecification channel appears in isolation, and the implications for diagonal weighting (Proposition (ref)). Section (ref) illustrates the resulting sensitivity and informativeness through canonical examples, including 2SLS and a normal mean with a misspecified variance restriction. Section (ref) applies the diagnostics to the BLP model of automobile demand, the BPP model of household consumption insurance, and the AJRY income-and-democracy regressions. Section (ref) concludes.
Consider i.i.d.\ random vectors \(X_1,\ldots,X_n\) with unknown distribution \(P_0\). Let \(g(X_i,\theta)\) be a \(q\times 1\) vector of moment functions, and let \(\theta \in \Theta \subset \mathbb{R}^p\) be a \(p\times 1\) parameter vector with \(q>p\). A GMM estimator is defined as \[ \widehat\theta = \arg\min_{\theta \in \Theta} \widehat g(\theta)' W \widehat g(\theta), \] where \(\widehat g(\theta) = n^{-1}\sum_{i=1}^n g(X_i,\theta)\) is the sample moment vector and \(W\) is a positive semidefinite weight matrix. For expositional clarity, we treat the weight matrix \(W\) as deterministic in this section and consider stochastic weight matrices in Section (ref). The moment condition model is said to be correctly specified if there exists \(\theta \in \Theta\) such that \(\mathbb{E}[g(X_i,\theta)] = 0\); otherwise the model is misspecified.
We denote by \(\theta_0\) the probability limit of the GMM estimator \(\widehat\theta\). Under standard regularity conditions, \(\theta_0\) is the unique minimizer of the population GMM criterion \[ Q(\theta) = \mathbb{E}[g(X_i,\theta)]' W \mathbb{E}[g(X_i,\theta)] . \] When the model is correctly specified, this minimizer satisfies \(\mathbb{E}[g(X_i,\theta_0)] = 0\) and therefore coincides with the true parameter. When the model is misspecified, \(\theta_0\) is the pseudo-true parameter that best fits the moment conditions in the GMM objective. Let \(g(\theta)=\mathbb{E}[g(X_i,\theta)]\) and \(g = g(\theta_0)\).
Suppose the standard regularity conditions hold, e.g., Theorem 1 of \citet*{hall2003large}. Then,
By properties of the multivariate normal distribution, the conditional expectation of \(\widetilde\theta\) given \(\widetilde g\) is linear and given by
We interpret the corresponding population regression coefficient as a measure of the sensitivity of the estimator to the moments.
The sensitivity matrix \(\Lambda\) is \(p\times q\). Its \((k,l)\) element measures how a local deviation in the \(l\)th moment affects the first-order behavior of the \(k\)th component of the estimator around its probability limit, holding the other moments fixed. The linear relationship in equation (ref) also appears in \citet*{andrews2017measuring}, who study sensitivity under correct specification and analyze the response of \(\widehat\theta\) to local perturbations of the data distribution. Our framework extends their sensitivity measure by allowing for misspecified moment conditions. Under correct specification, MRS nests the AGS sensitivity matrix (Section (ref)).
Sensitivity captures how the estimator responds locally to deviations in the realized moments. A complementary concept concerns the extent to which the moments are informative for the estimator. We adapt the notion of informativeness from \citet*{andrews2020informativeness} to the GMM setting under misspecification. In their framework, $\widehat c$ is a scalar structural estimator and $\widehat\gamma$ is a vector of descriptive statistics. Informativeness is the population $R^2$ from regressing $\widehat c$ on $\widehat\gamma$ under their joint asymptotic distribution. We specialize their setup by taking $\widehat c$ to be a component of the GMM estimator $\widehat\theta$ and $\widehat\gamma$ to be the sample moment vector $\widehat g(\theta_0)$.
By Definition (ref), $\sigma_{\theta_k g}\sigma_{gg}^{-1}$ is the $k$th row of $\Lambda$, so $\Delta_k$ in Definition (ref) is the population $R^2$ from the linear projection of $\widetilde\theta_k$ on $\widetilde g$. We interpret $\Delta_k$ as a measure of structural efficiency under misspecification. By “structural efficiency” we mean the proportion of the estimator's asymptotic variance attributable to sampling variation in the moments, as opposed to variation arising from the Jacobian or the estimated weight matrix. Lower values indicate that a larger share of the estimator's variance is left unexplained by the moments.
Under correct specification, the influence function of the GMM estimator is $\psi(X_i) = \Lambda_{AGS}\, g(X_i, \theta_0)$, where $\Lambda_{AGS} = -(G'WG)^{-1}G'W$ and $G = \mathbb E[\partial g(X_i, \theta_0)/\partial \theta']$ is the population Jacobian of the moments at $\theta_0$. Since the influence function is a deterministic linear combination of the moments, it is spanned by the moments themselves, and $\Delta_k = 1$ for every $k$. \citet*{andrews2020informativeness} make essentially this observation in their footnote 2 (p. 2234) and implement it in their Section 5.1, Step 4: when the descriptive statistic is a vector of estimation moments that fully determines the structural estimator, $\Delta_k = 1$ by construction. Under misspecification, the GMM influence function contains additional terms that are not in the linear span of the moments, and $\Delta_k$ is generally less than one. We derive these additional terms in Section (ref).
Structural efficiency is distinct from the test of overidentifying restrictions. The $J$-test asks whether the population moments are jointly compatible with zero at some parameter value. Structural efficiency instead asks how much of the estimator's asymptotic variation is linearly explained by sampling variation in the moments themselves. A rejected $J$-test does not by itself indicate whether the misspecification materially affects the precision or stability of a particular parameter estimate; $\Delta_k$ measures this channel. The two diagnostics are therefore complementary.
To implement the proposed sensitivity and informativeness measures, we require estimates of the asymptotic covariance matrix \(\Sigma\). We use influence function representations for this purpose. For an asymptotically linear estimator such as the GMM estimator \(\widehat\theta\), an influence function \(\psi(x)\) satisfies
Let \(\psi\) and \(\nu\) denote influence functions associated with the GMM estimator \(\widehat\theta\) and the sample moment vector \(\widehat g(\theta_0)\), respectively. Then
where \[ \nu(X_i) = g(X_i,\theta_0) - g. \] As a consequence, the misspecification-robust sensitivity defined in Definition (ref) can be written as
Similarly, the informativeness of the moments for the \(k\)th parameter, \(\Delta_k\) in Definition (ref), is given by
where \(\psi_k\) is the \(k\)th component of \(\psi\).
A closely related expression appears in \citet*{iskrev2019expect}, who measures the sensitivity of structural parameters to calibrated parameters by treating the latter as if estimated. \citet*{jorgensen2023sensitivity} argues that calibrated parameters are non-stochastic and should be held fixed, leading to a different sensitivity measure. Equation (ref) concerns sample moments, which are stochastic by construction, and so aligns formally with the Iskrev treatment though applied to a different object.
Equations (ref) and (ref) are expressed in terms of population moments and influence functions. In practice, the sensitivity matrix \(\Lambda\) and the informativeness measure \(\Delta_k\) can be consistently estimated using a plug-in approach based on estimated influence functions. We construct estimated influence functions, denoted \(\widehat{\psi}_i\) and \(\widehat{\nu}_i\), for each observation \(i=1,\ldots,n\), by replacing unknown population quantities with their sample counterparts evaluated at the GMM estimate \(\widehat{\theta}\).
First, the estimated influence function for the moments, \(\widehat{\nu}_i\), is obtained by centering the moment function evaluated at the GMM estimate, \[ \widehat{\nu}_i = g(X_i,\widehat{\theta}) - \widehat g(\widehat{\theta}). \] Second, the estimated influence function for the GMM estimator, \(\widehat{\psi}_i\), is obtained by substituting sample estimators into the analytic expressions for the influence function \(\psi(X_i)\) derived in Proposition (ref) below. The population influence function depends on the Jacobian of the moment conditions, the weight matrix, and additional terms involving derivatives of the GMM criterion. We replace these population quantities with their sample counterparts evaluated at \(\widehat\theta\) and compute \[ \widehat{\psi}_i = \widehat\psi(X_i;\widehat\theta). \] With these estimated influence functions in hand, the sensitivity matrix \(\widehat{\Lambda}\) can be computed using a sample analogue of equation (ref). Specifically, \(\widehat{\Lambda}\) is obtained by regressing \(\widehat{\psi}_i\) on \(\widehat{\nu}_i\), which yields \[ \widehat{\Lambda} = \left( \sum_{i=1}^n \widehat{\psi}_i \widehat{\nu}_i^{\prime} \right) \left( \sum_{i=1}^n \widehat{\nu}_i \widehat{\nu}_i^{\prime} \right)^{-1}. \]
Similarly, the informativeness measure \(\widehat{\Delta}_k\) can be estimated using a sample analogue of equation (ref), based on the sample variances and covariances of \(\widehat{\psi}_{ik}\) and \(\widehat{\nu}_i\). This implementation requires only the analytic form of the influence function \(\psi\), which is derived in the next section for a range of GMM estimators.
Equations (ref) and (ref) show that the misspecification-robust sensitivity, \(\Lambda\), and informativeness, \(\Delta_k\), can be computed directly from the influence functions of the estimator and the moments. In this section, we derive the misspecification-robust influence functions for GMM estimators. As shown in \citet*{hall2003large} and subsequent work, the asymptotic behavior of GMM under misspecification depends on both the choice of the weight matrix and the estimation procedure.
We consider a class of GMM estimators defined as minimizers of \[ \widehat Q_n(\theta) = \widehat g(\theta)' \widehat W \widehat g(\theta), \] with different choices of the weight matrix.
In one-step GMM, the weight matrix does not depend on the parameter. We consider both the case of a deterministic weight matrix, where \(\widehat W = W\) is non-stochastic, and the case of an estimated weight matrix satisfying \(\widehat W \xrightarrow{p} W\), where \(\widehat W\) is asymptotically linear and does not depend on \(\theta\).
In two-step efficient GMM, the weight matrix is evaluated at a preliminary estimator \(\widehat\phi\). Specifically, the second-step estimator minimizes the criterion using the efficient weight matrix \[ \widehat W(\widehat\phi) = \left( \frac{1}{n}\sum_{i=1}^{n} g(X_i,\widehat\phi) g(X_i,\widehat\phi)' \right)^{-1}. \]
Iterated GMM updates the weight matrix repeatedly by re-evaluating it at the current parameter estimate until convergence. At convergence, the estimator minimizes the criterion with weight matrix \(\widehat W(\widehat\theta)\).
We now introduce notation used in the influence function derivations. Let $\widehat W(\theta)$ be a parameter-dependent weight matrix estimator, and its probability limit be $W(\theta)$. Let $W = W(\theta_0)$.
The influence function derivations rely on the population curvature of the GMM objective. Let \[ R = \mathbb E\!\left[\frac{\partial}{\partial\theta'}\operatorname{vec}(G(X_i,\theta_0)')\right], \qquad H = (g'W\otimes I_p)R. \] These matrices are part of the one-step GMM influence function. Define \[ \gamma_i = \operatorname{vec}\{G(X_i,\theta_0)' - G'\}. \] We use the convention $\operatorname{vec}(G')$ rather than $\operatorname{vec}(G)$, under which Proposition (ref) relies on the identity $(G(X_i,\theta_0) - G)'Wg = (g'W\otimes I_p)\gamma_i$ for any weight matrix $W$. The influence functions of two-step and iterated GMM additionally involve \[ B = (g'\otimes G')S, \] where $S$ is the derivative of $\operatorname{vec} W(\cdot)$ with respect to its argument, evaluated at the relevant parameter value: $S = \partial \operatorname{vec} W(\phi)/\partial\phi'|_{\phi_0}$ for two-step GMM and $S = \partial \operatorname{vec} W(\theta)/\partial\theta'|_{\theta_0}$ for iterated GMM. The matrix $B$ captures the dependence of the first-order condition on the parameter inside the weight matrix, and governs the linearization of the iteration map (Proposition (ref)).
The matrices $H$ and $B$ depend on the misspecification vector $g$ and vanish under correct specification. The population curvature matrix is defined as \[ A = G'WG + H. \]
We assume the following regularity conditions.
Under Assumption (ref), the GMM estimators considered above are consistent for their pseudo-true probability limits and asymptotically linear, with influence functions given by Proposition (ref). Consistency under misspecification follows from standard arguments for misspecified GMM \citep*[e.g.,][]{hall2003large,hansen2021inference,hwang2022doubly}; these influence functions are the inputs to the misspecification-robust sensitivity and informativeness measures in equations (ref) and (ref).
Each bracket contains a moment channel ($G'W\nu_i$) and misspecification channels (Jacobian, weight-matrix, and first-step). For ease of reference, let \[ M_\nu = -A^{-1}G'W, \quad M_\gamma = -A^{-1}(g'W\otimes I_p), \quad M_\omega = -A^{-1}(g'\otimes G'). \] The iterated case differs from the one-step estimated-weight case only in replacing $A^{-1}$ with $(A+B)^{-1}$. The matrices $M_\gamma$, $M_\omega$, and $B$ each contain $g$ as a factor and vanish identically when $g = 0$.
Versions of these influence functions appear in \citet*{hall2003large}, \citet*{hwang2022doubly}, and \citet*{hansen2021inference}. Continuously Updating GMM (CUGMM), which jointly updates the parameter and the weight matrix, presents distinct theoretical challenges under misspecification, including non-convexity and the risk of variance collapse \citep*{kleibergen2025double}. We provide analysis of CUGMM, including its consistency and a derivation of the misspecification-robust influence function, in Appendix (ref).
Under $g = 0$, $M_\gamma = M_\omega = 0$, so $\psi(X_i) = M_\nu\nu_i$, $\Lambda$ collapses to $M_\nu = \Lambda_{AGS}$, and $\Delta_k = 1$ for every $k$. The condition $g = 0$ is sufficient for $\Delta_k = 1$ but not necessary. $\Delta_k$ can equal one under misspecification when the misspecification channels lie in the linear span of the moments. The next corollary states the necessary and sufficient condition.
Among the channels of Proposition (ref), the weight-matrix channel is the one the researcher controls through the choice of weight matrix. This choice is most transparent in minimum distance estimation, where the moment is the matching residual $g(X_i,\theta)=m_i-f(\theta)$: $m_i$ collects the data moments, with mean $\mu$ and covariance $V=\operatorname{Var}(m_i)$, and $f$ is the model-implied map. Minimum distance isolates the weight-matrix channel. The centered moment $\nu_i=m_i-\mu$ and the covariance $V$ do not depend on $\theta$, so the optimal weight $\widehat W=\widehat V^{-1}$, the inverse of the sample covariance of the data moments, involves no preliminary parameter estimate and the first-step channel is absent; the Jacobian $-\partial f(\theta)/\partial\theta'$ is nonrandom, so $\gamma_i=0$ and the Jacobian channel vanishes. Only the moment and weight-matrix channels remain, and the influence function of the optimal minimum distance (OMD) estimator can be characterized in full. Minimum distance is also the setting where the choice between the optimal and a diagonal weight is a long-standing practical concern \citep*{altonji1996small,cheng2026weight}, and the setting of the BPP application in Section (ref). Proposition (ref) derives the OMD influence function in closed form and relates the resulting variance and informativeness to the population overidentification criterion $g'Wg$.
The weight-matrix channel is the moment channel times the mean-zero scalar $\nu_i'Wg$, a product quadratic in $\nu_i$ that lies outside $\mathcal L(\nu_i)$ and, by Corollary (ref), lowers $\Delta_k$. The squared multiplier $(1-\nu_i'Wg)^2$ has mean exactly $1+g'Wg$. Relative to the benchmark $V_0$, estimating the optimal weight adds the population overidentification criterion times the moment-channel variance, plus the cumulant $\kappa_k$. Under a fixed weight, e.g., equally weighted minimum distance (EWMD) which fixes the weight at the identity, the channel is absent, $\psi(X_i)=M_\nu\nu_i$, and $\Delta_k=1$ regardless of misspecification. The gap is therefore the informativeness cost of estimating the efficient weight. This structure is specific to weights that invert the covariance matrix of the moments. A generic data-dependent weight has a weight-matrix channel too, but one that neither factors through a single scalar nor scales with the population overidentification criterion.
The diagonally weighted minimum distance (DWMD) estimator is another popular choice in practice. The DWMD weight is $\widehat D=(\operatorname{diag}\widehat V)^{-1}$, the inverse of the diagonal of the same sample covariance that OMD inverts in full: each moment is weighted by the inverse of its own sampling variance, $\widehat D_{kk}=1/\widehat V_{kk}$.\footnote{This inverse-variance weighting follows \citet*{chatterjee2021estimating} and \citet*{cheng2026weight}. It is distinct from $\operatorname{diag}(\widehat V^{-1})$, the diagonal of the optimal weight, which coincides with it only when $V$ is diagonal; either replaces the rank-one outer product with a full-rank diagonal.} Estimating this diagonal weight contributes a weight-matrix channel through the influence function $\omega_i^D$ of $\widehat D$, but one that does not simplify as in OMD. The $k$th diagonal entry $1/\widehat V_{kk}$ has influence function $-(\nu_{ik}^2-V_{kk})/V_{kk}^2$, where $\nu_{ik}$ is the $k$-th coordinate of $\nu_i$, so the channel is a sum of per-moment terms, \[ (g'\otimes G')\,\omega_i^D = -\sum_{k=1}^{q} G_{k\cdot}'\,g_k\,\frac{\nu_{ik}^2-V_{kk}}{V_{kk}^2} = -\sum_{k=1}^{q} G_{k\cdot}'\,g_k\,\frac{\nu_{ik}^2}{V_{kk}^2}, \] where $G_{k\cdot}$ is the $k$th row of $G$ and the second equality uses the DWMD (population) first-order condition $G'(\operatorname{diag}V)^{-1}g=0$. Each term is quadratic in a single coordinate $\nu_{ik}$ rather than the scalar $\nu_i'Wg$ behind the OMD factorization. It neither factors through one scalar nor scales with the population overidentification criterion $g'Wg$: by keeping only the diagonal of the weight, DWMD discards the cross-moment information that the efficient weight uses.
The influence function for two-step GMM derived in Proposition (ref)(iii) has an intuitive decomposition. It consists of a direct estimation effect and an indirect effect arising from the first-step estimator \(\widehat{\phi}\). This structure allows us to formally characterize the sensitivity of the second-step estimator to the first-step estimator.
Proposition (ref) shows that the indirect component of the two-step GMM influence function arises because first-step estimation uncertainty enters through the iteration map. The matrix \(A^{-1}B\) represents the linearization of a single iteration step, as in the two-step estimator. Iterated GMM repeatedly applies the same iteration map, so the local behavior of the iterated procedure is governed by successive applications of this operator. The contraction condition in Assumption (ref)(v) is imposed on the population iteration map underlying the two-step estimator. Linearizing this map around \(\theta_0\) yields the Jacobian \(-A^{-1}B\), so the contraction condition is equivalent to requiring \(\rho(-A^{-1}B)<1\), where \(\rho(\cdot)\) denotes the spectral radius.
Under this condition, the $s$-step influence functions follow a recursion driven by repeated application of $A^{-1}B$. Applying Proposition (ref)(iii) with the $(s-1)$-step estimator as the first step, \[ \psi^{ss}(X_i) = -A_s^{-1}\big[f_{2,s}(X_i) + B_s\,\psi^{(s-1)(s-1)}(X_i)\big], \qquad \psi^{11}(X_i)=\psi^{1s,W}(X_i), \] where $f_{2,s}$ collects the moment, Jacobian, and weight-matrix channels of Proposition (ref)(iii) and $A_s$, $B_s$, $f_{2,s}$ are evaluated at the step-specific pseudo-true values $(\theta^{(s)},\theta^{(s-1)})$. \citet*{hansen2021inference} establish that $\theta^{(s)}$ converges to the fixed point $\theta_0$ geometrically; Lemma (ref) in the Appendix states this result in the form used in the proof of Proposition (ref). The next proposition shows that the $s$-step influence function converges to the iterated influence function at the same geometric rate, yielding a fixed-point interpretation of iterated GMM under misspecification.
Propositions (ref) and (ref) together provide two practical diagnostics. Proposition (ref) gives $\Lambda_\phi = -A^{-1}B$, the sensitivity of the two-step estimator's probability limit to perturbations in the first-step probability limit. A researcher using two-step GMM can compute $\Lambda_\phi$ at $\widehat\theta$ to assess how much the second-step estimate depends on the first-step choice. Proposition (ref) makes the spectral radius $\rho(-A^{-1}B)$ the rate diagnostic: the $s$-step influence function, and with it the misspecification-robust variance and informativeness, approach their iterated limits geometrically at any rate above $\rho(-A^{-1}B)$. To leading order, each weight update shrinks the remaining gap by the factor $\rho$. At the values estimated in our iterating applications, $\rho=0.61$ in AJRY and $0.64$ in BLP (Sections (ref) and (ref)), the gap falls below $10\%$ of its initial size within six and seven steps, respectively. In BPP the OMD weight is parameter-free, so $B=0$, $\rho=0$, and the iteration is trivial. Both diagnostics are computable from the same quantities used to construct misspecification-robust standard errors and require no additional estimation.
When the weight matrix \(W\) is deterministic, the influence function is given by Proposition (ref)(i): \[ \psi^{1s}(X_i) = -A^{-1} \left[ G'W\nu_i + (g'W\otimes I_p)\gamma_i \right] = M_\nu\nu_i + M_\gamma\gamma_i. \] This structure reveals how misspecification alters the sensitivity relative to the correctly specified case. The population curvature matrix \(A = G'WG + H\) now includes the term \(H = (g'W \otimes I_p)R\), which arises from the curvature of the moment function evaluated at a parameter value where the population moments are nonzero.
Under misspecification, the influence function contains the Jacobian channel $M_\gamma\gamma_i$ and, when the weight matrix is estimated, the weight-matrix channel $M_\omega\omega_i$. By Corollary (ref), these channels reduce informativeness whenever their sum has a component orthogonal to $\mathcal L(\nu_i)$ in $L^2(P)$. Both channels are proportional to the misspecification vector $g$ and vanish under correct specification. They also vanish in one-step GMM with a deterministic weight matrix when the Jacobian is deterministic across observations ($\gamma_i = 0$): the influence function reduces to $\psi^{1s}(X_i) = M_\nu\nu_i$ and $\Delta_k = 1$ for every parameter, regardless of the Hansen $J$ statistic computed at the efficient weight. Minimum distance estimation \citep*{blundell2008consumption} provides the leading instance, examined in Section (ref). The 2SLS estimator illustrates the case where the Jacobian channel is active, examined in Example (ref).
Efficient GMM estimators use a weight matrix that depends on the parameter, \(W(\theta)=\Omega(\theta)^{-1}\) with \(\Omega(\theta)=\mathbb E[g(X_i,\theta)g(X_i,\theta)']\) the uncentered second-moment matrix. In the two-step GMM estimator, the weight matrix is evaluated at a preliminary estimator \(\widehat{\phi}\), so the second-step estimator inherits uncertainty from the first step. Proposition (ref)(iii) shows that the influence function takes the form \[ \psi^{2s}(X_i) = -A^{-1}\!\left[G'W\nu_i + (g'W\otimes I_p)\gamma_i + (g'\otimes G')\omega_i \right]-A^{-1}B\psi^\phi(X_i). \] The matrix \(B\) captures the sensitivity of the optimal weight matrix to the parameter, and the term \(-A^{-1}B\,\psi^\phi(X_i)\) quantifies how first-step estimation uncertainty enters the second-step estimator. Under correct specification, \(g=0\) and hence \(B=0\), recovering the standard result that the choice of the first-step estimator does not affect asymptotic efficiency.
Iterated GMM repeatedly updates the weight matrix until the parameter used in the weight matrix coincides with the estimated parameter. As shown in Proposition (ref)(iv), this alters the effective curvature of the problem from \(A\) to \(A+B\). The estimator no longer depends on an external preliminary estimator, but sensitivity to misspecification persists in the curvature through \(B\). Continuously updating GMM, which minimizes over the parameter inside the weight matrix as well, is treated in Appendix (ref).
The following example illustrates a central implication of these results: efficient GMM estimators can share the same local sensitivity while exhibiting substantially different informativeness under misspecification.
By Corollary (ref), the differences in informativeness across estimators arise from which misspecification channels lie outside $\mathcal L(\nu_i)$. For one-step GMM with $W=I$, the only active channel is the Jacobian channel: $r_k(X_i) = M_\gamma\gamma_i$, with $\gamma_i = (0,-2U_i)'$ and $\nu_i = (U_i, U_i^2-\sigma^2)'$, where $U_i = X_i - \mu$. The Jacobian channel reduces to a multiple of $U_i$, which is the first coordinate of $\nu_i$, so it adds no information beyond the moments themselves and $\Delta = 1$. For two-step and iterated GMM, the weight-matrix channel involves the influence function of the estimated second-moment matrix, which contains centered cubic and quartic terms in $U_i$ that are not in $\mathcal L(\nu_i)$; hence $\Delta<1$.
Figure (ref) plots informativeness for the three estimators as a function of $\sigma^2$ over the range $\sigma^2 \geq 1/2$ in which all three pseudo-true values coincide with $\mu$. The one-step estimator with identity weight has $\Delta = 1$ throughout. The two efficient estimators coincide with the one-step estimator at $\sigma^2 = 1$, where the variance restriction is correctly specified, and fall below one elsewhere. Their informativeness ranking reverses across $\sigma^2 = 1$: two-step dominates iterated for $\sigma^2 > 1$, while iterated dominates two-step for $\sigma^2 < 1$.
We apply the misspecification-robust sensitivity and informativeness diagnostics to three canonical structural settings, each chosen to illustrate a different combination of the influence-function channels in Proposition (ref). BLP combines a parameter-dependent efficient weight with a large Jacobian channel; BPP isolates the weight-matrix channel in a minimum-distance setting where the Jacobian is deterministic; and AJRY likewise has a parameter-dependent weight, but a smaller cluster-level Jacobian channel, and is the setting in which we examine the iteration path and the centering of the Hansen $J$ statistic. The BLP and AJRY sections examine how iterating the weight matrix reshapes the two diagnostics.
Our first application revisits the automobile demand and supply model of \citet*{berry1995automobile} (BLP), following the implementation in \citet*{andrews2017measuring} (AGS). The purpose of this application is to illustrate how allowing for misspecified moments alters sensitivity diagnostics relative to AGS, holding the model, data, and estimation procedure fixed.
The model is identical to that in AGS. Let $X_j$ denote the observable characteristics for product $j$ in a given market, $p_j$ its observed price, and $Z_j$ the vector of instruments. The unobservables $u_j(\theta) = (\xi_j(\theta),\eta_j(\theta))'$ collect the unobserved component of product quality and the unobserved component of marginal cost. Although $\xi_j(\theta)$ and $\eta_j(\theta)$ are not directly observed, they are recovered from the model at any candidate parameter $\theta$: $\xi_j(\theta)$ is obtained by inverting the market-share equation following \citet*{berry1994estimating}, and $\eta_j(\theta)$ is obtained from the firms' Bertrand-Nash pricing first-order conditions after recovering the model-implied marginal cost $mc_j(\theta)$. The moment function $g(X_j,\theta) = Z_j \otimes u_j(\theta)$ is therefore computable at each $\theta$, even though the structural unobservables themselves are not. Stacking the demand and supply moments yields a 31-dimensional sample moment vector $\widehat g(\theta) = n^{-1}\sum_{j=1}^n g(X_j,\theta)$.
Estimation is conducted by efficient two-step GMM. The second-step estimator $\widehat\theta$ minimizes \[ \widehat Q(\theta) = \widehat g(\theta)' \widehat W(\widehat\phi)\, \widehat g(\theta), \] where $\widehat W(\widehat\phi) = \bigl\{n^{-1}\sum_j g(X_j,\widehat\phi)g(X_j,\widehat\phi)'\bigr\}^{-1}$ is the estimated efficient weight matrix and $\widehat\phi$ is the first-step GMM estimator with identity weight matrix.
The parameter of interest is the average markup, \[ c(\theta) = \frac{1}{n}\sum_j \frac{p_j - mc_j(\theta)}{p_j}, \] where $p_j$ is the observed price and $mc_j(\theta)$ is the model-implied marginal cost recovered from the supply-side first-order conditions. The influence function of $c(\theta)$ follows from the chain rule: $\partial c(\theta)/\partial\theta'$ multiplied by the influence function of the GMM estimator.
AGS study sensitivity under local violations of instrument validity, modeled as direct effects of instruments entering demand or cost equations at rate \(1/\sqrt n\). We compute the corresponding normalized sensitivity coefficients using both AGS and our MRS measure for the same excluded instruments. The AGS sensitivities are taken directly from their replication files, while MRS is computed using the misspecification-robust influence function derived in Proposition (ref)(iii).
Figure (ref) compares the two sensitivity measures. The attenuation under MRS is most pronounced for the demand-side instruments. In the AGS calculation, several demand-side instruments have sizable positive sensitivity for the average markup. Under MRS, these demand-side sensitivities are close to zero and in some cases change sign. On the supply side, the comparison is more heterogeneous: MRS often reduces sensitivity, but some supply-side moments remain influential and a number of rankings change. The main empirical finding is that allowing for misspecification substantially reorders the sensitivity profile. Table (ref) revisits the asymptotic bias calculations in AGS using MRS, showing that violations related to economies of scope and to cross-firm demand spillovers have substantially smaller effects under MRS relative to AGS.
The Hansen $J$ statistic is $345$ on $25$ degrees of freedom with $p<0.001$, decisively rejecting the moment conditions. By Corollary (ref), $\widehat\Delta_{markup}<1$ reflects the part of the influence function outside the moment span $\mathcal L(\nu_i)$. Both channels are active in BLP: the Jacobian channel $M_\gamma\gamma_i$ because $G(X_j,\theta)$ varies non-trivially across products through the nonlinear demand system, and the weight-matrix channel $M_\omega\omega_i$ because $\widehat W$ depends on the first-step $\widehat\phi$. Their joint contribution gives $\widehat\Delta_{markup}=0.56$:\footnote{The MRS-based quantities for BLP are computed in MATLAB R2026a; under R2024a the average-markup informativeness is $0.58$ in place of $0.56$, a second-decimal toolchain difference with no effect on any conclusion.} even after conditioning on the full vector of moments, $44\%$ of the asymptotic variance of the markup estimator is unexplained by the linear projection on the moments. The $J$-test detects misspecification; $\Delta$ quantifies the share of structural efficiency it costs.
For completeness, we also compute the iterated GMM analogue using Proposition (ref)(iv). The spectral radius $\rho(-A^{-1}B)\approx 0.64$ remains below one along the iteration path, so the iterated estimator is locally well-defined; the iteration converges within a small number of weight updates, leaving the Hansen $J$ essentially unchanged at $\approx 345$, while the iterated $\widehat\Delta_{markup}=0.24$ falls well below the two-step value of $0.56$.
Our second application revisits the model of household income dynamics and consumption insurance of \citet*{blundell2008consumption} (BPP). BPP estimate the model by diagonally weighted minimum distance (DWMD); we report both that scheme and the optimal weight matrix, and the contrast between the two is itself diagnostic.
The BPP model also provides a clean illustration of Proposition (ref) and Corollary (ref). By construction, the Jacobian of the MD estimator is deterministic across households, so $\gamma_i = 0$ and the Jacobian channel vanishes. Under the optimal weight the total misspecification residual reduces to the weight-matrix channel $r_k(X_i) = M_\omega\omega_i$, the rank-one factorization of Proposition (ref), which generates structural-efficiency loss whenever this term has a nonzero component orthogonal to $\mathcal L(\nu_i)$. We estimate the BPP model with time-varying shock variances, comprising $35$ parameters and $325$ second-moment conditions. Even when the Jacobian contributes no sampling variation, the weight-matrix channel can drive informativeness substantially below one.
The BPP model decomposes log income into a permanent and a transitory component, with consumption growth depending on permanent and transitory income shocks through partial-insurance coefficients $\phi$ and $\psi$. The full structural equations and the construction of the non-collinear moment set are given in Appendix (ref). We estimate the specification with time-varying variances of the permanent shock, the transitory shock, and the consumption innovation, holding $\phi$ and $\psi$ constant across periods; the parameter vector $\beta$ collects these time-varying variances along with $\phi$, $\psi$, the MA coefficient, and the measurement-error variance. Let $m_i$ denote the household-level vector of sample cross-products of income and consumption growth,\footnote{The i.i.d. assumption is imposed at the household level. Although the moments use within-household time-series covariances, these are collected into the household-level vector $m_i$, and within-household serial dependence is absorbed into $m_i$. The asymptotic approximation treats households as independently sampled.} and let $m(\beta)$ denote the corresponding theoretical covariances. The sample moment conditions are $g(X_i, \beta) = m_i - m(\beta)$.
The OMD estimator minimizes $\widehat g(\beta)'\widehat W\widehat g(\beta)$ with $\widehat W = [n^{-1}\sum_i (m_i - \bar m)(m_i - \bar m)']^{-1}$ where $\bar{m}=n^{-1}\sum_{i}m_{i}$, the inverse of the sample covariance matrix of the data moments.\footnote{Our construction of this sample covariance follows \citet*[Appendix D]{blundell2008consumption}, except that for unbalanced panels each entry is divided by the product of the two moments' observation counts rather than by the number of households for whom both moments are observed. The two coincide for a balanced panel. For the cohort-A subsample BPP's normalization produces a covariance estimate that is not positive-definite and so cannot be used to form the OMD or DWMD weight matrix; ours fixes this with no other change to the construction.} Because $\widehat W$ is built from the data moments alone, it does not depend on $\beta$: OMD is one-step GMM with an estimated weight matrix, its influence function is given by Proposition (ref), and there is nothing to iterate.\footnote{In the unbalanced panel the per-moment subsamples leave a residual dependence of the weight on $\beta$ through the missing-data pattern, so the formally iterated estimator is not exactly the one-step; re-estimating with the updated weight converges in three steps with negligible movement in the estimates.}
We also estimate the model by DWMD, the scheme used by BPP, with weight $\widehat D = (\operatorname{diag}\widehat V)^{-1}$, weighting each moment by the inverse of its own sampling variance, and by EWMD, the equal-weighting scheme of \citet*{altonji1996small}, with the identity weight. The three estimators share the same moments and model map and differ only in how much of the weight is estimated: all of $\widehat V^{-1}$ for OMD, its diagonal for DWMD, none for EWMD.
Parameter estimates and standard errors are reported in Table (ref). The Hansen $J$ statistic for overidentifying restrictions is $538.7$ on $290$ degrees of freedom with $p<0.001$, indicating that the BPP moment conditions are decisively rejected.
Table (ref) also reports the misspecification-robust informativeness. Under OMD, $\widehat\Delta = 0.79$ for both $\phi$ and $\psi$; across all $35$ parameters $\widehat\Delta_k$ ranges from $0.70$ to $0.87$ with median $0.79$. The misspecification-robust standard errors exceed the conventional ones ($0.043$ versus $0.034$ for $\phi$); the gap is exactly the weight-matrix channel that the conventional formula omits.
The results under DWMD are markedly different. The partial-insurance coefficients are estimated at $\widehat\phi=0.68$ and $\widehat\psi=0.03$, close to the values reported by BPP and well away from the OMD estimates of $0.33$ and $0.07$; EWMD gives $\widehat\phi=0.61$ and $\widehat\psi=0.00$. Since the three weightings share a probability limit under correct specification, this divergence is itself a symptom of misspecification. The DWMD informativeness, however, is $\widehat\Delta\approx0.99$ for $\phi$, for $\psi$, and for every one of the $35$ parameters. Figure (ref) plots the full distribution under both weightings: the OMD values are scattered between $0.70$ and $0.87$, while the DWMD values cluster just below one.
The contrast is structural. OMD's weight inverts the sample covariance of the moments. By Proposition (ref), once the overidentifying restrictions are rejected, this construction makes the estimated weight contribute variation that the moments cannot account for, and the contribution grows with the strength of the rejection. With the BPP moments decisively rejected, roughly a fifth of the asymptotic variance of $\widehat\phi$ and $\widehat\psi$ reflects estimation of the weight matrix rather than the moments. DWMD keeps only the diagonal of that weight and discards the cross-moment terms, so it carries no comparable contribution; its variance stays almost entirely explained by the moments, and $\widehat\Delta\approx0.99$. EWMD is the limiting case: the identity weight estimates nothing, so by Proposition (ref) the weight-matrix channel is absent and $\Delta_k=1$ for every parameter, which the numerical computation reproduces exactly.
The simpler weights are not free: the misspecification-robust standard error of $\phi$ rises from $0.043$ under OMD to $0.113$ under DWMD and $0.109$ under EWMD, the efficiency-informativeness trade of Remark (ref). The loss of informativeness in BPP is thus a feature of the optimal weight matrix specifically, not of weighted minimum distance in general.
Our third application revisits the income-and-democracy regressions of \citet*{acemoglu2008income} (AJRY), a leading example of dynamic panel difference-GMM. AJRY estimate the autoregressive specification \[ d_{it} = \alpha\,d_{i,t-1} + \gamma\,y_{i,t-1} + \mu_t + \delta_i + u_{it}, \] where $d_{it}$ is a measure of democracy, $y_{i,t-1}$ is lagged log income per capita, $\mu_t$ are time effects, and $\delta_i$ are country fixed effects. First-differencing eliminates $\delta_i$, and the resulting moment conditions $\mathbb E[Z_i\Delta u_{i}] = 0$ use Arellano-Bond instruments built from lagged levels of $d$ and $y$. Because the data form a country panel, the cross-sectional unit is the country and within-country serial dependence is absorbed into the stacked cluster moment $Z_i'\Delta u_i$. The influence functions and standard errors are cluster-robust, computed at the country level as in Appendix (ref). The parameter of interest is $\gamma$, the effect of income on democracy.
We replicate the five-year panel specification of \citet*{hansen2021inference}, which reports the iterated GMM estimator of AJRY with misspecification-robust standard errors. The iterated estimate is $\widehat\alpha = 0.744$ and $\widehat\gamma = -0.009$, with misspecification-robust standard errors $0.128$ and $0.039$ respectively. The Hansen $J$ statistic at the iterated fixed point is $45.3$ on $44$ degrees of freedom ($p=0.42$), failing to reject at conventional levels.\footnote{The Hansen $J$ statistic is $J = n\,\widehat g'\widehat W\widehat g$, with $\widehat g$ the mean cluster moment vector and $\widehat W$ the inverse of the uncentered cluster second-moment matrix $n^{-1}\sum_i(Z_i'\Delta \widehat{u}_i)(Z_i'\Delta \widehat{u}_i)'$, evaluated at the relevant parameter estimate.} The one-step estimator that AJRY report uses the deterministic Arellano-Bond weight matrix and yields a Hansen $J$ of $71.3$ ($p=0.006$). The two-step efficient estimator yields $\widehat\gamma = -0.012$ with a robust standard error of $0.049$ and a Hansen $J$ of $61.4$ ($p=0.04$). All three estimators find the income coefficient statistically insignificant, but they deliver materially different specification-test verdicts.
AJRY report a one-step estimate of $\gamma$ that is statistically insignificant, concluding that income has no causal effect on democracy. Its misspecification-robust informativeness is $\widehat\Delta_\gamma = 0.93$. The Arellano-Bond weight does not depend on the parameter, so the one-step estimator activates the Jacobian channel but no weight-matrix channel; the modest loss is the Jacobian channel alone. The Jacobian channel is active because the per-unit Jacobian $G(X_i,\theta)$ varies across countries through the lagged democracy and income terms in $Z_i'\Delta X_i$. Adopting the efficient weight matrix lowers informativeness rather than raising it. The two-step efficient estimator has $\widehat\Delta_\gamma = 0.81$ and the iterated estimator $\widehat\Delta_\gamma = 0.77$. Figure (ref) plots the $J$ statistic and $\widehat\Delta_\gamma$ along the iteration path, and the two diagnostics give opposite verdicts on iteration. Iterating the weight matrix moves the overidentifying-restrictions test from rejection at the one-step and two-step estimators ($p=0.006$ and $p=0.04$) to non-rejection at the fixed point ($p=0.42$), while $\widehat\Delta_\gamma$ falls in parallel from $0.93$ to $0.77$.
The non-rejection, however, depends on how the weight matrix is centered. The statistic of footnote (ref) uses the uncentered second-moment matrix; with the centered covariance, the two statistics at the same parameter value are linked by $J = J_c/(1+J_c/n)$, by the Sherman--Morrison formula. Under correct specification the two are asymptotically equivalent and either yields a valid $\chi^2$ test, so the centering is usually treated as innocuous. Under misspecification they diverge by a first-order amount, and the centered version is the more powerful test \citep*{hall2000covariance}. Figure (ref) plots the two statistics along their respective iteration paths. The centered and uncentered iterations converge to the same fixed-point estimator, where the uncentered statistic is below the critical value ($45.3$, $p=0.42$) and the centered statistic far above it ($70.4$, $p=0.007$); the centered test rejects at every step of its path. A specification verdict that flips with the centering convention is not a reliable reading of misspecification. The informativeness, by contrast, is invariant to the convention: the centered and uncentered weights produce the same iterated estimator and influence function, so $\widehat\Delta_\gamma=0.77$ either way, and the structural-efficiency loss is a property of the estimator rather than of the test.
Figure (ref) traces the efficient-weight iteration from five first-step weights. Panel (a) reproduces Figure 1(b) of \citet*{hansen2021inference}: the income coefficient converges to a common fixed point from every start, consistent with the contraction rate $\rho(-A^{-1}B)\approx 0.61$ below one. The informativeness likewise reaches the common fixed-point value $\widehat\Delta_{\gamma}=0.77$ from every start (panel b). Because the $s$-step estimators along the path converge to distinct pseudo-true values, their step-wise informativeness values are not directly comparable, but the panel shows that the approach to the fixed point is non-monotone.
Iteration affects the two diagnostics differently across the applications. In AJRY both move, the test passing to non-rejection while $\widehat\Delta_\gamma$ falls from $0.81$ to $0.77$. In BLP, by contrast, iteration leaves the Hansen $J$ essentially unchanged at $\approx 345$ while $\widehat\Delta_{markup}$ falls sharply from $0.56$ to $0.24$ (Section (ref)). $J$ and $\widehat\Delta_k$ measure different functionals of the same misspecification: $J$ records how far the moments depart from zero at the pseudo-true value, and $\widehat\Delta_k$ the share of the estimator's variance the moments explain, so iteration reshapes the Jacobian and weight-matrix channels, large in BLP and modest in AJRY, without moving $J$ in step.
This paper develops a sensitivity and informativeness framework for GMM estimators that remains valid under general moment misspecification. By expressing sensitivity measures through influence functions, we provide a unified characterization of how deviations in moment conditions affect GMM estimators across one-step, two-step, and iterated procedures, with continuously updating GMM treated in Appendix (ref). Once the misspecification-robust influence function is in hand, the sensitivity matrix $\Lambda$ and the informativeness measure $\Delta_{k}$ follow at essentially no additional cost.
The informativeness measure $\Delta_{k}$ is the diagnostic contribution. It quantifies the share of an estimator's asymptotic variance driven by sampling variation in the moments, as opposed to variation arising from the Jacobian or the estimated weight matrix, and is complementary to the Hansen $J$-test: the $J$-test asks whether the moment restrictions are jointly violated; $\Delta_{k}$ asks how much of the estimator's variance is left unexplained by the moments after such a violation. The income-and-democracy application illustrates the complementarity. At the iterated fixed point the $J$-test verdict depends on the centering of the weight matrix ($p=0.42$ uncentered, $p=0.007$ centered), while $\widehat\Delta_\gamma = 0.77$ under either convention, so nearly a quarter of the income coefficient's asymptotic variance is unexplained by the moments regardless of how the test is run.
A practical implication concerns the choice of weight matrix. In minimum distance estimation, the optimal weight that delivers classical efficiency under correct specification carries a hidden cost under misspecification: a loss of informativeness, or structural efficiency, that is absent under correct specification. By Proposition (ref), the very construction that attains this efficiency, the inversion of the moments' second-moment matrix, is the one that makes the estimated weight contribute variation the moments cannot account for once the overidentifying restrictions are rejected, and this informativeness loss grows with the strength of the rejection. Simpler weights avoid it. Diagonally weighted minimum distance, which inverts only the moments' marginal variances, raises $\widehat\Delta$ for all $35$ parameters of our BPP application from a median of $0.79$ to about $0.99$, and equal weighting, which estimates no weight at all, attains $\Delta_{k}=1$ exactly. Pursuing classical efficiency under misspecification therefore trades informativeness for an efficiency gain the optimal weight no longer delivers.
We recommend that practitioners using over-identified GMM report $\Delta_{k}$ alongside misspecification-robust standard errors. Where $\Delta_{k}$ is well below one, the moments explain only a small share of the estimator's asymptotic variance, and any economic interpretation of the parameter should be made with that in mind. The framework does not correct for misspecification; it clarifies how and through which channels misspecification enters the estimator.