EconBase
← Back to paper

Repairing Locally Misspecified GMM: An Empirical Bayes Approach

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

156,528 characters

Repairing Locally Misspecified GMM: An Empirical Bayes Approach


\title{Repairing Locally Misspecified GMM: An Empirical Bayes Approach}
\author{Patrick Kline\thanks{This manuscript was prepared for the 2026 Sargan lecture. I am grateful
to Isaiah Andrews, Kevin Chen, Andres Santos, and Jinglin Yang for
helpful feedback on earlier versions of this draft. Luan Borelli, Claude Opus, and Refine.ink provided outstanding research assistance on this project.}\\
UC Berkeley}
\maketitle
\begin{abstract}
Econometric models offer parsimonious but inexact approximations to
data-generating processes. This paper studies the generalized method
of moments (GMM) when exchangeable specification errors of order $n^{-1/2}$
contaminate the moment conditions. I develop estimators for the mean
and variance of these specification errors, establishing their consistency
in an asymptotic framework where the number of overidentifying restrictions
grows with the sample size. These hyperparameter estimates are used
to develop a feasible bias-corrected estimator of target
parameters. I also propose an empirical Bayes estimator that weakly improves precision by subtracting a
best linear predictor of the first-order estimation error from the bias-corrected estimator.
Using a combinatorial central limit theorem, I establish asymptotic normality of both estimators and provide
variance estimators that enable misspecification-aware frequentist
inference. Simulation exercises indicate the procedures can meaningfully
improve on standard two-stage least squares estimation when exclusion
violations are present. Revisiting the influential study of \textcite{angrist1991does}, I
consider an instrument set where exchangeable excludability violations are plausible.
Repairing the two-stage least squares estimates of the returns to
schooling moves them in the direction of ordinary least squares and
reduces sensitivity to the specification of controls.
\end{abstract}
\newpage{}

The generalized method of moments (GMM) is a widely used technique
for fitting parsimonious semi-parametric models to data. The procedure
consists of specifying a function $g:\mathbb{R}^{p}\to\mathbb{R}^{m}$
whose zeros define $m$ moment restrictions on a parameter vector
$\theta\in\mathbb{R}^{p}$. A distinguishing feature of GMM, relative
to classical method of moments estimators, is that the number of moment
restrictions can exceed the number of parameters $\left(m\geq p\right)$.
The case where $m>p$ is known as an overidentified model \parencite{anderson1949estimation}.
In practice, GMM models often contain many overidentifying restrictions.
For instance, \textcite{angrist1991does} consider a specification
where $m-p=179$. The primary specification in \textcite{dellavigna2012testing}
has $m-p=55$, while \textcite{gourinchas2002consumption} consider
a specification with $m-p=36$.

Foundational work by \textcite{sargan1958estimation} and \textcite{hansen1982large}
developed a framework for testing these overidentifying restrictions
by comparing the  $J$-statistic to its limiting distribution under
proper specification. In practice, overidentifying restrictions are
often rejected. Such rejections are typically met with a degree of
ambivalence, as econometric models are ultimately approximations.\footnote{This ambivalence is not unique to the  $J$-test. In \textcite{hansen_sargent_2014_uncertainty},
Thomas Sargent recalls that ``Our good friend Robert E. Lucas, Jr.
told us in the early 1980s that our likelihood ratio tests and moment
matching tests were rejecting too many good models.'' \textcite{andrews2026misspecification} provide a decision-theoretic approach to dealing with potentially misspecified likelihood functions.} However, the nature of these approximation errors is rarely quantified
precisely. Statistical model rejections are also difficult
to explain to journal editors and referees, which may explain why
$J$-statistics are rarely reported in modern empirical work \parencite{andrews2025purpose}.

This paper explores how to use the information in overidentifying
restrictions to build improved estimates of the model parameters  $\theta$
that account for specification error. Following \textcite{andrews2017measuring},
I work within a local misspecification framework in which the $m$
moment conditions are perturbed by errors that shrink with the sample
size. Unlike past work on local misspecification, I treat these errors as exchangeable random
variables. I propose estimators of the mean and variance of the specification
errors, the latter of which captures overdispersion in the  $J$-statistic.
I then establish consistency of these hyperparameter estimates in
an asymptotic regime where both $m$ and $n$ grow large but
$m^{2}/n\rightarrow0$, which rules out settings where
many-instruments biases become first order \parencite{altonji1996small,newey2004higher}.
This moderately overidentified regime arguably captures a large share
of settings where GMM is employed.

I use the proposed hyperparameter estimates to develop two improved
estimators of structural parameters.
The first estimator is a plug-in bias correction that subtracts an estimate of the
GMM estimator's first-order misspecification bias. In the special case of linear instrumental variables models, this correction is closely related to ``visual IV'' diagnostics \parencite{angrist2009mostly}, wherein reduced forms are regressed on first stages. The bias-corrected estimator allows for a non-zero intercept in this relationship, which effectively purges the fitted slope of an omitted variables bias.

The second estimator is an empirical Bayes shrinkage
adjustment that subtracts a best linear predictor of the first-order
estimation error from the bias-corrected estimator. I show that the shrinkage adjustment weakly reduces leading-term estimation risk relative to bias correction by removing the predictable structure of the specification errors. This finding closely parallels the inadmissibility result of \textcite{brown1990ancillarity}, who established that adjusting for a high-dimensional nuisance regressor
can lower risk for a low-dimensional target.

Leveraging the exchangeable structure of the specification errors,
I establish asymptotic normality of both estimators via a combinatorial central limit theorem. I use these results to develop consistent standard error
estimators that enable \emph{misspecification-aware} frequentist inference. Confidence intervals based on these standard errors attain asymptotically valid unconditional coverage (i.e., averaged over the unknown distribution of the specification errors).

A few papers take related approaches. One is \textcite{andrews2024true}
who propose an approach to inference on structural parameters when priors over moment misspecification are rotation invariant.
The distributional assumptions on specification errors considered
here nest rotation-invariance but require bounded fourth moments.
In particular, rotation-invariant priors must have mean zero, a restriction
I relax. However, \textcite{andrews2024true} deliver fixed-$m$
Bayesian inference while the frequentist misspecification-aware inferences
proposed here require that $m$ grows with $n$.

A second related contribution is \textcite{chernozhukov2025plausible} who
propose a quasi-Bayes approach to GMM in which priors are specified
over both the structural parameters and the specification errors.
As in this paper, they consider an environment in which $m$ grows
with $n$. While their Bayesian calibration and coverage statements
are contingent on an assumed prior over misspecification (e.g., Gaussianity),
I rely on exchangeability of the specification errors, alongside standard moment and regularity conditions, for estimation and inference.

Finally, earlier work by \textcite{kolesar2015identification} studies linear instrumental variables models featuring excludability violations in an environment where the number of instruments can grow with the sample size. They propose a bias-corrected k-class estimator predicated on these violations being orthogonal to the first-stage coefficients. The exchangeability assumption I invoke implies non-correlation in expectation but not almost surely. While their approach relies on homoscedastic errors in the first-stage and structural equations, I work in a more general GMM framework that accommodates both nonlinear moment conditions and heteroscedasticity.

Though the framework developed here offers new opportunities for misspecification-aware
estimation and inference, the methods proposed are not entirely automatic.
Moment conditions need to be scaled carefully to ensure violations
of the nominal model are plausibly exchangeable. These scaling decisions
are inherently context dependent, requiring commitment to a particular
model of misspecification. This limitation reflects the general observation
that misspecification can only be studied with respect to larger encompassing
models that plausibly hold in the setting under study \parencite{armstrong2025misspecification}. I develop diagnostics for detecting departures from exchangeability that can aid
in assessing the suitability of a given encompassing model.

I illustrate the methods in a series of simulation exercises that
consider exclusion failures in linear IV specifications. Concerns about excludability are often heightened in applied work using multiple instruments. For example, \textcite{angrist2017leveraging,angrist2024credible} considered a parametric empirical Bayes approach to remedying such failures.
The simulation design I study departs from the usual i.i.d.\ assumptions
employed in empirical Bayes modeling by considering specification
errors that are exchangeable but not independent. I find that both
of the proposed corrections can yield sizable risk improvements even when
the degree of overidentification is modest. To assess the theory's
asymptotic predictions, I simulate a setting where the number of moments
grows with the sample size. The simulations confirm that tests based
on the proposed misspecification-aware variance estimators have correct
asymptotic size.

I then revisit the influential study of \textcite{angrist1991does},
examining schooling instruments based on season of birth by state
of birth interactions that plausibly satisfy exchangeability. I find
that moment conditions based on these instruments exhibit substantial
mean bias. The changing direction of the bias across control sets
is consistent with exogeneity failures of the sort posited by \textcite{rosenzweig2000natural}
and \textcite{buckles2013season}. The corrections consistently shift
the two-stage least squares (TSLS) estimates in the direction of ordinary
least squares (OLS), substantially narrowing the dispersion of point
estimates across alternative sets of controls. The corrected estimates
also exhibit larger standard errors than TSLS, reflecting the misspecification-aware
nature of the proposed variance estimators.

The rest of the paper is organized as follows. Section \ref{sec:GMM-preliminaries} reviews the basics of GMM under
proper specification and under local misspecification. Section \ref{sec:local_misspecification}
introduces the exchangeable misspecification framework and Section \ref{sec:Examples} discusses
some examples of models that fall into this framework.
Section \ref{sec:Identification-of-hyperparameter} studies identification
of the mean and variance of the specification errors. Section \ref{sec:A-differencing-approach}
discusses a GMM estimator based on differenced moment conditions,
which serves as a pedagogical bridge to the empirical Bayes estimators
of interest. Section \ref{sec:estimation} proposes estimators of
the mean and variance of the specification error distribution and
establishes their consistency. Section \ref{sec:bias-estimation}
proposes a plug-in bias-corrected estimator and an empirical Bayes
shrinkage estimator of target parameters that account for the specification
errors in different ways. Section \ref{sec:distribution-theory} develops
distribution theory for these estimators enabling misspecification-aware
inference. Section \ref{sec:monte-carlo} evaluates the finite sample
performance of the estimators in a series of simulation exercises.
Section \ref{sec:empirical} revisits the study of \textcite{angrist1991does}.
Section \ref{sec:Conclusion} concludes with directions for future
work.

\section{GMM preliminaries \protect\label{sec:GMM-preliminaries}}

In this section I briefly review GMM estimation theory and state some standard
regularity conditions that are maintained throughout the analysis.
Suppose there exists a unique parameter vector $\theta_{0}\in\Theta$
obeying $g\left(\theta_{0}\right)=0$, where the vector $g:\mathbb{R}^{p}\rightarrow\mathbb{R}^{m}$ of population moment
conditions is continuously
differentiable. Assume the parameter space $\Theta$ is compact and
convex and that the true $\theta_{0}$ lies strictly in its interior.
Let the $m\times p$ matrix  $G$ denote the Jacobian $\nabla_{\theta}g\left(\theta\right)$
evaluated at  $\theta_{0}$ and
assume that $G$ has rank $p\leq m$. I will focus on the overidentified
case, assuming $m-p\geq2$. The smoothness and rank conditions are maintained
throughout my analysis. However, the nominal moment condition $g(\theta_{0})=0$
will be relaxed by allowing for local misspecification.

GMM estimates are obtained by constructing empirical moment conditions
$\hat{g}\left(\theta\right)$ that converge in probability to $g\left(\theta\right)$.
The GMM estimator $\hat{\theta}$ minimizes the quadratic form  $\hat{g}\left(\theta\right)'\hat{W}\hat{g}\left(\theta\right)$,
where $\hat{W}$ is some positive definite weighting matrix that converges
in probability to a nonsingular $m\times m$ matrix $W$. I assume
that $G'WG$ is non-singular. Letting $\hat{G}$ denote the Jacobian
of $\hat{g}\left(\theta\right)$ evaluated at $\hat{\theta}$, the
GMM estimator satisfies the first order condition $\hat{G}'\hat{W}\hat{g}\left(\hat{\theta}\right)=0.$

\subsection{Influence and sensitivity \protect\label{subsec:Influence-and-sensitivity}}

Suppose that the scaled empirical moments, when evaluated at the true
parameter vector  $\theta_{0}$, converge to normals:
\[
\sqrt{n}\hat{g}\left(\theta_{0}\right)\overset{d}{\rightarrow}N\left(0,V\right).
\]
Standard arguments \parencite[e.g.,][Theorem 3.4]{newey1994large}
yield the influence function representation of GMM estimators
\[
\sqrt{n}\left(\hat{\theta}-\theta_{0}\right)=-\left(G'WG\right)^{-1}G'W\sqrt{n}\hat{g}\left(\theta_{0}\right)+o_{p}\left(1\right).
\]
Hence, the GMM estimator is, to first order, a linear combination
of normals. \textcite{andrews2017measuring} term these combination
weights the ``sensitivity matrix''
\[
\Lambda:=-\left(G'WG\right)^{-1}G'W,
\]
which summarizes the local influence of the empirical moment conditions
on the parameter estimates. The limiting distribution of GMM under
proper specification can be written in terms of the sensitivity matrix
as
\[
\sqrt{n}\left(\hat{\theta}-\theta_{0}\right)\overset{d}{\rightarrow}N\left(0,\Lambda V\Lambda'\right).
\]

Note that premultiplying a vector of moments by  $-\Lambda$ amounts
to a weighted least squares regression using the columns of the Jacobian
 $G$ as regressors. It will be useful to consider the corresponding
``hat'' matrix $H:=-G\Lambda.$ Premultiplying a vector of moments
by $H$ gives predicted values of the moment conditions generated
from a generalized least squares (GLS) fit. $H$ is an \textit{oblique}
projection matrix: it is idempotent but not generally symmetric. The
complementary residual maker matrix, which features prominently in
later sections, will be denoted $M:=I_{m}-H$. The sample analogs
of $\Lambda$,  $H$, and $M$ will be denoted by $\hat{\Lambda}=-\left(\hat{G}'\hat{W}\hat{G}\right)^{-1}\hat{G}'\hat{W}$,
 $\hat{H}=-\hat{G}\hat{\Lambda}$, and $\hat{M}=I_{m}-\hat{H}$ respectively.

\subsection{Local and global misspecification\protect\label{subsec:Misspecification-as-contaminatio}}

Describing misspecification requires enlarging the set of data generating
processes (DGPs) under consideration to include those that violate
the researcher's posited moment conditions. Fix a parameter vector
$\theta_{0}\in\Theta$ and denote by $\mathcal{P}_{0}$ the family
of data generating processes under which $g\left(\theta_{0}\right)=0$.
Note that for any DGP $P_{0}\in\mathcal{P}_{0}$ one can think of
the parameter vector $\theta_{0}$ as a functional $\theta_{0}=\theta\left(P_{0}\right)$.
Misspecification arises when the actual DGP lies outside $\mathcal{P}_{0}$,
in which case $g\left(\theta_{0}\right)\neq0$.

\textcite{andrews2017measuring} consider a family of DGPs that lie within an $n^{-1/2}$ neighborhood of $\mathcal{P}_{0}$. Under this notion of \textit{local
misspecification}, the moment conditions obey $g\left(\theta_{0}\right)=b/\sqrt{n}$ for a fixed vector $b\in\mathbb{R}^{m}$. They establish, under regularity conditions analogous to those in Section \ref{subsec:Influence-and-sensitivity},
that this environment yields
\[
\sqrt{n}\left(\hat{\theta}-\theta_{0}\right)\overset{d}{\rightarrow}N\left(\Lambda b,\Lambda V\Lambda'\right).
\]
Hence, the GMM estimator exhibits first order bias $\Lambda b$ but its asymptotic variance is unchanged.
If the violations were to shrink at a slower rate (e.g., $g\left(\theta_{0}\right)=b/n^{1/4}$),
the scaled bias would explode relative to the variance, precluding
convergence in distribution. Thus, local misspecification is a tool
for studying specification errors that generate modest first order
biases in GMM comparable in magnitude to sampling variation. This
sort of misspecification is difficult to detect with a standard
$J$-test: power does not grow with sample size.

In contrast, under \textit{global misspecification}, the violations do not shrink
with the sample size. Formally, $g\left(\theta_{0}\right)=b$ for a fixed nonzero
vector $b$. In this case, GMM is generally inconsistent for $\theta_{0}$,
converging instead to a pseudo-true value that depends on the weighting
matrix. Likewise, the $J$-test generally has power that grows with the sample size.

Since GMM's leading bias of $\Lambda b/\sqrt{n}$ under local misspecification
vanishes as $n$ grows large, it remains consistent for the true parameter
vector of interest $\theta_{0}$. Thus, the local misspecification
framework offers the potential to conduct inference on the true parameter
$\theta_{0}$, as opposed to a pseudo-true target, by augmenting
measures of sampling uncertainty to account for specification errors.

\section{An exchangeable model of local misspecification \protect\label{sec:local_misspecification}}

Building on the framework in \textcite{andrews2017measuring}, I allow
the moment conditions to be perturbed by a
vector of specification errors that shrinks with the sample size.
My key departure will be to view the specification errors as exchangeable random variables rather than fixed constants. Specifically, I assume $g\left(\theta_{0}\right)=b/\sqrt{n}$ for a
random vector $b\in\mathbb{R}^{m}$. The scaled sample moment conditions then
decompose as
\[
r_{n}:=\sqrt{n}\hat{g}\left(\theta_{0}\right)=b+\varepsilon_{n},
\]
where $\varepsilon_{n}:=\sqrt{n}\left(\hat{g}\left(\theta_{0}\right)-g\left(\theta_{0}\right)\right)$
captures noise in the moment conditions attributable to sampling variation.

The vector $b$ encodes the population-level structure of misspecification
across moment conditions. Exchangeability implies that all entries $b_j$ of $b$ have the same marginal distribution and that all pairs $(b_j, b_k)$,
for $j \neq k$, have the same joint distribution. This is a reasonable perspective
to take when the moment conditions share the same basic form ---
i.e., they are all measured in the same units and involve averages
over similar mathematical objects --- and one does not know which
moment conditions are likely to be misspecified a priori.

The following assumption formalizes this perspective and adds regularity conditions used in later sections. In what follows, $\Vert \cdot \Vert$ denotes the Euclidean norm.
\begin{assumption}[Independence and exchangeability]
\label{assu:random_b}The pair $\left(b,\varepsilon_{n}\right)$
obeys the following conditions:

i) \textup{$b\perp\varepsilon_{n}$ for each $n$, $\Vert E[\varepsilon_{n}]\Vert\to0$, and $\varepsilon_{n}\overset{d}{\rightarrow}\varepsilon\sim N\left(0,V\right)$,}

ii) \textup{the specification errors $b=\left(b_{1},\dots,b_{m}\right)'$ are exchangeable with bounded fourth moments.}
\end{assumption}
Assumption \ref{assu:random_b}.i imposes two conditions on the relationship between
specification errors and the moment-condition noise. First, $b$ and
$\varepsilon_{n}$ are independent at every sample size. This is a
substantive economic restriction that would be violated if, for example,
a data-dependent moment selection rule induced correlation between
$b$ and $\varepsilon_{n}$. Second, the moment-condition noise is asymptotically unbiased and normal with covariance $V$.

Assumption \ref{assu:random_b}.ii is the key exchangeability restriction. The bounded fourth moments assumption will prove useful for the asymptotic analysis of later sections. Together, these conditions ensure that the specification errors share a common marginal mean $\bar \mu:=E[b_j]$ and variance $\bar \sigma^2 := \mathrm{Var}(b_j)$, both of which are finite.

\subsection{Defining hyperparameters}
The marginal mean $\bar \mu$ and variance $\bar \sigma^2$ of the specification errors describe the properties of a hypothetical superpopulation from which the specification errors were sampled. Since the goal of this analysis will be to improve GMM estimates given the specification errors actually faced by the researcher, I will condition on the mean and variance of the realized specification errors.

Denote the realized mean and variance of the specification errors by
\[
\mu:=\frac{1}{m}\sum_{j=1}^{m}b_{j},\qquad \sigma^{2}:=\frac{1}{m-1}\sum_{j=1}^{m}\left(b_{j}-\mu\right)^{2}.
\]
In subsequent sections, I treat these objects as hyperparameters to be targeted in estimation. The following lemma derives the first two moments of the specification errors $b$ conditional on these hyperparameters.
\begin{lem}
\label{lem:cond-moments}Assumption \ref{assu:random_b}.ii implies that:
\[
E\left[b\mid\mu,\sigma^{2}\right]=\mu\,1_{m},\qquad \mathrm{Var}\left[b\mid\mu,\sigma^{2}\right]=\sigma^{2}Q,\qquad Q:=I_{m}-\frac{1}{m}1_{m}1_{m}'.
\]
\end{lem}
\begin{proof}
Since $\mu$ and $\sigma^{2}$ are symmetric functions of $b$, conditioning on them preserves exchangeability. Write $b=\mu1_{m}+Qb$ and note that $Qb=b-\mu1_{m}$ is conditionally exchangeable and orthogonal to $1_{m}$. Hence, the conditional mean of $Qb$ must be both proportional to and orthogonal to $1_{m}$. Consequently, its conditional mean must equal zero, implying $E[b\mid\mu,\sigma^{2}]=\mu1_{m}$. Likewise, the conditional second moment $E[Qb(Qb)'\mid\mu,\sigma^{2}]$ must be a permutation-invariant matrix---i.e., a linear combination of $I_{m}$ and $1_{m}1_{m}'$. Since $1_{m}'Qb=0$, this second moment matrix must also annihilate $1_{m}$, forcing it to be proportional to $Q$. Hence, $\mathrm{Var}[b\mid\mu,\sigma^{2}]=\mathrm{Var}[Qb\mid\mu,\sigma^{2}] \propto Q$. Since $(Qb)'Qb=\sum_{j}(b_{j}-\mu)^{2}=(m-1)\sigma^{2}$ is a function of the conditioning variables, $\mathrm{tr}(\mathrm{Var}[b\mid\mu,\sigma^{2}])=E[(Qb)'Qb\mid\mu,\sigma^{2}]=(m-1)\sigma^{2}$. Direct calculation yields $\mathrm{tr}(Q)=m-1$. Therefore, the constant of proportionality must be $\sigma^2$.
\end{proof}

The conditional pairwise correlation between entries of $b$ is $-1/(m-1)$, which reflects the adding up constraint imposed by conditioning on $\mu$. Thus, conditioning removes the nuisance parameter governing the marginal correlation between the entries of $b$. While conditioning on $(\mu,\sigma^{2})$ pins down the first two moments of $b$, the centered errors $b-\mu1_{m}$ may remain conditionally dependent at higher moments in ways that depend on additional parameters.


\begin{rem}[Linear transformations]\label{rem:equivariant}
    Exchangeability is preserved under permutation-equivariant linear transformations. Hence, if $r_n$ obeys Assumption~\ref{assu:random_b}, so does $(c_1 I_m+c_2 1_m1_m')r_n$ for finite constants $c_1\neq0$ and $c_2$. This transformation maps the hyperparameters as $\mu\mapsto(c_1+mc_2)\mu$ and $\sigma^2\mapsto c_1^2\sigma^2$.
\end{rem}


\subsection{Leading term and bias}

Assumption \ref{assu:random_b}, in conjunction with the standard
GMM regularity conditions discussed in Section \ref{sec:GMM-preliminaries},
imply the asymptotic representation
\[
\sqrt{n}\left(\hat{\theta}-\theta_{0}\right)=\Lambda r_{n}+o_{p}\left(1\right)\overset{d}{\rightarrow}\Lambda r,\qquad r:=b+\varepsilon.
\]
Note that Assumption \ref{assu:random_b}.ii ensures that the first four moments of $r$ are finite. Conditional on the hyperparameters, the leading term $\Lambda r=\Lambda b+\Lambda\varepsilon$ has mean
\[
E\left[\Lambda r\mid\mu,\sigma^{2}\right]=\mu\,\Lambda1_{m}.
\]
When the specification errors have nonzero mean, the asymptotic distribution of $\sqrt{n}(\hat{\theta}-\theta_{0})$ is centered at $\mu\Lambda1_{m}$ rather than at zero. Its variance splits into two terms,
\[
\mathrm{Var}\left[\Lambda r\mid\mu,\sigma^{2}\right]=\underbrace{\Lambda V\Lambda'}_{\text{moment uncertainty}}+\underbrace{\sigma^{2}\Lambda Q\Lambda'}_{\text{specification error}}=\Lambda\Sigma\Lambda',\qquad \Sigma:=V+\sigma^{2}Q.
\]

In Section~\ref{sec:bias-estimation}, I propose estimators that mitigate both the bias and the variance attributable to specification errors by leveraging estimates of $\mu$ and $\sigma^{2}$. To this end, I condition on $(\mu,\sigma^{2})$ throughout the analysis, treating them as fixed parameters. For brevity, I leave this conditioning implicit wherever possible.

\section{Examples \protect\label{sec:Examples}}

It is useful to walk through some examples of misspecified models that can be fit into the framework of the previous section. My discussion will center on
how Assumption~\ref{assu:random_b}.ii can be violated and
how to design moment conditions where it can plausibly be satisfied.

\begin{example}[Equal means]
\label{exa:mean}Let  $Y_{n}$ be an $m\times1$ vector of independent
sample means with $\mathrm{Var}\left(Y_{n}\right)=\mathrm{diag}\left(v_{1},\dots,v_{m}\right)/n=V/n$,
where  $n$ is the total sample size. The researcher considers a model
that restricts the $m\times1$ vector of population means $\vartheta_{n}$
to equal a common scalar $\theta$.

To map this example to the local misspecification framework, let $\vartheta_{n}:=\theta_{0}1_{m}+b/\sqrt{n}$
for some $\theta_{0}\in\Theta$ and specification errors $b\in\mathbb{R}^{m}$.
The $m$-vector of population moment conditions then takes the form
$g\left(\theta\right)=\vartheta_{n}-\theta1_{m}$, which implies the
population Jacobian is $G=-1_{m}$. Thus, GMM is consistent and the
scaled population moment conditions $\sqrt{n}g\left(\theta_{0}\right)=b$
are the specification errors.

If $b$ is exchangeable with bounded fourth moments, then Assumption~\ref{assu:random_b}.ii
is satisfied. In contrast, Assumption~\ref{assu:random_b}.ii
will be violated if either $E\left[b_{j}\right]$ or $\mathrm{Var}\left(b_{j}\right)$
varies with $j$. For instance, if noisier means have larger deviations
then one might expect $E\left[b_{j}\right]\propto\sqrt{v_{j}}$. In this case, residualizing the means against the noise levels---e.g., by using the methods proposed by \textcite{chen2026empirical}---can potentially restore exchangeability.
\end{example}

The following example considers an overidentified instrumental variables
model in which the plausibility of the exchangeability assumption hinges on moment scaling.
\begin{example}[Exclusion violations]
\label{exa:IV_example}A researcher has cross-sectional data
on  $m>2$ instruments  $Z_{i1},\dots,Z_{im}$ for a scalar endogenous
treatment  $T_{i}$,  where  $i$ indexes individual level observations.
Collect the instruments into the vector $Z_{i}:=\left(Z_{i1},\dots,Z_{im}\right)^{\prime}$
and denote the outcome by  $Y_{i}$. The triples  $\left\{ Y_{i},T_{i},Z_{i}'\right\} _{i=1}^{n}$
are presumed to be  i.i.d.\ samples from a DGP $P_{n}$ indexed
by the sample size to capture local misspecification. I use $E_{n}\left[\cdot\right]$
to denote expectations under this drifting DGP conditional on the realized $b$. The $E_{\infty}\left[\cdot\right]$ operator denotes
expectations under the limiting law $P_{\infty}:=\lim_{n\to\infty}P_{n}$ conditional on $b$.

The researcher's model stipulates the population moment restrictions
$g\left(\theta\right)=E_{n}\left[\left(Y_{i}-\theta T_{i}\right)Z_{i}\right]$
for scalar  $\theta$, implying $G=-E_{n}\left[T_{i}Z_{i}\right]$.
Now define $\vartheta_{n}:=E_{n}\left[Y_{i}Z_{i}\right]$ and suppose that for some vector $\gamma\in\mathbb{R}^{m}$
\[
\vartheta_{n}=\theta_{0}E_{n}\left[T_{i}Z_{i}\right]+\frac{1}{\sqrt{n}}E_{n}\left[Z_{i}Z_{i}'\right]\gamma\Rightarrow E_{n}\left[\left(Y_{i}-\theta_{0}T_{i}-\frac{1}{\sqrt{n}}Z_{i}'\gamma\right)Z_{i}\right]=0.
\]
Excludability violations of this form were considered by \textcite{conley2012plausibly}, who studied estimation and inference given priors on the direct effects. \textcite{kolesar2015identification} also considered excludability violations of this form in the case where the entries of $\gamma$ are uncorrelated with the instrument first stages.

Suppose the direct effects  $\left\{ \gamma_{j}\right\} _{j=1}^{m}$
are exchangeable. The associated scaled
population moment conditions satisfy
\[
\sqrt{n}g\left(\theta_{0}\right)=E_{n}\left[Z_{i}Z_{i}'\right]\gamma\to E_{\infty}\left[Z_{i}Z_{i}'\right]\gamma=b.
\]
In this case, $b$ is generally not exchangeable unless $E_{\infty}\left[Z_{i}Z_{i}'\right]$ takes the permutation-equivariant form described in Remark \ref{rem:equivariant}. Since $E_{\infty}\left[Z_{i}Z_{i}'\right]=\mathrm{Var}_\infty(Z_i)+E_\infty[Z_i]E_\infty[Z_i]'$, this effectively requires the instruments to share a common mean and exhibit an exchangeable (equicorrelated and homoscedastic) variance structure.

This difficulty can be circumvented by working with rescaled moment
conditions taking the form
\[
g\left(\theta\right)=E_{n}\left[Z_{i}Z_{i}'\right]^{-1}E_{n}\left[\left(Y_{i}-\theta T_{i}\right)Z_{i}\right]=\delta - \theta \pi,
\]
for $\delta:=E_{n}\left[Z_{i}Z_{i}'\right]^{-1}E_{n}\left[Z_{i}Y_{i}\right]$ the $m$-vector of population reduced form coefficients and $\pi:=E_{n}\left[Z_{i}Z_{i}'\right]^{-1}E_{n}\left[Z_{i}T_{i}\right]$ the corresponding vector of first stage coefficients. These rescaled moments have $b=\gamma$, satisfying Assumption \ref{assu:random_b}.ii whenever $\gamma$ is exchangeable with bounded fourth moments.
\end{example}

Because excludability violations are often cited as a first-order
concern in applied work with instrumental variables, I will focus
on variants of Example \ref{exa:IV_example} in
the applications of Section \ref{sec:monte-carlo} and \ref{sec:empirical}. As the following example illustrates, however, the methods developed in this paper are equally applicable to nonlinear models.

\begin{example}[A binary choice restriction]\label{exa:binarychoice}
    Consider the binary choice model $\Pr\left(Y_{i}=1|X_{i}\right)=F\left(X_{i}'\beta\right),$ where $\beta\in\mathbb{R}^{m}$ and the index function $F:\mathbb{R}\to(0,1)$ is strictly increasing and continuously differentiable with derivative $f:\mathbb{R}\to\mathbb{R}_{>0}$. Maximum likelihood estimation of such a model is equivalent to method of moments on the $m$-vector of log-likelihood scores, the population versions of which can be written
    \[
    s(\beta)= E_{n}\left[\frac{Y_{i}-F\left(X_{i}'\beta\right)}{F\left(X_{i}'\beta\right)\left[1-F\left(X_{i}'\beta\right)\right]}f\left(X_{i}'\beta\right)X_{i}\right].
    \]

    Now suppose a researcher imposes the coefficient restriction $\beta=\theta 1_{m}$  for $\theta\in\mathbb{R}$. For example, in the discrete choice experiment of \textcite{mas2019labor}, job applicants were randomly presented with one of $m$ pairs of job options involving distinct earnings and hours bundles. The outcome $Y_i$ measured whether individual $i$ chose the bundle with longer hours, while the vector $X_i$ measured the interaction between an indicator for the assigned menu and the implicit wage of the additional hours required by the longer option. In this case, the restriction would require the marginal rate of substitution between leisure and consumption to be equal across the $m$ distinct choice menus.

    It is natural to estimate this restricted model by GMM using the moment conditions
    \[
    s\left(\theta 1_{m}\right)=E_{n}\left[\frac{Y_{i}-F\left(\theta X_{i}'1_{m}\right)}{F\left(\theta X_{i}'1_{m}\right)\left[1-F\left(\theta X_{i}'1_{m}\right)\right]}f\left(\theta X_{i}'1_{m}\right)X_{i}\right].
    \]
    Suppose however that the true choice probability obeys a drifting DGP
    \[
        \Pr\left(Y_{i}=1|X_{i}\right)=F\left(\theta_{0}X_{i}'1_{m}\right)+\frac{1}{\sqrt{n}}X_{i}'b\cdot F\left(\theta_{0}X_{i}'1_{m}\right)\left[1-F\left(\theta_{0}X_{i}'1_{m}\right)\right],
    \]
    where $b\in\mathbb{R}^{m}$. When $X_i'b$ has bounded support (e.g., the menu indicators described above) this expression is almost surely a valid probability for large enough $n$. By iterated expectations,
    \[
        s\left(\theta_{0}1_{m}\right)=\frac{1}{\sqrt{n}}E_{n}\left[\left(X_{i}'b\right)f\left(\theta_0 X_{i}'1_{m}\right)X_{i}\right]=\frac{1}{\sqrt{n}}E_{n}\left[f\left(\theta_0 X_{i}'1_{m}\right) X_{i}X_{i}'\right]b.
    \]

    Since $f>0$, when $E_{n}[X_{i}X_{i}']$ is positive definite, the weighted matrix $E_{n}[f(\theta X_{i}'1_{m})X_{i}X_{i}']$ is also invertible. One can then work with the rescaled moment condition
    \[
     g\left(\theta\right)=E_{n}\left[f\left(\theta X_{i}'1_{m}\right) X_{i}X_{i}'\right]^{-1}s(\theta 1_{m}),
    \]
    which yields $\sqrt{n} g\left(\theta_{0}\right)=b$.
    When $b$ is exchangeable with bounded fourth moments, Assumption~\ref{assu:random_b}.ii is satisfied.
\end{example}

\section{Identification of hyperparameters \protect\label{sec:Identification-of-hyperparameter}}

This section studies identification of the hyperparameters $\left(\mu,\sigma^{2}\right)$, which are treated as fixed (i.e., conditioned on) throughout.
Recall that the GMM estimator sets $p$ linear combinations of the
sample moments equal to zero. Assumption~\ref{assu:random_b}
implies the scaled moment conditions can be written
\[
\hat{r}:=\sqrt{n}\hat{g}\left(\hat{\theta}\right)=Mr_{n}+o_{p}\left(1\right)\overset{d}{\rightarrow}Mr,
\]
where the sample and population annihilator matrices were defined
in Section \ref{subsec:Influence-and-sensitivity}. Intuitively, the
process of fitting GMM masks $p$ linear combinations of the  $b_{j}$'s
that lie in the nullspace of  $M$. The exchangeability assumption
allows one to infer moments of these hidden linear combinations
from the  $m-p$ identified linear combinations that lie in  $M$'s
column space.

\subsection{Moment equations}

Assumption~\ref{assu:random_b} implies
\[
E\left[Mr\right]=\mu M1_{m},\quad \mathrm{Var}\left[Mr\right]=M\Sigma M'.
\]
Provided $\left\Vert M1_{m}\right\Vert ^{2}=\left(M1_{m}\right)'\left(M1_{m}\right)>0$,
it follows that
\begin{equation}
E\left[\frac{\left(M1_{m}\right)'Mr}{\left\Vert M1_{m}\right\Vert ^{2}}\right]=\mu,\quad E\left[\frac{\left\Vert Mr-\mu M1_{m}\right\Vert ^{2}-\mathrm{tr}\left(MVM'\right)}{\mathrm{tr}\left(M'MQ\right)}\right]=\sigma^{2}.\label{eq:identification}
\end{equation}
Note that $\left\Vert M1_{m}\right\Vert ^{-2}\left(M1_{m}\right)'Mr$
is the coefficient from an OLS regression of $Mr$ on $M1_{m}$. The
ratio $\left[\left\Vert Mr-\mu M1_{m}\right\Vert ^{2}-\mathrm{tr}\left(MVM'\right)\right]/\mathrm{tr}\left(M'MQ\right)$
is a degrees-of-freedom adjusted estimator of the variance of the
specification errors. It subtracts the expected sampling noise $\mathrm{tr}\left(MVM'\right)$
from the residual sum of squares, then divides by the effective residual degrees of freedom $\mathrm{tr}\left(M'MQ\right)=\mathrm{tr}\left(MM'\right)-\left\Vert M1_{m}\right\Vert ^{2}/m$.
When $m-p\ge2$, this denominator is strictly positive and the ratio is well-defined. These population moment equations motivate the hyperparameter
estimators that will be proposed in Section \ref{sec:estimation},
which replace $Mr$ with the sample residuals $\hat{r}$ and population
matrices with their sample analogs.

\subsection{Identification failures \protect\label{subsec:Identification-failures}}

Some overidentified models imply $M1_{m}=0$, leading to non-identification
of  $\mu$. This problem arises whenever  $1_{m}$ lies in the column
space of $G$.

\subsubsection{Constant Jacobians}

Recall that $G=-1_{m}$ in Example \ref{exa:mean}, implying the mean
bias loads directly onto the parameter of interest: $E\left[\Lambda r\right]=\mu$.
Here the hyperparameter  $\mu$ is absorbed by the GMM fitting
process because it is collinear with the parameter being estimated.
Consequently, the $J$-test has no power to detect non-zero values
of $\mu$ \parencite{newey1985gmm}.

The same issue arises in the binary choice model of Example~\ref{exa:binarychoice} when the link function $F$ is logistic. Differentiating the rescaled moments at $\theta_{0}$ yields the Jacobian
\[
G=-E_{n}\left[f\left(\theta_{0}X_{i}'1_{m}\right)X_{i}X_{i}'\right]^{-1}E_{n}\left[\frac{f\left(\theta_{0}X_{i}'1_{m}\right)^{2}}{F\left(\theta_{0}X_{i}'1_{m}\right)\left[1-F\left(\theta_{0}X_{i}'1_{m}\right)\right]}X_{i}X_{i}'\right]1_{m},
\]
up to terms that vanish under the drifting DGP. For the logistic link, $f=F\left(1-F\right)$ makes the two expectations equal, yielding $G=-1_{m}$. Hence, the mean $\mu$ is unidentified in the logit model. In contrast, link functions with $f\neq F\left(1-F\right)$ generically produce a Jacobian with non-constant entries, preserving identification of $\mu$.

\subsubsection{Intercepts in linear models \protect\label{subsec:Intercepts-in-linear}}

In overidentified linear models, $\mu$ typically remains identified
even when an intercept is included. Consider a $p=2$ extension of Example~\ref{exa:IV_example} where  $\theta=\left(\alpha,\beta\right)'$,
with  $\alpha$ capturing an intercept and $\beta$ a slope. Moreover,
an intercept is added to the  $m$ fundamental instruments  $Z_{i}$
so that the rescaled moment conditions take the form  $g\left(\theta\right)=S^{-1} E_{n}\left[\left(Y_{i}-\left(1,T_{i}\right)\theta\right)\left(1,Z_{i}'\right)'\right]$, where $S:=E_{n}[\left(1,Z_{i}'\right)'\left(1,Z_{i}'\right)]$ is the second moment matrix of the augmented instruments.
In this case
\[
G=-S^{-1}\left[\begin{array}{cc}
1 & E_{n}\left[T_{i}\right]\\
E_{n}\left[Z_{i}\right] & E_{n}\left[T_{i}Z_{i}\right]
\end{array}\right] = - \pi,
\]
where $\pi$ is an $(m+1) \times 2$ matrix of first-stage coefficients. The first column of $\pi$ is necessarily a unit vector $(1, 0_m')'$ capturing the first-stage coefficients for the constant vector. Its second column collects the population coefficients from regressing $T_i$ on the augmented instruments: an intercept $\pi_{12}$ and slopes $\pi_{22},\dots,\pi_{m+1,2}$ on the $m$ fundamental instruments. Identification of $\mu$ fails only when these $m$ slopes share a common nonzero value.

In contrast, identification fails in an extension of Example~\ref{exa:IV_example}
with $\theta=\left(\alpha,\beta\right)'$ and mutually exclusive binary
instruments  $Z_{ij}\in\left\{ 0,1\right\} $ obeying $\sum_{j=1}^{m}Z_{ij}=1$,
a design that arises often in the recent literature on judge IV \parencite{chyn2025examiner}.
Note that the intercept construction of the previous paragraph is unavailable here: $\sum_{j=1}^{m}Z_{ij}=1$ makes the augmented instrument set $(1,Z_{i}')$ perfectly collinear, rendering $S$ singular.
To understand the problem, consider the rescaled moment conditions
\[
g\left(\theta\right)=E_{n}\left[Z_{i}Z_{i}'\right]^{-1}E_{n}\left[\left(Y_{i}-\left(1,T_{i}\right)\theta\right)Z_{i}\right].
\]
Letting  $q=E_{n}\left[Z_{i}\right]$, the Jacobian takes the form
\[
G=-\mathrm{diag}\left(q\right)^{-1}\left(E_{n}\left[Z_{i}\right],E_{n}\left[T_{i}Z_{i}\right]\right)=-\left(1_{m},\mathrm{diag}\left(q\right)^{-1}E_{n}\left[T_{i}Z_{i}\right]\right).
\]
From the first column,  $1_{m}$ lies in the span of $G$. Therefore
$M1_{m}=0$, implying $\mu$ is not identified. This failure stems fundamentally from the adding up constraint $\sum_{j}Z_{ij}=1$: a direct effect common to all judges is constant and is absorbed by the intercept.

\subsubsection{Which parameters are biased?}

When $\mu$ is not identified, GMM's leading bias term $\mu\Lambda1_{m}$
cannot be consistently estimated. Non-identification of $\mu$ therefore
implies that some linear combination of the entries in  $\hat{\theta}$
has a leading bias of order $O\left(n^{-1/2}\right)$ that cannot
be removed. However, other linear combinations may have no leading
bias. In the judge-IV example,  $\mu$
loads directly on the intercept parameter  $\alpha$. But  $\beta$
is insensitive to  $\mu$, yielding no leading-term bias in the GMM
point estimate $\hat{\beta}$:
\[
E\left[\Lambda_{\beta}r\right]=\mu\Lambda_{\beta}1_{m}=-\mu\Lambda_{\beta}G_{\alpha}=0.
\]
Here, $\Lambda_{\beta}$ is the row of $\Lambda$ corresponding to
$\beta$ and $G_{\alpha}=-1_{m}$ is the intercept column of $G$.
The final equality uses $\Lambda_{\beta}G_{\alpha}=0$, which follows
from $\Lambda G=-I_{p}$.

Though GMM exhibits no leading-term bias for $\beta$ in this example,
it does exhibit conditional bias $E\left[\Lambda_{\beta}r\mid b\right]=\Lambda_{\beta}b$.
Averaging over the exchangeable specification errors, this conditional
bias contributes variance  $\sigma^{2}\Lambda_{\beta}\Lambda_{\beta}'$
to the asymptotic distribution of $\hat{\beta}$:
\[
\mathrm{Var}\left[\Lambda_{\beta}r\right]=\underbrace{\Lambda_{\beta}V\Lambda_{\beta}'}_{\text{sampling error}}+\underset{\text{specification error}}{\underbrace{\sigma^{2}\Lambda_{\beta}\Lambda_{\beta}'}}.
\]
The specification-error term is $\sigma^{2}\Lambda_{\beta}Q\Lambda_{\beta}'$, which reduces to $\sigma^{2}\Lambda_{\beta}\Lambda_{\beta}'$ because $\Lambda_{\beta}1_{m}=0$.
In this scenario, the job of ``repairing'' GMM simplifies to reducing
the noise in  $\hat{\beta}$ attributable to specification error using
estimates of  $\sigma^{2}$, a task that can be accomplished via empirical
Bayes shrinkage.

\section{The limits of differencing \protect\label{sec:A-differencing-approach}}

Before moving on to empirical Bayes approaches that rely on estimating
the hyperparameters $\left(\mu,\sigma^{2}\right)$, it is useful to
consider what can be achieved when the hyperparameters are treated
as nuisance parameters. A natural strategy for dealing with the exchangeable
structure of the errors in the moment conditions is to difference
their shared mean $\mu$ out. As in Section \ref{sec:Identification-of-hyperparameter},
distributional statements throughout this section are conditional
on $\left(\mu,\sigma^{2}\right)$.

\subsection{Identification}

Let $D$ be a $\left(m-1\right)\times m$ differencing matrix obeying
$D1_{m}=0$ and $DD'=I_{m-1}$ (e.g., a Helmert contrast matrix).
The differenced sample moment conditions $D\hat{g}\left(\theta\right)$,
when evaluated at $\theta_{0}$, obey
\[
\sqrt{n}D\hat{g}\left(\theta_{0}\right)\overset{d}{\rightarrow}Dr=Db+D\varepsilon.
\]
Under Assumption \ref{assu:random_b}
\[
E\left[Dr\right]=\mu D1_{m}=0,\quad \mathrm{Var}\left(Dr\right)=\sigma^{2}DD'+DVD'=\sigma^{2}I_{m-1}+DVD'.
\]
Thus, differencing removes the mean bias of the moment conditions and consequently the leading term bias of the resulting GMM estimator, at the potential cost of inflating its variance. In fact, when $b$
is exchangeable, the prior expectation of the differenced violations
vanishes regardless of their scale: $E[Db]=DE[b]=0$. Hence, differencing
yields moment conditions whose superpopulation mean is zero even
under global misspecification of the sort described in Section \ref{subsec:Misspecification-as-contaminatio}.

Mirroring the discussion in Section \ref{subsec:Identification-failures},
one concern is that differencing may undermine identification of $\theta_{0}$
itself (e.g., by leading an intercept to drop out of the original
moment conditions). This concern is ruled out by the condition $M1_{m}\neq0$,
which ensures the differenced moment Jacobian $DG$ has full column
rank $p$. Under this rank condition, $\theta_{0}$ is locally identified from the differenced moments $Dg\left(\theta\right)$. These moments have mean zero at $\theta_{0}$, since $E[Dr]=\mu D1_{m}=0$. Global identification additionally requires that $\theta_{0}$ be the unique such solution.

\subsection{Leading term}

Let $\hat{\theta}_{D}$ denote the estimator obtained from conducting
GMM on the differenced moment conditions with $\left(m-1\right)\times\left(m-1\right)$
weighting matrix $W_{D}$. The sensitivity matrix of the differenced
GMM estimator is $\Lambda_{D}:=-\left(G'D'W_{D}DG\right)^{-1}G'D'W_{D}$.
The arguments above imply that when $M1_{m}\neq0$,
\[
\sqrt{n}\left(\hat{\theta}_{D}-\theta_{0}\right)=\Lambda_{D}Dr_{n}+o_{p}\left(1\right).
\]
Since $E\left[Dr\right]=0$, the leading-term bias has been removed.
However, the differenced estimator need not converge to a normal as
$n$ grows large.

To see why normality is not ensured in the asymptotic limit, observe
that
\[
\Lambda_{D}Dr=\underbrace{\Lambda_{D}D\varepsilon}_{\text{sampling variability}}+\underbrace{\Lambda_{D}Db}_{\text{specification errors}}.
\]
The first term, capturing sampling variability, is $N\left(0,\Lambda_{D}DVD'\Lambda_{D}'\right)$.
In contrast, the second term depends on the unknown exchangeable distribution
of the specification errors $b$.

\subsection{Limitations}

Recall that the asymptotics thus far have let $n$ grow large with
the number of moment conditions $m$ fixed. When $m$ is small, the
linear combination $\Lambda_{D}Db$ can be far from normal.
Assumption \ref{assu:random_b} guarantees that this term has
variance $\sigma^{2}\Lambda_{D}\Lambda_{D}'$, that four of its moments
exist, and that it is independent of the first term. However, a large
family of exchangeable distributions satisfy this property.

In the next section, I develop an asymptotic framework where the number
of moment conditions grows with the sample size $n$, which allows
me to consistently estimate the hyperparameters $\left(\mu,\sigma^{2}\right)$
at rates suitable for inference on $\theta_{0}$. Estimating the first-step
hyperparameters can provide two advantages beyond inference. One is
that knowledge of the hyperparameters can enable additional reductions
in the variability of the estimator. In particular, knowledge of $\sigma^{2}$
can be used to develop shrinkage estimators with certain efficiency
advantages.

Second, the hyperparameters are often interesting in their own right,
providing economically interpretable information about the nature
of the specification errors. As will be illustrated by the empirical
application of Section \ref{sec:empirical}, it is reasonable for
researchers to scrutinize the hyperparameter estimates as a way of
gauging the plausibility of the exchangeability framework on which
the estimators are predicated.

\section{\protect\label{sec:estimation}Estimating hyperparameters}

This section considers how to estimate the hyperparameters  $\left(\mu,\sigma^{2}\right)$
in an environment where  $m$ and $n$ grow jointly, while  $p$ remains
fixed. Intuitively, these results leverage the high degree of overidentification
 $m-p$ found in many empirical studies, which provide opportunities
to detect overdispersion in the empirical moment conditions attributable
to misspecification.

The results in this section are established under a joint asymptotic
framework in which $m$ grows more slowly than $\sqrt{n}$. This asymptotic regime,
which was also considered by \textcite{newey1990efficient} and \textcite{chernozhukov2025plausible},
allows me to state explicit rate conditions under which the feasible
estimators approximate their oracle counterparts. I collect the regularity
conditions used throughout below, using $\left\Vert \cdot\right\Vert _{2}$
to denote the spectral operator norm.
\begin{assumption}
\label{assu:regime}(Asymptotic Regime and Regularity). The following
conditions hold as  $n\to\infty$:

i)  $m=m(n)\to\infty$ with  $m^{2}/n\to0$,  and  $p$ is fixed.

ii) $\|W\|_{2}=1$ and $\|V\|_{2}=1$ (normalizations). The row sums
of $W$ and $V$ are uniformly bounded: $\max_{1\le j\le m}\sum_{k}|W_{jk}|=O(1)$
and $\max_{1\le j\le m}\sum_{k}|V_{jk}|=O(1)$.

iii) The first-stage estimators satisfy, uniformly in  $m$: $\|\hat{M}-M\|_{2}=O_{p}(\sqrt{m/n})$,
 $\|\hat{V}-V\|_{2}=O_{p}(\sqrt{m/n})$,  $\|\hat{\Lambda}-\Lambda\|_{2}=O_{p}(1/\sqrt{n})$,
and the linearization remainder satisfies $\|\sqrt{n}\hat{g}(\hat{\theta})-\sqrt{n}\hat{g}(\theta_{0})-\hat{G}\sqrt{n}(\hat{\theta}-\theta_{0})\|=O_{p}(\sqrt{m/n})$.

iv) $MVM'$ has rank $m-p$, and its smallest nonzero eigenvalue, denoted $\lambda_{\min}^{+}(MVM')$, satisfies $\lambda_{\min}^{+}(MVM')\cdot n/m\to\infty$.

v) The entries of $G$ are uniformly bounded, $\sup_{j,k}|G_{jk}|<\infty$, and $\left\Vert M\right\Vert _{2}=O(1)$.

vi) $m^{-1}\left\Vert M1_{m}\right\Vert ^{2}\to\kappa\in(0,\infty)$.

vii) $V_{n}:=\mathrm{Var}(\varepsilon_{n})$ exists with $\left\Vert V_{n}-V\right\Vert _{2}=O(\sqrt{m/n})$,
and for any $n$, $\mathrm{Var}(\varepsilon_{n}'M'M\varepsilon_{n})\le C\cdot\mathrm{tr}((M'MV)^{2})$.
\end{assumption}
Condition (i) specifies that the number of moment conditions grows
with the sample size but at a rate slower than $\sqrt{n}$, ensuring
that first-stage estimation errors remain asymptotically negligible.
When moment conditions are added at a faster rate, a many-instruments
bias emerges that compromises both GMM and the sensitivity-based
corrections that will be proposed. Thus, the asymptotic results that
follow apply to use cases where models are moderately overidentified,
but not so complex that techniques for finite-dimensional modeling
break down. In Section \ref{sec:empirical}, I explore a variant of
a specification considered by \textcite{angrist1991does} that fits
these criteria.

The first part of condition (ii) is without loss of generality: GMM
is invariant to rescaling of either the moment conditions $\hat g(\theta)$ or the weighting matrix $W$ by a scalar. These scalars fix the operator norms but leave the condition numbers of $W$ and $V$ unchanged. The order of $V$'s smallest eigenvalue $\lambda_{\min}(V)$, and hence of $\lambda_{\min}^+(MVM')$, is a feature of the design, not the normalization. Accordingly, the eigenvalue conditions below are stated as rates. The bounded row-sum conditions on $W$ and $V$ hold
automatically when these matrices are diagonal and more generally
when each of their rows has sparse or rapidly decaying entries. For
$V$, the condition controls pairwise dependence among the moment errors.

Condition (iii) requires that the first-stage estimators converge
uniformly in  $m$. The  $m\times m$ matrix $\hat{V}$ is required
to converge to $V$ at rate $O_{p}(\sqrt{m/n})$, which is delivered
by standard concentration inequalities under sub-Gaussian tails or
bounded support, but should be verified directly otherwise. The sensitivity matrix
$\hat{\Lambda}$ is required to converge at the parametric rate $O_{p}(1/\sqrt{n})$
in operator norm. Appendix~\ref{app:verify} verifies this rate in the cell-IV design, where the rows of $\Lambda$ average the first-stage errors across the $m$ moments. The operator norm of $\Lambda$ itself is bounded when the smallest eigenvalue of $G'WG$ is bounded away from zero, a bound imposed as Assumption~\ref{assu:rate}.ii below. The rate required of $\hat{M}$ can be traced to these inputs via the identity $M=I_{m}+G\Lambda$. When $\hat{G}$ converges at the $O_{p}(\sqrt{m/n})$ rate of $\hat{V}$, the decomposition $\hat{M}-M=(\hat{G}-G)\hat{\Lambda}+G(\hat{\Lambda}-\Lambda)$ is $O_{p}(\sqrt{m/n})$ in operator norm, using $\|G\|_{2}=O(\sqrt{m})$ from condition (v). Finally, the linearization remainder condition ensures
that the scaled residuals $\hat{r}=\sqrt{n}\hat{g}(\hat{\theta})$
are well-approximated by $\hat{M}\sqrt{n}\hat{g}(\theta_{0})$ under
joint asymptotics, which holds automatically for linear moment conditions
and follows from uniformly bounded Hessians in the nonlinear case.

Condition (iv) keeps the residualized moment covariance nondegenerate as the number of
moments grows. It requires the smallest nonzero eigenvalue of $MVM'$ to be of larger order than $m/n$. The rank-$(m-p)$ requirement makes $MVM'$ positive definite on the residualized subspace $\mathrm{col}(M)$. It does not require $\Sigma$ positive definite elsewhere. Because $\sigma^{2}Q$ only raises the form on $\mathrm{range}(M)$, the rank and floor carry over to $M\Sigma M'$ for every $\sigma^{2}\ge0$, including $\sigma^{2}=0$.

Condition (v) requires uniformly bounded entries of the Jacobian $G$,
ensuring that individual moment conditions do not exert disproportionate
influence on the parameter estimates, and additionally requires that
the projection matrix $M$ act as a bounded operator.
Condition (vi) requires that $1_{m}$ not lie asymptotically in the
column space of $G$, ensuring identification of the mean specification
error $\mu$. In the special case where $\left(M1_{m}\right)_{j}\to1$
uniformly across $j$, one can show that $m^{-1}G'W1_{m}\to0$ and
$\kappa=1$.

Condition (vii) bounds the variance of the quadratic form $\varepsilon_{n}'M'M\varepsilon_{n}$
that arises in the consistency argument for $\hat{\sigma}^{2}$ (Lemma~\ref{lem:sigma-consistency}).
Under i.i.d.\ moment contributions with finite fourth moments this
condition follows from standard moment computations on quadratic forms.
With weakly dependent or martingale-difference contributions the condition
follows from analogous moment-cumulant bounds. The first clause requires that the finite-sample covariance $V_{n}=\mathrm{Var}(\varepsilon_{n})$ exists and converges to its limit $V$ in operator norm at the rate $\sqrt{m/n}$ of the first-stage estimators in condition (iii). This is automatic in i.i.d.\ designs, where $V_{n}=V$ for all $n$; in the cell-IV design the gap is the smaller $O(1/n)$. It holds more generally whenever the moment-contribution covariance converges to its population analog at this rate.

In the remainder of this section, I will establish consistency of
estimators of the hyperparameters $\left(\mu,\sigma^{2}\right)$, which are treated as fixed throughout.
When Assumption \ref{assu:random_b} is invoked jointly with
Assumption \ref{assu:regime}, it is understood to hold at each $m$
in the sequence $m\left(n\right)$, with a fourth-moment bound that does not depend on $m$. The consistency arguments in this
section rely on the quadratic-form concentration of Lemma~\ref{lem:perm-quadratic}, while the empirical Bayes results of Section \ref{sec:bias-estimation} rely on the previously established
consistency of $\widehat{\sigma^{2}}$. In Section
\ref{sec:distribution-theory}, I leverage Assumption \ref{assu:random_b}
to establish asymptotic normality of both estimators.

\subsection{Estimating $\mu$ \protect\label{subsec:Estimating-mu}}

Suppose that $\left\Vert \hat{M}1_{m}\right\Vert>0$. As discussed in Section \ref{subsec:Identification-failures}, this
condition fails if a linear combination of the columns of $\hat{G}$
equals $1_{m}$. A plug-in estimator of $\mu$, motivated by \eqref{eq:identification},
is the regression coefficient
\[
\hat{\mu}=\left\Vert \hat{M}1_{m}\right\Vert ^{-2} \left(\hat{M}1_{m}\right)'\hat{r}.
\]

Under Assumption~\ref{assu:regime}.iii, the first-stage estimator  $\hat{M}$
converges to its population counterpart, while Assumption~\ref{assu:regime}.vi guarantees that $\left\Vert M1_{m}\right\Vert ^{2}$ is of order $m$ (i.e., that $\kappa>0$). The
following Lemma establishes the convergence rate of $\hat \mu$ when $m$ and  $n$ grow large.
\begin{lem}
\label{lem:mu_consistency}Under Assumptions~\ref{assu:random_b}
and \ref{assu:regime},
\[
\hat{\mu}-\mu=O_{p}\left(\frac{1}{\sqrt{m}}\right)+O_{p}\left(\sqrt{\frac{m}{n}}\right).
\]
\end{lem}
\begin{proof}
See appendix.
\end{proof}
The first term captures the error of an infeasible estimator that
uses the population $M$ and residual $r_{n}$, rather than $\hat{M}$ and $\hat{r}$, to estimate $\mu$. This term shrinks
at a $1/\sqrt{m}$ rate because it averages across $m$ moment conditions.
The second term captures the feasibility cost of replacing these population
quantities with their sample analogs, primarily the replacement of $M$ by $\hat{M}$, which converges to
$M$ in operator norm at rate $\sqrt{m/n}$ by Assumption \ref{assu:regime}.iii.
Thus,  $\hat{\mu}$ is consistent for $\mu$ when $m$ grows more
slowly than  $n$. Under Assumption \ref{assu:regime}.i, $m^2/n\to0$, implying the second term is dominated by the first.
\begin{rem}[Weak identification of $\mu$]
\label{rem:strong_id}The requirement $\kappa>0$ in Assumption~\ref{assu:regime}.vi places $\hat{\mu}$ in an asymptotic regime where its target $\mu$ is \emph{strongly identified}. An interesting question for future work is how to characterize the behavior of $\hat{\mu}$ in the setting where $m\kappa=O\left(1\right)$, under which the hyperparameter is \emph{weakly identified}.
\end{rem}

\subsection{Estimating  $\sigma^{2}$}

From \eqref{eq:identification},  a plug-in estimator of  $\sigma^{2}$
takes the form
\[
\widetilde{\sigma^{2}}=\frac{\left\Vert \hat{r}-\hat{\mu}\hat{M}1_{m}\right\Vert ^{2}-\mathrm{tr}\left(\hat{M}\hat{V}\hat{M}'\right)}{\mathrm{tr}\left(\hat{M}'\hat{M}Q\right)},
\]
 where $\hat{V}$ is an estimator of $V$ satisfying condition (iii) of Assumption~\ref{assu:regime}. Centering on the fitted $\hat{\mu}$ leaves $\widetilde{\sigma^{2}}$ with a small downward bias of order $O(1/m)$, negligible relative to its sampling error. To ensure the
estimate is non-negative, one can rely on the truncated estimator
\[
\widehat{\sigma^{2}}=\max\left\{ 0,\widetilde{\sigma^{2}}\right\} .
\]
The max operator introduces a mild upward bias near $\sigma^{2}=0$ of the same $O(1/\sqrt{m})$ order as the estimator's sampling error.


\subsubsection{Connection to the J-statistic}

The estimator  $\widetilde{\sigma^{2}}$ bears a close connection
to the usual  $J$-statistic. Assuming the variance estimator $\hat{V}$
is nonsingular, this statistic can be written
\[
\hat{J}=\hat{r}'\hat{V}^{-1}\hat{r},
\]
 where $\hat{r}=\sqrt{n}\hat{g}\left(\hat{\theta}\right)$ and $\hat{\theta}$
is constructed using $\hat{W}=\hat{V}^{-1}$.

In the special case where  $\hat{W}=\hat{V}=I$,  one can write
\[
\widetilde{\sigma^{2}}=\frac{\hat{J}-\left(m-p\right)-\hat{\mu}^{\,2}\,1_{m}'\hat{M}'\hat{M}1_{m}}{m-p-\tfrac{1}{m}1_{m}'\hat{M}'\hat{M}1_{m}}.
\]
Provided that the true moment covariance is also the identity ($V=I$), $m-p$ is the large-$n$ expected value of $\hat{J}$ under correct
specification. The term $-\hat{\mu}^{\,2}\,1_{m}'\hat{M}'\hat{M}1_{m}$
subtracts the additional noncentrality in $\hat{J}$ expected to arise
when $\mu=\hat{\mu}$ and $\sigma^{2}=0$. Hence, in this case, $\widetilde{\sigma^{2}}$ captures overdispersion in the moment conditions attributable to specification
error.

Because  $\widetilde{\sigma^{2}}$ weights moment deviations equally,
this intuition does not extend directly to other choices of  $\hat{W}$
and  $\hat{V}$, as the  $J$-statistic weights moments by their precision.
However, if one restricts attention to the optimally weighted case
in which  $\hat{W}=\hat{V}^{-1}$,  the following alternative estimator
of  $\sigma^{2}$ is also consistent, provided $\lambda_{\min}(V)$ is bounded away from zero:
\[
\widetilde{\sigma_{J}^{2}}:=\frac{\hat{J}-\left(m-p\right)-\hat{\mu}_W^{\,2}\,1_{m}'\hat{M}'\hat{V}^{-1}\hat{M}1_{m}}{\mathrm{tr}(\hat{V}^{-1}\hat{M}\hat{M}')-\tfrac{1}{m}1_{m}'\hat{M}'\hat{V}^{-1}\hat{M}1_{m}},
\]
where $\hat \mu_W:=(1_{m}' \hat{M}' \hat W \hat{M}1_{m}) ^{-1} 1_{m}' \hat{M}' \hat W \hat{r} $ is the $\hat W$-weighted analog of $\hat \mu$. Note that when  $\hat{V}=I$,  $\widetilde{\sigma_{J}^{2}}=\widetilde{\sigma^{2}}$. As with $\widetilde{\sigma^{2}}$, negative values are possible and the truncation $\max\{0,\widetilde{\sigma_{J}^{2}}\}$ would need to be applied in practice.

While the connection
of  $\widetilde{\sigma_{J}^{2}}$ to optimally-weighted GMM is appealing,
weighting squared moment deviations by  $\hat{V}^{-1}$ need not yield
uniformly more precise estimates of  $\sigma^{2}$ than  $\widetilde{\sigma^{2}}$.
Moreover, inverse variance weighting can yield finite sample biases
due to correlation between the weights and moment conditions \parencite{altonji1996small,newey2004higher}.

\subsubsection{Consistency}

Note that
\[
\mathrm{tr}\left(\hat{M}\hat{M}'\right)=\sum_{i,j=1}^{m}\hat{M}_{ij}^{2}\geq\sum_{i=1}^{m}\hat{M}_{ii}^{2}\geq\frac{\left(\mathrm{tr}\,\hat{M}\right)^{2}}{m}=\frac{\left(m-p\right)^{2}}{m},
\]
 where the second inequality follows from Cauchy-Schwarz and the final
equality follows from the fact that
\[
\mathrm{tr}\left(\hat{G}\left(\hat{G}'\hat{W}\hat{G}\right)^{-1}\hat{G}'\hat{W}\right)=\mathrm{tr}\left(\left(\hat{G}'\hat{W}\hat{G}\right)^{-1}\hat{G}'\hat{W}\hat{G}\right)=\mathrm{tr}\left(I_{p}\right)=p.
\]
Since $\|\hat{M}1_{m}\|^{2}/m=O_{p}(1)$ by Assumptions~\ref{assu:regime}.iii and~\ref{assu:regime}.vi,
the denominator $\mathrm{tr}(\hat{M}'\hat{M}Q)=\mathrm{tr}(\hat{M}\hat{M}')-\|\hat{M}1_{m}\|^{2}/m$
of $\widetilde{\sigma^{2}}$ exceeds $(m-p)^{2}/m-O_{p}(1)$ and so must grow with
$m$.

The following Lemma establishes the convergence rate of the variance
estimator under joint asymptotics.
\begin{lem}
\label{lem:sigma-consistency}Under Assumptions~\ref{assu:random_b}
and \ref{assu:regime},
\[
\widehat{\sigma^{2}}-\sigma^{2}=O_{p}\left(\frac{1}{\sqrt{m}}\right)+O_{p}\left(\sqrt{\frac{m}{n}}\right).
\]
\end{lem}
\begin{proof}
See appendix.
\end{proof}
As in Lemma \ref{lem:mu_consistency}, the rate is decomposable
into two terms. The first term is attributable to the performance
of an infeasible estimator using the oracle estimate of $\mu$ described after Lemma~\ref{lem:mu_consistency} along with the population
$M$ and $V$ matrices. The second term captures the first stage estimation
errors $\hat{\mu}-\mu$,  $\hat{M}-M$, and $\hat{V}-V$.

Writing $\hat{\Sigma}-\Sigma=(\hat{V}-V)+(\widehat{\sigma^{2}}-\sigma^{2})Q$, Assumption~\ref{assu:regime}.iii
controls the first term and Lemma~\ref{lem:sigma-consistency} the second. Both vanish as $m\to\infty$ and $m^{2}/n\to0$. Since $\|Q\|_{2}=1$, it follows immediately that $\hat{\Sigma}=\hat{V}+\widehat{\sigma^{2}}Q$ is consistent
for $\Sigma$.
\begin{rem}[Hyperparameter consistency with $m=o(n)$]
\label{rem:weaker-rate}While Assumption~\ref{assu:regime} stipulates
that  $m^{2}/n\to0$, the rates in Lemmas~\ref{lem:mu_consistency}
and~\ref{lem:sigma-consistency} require only  $m/n\to0$, under
which both the oracle error  $O_{p}(1/\sqrt{m})$ and the first-stage
error  $O_{p}(\sqrt{m/n})$ vanish. The stronger condition  $m^{2}/n\to0$
is driven by the  $m\times m$ matrix estimation problems that arise
with the shrinkage estimator studied in Section~\ref{sec:bias-estimation},
for which the  $O(m^{2})$ entries of  $\hat{M}$ must be estimated
precisely enough that they contribute downstream estimation error
of order $o_{p}\left(1/\sqrt{n}\right)$.
\end{rem}

\section{Bias correction and shrinkage \protect\label{sec:bias-estimation}}

In this section, I introduce two estimators that repair the damage to GMM introduced by misspecification. The first estimator removes the leading order bias by using the hyperparameter estimate $\hat \mu$. The second estimator uses the estimated hyperparameter $\widehat{\sigma^2}$ to form an empirical Bayes shrinkage prediction of the noise in the bias-corrected estimator. Subtracting this prediction from the bias-corrected estimator reduces its sensitivity to any remaining idiosyncratic specification errors. Each estimator is shown to retain GMM's $\sqrt{n}$ convergence rate despite reliance on plug-in hyperparameters. I show that the second estimator weakly improves on the first and provide conditions characterizing when the improvement is strict.

\subsection{Bias-corrected estimator}
Recall from Section \ref{sec:local_misspecification} that the leading misspecification bias of $\hat{\theta}$ is $\mu\Lambda1_{m}/\sqrt{n}$. Plugging in $\hat{\mu}$ for $\mu$ and $\hat{\Lambda}$ for $\Lambda$, then subtracting this leading term, yields the bias-corrected (BC) estimator
\[
\hat{\theta}_{BC}:=\hat{\theta}-\frac{\hat{\mu}}{\sqrt{n}}\cdot\hat{\Lambda}1_{m}.
\]

The BC estimator is a close cousin of the differencing estimator introduced in Section \ref{sec:A-differencing-approach}. Rather than differencing the leading term out, the BC estimator estimates it and subtracts it. When the moment conditions are linear and identity weights are used $(W=I_m, W_D=I_{m-1})$ the estimators $\hat \theta_{BC}$ and $\hat \theta_D$ are numerically equivalent.

The BC estimator is also closely related to a two-step minimum distance approach that \textcite{angrist2009mostly} term ``visual IV'' (VIV), wherein the reduced-form coefficients associated with dummy instruments are regressed on their corresponding first-stage estimates. The conventional TSLS fit is a weighted least squares regression of the reduced forms on the first stages, constrained to go through the origin. In contrast, the VIV estimator described below allows for an intercept, which should be approximately zero if the exclusion restriction is satisfied.

\begin{example}[BC as Visual IV \protect \label{exa:VIV}]
    Continuing the overidentified IV setting of Example~\ref{exa:IV_example}, suppose the instruments $Z_{i}\in\mathbb{R}^{m}$ are mutually exclusive cell dummies. In this case, the first-stage and reduced-form cell means can be written
    \[
        \hat{\pi} = \hat{W}^{-1}\sum_{i} Z_{i} T_{i}, \quad
        \hat{\delta} = \hat{W}^{-1}\sum_{i} Z_{i} Y_{i}, \quad
        \hat{W} := \sum_{i} Z_{i} Z_{i}',
    \]
    where $\hat{W}$ is the diagonal matrix of cell counts. The cell-dummy TSLS estimator
    \[
        \hat{\theta} = (\hat{\pi}'\hat{W}\hat{\pi})^{-1}\hat{\pi}'\hat{W}\hat{\delta},
    \]
    is the cell-size--weighted least squares slope of the reduced forms $\hat{\delta}$ on the first stages $\hat{\pi}$, constrained to pass through the origin. The sample sensitivity matrix of TSLS is $\hat{\Lambda} = (\hat{\pi}'\hat{W}\hat{\pi})^{-1}\hat{\pi}'\hat{W}$, while the vector $\hat{\Lambda}1_{m} = (\hat{\pi}'\hat{W}\hat{\pi})^{-1}\hat{\pi}'\hat{W}1_{m}$ gives the through-origin WLS coefficient of $1_{m}$ on $\hat{\pi}$.

    The omitted-variables bias formula relates the through-origin slope $\hat \theta$ to its intercept-inclusive counterpart $\hat \theta_{\mathrm{VIV}}$,
    \[
        \hat{\theta} = \hat{\theta}_{\mathrm{VIV}} + \frac{\hat{\mu}_{W}}{\sqrt{n}}\,\hat{\Lambda}1_{m},
    \]
    where $\hat{\theta}_{\mathrm{VIV}}$ is the slope and $\hat{\mu}_{W}/\sqrt{n}$ the intercept of the WLS regression of $\hat{\delta}$ on $(1_{m},\hat{\pi})$ with weight $\hat{W}$. The VIV intercept $\hat{\mu}_{W}$ agrees with $\hat \mu$ numerically when the cell counts are equal (i.e., when $\hat W \propto I_m$). In that case, the BC estimator $\hat \theta_{BC}$ also coincides with the slope $\hat{\theta}_{\mathrm{VIV}}$ of the VIV regression. Under unequal cell counts, the gap between the two estimators can be written
\[
\hat{\theta}_{BC}-\hat{\theta}_{\mathrm{VIV}}=\frac{\hat{\mu}_{W}-\hat{\mu}}{\sqrt{n}}\,\hat{\Lambda}1_{m}.
\]
Hence, the gap is entirely attributable to the two intercept estimates and scales with the vector $\hat{\Lambda}1_{m}$.
\end{example}

The same geometry extends to nonlinear models, provided the bias correction is read in the space of scaled moments.

\begin{example}[BC as VIV in a binary choice model \protect\label{exa:binarychoice_BC}]
    Recall the binary choice model of Example~\ref{exa:binarychoice}, with rescaled moments $g(\theta)=E_{n}[f(\theta X_{i}'1_{m})X_{i}X_{i}']^{-1}s(\theta1_{m})$ obeying $\sqrt{n}\,g(\theta_{0})=b$. The first-order expansion of these moment conditions around $\theta_{0}$ takes an IV-like form. Coordinate $j$ of the expansion can be written
    \[
        g_{j}(\theta)= \delta_{j}-\theta\,\pi_{j}+o_{p}(n^{-1/2}),
    \]
    where $\pi_j$ is the $j$th entry in the $m$-vector of population first stages
    \[
         \pi:=-G=E_{\infty}[f(\theta_{0}X_{i}'1_{m}) X_{i}X_{i}']^{-1}E_{\infty}\!\left[\frac{f(\theta_{0}X_{i}'1_{m})^{2}}{F(\theta_{0}X_{i}'1_{m})(1-F(\theta_{0}X_{i}'1_{m}))}X_{i}X_{i}'\right]1_{m},
    \]
    and $\delta_{j}=g_{j}(\theta_{0})+\theta_{0}\pi_{j}$ is the $j$th reduced form. This expansion holds uniformly over the $O_{p}(n^{-1/2})$ neighborhood of $\theta_{0}$ that contains $\hat{\theta}$.

    One can form a VIV-like estimator by replacing $E_{\infty}$ with the corresponding sample averages and $\theta_{0}$ with the GMM estimate $\hat{\theta}$. Doing so yields estimated first stages $\hat{\pi}$ and reduced forms $\hat{\delta}=\hat{g}(\hat{\theta})+\hat{\theta}\,\hat{\pi}$.
    As in Example~\ref{exa:VIV}, one can think of GMM as fitting a line to a scatter of moments via GLS and the bias correction as allowing for an intercept. With identity weights ($W=I_m$), the through-origin slope that arises from regressing $\hat{\delta}_{j}$ on $\hat{\pi}_{j}$ agrees with the GMM estimate $\hat{\theta}$ up to $o_p(n^{-1/2})$. Including an intercept yields a slope that agrees with $\hat{\theta}_{BC}$ to the same order. In contrast to Example~\ref{exa:VIV}, these equivalences are first-order rather than exact because the binary-choice moment is nonlinear and the line is fit to its linearization.
\end{example}


\subsection{Convergence rate of the BC estimator \protect\label{subsec:bc-rate}}

The following regularity conditions ensure that $\hat{\theta}_{BC}$
is well behaved as $n$ grows large.
\begin{assumption}
\label{assu:rate}(Rate conditions). The Jacobian $G$ and weighting
matrix $W$ satisfy:

i) $\left\Vert \Lambda1_{m}\right\Vert =O(1)$.

ii) $\left\Vert \Lambda\right\Vert _{2}=O(1)$.
\end{assumption}
Condition (i) bounds the magnitude of the bias correction
associated with a unit perturbation of all the moments. Condition
(ii) imposes an operator-norm bound on the GMM sensitivity matrix
$\Lambda$, ensuring that no single moment exerts unbounded influence
on the parameter estimate.

The following proposition establishes that $\hat{\theta}_{BC}$ converges
to $\theta_{0}$ at the usual parametric rate.
\begin{prop}
\label{prop:bc-rate}Under Assumptions~\ref{assu:random_b},
\ref{assu:regime}, and \ref{assu:rate}, the feasible bias-corrected
estimator satisfies
\[
\hat{\theta}_{BC}-\theta_{0}=O_{p}\left(\frac{1}{\sqrt{n}}\right).
\]
\end{prop}
\begin{proof}
See appendix.
\end{proof}
The BC adjustment $-(\hat{\mu}/\sqrt{n})\hat{\Lambda}1_{m}$ is $O_{p}(1/\sqrt{n})$, the order of the sampling noise, because $\hat{\mu}=O_{p}(1)$ by Lemma~\ref{lem:mu_consistency} and $\|\hat{\Lambda}1_{m}\|=O_{p}(1)$. This last bound combines the population bound $\|\Lambda1_{m}\|=O(1)$ of Assumption~\ref{assu:rate}.i with first-stage convergence: $\|(\hat{\Lambda}-\Lambda)1_{m}\|\le\sqrt{m}\,\|\hat{\Lambda}-\Lambda\|_{2}=\sqrt{m}\cdot O_{p}(1/\sqrt{n})=O_{p}(\sqrt{m/n})=o_{p}(1)$ by Assumption~\ref{assu:regime}.iii. Since $\hat{\mu}-\mu=o_{p}(1)$, replacing $\mu$ with $\hat{\mu}$ perturbs the adjustment by only $o_{p}(1/\sqrt{n})$, ensuring asymptotic equivalence with an estimator using the infeasible correction $-(\mu/\sqrt n)\hat\Lambda 1_m$.

\subsection{Empirical Bayes estimator \protect\label{subsec:eb-shrinkage}}

The BC estimator removes the leading bias of GMM but leaves a first-order
error with predictable structure. Removing this predictable component of the error can improve precision.

The proof of Lemma~\ref{lem:mu_consistency} establishes that $\hat{\mu}=w'r_{n}+o_{p}(1)$, where $w:=\left\Vert M1_{m}\right\Vert ^{-2}M'M1_{m}$ is a vector of population weights.
Consequently, the BC estimation error can be written
\[
\sqrt{n}(\hat{\theta}_{BC}-\theta_{0})=\Lambda r_{n}-\hat{\mu}\Lambda1_{m}+o_{p}(1)\overset{d}{\rightarrow}\Lambda M_{\mu}r,
\]
where $M_{\mu}:=I_{m}-1_{m}w'$ is the population centering matrix. Since $w'1_m=1$, $M_\mu$ annihilates the constant vector ($M_{\mu}1_{m}=0$). The BC estimation error is governed by the centered moment $\dot{r}:=M_{\mu}r$.
The GMM fitting process masks this quantity, revealing only $\ddot{r} :=M \dot{r}$, which has mean zero and covariance $M\Sigma_{\mu}M'$ with $\Sigma_{\mu}:=M_{\mu}\Sigma M_{\mu}'$.

The best linear predictor (BLP) of $\dot{r}$ given $\ddot{r}$ is $\Pi\ddot{r}$, where
\[
    \Pi:=\Sigma_{\mu}M'(M\Sigma_{\mu}M')^{+}M.
\]
The $+$ superscript denotes the Moore--Penrose inverse, which is required here because $M\Sigma_{\mu}M'$ is singular with rank $m-p-1$. Since the residual maker $M$ is idempotent, $\Pi M=\Pi$ and $\Pi\ddot{r}=\Pi\dot{r}=\Pi M_{\mu}r$. If $r$ is normally distributed, then $\Pi\ddot{r}$ is also the minimum mean squared error predictor of $\dot{r}$.

The BLP $\Pi\ddot{r}$ can be thought of as the fitted values from an infeasible regression of $\dot{r}$ on $\ddot{r}$ with no intercept. As noted by \textcite{stigler19901988}, these fitted values offer a shrunken, less variable image of the dependent variable $\dot{r}$. In this sense, $\Pi$ is a shrinkage operator. Accordingly, the residual vector $\dot{r} - \Pi \ddot{r}$ from this regression has lower variance than $\dot{r}$ whenever $\Sigma_\mu M'\neq0$ (i.e., whenever $\ddot{r}$ and $\dot{r}$ are correlated).

The corresponding BLP of the leading-term error $\Lambda \dot{r}$ is $\Lambda\Pi\ddot{r}$.  Subtracting this prediction from the scaled BC estimator yields an infeasible corrected estimator
\[
\hat{\theta}_{BC}-\frac{1}{\sqrt{n}}\Lambda\Pi\ddot{r} = \theta_0 + \frac{1}{\sqrt{n}} \Lambda \left(\dot{r} - \Pi \ddot{r}\right) + o_p(1 / \sqrt{n}).
\]
Since $\dot{r} - \Pi \ddot{r}$ exhibits weakly lower variance than $\dot{r}$, this infeasible estimator has weakly lower first-order variance than $\hat{\theta}_{BC}$.

To construct a feasible version of this estimator, I replace population quantities with sample analogues and follow the empirical Bayes principle of plugging in the estimated hyper-parameters $(\hat \mu, \widehat{\sigma^2})$ for their unknown counterparts. The matrix $M_{\mu}$ is replaced with $\hat{M}_{\mu}:=I_{m}-1_{m}\hat{w}'$, where
$\hat{w}:=\hat{M}'\hat{M}1_{m}/\|\hat{M}1_{m}\|^{2}$. Likewise, $\Sigma_{\mu}$ is replaced with $\hat{\Sigma}_{\mu}:=\hat{M}_{\mu}\hat{\Sigma}\hat{M}_{\mu}'$,
where $\hat{\Sigma}=\hat{V}+\widehat{\sigma^{2}}Q$ estimates the total
moment variance. Finally, the vector $M_{\mu}r$ has feasible counterpart $\hat{r}-\hat{\mu}1_{m}$, where $\hat{r}$ estimates the masked moment $Mr$. Hence the plug-in BLP is $\hat{\Pi}(\hat{r}-\hat{\mu}1_{m})$, where
\[
\hat{\Pi}:=\hat{\Sigma}_{\mu}\hat{M}'(\hat{M}\hat{\Sigma}_{\mu}\hat{M}')^{+}\hat{M}.
\]
The trailing $\hat{M}$ in $\hat{\Pi}$ accounts for the effects of the masking.

The corresponding plug-in predictor of the first-order error in the BC estimator is $\frac{1}{\sqrt{n}}\hat{\Lambda}\hat{\Pi}(\hat{r}-\hat{\mu}1_{m})$. Subtracting the predicted first-order error from
the BC estimate yields the empirical Bayes (EB) estimator
\[
\hat{\theta}_{EB}:=\hat{\theta}_{BC}-\frac{1}{\sqrt{n}}\hat{\Lambda}\hat{\Pi}(\hat{r}-\hat{\mu}1_{m}).
\]
The EB estimator is a regression adjustment of BC, or equivalently, of its influence function. Intuitively, removing the predictable component of the BC estimator's first-order error should improve its precision. I formalize this intuition in Section \ref{subsec:ordering}.



\begin{rem}[EB when $\mu$ is unidentified]\label{rem:eb-pure-shrinkage}
When $\|\hat{M}1_{m}\|^{2}\approx0$, the mean $\mu$ is unidentified. Once $\|\hat{M}1_{m}\|^{2}$ falls below a small threshold I set $\hat{\mu}:=0$ and $\hat{w}:=0$, replacing $\hat{M}_{\mu}$ by $I_{m}$. The BC estimator then coincides with GMM. The EB estimator remains well-defined, with shrinkage operator $\hat{\Pi}=\hat{\Sigma}\hat{M}'(\hat{M}\hat{\Sigma}\hat{M}')^{+}\hat{M}$, yielding a pure shrinkage correction with no mean adjustment. The theoretical results below will assume $\mu$ is identified.
\end{rem}

\subsection{Convergence rate of the EB estimator}

The EB estimator requires some additional regularity conditions to deal with the possibility that some moment directions become weakly identified after centering.
\begin{assumption}[Empirical Bayes regularity]\label{assu:eb-reg}
Define the projection matrix $\hat{\Pi}_{\sigma}:=\hat{\Sigma}_{\mu}^{\sigma}\hat{M}'(\hat{M}\hat{\Sigma}_{\mu}^{\sigma}\hat{M}')^{+}\hat{M}$, where $\hat{\Sigma}_{\mu}^{\sigma}:=\hat{M}_{\mu}(\hat{V}+\sigma^{2}Q)\hat{M}_{\mu}'$ is an oracle version of $\hat \Sigma_\mu$ using knowledge of $\sigma^2$. The centered moment covariance $M\Sigma_{\mu}M'$, $\hat{\Pi}_{\sigma}$, and $\Pi$ satisfy:

i) $\mathrm{tr}\big(((M\Sigma_{\mu}M')^{+})^{2}\big)=O(m)$;

ii) $E\big\|(\hat{\Pi}_{\sigma}-\Pi)M_{\mu}r_{n}\big\|^{2}=O(m^{2}/n)$;

iii) $\big\|(I_{m}-\Pi)M_{\mu}\big\|_{2}=O(1)$.
\end{assumption}
Condition (i) bounds the sum of the squared inverse eigenvalues of $M\Sigma_{\mu}M'$ to order $m$. This caps the weak identification left after centering: a growing number of mildly weak directions is allowed, but no single direction is allowed to vanish faster than $m^{-1/2}$. Condition (ii) requires $\hat{\Pi}_{\sigma}$ (the oracle version of $\hat{\Pi}$) to track its population counterpart $\Pi$ in mean square. This holds when the first-stage error avoids the weakly identified directions. By Markov's inequality, condition (ii) implies the predictor-error rate $\|(\hat{\Pi}_{\sigma}-\Pi)M_{\mu}r_{n}\|=O_{p}(m/\sqrt{n})$ that the appendix proofs invoke. Condition (iii) ensures the population shrinkage residual operator is bounded.

The following proposition establishes the approximation error of the
plug-in predictor $\hat{\Pi}(\hat{r}-\hat{\mu}1_{m})$ and the convergence rate of the empirical
Bayes estimator.
\begin{prop}
\label{prop:eb-rate}Under Assumptions~\ref{assu:random_b},
\ref{assu:regime}, \ref{assu:rate}, and \ref{assu:eb-reg}:\\
 (i) The feasible predictor satisfies
\[
\frac{1}{\sqrt{m}}\|\hat{\Pi}(\hat{r}-\hat{\mu}1_{m})-\Pi M_{\mu}r_{n}\|=O_{p}\left(\frac{1}{\sqrt{m}}\right)+O_{p}\left(\sqrt{\frac{m}{n}}\right).
\]
\\
(ii) The empirical Bayes estimator satisfies
\[
\hat{\theta}_{EB}-\theta_{0}=O_{p}\left(\frac{1}{\sqrt{n}}\right).
\]
\end{prop}
\begin{proof}
See appendix.
\end{proof}
Part (i) bounds the approximation error of the feasible predictor relative to an oracle that knows the population matrices $\left(\Pi,M_{\mu}\right)$ and the random vector $r_n$. This error converges at the same rate as the hyperparameters described in Lemmas~\ref{lem:mu_consistency} and \ref{lem:sigma-consistency}. As in those results, the approximation error scaled by $m^{-1/2}$ vanishes when $m$ grows more slowly than $n$.

Part
(ii) establishes that $\hat{\theta}_{EB}$ converges at the parametric
rate, matching the feasible bias-corrected estimator of
Proposition~\ref{prop:bc-rate}. The hyperparameter errors $\hat{\mu}-\mu$
and $\hat{\sigma}^{2}-\sigma^{2}$ are $o_{p}(1)$ by
Lemmas~\ref{lem:mu_consistency} and \ref{lem:sigma-consistency}. Because $\hat{\Pi}1_{m}=0$, the EB adjustment reduces to $-\hat{\Lambda}\hat{\Pi}\hat{r}/\sqrt{n}$. The mean error affects $\hat{\theta}_{EB}$ only through $\hat{\theta}_{BC}$. The variance error also enters the predictor, through $\hat{\Sigma}_{\mu}$.
The $1/\sqrt{n}$ scaling, combined with the $O_{p}(1)$ bound on
$\|\hat{\Lambda}\|_{2}$ from Assumption~\ref{assu:rate}.ii and the bound on the error in the feasible predictor established in part (i), ensures these hyperparameter errors do not compromise the
parametric rate.

\subsection{Variance ordering \protect \label{subsec:ordering}}

Define the leading-term risk of an estimator
as the MSE of the leading linear term in its $\sqrt{n}$-scaled estimation
error. Part (i) of the
following proposition establishes that the EB estimator weakly improves on the leading-term risk of the BC estimator. Part (ii) characterizes the leading term of the EB estimator under a mild nondegeneracy condition.
\begin{prop}
\label{prop:variance-ordering}Under Assumptions~\ref{assu:random_b} and \ref{assu:regime}:\\
(i) The bias-corrected and empirical Bayes estimators have leading-term
risks $V_{BC}=\Lambda\Sigma_{\mu}\Lambda'$ and
$V_{EB}=\Lambda(I-\Pi)\Sigma_{\mu}(I-\Pi)'\Lambda'$. The difference
\[
V_{BC}-V_{EB}=\Lambda\Sigma_{\mu}M'(M\Sigma_{\mu}M')^{+}M\Sigma_{\mu}\Lambda'
\]
is positive semidefinite. The leading-term risk improvement
is nonzero if and only if $M\Sigma_{\mu}\Lambda'\neq0$; when $p=1$,
nonzeroness and positive definiteness coincide, and positive definiteness
in general requires $m\geq2p+1$.\\
(ii) If, in addition, the eigenvalue $\lambda_{\min}(V)>0$, then the leading term of $\hat \theta_{EB}$ equals the influence function of efficient GMM applied to the centered moments $M_{\mu}g(\theta)$:
\[
\Lambda(I-\Pi)M_{\mu}=-\left(G'\Sigma_{\mu}^{+}G\right)^{-1}G'\Sigma_{\mu}^{+}M_{\mu},
\]
and consequently
\[
V_{EB}=\left(G'\Sigma_{\mu}^{+}G\right)^{-1}=\left(G'\Sigma^{-1}G-\frac{G'\Sigma^{-1}1_{m}1_{m}'\Sigma^{-1}G}{1_{m}'\Sigma^{-1}1_{m}}\right)^{-1}.
\]
Neither expression depends on the weighting matrix $W$.
\end{prop}
\begin{proof}
See appendix.
\end{proof}
The risk improvement offered by the EB estimator is governed by the matrix $M\Sigma_{\mu}\Lambda'$. Writing $\Sigma_{\mu}=M_{\mu}VM_{\mu}'+\sigma^{2}M_{\mu}M_{\mu}'$ splits this matrix into a noise term and a dispersion term. Each of these terms provides a potential source of risk improvement. For the dispersion term to contribute to EB's advantage over BC, the model must be misspecified ($\sigma^{2}>0$) with $M\Sigma_{\mu}\Lambda'\neq0$.

Since the noise term survives at $\sigma^{2}=0$, the proposition implies EB can improve on BC even under correct specification. Part (ii) of the proposition shows EB is asymptotically equivalent to efficient GMM applied to the centered moment conditions $M_{\mu}g(\theta)$ at every $\sigma^{2}\ge0$, regardless of the weighting matrix $W$. As a result, inefficient weighting inflates the leading-term variance of GMM and BC but not EB, widening the scope for risk improvements.

Another potential source of improvements under correct specification stems from the structure of the moment noise matrix $V$. With efficient GMM weighting, the EB risk improvement vanishes under homoscedastic moment noise ($V\propto I_m$), while heteroscedastic moment noise generally leaves it nonzero. This pattern reflects that the choice $W=V^{-1}$ is suboptimal for the centered moments unless the noise is homoscedastic.

A closely related variance ordering underlies the inadmissibility result of \textcite{brown1990ancillarity}, who considered the problem of estimating the intercept of a linear regression with random (ancillary) controls. In that context, as here, exploiting high-dimensional ancillary controls can reduce the risk of a low-dimensional target, even a scalar one. Indeed, the centered sample moment $\hat{r}-\hat{\mu}1_{m}$ is asymptotically ancillary for $\theta_0$, providing a formal connection with the framework of \textcite{brown1990ancillarity}.

\begin{rem}[Uniform integrability and moments]
Proposition~\ref{prop:variance-ordering} ranks $\hat{\theta}_{BC}$ and
$\hat{\theta}_{EB}$ by leading-term risk. When the errors $\sqrt{n}(\hat{\theta}_{EB}-\theta_{0})$ and
$\sqrt{n}(\hat{\theta}_{BC}-\theta_{0})$ are uniformly square-integrable, the two
estimators can be ranked in limiting MSE.
\end{rem}

\begin{rem}[Linear models and global misspecification]\label{rem:linear-global-short}
    When $g$ is affine in $\theta$ and the weighting matrix does not depend on $\theta_0$, the representations
    $\hat\theta_{BC}-\theta_0=\hat\Lambda\hat M_\mu\hat g(\theta_0)$ and
    $\hat\theta_{EB}-\theta_0=\hat\Lambda(I_m-\hat\Pi)\hat M_\mu\hat g(\theta_0)$ hold exactly. Since $\hat M_\mu 1_m=0$, both errors are invariant to common shifts $\hat g(\theta_0)\mapsto
    \hat g(\theta_0)+c\,1_m$. Consequently the common-mean component of the misspecification bias is
    purged, as with the differencing estimator of Section~\ref{sec:A-differencing-approach}. The variance ordering of Proposition~\ref{prop:variance-ordering} then holds exactly, regardless
    of the scale of the specification errors. However, the
    residual idiosyncratic bias is of the same order as that scale. Hence, consistency for $\theta_{0}$ still requires the specification errors to vanish as $n$ grows large.
    \end{rem}


\section{\protect\label{sec:distribution-theory}Distribution theory}

This section sharpens the results of Propositions~\ref{prop:bc-rate}
and~\ref{prop:eb-rate} to distributional convergence, which provides
a basis for misspecification-aware inference. The distributional convergence
results are established in Section \ref{subsec:CLTs}. Section \ref{subsec:Inference} discusses standard error estimation and inference. Section \ref{subsec:Diagnostic} develops two specification tests
that can be used to assess the plausibility of the exchangeability assumption.

\subsection{Central limit theorems \protect\label{subsec:CLTs}}

Define the centered moment vector $\eta_{n}:=r_{n}-\mu1_{m}=(b-\mu1_{m})+\varepsilon_{n}$. Both the BC and EB estimators studied below are, to leading order, studentized linear functions of $\eta_{n}$ formed from a centered sensitivity matrix. I first state a general central limit
theorem for any such estimator, then specialize it to the two of interest.

\subsubsection{General result}
Consider a generic estimator $\hat{\theta}_{\star}$ whose scaled error admits a linear expansion
$\sqrt{n}(\hat{\theta}_{\star}-\theta_{0})=\Lambda_{\star}\eta_{n}+\tilde{R}_{\star}$, where $\Lambda_{\star}$
is a $p\times m$ sensitivity matrix with $\Lambda_{\star}1_{m}=0$ and $\tilde{R}_{\star}$ is a remainder.
Write $V_{\star}:=\Lambda_{\star}\Sigma \Lambda_{\star}'$ for the variance of the leading term $\Lambda_{\star}\eta_{n}$.
The following assumption restricts how this leading-term variance grows with the sample size.
\begin{assumption}
\label{assu:lindeberg}The sensitivity matrix $\Lambda_{\star}$ satisfies $\Lambda_{\star}1_{m}=0$. Its
leading-term variance $V_{\star}$ and the realized variance $\sigma^{2}$ satisfy:

i) $V_{\star}$ and $\Lambda_{\star}V\Lambda_{\star}'$ are positive definite for sufficiently large $m$, and $\left\Vert \Lambda_{\star}\right\Vert _{2}=O(1)$.

ii) $(\Lambda_{\star}V\Lambda_{\star}')^{-1/2}\Lambda_{\star}\varepsilon_{n}\overset{d}{\rightarrow}N(0,I_{p})$.

iii) $\max_{1\le j\le m}\Lambda_{\star,j}'V_{\star}^{-1}\Lambda_{\star,j}\to0$,
where $\Lambda_{\star,j}$ denotes the $j$-th column of $\Lambda_{\star}$.

iv) $\lambda_{\min}(V_{\star})\cdot n/m\to\infty$, where $\lambda_{\min}\left(\cdot\right)$
denotes the smallest eigenvalue of the input matrix.

v) $\sigma^{2}=O_{p}(1)$, and $\lambda_{\min}(V_{\star})$ is bounded away from zero.
\end{assumption}
Condition (i) ensures that studentization is possible. Condition (ii)
is a high-level triangular-array central limit theorem for the noise vector $\Lambda_{\star}\varepsilon_{n}$, standardized by its own variance to a fixed $N(0,I_{p})$ limit. With i.i.d.\ moment contributions obeying the standard finite moment requirements
for cross-sectional GMM \parencite[e.g.,][]{newey1994large}, the
condition can be delivered by appeal to the multivariate Lindeberg-Feller
CLT. Under weakly dependent or time-series moment contributions, alternative
central limit theorems for mixingales or martingale-difference sequences
can deliver the result under analogous moment conditions. Condition
(iii) requires that no single moment dominate the studentized sensitivity
matrix $V_{\star}^{-1/2}\Lambda_{\star}$. This condition can equivalently be written
as $\max_{j}\Vert V_{\star}^{-1/2}\Lambda_{\star,j}\Vert^{2}\to0$, which
is the standard Lindeberg-type requirement for a central limit theorem
on weighted sums of exchangeable variables. Condition (iv) requires
the smallest eigenvalue of $V_{\star}$ to be of larger order than $m/n$,
ensuring first-stage noise is asymptotically negligible after studentization. This condition can be scrutinized in applications by inspecting the sample
analog $\lambda_{\min}(\hat{V}_{\star})\cdot n/m$. Condition (v) requires the realized variance to be bounded in probability and the leading-term variance to be nondegenerate. Because $\sigma^{2}\Lambda_{\star}Q\Lambda_{\star}'$ is positive semidefinite, $\lambda_{\min}(V_{\star})\ge\lambda_{\min}(\Lambda_{\star}V\Lambda_{\star}')$ at every realized variance. Conditions (iii) and (iv) can therefore be checked on nonrandom quantities even though $V_{\star}$ is random.

The following theorem shows that any estimator satisfying these conditions (with a negligible remainder) is asymptotically normal after studentization.
\begin{thm}
\label{thm:clt}Suppose $\sqrt{n}(\hat{\theta}_{\star}-\theta_{0})=\Lambda_{\star}\eta_{n}+\tilde{R}_{\star}$,
where $\Lambda_{\star}$ satisfies Assumption~\ref{assu:lindeberg} and $V_{\star}^{-1/2}\tilde{R}_{\star}=o_{p}(1)$.
Then under Assumptions~\ref{assu:random_b} and \ref{assu:regime},
\[
\sqrt{n}\,V_{\star}^{-1/2}\left(\hat{\theta}_{\star}-\theta_{0}\right)\overset{d}{\rightarrow}N\left(0,I_{p}\right).
\]
\end{thm}
\begin{proof}
See appendix.
\end{proof}
The theorem establishes that the leading term
$\Lambda_{\star}\eta_{n}$ of such an estimator is asymptotically normal. This term splits into a noise part and a specification part.
The noise part $\Lambda_{\star}\varepsilon_{n}$ is covered by condition (ii). The specification
part $\Lambda_{\star}(b-\mu1_{m})$ is covered by a central limit theorem for permutations when the realized variance is bounded away from zero, and is negligible after studentization when the realized variance shrinks to zero. The permutation argument uses the exchangeability and bounded fourth moments of Assumption~\ref{assu:random_b} and the no-dominant-column condition of Assumption~\ref{assu:lindeberg}.iii.

Note that the statistic in Theorem~\ref{thm:clt} is self-normalizing---the studentizer $V_{\star}$ is formed at the realized $\sigma^{2}$---and the limit is unconditional, with the probability integrating over both $b$ and $\varepsilon_{n}$. The realized variance may stay away from zero, drift toward zero, or fluctuate. The limit is $N(0,I_{p})$ in every case, and the normal approximation does not deteriorate near the correctly specified boundary $\sigma^{2}=0$. Section~\ref{subsec:Inference} discusses the sense in which the limit also holds conditional on $(\mu,\sigma^{2})$.

\subsubsection{Specialization to BC and EB}
The BC and EB estimators both admit leading-term expansions of the sort described by Theorem~\ref{thm:clt}. The BC
sensitivity matrix is $\Lambda_{BC}:=\Lambda M_{\mu}$ and the EB sensitivity matrix is
$\Lambda_{EB}:=\Lambda(I_{m}-\Pi)M_{\mu}$, with leading-term variances
$V_{BC}:=\Lambda_{BC}\Sigma\Lambda_{BC}'$ and $V_{EB}:=\Lambda_{EB}\Sigma\Lambda_{EB}'$.
Both satisfy $\Lambda_{BC}1_{m}=\Lambda_{EB}1_{m}=0$.

Since $V_{BC}=\Lambda_{BC}V\Lambda_{BC}'+\sigma^{2}\Lambda_{BC}\Lambda_{BC}'$ retains the pure sampling-noise term $\Lambda_{BC}V\Lambda_{BC}'$ at every $\sigma^{2}$, it is nondegenerate whenever the noise term is. Establishing nondegeneracy of the EB variance is subtler, as shrinkage (unlike bias correction) removes variance. The following lemma shows that $V_{EB}$ is also nondegenerate at any level of the realized variance and for any admissible weighting.
\begin{lem}\label{lem:eb-floor}
Suppose Assumptions~\ref{assu:random_b} and \ref{assu:regime} hold, $\lambda_{\min}(V)\ge c_{V}>0$, and $G'V^{-1}G=O(1)$. Then, for every $\sigma^{2}\ge0$ and every weighting matrix admitted by Assumption~\ref{assu:regime},
\[
\lambda_{\min}(V_{EB})\ge\frac{1}{\lambda_{\max}(G'V^{-1}G)}.
\]
The bound is free of $\sigma^{2}$ and of the weighting matrix.
\end{lem}
\begin{proof}
See appendix.
\end{proof}
The bound, which follows from standard Gauss--Markov reasoning, holds for every weighting matrix $W$, not only the efficient one. Thus, the EB estimator satisfies the nondegeneracy requirement of Assumption~\ref{assu:lindeberg}.v.

The following corollary gives conditions under which the BC and EB expansions have an asymptotically negligible remainder, in which case Theorem~\ref{thm:clt} applies.
\begin{cor}
\label{corr:bc-eb-clt}Suppose Assumptions~\ref{assu:random_b}, \ref{assu:regime},
and \ref{assu:rate} hold.
\begin{enumerate}
\item[(i)] If $\Lambda_{BC}$ satisfies Assumption~\ref{assu:lindeberg}, then
$\sqrt{n}\,V_{BC}^{-1/2}(\hat{\theta}_{BC}-\theta_{0})\overset{d}{\rightarrow}N(0,I_{p})$.

\item[(ii)] If, in addition, Assumption~\ref{assu:eb-reg} holds, $\Lambda_{EB}$ satisfies Assumption~\ref{assu:lindeberg},
$\|V_{EB}^{-1/2}\Lambda\|_{2}=O(1)$, and $\mathrm{tr}(V_{EB}^{-1}\Lambda_{EB}K\Lambda_{EB}')=O(1)$
with $K:=M_{\mu}'M'(M\Sigma_{\mu}M')^{+}MM_{\mu}$, then
$\sqrt{n}\,V_{EB}^{-1/2}(\hat{\theta}_{EB}-\theta_{0})\overset{d}{\rightarrow}N(0,I_{p})$.
\end{enumerate}
\end{cor}
\begin{proof}
See appendix.
\end{proof}
Part (ii) adds two conditions on the plug-in error in $\hat{\Pi}$. The condition $\|V_{EB}^{-1/2}\Lambda\|_{2}=O(1)$ controls the first-stage error in $\hat{M}$ and $\hat{V}$. The trace condition $\mathrm{tr}(V_{EB}^{-1}\Lambda_{EB}K\Lambda_{EB}')=O(1)$ controls the $\hat{\sigma}^{2}$-correction. These hold when $\lambda_{\min}(V_{EB})$ and $\lambda_{\min}(\Sigma)$ are bounded away from zero. Lemma~\ref{lem:eb-floor} delivers the first uniformly in $\sigma^{2}\ge0$ when $\lambda_{\min}(V)\ge c_{V}$ and $G'V^{-1}G=O(1)$. At $\sigma^{2}=0$ the second reduces to $\lambda_{\min}(V)\ge c_{V}$, since $\Sigma=V$. The conditions of Assumption~\ref{assu:eb-reg} are required across all values of $\sigma^{2}$. The squared-trace condition binds at $\sigma^{2}=0$, where the nonzero eigenvalues of $M\Sigma_{\mu}M'$ are smallest. These restrictions can be checked using sample analogs.

\subsection{Inference \protect\label{subsec:Inference}}

Feasible inference can be conducted using plug-in versions of the leading-term variances.
With $\hat{M}_{\mu}$ and $\hat{\Pi}$ as defined in Section~\ref{sec:bias-estimation},
let $\hat{\Lambda}_{BC}:=\hat{\Lambda}\hat{M}_{\mu}$
and $\hat{\Lambda}_{EB}:=\hat{\Lambda}\left(I_{m}-\hat{\Pi}\right)\hat{M}_{\mu}$.
The feasible variance estimators are $\hat{V}_{BC}:=\hat{\Lambda}_{BC}\hat{\Sigma}\hat{\Lambda}_{BC}'$
and $\hat{V}_{EB}:=\hat{\Lambda}_{EB}\hat{\Sigma}\hat{\Lambda}_{EB}'$,
with $\hat{\Sigma}:=\hat{V}+\widehat{\sigma^{2}}Q$.
\begin{cor}
\label{corr:feasible-inference}Suppose the assumptions of Corollary~\ref{corr:bc-eb-clt}(i)
hold, $\lambda_{\min}(V_{BC})\sqrt{n/m}\to\infty$, and $\|V_{BC}^{-1/2}\Lambda_{BC}\|_{2}=O(1)$. Then
\[
\sqrt{n}\,\hat{V}_{BC}^{-1/2}\left(\hat{\theta}_{BC}-\theta_{0}\right)\overset{d}{\rightarrow}N\left(0,I_{p}\right).
\]
Likewise, if the assumptions of Corollary~\ref{corr:bc-eb-clt}(ii) hold
and $\lambda_{\min}(V_{EB})\sqrt{n/m}\to\infty$, then
\[
\sqrt{n}\,\hat{V}_{EB}^{-1/2}\left(\hat{\theta}_{EB}-\theta_{0}\right)\overset{d}{\rightarrow}N\left(0,I_{p}\right).
\]
\end{cor}
\begin{proof}
See appendix.
\end{proof}
The rate conditions $\lambda_{\min}(V_{BC})\sqrt{n/m}\to\infty$
and $\lambda_{\min}(V_{EB})\sqrt{n/m}\to\infty$ hold when the smallest
eigenvalues of $V_{BC}$ and $V_{EB}$ are bounded away from zero. Lemma~\ref{lem:eb-floor} supplies the EB floor. For BC, $V_{BC}-\sigma^{2}\Lambda_{BC}\Lambda_{BC}'=\Lambda_{BC}V\Lambda_{BC}'$ is positive semidefinite at every $\sigma^{2}\ge0$, and $V_{BC}$ is bounded below whenever $V$ is bounded below on the row space of $\Lambda_{BC}$. The same floors give $\|V_{BC}^{-1/2}\Lambda_{BC}\|_{2}=O(1)$ and $\|V_{EB}^{-1/2}\Lambda_{EB}\|_{2}=O(1)$, which control the $\widehat{\sigma^{2}}$ contribution to the variance estimates.

Asymptotic $100(1-\gamma)\%$ confidence intervals for the $k$-th
component of $\theta_{0}$ take the form $\hat{\theta}_{BC,k}\pm z_{1-\gamma/2}\sqrt{\hat{V}_{BC,kk}/n}$
and $\hat{\theta}_{EB,k}\pm z_{1-\gamma/2}\sqrt{\hat{V}_{EB,kk}/n}$,
respectively. Under the regularity conditions of Corollary~\ref{corr:feasible-inference}, these intervals satisfy
\[
\begin{aligned}\Pr\!\left[\theta_{0,k}\in\hat{\theta}_{BC,k}\pm z_{1-\gamma/2}\sqrt{\hat{V}_{BC,kk}/n}\,\Big|\,\mu,\sigma^{2}\right] & \overset{p}{\rightarrow}1-\gamma,\\
\Pr\!\left[\theta_{0,k}\in\hat{\theta}_{EB,k}\pm z_{1-\gamma/2}\sqrt{\hat{V}_{EB,kk}/n}\,\Big|\,\mu,\sigma^{2}\right] & \overset{p}{\rightarrow}1-\gamma,
\end{aligned}
\]
for each $k=1,\ldots,p$ and $\gamma\in(0,1)$, where the probability
is taken over the joint distribution of $(b,\varepsilon_{n})$ given
$(\mu,\sigma^{2})$. Conditional coverage is random because the hyperparameters are computed from $b$. The convergence holds in probability: hyperparameter values at which coverage departs from $1-\gamma$ can exist but have vanishing probability. Averaging over the hyperparameters gives unconditional coverage of $1-\gamma$. The contributions $\widehat{\sigma^{2}}\hat{\Lambda}_{BC}\hat{\Lambda}_{BC}'$
and $\widehat{\sigma^{2}}\hat{\Lambda}_{EB}\hat{\Lambda}_{EB}'$ to the
variances make these intervals \textit{misspecification-aware}: they
capture the conditional cross-moment variability in $b$ that standard
GMM intervals ignore.
\begin{rem}[Uniform conditional coverage]\label{rem:pointwise-conditional}
Theorem~\ref{thm:clt} delivers conditional coverage in probability. Coverage at every value of $(\mu,\sigma^{2})$ requires additional restrictions ruling out the possibility that a large $\sigma^{2}$ reflects a single large specification error. A simple sufficient condition is the conditional moment bound $E[(b_{j}-\mu)^{4}\mid\mu,\sigma^{2}]\le C$.
\end{rem}

\subsection{Specification tests}\label{subsec:Diagnostic}

It is useful to have a diagnostic for whether the exchangeability restriction
of Assumption~\ref{assu:random_b}.ii is plausible in a given application. The machinery developed thus far suggests two simple tests.

\subsubsection{A Hausman-type test}
A first test for violations of exchangeability examines whether the shrinkage adjustment has the magnitude one would expect from exchangeable specification errors. Under the conditions of Corollary~\ref{corr:bc-eb-clt}(ii), the difference between the EB and BC estimators can be written:
\[
\hat{\theta}_{EB}-\hat{\theta}_{BC}=-\frac{1}{\sqrt{n}}\hat{\Lambda}\hat{\Pi}\left(\hat{r}-\hat{\mu}1_{m}\right) = -\frac{1}{\sqrt{n}} \Lambda\Pi M_{\mu}\eta_{n} + o_p(1/\sqrt{n}).
\]
Exchangeability restricts the term $M_{\mu}\eta_{n}$ to have mean zero, since $E[M_{\mu}b\mid\mu,\sigma^{2}]=\mu M_{\mu}1_{m}=0$. The leading-term variance of this contrast is the difference term of Proposition~\ref{prop:variance-ordering}(i):
\[
\Lambda\Pi\Sigma_{\mu}\Pi'\Lambda'=V_{BC}-V_{EB}.
\]
Hence, as in \textcite{hausman1978specification}, the leading-term variance of the difference equals the difference of the variances. This property follows from the EB estimator's representation in Proposition~\ref{prop:variance-ordering}(ii) as efficient GMM on the centered moment conditions.

These observations motivate the Hausman-type test statistic
\begin{align} \label{stat:Hausman}
    n\left(\hat{\theta}_{EB}-\hat{\theta}_{BC}\right)'\left(\hat{V}_{BC}-\hat{V}_{EB}\right)^{-1}\left(\hat{\theta}_{EB}-\hat{\theta}_{BC}\right),
\end{align}
where $\hat{V}_{BC}$ and $\hat{V}_{EB}$ are the variance estimators of Section~\ref{subsec:Inference}. The estimated variance difference is positive semidefinite by construction. When Assumption~\ref{assu:lindeberg} holds for the difference sensitivity $-\Lambda\Pi M_{\mu}$, Theorem~\ref{thm:clt} implies the statistic is approximately $\chi^{2}(p)$ under exchangeability. The test requires $V_{BC}-V_{EB}$ to be nondegenerate.

The statistic in \eqref{stat:Hausman} isolates the restriction that EB exploits beyond BC. Under exchangeability the shrinkage adjustment subtracts predictable noise and nothing else. In contrast, when $E[b]$ is non-exchangeable, the shrinkage adjustment shifts the difference by $-\Lambda\Pi M_{\mu}E[b]/\sqrt{n}$ to leading order. Hence, the null being tested is $H_0:-\Lambda\Pi M_{\mu}E[b]/\sqrt{n}=0$, while the alternative is $H_1:-\Lambda\Pi M_{\mu}E[b]/\sqrt{n}\neq 0$. Note that DGPs in $H_1$ will not induce a leading-term bias in $\hat \theta_{BC}$ if $\Lambda M_{\mu}E[b]=0$.

While \eqref{stat:Hausman} offers a computationally convenient diagnostic, the family of violations entertained by $H_1$ is rather narrow, indicating that this test will often be unable to detect interesting violations of exchangeability. I therefore consider a second test of exchangeability based on observed moment features.

\subsubsection{A test using moment features}
Exchangeability implies that the specification errors $b_{j}$ bear no systematic
relationship to observed features of the moment conditions that vary across $j$.
For example, in cell-based IV designs the cells typically differ in size. A significant covariance between the specification errors and cell size would therefore violate the implication of Assumption~\ref{assu:random_b}.ii that $E[b]$ is permutation invariant. The
distribution theory of Section~\ref{subsec:CLTs} suggests a simple test of this
sort of restriction.

Let $\Xi$ be an $m\times d$ matrix collecting $d$ observed moment-level features,
its $j$-th row holding the features of moment $j$, centered so that
$\Xi'1_{m}=0$. Under Assumption~\ref{assu:random_b}, conditioning on the features
leaves
\[
E\left[\Xi'b\mid\Xi\right]=\mu\,\Xi'1_{m}=0.
\]
The sample counterpart is $\Xi'\tilde{r}$, where
$\tilde{r}:=\hat{r}-\hat{\mu}\hat{M}1_{m}=\hat{M}\hat{M}_{\mu}\hat{r}$ is the
residual $\hat{r}$ after removing the estimated mean along $\hat{M}1_{m}$. This
vector should be close to zero under exchangeability.

Theorem~\ref{thm:clt} and
the arguments used in its proof imply that $\Xi'\tilde{r}$ is asymptotically
normal with variance $\Omega:=\Xi'MM_{\mu}\Sigma M_{\mu}'M'\Xi$ provided that no single moment dominates the weights $\Xi'MM_{\mu}$. This observation motivates the
test statistic
\begin{align}\label{stat:feature}
    \left(\Xi'\tilde{r}\right)'\hat{\Omega}^{-1}\left(\Xi'\tilde{r}\right),\qquad
    \hat{\Omega}:=\Xi'\hat{M}\hat{M}_{\mu}\hat{\Sigma}\hat{M}_{\mu}'\hat{M}'\Xi.
\end{align}
Under the null this statistic is approximately $\chi^{2}(d)$, and each feature
$a\in\{1,\dots,d\}$ can be assessed individually via the $t$-statistic
$\left(\Xi'\tilde{r}\right)_{a}/\sqrt{\hat{\Omega}_{aa}}$. The variance
$\hat{\Omega}$ treats $\Xi$ as fixed: this is exact when the features are
non-random and a leading-order approximation when they are precisely estimated
design quantities.

\begin{rem}[What the tests cannot detect]\label{rem:test-power}
Because $MM_{\mu}$ annihilates $\mathrm{span}(1_{m},G)$, the statistic $\Xi'\tilde{r}$ is blind to any component of the specification error lying in the span of the constant and the columns of the Jacobian. The statistic \eqref{stat:Hausman} shares this blind spot, as $M_{\mu}1_{m}=0$, $M_{\mu}G=G$, and $\Pi G=0$. A specification error proportional to a column of
$G$ is observationally equivalent to a shift in $\theta$. Hence, in the linear IV setting, neither test has power to detect a linear relationship between the specification errors and the first stage $\pi$.
\end{rem}

\section{Monte Carlo Evidence \protect\label{sec:monte-carlo}}

This section presents Monte Carlo evidence on the finite-sample performance
of the estimators developed above. The simulations are based on the
overidentified instrumental variables model of Example \ref{exa:IV_example}.
The results detail the relative performance of the GMM, BC, and EB
estimators under local misspecification.

\subsection{Simulation design\protect\label{subsec:Simulation-design}}

For $i=1 \dots n$, the data generating process (DGP) is
\begin{align*}
Y_{i} & =\alpha_{0}+\beta_{0}T_{i}+\frac{1}{\sqrt{n}}\left(\sum_{j=1}^{m}b_{j}Z_{ij}+b_{m+1}\right)+e_{i},\\
T_{i} & =\mathbf{1}\{Z_{i}^{\prime}\pi^{*}+\nu_{i}>0\},\quad e_{i}=\sigma_{e}\left(\rho\nu_{i}+\sqrt{1-\rho^{2}}\eta_{i}\right),\\
b & \sim t_{5}(\bar{\mu}\mathbf{1}_{m+1},\,\bar{\sigma}^{2}I_{m+1}),\quad\mathrm{Var}(b_{j})=\bar{\sigma}^{2},\quad\nu_{i}\overset{\mathrm{i.i.d.}}{\sim}N(0,1),\quad\eta_{i}\overset{\mathrm{i.i.d.}}{\sim}N(0,1),\\
 & \quad\quad\quad\quad\quad\quad\quad\nu_{i}\perp\eta_{i}\perp Z_{i}\perp b.
\end{align*}
The parameter  $\rho=\mathrm{Corr}(e_{i},\nu_{i})$ governs the degree
of endogeneity of  $T_{i}$, while the scale parameter  $\sigma_{e}$
controls the variance of the structural error. The vector $Z_{i}=(Z_{i1},\ldots,Z_{im})'$
is comprised of mutually exclusive binary instruments $Z_{ij}\in\{0,1\}$
obeying  $Z_{ij}Z_{i\ell}=0$ for  $j\neq\ell$. I specify
\[
\left(\begin{array}{c}
1-\sum_{j=1}^{m}Z_{ij}\\
Z_{i}
\end{array}\right)\sim\text{Multinomial}\left(1,\,\left(\begin{array}{c}
1-\sum_{j=1}^{m}q_{j}\\
q
\end{array}\right)\right),\quad q_{j}>0\text{ for \ensuremath{j=1,\dots,m}}.
\]

This design can be thought of as a setting where a single fundamental
instrument $\sum_{j=1}^{m}Z_{ij}$ has been fully interacted with
 $m$ group indicators. The term  $b_{m+1}$ in the DGP is a common baseline direct effect. Thus, the complete instrument set for the parameter vector $(\alpha_{0},\beta_{0})$ would be $\left(1,Z_{i}'\right)'=\left(1,Z_{i1},\ldots,Z_{im}\right)'$. To simplify the analysis, I will treat the intercept $\alpha_{0}$ as nuisance and target the slope $\beta_{0}=\theta_0$, making $p=1$.

Note that, unlike the judge-IV design discussed in Section \ref{subsec:Intercepts-in-linear}
--- where every observation belongs to exactly one of the $m$ instrument
categories and $\sum_{j=1}^{m}Z_{ij}=1$ --- here $\sum_{j=1}^{m}Z_{ij}\in\{0,1\}$
because the instruments are interactions of a binary fundamental instrument
with group indicators. Observations with $\sum_{j=1}^{m}Z_{ij}=0$
break the adding-up constraint.

The excludability
violations $\left(b_{1},\dots,b_{m+1}\right)$ follow a multivariate  $t$ distribution with 5 degrees
of freedom, location  $\bar{\mu}\mathbf{1}_{m+1}$, and covariance $\bar{\sigma}^{2}I_{m+1}$ (its scale matrix is set to $\frac{3}{5}\bar{\sigma}^{2}I_{m+1}$, so that the covariance equals $\bar{\sigma}^{2}I_{m+1}$). Therefore, these violations have exactly four moments. The
excludability violations are exchangeable but not  i.i.d.: the diagonal covariance matrix implies the entries are not correlated but allows for dependence at higher moments.


Both $\pi^{*}$ and  $q$ are deterministic functions of  $m$ and are
held fixed across all Monte Carlo replications. Group membership probabilities
are set to quantiles of a log-normal distribution restricted to a fixed range. I first generate  $q_{j}^{\mathrm{raw}}=\exp\!\bigl(\sigma_{q}\,\Phi^{-1}(u_{j})\bigr)$
with $\Phi$ the standard normal CDF and  $\sigma_{q}=0.83$, where the
 $u_{j}$ are evenly spaced over the fixed interval  $[\Phi(-2.231),\Phi(2.231)]$.
I then normalize the raw values so that  $\sum_{j=1}^{m}q_{j}=0.9$.
Because the range is held fixed, the cell-size spread  $\max_{j}q_{j}/\min_{j}q_{j}\approx40{:}1$
is the same for every  $m$. This is an infill design: as  $m$ grows, cells fill in
within a fixed range of sizes rather than spreading toward ever more extreme values.

The $\pi^{*}$ coefficients
use the same quantile scheme. I first generate  $\pi_{j}^{*,\mathrm{raw}}=\exp\!\bigl(\sigma_{\pi^{*}}\,\Phi^{-1}(u_{j})\bigr)$
with  $\sigma_{\pi^{*}}=0.67$, giving a fixed  $\approx20{:}1$ spread. I then rescale
the raw values so that  $\|\pi^{*}\|=3\sqrt{m/40}$,
which preserves average instrument strength as  $m$ grows. To break
correlation with group sizes,  $\pi^{*}$ values are assigned in an interleaved
pattern: odd-indexed groups receive  $\pi^{*}$ values in ascending order
and even-indexed groups in descending order.

In the baseline parameterization, I set  $\alpha_{0}=1$,  $\beta_{0}=1$,
 $\rho=0.1$,  $\bar{\mu}=8$,  $\bar{\sigma}=8$, and  $\sigma_{e}^{2}=1$.
Recall that $b$ is scaled by  $n^{-1/2}$. Hence, the coefficients
capturing excludability violations have mean  $8/\sqrt{n}$ and variance
 $64/n$.

\subsection{Moment conditions and estimator}\label{subsec:sim_moment}
To target the slope parameter $\beta_0$, I project out the nuisance intercept by demeaning the outcome, treatment, and instruments in sample to form $\tilde{Y}_{i}$, $\tilde{T}_{i}$, and $\tilde{Z}_{i}=Z_{i}-\hat q$, where $\hat q=n^{-1}\sum_{i=1}^n Z_i$. Continuing the convention introduced in Section \ref{sec:Examples}, let $E_{n}[\cdot]=E[\cdot \mid b_{1},\dots,b_{m+1}]$ denote expectations under the DGP conditional on the specification errors. The population moment condition I leverage for estimation can be written
\[
g\left(\theta\right)=S^{-1}E_{n}[(\tilde{Y}_{i}-\theta\tilde{T}_{i})\tilde{Z}_{i}], \qquad S:=E_{n}[\tilde{Z}_{i}\tilde{Z}_{i}'].
\]
Note that, following Example \ref{exa:IV_example}, I have rescaled by the inverse of the second moment matrix $S$ of residualized instruments, in order to avoid the requirement that this matrix takes a permutation equivariant form.

It is useful to rewrite this condition as $g\left(\theta\right)=\delta - \theta \pi$, where $\delta:=S^{-1}E_n[\tilde Z_i \tilde Y_i]$ is a vector of reduced form coefficients and $\pi:=S^{-1} E_n[\tilde Z_i \tilde T_i]$ the corresponding vector of first-stage coefficients. Note that
\[
E_n[\tilde Z_i \tilde Y_i] = \beta_0 E_n[\tilde Z_i \tilde T_i] + Sb / \sqrt{n}, \qquad b:=(b_{1},\dots,b_{m})',
\]
where the baseline violation $b_{m+1}$ was eliminated by demeaning. Thus,
\[
g(\beta_0) = \delta - \beta_0 \pi = b/ \sqrt{n}.
\]
The vector $b$ has a $t_5$ distribution, which is exchangeable with four finite moments. It therefore satisfies Assumption~\ref{assu:random_b}.ii. The realized hyperparameters $(\mu,\sigma^{2})$ targeted by the estimators $(\hat \mu, \hat \sigma^{2})$
are the sample mean and variance of $b$.

The population Jacobian is the $m$-vector $G=-\pi$. Since the entries of $\pi$ are distinct, $1_{m}$ is not in the column space of $G$ and $\mu$ is identified by equation \eqref{eq:identification} at each $m$. The infill design of Section~\ref{subsec:Simulation-design} holds the spread of the first stage fixed as $m$ grows. The normalized residual $m^{-1}\left\Vert M1_{m}\right\Vert ^{2}$ therefore stays bounded away from zero, converging to a positive $\kappa$. This confirms Assumption~\ref{assu:regime}.vi.

In conducting GMM estimation, I use the weighting matrix $\hat{W}=\hat S := n^{-1} \sum_{i=1}^n \tilde{Z}_{i}\tilde{Z}_{i}'$, which ensures the estimates are numerically equivalent to TSLS estimation of the residualized system. By the Frisch--Waugh--Lovell theorem, the TSLS estimate of $\beta$ in the residualized system is numerically equivalent to the TSLS slope from  estimation of $(\alpha,\beta)$ using the full instrument set $(1,Z_{i}')'$. Since the structural errors $e_i$ are homoscedastic, this TSLS weighting would be efficient under proper specification.

I verify analytically in Appendix~\ref{app:verify} that the DGP satisfies
the outstanding conditions of Assumptions~\ref{assu:regime} and \ref{assu:rate} required for consistency of the BC and EB estimators when $m$ grows with $n$. The conditions of Assumptions~\ref{assu:eb-reg} and \ref{assu:lindeberg} supporting inference are checked numerically.

\subsection{Baseline findings}

Table~\ref{tab:mc-main} reports the bias, standard deviation, and
root mean squared error (RMSE) of each estimator in a first simulation
design with  $n=10{,}000$ observations and  $m=40$ cell moments.
All results report estimation performance for the slope coefficient
 $\beta$ and are averaged across  $1{,}000$ Monte Carlo replications.
The feasible estimators use the hyperparameter estimates  $\hat{\mu}$
and  $\hat{\sigma}^{2}$, while the oracle estimators use the true
hyperparameters $\left(\mu,\sigma^{2}\right)$ based on the realized specification errors but continue to rely on the estimated matrices  $\hat{G}$,  $\hat{V}$, and  $\hat{W}$.

\begin{table}[H]
\caption{Estimator Performance ( $m=40$,  $n=10,000$) \protect\label{tab:mc-main}}

\begin{centering}
\begin{tabular}{lcccc}
\toprule
Estimator & Bias & Std & RMSE & Rej.\tabularnewline
\midrule
GMM & 0.220 & 0.185 & 0.287 & 0.386\tabularnewline
BC & 0.040 & 0.251 & 0.254 & 0.087\tabularnewline
EB & 0.038 & 0.239 & 0.242 & 0.088\tabularnewline
Oracle BC & 0.024 & 0.181 & 0.182 & 0.043\tabularnewline
Oracle EB & 0.023 & 0.178 & 0.180 & 0.046\tabularnewline
\bottomrule
\end{tabular}
\par\end{centering}
{\small\emph{Notes:}}{\small{} Bias, standard deviation, and RMSE of
 $\hat{\beta}$ across  $1{,}000$ replications,  $n=10{,}000$,  $m=40$.
Feasible estimators use the hyperparameter estimates described in
Section~\ref{sec:estimation}. ``Oracle'' estimators use the true
 $\left(\mu,\sigma^{2}\right)$ with sample matrices  $\hat{G}$,
 $\hat{V}$,  $\hat{W}$. Mean simulated  $J=90.7$ with  $39$ degrees
of freedom. The population leading-term risk ratio of EB to
BC, $V_{EB}/V_{BC}=1-\Lambda\Sigma_{\mu}M'(M\Sigma_{\mu}M')^{+}M\Sigma_{\mu}\Lambda'/(\Lambda\Sigma_{\mu}\Lambda')=0.936$,
is computed exactly from the design parameters. Rej. is the rejection rate of a two-sided
Wald test of the null hypothesis that  $\beta_{0}=1$ at the 5\% nominal
level.}{\small\par}
\end{table}

By Proposition~\ref{prop:variance-ordering}, the EB estimator improves
on the leading-term risk of BC whenever $M\Sigma_{\mu}\Lambda'\neq0$. With $p=1$, the ratio of the EB and BC leading-term risks is governed by a quadratic form in this quantity:
\[
\frac{V_{EB}}{V_{BC}}=1-\frac{\Lambda\Sigma_{\mu}M'(M\Sigma_{\mu}M')^{+}M\Sigma_{\mu}\Lambda'}{\Lambda\Sigma_{\mu}\Lambda'}.
\]
Evaluating this expression using the population matrices $(\Lambda,M,\Sigma_{\mu})$ yields 0.936, indicating a modest efficiency advantage of EB over BC in this design. The mean simulated  $J$ statistic is
90.7, far exceeding the 5\% critical value of 54.57. Faced with this
DGP, the overidentification test will tend to reject correct specification.


The sizeable specification errors present in this DGP lead the uncorrected
GMM estimator to exhibit a bias of 0.220 that accounts for the bulk
of its RMSE of 0.287. The bias-corrected estimator  $\hat{\beta}_{BC}$
reduces the bias to 0.040 and achieves an RMSE of 0.254, 12\% below
GMM. The feasible shrinkage estimator  $\hat{\beta}_{EB}$ improves
on this, attaining an RMSE of 0.242---16\% below GMM and about 5\%
below BC---by removing the predictable component of the bias-corrected
estimation error. The ratio of EB to BC Monte Carlo variances is $(0.239/0.251)^2=0.907$, close to the leading-term variance ratio of 0.936. Evidently, the asymptotic approximation provides an empirically useful guide to finite-sample behavior under this DGP.

The oracle estimators illustrate the gains available with
known hyperparameters: oracle BC reaches an RMSE of 0.182 and oracle
EB 0.180. Both oracle estimators exhibit small biases that capture the remaining TSLS many-instruments bias that would be present even under proper specification. The narrow gap in RMSE between the two oracle estimators
suggests that shrinkage yields limited gains when $\mu$ is already known. I document below that this advantage grows when the specification errors are more dispersed.

The ``Rej.'' column of Table~\ref{tab:mc-main} reports the rejection
rate of a two-sided Wald test of the null hypothesis that  $\beta_{0}=1$
at the 5\% nominal level. Thus, rejections correspond here to type
I errors. The oracle
BC and EB rejection rates are based on the following infeasible variance
matrices
\[
V_{OBC}=\hat{\Lambda}\left(\hat{V}+\sigma^{2}Q\right)\hat{\Lambda}',\quad V_{OEB}=\hat{\Lambda}\left(I_{m}-\hat{\Pi}_{\sigma}\right)\left(\hat{V}+\sigma^{2}Q\right)\left(I_{m}-\hat{\Pi}_{\sigma}\right)'\hat{\Lambda}',
\]
which use the sample sensitivity matrix  $\hat{\Lambda}$, the sample
moment-covariance matrix  $\hat{V}$, and the true {\small$\sigma^{2}$}. Here $\hat{\Pi}_{\sigma}$ is the oracle projection operator of Assumption~\ref{assu:eb-reg}.

The standard GMM Wald test over-rejects dramatically, with
a rejection rate of 38.6\%. This overrejection stems both from the
estimator's bias and use of the conventional sandwich variance estimator
that ignores misspecification. In contrast, the oracle BC and EB rejection rates are 4.3\% and 4.6\% respectively,
both close to the nominal level. The feasible BC and EB tests over-reject
modestly---8.7\% and 8.8\% respectively---reflecting the finite-sample cost of
estimating the hyperparameters  $\mu$ and  $\sigma^{2}$.

Appendix~\ref{app:differenced} reports parallel results for estimators
that difference the mean specification error away rather than estimate
it, implementing the strategy of Section~\ref{sec:A-differencing-approach},
along with the MBTSLS estimator of \textcite{kolesar2015identification}. Appendix Table \ref{tab:appc-differenced} shows that differenced GMM is nearly equivalent to levels BC and differenced EB is nearly equivalent to levels EB.  MBTSLS is biased in levels
because the nonzero mean $\bar{\mu}$ violates the orthogonality condition
it relies on. Applying MBTSLS to the differenced moment conditions yields the lowest bias of any estimator, which reflects that differenced MBTSLS removes not only the leading-term bias targeted by BC but the higher-order many-instruments bias. However, this extra bias reduction comes at a cost: the differenced MBTSLS estimator's standard deviation is nearly 50\% above that of EB, putting it at a severe RMSE disadvantage relative to both corrected estimators.

\subsection{Altering the misspecification hyperparameters}

Table \ref{tab:mc-rmse-vs-mu} examines how the relative performance of the estimators, as measured by RMSE, varies with the marginal mean specification error  $\bar{\mu}$ while holding $\bar{\sigma}=8$ and all other parameters fixed. Theoretically,
the BC, Oracle BC, EB, and Oracle EB estimators are invariant to $\bar{\mu}$ in large samples. A unit shift in $\bar{\mu}$ moves $\hat{\mu}$ by one in expectation and shifts the GMM estimate by $(1/\sqrt{n})\hat{\Lambda}1_{m}$. The bias correction $(\hat{\mu}/\sqrt{n})\hat{\Lambda}1_{m}$ removes this shift, and the shrinkage correction is likewise unaffected. Only the uncorrected GMM estimator's RMSE genuinely depends
on $\bar{\mu}$. The mild variation visible in the corrected-estimator
rows of the table reflects a mix of finite-sample deviations from asymptotic invariance and Monte Carlo noise.

\begin{table}[H]
\caption{RMSE by $\bar{\mu}$ ( $\bar{\sigma}=8$,  $n=10{,}000$,  $m=40$)
\protect\label{tab:mc-rmse-vs-mu}}

\begin{centering}
\begin{tabular}{lccccc}
\toprule
 & $\bar{\mu}=0$ & $\bar{\mu}=2$ & $\bar{\mu}=4$ & $\bar{\mu}=8$ & $\bar{\mu}=16$\tabularnewline
\midrule
GMM & 0.188 & 0.200 & 0.222 & 0.287 & 0.459\tabularnewline
BC & 0.256 & 0.256 & 0.255 & 0.254 & 0.246\tabularnewline
EB & 0.241 & 0.243 & 0.241 & 0.242 & 0.237\tabularnewline
Oracle BC & 0.185 & 0.186 & 0.182 & 0.182 & 0.181\tabularnewline
Oracle EB & 0.179 & 0.181 & 0.179 & 0.180 & 0.179\tabularnewline
\midrule
$J$-stat & 88.6 & 90.8 & 92.7 & 90.7 & 99.5\tabularnewline
$V_{EB}/V_{BC}$ & 0.936 & 0.936 & 0.936 & 0.936 & 0.936\tabularnewline
\bottomrule
\end{tabular}
\par\end{centering}
{\small\emph{Notes: }}{\small RMSE of $\hat{\beta}$ across 1,000 Monte
Carlo replications with $m=40$,  $n=10{,}000$, and $\bar{\sigma}=8$
held fixed. The final two rows report the mean simulated $J$ statistic
and the population leading-term risk ratio of EB to BC,
$V_{EB}/V_{BC}=1-\Lambda\Sigma_{\mu}M'(M\Sigma_{\mu}M')^{+}M\Sigma_{\mu}\Lambda'/(\Lambda\Sigma_{\mu}\Lambda')$,
computed exactly from the design parameters; the latter does not
depend on $\bar{\mu}$.}{\small\par}
\end{table}

At $\bar{\mu}=0$ the systematic component of the bias is absent. Oracle BC removes the realized mean without shrinking and barely improves on GMM, with an RMSE of 0.185. The further drop to 0.179 under Oracle EB reflects the advantages of shrinkage using the true $\sigma^2$. The feasible corrections pay the cost of
estimating the hyperparameters. This leaves GMM (0.188) ahead of BC (0.256)
and EB (0.241) by a wide margin. In contrast, at  $\bar{\mu}=8$,
the feasible corrections dominate: EB achieves 0.242 versus 0.287 for GMM.
At  $\bar{\mu}=16$, GMM degrades to 0.459 while BC and EB stay around
0.24--0.25. When the systematic bias is large, the corrected methods offer substantial improvements in RMSE.

Table~\ref{tab:mc-rmse-vs-sig2} examines how the relative performance
of the estimators depends on the marginal dispersion
$\bar{\sigma}$ of the specification errors while holding  $m=40$,
 $\bar{\mu}=8$, and all other parameters fixed. The mean simulated  $J$-statistic rises from 41.8 at  $\bar{\sigma}=0$, which is below the 5\% critical value of 54.57,
to 238.9 at  $\bar{\sigma}=16$. The final row reports the ratio of the EB and BC leading-term risks, $V_{EB}/V_{BC}$, computed exactly from the design parameters. The ratio is stable near 0.95 for $\bar{\sigma}\le8$ and falls to 0.861 at $\bar{\sigma}=16$, reflecting growing asymptotic gains from shrinkage as dispersion rises.

\begin{table}[H]
\caption{RMSE by $\bar{\sigma}$ ( $\bar{\mu}=8$,  $n=10{,}000$,  $m=40$)
\protect\label{tab:mc-rmse-vs-sig2}}

\begin{centering}
\begin{tabular}{lccccc}
\toprule
 & $\bar{\sigma}=0$ & $\bar{\sigma}=2$ & $\bar{\sigma}=4$ & $\bar{\sigma}=8$ & $\bar{\sigma}=16$\tabularnewline
\midrule
GMM & 0.264 & 0.266 & 0.272 & 0.287 & 0.345\tabularnewline
BC & 0.192 & 0.196 & 0.208 & 0.254 & 0.395\tabularnewline
EB & 0.188 & 0.192 & 0.204 & 0.242 & 0.345\tabularnewline
Oracle BC & 0.147 & 0.149 & 0.157 & 0.182 & 0.275\tabularnewline
Oracle EB & 0.153 & 0.155 & 0.161 & 0.180 & 0.247\tabularnewline
\midrule
$J$-stat & 41.8 & 45.3 & 55.7 & 90.7 & 238.9\tabularnewline
$V_{EB}/V_{BC}$ & 0.943 & 0.947 & 0.953 & 0.936 & 0.861\tabularnewline
\bottomrule
\end{tabular}
\par\end{centering}
{\small\emph{Notes: }}{\small RMSE of $\hat{\beta}$ across 1,000 Monte
Carlo replications with $m=40$,  $n=10{,}000$, and $\bar{\mu}=8$
held fixed. The column $\bar{\sigma}=8$ is the baseline. The final
two rows report the mean simulated $J$ statistic and the population
leading-term risk ratio of EB to BC,
$V_{EB}/V_{BC}=1-\Lambda\Sigma_{\mu}M'(M\Sigma_{\mu}M')^{+}M\Sigma_{\mu}\Lambda'/(\Lambda\Sigma_{\mu}\Lambda')$,
computed exactly from the design parameters.}{\small\par}
\end{table}

At $\bar{\sigma}=0$ (homogeneous specification error), the feasible
estimators are essentially tied (EB 0.188 vs.~BC 0.192), as are the oracle estimators (OEB 0.153 vs.~OBC 0.147). Recall from Section \ref{subsec:ordering} that with no idiosyncratic overdispersion to exploit and efficient weighting,
the precision gains from shrinkage derive entirely from heteroscedastic moment noise. The final row puts $V_{EB}/V_{BC}$ at 0.943, about a 3\% reduction in RMSE, consistent with the near tie. As the marginal variance
$\bar{\sigma}^2$ of specification errors grows, the signal from the specification errors
strengthens relative to sampling noise, making shrinkage increasingly
beneficial. At  $\bar{\sigma}=16$, the EB estimator has roughly 13\% lower RMSE than BC. In contrast, the asymptotic approximations suggest the RMSE ratio of these estimators should be driven entirely by their leading-term variance ratio, yielding a reduction of $1-\sqrt{0.861}\approx7\%$. Thus, the leading term underestimates the advantages of EB in the extreme heterogeneity design with $\bar{\sigma}=16$.

The RMSE of GMM increases monotonically with $\bar{\sigma}$. The advantage of BC over GMM decreases in $\bar{\sigma}$, as larger specification errors raise the precision cost of the first-order bias correction. At $\bar{\sigma}=16$, GMM dominates BC as the precision costs eventually outweigh the gains of bias correction. The EB estimator dampens this precision cost and is the overall best-performing feasible estimator in Table~\ref{tab:mc-rmse-vs-sig2}. However, in the presence of extreme idiosyncratic dispersion ($\bar{\sigma}=16$) its RMSE is tied with GMM.

Appendix Table \ref{tab:appc-sweeps} reports corresponding results for differenced estimators and MBTSLS. Differenced GMM and EB track levels BC and EB
closely at every setting of the hyperparameters. Evidently, estimating $\mu$
costs about as much as eliminating it. Differenced MBTSLS is dominated by differenced EB at all hyperparameter values, reflecting the more severe precision costs of correcting for higher-order many-instruments bias.

\subsection{Asymptotic behavior}\label{subsec:asymptotic}

To assess the rate predictions of Propositions~\ref{prop:bc-rate}
and \ref{prop:eb-rate}, I vary  $n$ and  $m$ jointly, setting  $m$
equal to  $n^{0.4}$ rounded to the nearest integer. This choice ensures
that  $m^{2}/n\to0$ as required by Assumption~\ref{assu:regime}.
Table~\ref{tab:mc-bias-vs-n} reports  the bias and RMSE of all five
estimators across seven sample sizes from  $n=5{,}000$ to  $n=500{,}000$. Since the specification errors are scaled by $1/\sqrt{n}$, all of the estimators are consistent. It is convenient then to multiply both the bias and RMSE by $\sqrt{n}$ so that differences in estimator performance do not collapse as the sample size grows.

\begin{table}[H]
\centering
\caption{Bias and RMSE as Functions of  $n$ with  $m=\mathrm{round}(n^{0.4})$\protect\label{tab:mc-bias-vs-n}}

\begin{centering}
\setlength{\tabcolsep}{3pt}
\begin{tabular}{rccccccccc ccccc}
\toprule
 &  &  &  & \multicolumn{5}{c}{$\sqrt{n}\times$Bias} &  & \multicolumn{5}{c}{$\sqrt{n}\times$RMSE}\tabularnewline
\midrule
$n$ & $m$ & $V_{EB}/V_{BC}$ &  & GMM & BC & EB & OBC & OEB &  & GMM & BC & EB & OBC & OEB\tabularnewline
5,000 & 30 & 0.924 &  & 21.3 & 4.2 & 4.5 & 3.4 & 3.5 &  & 28.8 & 25.0 & 23.8 & 19.1 & 18.6\tabularnewline
10,000 & 40 & 0.936 &  & 22.0 & 4.0 & 3.8 & 2.4 & 2.3 &  & 28.7 & 25.4 & 24.2 & 18.2 & 18.0\tabularnewline
20,000 & 53 & 0.945 &  & 24.0 & 4.7 & 4.2 & 2.9 & 2.3 &  & 29.5 & 24.2 & 23.2 & 17.3 & 17.3\tabularnewline
50,000 & 76 & 0.953 &  & 24.5 & 3.7 & 3.4 & 2.4 & 2.0 &  & 30.3 & 23.4 & 22.3 & 17.7 & 17.4\tabularnewline
100,000 & 100 & 0.956 &  & 24.5 & 2.3 & 2.4 & 1.7 & 1.7 &  & 29.9 & 23.4 & 22.9 & 17.1 & 17.5\tabularnewline
200,000 & 132 & 0.958 &  & 25.5 & 3.6 & 3.3 & 2.5 & 2.2 &  & 30.6 & 22.7 & 22.2 & 16.9 & 17.4\tabularnewline
500,000 & 190 & 0.959 &  & 25.6 & 3.2 & 3.1 & 2.2 & 2.0 &  & 30.3 & 22.3 & 21.8 & 16.3 & 16.7\tabularnewline
\bottomrule
\end{tabular}
\par\end{centering}
\raggedright{\small\emph{Notes:}}{\small{} Entries are  $\sqrt{n}$  times the bias and RMSE of  $\hat{\beta}$ across
1,000 replications per grid point. The population leading-term risk
ratio $V_{EB}/V_{BC}$ is computed exactly from the design parameters.}{\small\par}
\end{table}

The  scaled bias of GMM grows modestly with $n$ but levels off around $25$, indicating a limiting $\sqrt{n}$ rate. In contrast, the scaled  biases of BC and EB gradually drift towards zero, falling from about  $4$  at  $n=5{,}000$  to roughly  $3$
at  $n=500{,}000$. The oracle estimators do slightly better than their feasible counterparts but exhibit the same drift towards zero bias. These gradual drops in the scaled bias of the corrected estimators reflect the many-instruments bias of TSLS that would be present even under proper specification. Since this bias is of order $m/n$, the scaled bias reported in the table is of order $m/\sqrt{n}=n^{-0.1}$.

The scaled RMSE of GMM rises slightly before stabilizing around 30. The rise reflects its scaled bias, which grows with $n$ even as the scaled variance falls. In contrast, the scaled RMSE of the corrected estimators edges down as $n$ grows before stabilizing. The bias of these estimators is already small, so this decline primarily reflects a falling scaled variance. The leveling off in RMSE of the corrected estimators is
consistent with the parametric-rate predictions of Propositions \ref{prop:bc-rate}
and \ref{prop:eb-rate}.



\begin{figure}[H]
\begin{centering}
\caption{Wald test rejection rates vs. $m$ with $m=\mathrm{round}(n^{0.4})$\protect\label{fig:mc-rej-vs-m}}
\includegraphics[width=0.85\textwidth]{fig_mc_rej_vs_m.eps}
\par\end{centering}
\raggedright{\small\emph{Notes:}}{\small{} Figure depicts rejection
rate of Wald test that $\beta_{0}=1$ for each estimator. Nominal
level of each test is 5\%. 1,000 simulations used per grid point.}{\small\par}
\end{figure}

GMM has the highest RMSE at every sample size. By Proposition~\ref{prop:variance-ordering}, the risk ratios predict that EB improves on BC throughout the grid: $V_{EB}/V_{BC}$ rises from 0.924 at $m=30$ to 0.959 at $m=190$. Consistent with this prediction, EB exhibits lower RMSE than BC at all sample sizes, with a gap that narrows as $m$ grows.  The oracle estimators have the lowest RMSEs but exhibit comparable performance at large sample sizes.

Figure~\ref{fig:mc-rej-vs-m} plots rejection rates across the Table~\ref{tab:mc-bias-vs-n}
grid. GMM's bias leads to over-rejection at all sample sizes, hovering
between 38\% and 41\% type I error rates as  $m$ grows large. In
contrast, both the BC and EB rejection rates quickly approach the nominal
5\% level as $m$ grows large. This finding is consistent with Appendix Table \ref{tab:appb-diagnostics}, which verifies that the BC and EB estimators of $\beta$ satisfy the leverage requirements of Corollary~\ref{corr:feasible-inference}. The oracle BC and EB rejection rates also hover near the nominal
5\% level across all $m$, with OBC and OEB exhibiting rejection rates of 5.3\% and 4.3\% respectively at $m=190$.

\section{Empirical Application: Returns to Schooling \protect\label{sec:empirical}}

In this section, I revisit the work of \textcite{angrist1991does},
who used quarter-of-birth (QOB) indicators as instruments for years
of education in a log weekly wage equation to estimate the returns
to schooling. Table VII of their paper reports TSLS estimates based
on interactions between QOB and state-of-birth (SOB) as excluded instruments.
The same specification also excludes interactions between QOB and
year-of-birth (YOB), yielding 180 excluded instruments in total.

I will simplify this instrument set in two respects with the dual
aims of ensuring exchangeability and avoiding potential biases that
can arise under many instrument asymptotics. First, I collapse the
three QOB indicators into a single $\mathbf{1}\left\{ \text{QOB}>1\right\} $
indicator for being born in one of the last three quarters, interacted
with SOB. \textcite{angrist1991does} work with the same binarized
version of QOB in their Table III, which they show yields similar
results to specifications leveraging the complete set of QOB indicators
as instruments via TSLS. Second, I drop the season of birth interactions
with YOB. This yields 51 instruments (one per state).

A key identifying assumption of \textcite{angrist1991does}'s study
is that season of birth is correlated with earnings only through its
causal effect on educational attainment. This claim was questioned
early on by \textcite{bound1995problems}, who considered the potential
effects of small violations of QOB exogeneity. Later work by \textcite{buckles2013season}
provided evidence that economically disadvantaged families are more
likely to have children in the first quarter of the year, suggesting
the average earnings of children born in later quarters would be higher
even if their years of schooling were equalized. Since \textcite{angrist1991does}
establish that individuals born in later quarters tend to complete
more schooling, the family background differences highlighted by \textcite{buckles2013season}
may lead QOB-based IV estimates to overstate the returns to schooling
when instruments have a positive first stage. In what follows, I investigate
whether the BC and EB estimators reach similar conclusions about the
likely bias in TSLS.

\subsection{Estimation sample and moment conditions}

The estimation sample consists of men from the 1940--1949 birth cohorts
in the 1980 Census 5\% Public Use Micro Sample (PUMS), yielding $n=486{,}926$
observations. The instruments are $m=51$ mutually exclusive indicators
of the form $Z_{ij}=\mathbf{1}\left\{ \text{QOB}_{i}>1,\,\text{SOB}_{i}=j\right\} $
for each state $j=1,\dots,51$. Each instrument flags whether
individual $i$ was born in a non-Q1 quarter in a particular state.
Note that $\sqrt{n}\approx698$ and $m^{2}/n\approx0.005$, which
suggests the asymptotic regime of Assumption \ref{assu:regime}.i
is easily satisfied.

I consider four specifications of controls $X_{i}$ mirroring those
entertained by \textcite{angrist1991does}: (i) SOB dummies only,
(ii) SOB dummies with age (in quarters) and age-squared, (iii) SOB
dummies with individual-level covariates (indicators for race, marital
status, SMSA residence, and eight census region dummies), and (iv)
SOB dummies with all covariates plus age and age-squared. Rather than
include a constant, I always use 51 state-of-birth dummies that sum
to one.

The endogenous variable is completed years of schooling $T_{i}$.
The scalar parameter of interest is $\theta$, which measures the
returns to an additional year of schooling. Since the controls are
nuisance parameters, I partial them out. Thus, in all four specifications
the degree of overidentification is $m-p=50$. As the simulations
of the previous section demonstrate, this level of overidentification
can yield sufficiently accurate hyperparameter estimates for corrected
estimators to generate non-trivial improvements over GMM.

Let $Y_{i}$ denote log weekly earnings, and let $\tilde{Y}_{i}$,
$\tilde{T}_{i}$, and $\tilde{Z}_{i}$ denote the residuals from OLS
regressions of $Y_{i}$, $T_{i}$, and $Z_{i}$ on $X_{i}$. I work with
sample moment conditions of the form
\[
\hat{g}\left(\theta\right)=\hat{S}^{-1}\left(\frac{1}{n}\sum_{i=1}^{n}\left(\tilde{Y}_{i}-\theta\tilde{T}_{i}\right)\tilde{Z}_{i}\right),\quad \hat{S}=\frac{1}{n}\sum_{i}\tilde{Z}_{i}\tilde{Z}_{i}'.
\]
As explained in Section \ref{subsec:sim_moment}, rescaling by the inverse of the second moment matrix of residualized instruments $\hat{S}$ ensures that these moment conditions capture excludability violations when evaluated at the true parameter $\theta_0$. Here, the excludability violations capture direct effects of being born outside
the first quarter in state $j$, net of the controls $X_{i}$. Exchangeability
across states of these direct effects is plausible ex ante: the family-composition mechanism
of \textcite{buckles2013season} could operate similarly in each state.

The sample Jacobian is the rescaled first-stage moment $\hat{G}=-\hat{S}^{-1}(\frac{1}{n}\sum_{i}\tilde{T}_{i}\tilde{Z}_{i})$. Since non-Q1 births have more education on average, most of its entries are negative.
The vector $1_{m}$ is in the column space of $\hat{G}$ only if the
first stages are common across states. Fortunately, cross-state variation
in compulsory schooling laws yields substantial variation in first stage strength
across states, which was \textcite{angrist1991does}'s motivation for
studying SOB interactions.

\subsection{OLS and TSLS Results}

Table~\ref{tab:ak91-replication} reports OLS and TSLS estimates
for the non-Q1~$\times$~SOB instrument design. Without age controls,
the TSLS estimates of the return to education are around 4\%, below
the OLS estimates of approximately 5\%. Adding age controls raises
the TSLS estimates sharply, to 12--14\%, well above OLS. This sensitivity
to the inclusion of age reflects the mechanical correlation between
age and QOB: non-Q1 individuals are younger on average. \textcite{rosenzweig2000natural}
argued that this mechanical dependence led to excludability difficulties
as younger workers have less labor market experience and therefore
lower potential wages. The sizable increase in the TSLS estimates
after controlling for the age quadratic is consistent with this view.

\begin{table}[H]
\caption{TSLS Estimates: Returns to Education with nonQ1~$\times$~SOB Instruments
\protect\label{tab:ak91-replication}}

\begin{centering}
\begin{tabular}{lcccc}
\toprule
 & SOB & SOB + age & SOB + cov & SOB + cov + age\tabularnewline
\midrule
OLS & 0.0535 & 0.0555 & 0.0495 & 0.0513\tabularnewline
 & (0.0004) & (0.0004) & (0.0003) & (0.0003)\tabularnewline
TSLS & 0.0424 & 0.1413 & 0.0360 & 0.1240\tabularnewline
 & (0.0188) & (0.0246) & (0.0193) & (0.0251)\tabularnewline
$J$ {[}d.f.{]} & 67.32 {[}50{]} & 46.64 {[}50{]} & 58.66 {[}50{]} & 45.65 {[}50{]}\tabularnewline
$p$-value & 0.052 & 0.609 & 0.188 & 0.649\tabularnewline
\bottomrule
\end{tabular}
\par\end{centering}
{\small\emph{Notes:}}{\small{} Sample is 486,926 men from the 1940--1949
birth cohorts in the 1980 Census 5\% PUMS. Instruments are 51 $\mathbf{1}\left\{ \text{QOB}>1\right\} $~$\times$~SOB
interactions. Heteroscedasticity-robust standard errors in parentheses.
 $J$-statistic degrees of freedom in brackets.}{\small\par}
\end{table}

In contrast to the Monte Carlo simulations, these data exhibit minimal
evidence of misspecification. The $J$-test borderline rejects correct
specification in one specification without age controls ($p=0.052$)
and fails to reject in the other ($p=0.188$); both specifications
with age controls clearly fail to reject ($p=0.609$ and $p=0.649$).
I will show that misspecification-aware point estimates and standard
errors nonetheless offer additional insight in this environment.

To anchor the subsequent misspecification-aware analysis to these
TSLS specifications, I weight GMM by $\hat{W}=\hat{S}$. Because the
moments are already rescaled by $\hat{S}^{-1}$, this weight reproduces
the TSLS point estimates exactly and is efficient under correct specification
and homoscedasticity of the earnings errors. I estimate moment uncertainty
with the heteroscedasticity-robust sandwich estimator $\hat{V}$,
computed from the outer product of the rescaled moment contributions
to account for estimation of the controls.

\subsection{Scrutinizing exchangeability}\label{subsec:Scrutinizing}

To assess the plausibility of exchangeability in these data, I apply the
specification test in \eqref{stat:feature} of Section~\ref{subsec:Diagnostic} using as features the centered
inverse-root cell frequency $\hat{q}_{j}^{-1/2}$, together with centered
indicators for the four Census regions of state $j$ (Northeast, Midwest, South,
West). The first feature tests whether the specification errors covary with the
noise of the state cell moments---whose standard deviation scales with
$\hat{q}_{j}^{-1/2}$---as in Example~\ref{exa:mean}. The region indicators test
whether the direct effects of season of birth are systematically larger in some
parts of the country than others, which would arise, for example, if the
family-composition channel of \textcite{buckles2013season} operates with
systematically different intensity across regions.

Recall from Remark~\ref{rem:test-power} that the
test cannot detect violations involving features perfectly aligned with first-stage strength. Regressing $\hat{q}_{j}^{-1/2}$ on the constant and the first stage coefficients
$-\hat{G}$ leaves 94--97\% of its variance in the residual depending on the control specification. Corresponding regressions with each region indicator as an outcome leave 75--99\% of the variance in the residual. Thus,
basing the test on these features probes for violations in directions that should yield non-trivial power.


\begin{table}[H]
\caption{Exchangeability diagnostics\protect\label{tab:exch-diagnostic}}

\begin{centering}
\begin{tabular}{lcccc}
\toprule
 & SOB & SOB + age & SOB + cov & SOB + cov + age\tabularnewline
\midrule
Cell noise ($\hat{q}_{j}^{-1/2}$) & $-0.60$ & $-0.49$ & $-0.68$ & $-0.52$\tabularnewline
Northeast & 0.44 & 0.36 & 0.42 & 0.35\tabularnewline
Midwest & $-0.76$ & $-0.68$ & $-0.87$ & $-0.74$\tabularnewline
South & 0.29 & 0.27 & 0.65 & 0.50\tabularnewline
West & 0.11 & 0.11 & $-0.06$ & $-0.02$\tabularnewline
\midrule
$\chi^{2}(4)$ & 1.04 & 0.76 & 1.50 & 0.94\tabularnewline
$p$-value & 0.90 & 0.94 & 0.83 & 0.92\tabularnewline
\bottomrule
\end{tabular}
\par\end{centering}
{\small\emph{Notes:}}{\small{} The first five rows report individual feature
$t$-statistics $(\Xi'\tilde{r})_{a}/\sqrt{\hat{\Omega}_{aa}}$. The cell-noise
feature is the centered inverse-root cell frequency $\hat{q}_{j}^{-1/2}$, a proxy
for the noise scale of the state cell moment. The four region indicators, likewise centered, identify the Census region of state $j$. The joint statistic
$(\Xi'\tilde{r})'\hat{\Omega}^{-1}(\Xi'\tilde{r})\sim\chi^{2}(4)$ uses the
cell-noise feature and three region contrasts (West omitted, since the four
region indicators are collinear). Sample and instruments as in
Table~\ref{tab:ak91-replication}.}{\small\par}
\end{table}


Table~\ref{tab:exch-diagnostic} reports the results. The cell-noise feature
shows a mild negative association that is far from significant in every
specification ($|t|\le0.68$). No region indicator approaches significance
(the largest is $0.87$ in absolute value). The joint $\chi^{2}(4)$ statistic
lies well below its degrees of freedom in all four specifications, and none
rejects at conventional levels ($p\ge0.83$). The data thus provide no evidence against exchangeability.

\subsection{Misspecification-aware estimates}

Table~\ref{tab:ak91-bceb} presents estimates of the hyperparameters
$\left(\mu,\sigma^{2}\right)$ alongside the GMM, BC, and EB estimates. Again, the GMM point estimates and standard errors coincide with the
TSLS estimates reported in Table~\ref{tab:ak91-replication} by construction.
The reported $J$-statistics, however, differ slightly from those
in Table~\ref{tab:ak91-replication}. Both statistics are quadratic forms of the type $\hat{r}'\hat{V}^{-1}\hat{r}$ evaluated at the TSLS estimate, with heteroscedasticity-consistent variance matrix $\hat{V}$. The statistics reported here use the
FWL-residualized excluded-instrument moments $\hat{r}=\sqrt{n}\hat{g}(\hat{\theta})$. The Table~\ref{tab:ak91-replication}
statistics instead use the full unresidualized system, including the
control block.
Under the null of correct specification and conditional homoscedasticity, a
setting in which the TSLS weighting is asymptotically efficient, the two are
asymptotically equivalent.

Consistent with strong identification of $\mu$ (Assumption~\ref{assu:regime}.vi), $\hat{\kappa}:=m^{-1}\|\hat{M}1_{m}\|^{2}$
is far from zero, ranging across the four specifications from $0.68$
to $0.92$. Table~\ref{tab:ak91-bceb} also reports the empirical Lindeberg
conditions that Theorem~\ref{thm:clt} requires for the BC and EB
estimators. The BC values range
from $0.0006$ to $0.0015$, while the EB values range from $0.0010$ to $0.0047$. All of these values are very small, suggesting that the
no-dominant-moment condition of Assumption~\ref{assu:lindeberg}.iii holds for each estimator. Large
values would call into question the validity of the
misspecification-aware standard errors.

\begin{table}[H]
\caption{Bias-Corrected and Empirical Bayes Estimates: Returns to Education
\protect\label{tab:ak91-bceb}}

\begin{centering}
\begin{tabular}{lcccc}
\toprule
 & SOB & SOB + age & SOB + cov & SOB + cov + age\tabularnewline
\midrule
\emph{Hyperparameters} &  &  &  & \tabularnewline
$\hat{\mu}$ & $-5.39$ & 4.49 & $-3.67$ & 4.75\tabularnewline
$\hat{\sigma}^{2}$ & 50.97 & 0.00 & 0.10 & 0.00\tabularnewline
 &  &  &  & \tabularnewline
\midrule
\emph{Point estimates} &  &  &  & \tabularnewline
GMM & 0.0424 & 0.1413 & 0.0360 & 0.1240\tabularnewline
 & (0.0188) & (0.0246) & (0.0193) & (0.0251)\tabularnewline
BC & 0.0876 & 0.1108 & 0.0678 & 0.0927\tabularnewline
 & (0.0364) & (0.0318) & (0.0341) & (0.0314)\tabularnewline
EB & 0.0846 & 0.1103 & 0.0769 & 0.0935\tabularnewline
 & (0.0310) & (0.0270) & (0.0252) & (0.0271)\tabularnewline
 &  &  &  & \tabularnewline
\midrule
\emph{Diagnostics} &  &  &  & \tabularnewline
$J$ {[}d.f.{]} & 66.21 {[}50{]} & 46.14 {[}50{]} & 57.76 {[}50{]} & 45.14 {[}50{]}\tabularnewline
$\hat{\kappa}$ & 0.675 & 0.884 & 0.708 & 0.924\tabularnewline
$\max_{j}\hat{\Lambda}_{BC,j}'\hat{V}_{BC}^{-1}\hat{\Lambda}_{BC,j}$ & 0.0006 & 0.0012 & 0.0008 & 0.0015\tabularnewline
$\max_{j}\hat{\Lambda}_{EB,j}'\hat{V}_{EB}^{-1}\hat{\Lambda}_{EB,j}$ & 0.0010 & 0.0039 & 0.0046 & 0.0047\tabularnewline
Hausman $p$-value & 0.88 & 0.98 & 0.69 & 0.96\tabularnewline
\bottomrule
\end{tabular}
\par\end{centering}
{\small\emph{Notes:}}{\small{} Sample and instruments as in Table~\ref{tab:ak91-replication}.
All 51 nonQ1~$\times$~SOB moment conditions are treated as potentially
misspecified. GMM uses weighting matrix $\hat{W}=\hat{S}$
on $\hat{S}^{-1}$-rescaled moments, reproducing the TSLS estimator.
Heteroscedasticity-robust standard errors in parentheses. The Hausman row reports the $p$-value from comparing the test statistic in \eqref{stat:Hausman} to a $\chi^{2}(1)$ distribution.}{\small\par}
\end{table}



The estimated mean specification error $\hat{\mu}$ is negative in
the specifications without age controls, consistent with the experience
channel highlighted by \textcite{rosenzweig2000natural}: non-Q1 individuals
have less potential labor market experience, producing a negative
direct effect on earnings. With age controls,  $\hat{\mu}$ turns
positive, consistent with the family composition mechanism of \textcite{buckles2013season}:
once potential experience differences are absorbed, the remaining
direct effect of non-Q1 birth reflects better family backgrounds,
which raise earnings.

On the original log-wage scale, these estimates correspond to a mean
per-state direct effect $\hat{\mu}/\sqrt{n}$ ranging from $-0.0077$ to
$0.0068$ across the four specifications, implying exclusion restrictions
of modest magnitudes (between $0.53\%$ and $0.77\%$ in absolute
terms). Aggregating across the 51 states, the bias correction $-(\hat{\mu}/\sqrt{n})\hat{\Lambda}1_{m}$
ranges from $-0.031$ to $+0.045$ across the four specifications.
In the SOB specification, for instance, $\hat{\mu}=-5.39$ generates
a BC adjustment of approximately $+0.045$ that moves the schooling
return from $0.042$ (TSLS) to $0.088$ (BC).

Consistent with the $J$-statistics falling below their degrees of
freedom, the estimated variance $\hat{\sigma}^{2}$ is zero in both
specifications with age controls, indicating no excess dispersion
beyond the mean shift once experience differences are absorbed. In
the specifications without age controls, the $J$-statistics exceed
their expected values but $\hat{\sigma}^{2}$ is only meaningfully positive in the SOB-only design. Thus, there is
some heterogeneity in the direct effects across states but it also seems to be captured by the controls other than age.

Without age controls,  $\hat{\mu}<0$ and the BC and EB corrections
adjust upward. In the specifications with age controls,  $\hat{\mu}>0$
and the BC and EB corrections adjust downward. In all specifications,
the corrected estimates converge toward the 0.07--0.11 range, substantially
narrowing the cross-specification dispersion relative to GMM/TSLS.
While the corrections consistently shift estimates in the direction
of OLS, they overshoot in the specifications without age controls,
where the estimated misspecification is largest.

Comparing the difference between the BC and EB point estimates to the difference in their squared standard errors yields the Hausman-type test of \eqref{stat:Hausman}, reported in the bottom row of Table~\ref{tab:ak91-bceb}. Consistent with the earlier findings of Section \ref{subsec:Scrutinizing}, the test fails to reject in all four specifications. Evidently, shrinkage adjustments of this magnitude are to be expected under exchangeability given the hyperparameter estimates.

The BC standard
errors are larger than those of TSLS because they account
for two additional sources of uncertainty: the estimation of $\hat{\mu}$
and (when $\hat{\sigma}^{2}>0$) the cross-moment dispersion of specification
errors. The increases are especially stark in the SOB-only design, where both contributions are present.
As noted in Section \ref{subsec:ordering}, EB can improve on BC even when $\hat{\sigma}^{2}=0$, particularly when a suboptimal weighting matrix is used or the moments are heteroscedastic.
The EB standard errors are 14--26\% smaller than those of BC, including in the specifications with $\hat{\sigma}^{2}=0$, reflecting the non-trivial moment heteroscedasticity driven by cell-size variation. Notably, the EB standard errors are only slightly larger than TSLS in these specifications, indicating that the ultimate precision costs of bias-correction are low in this application when combined with shrinkage.

On net, the BC and EB corrections are consistent with season of birth
having direct effects on earnings whose sign depends on whether age
controls are included: the direct effects are negative when age controls
are omitted but positive when age controls are included. The corrections
bring TSLS estimates into closer agreement across specifications,
yielding returns to schooling in the range of 7--11\% per year. This
relative stability of the corrected estimates is reassuring, suggesting
that the misspecification-aware methods can detect and mitigate the
influence of omitted variables bias. Likewise, the elevated standard
errors provided by the BC and EB estimators provide a more honest
assessment of the composite uncertainty faced by researchers.

\subsection{Visual IV}

As a final exercise, it is useful to compare the results of Table~\ref{tab:ak91-bceb} to those that would have emerged from a simpler visual IV diagnostic of the form discussed in Example~\ref{exa:VIV}. Define the reduced form and first stage coefficient vectors
\[
    \hat \delta = \hat S^{-1} \left(\frac{1}{n}\sum_{i=1}^n \tilde Z_i \tilde Y_i\right) \qquad \hat{\pi} = \hat S^{-1} \left(\frac{1}{n}\sum_{i=1}^n \tilde Z_i \tilde T_i\right).
\]
The reduced-form entries $\hat{\delta}_{j}$ give
the covariate-adjusted log-wage gap for non-Q1 births in state $j$, while the
first-stages $\hat{\pi}_{j}$ give the adjusted schooling gap in state $j$. Figure~\ref{fig:ak91-viv} plots a scatter of these objects.

\begin{figure}[H]
    \begin{centering}
    \caption{Visual IV: reduced forms versus first stages across states\protect\label{fig:ak91-viv}}
    \par\end{centering}
    \centering{}\includegraphics[width=0.8\textwidth]{fig_ak91_viv.eps}
    \par\medskip{}
    \begin{centering}
    \begin{tabular}{lcccc}
    \toprule
     & SOB & SOB + age & SOB + cov & SOB + cov + age\tabularnewline
    \midrule
    Intercept $\hat{\mu}_{W}/\sqrt{n}$ & $-$0.0098 & 0.0055 & $-$0.0080 & 0.0055\tabularnewline
     & (0.0029) & (0.0023) & (0.0027) & (0.0021)\tabularnewline
    Implied $\hat{\mu}_{W}$ & $-$6.85 & 3.87 & $-$5.55 & 3.82\tabularnewline
     & (2.05) & (1.61) & (1.89) & (1.48)\tabularnewline
    Slope $\hat{\theta}_{\mathrm{VIV}}$ & 0.0998 & 0.1150 & 0.0841 & 0.0988\tabularnewline
     & (0.0256) & (0.0247) & (0.0249) & (0.0242)\tabularnewline
    \bottomrule
    \end{tabular}
    \par\end{centering}
    \raggedright{\small\emph{Notes:}}{\small{} Each marker is a state ($m=51$); the
    horizontal axis is the first-stage coefficient $\hat{\pi}_{j}$ and the vertical axis the
    reduced-form coefficient $\hat{\delta}_{j}$, both residualized against the controls of the
    corresponding specification in Table~\ref{tab:ak91-bceb}.
    Lines are GLS fits of
    $\hat{\delta}_{j}$ on $\hat{\pi}_{j}$ with weight $\hat{S}$, allowing a non-zero intercept. The slope $\hat{\theta}_{\mathrm{VIV}}$
    approximates the bias-corrected return and the intercept $\hat{\mu}_{W}/\sqrt{n}$ the
    mean exclusion violation. GLS standard errors in parentheses.}{\small\par}
\end{figure}

A through-origin GLS fit of reduced forms to first stages with weighting matrix $\hat{S}$ reproduces the TSLS slope exactly. In contrast, Figure~\ref{fig:ak91-viv} overlays
GLS fits that allow for a non-zero intercept. The fitted intercept
$\hat{\mu}_{W}/\sqrt{n}$ estimates the weighted mean direct effect of non-Q1 birth on earnings. As in Table~\ref{tab:ak91-bceb}, the VIV indicates these effects are
negative in the specifications without age controls and positive once they are included. The implied $\hat{\mu}_{W}$ are of comparable magnitude to the $\hat{\mu}$ estimates in Table~\ref{tab:ak91-bceb} and share their signs across all four specifications.

The fitted slopes give the visual-IV returns $\hat{\theta}_{\mathrm{VIV}}$, which are modestly larger than the BC estimates of Table~\ref{tab:ak91-bceb}. Like the BC estimates, the VIV returns are less dispersed across control specifications than the TSLS estimates, ranging from 8--12\% per year. However, the VIV standard errors are smaller than those of the BC estimator in Table~\ref{tab:ak91-bceb} and comparable to those of TSLS, which is to be expected given that the GLS standard errors assume the model is properly specified.

In this case, the impression one takes away from the VIV diagnostic about the excludability of the instruments and the magnitude of returns to schooling aligns closely with the findings from the more sophisticated BC and EB methods. A tentative conclusion is that plotting VIV fits including an intercept offers a useful diagnostic for gauging the average direction of exclusion violations. However, the standard errors from such methods fail to incorporate uncertainty due to idiosyncratic specification error, potentially leading to substantial overstatement of precision even after accounting for mean biases.

\section{Conclusion \protect\label{sec:Conclusion}}

Econometric models are almost always misspecified. The methods developed here leverage overidentifying restrictions to
correct GMM estimates of parameters of interest and their standard
errors for exchangeable misspecification. As evidenced by the SOB+age specification of Table \ref{tab:ak91-bceb},
these corrections can be non-trivial even when the realized $J$-statistics
are small. While the $J$-statistic distributes power across local
alternatives in $m-p$ directions, the hyperparameters $\left(\mu,\sigma^{2}\right)$
used in the correction pool information across these dimensions into
a low-dimensional summary.

As the simulations and empirical application illustrate, the proposed
correction methods are not entirely automatic. The researcher must
have some sense of which aspect of the moment conditions is misspecified.
Moreover, the moment conditions must share a common structure that
makes exchangeability plausible. As the examples of Section \ref{sec:Examples}
demonstrate, the choice of moment scaling determines the exchangeability
model and is therefore a substantive modeling decision; alternative
scalings may yield different corrections. This reflects the general
observation that misspecification can only be studied relative to
larger encompassing models that are plausibly operative in the setting
under study \parencite{armstrong2025misspecification}.

When a plausible scaling has been chosen, the empirical Bayes methods
proposed here offer a useful supplement to GMM estimates. Adoption
of these methods may carry the added benefit of improving research
transparency by reducing selective reporting of $J$ statistics and
other model diagnostics. Rather than fretting about statistical model
rejections, researchers should focus their efforts on repairing their
estimates of flawed, but ultimately useful, econometric models.

\printbibliography

\clearpage