Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
53,752 characters · 11 sections · 53 citation commands
A Misuse of Specification Tests
\abstract{ Empirical researchers often perform model specification tests, such as Hausman tests and overidentifying restrictions tests, to assess the validity of estimators rather than that of models. This paper examines the effectiveness of such specification pretests in detecting invalid estimators. We analyze the local asymptotic properties of test statistics and estimators and show that locally unbiased specification tests cannot determine whether asymptotically efficient estimators are asymptotically biased. In particular, an estimator may remain valid even when the null hypothesis of correct model specification is false, and it may be invalid even when the null hypothesis is true. The main message of the paper is that correct model specification and valid estimation are distinct issues: correct specification is neither necessary nor sufficient for asymptotically unbiased estimation. }
Model specification tests, such as Hausman tests (hau1978) and overidentifying restrictions tests (sar1958; han1982), are widely used in empirical studies, although finding a true model is rarely the primary goal. Empirical researchers often conduct these tests to guide the choice of estimator. For example, the Durbin-Wu-Hausman (DWH) test (dur1954; wu1973) is typically used to decide between the ordinary least squares (OLS) estimator and the two-stage least squares (2SLS) estimator. Failure to reject the null hypothesis is taken as evidence in favor of OLS, while rejection is viewed as support for 2SLS.
This paper examines the effectiveness of specification pretests in assessing the validity of estimators within a local misspecification framework. It is widely believed that an estimator based on a null model is valid if a specification test fails to reject the null and is invalid if it rejects it. We investigate whether a specification test rejects the null when its associated estimator is invalid and fails to reject it when the estimator is valid.
Our main contribution is to show that specification tests are generally uninformative about the validity of estimators. This result builds on the framework of che_san2018. We analyze how the asymptotic bias of an estimator and the power of a specification test respond to local deviations from a benchmark distribution. We formalize these deviations using a score function of a path, which can be decomposed into two orthogonal components. We show that the asymptotic bias of an efficient estimator depends only on one component, whereas the local power of a locally unbiased test depends only on the other. Because these components are orthogonal, a specification test provides no information about the presence of asymptotic bias.
Our orthogonality result generalizes the finding in Section 5.1.3 of hal2005, which examines the connection between the efficient GMM estimator and the $J$ test. hal2005 provided an orthogonal decomposition of moment restrictions into identifying and overidentifying restrictions. He showed that the asymptotic bias of the GMM estimator is affected only by violations of identifying restrictions, while the local power of the $J$ test is affected only by violations of overidentifying restrictions. This result implies that the $J$ test cannot detect the asymptotic bias of the GMM estimator. As we discuss in Section (ref), this decomposition directly corresponds to our more general orthogonal decomposition of the score function. Our approach is not limited to the GMM case and can be applied to a much broader class of semiparametric estimators and test statistics.
Another contribution of this paper is to clarify what Hausman tests are actually testing. While commonly covered in econometrics textbooks, their null and alternative hypotheses are often left implicit, giving readers a vague understanding of what the tests are formally evaluating. We characterize the precise null and alternative hypotheses under which Hausman tests are locally unbiased. Furthermore, contrary to the standard setup, we show that both the efficient and inefficient estimators used to construct the test statistic share the same asymptotic bias under certain local deviations. This result implies that Hausman tests cannot necessarily be relied upon as a robustness check, because failure to reject the null hypothesis does not imply that the efficient estimator is free of asymptotic bias.
Our findings build on the well-established theory of che_san2018 and the broader semiparametric estimation literature. A key contribution of this paper is to uncover a previously unrecognized implication of this theory: specification tests are uninformative about estimator validity. While existing research has examined the local behavior of estimators and tests, it has neither emphasized this fundamental disconnect nor analyzed it from the perspective of pretesting. By making this implication explicit, we provide empirical researchers with a clearer understanding of the limits of specification tests and their relationship to the validity of estimators.
Many studies have investigated the impact of pretesting or model selection on subsequent inference. Examples include jud_boc1978, pot1991, kab1995, lee_pot2005, lee_pot2006, and and_gug2009ecta, and_gug2009joe. gug2010et,gug2010joe, gug_kum2012, and dok_wan2021 specifically investigated the impact of specification tests on subsequent inference. gug_kum2012 also provide several examples of empirical studies in which overidentifying restrictions tests are used as a pretest. These studies show that inference ignoring the effect of pretesting can be highly misleading. In contrast to these previous studies, the contribution of the present paper is to question the very rationale for pretesting itself.
The study of the local power properties of specification tests can be traced back to new1985, new1985ecta. These papers pointed out that specification tests may fail to detect local deviations that induce asymptotic bias in estimators. However, they did not emphasize that specification tests may exhibit local power even when estimators remain valid. This distinction is an important contribution of the present study.
The local asymptotic framework is also widely used in the context of robust estimation and inference, particularly for models specified by moment restrictions. kit_et_al2013 proposed a robust point estimator when the sample is obtained from a local deviation from a benchmark distribution. arm_kol2021 proposed confidence intervals that take into account the potential bias resulting from a local deviation. See also and_et_al2017, bon_wei2022 and and_et_al2025 for related issues. Our results highlight that, when potential misspecification is suspected, it is important to use robust methods from the outset, rather than relying on specification tests.
The remainder of the paper is organized as follows. Section (ref) introduces the setting of our analysis. Section (ref) presents the main results of the paper. Sections (ref) and (ref) show that the results of Section (ref) hold in popular models in econometrics. Section (ref) examines the connection between the efficient GMM estimator and the $J$ test statistic, while Section (ref) investigates the properties of the DWH test. Section (ref) presents the results of Monte Carlo studies, illustrating our findings using specific models. Section (ref) concludes. The Appendix gives an auxiliary result to Section (ref).
Let $P$ be a probability distribution defined on a sample space $\mathcal{X}$, and let $\mathcal{M}$ denote the set of all distributions on $\mathcal{X}$. The distribution $P$ may be motivated by an underlying economic model or regarded as an idealized distribution from which researchers hope to draw a sample. We assume that $P$ belongs to a semiparametric model $\mathbf{P}$, which is represented as a set of distributions on $\mathcal{X}$. Researchers know only that $P \in \mathbf{P}$, without further structural information. The finite-dimensional parameter of interest, $\theta_0 \in \Theta$, is defined as $\theta_0 = \psi(P)$ for a functional $\psi: \mathbf{P} \to \Theta$. This parameter typically corresponds to a component of a structural model, capturing the economic mechanism of interest.
A random sample $\{X_1, \dots, X_n\}$ is drawn from a distribution $\mu$ on $\mathcal{X}$. Let $\hat{\theta}_n=\hat{\theta}_n(X_1, \dots, X_n)$ be an asymptotically efficient estimator for $\theta_0$ under $P$, meaning that it attains the minimum asymptotic variance among all regular estimators. We consider the possibility that the DGP $\mu$ deviates locally from $P$, which may arise from data contamination, measurement error, model misspecification, or other sources. If the deviation vanishes at the rate $n^{-1/2}$, then $\hat{\theta}_n$ remains consistent for $\theta_0$, but may be asymptotically biased in the sense that the asymptotic mean of $\sqrt{n}(\hat{\theta}_n - \theta_0)$ is nonzero.
To detect the local deviation, we consider a specification test $\phi_n: \{X_i \}_{i=1}^n \to [0,1]$ for the following null and alternative hypotheses:
Whether the test exhibits nontrivial power against the local deviation depends on the direction of the deviation from $P$.
We investigate the relationship between the estimator and the specification test. Our main concern is whether the test can detect local deviations that induce asymptotic bias in the estimator. As we shall see shortly, not all local deviations lead to asymptotic bias in $\hat{\theta}_n$. Accordingly, if the test is used to access the validity of the estimator, it should have nontrivial local power if and only if the estimator is asymptotically biased. The asymptotic behavior of $\hat{\theta}_n$ and $\phi_n$ depends on the direction of the local deviation. To formalize this notion, we adopt the framework of che_san2018.
The direction of the deviation from $P$ is defined in terms of the score function of a path. A path $t \to P_t$ is a function defined on $[0,\epsilon)$ for some $\epsilon >0$ such that $P_t \in \mathcal{M}$ for all $t \in [0,\epsilon)$ and $P_0 = P$. In other words, a path is a parametric model indexed by the scaler parameter $t$. We focus on paths $t \to P_{t,g} \in \mathcal{M}$ that satisfy
for some function $g : \mathcal{X} \mapsto \mathbb{R}$. If this holds, the path is said to be differentiable in quadratic mean at $t=0$. The function $g$ is referred to as the score function because it is typically given by \[ g(x) = \left. \frac{\partial}{\partial t} \log dP_{t,g}(x)\right|_{t=0}. \] Whenever such a $g$ exists, it must satisfy $\mathbb{E}[g(X)]=0$ and $\mathbb{E}[g^2(X)]<\infty$, where $\mathbb{E}$ denotes the expectation with respect to $P$. Thus, the set of all possible score functions is given by \[ L_0^2(P) \equiv \left\{ g: \mathcal{X} \to \mathbb{R}: \mathbb{E}[g(X)]=0 \ \text{and} \ \mathbb{E}[g^2(X)]<\infty \right\}. \]
The directions of deviation that are consistent with the null hypothesis are captured by the tangent set. We define \[ T(P)= \left\{ g \in L_0^2(P) : \eqref{hellinger} \ \text{holds for some} \ t \mapsto P_{t,g} \in \mathbf{P} \right\}, \] which is called the tangent set of model $\mathbf{P}$ at $P$. Intuitively, it consists of all directions along which a path $P_{t}$ can deviate from $P$ while remaining within the model $\mathbf{P}$. We consider the case where $T(P)$ is a linear space, which is typical in many semiparametric models. In that case, the tangent set is also called the tangent space.
As noted in che_san2018, any score $g \in L_0^2(P)$ can be decomposed into two orthogonal components when $T(P)$ is linear. Let $\bar{T}(P)$ be the closure of $T(P)$, and define its orthogonal complement as \[ \bar{T}(P)^\bot = \left\{ f \in L_0^2(P): \mathbb{E}[f(X)g(X)] =0 \ \text{for all} \ g \in \bar{T}(P) \right\}. \] Then, we obtain the orthogonal decomposition $L_0^2(P) = \bar{T}(P) \oplus \bar{T}(P)^\bot$. Consequently, any score $g \in L_0^2(P)$ can be decomposed as $g = \Pi_{T}(g) + \Pi_{T^\bot}(g)$, where $\Pi_T$ and $\Pi_{T^\bot}$ denote the projection onto $\bar{T}(P)$ and $\bar{T}(P)^\bot$, respectively.
che_san2018 showed that local overidentification is both necessarily and sufficient for the existence of an asymptotically efficient estimator for $\theta_0$ and a locally unbiased test for (ref). They define that $P$ is locally overidentified by $\mathbf{P}$ if $\bar{T}(P) \neq L_0^2(P)$. This definition generalizes the classical definition of overidentification, which compares the number of moments with the number of parameters. See also Theorem 2.1 of new1994 for the related issue.
Building on the framework above, we investigate the relationship between the asymptotically efficient estimator and the locally unbiased test. Specifically, we consider the case where the sample is drawn from $P_{1/\sqrt{n},g}$ for some $g \in L_0^2(P)$, so that the DGP represents a local deviation from $P$. We then analyze how the asymptotic bias of the estimator and the local power of the test depend on the choice of $g$.
To facilitate understanding of the framework above, we illustrate how the $J$ test fits within it through the following example.
This section shows that the directions of local deviation that can be detected by locally unbiased specification tests are orthogonal to those that induce asymptotic bias in asymptotically efficient estimators. This means that it is impossible to determine whether efficient estimators are asymptotically biased by using any locally unbiased specification tests. We also discuss the properties of Hausman tests, clarifying their null and alternative hypotheses.
To investigate the properties of specification tests, we introduce some additional definitions from che_san2018. A test $\phi_n$ for (ref) has local asymptotic level $\alpha$ if \[ \limsup_{n \to \infty} \int \phi_n dP_{1/\sqrt{n}, g}^n \le \alpha \quad \] for any path $t \mapsto P_{t,g} \in \mathbf{P}$, where $P_{1/\sqrt{n},g}^n$ denotes the joint distribution of $\{X_1, \dots, X_n\}$. The test has a local asymptotic power function $\pi: L_0^2(P) \to [0,1]$ if \[ \lim_{n \to \infty} \int \phi_n dP_{1/\sqrt{n},g}^n = \pi(g) \] for any path $t \to P_{t, g} \in \mathcal{M}$. Finally, the test is locally unbiased if it satisfies $\pi(g) \le \alpha$ for all $t \to P_{t,g} \in \mathbf{P}$ and $\pi(g) \ge \alpha$ for all $t \mapsto P_{t,g} \in \mathcal{M} \setminus \mathbf{P}$.
It is clear from the above definitions that locally unbiased tests do not have nontrivial local power if $g \in \bar{T}(P)$ because the local deviation is consistent with the null hypothesis. Thus, locally unbiased tests cannot distinguish between $P$ and $P_{1/\sqrt{n}, g}$ if $g \in \bar{T}(P)$. The power of locally unbiased tests depends only on $\Pi_{T^\bot}(g)$.
Local unbiasedness of a test implies that its test statistic is asymptotically composed of elements of $\bar{T}(P)^\bot$. For instance, if a test statistic $T_n$ is asymptotically chi-squared distributed under $P$, then it satisfies \[ T_n = \sum_{j=1}^K \left( \frac{1}{\sqrt{n}} \sum_{i=1}^n f_j(X_i) \right)^2 +o_P(1) \] where $f_1, \dots, f_K$ are orthonormal and $K$ determines the degrees of freedom. Moreover, since the path is differentiable in quadratic mean at $t=0$, $P_{t,g}$ satisfies
for all $g \in L_0^2(P)$ (Lemma 25.14 of van1998). The advantage of having (ref) hold is that it allows us to derive the asymptotic distribution of a statistic under $P_{1/\sqrt{n},g}$ using the asymptotic distribution under $P$. By Le Cam's third lemma, we obtain \[ T_n \stackrel{g}{\rightsquigarrow} \chi_K^2 \left(\sum_{j=1}^K \mathbb{E}[f_j(X)g(X)]^2 \right), \] where $\stackrel{g}{\rightsquigarrow}$ denotes the weak convergence under $P_{1/\sqrt{n}, g}$ and $\chi_k^2(a)$ denotes the noncentral chi-squared distribution with degrees of freedom $k$ and the noncentrality parameter $a$. Because the locally unbiased test has nontrivial local power only if $\Pi_{T^\bot}(g)(X) \neq 0$, it must be the case that $f_j \in \bar{T}(P)^\bot$ for all $j=1, \dots, K$.
Next, we see that the asymptotic bias of asymptotically efficient estimators depends only on $\Pi_T(g)$. Note that any asymptotically efficient estimator for $\theta_0$ is asymptotically linear and can be written as \[ \sqrt{n}(\hat{\theta}_n - \theta_0) = \frac{1}{\sqrt{n}} \sum_{i=1}^n \nu(X_i) + o_P(1), \] where $\nu$ is the efficient influence function. Hence, by Le Cam's third lemma, we obtain \[ \sqrt{n}(\hat{\theta}_n - \theta_0) \stackrel{g}{\rightsquigarrow} N(\mathbb{E}[\nu(X)g(X)], \mathbb{E}[\nu(X) \nu(X)']). \] This means that the asymptotic distribution of $\hat{\theta}_n$ is unaffected by the local deviation if $\nu(X)$ is uncorrelated with $g(X)$. It is known that each element of the efficient influence function belongs to $\bar{T}(P)$. Therefore, the asymptotic bias depends only on $\Pi_{T}(g)$. If $g \in \bar{T}(P)^\bot$, then the asymptotic distribution of $\hat{\theta}_n$ under $P_{1/\sqrt{n}, g}$ is the same as that of under $P$.
Combining these results, we obtain the following proposition.
Proposition (ref) states that the directions of local deviation detectable by locally unbiased specification tests are orthogonal to those that generate asymptotic bias in asymptotically efficient estimators. Consequently, specification tests provide no information about the validity of estimators. A test may happen to correctly reject the null hypothesis when an estimator is locally biased. However, no locally unbiased specification test can distinguish between the cases $\Pi_{T}(g) = 0$ and $\Pi_{T}(g) \neq 0$, so any detection of bias is coincidental.
One might think that if $\Pi_{T^\bot}(g) \neq 0$, then $\Pi_T(g)$ is also likely to be nonzero. Even if such a belief can be justified, the test still provide no information about the magnitude of the bias. The magnitude of the test statistic provides no guidance as to whether the estimator is more or less appropriate.
In the discussion so far, we have not specified any particular model under the alternative hypothesis. If a maintained hypothesis exists, we consider the following null and alternative hypotheses:
where $\mathbf{M}$ is another semiparametric model satisfying $\mathbf{P} \subset \mathbf{M} \subset \mathcal{M}$. That is, we assume $\mu \in \mathbf{M}$ under both the null and alternative hypotheses. The parameter of interest $\theta_0$ is defined by $\theta_0 = \psi(P) = \varphi(P)$ for some functionals $\psi:\mathbf{P} \to \Theta$ and $\varphi:\mathbf{M} \to \Theta$ such that $\psi(Q) = \varphi(Q)$ for all $Q \in \mathbf{P}$.
In this setting, we can decompose $L_0^2(P)$ using the tangent set of model $\mathbf{M}$ at $P$. Let \[ M(P) = \left\{g \in L_0^2(P): \eqref{hellinger} \ \text{holds for some} \ t \mapsto P_{t,g} \in \mathbf{M} \right\} \] and let $\bar{M}(P)$ be the closure of $M(P)$. Since $\bar{T}(P) \subset \bar{M}(P)$, we have \[ L_0^2(P) = \bar{T}(P) \oplus \left\{ \bar{T}(P)^\bot \cap \bar{M}(P) \right\} \oplus \bar{M}(P)^\bot. \] Thus, any score $g \in L_0^2(P)$ can be decomposed as $g = \Pi_T(g) + \Pi_{T^\bot \cap M}(g) + \Pi_{M^\bot}(g)$. The same decomposition is also discussed in che_san2018.
It is clear that locally unbiased tests for (ref) have nontrivial local power for the deviation $P_{1/\sqrt{n}, g}$ only if $\Pi_{T^\bot \cap M}(g) \neq 0$. In contrast, asymptotically efficient estimators for $\theta_0$ in $\mathbf{P}$ are asymptotically biased only if $\Pi_T(g) \neq 0$. Thus, again, specification tests cannot detect local deviations that cause asymptotic bias in efficient estimators.
Hausman tests are representative examples of tests that fall within the above setup. Let $\hat{\theta}_n$ and $\tilde{\theta}_n$ be asymptotically efficient estimators for $\theta_0 \in \Theta$ in models $\mathbf{P}$ and $\mathbf{M}$, respectively. Suppose that $\mathbf{P} \subset \mathbf{M}$, so that $\hat{\theta}_n$ is more efficient than $\tilde{\theta}_n$. Two estimators can be written as
for some $\nu \in \bar{T}(P)$ and $\tau \in \bar{M}(P)$. Here, $\nu \in \bar{T}(P)$ means that all elements of $\nu$ belong to $\bar{T}(P)$ and the same for $\tau \in \bar{M}(P)$. The test statistic is given by \[ T_n = n(\tilde{\theta}_n - \hat{\theta}_n)' \hat{V}^{-1} (\tilde{\theta}_n -\hat{\theta}_n), \] where $\hat{V}$ is a consistent estimator for the asymptotic variance of $\tilde{\theta}_n -\hat{\theta}_n$. If the inverse matrix does not exist, it is replaced with the Moore-Penrose generalized inverse.
Under the above conditions, the Hausman test is a locally unbiased test for (ref). The reason is as follows. It is well known that the efficient influence function is obtained by projecting any other influence function onto the tangent space (see, e.g., van1998). Since $\hat{\theta}_n$ is more efficient than $\tilde{\theta}_n$, we have $\nu =\Pi_T(\tau)$ and $\tau - \nu \in \bar{T}(P)^\bot \cap \bar{M}(P)$. This implies that the well-known fact that the asymptotic variance of $\tilde{\theta}_n - \hat{\theta}_n$ equals the difference of the asymptotic variances of $\tilde{\theta}_n$ and $\hat{\theta}_n$. Moreover, since $\tau -\nu \in \bar{T}(P)^\bot \cap \bar{M}(P)$, the test statistic can be expressed as \[ T_n = \sum_{j=1}^K \left( \frac{1}{\sqrt{n}} \sum_{i=1}^n f_j(X_i) \right)^2 +o_P(1), \] where $f_1, \dots, f_K \in \bar{T}(P)^\bot \cap \bar{M}(P)$ are orthonormal. Thus, we have \[ T_n \stackrel{g}{\rightsquigarrow} \chi_K^2 \left(\sum_{j=1}^p \mathbb{E}[f_j(X)g(X)]^2 \right). \] The test has nontrivial local power only if $\Pi_{T^\bot \cap M}(g) \neq 0$.
Next, we investigate the asymptotic properties of two estimators. Applying Le Cam's third lemma, we obtain
Since $\nu \in \bar{T}(P)$, the asymptotic bias of $\hat{\theta}_n$ depends only on $\Pi_T(g)$, and thus the Hausman test cannot detect it. Moreover, when $g \in \bar{T}(P)$, both efficient and inefficient estimators share the same asymptotic bias. This occurs because $\psi(Q) = \varphi(Q)$ must hold for all $Q \in \mathbf{P}$. Furthermore, if $g \in \bar{T}(P)^\bot \cap \bar{M}(P)$, then only $\tilde{\theta}_n$ can be asymptotically biased. Thus, the test may reject the null hypothesis due to the bias of the inefficient estimator rather than that of the efficient estimator.
The above result may seem incompatible with the standard setting of Hausman tests, which assumes that the inefficient estimator is valid in all cases. The result becomes compatible if the directions of deviation from $P$ are restricted to those for which the inefficient estimator remains valid. However, this amounts to considering only limited deviations. If this restriction is incorrect, it may lead to misleading conclusions. We therefore take an agnostic view of the directions of deviation and analyze the behavior of estimators and test statistics over a broad class of deviations.
Another important feature of the Hausman test is that the estimators used to construct the test statistic determine the substantive null and alternative hypotheses. In the case of the DWH test, where researchers compare the OLS estimator with the 2SLS estimator, the null model is the homoskedastic linear regression model, since OLS is asymptotically efficient under this model. In contrast, the alternative model is the homoskedastic linear instrumental variable (IV) model. A detailed analysis of the DWH test is provided in Section (ref).
In this section, we examine the relationship between the efficient GMM estimator and the $J$ test, building on the framework introduced in Sections (ref) and (ref). More generally, the result applies to all asymptotically efficient estimators and all locally unbiased tests in moment restriction models. For instance, an analogous relationship holds between the empirical likelihood estimator and the empirical likelihood ratio statistic of qin_law1994.
To facilitate the analysis, we first reformulate the model following the approach of sue2024. This formulation is convenient for obtaining the tangent set in the conventional manner (e.g., Section 25.4 of van1998). Since model (ref) involves two unknown parameters $\theta_0$ and $P$, we write the model as a set of distributions indexed by the finite-dimensional parameter $\theta \in \Theta$ and the infinite-dimensional nuisance parameter $\eta \in \mathcal{M}$. Specifically, for given $\theta \in \Theta$ and $\eta \in \mathcal{M}$, we define $P_{\theta, \eta}$ as the solution to
where $\mathbf{P}_\theta =\{ Q \in \mathcal{M} : \int m_\theta dQ =0\}$. That is, $P_{\theta, \eta}$ is the projection of $\eta$ onto $\mathbf{P}_\theta$ with respect to the Kullback--Leibler divergence. By a duality theorem, $P_{\theta, \eta}$ satisfies \[ \frac{dP_{\theta, \eta}}{d\eta} = \frac{\exp(\lambda_{\theta, \eta}' m_\theta)}{\int \exp(\lambda_{\theta, \eta}' m_\theta) d\eta}, \] where $\lambda_{\theta, \eta} = \arg \min_{\lambda \in \mathbb{R}^l} \int \exp(\lambda'm_\theta) d\eta$. See bor_lew1991 and kom_rag2016 for details. Our model $\mathbf{P}=\{P_{\theta, \eta} :\theta \in \Theta, \eta \in \mathcal{M} \}$ is the set of all distributions on $\mathcal{X}$ that satisfy moment restrictions for some $\theta \in \Theta$.
We specify the tangent set of $\mathbf{P}$ at $P$. To do this, we consider a path of the form $P_t = P_{\theta_0+t h, \eta_t}$, where $h \in \mathbb{R}^p$ and $\eta_t$ is a perturbation from $P$ that coincides with $P$ at $t=0$. Under certain conditions, $P_t$ satisfies \[ \lim_{t \to 0} \int \left( \frac{dP_t^{1/2} -dP^{1/2}}{t} -\frac{1}{2} (h'\dot{\ell}_{\theta_0, \eta_0} + \dot{l} ) dP^{1/2} \right)^2 =0 \] where \[ \dot{\ell}_{\theta_0, \eta_0}(x)= - \mathbb{E}[\nabla m_{\theta_0}(X) ]' \Sigma^{-1} m_{\theta_0}(x) \] with $\nabla m_\theta =\partial m_\theta/ \partial \theta' $ and $\Sigma = \mathbb{E}[m_{\theta_0}(X)m_{\theta_0}(X)']$. Moreover, $\dot{l}: \mathcal{X} \to \mathbb{R}$ is an element of the set $\dot{\mathbf{P}}_\eta = \{ \dot{l} \in L_0^2(P) : \mathbb{E}[m_{\theta_0}(X) \dot{l}(X)] =0 \}$. See sue2024 for details. The function $\dot{\ell}_{\theta_0, \eta_0}$ is interpreted as the score function for $\theta_0$ when $\eta_0$ is fixed while $\dot{l}$ is interpreted as the score function for $\eta_0$ when $\theta_0$ is fixed. The tangent set is given by $T(P) = \{ {\rm lin} \ \dot{\ell}_{\theta_0, \eta_0} + \dot{\mathbf{P}}_\eta \}$, where ${\rm lin}$ denotes the linear span.
Notice that $\dot{\ell}_{\theta_0, \eta_0}$ is orthogonal to all elements of $\dot{\mathbf{P}}_\eta$. Thus, $\dot{\ell}_{\theta_0, \eta_0}$ is indeed the efficient score function for estimating $\theta_0$. Because the efficient information matrix is given by \[ I_{\theta_0, \eta_0} = \mathbb{E}[\dot{\ell}_{\theta_0, \eta_0}(X) \dot{\ell}_{\theta_0, \eta_0}(X)'] = \mathbb{E}\left[ \nabla m_{\theta_0}(X) \right]' \Sigma^{-1} \mathbb{E}\left[ \nabla m_{\theta_0}(X) \right], \] the efficient influence function is $I_{\theta_0, \eta_0}^{-1} \dot{\ell}_{\theta_0, \eta_0}$.
Now, we investigate the local asymptotic property of the GMM estimator. The efficient GMM estimator satisfies \[ \sqrt{n}(\hat{\theta}_n - \theta_0) = \frac{1}{\sqrt{n}} \sum_{i=1}^n I_{\theta_0, \eta_0}^{-1} \dot{\ell}_{\theta_0, \eta_0}(X_i) + o_P(1). \] This shows that the efficient GMM estimator is best regular in $\mathbf{P}$. Moreover, it follows from Le Cam's third lemma that
for any $g \in L_0^2(P)$. Here, the efficient influence function clearly belongs to $\bar{T}(P)$ and is therefore orthogonal to $\Pi_{T^\bot}(g)$. Thus, the asymptotic bias depends only on $\Pi_T(g)$.
The GMM estimator can be asymptotically unbiased even moment restrictions are violated because it utilizes only a part of moment restrictions. It follows from (ref) that the GMM estimator is asymptotically unbiased if
Although $\mathbb{E}[m_{\theta_0}(X) g(X)] \neq 0$ in general, the left-hand side of (ref) can be 0 because the rank of $\mathbb{E}[\nabla m_{\theta_0}(X)]' \Sigma^{-1}$ is $p$. Note that the left-hand side of (ref) is approximately the derivative of the GMM population objective function under local deviation. Evidently, if the objective function is minimized at $\theta_0$, then the GMM estimator is valid even if the moment restrictions are not satisfied at $\theta_0$.
Next, we investigate the local asymptotic property of the $J$ test statistic. It can be expressed as \[ J_n = \left( \frac{1}{\sqrt{n}} \sum_{i=1}^n m_{\hat{\theta}_n}(X_i) \right)' \Sigma^{-1} \left( \frac{1}{\sqrt{n}} \sum_{i=1}^n m_{\hat{\theta}_n}(X_i) \right) +o_P(1). \] Some calculation yields that \[ \Sigma^{-1/2}\frac{1}{\sqrt{n}} \sum_{i=1}^n m_{\hat{\theta}_n}(X_i) \stackrel{g}{\rightsquigarrow} N((I-P(\theta_0)) \Sigma^{-1/2} \mathbb{E}[m_{\theta_0}(X)g(X)], I-P(\theta_0)), \] where $P(\theta_0) =\Sigma^{-1/2} \mathbb{E}[\nabla m_{\theta_0}(X)] \left( \mathbb{E}[\nabla m_{\theta_0}(X)]' \Sigma^{-1} \mathbb{E}[\nabla m_{\theta_0}(X)] \right)^{-1} \mathbb{E}[\nabla m_{\theta_0}(X)]' \Sigma^{-1/2}$ is a projection matrix. It follows that $J_n$ converges in distribution to the noncentral chi-square distribution with degrees of freedom $l-p$ and noncentrality parameter \[ \mathbb{E}[m_{\theta_0}(X) g(X) ]' \Sigma^{-1/2} (I-P(\theta_0)) \Sigma^{-1/2} \mathbb{E}[m_{\theta_0}(X) g(X)]. \] If $g \in \bar{T}(P)$, then $g$ can be written as $g = h'\dot{\ell}_{\theta_0, \eta_0} + \dot{l}$ for some $h \in \mathbb {R}^p$ and $\dot{l} \in \dot{\mathbf{P}}_\eta$. In this case,
which implies that the noncentrality parameter is zero. Therefore, the $J$ test is locally unbiased for testing $H_0: \mu \in \mathbf{P}$ against $H_1: \mu \in \mathcal{M} \setminus \mathbf{P}$. In particular, it fails to detect the asymptotic bias in the efficient GMM estimator.
hal2005 showed a similar orthogonality result from a different perspective. He considered a local deviation of the form \[ \Sigma^{-1/2} \mathbb{E}_n[m_{\theta_0}(X)] = \frac{\delta}{\sqrt{n}} \] for some $\delta \in \mathbb{R}^l$, where $\mathbb{E}_n$ is the expectation with respect to a local deviation. The vector $\Sigma^{-1/2} \mathbb{E}_n[m_{\theta_0}(X)]$ can be decomposed as
Since $P(\theta_0)$ is a projection matrix, the two terms in (ref) are orthogonal to each other. hal2005 defines that the identifying restrictions are satisfied if the first term of (ref) is 0 and the overidentifying restrictions are satisfied if the second term is 0. He showed that the GMM estimator is asymptotically biased only if the identifying restrictions are locally violated while the $J$ test has nontrivial local power only if the overidentifying restrictions are locally violated.
Our decomposition of the score function essentially yields the same result as hal2005. If the expectation is taken with respect to $P_{1/\sqrt{n},g}$, then $\mathbb{E}_n[m_{\theta_0}(X)]$ is nearly equal to $\mathbb{E}[m_{\theta_0}(X)g(X)]/\sqrt{n}$. Because (ref) holds if $g \in \bar{T}(P)^\bot$, the first term of (ref) can be nonzero only if $\Pi_T(g) \neq 0$. In contrast, the second term can be nonzero only if $\Pi_{T^\bot}(g) \neq 0$ because (ref) holds if $g \in \bar{T}(P)$. Therefore, the decomposition of the score function produces the same result with hal2005.
This section investigates the properties of the DWH test. The test statistic is given by \[ T_n =(\tilde{\theta}_n-\hat{\theta}_n )' \hat{V}^{-} (\tilde{\theta}_n-\hat{\theta}_n), \] where $\hat{\theta}_n$ and $\tilde{\theta}_n$ denote the OLS and 2SLS estimators for $\theta_0 \in \mathbb{R}^p$ in the linear model: \[ Y =X'\theta_0 + e = X_1'\theta_{01} + X_2'\theta_{02} +e, \] where $X_1$ is suspected to be endogenous. Here, $\hat{V}^{-}$ denote the generalized inverse of a consistent estimator for the asymptotic variance of $\tilde{\theta}_n-\hat{\theta}_n$.
As stated in Section (ref), estimators constructing the test statistic implicitly define the null and alternative hypotheses. Let $Z_1$ be a vector of instrumental variables for $X_1$ and define $Z=(Z_1', X_2')'$. The null and maintained models are specified as the sets of joint distributions of $(X_1', Y, Z')'$. Since the OLS estimator is efficient in the homoskedastic linear regression model and the 2SLS estimator is efficient in the homoskedastic linear IV model, we consider the following models:
where $\mathbb{E}_Q[\cdot]$ and $\mathbb{E}_Q[\cdot|\cdot]$ denote the unconditional and conditional expectations with respect to $Q \in \mathcal{M}$. The parameter of interest $\theta_0$ satisfies \[ \mathbb{E}[Y-X'\theta_0 \mid X_1,Z] = 0 \quad \text{and} \quad \mathbb{E}[Z(Y-X'\theta_0)]=0. \]
It is important to note that the set of conditioning variables in $\mathbf{P}$ is $(X_1', Z')'$ rather than $X$. While conditioning only on $X$ is sufficient if we do not consider the maintained model, including $Z_1$ is necessary when the maintained model exists to ensure that the exclusion restriction is satisfied in both models. Without the exclusion restrictions, the two estimators may estimate completely different parameters even under the null hypothesis.
The tangent set of $\mathbf{M}$ can be obtained by using the result of Section (ref) with $m_{\theta}(x_1,y,z) =z(y-x'\theta)$. If we ignore the homoskedasticity assumption, the efficient score for estimating $\theta_0$ is given by $\dot{\ell}_{\theta_0, \eta_0}^{M}(x_1, y,z)=\mathbb{E}[XZ']\mathbb{E}[ZZ'e^2]^{-1}z(y-x'\theta_0)$. Under homoskedasticity, we have $\mathbb{E}[ZZ'e^2]=\sigma_0^2 \mathbb{E}[ZZ']$ for some constant $\sigma_0^2$. Hence, the efficient score function reduces to \[ \dot{\ell}_{\theta_0, \eta_0}^M(x_1, y, z) = \mathbb{E}[XZ'] \mathbb{E}[ZZ']^{-1} z(y-x'\theta_0)/\sigma_0^2. \] The tangent set is given by $M(P)=\{ {\rm lin} \ \dot{\ell}_{\theta_0,\eta_0}^M + \dot{\mathbf{M}}_\eta \}$, where \[ \dot{\mathbf{M}}_\eta =\left\{ \dot{l}^M \in L_0^2(P): \mathbb{E}[Ze \dot{l}^M(X_1,Y,Z)] =0 \right\}. \]
To obtain the tangent set of $\mathbf{P}$, we introduce a new formulation of the conditional moment restrictions model. As in the case of $\mathbf{M}$, we first rewrite the model as the set of distributions indexed by the parameter of interest $\theta \in \Theta$ and an infinite-dimensional nuisance parameter $\eta \in \mathcal{M}$. Then, our model can be written in the form $\mathbf{P}=\{P_{\theta, \eta} : \theta \in \Theta, \eta \in \mathcal{M} \}$. Using an argument similar to that in Section (ref), we obtain the following efficient score function: $\dot{\ell}_{\theta_0, \eta_0}^P(x_1, y, z) = x(y-x'\theta_0)/\mathbb{E}[e^2|X_1=x_1, Z=z]$. See the Appendix for the derivation. Under homoskedasticity, this reduces to \[ \dot{\ell}_{\theta_0,\eta_0}^P(x_1, y,z) = x (y-x'\theta_0)/\sigma_0^2, \] which confirms that the OLS estimator is best regular in model $\mathbf{P}$. The tangent set is given by $T(P)=\{ {\rm lin} \ \dot{\ell}_{\theta_0, \eta_0}^P + \dot{\mathbf{P}}_\eta\}$, where \[ \dot{\mathbf{P}}_\eta = \left\{ \dot{l}^P \in L_0^2(P) : \mathbb{E}[h(X_1, Z) e \dot{l}^P(X_1,Z, Y)]=0 \ \text{for any function} \ h \ \text{of} \ (X_1, Z) \right\}. \]
Now we show that the DWH test is locally unbiased for testing
By Le Cam's third lemma, we have
Suppose first that $g \in \bar{T}(P)$. Then $g$ can be written as $g = h' \dot{\ell}_{\theta_0, \eta_0}^P + \dot{l}^P$ for some $h \in \mathbb{R}^p$ and $\dot{l}^P \in \dot{\mathbf{P}}_\eta$. It follows that \[ \mathbb{E}[XX']^{-1} \mathbb{E}[Xe g(X_1, Y, Z)] = \mathbb{E}[XX']^{-1} \mathbb{E}[XX' e^2] h / \sigma_0^2 = h, \] and similarly, \[ \mathbb{E}[XZ'] \mathbb{E}[ZZ']^{-1} \mathbb{E}[Z e g(X_1, Y, Z)] = h. \] Thus, the OLS and 2SLS estimators have the same asymptotic bias, and the test has no local power. This is natural because if $\mathbb{E}_\mu[Y - X'\theta \mid X_1, Z] = 0$ holds for some $\theta$, then we must also have $\mathbb{E}_\mu[Z(Y - X'\theta)] = 0$ for the same $\theta$. It would therefore be incoherent to regard 2SLS as asymptotically unbiased while OLS is biased; both must share the same bias when $g \in \bar{T}(P)$.
Next, if $g \in \bar{T}(P)^\perp \cap \bar{M}(P)$, then $g$ is orthogonal to $\dot{\ell}_{\theta_0, \eta_0}^P$, and the OLS estimator is asymptotically unbiased. At the same time, $g$ can be expressed as $g = h' \dot{\ell}_{\theta_0, \eta_0}^M + \dot{l}^M$ for some $h \in \mathbb{R}^p$ and $\dot{l}^M \in \dot{\mathbf{M}}_\eta$, which implies that the 2SLS estimator can be asymptotically biased with bias $h$. Since the two estimators have different asymptotic means, the DWH test has nontrivial local power if $h \neq 0$. Note that rejection of the null hypothesis occurs due to the asymptotic bias of the 2SLS estimator, even when the OLS estimator is asymptotically unbiased.
Finally, suppose that $g \in \bar{M}(P)^\perp$. Because $g$ is orthogonal to both $\dot{\ell}_{\theta_0,\eta_0}^P$ and $\dot{\ell}_{\theta_0, \eta_0}^M$, it does not affect the asymptotic bias of either estimator. Therefore, the DWH test does not have local power. Combining these results, we see that the DWH test is locally unbiased for testing (ref).
We now explain how the OLS estimator can be asymptotically unbiased when $g \in \bar{T}(P)^{\bot}$. Because the null model is misspecified in this case, there is no $\theta$ that satisfies $\mathbb{E}_\mu[Y-X'\theta|X_1, Z]=0$ and $\mathbb{E}_\mu[(Y-X'\theta)^2|X_1, Z]=\sigma^2$. However, the case $\mathbb{E}_\mu[X(Y-X'\theta_0)]=0$ is not excluded, so that the OLS can be asymptotically unbiased. In this case, $X_1$ is uncorrelated with $e$ under $\mu$ even though the null hypothesis is not true. Therefore, the DWH test does not necessarily serve as a test of “exogeneity.”
This section reports Monte Carlo simulation results for the $J$ test and the DWH test. Although the theoretical analysis presented in preceding sections is conducted within the local misspecification framework, the main implications extend to global misspecification as well. Accordingly, the simulations are performed under the global misspecification setting. Our objective is to illustrate cases in which estimators remain valid while specification tests reject the null hypothesis with high probability, as well as cases in which the estimators are invalid yet the tests lack power.
This subsection investigates the relationship between the $J$ (Sargan) test and the 2SLS estimator. We draw a random sample $\{ Y_i, X_i, Z_{1i}, Z_{2i}\}_{i=1}^n $ generated according to the following system:
The potential instruments $(Z_{1i}, Z_{2i})'$ are mutually independent and follow standard normal distributions. The error terms \((v_i, e_i)'\) are generated from a bivariate normal distribution with mean zero, unit variances, and correlation $0.5$. We consider three DGPs (DGP1–DGP3) corresponding to the settings below.
DGP1 corresponds to the ideal distribution $P$, while DGP2 and DGP3 represent deviations from $P$. For DGP2, there is no value of $\theta$ that makes the instruments orthogonal to $u_i$. Nevertheless, the population objective function of the 2SLS estimator attains its minimum at $\theta_0$, the true parameter value used to generate the data. Consequently, we expect that the 2SLS estimator is valid, whereas the $J$ test rejects the null hypothesis with high probability. In contrast, DGP3 is designed such that the moment restrictions are not satisfied at $\theta_0$ but are satisfied at $\theta = \theta_0 + 0.2$. Accordingly, we expect that the $J$ test does not reject the null hypothesis, whereas the 2SLS estimator is invalid.
Table (ref) reports the simulation results based on 1,000 iterations with a sample size of $n = 500$. The true parameter is set to $\theta_0 = 1.5$. The $J$ test is implemented at the $5\%$ nominal significance level. We report the mean of the 2SLS estimator, the coverage rate of the associated $95\%$ confidence interval, and the rejection rate of the $J$ test. The results highlight the contrast between the 2SLS estimator and the $J$ test across DGP2 and DGP3: the 2SLS estimator remains valid under DGP2 but is invalid under DGP3, whereas the $J$ test detects misspecification under DGP2 but fails to reject under DGP3, consistent with our theoretical expectations.
This subsection investigates the relationship between the DWH test and the OLS estimator. We draw a random sample $\{ Y_i, X_{1i}, X_{2i}, Z_i \}_{i=1}^n$ generated according to the following system:
where $X_{2i}$ and $Z_i$ are mutually independent standard normal random variables. The error terms $(v_i, e_i)'$ are generated from a bivariate normal distribution with mean zero, unit variances, and correlation $\rho_{ve}$. The variable $Z_i$ serves as a potential instrument for $X_{1i}$. We consider three DGPs corresponding to the settings below.
DGP1 corresponds to the ideal distribution $P$, while DGP2 and DGP3 represent deviations from $P$. For DGP2, $\rho_{ve}$ is chosen so that the covariance between $X_i$ and $u_i$ is zero, whereas the instrument $Z_i$ is not valid. Consequently, we expect the DWH test to reject the null hypothesis due to the invalidity of the 2SLS estimator, even though the OLS estimator remains valid. For DGP3, both $X_{1i}$ and $Z_i$ are correlated with $u_i$. We set $\rho_{ve}$ such that the OLS and 2SLS estimators share the same probability limit. Therefore, we expect the DWH test not to reject the null hypothesis, despite both the OLS and 2SLS estimators being invalid.
Table (ref) reports the simulation results based on 1,000 iterations with the sample size of $n=500$. We set $\theta_{10} =\theta_{20}= 1.5$. The DWH test is conducted at the 5% nominal significance level. We report the mean of the OLS estimator for $\theta_{10}$, the coverage rate of the associated $95\%$ confidence interval, and the rejection rate of the DWH test. The results are again consistent with our theoretical predictions.
The simulation settings described above are somewhat artificial and are chosen to make the results transparent. However, it is entirely plausible to encounter cases in practice where an estimator exhibits only a small bias yet the null hypothesis is rejected, or where an estimator is invalid but the test cannot reject the null hypothesis.
This paper studied the local asymptotic properties of specification tests and asymptotically efficient estimators based on the framework of che_san2018. Although many studies have examined these properties separately, few have investigated the connection between specification tests and estimators. We showed that the directions of local deviation detectable by locally unbiased specification tests are orthogonal to those that induce asymptotic bias in asymptotically efficient estimators. Consequently, locally unbiased specification tests cannot detect bias in asymptotically efficient estimators. Our findings show that, although specification tests such as overidentification tests and Hausman tests are often used as pretests, their outcomes should not be taken as evidence either for or against the validity of estimators.
It may not be entirely unreasonable to believe that when one of the orthogonal components is nonzero (zero), the other is also nonzero (zero). However, rather than conducting a pretest based on such a belief, it seems more reasonable to dispense with pretesting and instead employ misspecification-robust methods, such as kit_et_al2013 and arm_kol2021, from the outset.