Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
449,744 characters · 17 sections · 74 citation commands
Using Invalid Instruments on Purpose: Focused Moment Selection and Averaging for GMM
In finite samples, the addition of a slightly endogenous but highly relevant instrument can reduce estimator variance by far more than bias is increased. Building on this observation, I propose a novel moment selection criterion for generalized method of moments (GMM) estimation: the focused moment selection criterion (FMSC). Rather than selecting only valid moment conditions, the FMSC chooses from a set of potentially mis-specified moment conditions based on the asymptotic mean squared error (AMSE) of their associated GMM estimators of a user-specified scalar target parameter $\mu$. To ensure a meaningful bias-variance tradeoff in the limit, I employ a drifting asymptotic framework in which mis-specification, while present for any fixed sample size, vanishes asymptotically. In the presence of such locally mis-specified moment conditions, GMM remains consistent although, centered and rescaled, its limiting distribution displays an asymptotic bias. Adding an additional mis-specified moment condition introduces a further source of bias while reducing asymptotic variance. The idea behind the FMSC is to trade off these two effects in the limit as an approximation to finite sample behavior.\footnote{When finite-sample MSE is undefined, AMSE comparisons remain meaningful: see Online Appendix (ref).} I suppose that two blocks of moment conditions are available: one that is assumed correctly specified, and another that may not be. This mimics the situation faced by an applied researcher who begins with a “baseline” set of relatively mild maintained assumptions and must decide whether to impose any of a collection of stronger but also more controversial “suspect” assumptions. When the (correctly specified) baseline moment conditions identify the model, the FMSC provides an asymptotically unbiased estimator of AMSE, allowing us select over the suspect moment conditions.\footnote{When this is not the case, it remains possible to use the AMSE framework to carry out a sensitivity analysis: see Online Appendix (ref).}
The primary goal of the FMSC is to select estimators with low AMSE, but researchers typically wish to report confidence intervals along with parameter estimates. Unfortunately the usual procedures for constructing asymptotic confidence intervals for GMM fail when applied to estimators chosen using a moment selection procedure. A “na\"{i}ve” 95% confidence interval constructed from the familiar textbook formula will generally under-cover: it will contain the true parameter value far less than 95% of the time because it fails to account for the additional sampling uncertainty that comes from choosing an estimator based on the data. To address the challenging problem of inference post-moment selection, I continue under the local mis-specification framework to derive the limit distribution of “moment average estimators,” data-dependent weighted averages of estimators based on different moment conditions. These estimators are interesting in their own right and include post-moment selection estimators as a special case. I propose two simulation-based procedures for constructing confidence intervals for moment average and post-selection estimators, including the FMSC. First is a “2-Step” confidence interval. I prove that this interval guarantees asymptotically valid inference: the asymptotic coverage of a nominal $100 \times (1 - \alpha)\%$ interval cannot fall below this level. The price of valid inference, however, is conservatism: the actual coverage of the 2-Step interval typically exceeds its nominal level.\footnote{This is unavoidable given certain impossibility results concerning post-selection inference. See, e.g.\ LeebPoetscher2005.} As a compromise between the conservatism of the 2-Step interval and the severe under-coverage of the na\"{i}ve interval I go on to propose a “1-Step” confidence interval. This interval is easier to compute than its 2-Step counterpart and performs well in empirically relevant examples, as I show both theoretically and in simulations below. The 1-Step interval is far shorter than the corresponding 2-Step interval and, while it can under-cover, the magnitude of the size distortion is modest compared to that of the na\"{i}ve intervals typically reported in applied work.
While my methods apply to general GMM models, I focus on two simple but empirically relevant examples: choosing between ordinary least squares (OLS) and two-stage least squares (TSLS) estimators, and selecting instruments in linear instrumental variables (IV) models. In the OLS versus TSLS example the FMSC takes a particularly transparent form, providing a risk-based justification for the Durbin-Hausman-Wu test, and leading to a novel “minimum-AMSE” averaging estimator that combines OLS and TSLS. The FMSC, averaging estimator, and related confidence interval procedures work well in practice, as I demonstrate in a series of simulation experiments and an empirical example from development economics.
The FMSC and minimum-AMSE averaging estimator considered here are derived for a scalar parameter interest, as this is the most common situation encountered in applied work.\footnote{For an extension of the FMSC to vector target parameters, see Online Appendix (ref).} As a consequence, Stein-type results do not apply: it is impossible to construct an estimator with uniformly lower risk than the “valid” estimator that uses only the baseline moment conditions. Nevertheless, as my simulation results show, selection and averaging can substantially outperform the valid estimator over large regions of the parameter space, particularly when the “suspect” moment conditions are highly informative and nearly correct. This is precisely the situation for which the FMSC is intended.
My approach to moment selection is inspired by the focused information criterion of ClaeskensHjort2003, a model selection criterion for maximum likelihood estimation. Like ClaeskensHjort2003, I study AMSE-based selection under mis-specification in a drifting asymptotic framework. In contradistinction, however, I consider moment rather than model selection, and general GMM rather than maximum likelihood estimation. Schorfheide2005 uses a similar approach to select over forecasts constructed from mis-specified vector autoregression models, developed independently of the FIC. While the use of locally mis-specified moment conditions dates back at least as far as Newey1985, the idea of using this framework for AMSE-based moment selection, however, is novel.
The existing literature on moment selection primarily aims to consistently select all correctly specified moment conditions while eliminating all invalid ones\footnote{Under the local mis-specification asymptotics considered below, consistent moment selection criteria simply choose all available moment conditions. For details, see Theorem (ref).} This idea begins with Andrews1999 and is extended by AndrewsLu and HongPrestonShum. More recently, Liao proposes a shrinkage procedure for consistent GMM moment selection and estimation. In a similar vein, CanerHanLee extend and generalize earlier work by Caner2009 on LASSO-type model selection for GMM to carry out simultaneous model and moment selection via an adaptive elastic net penalty. Whereas these proposals examine only the validity of the moment conditions under consideration, the FMSC balances validity against relevance to minimize AMSE. Although HallPeixe2003 and ChengLiao do consider relevance, their aim is to avoid including redundant moment conditions after consistently eliminating invalid ones. Some other papers that propose choosing, or combining, instruments to minimize MSE include DonaldNewey2001, DonaldImbensNewey2009, and KuersteinerOkui2010. Unlike the FMSC, however, these papers consider the higher-order bias that arises from including many valid instruments rather than the first-order bias that arises from the use of invalid instruments.
Another distinguishing feature of the FMSC is focus: rather than a one-size-fits-all criterion, the FMSC is really a method of constructing application-specific moment selection criteria. Consider, for example, a dynamic panel model. If your target parameter is a long-run effect while mine is a contemporaneous effect, there is no reason to suppose a priori that we should use the same moment conditions in estimation, even if we share the same model and dataset. The FMSC explicitly takes this difference of research goals into account.
Like Akaike's Information Criterion (AIC), the FMSC is a conservative rather than consistent selection procedure, as it remains random even in the limit. While consistency is a desirable property in many settings, the situation is more complex for model and moment selection: consistent and conservative selection procedures have different strengths, but these strengths cannot be combined Yang2005. The goal of this paper is estimators with low risk. Viewed from this perspective consistent selection criteria suffer from a serious defect: they exhibit unbounded minimax risk LeebPoetscher2008. Conservative criteria such as the FMSC do not suffer from this shortcoming. Moreover, as discussed in more detail below, the asymptotics of consistent selection paint a misleading picture of the effects of moment selection on inference. For these reasons, the fact that the FMSC is conservative rather than consistent is an asset in the present context.
Because it studies inference post-moment selection, this paper relates to a vast literature on “pre-test” estimators. For an overview, see LeebPoetscher2005, LeebPoetscher2009. There are several proposals to construct valid confidence intervals post-model selection, including Kabaila1998, HjortClaeskens and KabailaLeeb2006. To my knowledge, however, this is the first paper to treat the problem in general for post-moment selection and moment average estimators in the presence of mis-specification. Some related results appear in Berkowitz2008, Berkowitz2012, Guggenberger2010, Guggenberger2012, GuggenbergerKumar, and Caner2014. While I developed the simulation-based, two-stage confidence interval procedure described below by analogy to a suggestion in ClaeskensHjortbook, Leeb kindly pointed out that similar constructions have appeared in Loh1985, Berger1994, and Silvapulle1996. More recently, McCloskey takes a similar approach to study a class of non-standard testing problems.
The framework within which I study moment averaging is related to the frequentist model average estimators of HjortClaeskens. Two other papers that consider weighting estimators based on different moment conditions are Xiao and ChenChavezLinton. Whereas these papers combine estimators computed using valid moment conditions to achieve a minimum variance estimator, I combine estimators computed using potentially invalid conditions with the aim of reducing estimator AMSE. A related idea underlies the combined moments (CM) estimator of Judge2007. For a different approach to combining OLS and TSLS estimators, similar in spirit to the Stein-estimator and developed independently of the work presented here, see HansenStein. ChengLiaoShi provide related results for Stein-type moment averaging in a GMM context with potentially mis-specified moment conditions.
The results presented here are derived under strong identification and abstract from the many instruments problem. Supplementary simulation results presented in Online Appendix (ref), however, suggest that the FMSC can nevertheless perform well when the “baseline” assumptions only weakly identify the target parameter. Extending the idea behind the FMSC to allow for weak identification and possibly a large number of moment conditions is a challenging topic that I leave for future research.
The remainder of the paper is organized as follows. Section (ref) describes the asymptotic framework and Section (ref) derives the FMSC, both in general and for two specific examples: OLS versus TSLS and choosing instrumental variables. Section (ref) studies moment average estimators and shows how they can be used to construct valid confidence intervals post-moment selection. Section (ref) presents simulation results and Section (ref) considers an empirical example from development economics. Proofs appear at the end of the document; computational details and additional material appear in an Online Appendix.
Let $f(\cdot,\cdot)$ be a $(p+q)$-vector of moment functions of a random vector $Z$ and an $r$-dimensional parameter vector $\theta$, partitioned according to $f(\cdot,\cdot) = \left(g(\cdot,\cdot)', h(\cdot,\cdot)' \right)'$ where $g(\cdot,\cdot)$ and $h(\cdot,\cdot)$ are $p$- and $q$-vectors of moment functions. The moment condition associated with $g$ is assumed to be correct whereas that associated with $h$ is locally mis-specified. More precisely,
For any fixed sample size $n$, the expectation of $h$ evaluated at the true parameter value $\theta_0$ depends on the unknown constant vector $\tau$. Unless all components of $\tau$ are zero, some of the moment conditions contained in $h$ are mis-specified. In the limit however, this mis-specification vanishes, as $\tau/\sqrt{n}$ converges to zero. Uniform integrability combined with weak convergence implies convergence of expectations, so that $E[g(Z_i, \theta_0)]=0$ and $E[h(Z_i, \theta_0)]=0$. Because the limiting random vectors $Z_i$ are identically distributed, I suppress the $i$ subscript and simply write $Z$ to denote their common marginal law, e.g.\ $E[h(Z,\theta_0)]=0$. Local mis-specification is not intended as a literal description of real-world datasets: it is merely a device that gives an asymptotic bias-variance trade-off that mimics the finite-sample intuition. Moreover, while I work with an iid triangular array for simplicity, the results presented here can be adapted to handle dependent random variables.
Define the sample analogue of the expectations in Assumption (ref) as follows: $$f_n(\theta) = \frac{1}{n}\sum_{i=1}^n f(Z_{ni},\theta) = \left[
\right]=\left[
\right]$$ where $g_n$ is the sample analogue of the correctly specified moment conditions and $h_n$ is that of the (potentially) mis-specified moment conditions. A candidate GMM estimator $\widehat{\theta}_S$ uses some subset $S$ of the moment conditions contained in $f$ in estimation. Let $|S|$ denote the number of moment conditions used and suppose that $|S|>r$ so the GMM estimator is unique.\footnote{Identifying $\tau$ requires futher assumptions, as discussed in Section \ref{sec:ident}.} Let $\Xi_S$ be the $|S| \times(p +q)$ \emph{moment selection matrix} corresponding to $S$. That is, $\Xi_S$ is a matrix of ones and zeros arranged such that $\Xi_S f_n(\theta)$ contains only the sample moment conditions used to estimate $\widehat{\theta}_S$. Thus, the GMM estimator of $\theta$ based on moment set $S$ is given by $$\widehat{\theta}_S = \underset{\theta \in \Theta}{arg min}\; \left[\Xi_S f_n(\theta)\right]' \widetilde{W}_S \; \left[ \Xi_S f_n(\theta)\right].$$ where $\widetilde{W}_S$ is an $|S|\times |S|$, positive definite weight matrix. There are no restrictions placed on $S$ other than the requirement that $|S| >r$ so the GMM estimate is well-defined. In particular, $S$ may \emph{exclude} some or all of the valid moment conditions contained in $g$. This notation accommodates a wider range of examples, including choosing between OLS and TSLS estimators.
To consider the limit distribution of $\widehat{\theta}_S$, we require some further notation. First define the derivative matrices $$G = E\left[\nabla_{\theta} \; g(Z,\theta_0)\right], \quad H = E\left[\nabla_{\theta} \; h(Z,\theta_0)\right], \quad F = (G', H')'$$ and let $\Omega = Var\left[ f(Z,\theta_0) \right]$ where $\Omega$ is partitioned into blocks $\Omega_{gg}$, $\Omega_{gh}$, $\Omega_{hg}$, and $\Omega_{hh}$ conformably with the partition of $f$ by $g$ and $h$. Notice that each of these expressions involves the limiting random variable $Z$ rather than $Z_{ni}$, so that the corresponding expectations are taken with respect to a distribution for which all moment conditions are correctly specified. Finally, to avoid repeatedly writing out pre- and post-multiplication by $\Xi_S$, define $F_S = \Xi_S F$ and $\Omega_S = \Xi_S \Omega\Xi_S'$. The following high level assumptions are sufficient for the consistency and asymptotic normality of the candidate GMM estimator $\widehat{\theta}_S$.
Although Assumption (ref) closely approximates the standard regularity conditions for GMM estimation, establishing primitive conditions for Assumptions (ref) (d), (e), (g) and (h) is slightly more involved under local mis-specification. Low-level sufficient conditions for the two running examples considered in this paper appear in Online Appendix (ref). For more general results, see Andrews1988 Theorem 2 and Andrews1992 Theorem 4. Notice that identification, (c), and continuity, (d), are conditions on the distribution of $Z$, the marginal law to which each $Z_{ni}$ converges.
As we see from Theorems (ref) and (ref), any candidate GMM estimator $\widehat{\theta}_S$ is consistent for $\theta_0$ under local mis-specification. Unless $S$ excludes all of the moment conditions contained in $h$, however, $\widehat{\theta}_S$ inherits an asymptotic bias from the mis-specification parameter $\tau$. The local mis-specification framework is useful precisely because it results in a limit distribution for $\widehat{\theta}_S$ with both a bias and a variance. This captures in asymptotic form the bias-variance tradeoff that we see in finite sample simulations. In constrast, fixed mis-specification results in a degenerate bias-variance tradeoff in the limit: scaling up by $\sqrt{n}$ to yield an asymptotic variance causes the bias component to diverge.
Any form of moment selection requires an identifying assumption: we need to make clear which parameter value $\theta_0$ counts as the “truth.” One approach, following Andrews1999, is to assume that there exists a unique, maximal set of correctly specified moment conditions that identifies $\theta_0$. In the notation of the present paper\footnote{Although Andrews1999, AndrewsLu, and HongPrestonShum consider fixed mis-specification, we can view this as a version of local mis-specification in which $\tau \rightarrow \infty$ sufficiently fast.} this is equivalent to the following:
AndrewsLu and HongPrestonShum take the same basic approach to identification, with appropriate modifications to allow for simultaneous model and moment selection. An advantage of Assumption (ref) is that, under fixed mis-specification, it allows consistent selection of $S_{max}$ without any prior knowledge of which moment conditions are correct. In the notation of the present paper this corresponds to having no moment conditions in the $g$ block. As Hallbook points out, however, the second part of Assumption (ref) can fail even in very simple settings. When it does fail, the selected GMM estimator may no longer be consistent for $\theta_0$. A different approach to identification is to assume that there is a minimal set of at least $r$ moment conditions known to be correctly specified. This is the approach I follow here, as do Liao and ChengLiao.\footnote{For a dicussion of why Assumption (ref) is necessary and how to proceed when it fails, see Online Appendix (ref).}
Assumption (ref) and Theorem (ref) immediately imply that the valid estimator shows no asymptotic bias.
Both Assumptions (ref) and (ref) are strong, and neither fully nests the other. In the context of the present paper, Assumption (ref) is meant to represent a situation in which an applied researcher chooses between two groups of assumptions. The $g$--block contains the “baseline” assumptions while the $h$--block contains a set of stronger, more controversial “suspect” assumptions. The FMSC is designed for settings in which the $h$--block is expected to contain a substantial amount of information beyond that already contained in the $g$--block. The idea is that, if we knew the $h$--block was correctly specified, we would expect a large gain in efficiency by including it in estimation. This motivates the idea of trading off the variance reduction from including $h$ against the potential increase in bias.
The FMSC chooses among the potentially invalid moment conditions contained in $h$ based on the estimator AMSE of a user-specified scalar target parameter.\footnote{Although I focus on the case of a scalar target parameter in the body of the paper, the same idea can be applied to a vector of target parameters. For details see Online Appendix (ref).} Denote this target parameter by $\mu$, a real-valued, $Z$-almost continuous function of the parameter vector $\theta$ that is differentiable in a neighborhood of $\theta_0$. Further, define the GMM estimator of $\mu$ based on $\widehat{\theta}_S$ by $\widehat{\mu}_S = \mu(\widehat{\theta}_S)$ and the true value of $\mu$ by $\mu_0 = \mu(\theta_0)$. Applying the Delta Method to Theorem (ref) gives the AMSE of $\widehat{\mu}_S$.
For the valid estimator $\widehat{\theta}_v$ we have $K_v = \left[G'W_{v}G\right]^{-1}G' W_{v}$ and $\Xi_v =\left[
\right]$. Thus, the valid estimator $\widehat{\mu}_v$ of $\mu$ has zero asymptotic bias. In contrast, any candidate estimator $\widehat{\mu}_S$ that includes moment conditions from $h$ inherits an asymptotic bias from the corresponding elements of $\tau$, the extent and direction of which depends both on $K_S$ and $\nabla_\theta\mu(\theta_0)$. The setting considered here, however, is one in which using moment conditions from $h$ in estimation will reduce the asymptotic variance. In the nested case, where moment conditions from $h$ are \emph{added} to those of $g$, this follows automatically. The usual proof that adding moment conditions cannot increase asymptotic variance under efficient GMM \citep[see for example][ch.\ 6]{Hallbook} continues to hold under local mis-specification, because all moment conditions are correctly specified in the limit. In non-nested examples, for example when $h$ contains OLS moment conditions and $g$ contains IV moment conditions, however, this result does not apply because one would use $h$ \emph{instead of} $g$. In such examples, one must establish an analogous ordering of asymptotic variances by direct calculation, as I do below for the OLS versus IV example.
Using this framework for moment selection requires estimators of the unknown quantities: $\theta_0$, $K_S$, $\Omega$, and $\tau$. Under local mis-specification, the estimator of $\theta$ under any moment set is consistent. A natural estimator is $\widehat{\theta}_v$, although there are other possibilities. Recall that $K_S = [F_S'W_SF_S]^{-1} F_S'W_S \Xi_S$. Because it is simply the selection matrix defining moment set $S$, $\Xi_S$ is known. The remaining quantities $F_S$ and $W_S$ that make up $K_S$ are consistently estimated by their sample analogues under Assumption (ref). Similarly, consistent estimators of $\Omega$ are readily available under local mis-specification, although the precise form depends on the situation.\footnote{See Sections (ref) and (ref) for discussion of this point for the two running examples.} The only remaining unknown is $\tau$. Local mis-specification is essential for making meaningful comparisons of AMSE because it prevents the bias term from dominating the comparison. Unfortunately, it also prevents consistent estimation of the asymptotic bias parameter. Under Assumption (ref), however, it remains possible to construct an asymptotically unbiased estimator $\widehat{\tau}$ of $\tau$ by substituting $\widehat{\theta}_v$, the estimator of $\theta_0$ that uses only correctly specified moment conditions, into $h_n$, the sample analogue of the potentially mis-specified moment conditions. In other words, $\widehat{\tau} = \sqrt{n} h_n(\widehat{\theta}_v)$.
Returning to Corollary $\ref{cor:target}$, however, we see that it is $\tau \tau'$ rather than $\tau$ that enters the expression for AMSE. Although $\widehat{\tau}$ is an asymptotically unbiased estimator of $\tau$, the limiting expectation of $\widehat{\tau} \widehat{\tau}'$ is not $\tau\tau'$ because $\widehat{\tau}$ has an asymptotic variance. Subtracting a consistent estimate of the asymptotic variance removes this asymptotic bias.
It follows that
provides an asymptotically unbiased estimator of AMSE. Given a set $\mathscr{S}$ of candidate specifications, the FMSC selects the candidate $S^*$ that minimizes the expression given in Equation (ref), that is $S^*_{FMSC} = \arg \min_{S\in \mathscr{S}} \;\mbox{FMSC}_n(S)$.
In summary, the FMSC aims to choose the moment conditions that provide the lowest risk estimator of a target parameter $\mu$ where risk is defined as MSE.\footnote{One could choose a different risk function and proceed similarly, although I do not consider this idea further below. See, e.g., Claeskens2006 and ClaeskensHjort2008.} Because finite-sample MSE is unavailable, AMSE in a local-to-zero asymptotic framework serves in its stead. Since no consistent estimator of AMSE exists in this setting, FMSC uses an asymptotically unbiased estimator. This is the same idea that underlies the classical AIC and TIC model selection criteria as well as more recent procedures such as those described in ClaeskensHjort2003 and Schorfheide2005.
The simplest interesting application of the FMSC is choosing between ordinary least squares (OLS) and two-stage least squares (TSLS) estimators of the effect $\beta$ of a single endogenous regressor $x$ on an outcome of interest $y$. The intuition is straightforward: because TSLS is a high-variance estimator, OLS will have a lower mean-squared error provided that $x$ isn't too endogenous.\footnote{Because the moments of the TSLS estimator only exist up to the order of overidentificiation Phillips1980, Kinal mean-squared error should be understood to refer to “trimmed” mean-squared error when the number of instruments is two or fewer. For details, see Online Appendix (ref).} To keep the presentation transparent, I work within an iid, homoskedastic setting for this example and assume, without loss of generality, that there are no exogenous regressors.\footnote{The homoskedasticity assumption concerns the limit random variables: under local mis-specification there will be heteroskedasticity for fixed $n$. See Assumption (ref) in Online Appendix (ref) for details.} Equivalently we may suppose that any exogenous regressors, including a constant, have been “projected out.” Low-level sufficient conditions for all of the results in this section appear in Assumption (ref) of Online Appendix (ref). The data generating process is
where $\beta$ and $\boldsymbol{\pi}$ are unknown constants, $\mathbf{z}_{ni}$ is a vector of exogenous and relevant instruments, $x_{ni}$ is the endogenous regressor, $y_{ni}$ is the outcome of interest, and $\epsilon_{ni}, v_{ni}$ are unobservable error terms. All random variables in this system are mean zero, or equivalently all constant terms have been projected out. Stacking observations in the usual way, the estimators under consideration are $\widehat{\beta}_{OLS} = \left(\mathbf{x}'\mathbf{x}\right)^{-1}\mathbf{x}'\mathbf{y}$ and $\widetilde{\beta}_{TSLS} = \left(\mathbf{x}'P_Z\mathbf{x}\right)^{-1}\mathbf{x}'P_Z\mathbf{y}$ where we define $P_Z = Z(Z'Z)^{-1}Z'$.
We see immediately that, as expected, the variance of the OLS estimator is always strictly lower than that of the TSLS estimator since $\sigma^2_\epsilon/\sigma_x^2 = \sigma^2_\epsilon/(\gamma^2 + \sigma_v^2)$. Unless $\tau = 0$, however, OLS shows an asymptotic bias. In contrast, the TSLS estimator is asymptotically unbiased regardless of the value of $\tau$. Thus, $$\mbox{AMSE(OLS)} = \frac{\tau^2}{\sigma_x^4} + \frac{\sigma_\epsilon^2}{\sigma_x^2},\quad \quad \mbox{AMSE(TSLS)} = \frac{\sigma_\epsilon^2}{\gamma^2}.$$ and rerranging, we see that the AMSE of the OLS estimator is strictly less than that of the TSLS estimator whenever $\tau^2 < \sigma_x^2 \sigma_\epsilon^2\sigma_v^2/\gamma^2$. To estimate the unknowns required to turn this inequality into a moment selection procedure, I set $$\widehat{\sigma}_x^2 = n^{-1}\mathbf{x}'\mathbf{x}, \quad \widehat{\gamma}^2 = n^{-1}\mathbf{x}'Z(Z'Z)^{-1}Z'\mathbf{x}, \quad \widehat{\sigma}_v^2 = \widehat{\sigma}_x^2 - \widehat{\gamma}^2$$ and define $$\widehat{\sigma}_\epsilon^2 = n^{-1}\left(\textbf{y} - \textbf{x}\widetilde{\beta}_{TSLS} \right)'\left(\textbf{y} - \textbf{x}\widetilde{\beta}_{TSLS} \right)$$ Under local mis-specification each of these estimators is consistent for its population counterpart.\footnote{While using the OLS residuals to estimate $\sigma_\epsilon^2$ also provides a consistent estimate under local mis-specification, the estimator based on the TSLS residuals should be more robust.} All that remains is to estimate $\tau^2$. Specializing Theorem (ref) and Corollary (ref) to the present example gives the following result.
It follows that $\widehat{\tau}^2 - \widehat{\sigma}_\epsilon^2\widehat{\sigma}_x^2 \left(\widehat{\sigma}_v^2/\widehat{\gamma}^2\right)$ is an asymptotically unbiased estimator of $\tau^2$ and hence, substituting into the AMSE inequality from above and rearranging, the FMSC instructs us to choose OLS whenever $\widehat{T}_{FMSC} = \widehat{\tau}^2/\widehat{V} < 2$ where $\widehat{V} = \widehat{\sigma}_v^2 \widehat{\sigma}_\epsilon^2 \widehat{\sigma}_x^2/\widehat{\gamma}^2$. The quantity $\widehat{T}_{FMSC}$ looks very much like a test statistic and indeed it can be viewed as such. By Theorem (ref) and the continuous mapping theorem, $\widehat{T}_{FMSC} \rightarrow_d \chi^2(1)$. Thus, the FMSC can be viewed as a test of the null hypothesis $H_0\colon \tau = 0$ against the two-sided alternative with a critical value of $2$. This corresponds to a significance level of $\alpha \approx 0.16$. But how does this novel “test” compare to something more familiar, say the Durbin-Hausman-Wu (DHW) test? It turns out that in this particular example, although not in general, the FMSC is numerically equivalent to using OLS unless the DHW test rejects at the 16% level.
The equivalence between FMSC selection and a DHW test in this example is helpful for two reasons. First, it provides a novel justification for the use of the DHW test to select between OLS and TSLS. So long as it is carried out with $\alpha \approx 16\%$, the DHW test is equivalent to selecting the estimator that minimizes an asymptotically unbiased estimator of AMSE. Note that this significance level differs from the more usual values of 5% or 10% in that it leads us to select TSLS more often: OLS should indeed be given the benefit of the doubt, but not by so wide a margin as traditional practice suggests. Second, this equivalence shows that the FMSC can be viewed as an extension of the idea behind the familiar DHW test to more general GMM environments.\footnote{Note that the FMSC in this example, characterized in Theorem (ref), chooses between OLS and IV to minimize estimator AMSE. If one wishes to carry out inference post-selection one must contend with the size distortions of the familiar “textbook” confidence interval procedure, as pointed out by Guggenberger2010. I discuss this point extensively below in Section (ref) ans propose possible remedies.}
The OLS versus TSLS example is really a special case of instrument selection: if $x$ is exogenous, it is clearly “its own best instrument.” Viewed from this perspective, the FMSC amounts to trading off endogeneity against instrument strength. I now consider instrument selection in general for linear GMM estimators in an iid setting. Consider the model:
where $y$ is an outcome of interest, $\mathbf{x}$ is an $r$-vector of regressors, some of which are endogenous, $\mathbf{z}^{(1)}$ is a $p$-vector of instruments known to be exogenous, and $\mathbf{z}^{(2)}$ is a $q$-vector of potentially endogenous instruments. The $r$-vector $\beta$, $p\times r$ matrix $\Pi_1$, and $q\times r$ matrix $\Pi_2$ contain unknown constants. Stacking observations in the usual way, we can write the system in matrix form as $\mathbf{y} = X\beta +\boldsymbol{\epsilon}$ and $X = Z \Pi + V$, where $Z = (Z_1, Z_2)$ and $\Pi = (\Pi_1', \Pi_2')'$.
In this example, the idea is that the instruments contained in $Z_2$ are expected to be strong. If we were confident that they were exogenous, we would certainly use them in estimation. Yet the very fact that we expect them to be strongly correlated with $\mathbf{x}$ gives us reason to fear that they may be endogenous. The exact opposite is true of $Z_1$: these are the instruments that we are prepared to assume are exogenous. But when is such an assumption plausible? Precisely when the instruments contained in $Z_1$ are not especially strong. Accordingly, the FMSC attempts to trade off a small increase in bias from using a slightly endogenous instrument against a larger decrease in variance from increased instrument strength. To this end, consider a general linear GMM estimator of the form $$\widehat{\beta}_S = (X'Z_S \widetilde{W}_S Z_S' X)^{-1}X'Z_S \widetilde{W}_S Z_S' \mathbf{y}$$ where $S$ indexes the instruments used in estimation, $Z_S' = \Xi_S Z'$ is the matrix containing only those instruments included in $S$, $|S|$ is the number of instruments used in estimation and $\widetilde{W}_S$ is an $|S|\times|S|$ positive definite weighting matrix.
To implement the FMSC for this example, we simply need to specialize Equation (ref). To simplify the notation, let
where $0_{p\times q}$ denotes a $p\times q$ matrix of zeros and $\mathbf{I}_q$ denotes the $q\times q$ identity matrix. Using this convention, $Z_1 = Z \Xi_1'$ and $Z_2 = Z \Xi_2'$. In this example the valid estimator, defined in Assumption (ref), is given by
and we estimate $\nabla_\beta \mu(\beta)$ with $\nabla_\beta \mu(\widehat{\beta}_v)$. Similarly, $$-\widehat{K}_S = n\left(X'Z \Xi_S' \widetilde{W}_S \Xi_S Z' X\right)^{-1}X' Z \Xi_S' \widetilde{W}_S$$ is the natural consistent estimator of $-K_S$ in this setting.\footnote{The negative sign is squared in the FMSC expression and hence disappears. I write it here only to be consistent with the notation of Theorem (ref).} Since $\Xi_S$ is known, the only remaining quantities from Equation (ref) are $\widehat{\boldsymbol{\tau}}$, $\widehat{\Psi}$ and $\widehat{\Omega}$. The following result specializes Theorem (ref) to the present example.
Using this result, I construct the asymptotically unbiased estimator $\widehat{\tau}\widehat{\tau}' - \widehat{\Psi}\widehat{\Omega} \widehat{\Psi}'$ of $\tau\tau'$ from $$\widehat{\Psi} = \left[
\right], \quad -\widehat{K}_v = n\left(X'Z_1 \widetilde{W}_v Z_1' X\right)^{-1}X'Z_1 \widetilde{W}_v$$
All that remains before substituting values into Equation (ref) is to estimate $\Omega$. In the simulation and empirical examples discussed below I examine the TSLS estimator, that is $\widetilde{W}_S = (\Xi_S Z'Z\Xi_S)^{-1}$, and estimate $\Omega$ as follows. For all specifications except the valid estimator $\widehat{\beta}_v$, I employ the centered, heteroskedasticity-consistent estimator
where $u_i(\beta) = y_i - \mathbf{x}_i'\beta$, $\widehat{\beta}_S = (X'Z_S(Z_S'Z_S)^{-1}Z_S'X)^{-1}X'Z_S(Z_S'Z_S)^{-1}Z_S'\mathbf{y}$, $\mathbf{z}_{iS} = \Xi_S \mathbf{z}_i$ and $Z_S' = \Xi_S Z'$. Centering allows moment functions to have non-zero means. While the local mis-specification framework implies that these means tend to zero in the limit, they are non-zero for any fixed sample size. Centering accounts for this fact, and thus provides added robustness. Since the valid estimator $\widehat{\beta}_v$ has no asymptotic bias, the AMSE of any target parameter based on this estimator equals its asymptotic variance. Accordingly, I use
rather than the $(p\times p)$ upper left sub-matrix of $\widehat{\Omega}$ to estimate this quantity. This imposes the assumption that all instruments in $Z_1$ are valid so that no centering is needed, providing greater precision.
Because it is constructed from $\widehat{\tau}$, the FMSC is a random variable, even in the limit. Combining Corollary (ref) with Equation (ref) gives the following.
This corollary implies that the FMSC is a “conservative” rather than “consistent” selection procedure. This lack of consistency is a desirable feature of the FMSC for two reasons. First, as discussed above, the goal of the FMSC is not to select only correctly specified moment conditions: it is to choose an estimator with a low finite-sample MSE as approximated by AMSE. The goal of consistent selection is very much at odds with that of controlling estimator risk. As explained by Yang2005 and LeebPoetscher2008, the worst-case risk of a consistent selection procedure diverges with sample size.\footnote{This fact is readily apparent from the results of the simulation study from Section (ref): the consistent criteria, GMM-BIC and HQ, have the highest worst-case RMSE, while the conservative criteria, FMSC and GMM-AIC, have the lowest.} Second, while we know from both simulation studies Demetrescu and analytical examples LeebPoetscher2005 that selection can dramatically change the sampling distribution of our estimators, invalidating traditional confidence intervals, the asymptotics of consistent selection give the misleading impression that this problem can be ignored.
There are two main problems with applying “textbook” confidence intervals post-moment selection. First is model selection uncertainty: if the data had been slightly different, we would have chosen a different set of moment conditions. Accordingly, any confidence interval that conditions on the selected model must be too short. Second, textbook confidence intervals ignore the fact that selection is carried out over potentially invalid moment conditions. Even if our goal were to consistently eliminate such moment conditions, for example by using a consistent criterion such as the GMM-BIC of Andrews1999, in finite-samples we would not always be successful. Because of this, our intervals will be incorrectly centered. Accounting for these two effects requires a limit theory that accommodates mixture distributions: post-selection estimators are randomly-weighted averages of the individual candidate estimators. Because they choose a single candidate with probability approaching one in the limit, consistent selection procedures make it impossible to represent this phenomenon. In contrast, conservative selection procedures remain random even as the sample size goes to infinity, allowing us to derive a mixture-of-normals limit distribution and, ultimately, to carry out valid inference post-moment selection. In the remainder of this section, I derive the asymptotic distribution of generic “moment average” estimators and use them to propose simulation-based procedures for post-moment selection inference. For certain examples it is possible to analytically characterize the limit distribution of a post-FMSC estimator without resorting to simulation-based methods. I explore this possibility in detail for my two running examples: OLS versus TSLS and choosing instrumental variables. I also briefly consider a minimum-AMSE averaging estimator that combines OLS and TSLS.
A generic moment average estimator takes the form
where $\widehat{\mu}_S = \mu(\widehat{\theta}_S)$ is the estimator of the target parameter $\mu$ under moment set $S$, $\mathscr{S}$ is the collection of all moment sets under consideration, and $\widehat{\omega}_S$ is shorthand for the value of a data-dependent weight function $\widehat{\omega}_S=\omega(\cdot, \cdot)$ evaluated at moment set $S$ and the sample observations $Z_{n1}, \hdots, Z_{nn}$. As above $\mu(\cdot)$ is a $\mathbb{R}$-valued, $Z$-almost surely continuous function of $\theta$ that is differentiable in an open neighborhood of $\theta_0$. When $\widehat{\omega}_S$ is an indicator, taking on the value one at the moment set moment set that minimizes some moment selection criterion, $\widehat{\mu}$ is a post-moment selection estimator. To characterize the limit distribution of $\widehat{\mu}$, I impose the following mild conditions on $\widehat{\omega}_S$, requiring that they sum to one and are “well-behaved” in the limit so that I may apply the continuous mapping theorem.
Notice that the limit random variable from Corollary (ref), denoted $\Lambda(\tau)$, is a randomly weighted average of the multivariate normal vector $M$. Hence, $\Lambda(\tau)$ is non-normal. This is precisely the behavior for which we set out to construct an asymptotic representation. The conditions of Assumption (ref) are fairly mild. Requiring that the weights sum to one ensures that $\widehat{\mu}$ is a consistent estimator of $\mu_0$ and leads to a simpler expression for the limit distribution. While somewhat less transparent, the second condition is satisfied by weighting schemes based on a number of familiar moment selection criteria. It follows immediately from Corollary (ref), for example, that the FMSC converges in distribution to a function of $\tau$, $M$ and consistently estimable constants only. The same is true for weights based on the $J$-test statistic, as seen from the following result.
Post-selection estimators are merely a special case of moment average estimators. To see why, consider the weight function $$\widehat{\omega}_S^{MSC} = \mathbf{1}\left\{\mbox{MSC}_n(S) = \min_{S'\in \mathscr{S}} \mbox{MSC}_n(S')\right\}$$where $\mbox{MSC}_n(S)$ is the value of some moment selection criterion evaluated at the sample observations $Z_{n1}\hdots, Z_{nn}$. Now suppose $\mbox{MSC}_n(S) \rightarrow_d\mbox{MSC}_S(\tau,M)$, a function of $\tau$, $M$ and consistently estimable constants only. Then, so long as the probability of ties, $P\left\{\mbox{MSC}_S(\tau,M) = \mbox{MSC}_{S'}(\tau,M) \right\}$, is zero for all $S\neq S'$, we have $$\widehat{\omega}_S^{MSC} \rightarrow_d \mathbf{1}\left\{\mbox{MSC}_S(\tau,M) = \min_{S'\in \mathscr{S}} \mbox{MSC}_{S'}(\tau,M)\right\}$$ satisfying Assumption (ref) (b). Thus, post-selection estimators based on the FMSC, a downward $J$-test procedure, or the GMM moment selection criteria of Andrews1999 all fall within the ambit of (ref). The consistent criteria of Andrews1999, however, are not particularly interesting under local mis-specification.\footnote{For more discussion of these criteria, see Section (ref) below.} Intuitively, because they aim to select all valid moment conditions w.p.a.1, we would expect that under Assumption (ref) they choose the full moment set in the limit. The following result shows that this intuition is correct.\footnote{This result is a special case of a more general phenomenon: consistent selection procedures cannot detect model violations of order $O(n^{-1/2})$.}
When competing moment sets have similar criterion values in the population, sampling variation can be magnified in the selected estimator. This motivates the idea of averaging estimators based on different moment conditions rather than selecting them. To illustrate this idea, I now briefly revisit the OLS versus TSLS example from Section (ref) and derive an AMSE-optimal weighted average of the two estimators. Let $\widetilde{\beta}(\omega)$ be a convex combination of the OLS and TSLS estimators, namely
where $\omega \in [0,1]$ is the weight given to the OLS estimator.
The preceding result has several important consequences. First, since the variance of the TSLS estimator is always strictly greater than that of the OLS estimator, the optimal value of $\omega$ cannot be zero. No matter how strong the endogeneity of $x$, as measured by $\tau$, we should always give some weight to the OLS estimator. Second, when $\tau = 0$ the optimal value of $\omega$ is one. If $x$ is exogenous, OLS is strictly preferable to TSLS. Third, the optimal weights depend on the strength of the instruments $\mathbf{z}$ as measured by $\gamma$. All else equal, the stronger the instruments, the less weight we should give to OLS. To operationalize the AMSE-optimal averaging estimator suggested from Theorem (ref), I propose the plug-in estimator
where
This expression employs the same consistent estimators of $\sigma_x^2, \gamma$ and $\sigma_{\epsilon}$ as the FMSC expressions from Section (ref). To ensure that $\widehat{\omega}^*$ lies in the interval $[0,1]$, however, I use a positive part estimator for $\tau^2$, namely $\max\{0, \; \widehat{\tau}^2 - \widehat{V}\}$ rather than $\widehat{\tau}^2 - \widehat{V}$.\footnote{While $\widehat{\tau}^2 - \widehat{V}$ is an asymptotically unbiased estimator of $\tau^2$ it can be negative.} In the following section I show how one can construct confidence intervals for $\widehat{\beta}^*$ and related estimators.
Suppose that $K_S$, $\varphi_S$, $\theta_0$, $\Omega$ and $\tau$ were all known. Then, by simulating from $M$, as defined in Theorem (ref), the distribution of $\Lambda(\tau)$, defined in Corollary (ref), could be approximated to arbitrary precision. This is the basic intuition that I use to devise inference procedures for moment-average and post-selection estimators.
To operationalize this idea, first consider how we would proceed if we knew only the value of $\tau$. While $K_S$, $\theta_0$, and $\Omega$ are unknown this presents only a minor difficulty: in their place we can simply substitute the consistent estimators that appeared in the expression for the FMSC above. To estimate $\varphi_S$, we first need to derive the limit distribution of $\widehat{\omega}_S$, the data-based weights specified by the user. As an example, consider the case of moment selection based on the FMSC. Here $\widehat{\omega}_S$ is simply the indicator function
Substituting estimators of $\Omega$, $K_S$ and $\theta_0$ into $\mbox{FMSC}_S(\tau,M)$, defined in Corollary (ref), gives
where $\widehat{\mathcal{B}}(\tau,M) = (\widehat{\Psi} M + \tau)(\widehat{\Psi} M + \tau)' - \widehat{\Psi} \widehat{\Omega} \widehat{\Psi}$. Combining this with Equation (ref),
For GMM-AIC moment selection or selection based on a downward $J$-test, $\varphi_S(\cdot,\cdot)$ may be estimated analogously, following Theorem (ref). Continuing to assume for the moment that $\tau$ is known, consider the following algorithm:
Given knowledge of $\tau$, Algorithm (ref) yields valid inference for $\mu$. The problem, of course, is that $\tau$ is unknown and cannot even be consistently estimated. One idea would be to substitute the asymptotically unbiased estimator $\widehat{\tau}$ from (ref) in place of the unknown $\tau$. This gives rise to a procedure that I call the “1-Step” confidence interval:
The 1-Step interval defined in Algorithm (ref) is conceptually simple, easy to compute, and can perform well in practice, as I explore below. But as it fails to account for sampling uncertainty in $\widehat{\tau}$, it does not necessarily yield asymptotically valid inference for $\mu$. Fully valid inference requires the addition of a second step to the algorithm and comes at a cost: conservative rather than exact inference. In particular, the following procedure is guaranteed to yield an interval with asymptotic coverage probability of at least $(1-\alpha-\delta)\times 100\%$.
The preceding section presented two confidence interval that account for the effects of moment selection on subsequent inference. The 1-Step interval is intuitive and computationally straightforward but lacks theoretical guarantees, while the 2-Step interval guarantees asymptotically valid inference at the cost of greater computational complexity and conservatism. To better understand these methods and the trade-offs involved in deciding between them, I now specialize them to the two examples of FMSC selection that appear in the simulation studies described below. The structure of these examples allows us to bypass Algorithm (ref) and characterize the asymptotic properties of various proposals for post-FMSC without resorting to Monte Carlo simulations. Because this section presents asymptotic results, I treat any consistently estimable quantity that appears in a limit distribution as known.
In both the OLS versus IV example from Section (ref) and the slightly simplified version of the choosing instrument variables example implemented in Section (ref), the post-FMSC estimator $\widehat{\beta}_{FMSC}$ converges to a very convenient limit experiment.\footnote{The simplified version of the choosing instrumental variables example considers a single potentially endogenous instrument and imposes homoskedasticity. For more details see Section (ref) and Online Appendix (ref).} In particular,
with
where $Z_1, Z_2$ are independent standard normal random variables, $\eta$, $\sigma$ and $c$ are consistently estimable constants, and $\tau$ is the local mis-specification parameter. This representation allows us to tabulate the asymptotic distribution, $F_{FMSC}$ as follows:
where $\Phi$ is the CDF and $\varphi$ the pdf of a standard normal random variable. Note that the limit distribution of the post-FMSC distribution depends on $\tau$ in addition to the consistently estimable quantities $\sigma, \eta, c$ although I suppress this dependence to simplify the notation. While these expressions lack a closed form $G$, $H_1$ and $H_2$ are easy to compute, allowing us to calculate both $F_{FMSC}$ and the corresponding quantile function $Q_{FMSC}$\footnote{I provide code to evaluate both $F_{FMSC}$ and $Q_{FMSC}$ in my R package fmscr, available at \url{https://github.com/fditraglia/fmscr}.}
The ability to compute $F_{FMSC}$ and $Q_{FMSC}$ allows us to answer a number of important questions about post-FMSC inference. First, suppose that we were to carry out FMSC selection and then construct a $(1 - \alpha) \times 100\%$ confidence interval conditional in the selected estimator, completely ignoring the effects of the moment selection step. What would be the resulting asymptotic coverage probability and width of such a “na\"{i}ve” confidence interval procedure? Using calculations similar to those used above in the expression for $F_{FMSC}$, we find that the coverage probability of this na\"{i}ve interval is given by
where $G$, $H_1$, $H_2$ are as defined in Equations (ref)--(ref). And since the width of this na\"{i}ve CI equals that of the textbook interval for $\widehat{\beta}$ when $|\widehat{\tau}|<\sigma\sqrt{2}$ and that of the textbook interval for $\widetilde{\beta}$ otherwise, we have
where $\mbox{Width}_{Valid}(\alpha)$ is the width of a standard, textbook confidence interval for $\widetilde{\beta}$.
To evaluate these expressions we need values for $c, \eta^2, \sigma^2$ and $\tau$. For the remainder of this section I will consider the parameter values that correspond to the simulation experiments presented below in Section (ref). For the OLS versus TSLS example we have $c=1$, $\eta^2=1$ and $\sigma^2 = (1-\pi^2)/\pi^2$ where $\pi^2$ denotes the population first-stage R-squared for the TSLS estimator. For the choosing IVs example we have $c =\gamma/(\gamma^2 +1/9)$, $\eta^2 = 1/(\gamma^2 + 1/9)$ and $\sigma^2 = 1 + 9\gamma^2$ where $\gamma^2$ is the increase in the population first-stage R-squared of the TSLS estimator from adding $w$ to the instrument set.\footnote{The population first-stage R-squared with only $\mathbf{z}$ in the instument set is $1/9$.}
Table (ref) presents the asymptotic coverage probability and Table (ref) the expected relative width of the na\"{i}ve confidence interval procedure for a variety of values of $\tau$ and $\alpha$ for each of the two examples. For the OLS versus TSLS example, I allow $\pi^2$ to vary while for the choosing IVs example I allow $\gamma^2$ to vary. Note that the relative expected width does not depend on $\alpha$. In terms of coverage probability, the na\"{i}ve interval performs very poorly: in some regions of the parameter space the actual coverage is very close to the nominal level, while in others it is far lower. These striking size distortions, which echo the findings of Guggenberger2010 and Guggenberger2012, provide a strong argument against the use of the na\"{i}ve interval. Its attraction, of course, is width: the na\"{i}ve interval can be dramatically shorter than the corresponding “textbook” confidence interval for the valid estimator.
Is there any way to construct a post-FMSC confidence interval that does not suffer from the egregious size distortions of the na\"{i}ve interval but is still shorter than the textbook interval for the valid estimator? As a first step towards answering this question, Table (ref) presents the relative width of the shortest possible infeasible post-FMSC confidence interval constructed directly from $Q_{FMSC}$. This interval has asymptotic coverage probability exactly equal to its nominal level as it correctly accounts for the effect of moment selection on the asymptotic distribution of the estimators. Unfortunately it cannot be used in practice because it requires knowledge of $\tau$, for which no consistent estimator exists. As such, this interval serves as a benchmark against which to judge various feasible procedures that do not require knowledge of $\tau$. For certain parameter values this interval is shorter than the valid interval but the improvement is not uniform and indeed cannot be. Just as the FMSC itself cannot provide a uniform reduction in AMSE relative to the valid estimator, the infeasible post-FMSC cannot provide a corresponding reduction in width. In both cases, however, improvements are possible when $\tau$ is expected to be small, the setting in which this paper assumes that an applied researcher finds herself. The potential reductions in width can be particularly dramatic for larger values of $\alpha$. The question remains: is there any way to capture these gains using a feasible procedure?
Now, consider the 2-Step confidence interval procedure from Algorithm (ref). We can implement an equivalent procedure without simulation as follows. First we construct a $(1-\alpha_1)\times 100\%$ confidence interval for $\widehat{\tau}$ using $T = \sigma Z_1 + \tau$ where $Z_1$ is standard normal. Next we construct a $(1-\alpha_2)\times 100\%$ based on $Q_{FMSC}$ for each $\tau^*$ in this interval. Finally we take the upper and lower bounds over all of the resulting intervals. This interval is guaranteed to have asymptotic coverage probability of at least $1 - (\alpha_1 + \alpha_2)$ by an argument essentially identical to the proof of Theorem (ref). Protection against under-coverage, however, comes at the expense of extreme conservatism, particularly for larger values of $\alpha$. Numerical values for the coverage and median expected with of this interval appear in Online Appendix (ref). From both the numerical calculations and the theoretical result given in Theorem (ref) we see that the 2-Step systematically over-covers and hence cannot produce an interval shorter than the textbook CI for the valid estimator.
Now consider the 1-Step confidence interval from Algorithm (ref). Rather than first constructing a confidence region for $\tau$ and then taking upper and lower bounds, 1-Step interval simply takes $\widehat{\tau}$ in place of $\tau$ and then constructs a confidence interval from $Q_{FMSC}$ exactly as in the infeasible interval described above.\footnote{As in the construction of the na\"{i}ve interval, I take the shortest possible interval based on $Q_{FMSC}$ rather than an equal-tailed interval. Additional results for an equal-tailed version of this one-step procedure are available upon request. Their performance is similar.} Unlike its 2-Step counterpart, this interval comes with no generic theoretical guarantees, so I use the characterization from above to directly calculate its asymptotic coverage and expected relative width. The results appear in Tables (ref) and (ref). The 1-Step interval effectively “splits the difference” between the two-step interval and the na\"{i}ve procedure. While it can under-cover, the size distortions are quite small, particularly for $\alpha=0.1$ and $0.05$. At the same time, when $\tau$ is relatively small this procedure can yield shorter intervals. While a full investigation of this phenomenon is beyond the scope of the present paper, these calculations suggest a plausible way forward for post-FMSC inference that is less conservative than the two-step procedure from Algorithm (ref) by directly calculating the relevant quantities from the limit distribution of interest. This is possible because $\pi$ and $\gamma^2$ are both consistently estimable. And for any particular value of these parameters, the worst-case value of $\tau$ is interior. Using this idea, one could imagine specifying a maximum allowable size distortion and then designing a confidence interval to minimize width, possibly incorporating some prior restriction on the likely magnitude of $\tau$. Just as the FMSC aims to achieve a favorable trade-off between bias and variance, such a confidence interval procedure could aim to achieve a favorable trade-off between width and coverage. It would also be interesting to pursue analogous calculations for the minimum AMSE averaging estimator from Section (ref).
I begin by examining the performance of the FMSC and averaging estimator in the OLS versus TSLS example. All calculations in this section are based on the formulas from Sections (ref) and (ref) with 10,000 simulation replications. The data generating process is given by
with $(\epsilon_i, v_i, z_{1i}, z_{2i}, z_{3i}) \sim \mbox{ iid } N(0, \mathcal{S})$
for $i= 1, \hdots, N$ where $N$, $\rho$ and $\pi$ vary over a grid. The goal is to estimate the effect of $x$ on $y$, in this case 0.5, with minimum MSE either by choosing between OLS and TSLS estimators or by averaging them. To ensure that the finite-sample MSE of the TSLS estimator exists, this DGP includes three instruments leading to two overidentifying restrictions Phillips1980.\footnote{Alternatively, one could use fewer instruments in the DGP and work with trimmed MSE, as described in Online Appendix (ref).} This design satisfies regularity conditions that are sufficient for Theorem (ref) -- in particular it satisfies Assumption (ref) from Online Appendix (ref) -- and keeps the variance of $x$ fixed at one so that $\pi = Cor(x_i, z_{1i} + z_{2i} + z_{3i})$ and $\rho = Cor(x_i,\epsilon_i)$. The first-stage R-squared is simply $1 - \sigma_v^2/\sigma_x^2 = \pi^2$ so that larger values of $|\pi|$ reduce the variance of the TSLS estimator. Since $\rho$ controls the endogeneity of $x$, larger values of $|\rho|$ increase the bias of the OLS estimator.
Figure (ref) compares the root mean-squared error (RMSE) of the post-FMSC estimator to those of the OLS and TSLS estimators.\footnote{Note that, while the first two moments of the TSLS estimator exist in this simulation design, none of its higher moments do. This can be seen from the simulation results: even with 10,000 replications, the RMSE of the TSLS estimator shows a noticeable degree of simulation error.} For any values of $N$ and $\pi$ there is a value of $\rho$ below which OLS outperforms TSLS: as $N$ and $\pi$ increase this value approaches zero; as they decrease it approaches one. In practice, of course, $\rho$ in unknown so we cannot tell which of OLS and TSLS is to be preferred a priori. If we make it our policy to always use TSLS we will protect ourselves against bias at the potential cost of very high variance. If, on the other hand, we make it our policy to always use OLS then we protect ourselves against high variance at the potential cost of severe bias. FMSC represents a compromise between these two extremes that does not require advance knowledge of $\rho$. When the RMSE of TSLS is high, the FMSC behaves more like OLS; when the RMSE of OLS is high it behaves more like TSLS. Because the FMSC is itself a random variable, however, it sometimes makes moment selection mistakes.\footnote{For more discussion of this point, see Section (ref).} For this reason it does not attain an RMSE equal to the lower envelope of the OLS and TSLS estimators. The larger the RMSE difference between OLS and TSLS, however, the closer the FMSC comes to this lower envelope: costly mistakes are rare.
As shown above, the FMSC takes a very special form in this example: it is equivalent to a DHW test with $\alpha \approx 0.16$. Accordingly, Figure (ref) compares the RMSE of the post-FMSC estimator to those of DHW pre-test estimators with significance levels $\alpha = 0.05$ and $\alpha = 0.1$, indicated in the legend by DHW95 and DHW90. Since these three procedures differ only in their critical values, they show similar qualitative behavior. When $\rho$ is sufficiently close to zero, we saw from Figure (ref) that OLS has a lower RMSE than TSLS. Since DHW95 and DHW90 require a higher burden of proof to reject OLS in favor of TSLS, they outperform FMSC in this region of the parameter space. When $\rho$ crosses the threshold beyond which TSLS has a lower RMSE than OLS, the tables are turned: FMSC outperforms DHW95 and DHW90. As $\rho$ increases further, relative to sample size and $\pi$, the three procedures become indistinguishable in terms of RMSE. In addition to comparing the FMSC to DHW pre-test estimators, Figure (ref) also presents the finite-sample RMSE of the minimum-AMSE moment average estimator presented in Equations (ref) and (ref). The performance of the moment average estimator is very strong: it provides the lowest worst-case RMSE and improves uniformly on the FMSC for all but the largest values of $\rho$.
Because this example involves a scalar target parameter, no selection or averaging scheme can provide a uniform improvement over the minimax estimator, namely TSLS. But the cost of protection against the worst case is extremely poor performance when $\pi$ and $N$ are small. When this is the case, there is a strong argument for preferring the FMSC or minimum-AMSE estimator: we can reap the benefits of OLS when $\rho$ is small without risking the extremely large biases that could result if $\rho$ is in fact large.
Further simulation results for $\pi \in \left\{ 0.01, 0.05, 0.1 \right\}$ appear in Online Appendix (ref). For these parameter values the TSLS estimator suffers from a weak instrument problem leading the FMSC to substantially outperform the TSLS estimator. See Online Appendix (ref) for a more detailed discussion.
I now evaluate the performance of FMSC in the instrument selection example described in Section (ref) using the following simulation design:
for $i=1, 2, \hdots, N$ where $(\epsilon_i, v_i, w_i, z_{i1}, z_{2i}, z_{3i})' \sim \mbox{ iid } N(0,\mathcal{V})$ with
This setup keeps the variance of $x$ fixed at one and the endogeneity of $x$, $Cor(x, \epsilon)$, fixed at $0.5$ while allowing the relevance, $\gamma = Cor(x,w)$, and endogeneity, $\rho = Cor(w, \epsilon)$, of the instrument $w$ to vary. The instruments $z_1, z_2, z_3$ are valid and exogenous: they have first-stage coefficients of $1/3$ and are uncorrelated with the second stage error $\epsilon$. The additional instrument $w$ is only relevant if $\gamma \neq 0$ and is only exogenous if $\rho = 0$. Since $x$ has unit variance, the first-stage R-squared for this simulation design is simply $1 - \sigma_v^2 = 1/9 + \gamma^2$. Hence, when $\gamma = 0$, so that $w$ is irrelevant, the first-stage R-squared is just over 0.11. Increasing $\gamma$ increases the R-squared of the first-stage. This design satisfies the sufficient conditions for Theorem (ref) given in Assumption (ref) from Online Appendix (ref). When $\gamma = 0$, it is a special case of the DGP from Section (ref).
As in Section (ref), the goal of moment selection in this exercise is to estimate the effect of $x$ on $y$, as before 0.5, with minimum MSE. In this case, however, the choice is between two TSLS estimators rather than OLS and TSLS: the valid estimator uses only $z_1, z_2,$ and $z_3$ as instruments, while the full estimator uses $z_1, z_2, z_3,$ and $w$. The inclusion of $z_1, z_2$ and $z_3$ in both moment sets means that the order of over-identification is two for the valid estimator and three for the full estimator. Because the moments of the TSLS estimator only exist up to the order of over-identification Phillips1980, this ensures that the small-sample MSE is well-defined.\footnote{Alternatively, one could use fewer instruments for the valid estimator and compare the results using trimmed MSE. For details, see Online Appendix (ref).} All estimators in this section are calculated via TSLS without a constant term using the expressions from Section (ref) and 20,000 simulation replications.
Figure (ref) presents RMSE values for the valid estimator, the full estimator, and the post-FMSC estimator for various combinations of $\gamma$, $\rho$, and $N$. The results are broadly similar to those from the OLS versus TSLS example presented in Figure (ref). For any combination $(\gamma,N)$ there is a positive value of $\rho$ below which the full estimator yields a lower RMSE than the full estimator. As the sample size increases, this cutoff becomes smaller; as $\gamma$ increases, it becomes larger. As in the OLS versus TSLS example, the post-FMSC estimator represents a compromise between the two estimators over which the FMSC selects. Unlike in the previous example, however, when $N$ is sufficiently small there is a range of values for $\rho$ within which the FMSC yields a lower RMSE than both the valid and full estimators. This comes from the fact that the valid estimator is quite erratic for small sample sizes. Such behavior is unsurprising given that its first stage is not especially strong, $\mbox{R-squared}\approx 11\%$, and it has only two moments. In contrast, the full estimator has three moments and a stronger first stage. As in the OLS versus TSLS example, the post-FMSC estimator does not uniformly outperform the valid estimator for all parameter values, although it does for smaller sample sizes. The FMSC never performs much worse than the valid estimator, however, and often performs substantially better, particularly for small sample sizes.
I now compare the FMSC to the GMM moment selection criteria of Andrews1999, which take the form $MSC(S) = J_n(S) - h(|S|)\kappa_n$, where $J_n(S)$ is the $J$-test statistic under moment set $S$ and $-h(|S|)\kappa_n$ is a “bonus term” that rewards the inclusion of more moment conditions. For each member of this family we choose the moment set that minimizes $MSC(S)$. If we take $h(|S|) = (p + |S| - r)$, then $\kappa_n = \log{n}$ gives a GMM analogue of Schwarz's Bayesian Information Criterion (GMM-BIC) while $\kappa_n = 2.01 \log{\log{n}}$ gives an analogue of the Hannan-Quinn Information Criterion (GMM-HQ), and $\kappa_n = 2$ gives an analogue of Akaike's Information Criterion (GMM-AIC). Like the maximum likelihood model selection criteria upon which they are based, the GMM-BIC and GMM-HQ are consistent provided that Assumption (ref) holds, while the GMM-AIC, like the FMSC, is conservative. Figure (ref) gives the RMSE values for the post-FMSC estimator alongside those of the post-GMM-BIC, HQ and AIC estimators. I calculate the $J$-test statistic using a centered covariance matrix estimator, following the recommendation of Andrews1999. For small sample sizes, the GMM-BIC, AIC and HQ are quite erratic: indded for $N = 50$ the FMSC has a uniformly smaller RMSE. This problem comes from the fact that the $J$-test statistic can be very badly behaved in small samples.\footnote{For more details, see Online Appendix (ref).} As the sample size becomes larger, the classic tradeoff between consistent and conservative selection emerges. For the smallest values of $\rho$ the consistent criteria outperform the conservative criteria; for moderate values the situation is reversed. The consistent criteria, however, have the highest worst-case RMSE. For a discussion of a combined strategy based on the GMM information criteria of Andrews1999 and the canonical correlations information criteria of HallPeixe2003, see Online Appendix (ref). For a comparison with the downward $J$-test, see Online Appendix (ref). Online Appendix (ref) presents results for a modified simulation experiment in which the valid estimator suffers from a weak instrument problem. The FMSC performs very well in this case.