Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
91,980 characters · 10 sections · 38 citation commands
A Generalized Focused Information Criterion for GMM
An econometric model is a tool for answering a particular research question: different questions may suggest different models for the same data. And the fact that a model is wrong, as the old saying goes, does not prevent it from being useful. This paper proposes a novel selection criterion for GMM estimation that takes both of these points to heart: the generalized focused information criterion (GFIC). Rather than attempting to identify the correct specification, the GFIC chooses from a set of potentially mis-specified moment conditions and parameter restrictions to yield the smallest mean squared error (MSE) estimator of a user-specified scalar target parameter. We derive the GFIC under local mis-specification, using asymptotic mean squared error (AMSE) to approximate finite-sample MSE. In this framework mis-specification, while present for any fixed sample size, disappears in the limit so that asymptotic variance and squared bias remain comparable. GMM estimators remain consistent under local mis-specification but their limit distributions show an asymptotic bias. Adding an additional moment condition or imposing a parameter restriction generally reduces asymptotic variance but, if incorrectly specified, introduces a source of bias. The GFIC trades off these two effects in the first-order asymptotic expansion of an estimator to approximate its finite sample behavior.
The GFIC takes its motivation from a situation that is common in empirical practice. A researcher who hopes to estimate a parameter of interest $\mu$ must decide which assumptions to use. On the one hand is a set of relatively uncontroversial “baseline” assumptions. We suppose that the baseline assumptions are correct and identify $\mu$. But the very fact that the baseline assumptions do not raise eyebrows suggests that they may not be especially informative about $\mu$. On the other hand are one or more stronger controversial “suspect” assumptions. These stronger assumptions are expected to be much more informative about $\mu$. If we were certain that they were correct, we would definitely choose to impose them in estimation. Indeed, by continuity, even if they were nearly correct, imposing the suspect assumptions could yield a favorable bias-variance tradeoff. This is the essential idea behind the GFIC. When the baseline assumptions identify the model, the GFIC provides an asymptotically unbiased estimator of AMSE.
The GFIC is an extension of the focused moment selection criterion (FMSC) of DiTraglia2016. While the FMSC considers the problem of selecting moment conditions while holding the model specification fixed, the GFIC allows us to select over both aspects of our specification simultaneously. This extension is particularly valuable in panel data applications, where we may, for example, wish to carry out selection over the lag specification as well as the exogeneity assumptions used to estimate a dynamic panel model. We specialize the GFIC to such a dynamic panel example below, and provide simulation evidence of its performance. Online Appendix (ref) provides two additional examples: selecting between random and fixed effects estimators, and choosing between pooled OLS and mean-group estimators of an average effect in the presence of heterogeneity. In addition to extending the FMSC to a broader class of problems, we also extend the results of DiTraglia2016 on post-selection and moment-averaging estimators to the more general setting of the GFIC. We conclude with an empirical example modelling the price elasticity of cigarette demand.
As its name suggests, the GFIC is related to the focused information criterion (FIC) of ClaeskensHjort2003, a model selection procedure for maximum likelihood estimators that uses local mis-specification to approximate the MSE of a target parameter. Like the FIC and related proposals, e.g.\ Schorfheide2005, the GFIC uses local mis-specification to derive a risk-based selection criterion. Unlike them, however, the GFIC provides both moment and model selection for general GMM estimators. If the moment conditions used in estimation are the score of a maximum likelihood model and we consider model selection only, then the GFIC reduces to the FIC. Thus, the GFIC extends both the FIC and the FMSC of DiTraglia2016. Comparatively few papers propose criteria for simultaneous GMM model and moment selection under mis-specification.\footnote{See Smith1992 for an approach to GMM model selection based on non-nested hypothesis testing. For a detailed discussion of the literature on moment selection, see DiTraglia2016.} AndrewsLu propose a family of selection criteria by adding appropriate penalty and “bonus” terms to the J-test statistic, yielding analogues of AIC, BIC, and the Hannan-Quinn information criterion. HongPrestonShum extend this idea to generalized empirical likelihood (GEL). The principal goal of both papers is consistent selection: they state conditions under which the correct model and all correct moment conditions are chosen in the limit. As a refinement to this approach, LaiSmallLiu suggest a two-step procedure: first consistently eliminate incorrect models using an empirical log-likelihood ratio criterion, and then select from the remaining models using a bootstrap covariance matrix estimator. The point of the second step is to address a shortcoming in the standard limit theory. While first-order asymptotic efficiency requires that we use all available correctly specified moment conditions, this can lead to a deterioration in finite sample performance if some conditions are only weakly informative. HallPeixe2003 make a similar point about the dangers of including “redundant” moment conditions while Caner2009 proposes a lasso-type GMM estimator to consistently remove redundant parameters.
In contrast to these suggestions, the GFIC does not aim to identify the correct model and moment conditions: its goal is a low MSE estimate of a quantity of interest, even if this entails using a specification that is not exactly correct. As such, the GFIC is an “efficient” rather than a consistent selection criterion. There is an unavoidable trade-off between consistent selection and estimators with desirable risk properties Yang2005. Indeed, the worst-case risk of any consistent selection procedure is unbounded as sample size tends to infinity LeebPoetscher2008. In this sense, the fact that the GFIC is not consistent is a benefit rather than a liability. As we show in simulations below, its worst-case performance is much better than that of competing selection procedures.
Although not strictly a selection procedure, the combined moments (CM) estimator of JudgeMittelhammer takes a similar perspective to that of the GFIC, emphasizing that incoporating the information from an incorrect specification could lead to a favorable bias-variance tradeoff under the right circumstances. Unlike the GFIC, however, the CM estimator is not targeted to a particular research goal. A key point of the GFIC is that two researchers using the same dataset but interested in different target parameters may find it optimal, in a minimum MSE sense, to choose different model specifications. We explore this idea further in our dynamic panel example below.
The remainder of this paper is organized as follows. Section (ref) derives the asymptotic distribution of GMM estimators under locally mis-specified moment conditions and parameter restrictions. Section (ref) uses this information to calculate the AMSE of a user-specified target parameter and provides asymptotically unbiased estimators of the required bias parameters, yielding the GFIC. Section (ref) extends the results on averaging estimators and post-selection inference from DiTraglia2016 to the more general setting of this paper. Section (ref) specializes the GFIC to a dynamic panel example, and Section (ref) presents simulation results. Finally, Section (ref) presents our empirical example and Section (ref) concludes. Proofs and supplementary simulation results appear in the Appendix. Further examples and simulation results appear in Online Appendices (ref) and (ref).
Let $f(\cdot, \cdot)$ be a $(p+q)$-vector of moment functions of a random vector $Z$ and an $(r+s)$-dimensional parameter vector $\beta$. To represent moment selection, we partition the moment functions according to $f(\cdot,\cdot) = \left(g(\cdot, \cdot)', h(\cdot, \cdot)'\right)'$ where $g(\cdot, \cdot)$ and $h(\cdot, \cdot)$ are $p$- and $q$-vectors. The moment condition associated with $g(\cdot, \cdot)$ is assumed to be correct, while that associated with $h(\cdot,\cdot)$ is locally mis-specified. The moment selection problem is to choose which, if any, of the elements of $h$ to use in estimation. To represent model selection, we partition the full parameter vector according to $\beta = \left(\theta', \gamma'\right)'$, where $\theta$ is an $s$-vector and $\gamma$ an $r$-vector of parameters. The model selection problem is to decide which if any of the elements of $\gamma$ to estimate, and which to set equal to the corresponding elements of $\gamma_0$, an $r$-vector of known constants. The parameters contained in $\theta$ are those that we always estimate, the “protected” parameters. Any specification that does not estimate the full parameter vector $\beta$ is locally mis-specified. The precise form of the local mis-specification, over parameter restrictions and moment conditions, is as follows.
Assumption (ref) specifies a triangular array data generating process in which the the true parameter vector $\beta_n = \left(\theta_0', \gamma_n'\right)'$, changes with sample size but converges to $\beta_0 = \left(\theta_0', \gamma_0'\right)'$ as $n\rightarrow \infty$.\footnote{For simplicity, and because it is the case for all examples we consider below, we assume that the triangular array from Assumption (ref) is iid within each row. Note however, that the $Z_{ni}$ are not iid across rows: $\gamma_n$ and $\tau_n$ change with $n$. As such, the triangular array machinery is still required to describe our results.} Unless some elements of $\delta$ are zero, any estimator that restricts $\gamma$ is mis-specified for fixed $n$. In the limit, however, the restriction $\gamma = \gamma_0$ holds. Similarly, for any fixed sample size $n$, the expectation of $h$ evaluated at the true parameter value $\beta_n$ depends on the unknown constant vector $\tau$, but this source of mis-specification disappears in the limit. Thus, under Assumption (ref), only estimators that use moment conditions from $g$ to estimate the full parameter vector $\beta$ are correctly specified. In the limit, however, every estimator is correctly specified, regardless of which elements of $\gamma$ it restricts and which elements of $h$ it includes. The purpose of local mis-specification is to ensure that squared asymptotic bias is of the same order as asymptotic variance: Assumption (ref) is a device rather than literal description of real-world data. Note that, by Assumption (ref), the limiting random variable $Z_i$ satisfies the population moment condition $E[f\left(Z_i,\theta_0, \gamma_0\right)]=0$. Since the $Z_i$ are assumed to have a common marginal law, we will use the shorthand $Z$ for $Z_i$ throughout. Accordingly, define:
along with
Each of these expressions involves the limiting random variable $Z$ rather than $Z_{ni}$. Thus, the corresponding expectations are taken with respect to a distribution for which all moment conditions have expectation zero evaluated at $(\theta_0, \gamma_0)$.
Before defining the estimators under consideration, we require some further notation. Let $b$ be a model selection vector, an $r$-vector of ones and zeros indicating which elements of $\gamma$ we have chosen to estimate. When $b = 1_r$, where $1_m$ represents an $m$-vector of ones, we estimate both $\theta$ and the full vector $\gamma$. When $b = 0_r$, where $0_m$ denotes an $m$-vector of zeros, we estimate only $\theta$, setting $\gamma=\gamma_0$. More generally, we estimate $|b|$ components of $\gamma$ and set the others equal to the corresponding elements of $\gamma_0$. Let $\gamma^{(b)}$ be the $|b|$-dimensional subvector of $\gamma$ corresponding to those elements selected for estimation. Similarly, let $\gamma^{(-b)}_0$ denote the $(r-|b|)$-dimensional subvector containing the values to which we set those components of $\gamma$ that are not estimated. Analogously, let $c=\left(c_g', c_h'\right)'$ be a moment selection vector, a $(p+q)$-vector of ones and zeros indicating which of the moment conditions we have chosen to use in estimation. We denote by $|c|$ the total number of moment conditions used in estimation. Let $\mathcal{BC}$ denote the collection of all model and moment selection pairs $(b,c)$ under consideration. To express moment and model selection in matrix form, we define the selection matrices $\Xi_b$ and $\Xi_c$. Multiplying $\beta$ by the $(|b| + s)\times(r+s)$ model selection matrix $\Xi_b$ extracts the elements corresponding to $\theta$ and the subset of $\gamma$ indicated by the model selection vector $b$. Thus $\Xi_b \beta = \left(\theta', \gamma^{(b)'} \right)'$. Similarly, multiplying a vector by the $|c|\times(p+q)$ moment selection matrix $\Xi_c$ extracts the components corresponding to the moment conditions indicated by the moment selection vector $c$. To simplify the notation, we adopt the shorthand $F(b,c) = \Xi_c F \Xi_b'$ and $\Omega_c = \Xi_c \Omega \Xi_c'$ throughout.
To express the estimators themselves, define the sample analogue of the expectations in Assumption (ref) as follows,
and let $\widetilde{W}$ be a $(p+q)\times(p+q)$ positive semi-definite weighting matrix with blocks $\widetilde{W}_{gg}, \widetilde{W}_{gh}, \widetilde{W}_{hg}$ and $\widetilde{W}_{hh}$, partitioned conformably to the partition of $f(Z,\beta)$ by $g(Z,\beta)$ and $h(Z,\beta)$. Each model and moment selection pair $(b,c)\in \mathcal{BC}$ defines a $(|b|+s)$-dimensional estimator $\widehat{\beta}(b,c)=(\widehat{\theta}(b,c)', \widehat{\gamma}^{(b)}(b,c)')'$ of $\beta^{(b)}= \left(\theta', \gamma^{(b)'} \right)'$ according to
We now state a number of standard high-level regularity conditions that will be assumed throughout our derivations below.
A particularly important special case is the estimator using only the moment conditions in $g$ to estimate the full parameter vector $\beta = \left(\theta', \gamma'\right)'$. We call this the valid estimator and denote it by $\widehat{\beta}_v$. Because it is assumed to be correctly specified both for finite $n$ and in the limit, the valid estimator contains the information we use to identify $\tau$ and $\delta$, and thus carry out moment and model selection. We assume that the valid estimator is identified.
Because Assumption (ref) ensures that they are correctly specified in the limit, all candidate specifications $(b,c)\in \mathcal{BC}$ provide consistent estimators of $\theta_0$ under standard, high level regularity conditions (see Assumption (ref)). Essential differences arise, however, when we consider their respective asymptotic distributions. Under Assumption (ref), both $\delta$ and $\tau$ induce a bias term in the limiting distribution of $\sqrt{n}\left(\widehat{\beta}(b,c) - \beta_0^{(b)}\right)$.
Because it employs the correct specification, the valid estimator of $\theta$ shows no asymptotic bias. Moreover, the valid estimator of $\gamma$ has an asymptotic distribution that is centered around $\delta$, suggesting an estimator of this bias parameter.
The GFIC chooses among potentially incorrect moment conditions and parameter restrictions to minimize estimator AMSE for a scalar target parameter. Denote this target parameter by $\mu = \varphi(\theta, \gamma)$, where $\varphi$ is a real-valued, almost surely continuous function of the underlying model parameters $\theta$ and $\gamma$. Let $\mu_n = \varphi(\theta_0,\gamma_n)$ and define $\mu_0$ and $\widehat{\mu}(b,c)$ analogously. By Theorem (ref) and the delta method, we have the following result.
The true value of $\mu$, however, is $\mu_n$ rather than $\mu_0$ under Assumption (ref). Accordingly, to calculate AMSE we recenter the limit distribution as follows.
We see that the limiting distribution of $\widehat{\mu}(b,c)$ is not, in general, centered around zero: both $\tau$ and $\delta$ induce an asymptotic bias. Note that, while $\tau$ enters the limit distribution only once, $\delta$ has two distinct effects. First, like $\tau$, it shifts the limit distribution of $\sqrt{n}f_n(\theta_0, \gamma_0)$ away from zero, thereby influencing the asymptotic behavior of $\sqrt{n}\left(\widehat{\mu}(b,c) - \mu_0 \right)$. Second, unless the derivative of $\varphi$ with respect to $\gamma$ is zero at $(\theta_0,\gamma_0)$, $\delta$ induces a second source of bias when $\widehat{\mu}(b,c)$ is recentered around $\mu_n$. Crucially, this second source of bias exactly cancels the asymptotic bias present in the limit distribution of $\widehat{\gamma}_v$. Thus, the valid estimator of $\mu$ is asymptotically unbiased and its AMSE equals its asymptotic variance.
Using Corollary (ref), the AMSE of $\widehat{\mu}(b,c)$ is as follows,
where
The idea behind the GFIC is to construct an estimate $\widehat{\mbox{AMSE}}\left(\widehat{\mu}(b,c)\right)$ and choose the specification $(b^*,c^*)\in\mathcal{BC}$ that makes this quantity as small as possible. As a side-effect of the consistency of the estimators $\widehat{\beta}(b,c)$, the usual sample analogues provide consistent estimators of $K(b,c)$ and $F_{\gamma}' = (G_\gamma', H_\gamma')$ under Assumption (ref), and $\varphi(\widehat{\theta}_v,\gamma_0)$ is consistent for $\varphi_0$. Consistent estimators of $\Omega$ are also readily available under local mis-specification although the best choice may depend on the situation.\footnote{We discuss this in more detail for our dynamic panel example in Section (ref) below.} Since $\gamma_0$ is known, as are $\Xi_b$ and $\Xi_c$, only $\delta$ and $\tau$ remain to be estimated. Unfortunately, neither of these quantities is consistently estimable under local mis-specification. Intuitively, the data become less and less informative about $\tau$ and $\delta$ as the sample size increases since each term is divided by $\sqrt{n}$. Multiplying through by $\sqrt{n}$ counteracts this effect, but also stabilizes the variance of our estimators. Hence, the best we can do is to construct asymptotically unbiased estimators of $\tau$ and $\delta$. Corollary (ref) provides the required estimator for $\delta$, namely $\widehat{\delta} = \sqrt{n}\left(\widehat{\gamma}_v - \gamma_0\right)$, while Lemma (ref) provides an asymptotically unbiased estimator of $\tau$ by plugging $\widehat{\beta}_v$ into the sample analogue of the $h$-block of moment conditions.
Combining Corollary (ref) and Lemma (ref), gives the joint distribution of $\widehat{\delta}$ and $\widehat{\tau}$.
Now, we see immediately from Equation (ref) that $$BIAS\left(\widehat{\mu}\left(b,c\right)\right)^2 = \nabla_\beta \varphi_0' M(b,c) \left[
\right] M(b,c)' \nabla_\beta \varphi_0$$ Thus, the bias parameters $\tau$ and $\delta$ enter the AMSE expression in Equation \ref{eq:AMSE} as outer products: $\tau\tau'$, $\delta\delta'$ and $\tau\delta'$. Although $\widehat{\tau}$ and $\widehat{\delta}$ are asymptotically unbiased estimators of $\tau$ and $\delta$, it does \emph{not} follow that $\widehat{\tau}\widehat{\tau}'$, $\widehat{\delta}\widehat{\delta}'$ and $\widehat{\tau}\widehat{\delta}'$ are asymptotically unbiased estimators of $\tau\tau'$, $\delta\delta'$, and $\tau\delta'$. The following result shows how to adjust these quantities to provide the required asymptotically unbiased estimates.
Combining Corollary (ref) with consistent estimates of the remaining quantities yields the GFIC, an asymptotically unbiased estimator of the AMSE of our estimator of a target parameter $\mu$ under each specification $(b,c)\in \mathcal{BC}$
We choose the specification $(b^*,c^*)$ that minimizes the GFIC over the candidate set $\mathcal{BC}$.
While we are primarily concerned in this paper with the mean-squared error performance of our proposed selection techniques, it is important to have tools for carrying out inference post-selection. In this section we briefly present results that can be used to carry out valid inference for a range of model averaging and post-selection estimators, including the GFIC.\footnote{We direct the reader to DiTraglia2016 and the references contained therein for a more detailed discussion of inference post-selection.}
The GFIC is an efficient rather than consistent selection criterion: it aims to estimate a particular target parameter with minimum AMSE rather than selecting the correct specification with probability approaching one in the limit. As pointed out by Yang2005, among others, there is an unavoidable trade-off between consistent selection and desirable risk properties. Faced with this dilemma, the GFIC sacrifices consistency in the interest of low AMSE. Because it is not a consistent criterion, the GFIC remains random even in the limit. We can see this from Equation (ref) in Section (ref) and Corollary (ref). While the quantities $\nabla_\beta \widehat{\varphi}_0$, $\widehat{K}(b,c)$, and $\widehat{\Omega}_c$ are consistent estimators of their population counterparts, $\widehat{B}$ is only an asymptotically unbiased estimator of $B$ and thus has a limiting distribution. In particular $\widehat{B} \rightarrow_d \mathscr{B}(\mathscr{N}, \delta, \tau)$ where
Accordingly, to carry out inference post-GFIC, we need a limiting theory that is rich enough to accommodate randomly-weighted averages of the candidate estimators $\widehat{\mu}(b,c)$. To this end, consider an estimator of the form $\widehat{\mu} = \sum_{(b,c) \in \mathcal{BC}} \widehat{\omega}(b,c) \widehat{\mu}(b,c)$ where $\widehat{\mu}(b,c)$ denotes the target parameter under the moment conditions and parameter restrictions indexed by $(b,c)$, $\mathcal{BC}$ denotes the full set of candidate specifications, and $\widehat{\omega}(b,c)$ denotes a collection of data-dependent weights. These could be zero-one weights correponding to a moment or model selction criterion, e.g.\ select the estimator that minimizes GFIC, or model averaging weights.\footnote{For an example that averages over fixed and random effects estimators, see Section (ref).} We impose the following mild restrictions on the weights $\widehat{\omega}$.
Under the preceding conditions, we can derive the limit distribution of $\widehat{\mu}$ shown in the following Corollary.
Note that the limit distribution from the preceding corollary is highly non-normal: it is a randomly weighted average of a normal random vector, $\mathscr{N}$. To tabulate this distribution for the purposes of inference, we will in general need to resort to simulation. If $\tau$ and $\delta$ were known, the story would end here. In this case we could simply substitute consistent estimators of $K$ and $M$ and then repeatedly draw $\mathscr{N} \sim N(0, \widehat{\Omega})$, where $\widehat{\Omega}$ is a consistent estimator of $\Omega$, to tabulate the distribution of $\Lambda$ to arbitrary precision as follows.
Unfortunately, no consistent estimators of $\tau$ or $\delta$ exist: all we have at our disposal are asymptotically unbiased estimators. The following “1-step” confidence interval is constructed by substituting these into Algorithm (ref).
The 1-Step interval defined in Algorithm (ref) is conceptually simple, easy to compute, and can perform well in practice.\footnote{For more discussion on this point, see DiTraglia2016.} But as it fails to account for sampling uncertainty in $\widehat{\tau}$, $\mbox{CI}_1$ does not necessarily yield asymptotically valid inference for $\mu$. Fully valid inference requires the addition of a second step to the algorithm and comes at a cost: conservative rather than exact inference. In particular, the two-step procedure described in the following algorithm is guaranteed to yield an interval with asymptotic coverage probability of no less than $(1- \alpha_1 - \alpha_2)\times 100\%$.
We now specialize the GFIC to a dynamic panel model of the form
where $i = 1, \hdots, n$ indexes individuals and $t=1, \hdots, T$ indexes time periods. For simplicity, and without loss of generality, we suppose that there are no exogenous time-varying regressors and that all random variables are mean zero.\footnote{Alternatively, we can simply de-mean and project out any time-varying exogenous covariates after taking first-differences.} The unobserved error $\eta_i$ is a correlated individual effect: $\sigma_{x\eta}\equiv \mathbb{E}\left[ x_{it}\eta_i \right]$ may not equal zero. The endogenous regressor $x_{it}$ is assumed to be predetermined but not necessarily strictly exogenous: $\mathbb{E}[x_{it} v_{is}]=0$ for all $s \geq t$ but may be nonzero for $s < t$. We assume throughout that $y_{it}$ is stationary, which requires both $x_{it}$ and $u_{it}$ to be stationary and $|\boldsymbol{\gamma}| < 1$ where $\boldsymbol{\gamma} = (\gamma_1, \dots, \gamma_k)'$. Our goal is to estimate one of the following two target parameters with minimum MSE:
where $\mu_{\text{SR}}$ denotes the short-run effect and $\mu_{\text{LR}}$ the long-run effect of $x$ on $y$.
The question is which assumptions to use in estimation. Naturally, the answer may depend on whether our target is $\mu_{SR}$ or $\mu_{LR}$. Our first decision is what assumption to impose on the relationship between $x_{it}$ and $v_{it}$. This is the moment selection decision. We assumed above that $x$ is predetermined. Imposing the stronger assumption of strict exogeneity gives us more and stronger moment conditions, but using these in estimation introduces a bias if $x$ is not in fact strictly exogenous. Our second decision is how many lags of $y$ to use in estimation. This is the model selection decision. The true model contains $k$ lags of $y$. If we estimate only $r < k$ lags we not only have more degrees of freedom but more observations: every additional lag of $y$ requires us to drop one time period from estimation. In the short panel datasets common in microeconomic applications, losing even one additional time period can represent a substantial loss of information. At the same time, unless $\gamma_{r+1} = \cdots = \gamma_k = 0$, failing to include all $k$ lags in the model introduces a bias.
To eliminate the individual effects $\eta_i$ we work in first differences. Defining $\Delta$ in the usual way, so that $\Delta y_{it} = y_{it} - y_{it-1}$ and so on, we can write Equation (ref) as
For simplicity and to avoid many instruments problems -- see e.g.\ Roodman -- we focus here on estimation using the instrument sets
similar to AndersonHsiao. Modulo a change in notation, one could just as easily proceed using the instrument sets suggested by ArellanoBond. We use $\ell$ as a placeholder for the lag length used in estimation. If $\ell = 0$, $\mathbf{z}'_{it}(0,\text{P}) = x_{it-1}$ and $\mathbf{z}'_{it}(0,\text{S}) = (x_{it-1}, x_{it})$. Given these instrument sets, we have $(\ell + 1)\times (T -\ell - 1)$ moment conditions if $x$ is assumed to be predetermined versus $(\ell + 2)\times (T - \ell - 1)$ if it is assumed to be strictly exogenous, corresponding to the instrument matrices $Z_i(\ell,\text{P}) = \mbox{diag}\left\{ \mathbf{z}'_{it}(\ell,\text{P}) \right\}_{t = \ell + 2}^T$ and $Z_i(\ell,\text{S}) = \mbox{diag}\left\{\mathbf{z}'_{it}(\ell,\text{S}) \right\}_{t = \ell +2}^T$. To abstract for a moment from the model selection decision, suppose that we estimate a model with the true lag length: $\ell = k$. The only difference between the P and S sets of moment conditions is that the latter adds over-identifying information in the form of $E[x_{it}\Delta v_{it}]$. If $x$ is strictly exogenous, this expectation equals zero, but if $x$ is only predetermined, then $E[x_{it}\Delta v_{it}] = -E[x_{it}v_{it-1}] \neq 0$ so the over-identifying moment condition is invalid. Given our instrument sets, this is the only violation of strict exogeneity that is relevant for our moment selection decision so we take $E[x_{it}v_{it-1}] = -\tau/\sqrt{n}$.
In the examples and simulations described below we consider two-stage least squares (TSLS) estimation of $\mu_{SR}$ and $\mu_{LR}$ using the instruments defined in Equation (ref). Without loss of generality, we select between two lag length specifications: the first is correct, $\ell = k$, and the second includes $m$ lags too few: $\ell = r$ where $r = k-m$. Accordingly, we make the coefficients associated with the $(r+1)$\textsuperscript{th}, $\ldots, k$\textsuperscript{th} lags local to zero. Let $\boldsymbol{\gamma}' = (\gamma_1, \cdots, \gamma_{k-1}, \gamma_{k})$ denote the full vector of lag coefficients and $\boldsymbol{\gamma}_{r}' = (\gamma_1, \cdots, \gamma_{r})$ denote the first $r = k-m$ lag coefficients. Then, the true parameter vector is $\beta_n = (\theta, \boldsymbol{\gamma}'_{r}, \boldsymbol{\delta}'/\sqrt{n})'$ which becomes, in the limit, $\beta = (\theta, \boldsymbol{\gamma}'_r, \boldsymbol{0}')'$. Both $\boldsymbol{\delta}$ and $\boldsymbol{0}$ are of length $m$. To indicate the subvector of $\beta$ that excludes the $(r+1)$\textsuperscript{th}, $\ldots, k$\textsuperscript{th} lag coefficients, let $\beta_{r} = (\theta, \boldsymbol{\gamma}_r')'$.
Because the two lag specifications we consider use different time periods in estimation, we require some additional notation to make this clear. First let $\Delta \mathbf{y}_{i} = [\Delta y_{i,k+2}, \cdots, \Delta y_{iT}]'$ and $\Delta \mathbf{y}^+_{i} = [\Delta y_{i,k+2-m}, \Delta y_{i,k+2-(m-1)}, \cdots, \Delta y_{iT}]'$ where the superscript “+” indicates the inclusion of $m$ additional time periods: $t = k+2-m, \ldots, k+1$. Define $\Delta \mathbf{x}_i$, $\Delta \mathbf{x}_{i}^{+}$, $\Delta \mathbf{v}_i$, and $\Delta \mathbf{v}_{i}^{+}$ analogously. Next, define $L^{r+1}\Delta \mathbf{y}_i^{+} = [\Delta y_{i1}, \Delta y_{i2}, \cdots, \Delta y_{iT-(r+1)}]'$ where $L^{r+1}$ denotes the element-wise application of the $(r+1)$\textsuperscript{th} order lag operator. Note that the first element of $L^{r+1}\Delta \mathbf{y}_{i}^{+}$ is unobserved since $\Delta y_{i1} = y_{i1} - y_{i0}$ but $t=1$ is the first time period. Now we define the matrices of regressors for the two specifications:
Note that $W_i^{+}(r)$ contains $m$ more rows than $W_i(k)$ but $W_i(k)$ contains $m$ more columns than $W_i^{+}(r)$: removing the $(r+1)$\textsuperscript{th}, $\ldots, k$\textsuperscript{th} lags from the model by setting $\ell = r = k-m$ allows us to use $m$ additional time periods in estimation and reduces the number of regressors by $m$. Stacking over individuals, let $\Delta \mathbf{y} = [\Delta \mathbf{y}'_1 \cdots \Delta \mathbf{y}'_n]'$, $W_\ell = [W_1(\ell) \cdots W_n(\ell)]'$ and define $\Delta \mathbf{y}^{+}$ and $W_\ell^{+}$ analogously, where $\ell$ denotes the lag length used in estimation. Finally, let $Z'(\ell,\cdot) = [Z'_1(\ell,\cdot) \cdots Z'_n(\ell,\cdot)]$ where $(\cdot)$ is $\text{P}$ or $\text{S}$ depending on the instrument set in use. Using this notation, under local mis-specification the true model is
Using the shorthand $\widehat{Q} \equiv n[W' Z(Z'Z)^{-1} Z'W]^{-1}W'Z(Z'Z)^{-1}$ our candidate estimators are
where $(\cdot)$ is either $\text{P}$ or $\text{S}$ depending on which instrument set is used and $r = k-m$, $m$ lags fewer than the true lag length $k$. The following result describes the limit distribution of $\widehat{\beta}(k,\text{P})$, $\widehat{\beta}(k,\text{S})$, $\widehat{\beta}(r,\text{P})$, and $\widehat{\beta}(r,\text{S})$ which we will use to construct the GFIC.
To operationalize the GFIC, we need to provide appropriate estimators of all quantities that appear in Theorem (ref). To estimate ${Q}(k,\text{P})$, ${Q}(k,\text{S})$, ${Q}(r,\text{P})$, and ${Q}(r,\text{S})$ we employ the usual sample analogues $\widehat{Q}(\cdot,\cdot)$ given above, which remain consistent under local mis-specification. There are many consistent estimators for the variance matrices $\mathcal{V}(k,\text{P})$, $\mathcal{V}(k,\text{S})$, $\mathcal{V}(r,\text{P})$, $\mathcal{V}(r,\text{S})$ under local mis-specification. In our simulations below, we employ the usual heteroskedasticity-consistent, panel-robust variance matrix estimator. Because $E[\mathbf{z}_{it}(\ell,\text{S})\Delta v_{it}]\neq 0$, we center our estimators of $\mathcal{V}(\ell, \text{S})$ by subtracting the sample analogue of this expectation when calculating the sample variance. We estimate $\boldsymbol{\psi}_{\text{P}}$ and $\boldsymbol{\psi}_{\text{S}}$ as follows \[ \widehat{\boldsymbol{\psi}}_{P}' = \frac{1}{nT_k}
, \quad \widehat{\boldsymbol{\psi}}_{S}' = \frac{1}{nT_k}
\] where $T_k = T-k-1$. These estimates use our assumption of stationarity from above. The only remaining quantities we need to construct the GFIC involve the bias parameters $\boldsymbol{\delta}$ and $\tau$. We can read off an asymptotically unbiased estimator of $\boldsymbol{\delta}$ directly from Theorem (ref), namely $\widehat{\boldsymbol{\delta}} = \sqrt{n}\; (\widehat{\gamma}_{r+1}(k,\text{P}), \ldots, \widehat{\gamma}_k(k,\text{P}))'$ based on the instrument set that assumes only that $x$ is pre-determined rather than strictly exogenous. To construct an asymptotically unbiased estimator of $\tau$, we use the residuals from the specification that uses both the correct moment conditions and the correct lag specification, specifically
where $X' = [X_1 \cdots X_n]$ and $X_i = \mbox{diag}\left\{ x_{it} \right\}_{t = k + 2}^{T}$. The following result gives the joint limiting behavior of $\widehat{\boldsymbol{\delta}}$ and $\widehat{\tau}$, which we will use to construct the GFIC.
To provide asymptotically unbiased estimators of the quantities $\tau^2$, $\boldsymbol{\delta}\boldsymbol{\delta}'$ and $\boldsymbol{\delta}\tau$ that appear in the AMSE expressions for our estimators, we apply a bias correction to the asymptotically unbiased estimators of $\boldsymbol{\delta}$ and $\tau$ from Theorem (ref).
We have already discussed consistent estimation of $\widehat{\mathcal{V}}(k,\text{S})$. Since $\Pi$ is a known permutation matrix, it remains only to propose a consistent estimator of $\Psi$. The matrix $\Psi$, in turn, depends only on $Q(k,\text{P})$, and $\boldsymbol{\xi}'$. The sample analogue $\widehat{Q}(k,\text{P})$ is a consistent estimator for $Q(k,\text{P})$, as mentioned above, and
is consistent for $\xi'$. We now have all the quantities needed to construct the GFIC for $\mu_{SR}$, the short-run effect of $x$ on $y$. Since $\mu_{SR} = \theta$, we can read off the AMSE expression for this parameter directly from Theorem (ref). For the long-run effect $\mu_{LR}$, however, we need to formally apply the Delta-method and account for the fact that the true value of $(\gamma_{r+1}, \ldots, \gamma_k)'$ is $\boldsymbol{\delta}/\sqrt{n}$. Expressed as a function $\varphi$ of the underlying model parameters, \[ \mu_{LR} = \varphi(\theta, \boldsymbol{\gamma}_r, \boldsymbol{\gamma}_{-r}) = \theta / \left[1 - \boldsymbol{\iota}_r' \boldsymbol{\gamma}_r - \boldsymbol{\iota}_m'\boldsymbol{\gamma}_{-r}\right] \] where we define $\boldsymbol{\gamma}_{-r} \equiv (\gamma_{r+1}, \ldots, \gamma_k)'$. The derivatives of $\varphi$ are \[ \nabla \varphi \equiv \left[
\right] = \left( \frac{1}{1 - \boldsymbol{\iota}_r' \boldsymbol{\gamma}_r - \boldsymbol{\iota}_m' \boldsymbol{\gamma}_{-r}} \right)^2 \left[
\right]. \] Using this notation, the limiting value of $\mu_{LR}$ is $\mu_{LR}^{0} = \varphi(\theta, \boldsymbol{\gamma}_r, \boldsymbol{0}_m')$ while the true value is $\mu_{LR}^{n} = \varphi(\theta, \boldsymbol{\gamma}_r, \boldsymbol{\delta}/\sqrt{n})$. Similarly, the limiting value of $\nabla\varphi$ is $\nabla \varphi_0 = \nabla \varphi(\theta, \boldsymbol{\gamma}_r, \boldsymbol{0}_m')'$, obtained by putting zero in place of $\boldsymbol{\gamma}_{-r}$. We estimate this quantity consistently by plugging in the estimates from $\widehat{\beta}(k, \text{P})$.
We now consider two simulation experiments based on section (ref), applying the GFIC to a dynamic panel model. For both experiments our data generating process is similar to that of AndrewsLu, specifically
where $\mathbf{0}_m$ denotes an $m$-vector of zeros, $I_m$ the $(m\times m)$ identity matrix, and $\iota_m$ an $m$-vector of ones. Under this covariance matrix structure $\eta_i$ and $v_{i}$ are uncorrelated with each other, but both are correlated with $x_{i}$: $E[x_{it}\eta_i]=\sigma_{x\eta}$ and $x_{it}$ is predetermined but not strictly exogenous with respect to $v_{it}$. Specifically, $E[x_{it}v_{it-1}]=\sigma_{xv}$, while $E[x_{it}v_{is}]=0$ for $s\neq t-1$. We initialize the pre-sample observations of $y$ to zero, the mean of their stationary distribution, and generate the remaining time periods according to Equation (ref) with $\theta = 0.5$ and $\sigma_{x\eta} = 0.2$. The true lag length differs in our two examples as does the target parameter, so we explain these features of the simulation designs below. Unlike AndrewsLu we do not generate extra observations to keep the time dimension fixed across estimators with different lag specifications. This is for two reasons. First, in real-world applications such additional observations would not be available. Second, we are explicitly interested in trading off the efficiency gain from including additional time periods in estimation against the bias that arises from estimating an incorrect lag specification.
Consider two different researchers who happen to be working with the same panel dataset. One wishes to estimate the short-run effect of $x$ on $y$ while the other wishes to estimate the long-run effect. Should they use the same model specification? We now present an example showing that the answer, in general, is no. Suppose that the true model is \[ y_{it} = \theta x_{it} + \gamma_1 y_{it-1} + \gamma_2 y_{it-2} + \eta_i + v_{it} \] where $i = 1, \dots, n = 250$ and $t = 1, \dots, T=5$ and the regressor, individual effect and error term are generated according to Equation (ref), as described in the preceding section. Our model selection decision in this example is whether to set $\gamma_2 = 0$ and estimate a specification with one lag only. We denote this one-lag specification by $\mbox{L1}$ and the true specification, including both lags, by $\mbox{L2}$. To focus on the model selection decision, we fix the instrument set in this experiment to $\mathbf{z}_{it}(\ell,\text{P})$, defined in Equation (ref). Because this instrument set is valid when $x$ is pre-determined, it does not introduce bias into our estimation. Thus, bias only emerges if we estimate $\mbox{L1}$ when $\gamma_2\neq 0$. Our simulation design takes $\theta = 0.5, \gamma_1 = 0.4, \sigma_{x\eta} = 0.2$, and $\sigma_{xv} = 0.1$ and varies $\gamma_2$ over the range $\{0.10, 0.11, \dots, 0.19, 0.20\}$.
Table (ref) presents the results of the simulation, based on 1000 replications at each grid point. Because they are based on ratios of estimators of $\theta$ and $\gamma_1, \gamma_2$, estimators of the long-run effect may not have finite moments, making finite-sample MSE undefined. The usual solution to this problem in simulation settings is to work with so-called “trimmed” MSE by discarding observations that fall outside, say, a range $[-M, M]$ before calculating MSE.\footnote{Note that asymptotic MSE remains well-defined even for estimators that do not possess finite-sample moments so that GFIC comparisons remain meaningful. By taking the trimming constant $M$ to infinity, one can formalize the notion that asymptotic MSE comparisons can be used to “stand in” for finite-sample MSE even when the latter does not exist. For more details, See HansenShrink and online appendix C of DiTraglia2016.} Because there is no clear way to set the trimming constant $M$, it can be difficult to interpret results based on trimmed MSE unless one considers a variety of values of $M$. To avoid this issue, Table (ref) reports simulation results for median absolute deviation (MAD). Results for trimmed MSE with different choices of $M$ are similar and are available upon request.
The columns of Table (ref) labeled $\mbox{L1}$ and $\mbox{L2}$ give the MAD of estimators that fix the lag length to one and two, while those labeled GFIC give the MAD of an estimator that selects lag length via the GFIC. Notice that throughout the table $\gamma_2 \neq 0$ so that $\mbox{L1}$ is mis-specified. Nonetheless, $\mbox{L1}$ yields lower MAD estimators of both the short-run and long-run effects when $\gamma_2$ is sufficiently small and the difference can be substantial. When $\gamma_2 = 0.2$, for example, MAD for is 0.582 for the long-run effect estimator based on $\mbox{L1}$ versus 0.801 for that based on $\mbox{L2}$. Note moreover that the point at which $\gamma_2$ becomes large enough for $\mbox{L2}$ to be preferred depends on which effect we seek to estimate. When $\gamma_2$ equals 0.15 or 0.16, $\mbox{L1}$ gives a lower MAD for the short-run effect while $\mbox{L2}$ gives a lower MAD for the long-run effect. Because it is subject to random model selection errors, the GFIC can never outperform the oracle estimator that uses $\mbox{L1}$ when it is optimal in terms of MAD and $\mbox{L2}$ otherwise. Instead, the GFIC represents a compromise between two extremes: its MAD is never as large as that of the worst specification and never as small as that of the best specification. When there are large MAD differences between $\mbox{L1}$ and $\mbox{L2}$, however, GFIC is generally close to the optimum.
We now consider a more complicated simulation experiment that simultaneously selects over lag specification and endogeneity assumptions. In this simulation our target parameter is the short-run effect of $x$ on $y$, as in our empirical example below and the true model contains one lag. Specifically, \[ y_{it} = \theta x_{it} + \gamma y_{it-1} + \eta_i + v_{it} \] where $i = 1, \dots, n$ and $t = 1, \dots, T$ and the regressor, individual effect and error term are generated according to Equation (ref). Our model selection decision example is whether to set $\gamma = 0$ and estimate a specification wihout the lagged dependent variable, while our moment selection decision is whether to use only the instrument set $\mathbf{z}_{it}(\ell,\text{P})$, which assumes that $x$ is predetermined, or the instrument set $\mathbf{z}_{it}(\ell,\text{S})$ which assumes that it is strictly exogenous. Both instrument sets are defined in Equation (ref). We consider four specifications, each estimated by TSLS using the expressions from section (ref). The correct specification, $\text{LP}$, estimates both $\gamma$ and $\theta$ using only the “predetermined” instrument set. In contrast, $\text{LS}$ estimates both parameters using the “strict exogeneity” instrument set. The specifications $\text{P}$ and $\text{S}$ set $\gamma=0$ and estimate only $\theta$, using the predetermined and strictly exogenous instrument sets, respectively. Our simulation design sets $\theta = 0.5$, $\sigma_{x\eta}=0.2$ and varies $\gamma$, $\sigma_{xv}$, $T$ and $n$ over a grid. Specifically, we take $\gamma, \sigma_{xv} \in \{0, 0.005, 0.01, \hdots, 0.195, 0.2\}$, $n \in \{250,500\}$, $T \in \{4,5\}$.\footnote{Setting $T$ no smaller than 4 ensures that MSE exists for all four estimators: the finite sample moments of the TSLS estimator only exist up to the order of over-identification.} All values are computed based on 2000 simulation replications.
Table (ref), presents RMSE values multiplied by 1000 for ease of reading for each of the fixed specifications -- LP, LS, P, and S -- and for the various selection procedures.\footnote{In the interest of space, Table (ref) uses a coarser simulation grid than Figure (ref). The supplementary figures in Online Appendix (ref) present results over the full simulation grid.} We see that there are potentially large gains to be had by intentionally using a mis-specified estimator. Indeed, the correct specification, $\text{LP}$, is only optimal when both $\rho_{xv}$ and $\gamma$ are fairly large relative to sample size. When $T=4$ and $n=250$, for example, $\gamma$ and $\rho_{xv}$ must both exceed 0.10 before $\text{LP}$ has the lowest RMSE. Moreover, the advantage of the mis-specified estimators can be substantial.
In practice, of course, we do not know the values of $\gamma$, $\rho$, $\theta$, or the other parameters of the DGP so this comparison of finite-sample RMSE values is infeasible. Instead, we consider using the GFIC to select between the four specifications. Clearly there are gains to be had from estimating a mis-specified model in certain situations. The questions remains, can the GFIC identify them? Because it is an efficient rather than consistent selection criterion, the GFIC remains random, even in the limit. This means that the GFIC can never outperform the “oracle” estimator that uses whichever specification gives the lowest finite-sample RMSE. Moreover, because our target parameter is a scalar, Stein-type results do not apply: the post-GFIC estimator cannot provide a uniform improvement over the true specification $\text{LP}$. Nevertheless, the post-GFIC estimator can provide a substantial improvement over $\text{LP}$ when $\rho_{xv}$ and $\gamma$ are relatively small relative to sample size, as shown in the two leftmost panes of the top panel in Table (ref). This is precisely the situation for which the GFIC is intended: a setting in which we have reason to suspect that mis-specification is fairly mild. Moreover, in situations where $\text{LP}$ has a substantially lower RMSE than the other estimators, the post GFIC-estimator's performance is comparable.
To provide a basis for comparison, we now consider results for a number of alternative selection procedures. The first is a “Downward J-test,” which is intended to approximate what applied researchers may do in practice when faced with a model and moment selection decision such as this one. The Downward J-test selects the most restrictive specification that is not rejected by a J-test test with significance level $\alpha \in \left\{ 0.05, 0.1 \right\}$. We test the specifications $\left\{\text{S}, \text{P}, \text{LS}, \text{LP}\right\}$ in order and report the first that is not rejected. This means that we only report $\text{LP}$ if all the other specifications have been rejected. This procedure is, of course, somewhat ad hoc because the significance threshold $\alpha$ is chosen arbitrarily rather than with a view towards some kind of selection optimality. We also consider the GMM model and moment selection criteria of AndrewsLu. These take the form \[ MMSC_n(b,c) = J_n(b,c) - (|c|-|b|) \kappa_n \] where $|b|$ is the number of parameters estimated, $|c|$ the number of moment conditions used, and $\kappa_n$ is a function of $n$. Setting $\kappa_n = \log n$ gives the GMM-BIC, while $\kappa_n = 2$ gives the GMM-AIC and $\kappa_n = 2.01 (\log \log n)$ gives the GMM-HQ. Under certain assumptions, it can be shown that both the GMM-BIC and GMM-HQ are consistent: they select the maximal correctly specified estimator with probability approaching one in the limit. To implement these criteria, we calculate the J-test based on the optimal, two-step GMM estimator with a panel robust, heteroscedasticity-consistent, centered covariance matrix estimator for each specification. To compare selection procedures we use the same simulation grid as above, namely $\gamma, \sigma_{xv} \in \{0, 0.005, 0.01, \hdots, 0.195, 0.20\}$. Again, each point on the simulation grid is calculated from 2000 simulation replications. The bottom panel of Table (ref) presents results for these alternative selection procedures. There is no clear winner in point-wise RMSE comparisons between the GFIC and its competitors. A substantial difference between the GFIC and its competitors emerges, however, when we examine worst-case RMSE. Here the GFIC clearly dominates, providing the lowest worst-case RMSE across all configurations of $T$ and $n$. The differences are particularly stark for larger sample sizes. For example, when $T=5$ and $n=500$ the worst-case RMSE of GFIC is approximately half that of its nearest competitor: GMM-AIC. The consistent criteria, GMM-BIC and GMM-HQ, perform particularly poorly in terms of worst-case RMSE. This is unsurprising given that the worst-case risk of any consistent selection criteria diverges as sample size increases.\footnote{See, e.g., LeebPoetscher2008.}
We now consider an empirical example illustrating the GFIC in the dynamic panel setting introduced in Section (ref). Our exercise is based on BaltagiEtAl2000 who study the demand for cigarettes using panel data for 46 U.S.\ states between 1963 and 1992. Their model is \[ \ln C_{it} = \gamma \ln C_{i,t-1} + \theta \ln P_{it} + \beta_1 \ln Y_{it} + \beta_2 \ln Pn_{it} + \eta_i + \lambda_t + v_{it} \] where $C_{it}$ is the number of packs of cigarettes sold per person aged 16 and above, $P_{it}$ is the real average retail price of a pack of cigarettes, $Y_{it}$ is per capita disposable income, $Pn_{it}$ is the minimum average price of a pack of cigarettes in any state that neighbors state $i$, $\eta_i$ is a state fixed effect, and $\lambda_t$ is a time fixed effect. The lagged dependent variable in this model is meant to capture habit-persistence in cigarette consumption but it is the price elasticity not the habit-persistence per se that is of primary interest.
BaltagiEtAl2000 consider an exhaustive list of possible estimators for two target parameters, the short-run price elasticity $\theta$ and the long-run price elasticity $\theta/(1 - \gamma)$, and explore how the resulting estimates vary. Here we consider selecting between four alternative estimators of the short-run price elasticity $\theta$, as in the second simulation experiment from Section (ref). Each specification is estimated by TSLS in first differences, using the expressions from section (ref).\footnote{Our baseline specification is similar to the estimator that BaltagiEtAl2000 refer to as FD2SLS. There are two differences however. First, whereas they use lags of the exogenous controls $\ln Y_{it}$ and $\ln Pn_{it}$, our instrument set follows AndersonHsiao. Second, whereas BaltagiEtAl2000 appear to have inadvertantly omitted the time dummies from their FD2SLS specification, we include them.} For simplicity, we assume that the controls $\ln Y_{it}$ and $\ln Pn_{it}$ are exogenous with respect to $v_{it}$. We focus on two questions. First: what exogeneity assumption should we impose on $\ln P_{it}$? Second: should we allow for habit persistence by estimating $\gamma$?
Our baseline specification, $\text{LP}$, estimates both $\gamma$ and $\theta$ and assumes only that $\ln P_{it}$ is predetermined with respect to $v_{it}$. This estimator uses the instrument set $\mathbf{z}_{it}(\ell, \text{P})$ with $\ell = 1$ from Equation (ref). For the purposes of this exercise we assume that $\text{LP}$ is correctly specified. The specification $\text{LS}$ also estimates $\gamma$, but uses the expanded instrument set $\mathbf{z}_{it}(\ell, \text{S})$ with $\ell=1$ from equation (ref). The additional instruments used in $\text{LS}$ are only valid if $\ln P_{it}$ is strictly exogenous with respect to $v_{it}$. Like $\text{LP}$ and $\text{LS}$, the specifications $\text{P}$ and $\text{S}$ differ in whether or not they impose that $\ln P_{it}$ is predetermined or strictly exogenous. In contrast, however, they set $\gamma = 0$ and estimate a model with no habit persistence ($\ell =0$). This increases the number of time periods available for estimation.
Estimates and GFIC results for all specifications appear in Table (ref). In Panel (ref) we use data from 1975--1980 only ($T=6$). After first-differencing, this leaves 4 time periods for estimation in specifications that include a lag ($\text{LP}$ and $\text{LS}$) versus 5 for those that do not ($\text{P}$ and $\text{S}$). In panel (ref) we use data from 1975--1985 ($T=11$). After first-differencing this leaves 9 time periods for estimation without a lag versus 10 for estimation with a lag. We choose to artificially shorten the time dimension of the panel from BaltagiEtAl2000 to illustrate a key feature of the GFIC, namely that it takes into account the number of available time periods when selecting over parameter restrictions and moment conditions.\footnote{Appendix (ref) presents additional results, including the full-sample estimates, and some further discussion.} Each column in Table (ref) refers to a particular specification: $\text{LP}$, $\text{LS}$, $\text{P}$ or $\text{S}$. The first row of each panel gives the associated estimate of the target parameter $\theta$ while the second gives the estimated asymptotic variance of $\sqrt{n}(\widehat{\theta} - \theta_0)$, one of the two ingredients of the GFIC.\footnote{We estimate the asymptotic variance matrix as in BaltagiEtAl2000.} Unsurprisingly the asymptotic variance decreases with the number of time periods available for estimation: for a given sample period $\text{LP}$ and $\text{LS}$ have a higher asymptotic variance than $\text{P}$ and $\text{S}$, and all the estimates based on the 1975--1985 sample are more precise than their counterparts for the 1975-1980 sample. Moreover, for a given sample period, the estimators with a large instrument set show a lower asymptotic variance: $\text{LS}$ is more precisely estimated than $\text{LP}$ and $\text{S}$ is more precisely estimated than $\text{P}$.
The third row of each panel gives our estimate of the squared asymptotic bias of the various estimators of $\theta$. The “---” entry for the $\text{LP}$ estimator indicates that the GFIC is constructed under the assumption that this specification has no asymptotic bias. Note that the squared bias estimator is negative for $\text{LS}$ and $\text{S}$ in the 1975--1980 sample. This occurs when the second term in the estimate of the bias matrix $\widehat{B}$ from Corollary (ref), namely $\widehat{\Psi}\widehat{\Omega}\widehat{\Psi}'$, is larger than the first. Accordingly, we consider two alternative ways of constructing the GFIC from the asymptotic variance and squared bias estimates. The first, labelled “GFIC,” simply adds squared bias and variance. The second, labelled “GFIC+,” first truncates a negative squared bias estimate to zero, and then adds the result to the variance estimate. For the 1975--1980 sample period we find no evidence of appreciable bias in any of the three “suspect” specifications: $\text{LS}$, $\text{P}$, and $\text{LP}$. In contrast, each of these has a substantially small asymptotic variance than $\text{LP}$, so we would select either $\text{LS}$ or $\text{S}$ depending on whether we prefer to use GFIC or GFIC+.\footnote{When GFIC and GFIC+ disagree, we prefer GFIC+ for the same reason that the positive-part Stein estimator is preferred to the “plain-vanilla” Stein estimator.} In the 1975--1985 sample, the situation changes drastically. Over this longer time period, the relative advantage of $\text{LS}$, $\text{P}$ and $\text{S}$ over $\text{LP}$ in asymptotic variance decreases substantially and we find evidence of substantial bias in both the $\text{LS}$ and $\text{S}$ specifications. It appears that over these additional time periods, the assumption that $\ln P_{it}$ is strictly exogenous fails. Interestingly the difference in GFIC values between $\text{LP}$ and $\text{P}$ in this longer sample is negligible. As $\text{P}$ has a lower GFIC value than $\text{LP}$ in the 1975--1980 sample, it appears that accounting for habit persistence is relatively unimportant in estimating the short-run price elasticity of cigarette demand.
This paper has introduced the GFIC, a proposal to choose moment conditions and parameter restrictions based on the quality of the estimates they provide. The GFIC performs well in simulations for our dynamic panel example. While we focus here on applications to panel data, the GFIC can be applied to any GMM problem in which a minimal set of correctly specified moment conditions identifies an unrestricted model. A possible extension of this work would be to consider risk functions other than MSE, by analogy to ClaeskensCroux2006 and ClaeskensHjort2008. Another possibility would be to derive a version of the GFIC for GEL estimators. Although first-order equivalent to GMM, GEL estimators often exhibit superior finite-sample properties and may thus improve the quality of the selection criterion NeweySmith.
\singlespacing