Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
100,317 characters · 22 sections · 85 citation commands
Sensitivity Analysis using Approximate Moment Condition Models
Economic models are typically viewed as approximations of reality. However, conventional approaches to estimation and inference assume that a model holds exactly. In this paper, we weaken this assumption, and consider inference in a class of models characterized by moment conditions which are only required to hold in an approximate sense. The failure of the moment conditions to hold exactly may come from failure of exclusion restrictions (e.g.\ through omitted variable bias or because instruments enter the structural equation directly in an IV model), functional form misspecification, or other sources such as measurement error, or data contamination.
We assume that we have a model characterized by a set of population moment conditions $g(\theta)$. In the generalized method of moments (GMM) framework, for instance, $g(\theta)=E[g(w_{i}, \theta)]$, which can be estimated by the sample analog $\frac{1}{n}\sum_{i=1}^{n}g(w_{i}, \theta)$, based on the sample $\{w_{i}\}_{i=1}^{n}$. When evaluated at the true parameter value $\theta_{0}$, the population moment condition lies in a known set specified by the researcher,
The set $\mathcal{C}$ formalizes the way in which the moment conditions may fail, and it can then be varied as a form of sensitivity analysis, with $\mathcal{C}=\{0\}$ reducing to the correctly specified case.
We focus on local misspecification: the scaling by $\sqrt{n}$ implies that the specification error and the sampling error are of the same order of magnitude, and it ensures that the asymptotic approximation captures the fact that it may not be clear from the sample at hand whether the model is correctly specified. It also leads to increased tractability, allowing us to deliver a simple method for inference on a structural parameter of interest $h(\theta_{0})$, rather than a pseudo-true parameter. This tractability has made local misspecification a popular tool for sensitivity analysis in applied work, especially following the recent influential paper by andrews_measuring_2017.\footnote{For recent empirical examples using local sensitivity analysis, see gayle2019, or duflo2018.} As with any asymptotic device, our modeling of misspecification as local should not be taken to mean that we literally believe that the model would be closer to correct if we had more data. Rather, its usefulness should be judged by whether it yields accurate approximations to the finite-sample behavior of estimators and confidence intervals, which in our case requires that the set $\mathcal{C}/\sqrt{n}$ be small relative to sampling uncertainty.
We propose a simple method for constructing asymptotically valid confidence intervals (CIs) under this setup: one takes a standard estimator, such as the GMM estimator, and adds and subtracts its standard error times a critical value that takes into account the potential asymptotic bias of the estimator, in addition to its variance. A key insight of this paper is that because the CIs must be widened to take into account the potential bias, the optimal weighting matrix for the correctly specified case (the inverse of the variance matrix of the moments) is generally no longer optimal under local misspecification. Rather, the optimal weighting matrix takes into account potential misspecification in the moments in addition to the variance of their estimates: it places less weight on moments that are allowed to be further from zero according the researcher's specification of the set $\mathcal{C}$. We also show that an analogous result holds for other performance criteria, such as estimation under the mean-squared error: the optimal weighting matrix again trades off the potential misspecification of the moments against their precision, although the optimal tradeoff is different.
To illustrate the practical importance of this result, we apply our methods to form misspecification-robust CIs in an empirical model of automobile demand based on blp95. We consider sets $\mathcal{C}$ motivated by the forms of local misspecification considered in andrews_measuring_2017, who calculate the asymptotic bias of the usual GMM estimator in this model. We find that adjusting the weighting matrix to account for potential misspecification substantially reduces the potential bias of the estimator and, as a result, leads to large efficiency improvements of the optimal CI relative to a CI based on the GMM estimator that is optimal under correct specification: it shrinks the CI by up to a factor of $3$ or more in our main specifications. As a result, we obtain informative CIs in this model even under moderate amounts of misspecification.
We show that the CIs we propose are near-optimal when the set $\mathcal{C}$ is convex and centrosymmetric ($c\in\mathcal{C}$ implies $-c\in\mathcal{C}$). To this end, we argue that the relevant “limiting experiment” for the locally misspecified GMM model is isomorphic to an approximately linear model of SaYl78, which falls under a general framework studied by, among others, donoho94, CaLo04 and ArKo18optimal. We derive asymptotic efficiency bounds for CIs in the locally misspecified GMM model that formally translate bounds from the approximately linear limiting experiment to the locally misspecified GMM setting. In particular, these bounds imply that the scope for improvement over our CIs by optimizing expected length at a particular value of $\theta_{0}$ and $c=0$ (while still maintaining coverage over the whole parameter space for $\theta$ and $\mathcal{C}$) is limited, even if one optimizes expected length at the true values of $\theta_{0}$ and $c$.
These efficiency bounds address an important criticism of our CIs: they require a priori specification of the set $\mathcal{C}$ that defines misspecification, including both the magnitude of misspecification and which moments are misspecified. In particular, one cannot substantively improve upon our CI by, say, trying to use data-driven methods that gauge misspecification magnitude or try to determine which moments are misspecified. These bounds have implications for procedures proposed by andrews_hybrid_2009, ditraglia_using_2016 and mccloskey_2017, who consider the case where some moments are known to be correctly specified and no a priori bound is placed on the magnitude of misspecification of the remaining moments. As we discuss in (ref), in this case our CI reduces to the usual CI based on the $k_1$ correctly specified moments, and our efficiency bounds show that CIs proposed in these papers cannot substantively improve upon it.
Because we cannot use the data to determine the magnitude $M$ of the set $\mathcal{C}$, we recommend plotting our CIs as a function of the potential misspecification magnitude $M$, or reporting the smallest value of $M$ for which a particular finding breaks down. Such sensitivity analysis is easy to conduct under out proposed implementation. In particular, we show that, when the set $\mathcal{C}$ is characterized by $\ell_{p}$ constraints, the class of weightings that trace out the optimal bias-variance tradeoff as a function of how much relative weight we put on the bias can be easily computed by recasting the problem as a penalized regression problem. By exploiting this analogy, we develop a simple algorithm for computing this class under $\ell_{\infty}$ constraints that is similar to the LASSO/LAR algorithm ehjt04,rosset_piecewise_2007; under $\ell_{2}$ constraints, the solution admits a closed form.\footnote{An R package implementing our CIs under $\ell_{p}$ constraints is available at \url{https://github.com/kolesarm/GMMSensitivity}.} Furthermore, as we discuss in (ref), this class of weightings is entirely determined by the shape of $\mathcal{C}$; its magnitude $M$ only determines the optimal relative weight we should put on the bias. Thus, tracing out the optimal weighting as a function of $M$ can be done at essentially no additional computational cost. Furthermore, to avoid having to reoptimize the objective function with respect to the new weighting matrix, one can also form the CIs by adding and subtracting our critical value from a one-step estimator newey_large_1994 based on any initial estimate that is $\sqrt{n}$-consistent under correct specification. We illustrate this approach in our empirical application in (ref).
Our paper is related to several strands of literature. Our efficiency results are related to those in chamberlain_asymptotic_1987 for point estimation in the correctly specified setting (see also hansen85) and, more broadly, semiparametric efficiency theory in correctly specified settings van_der_vaart_asymptotic_1998. As we discuss in (ref), some of our efficiency results are novel even in the correctly specified case, and may be of independent interest. kitamura_robustness_2013 consider efficiency of point estimators satisfying certain regularity conditions when the misspecification is bounded by the Hellinger distance. As we discuss in more detail in (ref), our results imply that under this form of misspecification, the optimal weighting matrix remains the same as under correct specification; both the usual GMM estimator and the estimator proposed by kitamura_robustness_2013 can thus be used to form near-optimal CIs, and both estimators have the same local asymptotic minimax properties.
Local misspecification has been used in a number of papers, which include, among others, newey_generalized_1985, berkowitz_validity_2012, conley_plausibly_2012, guggenberger_asymptotic_2012, and bugni_inference_2018, and has antecedents in the literature on robust statistics huber_robust_2011. andrews_measuring_2017 consider this setting and note that asymptotic bias of a regular estimator can be calculated using influence function weights, which they call the sensitivity, and show how such calculations can be used for sensitivity analysis in applications (see also extensions of these ideas in andrews_informativeness_2018 and mukhin_sensitivity_2018). Our results imply that, if one is interested in inference, conclusions of such sensitivity analysis may be substantially sharpened by using the misspecification-optimal weighting matrix, or, equivalently, the misspecification-optimal sensitivity. In independent work, bonhomme_minimizing_2018 provide a framework for estimation and inference in misspecified likelihood models when the misspecification set $\mathcal{C}$ is defined with respect to a larger class of models using statistical notions of distance. Our focus is on overidentified moment condition models, as in andrews_measuring_2017, and we are agnostic about how $\mathcal{C}$ is determined. The proposal to use estimators that optimize an asymptotic bias-variance tradeoff using the influence function is common to both papers. The efficiency bounds in (ref) are unique to the present paper.
The rest of this paper is organized as follows. (ref) presents our misspecification robust CIs. (ref) gives step-by-step instructions for computing our CIs, along with discussion of other practical issues. (ref) presents efficiency bounds for CIs in locally misspecified models; it can be skipped by readers interested only in implementing the methods. (ref) discusses applications to particular moment condition models. (ref) presents an empirical application. Additional results and proofs are collected in appendices and an online supplement.
We have a model that maps a vector of parameters $\theta\in\Theta\subseteq \mathbb{R}^{d_{\theta}}$ to a $d_{g}$-dimensional population moment condition $g(\theta)$ that restricts the distribution of the observed data $\{w_{i}\}_{i=1}^{n}$. We allow the moment condition model to be locally misspecified, so that at the true value $\theta_{0}$, the population moment condition is not necessarily zero, but instead lies in a $\sqrt{n}$-neighborhood of $0$:
where $\mathcal{C}\subseteq\mathbb{R}^{d_{g}}$ is a known set. The set $\mathcal{C}$ may allow for misspecification in potentially all moment conditions; we do not require that some elements of $c$ are zero. Our goal is to construct a CI for a scalar $h(\theta_0)$, where $h\colon \mathbb{R}^{d_\theta}\to\mathbb{R}$ is a known function. For example, if we are interested in one of the elements $\theta_{j}$ of $\theta$, we would take $h(\theta)=\theta_{j}$. More generally, the function $h$ will be nonlinear, as is, for example, generally the case when $\theta$ is a vector of supply or demand parameters, and $h(\theta)$ is an elasticity, or some counterfactual.
This setup allows (but does not require) both $\theta_{0}$ and $h(\theta_{0})$ to have the same interpretation as in the correctly specified case, so that our CIs may still be interpreted as CIs for the structural parameter, elasticity, or counterfactual of interest. For this interpretation, one typically needs to rule out forms of misspecification that affect the mapping $\theta\mapsto h(\theta)$. While we do not formally consider cases in which this mapping itself is misspecified, such cases are covered under a mild generalization of our framework, in which $h$ is a function of both $\theta$ and $c$.
Note that the interpretation of $h(\theta_{0})$, and the conceptual framework defining $\theta_{0}$ is not affected by our modeling of the misspecification as local: given a set ${\widetilde{\mathcal{C}}}_{n}=\mathcal{C}/\sqrt{n}$, the moment conditions $g(\theta_{0})\in{\widetilde{\mathcal{C}}}_{n}$ describe the restrictions that the data generating process and the researcher's modeling assumptions place on $\theta_0$.\footnote{Formally, for a given sample size $n$, $\theta_{0}$ may be set identified, and the identified set under a distribution $P$ is defined as the set of parameters $\theta_0$ that satisfy the moment conditions $E_{P}g(w_i, \theta_0)\in {\widetilde{\mathcal{C}}}_{n}$ where $E_P$ denotes expectation under $P$. We construct CIs that cover $h(\theta_0)$ for points $\theta_0$ in the identified set im04. See (ref) and (ref) for formal definitions of coverage and optimality of our CIs.} The plausibility of these restrictions is evaluated for a given sample size at hand; it doesn't depend on assumptions about how ${\widetilde{\mathcal{C}}}_{n}$ changes with $n$. While we focus on sequences ${\widetilde{\mathcal{C}}}_n=\mathcal{C}/\sqrt{n}$, we discuss in (ref) how our insights can be used to construct CIs that are valid global misspecification, when ${\widetilde{\mathcal{C}}}_{n}$ is fixed with $n$.
To formalize the notion of asymptotic validity and efficiency of CIs, we will need to allow the true parameter value $\theta_0$ as well as the vector $c$ and the data generating process (and hence the map $\theta\mapsto g(\theta)$) to vary with the sample size. For clarity of exposition, we focus here on the case in which these parameters are fixed. See (ref) and (ref) for the general case. Under some forms of misspecification, such as functional form misspecification, there may be additional higher-order terms on the right-hand side of (ref); our results remain unchanged if this is the case. Again, for clarity of exposition, we focus on the case in which (ref) holds exactly.
We assume that the sample moment condition $\hat{g}(\theta)$, constructed using the data $\{w_{i}\}_{i=1}^{n}$, satisfies
where $\stackrel{d}{\to}$ denotes convergence in distribution as $n\to\infty$. In the GMM model, the population and sample moment conditions are given by $g(\theta)=E[g(w_{i}, \theta)]$ and $\hat{g}(\theta)=\frac{1}{n}\sum_{i=1}^{n}g(w_{i}, \theta)$, respectively, where $g(\cdot, \cdot)$ is a known function. However, to cover other minimum distance problems, we do not require that the moment conditions necessarily take this form. We further assume that the moment condition is smooth enough so that
where $\Gamma$ is the $d_{g}\times d_{\theta}$ derivative matrix of $g$ at $\theta_{0}$. Conditions (ref) and (ref) are standard regularity conditions in the literature on linear and nonlinear estimating equations; see newey_large_1994 for primitive conditions. Finally, we also assume that $h$ is continuously differentiable with the $1\times d_{\theta}$ derivative matrix at $\theta_{0}$ given by $H$.
Under correct specification, when $\mathcal{C}=\{0\}$, standard estimators $\hat{h}$ of $h(\theta)$ are asymptotically linear in $\hat{g}(\theta_{0})$. This will typically extend to our locally misspecified case, so that for some vector $k\in\mathbb{R}^{d_g}$,
where the convergence in distribution follows by (ref) and (ref). If in addition, the estimator is regular (so that equality in (ref) holds uniformly for $\theta$ in a $\sqrt{n}$-neighborhood of $\theta_{0}$), then $k$ will satisfy (see e.g. Section 2 in newey_semiparametric_1990)
For example, in a GMM model, if we take $\hat h=h(\hat{\theta}_{W})$ where
is the GMM estimator with weighting matrix $W$, (ref) will hold with $k'=-H(\Gamma' W\Gamma)^{-1}\Gamma' W$ newey_generalized_1985. Because the vector $k$ determines the local asymptotic bias of the estimator, we follow andrews_measuring_2017, and refer to $k$ as the sensitivity of $\hat{h}$.
We now show how to construct misspecification-robust CIs based on an asymptotically linear estimator $\hat{h}$ with a given sensitivity $k$. In (ref), we show how to choose this sensitivity optimally, to achieve the shortest CI among those based on regular asymptotically linear estimators. In (ref), we will show that, under this choice of $k$, the resulting CI is (near) optimal not only within the class of CIs based on regular asymptotically linear estimators, but among all CIs that satisfy the asymptotic coverage requirement.
Let $\hat{k}$ and $\hat{\Sigma}$ be consistent estimates of $k$ and $\Sigma$. Then by Slutsky's theorem,
Under correct specification, the right-hand side corresponds to a standard normal distribution, and we can form a CI with asymptotic coverage $100\cdot (1-\alpha)\%$ as $\hat{h}\pm z_{1-\alpha/2}\sqrt{\hat{k}'\hat{\Sigma}\hat{k}/n}$, where $z_{1-\alpha/2}$ is the $1-\alpha/2$ quantile of a $\mathcal{N}(0,1)$ distribution; this is the usual Wald CI\@.
When we allow for misspecification, the Wald CI will no longer be valid. However, note that the asymptotic bias ${k'c}/{\sqrt{k'\Sigma k}}$ is bounded in absolute value by ${\operatorname{\overline{bias}}_{\mathcal{C}}(k)}/{\sqrt{k'\Sigma k}}$ where $\operatorname{\overline{bias}}_{\mathcal{C}}(k)\equiv \sup_{c\in\mathcal{C}}\abs{k'c}$. Therefore, given $c$, the $z$-statistic in the preceding display is asymptotically $\mathcal{N}(t,1)$ where $\abs{t}\le \operatorname{\overline{bias}}_{\mathcal{C}}(k)/\sqrt{k'\Sigma k}$. This leads to the CI
where $\operatorname{cv}_\alpha(\overline{t})$ is the $1-\alpha$ quantile of $\abs{Z}$, with $Z\sim \mathcal{N}(\overline{t},1)$. In particular, $\operatorname{cv}_{\alpha}(0)=z_{1-\alpha/2}$, so that in the correctly specified case, (ref) reduces to the usual Wald CI\@. As we discuss in (ref), in the limiting experiment, this CI becomes equivalent to the fixed-length CI proposed by donoho94.
To form a one-sided CI based on an estimator $\hat{h}$ with sensitivity $k$, one can simply subtract its maximum bias, in addition to the standard error:
One could also form a valid two-sided CI by adding and subtracting the worst-case bias $\operatorname{\overline{bias}}_{\mathcal{C}}(\hat k)/\sqrt{n}$ from $\hat{h}$, in addition to adding and subtracting $z_{1-\alpha/2}\sqrt{\hat{k}'\hat\Sigma\hat k/n}$; however, since $\hat{h}$ cannot simultaneously have a large positive and a large negative bias, such CI will be conservative, and longer than the CI in (ref).
The asymptotic length of the CI in (ref) is given by
To attain the shortest possible CI, we therefore need to use an estimator with sensitivity that minimizes this expression. We restrict attention to asymptotically linear estimators that are regular, so that we need to minimize (ref) subject to (ref). The CI length in (ref) depends on $\theta$ only through $\Sigma$. Furthermore, it depends on the sensitivity only through the maximum bias, $\operatorname{\overline{bias}}_{\mathcal{C}}(k)$, and the variance $k'\Sigma k$. Therefore, rather than minimizing (ref) directly over all sensitivities $k$, one can first minimize the variance subject to a bound $\overline{B}$ on the worst-case bias,
and then vary the bound $\overline{B}$ to find the bias-variance trade-off that leads to the shortest CI\@. In our implementation in (ref), we focus on the case where $\mathcal{C}$ is characterized by $\ell_p$ constraints, in which case a closed-form expression for the worst-case bias $\sup_{c\in\mathcal{C}}\abs{k'c}$ is available, and it is computationally trivial to solve (ref) directly or in Lagrange multiplier form. In general, when the set $\mathcal{C}$ is convex, one can reformulate (ref) as a convex optimization problem, leading to a computationally tractable solution (see (ref)). One can also use (ref) to determine the optimal sensitivity for constructing one-sided CIs, if we use quantiles of excess length as the criterion for choosing a CI\@. We provide details in (ref).
Once the optimal sensitivity has been determined, we can implement an estimator with this sensitivity as a one-step estimator. In particular, let $\hat\theta_{\text{initial}}$ be an initial $\sqrt{n}$-consistent estimator of $\theta_0$, let $\hat{k}= k+o_{P}(1)$ be a consistent estimator of the desired sensitivity $k$. Then the one-step estimator
will have the desired sensitivity. This follows from the Taylor expansion
where the second line follows from (ref). It then follows from (ref) that the first term converges in probability to zero, and $\hat{h}$ satisfies (ref).
We now give step-by-step instructions for computing our CI\@. To make it easy to determine the sensitivity of the CI to the magnitude of misspecification, we consider sets of the form $\mathcal{C}=\mathcal{C}(M)=\{Mc\colon c\in\mathcal{C}(1)\}$, where the scalar $M$ measures the magnitude of misspecification. We discuss the exact specification of the set $\mathcal{C}(M)$ in (ref) below.
The fact that $M$ simply scales the potential magnitude of misspecification leads to a simplification when tracing out the optimal CI as a function of $M$. In particular, let $\{k_{\lambda}\}_{\lambda\geq 0}$ be the bias-variance optimizing class of sensitivities that traces out the solutions to (ref) as we vary the bound $\overline{B}$ when $\mathcal{C}=\mathcal{C}(1)$. The index $\lambda$ determines the relative weight on the bias; it can correspond to the Lagrange multiplier in a Lagrangian formulation of (ref), or we can simply take $\lambda=\overline B$ if we are solving (ref) directly. Let $\overline{B}_{\lambda}=\operatorname{\overline{bias}}_{\mathcal{C}(1)}(k_\lambda)$. It then follows by a change-of-variables argument that $\operatorname{\overline{bias}}_{\mathcal{C}(M)}(k_\lambda)= M\overline B_\lambda$, and that $k_\lambda$ minimizes the asymptotic variance subject to this bound on worst-case bias over $\mathcal{C}(M)$. Thus, $\{k_\lambda\}_{\lambda\geq 0}$ is also a bias-variance optimizing class of sensitivities for $\mathcal{C}(M)$. We therefore only need to compute the class $\{k_{\lambda}\}_{\lambda\geq 0}$ only once, even when a range of values $M$ is considered.
With this simplification, we can construct CIs for a range of values of $M$ as follows:
The CI given in (ref) has the apparent defect that the local misspecification vector $c$ is reflected in the length of the CI only through the a priori restriction $\mathcal{C}$ imposed by the researcher. Thus, if the researcher is conservative about misspecification, the CI will be wide, even if it turns out that $c$ is in fact much smaller than the a priori bounds defined by $\mathcal{C}$. Moreover, this approach requires the researcher to explicitly specify the set $\mathcal{C}$, including any tuning parameters such as the parameter $M$ if the set takes the form $\mathcal{C}=\mathcal{C}(M)$ that we considered in (ref). One may therefore seek to improve upon this CI by estimating the magnitude of $c$, or by estimating the tuning parameters, and constructing a CI that is shorter if these estimates indicate that misspecification is mild. Similarly, it may be restrictive to require that the CI be centered at an asymptotically linear estimator: this rules out, for example, using a $J$-test to decide which moments to use.
The main result of this section shows that, when $\mathcal{C}$ is convex and centrosymmetric ($c\in\mathcal{C}$ implies $-c\in\mathcal{C}$), the scope for improving on the CI in (ref) is nonetheless limited: no sequence of CIs that maintain coverage under all local misspecification vectors $c\in\mathcal{C}$ can be substantially tighter, even under correct specification. This result can be interpreted as translating results from a “limiting experiment” that is an extension of the linear regression model. We first give a heuristic derivation of this limiting experiment and explain our result in the context of this limiting experiment. We then present the formal asymptotic result, and discuss its implications in some familiar settings. Readers who are interested only in implementing the methods, rather than efficiency results, can skip this section.
We restrict attention in this section to the GMM model, in which $\hat g(\theta)=\frac{1}{n}\sum_{i=1}^{n}g(w_i, \theta)$, and we further restrict the data $\{w_{i}\}_{i=1}^{n}$ to be independent and identically distributed (i.i.d.). Similar to semiparametric efficiency theory in the standard, correctly specified case, this facilitates parts of the formal statements and proofs, such as the definition of the set of distributions under which coverage is required and the construction of least favorable submodels. We expect that analogous results could be obtained in other settings.
As discussed in (ref), we can form CIs based on linear estimators with asymptotic distribution $\mathcal{N}(k'c, k'\Sigma k)$. This suggests that the problem of constructing an asymptotically valid CI for $h(\theta)$ in the model (ref) is asymptotically equivalent to the problem of constructing a CI for the parameter $H\theta$ in the approximately linear model
where $\Gamma$, $H$ and $\Sigma^{1/2}$ are known, and we observe $Y$. One can think of this model as an “approximately” linear regression model, with $-\Gamma$ playing the role of the design matrix of the (fixed) regressors, and $c$ giving the approximation error. This model dates back at least to SaYl78, who considered estimation in this model when $\mathcal{C}$ is a rectangular set and $\Sigma$ is diagonal. The analog of the asymptotically linear estimator $\hat{h}$ in (ref) is the linear estimator $k'Y$. To see the analogy, note that $k'Y-H\theta$ is distributed $\mathcal{N}((-k'\Gamma-H)\theta+k'c, k'\Sigma k)$, and restricting ourselves to estimators that do not have infinite worst-case bias when $\theta$ is unrestricted gives the condition (ref). In the limiting experiment, the analog of the CI (ref) is given by the CI $k'Y\pm \operatorname{cv}_\alpha(\operatorname{\overline{bias}}_{\mathcal{C}}(k)/\sqrt{k'\Sigma k})\cdot \sqrt{k'\Sigma k}$. Finding weights that minimize the length of this CI is isomorphic to the problem of finding the sensitivity that minimizes the asymptotic CI length in (ref).
For a general convex set $\mathcal{C}$, the bias-variance optimization problem in (ref) can be reformulated as a convex programming problem, as shown in low95. In particular, when the set $\mathcal{C}$ is centrosymmetric (see (ref) for the general case), the bias-variance optimizing class of weights $\{k_{\lambda}\}_{\lambda> 0}$ is given by the class $\{k_{\delta}\}_{\delta> 0}$, where
and, for each $\delta$, $c_\delta, \theta_\delta$ are the solutions to the convex program
It then follows from donoho94 that among fixed-length CIs based on linear estimators (CIs that take the form $k'Y \pm \chi$ for some constant $\chi$), the shortest CI in the limiting experiment takes the form
where $\operatorname{\overline{bias}}_{\mathcal{C}}(k_\delta) =-k_\delta' c_\delta$, and $\delta^{*}=\operatorname*{argmin}_{\delta>0} 2\operatorname{cv}_\alpha(\operatorname{\overline{bias}}_{\mathcal{C}}(k_{\delta}) / \sqrt{k_{\delta}' \Sigma k_{\delta}})\cdot \sqrt{k_{\delta}'\Sigma k_{\delta}}$ is chosen to minimize the CI length. The CI in (ref) is an analog of this CI, with $\delta$ playing the role of the index $\lambda$.
The CI in (ref) takes a familiar form in the special case in which $\mathcal{C}$ is a linear subspace of $\mathbb{R}^{d_{g}}$, so that for some $d_{g}\times d_{\gamma}$ full-rank matrix $B$ with $d_{\gamma}\leq d_{g}-d_{\theta}$, $\mathcal{C}=\{B\gamma\colon \gamma\in\mathbb{R}^{d_{\gamma}}\}$. Let $B_{\perp}$ denote a $d_{g}\times (d_{g}-d_{\gamma})$ matrix that's orthogonal to $B$. Then for any $\delta>0$, $k_{\delta}'={k}'_{LS, B}$, where
is the sensitivity of the GLS estimator after pre-multiplying (ref) by $B_{\perp}'$, (which effectively picks out the observations with zero misspecification). Since this estimator is unbiased, the CI in (ref) becomes ${k}_{LS, B}'Y\pm z_{1-\alpha/2}\sqrt{k_{LS, B}'\Sigma k_{LS, B}}$.
Like the asymptotic CI (ref), the CI in (ref) has the potential drawback that its length is determined by the worst possible misspecification in $\mathcal{C}$, leaving open the possibility of efficiency improvements when $c$ turns out to be close to zero. As a best-case scenario for such improvements, consider the problem: among confidence sets with coverage at least $1-\alpha$ for all $\theta\in\mathbb{R}^{d_\theta}$ and $c\in\mathcal{C}$, minimize expected length when $\theta=\theta^*$ and $c=0$. Note that this setup is even more favorable for potential improvements on our CI, since it allows the researcher to guess correctly that $\theta$ is equal to some $\theta^*$, and it allows for confidence sets that are not intervals (in this case, length is defined as Lebesgue measure). Let $\kappa_{*}(H, \Gamma, \Sigma, \mathcal{C})$ denote the ratio of this optimized expected length relative to the length of the CI in (ref) (it can be shown that this ratio does not depend on $\theta^*$).
If $\mathcal{C}$ is convex, a formula for $\kappa_{*}(H, \Gamma, \Sigma, \mathcal{C})$ follows from applying the general results in Corollary 3.3 in ArKo18optimal to the limiting model. If $\mathcal{C}$ is also centrosymmetric, then
where $Z\sim \mathcal{N}(0,1)$ and $\omega(\delta)$ is two times the optimized value of (ref). Furthermore, we show in (ref) that the right-hand side is lower-bounded by $(z_{1-\alpha}(1-\alpha)-\tilde{z}_{\alpha}\Phi(\tilde{z}_{\alpha})+ \phi(z_{1-\alpha})-\phi(\tilde{z}_{\alpha}))/ z_{1-\alpha/2}$, where $\tilde{z}_{\alpha}=z_{1-\alpha}-z_{1-\alpha/2}$ for any $H$, $\Gamma$, $\Sigma$ and $\mathcal{C}$, where $\phi(\cdot)$ denotes the standard normal density. For $\alpha=0.05$, this universal lower bound evaluates to 71.7%. Evaluating $\kappa_{*}$ for particular choices of $H$, $\Gamma$, $\Sigma$, and $\mathcal{C}$ often yields even higher efficiency.
If $\mathcal{C}$ is a linear subspace, then $\omega(\delta)$ is linear, and
where the lower bound follows since $\phi(z_{1-\alpha})\geq \alpha z_{1-\alpha}$ by the Gaussian tail bound $1-\Phi(x)\leq \phi(x)/x$ for $x>0$. This bound corresponds to that in pratt61 for the case of a univariate normal mean. The potential efficiency improvement essentially comes from using prior knowledge of $\theta^*$ to turn a two-sided critical value into a one-sided critical value. Furthermore, it follows from joshi69 that the CI ${k}_{LS, B}'Y\pm z_{1-\alpha/2}\sqrt{k_{LS, B}'\Sigma k_{LS, B}}$ is the unique CI that achieves minimax expected length. Thus, not only is the scope for improvement at a particular $\theta^{*}$ bounded by (ref), any CI with shorter expected length at some $\theta^{*}$ must necessarily perform worse elsewhere in the parameter space.
For the one-sided CI (ref), the analogous CI in the limiting experiment is $\hor{k'Y-\operatorname{\overline{bias}}_{\mathcal{C}}(k)-z_{1-\alpha}\sqrt{k'\Sigma k}, \infty}$, and, as we discuss in (ref), to choose the optimal sensitivity $k$, one can consider optimizing a given quantile of its worst-case excess length. The results in ArKo18optimal again give an efficiency bound for improvement at $c=0$ and a particular $\theta^*$, analogous to (ref) for the two-sided case. See (ref) for details. If $\mathcal{C}$ is a linear subspace, then optimizing quantiles of worst-case excess length yields the CI $\hor{k_{LS, B}'Y-z_{1-\alpha}\sqrt{k_{LS, B}'\Sigma k_{LS, B}}, \infty}$, independently of the quantile one is optimizing. Furthermore, the efficiency bound implies that this one-sided CI is in fact fully optimal over all quantiles of excess length and all values of $\theta, c$ in the local parameter space.
These efficiency results for the CI (ref) in the limiting experiment suggest that the scope for improvement over the CI in (ref) should be limited in large samples. (ref), stated in the next section, uses the analogy with the approximately linear model (ref) along with Le Cam-style arguments involving least favorable submodels to show that this bound indeed translates to the locally misspecified GMM model. For one-sided CIs, we state an analogous result in (ref). We discuss the implications of these results in (ref).
To make precise our statements about coverage and efficiency, we need the notion of uniform (in the underlying distribution) coverage of a confidence interval. This requires additional notation, which we now introduce. Let $\mathcal{P}$ denote a set of distributions $P$ of the data $\{w_{i}\}_{i=1}^{n}$, and let $\Theta_n\subseteq\mathbb{R}^{d_\theta}$ denote the parameter space for $\theta$. We require coverage for all pairs $(\theta, P)\in \Theta_n\times \mathcal{P}$ such that $\sqrt{n}g_P(\theta)\in\mathcal{C}$, where the subscript $P$ on the population moment condition makes it explicit that it depends on the distribution of the data.\footnote{To be precise, we should also subscript all other quantities such as $\Gamma$ and $\Sigma$ by $P$. To prevent notational clutter, we drop this index in the main text unless it causes confusion.} Letting $\mathcal{S}_n=\{(\theta, P)\in \Theta_n\times\mathcal{P}: \sqrt{n}g_P(\theta)\in\mathcal{C}\}$ denote this set, the condition for coverage at confidence level $1-\alpha$ can be written
We say that a confidence set $\mathcal{I}_{n}$ is asymptotically valid (uniformly over $\mathcal{S}_{n}$) at confidence level $1-\alpha$ if this condition holds.\footnote{In general, $\theta_0$ and $h(\theta_0)$ may be set identified for a given sample size $n$ (although our assumptions imply that the identified set will shrink at a root-$n$ rate). The coverage requirement (ref) states that the CI must cover points in the identified set for $h(\theta)$, as in im04; see (ref).}
Among two-sided CIs of the form $\hat{h}\pm \hat{\chi}$ that are asymptotically valid, we prefer CIs with shorter expected length. To avoid issues with convergence of moments, we use truncated expected length, and define the asymptotic expected length of a two-sided CI at $P_{n}\in\mathcal{P}$ as $\liminf_{T\to\infty}\liminf_{n\to\infty}E_{P_{n}}\min\{\sqrt{n}\cdot 2\hat{\chi}, T\}$, where $E_{P}$ denotes expectation under $P$.
We are now ready to state the main efficiency result.
The proof for this theorem is given in (ref), which also gives an analogous result for one-sided confidence intervals. (ref), stated in (ref), require that the conditions in (ref) hold in a uniform sense over the class $\mathcal{P}$. In the supplemental materials, we give primitive conditions for these assumptions in the misspecified linear IV model. (ref), also stated in the appendix, requires that the class $\mathcal{P}$ be rich enough to contain a submodel that is least favorable for the GMM problem, so that the class doesn't implicitly impose any other conditions that could be used to make inference easier. In the supplemental materials, we provide a general way of constructing a submodel satisfying these conditions.
The universal lower bound on $\kappa_{*}$ is new and may be of independent interest. For $\alpha=0.05$, it evaluates to 71.7%. The universal lower bound is sharp in the sense that there exist $\Gamma, \Sigma, H$ and $\mathcal{C}$ for which $\kappa_{*}$ equals this lower bound. In particular applications, the efficiency bound $\kappa_{*}$ can be computed at estimates of $\Gamma$, $\Sigma$ and $H$, and often, this gives much higher efficiencies. We illustrate these bounds in the empirical application in (ref).
To help build intuition for the efficiency bound in (ref), and to relate this result to the literature, we now consider some special cases. We first discuss the (standard) correctly specified case. Second, we consider the case in which some moments are known to be valid, and the misspecification in the remaining moments is unrestricted. This case may be of interest in its own right. We then discuss the general case. Finally, we discuss the connection to certain statistical measures of distance considered in the literature.
Suppose that $\mathcal{C}=\{0\}$. This is in particular a linear subspace of $\mathbb{R}^{d_{g}}$, with $B=0$, and $B_{\perp}=I$, the $d_{g}\times d_{g}$ identity matrix. Thus, in the limiting experiment, the optimal CI uses the GLS estimator $k_{LS,0}'Y$, with $k_{LS,0}$ given in (ref) (with $B=0$). For testing the null hypothesis $H\theta=h_{0}$ against the one-sided alternative $H\theta\geq h_{0}$, the one-sided $z$-statistic based on ${k}_{LS,0}'Y$ is uniformly most powerful van_der_vaart_asymptotic_1998. Inverting these tests yields the CI $\hor{k_{LS,0}'Y-z_{1-\alpha}\sqrt{k_{LS,0}'\Sigma k_{LS,0}}, \infty}$. Since the underlying tests are uniformly most powerful, this CI achieves the shortest excess length, simultaneously for all quantiles and all possible values of the parameter $\theta$. For two-sided CIs, the results described in (ref) imply that the CI $h_{LS,0}'Y\pm z_{1-\alpha/2}\sqrt{k_{LS,0}'\Sigma k_{LS,0}}$ is the unique CI that achieves minimax expected length, and the efficiency of this CI relative to a CI that optimizes its expected length at a single value $\theta^*$ of $\theta$ when indeed $\theta=\theta^*$ is given in (ref). It evaluates to 84.99% at $\alpha=0.05$.
Applying (ref) to the case $\mathcal{C}=\{0\}$ gives an asymptotic version of the two-sided efficiency bound. Furthermore, the CI in (ref) reduces to the usual two-sided CI based on $\hat{\theta}_{\Sigma^{-1}}$. Thus, in this case, (ref) shows that very little can be gained over the usual two-sided CI by optimizing the CI relative to a particular distribution $P_0$. Results in the appendix give an analogous result for one-sided CIs. In the one-sided case, this asymptotic result is essentially a version of a classic result from the semiparametric efficiency literature for one-sided tests, applied to CIs (see Chapter 25.6 in van_der_vaart_asymptotic_1998). In the two-sided case, the result is, to our knowledge, new.
Consider now the case in which the first $d_{g}-d_{\gamma}$ moments are known to be valid, with the potential misspecification for the remaining $d_{\gamma}$ moments unrestricted. Then $\mathcal{C}=\{(0', \gamma')'\colon \gamma\in\mathbb{R}^{d_{\gamma}}\}$ corresponds to a linear subspace with $B$ given by the last $d_{\gamma}$ columns of the identity matrix, and $B_{\perp}$ given by the first $d_{g}-d_{\gamma}$ columns. Optimal CIs in the limiting experiment therefore use the estimator $k_{LS, B}'Y$, which is the GLS estimator based only on the observations with no misspecification.
The one-sided CI based on $k_{LS, B}'Y$ achieves the shortest excess length, simultaneously for all quantiles and all possible values of the parameter $\theta$. The two-sided CI $k_{LS, B}'Y\pm z_{1-\alpha/2}\sqrt{k_{LS, B}'\Sigma k_{LS, B}}$ is optimal in the same sense as the usual CI in (ref): it achieves minimax expected length, and its efficiency, relative to a CI that optimizes its length at a single $\theta^{*}$ and $\gamma=0$, is lower-bounded by $z_{1-\alpha}/z_{1-\alpha/2}$. (ref) formally translates the efficiency bound from the limiting model to the GMM model, so that the usual two-sided CI based on $h(\hat{\theta}_{W(B)})$ is asymptotically efficient in the same sense as the usual CI based on $h(\hat{\theta}_{\Sigma^{-1}})$ discussed in (ref) under correct specification. Just as with the results in (ref), this asymptotic result is, to our knowledge, new. The one-sided analog follows from the results in (ref). These results stand in sharp contrast to the results for estimation, where the MSE improvement at small values of $\gamma$ may be substantial.
An important consequence of these results is that asymptotically valid one-sided CIs based on shrinkage or model-selection procedures, such as one-sided versions of the CIs proposed in andrews_hybrid_2009, ditraglia_using_2016 or mccloskey_2017 must have worse excess length performance than the usual one-sided CI based on the GMM estimator $h(\hat{\theta}_{W(B)})$ that uses valid moments only. While it is possible to construct two-sided CIs that improve upon the usual CI based on $h(\hat{\theta}_{W(B)})$ at particular values of $\theta$ and $\gamma$, the scope for such improvement is smaller than the ratio of one- to two-sided critical values. Furthermore, any such improvement must come at the expense of worse performance at other points in the parameter space.\footnote{Consistently with these results, in a simulation study considered in ditraglia_using_2016, the post-model selection CI that he proposes is shown to be wider on average than the usual CI around a GMM estimator that uses valid moments only.} Therefore, in order to tighten CIs based on valid moments only, it is necessary to make a priori restrictions on the potential misspecification of the remaining moments.
According to the results in (ref), one must place a priori bounds on the amount of misspecification in order to use misspecified moments. This leads us to the general case, where we place the local misspecification vector $c$ in some set $\mathcal{C}$ that is not necessarily a linear subspace. One can then form a CI centered at an estimate formed from these misspecified moments using the methods in (ref). In the case where $\mathcal{C}$ is convex and centrosymmetric, (ref) shows that this CI is near optimal, in the sense that no other CI can improve upon it by more than a factor of $\kappa_{*}$, even in the favorable case of correct specification. Since the width of the CI is asymptotically constant under local parameter sequences $\theta_n\to\theta^*$ and sufficiently regular probability distributions $P_n\to P_0$ (for example, $P_n\to P_0$ along submodels satisfying (ref)), this also shows that the CI is near optimal in a local minimax sense. In the general case, (ref), as well as the analogous results for one-sided CIs in (ref) are, to our knowledge, new.
As we discuss in (ref), we recommend reporting results for a range of sets $\mathcal{C}(M)$ indexed by a scalar $M$ that bounds the magnitude of misspecification. One may instead wish to report a single CI based on a data-driven estimate of $M$, for example, by using a first-stage $J$ test to assess plausible magnitudes of misspecification. Formally, one would seek a CI that is valid over $\mathcal{C}(\overline{M})$ while improving length when in fact $\norm{\gamma}\ll \overline{M}$, where $\overline{M}$ is some initial conservative bound. When $\mathcal{C}$ is convex and centrosymmetric, (ref) shows that the scope for such improvements is limited: the average length of any such CI cannot be much smaller than the CI that uses the most conservative choice $\overline{M}$, even when $c=0$. The impossibility of choosing $M$ based on the data is related to the impossibility of using specification tests to form an upper bound for $M$. On the other hand, it is possible to obtain a lower bound for $M$ using such tests. We develop lower CIs for $M$ in (ref).
andrews_informativeness_2018 have shown that defining misspecification in terms of the magnitude of any divergence in the cressie_multinomial_1984 family leads to a set $\mathcal{C}$ that asymptotically takes the form $\mathcal{C}=\{\Sigma^{1/2}\gamma:\|\gamma\|_2\le M\}=\{c\colon c'\Sigma^{-1}c\le M^2\}$. The Cressie-Read family includes the Hellinger distance used by kitamura_robustness_2013, who consider minimax point estimation among estimators satisfying certain regularity conditions. Since this set $\mathcal{C}$ takes the form discussed in (ref) with $p=2$ and $B=\Sigma^{1/2}$, it follows from the discussion in (ref) that the optimal sensitivity corresponds to the GMM estimator with weighting matrix $(\lambda BB'+\Sigma)^{-1}=(\lambda+1)^{-1}\Sigma^{-1}$. Since this is proportional to the weighting matrix $\Sigma^{-1}$ that is optimal under correct specification, we obtain the same optimal sensitivity $k_{LS,0}'=-H(\Gamma' \Sigma^{-1}\Gamma)^{-1}\Gamma'\Sigma^{-1}$ as in the correctly specified case discussed in (ref). As we show in (ref), this form of $\mathcal{C}$ leads to a closed form solution for the efficiency bound $\kappa_*$.
The results above imply that any estimator with sensitivity $k_{LS,0}$ is near-optimal for CI construction. In line with these results, the estimator in kitamura_robustness_2013 has sensitivity $k_{LS,0}$. Thus, the usual GMM estimator $h(\hat{\theta}_{\Sigma^{-1}})$ and the estimator in kitamura_robustness_2013 are both near-optimal for CI construction, even if one allows for arbitrary CIs that are not necessarily centered at estimators that satisfy the regularity conditions in kitamura_robustness_2013. Also, because they have the same sensitivity, under this form of misspecification, the usual GMM estimator $h(\hat{\theta}_{\Sigma^{-1}})$ and the estimator in kitamura_robustness_2013 have the same local asymptotic minimax properties.
If the set $\mathcal{C}$ is convex but asymmetric (such as when $\mathcal{C}$ includes bounds on a norm as well as sign restrictions, or when $\mathcal{C}$ includes equality and sign restrictions, as in moon_estimation_2009), one can still apply bounds from ArKo18optimal to the limiting model described in (ref). Our general asymptotic efficiency bounds in (ref) translate these results to the locally misspecified GMM model so long as $\mathcal{C}$ is convex. Since the negative implications for efficiency improvements under correct specification use centrosymmetry of $\mathcal{C}$, introducing asymmetric restrictions, such as sign restrictions, is one possible way of getting efficiency improvements at some smaller set $\mathcal{D}\subseteq\mathcal{C}$ while maintaining coverage over $\mathcal{C}$. We derive efficiency bounds and optimal CIs for this problem in (ref). Interestingly, the scope for efficiency improvements can be different for one- and two-sided CIs, and can depend on the direction of the CI in this case. To get some intuition for this, note that, in the instrumental variables model with a single instrument and single endogenous regressor, sign restrictions on the covariance of an instrument with the error term can be used to sign the direction of the bias of the instrumental variables estimator, which is useful for forming a one-sided CI only in one direction.
Finally, while we focus on restrictions on $c$, one can also incorporate local restrictions on $\theta$. Our general results in (ref) give efficiency bounds that cover this case. Similar to the discussion above, these results have implications for using prior information about $\theta$ to determine the amount of misspecification, or to shrink the width of a CI directly. In particular, while it is possible to use prior information on $\theta$ (say, an upper bound on $\|\theta\|$ for some norm $\|\cdot\|$) to shrink the width of the CI, the width of the CI and the estimator around which it is centered must depend on the a priori upper bounds on the magnitude of $\theta$ and $c$ when this prior information takes the form of a convex, centrosymmetric set for $(\theta', c')'$. This rules out, for example, choosing the moments based on whether the resulting estimate for $\theta$ is in a plausible range.
This section describes particular applications of our approach, along with a discussion of implementation details appropriate to each application.
The single equation linear instrumental variables (IV) model is given by
where, in the correctly specified case, $E\varepsilon_{i}z_{i}=E(y_{i}-x_i'\theta_0)z_{i}=0$, with $z_{i}$ a $d_{g}$-vector of instruments. This is an instance of a GMM model with $g(\theta)=E(y_{i}-x_i'\theta)z_{i}$ and $\hat{g}(\theta)=\frac{1}{n}\sum_{i=1}^{n}z_{i}(y_{i}-x_{i}'\theta)$.
One common reason for misspecification in this model is that the instruments do not satisfy the exclusion restriction, because they appear directly in the structural equation (ref), so that $\varepsilon_{i}=z_{Ii}'\gamma/\sqrt{n}+\eta_{i}$, where $E[z_{i}\eta_{i}]=0$, and $z_{Ii}$ corresponds to a subset $I$ of the instruments, the validity of which one is worried about. This form of misspecification has previously been considered in a number of papers, including hahn_iv_2005, conley_plausibly_2012, and andrews_measuring_2017, among others. Bounding the norm of $\gamma$ using some norm $\norm{\cdot}$ then leads to the set given in (ref), with $B=E[z_{i}z_{Ii}']$.
Although the matrix $B$ is unknown, for the purposes of estimating the optimal sensitivity and constructing asymptotically valid CIs, it can be replaced by the sample analog $\hat{B}=n^{-1}\sum_{i=1}^{n}z_{i}z_{Ii}'$. This does not affect the asymptotic validity or coverage properties of the resulting CI\@. Under this setup, the parameter $M$ bounds that magnitude of $\gamma$, the direct effect of the instruments on the outcome. Therefore, the appropriate choice of $M$ will depend on the plausible magnitude of these direct effects---see, for example, conley_plausibly_2012 for examples and a discussion.
The linearity of the moment condition leads to simplifications in our implementation in (ref). In Step (ref), as the initial estimator, one can use the two-stage least squares (2SLS) estimator
This leads to the estimates $\hat{\Gamma}=-\frac{1}{n}\sum_{i=1}^{n}z_{i}x_i'$ and $\hat{\Sigma}=\frac{1}{n}\sum_{i=1}^{n} (y_{i}-x_i'\hat\theta_{\text{initial}})^2z_{i}z_{i}'$. Alternatively, if we assume homoskedasticity, we can use the estimator $\hat\Sigma_{H}=\frac{1}{n}\sum_{i=1}^{n}(y_{i}-x_i'\hat\theta_{\text{initial}})^2\cdot \frac{1}{n}\sum_{i=1}^{n}z_{i}z_{i}'$. In the correctly specified case, the 2SLS estimator is only optimal under homoskedasticity, while the GMM estimator with weighting matrix $\hat\Sigma^{-1}$ is optimal in general. Due to concerns with finite sample performance, however, it is common to use the 2SLS estimator along with standard errors based on a robust variance estimate, even when heteroskedasticity is suspected. Mirroring this practice, one can use $\hat{\Sigma}_{H}$ when forming the optimal sensitivity $\hat{k}_{\lambda^{*}_{M}}$ and worst-case bias in Step (ref), but use the robust variance estimate $\hat{\Sigma}$ in Step (ref) when forming the final CI in (ref). The CI will then be optimal under homoskedasticity, but it will remain valid under heteroskedasticity, just like the usual CI based on 2SLS with robust standard errors in the correctly specified case.
If the parameter of interest linear in $\theta$, $h(\theta)=H\theta$, then the one-step estimator $\hat{h}_{\lambda^{*}_{M}}$ in Step (ref) does not depend on the choice of the initial estimator (except possibly through the estimate of $\Sigma$ when forming the desired sensitivity):
where the second line follows since the sensitivities $\hat{k}_{\lambda^{*}_{M}}$ satisfy $H=-\hat{k}_{\lambda^{*}_{M}}'\hat\Gamma= \hat{k}_{\lambda^{*}_{M}}'\frac{1}{n}\sum_{i=1}^{n}z_{i}x_i'$. Since the estimator $\hat{h}_{\lambda^{*}_{M}}$ is linear, the worst-case bias calculations are the same under global misspecification, when the magnitude $M$ of $\gamma$ in (ref) grows at the rate $\sqrt{n}$. By using variance estimates that are valid under global misspecification in place of the variance estimate $\hat{k}_{\lambda^{*}_{M}}'\hat\Sigma \hat{k}_{\lambda^{*}_{M}}$ in the CI construction, one can ensure that the resulting CI also remains valid under global misspecification. See (ref) for details.
Specializing to the case where $z_{i}=x_i$, the misspecified IV model of (ref) gives a misspecified linear regression model as a special case. This can be used to assess sensitivity of regression results to issues such as omitted variables bias. In particular, consider the linear regression model
where $x_i$ and $y_{i}$ are observed and $w_i^*$ is a (possibly unobserved) omitted variable. Correlation between $w_i^*$ and $x_i$ will lead to omitted variables bias in the OLS regression of $y_{i}$ on $x_i$. If $w_{i}^{*}$ is unobserved, then we obtain our framework by making the assumption $\sqrt{n}Ew_{i}^{*}x_{i}\in \mathcal{C}$, for some set $\mathcal{C}$, and letting $\hat{g}(\theta)=\frac{1}{n}\sum_{i=1}^{n}x_{i}(y_{i}-x_{i}'\theta)$. This setup can also cover choosing between different sets of control variables. Suppose that $w_{i}^{*}=w_{i}'\gamma$, where $w_{i}$ is a vector of observed control variables that the researcher is considering not including in the regression. If $\gamma$ is unrestricted, then by the results in (ref), the long regression of $y_{i}$ on both $x_{i}$ and $w_{i}$ yields nearly optimal CIs. If one is willing to restrict the magnitude of $\gamma$, it is possible to tighten these CIs, with the setting reducing to that in (ref), with $z_{i}=(x_{i}, w_{i}')'$. The same framework can be used to incorporate selection bias by defining $w_i^*$ to be the inverse Mills ratio term in the formula for $E[y_{i}\mid x_i, i\;\text{observed}]$ in heckman_sample_1979.
Our setup allows for misspecification in moment conditions arising from functional form misspecification. To apply our setup, one must relate this misspecification to the bounds $\mathcal{C}$ on the moment conditions at the true parameter value. One approach to bounding functional form misspecification is to use smoothness conditions from the nonparametric statistics literature, such as bounds on derivatives tsybakov09. Since these sets are typically convex (taking a convex combination of two functions that satisfy a given bound on a given derivative gives a function that also satisfies this bound), they typically lead to convex sets $\mathcal{C}$, so that our framework can be applied.
As a simple example, consider a nonparametric IV model with discrete covariates:
Suppose $x$ takes values in the finite set $\mathcal{X}=\{\tilde x_1,\ldots, \tilde x_{N_x}\}$ and $z_i$ takes values in the finite set $\mathcal{Z}=\{\tilde z_1,\ldots, \tilde z_{N_z}\}$. This setting was considered by freyberger_identification_2015, who place only nonparametric smoothness or shape restrictions on the unknown function $m$. To see the connection with our setting, we note that such restrictions can be interpreted as bounds on specification error from a parametric model. If one models these restrictions as local to a parametric family, one obtains our setting. In particular, let $m(x_i)=f(x_i, \theta_0)+n^{-1/2}r(x_i)$, $r\in\mathcal{R}$, where $\mathcal{R}$ is a nonparametric smoothness class. For example, if $x_i$ is univariate, we can let $f(x_i, \theta)=\theta_1+\theta_2x_i$ and define $\mathcal{R}$ to be the class of functions with $r(0)=r'(0)=r''(0)$ and second derivative bounded by some constant $M$. This is equivalent to placing the bound $n^{-1/2}M$ on the second derivative of $m(\cdot)$, which corresponds to a Hölder smoothness class. We can then map this to a misspecified GMM model, with the $j$th element of the moment function given by $g_j(x_i, y_i, \theta)=(y_i-f(x_i, \theta_0))I(z_i=\tilde z_j)$ and $j$th element of the misspecification vector $c$ given by $Er(x_i)I(z_i=\tilde z_j)=\sum_{\tilde x\in\mathcal{X}} r(\tilde x) P(x_i=\tilde{x}, z_i=\tilde z_j)$. Stacking these equations, we see that $c=B\gamma$ where $B$ is a matrix composed of the elements $P(x_i=\tilde x, z_i=\tilde z_j)$ and $\gamma=(r(\tilde x_1), \ldots, r(\tilde x_{N_x}))'$. As with the IV setting in (ref), $B$ is unknown, but can be replaced by a consistent estimate based on the sample analogue. So long as the set $\mathcal{R}$ is convex, we obtain convex restrictions on $\gamma$ and therefore $c$, so that our framework applies.
This example brings up an important point about the interpretation of $h(\theta)$. If the object of interest is a functional of $m(x)=f(x, \theta_0)+n^{-1/2}r(x)$, then we will need to allow the object of interest $h(\cdot)$ to depend on the misspecification vector directly, as well as on $\theta$. As discussed at the beginning of (ref), this falls into a mild extension of our framework. Alternatively, under a suitable parametrization of $f$ and $r$, it is often possible to define the object of interest to be function of $\theta$ alone. For example, if we are interested in the derivative $m'(x_0)$ at a particular point $x_0$ under a bound on the second derivative of $m(\cdot)$, we can let $f(x, \theta)=\theta_1+\theta_2 x$ and define $\mathcal{R}$ to be the class of functions with $r(x_0)=r'(x_0)=r''(x_0)=0$ and second derivative bounded by $M$. Then $m'(x_0)=\theta_2$.
Often, the average effect of a counterfactual policy on a particular subset of a population is of interest, and we would like to weaken the assumptions under which this effect is point-identified. We have available estimates $\hat{\tau}=(\hat\tau_{1}, \dotsc, \hat{\tau}_{m})'$ of the parameter $\tau$, with $\hat{g}(\theta)=\hat{\tau}-A\theta$, and $g(\theta)=\tau-A\theta$ for a known matrix $A$. We would like to extrapolate from these estimates to learn about the parameter of interest $h(\theta_{0})=H\theta_{0}$. The potential extrapolation bias is captured by the assumption that $g(\theta_{0})\in{\widetilde{\mathcal{C}}}_{n}$, some convex set.
Note that, because the moment condition is linear in the parameter of interest, asymptotic validity of our CIs does not require that the set ${\widetilde{\mathcal{C}}}_{n}$ takes the form ${\widetilde{\mathcal{C}}}_{n}=\mathcal{C}/\sqrt{n}$. CIs given in (ref) based on linear estimators of the form $\hat{h}=k'\hat{\tau}$ (such as minimum distance estimators), with $\mathcal{C}=\sqrt{n}{\widetilde{\mathcal{C}}}_{n}$ are valid under both local and global misspecification (i.e.\ under the assumption that the set ${\widetilde{\mathcal{C}}}_{n}$ is fixed as $n\to\infty$).
One example that falls into this setup are differences-in-differences designs when the parallel trends assumption is violated. Here there are $m$ time periods, with treatment taking place in period $T_{0}$. The $(m-T_{0})$-vector $\theta$ corresponds to a vector of dynamic treatment effects on the treated, $A\theta=(0', \theta')'$, and $g(\theta_{0})$ is a vector of trend differences between the treated and untreated, with $g(\theta_{0})=0$ if the parallel trends assumption holds. RaRo19 build on the framework in this paper to develop CIs in this setting.
Another example that has been of recent interest involves nonseparable models with endogeneity. Under conditions in imbens_identification_1994 and heckman_structural_2005, instrumental variables estimates $\hat{\tau}_{m}$ with different instruments are consistent for average treatment effects for different subpopulations. A recent literature kowalski_doing_2016,brinch_beyond_2017,mogstad_using_2017 has focused on using assumptions on treatment effect heterogeneity to extrapolate these estimates to other populations. Our framework applies if these assumptions amount to placing the differences between the estimated treatment effects and the effect of interest in a known convex set.
This section illustrates the confidence intervals developed in (ref) in an empirical application to automobile demand based on the data and model in blp95. We use the version of the model as implemented by andrews_measuring_2017, who calculate the asymptotic bias of the GMM estimator with weighting matrix $\Sigma^{-1}$ under local misspecification in this setting.\footnote{The dataset for this empirical application has been downloaded from the andrews_measuring_2017 replication files, available at \url{https://doi.org/10.7910/DVN/LLARSN}.}
In this model, the utility of consumer $i$ from purchasing a vehicle $j$, relative to the outside option, is given by a random-coefficient logit model $U_{ij}=\sum_{k=1}^{K}x_{jk}({\beta}_{k}+\sigma_{k}v_{ik})-\alpha p_{j}/y_{i}+\xi_{j}+\epsilon_{ij}$, where $p_{j}$ is the price of the vehicle, $x_{jk}$ the $k$th observed product characteristic, $\xi_{j}$ is an unobserved product characteristic, and $\epsilon_{ij}$ is has an i.i.d.\ extreme value distribution. The income of consumer $i$ is assumed to be log-normally distributed, $y_{i}=e^{m+\varsigma v_{i0}}$, where the mean $m$ and the variance $\varsigma$ of log-income are assumed to be known and set to equal to estimates from the Current Population Survey. The unobservables $v_{i}=(v_{i0}, \dotsc, v_{iK})$ are i.i.d.\ standard normal, while the distribution of the unobserved product characteristic $\xi_{j}$ is unrestricted.
The marginal cost $mc_{j}$ for producing vehicle $j$ is given by $\log(mc_{j})=w_{j}'\nu+\omega_{j}$, where $w_{j}$ are observable characteristics, and $\omega_{j}$ is an unobservable characteristic. The full vector of model parameters is given by $\theta=(\sigma', \alpha, \beta', \nu')'$. Given this vector, and given a vector of unobservable characteristics, one can compute the market shares implied by utility maximization, which can be inverted to yield the unobservable characteristic as a function of $\theta$, $\xi_{j}(\theta)$. One can similarly invert the unobserved cost component, writing it as a function of $\theta$, $\omega_{j}(\theta)$, under the assumption that firms set prices to maximize profits in a Bertrand-Nash equilibrium. Given a vector $z_{dj}$ of demand-side instruments, and a vector $z_{sj}$ of supply-side instruments, this yields the moment condition $g(\theta)=E[\hat{\gamma}(\theta)]$, where
The BLP data spans the period 1971 to 1990, and includes information on essentially all $n=999$ models sold during that period (for simplicity, we have suppressed the time dimension in the description above). There are 5 observable characteristics $x_{j}$: a constant, horsepower per 10 pounds of weight (HPWt), a dummy for whether air-conditioning is standard (Air), mileage per 10 dollars (MP\$) defined as MPG over average gas price in a given year, and car size (Size), defined as length times width. The vector $z_{dj}$ consists of $x_{j}$, plus the sum of $x_{j}$ across models other than $j$ produced by the same firm, and for rival firms. There are 6 cost variables $w_{j}$: a constant, log of HPWt, Air, log of MPG, log of Size, and a time trend. The vector $z_{sj}$ consists of these variables, MP\$, and the sums of $w_{j}$ for own-firm products other than $j$, and for rival firms. After excluding collinear instruments, this gives a total of $d_{g}=31$ instruments, 25 of which are excluded to identify $d_{\theta}=17$ model parameters. The parameter of interest is average markup, $h(\theta)=\frac{1}{n}\sum_{j}(p_{j}-mc_{j}(\theta))/p_{j}$.
One may worry that some of these instruments are invalid, because elements of $z_{dj}$ or $z_{sj}$ may appear directly in the utility or cost function with the coefficient on the $\ell$th element given by $\delta_{d\ell}\gamma_{d\ell}/\sqrt{n}$ or $\delta_{s\ell}\gamma_{s\ell}/\sqrt{n}$, respectively. Here $\delta_{d\ell}$ and $\delta_{s\ell}$ are scaling constants so that, given the sample size $n=999$ at hand, $\gamma_{d\ell}$ has the interpretation that the consumer willingness to pay for one standard deviation change in the $\ell$th demand-side instrument $z_{dj\ell}$ is $\gamma_{d\ell}\%$ of the average 1980 car price, and changing the $\ell$th supply-side instrument $z_{sj\ell}$ by one standard deviation changes the marginal cost by $\gamma_{s\ell}$% of the average car price. andrews_measuring_2017 use this scaling in their sensitivity analysis, and they discuss economic motivation for concerns about this form of misspecification. By way of comparison, the estimates of the parameters $\beta$ and $\nu$ in the utility and cost function imply that consumers are on average willing to pay between 2.2 and 10.0% of the average car price for a standard deviation change in one of the included car characteristics, and that a standard deviation change in the included cost characteristics changes the marginal cost by between 3.8 and 11.1% of the average car price. We therefore interpret specifications of the set $\mathcal{C}$ that allow for $\abs{\gamma_{s\ell}}\approx1\text{--}2$ (or $\abs{\gamma_{d\ell}}\approx1\text{--}2$) as allowing for moderate amounts of misspecification in the $\ell$th supply-side (or demand-side) instrument.
We follow the implementation in (ref). Given a set $I$ of potentially invalid instruments, we follow (ref) and consider sets $\mathcal{C}$ of the form (ref), with $\norm{\cdot}$ corresponding to an $\ell_{p}$ norm with $p\in\{2,\infty\}$, and $B=\tilde{B}_{I}\cdot \#I^{1/p}$, where $\tilde{B}_{I}$ is given by the columns of
and $\#{I}$ is the number of potentially invalid instruments. The scaling by $(\#{I})^{1/p}$ ensures that the vector $\gamma=M(1,\dotsc,1)'$ is always included in the set. andrews_measuring_2017 report the sensitivity of the usual GMM estimator under this form of misspecification, considering misspecification in each instrument individually (so that $I$ contains a single element), and setting $M=1$. However, if one is concerned about the validity of several instruments, it is natural to allow $I$ to contain all instruments the validity of which is questionable. In our analysis, we vary the set of potentially misspecified instruments. We also vary $M$ in order to assess the sensitivity of conclusions to different amounts of misspecification. As we will see below, different choices of $\mathcal{C}$ lead to different sensitivities for the optimal estimator, and using the optimal sensitivity can reduce the width of the CI substantially relative to CIs based on the usual GMM estimator.
We use the estimate $\hat{\theta}_{\text{initial}}$ that corresponds to the GMM estimator based on the weight matrix that's optimal under correct specification, as reported in andrews_measuring_2017, and the estimates $\hat{\Gamma}$, $\hat{H}$, and $\hat{\Sigma}$ are computed following Step (ref) of the implementation.
To illustrate that using the sensitivity that is optimal under local misspecification can yield substantially tighter CIs, (ref) plots the confidence intervals based on the optimal sensitivity, as well as those based on $\hat{\theta}_{\text{initial}}$ under different sets $I$ of potentially invalid instruments and $\ell_{2}$ constraints on $\gamma$. It is clear from the figure that using the optimal sensitivity yields substantially tighter confidence intervals, relative to simply adjusting the usual CI by using the critical value $\operatorname{cv}_{\alpha}(\cdot)$ to take into account the potential bias of $h(\hat{\theta}_{\text{initial}})$, by as much as a factor of 3.4. The intuitive reason for this is that by adjusting the sensitivity of the estimator, it is possible to substantially reduce its bias at little cost in terms of an increase in variance. Thus, for example, while the CI for the average markup based on the estimate $\hat{\theta}_{\text{initial}}$ is essentially too wide to be informative when the set of potentially invalid instruments corresponds to all excluded instruments, the CI based on the optimal sensitivity, $[46.0, 66.0]\%$, is still quite tight.
If a researcher is ex ante unsure what form of misspecification one should worry about, as a sensitivity check, it is useful to consider the effects of different forms of misspecification. In (ref), we plot the optimal confidence intervals for different subsets of invalid instruments, under both $\ell_{2}$ and $\ell_{\infty}$ norms for $\gamma$. Although the choice of norm matters when the number of potentially misspecified instruments is greater than one, the results are qualitatively similar. Comparing the results for different choices of the set of potentially invalid instruments suggests that allowing supply-side instruments to be invalid generally increases the average markup estimate, while allowing demand-side instruments to be invalid has the opposite effect.
As it may be ex ante unclear what magnitude of misspecification is reasonable to allow for, as discussed in (ref), it is useful to plot the optimal CI for multiple choices of $M$. We do this in (ref) for $p=2$, and we allow all excluded instruments to be potentially invalid. One can see that while the CI is unstable for values of $M$ smaller than about $0.4$, for larger values of $M$, the estimate is quite stable and equal to about 50%. Even at $M=2$, one rejects the hypothesis that the optimal markup is equal to the initial estimate $h(\hat{\theta}_{\text{initial}})=32.7\%$. This suggests that ignoring misspecification in the BLP model likely leads to a downward bias in the estimate of the average markup. At the same time, it is possible to obtain reasonably tight CIs for the average markup even under a moderate amount of misspecification.
The $J$-statistic for testing the hypothesis that all moments are correctly specified equals 426.7. Consequently, the hypothesis is rejected at the usual significance levels. Furthermore, it can be seen from (ref) that the CIs for “all excluded” (that allow all excluded instruments to be invalid at $M=1$), and “all excluded demand” (that assume validity of supply-side instruments) do not overlap. This implies that either the misspecification in the demand-side instruments must be greater than 1% of the average care price ($M=1$), or else the supply-side instruments must also be invalid. (ref) implements the specification test from (ref) that gives lower CI $[M_{\min}, \infty]$ for $M$. The results suggest that if one assumes only a subset of the instruments is invalid, the misspecification in the potentially invalid instruments must be quite large. For example, if we assume that all instruments are valid except potentially the demand-side instruments based on rival firms' product characteristics, then the misspecification in these instruments must be greater than $M=5.36$. If we allow all instruments to be invalid, then $M\geq 1.13$.
Finally, to illustrate the implication of (ref) that one cannot substantively improve upon the CIs that we construct, we calculate the efficiency bound $\kappa_{*}$ for these CIs in (ref). The table shows that the bound is at least as high as the efficiency bound for the usual CI under correct specification (given in (ref) and equal to 84.99% at $\alpha=0.05$). Thus, the asymptotic scope for improvement over the CIs reported in (ref) at particular values of $\theta$ and $c=0$ is even smaller than the scope for improvement over the usual CI at particular values of $\theta$ under correct specification.