EconBase
← Back to paper

Sensitivity Analysis using Approximate Moment Condition Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

100,317 characters · 22 sections · 85 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Sensitivity Analysis using Approximate Moment Condition Models

abstractWe consider inference in models defined by approximate moment conditions. We show that near-optimal confidence intervals (CIs) can be formed by taking a generalized method of moments (GMM) estimator, and adding and subtracting the standard error times a critical value that takes into account the potential bias from misspecification of the moment conditions. In order to optimize performance under potential misspecification, the weighting matrix for this GMM estimator takes into account this potential bias, and therefore differs from the one that is optimal under correct specification. To formally show the near-optimality of these CIs, we develop asymptotic efficiency bounds for inference in the locally misspecified GMM setting. These bounds may be of independent interest, due to their implications for the possibility of using moment selection procedures when conducting inference in moment condition models. We apply our methods in an empirical application to automobile demand, and show that adjusting the weighting matrix can shrink the CIs by a factor of 3 or more.

Introduction

Economic models are typically viewed as approximations of reality. However, conventional approaches to estimation and inference assume that a model holds exactly. In this paper, we weaken this assumption, and consider inference in a class of models characterized by moment conditions which are only required to hold in an approximate sense. The failure of the moment conditions to hold exactly may come from failure of exclusion restrictions (e.g.\ through omitted variable bias or because instruments enter the structural equation directly in an IV model), functional form misspecification, or other sources such as measurement error, or data contamination.

We assume that we have a model characterized by a set of population moment conditions $g(\theta)$. In the generalized method of moments (GMM) framework, for instance, $g(\theta)=E[g(w_{i}, \theta)]$, which can be estimated by the sample analog $\frac{1}{n}\sum_{i=1}^{n}g(w_{i}, \theta)$, based on the sample $\{w_{i}\}_{i=1}^{n}$. When evaluated at the true parameter value $\theta_{0}$, the population moment condition lies in a known set specified by the researcher,

equation*[equation* omitted — 68 chars of source]

The set $\mathcal{C}$ formalizes the way in which the moment conditions may fail, and it can then be varied as a form of sensitivity analysis, with $\mathcal{C}=\{0\}$ reducing to the correctly specified case.

We focus on local misspecification: the scaling by $\sqrt{n}$ implies that the specification error and the sampling error are of the same order of magnitude, and it ensures that the asymptotic approximation captures the fact that it may not be clear from the sample at hand whether the model is correctly specified. It also leads to increased tractability, allowing us to deliver a simple method for inference on a structural parameter of interest $h(\theta_{0})$, rather than a pseudo-true parameter. This tractability has made local misspecification a popular tool for sensitivity analysis in applied work, especially following the recent influential paper by andrews_measuring_2017.\footnote{For recent empirical examples using local sensitivity analysis, see gayle2019, or duflo2018.} As with any asymptotic device, our modeling of misspecification as local should not be taken to mean that we literally believe that the model would be closer to correct if we had more data. Rather, its usefulness should be judged by whether it yields accurate approximations to the finite-sample behavior of estimators and confidence intervals, which in our case requires that the set $\mathcal{C}/\sqrt{n}$ be small relative to sampling uncertainty.

We propose a simple method for constructing asymptotically valid confidence intervals (CIs) under this setup: one takes a standard estimator, such as the GMM estimator, and adds and subtracts its standard error times a critical value that takes into account the potential asymptotic bias of the estimator, in addition to its variance. A key insight of this paper is that because the CIs must be widened to take into account the potential bias, the optimal weighting matrix for the correctly specified case (the inverse of the variance matrix of the moments) is generally no longer optimal under local misspecification. Rather, the optimal weighting matrix takes into account potential misspecification in the moments in addition to the variance of their estimates: it places less weight on moments that are allowed to be further from zero according the researcher's specification of the set $\mathcal{C}$. We also show that an analogous result holds for other performance criteria, such as estimation under the mean-squared error: the optimal weighting matrix again trades off the potential misspecification of the moments against their precision, although the optimal tradeoff is different.

To illustrate the practical importance of this result, we apply our methods to form misspecification-robust CIs in an empirical model of automobile demand based on blp95. We consider sets $\mathcal{C}$ motivated by the forms of local misspecification considered in andrews_measuring_2017, who calculate the asymptotic bias of the usual GMM estimator in this model. We find that adjusting the weighting matrix to account for potential misspecification substantially reduces the potential bias of the estimator and, as a result, leads to large efficiency improvements of the optimal CI relative to a CI based on the GMM estimator that is optimal under correct specification: it shrinks the CI by up to a factor of $3$ or more in our main specifications. As a result, we obtain informative CIs in this model even under moderate amounts of misspecification.

We show that the CIs we propose are near-optimal when the set $\mathcal{C}$ is convex and centrosymmetric ($c\in\mathcal{C}$ implies $-c\in\mathcal{C}$). To this end, we argue that the relevant “limiting experiment” for the locally misspecified GMM model is isomorphic to an approximately linear model of SaYl78, which falls under a general framework studied by, among others, donoho94, CaLo04 and ArKo18optimal. We derive asymptotic efficiency bounds for CIs in the locally misspecified GMM model that formally translate bounds from the approximately linear limiting experiment to the locally misspecified GMM setting. In particular, these bounds imply that the scope for improvement over our CIs by optimizing expected length at a particular value of $\theta_{0}$ and $c=0$ (while still maintaining coverage over the whole parameter space for $\theta$ and $\mathcal{C}$) is limited, even if one optimizes expected length at the true values of $\theta_{0}$ and $c$.

These efficiency bounds address an important criticism of our CIs: they require a priori specification of the set $\mathcal{C}$ that defines misspecification, including both the magnitude of misspecification and which moments are misspecified. In particular, one cannot substantively improve upon our CI by, say, trying to use data-driven methods that gauge misspecification magnitude or try to determine which moments are misspecified. These bounds have implications for procedures proposed by andrews_hybrid_2009, ditraglia_using_2016 and mccloskey_2017, who consider the case where some moments are known to be correctly specified and no a priori bound is placed on the magnitude of misspecification of the remaining moments. As we discuss in (ref), in this case our CI reduces to the usual CI based on the $k_1$ correctly specified moments, and our efficiency bounds show that CIs proposed in these papers cannot substantively improve upon it.

Because we cannot use the data to determine the magnitude $M$ of the set $\mathcal{C}$, we recommend plotting our CIs as a function of the potential misspecification magnitude $M$, or reporting the smallest value of $M$ for which a particular finding breaks down. Such sensitivity analysis is easy to conduct under out proposed implementation. In particular, we show that, when the set $\mathcal{C}$ is characterized by $\ell_{p}$ constraints, the class of weightings that trace out the optimal bias-variance tradeoff as a function of how much relative weight we put on the bias can be easily computed by recasting the problem as a penalized regression problem. By exploiting this analogy, we develop a simple algorithm for computing this class under $\ell_{\infty}$ constraints that is similar to the LASSO/LAR algorithm ehjt04,rosset_piecewise_2007; under $\ell_{2}$ constraints, the solution admits a closed form.\footnote{An R package implementing our CIs under $\ell_{p}$ constraints is available at \url{https://github.com/kolesarm/GMMSensitivity}.} Furthermore, as we discuss in (ref), this class of weightings is entirely determined by the shape of $\mathcal{C}$; its magnitude $M$ only determines the optimal relative weight we should put on the bias. Thus, tracing out the optimal weighting as a function of $M$ can be done at essentially no additional computational cost. Furthermore, to avoid having to reoptimize the objective function with respect to the new weighting matrix, one can also form the CIs by adding and subtracting our critical value from a one-step estimator newey_large_1994 based on any initial estimate that is $\sqrt{n}$-consistent under correct specification. We illustrate this approach in our empirical application in (ref).

Our paper is related to several strands of literature. Our efficiency results are related to those in chamberlain_asymptotic_1987 for point estimation in the correctly specified setting (see also hansen85) and, more broadly, semiparametric efficiency theory in correctly specified settings van_der_vaart_asymptotic_1998. As we discuss in (ref), some of our efficiency results are novel even in the correctly specified case, and may be of independent interest. kitamura_robustness_2013 consider efficiency of point estimators satisfying certain regularity conditions when the misspecification is bounded by the Hellinger distance. As we discuss in more detail in (ref), our results imply that under this form of misspecification, the optimal weighting matrix remains the same as under correct specification; both the usual GMM estimator and the estimator proposed by kitamura_robustness_2013 can thus be used to form near-optimal CIs, and both estimators have the same local asymptotic minimax properties.

Local misspecification has been used in a number of papers, which include, among others, newey_generalized_1985, berkowitz_validity_2012, conley_plausibly_2012, guggenberger_asymptotic_2012, and bugni_inference_2018, and has antecedents in the literature on robust statistics huber_robust_2011. andrews_measuring_2017 consider this setting and note that asymptotic bias of a regular estimator can be calculated using influence function weights, which they call the sensitivity, and show how such calculations can be used for sensitivity analysis in applications (see also extensions of these ideas in andrews_informativeness_2018 and mukhin_sensitivity_2018). Our results imply that, if one is interested in inference, conclusions of such sensitivity analysis may be substantially sharpened by using the misspecification-optimal weighting matrix, or, equivalently, the misspecification-optimal sensitivity. In independent work, bonhomme_minimizing_2018 provide a framework for estimation and inference in misspecified likelihood models when the misspecification set $\mathcal{C}$ is defined with respect to a larger class of models using statistical notions of distance. Our focus is on overidentified moment condition models, as in andrews_measuring_2017, and we are agnostic about how $\mathcal{C}$ is determined. The proposal to use estimators that optimize an asymptotic bias-variance tradeoff using the influence function is common to both papers. The efficiency bounds in (ref) are unique to the present paper.

The rest of this paper is organized as follows. (ref) presents our misspecification robust CIs. (ref) gives step-by-step instructions for computing our CIs, along with discussion of other practical issues. (ref) presents efficiency bounds for CIs in locally misspecified models; it can be skipped by readers interested only in implementing the methods. (ref) discusses applications to particular moment condition models. (ref) presents an empirical application. Additional results and proofs are collected in appendices and an online supplement.

Misspecification-robust CIs

We have a model that maps a vector of parameters $\theta\in\Theta\subseteq \mathbb{R}^{d_{\theta}}$ to a $d_{g}$-dimensional population moment condition $g(\theta)$ that restricts the distribution of the observed data $\{w_{i}\}_{i=1}^{n}$. We allow the moment condition model to be locally misspecified, so that at the true value $\theta_{0}$, the population moment condition is not necessarily zero, but instead lies in a $\sqrt{n}$-neighborhood of $0$:

equation[equation omitted — 100 chars of source]

where $\mathcal{C}\subseteq\mathbb{R}^{d_{g}}$ is a known set. The set $\mathcal{C}$ may allow for misspecification in potentially all moment conditions; we do not require that some elements of $c$ are zero. Our goal is to construct a CI for a scalar $h(\theta_0)$, where $h\colon \mathbb{R}^{d_\theta}\to\mathbb{R}$ is a known function. For example, if we are interested in one of the elements $\theta_{j}$ of $\theta$, we would take $h(\theta)=\theta_{j}$. More generally, the function $h$ will be nonlinear, as is, for example, generally the case when $\theta$ is a vector of supply or demand parameters, and $h(\theta)$ is an elasticity, or some counterfactual.

This setup allows (but does not require) both $\theta_{0}$ and $h(\theta_{0})$ to have the same interpretation as in the correctly specified case, so that our CIs may still be interpreted as CIs for the structural parameter, elasticity, or counterfactual of interest. For this interpretation, one typically needs to rule out forms of misspecification that affect the mapping $\theta\mapsto h(\theta)$. While we do not formally consider cases in which this mapping itself is misspecified, such cases are covered under a mild generalization of our framework, in which $h$ is a function of both $\theta$ and $c$.

Note that the interpretation of $h(\theta_{0})$, and the conceptual framework defining $\theta_{0}$ is not affected by our modeling of the misspecification as local: given a set ${\widetilde{\mathcal{C}}}_{n}=\mathcal{C}/\sqrt{n}$, the moment conditions $g(\theta_{0})\in{\widetilde{\mathcal{C}}}_{n}$ describe the restrictions that the data generating process and the researcher's modeling assumptions place on $\theta_0$.\footnote{Formally, for a given sample size $n$, $\theta_{0}$ may be set identified, and the identified set under a distribution $P$ is defined as the set of parameters $\theta_0$ that satisfy the moment conditions $E_{P}g(w_i, \theta_0)\in {\widetilde{\mathcal{C}}}_{n}$ where $E_P$ denotes expectation under $P$. We construct CIs that cover $h(\theta_0)$ for points $\theta_0$ in the identified set im04. See (ref) and (ref) for formal definitions of coverage and optimality of our CIs.} The plausibility of these restrictions is evaluated for a given sample size at hand; it doesn't depend on assumptions about how ${\widetilde{\mathcal{C}}}_{n}$ changes with $n$. While we focus on sequences ${\widetilde{\mathcal{C}}}_n=\mathcal{C}/\sqrt{n}$, we discuss in (ref) how our insights can be used to construct CIs that are valid global misspecification, when ${\widetilde{\mathcal{C}}}_{n}$ is fixed with $n$.

To formalize the notion of asymptotic validity and efficiency of CIs, we will need to allow the true parameter value $\theta_0$ as well as the vector $c$ and the data generating process (and hence the map $\theta\mapsto g(\theta)$) to vary with the sample size. For clarity of exposition, we focus here on the case in which these parameters are fixed. See (ref) and (ref) for the general case. Under some forms of misspecification, such as functional form misspecification, there may be additional higher-order terms on the right-hand side of (ref); our results remain unchanged if this is the case. Again, for clarity of exposition, we focus on the case in which (ref) holds exactly.

We assume that the sample moment condition $\hat{g}(\theta)$, constructed using the data $\{w_{i}\}_{i=1}^{n}$, satisfies

equation[equation omitted — 120 chars of source]

where $\stackrel{d}{\to}$ denotes convergence in distribution as $n\to\infty$. In the GMM model, the population and sample moment conditions are given by $g(\theta)=E[g(w_{i}, \theta)]$ and $\hat{g}(\theta)=\frac{1}{n}\sum_{i=1}^{n}g(w_{i}, \theta)$, respectively, where $g(\cdot, \cdot)$ is a known function. However, to cover other minimum distance problems, we do not require that the moment conditions necessarily take this form. We further assume that the moment condition is smooth enough so that

equation[equation omitted — 193 chars of source]

where $\Gamma$ is the $d_{g}\times d_{\theta}$ derivative matrix of $g$ at $\theta_{0}$. Conditions (ref) and (ref) are standard regularity conditions in the literature on linear and nonlinear estimating equations; see newey_large_1994 for primitive conditions. Finally, we also assume that $h$ is continuously differentiable with the $1\times d_{\theta}$ derivative matrix at $\theta_{0}$ given by $H$.

CIs based on asymptotically linear estimators

Under correct specification, when $\mathcal{C}=\{0\}$, standard estimators $\hat{h}$ of $h(\theta)$ are asymptotically linear in $\hat{g}(\theta_{0})$. This will typically extend to our locally misspecified case, so that for some vector $k\in\mathbb{R}^{d_g}$,

equation[equation omitted — 168 chars of source]

where the convergence in distribution follows by (ref) and (ref). If in addition, the estimator is regular (so that equality in (ref) holds uniformly for $\theta$ in a $\sqrt{n}$-neighborhood of $\theta_{0}$), then $k$ will satisfy (see e.g. Section 2 in newey_semiparametric_1990)

equation[equation omitted — 49 chars of source]

For example, in a GMM model, if we take $\hat h=h(\hat{\theta}_{W})$ where

equation[equation omitted — 114 chars of source]

is the GMM estimator with weighting matrix $W$, (ref) will hold with $k'=-H(\Gamma' W\Gamma)^{-1}\Gamma' W$ newey_generalized_1985. Because the vector $k$ determines the local asymptotic bias of the estimator, we follow andrews_measuring_2017, and refer to $k$ as the sensitivity of $\hat{h}$.

We now show how to construct misspecification-robust CIs based on an asymptotically linear estimator $\hat{h}$ with a given sensitivity $k$. In (ref), we show how to choose this sensitivity optimally, to achieve the shortest CI among those based on regular asymptotically linear estimators. In (ref), we will show that, under this choice of $k$, the resulting CI is (near) optimal not only within the class of CIs based on regular asymptotically linear estimators, but among all CIs that satisfy the asymptotic coverage requirement.

Let $\hat{k}$ and $\hat{\Sigma}$ be consistent estimates of $k$ and $\Sigma$. Then by Slutsky's theorem,

equation*[equation* omitted — 164 chars of source]

Under correct specification, the right-hand side corresponds to a standard normal distribution, and we can form a CI with asymptotic coverage $100\cdot (1-\alpha)\%$ as $\hat{h}\pm z_{1-\alpha/2}\sqrt{\hat{k}'\hat{\Sigma}\hat{k}/n}$, where $z_{1-\alpha/2}$ is the $1-\alpha/2$ quantile of a $\mathcal{N}(0,1)$ distribution; this is the usual Wald CI\@.

When we allow for misspecification, the Wald CI will no longer be valid. However, note that the asymptotic bias ${k'c}/{\sqrt{k'\Sigma k}}$ is bounded in absolute value by ${\operatorname{\overline{bias}}_{\mathcal{C}}(k)}/{\sqrt{k'\Sigma k}}$ where $\operatorname{\overline{bias}}_{\mathcal{C}}(k)\equiv \sup_{c\in\mathcal{C}}\abs{k'c}$. Therefore, given $c$, the $z$-statistic in the preceding display is asymptotically $\mathcal{N}(t,1)$ where $\abs{t}\le \operatorname{\overline{bias}}_{\mathcal{C}}(k)/\sqrt{k'\Sigma k}$. This leads to the CI

equation[equation omitted — 229 chars of source]

where $\operatorname{cv}_\alpha(\overline{t})$ is the $1-\alpha$ quantile of $\abs{Z}$, with $Z\sim \mathcal{N}(\overline{t},1)$. In particular, $\operatorname{cv}_{\alpha}(0)=z_{1-\alpha/2}$, so that in the correctly specified case, (ref) reduces to the usual Wald CI\@. As we discuss in (ref), in the limiting experiment, this CI becomes equivalent to the fixed-length CI proposed by donoho94.

To form a one-sided CI based on an estimator $\hat{h}$ with sensitivity $k$, one can simply subtract its maximum bias, in addition to the standard error:

equation[equation omitted — 173 chars of source]

One could also form a valid two-sided CI by adding and subtracting the worst-case bias $\operatorname{\overline{bias}}_{\mathcal{C}}(\hat k)/\sqrt{n}$ from $\hat{h}$, in addition to adding and subtracting $z_{1-\alpha/2}\sqrt{\hat{k}'\hat\Sigma\hat k/n}$; however, since $\hat{h}$ cannot simultaneously have a large positive and a large negative bias, such CI will be conservative, and longer than the CI in (ref).

Optimal CIs

The asymptotic length of the CI in (ref) is given by

equation[equation omitted — 188 chars of source]

To attain the shortest possible CI, we therefore need to use an estimator with sensitivity that minimizes this expression. We restrict attention to asymptotically linear estimators that are regular, so that we need to minimize (ref) subject to (ref). The CI length in (ref) depends on $\theta$ only through $\Sigma$. Furthermore, it depends on the sensitivity only through the maximum bias, $\operatorname{\overline{bias}}_{\mathcal{C}}(k)$, and the variance $k'\Sigma k$. Therefore, rather than minimizing (ref) directly over all sensitivities $k$, one can first minimize the variance subject to a bound $\overline{B}$ on the worst-case bias,

equation[equation omitted — 173 chars of source]

and then vary the bound $\overline{B}$ to find the bias-variance trade-off that leads to the shortest CI\@. In our implementation in (ref), we focus on the case where $\mathcal{C}$ is characterized by $\ell_p$ constraints, in which case a closed-form expression for the worst-case bias $\sup_{c\in\mathcal{C}}\abs{k'c}$ is available, and it is computationally trivial to solve (ref) directly or in Lagrange multiplier form. In general, when the set $\mathcal{C}$ is convex, one can reformulate (ref) as a convex optimization problem, leading to a computationally tractable solution (see (ref)). One can also use (ref) to determine the optimal sensitivity for constructing one-sided CIs, if we use quantiles of excess length as the criterion for choosing a CI\@. We provide details in (ref).

Once the optimal sensitivity has been determined, we can implement an estimator with this sensitivity as a one-step estimator. In particular, let $\hat\theta_{\text{initial}}$ be an initial $\sqrt{n}$-consistent estimator of $\theta_0$, let $\hat{k}= k+o_{P}(1)$ be a consistent estimator of the desired sensitivity $k$. Then the one-step estimator

equation*[equation* omitted — 99 chars of source]

will have the desired sensitivity. This follows from the Taylor expansion

equation*[equation* omitted — 313 chars of source]

where the second line follows from (ref). It then follows from (ref) that the first term converges in probability to zero, and $\hat{h}$ satisfies (ref).

Practical implementation

We now give step-by-step instructions for computing our CI\@. To make it easy to determine the sensitivity of the CI to the magnitude of misspecification, we consider sets of the form $\mathcal{C}=\mathcal{C}(M)=\{Mc\colon c\in\mathcal{C}(1)\}$, where the scalar $M$ measures the magnitude of misspecification. We discuss the exact specification of the set $\mathcal{C}(M)$ in (ref) below.

The fact that $M$ simply scales the potential magnitude of misspecification leads to a simplification when tracing out the optimal CI as a function of $M$. In particular, let $\{k_{\lambda}\}_{\lambda\geq 0}$ be the bias-variance optimizing class of sensitivities that traces out the solutions to (ref) as we vary the bound $\overline{B}$ when $\mathcal{C}=\mathcal{C}(1)$. The index $\lambda$ determines the relative weight on the bias; it can correspond to the Lagrange multiplier in a Lagrangian formulation of (ref), or we can simply take $\lambda=\overline B$ if we are solving (ref) directly. Let $\overline{B}_{\lambda}=\operatorname{\overline{bias}}_{\mathcal{C}(1)}(k_\lambda)$. It then follows by a change-of-variables argument that $\operatorname{\overline{bias}}_{\mathcal{C}(M)}(k_\lambda)= M\overline B_\lambda$, and that $k_\lambda$ minimizes the asymptotic variance subject to this bound on worst-case bias over $\mathcal{C}(M)$. Thus, $\{k_\lambda\}_{\lambda\geq 0}$ is also a bias-variance optimizing class of sensitivities for $\mathcal{C}(M)$. We therefore only need to compute the class $\{k_{\lambda}\}_{\lambda\geq 0}$ only once, even when a range of values $M$ is considered.

With this simplification, we can construct CIs for a range of values of $M$ as follows:

enumerate• Obtain an initial estimate $\hat\theta_{\text{initial}}$ and estimates $\hat H$, $\hat \Gamma$ and $\hat \Sigma$ of $H$, $\Gamma$ and $\Sigma$. In particular, for the GMM model, when $\hat{g}(\theta) =\frac{1}{n}\sum_{i=1}^{n} g(w_i, \theta)$, we can take $\hat\theta_{\text{initial}}$ to be the GMM estimator $\hat{\theta}_{W}=\operatorname*{argmin}_{\theta} \hat g(\theta)'W\hat g(\theta)$ for some weight matrix $W$. The remaining objects are the usual quantities used to estimate the asymptotic variance of this estimator: $\hat{\Sigma}= \frac{1}{n}\sum_{i=1}^{n} g(w_i, \hat\theta_{\text{initial}}) g(w_i, \hat\theta_{\text{initial}})'$ (or, in the case of dependent observations, an autocorrelation robust version of this estimate), $\hat \Gamma=\frac{d}{d\theta'} \hat{g}(\theta) \big|_{\theta=\hat\theta_{\text{initial}}}$ (or, if $g(\theta, w)$ is nonsmooth, a numerical derivative as in hong_extremum_2015, or Section 7.3 of newey_large_1994) and $\hat{H}=\frac{d}{d\theta'}h(\theta)\big|_{\theta=\hat\theta_{\text{initial}}}$. • Compute the bias-variance optimizing class $\{\hat{k}_{\lambda}\}_{\lambda\geq 0}$ that solves (ref) with $\mathcal{C}=\mathcal{C}(1)$ and with $\hat\Sigma$ in place of $\Sigma$. Algorithms and closed-form solutions for computing $\{\hat{k}_{\lambda}\}_{\lambda \geq 0}$ for particular choices of $\mathcal{C}(M)$ are discussed in (ref). Let $\overline B_\lambda=\sup_{c\in\mathcal{C}(1)} |\hat k_\lambda'c|$. For each $M$, let $\lambda^*_{M}$ minimize the CI length\footnote{The critical value $\operatorname{cv}_{\alpha}(b)$ can easily be computed in statistical software as the square root of the $1-\alpha$ quantile of a non-central $\chi^{2}$ distribution with $1$ degree of freedom and non-centrality parameter $b^{2}$.} $2\operatorname{cv}_\alpha(M \overline B_\lambda/ \sqrt{\hat{k}_{\lambda}' \hat\Sigma \hat k_{\lambda}})\cdot \sqrt{\hat{k}_{\lambda}' \hat\Sigma \hat{k}_{\lambda}}$ over $\lambda$. • For each $M$, construct the one-step estimator $\hat{h}_{\lambda^{*}_{M}} = h(\hat\theta_{\text{initial}})+\hat{k}_{\lambda^{*}_{M}}' \hat{g}(\hat\theta_{\text{initial}})$, and report the misspecification-robust CI under $\mathcal{C}(M)$ \begin{equation} \hat{h}_{\lambda_{M}^{*}}\pm \operatorname{cv}_{\alpha} \left(M \overline B_{\lambda^{*}_{M}}/ \sqrt{\hat{k}_{\lambda_{M}^*}' \hat\Sigma \hat k_{\lambda_{M}^*}}\right)\cdot \sqrt{\hat{k}_{\lambda_{M}^*}' \hat\Sigma \hat k_{\lambda^*} / n}. \end{equation}
remark[Choice of $\mathcal{C}(M)$] A simple and flexible way of forming the set $\mathcal{C}$ is to take \begin{equation} \mathcal{C}=\mathcal{C}(M)=\{B\gamma\colon \norm{\gamma}\le M\}, \end{equation} where $B$ is a $d_g\times d_\gamma$ matrix and $\norm{\cdot}$ is some norm. The matrix $B$ can be used to standardize moments, account for their correlations, or to pick out which moments are believed to be misspecified. For instance, setting $B$ to the last $d_{\gamma}$ columns of the $d_{g}\times d_{g}$ identity matrix allows for misspecification in the last $d_{\gamma}$ moments, while maintaining that the first $d_{g}-d_{\gamma}$ moments are valid. In light of our result in (ref) that it is not possible to determine the set $\mathcal{C}$ in a data-driven way, the normalizing matrix $B$ and the baseline misspecification magnitude $M$ used should be chosen to reflect application-specific arguments about which forms of misspecification are plausible; we can then vary $M$ over other plausible choices as a form of sensitivity analysis. We illustrate this in the context of our application in (ref), and we refer the reader to conley_plausibly_2012 and andrews_measuring_2017 for additional examples and discussion. Alternatively, one can also use measures of statistical distance such as the probability of detecting that the model is misspecified to aid with interpretation of $M$, as suggested in HaSa08robustness or bonhomme_minimizing_2018. While it is not possible to determine $M$ automatically, it is possible to obtain a lower CI $[M_{\min}, \infty]$ for $M$, which can be used as a diagnostic check verifying that the values of $M$ considered are not too small. We develop such tests by generalizing the $J$-test of overidentifying restrictions in (ref). We recommend reporting the lower bound $M_{\min}$ along with the plot of the optimal CI as a function of $M$. The norm $\norm{\cdot}$ determines how the researcher's bounds on each element of $\gamma$ interact. With the $\ell_{\infty}$ norm, one places separate bounds on each element of $\gamma$, which leads to a simple interpretation: no single element of $\gamma$ can be greater than $M$. Under an $\ell_{p}$ norm with $1\leq p< \infty$, the bounds on each element of $\gamma$ interact with each other, so that larger amounts of misspecification in one element is allowed if other elements are correctly specified. Depending on whether such interactions are desirable, we recommend setting $p=2$, or $p=\infty$. For these choices of the norm, computing the class of optimal sensitivities $\{\hat{k}_{\lambda}\}_{\lambda\geq 0}$ is particularly simple. In particular, when $\norm{\cdot}$ corresponds to an $\ell_{p}$ norm, the worst-case bias has a closed form, since by Hölder's inequality, $\operatorname{\overline{bias}}_{\mathcal{C}(M)}(k)=\sup_{\norm{\gamma}_{p}\leq 1}M\abs{k'B\gamma}=M\norm{B'k}_{p'}$, where $p'$ is the Hölder complement of $p$ ($p'=1$ if $p=\infty$, while $p'=2$ if $p=2$), and the optimal sensitivities $\{\hat{k}_{\lambda}\}_{\lambda\geq 0}$ can be computed by casting the problem as a penalized regression problem. We explain the connection to penalized regression, and provide details in (ref). When $p=2$, so that $\norm{\cdot}$ corresponds to the Euclidean norm, the problem is analogous to ridge regression, and the optimal sensitivities in Step (ref) of the implementation take the form $\hat{k}_{\lambda}'=-H(\Gamma'W_{\lambda}\Gamma)^{-1}\Gamma'W_{\lambda}$, where $W_{\lambda}=(\lambda BB'+\hat{\Sigma})^{-1}$, with $\overline{B}_\lambda=\norm{B'\hat{k}_{\lambda}}_{2}$. As an alternative to using the one-step estimator in Step (ref) of the implementation, one can implement this sensitivity directly as a GMM estimator with weighting matrix $W_{\lambda}$ (see also (ref) below). Relative to the optimal weighting matrix $\Sigma^{-1}$ under correct specification, the matrix $W_{\lambda}$ trades off precision of the moments against their potential misspecification. When $p=\infty$, the penalized regression analogy leads a simple algorithm for computing the optimal sensitivities $\{\hat{k}_{\lambda}\}_{\lambda\geq 0}$ that is similar to the LASSO/LAR algorithm ehjt04. We give details on the algorithm in (ref). It follows from this algorithm if $B$ corresponds to columns of the identity matrix, as $M$ grows, the optimal sensitivity successively drops the “least informative” moments, so that in the limit, if $d_{g}\leq d_{\gamma}+d_{\theta}$, the optimal sensitivity corresponds to that of an exactly identified GMM estimator based on the $d_{\theta}$ “most informative” moments only, where “informativeness” is given by both the variability of a given moment, and its potential misspecification. If $d_{g}>d_{\gamma}+d_{\theta}$, one simply drops all invalid moments in the limit.
remarkIn Step (ref) of our implementation, we use a one-step estimator $\hat{h}_{\lambda}$ to compute a CI that is asymptotically valid and optimal. Due to concerns about finite-sample behavior (analogous to concerns about finite sample behavior of one-step estimators in the correctly specified case), one may prefer using a different estimator that is asymptotically equivalent to $\hat{h}_{\lambda}$. In general, one can implement an estimator with sensitivity $k$ as a GMM or minimum distance estimator by using an appropriate weighting matrix, so that one can in particular replace $\hat{h}_{\lambda}$ by $h(\hat{\theta}_{W})$, with the weighting matrix $W$ appropriately chosen. To give the formula for the weighting matrix, let $\Gamma_{\perp}$ denote a $d_{g}\times (d_{g}-d_{\theta})$ matrix that's orthogonal to $\Gamma$, so that $\Gamma_{\perp}'\Gamma=0$, and let $\hat{\Gamma}_{\perp}$ denote a consistent estimate. Let $S$ denote a $d_{g}\times d_{\theta}$ matrix that satisfies $S'\hat{\Gamma}=-I$ and $\hat{k}_{\lambda}=S\hat{H}'$. Then we can set $W=S W_{1} S'+\hat{\Gamma}_{\perp}W_{2}\hat{\Gamma}_{\perp}'$ for some non-singular matrix $W_{1}$, and an arbitrary conformable matrix $W_{2}$. It can be verified by simple algebra that $\hat{\theta}_{W}$ will have sensitivity $k_{\lambda}$.
remark[Global misspecification] While we focus on the local misspecification setting, in which the set $\mathcal{C}/\sqrt{n}$ shrinks with $n$ at a $1/\sqrt{n}$, one can use our insights about optimal weighting to construct a CI that retains the near-optimality properties of the above CI under local misspecification, while having correct coverage under asymptotics in which this set shrinks more slowly or stays fixed with the sample size (the latter is termed “global misspecification” in the literature). Let $W$ be a weighting matrix that leads to the optimal sensitivity, as described in (ref) above, and let $\mathcal{I}_{\tilde{c}}$ be a CI constructed from the GMM estimator with moment conditions $\theta\mapsto g(w_i, \theta)-{\tilde{c}}$. Let $\mathcal{I}=\cup_{{\tilde{c}}\in \mathcal{C}/\sqrt{n}}\mathcal{I}_{\tilde{c}}$ be the union of these CIs over possible values of ${\tilde{c}}$ in the set $\mathcal{C}/\sqrt{n}$. Such an approach was suggested in the context of misspecified linear IV by conley_plausibly_2012, although they did not consider adjusting the weighting matrix. The resulting CI has correct asymptotic coverage under both local and global misspecification, and, for one-sided CI construction, is asymptotically equivalent under local misspecification to the CI discussed above. We provide further details in (ref). In the \namecref{global_misspec_sec_append}, we also discuss a second approach to constructing CIs valid under global misspecification based on misspecification-robust standard errors HaIn03, which is applicable if the estimate of the worst-case bias under global misspecification is asymptotically normal. The resulting one- and two-sided CIs are asymptotically equivalent under local misspecification to the optimal CIs discussed above.
remark[Other performance criteria] In addition to constructing a CI, one may be interested in a point estimate of $h(\theta_0)$, using mean squared error (MSE) as the criterion. The steps to forming the MSE optimal point estimate are exactly the same as above, except that, rather than minimizing CI length in Step (ref), we choose $\lambda$ to minimize $\operatorname{\overline{bias}}_{\mathcal{C}}(\hat{k}_\lambda)^2+ \hat{k}_\lambda'\hat{\Sigma} \hat{k}_\lambda=M \overline B_{\lambda} + \hat{k}_\lambda'\hat{\Sigma} \hat{k}_\lambda$. Similar ideas apply to other criteria, such as mean absolute deviation or quantiles of excess length of one-sided CIs (discussed in (ref)). If $\lambda$ is chosen differently in Step (ref) the CI computed in Step (ref) will be longer than the one computed at $\lambda_{M}^*$, but it will still have correct coverage.

Efficiency bounds and near optimality

The CI given in (ref) has the apparent defect that the local misspecification vector $c$ is reflected in the length of the CI only through the a priori restriction $\mathcal{C}$ imposed by the researcher. Thus, if the researcher is conservative about misspecification, the CI will be wide, even if it turns out that $c$ is in fact much smaller than the a priori bounds defined by $\mathcal{C}$. Moreover, this approach requires the researcher to explicitly specify the set $\mathcal{C}$, including any tuning parameters such as the parameter $M$ if the set takes the form $\mathcal{C}=\mathcal{C}(M)$ that we considered in (ref). One may therefore seek to improve upon this CI by estimating the magnitude of $c$, or by estimating the tuning parameters, and constructing a CI that is shorter if these estimates indicate that misspecification is mild. Similarly, it may be restrictive to require that the CI be centered at an asymptotically linear estimator: this rules out, for example, using a $J$-test to decide which moments to use.

The main result of this section shows that, when $\mathcal{C}$ is convex and centrosymmetric ($c\in\mathcal{C}$ implies $-c\in\mathcal{C}$), the scope for improving on the CI in (ref) is nonetheless limited: no sequence of CIs that maintain coverage under all local misspecification vectors $c\in\mathcal{C}$ can be substantially tighter, even under correct specification. This result can be interpreted as translating results from a “limiting experiment” that is an extension of the linear regression model. We first give a heuristic derivation of this limiting experiment and explain our result in the context of this limiting experiment. We then present the formal asymptotic result, and discuss its implications in some familiar settings. Readers who are interested only in implementing the methods, rather than efficiency results, can skip this section.

We restrict attention in this section to the GMM model, in which $\hat g(\theta)=\frac{1}{n}\sum_{i=1}^{n}g(w_i, \theta)$, and we further restrict the data $\{w_{i}\}_{i=1}^{n}$ to be independent and identically distributed (i.i.d.). Similar to semiparametric efficiency theory in the standard, correctly specified case, this facilitates parts of the formal statements and proofs, such as the definition of the set of distributions under which coverage is required and the construction of least favorable submodels. We expect that analogous results could be obtained in other settings.

Limiting experiment

As discussed in (ref), we can form CIs based on linear estimators with asymptotic distribution $\mathcal{N}(k'c, k'\Sigma k)$. This suggests that the problem of constructing an asymptotically valid CI for $h(\theta)$ in the model (ref) is asymptotically equivalent to the problem of constructing a CI for the parameter $H\theta$ in the approximately linear model

equation[equation omitted — 154 chars of source]

where $\Gamma$, $H$ and $\Sigma^{1/2}$ are known, and we observe $Y$. One can think of this model as an “approximately” linear regression model, with $-\Gamma$ playing the role of the design matrix of the (fixed) regressors, and $c$ giving the approximation error. This model dates back at least to SaYl78, who considered estimation in this model when $\mathcal{C}$ is a rectangular set and $\Sigma$ is diagonal. The analog of the asymptotically linear estimator $\hat{h}$ in (ref) is the linear estimator $k'Y$. To see the analogy, note that $k'Y-H\theta$ is distributed $\mathcal{N}((-k'\Gamma-H)\theta+k'c, k'\Sigma k)$, and restricting ourselves to estimators that do not have infinite worst-case bias when $\theta$ is unrestricted gives the condition (ref). In the limiting experiment, the analog of the CI (ref) is given by the CI $k'Y\pm \operatorname{cv}_\alpha(\operatorname{\overline{bias}}_{\mathcal{C}}(k)/\sqrt{k'\Sigma k})\cdot \sqrt{k'\Sigma k}$. Finding weights that minimize the length of this CI is isomorphic to the problem of finding the sensitivity that minimizes the asymptotic CI length in (ref).

For a general convex set $\mathcal{C}$, the bias-variance optimization problem in (ref) can be reformulated as a convex programming problem, as shown in low95. In particular, when the set $\mathcal{C}$ is centrosymmetric (see (ref) for the general case), the bias-variance optimizing class of weights $\{k_{\lambda}\}_{\lambda> 0}$ is given by the class $\{k_{\delta}\}_{\delta> 0}$, where

equation[equation omitted — 219 chars of source]

and, for each $\delta$, $c_\delta, \theta_\delta$ are the solutions to the convex program

equation[equation omitted — 175 chars of source]

It then follows from donoho94 that among fixed-length CIs based on linear estimators (CIs that take the form $k'Y \pm \chi$ for some constant $\chi$), the shortest CI in the limiting experiment takes the form

equation[equation omitted — 253 chars of source]

where $\operatorname{\overline{bias}}_{\mathcal{C}}(k_\delta) =-k_\delta' c_\delta$, and $\delta^{*}=\operatorname*{argmin}_{\delta>0} 2\operatorname{cv}_\alpha(\operatorname{\overline{bias}}_{\mathcal{C}}(k_{\delta}) / \sqrt{k_{\delta}' \Sigma k_{\delta}})\cdot \sqrt{k_{\delta}'\Sigma k_{\delta}}$ is chosen to minimize the CI length. The CI in (ref) is an analog of this CI, with $\delta$ playing the role of the index $\lambda$.

The CI in (ref) takes a familiar form in the special case in which $\mathcal{C}$ is a linear subspace of $\mathbb{R}^{d_{g}}$, so that for some $d_{g}\times d_{\gamma}$ full-rank matrix $B$ with $d_{\gamma}\leq d_{g}-d_{\theta}$, $\mathcal{C}=\{B\gamma\colon \gamma\in\mathbb{R}^{d_{\gamma}}\}$. Let $B_{\perp}$ denote a $d_{g}\times (d_{g}-d_{\gamma})$ matrix that's orthogonal to $B$. Then for any $\delta>0$, $k_{\delta}'={k}'_{LS, B}$, where

equation[equation omitted — 207 chars of source]

is the sensitivity of the GLS estimator after pre-multiplying (ref) by $B_{\perp}'$, (which effectively picks out the observations with zero misspecification). Since this estimator is unbiased, the CI in (ref) becomes ${k}_{LS, B}'Y\pm z_{1-\alpha/2}\sqrt{k_{LS, B}'\Sigma k_{LS, B}}$.

Like the asymptotic CI (ref), the CI in (ref) has the potential drawback that its length is determined by the worst possible misspecification in $\mathcal{C}$, leaving open the possibility of efficiency improvements when $c$ turns out to be close to zero. As a best-case scenario for such improvements, consider the problem: among confidence sets with coverage at least $1-\alpha$ for all $\theta\in\mathbb{R}^{d_\theta}$ and $c\in\mathcal{C}$, minimize expected length when $\theta=\theta^*$ and $c=0$. Note that this setup is even more favorable for potential improvements on our CI, since it allows the researcher to guess correctly that $\theta$ is equal to some $\theta^*$, and it allows for confidence sets that are not intervals (in this case, length is defined as Lebesgue measure). Let $\kappa_{*}(H, \Gamma, \Sigma, \mathcal{C})$ denote the ratio of this optimized expected length relative to the length of the CI in (ref) (it can be shown that this ratio does not depend on $\theta^*$).

If $\mathcal{C}$ is convex, a formula for $\kappa_{*}(H, \Gamma, \Sigma, \mathcal{C})$ follows from applying the general results in Corollary 3.3 in ArKo18optimal to the limiting model. If $\mathcal{C}$ is also centrosymmetric, then

equation[equation omitted — 291 chars of source]

where $Z\sim \mathcal{N}(0,1)$ and $\omega(\delta)$ is two times the optimized value of (ref). Furthermore, we show in (ref) that the right-hand side is lower-bounded by $(z_{1-\alpha}(1-\alpha)-\tilde{z}_{\alpha}\Phi(\tilde{z}_{\alpha})+ \phi(z_{1-\alpha})-\phi(\tilde{z}_{\alpha}))/ z_{1-\alpha/2}$, where $\tilde{z}_{\alpha}=z_{1-\alpha}-z_{1-\alpha/2}$ for any $H$, $\Gamma$, $\Sigma$ and $\mathcal{C}$, where $\phi(\cdot)$ denotes the standard normal density. For $\alpha=0.05$, this universal lower bound evaluates to 71.7%. Evaluating $\kappa_{*}$ for particular choices of $H$, $\Gamma$, $\Sigma$, and $\mathcal{C}$ often yields even higher efficiency.

If $\mathcal{C}$ is a linear subspace, then $\omega(\delta)$ is linear, and

equation[equation omitted — 203 chars of source]

where the lower bound follows since $\phi(z_{1-\alpha})\geq \alpha z_{1-\alpha}$ by the Gaussian tail bound $1-\Phi(x)\leq \phi(x)/x$ for $x>0$. This bound corresponds to that in pratt61 for the case of a univariate normal mean. The potential efficiency improvement essentially comes from using prior knowledge of $\theta^*$ to turn a two-sided critical value into a one-sided critical value. Furthermore, it follows from joshi69 that the CI ${k}_{LS, B}'Y\pm z_{1-\alpha/2}\sqrt{k_{LS, B}'\Sigma k_{LS, B}}$ is the unique CI that achieves minimax expected length. Thus, not only is the scope for improvement at a particular $\theta^{*}$ bounded by (ref), any CI with shorter expected length at some $\theta^{*}$ must necessarily perform worse elsewhere in the parameter space.

For the one-sided CI (ref), the analogous CI in the limiting experiment is $\hor{k'Y-\operatorname{\overline{bias}}_{\mathcal{C}}(k)-z_{1-\alpha}\sqrt{k'\Sigma k}, \infty}$, and, as we discuss in (ref), to choose the optimal sensitivity $k$, one can consider optimizing a given quantile of its worst-case excess length. The results in ArKo18optimal again give an efficiency bound for improvement at $c=0$ and a particular $\theta^*$, analogous to (ref) for the two-sided case. See (ref) for details. If $\mathcal{C}$ is a linear subspace, then optimizing quantiles of worst-case excess length yields the CI $\hor{k_{LS, B}'Y-z_{1-\alpha}\sqrt{k_{LS, B}'\Sigma k_{LS, B}}, \infty}$, independently of the quantile one is optimizing. Furthermore, the efficiency bound implies that this one-sided CI is in fact fully optimal over all quantiles of excess length and all values of $\theta, c$ in the local parameter space.

These efficiency results for the CI (ref) in the limiting experiment suggest that the scope for improvement over the CI in (ref) should be limited in large samples. (ref), stated in the next section, uses the analogy with the approximately linear model (ref) along with Le Cam-style arguments involving least favorable submodels to show that this bound indeed translates to the locally misspecified GMM model. For one-sided CIs, we state an analogous result in (ref). We discuss the implications of these results in (ref).

Asymptotic efficiency bound

To make precise our statements about coverage and efficiency, we need the notion of uniform (in the underlying distribution) coverage of a confidence interval. This requires additional notation, which we now introduce. Let $\mathcal{P}$ denote a set of distributions $P$ of the data $\{w_{i}\}_{i=1}^{n}$, and let $\Theta_n\subseteq\mathbb{R}^{d_\theta}$ denote the parameter space for $\theta$. We require coverage for all pairs $(\theta, P)\in \Theta_n\times \mathcal{P}$ such that $\sqrt{n}g_P(\theta)\in\mathcal{C}$, where the subscript $P$ on the population moment condition makes it explicit that it depends on the distribution of the data.\footnote{To be precise, we should also subscript all other quantities such as $\Gamma$ and $\Sigma$ by $P$. To prevent notational clutter, we drop this index in the main text unless it causes confusion.} Letting $\mathcal{S}_n=\{(\theta, P)\in \Theta_n\times\mathcal{P}: \sqrt{n}g_P(\theta)\in\mathcal{C}\}$ denote this set, the condition for coverage at confidence level $1-\alpha$ can be written

equation[equation omitted — 139 chars of source]

We say that a confidence set $\mathcal{I}_{n}$ is asymptotically valid (uniformly over $\mathcal{S}_{n}$) at confidence level $1-\alpha$ if this condition holds.\footnote{In general, $\theta_0$ and $h(\theta_0)$ may be set identified for a given sample size $n$ (although our assumptions imply that the identified set will shrink at a root-$n$ rate). The coverage requirement (ref) states that the CI must cover points in the identified set for $h(\theta)$, as in im04; see (ref).}

Among two-sided CIs of the form $\hat{h}\pm \hat{\chi}$ that are asymptotically valid, we prefer CIs with shorter expected length. To avoid issues with convergence of moments, we use truncated expected length, and define the asymptotic expected length of a two-sided CI at $P_{n}\in\mathcal{P}$ as $\liminf_{T\to\infty}\liminf_{n\to\infty}E_{P_{n}}\min\{\sqrt{n}\cdot 2\hat{\chi}, T\}$, where $E_{P}$ denotes expectation under $P$.

We are now ready to state the main efficiency result.

theoremSuppose that $\mathcal{C}$ is convex and centrosymmetric. Let $\hat h_{\lambda^*}$ and $\hat\chi^*_{\lambda^*}$ be formed as in (ref). Suppose that (ref) in (ref) hold. Suppose that the data $\{w_{i}\}_{i=1}^{n}$ are i.i.d.\ under all $P\in\mathcal{P}$. Let $(\theta^*, P_0)$ be correctly specified (i.e. $g_{P_0}(\theta^*)=0$) such that $\mathcal{P}$ contains a submodel through $P_0$ satisfying (ref). Then: \begin{enumerate}[label=(\roman{enumi})] • The CI $\hat{h}_{\lambda^{*}}\pm \hat \chi^*_{\lambda^*}$ is asymptotically valid, and its half-length $\hat\chi^*_{\lambda^*}$ satisfies $\sqrt{n}\hat \chi^*_{\lambda^*}=\chi(\theta, P)+o_P(1)$ uniformly over $(\theta, P)\in\mathcal{S}_n$ where \begin{equation*} \chi(\theta, P)=\min_k \operatorname{cv}_\alpha(\operatorname{\overline{bias}}_{\mathcal{C}}(k)/ \sqrt{k'\Sigma_{\theta, P}k})\sqrt{k'\Sigma_{\theta, P}k} \end{equation*} with $\operatorname{\overline{bias}}_{\mathcal{C}}(k)$ calculated with $\Gamma=\Gamma_{\theta, P}$ and $H=H_{\theta}$. • For any other asymptotically valid CI $\hat{h}\pm \hat{\chi}$, \begin{equation*} \frac{\liminf_{T\to\infty}\liminf_{n\to\infty}E_{P_0}\min\{\sqrt{n}\cdot 2\hat{\chi}, T\}} {2\chi(\theta^*, P_0)} \ge \kappa_{*}(H_{\theta^*}, \Gamma_{\theta^*, P_0}, \Sigma_{\theta^*, P_0}, \mathcal{C}), \end{equation*} where $\kappa_{*}(H, \Gamma, \Sigma, \mathcal{C})$ is defined in (ref). Furthermore, for any $H, \Sigma, \Gamma$, and $\mathcal{C}$, $\kappa_{*}$ admits the universal lower bound $(z_{1-\alpha}(1-\alpha)-\tilde{z}_{\alpha}\Phi(\tilde{z}_{\alpha})+ \phi(z_{1-\alpha})-\phi(\tilde{z}_{\alpha}))/ z_{1-\alpha/2}$, where $\tilde{z}_{\alpha}=z_{1-\alpha}-z_{1-\alpha/2}$ and $\phi(\cdot)$ denotes the standard normal density. \end{enumerate}

The proof for this theorem is given in (ref), which also gives an analogous result for one-sided confidence intervals. (ref), stated in (ref), require that the conditions in (ref) hold in a uniform sense over the class $\mathcal{P}$. In the supplemental materials, we give primitive conditions for these assumptions in the misspecified linear IV model. (ref), also stated in the appendix, requires that the class $\mathcal{P}$ be rich enough to contain a submodel that is least favorable for the GMM problem, so that the class doesn't implicitly impose any other conditions that could be used to make inference easier. In the supplemental materials, we provide a general way of constructing a submodel satisfying these conditions.

The universal lower bound on $\kappa_{*}$ is new and may be of independent interest. For $\alpha=0.05$, it evaluates to 71.7%. The universal lower bound is sharp in the sense that there exist $\Gamma, \Sigma, H$ and $\mathcal{C}$ for which $\kappa_{*}$ equals this lower bound. In particular applications, the efficiency bound $\kappa_{*}$ can be computed at estimates of $\Gamma$, $\Sigma$ and $H$, and often, this gives much higher efficiencies. We illustrate these bounds in the empirical application in (ref).

Discussion

To help build intuition for the efficiency bound in (ref), and to relate this result to the literature, we now consider some special cases. We first discuss the (standard) correctly specified case. Second, we consider the case in which some moments are known to be valid, and the misspecification in the remaining moments is unrestricted. This case may be of interest in its own right. We then discuss the general case. Finally, we discuss the connection to certain statistical measures of distance considered in the literature.

Correctly specified case

Suppose that $\mathcal{C}=\{0\}$. This is in particular a linear subspace of $\mathbb{R}^{d_{g}}$, with $B=0$, and $B_{\perp}=I$, the $d_{g}\times d_{g}$ identity matrix. Thus, in the limiting experiment, the optimal CI uses the GLS estimator $k_{LS,0}'Y$, with $k_{LS,0}$ given in (ref) (with $B=0$). For testing the null hypothesis $H\theta=h_{0}$ against the one-sided alternative $H\theta\geq h_{0}$, the one-sided $z$-statistic based on ${k}_{LS,0}'Y$ is uniformly most powerful van_der_vaart_asymptotic_1998. Inverting these tests yields the CI $\hor{k_{LS,0}'Y-z_{1-\alpha}\sqrt{k_{LS,0}'\Sigma k_{LS,0}}, \infty}$. Since the underlying tests are uniformly most powerful, this CI achieves the shortest excess length, simultaneously for all quantiles and all possible values of the parameter $\theta$. For two-sided CIs, the results described in (ref) imply that the CI $h_{LS,0}'Y\pm z_{1-\alpha/2}\sqrt{k_{LS,0}'\Sigma k_{LS,0}}$ is the unique CI that achieves minimax expected length, and the efficiency of this CI relative to a CI that optimizes its expected length at a single value $\theta^*$ of $\theta$ when indeed $\theta=\theta^*$ is given in (ref). It evaluates to 84.99% at $\alpha=0.05$.

Applying (ref) to the case $\mathcal{C}=\{0\}$ gives an asymptotic version of the two-sided efficiency bound. Furthermore, the CI in (ref) reduces to the usual two-sided CI based on $\hat{\theta}_{\Sigma^{-1}}$. Thus, in this case, (ref) shows that very little can be gained over the usual two-sided CI by optimizing the CI relative to a particular distribution $P_0$. Results in the appendix give an analogous result for one-sided CIs. In the one-sided case, this asymptotic result is essentially a version of a classic result from the semiparametric efficiency literature for one-sided tests, applied to CIs (see Chapter 25.6 in van_der_vaart_asymptotic_1998). In the two-sided case, the result is, to our knowledge, new.

Some valid and some invalid moments

Consider now the case in which the first $d_{g}-d_{\gamma}$ moments are known to be valid, with the potential misspecification for the remaining $d_{\gamma}$ moments unrestricted. Then $\mathcal{C}=\{(0', \gamma')'\colon \gamma\in\mathbb{R}^{d_{\gamma}}\}$ corresponds to a linear subspace with $B$ given by the last $d_{\gamma}$ columns of the identity matrix, and $B_{\perp}$ given by the first $d_{g}-d_{\gamma}$ columns. Optimal CIs in the limiting experiment therefore use the estimator $k_{LS, B}'Y$, which is the GLS estimator based only on the observations with no misspecification.

The one-sided CI based on $k_{LS, B}'Y$ achieves the shortest excess length, simultaneously for all quantiles and all possible values of the parameter $\theta$. The two-sided CI $k_{LS, B}'Y\pm z_{1-\alpha/2}\sqrt{k_{LS, B}'\Sigma k_{LS, B}}$ is optimal in the same sense as the usual CI in (ref): it achieves minimax expected length, and its efficiency, relative to a CI that optimizes its length at a single $\theta^{*}$ and $\gamma=0$, is lower-bounded by $z_{1-\alpha}/z_{1-\alpha/2}$. (ref) formally translates the efficiency bound from the limiting model to the GMM model, so that the usual two-sided CI based on $h(\hat{\theta}_{W(B)})$ is asymptotically efficient in the same sense as the usual CI based on $h(\hat{\theta}_{\Sigma^{-1}})$ discussed in (ref) under correct specification. Just as with the results in (ref), this asymptotic result is, to our knowledge, new. The one-sided analog follows from the results in (ref). These results stand in sharp contrast to the results for estimation, where the MSE improvement at small values of $\gamma$ may be substantial.

An important consequence of these results is that asymptotically valid one-sided CIs based on shrinkage or model-selection procedures, such as one-sided versions of the CIs proposed in andrews_hybrid_2009, ditraglia_using_2016 or mccloskey_2017 must have worse excess length performance than the usual one-sided CI based on the GMM estimator $h(\hat{\theta}_{W(B)})$ that uses valid moments only. While it is possible to construct two-sided CIs that improve upon the usual CI based on $h(\hat{\theta}_{W(B)})$ at particular values of $\theta$ and $\gamma$, the scope for such improvement is smaller than the ratio of one- to two-sided critical values. Furthermore, any such improvement must come at the expense of worse performance at other points in the parameter space.\footnote{Consistently with these results, in a simulation study considered in ditraglia_using_2016, the post-model selection CI that he proposes is shown to be wider on average than the usual CI around a GMM estimator that uses valid moments only.} Therefore, in order to tighten CIs based on valid moments only, it is necessary to make a priori restrictions on the potential misspecification of the remaining moments.

General case

According to the results in (ref), one must place a priori bounds on the amount of misspecification in order to use misspecified moments. This leads us to the general case, where we place the local misspecification vector $c$ in some set $\mathcal{C}$ that is not necessarily a linear subspace. One can then form a CI centered at an estimate formed from these misspecified moments using the methods in (ref). In the case where $\mathcal{C}$ is convex and centrosymmetric, (ref) shows that this CI is near optimal, in the sense that no other CI can improve upon it by more than a factor of $\kappa_{*}$, even in the favorable case of correct specification. Since the width of the CI is asymptotically constant under local parameter sequences $\theta_n\to\theta^*$ and sufficiently regular probability distributions $P_n\to P_0$ (for example, $P_n\to P_0$ along submodels satisfying (ref)), this also shows that the CI is near optimal in a local minimax sense. In the general case, (ref), as well as the analogous results for one-sided CIs in (ref) are, to our knowledge, new.

As we discuss in (ref), we recommend reporting results for a range of sets $\mathcal{C}(M)$ indexed by a scalar $M$ that bounds the magnitude of misspecification. One may instead wish to report a single CI based on a data-driven estimate of $M$, for example, by using a first-stage $J$ test to assess plausible magnitudes of misspecification. Formally, one would seek a CI that is valid over $\mathcal{C}(\overline{M})$ while improving length when in fact $\norm{\gamma}\ll \overline{M}$, where $\overline{M}$ is some initial conservative bound. When $\mathcal{C}$ is convex and centrosymmetric, (ref) shows that the scope for such improvements is limited: the average length of any such CI cannot be much smaller than the CI that uses the most conservative choice $\overline{M}$, even when $c=0$. The impossibility of choosing $M$ based on the data is related to the impossibility of using specification tests to form an upper bound for $M$. On the other hand, it is possible to obtain a lower bound for $M$ using such tests. We develop lower CIs for $M$ in (ref).

Cressie-Read divergences

andrews_informativeness_2018 have shown that defining misspecification in terms of the magnitude of any divergence in the cressie_multinomial_1984 family leads to a set $\mathcal{C}$ that asymptotically takes the form $\mathcal{C}=\{\Sigma^{1/2}\gamma:\|\gamma\|_2\le M\}=\{c\colon c'\Sigma^{-1}c\le M^2\}$. The Cressie-Read family includes the Hellinger distance used by kitamura_robustness_2013, who consider minimax point estimation among estimators satisfying certain regularity conditions. Since this set $\mathcal{C}$ takes the form discussed in (ref) with $p=2$ and $B=\Sigma^{1/2}$, it follows from the discussion in (ref) that the optimal sensitivity corresponds to the GMM estimator with weighting matrix $(\lambda BB'+\Sigma)^{-1}=(\lambda+1)^{-1}\Sigma^{-1}$. Since this is proportional to the weighting matrix $\Sigma^{-1}$ that is optimal under correct specification, we obtain the same optimal sensitivity $k_{LS,0}'=-H(\Gamma' \Sigma^{-1}\Gamma)^{-1}\Gamma'\Sigma^{-1}$ as in the correctly specified case discussed in (ref). As we show in (ref), this form of $\mathcal{C}$ leads to a closed form solution for the efficiency bound $\kappa_*$.

The results above imply that any estimator with sensitivity $k_{LS,0}$ is near-optimal for CI construction. In line with these results, the estimator in kitamura_robustness_2013 has sensitivity $k_{LS,0}$. Thus, the usual GMM estimator $h(\hat{\theta}_{\Sigma^{-1}})$ and the estimator in kitamura_robustness_2013 are both near-optimal for CI construction, even if one allows for arbitrary CIs that are not necessarily centered at estimators that satisfy the regularity conditions in kitamura_robustness_2013. Also, because they have the same sensitivity, under this form of misspecification, the usual GMM estimator $h(\hat{\theta}_{\Sigma^{-1}})$ and the estimator in kitamura_robustness_2013 have the same local asymptotic minimax properties.

Extensions: asymmetric constraints and constraints on \texorpdfstring{$\theta$}{theta}

If the set $\mathcal{C}$ is convex but asymmetric (such as when $\mathcal{C}$ includes bounds on a norm as well as sign restrictions, or when $\mathcal{C}$ includes equality and sign restrictions, as in moon_estimation_2009), one can still apply bounds from ArKo18optimal to the limiting model described in (ref). Our general asymptotic efficiency bounds in (ref) translate these results to the locally misspecified GMM model so long as $\mathcal{C}$ is convex. Since the negative implications for efficiency improvements under correct specification use centrosymmetry of $\mathcal{C}$, introducing asymmetric restrictions, such as sign restrictions, is one possible way of getting efficiency improvements at some smaller set $\mathcal{D}\subseteq\mathcal{C}$ while maintaining coverage over $\mathcal{C}$. We derive efficiency bounds and optimal CIs for this problem in (ref). Interestingly, the scope for efficiency improvements can be different for one- and two-sided CIs, and can depend on the direction of the CI in this case. To get some intuition for this, note that, in the instrumental variables model with a single instrument and single endogenous regressor, sign restrictions on the covariance of an instrument with the error term can be used to sign the direction of the bias of the instrumental variables estimator, which is useful for forming a one-sided CI only in one direction.

Finally, while we focus on restrictions on $c$, one can also incorporate local restrictions on $\theta$. Our general results in (ref) give efficiency bounds that cover this case. Similar to the discussion above, these results have implications for using prior information about $\theta$ to determine the amount of misspecification, or to shrink the width of a CI directly. In particular, while it is possible to use prior information on $\theta$ (say, an upper bound on $\|\theta\|$ for some norm $\|\cdot\|$) to shrink the width of the CI, the width of the CI and the estimator around which it is centered must depend on the a priori upper bounds on the magnitude of $\theta$ and $c$ when this prior information takes the form of a convex, centrosymmetric set for $(\theta', c')'$. This rules out, for example, choosing the moments based on whether the resulting estimate for $\theta$ is in a plausible range.

Applications

This section describes particular applications of our approach, along with a discussion of implementation details appropriate to each application.

Instrumental variables

The single equation linear instrumental variables (IV) model is given by

equation[equation omitted — 74 chars of source]

where, in the correctly specified case, $E\varepsilon_{i}z_{i}=E(y_{i}-x_i'\theta_0)z_{i}=0$, with $z_{i}$ a $d_{g}$-vector of instruments. This is an instance of a GMM model with $g(\theta)=E(y_{i}-x_i'\theta)z_{i}$ and $\hat{g}(\theta)=\frac{1}{n}\sum_{i=1}^{n}z_{i}(y_{i}-x_{i}'\theta)$.

One common reason for misspecification in this model is that the instruments do not satisfy the exclusion restriction, because they appear directly in the structural equation (ref), so that $\varepsilon_{i}=z_{Ii}'\gamma/\sqrt{n}+\eta_{i}$, where $E[z_{i}\eta_{i}]=0$, and $z_{Ii}$ corresponds to a subset $I$ of the instruments, the validity of which one is worried about. This form of misspecification has previously been considered in a number of papers, including hahn_iv_2005, conley_plausibly_2012, and andrews_measuring_2017, among others. Bounding the norm of $\gamma$ using some norm $\norm{\cdot}$ then leads to the set given in (ref), with $B=E[z_{i}z_{Ii}']$.

Although the matrix $B$ is unknown, for the purposes of estimating the optimal sensitivity and constructing asymptotically valid CIs, it can be replaced by the sample analog $\hat{B}=n^{-1}\sum_{i=1}^{n}z_{i}z_{Ii}'$. This does not affect the asymptotic validity or coverage properties of the resulting CI\@. Under this setup, the parameter $M$ bounds that magnitude of $\gamma$, the direct effect of the instruments on the outcome. Therefore, the appropriate choice of $M$ will depend on the plausible magnitude of these direct effects---see, for example, conley_plausibly_2012 for examples and a discussion.

The linearity of the moment condition leads to simplifications in our implementation in (ref). In Step (ref), as the initial estimator, one can use the two-stage least squares (2SLS) estimator

equation*[equation* omitted — 310 chars of source]

This leads to the estimates $\hat{\Gamma}=-\frac{1}{n}\sum_{i=1}^{n}z_{i}x_i'$ and $\hat{\Sigma}=\frac{1}{n}\sum_{i=1}^{n} (y_{i}-x_i'\hat\theta_{\text{initial}})^2z_{i}z_{i}'$. Alternatively, if we assume homoskedasticity, we can use the estimator $\hat\Sigma_{H}=\frac{1}{n}\sum_{i=1}^{n}(y_{i}-x_i'\hat\theta_{\text{initial}})^2\cdot \frac{1}{n}\sum_{i=1}^{n}z_{i}z_{i}'$. In the correctly specified case, the 2SLS estimator is only optimal under homoskedasticity, while the GMM estimator with weighting matrix $\hat\Sigma^{-1}$ is optimal in general. Due to concerns with finite sample performance, however, it is common to use the 2SLS estimator along with standard errors based on a robust variance estimate, even when heteroskedasticity is suspected. Mirroring this practice, one can use $\hat{\Sigma}_{H}$ when forming the optimal sensitivity $\hat{k}_{\lambda^{*}_{M}}$ and worst-case bias in Step (ref), but use the robust variance estimate $\hat{\Sigma}$ in Step (ref) when forming the final CI in (ref). The CI will then be optimal under homoskedasticity, but it will remain valid under heteroskedasticity, just like the usual CI based on 2SLS with robust standard errors in the correctly specified case.

If the parameter of interest linear in $\theta$, $h(\theta)=H\theta$, then the one-step estimator $\hat{h}_{\lambda^{*}_{M}}$ in Step (ref) does not depend on the choice of the initial estimator (except possibly through the estimate of $\Sigma$ when forming the desired sensitivity):

equation*[equation* omitted — 457 chars of source]

where the second line follows since the sensitivities $\hat{k}_{\lambda^{*}_{M}}$ satisfy $H=-\hat{k}_{\lambda^{*}_{M}}'\hat\Gamma= \hat{k}_{\lambda^{*}_{M}}'\frac{1}{n}\sum_{i=1}^{n}z_{i}x_i'$. Since the estimator $\hat{h}_{\lambda^{*}_{M}}$ is linear, the worst-case bias calculations are the same under global misspecification, when the magnitude $M$ of $\gamma$ in (ref) grows at the rate $\sqrt{n}$. By using variance estimates that are valid under global misspecification in place of the variance estimate $\hat{k}_{\lambda^{*}_{M}}'\hat\Sigma \hat{k}_{\lambda^{*}_{M}}$ in the CI construction, one can ensure that the resulting CI also remains valid under global misspecification. See (ref) for details.

remarkThis framework can also be used to incorporate a priori restrictions on the magnitude of coefficients on control variables in an instrumental variables regression. Suppose that we have a set of controls $w_{i}$, that appear in the structural equation (ref), so that $y_{i}=x_i'\theta+w_i'\gamma/\sqrt{n}+\epsilon_{i}$, and $\epsilon_{i}$ is uncorrelated with $w_{i}$ as well as vector of instruments $\tilde{z}_{i}$. If one is willing to restrict the magnitude of the coefficient vector $\gamma$, so that $\norm{\gamma}\leq M$, then one can add $w_{i}$ to the original vector of instruments $\tilde{z}_{i}$, $z_{i}=(\tilde{z}_{i}', w_{i}')'$. For example, if one is concerned with functional form misspecification, one can define the control variables to be higher order series terms. We then obtain the misspecified IV model with the set $\mathcal{C}$ given by (ref), with $B=E[z_{i}w_{i}']$. Thus, we can interpret this model as a locally misspecified version of a model with $w_{i}$ used as an excluded instrument.
remarkInstead of bounding the coefficient vector $\gamma$, one can alternatively bound the magnitude of the direct effect $z_{Ii}'\gamma$. If all instruments are potentially invalid, $z_{Ii}=z_{i}$, and one sets $\mathcal{C}=\{\gamma\colon E[(z_{i}'\gamma)^{2}]\leq M\}$, then under homoscedasticity, this corresponds to the case discussed in (ref), where the uncertainty from potential misspecification is exactly proportional to the asymptotic sampling uncertainty in $\hat{g}(\theta)$. Consequently, in this case the optimal sensitivity is the same as that given by the 2SLS estimator.

Omitted variables bias in linear regression

Specializing to the case where $z_{i}=x_i$, the misspecified IV model of (ref) gives a misspecified linear regression model as a special case. This can be used to assess sensitivity of regression results to issues such as omitted variables bias. In particular, consider the linear regression model

equation*[equation* omitted — 95 chars of source]

where $x_i$ and $y_{i}$ are observed and $w_i^*$ is a (possibly unobserved) omitted variable. Correlation between $w_i^*$ and $x_i$ will lead to omitted variables bias in the OLS regression of $y_{i}$ on $x_i$. If $w_{i}^{*}$ is unobserved, then we obtain our framework by making the assumption $\sqrt{n}Ew_{i}^{*}x_{i}\in \mathcal{C}$, for some set $\mathcal{C}$, and letting $\hat{g}(\theta)=\frac{1}{n}\sum_{i=1}^{n}x_{i}(y_{i}-x_{i}'\theta)$. This setup can also cover choosing between different sets of control variables. Suppose that $w_{i}^{*}=w_{i}'\gamma$, where $w_{i}$ is a vector of observed control variables that the researcher is considering not including in the regression. If $\gamma$ is unrestricted, then by the results in (ref), the long regression of $y_{i}$ on both $x_{i}$ and $w_{i}$ yields nearly optimal CIs. If one is willing to restrict the magnitude of $\gamma$, it is possible to tighten these CIs, with the setting reducing to that in (ref), with $z_{i}=(x_{i}, w_{i}')'$. The same framework can be used to incorporate selection bias by defining $w_i^*$ to be the inverse Mills ratio term in the formula for $E[y_{i}\mid x_i, i\;\text{observed}]$ in heckman_sample_1979.

Functional form misspecification

Our setup allows for misspecification in moment conditions arising from functional form misspecification. To apply our setup, one must relate this misspecification to the bounds $\mathcal{C}$ on the moment conditions at the true parameter value. One approach to bounding functional form misspecification is to use smoothness conditions from the nonparametric statistics literature, such as bounds on derivatives tsybakov09. Since these sets are typically convex (taking a convex combination of two functions that satisfy a given bound on a given derivative gives a function that also satisfies this bound), they typically lead to convex sets $\mathcal{C}$, so that our framework can be applied.

As a simple example, consider a nonparametric IV model with discrete covariates:

equation*[equation* omitted — 43 chars of source]

Suppose $x$ takes values in the finite set $\mathcal{X}=\{\tilde x_1,\ldots, \tilde x_{N_x}\}$ and $z_i$ takes values in the finite set $\mathcal{Z}=\{\tilde z_1,\ldots, \tilde z_{N_z}\}$. This setting was considered by freyberger_identification_2015, who place only nonparametric smoothness or shape restrictions on the unknown function $m$. To see the connection with our setting, we note that such restrictions can be interpreted as bounds on specification error from a parametric model. If one models these restrictions as local to a parametric family, one obtains our setting. In particular, let $m(x_i)=f(x_i, \theta_0)+n^{-1/2}r(x_i)$, $r\in\mathcal{R}$, where $\mathcal{R}$ is a nonparametric smoothness class. For example, if $x_i$ is univariate, we can let $f(x_i, \theta)=\theta_1+\theta_2x_i$ and define $\mathcal{R}$ to be the class of functions with $r(0)=r'(0)=r''(0)$ and second derivative bounded by some constant $M$. This is equivalent to placing the bound $n^{-1/2}M$ on the second derivative of $m(\cdot)$, which corresponds to a Hölder smoothness class. We can then map this to a misspecified GMM model, with the $j$th element of the moment function given by $g_j(x_i, y_i, \theta)=(y_i-f(x_i, \theta_0))I(z_i=\tilde z_j)$ and $j$th element of the misspecification vector $c$ given by $Er(x_i)I(z_i=\tilde z_j)=\sum_{\tilde x\in\mathcal{X}} r(\tilde x) P(x_i=\tilde{x}, z_i=\tilde z_j)$. Stacking these equations, we see that $c=B\gamma$ where $B$ is a matrix composed of the elements $P(x_i=\tilde x, z_i=\tilde z_j)$ and $\gamma=(r(\tilde x_1), \ldots, r(\tilde x_{N_x}))'$. As with the IV setting in (ref), $B$ is unknown, but can be replaced by a consistent estimate based on the sample analogue. So long as the set $\mathcal{R}$ is convex, we obtain convex restrictions on $\gamma$ and therefore $c$, so that our framework applies.

This example brings up an important point about the interpretation of $h(\theta)$. If the object of interest is a functional of $m(x)=f(x, \theta_0)+n^{-1/2}r(x)$, then we will need to allow the object of interest $h(\cdot)$ to depend on the misspecification vector directly, as well as on $\theta$. As discussed at the beginning of (ref), this falls into a mild extension of our framework. Alternatively, under a suitable parametrization of $f$ and $r$, it is often possible to define the object of interest to be function of $\theta$ alone. For example, if we are interested in the derivative $m'(x_0)$ at a particular point $x_0$ under a bound on the second derivative of $m(\cdot)$, we can let $f(x, \theta)=\theta_1+\theta_2 x$ and define $\mathcal{R}$ to be the class of functions with $r(x_0)=r'(x_0)=r''(x_0)=0$ and second derivative bounded by $M$. Then $m'(x_0)=\theta_2$.

Treatment effect extrapolation

Often, the average effect of a counterfactual policy on a particular subset of a population is of interest, and we would like to weaken the assumptions under which this effect is point-identified. We have available estimates $\hat{\tau}=(\hat\tau_{1}, \dotsc, \hat{\tau}_{m})'$ of the parameter $\tau$, with $\hat{g}(\theta)=\hat{\tau}-A\theta$, and $g(\theta)=\tau-A\theta$ for a known matrix $A$. We would like to extrapolate from these estimates to learn about the parameter of interest $h(\theta_{0})=H\theta_{0}$. The potential extrapolation bias is captured by the assumption that $g(\theta_{0})\in{\widetilde{\mathcal{C}}}_{n}$, some convex set.

Note that, because the moment condition is linear in the parameter of interest, asymptotic validity of our CIs does not require that the set ${\widetilde{\mathcal{C}}}_{n}$ takes the form ${\widetilde{\mathcal{C}}}_{n}=\mathcal{C}/\sqrt{n}$. CIs given in (ref) based on linear estimators of the form $\hat{h}=k'\hat{\tau}$ (such as minimum distance estimators), with $\mathcal{C}=\sqrt{n}{\widetilde{\mathcal{C}}}_{n}$ are valid under both local and global misspecification (i.e.\ under the assumption that the set ${\widetilde{\mathcal{C}}}_{n}$ is fixed as $n\to\infty$).

One example that falls into this setup are differences-in-differences designs when the parallel trends assumption is violated. Here there are $m$ time periods, with treatment taking place in period $T_{0}$. The $(m-T_{0})$-vector $\theta$ corresponds to a vector of dynamic treatment effects on the treated, $A\theta=(0', \theta')'$, and $g(\theta_{0})$ is a vector of trend differences between the treated and untreated, with $g(\theta_{0})=0$ if the parallel trends assumption holds. RaRo19 build on the framework in this paper to develop CIs in this setting.

Another example that has been of recent interest involves nonseparable models with endogeneity. Under conditions in imbens_identification_1994 and heckman_structural_2005, instrumental variables estimates $\hat{\tau}_{m}$ with different instruments are consistent for average treatment effects for different subpopulations. A recent literature kowalski_doing_2016,brinch_beyond_2017,mogstad_using_2017 has focused on using assumptions on treatment effect heterogeneity to extrapolate these estimates to other populations. Our framework applies if these assumptions amount to placing the differences between the estimated treatment effects and the effect of interest in a known convex set.

Empirical application

This section illustrates the confidence intervals developed in (ref) in an empirical application to automobile demand based on the data and model in blp95. We use the version of the model as implemented by andrews_measuring_2017, who calculate the asymptotic bias of the GMM estimator with weighting matrix $\Sigma^{-1}$ under local misspecification in this setting.\footnote{The dataset for this empirical application has been downloaded from the andrews_measuring_2017 replication files, available at \url{https://doi.org/10.7910/DVN/LLARSN}.}

Model description and implementation

In this model, the utility of consumer $i$ from purchasing a vehicle $j$, relative to the outside option, is given by a random-coefficient logit model $U_{ij}=\sum_{k=1}^{K}x_{jk}({\beta}_{k}+\sigma_{k}v_{ik})-\alpha p_{j}/y_{i}+\xi_{j}+\epsilon_{ij}$, where $p_{j}$ is the price of the vehicle, $x_{jk}$ the $k$th observed product characteristic, $\xi_{j}$ is an unobserved product characteristic, and $\epsilon_{ij}$ is has an i.i.d.\ extreme value distribution. The income of consumer $i$ is assumed to be log-normally distributed, $y_{i}=e^{m+\varsigma v_{i0}}$, where the mean $m$ and the variance $\varsigma$ of log-income are assumed to be known and set to equal to estimates from the Current Population Survey. The unobservables $v_{i}=(v_{i0}, \dotsc, v_{iK})$ are i.i.d.\ standard normal, while the distribution of the unobserved product characteristic $\xi_{j}$ is unrestricted.

The marginal cost $mc_{j}$ for producing vehicle $j$ is given by $\log(mc_{j})=w_{j}'\nu+\omega_{j}$, where $w_{j}$ are observable characteristics, and $\omega_{j}$ is an unobservable characteristic. The full vector of model parameters is given by $\theta=(\sigma', \alpha, \beta', \nu')'$. Given this vector, and given a vector of unobservable characteristics, one can compute the market shares implied by utility maximization, which can be inverted to yield the unobservable characteristic as a function of $\theta$, $\xi_{j}(\theta)$. One can similarly invert the unobserved cost component, writing it as a function of $\theta$, $\omega_{j}(\theta)$, under the assumption that firms set prices to maximize profits in a Bertrand-Nash equilibrium. Given a vector $z_{dj}$ of demand-side instruments, and a vector $z_{sj}$ of supply-side instruments, this yields the moment condition $g(\theta)=E[\hat{\gamma}(\theta)]$, where

equation*[equation* omitted — 153 chars of source]

The BLP data spans the period 1971 to 1990, and includes information on essentially all $n=999$ models sold during that period (for simplicity, we have suppressed the time dimension in the description above). There are 5 observable characteristics $x_{j}$: a constant, horsepower per 10 pounds of weight (HPWt), a dummy for whether air-conditioning is standard (Air), mileage per 10 dollars (MP\$) defined as MPG over average gas price in a given year, and car size (Size), defined as length times width. The vector $z_{dj}$ consists of $x_{j}$, plus the sum of $x_{j}$ across models other than $j$ produced by the same firm, and for rival firms. There are 6 cost variables $w_{j}$: a constant, log of HPWt, Air, log of MPG, log of Size, and a time trend. The vector $z_{sj}$ consists of these variables, MP\$, and the sums of $w_{j}$ for own-firm products other than $j$, and for rival firms. After excluding collinear instruments, this gives a total of $d_{g}=31$ instruments, 25 of which are excluded to identify $d_{\theta}=17$ model parameters. The parameter of interest is average markup, $h(\theta)=\frac{1}{n}\sum_{j}(p_{j}-mc_{j}(\theta))/p_{j}$.

One may worry that some of these instruments are invalid, because elements of $z_{dj}$ or $z_{sj}$ may appear directly in the utility or cost function with the coefficient on the $\ell$th element given by $\delta_{d\ell}\gamma_{d\ell}/\sqrt{n}$ or $\delta_{s\ell}\gamma_{s\ell}/\sqrt{n}$, respectively. Here $\delta_{d\ell}$ and $\delta_{s\ell}$ are scaling constants so that, given the sample size $n=999$ at hand, $\gamma_{d\ell}$ has the interpretation that the consumer willingness to pay for one standard deviation change in the $\ell$th demand-side instrument $z_{dj\ell}$ is $\gamma_{d\ell}\%$ of the average 1980 car price, and changing the $\ell$th supply-side instrument $z_{sj\ell}$ by one standard deviation changes the marginal cost by $\gamma_{s\ell}$% of the average car price. andrews_measuring_2017 use this scaling in their sensitivity analysis, and they discuss economic motivation for concerns about this form of misspecification. By way of comparison, the estimates of the parameters $\beta$ and $\nu$ in the utility and cost function imply that consumers are on average willing to pay between 2.2 and 10.0% of the average car price for a standard deviation change in one of the included car characteristics, and that a standard deviation change in the included cost characteristics changes the marginal cost by between 3.8 and 11.1% of the average car price. We therefore interpret specifications of the set $\mathcal{C}$ that allow for $\abs{\gamma_{s\ell}}\approx1\text{--}2$ (or $\abs{\gamma_{d\ell}}\approx1\text{--}2$) as allowing for moderate amounts of misspecification in the $\ell$th supply-side (or demand-side) instrument.

We follow the implementation in (ref). Given a set $I$ of potentially invalid instruments, we follow (ref) and consider sets $\mathcal{C}$ of the form (ref), with $\norm{\cdot}$ corresponding to an $\ell_{p}$ norm with $p\in\{2,\infty\}$, and $B=\tilde{B}_{I}\cdot \#I^{1/p}$, where $\tilde{B}_{I}$ is given by the columns of

equation*[equation* omitted — 202 chars of source]

and $\#{I}$ is the number of potentially invalid instruments. The scaling by $(\#{I})^{1/p}$ ensures that the vector $\gamma=M(1,\dotsc,1)'$ is always included in the set. andrews_measuring_2017 report the sensitivity of the usual GMM estimator under this form of misspecification, considering misspecification in each instrument individually (so that $I$ contains a single element), and setting $M=1$. However, if one is concerned about the validity of several instruments, it is natural to allow $I$ to contain all instruments the validity of which is questionable. In our analysis, we vary the set of potentially misspecified instruments. We also vary $M$ in order to assess the sensitivity of conclusions to different amounts of misspecification. As we will see below, different choices of $\mathcal{C}$ lead to different sensitivities for the optimal estimator, and using the optimal sensitivity can reduce the width of the CI substantially relative to CIs based on the usual GMM estimator.

We use the estimate $\hat{\theta}_{\text{initial}}$ that corresponds to the GMM estimator based on the weight matrix that's optimal under correct specification, as reported in andrews_measuring_2017, and the estimates $\hat{\Gamma}$, $\hat{H}$, and $\hat{\Sigma}$ are computed following Step (ref) of the implementation.

Results

To illustrate that using the sensitivity that is optimal under local misspecification can yield substantially tighter CIs, (ref) plots the confidence intervals based on the optimal sensitivity, as well as those based on $\hat{\theta}_{\text{initial}}$ under different sets $I$ of potentially invalid instruments and $\ell_{2}$ constraints on $\gamma$. It is clear from the figure that using the optimal sensitivity yields substantially tighter confidence intervals, relative to simply adjusting the usual CI by using the critical value $\operatorname{cv}_{\alpha}(\cdot)$ to take into account the potential bias of $h(\hat{\theta}_{\text{initial}})$, by as much as a factor of 3.4. The intuitive reason for this is that by adjusting the sensitivity of the estimator, it is possible to substantially reduce its bias at little cost in terms of an increase in variance. Thus, for example, while the CI for the average markup based on the estimate $\hat{\theta}_{\text{initial}}$ is essentially too wide to be informative when the set of potentially invalid instruments corresponds to all excluded instruments, the CI based on the optimal sensitivity, $[46.0, 66.0]\%$, is still quite tight.

If a researcher is ex ante unsure what form of misspecification one should worry about, as a sensitivity check, it is useful to consider the effects of different forms of misspecification. In (ref), we plot the optimal confidence intervals for different subsets of invalid instruments, under both $\ell_{2}$ and $\ell_{\infty}$ norms for $\gamma$. Although the choice of norm matters when the number of potentially misspecified instruments is greater than one, the results are qualitatively similar. Comparing the results for different choices of the set of potentially invalid instruments suggests that allowing supply-side instruments to be invalid generally increases the average markup estimate, while allowing demand-side instruments to be invalid has the opposite effect.

As it may be ex ante unclear what magnitude of misspecification is reasonable to allow for, as discussed in (ref), it is useful to plot the optimal CI for multiple choices of $M$. We do this in (ref) for $p=2$, and we allow all excluded instruments to be potentially invalid. One can see that while the CI is unstable for values of $M$ smaller than about $0.4$, for larger values of $M$, the estimate is quite stable and equal to about 50%. Even at $M=2$, one rejects the hypothesis that the optimal markup is equal to the initial estimate $h(\hat{\theta}_{\text{initial}})=32.7\%$. This suggests that ignoring misspecification in the BLP model likely leads to a downward bias in the estimate of the average markup. At the same time, it is possible to obtain reasonably tight CIs for the average markup even under a moderate amount of misspecification.

The $J$-statistic for testing the hypothesis that all moments are correctly specified equals 426.7. Consequently, the hypothesis is rejected at the usual significance levels. Furthermore, it can be seen from (ref) that the CIs for “all excluded” (that allow all excluded instruments to be invalid at $M=1$), and “all excluded demand” (that assume validity of supply-side instruments) do not overlap. This implies that either the misspecification in the demand-side instruments must be greater than 1% of the average care price ($M=1$), or else the supply-side instruments must also be invalid. (ref) implements the specification test from (ref) that gives lower CI $[M_{\min}, \infty]$ for $M$. The results suggest that if one assumes only a subset of the instruments is invalid, the misspecification in the potentially invalid instruments must be quite large. For example, if we assume that all instruments are valid except potentially the demand-side instruments based on rival firms' product characteristics, then the misspecification in these instruments must be greater than $M=5.36$. If we allow all instruments to be invalid, then $M\geq 1.13$.

Finally, to illustrate the implication of (ref) that one cannot substantively improve upon the CIs that we construct, we calculate the efficiency bound $\kappa_{*}$ for these CIs in (ref). The table shows that the bound is at least as high as the efficiency bound for the usual CI under correct specification (given in (ref) and equal to 84.99% at $\alpha=0.05$). Thus, the asymptotic scope for improvement over the CIs reported in (ref) at particular values of $\theta$ and $c=0$ is even smaller than the scope for improvement over the usual CI at particular values of $\theta$ under correct specification.