Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
157,665 characters · 24 sections · 0 citation commands
Minimizing Sensitivity to Model Misspecification
\vskip 3cm
\baselineskip21pt
\setcounter{page}{0}\thispagestyle{empty}
Although economic models are intended as plausible approximations to a complex economic reality, econometric inference often relies on the model being an exact description of the population environment. To account for the possibility that their models are misspecified, economists have developed a number of approaches such as specification tests, semi-parametric and nonparametric methods, and more recently bounds approaches. Implementing those approaches typically requires estimating a more general model than the original specification, possibly involving nonparametric and partially identified components.
In this paper, we consider a different approach, which consists in quantifying how model misspecification affects the parameter of interest, and in modifying the estimate in order to minimize the impact of misspecification. The goal of the analysis is twofold. First, we provide simple adjustments, which do not require re-estimating the model, and provide guarantees on performance when the model is misspecified. Second, we construct confidence intervals that account for model misspecification error in addition to sampling uncertainty.
In our approach, we consider deviations from a {reference specification} of the model{, in a particular class}. The reference model is parametric and fully specified given covariates. It may, for example, correspond to the empirical specification of a structural economic model. We do not assume that the reference model is correctly specified, and allow for local deviations from it within a larger class of models. Relative to other approaches, a local analysis presents important advantages in terms of tractability.
We construct {minimax} estimators which minimize worst-case mean squared error (MSE) in a given neighborhood of the reference model. The worst case is influenced by the directions of model misspecification which matter most for the parameter of interest. We focus in particular on two types of neighborhoods, for two leading classes of applications: Euclidean neighborhoods, in settings where the larger class of models containing the reference specification is parametric, and Kullback-Leibler neighborhoods, in semi-parametric mixture models where misspecification of functional forms is measured by the Kullback-Leibler divergence between density functions.
The framework we propose is inspired by Hansen and Sargent's (2001, 2008) work on robust decision making under uncertainty and ambiguity. As in their work, optimal decisions depend on the size of the neighborhood around the reference model. In this paper, we do not attempt to provide a data-driven choice for the neighborhood size. Instead, we take the size as given and derive formulas for optimal estimation in neighborhoods of a given size. We discuss how to interpret the magnitude of the neighborhood size in various parametric and semi-parametric examples. In addition, we show that the neighborhood size can be mapped to the local power --- in certain directions --- of a likelihood-ratio test of correct specification of the reference model.
Our approach delivers a class of estimators that can be used for systematic sensitivity analysis. In addition, we show how to construct confidence intervals which asymptotically contain the population parameter of interest with pre-specified probability, both under correct specification and local misspecification. We show that acknowledging misspecification leads to easy-to-compute enlargements of conventional confidence intervals. Such confidence intervals are “honest”, in the sense that they account for the bias of the estimator (e.g., Donoho, 1994, Armstrong and Koles\'ar, 2020).
Our local approach leads to tractable expressions for worst-case bias and MSE, as well as for minimum-MSE estimators in a given neighborhood of the reference model. A minimum-MSE estimator takes the form of a one-step adjustment of the estimator based on the reference model by a term which reflects the impact of model misspecification, in addition to a more standard term which adjusts the estimate in the direction of the efficient estimator based on the reference model. Implementing the optimal estimator only requires computing the score and Hessian of a larger model, evaluated at the reference model. The large model never needs to be estimated. This feature of our approach is reminiscent of the logic of Lagrange Multiplier (LM) testing.
We illustrate our approach using three examples. We first study the evaluation of the PROGRESA program in Mexico, which provides income transfers to households subject to the condition that the child attends school. Todd and Wolpin (2006) estimate a structural model of education choice on villages that were initially randomized out. They compare the predictions of the structural model with the estimated experimental impact. As emphasized by Todd and Wolpin (2008) and Attanasio et al. (2012), the ability to predict the effects of the program based solely on control villages imposes restrictions on the economic model. Within a simple static model of education choice, we assess the sensitivity of counterfactual predictions to a form of misspecification under which program participation may have a direct “stigma” effect on the marginal utility of schooling (Wolpin, 2013).
We next study the impact of misspecification of the error distribution in a cross-sectional binary choice model. Our aim is to estimate the outcome probabilities under different values of the covariates. While point-identification can be achieved under independence and sufficiently rich support of covariates (Manski, 1988), the quantities of interest are partially identified in our setting. Relying on a normal (probit) reference model, we show how our estimators and confidence intervals can be used for sensitivity analysis, when the researcher is concerned about misspecification of the normal distribution.
Our third and last example is a dynamic binary choice model in short panel data. We assume that time-varying errors are i.i.d. normal, but leave the distribution of individual heterogeneity given initial conditions unrestricted. In this setting also, common parameters and average effects often fail to be point-identified (Chamberlain, 2010, Honor\'e and Tamer, 2006, Chernozhukov et al., 2013), thus motivating a sensitivity analysis approach. We show that minimizing worst-case MSE in such panel data settings leads to a Tikhonov-regularized estimator, where the penalization reflects the degree of misspecification allowed for. In simulations, we illustrate that our estimator can provide substantial bias and MSE reduction relative to commonly used estimators.
\paragraph{Related work and outline.}
As in the literature on robust statistics (Huber, 1964, Huber and Ronchetti, 2009, Hampel et al., 1986, and especially Rieder, 1994), we rely on a minimax approach and aim to minimize the worst-case impact of misspecification in a neighborhood of a model. A difference with this work is that we focus on misspecification of {specific aspects} of a model, by considering {parametric or semi-parametric} classes of models around the reference specification. By contrast, the robust statistics literature has mostly focused on fully {nonparametric} classes, motivated by data contamination issues.
A related literature studies orthogonalization and locally robust moment functions; see Neyman (1959), Newey (1994), Chernozhukov et al. (2018), Chernozhukov et al. (2020), and also Fraser (1964). Here we account for both bias and variance, weighting them by the size of the neighborhood around the reference model. In addition, our approach does not require the larger model to be point-identified. Our analysis also connects to Bayesian robustness (e.g., Berger and Berliner, 1986, Gustafson, 2000, Vidakovic, 2000, Mueller, 2012), although our minimum-MSE estimators and confidence intervals have a frequentist interpretation.
Also related are the literatures on statistical decision theory (e.g., Wald, 1950, Chamberlain 2000, Watson and Holmes, 2016, Hansen and Marinacci, 2016, and especially Hansen and Sargent, 2008) and the literature on sensitivity analysis in statistics and economics (e.g., Rosenbaum and Rubin, 1983, Leamer, 1985, Imbens, 2003, Altonji et al., 2005, Nevo and Rosen, 2012, Oster, 2019, Masten and Poirier, 2020, 2021). Our analysis of minimum-MSE estimation and sensitivity in the OLS/IV example is related to Hahn and Hausman (2005) and Angrist et al. (2017). Our approach based on local misspecification has a number of precedents, such as Newey (1985), Conley et al. (2012), Guggenberger (2012), Bugni et al. (2012), Kitamura et al. (2013), and Bugni and Ura (2019). Also related is Claeskens and Hjort's (2003) work on the focused information criterion.
Recent papers rely on a local approach to misspecification to provide tools for sensitivity analysis. Andrews et al. (2017) propose a measure of sensitivity of parameter estimates to the moments used in estimation. Andrews et al. (2020) introduce a measure of informativeness of descriptive statistics in the estimation of structural models; see also Mukhin (2018). Our goal is different, in that we aim to provide a framework for estimation and inference in the presence of misspecification. Armstrong and Koles\'ar (2021) study models defined by over-identified systems of moment conditions that are approximately satisfied at true values, up to an additive term that vanishes asymptotically, and derive results for optimal estimation and inference. In this paper, we seek to ensure robustness to misspecification of a reference model within a larger class of models.
Our focus on {specific} forms of model misspecification also relates to recent approaches to estimate partially identified models (Chen et al., 2011, Norets and Tang, 2014, Schennach, 2013, Giacomini and Kitagawa, 2021). Christensen and Connault (2019) consider structural models defined by equilibrium conditions, and develop inference methods on the identified set of counterfactual predictions subject to restrictions on the distance between the true model and a reference specification. Our local approach is complementary to these methods. It allows tractability in complex models, such as structural economic models, since implementation does not require estimating a larger model. In our framework, we view the parametric reference model as a useful benchmark, although its predictions need to be modified in order to minimize the impact of misspecification. This aspect relates our paper to shrinkage methods (e.g., Hansen, 2016, 2017, Fessler and Kasy, 2019, Maasoumi, 1978), with the difference that here we are interested in a single parameter.
The plan of the paper is as follows. In Section (ref), we describe our framework and derive the main results. In Section (ref), we apply our framework to parametric and semi-parametric mixture models. In Section (ref), we discuss how to use our approach for sensitivity analysis, with a focus on the interpretation of neighborhood size. In Sections (ref) and (ref), we present our illustrations. Finally, we conclude in Section (ref). The supplementary material available \href{https://sites.google.com/site/stephanebonhommeresearch/}{online} contains an appendix and codes for replication.
In this section, we describe the main elements of our approach in a general setting. In the next section, we will specialize the analysis to a locally quadratic setting, which includes both parametric misspecification and semi-parametric misspecification of distributional functional forms.
We observe a random sample $( Y_i \, : \, i=1,\ldots,n)$ from a density $f_{{\Greekmath 010C},{\Greekmath 0119}}(y)$ (with respect to a continuous or discrete measure), where ${\Greekmath 010C} \in {\cal B}$ is a finite-dimensional parameter, and ${\Greekmath 0119} \in \Pi$ is a finite- or infinite-dimensional parameter. Throughout the paper, the parameter of interest is ${\Greekmath 010E}_{{\Greekmath 010C},{\Greekmath 0119}} $, a scalar function or functional of ${\Greekmath 010C}$ and ${\Greekmath 0119}$. We assume that ${\Greekmath 010E}_{{\Greekmath 010C},{\Greekmath 0119}}$ and $f_{{\Greekmath 010C},{\Greekmath 0119}}$ are known, smooth functions of ${\Greekmath 010C}$ and ${\Greekmath 0119}$. Examples of functionals of interest in economic applications include counterfactual policy effects in structural models, and average effects in panel data settings. The true parameter values ${\Greekmath 010C}_0$ and ${\Greekmath 0119}_0$ that generate the observed data $Y_1,\ldots,Y_n$ are unknown to the researcher. Our goal is to estimate ${\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}$ and construct confidence intervals for it. We abstract from covariates to simplify the presentation, but it is straightforward to extend our results to conditional models of the form $f_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}(y\,|\, x)$; see Subsection (ref).
Our starting point is that the researcher has chosen a reference model ${\Greekmath 0119}({\Greekmath 010D})$, which parameterizes the unknown ${\Greekmath 0119} \in \Pi$ in terms of a finite-dimensional parameter ${\Greekmath 010D} \in {\mathcal G} $. We say that the reference model is {correctly specified} if there exists a value ${\Greekmath 010D} \in {\mathcal G}$ such that ${\Greekmath 0119}_0 = {\Greekmath 0119}({\Greekmath 010D})$. Otherwise, we say that the model is {misspecified}. {To measure misspecification, we rely on a distance measure $d$ on $\Pi$, and we denote the maximal amount of misspecification as ${\Greekmath 010F}\geq 0$.}
In our theory, we consider an asymptotic sequence where ${\Greekmath 010F} = {\Greekmath 010F}_n $ tends to zero as $n $ tends to infinity, so the maximal amount of misspecification gets smaller as the sample size increases. The reason for focusing on ${\Greekmath 010F}$ tending to zero is tractability, as a small-${\Greekmath 010F}$ analysis allows us to rely on linearization techniques and obtain simple, explicit expressions. Moreover, when estimating ${\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}$, the estimation bias due to misspecification (of order ${\Greekmath 010F}^{1/2}$) and the standard deviation (of order $n^{-1/2}$) are asymptotically comparable, so both play a role in the mean squared error. This local asymptotic approach has a number of precedents in the literature, notably Rieder (1994). Along the sequence, the true parameter ${\Greekmath 0119}_0 = {\Greekmath 0119}_{0,n}$ depends on $n$, and we assume that, for a fixed parameter ${\Greekmath 010D}_* $, $d( {\Greekmath 0119}_{0,n} , {\Greekmath 0119}({\Greekmath 010D}_*) )\leq {\Greekmath 010F}_n$ for all $n$. This implies that $\lim_{n \rightarrow \infty} d( {\Greekmath 0119}_{0,n}, {\Greekmath 0119}({\Greekmath 010D}_*) ) = 0$; that is, ${\Greekmath 0119}_{0,n}$ converges to ${\Greekmath 0119}({\Greekmath 010D}_*)$ as $n $ tends to infinity. Hereafter we drop the indices $n$ and do not make the sample size dependence of ${\Greekmath 010F}$ and ${\Greekmath 0119}_0$ explicit. For example, we simply write $d( {\Greekmath 0119}_0 , {\Greekmath 0119}({\Greekmath 010D}_*) )\leq {\Greekmath 010F}$.
Given the distance measure $d$, and some ${\Greekmath 010F}>0$, we define an ${\Greekmath 010F}$-{neighborhood} around ${\Greekmath 0119}({\Greekmath 010D}_*)$ as
We assume that the true ${\Greekmath 0119}_0$ that generates the data satisfies ${\Greekmath 0119}_0 \in \Gamma_{{\Greekmath 010F}}({\Greekmath 010D}_*) $. Later we will assume that ${\Greekmath 010D}_*$ can be estimated consistently by some preliminary estimator $\widehat {\Greekmath 010D} $. The distance measure $d$, the misspecification bound ${\Greekmath 010F}$, and the preliminary estimator $\widehat {\Greekmath 010D}$ are chosen by the researcher.\footnote{Instead of defining the ${\Greekmath 010F}$-neighborhood of misspecified models around a fixed point ${\Greekmath 0119}({\Greekmath 010D}_*)$, one could alternatively consider all ${\Greekmath 0119}_0$ in the set $\Gamma_{{\Greekmath 010F}} = \cup_{{\Greekmath 010D} \in {\cal G}} \, \Gamma_{{\Greekmath 010F}}({\Greekmath 010D})$, which is the ${\Greekmath 010F}$-neighborhood around the manifold of reference models ${\Greekmath 0119}({\Greekmath 010D})$, ${\Greekmath 010D} \in {\cal G}$. This alternative definition would avoid having to define ${\Greekmath 010D}_*$, and we employed it in the first version of this paper to justify the same local approximation to the worst-case MSE optimal estimator derived below (see Subsection 2.3 in Bonhomme and Weidner, 2018). In the current presentation, we only introduce the ${\Greekmath 010F}$-neighborhood $\Gamma_{{\Greekmath 010F}}({\Greekmath 010D}_*)$ around a fixed ${\Greekmath 010D}_*$, however this has no effect on the minimax misspecification adjustments that we derive. By fixing ${\Greekmath 010D}_*$, this presentation also aligns with the way in which local minimax results are typically discussed in statistics (see, e.g., Theorem 8.11 in Van der Vaart, 2007).}
\paragraph{Examples.} As a first example, consider a parametric model defined by Euclidean parameters ${\Greekmath 010C}$ and ${\Greekmath 0119}$, where ${\Greekmath 0119}=0$ under the reference model. For example, ${\Greekmath 0119}$ can represent the effect of an omitted control variable in a regression, or the degree of endogeneity of a regressor as in the example we analyze in Subsection (ref). Suppose that the researcher is interested in the parameter ${\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}=c'{\Greekmath 010C}_0$ for a known vector $c$, such as one component of ${\Greekmath 010C}_0$. In this case, we will take the weighted Euclidean (squared) distance $d({\Greekmath 0119}_0,{\Greekmath 0119})=\|{\Greekmath 0119}_0-{\Greekmath 0119}\|_{\Omega}^2=({\Greekmath 0119}_0-{\Greekmath 0119})'\Omega ({\Greekmath 0119}_0-{\Greekmath 0119})$, for a positive-definite matrix $\Omega$.
As a second example, consider a semi-parametric mixture model whose likelihood depends on a finite-dimensional parameter vector ${\Greekmath 010C}$ and a nonparametric density ${\Greekmath 0119}$ of unobservables $A\in{\cal{A}}$, abstracting from conditioning covariates for simplicity. The joint density of $(Y,A)$ is $g_{{\Greekmath 010C}_0}(y\,|\, a){\Greekmath 0119}_0(a)$, for some known function $g$. Suppose that the researcher's goal is to estimate an average effect ${\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}=\mathbb{E}_{{\Greekmath 0119}_0}\Delta(A,{\Greekmath 010C}_0)$, for a known function $\Delta$. It is common to estimate the model by parameterizing the unknown density as ${\Greekmath 0119}({\Greekmath 010D})$, where ${\Greekmath 010D}$ is finite-dimensional. We focus on situations where, although the researcher thinks of ${\Greekmath 0119}({\Greekmath 010D})$ as a plausible approximation to the population distribution ${\Greekmath 0119}_0$, she is not willing to rule out that it may be misspecified. In this case we use the Kullback-Leibler divergence to define semi-parametric neighborhoods, and we take $d({\Greekmath 0119}_0,{\Greekmath 0119})=2\int_{\cal{A}} \log \left(\frac{{\Greekmath 0119}_0(a)}{{\Greekmath 0119}(a)}\right){\Greekmath 0119}_0(a)da$. $\square$
\vskip .3cm
We focus on asymptotically linear estimators $\widehat{{\Greekmath 010E}}=\widehat{{\Greekmath 010E}}(Y_1,\ldots,Y_n)$ that admit a stochastic expansion of the form
where this expansion holds uniformly for all $P_0=P_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}$ such that ${\Greekmath 0119}_0\in\Gamma_{{\Greekmath 010F}}({\Greekmath 010D}_*)$, in a sense that we will discuss below and make precise in Theorem (ref). Along the sequence we consider, the product ${\Greekmath 010F} n$ tends to a positive constant, so the remainder in ((ref)) is $o_{P_0}(n^{-\frac{1}{2}} )$. Although asymptotic linearity is satisfied by many econometric estimators, it can fail in certain semi-parametric problems (e.g., Cattaneo et al., 2014) and in problems involving model selection or shrinkage (e.g., Liao, 2013, Cheng and Liao, 2015), for example.
Equation ((ref)) is a form of local regularity of the estimator $\widehat{{\Greekmath 010E}}$. Consider first the correctly specified case, where ${\Greekmath 010F}=0$ and ${\Greekmath 0119}_0 = {\Greekmath 0119}({\Greekmath 010D}_*)$. In this case $h(\cdot,{\Greekmath 010C}_0,{\Greekmath 010D}_*)$ is the influence function of $\widehat{{\Greekmath 010E}}$. We assume that the following conditions are satisfied,
and
where $\mathbb{E}_{{\Greekmath 010C},{\Greekmath 0119}}$ denotes the expectation under $f_{{\Greekmath 010C},{\Greekmath 0119}}$, and ${\Greekmath 0272}_{{\Greekmath 010C}{\Greekmath 010D}}$ denotes the derivative with respect to the vector $({\Greekmath 010C}',{\Greekmath 010D}')'$. Both (ref) and (ref) are standard properties of influence functions of regular asymptotically linear estimators.
We will refer to (ref) as unbiasedness, since it guarantees that $ \widehat{{\Greekmath 010E}}$ is asymptotically unbiased for ${\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}$ under correct specification of the reference model. We assume that unbiasedness holds at all possible values of ${\Greekmath 010C}_0$ and ${\Greekmath 010D}_*$. Then, by differentiating (ref) with respect to ${\Greekmath 010C}_0$ and ${\Greekmath 010D}_*$ and plugging the resulting equations into (ref), we obtain
Under unbiasedness, (ref) and (ref) are equivalent. We will later work with (ref), since it only features $h(Y, {\Greekmath 010C}_0,{\Greekmath 010D}_*)$ and not its gradient. Under suitable conditions, (ref) is necessary and sufficient for the asymptotically linear estimator $\widehat{{\Greekmath 010E}}$ to be regular; see, e.g., Newey (1990). As an example, for m-estimators, (ref) can be interpreted as the generalized information matrix equality. Asymptotic linearity and regularity are commonly imposed in the semi-parametric efficiency literature (Bickel et al., 1993). These conditions rule out, for example, superefficient estimators such as Hodges' estimator. We will refer to (ref), or alternatively (ref), as local robustness, using a terminology introduced by Chernozhukov et al. (2020).\footnote{ While in Chernozhukov et al. (2020) local robustness is imposed as a substantive restriction on more general moment functions, in our setting ((ref)) and ((ref)) are regularity conditions given unbiasedness.}
Consider now the misspecified case, where ${\Greekmath 010F}>0$. In this case, we strengthen the condition of asymptotic linearity, and require that $\widehat{{\Greekmath 010E}}$ be locally asymptotically linear; see, e.g., Klaassen (1987). Formally, under local, small-${\Greekmath 010F}$ misspecification, we assume the stochastic expansion (ref) continues to hold, but now uniformly for all ${\Greekmath 0119}_0\in\Gamma_{{\Greekmath 010F}}({\Greekmath 010D}_*)$.\footnote{To see why (ref) is a plausible way of imposing asymptotic linearity here, let ${\Greekmath 011E}(Y_i,{\Greekmath 010C}_0,{\Greekmath 0119}_0)$ be the influence function of $ \widehat{{\Greekmath 010E}}$. Expanding as $n\rightarrow\infty$ and ${\Greekmath 010F}\rightarrow 0$ we have
In this expansion the term linear in ${\Greekmath 0119}_0 - {\Greekmath 0119}({\Greekmath 010D}_*)$ vanishes, whenever ${\Greekmath 011E}(Y,{\Greekmath 010C}_0,{\Greekmath 0119}_0)$ satisfies an influence function regularity condition analogous to (ref), and the term quadratic in ${\Greekmath 0119}_0 - {\Greekmath 0119}({\Greekmath 010D}_*)$ gives a contribution $ o_{P_0}({\Greekmath 010F}^{\frac{1}{2}}) $. } In the following, we focus on locally asymptotically linear estimators that satisfy (ref), under the conditions (ref) and (ref). Notice that, under local misspecification, the influence function $h(Y,{\Greekmath 010C}_0,{\Greekmath 010D}_*)$ has no longer mean zero under $P_0$ in general.
Our goal in this paper is twofold. First, we will construct confidence intervals for the target parameter ${\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}$ which are uniformly asymptotically valid on $\Gamma_{{\Greekmath 010F}}({\Greekmath 010D}_*)$. Second, an important goal of the analysis is to construct estimators $\widehat{{\Greekmath 010E}}=\widehat{{\Greekmath 010E}}(Y_1,\ldots,Y_n)$ that are asymptotically optimal in a minimax sense. For this purpose, we will show how to compute a function $h$ such that the (trimmed) worst-case mean squared error (MSE) $\mathbb{E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0} [ ( \widehat {\Greekmath 010E} - {\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0} )^2 ]$ in the ${\Greekmath 010F}$-neighborhood $\Gamma_{{\Greekmath 010F}}({\Greekmath 010D}_*)$ of the reference model, among estimators of the form
is minimized under our local asymptotic analysis. In fact, we will show how to compute estimators that minimize (trimmed) worst-case MSE among asymptotically linear estimators; see Theorem (ref) below for a precise statement. Here $\widehat{{\Greekmath 010C}}$ and $\widehat{{\Greekmath 010D}}$ are preliminary estimators of ${\Greekmath 010C}_0$ and ${\Greekmath 010D}_*$ that are {root-$n$} consistent under correct specification. For example, $\widehat{{\Greekmath 010C}}$ and $\widehat{{\Greekmath 010D}}$ may be maximum likelihood estimators (MLE) based on the reference model. It follows from (ref) and (ref) that, under regularity conditions on the preliminary estimators, the form of the minimum-MSE $h$ function is not affected by the choice of $\widehat{{\Greekmath 010C}}$ and $\widehat{{\Greekmath 010D}}$.
\paragraph{Examples (cont.)}
In our parametric example, a natural estimator is the MLE of $c'{\Greekmath 010C}_0$ based on the reference specification, such as the OLS estimator under the assumption that ${\Greekmath 0119}=0$; e.g., that the coefficient of an omitted control variable is zero. In a correctly specified likelihood setting, such an estimator will be consistent and efficient. However, when the reference model is misspecified, it may be dominated in terms of bias or MSE by other regular estimators.
In our semi-parametric mixture example, a commonly used (“random-effects”) estimator of ${\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}=\mathbb{E}_{{\Greekmath 0119}_0}\Delta(A,{\Greekmath 010C}_0)$ is obtained by replacing the population average by an integral with respect to the parametric distribution ${\Greekmath 0119}(\widehat{{\Greekmath 010D}})$, where $\widehat{{\Greekmath 010D}}$ is the MLE of ${\Greekmath 010D}$. Another popular (“empirical Bayes”) estimator is obtained by substituting an integral with respect to the posterior distribution of $A$ based on ${\Greekmath 0119}(\widehat{{\Greekmath 010D}})$. We will compare the finite-sample performance of these estimators to that of our minimum-MSE estimator in our panel data illustration in Section (ref). $\square$
In this subsection, we provide heuristic derivations for the worst-case bias and the minimum-MSE estimator. This will lead to the main expressions in equations ((ref)) and ((ref)) below. In the next subsection, we will provide regularity conditions under which these derivations are formally justified.
We assume that $\Gamma_{\Greekmath 010F}({\Greekmath 010D}_*)$ is a convex set. For any linear map $u:\Pi \rightarrow \mathbb{R}$, we define\footnote{The definition in ((ref)) is stated for finite-dimensional ${\Greekmath 0119}$. This definition can be generalized to the case where ${\Greekmath 0119}$ is infinite-dimensional; see Appendix (ref) and Section (ref). }
We assume that the distance measure $d$ is chosen such that $ \left\| \cdot \right\|_{{\Greekmath 010D}_*}$ is unique and well-defined, and that it is a norm, dual to a local approximation to $d( {\Greekmath 0119}_0, {\Greekmath 0119}({\Greekmath 010D}_*))$ for fixed ${\Greekmath 0119}({\Greekmath 010D}_*)$. Both our examples of distance measures -- weighted Euclidean distance and Kullback-Leibler divergence -- satisfy these assumptions.
We focus on estimators $\widehat{{\Greekmath 010E}}$ that satisfy ((ref)) for a suitable $h$ function for which (ref) and (ref) hold. Under appropriate regularity conditions, the worst-case bias of $\widehat{{\Greekmath 010E}}$ in the neighborhood $\Gamma_{\Greekmath 010F}({\Greekmath 010D}_*)$ can be expanded for small ${\Greekmath 010F}$ and large $n$ as
where
for $\|\cdot\|_{{\Greekmath 010D}_*}$ the dual norm defined in ((ref)).\footnote{When ${\Greekmath 0119}$ is infinite-dimensional ${\Greekmath 0272}_{{\Greekmath 0119}}$ denotes a G\^ateaux derivative.}
Then, the worst-case MSE in $\Gamma_{\Greekmath 010F}({\Greekmath 010D}_*)$ can be expanded as follows, again under appropriate regularity conditions (see Lemma (ref) in the appendix),
We therefore define the {minimum-MSE} function $h_{{\Greekmath 010F}}^{\rm MMSE}(y,{\Greekmath 010C}_0,{\Greekmath 010D}_*)$ as
Finally, let $\widehat {\Greekmath 010C}$ and $\widehat {\Greekmath 010D}$ be preliminary estimators that are consistent for ${\Greekmath 010C}_0$ and ${\Greekmath 010D}_*$ under the reference model $f_{{\Greekmath 010C}_0,{\Greekmath 0119}({\Greekmath 010D}_*)}$. Then, the {minimum-MSE} estimator of ${\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}$ is given by
This estimator minimizes an asymptotic approximation to the worst-case MSE in $\Gamma_{\Greekmath 010F}({\Greekmath 010D}_*)$. Using a small-${\Greekmath 010F}$ approximation is crucial for analytic tractability, since the variance term in ((ref)) only needs to be calculated under the reference model, and the optimization problem (ref) is convex. In practice, ((ref)) only needs to be solved at $ \widehat {\Greekmath 010C}$ and $ \widehat{{\Greekmath 010D}}$. In addition, as we already pointed out, the form of the minimum-MSE estimator is not affected by the choice of the preliminary estimators $\widehat {\Greekmath 010C}$ and $\widehat{{\Greekmath 010D}}$.
The constraints on $h_{{\Greekmath 010F}}^{\rm MMSE}( \cdot ,{\Greekmath 010C}_0,{\Greekmath 010D}_*) $ imposed in (ref) are the unbiasedness condition (ref) and the local robustness condition ((ref)). As we discussed above, given unbiasedness, local robustness is a regularity condition, and unbiasedness is a substantive condition that implies that our estimator $ \widehat {\Greekmath 010E}\,^{\rm MMSE}_{\Greekmath 010F} $ is only optimal within the class of estimators that are asymptotically unbiased for ${\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}$ under the reference model.
\paragraph{Special cases.}
To provide intuition about the minimum-MSE function $h^{\rm MMSE}_{{\Greekmath 010F}}$, let us define two Hessian matrices $H_{{\Greekmath 010C}{\Greekmath 010D}}$, of size $\limfunc{dim}{\Greekmath 010C}+\limfunc{dim}{\Greekmath 010D}$, and ${\cal{H}}_{{\Greekmath 010C}{\Greekmath 0119}}$, of size $\limfunc{dim}{\Greekmath 010C}+\limfunc{dim}{\Greekmath 0119}$, as\footnote{The definition of ${\cal{H}}_{{\Greekmath 010C}{\Greekmath 0119}}$ generalizes to the infinite-dimensional ${\Greekmath 0119}$ case; see Appendix (ref) and Section (ref). }
Throughout our analysis, we assume that $H_{{\Greekmath 010C}{\Greekmath 010D}}$ is invertible. This requires that the Hessian matrix of the parametric reference model be non-singular, thus requiring that ${\Greekmath 010C}_0$ and ${\Greekmath 010D}_*$ be identified under the reference model. When ${\Greekmath 010F}=0$ we find that
Thus, under the assumption that the parametric reference model is correctly specified, $\widehat {\Greekmath 010E}^{\rm MMSE}_{\Greekmath 010F}$ is simply the one-step approximation to the MLE for ${\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}$ that maximizes the likelihood with respect to the “small” parameter $({\Greekmath 010C}',{\Greekmath 010D}')'$. This “one-step efficient” adjustment is purely based on efficiency considerations. Such one-step approximations are classical estimators in statistics (e.g., Bickel et al., 1993).
Another interesting special case of the minimum-MSE $h$ function arises in the limit ${\Greekmath 010F} \rightarrow \infty$, when the matrix or operator ${\cal{H}}_{{\Greekmath 010C}{\Greekmath 0119}}$ is invertible. Note that invertibility of ${\cal{H}}_{{\Greekmath 010C}{\Greekmath 0119}}$, which may fail when ${\Greekmath 0119}_0$ is not identified, is not needed in our analysis and we only use it to analyze this limiting case. We then have that
In this limit $\widehat {\Greekmath 010E}\,^{\rm MMSE}_{\Greekmath 010F}$ is simply the one-step approximation to the MLE for ${\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}$ that maximizes the likelihood with respect to the “large” parameter $({\Greekmath 010C}',{\Greekmath 0119}')'$. For any ${\Greekmath 010F}$, the estimator $ \widehat {\Greekmath 010E}\,^{\rm MMSE}_{\Greekmath 010F} $ is a nonlinear interpolation between the one-step MLE approximation to the parametric reference model and the one-step MLE approximation to the large model. We obtain one-step approximations in our approach, since (ref) is only a {local} approximation to the full MSE-minimization problem.
However, an estimator based on ((ref)) may be ill-behaved in non point-identified problems, or in problems where the identification of ${\Greekmath 0119}_0$ is irregular. By contrast, $\widehat {\Greekmath 010E}\,^{\rm MMSE}_{\Greekmath 010F}$ is always well-defined, since the variance of $h(Y,{\Greekmath 010C}_0,{\Greekmath 010D}_*)$ acts as a sample size-dependent regularization. The form of $\widehat {\Greekmath 010E}\,^{\rm MMSE}_{\Greekmath 010F}$ is thus based on {both} efficiency and robustness. In addition, note that, while neither ((ref)) nor ((ref)) involve the particular choice of distance measure with respect to which neighborhoods are defined, for given ${\Greekmath 010F}>0$ the minimum-MSE estimator will depend on the chosen distance measure.
Lastly, it is common in applications with covariates to model the conditional distribution of outcomes $Y$ given covariates $X$ as $f_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}(y\,|\, x)$, while leaving the marginal distribution of $X$, $f_X(x)$, unspecified. Our approach can easily be adapted to deal with such {conditional} models, as we will describe in Subsection (ref) in a locally quadratic setting.
In this subsection, we provide a formal characterization of the minimum-MSE estimator by showing that it achieves minimum worst-case MSE in a class of regular asymptotically linear estimators, as $n$ tends to infinity and ${\Greekmath 010F} n$ tends to a constant. All sequences can thus be equivalently indexed by ${\Greekmath 010F}$ or $n$; for example, $h_{{\Greekmath 010F}}$ in the following theorem could equivalently be indexed by $n$. Moreover, under the stated assumptions, the heuristic derivations of the previous subsection are formally justified. All proofs are in the appendix.
\vskip .3cm
We establish Theorem (ref) in a joint asymptotic sequence where ${\Greekmath 010F} $ tends to zero as $n$ tends to infinity and ${\Greekmath 010F} n $ tends to a finite positive constant. Under this sequence, the leading term in the worst-case MSE is of order ${\Greekmath 010F}$ (squared bias), or equivalently of order $1/n$ (variance). The theorem considers a trimmed MSE to allow for the possibility that the estimators for $ {\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}$ do not have moments. The trimming cutoff $m_n$ shrinks to zero at a rate slower than $n^{-1/2}$ (or equivalently ${\Greekmath 010F}^{1/2}$), so that for estimators without heavy tails the leading-order bias and standard deviation should not be affected by the trimming. The theorem states that the leading-order worst-case trimmed MSE achieved by our minimum-MSE estimator $\widehat {\Greekmath 010E}^{\, \rm MMSE}_{{\Greekmath 010F}}$ is at least as good as the one achieved by any other sequence of estimators satisfying our regularity conditions. All the assumptions on $\widehat {\Greekmath 010E}_{{\Greekmath 010F}}$ and $h_{{\Greekmath 010F}}(\cdot, {\Greekmath 010C},{\Greekmath 010D})$ that we require for this result are listed in the statement of the theorem. Note that in Theorem (ref) we {assume} that ((ref)) holds, subject to (ref) and (ref). On might conjecture that the result holds absent these conditions. However, our current proof crucially relies on them.\footnote{Condition ((ref)) is a form of local regularity of the sequence of estimators $\widehat{{\Greekmath 010E}}_{{\Greekmath 010F}}$. The additional regularity conditions in Assumptions (ref) and (ref) are smoothness conditions on $f_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}(y)$, ${\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}$, ${\Greekmath 0119}({\Greekmath 010D})$, and $d({\Greekmath 0119}_0,{\Greekmath 0119}({\Greekmath 010D}))$ as functions of ${\Greekmath 010C}_0$, ${\Greekmath 0119}_0$, and ${\Greekmath 010D}$, and an appropriate rate condition on the preliminary estimators $\widehat {\Greekmath 010C}$ and $\widehat {\Greekmath 010D}$. In particular, in Assumption (ref), we require the preliminary estimators $\widehat {\Greekmath 010C}$ and $\widehat {\Greekmath 010D}$ to have moments of order larger than two. This may require modifying the preliminary estimators to ensure that they have finite moments, as in for example Hausman et al. (2011), who focus on GMM estimators.} {We will provide explicit expressions for $h_{{\Greekmath 010F}}^{\rm MMSE}$ in various models in the next section.}
In addition to point estimates, our framework allows us to compute confidence intervals that contain ${\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}$ with pre-specified probability under our local asymptotic analysis. To see this, let $\widehat{{\Greekmath 010E}}$ be an estimator satisfying ((ref)), ((ref)), and ((ref)). For a given confidence level ${\Greekmath 010B}\in (0,1)$, let us define the following interval
where $b_{{\Greekmath 010F}}\left(\cdot\right)$ is given by ((ref)), $\widehat{{\Greekmath 011B}}_h^2$ is the sample variance of $h(Y_1,\widehat{{\Greekmath 010C}},\widehat{{\Greekmath 010D}}),\ldots , h(Y_n,\widehat{{\Greekmath 010C}},\widehat{{\Greekmath 010D}})$, and $c_{1-{\Greekmath 010B}/2}=\Phi^{-1}(1-{\Greekmath 010B}/2)$ is the $(1-{\Greekmath 010B}/2)$-standard normal quantile. Under suitable regularity conditions, the interval $CI_{{\Greekmath 010F}}(1-{\Greekmath 010B},\widehat{{\Greekmath 010E}})$ contains ${\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}$ with probability {at least} $1-{\Greekmath 010B}$ as $n$ tends to infinity and ${\Greekmath 010F} n$ tends to a constant, both under correct specification and under local misspecification of the reference model. Formally, we have the following result.
\vskip .3cm
Such “fixed-length” confidence intervals, which take into account both misspecification bias and sampling uncertainty, have been studied in different contexts (e.g., Donoho, 1994, Armstrong and Koles\'ar, 2020, 2021).\footnote{A variation suggested by these authors, which reduces the length of the interval, is to compute the interval as $\widehat{{\Greekmath 010E}}\pm$ $b_{{\Greekmath 010F}}(h,\widehat{{\Greekmath 010C}},\widehat{{\Greekmath 010D}})$ times the $(1-{\Greekmath 010B})$-quantile of $\left|{\cal{N}}\left(1,\widehat{{\Greekmath 011B}}_h^2/(nb_{{\Greekmath 010F}}(h,\widehat{{\Greekmath 010C}},\widehat{{\Greekmath 010D}})^2)\right)\right|$.}
In this section, we first derive explicit expressions for minimum-MSE estimators in a class of models that have a locally quadratic structure; that is, where the dual norm $\|\cdot\|_{{\Greekmath 010D}_*}$ can be associated with an inner product. We then apply these results to parametric models and semi-parametric mixture models. Our motivation for focusing on these settings is that they allow us to develop practical implementation methods. We have not explored implementation in other models with infinite-dimensional ${\Greekmath 0119}$ parameters.
Formally, let us define the tangent space $\overline{{\cal{T}}}$ of the parameter space $\Pi$ at ${\Greekmath 0119}({\Greekmath 010D}^*)$, where for simplicity we ignore the dependence on ${\Greekmath 010D}^*$ in the notation.\footnote{The tangent space at ${\Greekmath 0119}({\Greekmath 010D}^*)$ includes all directions in which one can pass tangentially through ${\Greekmath 0119}({\Greekmath 010D}^*)$. See Barden and Thomas (2003) for a formal exposition.} Let us then define the cotangent space ${\cal T}$ as the set of linear maps $u:\overline{{\cal{T}}}\rightarrow \mathbb{R}$. Throughout this section, we assume that ${\cal T}$ is a Hilbert space equipped with the norm $\left\|\cdot\right\|_{{\Greekmath 010D}_*}$. This locally quadratic structure characterizes our two leading examples of parametric and semi-parametric mixture models. In such cases the tangent space (at an interior ${\Greekmath 0119}({\Greekmath 010D}^*)$) is simply $\Pi$, and the cotangent space is the set of linear maps $u:\Pi\rightarrow \mathbb{R}$.
Consider the case where the square of the local dual norm defined in (ref) can be written as $\left\| u \right\|^2_{{\Greekmath 010D}_*} = u^{\top} u $, where $u^{\top} w$ represents some inner product of elements $u$ and $w$ of the cotangent space ${\cal T}$ of $\Pi$ at ${\Greekmath 0119}({\Greekmath 010D}_*)$. For conciseness, from now on we will remove the subscripts ${\Greekmath 010C}_0$, ${\Greekmath 010D}_{*}$, and ${\Greekmath 0119}({\Greekmath 010D}_*)$ throughout, unless there is a risk of confusion. In particular, unless otherwise noted, all expectations will be evaluated under the reference model. Here, ${\Greekmath 0119}$ can be finite-dimensional as in parametric models (which we analyze in Subsection (ref)), or infinite-dimensional as in semi-parametric mixture models where ${\Greekmath 0119}$ is a density (studied in Subsection (ref)).
Let us start by introducing some notation. Let $s_{{\Greekmath 010C}{\Greekmath 010D}}(y)={\Greekmath 0272}_{{\Greekmath 010C}{\Greekmath 010D}} \log f(y) $ and $s_{{\Greekmath 0119}}(y)={\Greekmath 0272}_{{\Greekmath 0119}} \log f(y) $ denote the components of the score. We define the Hessian operators $H_{{\Greekmath 0119}} : {\cal T} \rightarrow {\cal T}$, $H_{{\Greekmath 0119},{\Greekmath 010C}{\Greekmath 010D}}:\mathbb{R}^{\limfunc{dim}{\Greekmath 010C}+\limfunc{dim}{\Greekmath 010D}}\rightarrow{\cal{T}}$, and $H_{{\Greekmath 010C}{\Greekmath 010D},{\Greekmath 0119}}:{\cal{T}}\rightarrow\mathbb{R}^{\limfunc{dim}{\Greekmath 010C}+\limfunc{dim}{\Greekmath 010D}}$ by\footnote{ Formally, $u^{\top}$ is an element of the tangent space of $\Pi$ at ${\Greekmath 0119}({\Greekmath 010D}_*)$; that is, $u \mapsto u^{\top}$ represents a linear mapping from the cotangent space ${\cal T}$ to the tangent space $\overline {\cal T}$. } $$ H_{{\Greekmath 0119}} = \mathbb{E} s_{\Greekmath 0119} (Y) s_{\Greekmath 0119} (Y)^{\top},\quad H_{{\Greekmath 0119},{\Greekmath 010C}{\Greekmath 010D}} = \mathbb{E} s_{{\Greekmath 0119}} (Y) s_{{\Greekmath 010C}{\Greekmath 010D}} (Y)' ,\quad H_{{\Greekmath 010C}{\Greekmath 010D},{\Greekmath 0119}} = \mathbb{E} s_{{\Greekmath 010C}{\Greekmath 010D}} (Y)s_{\Greekmath 0119}(Y) ^{\top}.$$ In addition, we define the following projected versions of the gradient $\widetilde {\Greekmath 0272}_{{\Greekmath 0119}} = {\Greekmath 0272}_{{\Greekmath 0119}} - H_{{\Greekmath 0119},{\Greekmath 010C}{\Greekmath 010D}} H_{{\Greekmath 010C}{\Greekmath 010D}}^{-1} {\Greekmath 0272}_{{\Greekmath 010C}{\Greekmath 010D}} $, score $\widetilde s_{{\Greekmath 0119}}(y)=s_{{\Greekmath 0119}}(y)-H_{{\Greekmath 0119},{\Greekmath 010C}{\Greekmath 010D}} H_{{\Greekmath 010C}{\Greekmath 010D}}^{-1} s_{{\Greekmath 010C}{\Greekmath 010D}}(y) $, and Hessian $ \widetilde H_{{{\Greekmath 0119}}} = H_{{\Greekmath 0119}} - H_{{\Greekmath 0119},{\Greekmath 010C}{\Greekmath 010D}} H_{{\Greekmath 010C}{\Greekmath 010D}}^{-1}H_{{\Greekmath 010C}{\Greekmath 010D},{\Greekmath 0119}}$.
The next lemma characterizes the minimum-MSE $h$ function in the locally quadratic case.
So far in our presentation, we have abstracted from covariates. We now consider the case where in addition to the outcomes $Y_i$ we observe a vector of covariates $X_i$. We assume that $(Y_i,X_i)$ are randomly drawn from a conditional distribution of $Y_i$ given $X_i$ given by the model $f_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}(y\,|\,x)$, and an unrestricted marginal distribution $f_X$ of $X_i$. Our parameter of interest is ${\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0,f_X} = \mathbb{E}_{f_X} {\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}(X)$, where $\mathbb{E}_{f_X}$ denotes an expectation over $f_X$. We consider estimators of the form $$\widehat{{\Greekmath 010E}}_h= \frac 1 n \sum_{i=1}^n \, {\Greekmath 010E}_{\widehat {\Greekmath 010C},{\Greekmath 0119}(\widehat {\Greekmath 010D})}(X_i)+\frac 1 n \sum_{i=1}^n \,h (Y_i,X_i,\widehat {\Greekmath 010C},\widehat {\Greekmath 010D},\widehat f_X ) ,$$ where $\widehat {\Greekmath 010C}$ and $\widehat {\Greekmath 010D}$ are preliminary estimates whose probability limits are ${\Greekmath 010C}_0$ and ${\Greekmath 010D}_*$, and $ \widehat f_X$ is the empirical distribution of $X_i$ in the sample. While $f_X$ is unknown and infinite-dimensional, it only enters into our object of interest (and the expression of $h_{{\Greekmath 010F}}^{\rm MMSE}$ below) as an expectation, and the corresponding sample average is still estimated at the $\sqrt{n}$-rate. We have the following characterization of the minimum-MSE influence function.
A first difference between (ref) and (ref) is that various expectations over $f_X$ occur here, which we will replace by sample averages when calculating the estimator $\widehat{{\Greekmath 010E}}^{\rm MMSE}_{\Greekmath 010F}$ in practice. A second difference comes from the term $ {\Greekmath 010E}(x)-\mathbb{E}_{f_X}\,{\Greekmath 010E}(X)$. However, this term does not contribute to $\widehat{{\Greekmath 010E}}^{\rm MMSE}_{\Greekmath 010F}$, since its sample average is zero once we replace $f_X$ by the empirical distribution $\widehat f_X$.\footnote{The term $ {\Greekmath 010E}(x)-\mathbb{E}_{f_X}\,{\Greekmath 010E}(X)$ ensures that $h_{{\Greekmath 010F}}^{\rm MMSE}$ is locally robust with respect to $f_X$ in the sense of ((ref)).}
A simple locally quadratic example is a parametric model where ${\Greekmath 0119}$ is finite-dimensional, and the distance measure over ${\Greekmath 0119}$ is based on a weighted Euclidean metric $\|\cdot\|_{\Omega}$ for a positive definite weight matrix $\Omega$. Here we treat $\Omega$ and the neighborhood size ${\Greekmath 010F}$ as known. We will discuss the choice of $\Omega$, and the interpretation of ${\Greekmath 010F}$, in Section (ref).
The small-${\Greekmath 010F}$ approximation to the bias of $\widehat{{\Greekmath 010E}}$ is given by ((ref)), with $\|\cdot\|_{{\Greekmath 010D}_*}=\|\cdot\|_{\Omega^{-1}}$, where $\Omega^{-1}$ is the inverse of $\Omega$. In this case, for vectors $u,w \in \mathbb{R}^{\dim {\Greekmath 0119}}$ we have $ u^\top \, w = u' \, \Omega^{-1} w $. Let $$\mathbb{H}_{\Greekmath 0119}= \mathbb{E}\left[ s_{{\Greekmath 0119}} (Y) s_{{\Greekmath 0119}} (Y)' \right] ,\quad \widetilde{\mathbb{H}}_{\Greekmath 0119}= \mathbb{E}\left[ s_{{\Greekmath 0119}} (Y) s_{{\Greekmath 0119}} (Y)' \right]-\mathbb{E}\left[ s_{{\Greekmath 0119}} (Y) s_{{\Greekmath 010C}{\Greekmath 010D}} (Y)' \right]H_{{\Greekmath 010C}{\Greekmath 010D}}^{-1}\mathbb{E}\left[ s_{{\Greekmath 010C}{\Greekmath 010D}} (Y)s_{{\Greekmath 0119}} (Y)' \right],$$ be the usual parametric Hessian matrices. We have $H_{{\Greekmath 0119}} = {\mathbb{H}}_{\Greekmath 0119}\Omega^{-1}$, $ H_{{\Greekmath 0119},{\Greekmath 010C}{\Greekmath 010D}} = \mathbb{E}\left[ s_{{\Greekmath 0119}} (Y) s_{{\Greekmath 010C}{\Greekmath 010D}} (Y)' \right] $, $H_{{\Greekmath 010C}{\Greekmath 010D},{\Greekmath 0119}}=H_{{\Greekmath 0119},{\Greekmath 010C}{\Greekmath 010D}}'\Omega^{-1}$, $ \widetilde s_{{\Greekmath 0119}} = s_{{\Greekmath 0119}} - H_{{\Greekmath 0119},{\Greekmath 010C}{\Greekmath 010D}} H_{{\Greekmath 010C}{\Greekmath 010D}}^{-1} s_{{\Greekmath 010C}{\Greekmath 010D}} $, and $\widetilde H_{{{\Greekmath 0119}}} = \widetilde{\mathbb{H}}_{\Greekmath 0119}\Omega^{-1} $. From ((ref)) we then obtain the following.
In addition to the “one-step efficient” adjustment $h_{0}^{\rm MMSE} = s_{{\Greekmath 010C}{\Greekmath 010D}} ' H_{{\Greekmath 010C}{\Greekmath 010D}} ^{-1} {\Greekmath 0272}_{{\Greekmath 010C}{\Greekmath 010D}} {\Greekmath 010E}$, the minimum-MSE function $h_{{\Greekmath 010F}}^{\rm MMSE}$ in Corollary (ref) thus provides a further adjustment that is motivated by robustness concerns. It is easy to generalize this formula to account for conditioning covariates whose distribution is unspecified, as in Lemma (ref).
It is interesting to compute the limit of the MSE-minimizing $h$ function as ${\Greekmath 010F}$ tends to infinity in the case where $H_{\Greekmath 0119}$ is invertible. This leads to the following expression, which is identical to (ref),
where $\widetilde{\mathbb{H}}_{\Greekmath 0119}^\dagger$ denotes the Moore-Penrose generalized inverse of $\widetilde{\mathbb{H}}_{\Greekmath 0119}$. Comparing ((ref)) and Corollary (ref) shows that the optimal $\widehat {\Greekmath 010E}\,^{\rm MMSE}_{\Greekmath 010F}$ is a Ridge-regularized version of the one-step full MLE, where $({\Greekmath 010F} n)^{-1} I $ regularizes the projected Hessian matrix $\widetilde{ H}_{{\Greekmath 0119}}=\widetilde{\mathbb{H}}_{\Greekmath 0119}\Omega^{-1} $. Our “robust” adjustment remains well-defined under singularity, and it accounts for small or zero eigenvalues of the Hessian in an MSE-optimal way.
\paragraph{A linear regression example.}
Studying a linear regression model helps to illustrate some of the main features of our approach. Consider the model
where $Y$ is a scalar outcome, and $X$ and $Z$ are random vectors of covariates and instruments, respectively, ${\Greekmath 010C}$ is a $\limfunc{dim}X$ parameter vector, and $C$ is a $\limfunc{dim}X\times \limfunc{dim}Z$ matrix. We assume that $U={\Greekmath 0119}' V+{\Greekmath 0118}$, where ${\Greekmath 0118}$ is normal with zero mean and variance ${\Greekmath 011B}^2$, independent of {$X,Z,V$}, and $V$ is normal with zero mean and non-singular covariance matrix $\Sigma_V$, independent of $Z$. Let $\Sigma_Z $ be the covariance matrix of $Z$, and let $\Sigma_X=C \Sigma_ZC'+\Sigma_V$. For simplicity we assume that $C$, $\Sigma_V$, $\Sigma_Z$, and ${\Greekmath 011B}^2$ are known, and we take $\Omega=I$ to be the identity matrix. The parameters are thus ${\Greekmath 010C}$ and ${\Greekmath 0119}$. In the reference model we take ${\Greekmath 0119}=0$, hence treating $X$ as exogenous whereas the larger model allows for endogeneity. The target parameter is ${\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}=c'{\Greekmath 010C}_0$, for a known $\limfunc{dim}{\Greekmath 010C}\times 1$ vector $c$.
From ((ref)) we have\footnote{In this case, there is no ${\Greekmath 010D}$ parameter, $s_{\Greekmath 010C} (y,x\,|\,z)=\frac{1}{{\Greekmath 011B}^2}x(y-x'{\Greekmath 010C}_0)$, $s_{\Greekmath 0119} (y,x\,|\,z)=\frac{1}{{\Greekmath 011B}^2}(x-C z)(y-x'{\Greekmath 010C}_0)$, ${\mathbb{E}}_{f_Z} H_{{\Greekmath 010C}}(Z) = \frac{1}{{\Greekmath 011B}^2}\Sigma_X$, $\widetilde{\Greekmath 0272}_{{\Greekmath 0119}} = {\Greekmath 0272}_{\Greekmath 0119} - \Sigma_V\Sigma_X^{-1}{\Greekmath 0272}_{\Greekmath 010C}$, and ${\mathbb{E}}_{f_Z} \widetilde{H}_{{\Greekmath 0119}}(Z) = \frac{1}{{\Greekmath 011B}^2}(\Sigma_V-\Sigma_V\Sigma_X^{-1}\Sigma_V)$.}
Hence, when ${\Greekmath 010F}=0$ the minimum-MSE estimator of $c'{\Greekmath 010C}_0$ is the “one-step efficient” adjustment in the direction of the OLS estimator, with influence function $h_{0}^{\rm MMSE}(y,x,z) = (y-x'{\Greekmath 010C}_0)x' \Sigma_X ^{-1} \, c$. As ${\Greekmath 010F} $ tends to infinity, assuming $C \Sigma_ZC'$ is invertible, it follows from ((ref)) that
which is the influence function of the IV estimator.
For given ${\Greekmath 010F}>0$ and $n$, our adjustment remains well-defined even when $C \Sigma_ZC'$ is singular.\footnote{Note that in the absence of an instrument $Z$, the minimum-MSE estimator coincides with the one-step efficient adjustment in the direction of the OLS estimator.} When $c'{\Greekmath 010C}_0$ is identified (that is, when $c$ belongs to the range of $C$), the minimum-MSE estimator remains well-behaved as ${\Greekmath 010F} n$ tends to infinity, otherwise setting a finite ${\Greekmath 010F}$ value is needed to control the increase in variance. The term $({\Greekmath 010F} n)^{-1}$ in ((ref)) acts as a form of regularization, akin to Ridge regression. In Appendix (ref), we show how to extend the parametric setting of this subsection to models defined by moment restrictions, and we revisit this example while dropping the normality assumptions. $\square$
We now consider a class of semi-parametric models, where the distribution of outcomes $Y$ conditional on unobserved latent variables $A\in {\cal A}$ is described parametrically by $Y \, \big|\, A \sim g_{{\Greekmath 010C}_0}(\cdot\, |\,A)$, with finite-dimensional unknown parameter ${\Greekmath 010C}_0 \in {\cal B}$, while the distribution of $A \sim {\Greekmath 0119}_0$ is left unrestricted in the “large” correctly specified model. Here $\Pi$ is the set of probability distributions over $ {\cal A}$. The distribution of observed outcomes as a function of the unknown parameters ${\Greekmath 010C}_0 \in {\cal B}$ and ${\Greekmath 0119}_0 \in \Pi$ is given by
The parameter of interest is a functional of ${\Greekmath 010C}_0$ and ${\Greekmath 0119}_0$, which takes the form of an expectation over $A$; that is, $${\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0} = \mathbb{E}_{{\Greekmath 0119}_0} \, \Delta_{{\Greekmath 010C}_0}(A) = \int_{\cal A} \, \Delta_{{\Greekmath 010C}_0}(a) \, {\Greekmath 0119}_0(a) \, da , $$ where $ \Delta_{{\Greekmath 010C}_0}(a)$ is a known function of ${\Greekmath 010C}_0$ and $a$.
In Section (ref), we will illustrate this setup in two binary choice models: a cross-sectional model, and a dynamic panel data model. In the first case, $A$ is an error term independent of covariates, normally distributed under the reference model. In the second case, $A$ is a latent individual effect correlated with initial conditions, specified using a parametric correlated random-effects reference model (Chamberlain, 1984). In both models, we will estimate average effects, which are expectations with respect to the distribution of $A$. Our approach will provide insurance against misspecification of the parametric functional forms.
Let us specify a parametric reference model for the distribution of the latent variables $A$, and denote the reference density by ${\Greekmath 0119}({\Greekmath 010D})$, where ${\Greekmath 010D}$ is a finite-dimensional parameter. Under the reference model, the distribution of outcomes is given by $f_{{\Greekmath 010C}_0,{\Greekmath 0119}({\Greekmath 010D}_*)}(y) = \int_{\cal A} \, g_{{\Greekmath 010C}_0}(y\,|\,a) \, {\Greekmath 0119}(a\,|\,{\Greekmath 010D}_*) \, da $. However, this model may be misspecified, and we assume that the true distribution ${\Greekmath 0119}_0$ belongs to the neighborhood $\Gamma_{{\Greekmath 010F}}({\Greekmath 010D}_*) = \{{\Greekmath 0119}_0 \in \Pi \,:\,d( {\Greekmath 0119}_0, {\Greekmath 0119}({\Greekmath 010D}_*) )\leq {\Greekmath 010F}\}$, which we define here in terms of the Kullback-Leibler (KL) divergence $d( {\Greekmath 0119}_0, {\Greekmath 0119}({\Greekmath 010D}_*) ) = 2 \, \mathbb{E}_{{\Greekmath 0119}_0} \log[{\Greekmath 0119}_0(A) / {\Greekmath 0119}(A\,|\,{\Greekmath 010D}_*)]$.
We are going to derive the expression of the minimum-MSE estimator by applying ((ref)). In this setting, elements of the cotangent space ${\cal T}$ of ${\Greekmath 0119}({\Greekmath 010D})$ at ${\Greekmath 010D}_*$ can be represented by functions $u : {\cal A} \mapsto \mathbb{R}$. For example, the gradient ${\Greekmath 0272}_{{\Greekmath 0119}} {\Greekmath 010E}_{{\Greekmath 010C},{\Greekmath 0119}} $ is a cotangent element, which can be represented by the function $\Delta_{\Greekmath 010C}(\cdot)$.\footnote{Note that, since ${\Greekmath 0119}$ integrates to one (and therefore tangent space elements integrate to zero), one can equivalently represent ${\Greekmath 0272}_{{\Greekmath 0119}} {\Greekmath 010E}_{{\Greekmath 010C},{\Greekmath 0119}} $ as $\Delta_{\Greekmath 010C}(\cdot)-c$ for any constant $c$. A possible choice is $c=\mathbb{E}_{{\Greekmath 0119}({\Greekmath 010D}_*)}\Delta_{\Greekmath 010C}(A)$. } For elements $u,w \in {\cal T}$ we define their inner product by $ u^\top \, w = {\rm Cov}_{{\Greekmath 0119}({\Greekmath 010D}_*)}\left[ u(A) , w(A) \right] $; that is, the corresponding squared norm in (ref) is $ \left\| u \right\|^2_{{\Greekmath 010D}_*} = {\rm Var}_{{\Greekmath 0119}({\Greekmath 010D}_*)} \left[ u(A) \right] $. One can show that this norm is indeed the dual to a suitable local approximation to the KL divergence as defined in (ref); see Appendix (ref).
Let us omit again parameter subscripts from the notation for conciseness. From (ref), we see that $s_{{\Greekmath 0119}}(y)={\Greekmath 0272}_{{\Greekmath 0119}} \log f(y)$ can be represented by the function $g(y\,|\,a) / f(y) $. As a result, for $u \in {\cal T}$ we have
where we have used that $\mathbb{E} g(y\,|\,A) / f(y) =1$. In addition, we have
and, for any function $h$, $\mathbb{E}[h(Y)s_{{\Greekmath 0119}}(Y)]$ can be represented by the function $\mathbb{E}[h(Y)\,|\, A=a]$.
Rewriting the first-order condition in equation (ref), we thus obtain the following result, which shows that $h_{{\Greekmath 010F}}^{\rm MMSE}$ is the solution to a linear system.
Corollary (ref) can be generalized to account for conditioning covariates $X$. We now apply Lemma (ref) to provide two generalizations, which we will use in the two examples in Section (ref). In the first one, we assume that $A$ and $X$ are independent under ${\Greekmath 0119}_0$. This is the case in our cross-sectional illustration in Subsection (ref), where $A$ is an error term independent of $X$. In this case, ${\Greekmath 0119}_0$ is the marginal distribution of $A$. We then have the following characterization.
In the second generalization, we leave the joint distribution of $(A,X)$ unrestricted under ${\Greekmath 0119}_0$. This is the case in our panel data illustration in Subsection (ref), where $A$ is an individual effect that may be correlated with $X$. In this case, ${\Greekmath 0119}_0$ is the conditional distribution of $A$ given $X$, and we measure the distance between conditional distributions using $d( {\Greekmath 0119}_0, {\Greekmath 0119}({\Greekmath 010D}_*) ) = 2 \, \mathbb{E}_{f_X}\mathbb{E}_{{\Greekmath 0119}_0} \log[{\Greekmath 0119}_0(A\,|\, X) / {\Greekmath 0119}(A\,|\,X,{\Greekmath 010D}_*)]$. We then have the following characterization.
To provide intuition about the form of the solution in semi-parametric mixture models, let us start by considering a setting where ${\Greekmath 010C}_0$ and ${\Greekmath 010D}_*$ are known to the researcher, while abstracting from covariates for simplicity. Let $\mathbb{E}_{{\cal{Y}}\,|\, {\cal{A}}}$ and $\mathbb{E}_{{\cal{A}}\,|\, {\cal{Y}}}$ denote the conditional expectation operators of $Y$ given $A$ and $A$ given $Y$, respectively. Corollary (ref) implies that (see Appendix (ref) for a derivation)
where $\circ$ denotes the composition operator, and $\mathbb{I}_{\cal{A}}$ denotes the identity operator; that is, $\mathbb{I}_{\cal{A}}{\Greekmath 0119}={\Greekmath 0119}$ for ${\Greekmath 0119}:{\cal{A}}\rightarrow\mathbb{R}$. In semi-parametric mixture settings such as panel data models, average effects are often only partially identified or not root-$n$ estimable due to ill-posedness.\footnote{See, e.g., Chernozhukov et al. (2013), Pakes and Porter (2013), Severini and Tripathi (2012), and Bonhomme and Davezies (2017).} The presence of the Tikhonov penalty $({\Greekmath 010F} n)^{-1}$ in ((ref)) bypasses these issues by making the operator $[\mathbb{E}_{{\cal{Y}}\,|\, {\cal{A}}}\circ\mathbb{E}_{{\cal{A}}\,|\, {\cal{Y}}}+({\Greekmath 010F} n)^{-1}\mathbb{I}_{\cal{A}}]$ non-singular. By focusing on a shrinking neighborhood of the reference distribution, as opposed to entertaining any possible distribution, our approach avoids issues of non-identification and ill-posedness while guaranteeing MSE-optimality within that neighborhood.\footnote{Note that regular estimation is possible when there exists a function ${\Greekmath 0120}(y)$ such that $\Delta(A)-{\Greekmath 010E}=\mathbb{E}[{\Greekmath 0120}(Y)\,|\, A]$. In this case ${\limfunc{lim}}_{{\Greekmath 010F}\rightarrow \infty} \,\widehat{{\Greekmath 010E}}_{{\Greekmath 010F}}^{\rm MMSE}$ is consistent for $\mathbb{E}_{{\Greekmath 0119}_0}\, \Delta(A)={\Greekmath 010E}+\mathbb{E}_{{\Greekmath 0119}_0}\, \mathbb{E}[{\Greekmath 0120}(Y)\,|\, A]$ for all ${\Greekmath 0119}_0$. }
Next, consider the estimation of $c'{\Greekmath 010C}_0$, for $c$ a $\limfunc{dim}{\Greekmath 010C}\times 1$ vector, and assume ${\Greekmath 010D}_*$ known for simplicity. Let $\mathbb{I}_{\cal{Y}}h=h$ for $h:{\cal{Y}}\rightarrow\mathbb{R}$, and let
It follows from Corollary (ref) that
As ${\Greekmath 010F}$ tends to infinity, $\mathbb{W}^{{\Greekmath 010F}}$ approximates the functional differencing projection operator $\mathbb{W}=\mathbb{I}_{\cal{Y}}-\mathbb{E}_{{\cal{A}}\,|\, {\cal{Y}}} \mathbb{E}_{{\cal{A}}\,|\, {\cal{Y}}}^{\dagger}$, where $\mathbb{E}_{{\cal{A}}\,|\, {\cal{Y}}}^{\dagger}$ denotes the Moore-Penrose generalized inverse of $\mathbb{E}_{{\cal{A}}\,|\, {\cal{Y}}}$ (see Bonhomme, 2012). In this limit, the minimum-MSE estimator is the one-step approximation to the semi-parametric efficient estimator of $c'{\Greekmath 010C}_0$. Yet, the efficient estimator fails to exist when the matrix denominator in ((ref)) is singular.\footnote{In discrete choice panel data models, common parameters are generally not point-identified (Chamberlain, 2010, Honor\'e and Tamer, 2006). In panel data models with continuous outcomes, identification and regularity require high-level “non-surjectivity” conditions which may be hard to verify (Bonhomme, 2012).} Here the term $({\Greekmath 010F} n)^{-1}$ acts as a regularization of the functional differencing projection, which makes $h^{\rm MMSE}_{{\Greekmath 010F}}$ well-defined irrespective of the nature of identification.
Lastly, consider a model with covariates $X$ that are independent of the latent variables $A$, as in our illustration in Subsection (ref). Assuming that ${\Greekmath 010C}_0$ and ${\Greekmath 010D}_*$ are known, and that $\Delta(A)$ does not depend on $X$, and letting $\mathbb{E}_{{\cal{Y}},{\cal{X}}\,|\, {\cal{A}}}$ and $\mathbb{E}_{{\cal{A}}\,|\, {\cal{Y}},{\cal{X}}}$ denote the conditional expectation operators of $(Y,X)$ given $A$ and $A$ given $(Y,X)$, respectively, Corollary (ref) implies
The solution is similar to ((ref)), with the difference that here, due to independence, both $Y$ and $X$ are informative about the latent $A$.
To implement the method in parametric settings, the researcher needs to compute the score and Hessian of the larger model. Since we focus on smooth models, methods based on numerical derivatives or simulation-based approximations can be used. Minimum-MSE estimators are generally not available in closed form in semi-parametric mixture models. {Nevertheless, in these models $h_{{\Greekmath 010F}}^{\rm MMSE}$ can be computed by minimizing the quadratic objective (abstracting from covariates for simplicity)
with respect to $h$, subject to the linear constraints $ \mathbb{E} h(Y) =0 $ and $\mathbb{E} \, h(Y) s_{{\Greekmath 010C}{\Greekmath 010D}} (Y) = {\Greekmath 0272}_{{\Greekmath 010C}{\Greekmath 010D}} {\Greekmath 010E}$. This is a regularized linear inverse problem (see, e.g., Engl et al., 2000, and Kress, 2014), which is well-posed given the presence of the Tikhonov penalty $({\Greekmath 010F} n)^{-1}$. Numerous numerical approaches have been developed to solve linear inverse problems. In the illustrations we implement a simulation-based method that relies on matrix operations. In Appendix (ref) we describe this method, and also explain how we compute confidence intervals. Note that, given initial estimates $\widehat{{\Greekmath 010C}}$ and $\widehat{{\Greekmath 010D}}$, computing minimum-MSE estimators and confidence intervals does not require nonlinear optimization.}
In this section, we discuss how to apply our approach in practice. So far, we have shown how to compute minimum-MSE estimators and confidence intervals for different values of ${\Greekmath 010F}$. We now describe a strategy to choose an interpretable range for ${\Greekmath 010F}$, and report the results of the estimation on this range, in the spirit of sensitivity analysis.
In order to apply our approach, an important first step is to form intuition about orders of magnitudes of ${\Greekmath 010F}$ in the model under study. The literature on sensitivity analysis that we referred to in the introduction provides intuition in certain models; see for example the analysis of linear IV models in Conley et al. (2012). In Sections (ref) and (ref), we will discuss the magnitudes of ${\Greekmath 010F}$ in our examples. In our evaluation of a conditional cash transfer program in Section (ref), misspecification stems from omitted stigma effects in households' preferences. In this case, we will show that ${\Greekmath 010F}$ can be mapped to the ratio of the marginal utility of the subsidy (that is, the “stigma” effect) to that of consumption. Economic intuition can then suggest whether a given ${\Greekmath 010F}$ is large or small. {In this setting, one can motivate the assumption that ${\Greekmath 010F}$ shrinks as $n$ increases as reflecting that the econometrician's uncertainty about the presence of stigma effects diminishes when the sample gets larger.}
In the binary choice models we study in Section (ref), ${\Greekmath 0119}$ is infinite-dimensional, and ${\Greekmath 010F}/2$ is the squared radius of a Kullback-Leibler ball around a normal density. To visualize the implications of assuming that the true ${\Greekmath 0119}$ belongs to an ${\Greekmath 010F}$-neighborhood of the normal, we will plot worst-case probability bounds, and show how setting ${\Greekmath 010F}$ to a particular value imposes ex-ante restrictions on the parameter of interest. Local approximations to bounds on functionals of ${\Greekmath 0119}$ are easy to compute. Alternatively, one may report estimated worst-case distributions --- i.e., a ${\Greekmath 0119}_0$ that achieves the supremum in ((ref)) --- in the spirit of Christensen and Connault (2019).
As a complement to forming context-specific intuition about magnitudes, here we outline a generic interpretation based on statistical testing. Specifically, we show that setting ${\Greekmath 010F}$ is isomorphic to setting a lower bound on the local power of a likelihood-ratio test of the reference model, against alternatives outside the neighborhood $\Gamma_{{\Greekmath 010F}}({\Greekmath 010D}_*)$ in certain directions. The ${\Greekmath 010F}$-neighborhood will thus contain all models that are hard to statistically distinguish from the reference model in those directions. This logic has antecedents in robust statistics (Huber and Ronchetti, 2009), and robust control in economics (Hansen and Sargent, 2008).
To proceed, let us focus on the parametric case of Subsection (ref) with identity weight matrix $\Omega$. Let $v$ be a unitary vector, and consider a likelihood-ratio test of the null hypothesis $H_0: \, {\Greekmath 0119}_0={\Greekmath 0119}({\Greekmath 010D}_*)$ against the local alternative $H_1:\, {\Greekmath 0119}_0={\Greekmath 0119}({\Greekmath 010D}_*)+{\Greekmath 0118} v/\sqrt{n}$, for some constant ${\Greekmath 0118}>0$. Let the size of the test be ${\Greekmath 010B}\in(0,1)$. The local power of the test is then $p=\Pr\left(Z[{\Greekmath 0116}]>\widetilde{c}_{{\Greekmath 010B}}\right)$, where $\widetilde{c}_{{\Greekmath 010B}}$ is the $(1-{\Greekmath 010B})$-quantile of the chi-squared distribution with one degree of freedom, and $Z[{\Greekmath 0116}]$ follows a non-central chi-squared distribution with one degree of freedom and non-centrality parameter ${\Greekmath 0116}=\|\widetilde{H}_{{\Greekmath 0119}}^{1/2} v\|{\Greekmath 0118}$; see, e.g., Van der Vaart (2007, page 237).\footnote{Here $\widetilde{H}_{{\Greekmath 0119}}$ is the usual parametric (projected) Hessian matrix, since $\Omega$ is the identity.} For given ${\Greekmath 010B}$ and $p$ values, let ${\Greekmath 0116}({\Greekmath 010B},p)$ be such that $\Pr\left(Z[{\Greekmath 0116}({\Greekmath 010B},p)]>\widetilde{c}_{{\Greekmath 010B}}\right)=p$.\footnote{${\Greekmath 0116}({\Greekmath 010B},p)$ is implicitly defined by $\Phi({\Greekmath 0116}({\Greekmath 010B},p)+\Phi^{-1}({\Greekmath 010B}/2))+\Phi(-{\Greekmath 0116}({\Greekmath 010B},p)+\Phi^{-1}({\Greekmath 010B}/2))=p$, where $\Phi$ is the standard normal cumulative distribution function.} It follows that $ {\Greekmath 0116}({\Greekmath 010B},p)=\|\widetilde{H}_{{\Greekmath 0119}}^{1/2} v\|{\Greekmath 0118}$. Hence, noting that ${\Greekmath 0116}({\Greekmath 010B},p)$ is increasing in $p$, and defining$${\Greekmath 010F}(v)=\frac{{\Greekmath 0116}({\Greekmath 010B},p)^2}{nv'\widetilde{H}_{{\Greekmath 0119}}v},$$ taking ${\Greekmath 010F}\geq {\Greekmath 010F}(v)$ ensures that local power in direction $v$ is at least $p$ whenever ${\Greekmath 0118}/\sqrt{n}\geq {\Greekmath 010F}^{\frac{1}{2}}$.\footnote{This definition is easy to extend to the general locally quadratic case of Subsection (ref). Let $v$ be a unitary direction in the tangent space $\overline{\cal{T}}$ of ${\Greekmath 0119}({\Greekmath 010D})$ at ${\Greekmath 010D}_*$, and let $\Omega_{{\Greekmath 010D}_*} \, : \, \overline {\cal T} \rightarrow {\cal T}$ be the linear operator defined in the supplementary appendix. In the parametric case $\Omega_{{\Greekmath 010D}_*}$ is simply the matrix $\Omega$. In the general setup the non-centrality parameter is $\langle v,\widetilde{H}_{{\Greekmath 0119}}\Omega_{{\Greekmath 010D}_*}v\rangle^{1/2} \, {\Greekmath 0118}$, and ${\Greekmath 010F}(v)={\Greekmath 0116}({\Greekmath 010B},p)^2/(n\langle v,\widetilde{H}_{{\Greekmath 0119}}\Omega_{{\Greekmath 010D}_*}v\rangle)$, for $\left\langle v,u\right\rangle\in\mathbb{R}$ the scalar product between $v\in\overline{{\cal{T}}}$ and $u\in{\cal{T}}$.} Note that, for fixed ${\Greekmath 010B}$ and $p$, the product ${\Greekmath 010F}(v) n$ is independent of $n$.
To ensure power larger than $p$ outside the neighborhood in all directions $v$, one could compute the supremum of ${\Greekmath 010F}(v)$ over all directions.\footnote{It is sufficient to consider directions that are orthogonal to the directions ${\Greekmath 0272}_{{\Greekmath 010D}_*}{\Greekmath 0119}'$ of the reference model.} Setting ${\Greekmath 010F}$ larger than all ${\Greekmath 010F}(v)$'s is motivated by a desire to calibrate the fear of misspecification of the researcher: when $p$ is large, say $80\%$, all alternatives outside the neighborhood $\Gamma_{{\Greekmath 010F}}({\Greekmath 010D}_*)$ are then easy to statistically distinguish from the reference model based on a sample of $n$ observations. However, for the supremum of ${\Greekmath 010F}(v)$ to be finite, $\widetilde{H}_{{\Greekmath 0119}}$ needs to be non-singular, which precludes models with partial or irregular identification. As an example, in the linear model of Subsection (ref), the supremum of ${\Greekmath 010F}(v)$ is infinite whenever ${\Sigma}_X-{\Sigma}_V=C \Sigma_ZC'$ is singular; that is, whenever the IV model is under-identified. In such a case, there thus exist certain directions along which the specification test has no power, no matter how large ${\Greekmath 010F}$ is. Likewise, in semi-parametric models, the eigenvalues of the infinite-dimensional operator $\widetilde{H}_{{\Greekmath 0119}}$ may not be bounded away from zero due to ill-posedness.
In partially or irregularly identified models, given some fixed values of ${\Greekmath 010B}$ and $p$, a possibility is to report several ${\Greekmath 010F}$ value: a first value ${\Greekmath 010F}_1$ that corresponds to the infimum of ${\Greekmath 010F}(v)$ over all directions $v$ --- hence to the most favorable direction; a second value ${\Greekmath 010F}_2\geq {\Greekmath 010F}_1$, such that power is at least $p$ outside the neighborhood in the most favorable direction in the subspace of directions orthogonal to the most favorable one; a third value ${\Greekmath 010F}_3\geq {\Greekmath 010F}_2$ that provides power guarantees along the most favorable direction orthogonal to the previous two ones, and so on. Letting ${\Greekmath 0115}_k(B)$ denote the $k$-th largest eigenvalue of $B$, we have
We will report the first few ${\Greekmath 010F}_k$ values in our illustrations --- taking ${\Greekmath 010B}=5\%$ and $p=80\%$ --- as a complement to context-specific interpretations of orders of magnitude.\footnote{In Appendix (ref), we describe how to compute ${\Greekmath 010F}_k$ in semi-parametric mixture models using a simulation-based approach.}
By providing intuition about ${\Greekmath 010F}$, either through an interpretation of magnitudes in the context under study (see Subsection (ref)), and/or through a generic approach based on statistical testing (see Subsection (ref)), the researcher selects a range of possible values for ${\Greekmath 010F}$. We then recommend plotting the minimum-MSE estimator, and its associated 95% confidence interval, as a function of ${\Greekmath 010F}$ on this range. In our illustrations, we will use this device to report results, and we will indicate particular values of ${\Greekmath 010F}$ on the x-axis to facilitate interpretation.
By reporting those minimum-MSE estimates and bias-adjusted confidence intervals, we learn about the fragility of the estimation results under the reference model parameterized by ${\Greekmath 010D}$, relative to the larger model parameterized by ${\Greekmath 0119}$. Exploring the sensitivity to model misspecification in that way is in line with the traditional suggestion of comparing estimation results obtained from different model specifications (see, e.g., Leamer, 1983, 1985). However, our local approach does not require the researcher to estimate the larger model, which is particularly relevant in situations where the latter is partially or irregularly identified, or computationally hard to estimate.
Implementing our approach requires choosing a norm on $\Pi$, which governs the shape of $\Gamma_{{\Greekmath 010F}}({\Greekmath 010D}_*)$. In parametric models, the researcher may have a preferred weight matrix $\Omega$, thus putting more weight on certain elements of the vector ${\Greekmath 0119}$. An automatic weighting scheme is to set $\Omega$ to be equal to the diagonal of the projected Hessian matrix $ \widetilde{\mathbb{H}}_{\Greekmath 0119}$. This choice can be motivated using a statistical testing logic as in Subsection (ref), focusing on component-wise directions in the canonical basis of $\mathbb{R}^{\limfunc{dim}{\Greekmath 0119}}$. Taking the diagonal, instead of the entire matrix $ \widetilde{\mathbb{H}}_{\Greekmath 0119}$, as a weight is in line with our aim to cover models where the parameter of interest may not be regularly estimable.\footnote{In applications, other norms may have particular appeal. For example, measuring deviations according to the supremum norm will lead to an $\ell^1$ dual norm in ((ref)), in the spirit of Armstrong and Koles\'ar (2021). While our estimators and confidence intervals remain well-defined in this case, that setting is not locally quadratic.}
In semi-parametric mixture models, where $\Pi$ is a set of densities, we rely on the Kullback-Leibler divergence for computational convenience. KL is locally quadratic, and this choice allows us to obtain the explicit characterizations of Lemma (ref), and to compute minimum-MSE estimators by solving linear systems. Note, however, that the KL divergence does not impose shape or smoothness restrictions on the densities inside the neighborhood.
The goal of this section is to predict program impacts in the context of the PROGRESA conditional cash transfer program, building on the structural evaluation of the program in Todd and Wolpin (2006, TW hereafter) and Attanasio et al. (2012, AMS). We estimate a simple model in the spirit of TW, and adjust its predictions against a specific form of misspecification under which the program may have a “stigma” effect on preferences.
Following TW and AMS, we focus on PROGRESA's education component, which consists of cash transfers to families conditional on children attending school. Those represent substantial amounts as a share of total household income. The implementation of the policy was preceded by a village-level randomized evaluation in 1997-1998. As TW and AMS point out, the randomized control trial is silent about the effect that other, related policies could have, such as higher subsidies or unconditional income transfers, which motivates the use of structural methods.
To analyze this question, we consider a simplified version of TW's model described in Wolpin (2013), which is a static, one-child model with no fertility decision. To describe this model, let $U(C,S,{\Greekmath 011C},v)$ denote the utility of a unitary household, where $C$ is consumption, $S\in\{0,1\}$ denotes the schooling attendance of the child, ${\Greekmath 011C}$ is the level of the PROGRESA subsidy, and $v$ are taste shocks. Utility may also depend on characteristics $X$, which we abstract from for conciseness in the presentation.\footnote{Empirically, we include as covariates the age of the child and her parents, distance to the nearest school, eligibility and year indicators, and the highest grade obtained. We perform estimation separately by gender.} Note the direct presence of the subsidy ${\Greekmath 011C}$ in the utility function, which may reflect a stigma effect. This direct effect plays a key role in the analysis. The budget constraint is: $C=Y+W(1-S)+{\Greekmath 011C} S$, where $Y$ is household income and $W$ is the child's wage. This is equivalent to: $C=Y+{\Greekmath 011C}+(W-{\Greekmath 011C})(1-S)$. Hence, {in the absence of a direct effect on utility}, the program's impact is equivalent to an increase in income and a decrease in the child's wage.
Following Wolpin (2013) we parameterize the utility function as $$U(C,S,{\Greekmath 011C},v)=aC+bS+dCS+{\Greekmath 0115} {\Greekmath 011C} S+Sv,$$ where ${\Greekmath 0115}$ denotes the direct (stigma) effect of the program. The schooling decision is then $$S=\boldsymbol{1}\{U(Y+{\Greekmath 011C},1,{\Greekmath 011C},v)>U(Y+W,0,0,v)\}=\boldsymbol{1}\{v>a(Y+W)-(a+d)(Y+{\Greekmath 011C})-{\Greekmath 0115} {\Greekmath 011C}-b\}.$$ Assuming that $v$ is standard normal, independent of wages, income, and program status (that is, of the subsidy ${\Greekmath 011C}$), we obtain $$\Pr(S=1\,|\, y,w,{\Greekmath 011C})=1-\Phi\left[a(y+w)-(a+d)(y+{\Greekmath 011C})-{\Greekmath 0115} {\Greekmath 011C}-b\right],$$ where $\Phi$ is the standard normal cdf.
We use the specification with ${\Greekmath 0115}= 0$ as our reference model, and estimate it on control villages. When ${\Greekmath 0115}=0$, the average effect of the subsidy on school attendance is
As Wolpin (2013) notes, data under the subsidy regime (${\Greekmath 011C}={\Greekmath 011C}^{\rm treat}$) is {not} needed to construct an empirical counterpart to this quantity, since treatment status is independent of $Y,W$.\footnote{AMS make a related point (albeit in a different model), and use both control and treated villages to estimate their structural model. AMS also document the presence of general equilibrium effects of the program on wages. We abstract from such effects in our analysis.}
We contrast two strategies to predict the effect of the program and other counterfactual policies, while accounting for misspecification of the reference model due to the presence of stigma effects. The first strategy --- which we refer to as ex-ante policy prediction --- is only based on data from control villages, whereas the second strategy --- ex-post prediction --- combines both control and treated villages. In both cases, we allow for ${\Greekmath 0115}\neq 0$ in the larger model. While in the present simple static context one could easily estimate a version of the larger model, in dynamic structural models such as the one estimated by TW, estimating a different model in order to assess the impact of any given form of misspecification may be computationally prohibitive. This highlights an advantage of our approach, which does not require the researcher to estimate the parameters under a new model.
To cast this setting into our framework, let ${\Greekmath 010C}=(a,b,d)$, ${\Greekmath 0119}={\Greekmath 0115}$, and $${\Greekmath 010E}_{{\Greekmath 010C},{\Greekmath 0119}}=\mathbb{E}\left(\Phi\left[a(Y+W)-(a+d)(Y+{\Greekmath 011C}^{\rm treat})-{\Greekmath 0115} {\Greekmath 011C}^{\rm treat}-b\right]-\Phi\left[a(Y+W)-(a+d)Y-b\right]\right).$$ We focus on the effect on eligible (i.e., poorer) households. We add covariates to gender-specific school attendance equations, which include the age of the child and her parents, year indicators, distance to school, and an eligibility indicator. We report estimates of ${\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}$ as well as confidence intervals.
We use the sample from TW. We drop observations with missing household income, and focus on 1219 boys and 1089 girls aged 12 to 15.\footnote{Children's wages are only observed for those who work. We {impute potential wages} to all children based on a linear regression that in particular exploits province-level variation and variation in distance to the nearest city, similar to AMS.} Descriptive statistics on the sample show that average weekly household income is 242 pesos, the average weekly wage is 132 pesos, and the PROGRESA subsidy ranges between 31 and 59 pesos per week depending on grade and gender. Average school attendance drops from 90% at age 12 to between 40% and 50% at age 15.
We start by providing intuition regarding the values of ${\Greekmath 010F}$ in the present context. In the structural model, ${\Greekmath 0115}$ is the marginal utility of the subsidy for households sending their child to school. In turn, the marginal utility of consumption is given by $a+d$. Hence, bounding ${\Greekmath 0115}^2$ by ${\Greekmath 010F}$ is equivalent to bounding the ratio of marginal utility of the subsidy to marginal utility of consumption by $\sqrt{{\Greekmath 010F}}/(a+d)$. For example, households valuing the subsidy as much as consumption in absolute value --- arguably an upper bound on the stigma effect --- corresponds to $\overline{{\Greekmath 010F}}=(a+d)^2$.
In Figure (ref), we show the minimum-MSE estimator of the impact of the PROGRESA subsidy on eligible households, together with 95% confidence intervals, for a range of values around $\overline{{\Greekmath 010F}}=(a+d)^2$ (which we show in the dashed vertical line). In the horizontal dotted line, we show estimates based on the reference model. In the top panel, we show ex-ante prediction results based on control villages only. We see that the minimum-MSE estimator and the one based on the reference model are equal in this case. This is intuitive, since ${\Greekmath 0119}={\Greekmath 0115}$ is scalar, and control villages provide no information about it.\footnote{This is analogous to the case of a linear regression with endogeneity and no instrument, which we mentioned in footnote (ref).} However, the confidence intervals --- which account for model misspecification --- are large. When ${\Greekmath 010F}=\overline{{\Greekmath 010F}}$, the 95% intervals include zero for both genders. This quantifies the uncertainty associated with ex-ante prediction when the researcher does not rule out the presence of stigma.
In the bottom panel of Figure (ref), we show the results of ex-post prediction based on both control and treated villages. In the sample of boys, the minimum-MSE estimator is lower than the one based on the reference model, suggesting that the reference model is misspecified. In contrast, the two estimators are close to each other in the sample of girls.\footnote{Note that, for boys, minimum-MSE estimates at all ${\Greekmath 010F}$ values --- including ${\Greekmath 010F}=0$ --- are lower than the estimate from the reference model {(note that here we estimate the reference model using control villages only, and use both controls and treated to compute the minimum-MSE estimator)}. This suggests that the functional form of the schooling decision is {not} invariant to treatment status, highlighting that predictions based off control villages are less satisfactory for boys (as also found by TW).} In addition, the 95% confidence intervals are substantially tighter than when using control villages only. When ${\Greekmath 010F}=\overline{{\Greekmath 010F}}$ (shown in the vertical dashed line), the program estimates on school attendance are positive and marginally significant at the 5% level for girls, and positive and marginally significant at 10% for boys. In the vertical solid line, we highlight the value ${\Greekmath 010F}_1$ given by ((ref)) for $k=1$. Since ${\Greekmath 0119}={\Greekmath 0115}$ is scalar, setting ${\Greekmath 010F}\geq {\Greekmath 010F}_1$ ensures that, for all models outside the neighborhood, a 5%-likelihood ratio specification test has local power larger than 80%. Taking ${\Greekmath 010F}= {\Greekmath 010F}_1$ implies that the ratio of marginal utility of the subsidy to marginal utility of consumption is bounded by 1.8 (girls) and 1.3 (boys). While ${\Greekmath 010F}_1$ is larger than $\overline{{\Greekmath 010F}}$, the implied minimum-MSE estimators and confidence intervals are similar.\footnote{Note that ${\Greekmath 010F}_1$ is infinite in the ex-ante case (top panel of Figure (ref)). This is due to control villages not providing any information about ${\Greekmath 0119}$ in this case. }
In Table (ref), we report estimates of the program impacts, as well as predictions of counterfactual policies. The left two columns correspond to ex-ante prediction based on control villages only, while the right two columns correspond to ex-post prediction based on both controls and treated. We show the results for ${\Greekmath 010F}=\overline{{\Greekmath 010F}}$, corresponding to equal marginal utilities of subsidy and consumption in absolute value. In the top panel, we focus on the impact of the PROGRESA subsidy on eligible households. We see that PROGRESA has a positive impact on attendance of both boys and girls. The impacts predicted by the reference model are large, approximately 8 percentage points, and are quite close to the results reported in Todd and Wolpin (2006, 2008). However, the confidence intervals which account for model misspecification (third row, left two columns) are very large for both genders.
When adding treated villages to the sample (right two columns in Table (ref)), confidence intervals accounting for misspecification are tighter. Moreover, the minimum-MSE point estimates and those based on the reference model differ in this case. For boys, the minimum-MSE estimate is substantially lower than the one based on the reference model (3.6% versus 7.8%), while for girls the effects are similar. Interestingly, for boys the minimum-MSE estimates are closer to the experimental differences in means between treated and control villages.
Lastly, in the middle and bottom panels of Table (ref), we show estimates of the effects of two counterfactual policies: doubling the PROGRESA subsidy, and removing the conditioning of the income transfer on school attendance. Unlike for the main PROGRESA impacts, there is no experimental counterpart to such counterfactuals. While ex-ante predictions are associated with wide confidence intervals, ex-post minimum-MSE estimates based on both control and treated villages predict a substantial effect of doubling the subsidy on girls' attendance, and a more moderate effect on boys. By contrast, we find insignificant effects of an unconditional income transfer.
In this section, we apply our approach to cross-sectional and panel data binary choice models, where we allow for misspecification of the distribution of unobservables. {In both applications there is a substantial amount of misspecification, and we use simulations to assess the behavior of the minimum-MSE estimator --- which is theoretically justified under local misspecification --- in these settings of practical relevance.}
Consider the binary choice model
where $A$ follows a distribution ${\Greekmath 0119}_0$, independent of $X$. We are interested in estimating the prediction function ${{\Greekmath 010E}}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}=\mathbb{E}_{{\Greekmath 0119}_0}[\mathbbm{1}\{x_0'{\Greekmath 010C}_0+A\geq 0\}]$, at some $x_0$ not necessarily in the support of $X$. We focus on the reference specification $A\sim {\cal{N}}(0,1)$, independent of $X$. We allow for the possibility that this parametric model is misspecified, while maintaining independence between $A$ and $X$ under ${\Greekmath 0119}_0$. We observe an i.i.d. sample $(Y_i,X_i)$ for $i=1,...,n$.
In neighborhoods that consist of distributions of $A$ independent of $X$, the minimum-MSE influence function is given by Corollary (ref), with ${\Greekmath 0272}_{{\Greekmath 010C}}{\Greekmath 010E}=x_0{\Greekmath 011E}(x_0'{\Greekmath 010C}_0)$ for ${\Greekmath 011E}$ the standard normal density, $\Delta(a)=\mathbbm{1}\{x_0'{\Greekmath 010C}_0+a\geq 0\}$, and without ${\Greekmath 010D}$ parameter. Given a preliminary estimator $\widehat{{\Greekmath 010C}}$ (e.g., obtained by probit), an empirical counterpart to $\overline{h}^{\rm MMSE}_{{\Greekmath 010F}}(a)$ is $$\frac{1}{n}\sum_{i=1}^n \mathbbm{1}\{X_i'\widehat{\Greekmath 010C}+a\geq 0\}h_{{\Greekmath 010F}}^{\rm MMSE}(1,X_i)+(1-\mathbbm{1}\{X_i'\widehat{\Greekmath 010C}+a\geq 0\})h_{{\Greekmath 010F}}^{\rm MMSE}(0,X_i).$$ We compute $h_{{\Greekmath 010F}}^{\rm MMSE}(1,X_i)$ and $h_{{\Greekmath 010F}}^{\rm MMSE}(0,X_i)$, for $i=1,...,n$, based on Corollary (ref) by solving a linear system. In this model, all conditional expectations are available in closed form.
In model ((ref)), under independence between $A$ and $X$, ${\Greekmath 010C}_0$ and ${\Greekmath 0119}_0$ are point-identified up to scale under sufficiently rich support of $X$ (Manski, 1988). Under such conditions, ${{\Greekmath 010E}}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}$ is identified. More generally, it is partially identified. We now set up a simulation where the support of $X$ is discrete, and we vary the number of support points and the target $x_0$. In this way, we learn how our estimators and confidence intervals perform in settings where the support of $X$, and hence the size of the identified set, vary.
We will show estimates in data generating processes (DGPs) with a scalar covariate and an intercept, and ${\Greekmath 010C}_0=(2,-1)'$, where the second element corresponds to the intercept. We draw $1000$ simulated samples of size $n=500$, where $A$ has mean zero and variance one, and is distributed as a mixture of two normals whose means are approximately two standard deviations apart. Covariates are discrete uniform on $[0,1]$, with either $n_X=4$ or $n_X=20$ points of support. We show the densities of $X$ and $A$ in panels (a) and (b) of Figure (ref). We focus on the predicted values at $x_0=(0.5,1)'$ and $x_0=(-0.5,1)'$, respectively. We refer to the first case as interpolation, and to the second one as extrapolation.
We report minimum-MSE estimates and confidence intervals on a range of ${\Greekmath 010F}$ values. To provide intuition about orders of magnitude, one can compute (local approximations to twice) the KL divergence between the standard normal and other common distributions. For example, a scaled student-$t$ distribution with unitary variance and $5$, $3$, or $2.1$ degrees of freedom, respectively, corresponds to a distance of $0.14$, $0.24$, and $1.5$; the true bimodal ${\Greekmath 0119}_0$ in the DGP corresponds to a distance of $1.6$; and the standardized logistic corresponds to a distance of $0.07$. Restricting ${\Greekmath 0119}_0$ to belong to an ${\Greekmath 010F}$-neighborhood of the normal also has implications for its functionals. As an example, in Figure (ref), we show pointwise bounds on $\mathbb{E}_{{\Greekmath 0119}_0}[\boldsymbol{1}\{a+A\geq 0\}]$, as a function of $\Phi(a)$, computed using small-${\Greekmath 010F}$ approximations. We see that taking ${\Greekmath 010F}=0.1$ tightly restricts possible values that the parameter can take. By contrast, when ${\Greekmath 010F}=1$, the a priori bounds on the parameter are very wide for $a$ close to $0.5$ --- which is relevant for the interpolation case --- but the neighborhood does restrict the parameter value when $a$ in close to zero or one --- which is relevant for the extrapolation case. Given this, we will interpret ${\Greekmath 010F}$ values of the order of 0.1 or lower as reflecting “mild” misspecification, and values of the order of 1 or larger as corresponding to “large” misspecification.
In this example, it is also informative to interpret ${\Greekmath 010F}$ by relating it to the power of a specification test, as we described in Subsection (ref). When $X$ has 4 support points, $\widetilde{H}_{\Greekmath 0119}$ has only $n_X-1=3$ non-zero eigenvalues (since the $X'{\Greekmath 010C}$ partition the real line into $n_X+1$ intervals, and the two elements in ${\Greekmath 010C}$ are estimated), two of them corresponding to non-constant eigenfunctions. In this case, ${\Greekmath 010F}_1$ and ${\Greekmath 010F}_2$ given by ((ref)) are approximately $0.2$ and $0.4$ on average, where we set size to ${\Greekmath 010B}=5\%$ and power to $p=80\%$. When $X$ has 20 support points, $\widetilde{H}_{{\Greekmath 0119}}$ has $n_X-1=19$ non-zero eigenvalues corresponding to non-constant eigenfunctions. The first three values ${\Greekmath 010F}_1$, ${\Greekmath 010F}_2$, and ${\Greekmath 010F}_3$ are approximately $0.3$, $0.6$, and $1.1$ on average. In contrast with the parametric case of Section (ref), here setting ${\Greekmath 010F}\geq{\Greekmath 010F}_k$ only provides power guarantees along particular directions. In Appendix (ref), we plot those directions, and we provide additional intuition about the interpretation of ${\Greekmath 010F}$ based on statistical testing in this example.
We show the results of the simulation in Figure (ref). Consider first the top panel, where we wish to interpolate the prediction function at $x_0=0.5$. When $X$ has 4 support points, we see that the probit estimator based on the reference model, indicated by the solid horizontal line, is substantially biased. By contrast, the bias of the minimum-MSE estimator is smaller, and it decreases as ${\Greekmath 010F}$ increases. We see that the minimum-MSE estimator is close to unbiased for both ${\Greekmath 010F}_1$ and ${\Greekmath 010F}_2$. Moreover, the dispersion of the estimator is stable as ${\Greekmath 010F}$ increases. In addition, we compute the identified set for ${\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}$ in the DGP using linear programming and a grid of ${\Greekmath 010C}_0$ values. We find $[0.334,0.345]$, which shows that the identified set is not wide in this DGP.
The case where $X$ has 20 support points is overall quite similar, but with several differences. We see that the minimum-MSE estimator is virtually unbiased when ${\Greekmath 010F}\geq {\Greekmath 010F}_1$. In this case, the identified set for ${\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}$ is essentially a singleton: $[0.334,0.335]$. Moreover, we see that the variance of the minimum-MSE estimator increases with ${\Greekmath 010F}$. Such a variance increase, and the associated regularization role of ${\Greekmath 010F}$, also characterize models with continuously distributed covariates and other ill-posed inverse problems.
Lastly, consider the lower panel in Figure (ref). This extrapolation case is very different from the interpolation one. Indeed, the data provide little information about the value of the prediction function at $x_0=-0.5$. To illustrate, the identified set for ${\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}$ is $[0,0.3219]$ (respectively, $[0,0.2956]$) when $X$ has 4 (resp., 20) points of support. We see that the minimum-MSE estimator has approximately the same bias as the probit estimator in this case.
We show additional information about the simulation results in Tables (ref) and (ref) in the appendix. In particular, we report the lengths of our 95% confidence intervals (CI) for ${\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}$, which are asymptotically valid under ${\Greekmath 010F}$-misspecification, and the associated coverage probabilities. In all DGPs, we find that, when taking ${\Greekmath 010F}\geq {\Greekmath 010F}_1$, the confidence intervals contain the true value with a probability that exceeds 95%.\footnote{While this finding is interesting, note that our CI construction has coverage guarantees only when ${\Greekmath 0119}_0$ belongs to an ${\Greekmath 010F}$-neighborhood of ${\Greekmath 0119}({\Greekmath 010D}_*)$, which is not the case here since the true distribution of $A$ lies outside the neighborhoods for the range of ${\Greekmath 010F}$ that we consider.}
In this subsection, we present simulations in the following dynamic panel data probit model with individual effects
where $U_{1},...,U_{T}$ are i.i.d. standard normal, independent of $A$ and $Y_{0}$. Here $Y_{0}$ is observed, so there are effectively $T+1$ time periods. We focus on the average {state dependence} effect ${\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}=\mathbb{E}_{{\Greekmath 0119}_0}\left[\Phi({\Greekmath 010C}_0+A)-\Phi(A)\right]$, and we will also report estimates of the autoregressive parameter ${\Greekmath 010C}_0$. We assume that the probit conditional likelihood given individual effects and lagged outcomes is correctly specified. However, we do not assume knowledge of ${\Greekmath 0119}_0$ or its functional form. We specify a normal reference density for $A$ given $Y_{0}$, with mean ${\Greekmath 0116}_1+{\Greekmath 0116}_2 Y_{0}$ and variance ${\Greekmath 011B}^2$; hence here ${\Greekmath 010D}=({\Greekmath 0116}_1,{\Greekmath 0116}_2,{\Greekmath 011B}^2)'$. Binary choice panel data models are often partially identified for fixed $T$ (Chamberlain, 2010, Honor\'e and Tamer, 2006), and no semi-parametrically consistent estimators of ${\Greekmath 010C}_0$ and ${\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}$ in the dynamic probit model are available in the literature. Here we report simulation results suggesting that minimum-MSE estimators can perform well under sizable misspecification of the reference density.
In the simulation, we set a bimodal distribution that has modes $\{-1,2\}$ when $Y_0=0$ and $\{-3,0\}$ when $Y_0=1$, with some asymmetry between the two modes; see panel (c) of Figure (ref). {We draw $Y_0$ from a Bernoulli distribution with probability $0.5$.} We take $n=500$, and show the results for $T=5$, $10$, and $20$, based on $1000$ simulations. In neighborhoods that consist of unrestricted joint distributions ${\Greekmath 0119}_0$ of $(A,X)$, the minimum-MSE $h$ function is given by Corollary (ref), for $X=Y_{0}$, and either $\Delta(a)=\Phi({\Greekmath 010C}_0+a)-\Phi(a)$ or $\Delta(a)={\Greekmath 010C}_0$, depending on the quantity of interest. We use $S=1000$ simulated draws to compute the minimum-MSE estimators, since no closed-form solution is available in this case {(see Appendix (ref))}.
In Figure (ref), we see that the parametric random-effects dynamic probit estimates of ${\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}$ and ${\Greekmath 010C}_0$ are substantially biased for $T=5$ and $T=10$, whereas the bias is smaller when $T=20$. By contrast, the minimum-MSE estimator performs better in terms of bias for both quantities of interest, in particular when taking ${\Greekmath 010F}$ to be one of the first few ${\Greekmath 010F}_k$'s given by ((ref)). In the top panel of Table (ref), we show the bias and root MSE of various estimators of ${\Greekmath 010E}_{{\Greekmath 010C}_0,{\Greekmath 0119}_0}$: the random-effects estimator based on the normal reference model, an empirical Bayes estimator, the linear probability estimator, and the minimum-MSE estimators based on ${\Greekmath 010F}_1,{\Greekmath 010F}_2,{\Greekmath 010F}_3$.\footnote{The random-effects and empirical Bayes estimators are given by $\frac{1}{n}\sum_{i=1}^n\mathbb{E}_{{\Greekmath 0119}(\widehat{{\Greekmath 010D}})}[\Phi(\widehat{{\Greekmath 010C}}+A)-\Phi(A)]$ and $\frac{1}{n}\sum_{i=1}^n\mathbb{E}_{{\Greekmath 0119}(\widehat{{\Greekmath 010D}})}[\Phi(\widehat{{\Greekmath 010C}}+A)-\Phi(A)\,|\, Y=Y_i]$, respectively. In fixed-lengths panels both estimators are consistent under the parametric reference specification, and the random-effects estimator is efficient. However, the two estimators are generally biased under misspecification. { Bonhomme and Weidner (2021) show that the empirical Bayes estimator has minimum local asymptotic worst-case specification error, albeit in neighborhoods of the reference model where the probit conditional likelihood given individual effects and lagged outcomes may be incorrectly specified.} } We see that the minimum-MSE estimator dominates all other estimators, for these ${\Greekmath 010F}_k$ values, when $T=5$ and $T=10$. In the bottom panel of Table (ref), we show the results for the random-effects MLE and minimum-MSE estimators of ${\Greekmath 010C}_0$. The results are similar to the case of average state dependence. In this DGP, minimum-MSE estimators achieve bias reduction under misspecification, even when $T$ is quite small. Bias reduction comes with some increase in variance, yet, the overall MSE is lower for minimum-MSE estimators compared to the MLE. Lastly, in Tables (ref) and (ref) in the appendix, we show additional information about the simulation results, for the autoregressive parameter and the average state dependence parameter, respectively.
We propose a framework for estimation and inference in the presence of model misspecification. This allows researchers to perform sensitivity analysis for existing estimators, and to construct improved estimators and confidence intervals that are less sensitive to model assumptions. Our approach is based on a minimax mean squared error rule, which consists of a one-step adjustment of the initial estimate. This adjustment is motivated by both robustness and efficiency, and it remains valid when the identification of the “large” model is irregular or point-identification fails. Hence, our approach provides a {complement to partial identification methods}, when the researcher sees her reference model as a plausible, albeit imperfect, approximation to reality. Given a parametric reference model, implementing our estimators and confidence intervals {does not} require estimating a {larger model}. This is an attractive feature in complex models such as dynamic structural models, for which sensitivity analysis methods are needed. {Lastly, while our theory applies quite generally, we have provided explicit expressions and described implementation in two specific classes of problems: parametric models, and semi-parametric likelihood models with a mixture structure. Generalizing the applicability of the approach to other semi-parametric models is an important task for future work.}
\baselineskip12pt