EconBase
← Back to paper

Posterior Average Effects

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

98,952 characters · 17 sections · 0 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Posterior Average Effects

\vskip 3cm

abstractEconomists are often interested in estimating averages with respect to distributions of unobservables, such as moments of individual fixed-effects, or average partial effects in discrete choice models. For such quantities, we propose and study posterior average effects (PAE), where the average is computed {conditional} on the sample, in the spirit of empirical Bayes and shrinkage methods. While the usefulness of shrinkage for prediction is well-understood, a justification of posterior conditioning to estimate population averages is currently lacking. We show that PAE have minimum worst-case specification error under various forms of misspecification of the parametric distribution of unobservables. In addition, we introduce a measure of informativeness of the posterior conditioning, which quantifies the worst-case specification error of PAE relative to parametric model-based estimators. As illustrations, we report PAE estimates of distributions of neighborhood effects in the US, and of permanent and transitory components in a model of income dynamics. JEL codes:\ C13, C23. Keywords:\ model misspecification, robustness, sensitivity analysis, empirical Bayes, posterior conditioning, latent variables.

\baselineskip21pt

\setcounter{page}{0}\thispagestyle{empty}

Introduction

In many settings, applied researchers wish to estimate population averages with respect to a distribution of unobservables. This includes moments of individual fixed-effects in panel data, and average partial effects in discrete choice models, which are expectations with respect to some distribution of shocks or heterogeneity. The standard approach in applied work is to assume a parametric form for the distribution of unobservables, and to compute the average effect under that assumption. For example, in binary choice, researchers often assume normality of the error term, and compute average partial effects under normality. This “model-based” estimation of average effects is justified under the assumption that the parametric model is {correctly specified}.

In this paper, we consider a different approach, where the average effect is computed {conditional on the observation sample}. We refer to such estimators as “posterior average effects” (PAE). Posterior averaging is appealing for prediction purposes, and it plays a central role in Bayesian and empirical Bayes approaches (e.g., Berger, 1980, Morris, 1983). Here we focus instead on the estimation of population expectations. Our goal is twofold: to propose a novel class of estimators, and to provide a frequentist framework to understand when and why posterior conditioning may be useful in estimation. Our main result will show that PAE have robustness properties when the parametric model is {misspecified}.

PAE are closely related to empirical Bayes (EB) estimators, which are increasingly popular in applied economics. Consider a fixed-effects model of teacher quality, which is our main example. When the number of observations per teacher is small, the dispersion of teacher fixed-effects is likely to overstate that of true teacher quality, since teacher effects are estimated with noise. An alternative approach is to postulate a prior distribution for teacher quality --- typically, a normal --- and report posterior estimates, holding fixed the values of the mean and variance parameters. The hope is that such EB estimates, which are shrunk toward the prior, are less affected by noise than the teacher fixed-effects (e.g., Kane and Staiger, 2008, Chetty et al., 2014, Angrist et al., 2017). However, while EB estimates are well-justified predictors of the quality of individual teachers, it is not obvious how to aggregate them across teachers when the goal is to estimate a population average such as a moment or a distribution function.

As an example, suppose we wish to estimate the distribution function of teacher quality evaluated at a point. Since this quantity is an average of indicator functions, the PAE is simply an average of posterior means --- that is, of EB estimates --- {of the indicator functions}. This estimator is available in closed form. However, the PAE differs from the empirical distribution of the EB estimates of teacher effects. In particular, while the variance of EB estimates is too small relative to that of latent teacher quality, the PAE has the correct variance. Related applications of PAE include settings involving neighborhood/place effects (Chetty and Hendren, 2017, Finkelstein et al., 2017) or hospital quality (Hull, 2018).

Importantly, although posterior averages have desirable properties for predicting individual parameters, their usefulness for estimating {population average quantities} is not evident. For example, suppose that teacher quality is normally distributed. In this case, a model-based normal estimator of the distribution of teacher quality is consistent. Moreover, it is asymptotically efficient when means and variances are estimated by maximum likelihood. Hence, in the correctly specified case, there is no reason to deviate from the standard model-based approach and compute posterior estimators. The main insight of this paper is that, under {misspecification} --- e.g., when teacher quality is not normally distributed --- conditioning on the data using PAE can be beneficial.

To study estimators under misspecification, we focus on specification error, which is the population discrepancy between the probability limit of an estimator and the true parameter value. In our main results, we show that PAE have {minimum worst-case specification error}, where the worst case is computed in a nonparametric neighborhood of the reference parametric distribution (e.g., a normal). Specifically, we show that, when neighborhoods are defined in terms of the Pearson chi-squared divergence, PAE have minimum worst-case specification error within a large class of estimators, for any neighborhood size smaller than a threshold value that we characterize. In addition, when broadening the class of neighborhoods to ${\Greekmath 011E}$-divergences, we show that, while PAE do not have minimum worst-case specification error in general in fixed-size neighborhoods, they achieve minimum worst-case specification error under {local} misspecification, i.e. when the size of the neighborhood tends to zero.

In our examples and illustrations, we find that the information contained in the posterior conditioning is setting-specific. This is intuitive, since although PAE have minimum worst-case specification error under our conditions, the specification error is not zero in general and it varies between applications. PAE tend to behave better when the realizations of outcome variables (such as test scores) are more informative about the values of the unobservables (such as the quality of a teacher). Consistently with this intuition, our local result suggests quantifying the “informativeness” of the posterior conditioning using an easily computable $R^2$ coefficient.

While our theoretical results focus on population specification error, in practice PAE are also affected by sampling error, due to the fact that the sample size --- e.g., the number of teachers --- is not infinite. A common approach to account for both sampling variability and specification error is to focus on mean squared error. In general, PAE do not have minimum mean squared error: indeed, in finite samples, model-based estimators can have smaller mean squared error than PAE. In Bonhomme and Weidner (2018), we show how to construct estimators that minimize mean squared error under local asymptotic misspecification. However, such estimators depend on the neighborhood size. In contrast, PAE do {not} require taking a stand on the degree of misspecification through the size of the neighborhood, and they are simple to implement and do not depend on tuning parameters. To complement the theory, we report the results of a Monte Carlo simulation, where we compare the performance of the PAE to those of a model-based estimator and a nonparametric deconvolution-based estimator. We find that, while the model-based estimator tends to perform best under correct specification, the performance of the PAE appears less sensitive to misspecification than those of the model-based and nonparametric estimators.

To illustrate the scope of PAE for applications, we then consider two empirical settings. In the first one, we study the estimation of neighborhood/place effects in the US. Chetty and Hendren (2017) report estimates of the variance of neighborhood effects, as well as EB estimates of those effects. Our goal is to estimate the distribution of effects across neighborhoods. We find that, when using a normal prior as in Chetty and Hendren (2017), our posterior estimator of the distribution function of neighborhood effects across commuting zones is not normal. However, we also show through simulations and computation of our posterior informativeness measure that the signal-to-noise ratio in the data is not high enough to be confident about the exact shape of the distribution. Hence, in this setting, PAE inform our knowledge of the distribution of neighborhood effects, and motivate future analyses using more flexible model specifications and individual-level data.

In the second empirical illustration, our goal is to estimate the distributions of latent components in a permanent-transitory model of income dynamics (e.g., Hall and Mishkin, 1982, Blundell et al., 2008), where log-income is the sum of a random-walk component and a component that is independent over time. Researchers often estimate the covariance structure of the latent components in a first step. Then, in order to document distributions or to use the income process in a consumption-saving model, they often assume Gaussianity. However, there is increasing evidence that income components are not Gaussian (e.g., Geweke and Keane, 2000, Hirano, 2002, Bonhomme and Robin, 2010, Guvenen et al., 2016). We estimate posterior distribution functions of permanent and transitory income components using recent waves from the Panel Study of Income Dynamics (PSID). Our PAE estimates suggest some departure from Gaussianity, especially for the transitory income component.

We analyze several extensions. First, we describe the form of PAE in several models, including binary choice and censored regression. Second, we discuss how to construct confidence intervals and specification tests based on PAE. Lastly, we revisit the question of optimality of EB estimates for {predicting} individual parameters. By extending our misspecification analysis from worst-case specification error of sample averages to worst-case mean squared prediction error, we show that EB estimators remain optimal, up to small-order terms, under local deviations from normality.

\paragraph{Related literature and outline.}

PAE are closely related to parametric EB estimators (Efron and Morris, 1973, Morris, 1983). For recent econometric applications of shrinkage methods (James and Stein, 1961, Efron, 2012), see Hansen (2016), Fessler and Kasy (2018), and Abadie and Kasy (2018). Recent contributions to nonparametric EB methods are Koenker and Mizera (2014) and Ignatiadis and Wager (2019).

Our analysis is also related to deconvolution and other nonparametric approaches. However, in our framework we allow for forms of misspecification under which the quantity of interest is not consistently estimable, and we search for estimators that have the smallest specification error.

In panel data settings, Arellano and Bonhomme (2009) study the asymptotic properties of random-effects estimators of averages of functions of covariates and individual effects. They show that, when the distribution of individual effects is misspecified whereas the other features of the model are correctly specified, PAE are consistent as $n$ and $T$ tend to infinity. By contrast, in our setup, only $n$ tends to infinity, and misspecification may affect the entire joint distribution of unobservables.

Our analysis also connects to the literature on robustness to model misspecification (e.g., Huber and Ronchetti, 2009, Kitamura et al., 2013, Andrews et al., 2017, 2020, Armstrong and Koles\'ar, 2018, Bonhomme and Weidner, 2018, Christensen and Connault, 2019). Here our aim is to propose and justify a class of simple, practical estimators.

The plan of the paper is as follows. In Section (ref) we motivate the analysis by considering a fixed-effects model of teacher quality. In Section (ref) we present our framework and derive our main theoretical results. In Section (ref) we illustrate the use of PAE in two empirical settings. In Section (ref) we describe several extensions. Finally, we conclude in Section (ref). Replication codes are available as \href{https://sites.google.com/site/stephanebonhommeresearch/}{\color{blue}{online material}}.

Motivating example: a fixed-effects model

To motivate the analysis, we start by considering the following model

equation[equation omitted — 118 chars of source]

To fix ideas, we will think of $Y_{ij}$ as an average test score of teacher $i$ in classroom $j$, ${\Greekmath 010B}_i$ as the quality of teacher $i$, and ${\Greekmath 0122}_{ij}$ as a classroom-specific shock. There are $n$ teachers and $J$ observations per teacher. For simplicity, we abstract away from covariates (such as students' past test scores), but those will be present in the framework we will introduce in the next section. Although here we focus on teacher effects, this model is of interest in other settings, such as the study of neighborhood effects, school effectiveness, or hospital quality, for example.

Suppose we wish to estimate a feature of the distribution of teacher quality ${\Greekmath 010B}$. As an example, here we consider the distribution function of ${\Greekmath 010B}$ at a particular point $a$, $$F_{{\Greekmath 010B}}(a)=\mathbb{E}\left[ \boldsymbol{1}\{{\Greekmath 010B}\leq a\}\right],$$ which is the percentage of teachers whose quality is below $a$. A first estimator is the empirical distribution of the fixed-effects estimates $\widehat{{\Greekmath 010B}}_i=\overline{Y}_i=\frac{1}{J}\sum_{j=1}^JY_{ij}$, for all teachers $i=1,...,n$; that is,

equation[equation omitted — 139 chars of source]

where FE stands for “fixed-effects”. An obvious issue with this estimator is that $\overline{Y}_i={\Greekmath 010B}_i+\overline{{\Greekmath 0122}}_i$ is a noisy estimate of ${\Greekmath 010B}_{i}$, where $\overline{{\Greekmath 0122}}_i=\frac{1}{J}\sum_{j=1}^J{\Greekmath 0122}_{ij}$. Indeed, due to the presence of noise, for fixed $J$ and $n$ tends to infinity the distribution $\widehat{F}^{\rm FE}_{{\Greekmath 010B}}$ tends to be {too dispersed} relative to $F_{{\Greekmath 010B}}$ (although one can show that $\widehat{F}^{\rm FE}_{{\Greekmath 010B}}(a)$ is consistent for $F_{{\Greekmath 010B}}(a)$ as $J$ tends to infinity jointly with $n$ under mild conditions, see Jochmans and Weidner, 2018).

A different strategy is to model the joint distribution of ${\Greekmath 010B},{\Greekmath 0122}_{1},...,{\Greekmath 0122}_{J}$. A simple specification is a multivariate normal distribution with means ${\Greekmath 0116}_{{\Greekmath 010B}}$ and ${\Greekmath 0116}_{{\Greekmath 0122}}=0$, and variances $s_{{\Greekmath 010B}}^2$ and $s_{{\Greekmath 0122}}^2$. This specification can easily be made more flexible by allowing for different $s_{{\Greekmath 0122}_j}^2$'s across $j$, for correlation between the different ${\Greekmath 0122}_j$'s, or for means and variances being functions of covariates, for example. Under the assumption that all components are uncorrelated, ${\Greekmath 0116}_{\Greekmath 010B}$, $s_{{\Greekmath 010B}}^2$ and $s^2_{{\Greekmath 0122}}$ can be consistently estimated for fixed $J$ as $n$ tends to infinity, using quasi-maximum likelihood or minimum distance based on mean and covariance restrictions.

Given estimates $\widehat{{\Greekmath 0116}}_{{\Greekmath 010B}}$, $\widehat{s}_{{\Greekmath 010B}}^2$, $\widehat{s}_{{\Greekmath 0122}}^2$, we can compute empirical Bayes (EB) estimates (Morris, 1983) of the ${\Greekmath 010B}_i$ as

equation[equation omitted — 228 chars of source]

where the expectation is taken with respect to the posterior distribution of ${\Greekmath 010B}$ given $Y=Y_i$ for $\widehat{{\Greekmath 0116}}_{{\Greekmath 010B}}$, $\widehat{s}_{{\Greekmath 010B}}^2$, $\widehat{s}_{{\Greekmath 0122}}^2$ fixed, and $\widehat{{\Greekmath 011A}}=\frac{\widehat{s}_{{\Greekmath 010B}}^2}{\widehat{s}_{{\Greekmath 010B}}^2+\widehat{s}_{{\Greekmath 0122}}^2/J}$ is a shrinkage factor. Here, $Y_i$ are vectors containing all $Y_{ij}$, $j=1,...,J$. The EB estimates in ((ref)) are well-justified as predictors of the ${\Greekmath 010B}_i$, since (when treating $\widehat{{\Greekmath 0116}}_{{\Greekmath 010B}}$, $\widehat{s}_{{\Greekmath 010B}}^2$, $\widehat{s}_{{\Greekmath 0122}}^2$ as fixed) $\widehat{{\Greekmath 0116}}_{{\Greekmath 010B}}+\widehat{{\Greekmath 011A}}(\overline{Y}_i-\widehat{{\Greekmath 0116}}_{{\Greekmath 010B}})$ is the minimum mean squared error predictor of ${\Greekmath 010B}_i$ under normality.

Given their rationale for prediction purposes, it is appealing to try and aggregate the EB estimates in order to estimate our target quantity $F_{{\Greekmath 010B}}(a)$. A possible estimator is

equation[equation omitted — 274 chars of source]

where PM stands for “posterior means”. For fixed $J$ as $n$ tends to infinity, the EB estimates tend to be {less dispersed} than the true ${\Greekmath 010B}_i$, and $\widehat{F}^{\rm PM}_{{\Greekmath 010B}}(a)$ is inconsistent in general. Indeed, while in large samples the variance of the fixed-effects estimates is ${\Greekmath 011A}^{-1}s_{{\Greekmath 010B}}^2>s_{{\Greekmath 010B}}^2$, the variance of the EB estimates is ${\Greekmath 011A} s_{{\Greekmath 010B}}^2<s_{{\Greekmath 010B}}^2$, where ${\Greekmath 011A}=\frac{s_{{\Greekmath 010B}}^2}{s_{{\Greekmath 010B}}^2+s_{{\Greekmath 0122}}^2/J}$.

Instead of computing the distribution of EB estimates as in ((ref)), a related idea is to compute the posterior distribution estimator $$ \widehat{F}^{\rm P}_{{\Greekmath 010B}}(a)=\frac{1}{n}\sum_{i=1}^n \mathbb{E}\,\left[ \boldsymbol{1}\{{\Greekmath 010B}\leq a\}\,|\, Y=Y_i\right],$$ where P stands for “posterior”. Using the normality assumption, we obtain

equation[equation omitted — 336 chars of source]

where $\Phi$ denotes the distribution function of the standard normal. $\widehat{F}^{\rm P}_{{\Greekmath 010B}}(a)$ is an example of a {posterior average effect} (PAE). One can check that it is consistent for fixed $J$ as $n$ tends to infinity, when the distribution of ${\Greekmath 010B},{\Greekmath 0122}_{1},...,{\Greekmath 0122}_{J}$ is normal. Under non-normality, $\widehat{F}^{\rm P}_{{\Greekmath 010B}}(a)$ is generally inconsistent for fixed $J$ as $n$ tends to infinity. Moreover, the mean and variance of $\widehat{F}^{\rm P}_{{\Greekmath 010B}}$ are $(1-\widehat{{\Greekmath 011A}})\widehat{{\Greekmath 0116}}_{{\Greekmath 010B}}+\widehat{{\Greekmath 011A}}\frac{1}{n}\sum_{i=1}^n\overline{Y}_i$ and $(1-\widehat{{\Greekmath 011A}})\widehat{s}_{{\Greekmath 010B}}^2+\widehat{{\Greekmath 011A}}^2\left[\frac{1}{n}\sum_{i=1}^n\overline{Y}_i^2-(\frac{1}{n}\sum_{i=1}^n\overline{Y}_i)^2\right]$, respectively, which are consistent for ${\Greekmath 0116}_{{\Greekmath 010B}}$ and $s_{{\Greekmath 010B}}^2$ for fixed $J$ as $n$ tends to infinity.

The last estimator we consider here is directly based on the normal specification for ${\Greekmath 010B}$,

equation[equation omitted — 187 chars of source]

where M stands for “model”. This estimator enjoys attractive properties when the distribution of ${\Greekmath 010B},{\Greekmath 0122}_{1},...,{\Greekmath 0122}_{J}$ is indeed normal. In this case, $\widehat{F}^{\rm M}_{{\Greekmath 010B}}(a)$ is consistent for fixed $J$ as $n$ tends to infinity, and it is efficient when $\widehat{{\Greekmath 0116}}_{{\Greekmath 010B}}$ and $\widehat{s}_{{\Greekmath 010B}}^2$ are maximum likelihood estimates. Moreover, the mean and variance of $\widehat{F}^{\rm M}_{{\Greekmath 010B}}$ are $\widehat{{\Greekmath 0116}}_{{\Greekmath 010B}}$ and $\widehat{s}_{{\Greekmath 010B}}^2$, which are consistent irrespective of normality. However, when ${\Greekmath 010B},{\Greekmath 0122}_{1},...,{\Greekmath 0122}_{J}$ is {not} normally distributed, $\widehat{F}^{\rm M}_{{\Greekmath 010B}}(a)$ is generally inconsistent for fixed $J$ as $n$ tends to infinity. Moreover, $\widehat{F}^{\rm M}_{{\Greekmath 010B}}(a)$ only depends on the data through the mean $\widehat{{\Greekmath 0116}}_{{\Greekmath 010B}}$ and the variance $\widehat{s}^2_{{\Greekmath 010B}}$. In particular, $\widehat{F}^{\rm M}_{{\Greekmath 010B}}$ is always normal, even when the data show clear evidence of non-normality.

Which one of these estimators should one use? The answer is not obvious, since they are all inconsistent as $n$ tends to infinity for fixed $J$ in general. In a framework that allows for misspecification of the normal distribution of ${\Greekmath 010B},{\Greekmath 0122}_{1},...,{\Greekmath 0122}_{J}$, we will show that the PAE $\widehat{F}^{\rm P}_{{\Greekmath 010B}}(a)$ has minimum worst-case specification error in certain neighborhoods around the normal reference distribution. To our knowledge, unlike the other three estimators above, posterior estimators of distributions are novel to practitioners. They are easy to implement, and do not depend on additional tuning parameters. Our characterization provides a rationale for reporting them in applications, alongside other parametric and semi-parametric estimators.

Note that one may wish to relax normality by making the specification of ${\Greekmath 010B}$, and possibly ${\Greekmath 0122}_j$, more flexible. Deconvolution and nonparametric maximum likelihood estimators are often used for this purpose (e.g., Delaigle et al., 2008, Bonhomme and Robin, 2010, Koenker and Mizera, 2014). While these estimators may be consistent even when ${\Greekmath 010B}$ is not normal, consistency relies on additional restrictions on the model. For example, the assumptions in Kotlarski (1967) require that ${\Greekmath 010B}$, ${\Greekmath 0122}_1$, \ldots, ${\Greekmath 0122}_J$ be mutually {independent}. By contrast, we do {not} impose any such additional conditions in our framework. In Section (ref), we will show that asymptotically linear estimators have larger specification error than PAE under the form of misspecification that we consider.

To illustrate that an independence assumption among ${\Greekmath 010B}$, ${\Greekmath 0122}_1$, \ldots, ${\Greekmath 0122}_J$ can be restrictive, consider a situation where the researcher is concerned that the variance of ${\Greekmath 0122}_j$ depends on ${\Greekmath 010B}$. For instance, the variance of classroom-level shocks may depend on teacher quality. The presence of such conditional heteroskedasticity would invalidate conventional nonparametric deconvolution estimators. By contrast, we will show that $\widehat{F}^{\rm P}_{{\Greekmath 010B}}(a)$ has minimum specification error in neighborhoods of distributions that allow for conditional heteroskedasticity. In Section (ref) and the appendix, we will compare the finite-sample behavior of the parametric model-based estimator, the PAE, and a nonparametric deconvolution estimator, in data simulated from various specifications of model ((ref)).

In model ((ref)), the researcher may be interested in estimating other quantities. As an example, consider the coefficient in the population regression of teacher quality ${\Greekmath 010B}$ on a vector of covariates $W$; that is,

equation[equation omitted — 128 chars of source]

In applications, it is common to regress fixed-effects estimates on covariates to help interpret them (as in Dobbie and Fryer, 2013, among many others), and to compute

equation[equation omitted — 148 chars of source]

Alternatively, one may regress the EB estimates of ${\Greekmath 010B}_i$, as given by ((ref)), on covariates (as in Angrist et al., 2017, and Hull, 2018, for example), and compute

equation[equation omitted — 282 chars of source]

which is a PAE based on a normal reference specification for ${\Greekmath 010B}$. We will see that, in our framework, the rationale for reporting $\widehat{{\Greekmath 010E}}^{\rm P}$ or $\widehat{{\Greekmath 010E}}^{\rm FE}$ depends on the form of misspecification that the researcher is concerned about.

The framework we describe next applies to the estimation of different quantities in a variety of settings. In Section (ref) we apply PAE to model ((ref)) and estimate the distribution of neighborhood/place effects in the US (Chetty and Hendren, 2017). In addition, we show that the permanent-transitory model of income dynamics (e.g., Hall and Mishkin, 1982) has a structure similar to model ((ref)), and we report PAE estimates in this context. Lastly, in other models --- such as static or dynamic discrete choice models and models with censored outcomes --- our results motivate the use of PAE as complements to other estimators that researchers commonly report, and we provide examples in Section (ref) and analyze them in the appendix.

Framework and main results

In this section we describe our framework to study PAE, and present our main results.

Model-based estimators and PAE

We consider the following class of models,

equation[equation omitted — 66 chars of source]

where outcomes $Y_i$ and covariates $X_i$ are observed by the researcher, and $U_i$ are unobserved. The function $g_{{\Greekmath 010C}}$ is known up to the finite-dimensional parameter ${\Greekmath 010C}$. Our aim is to estimate an average effect of the form

equation[equation omitted — 137 chars of source]

where ${\Greekmath 010E}_{{\Greekmath 010C}}$ is scalar, and known given ${\Greekmath 010C}$. Here $f_0$ denotes the true density of $U\,|\,X$. The expectation is taken with respect to the product $f_0 f_X$, where $f_X$ is the marginal density of $X$. For conciseness we leave the dependence on $f_X$ implicit. While we focus on a scalar ${\Greekmath 010E}_{{\Greekmath 010C}}$, our results continue to hold in the vector-valued case, as we show at the end of this section. In Appendix (ref), we discuss how to estimate quantities that depend on $f_0$ nonlinearly.

While the researcher does not know the true $f_0$, she has a reference parametric density $f_{{\Greekmath 011B}}$ for $U\,|\, X$, which depends on a finite-dimensional parameter ${\Greekmath 011B}$. We will allow $f_{{\Greekmath 011B}}$ to be misspecified, in the sense that $f_0$ may not belong to $\{f_{{\Greekmath 011B}}\}$. However, we will always assume that $g_{{\Greekmath 010C}}$ is correctly specified. In other words, misspecification will only affect the distribution of $U$ and its dependence on $X$, not the structural link between $(U,X)$ and outcomes.

To estimate $\overline{{\Greekmath 010E}}$ in ((ref)), we assume that the researcher has an estimator $\widehat{{\Greekmath 010C}}$ that remains consistent for ${\Greekmath 010C}$ under misspecification of $f_{{\Greekmath 011B}}$. More precisely, we will only consider potential true densities $f_0$ such that $\widehat{{\Greekmath 010C}}$ tends in probability to the true value ${\Greekmath 010C}$ under $f_0$. For example, in the fixed-effects model ((ref)), consistent estimates of means and variances can be obtained in the absence of normality.

To map model ((ref)) to the general notation of this section, note that in this case there are no covariates $X$, and the vector of unobservables $U$ is $$U=\left(\frac{{\Greekmath 010B}-{\Greekmath 0116}_{{\Greekmath 010B}}}{s_{{\Greekmath 010B}}},\frac{{\Greekmath 0122}_1}{s_{{\Greekmath 0122}}},...,\frac{{\Greekmath 0122}_J}{s_{{\Greekmath 0122}}}\right)'.$$ The vector ${\Greekmath 010C}$ is ${\Greekmath 010C}=({\Greekmath 0116}_{{\Greekmath 010B}},s_{{\Greekmath 010B}}^2,s_{{\Greekmath 0122}}^2)'$. The reference distribution for $U$ is a standard multivariate normal, so the reference density $f_{{\Greekmath 011B}}$ is known in this case --- in other words, the parameter ${\Greekmath 011B}$ in $f_{{\Greekmath 011B}}$ can be omitted. We assume that the researcher has computed an estimator $\widehat{{\Greekmath 010C}}$, for example by quasi-maximum likelihood or minimum distance, which remains consistent for ${\Greekmath 010C}$ when $U$ is not normally distributed.

In certain applications, the reference density depends on some parameters ${\Greekmath 011B}$ that cannot be consistently estimated absent parametric assumptions. In Appendix (ref), we describe discrete choice and censored regression models that have this structure. In such settings, we assume that the researcher has an estimator $\widehat{{\Greekmath 011B}}$ that tends in probability to some ${\Greekmath 011B}_*$ under $f_0$. Unlike ${\Greekmath 010C}$, the parameter ${\Greekmath 011B}_*$ is a model-specific “pseudo-true value” that is not assumed to have generated the data. However, in our leading example of model ((ref)), as well as in the model's generalizations that we study in our empirical illustrations in Section (ref), the references to $\widehat{{\Greekmath 011B}}$ and ${\Greekmath 011B}_*$ can be omitted from all subsequent statements and derivations.

Given $\widehat{{\Greekmath 010C}}$, $\widehat{{\Greekmath 011B}}$, a sample $\{Y_i,X_i,\, i=1,...,n\}$ from $(Y,X)$, and the parametric density $f_{{\Greekmath 011B}}$, a {model-based} estimator of $\overline{{\Greekmath 010E}}$ is

equation[equation omitted — 225 chars of source]

where, with some abuse of notation, the expectation with respect to $f_{\widehat{{\Greekmath 011B}}}$ is computed only over $U$. When not available in closed form, this estimator can be computed by numerical integration or simulation under the parametric density $f_{\widehat{{\Greekmath 011B}}}$. It is easy to see that, under standard conditions, $\widehat{{\Greekmath 010E}}^{\rm M}$ is consistent for $\overline{{\Greekmath 010E}}$ under correct specification; that is, when $f_{{\Greekmath 011B}_*}$ is the true density of $U\,|\, X$.

To construct a posterior estimator, consider the posterior density $p_{{\Greekmath 010C},{\Greekmath 011B}}$ of $U\,|\, Y,X$. This posterior density is computed using Bayes rule, based on the prior $f_{{\Greekmath 011B}}$ on $U\,|\, X$ and the likelihood of $Y\,|\, U,X$ implied by $g_{{\Greekmath 010C}}$. Formally, let ${\cal{U}}(y,x,{\Greekmath 010C})=\{u\,:\, y=g_{{\Greekmath 010C}}(u,x)\}$. We define, whenever the denominator is non-zero,

equation[equation omitted — 275 chars of source]

We will compute $p_{{\Greekmath 010C},{\Greekmath 011B}}$ analytically in our examples. In Appendix (ref) we describe a simulation-based computational approach when an analytical expression is not available. We define the {posterior average effect} (PAE) as the posterior estimator

equation[equation omitted — 263 chars of source]

where, again, the expectation is only taken over $U$. Under standard regularity conditions, it is easy to see that, like $\widehat{{\Greekmath 010E}}^{\rm M}$, the PAE $\widehat{{\Greekmath 010E}}^{\rm P}$ is consistent for $\overline{{\Greekmath 010E}}$ under correct specification.

From a Bayesian perspective, $\widehat{{\Greekmath 010E}}^{\rm P}$ is a natural estimator to consider when ${\Greekmath 010C}$ and ${\Greekmath 011B}$ are known. Indeed, $\widehat{{\Greekmath 010E}}^{\rm P}$ is then the posterior mean of $\frac{1}{n}\sum_{i=1}^n{\Greekmath 010E}_{{\Greekmath 010C}} (U_i,X_i)$, where the prior on $U_i$ is $f_{{\Greekmath 011B}}$, independent across $i$. An alternative Bayesian interpretation is obtained by specifying a nonparametric prior on $f_0$, and computing the posterior mean of $\overline{{\Greekmath 010E}}$ under this prior, as we discuss in Appendix (ref) in the case where $U$ has finite support. However, a frequentist justification for $\widehat{{\Greekmath 010E}}^{\rm P}$ appears to be lacking in the literature. Indeed, under correct specification of $f_{{\Greekmath 011B}}$, both estimators $\widehat{{\Greekmath 010E}}^{\rm P}$ and $\widehat{{\Greekmath 010E}}^{\rm M}$ are consistent, and, as we pointed out in the previous section, $\widehat{{\Greekmath 010E}}^{\rm P}$ may have a higher variance than $\widehat{{\Greekmath 010E}}^{\rm M}$. The key difference between model-based and posterior estimators is that $\widehat{{\Greekmath 010E}}^{\rm P}$ is conditional on the observation sample. An intuitive rationale for the conditioning is the recognition that realizations $Y_i$ may be informative about the values of the unknown $U_i$'s. We next formalize this intuition in a framework that accounts for specification error.

Neighborhoods, estimators, and worst-case specification error

Let $P({\Greekmath 010C},f_0)$ denote the true density of $(Y,U,X)$, where as before we omit the reference to the marginal density of $X$ for conciseness. We assume that, under $P({\Greekmath 010C},f_0)$, $\widehat{{\Greekmath 010C}}$ is consistent for the true ${\Greekmath 010C}$, and $\widehat{{\Greekmath 011B}}$ is consistent for a model-specific “pseudo-true” value ${\Greekmath 011B}_{*}$, where $\mathbb{E}_{P({\Greekmath 010C},f_0)} [{\Greekmath 0120}_{{\Greekmath 010C},{\Greekmath 011B}_{*}}(Y,X)]=0$ for some moment function ${\Greekmath 0120}$. For example, $\widehat{{\Greekmath 010C}}$ and $\widehat{{\Greekmath 011B}}$ may be the method-of-moments estimators that solve $\sum_{i=1}^n {\Greekmath 0120}_{\widehat {\Greekmath 010C},\widehat {\Greekmath 011B}}(Y_i,X_i) = 0$. In models with no ${\Greekmath 011B}$ parameters, such as model ((ref)) and its generalizations, we only assume that $\widehat{{\Greekmath 010C}}$ is consistent for ${\Greekmath 010C}$, and that $\mathbb{E}_{P({\Greekmath 010C},f_0)} [{\Greekmath 0120}_{{\Greekmath 010C}}(Y,X)]=0$ for some ${\Greekmath 0120}$. Throughout, we take the estimators $\widehat{{\Greekmath 010C}}$ (and possibly $\widehat{{\Greekmath 011B}}$), and the moment function ${\Greekmath 0120}$, as given. In particular, we do not address the question of optimal estimation of ${\Greekmath 010C}$ under misspecification.

Given a distance measure $d$ and a scalar ${\Greekmath 010F}\geq 0$, we define the following {neighborhood} of the reference density $f_{{\Greekmath 011B}}$: $$\Gamma_{{\Greekmath 010F}}=\left\{f_0\,:\, d(f_0,f_{{\Greekmath 011B}_*})\leq {\Greekmath 010F},\,\,\, \mathbb{E}_{P({\Greekmath 010C},f_0)} [{\Greekmath 0120}_{{\Greekmath 010C},{\Greekmath 011B}_*}(Y,X)]=0\right\}.$$ This neighborhood consists of densities of $U\,|\, X$ that are at most ${\Greekmath 010F}$ away from $f_{{\Greekmath 011B}_*}$, and under which $\widehat{{\Greekmath 010C}}$ and $\widehat{{\Greekmath 011B}}$ converge asymptotically to ${\Greekmath 010C}$ and ${\Greekmath 011B}_*$, respectively. The case ${\Greekmath 010F}=0$ corresponds to correct specification of the reference density, whereas ${\Greekmath 010F}>0$ corresponds to misspecification.

For ease of notation we omit the dependence of $\Gamma_{{\Greekmath 010F}}$ on ${\Greekmath 010C}$, ${\Greekmath 011B}_*$, and ${\Greekmath 0120}$, all of which we consider fixed and given in this section. Indeed, we assume that the researcher has chosen an estimator $\widehat{{\Greekmath 010C}}$, and, depending on the setting, an estimator $\widehat{{\Greekmath 011B}}$ --- our theory is silent about where these choices come from --- and that she has already observed their realized values in a large sample. The moment function ${\Greekmath 0120}$ is determined by this choice of estimators. Moreover, in large samples, the population values ${\Greekmath 010C}$ and ${\Greekmath 011B}_*$ are arbitrarily close to the observed values $\widehat{{\Greekmath 010C}}$ and $\widehat{{\Greekmath 011B}}$. In our setup, we only consider densities of unobservables $f_0$ that are consistent with those values, in the sense that the moment restriction $\mathbb{E}_{P({\Greekmath 010C},f_0)} [{\Greekmath 0120}_{{\Greekmath 010C},{\Greekmath 011B}_*}(Y,X)]=0$ holds. This large-sample logic is consistent with our focus on specification error; see ((ref)) below.

Note that the same logic might suggest imposing that other features of the joint population distribution of the data $(Y,X)$, such as means, covariances, higher-order moments, or even the entire distribution, be kept constant for all $f_0\in\Gamma_{{\Greekmath 010F}}$. Restricting neighborhoods in this way does not affect the results in this section, because those are valid for all possible ${\Greekmath 0120}$, and one could thus impose additional moment restrictions on $f_0$.

Let us denote the supports of $X$ and $U$ as ${\cal{X}}$ and ${\cal{U}}$, respectively. We assume that $d$ is a ${\Greekmath 011E}${-divergence} of the form $$d(f_0,f_{{\Greekmath 011B}})=\int_{{\cal{X}}}\int_{{\cal{U}}} {\Greekmath 011E} \left(\frac{f_0(u\,|\, x)}{f_{{\Greekmath 011B}}(u\,|\, x)}\right)f_{{\Greekmath 011B}}(u\,|\, x) \, f_X(x) \, du \, dx,$$ where ${\Greekmath 011E}$ is a convex function that satisfies ${\Greekmath 011E}(1)=0$ and ${\Greekmath 011E}''(1)>0$. This family contains as special cases the ${\Greekmath 011F}^2$ divergence (averaged over $X$), the Kullback-Leibler divergence, the Hellinger distance, and more generally the members of the Cressie-Read family of divergences (Cressie and Read, 1984). It is commonly used to measure misspecification, see Andrews et al. (2020) and Christensen and Connault (2019) for recent examples.

We focus on asymptotically linear estimators of $\overline{{\Greekmath 010E}}$ that satisfy, for a scalar non-stochastic function ${\Greekmath 010D}$ and as $n$ tends to infinity,

equation[equation omitted — 221 chars of source]

Note that $\widehat{{\Greekmath 010E}}_{{\Greekmath 010D}}$ depends on $\widehat{{\Greekmath 010C}},\widehat{{\Greekmath 011B}}$, but for conciseness we leave the dependence implicit in the notation. Many estimators can be written in this form (see, e.g., Bickel et al., 1993). Given an estimator $\widehat{{\Greekmath 010E}}_{{\Greekmath 010D}}$, we define its ${\Greekmath 010F}$-worst-case {specification error} as

align[align omitted — 305 chars of source]

We will take the worst-case specification error $b_{{\Greekmath 010F}}({\Greekmath 010D})$ to be our measure of how well an estimator $\widehat{{\Greekmath 010E}}_{{\Greekmath 010D}}$ performs under misspecification. It quantifies the maximum discrepancy, under any possible $f_0$ in the neighborhood $\Gamma_{{\Greekmath 010F}}$, between the probability limit of the estimator and the true parameter value. Under suitable regularity conditions, $\mathbb{E}_{P({\Greekmath 010C},f_0)}[{\Greekmath 010D}_{{{\Greekmath 010C}},{{\Greekmath 011B}}_{*}}(Y,X)]-\mathbb{E}_{f_0}[{\Greekmath 010E}_{{\Greekmath 010C}}(U,X)]$ in ((ref)) is the asymptotic bias of $\widehat{{\Greekmath 010E}}_{{\Greekmath 010D}}$ under $P({\Greekmath 010C},f_0)$.

By focusing on the worst-case specification error $b_{{\Greekmath 010F}}({\Greekmath 010D})$, we abstract from other sources of estimation error. Importantly, we do not account for sampling variability. In Bonhomme and Weidner (2018), we study an alternative approach that consists in minimizing worst-case mean squared error under a local asymptotic --- i.e., as ${\Greekmath 010F}$ tends to zero, $n$ tends to infinity, and ${\Greekmath 010F} n$ tends to a positive constant. Applying this approach to the present case gives estimators that have a smaller worst-case mean squared error than PAE in general. However, unlike PAE, minimum-MSE estimators depend on ${\Greekmath 010F}$, as we will discuss Subsection (ref) below. Relative to such estimators, PAE do not require the researcher to take a stand on the degree of misspecification ${\Greekmath 010F}$, and they are easy to implement.

Result under small-${\Greekmath 010F}$ misspecification

Before stating our first main result, we first characterize the worst-case specification error $b_{{\Greekmath 010F}}({\Greekmath 010D})$ of estimators $\widehat{{\Greekmath 010E}}_{{\Greekmath 010D}}$ for small ${\Greekmath 010F}$. For conciseness, in the remainder of this section we suppress the reference to ${\Greekmath 010C},{\Greekmath 011B}_*$ from the notation, and we denote as $\mathbb{E}_*$ and $\limfunc{Var}_*$ expectations and variances that are taken under the reference model $P({\Greekmath 010C},f_{{\Greekmath 011B}_*})$. All proofs are in Appendix (ref).

lemmaLet $\widetilde {\Greekmath 0120}(y,x) = {\Greekmath 0120}(y,x) - \mathbb{E}_* \left[ {\Greekmath 0120}(Y,X) \big| X=x \right]$. Suppose that one of the following conditions holds: \begin{itemize} • ${\Greekmath 011E}(1)=0$, ${\Greekmath 011E}(r)$ is four times continuously differentiable with ${\Greekmath 011E}''(r)>0$ for all $r >0$, $\mathbb{E}_* [{\Greekmath 0120}(Y,X)] = 0$, $\mathbb{E}_* \big[ \widetilde {\Greekmath 0120}(Y,X) \, \widetilde {\Greekmath 0120}(Y,X)'\big] >0$, and $\left| {\Greekmath 010D}(y,x) \right|$, $\left| {\Greekmath 010E}(u,x) \right| $, $\left| {\Greekmath 0120}(y,x) \right|$ are bounded over the domain of $Y$, $U$, $X$. • Condition (ii) of Lemma (ref) in Appendix (ref) holds (this alternative condition allows for unbounded ${\Greekmath 010D}$, ${\Greekmath 010E}$, ${\Greekmath 0120}$, but at the cost of stronger assumptions on ${\Greekmath 011E}(r)$). \end{itemize} Then, as ${\Greekmath 010F}$ tends to zero we have \begin{align*} &b_{{\Greekmath 010F}}({\Greekmath 010D})=\left|\mathbb{E}_{*}[{\Greekmath 010D}(Y,X)-{\Greekmath 010E}(U,X)]\right| \\ & \,\,\, +{\Greekmath 010F}^{\frac{1}{2}}\bigg\{\frac{2}{{\Greekmath 011E}”(1)} {\rm Var}_*\Big( {\Greekmath 010D}(Y,X)- {\Greekmath 010E}(U,X) - \mathbb{E}_{*} \left[{\Greekmath 010D}(Y,X)-{\Greekmath 010E}(U,X) \, | \, X \right] -{\Greekmath 0115}' \widetilde {\Greekmath 0120}(Y,X) \Big) \bigg\}^{\frac{1}{2}} +{\cal O}({\Greekmath 010F}), \end{align*} where ${\Greekmath 0115}=\left\{\mathbb{E}_* \big[ \widetilde {\Greekmath 0120}(Y,X) \, \widetilde {\Greekmath 0120}(Y,X)'\big]\right\}^{-1} \mathbb{E}_*\left[\left({{\Greekmath 010D}}(Y,X)-{{\Greekmath 010E}}(U,X)\right) \widetilde {\Greekmath 0120}(Y,X)\right]$.

\vskip .3cm

To derive the formula for the worst-case specification error in Lemma (ref), we maximize the specification error with respect to $f_0$ subject to three contraints: $f_0$ belongs to an ${\Greekmath 010F}$-neighborhood of $f_*$, it is such that the moment condition is satisfied at $({\Greekmath 010C},{\Greekmath 011B}_*)$, and it is a density. In part $(i)$ we focus on the case where ${\Greekmath 010D}$, ${\Greekmath 010E}$ and ${\Greekmath 0120}$ are bounded. This is satisfied, for example, if those functions and $g(u,x)$ are all continuous, and the domain of $U$ and $X$ is bounded. To accommodate situations where supports are unbounded, such as the example of Section (ref), in part $(ii)$ we allow for unbounded functions ${\Greekmath 010D}$, ${\Greekmath 010E}$ and ${\Greekmath 0120}$, which only requires existence of third moments under the reference distribution. To guarantee that $b_{{\Greekmath 010F}}({\Greekmath 010D})$ is well-defined in the unbounded case, we require a regularization of the function ${\Greekmath 011E}(r)$ for large values of $r$.

Lemma (ref) implies that the small-${\Greekmath 010F}$ specification error of the PAE is, up to smaller-order terms, proportional to the within-$(Y,X)$ standard deviation of ${\Greekmath 010E}(U,X)$ under the reference model: $$b_{{\Greekmath 010F}}({\Greekmath 010D}^{\rm P})={\Greekmath 010F}^{\frac{1}{2}}\left\{\frac{2}{{\Greekmath 011E}''(1)}{\limfunc{Var}}_{*}\left({\Greekmath 010E}(U,X)-\mathbb{E}_{*}[{\Greekmath 010E}(U,X)\,|\, Y,X]\right)\right\}^{\frac{1}{2}}+{\cal O}({\Greekmath 010F}).$$ In the fixed-effects model ((ref)) of teacher quality, the worst-case specification error of the PAE $\widehat{F}^{\rm P}_{{\Greekmath 010B}}(a)$ is $$b_{{\Greekmath 010F}}({\Greekmath 010D}^{\rm P})={\Greekmath 010F}^{\frac{1}{2}}\left\{\frac{4}{{\Greekmath 011E}''(1)}T\left(\frac{a-{{\Greekmath 0116}}_{{\Greekmath 010B}}}{{{\Greekmath 011B}}_{{\Greekmath 010B}}},\sqrt{\frac{1-{\Greekmath 011A}}{1+{\Greekmath 011A}}}\right)\right\}^{\frac{1}{2}}+{\cal O}({\Greekmath 010F}),$$ where $T(a,b)={\Greekmath 0127}(a)\int_0^b\frac{{\Greekmath 0127}(az)}{1+z^2}dz$ is Owen's T function (Owen, 1956), and ${\Greekmath 0127}$ is the standard normal density. The specification error decreases as the number $J$ of observations per teacher increases, and tends to zero as $J$ tends to infinity and the shrinkage factor ${\Greekmath 011A}$ tends to one.

The next theorem, which holds for all functions ${\Greekmath 010D}(y,x)$, subject to regularity conditions, shows that the PAE has minimum worst-case specification error locally.

theoremSuppose that the conditions of Lemma (ref) hold, and let \begin{equation}{\Greekmath 010D}^{\rm P}(y,x)=\mathbb{E}_*[{\Greekmath 010E}(U,X)\,|\, Y=y,X=x].\end{equation} Then, as ${\Greekmath 010F}$ tends to zero we have $$b_{{\Greekmath 010F}}({\Greekmath 010D}^{\rm P})\leq b_{{\Greekmath 010F}}({\Greekmath 010D})+{\cal O}({\Greekmath 010F}).$$

Result under fixed-${\Greekmath 010F}$ misspecification

To show our second main result, let us now focus on the case ${\Greekmath 011E}(t)=\frac 1 2 (t-1)^2$; that is, we choose the distance measure $d(f_0,f_{{\Greekmath 011B}})$ to be the Pearson ${\Greekmath 011F}^2$ divergence. For this quadratic distance measure, we show that PAE satisfy a fixed-${\Greekmath 010F}$ optimality result, which is valid for all values of ${\Greekmath 010F}$ that are smaller than

align[align omitted — 273 chars of source]

where ${\Greekmath 010D}^{\rm P}(y,x)$ is given by (ref).

theoremAssume that $\mathbb{E}_* [{\Greekmath 0120}(Y,X)] = 0$, ${\Greekmath 011E}(t)=\frac 1 2 (t-1)^2$, and that ${\Greekmath 010D}(Y,X)$ and ${\Greekmath 010E}(U,X)$ have finite second moments under the reference model. Then, for $0< {\Greekmath 010F} \leq \overline {\Greekmath 010F}$, we have $$b_{{\Greekmath 010F}}({\Greekmath 010D}^{\rm P})\leq b_{{\Greekmath 010F}}({\Greekmath 010D}).$$

In Theorem (ref) we show that ${\Greekmath 010D}^{\rm P}$ is an exact minimizer of the function $ b_{{\Greekmath 010F}}({\Greekmath 010D})$. This is in contrast with Theorem (ref), where we relied on a small-${\Greekmath 010F}$ approximation. The condition $ {\Greekmath 010F} \leq \overline {\Greekmath 010F}$ guarantees that, for ${\Greekmath 010D}={\Greekmath 010D}^{\rm P}$, the constraint $f_0(u\, |\, x) \geq 0$ is non-binding in the optimization problem over $f_0$ in (ref), implying that the problem has a simple analytic solution. Although, in many settings such as model ((ref)), the parameter of interest $\overline{{\Greekmath 010E}}$ is not consistently estimable under our assumptions, Theorem (ref) shows that PAE achieve the smallest possible worst-case specification error when the true distribution $f_0$ lies sufficiently close to the reference distribution $f_{{\Greekmath 011B}_*}$, as measured according to the ${\Greekmath 011F}^2$ divergence.

If the distance measure $d(f_0,f_{{\Greekmath 011B}})$ is not a ${\Greekmath 011F}^2$-divergence, or if ${\Greekmath 010F} > \overline {\Greekmath 010F}$, then ${\Greekmath 010D}^{\rm P}$ is not the exact minimizer of worst-case specification error $ b_{{\Greekmath 010F}}({\Greekmath 010D})$. Moreover, in such cases the estimator with minimum worst-case specification error depends on ${\Greekmath 010F}$ in general. However, one can still establish a fixed-${\Greekmath 010F}$ bound on worst-case specification error, as the next result shows.

theoremLet ${\Greekmath 010D}^{\rm P}$ be as in ((ref)), and assume that ${\Greekmath 011E}(r)$ is convex with ${\Greekmath 011E}(1)=0$. Then, for all ${\Greekmath 010F}>0$, $$ b_{{\Greekmath 010F}}({\Greekmath 010D}^{\rm P})\leq 2 \, \limfunc{inf}_{{\Greekmath 010D}}\, b_{{\Greekmath 010F}}({\Greekmath 010D}).$$

\vskip .3cm

In Theorem (ref) we establish a fixed-${\Greekmath 010F}$ bound on the worst-case specification error of PAE, which holds for all ${\Greekmath 010F}>0$ and all ${\Greekmath 011E}$-divergences such that ${\Greekmath 011E}$ is convex with ${\Greekmath 011E}(1)=0$. The infimum is taken over all possible functions ${\Greekmath 010D}(y,x)$, subject to measurability conditions, which we implicitly assume throughout the paper. Although $\widehat{{\Greekmath 010E}}^{\rm P}$ may not minimize worst-case specification error for finite ${\Greekmath 010F}$, Theorem (ref) shows that its worst-case specification error is never larger than twice the minimum worst-case specification error. In addition, the factor two in Theorem (ref) cannot be improved upon in general, as we show in Appendix (ref) in the context of a simple binary choice model.

Discussion

In this subsection, we discuss several features and implications of our main results given by Theorems (ref) and (ref).

\paragraph{Uniqueness.}

In the absence of covariates and for known parameters ${\Greekmath 010C}$, ${\Greekmath 011B}_*$, the proof of Theorem (ref) shows that ${\Greekmath 010D}^{\rm P}$ is the unique minimizer of the first-order worst-case specification error. Likewise, ${\Greekmath 010D}^{\rm P}$ is also unique in Theorem (ref). More generally, if covariates are present and the parameters ${\Greekmath 010C}$, ${\Greekmath 011B}_*$ are estimated, then the leading order contribution of $b_{{\Greekmath 010F}}({\Greekmath 010D})$ is minimized if and only if ${\Greekmath 010D}(Y,X) = {\Greekmath 010D}^{\rm P}(Y,X) + {\Greekmath 0121}(X) + {\Greekmath 0115}'{\Greekmath 0120}(Y,X) + o_{P_*}(1) $, for some ${\Greekmath 0115}$ and ${\Greekmath 0121}$ such that $\mathbb{E}_{f_X} [{\Greekmath 0121}(X)] = 0$ --- see part (ii) of Theorem (ref) in Appendix (ref) for a formal statement. Hence, while the PAE is not the unique minimizer of the local worst-case specification error in this case, any minimizer differs from the PAE by a zero-mean function of $X$ and a linear combination of the moment function ${\Greekmath 0120}$. In addition, $\widehat{{\Greekmath 010E}}^{\rm P}$ has smallest variance within the class of minimum worst-case specification error estimators.

\paragraph{Form of misspecification.} Theorems (ref) and (ref) rely on specific distance measures, ${\Greekmath 011F}^2$ divergence for the latter and any member of the ${\Greekmath 011E}$-divergence family for the former. Under other distance measures, the PAE will not have minimum worst-case specification error in general.

Given a distance measure, the theorems are based on nonparametric neighborhoods that consist of unrestricted distributions of $U\,|\, X$, except for the moment conditions that pin down ${\Greekmath 010C}$ and ${\Greekmath 011B}_*$. However, if one is willing to make additional assumptions on $f_0$ that further restrict the neighborhood, then one can construct estimators that are more robust than $\widehat{{\Greekmath 010E}}^{\rm P}$ within a particular class. As an example, consider the fixed-effects model ((ref)). Suppose that, in addition to assuming that ${\Greekmath 010B}$, ${\Greekmath 0122}_1$, ..., ${\Greekmath 0122}_J$ are mutually uncorrelated, the researcher is willing to assume that they are fully independent. In that case, the distribution of ${\Greekmath 010B}$ can be consistently estimated under suitable regularity conditions, provided $J\geq 2$ (Kotlarski, 1967, Li and Vuong, 1998). However, the PAE in ((ref)) is inconsistent for fixed $J$ as $n$ tends to infinity. As a consequence, the PAE does not minimize worst-case specification error in a semi-parametric neighborhood that consists of distributions with independent marginals.

To elaborate further on this point, consider the coefficient $\overline{{\Greekmath 010E}}$ in the population regression of ${\Greekmath 010B}$ on a covariates vector $W$, see (ref). A possible estimator is the coefficient $\widehat{{\Greekmath 010E}}^{\rm FE}$ in the regression of the fixed-effects estimates $\overline{Y}_i$ on $W_i$, see ((ref)). Under correct specification of the reference model, $\widehat{{\Greekmath 010E}}^{\rm FE}$ is consistent for $\overline{{\Greekmath 010E}}$. However, $\widehat{{\Greekmath 010E}}^{\rm FE}$ may be inconsistent under the type of misspecification that we allow for, since ${\Greekmath 0122}_j$ and $W$ may be correlated under $f_0$. For example, $W$ (e.g., teacher absenteeism) may be influenced by ${\Greekmath 010B}$ and factors that correlate with ${\Greekmath 0122}_j$. Theorem (ref) shows that, under such misspecification, the PAE $\widehat{{\Greekmath 010E}}^{\rm P}$ in ((ref)) has minimum worst-case specification error locally. Nevertheless, if the researcher is confident that $W$ should not enter the outcome equation, and that it is independent of ${\Greekmath 0122}_j$, then it is natural to report the consistent estimator $\widehat{{\Greekmath 010E}}^{\rm FE}$.

\paragraph{Posterior informativeness.} Our small-${\Greekmath 010F}$ calculations can be used to compare the worst-case specification errors of the PAE $\widehat{{\Greekmath 010E}}^{\rm P}$ to that of the model-based estimator $\widehat{{\Greekmath 010E}}^{\rm M}$. To see this, let ${\Greekmath 010D}^{\rm M}_{{\Greekmath 010C},{\Greekmath 011B}}(x) = \mathbb{E}_{f_{\Greekmath 011B}} [ {\Greekmath 010E}_{{\Greekmath 010C}}(U,X) \,|\, X=x ]$. Using Lemma (ref), the ratio of the two worst-case specification errors satisfies

equation[equation omitted — 379 chars of source]

where $v(U,X)$ is the population residual of $ ({\Greekmath 010E}(U,X) - {{\Greekmath 010D}}^{\rm M}(X))$ on $\widetilde{{\Greekmath 0120}}(Y,X)$, under the parametric reference model; that is, $v(u,x) = {{\Greekmath 010E}}(u,x) - {{\Greekmath 010D}}^{\rm M}(x) + {\Greekmath 0115}' \widetilde{{\Greekmath 0120}}(g(u,x),x)$, where all functions are evaluated at ${\Greekmath 010C},{\Greekmath 011B}_*$, and ${\Greekmath 0115}$ is as defined in Lemma (ref) for the case ${\Greekmath 010D}={\Greekmath 010D}^{\rm M}$. Intuitively, the robustness of $\widehat{{\Greekmath 010E}}^{\rm P}$ relative to $\widehat{{\Greekmath 010E}}^{\rm M}$ depends on how informative the outcome values $Y_i$ are for the latent individual parameters ${\Greekmath 010E}(U_i,X_i)$.

In practice, we will report an empirical counterpart to the small-${\Greekmath 010F}$ limit of $1-\frac{b_{{\Greekmath 010F}}^2({\Greekmath 010D}^{\rm P})}{b_{{\Greekmath 010F}}^2({\Greekmath 010D}^{\rm M})}$. This quantity can be simply expressed as the $R^2$ in the population nonparametric regression of $v(U,X)$ on $Y,X$ under the reference model; that is,

equation[equation omitted — 141 chars of source]

where with some abuse of notation here $v(U,X)$ denotes the sample residual of $ ( {{\Greekmath 010E}}_{\widehat{{\Greekmath 010C}}}(U,X) - {{\Greekmath 010D}}_{\widehat{{\Greekmath 010C}},\widehat{{\Greekmath 011B}}}^{\rm M}(X))$ on $\widetilde{{\Greekmath 0120}}_{\widehat{{\Greekmath 010C}},\widehat{{\Greekmath 011B}}}(Y,X)$, and expectations and variances are taken with respect to $P(\widehat{{\Greekmath 010C}},f_{\widehat{{\Greekmath 011B}}})$. Using a term from Andrews et al. (2020) --- albeit in a different setting --- we refer to $R^2$ in ((ref)) as a measure of the “informativeness” of the posterior conditioning, and we will report it in our illustrations. As an example, for $\widehat{F}^{\rm P}_{{\Greekmath 010B}}(a)$ in model ((ref)), the informativeness of the posterior conditioning is

equation[equation omitted — 454 chars of source]

In this case the $R^2$ increases with the number $J$ of observations per teacher, and it tends to one as $J$ tends to infinity.

\paragraph{Multi-dimensional PAE.} For simplicity, in this section we have focused on the case where the target parameter $\overline {\Greekmath 010E}$ in (ref) is scalar. However, our results can be extended to multi-dimensional parameters. The definition of worst-case specification error in (ref) is then modified to $$b_{{\Greekmath 010F}}({\Greekmath 010D})=\limfunc{sup}_{f_0\in\Gamma_{{\Greekmath 010F}}}\, \left\| \mathbb{E}_{P({\Greekmath 010C},f_0)}[{\Greekmath 010D}(Y,X) ]-\mathbb{E}_{f_0}[ {\Greekmath 010E}(U,X)]\right\|,$$ where $\| \cdot \|$ is a norm over the vector space in which ${\Greekmath 010D}(Y,X)$ and ${\Greekmath 010E}(U,X)$ take values.

If $\| \cdot \|_*$ denotes the corresponding dual norm, then we can rewrite $b_{{\Greekmath 010F}}({\Greekmath 010D})= \sup_{\|v\|_* = 1} b_{{\Greekmath 010F}}({\Greekmath 010D},v)$, where $b_{{\Greekmath 010F}}({\Greekmath 010D},v) = \limfunc{sup}_{f_0\in\Gamma_{{\Greekmath 010F}}}\, \big| \mathbb{E}_{P({\Greekmath 010C},f_0)}[v' {\Greekmath 010D}(Y,X)] - \mathbb{E}_{f_0}[v' {\Greekmath 010E}(U,X)]\big|$. Our minimum worst-case specification error results for PAE for scalar $\overline {\Greekmath 010E}$ then apply to $b_{{\Greekmath 010F}}({\Greekmath 010D},v)$ for every given vector $v$, and the minimum-specification error properties are preserved after taking the supremum over the set of vectors $v$ with $ \|v\|_* = 1$. Thus, in the multi-dimensional case, PAE minimize worst-case specification error for small ${\Greekmath 010F}$ in the sense of Theorem (ref), and for fixed ${\Greekmath 010F}$ under the conditions of Theorem (ref). In our leading example of Section (ref), suppose we are interested in the entire distribution function $F_{{\Greekmath 010B}}$. In this case, the average effect is a function indexed by $a$. Taking the supremum norm $\|\cdot\|_\infty$ over distribution functions, we obtain that, as an estimator of $F_{{\Greekmath 010B}}$, the PAE minimizes worst-case specification error under suitable conditions.

\paragraph{Mean squared error.} While we have shown that PAE minimize worst-case specification error locally under the conditions of Theorem (ref), and for fixed ${\Greekmath 010F}$ under the conditions of Theorem (ref), PAE generally do not have minimum mean squared error (MSE). To see this, let us assume that ${\Greekmath 010C}$ and ${\Greekmath 011B}_*$ are known. In a local asymptotic framework where $n$ tends to infinity, ${\Greekmath 010F}$ tends to zero, and $n{\Greekmath 010F}$ tends to a positive constant, and under suitable regularity conditions, we show in Appendix (ref) that the estimator with minimum worst-case MSE is given by

align[align omitted — 327 chars of source]

which is a linear combination between the model-based estimator and the PAE. The model-based estimator $\widehat{{\Greekmath 010E}}^{\rm M}$, which has the smallest asymptotic variance, will be preferred when ${\Greekmath 010F}$ is small relative to $1/n$, while the PAE, which has smallest specification error, will be preferred when ${\Greekmath 010F}$ is large relative to $1/n$. However, in order to implement such estimators $\widehat{{\Greekmath 010E}}^{\rm MMSE}$ that minimize worst-case MSE, knowledge of ${\Greekmath 010F}$ is required. See Bonhomme and Weidner (2018) for an approach to minimum-MSE estimation.

Simulations and empirical illustrations

In this section, we study two empirical applications: we estimate the distribution of income neighborhood effects in the US, and the distributions of permanent and transitory earnings components in the PSID. We start the section by summarizing the results of a Monte Carlo simulation exercise, in samples generated from various specifications of model ((ref)).

Monte Carlo simulation: summary of results

While Theorems (ref) and (ref) show that PAE minimize worst-case specification error under small-${\Greekmath 010F}$ and fixed-${\Greekmath 010F}$ misspecification, respectively, they are silent about other forms of estimation error. In Appendix (ref) we report the results of a Monte Carlo simulation exercise, where we compare the performance of PAE and other estimators in finite sample in the fixed-effects model ((ref)), for various specifications. Here we briefly summarize the results from the simulation exercise.

We compare the performance of four estimators: the fixed-effects estimator given by ((ref)), the PAE given by ((ref)), the model-based estimator given by ((ref)), and a nonparametric kernel deconvolution estimator with normal errors (Stefanski and Carroll, 1990). We analyze two sets of data generating processes. When the reference normal distribution for ${\Greekmath 010B}_i$ is correctly specified, the model-based estimator performs best, as expected. We find that, while the PAE has both larger bias and variance than the model-based estimator in this case, it is less biased and less variable than both the nonparametric deconvolution estimator and the fixed-effects estimator, especially when the number of measurements $J$ is small (see Appendix Figure (ref)).

We next turn to data generating processes where ${\Greekmath 010B}_i$ is not normal, drawn from a skewed Beta distribution. We find that the model-based estimator is substantially biased in this case. The nonparametric deconvolution estimator has smallest bias when errors are normally distributed, but it is heavily biased when errors are non-normal. By contrast, although it has no consistency guarantees in these settings, the PAE tends to perform comparatively well in all situations, for bias and variance (see Appendix Figure (ref)).

Overall, the simulations complement our theory by highlighting that, beyond specification error, other sources of estimation error matter in practice. Under correct specification of the reference distribution, the model-based estimator should be preferred. At the same time, our results suggest that, at least in the particular settings we focus on, the performance of the PAE appears less sensitive to misspecification than those of the model-based and nonparametric deconvolution estimators. Moreover, we find that the robustness gains provided by the PAE depend on the signal-to-noise ratio and the informativeness of the posterior conditioning. We provide details on the simulations in Appendix (ref).

Neighborhood effects

In this subsection and the next, we revisit two applications of models with latent variables. In our first illustration, we focus on a model of neighborhood effects following Chetty and Hendren (2017), using data for the US that these authors made public. In our second illustration, we study a permanent-transitory model of income dynamics (Hall and Mishkin, 1982, Blundell et al., 2008) using the PSID. In both cases, we rely on a normal reference specification and assess how and by how much the posterior conditioning informs the estimates of the parameters of interest.

Here we start with estimates of neighborhood (or “place”) effects reported in Chetty and Hendren (2017, CH hereafter). Those were obtained using individuals who moved between different commuting zones at different ages. The outcome variable that we focus on is the causal estimate of the income rank at age 26 of a child whose parents are at the 25 percentile of the income distribution. This is CH's preferred measure of place effect.

CH report an estimate of the variance of neighborhood effects, corrected for noise. In addition, they report individual predictors. Here we are interested in documenting the entire distribution of place effects. To do so, we consider the model $\widehat{{\Greekmath 0116}}_{c}={\Greekmath 0116}_c+\overline{{\Greekmath 0122}}_c$, for each commuting zone $c$, where $\widehat{{\Greekmath 0116}}_{c}$ is a neighborhood-specific fixed-effects reported by CH, ${\Greekmath 0116}_c$ is the true effect of neighborhood $c$, and $\overline{{\Greekmath 0122}}_c$ is additive estimation noise. CH also report estimates $\widehat{s}_{c}^2$ of the variances of $\overline{{\Greekmath 0122}}_c$ for every $c$. When weighted by population, the fixed-effects estimates $\widehat{{\Greekmath 0116}}_{c}$ have mean zero. We treat neighborhoods as independent observations. The statistics we use for calculations are available at: https://opportunityinsights.org/paper/neighborhoodsii/. Given the aggregate data at hand, we necessarily need to assume that estimates $\widehat{{\Greekmath 0116}}_{c}$ are independent across neighborhoods $c$, although this might be restrictive in this setting.

We first estimate the variance of place effects ${\Greekmath 0116}_c$, following CH. We trim the top 1% percentile of $\widehat{s}_{c}^2$, and weigh all results by population weights. While this differs slightly from CH's approach, which is based on $1/\widehat{s}_{c}^2$ precision weights and no trimming, we replicated the analysis using precision weights in the un-trimmed sample and found similar results. We have information about place effects in $C=590$ commuting zones $c$ in our sample, compared to 595 in the sample without trimming. We estimate a sizable variance of neighborhood fixed-effects: $\limfunc{Var}(\widehat{{\Greekmath 0116}}_{c})=.077$. In turn, the mean of $ \widehat{s}_{c}^2$ weighted by population is $\widehat{s}_{\overline{{\Greekmath 0122}}}^2=.047$. Given those, we estimate the variance of place effects as $\widehat{s}_{{\Greekmath 0116}}^2=\limfunc{Var}(\widehat{{\Greekmath 0116}}_{c})-\widehat{s}_{\overline{{\Greekmath 0122}}}^2=.030$. In this setting, the shrinkage factor $\widehat{{\Greekmath 011A}}_c=\widehat{s}_{{\Greekmath 0116}}^2/(\widehat{s}_{{\Greekmath 0116}}^2+\widehat{s}_{c}^2)$ exhibits substantial heterogeneity across commuting zones. Indeed, the mean of $\widehat{{\Greekmath 011A}}_c$ is .62, and its 10% and 90% percentiles are .21 and .93, respectively.

We use a normal with zero mean and variance $\widehat{s}_{{\Greekmath 0116}}^2$ as a prior for ${\Greekmath 0116}_c$. Then, we estimate the distribution function of neighborhood effects ${{\Greekmath 0116}}_c$ using the PAE given by ((ref)); that is,

equation*[equation* omitted — 289 chars of source]

where ${\Greekmath 0119}_c$ are population weights. In addition, in order to ease the visualization of the results, we will also report estimates of densities, which are the derivatives of the PAE of distribution functions. Note that the density of ${\Greekmath 0116}$ at $a$ can be approximated for arbitrarily small $h>0$ by the expectation of $\boldsymbol{1}\{|{\Greekmath 0116}-a|/h\}/2h$. Taking the limit of the corresponding PAE as $h$ tends to zero gives the derivative of $\widehat{F}^{\rm P}_{{\Greekmath 0116}}$ at $a$. We thus expect derivatives of PAE of distribution functions to enjoy similar minimum-worst-case specification error properties as PAE, but we do not formalize the required assumptions here.

In the top panel of Figure (ref), we report several estimates of distribution functions. In the bottom panel, we report the corresponding density estimates. In the left graphs, we show nonparametric kernel estimates of the distribution function (respectively, density) of the fixed-effects $\widehat{{\Greekmath 0116}}_{c}$, weighted by population (in solid), together with the best-fitting normal (in dashed). The graphs show substantial non-normality of the fixed-effects estimates. In particular, the large variance appears to be driven by some large positive and negative estimates $\widehat{{\Greekmath 0116}}_{c}$. In the right graphs, we report the PAE $\widehat{F}^{\rm P}_{{\Greekmath 0116}}$ of the distribution function of true place effects ${\Greekmath 0116}_c$, with the associated density (in solid). In addition, we show the normal prior, with zero mean and variance $\widehat{s}_{{\Greekmath 0116}}^2$ (in dashed). The posterior distribution of neighborhood effects differs from the normal prior, although the two estimators have the same variance by construction. In comparison, neighborhood-specific empirical Bayes estimates have a substantially lower dispersion. In Appendix Figure (ref) we report an estimate of their distribution function $\widehat{F}^{\rm PM}_{{\Greekmath 0116}}$ and associated density. While $\widehat{s}_{{\Greekmath 0116}}^2=.030$ and the variance associated with $\widehat{F}^{\rm P}_{{\Greekmath 0116}}$ is $.030$, the variance of the empirical Bayes estimates is only $.010$. In addition, a specification test that compares model-based estimator and PAE, which we describe in Appendix (ref), suggests that these differences are statistically significant. Indeed, assuming independence across commuting zones, we obtain p-values below .01 at all deciles except the bottom two.

figure[figure omitted — 1,039 chars of source]

To assess how likely it is that the posterior estimator approximates the shape of the distribution of true neighborhood effects, we next perform two different exercises, based on a simulation and on numerical calculations motivated by our theory. We start with a Monte Carlo simulation, where ${\Greekmath 0116}_c$, for $c=1,...,C_{\rm sim}$, are {log-normally} distributed with zero mean and variance $\widehat{s}_{{\Greekmath 0116}}^2$, and $\overline{{\Greekmath 0122}}_c$ are normally distributed independent of ${\Greekmath 0116}_c$ with zero mean. We consider three scenarios for the noise variances $\widehat{s}_c^2$: the estimates from CH, one-third of those values, and one-tenth of those values. In this exercise we again weigh by population. We show the results for $C_{\rm sim}=100,000$ simulated neighborhoods. In the left graphs of Figure (ref) we see that, when the noise variances are the ones from the data, the posterior density is more skewed than the normal, yet the posterior shape is quite different from the true log-normal distribution of ${\Greekmath 0116}_c$. When reducing the noise variances in the middle and right graphs, the posterior distribution function and density estimates get closer to the log-normal ones. In the right graphs, where the shrinkage factor is .90 on average (as opposed to .62 in the data), the posterior distribution function and density approximate the highly non-normal shape of the true distribution of neighborhood effects very well.

figure[figure omitted — 1,356 chars of source]

We next turn to our posterior informativeness measure, which is given by equation ((ref)). Note the $R^2$ coefficient varies along the distribution. We find that the weighted average $R^2$ across values of $a$ is 28%, where we weigh across cutoff values $a$ by the reference distribution for ${\Greekmath 010B}$. This value is consistent with the message of Figure (ref), since it suggests that, while the posterior conditioning informs the shape of the distribution of neighborhood effects, the signal-to-noise ratio is not high enough to be confident about the exact shape.

Lastly, we perform two additional exercises as robustness checks. Firstly, we incorporate the mean income $\overline{y}_c$ of permanent residents in county $c$ at the 25% percentile as a covariate. CH rely on information on permanent residents' income to improve the accuracy of individual predictions. Here we use it to refine the reference distribution and to improve the estimation of the distribution of neighborhood effects. Specifically, our reference model for ${\Greekmath 0116}_c$ is then a correlated random-effects specification, where the mean depends on $\overline{y}_c$ linearly. Appendix Figure (ref) shows small differences with our baseline estimates. Secondly, we re-do our main analysis at the county level, instead of the commuting zone level. In that case, the signal-to-noise ratio is lower, our posterior informativeness $R^2$ measure is 17% on average, and Appendix Figure (ref) shows that the normal prior and the posterior distributions are closer to each other than in the case of commuting zones.

Income dynamics

In this subsection, we consider the following permanent-transitory model of household log-income, $$ Y_{it}={\Greekmath 0111}_{it}+{\Greekmath 0122}_{it}, \qquad {\Greekmath 0111}_{it}={\Greekmath 0111}_{i,t-1}+V_{it},\qquad i=1,...,n,\quad t=1,...,T, $$ where ${\Greekmath 0122}_{it}$ and $V_{it}$ are independent at all lags and leads, and independent of ${\Greekmath 0111}_{i0}$. This process is commonly used as an input for life-cycle consumption/savings models. Researchers often estimate covariances in a first step using minimum distance, and then impose a normality assumption for further analysis. However, there is increasing evidence that income components are {not} normally distributed. Instead of using a more flexible model --- as has been done by Carlton and Hall (1978) and a large subsequent literature --- here we compute posterior average effects. The advantages of this approach are that no additional assumptions are needed, and that implementation is straightforward.

We focus on six recent waves of the PSID 1999-2009 (every other year), see Blundell et al. (2016) for a description of the data. We use the same sample selection as in Arellano et al. (2017), and work with a balanced panel of $n=792$ households over $T=6$ periods. $Y_{it}$ are residuals of log total pre-tax household labor earnings on a set of demographics, which include cohort interacted with education categories for both household members, race, state, and large-city dummies, a family size indicator, number of kids, a dummy for income recipient other than husband and wife, and a dummy for kids out of the household. Our aim is to estimate the distributions of ${\Greekmath 0111}_{it}$ and ${\Greekmath 0122}_{it}$. To do so, we compare normal model-based estimates with posterior estimates, by plotting distribution functions as well as the implied densities. The model's structure is similar to that of the fixed-effects model ((ref)), and analytical expressions for posterior estimators are easy to derive.

figure[figure omitted — 961 chars of source]

In the left graphs of Figure (ref), we show the distribution of the permanent component ${\Greekmath 0111}_{it}$. In the right graphs, we show the distribution of the transitory component ${\Greekmath 0122}_{it}$. We show PAE in solid, and model-based estimators in dashed. In the top panel we report estimates of distribution functions, and in the bottom panel we report the implied density estimates. The estimates show mild deviation from Gaussianity for the permanent component, and stronger evidence of non-Gaussianity for the transitory component. In particular, the latter shows excess kurtosis (i.e., “peakedness”) relative to the normal.

Several papers have already documented the presence of excess kurtosis in income components, particularly in transitory innovations, using parametric or semi-parametric methods. The estimates in Figure (ref) share some qualitative similarities with recent findings in the literature. For example, the estimates of a flexible non-normal and non-linear model in Arellano et al. (2017, Figure 3) are quite similar to the PAE estimates in Figure (ref) for permanent components. At the same time, their estimates of the distribution of transitory components show substantially more pronounced non-Gaussianity and excess kurtosis relative to PAE. This finding is in agreement with our posterior informativeness measure R$^2$, which is 12% on average along the distribution for the permanent component, and 8% on average for the transitory component. This degree of informativeness suggests that posterior estimates may suffer from substantial specification error when the reference distribution is misspecified.

Overall, these empirical illustrations give two examples where, starting from a normal prior, the posterior conditioning is informative about the true unknown distributions. In both settings, PAE are not normal. Yet, as indicated by the $R^2$ values we report, the signal-to-noise ratios are not high enough to be certain about the exact shapes of the distributions of interest, thus motivating further analyses using non-normal specifications. PAE should be useful in other environments where model ((ref)) and its extensions are widely used, for example in teacher value-added applications, where the signal-to-noise ratio is driven by the number of observations per teacher. Moreover, PAE are also applicable to other --- nonlinear --- econometric models, as we describe in the next section.

Complements and extensions

In this section, we outline several complements and extensions that we analyze in detail in the appendix.

PAE in other models

PAE are applicable to a variety of settings. In many econometric models, semi-parametric estimators --- i.e., robust to distributional assumptions on unobservables --- of ${\Greekmath 010C}$ parameters are available; see Powell (1994) for examples. In such models, PAE provide estimators of average effects that enjoy robustness properties when parametric assumptions are violated. In Appendix (ref) we study static binary and ordered choice models, censored regression models, and panel data binary choice models. We also show how the White (1980) formula for robust standard errors in linear regression can be interpreted as a PAE.

Confidence intervals and specification test

Under correct specification of the reference model, it is easy to derive the asymptotic distributions of $\widehat{{\Greekmath 010E}}^{\rm M}$ and $\widehat{{\Greekmath 010E}}^{\rm P}$ using standard arguments. Moreover, under local misspecification, confidence intervals that account for both model uncertainty and sampling uncertainty can be constructed following Armstrong and Koles\'ar (2018) and Bonhomme and Weidner (2018). However, such confidence intervals require the researcher to set a value for the degree of misspecification ${\Greekmath 010F}$. In Appendix (ref), we provide details on confidence intervals calculations. In addition, we explain how to construct a specification test of the reference model based on the difference $\widehat{{\Greekmath 010E}}^{\rm P}-\widehat{{\Greekmath 010E}}^{\rm M}$.

Robustness in prediction

In applications such as the fixed-effects model ((ref)) of teacher quality, researchers are often interested in {predicting} the quality ${\Greekmath 010B}_i$ of teacher $i$. Although our focus in this paper is on the estimation of population averages, it is interesting to see how different predictors perform under misspecification of the reference distribution. It is well-known that EB estimators minimize mean squared prediction error when the normal reference model is correctly specified. However, when normality fails, the best predictor is a different posterior mean, which does {not} generally coincide with the EB estimate based on a normal prior. Intuitively, conditioning on nonlinear functions of the data may improve prediction accuracy.

In Appendix (ref) we use our framework --- applied to worst-case mean squared prediction error instead of worst-case specification error of a sample average --- to provide results on the robustness of EB estimators in the presence of misspecification. We show that EB estimators have minimum worst-case mean squared prediction error, up to smaller-order terms, under local deviations from normality. In addition, we derive a fixed-${\Greekmath 010F}$, non-local risk bound in the spirit of Theorem (ref).

Conclusion

Posterior averages are commonly used to predict individual parameters, such as teacher quality or neighborhood effects, and they play a central role in Bayesian and empirical Bayes approaches. In this paper, we have provided a frequentist justification for posterior conditioning when the goal of the researcher is to estimate a population average quantity. We have shown that posterior average effects (PAE) have minimum worst-case specification error under various forms of misspecification of parametric assumptions. PAE are simple to implement, and our analysis provides a rationale for reporting them in applications alongside other parametric and semi-parametric estimators, as well as a simple way to assess the informativeness of the posterior conditioning. As an example, Arnold et al. (2020) recently reported PAE to document judge heterogeneity in the context of bail decisions. While we have used a linear fixed-effects model as a running example due to its popularity, there are other possible applications, some of which we discuss in the appendix.

\vskip 1cm

\paragraph{Acknowledgments.} We would like to thank two anonymous referees, Manuel Arellano, Tim Armstrong, Raj Chetty, Tim Christensen, Nathan Hendren, Peter Hull, Max Kasy, Derek Neal, Jesse Shapiro, Xiaoxia Shi, Danny Yagan, and audiences at various places for comments. Bonhomme acknowledges support from the NSF, Grant SES-1658920. Weidner acknowledges support from the Economic and Social Research Council through the ESRC Centre for Microdata Methods and Practice grant RES-589-28-0001 and from the European Research Council grants ERC-2014-CoG-646917-ROMIA and ERC-2018-CoG-819086-PANEDA.

thebibliography{99} \bibitem Abadie, A., and M. Kasy (2018): “The Risk of Machine Learning,” to appear in the Review of Economics and Statistics. \bibitem Andrews, I., M. Gentzkow, and J. M. Shapiro (2017): “Measuring the Sensitivity of Parameter Estimates to Estimation Moments,” Quarterly Journal of Economics, 132(4), 1553--1592. \bibitem Andrews, I., M. Gentzkow, and J. M. Shapiro (2020): “On the Informativeness of Descriptive Statistics for Structural Estimates,” Econometrica, 88(6), 2231--2258. \bibitem Angrist, J. D., P. D. Hull, P. A. Pathak, and C. R. Walters (2017): “Leveraging Lotteries for School Value-Added: Testing and Estimation,” Quarterly Journal of Economics, 132(2), 871--919. \bibitem Arellano, M., Blundell, R., and S. Bonhomme (2017): “Earnings and Consumption Dynamics: A Nonlinear Panel Data Framework,” Econometrica, 85(3), 693--734. \bibitem Arellano, M., and S. Bonhomme, S. (2009): “Robust Priors in Nonlinear Panel Data Models,” Econometrica, 77(2), 489--536. \bibitem Armstrong, T. B., and M. Koles\'ar (2018): “Sensitivity Analysis Using Approximate Moment Condition Models,” arXiv preprint arXiv:1808.07387. \bibitem Arnold, D., W. S. Dobbie, and P. Hull (2020): “Measuring Racial Discrimination in Bail Decisions,” (No. w26999). National Bureau of Economic Research. \bibitem Berger, J. (1980): \textit{Statistical Decision Theory: Foundations, Concepts, and Methods}. Springer. \bibitem Bickel, P. J., C. A. J. Klaassen, Y. Ritov, and J. A. Wellner (1993): \textit{Efficient and Adaptive Inference in Semiparametric Models.} Johns Hopkins University Press. \bibitem Blundell, R., L. Pistaferri, and I. Preston (2008): \textquotedblleft Consumption Inequality and Partial Insurance,\textquotedblright\ \textit{American Economic Review}, 98(5): 1887--1921. \bibitem Blundell, R., L. Pistaferri, and I. Saporta-Eksten (2016): “Consumption Smoothing and Family Labor Supply,” \textit{American Economic Review}, 106(2), 387--435. \bibitem { Bonhomme, S., and J. M. Robin (2010): \textquotedblleft Generalized Nonparametric Deconvolution with an Application to Earnings Dynamics,\textquotedblright\ \textit{Review of Economic Studies}, 77(2), 491--533. } \bibitem Bonhomme, S., and Weidner, M. (2018): “Minimizing sensitivity to model misspecification,” arXiv preprint arXiv:1807.02161. \bibitem Carlton, D. W., and R. E. Hall (1978): “The Distribution of Permanent Income,” in \textit{Income Distribution and Economic Inequality}. New York: Halsted. \bibitem Chetty, R., Friedman, J. N., and Rockoff, J. E. (2014): “Measuring the impacts of teachers I: Evaluating bias in teacher value-added estimates,” \textit{American Economic Review}, 104(9), 2593-2632. \bibitem Chetty, R., and N. Hendren (2018): “The Impacts of Neighborhoods on Intergenerational Mobility: County-Level Estimates,” \textit{Quarterly Journal of Economics}, 133(2), 1163-1228. \bibitem Christensen, T., and B. Connault (2019): “Counterfactual Sensitivity and Robustness,” unpublished manuscript. \bibitem Cressie, N., and T. R. C. Read (1984): “Multinomial Goodness-of-Fit Tests,” \textit{Journal of the Royal Statistical Society Series B}, 46(3), 440--464. \bibitem Delaigle, A., P. Hall, and A. Meister (2008): “On Deconvolution with Repeated Measurements,” \textit{Annals of Statistics}, 36(2), 665--685. \bibitem Dobbie, W., and R. G. Fryer Jr (2013): “Getting Beneath the Veil of Effective Schools: Evidence from New York City,” \textit{American Economic Journal: Applied Economics}, 5(4), 28--60. \bibitem Efron, B. (2012): \textit{Large-Scale Inference: Empirical Bayes Methods for Estimation, Testing, and Prediction.} Vol. 1. Cambridge University Press. \bibitem Efron, B., and C. Morris (1973): “Stein's Estimation Rule and its Competitors -- An Empirical Bayes Approach,” \textit{Journal of the American Statistical Association}, 68(341), 117-130. \bibitem Fessler, P., and M. Kasy (2018): “How to Use Economic Theory to Improve Estimators,” to appear in the \textit{Review of Economics and Statistics}. \bibitem Finkelstein, A., M. Gentzkow, P. Hull, and H. Williams (2017): “Adjusting Risk Adjustment -- Accounting for Variation in Diagnostic Intensity,” \textit{New England Journal of Medicine}, 376, 608--610. \bibitem Geweke, J., and M. Keane (2000): “An Empirical Analysis of Earnings Dynamics Among Men in the PSID: 1968-1989,” \textit{Journal of Econometrics}, 96(2), 293--356. \bibitem Guvenen, F., F. Karahan, S. Ozcan, and J. Song (2016): “What Do Data on Millions of U.S. Workers Reveal about Life-Cycle Earnings Risk?” to appear in \textit{Econometrica}. \bibitem Hall, R., and F. Mishkin (1982): \textquotedblleft The sensitivity of Consumption to Transitory Income: Estimates from Panel Data of Households,\textquotedblright\ \textit{Econometrica}, 50(2): 261--81. \bibitem Hansen, B. E. (2016): “Efficient Shrinkage in Parametric Models,” \textit{Journal of Econometrics}, 190(1), 115--132. \bibitem Hirano, K. (2002): “Semiparametric Bayesian Inference in Autoregressive Panel Data Models,” \textit{Econometrica}, 70(2), 781--799. \bibitem Huber, P. J., and E. M. Ronchetti (2009): \textit{Robust Statistics}. Second Edition. Wiley. \bibitem Hull, P. (2018): “Estimating Hospital Quality with Quasi-Experimental Data,” unpublished manuscript. \bibitem Ignatiadis, N., and S. Wager (2019): “Bias-Aware Confidence Intervals for Empirical Bayes Analysis,” arXiv preprint arXiv:1902.02774. \bibitem James, W., and C. Stein (1961): “Estimation with Quadratic Loss,” in \textit{Proc. Fourth Berkeley Symp. Math. Statist. Prob.}, 1, 361--379. Univ. of California Press. \bibitem Jochmans, K., and Weidner, M. (2018): “Inference on a distribution from noisy draws,” arXiv preprint arXiv:1803.04991. \bibitem Kane, T. J., and Staiger, D. O. (2008): “Estimating Teacher Impacts on Student Achievement: An Experimental Evaluation”, National Bureau of Economic Research (No. w14607). \bibitem Kitamura, Y., Otsu, T., and Evdokimov, K. (2013): “Robustness, infinitesimal neighborhoods, and moment restrictions”, \textit{Econometrica}, 81(3), 1185-1201. \bibitem Koenker, R., and I. Mizera (2014): “Convex Optimization, Shape Constraints, Compound Decisions, and Empirical Bayes Rules,” \textit{Journal of the American Statistical Association}, 109(506), 674--685. \bibitem Kotlarski, I. (1967): “On Characterizing the Gamma and the Normal Distribution,” \textit{Pacific Journal of Mathematics}, 20(1), 69--76. \bibitem Li, T., and Q. Vuong (1998): “Nonparametric Estimation of the Measurement Error Model Using Multiple Indicators,” \textit{Journal of Multivariate Analysis}, 65(2), 139--165. \bibitem Morris, C. N. (1983): “Parametric Empirical Bayes Inference: Theory and Applications,” \textit{Journal of the American Statistical Association}, 78(381), 47--55. \bibitem Owen, D. B. (1956): “Tables for Computing Bivariate Normal Probabilities,” \textit{The Annals of Mathematical Statistics}, 27(4), 1075--1090. \bibitem Powell, J. L. (1994): “Estimation of Semiparametric Models,” \textit{Handbook of Econometrics}, 4, 2443--2521. \bibitem Stefanski, L. A., and R. J. Carroll (1990): “Deconvolving Kernel Density Estimators,” \textit{Statistics}, 21(2), 169--184.

\baselineskip21pt