Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
110,318 characters · 21 sections · 91 citation commands
Quasi-Bayes in Latent Variable Models
A large class of models explicitly account for the influence of latent unobservables on observed economic behavior. These models are constructed to be informative at the population in that the distribution of observables uniquely identifies the latent distribution of interest. However, when the link between observables and latent heterogeneity is tenuous, the finite sample identifying strength can be very weak. As a consequence, classical nonparametric estimators often display several undesirable properties, such as large finite-sample variability, irregularity and extreme sensitivity to small perturbations of the data and user-selected tuning parameters.
In empirical settings, it is common practice to impose strong distributional restrictions on the unobservables. Intuitively, these restrictions limit the model complexity, and thereby require weaker information content from the data. These approaches are attractive in that they significantly reduce finite sample uncertainty and provide a researcher with a tractable distribution (e.g. Gaussian) to use for counterfactual analysis. On the downside, they often introduce significant misspecification bias if the restrictions are incorrect. This bias is particularly evident in the descriptive evidence and estimates of non-Gaussian models (bonhomme2010generalized; \citealp*{guvenen2014nature}; \citealp*{arellano2017earnings}).
Motivated by these concerns, this paper proposes a class of nonparametric estimators that arise as solutions to quasi-Bayes decision rules. Traditional quasi-Bayes (chernozhukov2003mcmc) combines a GMM objective function $Q_n(\theta)$ and a prior $\nu(\theta)$ for a finite dimensional parameter $\theta \in \mathbb{R}^k$ to create a quasi-posterior distribution:
In this paper, we generalize the representation in ((ref)) to settings where the parameter is an infinite dimensional latent distribution of interest. We maintain the classical quasi-Bayes interpretation in that $Q_n(.)$ only depends on the observables through a collection of identifying population moments. In this setup, $Q_n(.)$ can be viewed as a generalized GMM objective and $\nu(.)$ is a nonparametric prior on distributions. As is standard, the effect of the prior is negligible in settings with strong identification. However, with weakly informative data and a nearly flat objective function $Q_n(.)$, the prior provides significant regularization and allows researchers to incorporate additional information (e.g. moment and smoothness restrictions) that improve the finite-sample information content.
We propose a theoretically motivated class of priors that are supported on nonparametric Gaussian mixtures. Our focus on this class is partially motivated from the observation that finite Gaussian mixtures are widely used to model rich forms of heterogeneity. As Gaussian mixtures can be efficiently sampled from, they provide a researcher with a convenient framework to perform counterfactual analysis. One consequence of building a framework around this structure is that we also obtain decision rules with particularly desirable theoretical guarantees in settings where the true data generating process closely resembles a (possibly infinite) Gaussian mixture, while remaining viable in the broader fully nonparametric setting. As we illustrate in our analysis, this choice is also computationally attractive in that it greatly facilitates the interplay analysis between the identifying moments and the prior.
The main theoretical contributions of this paper are as follows. We provide a unified treatment of quasi-Bayes for a large class of nonparametric latent variable models. This includes all models in which characteristic function based moment restrictions are employed to identify the latent distribution. As we illustrate in Section (ref), this encompasses a large class of commonly used latent variable models in economics. In each model, we take these identifying restrictions as given and use them to build a moment-based likelihood. We then derive posterior contraction rates for the quasi-Bayes posterior in a variety of metrics (e.g. Wasserstein and $L^2$). As our work is the first quasi-Bayes approach to nonparametric latent heterogeneity, we expect the analysis to be of independent interest and useful towards a wide range of related extensions.
As an application, we use our methodology to model individual earnings data from the Panel Study of Income Dynamics (PSID), revisiting the seminal study in bonhomme2010generalized. Consistent with the existing literature (e.g. lilliard1978dynamic; arellano2017earnings), we find that the distribution of U.S. wage shocks is highly non-Gaussian and leptokurtic. Overall, we find that our approach accurately reproduces the higher-order moments observed in U.S. wage growth data. Finally, we use our framework to examine the influence of permanent and transitory shocks on wage mobility.
The paper is organized as follows. Section (ref) introduces the quasi-Bayes framework, provides a detailed description of our procedures and outlines the class of models to which our methodology applies. Section (ref) provides the main assumptions and states the main results. In Section (ref), we apply our methodology to study earnings dynamics in U.S wages, using data from the PSID. Section (ref) verifies the main assumptions for several latent variable models and further develops the main results in a model specific context. Section (ref) provides simulation evidence on quasi-Bayes estimators relative to common alternatives. Section (ref) concludes.
Latent variable models have a long history in econometrics and statistics. In Section (ref), we provide an overview of several commonly used models in the literature. For a more comprehensive overview, see \citet*{chen2011nonlinear}; schennach2022measurement.
chernozhukov2003mcmc developed the foundational quasi-Bayes limit theory for parametric models that are strongly identified through a finite set of moments. In this article, we focus on models with latent heterogeneity. Importantly, in such models, the parameter is infinite dimensional, and its recovery is statistically ill-posed. Recently, quasi-Bayes procedures have also been applied in parametric models with weak or non-standard identification (\citealp*{chen2018monte}; andrews2022optimal). As the finite sample information content in latent variable models is usually limited, we view our work as complementary to this line of research and particularly informative for the potential of quasi-Bayes decision rules in nonparametric moment models.
Our work contributes to a growing literature on Bayesian inference in density estimation and deconvolution. For Bayesian density estimation, see ghosal2001entropies, ghosal2007posterior, shen2013adaptive, among many others. For Bayesian estimation in deconvolution with a known error distribution, see donnet2018posterior; rousseau2024wasserstein. To the best of our knowledge, our paper is the first to develop a quasi-Bayes framework to estimate a density and latent distribution. Within this framework, many of our results also represent the first nonparametric Bayes guarantees for the models under consideration.
The set of Borel probability measures on $\mathbb{R}^d$ is denoted by $\mathcal{B}(\mathbb{R}^d)$. Convolution of a function $g: \mathbb{R}^d \rightarrow \mathbb{R}$ with a measure $\mu \in \mathcal{B}(\mathbb{R}^d) $ is denoted by $g \star \mu (x) = \int_{\mathbb{R}^d} g(x-z) d \mu (z)$. Given two measures $(\lambda,\mu) $, we denote the product measure by $\nu = \lambda \otimes \mu$. Let $ \mathbf{H}^{p} = ( \mathbf{H}^p, \| . \|_{\mathbf{H}^p}) $ denote the usual $p$-Sobolev space of functions on $\mathbb{R}^d$.
Let $Y \in \mathbb{R}^q$ denote a vector of observables. We are interested in the distribution of a latent unobservable $X \in \mathbb{R}^d$ that influences $Y$ through a known economic model. In many natural settings, identification of the latent distribution is established by verifying that the latent characteristic moments $\varphi_{X}(t) = \mathbb{E}[e^{\mathbf{i}t'X}]$, or its derivatives, can be identified through suitable population moments of $Y$. To fix ideas, suppose $\mathcal{G}(.) $ is a known function valued operator. Then, the latent distribution is assumed to satisfy a restriction of the form
The argument of $\mathcal{G}(.)$ consists of two inputs, a probability distribution $P$ on $Y$ and a characteristic function $\varphi_F$. In most cases, we can view ((ref)) as a generalized moment restriction in that the operator $\mathcal{G}(.)$ only depends on $P$ through a collection of moments on $Y$. We say the restriction is an identifying restriction if
In this paper, we focus on a general class of latent variable models for which a generalized moment restriction of the form in ((ref)) is known to hold. As the following examples illustrate, this encompasses a large class of commonly used models in economics. Furthermore, the map $ \varphi_F \rightarrow \mathcal{G}(\mathbb{P}_Y,\varphi_F) $ is usually an elementary transformation of $\varphi_F$ or its derivatives.
The preceding examples and their extensions cover a broad class of latent variable models commonly used in economics, though the list is far from exhaustive. Similar characteristic moment restrictions arise in a variety of other settings, including models with mismeasured regressors (schennach2004estimation,schennach2007instrumental; \citealp*{ben2017identification}), nonparametric panel data arellano2012identifying,wilhelm2015identification, and workhorse models for consumption dynamics arellano2017earnings,arellano2024heterogeneity.
In nearly all the settings described above, the standard approach to estimating the latent distribution involves inverting the identifying restrictions. This typically entails solving a finite-sample analog of \(\mathcal{G}(\mathbb{P}_Y, \varphi_X) = \mathbf{0}\) to obtain an estimate \(\widehat{\varphi}_X\). The estimated latent density \(\widehat{f}_X\) is then constructed using an inverse Fourier transform.\footnote{For applications of this approach to specific models, see horowitz1996semiparametric; bonhomme2010generalized; krasnokutskaya2011identification; gautier2013nonparametric; masten2018random; herbst2024opportunity, among many others.} As noted in the literature (e.g. \citealp*{efron2016empirical}; \citealp*{arellano2023recovering}), these solutions often exhibit significant finite sample variability and are highly sensitive to small perturbations in the data and user-selected tuning parameters. Intuitively, the large finite sample variability of these estimators arises from their representation as an inverse of an ill-posed objective function over an excessively large parameter space.\footnote{These estimators allow for any square integrable function- the effective parameter space is $L^2(\mathbb{R}^d)$. As a consequence, the estimator $\widehat{f}_X$ is often negative on a subset of the domain and $\int \widehat{f}_X(t) dt \neq 1$.}
Two natural questions arise from the preceding discussion: $(i)$ Can the information content in the moments be utilized more efficiently for the class of nonparametric latent variable models described above? $(ii)$ Is the identifying strength in the moments significantly amplified in settings where some regularity is imposed on the parameter space? To address these questions, this paper proposes and examines a class of nonparametric estimators that arise as solutions to quasi-Bayesian decision rules.
In parametric models, where a finite-dimensional parameter $\theta$ is identified through finite moment restrictions, the classical quasi-Bayes approach (e.g., chernozhukov2003mcmc) treats a monotonic transformation of the GMM objective function as a likelihood for $\theta$. In our setting, we view the latent distribution $F_X$ as an infinite-dimensional parameter of interest. The identifying restriction $ \mathcal{G}(\mathbb{P}_Y,\varphi_{\mathbf{X}}) = \mathbf{0} $ then motivates a quasi-Bayes likelihood of the form
While ideal, this choice is infeasible as the true distribution $\mathbb{P}_{Y}$ is unknown and must be estimated from the data. In this paper, we proceed by replacing $\mathbb{P}_Y$ with the empirical measure $\mathbb{P}_{n,Y} = n^{-1} \sum_{i=1}^n \delta_{Y_i}$. As the preceding examples illustrate, the explicit dependence of $\mathcal{G}(\mathbb{P}_Y, \varphi_F)$ on $\mathbb{P}_Y$ is usually through an elementary transformation of $\varphi_{Y}(t) = \mathbb{E}[e^{\mathbf{i}t' Y}]$. In particular, the quantity $\mathcal{G}(\mathbb{P}_{n,Y}, \varphi_F)(t)$ will typically depend on $\widehat{\varphi}_{Y}(t) = \mathbb{E}_n[ e^{\mathbf{i}t'Y}]$ and as a consequence, it cannot be consistently estimated uniformly over $t \in \mathbb{R}^q$.\footnote{Under standard regularity conditions, we have $ \sup_{\| t \|_{\infty} \leq T} \left| \widehat{\varphi}_Y(t) - \varphi_Y(t) \right| = O_{\mathbb{P}}( n^{-1/2} \sqrt{\log T} )$. This is Lemma 1 in the Appendix. } To control the finite sample uncertainty, we follow the usual approach and consider a bounded set of restrictions that grows with the sample size $n$. Specifically, given any fixed $T > 0$, we denote the ball of radius $T$ and the $L^2$ magnitude (norm) of a function $f : \mathbb{R}^q \rightarrow \mathbb{C}$ over the ball by
Denote the observed data by $\mathcal{D}_n = \{ Y_1,\dots,Y_n \}$. Given a deterministic, or possibly data driven, sequence of constants $T_n \uparrow \infty$ that grows with the sample size $n$, we define the characteristic likelihood by
When combined with a (possibly data dependent) prior $\nu $ that is supported on a set of distributions $\mathcal{F}$, we obtain the quasi-Bayes posterior:
In principle, many distinct families of priors are suitable to construct the quasi-Bayes posterior in ((ref)). In Section (ref), we provide general results that are applicable to a wide class of priors. However, in the interest of sampling efficiently from the quasi-posterior and using the samples in a tractable form, it is convenient to focus on priors that satisfy two basic properties. First, they have a known and tractable characteristic function. This ensures that the quasi-likelihood in ((ref)) and suitable derivatives of it can be easily evaluated. Second, the prior is supported on distributions that are tractable to sample from. This property is convenient for any post estimation analysis, especially if the overall goal is to estimate complex functionals of the latent distribution. In the remainder of this section, we describe a class of priors that satisfy both of these requirements.
Given a covariance matrix $\Sigma \in \mathbb{R}^{d \times d}$, denote by $\phi_{\Sigma}$ the $N(\mathbf{0},\Sigma)$ density on $\mathbb{R}^d$. Given a probability distribution $P \in \mathcal{B}(\mathbb{R}^d)$, we denote the Gaussian mixture with mixing distribution $P$ and covariance matrix $\Sigma $ by
If $P$ is a discrete distribution, i.e $P = \sum_{j=1}^{\infty} p_j \delta_{\mu_j} $ which assigns probability mass $p_j$ to a point $\mu_j \in \mathbb{R}^d$, the preceding definition reduces to
Let $\varphi_F(t) = \int e^{\mathbf{i}t' x} d F(x)$ denotes the characteristic function (CF) of a distribution $F$. If $F = \phi_{P,\Sigma}$, we denote it by $\varphi_{P,\Sigma}$. By elementary properties of the CF, we have
For a discrete distribution $P = \sum_{j=1}^{\infty} p_j \delta_{\mu_j} $, the preceding expression simplifies to
A Gaussian mixture is completely characterized by its mixing distribution $P$ and covariance matrix $\Sigma$. In particular, a prior on $(P,\Sigma)$ is equivalent to placing a prior on infinite Gaussian mixtures. For a prior on $\Sigma$, a common choice is the Inverse-Wishart distribution. In Section (ref), we show that a large class of covariance priors can be accommodated for the models considered in this paper. For the mixing distribution $P$, we focus on the canonical choice of a prior for a distribution, the Dirichlet process.
An alternative definition, particularly useful for simulation, is given below.
A Dirichlet process prior can be used to build a prior on nonparametric Gaussian mixtures. Specifically, given a Dirichlet process prior $\text{DP}_{\alpha}$ and an independent prior $G$ on the set of positive definite matrices, the induced DP-GM prior on Gaussian mixtures is
An advantage of focusing on nonparametric Gaussian mixture priors is that they greatly facilitate the interplay analysis between the prior and the moments. As shown in ((ref)), Gaussian mixtures admit a simple characteristic function, which in turn leads to a tractable quasi-likelihood. Importantly, priors on Gaussian mixtures translate seamlessly to priors on characteristic moments and vice versa.
An informative prior may set $\alpha$ to be a reasonable guess for the latent distribution, for instance, via a preliminary estimation based on a parametric model. A theoretical discussion of this is provided in Remark (ref). In the remaining sections, we will frequently refer to the prior in ((ref)), which is also the prior used in our application and simulations. For notational convenience, we denote this product prior by $\nu_{\alpha,G} = \text{DP}_{\alpha} \otimes G$.
We begin by outlining the main regularity conditions in Section (ref). Section (ref) presents and discusses the main results for the general quasi-Bayes posterior in ((ref)). As our framework accommodates a wide range of models, we state the conditions in a suitably general form. In Section (ref), we verify these assumptions for the models in Section (ref) using low-level conditions and discuss their main implications in model specific detail.
We state and discuss the assumptions that we impose on the model and prior. Throughout this section, let $\nu_n$ denote a, possibly data dependent, prior that is supported on a set of distributions $\mathcal{F}_n \subseteq \mathcal{B}(\mathbb{R}^d)$. Let $T_n \uparrow \infty$ denote a deterministic sequence of positive constants. Let $(\epsilon_n)_{n=1}^{\infty} $ denote a deterministic sequence of positive constants that converge to zero at a slower than parametric rate : $\epsilon_n \downarrow 0 $ and $n \epsilon_n^2 \uparrow \infty$.
Assumption (ref) provides bounds on the sampling uncertainty that arises from the true population distribution $\mathbb{P}_{Y}$ being unknown. Typically, $\mathcal{S}_n$ represents a ball (in an suitable metric) centered around a fixed distribution $F_n$. In certain settings, such as Examples (ref) and (ref), $\mathcal{S}_n$ may also include additional constraints (e.g. moment bounds) on $F$, which help in controling the sampling uncertainty.
The set \( \mathcal{R}_n \) in Assumption (ref) is introduced to provide some flexibility when direct verification of a local concentration bound is challenging for the $\mathcal{S}_n$ in Assumption (ref). In such cases, $\mathcal{R}_n$ relaxes certain restrictions (e.g. moment bounds) imposed on \( \mathcal{S}_n \). Assumption (ref)\((ii)\) further requires that the subset of \( \mathcal{R}_n \) where these restrictions fail to hold is sufficiently negligible. Typically, \( \mathcal{R}_n \) is a small ball (in a suitable metric) around $F_n$. Assumption (ref)\((i)\) then imposes a standard small ball local concentration condition on the prior.
In this section, we verify that the quasi-Bayes posterior in ((ref)) asymptotically concentrates on local neighborhoods of the latent distribution.
Theorem (ref) establishes contraction with the respect to the weak metric $$d_{\mathcal{G}} (F,F_X) = \| \mathcal{G}(\mathbb{P}_{n,Y}, \varphi_{F}) - \mathcal{G}(\mathbb{P}_{Y}, \varphi_{X}) \|_{\mathbb{B}(T_n)} .$$The interpretation of this convergence varies across models. Typically, $\epsilon_n \downarrow 0$ at a fast rate, and the contraction can be used to deduce fast rates for a class of well-behaved functionals of the latent distribution (see Remark (ref) and the discussion in Section (ref)). More generally, it can be interpreted as a preliminary contraction for a suitable “direct problem”, which can then be leveraged to obtain convergence in a stronger distance. In particular, if ((ref)) holds and the bulk of the posterior mass is contained in a well-behaved subset, it is often possible to deduce results in a stronger metric like $ d_2(F,F_X) = \| f- f_X \|_{L^2} $. To fix ideas, given a distance metric $d(.)$ and a class of distributions $\mathcal{H}_n$, we define the modulus of continuity by
The modulus of continuity is frequently used to characterize the convergence rate in inverse problems (e.g. chen2012estimation). The following result is a straightforward consequence of Theorem (ref).
The constant $D'$, which regulates the decay of mass on $\mathcal{H}_n^c$, is required to be larger than some of the preceding constants that appear in Assumption (ref) - (ref). Usually, the set $\mathcal{H}_n$ is chosen as a function of $D'$ so as to ensure the desired bound holds trivially.
Corollary (ref) provides contraction rates in terms of the modulus $\omega_n(d,\mathcal{H}_n,L \epsilon_n)$. Intuitively, as $\omega_n \rightarrow 0$, the posterior concentrates on local neighborhoods of the latent distribution. The posterior mean naturally serves as an estimator, while the variability in the samples provides a measure of uncertainty.
In Section (ref), we verify the main assumptions for several latent variable models using the DP-GM prior introduced in Section (ref), and we further develop the main results in a model-specific context. In particular, we establish upper bounds on the modulus \( \omega_n(\cdot) \) and derive posterior contraction rates under a broad class of metrics, including the Wasserstein and \( L^2 \) distances. As we demonstrate in Section (ref), many of our results provide the first nonparametric Bayes guarantees for the models under consideration.
In this section, we apply our methodology to study the latent structure of permanent and transitory components in the Panel Study of Income Dynamics (PSID) data. Our analysis closely follows the seminal work in bonhomme2010generalized, in which deconvolution-based methods were used to nonparametrically estimate the latent structure.
Our objective in this section is to demonstrate the viability of quasi-Bayes procedures in a non-trivial empirical setting. The dataset consists of a panel of 624 individuals observed over 10 time periods. The model contains 17 latent distributions. Given the relatively small sample size and the presence of several nonparametric latent distributions, it is inevitable that results will exhibit some dependence on tuning parameters. Indeed, as the empirical analysis in bonhomme2010generalized; arellano2017earnings illustrate, even with carefully chosen tuning parameters, obtaining reasonable estimates through direct frequentist estimation is a very challenging task.
We begin with a brief summary of the data and model. The data is from the PSID, between 1978 and 1987. Let $y_{i,t}$ denotes annual log earnings for individual $i$ at time period $t$ and $x_{i,t}$ an associated set of regressors. The regressors are a quadratic polynomial in age and indicators for education, race, geography and year. The OLS residuals of $ y_{i,t} $ on $ x_{i,t} $ are denoted by $ w_{i,t}$. The residual differences are denoted by $\Delta w_{i,t} = w_{i,t} - w_{i,t-1}$. After restricting the sample to male workers with no missing observations for $ \Delta w_{i,t}$ and a wage growth that does not exceed $150 \%$ in absolute value, our sample size consists of $n=624$ individuals. For each individual, we observe wages between 1978 and 1987, for a total of $M=10$ time periods. The standard model is given by
Here, $f_i$ is an individual-level fixed effect, and $\{\epsilon_{i,t} \}_{t=1}^M$, $\{\eta_{i,t} \}_{t=2}^{M-1}$ are mean-zero latent factors. The model in ((ref)) and various generalizations of it are widely used to study income dynamics (hall1982sensitivity; abowd1989covariance). We think of $w_{i,t}^P$ as the permanent component and $\eta_{i,t}$ as the transitory component. Thus, the distribution of permanent income is determined by $\epsilon_{i,t}$ and the distribution of transitory income by $\eta_{i,t}$. From first differencing the model in ((ref)), we can write
We view $ \mathbf{Y}_i = (\Delta w_{i,2}, \dots , \Delta w_{i,M} ) $ as the observations for $i=1,\dots,n$. The model in ((ref)) is a special case of the multi-factor model $ \mathbf{Y}_i = \mathbf{A} \mathbf{X}_i $ in Section (ref). In this setting, the latent factors are $$ \mathbf{X}_i = \{ \eta_{i,2}, \dots , \eta_{i,M-1} , \epsilon_{i,2} , \dots , \epsilon_{i,M} \}.$$ As noted in the literature (e.g. geweke2000empirical; bonhomme2010generalized), assuming a Gaussian model for the latent factors implies a distribution for the wage growth $\Delta w$ that is inconsistent with the observed higher-order moments (e.g. kurtosis) in the data. This raises the concern that some components of the wage shocks $\mathbf{X}_i$ may be non-Gaussian. Over $M=10$ time periods, there are $17$ (the dimension of $\mathbf{X}_i$) latent distributions that must be estimated. All implementation details are provided in the Appendix.
Figure (ref) illustrates the posterior distribution of the latent variance structure, with the solid black dashed line indicating the posterior mean. The posterior mean estimates of the variance are $\widehat{\sigma}_{\eta}^2 = 0.014$ and $\widehat{\sigma}_{\epsilon}^2 = 0.015$. The corresponding posterior means of the kurtosis are $\widehat{\kappa}_{\eta} = 20.46$ and $\widehat{\kappa}_{\epsilon} = 12.19$. Figure (ref) plots the posterior mean of the average standardized latent density, with a standard Gaussian density overlaid for comparison. These results indicate that the latent factors are non-Gaussian, symmetric, and leptokurtic.
Noticeably, the transitory component shows a larger degree of kurtosis and departure from standard Gaussianity. These distributional observations are consistent with previous studies of U.S earning growth data (e.g. \citealp*{arellano2017earnings}; \citealp*{guvenen2021data}). Furthermore, we note that quasi-Bayes estimates do not contain the extreme tail oscillations that frequently arise in deconvolution-based estimators (e.g. horowitz1996semiparametric; bonhomme2010generalized). This suggests that such oscillations originate from the choice of estimator, rather than from the information content in the moments.
Table (ref) compares the observed moments of the wage growth residuals, $\Delta w$, with the model implied moments. Here, $\textbf{Normal}$ indicates a maximum likelihood estimate with normal latent factors. Since a normal distribution for the latent structure implies normality in the wage growth residuals, $\Delta w$, the model-implied kurtosis substantially underestimates the observed value. In contrast, quasi-Bayes does not match the variance exactly but achieves a significantly better fit for the higher moments. These higher moments are particularly important for counterfactuals that arise in consumption and earnings dynamics.
Since our posterior is supported on nonparametric Gaussian mixtures, posterior samples can be easily utilized for any downstream application. As an illustration, we examine the influence of permanent and transitory shocks on wage mobility. For each period $t$, this influence can be characterized by the functionals:
These functionals resemble some of the conditional means that arise in Gaussian measurement error models within the empirical Bayes literature (e.g. gu2023invidious). In our framework, they can be estimated in general multifactor models, including settings with non-Gaussian latent factors.
Our objective is to examine how permanent and transitory components contribute to wage mobility. For instance, are large wage changes driven predominantly by transitory or permanent shocks? Moreover, at what threshold does this change? Averaging over time, we define $\bar{\gamma}_{\epsilon}(s) = \frac{1}{M-1} \sum_{t=2}^M \gamma_{t,\epsilon}(s)$ and $\bar{\gamma}_{\eta}(s) = \frac{1}{M-1} \sum_{t=2}^M \gamma_{t,\eta}(s)$.
In Figure (ref), we see that the posterior distributions of $\bar{\gamma}_{\epsilon}$ and $\bar{\gamma}_{\eta}$ are tightly concentrated around their posterior means. The convergence of these functionals depends only on contraction in the weak metric (Theorem (ref)) and is therefore much faster than direct recovery of the latent factors.\footnote{For further discussion, see Remark (ref) and the discussion in Section (ref).} As a consequence, estimation of these functionals is more precise and posterior uncertainty is low. At $s = 1$, the posterior means are \[ \mathbb{E}\big[\bar{\gamma}_{\epsilon}(1) \,\big|\, \mathcal{D}_n\big] = 0.279 \quad\text{and}\quad \mathbb{E}\big[\bar{\gamma}_{\eta}(1) \,\big|\, \mathcal{D}_n\big] = 0.721 \] Thus, a log wage growth of \(+100\%\) is expected to be of transitory origin for \(72.1\%\) of its value. Figure (ref) presents the transitory--permanent decomposition for the entire wage mobility path, covering growth rates between \(\pm 100\%\). In particular, moderate to large wage changes are predominantly of transitory origin.
In this section, we verify the main assumptions outlined in Section (ref) and explore their implications for various latent variable models. In all cases, we focus on the DP-GM prior discussed in Section (ref). In this case, the prior is supported on nonparametric Gaussian mixtures through the mixture representation
This induces a prior $\nu$ on the space of probability distributions on $\mathbb{R}^d$. For ease of notation, we will reference the induced prior using the same notation as above, i.e $ \nu = \text{DP}_{\alpha} \otimes G$. Before proceeding further, we state and discuss the main assumptions that we impose on the Dirichlet process base measure $\alpha$ and covariance prior $G$.
Since the researcher chooses the prior, Assumptions (ref) and (ref) can always be satisfied. Restricting $\alpha$ to be a finite Gaussian mixture in Assumption (ref) facilitates the computation as it is necessary to sample from the base measure. It is not to be interpreted as restrictive, as any distribution can be approximated arbitrarily well by a Gaussian mixture with sufficiently many components. Assumption (ref) $(i-ii)$ allows for a large class of covariance priors and is commonly used in density estimation (shen2013adaptive; ghosal2017fundamentals). In dimension $d=1$, it holds with $\kappa=1/2$ if $G$ is the distribution of the square of an inverse-Gamma distribution. In dimension $d > 1$, it holds with $\kappa = 1$ if $G$ is the inverse-Wishart distribution.\footnote{In higher dimensions, a computationally attractive choice is the class of Lewandowski-Kurowicka-Joe (LKJ) priors lewandowski2009generating.} Assumption (ref) $(iii)$, for $v_3 > 0$, is equivalent to Assumption 4.2 in rousseau2024wasserstein. In $d=1$, it holds with a truncated inverse-Gamma prior.
In some cases, researchers may have prior estimates of the latent distribution from a preliminary parametric model, or certain features (e.g. moments) may be estimable using the observables. Similar to the analysis in mcauliffe2006nonparametric, this can be accomodated through a data-dependent (empirical Bayes) $\text{DP}_{\alpha}$ prior. For example, in measurement error models (Example (ref)), a reasonable choice is $\widehat{\alpha} = \mathcal{N}(\mathbb{E}_n(Y), \widehat{\text{Var}}(X))$. The following remark provides some clarification on the effect of data-dependent priors.
In the remainder of this section, we focus on specific results for the latent variable models introduced in Section (ref). For each model, we examine the following questions in detail: (i) What are the minimal conditions on $F_X$ that ensure quasi-Bayes consistency? (ii) How do convergence rates depend on the smoothness of $F_X$? (iii) Are rates faster for certain classes of distributions (e.g. infinite Gaussian mixtures), and what is the interplay with the model ill-posedness?
Consider the classical measurement error model, introduced in Example (ref), where we observe a random sample from
In this setting, $W \in \mathbb{R}^d $ is the observed vector, $X \in \mathbb{R}^d$ is an unobserved vector, and $\epsilon$ is a nuisance error that is independent of $X$. We are interested in recovering the latent distribution of $X$. As the individual contributions of $X$ and $\epsilon$ cannot be separately identified from an observation of $W$, identification of the latent distribution of $X$ generally requires some further auxiliary information on the distribution $F_{\epsilon}$ of the unobserved error $\epsilon$. As in kato2018uniform; arellano2023recovering, we consider the setting where a researcher has auxiliary information in the form of a random sample ${\epsilon}_1^{\star}, \dots , {\epsilon}_n^{\star} \stackrel{i.i.d}{\sim} F_{\epsilon} $.\footnote{If the auxiliary sample has size $m_n = o(n)$, the results remain valid, albeit with a possibly slower rate of convergence. There are no restrictions on the dependence of the auxiliary sample with the original sample.} As a basic premise, we impose the following low level condition on the observable $Y = (W , \epsilon^{\star})$.\begin{condition2}[Data] $(i)$ $(W_i)_{i=1}^n $ is a sequence of independent and identically distributed random vectors. $(ii)$ $(\epsilon_i^{\star})_{i=1}^n \stackrel{i.i.d}{\sim} F_{\epsilon} $ where $F_{\epsilon}$ is the distribution of the unknown error $\epsilon$ in ((ref)). $(iii)$ $(Y,\epsilon)$ have finite second moments. \end{condition2} We now proceed with stating the main results for the model in detail. From the discussion in Example (ref), the population characteristic restriction is given by
With $Y = (W , \epsilon^{\star})$ as the observable, the feasible analog $\mathcal{G}(\mathbb{P}_{n,Y} ,\varphi_X)$, obtained by using the empirical measure $\mathbb{P}_{n,Y}$, is given by
Our first result examines the limiting behavior of quasi-Bayes under minimal distributional assumptions on the data. Specifically, we avoid imposing any strong distributional structure on $(X,\epsilon)$, instead relying solely on primitive moment and tail bounds. As these basic assumptions will not preclude distributions with irregular components (e.g. mass points), it is necessary to quantify the behavior through a metric that is well-defined for a broad class of probability measures. To that end, we make use of the Wasserstein metric:
Among other desirable properties, it is well known (see, e.g., villani2009optimal) that convergence in $\mathbf{W}_1(.)$ implies convergence in distribution: $\mathbf{W}_1(F_n, F_X) \to 0 \implies F_n \rightsquigarrow F_X.$ The following result demonstrates that the quasi-Bayes posterior contracts to the latent distribution in Wasserstein distance.
Theorem (ref) is very general in its scope, as it does not require any strong distributional assumptions on $(X,\epsilon)$ and relies only on an auxiliary sample of $\epsilon$. It remains valid even when $(X,\epsilon)$ contain discrete components. Intuitively, regardless of the underlying structure, the moment in ((ref)) are always valid identifying restrictions. This result appears to be the first in the literature to establish a deconvolution estimation guarantee under such minimal assumptions.
The exponential lower bound on $ \varsigma_T = \inf_{\|t\| \leq T} \left| \varphi_{\epsilon}(t) \right| $ allows for measurement errors whose characteristic functions decay rapidly. For instance, if $\epsilon$ follows a Gaussian distribution, the lower bound is attained exactly at $\zeta = 2$. Remarkably, the contraction result in Theorem (ref) holds for any sequence $T_n = o(n^{1/5})$, regardless of the distribution of $\epsilon$. This result extends to dimensions $d > 1$, although in such cases additional constraints on the growth rate of $T_n$ are required.
Finally, we note that the slow convergence rate in Theorem (ref) is typical in deconvolution problems (see, e.g., carroll1988optimal). Without further regularity conditions on $(F_X, F_{\epsilon})$, the rate in Theorem (ref) cannot be improved; it is minimax optimal \citep*{dedecker2013minimax}.
In the remainder of this section, we focus on establishing contraction rates in settings where the latent distribution or measurement error are known to be sufficiently regular, either through Sobolev type constraints or restrictions on the nonparametric functional form. An important example of such a setting is when the latent distribution $F_X$ is a nonparametric Gaussian mixture:
Here, $\Sigma_0 \in \mathbf{S}_+^d$ is a positive definite matrix and $F_0$ is a mixing distribution. If $F_0$ is a discrete distribution with finite support, this reduces to a Gaussian mixture with finitely many components. Our focus on this class is partially motivated by the observation that finite Gaussian mixtures are widely used to model complex heterogeneity in latent variable models (e.g. cunha2010estimating; attanasio2020human). Importantly, if the mixing distribution $F_0$ is continuously distributed or has unbounded support, the specification in ((ref)) allows for Gaussian mixtures with infinite components.
As the specification in ((ref)) lies between a parametric and fully nonparametric model, a natural question is whether fast rates of convergence can be expected in this setup and, in particular, how is this influenced by the model ill-posedness. Intuitively, as the prior concentrates on infinite Gaussian mixtures, the expectation is that the prior should exhibit negligible bias in approximating a latent density with Gaussian mixture structure as in ((ref)). In the next result, we formalize this intuition. To this end, our main assumption regarding the Gaussian mixture specification is as follows. \begin{condition2}[Nonparametric Mixture] $f_X = \phi_{\Sigma_0} \star F_0$ where $\Sigma_0 \in \mathbb{R}^{d \times d}$ is a positive definite matrix and $F_0$ is a mixing distribution that satisfies $ F_0 \big( t \in \mathbb{R}^d : \| t \| > z \big) \leq C \exp(- C' z ^{\chi}) $ for universal constants $C,C',\chi > 0$. \end{condition2} Condition ((ref)) imposes an exponential tail on the mixing distribution $F_0$. Importantly, this covers all Gaussian mixtures that have (possibly infinite) modes contained in a compact subset of $\mathbb{R}^d$. The following result demonstrates that, with an appropriately scaled covariance prior, the quasi-Bayes posterior contracts towards the true latent density at a polynomial (and sometimes nearly parametric) rate.
Under a mildly ill-posed regime, the quasi-Bayes posterior attains nearly (up to a logarithmic factors) parametric rates of convergence. If the model is severely ill-posed at a rate $\zeta$ that is lower than Gaussian type ill-posedness, nearly parametric rates in the form $n^{-1/2 + \epsilon}$ for any $\epsilon > 0$ can be obtained. In the extreme case where the model exhibits Gaussian ill-posedness, the model bias (which is determined by the decay rate of a Gaussian mixture characteristic function) has similar order to the variance. In this case, the rate is still polynomial but the exponent $V$ depends on second order factors such as the constant $B$ and the eigenvalues of the matrix $\Sigma_0$ in representation ((ref)). In a univariate framework with a kernel deconvolution based estimator, a similar finding is established in butucea2008sharp2.
As Gaussian mixtures can approximate a density arbitrarily well, it is straightforward to modify the preceding results to obtain consistency for a general density.\footnote{Here, consistency in a norm $\| . \|$, is defined as $ \nu_{} \big( \phi_{P,\Sigma} : \| f_X - \phi_{P,\Sigma} \|_{} > \epsilon \:\big| \: \mathcal{D}_n \big) = o_{\mathbb{P}}(1) \; \; \forall \; \epsilon > 0$} In the interest of comparison with existing results in the literature, we focus on the more difficult task of establishing a rate of convergence. To that end, consider the case where the true latent density $f_X$ is not an exact Gaussian mixture and thus, the prior model exhibits non-negligible bias. In this model, the Gaussian mixture bias is closely related to the quantity
Intuitively, the smoothness of the convolution $ f_X \star f_{\epsilon} $ is the sum of the smoothness of the individual functions. As such, we expect the Gaussian mixture bias to decay at a rate proportional to the combined Sobolev smoothness (or equivalently the characteristic function decay) of $f_X$ and $f_{\epsilon}$. To derive the limit theory, we consider the following popular characterization of the Gaussian mixture bias. \begin{condition2}[Gaussian Mixture Bias] $f_X \in \mathbf{H}^p$ for some $p > 1/2$. For all $\sigma > 0 $ sufficiently small and some $\chi > 0$, there exists $(i)$ a mixing distribution $F_{\sigma} = F_{X,\epsilon,\sigma}$ supported on the cube $ I_{\sigma} = [-C (\log \sigma^{-1})^{1/\chi},C (\log \sigma^{-1})^{1/\chi} ]^d$ and $(ii)$ a covariance matrix $\Sigma_{\sigma} = \Sigma_{X,\epsilon,\sigma} $ with $\lambda_{\min}(\Sigma_{\sigma}) \geq \sigma^2$ such that the Gaussian mixture $\phi_{F_{\sigma},\Sigma_{\sigma}}$ satisfies $\| f_X \star F_{\epsilon} - \phi_{F_{\sigma},\Sigma_{\sigma}} \star F_{\epsilon} \|_{L^2} \leq D \sigma^{p + \zeta} $ for some $\zeta > 0$, where $C,D < \infty$ are universal constants. \end{condition2} Variations of Condition (ref), usually with additional smoothness constraints on $\log f_X$, are commonly used in Bayesian density estimation and deconvolution. For an explicit construction of $(F_{\sigma},\Sigma_{\sigma})$ in Condition (ref), see shen2013adaptive; ghosal2017fundamentals for density estimation and rousseau2024wasserstein for deconvolution. We note that, at least in our quasi-Bayes setup, Condition (ref) can be weakened further. In particular, it suffices that the approximation holds with respect to the weaker Fourier norm $\| . \|_{\mathbb{B}(T_n)}$. Finally, we note that our procedures do not require knowledge of the optimal mixing distribution $F_{\sigma}$ in Condition (ref) for implementation. The existence is purely a proof device towards obtaining theoretical guarantees.
The following result derives the quasi-Bayes contraction rate when the Gaussian mixture bias is as in Condition (ref) and $ \varphi_{\epsilon}(.) $ decays at a mildly ill-posed rate of order $\zeta \geq 0$.
The obtained rates are minimax-optimal, up to a logarithmic factor. When $\zeta=0$, we recover the usual optimal rates for multivariate density estimation. In dimension $d=1$, the rates agree with the univariate Bayesian deconvolution rates (with known error distribution) obtained in donnet2018posterior.
Interestingly, when $d=1$ and $\zeta=0$, the $L^{\infty}$ quasi-Bayes contraction rates in Corollary (ref) coincide with the univariate density pure-Bayes rates obtained in gine2011rates. For $d>1$ or $\zeta > 0$, the contraction rates appear to be novel; in particular, we could not find any pure-Bayes results for these cases in the literature. For simplicity of exposition, all remaining results in Section (ref) are stated exclusively for the Wasserstein and $\| \cdot \|_{L^2}$ metrics. By analogous arguments to those used in Corollary (ref), these results can be extended to other metrics.
Consider the repeated measurements model, introduced in Example (ref), where we observe a random sample from the model:
In this setting, $Y_1,Y_2 \in \mathbb{R}$ are observed measurements, $X \in \mathbb{R}$ is unobserved, and $\epsilon_1,\epsilon_2$ are unobserved nuisance errors. We assume $\mathbb{E}[\epsilon_1|X,\epsilon_2] = 0$ and that $\epsilon_2 $ is independent of $X$. The primitive assumption that we impose on the observations is summarized in the following condition. \begin{condition2}[Data] $(i)$ $(Y_{1,i},Y_{2,i})_{i=1}^n $ is a sequence of independent and identically distributed random variables. $(ii)$ $Y_1$ and $Y_2$ have finite second moments. \end{condition2} From the discussion in Example (ref), the population characteristic restriction is
With $Y = (Y_1 , Y_2)$ as the observable, the feasible analog $\mathcal{G}(\mathbb{P}_{n,Y} ,\varphi_X)$, obtained by using the empirical measure $\mathbb{P}_{n,Y}$, is given by
Similar to the analysis in Section (ref), our first result examines the limiting behavior of quasi-Bayes under minimal distributional assumptions on the data.
The exponential lower bound on $ \inf_{\| t \|_{\infty} \leq T}\left|\varphi_{Y_2} (t)\right|$ is a very weak restriction. Analogous to the identification restrictions with measurement error, if the term $\varphi_{Y_2}(t) \partial_t \log \varphi_X(t)$ in the population objective function vanished, $\varphi_X$ would not be identifiable. The upper bound on $ \sup_{\|t \|_{\infty} \leq T}\left| \partial_t \log \varphi_X(t) \right|$ is very conservative—the quantity is typically uniformly bounded when $\varphi_X$ decays at a polynomial rate, and it grows at rate $T$ in the extreme case of a Gaussian decay. To the best of our knowledge, Theorem (ref) is the first Wasserstein convergence guarantee for the repeated measurements model.\footnote{We could not find any prior results, whether frequentist or Bayesian, in the literature.} Additionally, it is the first nonparametric Bayes contraction rate for this model.
Next, similar to the analysis in Section (ref), we focus on establishing further contraction rates in settings where the latent distribution is known to be sufficiently regular. Specifically, we consider the case when the latent distribution is a nonparametric Gaussian mixture:
Here, $\sigma_0 > 0$ and $F_0$ is a mixing distribution. If the mixing distribution $F_0$ is continuously distributed or has unbounded support, the specification in ((ref)) allows for Gaussian mixtures with infinite components. Our main assumption on the Gaussian mixture specification is as follows. \begin{condition2}[Nonparametric Mixture] $(i) $ $f_X = \phi_{\sigma_0^2} \star F_0$ for some $ \sigma_0 > 0$ and mixing distribution $F_0$ that satisfies $ F_0 \big( t \in \mathbb{R} : \left| t \right| > z \big) \leq C \exp(- C' z ^{\chi}) $ for some $\chi > 0$ and universal constants $C,C' > 0$. $(ii)$ $ \sup_{\| t \| \leq T} \left| \partial_t \log \varphi_{F_0} (t) \right| \leq C (1+T) $ for a universal constant $C> 0$. \end{condition2} Condition $(i)$ imposes an exponential tail on the mixing distribution $F_0$. Importantly, this covers all Gaussian mixtures that have (possibly infinite) modes contained in a compact subset of $\mathbb{R}$. Condition $(ii)$ is similar to the hypothesis of Theorem (ref). The following result demonstrates that, with an appropriately scaled covariance prior, the quasi-Bayes posterior contracts towards the true latent density at a polynomial rate.
The main difference in this setting, compared to Theorem (ref), is that a nonparametric Gaussian mixture specification on $F_X$ influences both the regularity and ill-posedness in the model. To see this, note that, up to sampling uncertainty, the quasi-Bayes objective function is $$ Q(F) = \| \varphi_{Y_2} ( \partial_t \log \varphi_X - \partial_t \log \varphi_F) \|_{\mathbb{B}(T)}. $$ Thus, the model ill-posedness is driven by the decay of $\varphi_{Y_2}$. Since $\varphi_{Y_2} = \varphi_{X} \times \varphi_{\epsilon_2}$, a rapid decay of $\varphi_X$ also influences $\varphi_{Y_2}$. The contraction rate in Theorem (ref) is obtained through examining the interplay and dependence between $\varphi_{Y_2}$ and $\varphi_{X}$. As expected, similar to the discussion following Theorem (ref), the exponent $V$ depends on second order factors, e.g. the constants $(\sigma_0^2, c')$.
Following bonhomme2010generalized, consider the multi-factor model, where we observe a random sample from the model
In this setting, $\mathbf{Y} = (Y_1,\dots,Y_L)' \in \mathbb{R}^L$ is a vector of $L$ measurements, $\mathbf{X} = (X_1,\dots,X_K)'$ is a vector of $K$ latent and mutually independent factors and $\mathbf{A}$ is a known $L \times K$ matrix of factor loadings. The primitive assumption that we impose on the observations is summarized in the following condition. \begin{condition2}[Data] $(i)$ $( \mathbf{Y}_i )_{i=1}^n $ is a sequence of independent and identically distributed random vectors. $(ii)$ Finite second moment: $\mathbb{E}\big( \| \mathbf{Y} \|^2 \big) < \infty $. $(iii)$ The factors $X_1,\dots,X_K$ are mutually independent and demeaned, i.e $\mathbb{E}[X_k] = 0$ for $k=1,\dots,K$. \end{condition2} Suppose we are interested in the latent distribution $F_{X_k}$ for some $k \in \{ 1 , \dots , K \}$. Let $\mathcal{V}_{ \mathbf{Y} } (t)$ denote the vector of upper triangular elements of $\nabla \nabla' \log \varphi_{\mathbf{Y}} (t) $. Let $V(\mathbf{A}_k)$ denote the vector of upper triangular elements of $\mathbf{A}_k \mathbf{A}_k'$ and define $\mathbf{Q} = [ V(\mathbf{A}_1) , \dots , V(\mathbf{A}_K) ] $ to be the matrix with columns given by those vectors. From the discussion in Example (ref), if $\mathbf{Q}^* = (\mathbf{Q}'\mathbf{Q})^{-1} \mathbf{Q}'$ and $\mathbf{Q}_k^*$ denotes the $k^{th}$ row of $\mathbf{Q}^*$, the population characteristic restriction is given by
With $\mathbf{Y}$ as the observable, the feasible analog $\mathcal{G}(\mathbb{P}_{n,Y} ,\varphi_X)$, obtained by using the empirical measure $\mathbb{P}_{n,Y}$, is given by
where $\widehat{\varphi}_{\mathbf{Y}} = \mathbb{E}_n[e^{\mathbf{i} t' \mathbf{Y}}] $ and $\widehat{\mathcal{V}}_{\mathbf{Y}}(t)$ is the vector of upper triangular elements of $\nabla \nabla ' \log \widehat{\varphi}_{\mathbf{Y}}(t)$.
In this model, the identifying restriction $\mathcal{G}(\mathbb{P}_Y, \varphi_{X_k}) = \mathbf{0}$ only identifies the latent distribution up to an unknown mean. Intuitively, this is because the quantity $(\varphi_{X_k})''(t)$ does not depend on $\mathbb{E}[X_k]$. Thus, the assumption that $\mathbb{E}[X_k] = 0$ is a normalization that pins down the distribution uniquely. To account for this mean-zero restriction in our setup, we simply demean samples from the usual quasi-Bayes posterior. Specifically, we consider
Similar to the analysis in Section (ref), our first result examines the limiting behavior of quasi-Bayes under minimal distributional assumptions on the data.
The exponential lower bound on \( \inf_{\| t \|_{\infty} \leq T}\left|\varphi_{\mathbf{Y}} (t)\right| \) is a very weak restriction and is similar to the condition imposed in Theorem (ref). The upper bound on \( \sup_{\|t \|_{\infty} \leq T}\left| \big( \log \varphi_X \big)''(t) \right| \) is very conservative—the quantity is typically uniformly bounded, even in extreme case of Gaussian decay. To the best of our knowledge, Theorem (ref) is the first nonparametric Bayes and Wasserstein convergence guarantee for the Multi-Factor model.
Next, similar to the analysis in Section (ref), we focus on establishing further contraction rates in settings where the latent distribution is known to be sufficiently regular. Specifically, we consider the case when the latent distribution is a nonparametric Gaussian mixture. Our main assumption on the nonparametric Gaussian mixture specification is as follows. \begin{condition2}[Nonparametric Mixture] $(i) $ $f_{X_k} = \phi_{\sigma_0^2} \star F_0$ for some $ \sigma_0 > 0$ and mixing distribution $F_0$ that satisfies $ F_0 \big( t \in \mathbb{R} : \left| t \right| > z \big) \leq C \exp(- C' z ^{\chi}) $ for some $\chi > 0$ and universal constants $C,C' > 0$. $(ii)$ $ \sup_{\| t \| \leq T} \left| ( \log \varphi_{F_0})'' (t) \right| \leq C T $ for some universal constant $C > 0$. \end{condition2} The following result demonstrates that, with an appropriately scaled covariance prior, the quasi-Bayes posterior contracts towards the true latent density at a polynomial rate.
Similar to the discussion following Theorem (ref), the exponent $V$ depends on second order factors such as the constants $(\sigma_0^2, c')$.
In this section, we provide simulation evidence on the finite sample performance of quasi-Bayes posteriors. The setup and designs used follows bonhomme2010generalized. Two measurements, \( (Y_1, Y_2) \), are generated according to the specification:
where $(X, \epsilon_1,\epsilon_2 )$ are mutually independent and $X \sim F_X$ is a latent variable whose distribution is of interest. Similar to the analysis in bonhomme2010generalized, we can difference the measurements to obtain
If the nuisance errors are assumed to be symmetric, we can use the observations in ((ref)) as an auxiliary sample for the error in ((ref)). This provides us with a sample to construct the measurement error quasi-Bayes posterior of Section (ref).
HM, LV, BR refer to the estimators proposed in horowitz1996semiparametric, li1998nonparametric and bonhomme2010generalized, respectively. Decon denotes the classical kernel deconvolution estimator carroll1988optimal. Following bonhomme2010generalized; arellano2023recovering, the tuning parameters for these estimators are chosen using the optimal bandwidth selection plug-in method proposed by delaigle2004practical. This tuning selection procedure is data-driven and so it varies across designs and data realizations. Let $\textbf{QB}$ denote the quasi-Bayes posterior mean. In contrast to the other estimators, our implementation of quasi-Bayes (QB) does not vary any tuning parameters—the same choice of \( T \) and prior is applied across all designs and data realizations.
Samples from the quasi-Bayes posterior are generated using Stan. The concentration parameter of the $\text{DP}_{\alpha}$ prior is $\beta=1$ and the base measure is $\alpha = \mathcal{N}(\mathbb{E}_n(\widetilde{Y}_1), \widehat{Var}(\widetilde{Y}_1))$. The scale prior is $ \sigma \sim \text{Inv-Gamma}(1,1)$. We use the same prior for all designs. Further implementation details are provided in the Appendix.
Given an estimator $\widehat{f}$ of the density $f_X$, we denote the MISE by \[ \mathrm{MISE} = \mathbb{E} \left[ \int \big( \widehat{f}(x) - f_X(x) \big)^2 \, dx \right]. \] In Table (ref), we report the MISE for the density $f_X$ with Gaussian error distributions. The QB estimator significantly outperforms alternatives, even when the tuning parameter $T$ is held fixed across designs and data realizations. This stands in contrast to traditional Fourier-based deconvolution methods, which are highly sensitive and require careful tuning. One possible interpretation is that the identifying strength in the Fourier moments is significantly amplified when non-trivial regularity, introduced here via a nonparametric prior, is introduced into the parameter space. In Table (ref), we provide additional results under Laplace errors.
The results in Tables (ref) and (ref) also demonstrate that convergence is exceptionally fast when the latent distribution has a nonparametric Gaussian mixture structure, consistent with the statement of Theorem (ref).
In Table (ref), we report the $L^1$ Wasserstein distance.\footnote{An advantage of focusing on the $\mathbf{W}_1$ metric is that it is well-defined even between mutually singular distributions (e.g. discrete and continuous).} For comparison, we consider the empirical Bayes (denoted as EB) non-parametric maximum likelihood estimator (NPMLE) of $F_X$. The likelihood defining the EB estimator assumes Gaussian noise—we estimate the noise scale using the standard deviation of the measurements in ((ref)). In addition to the distributions mentioned above, we include two additional discrete designs. The first is a two-point discrete distribution, $ F_D = \frac{1}{8} \delta_{x=0} + \frac{7}{8} \delta_{x=2}. $ This distribution was previously used in koenker2019comment to assess the $\mathbf{W}_1$ performance of the EB estimator. The second is a discrete uniform distribution, \( U(\mathcal{S}) \), which assigns equal probability mass to the set of points $ \mathcal{S} = \{-0.5, -0.25, 0, 0.25, 0.5\}. $
The QB estimator performs remarkably well across all designs. As expected, under misspecification, the EB Wasserstein risk increases substantially when the true errors are non-Gaussian. Since the quasi-Bayes likelihood only depends on moments of the observable $(Y_1,Y_2)$, there is no misspecification, and the results are very similar across the two error designs. Consistent with the findings in koenker2019comment, the EB performs exceptionally well for the two-point discrete distribution $F_D$ but experiences a significant decline in accuracy with continuous designs. As the distribution $U(\mathcal{S})$ illustrates, the accuracy loss can also be significant for discrete distributions with a modest number of support points. In this case, the worse performance was primarily due to the EB concentrating nearly all its mass on only two support points.
This paper introduces and develops a quasi-Bayes framework for a large class of latent variable models. Simulations demonstrate that our quasi-Bayes procedures are viable and perform remarkably well relative to existing alternatives. As an application, we used our method to model the latent structure in U.S earnings data. In this section, we discuss several remarks and extensions concerning optimal weighting, the estimation of functionals, and frequentist inference.
The quasi-Bayes characteristic likelihood in ((ref)) is based on identity weighted moment conditions. This framework admits a straightforward extension to allow for optimal weighting. Interestingly, in this case, the objective function shares similarities with the definition proposed in carrasco2000generalization. In practice, the integral in ((ref)) is typically approximated using discrete measure over grid points ${t_1,\dots,t_M} \in \mathbb{B}(T)$. As such, the classical optimally weighted GMM objective at these points can be used to approximate the optimally weighted characteristic likelihood, albeit at the cost of some discretization bias. A formal treatment that accounts for the computational discretization bias would be useful. Ideally, one can also exploit the structure of the problem for greater computational efficiency. For instance, since $\left| \varphi_{Y}(t) \right| \rightarrow 0$ quickly as $\|t \| \rightarrow \infty$, it is reasonable to expect that placing more emphasis on accurate optimal weighting at lower frequency levels will be more beneficial.\footnote{I thank Victor Chernozhukov for this observation; see also ben2021identification.}
An important question concerns the relevance of these results for downstream applications that only depend on a functional of the latent distribution. As discussed in Remark (ref) and Section (ref), when the functional of interest is sufficiently regular, our derived rates admit significantly faster counterparts for that functional. Indeed, this will usually hold whenever the functional is a conditional mean that conditions on information from the observables. In general, however, the regularity of an arbitrary functional is often unknown and may vary substantially with the distribution of the latent unobservables. For this reason, our exposition focuses primarily on the direct estimation of the latent distribution. This is especially important in non-Gaussian models (e.g. guvenen2014nature; arellano2017earnings), where the rates for functionals are typically much slower than in the Gaussian setting.
Quasi-Bayes is an attractive tool in that it not only provides point estimates but also an automatic measure of uncertainty through the posterior distribution. It is natural to question whether such credible sets have valid frequentist guarantees. We note that inferential guarantees are usually difficult to obtain in any model with an infinite dimensional latent distribution. For example, downstream empirical applications of the NPMLE empirical Bayes framework (e.g. gu2023invidious; kline2024discrimination) typically rely exclusively on point estimation, as no distributional theory is available. Similarly, common models for earnings dynamics (e.g. arellano2017earnings; arellano2024heterogeneity) provide inferential guarantees only under correct parametric specification. For suitable parametric versions of our models and priors, classical results by chernozhukov2003mcmc imply the frequentist validity of optimally weighted quasi-Bayes credible sets. Extending these results to our nonparametric setting is an important topic for further investigation.\footnote{Recently, kankanala2023gaussian showed that optimally weighted quasi‐Bayes credible have valid frequentist coverage for regular functionals in a nonparametric conditional moment restriction model. }
Finally, we note that our work provides the first analysis of infinite dimensional latent heterogeneity within a quasi-Bayes framework. Many of our results also provide the first nonparametric Bayes guarantees for the models under consideration. The quasi-Bayes approach is particularly appealing in these settings because the models are often highly ill-posed and require non-trivial forms of regularization. A key takeaway from our findings is that classical nonparametric Bayes priors can significantly enhance the finite sample information content in the moments. This estimation framework complements much of the foundational analysis in the literature (e.g. bonhomme2010generalized) that derived the identifying moments. In this article, we focused on the class of models characterized by Fourier moment conditions, given their widespread use in the literature. The analysis could be extended to many other settings that use moments to characterize infinite dimensional latent heterogeneity (e.g. bonhomme2012functional).