EconBase
← Back to paper

Bayesian Estimation and Comparison of Conditional Moment Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

90,039 characters · 17 sections · 21 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

-.1inBayesian Estimation and Comparison of Conditional Moment Models

abstractWe consider the Bayesian analysis of models in which the unknown distribution of the outcomes is specified up to a set of conditional moment restrictions. The nonparametric exponentially tilted empirical likelihood function is constructed to satisfy a sequence of unconditional moments based on an increasing (in sample size) vector of approximating functions (such as tensor splines based on the splines of each conditioning variable). For any given sample size, results are robust to the number of expanded moments. We derive Bernstein-von Mises theorems for the behavior of the posterior distribution under both correct and incorrect specification of the conditional moments, subject to growth rate conditions (slower under misspecification) on the number of approximating functions. A large-sample theory for comparing different conditional moment models is also developed. The central result is that the marginal likelihood criterion selects the model that is less misspecified. We also introduce sparsity-based model search for high-dimensional conditioning variables, and provide efficient MCMC computations for high-dimensional parameters. Along with clarifying examples, the framework is illustrated with real-data applications to risk-factor determination in finance, and causal inference under conditional ignorability.

Keywords: Bayesian inference, Bernstein-von Mises theorem, Conditional moment restrictions, Exponentially tilted empirical likelihood, Marginal likelihood, Misspecification, Posterior consistency.

\thispagestyle{empty}

{21.0pt}

Introduction

\sloppy

We tackle the problem of prior-posterior inference when the only available information about the unknown parameter ${\Greekmath 0112} \in \Theta \subset \mathbb{R}^{p}$ is supplied by a set of conditional moment (CM) restrictions

equation[equation omitted — 105 chars of source]

where ${\Greekmath 011A} (X,{\Greekmath 0112} )$ is a $d$-vector of known functions of a $\mathbb{R}^{d_{x}}$-valued random vector $X$ and the unknown ${\Greekmath 0112} $, and $P$ is the unknown conditional distribution of $X$ given a $\mathbb{R}^{d_{z}}$-valued random vector $Z$. Such models are important because many standard models in statistics can be recast in terms of CM restrictions. These models also arise naturally in causal inference, missing data problems, and in models derived from theory in economics and finance. Because the CM conditions constrain the set of possible distributions $P$, we say that the model is correctly specified if the true data generating process $P_{\ast }$ is in the set of distributions constrained to satisfy these moment conditions for some ${\Greekmath 0112} \in \Theta $, while the model is misspecified if $P_{\ast }$ is not in the set of implied distributions for any ${\Greekmath 0112} \in \Theta $.

A different starting point is when one is given the unconditional moments, say $\mathbf{E}^{P}[g(X,{\Greekmath 0112} )]=0$. Prior-posterior analysis can then be based on the empirical likelihood, for example, \citet*{Lazar} and many others, or the exponentially tilted empirical likelihood (ETEL), as in \citet*{Schennach2005 } and \citet*{ChibShinSimoni2018}. Developing a Bayesian framework for CM models is important. While it is true that the conditional moments imply that ${\Greekmath 011A} (X,{\Greekmath 0112} )$ is uncorrelated with $Z$, i.e., $\mathbf{E}^{P}[{\Greekmath 011A} (X, {\Greekmath 0112} ) \otimes Z]=0$, where $\otimes $ is the Kronecker product operator, the conditional moments assert even more, that ${\Greekmath 011A} (X,{\Greekmath 0112} )$ is uncorrelated with any measurable, bounded function of $Z$. Thus, there is an efficiency loss if this information is ignored.

We approach this problem by first constructing $K$ unconditional moments

equation[equation omitted — 131 chars of source]

based on an increasing (in sample size) vector of approximating functions, $q^{K}(Z)\coloneqq (q_{1}^{K}(Z),\ldots ,q_{K}^{K}(Z))'$, obtained, for instance, from splines of each variable in $Z$ \citep*{DonaldImbensNewey2003}. Efficiency loss is avoided as the number of moments increases with sample size. Next, for each sample size and for each ${\Greekmath 0112}$, the nonparametric exponentially tilted empirical likelihood (ETEL) function is constructed to satisfy these unconditional moments. Unlike the empirical likelihood, the ETEL function has a fully Bayesian interpretation. It is the likelihood that emerges from integrating out $P$ with respect to a nonparametric prior that satisfies the CMs. The posterior of interest is then this nonparametric likelihood multiplied by a prior distribution of the parameters. Due to the fact that the nonparametric likelihood is limited to a set $H_{n,K}$ of ${\Greekmath 0112}$ values for which the empirical counterpart of the moment conditions (ref) are equal to 0, the posterior (equivalently, the prior) is truncated to the set $H_{n,K}$.

We study the prior-posterior mapping on many fronts, taking up the question of misspecified models, model comparisons, and computations, combining careful theoretical work with the needs of applications. The posterior distribution is shown to satisfy Bernstein-von Mises (BvM) theorems in both the correct and misspecified cases. In the former case the growth of $K$ (for approximating functions given by splines) is at most $n^{1/6}$, where $n$ is the sample size. The asymptotic posterior variance is then equal to the semiparametric efficiency bound derived in \citet*{Chamberlain1987}. In the latter case, in parallel with \citet*{kleijn2012}, the posterior distribution of the centered and scaled parameter $\sqrt{n}({\Greekmath 0112} -{\Greekmath 0112} _{\circ })$, where ${\Greekmath 0112} _{\circ }$ is the pseudo-true value, converges to a Normal distribution with variance that now is different from the variance of the frequentist estimator. Interestingly, this convergence holds only if $K$ increases more slowly than in the correctly specified case. This can be interpreted as limiting the number of implied unconditional moments to limit the magnification of the misspecification.

We informally use these rate conditions from the theoretical analysis to guide the range of choice of $K$ for any given $n$. Due to the fact that for a fixed $n$ the volume (prior probability content) of the region of truncation $H_{n,K}$ decreases with $K$ (a result of more restrictions), values of $K$ beyond the range recommended by the theory amplify the Bayesian bias, and, hence, should be avoided. Large values of $K$ can also produce rank-deficiency of the approximating functions basis matrix and, in the event of a misspecified model, increase misspecification. Around the values of $K$ we recommend, the posterior distribution is generally robust to $K$, and little fine-tuning is necessary.

Finite sample summaries of the posterior distribution are obtained by Markov chain Monte Carlo (MCMC) methods. Since the posterior is underpinned by a non-parametric likelihood, and the effective prior is truncated, efficient sampling is not automatic. However, after extensive study, we have produced a near-black-box MCMC approach (available as a R-package) that is based on the tailored Metropolis-Hastings (M-H) algorithm of \citet*{ChibGreenberg1995} and its randomized version in \citet*{ChibRamamurthy2010}.

The entire paper is interspersed with examples of pedagogical importance and practical relevance. Real data applications to risk-factor determination in finance, and causal inference under conditional ignorability, are included.

It is worth noting that previous Bayesian work on conditional moments, for example, \citet*{liao2011}, \citet*{FS2012JoE, florens_simoni_2016}, \citet*{kato2013}, \citet*{chen2017monte} and LIAOSimoni2019, has little overlap with the discussion here. A major difference is that none of these papers adopt the fully Bayesian ETEL framework. Another is that these papers examine a different class of CM models. Finally, none of these papers takes up the question of model comparisons. Nonetheless, these papers and the current work, taken together, represent an important broadening of the Bayesian enterprise to new classes of models.

The rest of the paper is organized as follows. Section (ref) has the sketch of the conditional moment setting. Section (ref) discusses the prior-posterior analysis and the large-sample properties of the posterior distribution. Section (ref) is concerned with the problem of comparing CM models via marginal likelihoods. In Section (ref) two extensions are considered and Section (ref) has real data applications to finance and causal inference. Section (ref) concludes. Proofs are in the online supplementary appendix.

Setting and Motivation

Let $X\coloneqq(X_{1}^{\prime },X_{z}^{\prime })^{\prime }$ be an $\mathbb{R} ^{d_{x}}$-valued random vector and $Z\coloneqq(Z_{1}^{\prime },X_{z}^{\prime })^{\prime }$ be an $\mathbb{R}^{d_{z}}$-valued random vector. The vectors $ Z $ and $X$ have elements in common if the dimension of the subvector $X_{z}$ is non-zero. Moreover, we denote $W\coloneqq(X^{\prime },Z_{1}^{\prime })^{\prime }\in \mathbb{R}^{d_{w}}$ and its (unknown) joint distribution by $P$. By abuse of notation, let $P$ also denote the associated conditional distribution. Suppose that we are given a random sample $ W_{1:n}=(W_{1},\ldots ,W_{n})$ of $W$. Hereafter, $\mathbf{E} ^{P}[\cdot ]$ is the expectation with respect to $P$ and $\mathbf{E} ^{P}[\cdot |\cdot ]$ is the conditional expectation with respect to the conditional distribution associated with $P$.

The parameter of interest is ${\Greekmath 0112} \in \Theta \subset \mathbb{R}^{p}$, which is related to the conditional distribution $P$ through the conditional moment restrictions

equation[equation omitted — 97 chars of source]

where ${\Greekmath 011A} (X,{\Greekmath 0112} )$ is a $d$-vector of known functions. Many interesting and important models in statistics fall into this framework.

\begin{example1} (Linear model with heteroscedasticity of unknown form) Suppose that

equation[equation omitted — 111 chars of source]

where ${\Greekmath 011A} (X,{\Greekmath 0112} )=\left( Y-{\Greekmath 0112} _{0}-{\Greekmath 0112} _{1}X\right) $, $Z=(1,X)$ and $d=1$. This CM model is consistent with the data generating process (DGP) $Y={\Greekmath 0112} _{0}+{\Greekmath 0112} _{1}X+{\Greekmath 0122} $, where ${\Greekmath 0122} = h(X)U$, and $(X,U)$ (independent) follow some unknown distribution $P$, with $E(U)=0$, and the heteroscedasticity function $h(X)$ is unknown. The restrictions

equation[equation omitted — 306 chars of source]

where now ${\Greekmath 011A} (X,{\Greekmath 0112} )$ is a $(2\times 1)$ vector of functions, additionally impose that ${\Greekmath 0122} $ is conditionally symmetric. \end{example1}

Note that in the foregoing example, the two unconditional moment conditions

equation[equation omitted — 134 chars of source]

which assert that: (i) ${\Greekmath 0122} $ has mean zero and (ii) ${\Greekmath 0122}$ is uncorrelated with $X$, are weaker but, if the CM model is correct, less informative about ${\Greekmath 0112}$.

Prior-Posterior Analysis

Expanded Moment Conditions

The starting point, as in the frequentist approaches of \citet*{DonaldNewey2001}, \citet*{AiChen2003} and carrasco_florens_2000, is a transformation of the CM restrictions into unconditional moment restrictions. Following \citet*{DonaldImbensNewey2003}, let $q^{K}(Z) \coloneqq (q_{1}^{K}(Z),\ldots ,q_{K}^{K}(Z))^{\prime }$, $K>0$, denote a $K$-vector of real-valued functions of $Z$, for instance, splines. Suppose that these functions satisfy the following condition for the distribution $P$.

assumFor all $K$, $\mathbf{E}^P[q^K(Z)^{\prime} q^K(Z)]$ is finite, and for any function $a(z): \R^{d_{z}} \rightarrow \R $ with $ \mathbf{E}^P[a(Z)^2]< \infty$ there are $K\times 1$ vectors ${\Greekmath 010D}_K$ such that as $K\rightarrow \infty$, \begin{equation*} \mathbf{E}^P[(a(Z)-q^K(Z)^{\prime }{\Greekmath 010D}_K)^2]\rightarrow 0. \end{equation*}

Now, let ${\Greekmath 0112}_*$ be the value of ${\Greekmath 0112}$ that satisfies (ref) for the true $P$. If $\mathbf{E}^{P}[{\Greekmath 011A} (X,{\Greekmath 0112}_*)^{\prime_*}{\Greekmath 011A} (X,{\Greekmath 0112} )]<\infty $, then \citet*[Lemma 2.1]{DonaldImbensNewey2003} established that: (1) if equation (ref) is satisfied with ${\Greekmath 0112} ={\Greekmath 0112}_{\ast }$, then $\mathbf{E}^{P}[{\Greekmath 011A} (X,{\Greekmath 0112} _{\ast })\otimes q^{K}(Z)]=0$ for all $K$; (2) if equation (ref) is not satisfied by ${\Greekmath 0112} ={\Greekmath 0112}_{\ast }$, then $\mathbf{E} ^{P}[{\Greekmath 011A} (X,{\Greekmath 0112} _{\ast })\otimes q^{K}(Z)]\neq 0$, for all large enough $ K$.

Henceforth, we let $g(W,{\Greekmath 0112} )\coloneqq{\Greekmath 011A} (X,{\Greekmath 0112} )\otimes q^{K}(Z)$ denote the expanded functions and refer to

equation[equation omitted — 80 chars of source]

as the expanded moments. Under the stated assumptions, the expanded moments are equivalent to the CM restrictions (ref), as $K\rightarrow \infty$.

In our numerical examples, we construct $q^{K}(Z)$ using the natural cubic spline basis of \citet*{Chib2010}, with $K$ fixed at a given value, as in sieve estimation. If $Z$ consists of more than one element, say $(Z_{1},Z_{2},Z_{3})$ where $Z_{1}$ and $Z_{2}$ are continuous variables and $Z_{3}$ is binary, then the basis matrix $B$ is constructed as follows. Let $\boldsymbol{z}_{j}$ denote the $n\times 1$ sample data on $Z_{j}$ $\left( j\leq 3\right) $. Let $\boldsymbol{Z}=\left( \boldsymbol{z}_{1},\boldsymbol{z }_{2},\boldsymbol{z}_{1}\odot \boldsymbol{z}_{2},\boldsymbol{z}_{1}\odot \boldsymbol{z}_{3},\boldsymbol{z}_{2}\odot \boldsymbol{z}_{3}\right) $ denote the $n\times 5$ matrix of the continuous data and interactions of the continuous data and the binary data. Now suppose $\left( {\Greekmath 011C} _{j1},\ldots ,{\Greekmath 011C} _{jK}\right) $, for $j=1,\ldots,5$ are $K$ knots based on each column of $\boldsymbol{Z}$ and let $B_{j}$ denote the corresponding $n\times K$ matrix of cubic spline basis functions. Then, $B$ is given by

equation*[equation* omitted — 165 chars of source]

where $B_{j}^{\ast }$ $\left( j=2,3,4,5\right) $ is the $n\times (K-1)$ matrix in which each column of $B_{j}$ is subtracted from its first and then the first column is dropped, see \citet*{Chib2010}. Thus, the dimension of this $B$ matrix is $n\times K^{\ast}$, where $K^{\ast} = \left( 5K-4+1\right)$. If $K^{\ast}$ is large, in relation to $n$, data-compression methods can be employed. Specifically, let $R$ denote the $K^{\ast} \times K^{\ast}$ orthogonal matrix of eigenvectors from the singular value decomposition of $B$, and let $e$ denote the corresponding $K^{\ast} \times 1$ vector of eigenvalues. Then, after employing the rotation $BR$, the columns of $BR$ corresponding to small values of $e$ are dropped, and the resulting column-reduced $BR$ matrix is taken as the basis matrix. We refer to this as the rotated column reduced basis matrix. To define the expanded functions, let ${\Greekmath 011A} _{l}(\boldsymbol{X},{\Greekmath 0112} )$ $(l\leq d)$ denote a $ n\times 1$ vector of the $l$th element of ${\Greekmath 011A} (X,{\Greekmath 0112} )$ evaluated at the sample data matrix $\boldsymbol{X}$. Then, the expanded functions for the sample observations are obtained by multiplying ${\Greekmath 011A} _{l}(\boldsymbol{X} ,{\Greekmath 0112} )$ by the matrix $B$ (or by the rotated column reduced $B$) and concatenating. We use versions of this approach in our examples.

Posterior distribution

We base the prior-posterior mapping, for each sample size and ${\Greekmath 0112}$, on the nonparametric exponentially tilted empirical likelihood (ETEL) function. The ETEL has a fully Bayesian interpretation \citep*{Schennach2005} as an integrated likelihood, integrated over the prior on the data distribution $P$ that satisfies the given moments. Other such priors exist, for example, \citet*{KitamuraOtsu2011}, \citet*{Shin2014} and \citet*{FlorensSimoni2015}, that lead to different integrated likelihoods.

The ETEL function takes the form

equation[equation omitted — 99 chars of source]

where $\{\widehat{p}_{i}({\Greekmath 0112} ),i=1,\ldots ,n\}$ are the probabilities that minimize the Kullback-Leibler divergence between the probabilities $(p_{1},\ldots ,p_{n})$ assigned to each sample observation and the empirical probabilities $(\frac{1}{n},\ldots ,\frac{1}{n})$, subject to the conditions that the probabilities $(p_{1},\ldots ,p_{n})$ sum to one and that the expectation under these probabilities satisfy the given unconditional moment conditions (ref).

Specifically, $\{\widehat{p} _{i}({\Greekmath 0112} ),i=1,\ldots ,n\}$ are the solution of the following problem:

equation[equation omitted — 270 chars of source]

(see \citet*{Schennach2005} for a proof). In practice, the solution of this problem emerges from the dual (saddlepoint) representation (see e.g. \citet*{csiszar1984}) as

equation[equation omitted — 309 chars of source]

where $\widehat{{\Greekmath 0115} }({\Greekmath 0112} )=\arg \min_{{\Greekmath 0115} \in \mathbb{R}^{dK}} \frac{1}{n}\sum_{i=1}^{n}e^{{\Greekmath 0115} ^{\prime }g(w_{i},{\Greekmath 0112} )}$ is the estimated tilting parameter.

Let $\mathcal{C}o({\Greekmath 0112}) \coloneqq \left\{ \sum_{i=1}^{n}p_{i}g(w_{i},{\Greekmath 0112}),p_{i}\geq 0, \sum_{i=1}^n p_i = 1\right\}$ be the convex hull of $\{g(w_{i},{\Greekmath 0112} )\}_{i=1}^n$ for a given ${\Greekmath 0112}$ and $\overline{\mathcal{C}o({\Greekmath 0112})}$ denote its interior. Let $H_{n,K}\coloneqq\left\{ {\Greekmath 0112} \in \Theta ; 0\in \overline{\mathcal{C}o({\Greekmath 0112})}\right\} $ denote the set of ${\Greekmath 0112} $ values for which the empirical moment conditions hold. Then, the posterior distribution is the truncated distribution given by

equation[equation omitted — 389 chars of source]

where $1\{\cdot\}$ is the indicator function.

Combining the indicator function with the prior, we see that the (effective) prior is truncated to ${\Greekmath 0112}\in H_{n,K}$. This fact can be used to argue that, for fixed $n$, it is not desirable to have a large $K$. This is because as $K$ increases for a given $n$, the support of the prior shrinks (equivalently, the prior probability content of the region of truncation decreases), due to the fact that more restrictions are imposed. We refer to this prior probability content by the shorthand, volume. Reduction in the volume tends to increase the Bayesian bias and reduce the posterior spread, without any change in the data, with deleterious impact on the posterior. In practice, we use the rule $2n^{1/6}$ to fix $K$. Larger values than this can, of course, be tried, but one should make sure that the volume of $H_{n,K}$ does not become much smaller than one. Around the values of $K$ we recommend, the posterior distribution is generally robust to $K$, and little fine-tuning is necessary.

\begin{contexample1} To illustrate the role of $K$ in the prior-posterior analysis, and its impact on the volume (prior probability content) of $H_{n,K}$, we create a set of simulated data $\{y_{i},x_{i}\}_{i=1}^{n}$, $n=250$, with covariates $X\sim \mathcal{U}(-1,2.5)$, intercept ${\Greekmath 0112} _{0}=1$, slope ${\Greekmath 0112} _{1}=1$, and ${\Greekmath 0122} _{i}$ is distributed according to ${\Greekmath 0122} _{i}\sim\mathcal{SN}(m(x_{i}),h(x_{i}),s(x_{i}))$, where $\mathcal{SN}(m,h,s)$ is the skew normal distribution with location, scale, and shape parameters given by $(m,h,s)$, each depending on $x_{i}$. When $s$ is zero, $ {\Greekmath 0122} _{i}$ is normal with mean $m$ and standard deviation $h$. We set $m(x_{i})=-h(x_{i})\sqrt{2/{\Greekmath 0119} }s(x_{i})/(\sqrt{1+s(x_{i})^{2}})$, so that $\mathbf{E}^{P}[{\Greekmath 0122} |X]=0$.

Suppose that $h(x)=\sqrt{\exp (1+0.7x+0.2x^{2})}$ and $s(x)=1+x^{2}$. Model parameters are estimated solely from the condition $\mathbf{E}^{P}[{\Greekmath 0122} |Z]=0$, $Z=(1,X)$. The prior is the default independent student-$t$ distribution with location 0, dispersion 5, and degrees of freedom 2.5, truncated to $H_{n,K}$. The posterior is computed for $K$ given by 2, $2n^{1/6}$, 9 and 20 (the value $2n^{1/6}$ is based on the theory below for splines approximating functions). The results are shown in Table (ref). Importantly, when $K$ is close to the value suggested by theory, the posterior distribution is robust to $K$. However, when $K=20$, quite different from the recommended value, the Bayesian bias is larger and the posterior standard deviation is smaller, without any change in the data. This is due to the effect of the prior, in the following way. As $K$ increases for a fixed $n$, the volume of $H_{n,K}$ decreases, equivalently, the support of the prior distribution shrinks, as illustrated in Figure 1. This explains why values of $K$ close to the recommended value are preferred.

As an aside, if the true model was unconditional (i.e., without conditional heteroskedasticity), then there is little loss in using more expanded moments - the extra moments are superfluous and hence do not change the effective support of the prior. In that case, no tangible cost in imposed, apart from the computational burden of carrying those moments along.

table[table omitted — 1,691 chars of source]
figure[figure omitted — 1,021 chars of source]

\end{contexample1}

Asymptotic properties

Consider now the large sample behavior of the posterior distribution of ${\Greekmath 0112}$. We let ${\Greekmath 0112}_*$ and $P_*$, respectively, denote the true value of ${\Greekmath 0112}$ and of the data distribution $P$. As notation, when the true distribution $P_{*}$ is involved, expectations $\mathbf{E}^{P}[\cdot ]$ (resp. $\mathbf{E}^{P}[\cdot |\cdot ]$) are taken with respect to $P_{*}$ (resp. the conditional distribution associated with $P_{*}$). In addition, we denote $\ell_{n,{\Greekmath 0112}}(W_i) \coloneqq \log \widehat p_{i}({\Greekmath 0112}) = \log\frac{e^{\widehat{{\Greekmath 0115}}({\Greekmath 0112})'g(W_i,{\Greekmath 0112})}}{\sum_{j=1}^n e^{\widehat{{\Greekmath 0115}}({\Greekmath 0112})'g(W_j,{\Greekmath 0112})}}$, $${\Greekmath 011A}_{{\Greekmath 0112}}(X,{\Greekmath 0112}) \coloneqq \frac{\partial {\Greekmath 011A}(X,{\Greekmath 0112})}{\partial{\Greekmath 0112}^{\prime }},\qquad D(Z) \coloneqq \mathbf{E}^P[{\Greekmath 011A}_{{\Greekmath 0112}}(X,{\Greekmath 0112}_*)|Z],$$ $$\Sigma(Z)\coloneqq\mathbf{E}^P[{\Greekmath 011A}(X,{\Greekmath 0112}_*){\Greekmath 011A}(X,{\Greekmath 0112}_*)'|Z],\quad \textrm{ and }\quad {\Greekmath 011A}_{j{\Greekmath 0112}{\Greekmath 0112}}(X,{\Greekmath 0112}_*) \coloneqq \partial^2{\Greekmath 011A}_j(X,{\Greekmath 0112}_*)/\partial{\Greekmath 0112}\partial {\Greekmath 0112}'.$$ For a vector $a$, $\|a\|$ denotes the Euclidean norm. For a matrix $A$, $\|A\|$ denotes the operator norm (the largest singular value of the matrix). Finally, let $\mathcal{Z}\coloneqq\mathrm{supp}(Z)$ denote the support of $Z$.\\ The first assumption is a normalization for the second moment matrix of the approximating functions which is standard in the literature, see e.g. NEWEY1997 and DonaldImbensNewey2003.

assumFor each $K$ there is a constant scalar ${\Greekmath 0110}(K)$ such that $\sup_{z\in\mathcal{Z}}\|q^K(z)\| \leq {\Greekmath 0110}(K)$, $\mathbf{E}^P[q^K(Z) q^K(Z)']$ has smallest eigenvalue bounded away from zero uniformly in $K$, and $\sqrt{K}\leq {\Greekmath 0110}(K)$.

The bound ${\Greekmath 0110}(K)$ is known explicitly in a number of cases depending on the approximating functions we use. DonaldImbensNewey2003 provide a discussion and explicit formulas for ${\Greekmath 0110}(K)$ in the case of splines, power series and Fourier series. We also refer to NEWEY1997 for primitive conditions for regression splines and power series.

assum(a) There exists a unique ${\Greekmath 0112}_*\in\Theta$ that satisfies $\mathbf{E}^P[{\Greekmath 011A}(X,{\Greekmath 0112})|Z] = 0$ for the true $P_*$; (b) the data $W_i\coloneqq(X_i,Z_{1i})$, $i=1,\ldots, n$ are i.i.d. according to $P_*$; (c) $\mathbf{E}^P[\sup_{{\Greekmath 0112}\in\Theta}\|{\Greekmath 011A}(X,{\Greekmath 0112})\|^2|Z]$ is bounded.

This assumption is the same as DonaldImbensNewey2003. The following three assumptions are also the same as the ones required by DonaldImbensNewey2003 to establish asymptotic normality of the Generalized Empirical Likelihood (GEL) estimator.

assum(a) ${\Greekmath 0112}_*\in int(\Theta)$; (b) ${\Greekmath 011A}(X,{\Greekmath 0112})$ is twice continuously differentiable in a neighborhood $\mathcal{U}$ of ${\Greekmath 0112}_*$, $\mathbf{E}^P[\sup_{{\Greekmath 0112}\in\mathcal{U}}\|{\Greekmath 011A}_{{\Greekmath 0112}}(X,{\Greekmath 0112})\|^2|Z]$ and $\mathbf{E}^P[\sup_{{\Greekmath 0112}\in\mathcal{U}}\|{\Greekmath 011A}_{j{\Greekmath 0112}{\Greekmath 0112}}(X,{\Greekmath 0112}_*)\|^2|Z]$, $j=1,\ldots d$, are bounded on $\mathcal{Z}$; (c) $\mathbf{E}^P[D(X)D(X)']$ is nonsingular.
assum(a) $\Sigma(Z)$ has smallest eigenvalue bounded away from zero; (b) for a neighborhood $\mathcal{U}$ of ${\Greekmath 0112}_*$, $\mathbf{E}^P[\sup_{{\Greekmath 0112}\in\mathcal{U}}\|{\Greekmath 011A}(X,{\Greekmath 0112})\|^4|z]$ is bounded, and for all ${\Greekmath 0112}\in\mathcal{U}$, $\|{\Greekmath 011A}(X,{\Greekmath 0112}) - {\Greekmath 011A}(X,{\Greekmath 0112}_*)\| \leq {\Greekmath 010E}(X)\|{\Greekmath 0112} - {\Greekmath 0112}_*\|$ and $\mathbf{E}^P[{\Greekmath 010E}(X)^2|Z]$ is bounded.
assumThere is ${\Greekmath 010D} > 2$ such that $\mathbf{E}^P[\sup_{{\Greekmath 0112}\in\Theta}\|{\Greekmath 011A}(X,{\Greekmath 0112})\|^{{\Greekmath 010D}}] < \infty$ and ${\Greekmath 0110}(K)^2 K/ n^{1 - 2/{\Greekmath 010D}} \rightarrow 0$.

Part (b) of Assumption (ref) imposes a Lipschitz condition which allows application of uniform convergence results. The last assumption is about the prior distribution of ${\Greekmath 0112}$ and is standard in the Bayesian literature on frequentist asymptotic properties of Bayes procedures.

assum(a) ${\Greekmath 0119}$ is a continuous probability measure that admits a density with respect to the Lebesgue measure; (b) ${\Greekmath 0119}$ is positive on a neighborhood of ${\Greekmath 0112}_*$.

We are now able to state our first major result in which we establish the asymptotic normality and efficiency of the posterior distribution of the local parameter $h \coloneqq \sqrt{n}({\Greekmath 0112} - {\Greekmath 0112}_*)$.

thm[Bernstein-von Mises] Under Assumptions (ref)-(ref), if $K\rightarrow \infty$, ${\Greekmath 0110}(K) K^2/\sqrt{n} \rightarrow 0$, and if for any ${\Greekmath 010E} > 0$, $\exists{\Greekmath 010F}>0$ such that as $n\rightarrow \infty$ \begin{equation} P\left(\sup_{\|{\Greekmath 0112} - {\Greekmath 0112}_*\| > {\Greekmath 010E}}\frac{1}{n}\sum_{i=1}^n \left(\ell_{n,{\Greekmath 0112}}(W_i) - \ell_{n,{\Greekmath 0112}_*}(W_i)\right) \leq - {\Greekmath 010F}\right) \rightarrow 1, \end{equation} then the posterior distribution ${\Greekmath 0119}(\sqrt{n}({\Greekmath 0112} - {\Greekmath 0112}_*)|W_{1:n})$ converges in total variation towards a random Normal distribution, that is, \begin{equation} \sup_{B}\left|{\Greekmath 0119}(\sqrt{n}({\Greekmath 0112}-{\Greekmath 0112}_*)\in B|W_{1:n},K) - \mathcal{N}_{\Delta_{n,{\Greekmath 0112}_*},V_{{\Greekmath 0112}_*}}(B)\right|\overset{p}{\to} 0, \end{equation} where $B\subseteq \Theta$ is any Borel set, $\Delta_{n,{\Greekmath 0112}_*} \coloneqq -\frac{1}{\sqrt{n}}\sum_{i=1}^n V_{{\Greekmath 0112}_*}D(Z_i)'\Sigma(Z_i)^{-1}{\Greekmath 011A}(X_i,{\Greekmath 0112}_*)$ is bounded in probability and $V_{{\Greekmath 0112}_*} \coloneqq \left(\mathbf{E}^P[D(Z)' \Sigma(Z)^{-1} D(Z)]\right)^{-1}$.

We note that the centering $\Delta_{n,{\Greekmath 0112}_*}$ of the limiting normal distribution satisfies $\frac{1}{\sqrt{n}}\sum_{i=1}^n\frac{d\log\widehat p_i({\Greekmath 0112}_*)}{d{\Greekmath 0112}} - V_{{\Greekmath 0112}_*}^{-1}\Delta_{n,{\Greekmath 0112}_*} \overset{p}{\to} 0$. We also note that the condition ${\Greekmath 0110}(K) K^2/\sqrt{n} \rightarrow 0$ in the theorem implies $K/n\rightarrow 0$, which is a classical condition in the sieve literature. This condition is required to establish a stochastic Local Asymptotic Normality (LAN) expansion, which is an intermediate step to prove the BvM result, as we explain below. The LAN expansion is not required to establish asymptotic normality of the GEL estimators, which explains why our condition is slightly stronger than the condition ${\Greekmath 0110}(K) K/\sqrt{n} \rightarrow 0$ required by \citet*{DonaldImbensNewey2003}. On the other hand, our condition is weaker than the condition ${\Greekmath 0110}(K)^2 K^2/\sqrt{n} \rightarrow 0$ required by \citet*{DONALD2009} to establish the mean square error of the GEL estimators. The asymptotic covariance of the posterior distribution coincides with the semiparametric efficiency bound given in Chamberlain1987 for conditional moment condition models. This means that, for every ${\Greekmath 010B}\in (0,1)$, $(1 - {\Greekmath 010B})$-credible regions constructed from the posterior of ${\Greekmath 0112}$ are $(1 - {\Greekmath 010B})$-confidence sets asymptotically.

The proof of this theorem is given in the supplementary appendix and consists of three steps. In the first step we show consistency of the posterior distribution of ${\Greekmath 0112}$, namely:

equation[equation omitted — 188 chars of source]

for any $M_n\rightarrow \infty$, as $n\rightarrow \infty$. To show this, the identification assumption (ref) is used. In the second step we show that the ETEL function satisfies a stochastic LAN expansion:

equation[equation omitted — 321 chars of source]

where $\mathcal{H}$ denotes a compact subset of $\mathbb{R}^p$ and $V_{{\Greekmath 0112}_*}^{-1}\Delta_{n,{\Greekmath 0112}_*}\overset{d}{\to} \mathcal{N}(0, V_{{\Greekmath 0112}_*}^{-1})$. As the ETEL function is an integrated likelihood, expansion (ref) is better known as integral LAN in the semiparametric Bayesian literature, see e.g. BickelKleijn2012. In the third step of the proof we use arguments as in the proof of VanDerVaart2000 to show that (ref) and (ref) imply asymptotic normality of ${\Greekmath 0119}(\sqrt{n}({\Greekmath 0112}-{\Greekmath 0112}_*)\in B|W_{1:n},K)$. While these three steps are classical in proving the Bernstein-von Mises phenomenon, establishing (ref) raises challenges that are otherwise absent. This is because the ETEL function is a nonstandard likelihood that involves estimated parameters $\widehat{\Greekmath 0115}({\Greekmath 0112}_*)$ whose dimension is $d K$, which increases with $n$. While $\|\widehat{\Greekmath 0115}({\Greekmath 0112}_*)\|$ and $\|\frac{1}{n}\sum_{i=1}^n g(W_i,{\Greekmath 0112}_*)\|$ are expected to converge to zero in the correctly specified case, the rate of convergence is slower than $n^{-1/2}$. In the supplementary appendix we show that this rate is $\sqrt{K/n}$ under the previous assumptions.

Misspecified model

We now generalize the preceding BvM result for the important class of misspecified conditional moment models.

df[Misspecified model] We say that the conditional moment conditions model $\mathbf{E}^P[{\Greekmath 011A}(X,{\Greekmath 0112})|Z] = 0$ is misspecified if the set of probability measures implied by the moment restrictions does not contain the true data generating process $P_*$ for any ${\Greekmath 0112}\in \Theta$, that is, $P_*\notin \mathcal{P}$ where $\mathcal{P} \coloneqq \bigcup_{{\Greekmath 0112}\in\Theta }\widetilde{\mathcal{P}}_{{\Greekmath 0112}}$ and $\widetilde{\mathcal{P}}_{{\Greekmath 0112}} = \{Q\in \mathbb{M}_{X|Z}; \,\mathbf{E}^Q[{\Greekmath 011A}(X,{\Greekmath 0112})|Z] = 0 \; a.s.\}$ with $\mathbb{M}_{X|Z}$ the set of all conditional probability measures of $X|Z$.

In essence, if (ref) is misspecified then there is no ${\Greekmath 0112}\in\Theta$ such that $\mathbf{E}^P[{\Greekmath 011A}(X,{\Greekmath 0112})\otimes q^K(Z)] = 0$ almost surely for every $K$ large enough. Now, for every ${\Greekmath 0112}\in\Theta$ define $Q^*({\Greekmath 0112})$ as the minimizer of the Kullback-Leibler divergence of $P_*$ to the model $\mathcal{P}_{{\Greekmath 0112}} \coloneqq \{Q\in \mathbb{M}; \mathbf{E}^Q[g(W,{\Greekmath 0112})]=0\}$, where $\mathbb{M}$ denotes the set of all the probability measures on $\mathbb{R}^{d_w}$. That is, $Q^*({\Greekmath 0112}) \coloneqq \mathrm{arginf}_{Q\in\mathcal{P}_{{\Greekmath 0112}}}\mathbb{K}(Q||P_*)$, where $\mathbb{K}(Q||P_*)\coloneqq\int\log(dQ/dP_*)dQ$. If we suppose that the dual representation of the Kullback-Leibler minimization problem holds, then the $P_*$-density of $Q^*({\Greekmath 0112})$ has the closed form: $[dQ^*({\Greekmath 0112})/dP_*](W_i) = \frac{e^{{\Greekmath 0115}_{\circ}'g(W_i,{\Greekmath 0112})}}{\mathbf{E}^P[ e^{{\Greekmath 0115}_{\circ}'g(W_j,{\Greekmath 0112})}]}$, where ${\Greekmath 0115}_\circ$ denotes the tilting parameter and is defined in the same way as in the correctly specified case:

equation[equation omitted — 252 chars of source]

We also impose a condition to ensure that the probability measures $\mathcal{P}\coloneqq\bigcup_{{\Greekmath 0112}\in{\Greekmath 0112}}\mathcal{P}_{\Greekmath 0112}$, which are implied by the model, are dominated by the true probability measure $P_*$. This is required for the validity of the dual theorem. Therefore, following Sueishi2013, we replace Assumption (ref) (a) by the following.

assumFor every ${\Greekmath 0112}\in\Theta$, there exists $Q\in \mathcal{P}_{{\Greekmath 0112}}$ such that $Q$ is mutually absolutely continuous with respect to $P_*$, where $\mathcal{P}_{{\Greekmath 0112}}\coloneqq\{Q\in \mathbb{M}; \mathbf{E}^Q[g(W,{\Greekmath 0112})]=0\}$ and $\mathbb{M}$ denotes the set of all the probability measures on $\mathbb{R}^{d_w}$.

This assumption implies that $\mathcal{P}_{{\Greekmath 0112}}$ is non-empty. A similar assumption is also made by kleijn2012 and \citet*{ChibShinSimoni2018} to establish the BvM under misspecification. The pseudo-true value of the parameter ${\Greekmath 0112}\in \Theta$ is denoted by ${\Greekmath 0112}_\circ$ and is defined as the minimizer of the Kullback-Leibler divergence between the true $P_*$ and $Q^*({\Greekmath 0112})$:

equation[equation omitted — 163 chars of source]

where $\mathbb{K}(P_*||Q^*({\Greekmath 0112}))\coloneqq\int \log(dP_*/dQ^*({\Greekmath 0112})) dP_*$. Under the preceding absolute continuity assumption, the pseudo-true value ${\Greekmath 0112}_\circ$ is available as

equation[equation omitted — 282 chars of source]

Note that ${\Greekmath 0115}_\circ ({\Greekmath 0112}_\circ)$, the value of the tilting parameter at the pseudo-true value ${\Greekmath 0112}_\circ$, is nonzero because the moment conditions do not hold.\\ Assumption (ref) implies that $\mathbb{K}(Q^*({\Greekmath 0112}_\circ)||P_*)<\infty$. We supplement this with the assumption that $\mathbb{K}(P_*||Q^*({\Greekmath 0112}))<\infty$, $\forall {\Greekmath 0112}\in\Theta$ (so that $\mathbb{K}(P_*||Q^*({\Greekmath 0112}_\circ))<\infty$). Because consistency in misspecified models is defined with respect to the pseudo-true value ${\Greekmath 0112}_\circ$, we need to replace Assumption (ref) (b) by the following Assumption (ref) (b) which, together with Assumption (ref) (a), requires the prior to put enough mass to balls around ${\Greekmath 0112}_\circ$.

assum(a) ${\Greekmath 0119}$ is a continuous probability measure that admits a density with respect to the Lebesgue measure; (b) The prior distribution ${\Greekmath 0119}$ is positive on a neighborhood of ${\Greekmath 0112}_\circ$, where ${\Greekmath 0112}_\circ$ is as defined in (ref).

Hereafter, we use the sub/super index $Q^*({\Greekmath 0112}_\circ)$ to denote an expectation, a variance or covariance taken with respect to the probability $Q^*({\Greekmath 0112}_\circ)$. The following assumption is analogous to the second part of Assumption (ref) for the $P_*$-density of $Q_*({\Greekmath 0112})$ replacing $P$.

assumFor each $K$ the matrix $\mathbf{E}^{Q^*({\Greekmath 0112}_\circ)}[q^K(Z) q^K(Z)']$ has smallest eigenvalue bounded away from zero uniformly in $K$.

In the next assumption we denote by $int(\Theta)$ the interior of $\Theta$ and by $\mathcal{U}$ a ball centered at ${\Greekmath 0112}_\circ$ with radius $h/\sqrt{n}$ for some $h\in\mathcal{H}$ and $\mathcal{H}$ a compact subset of $\mathbb{R}^p$.

assum(a) The data $W_i\coloneqq(X_i,Z_i)$, $i=1,\ldots, n$ are i.i.d. according to $P_*$ and \\ (b) The pseudo-true value ${\Greekmath 0112}_\circ\in int(\Theta)$ is the unique maximizer of $${\Greekmath 0115}_\circ({\Greekmath 0112})'\mathbf{E}^P[g(W,{\Greekmath 0112})] - \log\mathbf{E}^P[\exp\{{\Greekmath 0115}_\circ({\Greekmath 0112})'g(W,{\Greekmath 0112})\}],$$ where $\Theta\subset\mathbb{R}^p$;\% is compact;\\ (c) ${\Greekmath 011A}(X,{\Greekmath 0112})$ is continuous at each ${\Greekmath 0112}\in\Theta$ with probability one;\\ (d) ${\Greekmath 011A}(X,{\Greekmath 0112})$ is twice continuously differentiable in a neighborhood $\mathcal{U}$ of ${\Greekmath 0112}_\circ$ with probability one and for ${\Greekmath 0114} = 0,1$, $\mathbf{E}^P\left[\sup_{{\Greekmath 0112}\in\mathcal{U}}\left|\frac{dQ^*({\Greekmath 0112})}{dP_*}(W)\right|^{{\Greekmath 0114}}\|{\Greekmath 011A}_{j{\Greekmath 0112}{\Greekmath 0112}}(X,{\Greekmath 0112})\|^2\|q^K(Z)\|^2\right] = \mathcal{O}(K)$, $j=1,\ldots d$;\%, is bounded;\\ (e) for a neighborhood $\mathcal{U}$ of ${\Greekmath 0112}_\circ$ and for ${\Greekmath 0114} =0,1,2$, $j=2,4$ it holds that $$\mathbf{E}^P\left[\sup_{{\Greekmath 0112}\in\mathcal{U}}\left|\frac{dQ^*({\Greekmath 0112})}{dP_*}(W_i)\right|^{{\Greekmath 0114}}\|{\Greekmath 011A}(X,{\Greekmath 0112})\|^j\|q^K(Z)\|^j\right] = \mathcal{O}({\Greekmath 0110}(K)^{j-2} K),$$ where ${\Greekmath 0110}(K)$ is as defined in Assumption (ref);\\ (f) for a neighborhood $\mathcal{U}$ of ${\Greekmath 0112}_\circ$ and for ${\Greekmath 0114} = 0,1,2$, $j=1,2,4$ it holds that $$\mathbf{E}^P\left[\sup_{{\Greekmath 0112}\in\mathcal{U}}\left|\frac{dQ^*({\Greekmath 0112})}{dP_*}(W_i)\right|^{{\Greekmath 0114}}\|{\Greekmath 011A}_{{\Greekmath 0112}}(X_i,{\Greekmath 0112})\|^j\|q^K(Z)\|^j\right] = \mathcal{O}({\Greekmath 0110}(K)^{\max\{j-2,0\}} K),$$ where ${\Greekmath 0110}(K)$ is as defined in Assumption (ref);\\ (g) the matrix $\mathbf{E}^{Q^*({\Greekmath 0112}_\circ)}\left[\left.{\Greekmath 011A}(X,{\Greekmath 0112}_\circ){\Greekmath 011A}(X,{\Greekmath 0112}_\circ)'\right|Z\right]$ has smallest (resp. largest) eigenvalue bounded away from zero (resp. infinity);\\ (h) for a neighborhood $\mathcal{U}$ of ${\Greekmath 0112}_\circ$ it holds that $\mathbf{E}^P\left[\sup_{{\Greekmath 0112}\in\mathcal{U}}\left|\frac{dQ^*({\Greekmath 0112})}{dP_*}(W_i)\right|^{2}\right]$ is bounded.

Assumption (ref) (b) guarantees uniqueness of the pseudo-true value and is a standard assumption in the literature on misspecified models (see e.g. White1984). Assumptions (ref) (d)-(f) are the counterparts of Assumptions (ref) (b) and (ref) (b), respectively, for the misspecified case. It is important to notice that they implicitly contain the first part of Assumption (ref). The reason why we cannot separate the part involving the moment function ${\Greekmath 011A}(X,{\Greekmath 0112})$ (or its derivative) and the one involving $q^K(Z)$ in the assumption, as we do for the correctly specified model, is that the $P_*$-density of $Q_*({\Greekmath 0112})$ cannot be factorized in a conditional density of $X$ given $(Z,{\Greekmath 0112})$ and a marginal density of $Z$ independent of ${\Greekmath 0112}$. In particular, in the misspecified case the pseudo-true value of the tilting parameter ${\Greekmath 0115}_\circ({\Greekmath 0112}_\circ)$ is not equal to zero as it is the tilting parameter in the correctly specified case. Assumption (ref) (g) is the counterpart of Assumption (ref) (b) for the misspecified case.

The BvM theorem for misspecified models now follows. Let $G_i({\Greekmath 0112}_\circ) \coloneqq {\Greekmath 011A}_{{\Greekmath 0112}}(X_i,{\Greekmath 0112}_\circ)\otimes q^K(Z_i)$, $D^{\dagger}_\circ(Z) \coloneqq \mathbf{E}^{Q^*({\Greekmath 0112}_\circ)}[{\Greekmath 011A}_{{\Greekmath 0112}}(X,{\Greekmath 0112}_\circ)|Z]$ and $\Sigma_{\circ}(Z)\coloneqq\mathbf{E}^{Q^*({\Greekmath 0112}_\circ)}[{\Greekmath 011A}(X,{\Greekmath 0112}_\circ){\Greekmath 011A}(X,{\Greekmath 0112}_\circ)'|Z]$. Moreover, let $\mathcal{H}$ denote a compact subset of $\mathbb{R}^p$ and ${\Greekmath 0112}_h \coloneqq {\Greekmath 0112}_\circ + h/\sqrt{n}$, with $h\in\mathcal{H}$.

thm[Bernstein-von Mises (misspecified)] Let Assumptions (ref), (ref), (ref) - (ref) hold. Assume that there exists a constant $C>0$ such that for any sequence $M_n\rightarrow \infty$, \begin{equation} P_*\left(\sup_{\|{\Greekmath 0112} - {\Greekmath 0112}_\circ\| > M_n/\sqrt{n}}\frac{1}{n}\sum_{i=1}^n \left(\ell_{n,{\Greekmath 0112}}(W_i) - \ell_{n,{\Greekmath 0112}_\circ}(W_i)\right) \leq - CM_n^2/n\right) \rightarrow 1, \end{equation} as $n\rightarrow \infty$. If $K\rightarrow \infty$ and $\sup_{{\Greekmath 0112}\in\mathcal{U}}\|{\Greekmath 0115}_\circ({\Greekmath 0112})\|^2\max\{{\Greekmath 0110}(K),K\}K\sqrt{K/n} \rightarrow 0$ then, the posteriors converge in total variation towards a Normal distribution, that is, \begin{equation} \sup_{B}\left|{\Greekmath 0119}(\sqrt{n}({\Greekmath 0112}-{\Greekmath 0112}_\circ)\in B|W_{1:n},K) - \mathcal{N}_{\Delta_{n,{\Greekmath 0112}_\circ},V_{{\Greekmath 0112}_\circ}}(B)\right|\overset{p}{\to} 0, \end{equation} where $B\subseteq \Theta$ is any Borel set, $\Delta_{n,{\Greekmath 0112}_\circ}$ is a random vector bounded in probability and $V_{{\Greekmath 0112}_\circ}$ is a positive definite matrix equal to the inverse of: { \begin{multline*} V_{{\Greekmath 0112}_\circ}^{-1} = \mathbf{E}^P\left[D^{\dagger}_\circ(Z)\Sigma_{\circ}(Z)^{-1}D_\circ(Z)\right] + \mathbf{E}^{Q^*({\Greekmath 0112}_\circ)} \left[G_i({\Greekmath 0112}_\circ)'{\Greekmath 0115}_\circ({\Greekmath 0112}_\circ) g(W_i,{\Greekmath 0112}_\circ)'\right]\Omega_\circ^{-1}\mathbf{E}^P[G_i({\Greekmath 0112}_\circ)]\\ - \frac{d{\Greekmath 0115}_{\circ}({\Greekmath 0112}_\circ)'}{d{\Greekmath 0112}}\left(\mathbf{E}^{P}\left[G_i({\Greekmath 0112}_{\circ})\right] - \mathbf{E}^{Q^*({\Greekmath 0112}_\circ)}\left[G_i({\Greekmath 0112}_{\circ})\right]\right) - \sum_{j=1}^{d} \frac{d^2{\Greekmath 0115}_{\circ,j}({\Greekmath 0112}_\circ)'}{d{\Greekmath 0112} d{\Greekmath 0112}'}\mathbf{E}^P[{\Greekmath 011A}_{j}(X_i,{\Greekmath 0112}_\circ)q^K(Z)]\\ - \sum_{j=1}^{d} \left(\mathbf{E}^P\left[{\Greekmath 011A}_{j{\Greekmath 0112}{\Greekmath 0112}}(X_i,{\Greekmath 0112}_\circ)q^K(Z_i)\right] - \mathbf{E}^{Q^*({\Greekmath 0112}_\circ)}\left[{\Greekmath 011A}_{j{\Greekmath 0112}{\Greekmath 0112}}(X_i,{\Greekmath 0112}_\circ)q^K(Z_i)\right]\right) {\Greekmath 0115}_{\circ,j}({\Greekmath 0112}_\circ)\\ + Var_{Q^*({\Greekmath 0112}_\circ)}[G_i({\Greekmath 0112}_\circ)'{\Greekmath 0115}_\circ({\Greekmath 0112}_\circ)] + \frac{d{\Greekmath 0115}_\circ({\Greekmath 0112}_\circ)'}{d{\Greekmath 0112}}Cov_{Q^*({\Greekmath 0112}_\circ)}\left(g(W_i,{\Greekmath 0112}_\circ),G_i({\Greekmath 0112}_\circ)'{\Greekmath 0115}_\circ({\Greekmath 0112}_\circ)\right). \end{multline*} }

Just as in kleijn2012, this theorem establishes that the posterior distribution of the centered and scaled parameter $\sqrt{n}({\Greekmath 0112}-{\Greekmath 0112}_\circ)$ converges to a Normal distribution with a random mean that is bounded in probability. The rate restriction $\sup_{{\Greekmath 0112}\in\mathcal{U}}\|{\Greekmath 0115}_\circ({\Greekmath 0112})\|^2\max\{{\Greekmath 0110}(K),K\}K\sqrt{K/n} \rightarrow 0$ is much stronger than the one in Theorem (ref). The slower rate condition on $K$ is intuitive. When the conditional moment conditions are misspecified, limiting the number of implied unconditional moments serves to limit the magnification of the misspecification. In the (completely) hypothetical situation where one knew that the conditional moment conditions are misspecified, one would either discard the misspecified moment conditions or take a small and fixed number of expanded moment conditions (for instance with $K=1$). In practice, of course, this strategy cannot be implemented (because one does not know whether the model is correctly specified or misspecified) and, therefore, $K$ must always go to $\infty$, but slower than under correct specification.

An additional remark is that, because of misspecification, the covariance matrix $V_{{\Greekmath 0112}_\circ}$ appearing in Theorem (ref) is expected to be different from the asymptotic covariance matrix of the corresponding frequentist point estimator in AI2007. Therefore, while correctly centered, the $(1 - {\Greekmath 010B})$ posterior credible sets are not in general $(1 - {\Greekmath 010B})$ confidence sets. The coverage rate can be less, or more, than the nominal level depending on the true data generating process and the extent of misspecification.

The strategy of the proof of this result is generally similar to the proof of Theorem (ref) with ${\Greekmath 0112}_*$ replaced by the pseudo-true value ${\Greekmath 0112}_\circ$. However, proving that the ETEL function satisfies a stochastic LAN expansion is more complex, for the following reasons. First, the limit of $\widehat{\Greekmath 0115}({\Greekmath 0112}_\circ)$ is ${\Greekmath 0115}_\circ({\Greekmath 0112}_\circ)$, which is not zero. Therefore, several terms that were equal to zero in the LAN expansion under correct specification, are non-zero in the misspecified case. The limit in distribution of these terms has to be derived. This explains our stronger assumptions with respect to the correctly specified case. Second, the quantity $\frac{1}{\sqrt{n}}\sum_{i=1}^n g(W_i,{\Greekmath 0112}_\circ)$ is no longer centered on zero, which leads to an additional bias term.

\begin{contexample1}

Consider now the impact of $K$ in misspecified models. For simplicity, we fix ${\Greekmath 0112}_{0}$ to a certain value and we only estimate ${\Greekmath 0112}_{1}$. We generate the sample data as above with $({\Greekmath 0112}_{0}=1,{\Greekmath 0112}_1 = 1)$ and then estimate ${\Greekmath 0112}_1$, fixing ${\Greekmath 0112}_{0}$ at 0.5. Therefore, the moment condition $\mathbf{E}^{P}[{\Greekmath 0122} |Z]=0$ is misspecified. The pseudo-true value of ${\Greekmath 0112}_{1}$ is obtained from a side calculation. More specifically, we generate 5 million observations from the true model. We assume that this sample size is large enough to represent the $n \rightarrow \infty$ situation. Then, we misspecify the CM conditions by setting ${\Greekmath 0112}_{0}=0.5$, construct the expanded moments with these 5 millions observations setting $K = 26$, and numerically find the value of ${\Greekmath 0112}_{1}$ that maximizes the BETEL posterior distribution. The pseudo true value is this maximized value. It is 1.004. The adverse impact of increasing $K$ (for a fixed $n$) on the Bayesian bias and posterior standard deviation is reported in Table (ref). As in the case of the correctly specified models, relatively small values of $K$ (around $K=5$ for $n=250$) lead to the best results and, in addition, the value of the Bayesian bias increase when $K$ increases beyond the recommended value. The posterior sd declines more sharply as K increases. As pointed out before, these effects are due to the reduction in the support of the prior, now magnified by moment misspecification. Finally, we calculate frequentist coverage rate of the equal-tailed 90% credible set based on 100 repetitions for both the correctly specified model (${\Greekmath 0112}_0$ now set equal to 1) and misspecified model with $K\approx2n^{1/6}$. Consistent with our theory the BETEL credible set is different from the frequentist confidence set when conditional moment conditions are misspecified. When $n=250$ the coverage rates are 91% for the correctly specified case and 86% for the misspecified case. When the number of observations increases to 1,000, the coverage rate for the misspecified case further moves down to 80% while the coverage rate for the correctly specified case remains around 90%.

table[table omitted — 685 chars of source]

\end{contexample1}

Model Comparisons

In practice, we can be unsure about elements of the conditional moment model. For instance, we can be faced with a large number of variables in $Z$, but only some of which are relevant. In such cases, any specific model may be considered to be misspecified, and the goal is to find the best model given the data.

Let $M_{\ell }$ denote the $\ell $th model in the model space. Each model is characterized by a parameter ${\Greekmath 0112} ^{\ell }$ and an extended set of moment functions given by $g^{\ell }(W,{\Greekmath 0112} ^{\ell })$. In addition, each model $M_{\ell }$ is described by a prior distribution for ${\Greekmath 0112} ^{\ell }$. The posterior distribution is obtained based on (ref). The aim is to compare these models by marginal likelihoods, denoted by $m(W_{1:n}|M_{\ell},K)$. These are each calculated by the marginal likelihood identity of Chib1995 (where we explicit the dependence on $M_{\ell}$ in the notation):

equation[equation omitted — 275 chars of source]

and by the method of ChibJeliazkov2001. In this expression, $\tilde{{\Greekmath 0112}}^{\ell }$ is any point in the support of the posterior (such as the posterior mean).

remComparison of CM condition models differs in one important aspect from the framework for comparing unconditional moment condition models that was established in \citet*{ChibShinSimoni2018}, where it is shown that to make the unconditional moment condition models comparable it is necessary to linearly transform the moment functions so that all the transformed moments are included in each model. This linear transformation consists of adding an extra parameter different from zero to the components of the vector $g({\Greekmath 0112} ,W)$ that correspond to the restrictions not included in a specific model. When comparing conditional moment models, however, this transformation is not necessary because the convex hulls associated with different expanded models have the same dimension asymptotically.

Model selection consistency

Suppose that there are $J$ contending models. Suppose also that at least $J-1$ of these models are misspecified and the remaining one can be either misspecified or correctly specified. Moreover, suppose that a model $M_\ell$ is selected by the size of the marginal likelihoods. Then, in Theorem (ref) we show that this criterion in the limit picks the model $M_\ell$ with the smallest Kullback-Leibler divergence between $P_*$ and the corresponding $Q^*({\Greekmath 0112}^\ell)$, where $Q^*({\Greekmath 0112}^\ell) = \mathrm{arginf}_{Q\in\mathcal{P}_{{\Greekmath 0112}^{\ell}}}\mathbb{K}(Q||P_*)$ and $\mathcal{P}_{{\Greekmath 0112}^{\ell}}$ is defined in Section (ref).

thmLet the assumptions of Theorem (ref) hold. Let us consider the comparison of $J<\infty$ models $M_\ell$, $\ell=1,\ldots,J$, such that $J-1$ of these models each has at least one misspecified moment condition and model $M_j$ can be either correctly specified or contain some misspecified moment condition, that is, $M_{\ell}$ does not satisfy Assumption (ref) (a), $\forall \ell \neq j$. Then, \begin{equation*} \lim_{n\rightarrow \infty}P_*\left(\log m(W_{1:n}|M_{j},K) > \max_{\ell\neq j}\log m(W_{1:n}|M_{\ell}.K)\right) = 1 \end{equation*} if and only if $\mathbb{K}(P_*||Q^*({\Greekmath 0112}_{\circ}^j))< \min_{\ell\neq j}\mathbb{K}(P_*||Q^*({\Greekmath 0112}_{\circ}^{\ell}))$, where $\mathbb{K}(P||Q):=\int \log(dP/dQ)dP$.

Note that if one model in the contending set of models is correctly specified, then this model will have zero Kullback-Leibler divergence and, therefore, according to Theorem (ref), that model will have the largest marginal likelihood and will be selected by our procedure.

To understand the ramifications of the preceding result, suppose that we are interested in comparing models with the same moment conditions but different conditioning variables:

equation[equation omitted — 343 chars of source]

where $Z_{1}$ and $Z_{2}$ may have some elements in common, in particular $ Z_{2}$ might be a subvector of $Z_{1}$ (or vice versa). A situation of this type, where we are unsure about the validity of instrumental variables, is the following.

\begin{example2} (Comparing IV models) Consider the following model with three instruments $ \left( Z_{1},Z_{2},Z_{3}\right) $:

equation*[equation* omitted — 293 chars of source]

where $(e_{1},e_{2})^{\prime }$ are non-Gaussian and correlated. Thus, $X$ in the outcome model is correlated with the error $e_{1}$. Let true ${\Greekmath 0112} =\left( {\Greekmath 0112} _{0},{\Greekmath 0112} _{1}\right) $ equal $(1,1)$. Moreover, suppose that the $Z_{j}$'s are relevant instruments, that is, $cov(X,Z_{j})\neq 0$ for $j\leq 3$, and

equation[equation omitted — 150 chars of source]

In practice, some instruments can be valid and some not, and the goal is to select the valid instruments. To this end, we generate $(e_{1},e_{2},Z_{1})$ from a Gaussian copula whose covariance matrix is $\Sigma =[1,0.7,0.7;0.7,1,0;0.7,0,1]$ such that the marginal distribution of $e_{1}$ is the skewed mixture of two normal distributions $0.5\mathcal{N}(0.5,0.5^{2})+0.5\mathcal{N}(-0.5,1.118^{2})$ and the marginal distribution of $e_{2}$ is $\mathcal{N}(0,1)$. Under this setup, $Z_{1}$ is an invalid instrument. Consider the following three models

align[align omitted — 324 chars of source]

Because $Z_{1}$ is an invalid instrument, models $\mathcal{M}_{1}$ and $ \mathcal{M}_{2}$ are misspecified.

In $\mathcal{M}_{1}$, the basis matrix $B$ is made from the variables $ \left( \boldsymbol{z}_{1},\boldsymbol{z}_{2},\boldsymbol{z}_{1}\odot \boldsymbol{z}_{2},\boldsymbol{z}_{1} \odot\boldsymbol{z}_{3},\boldsymbol{z} _{2}\odot\boldsymbol{z}_{3}\right)$, each using $K$ knots, concatenated with the vector $\boldsymbol{z}_{3}$. In $\mathcal{M}_{2}$, $B$ is made from the variables $\left( \boldsymbol{z}_{1},\boldsymbol{z}_{1} \odot\boldsymbol{z}_{3}\right)$, each using $K$ knots, concatenated with the vector $\boldsymbol{z}_{3}$. In $\mathcal{M}_{3}$, $B$ is made from the variables $\left( \boldsymbol{z}_{2},\boldsymbol{z}_{2} \odot\boldsymbol{z}_{3}\right)$, each using $K$ knots, concatenated with the vector $\boldsymbol{z}_{3}$. The number of columns in the $B$ matrix is $ 5(K-1)+2$ for $\mathcal{M}_{1}$, and $2(K-1)+2$ for $\mathcal{M}_{2}$ and $\mathcal{M}_{3}$. The prior for ${\Greekmath 0112} _{0}$ and ${\Greekmath 0112} _{1}$ is the product of student-$t$ distributions with mean zero, dispersion 5, and degrees of freedom equal to 2.5. A repeated sampling experiment is conducted. The marginal likelihood of each model is calculated in 200 repeated samples. Table (ref) reports the model selection frequency for $(n=100,K =4)$, $(n=250,K=5)$, and $(n=1000,K=6)$, where $K$ is based on $K = 2 n^{1/6}$. Note that the model with the valid instruments, i.e., $\mathcal{M}_{3}$, is selected more frequently as the number of observation gets larger, in conformity with the theory.

table[table omitted — 649 chars of source]

\end{example2}

Additional topics

High dimensional $Z$

We now consider the case where $Z$ lies in a high-dimensional space. If all the elements of $Z$ are relevant, then the situation can be challenging, but there is an interesting sub-case that is worth discussing. Suppose that the conditional expectation depends only on a few elements of $Z$ or, in other words, most of the elements of $Z$ are redundant. In this case, one can find the relevant elements of $Z$ by estimating and comparing models that condition on different subsets of $Z$, where the cardinality of these subsets is say 2 or 3. The relevant elements of $Z$ correspond to the model with the largest marginal likelihood. We refer to this procedure as sparsity-based model selection. The next example provides an illustration.

\begin{example3} (Sparsity-based model selection) Recall our Example 2, but assume that one has nine additional potential $Z$'s

equation*[equation* omitted — 133 chars of source]

for $j=4,5,...,12$. Recall, $Z_{1}$ is an invalid instrument. Therefore, $ Z_{j}$'s for $j=4,...,12$ are also invalid. Suppose that $Z_{3}$ affects the conditional expectation, but that one is unsure about the remaining elements of $Z$. Suppose one believes that at most three elements of $Z$ affect the conditional expectation (the sparsity assumption). In this situation one can compute marginal likelihoods of the following 66 models:

equation[equation omitted — 136 chars of source]

and

equation[equation omitted — 226 chars of source]

with the correct model given by

equation[equation omitted — 96 chars of source]

Sample data (size $n=250$) is generated from the design in Example 2. Estimation and marginal likelihood computations are based on expanded moments from $K=3$ basis functions for each conditioning $Z_j$. A summary of the marginal likelihood results appears in Figure (ref), sorted by the size of the marginal likelihood. The top ranked model is the true model. As shown in the right panel of the same figure, which reports the posterior model probabilities (under a uniform prior on model space), the support for the true model is decisive.

figure[figure omitted — 948 chars of source]

\end{example3}

High dimensional $\protect{\Greekmath 0112}$: TaRB-MH

It is also important to consider the estimation of conditional moment models that contain a high-dimensional ${\Greekmath 0112}$. While the regularizing role of the Bayesian prior is important, MCMC sampling of the posterior simulation becomes more complicated. From our experience, the one-block M-H algorithm that we have used above tends to be inefficient. A straightforward alternative sampling scheme is offered by the Tailored Randomized Block MH algorithm of ChibRamamurthy2010. This algorithm, which has proved useful in several similarly complex settings, trades more computations for gains in simulation efficiency.

\begin{contexample4} (IV regression with additional exogenous regressors). Consider the previous IV regression model, but now with 18 additional exogenous regressors $w$

equation*[equation* omitted — 121 chars of source]

where $w_{i} = [w_{i}^{(1)'}, w_{i}^{(2)'}, w_{i}^{(3)'}]'$, and each group $w_{i}^{(j)'}$, $j \leq 3$, are identically and independently drawn from $\mathcal{N}_6(0, \Sigma({\Greekmath 011A}))$, where $\Sigma({\Greekmath 011A})$ is a $6 \times 6$ matrix in correlation form with each off-diagonal element set equal to 0.97. In addition, ${\Greekmath 010D}$ is a vector of ones. In total, there are 20 unknown parameters. Other elements of the DGP are unchanged. Suppose one has 1500 observations from this DGP, and we estimate $[{\Greekmath 0112}_{0}, {\Greekmath 0112}_{1}, {\Greekmath 010D}']'$ from $\mathbf{E}^{P}[(Y-{\Greekmath 0112} _{0}-{\Greekmath 0112} _{1}X)|Z_{2},Z_{3}, W]=0$ with the expanded moment conditions similar to $\mathcal{M}_{3}$ in Example (ref). The basis function matrix is formed with $Z_{2}$, $Z_{3}$, and $Z_{2} Z_{3}$, concatenated with columns in $W$. We set $K = 6$, following our recommendation, which leads to 36 expanded moment conditions. The training sample prior is based on the first 10% of the sample, and estimation on the remaining 90%. The prior is a product of independent student-$t$ distributions with 5 degrees of freedom, centered on the two-stage least squares (2SLS) estimate, and scale equal to two times the 2SLS standard error.

table[table omitted — 2,330 chars of source]

\end{contexample4}

The results appear in Table (ref). In implementing the TaRB-MH sampling scheme, the probability of starting a new block is set to 0.3, so that the each block, within each MCMC iteration, contains 6 parameters on average. For comparison, results from the single-block sampling scheme (on the same conditional moments) are also included. It is evident that the two MCMC samplers produce identical posterior moments, but that the TaRB-MH sampler dominates the one-block MH sampler in terms of simulation efficiency as measured by the inefficiency factor (the ratio of the numerical variance of the mean to the variance of the mean assuming independent draws). An inefficiency factor close to 1 indicates that the MCMC draws are essentially independent. Therefore, armed with the TaRB-MH sampler, computational efficiency is retained, even in higher-dimensional ${\Greekmath 0112}$ problems.

Applications

Asset pricing

A key question in finance concerns the makeup of the pricing kernel, or the stochastic discount factor (SDF). Factors in the SDF are the risk factors that explain the cross-section of expected equity returns and, for this reason, establishing the identity of these risk factors has been a long-standing quest in finance.

Following notation from ChibZeng2020, write the SDF $M_{t}$ at time (month) $t$ as

equation[equation omitted — 65 chars of source]

where $x_{t}$ is a $(k_{x}\times 1)$ vector of risk factors (empirically these are the excess returns on portfolios of stocks), and $b$ is the unknown risk-factor premia and ${\Greekmath 0116} _{x}=$ $\mathbf{E}^{P}(x_{t})$.\ The parameters $(b,{\Greekmath 0116} _{x})$ are unknown. Suppose that there are other factors (excess returns on other portfolios) that are collected in a $(k_{w}\times 1)$ vector $w_{t}$. Let $f_{t} \coloneqq(x_{t}^{\prime },w_{t}^{\prime})^{\prime }$ be a $(k_{f}\times 1)$-vector, where $k_{f}=k_{x}+k_{w}$. If $x_{t}$ are risk factors, then finance theory dictates that the restriction $\mathbf{E}^{P}(M_{t}f_{t})=0$ holds. Given a sample of observations $ \{f_{t}\}_{t=1}^{n}$, one can estimate $(b,{\Greekmath 0116} _{x})$ based on the following moment conditions

equation*[equation* omitted — 162 chars of source]

where the second conditional moment restriction identifies ${\Greekmath 0116} _{x}$.

As an example, consider the data at \url{http://apps.olin.wustl.edu/faculty/chib/rpackages/czfactor/czfactor.pdf} on monthly excess returns (Jan 1974 -- Dec 2018) on $k_{f}=12$ potential risk-factors. Thus, in this situation, there are 12 conditioning variables, an illustration of a modestly high-dimensional $Z$. Let $x_{t}$ be the excess return on the market portfolio (denoted Mkt in the data).

Now construct the expanded moment conditions as

equation[equation omitted — 178 chars of source]

where $q^{K}(f_{1,t-1})$ consist of $K=3$ basis functions, and $\tilde{q} ^{K}(f_{j,t-1})$ $(j\geq 2)$ each consist of 2 basis functions derived from $ q^{K}(f_{j,t-1})$ by subtracting the second and third columns from the first and then dropping the first. Along with these 25 expanded moment conditions, the pricing conditions supply an additional twelve, for a total of 37 moment conditions.

For the prior, one can employ the training sample approach. From the first 80 observations (the training sample) the hyperparameters of the independent student-t distribution of $(b,{\Greekmath 0116} _{x})$ with 2.5 degrees of freedom are set as follows. The center of the prior density is set to the Generalized Method of Moments (GMM) estimate, and the scale to two times the GMM\ standard error. This black-box prior is particularly convenient if the analysis has to be repeated for different possible variables in the SDF. The remaining 459 observations are used to construct a joint posterior distribution of $b$ and ${\Greekmath 0116}_{x}$.

table[table omitted — 411 chars of source]

The posterior summary of $(b,{\Greekmath 0116}_{x})$ from 50,000 MCMC draws is given in Table (ref), and the prior and posterior density of $b$ are presented in Figure (ref). This summary, specifically the lower and upper limits of the marginal posterior of $b$ confirm that $b$ is non-zero, and hence that the Mkt variable is a risk-factor.

figure[figure omitted — 348 chars of source]

ATE under conditional ignorability

For another important application of the methods in this paper, consider the problem of estimating the average treatment effect (ATE) under the assumption of conditional ignorability. In the frequentist literature, the ATE is commonly estimated by propensity score methods and, on the Bayesian side, from models of the potential outcomes. These models generally have a nonparametric mean function, but parametric noise. By adopting the conditional moment perspective, however, one can evade the burden of distributional assumptions.

Data is from the 1997 Child Development Supplement to the Panel Study of Income Dynamics GuoFraser15, where the ATE is calculated by the propensity score. The research question is the effect of childhood welfare dependency on academic achievement. The latter, the dependent variable $y$, is measured by the child's score on the \textquotedblleft letter-word identification\textquotedblright\ section of the Woodcock-Johnson Revised Tests of Achievement. The treatment variable $x$ equals one if the child received AFDC (Aid to Families with Dependent Children) benefits at any time from birth to 1997 (the survey year) and equals zero if the child never received benefits during that period. It is assumed that the potential outcomes are independent of $x$, conditioned on $ z_{1}$,$z_{2},...,z_{6}$ (the assumption of conditional ignorability), where

itemize$z_{1}$: mratio97, the ratio of family income to the poverty line in 1997 • $z_{2}$: pcged97, the caregiver's years of schooling • $z_{3}$: pcg_adc, the number of years in which the caregiver received AFDC in her childhood • $z_{4}$: age97, the child's age in 1997 • $z_{5}$: race, one for African-American children and zero for other • $z_{6}$: male, one if the child is male and zero if female.

Two observations from the sample are dropped. These have values of mratio97 larger than $9$ standard deviation from the mean of mratio97. Apart from mratio97 and age97, the other variables are categorical. There are $n_{0}=727 $ control subjects and $n_{1}=274$ treated subjects. The ATE is expected to be negative, reflecting the hypothesis that welfare dependency has an adverse effect on academic achievement.

To answer the research question, suppose that the potential outcomes for the controls satisfy the conditional moments

equation*[equation* omitted — 217 chars of source]

and those for the treated satisfy the conditional moments

equation*[equation* omitted — 217 chars of source]

where $\{h_{01},h_{04},h_{11},h_{14}\}$ are four non-parametric functions. These are each modeled by natural cubic splines with 5 knots. Thus, the parameters ${\Greekmath 0112} _{j}$ of the $j$th potential outcome model consist of $ ({\Greekmath 010C} _{j0},{\Greekmath 010C} _{j2},{\Greekmath 010C} _{j3},{\Greekmath 010C} _{j5},{\Greekmath 010C} _{j6})$ plus the eight spline coefficients. Special cases of this model, mentioned below, are considered and evaluated by marginal likelihoods. For example, models in which the $h$ functions are linear are of interest.

The expanded moments are constructed as follows. The basis matrix has cubic spline basis functions for $(z_{1},z_{4})$, each with 5 knots, concatenated with $(z_{2},z_{3},z_{5},z_{6})$ (as is) because the latter variables are all essentially categorical. In total, this produces 13 expanded unconditional moments for the estimation of the $y_{0}$ and $y_{1}$ models. The prior distribution on the parameters is a product of student-t distributions with 2.5 degrees of freedom with mean of the intercept equal to the mean of the first 50 observations, the mean of the slopes equal to 0, and dispersion equal to 5.

Four models are estimated and evaluated. In the baseline model, the $h$ functions are linear. In the second model, only the effect of $z_{1}$ is nonparametric. In the third model, only the effect of $z_{4}$ is assumed to be nonparametric and, finally, in the fourth model, both $z_{1}$ and $z_{4}$ are nonparametric. The results given in Table (ref) show that the model best supported by these data is the third.

table[table omitted — 521 chars of source]

Consider now posterior inference on the ATE. By definition, the sample version of the ATE is

equation*[equation* omitted — 259 chars of source]

where, in the model selected by the preceding comparison,

equation*[equation* omitted — 271 chars of source]

Clearly, if we evaluate the latter expression at each posterior draw of $\left( {\Greekmath 0112}_{0},{\Greekmath 0112}_{1}\right)$, we produce a sample of the ATE from its posterior distribution. We summarize this sample in Table (ref) and Figure (ref).

table[table omitted — 533 chars of source]

One can see that the ATE posterior point estimate is similar in size to the propensity score estimate, but the posterior standard deviation is smaller, leading to a less dispersed interval estimate. As a takeaway, it is striking that the Bayesian analysis of this important problem can be prosecuted under such minimal assumptions.

figure[figure omitted — 377 chars of source]

Conclusion

In this paper we have developed perhaps the first Bayesian framework for analyzing an important and broad class of semiparametric models in which the distribution of the outcomes is defined only up to a set of conditional moments, some of which may be misspecified. We have derived BvM theorems for the behavior of the posterior distribution under both correct and incorrect specification of the conditional moments, and developed the theory for comparing different conditional moment models through a comparison of model marginal likelihoods. In addition, we have discussed settings with a high-dimensional $Z$ and ${\Greekmath 0112}$, the former addressed by a sparsity-based model search procedure, and the latter by the TaRB-MH MCMC algorithm for efficient posterior sampling.

Our theory and various examples, taken together, show that the developments in this paper make possible, for the first time, the formal (and practical) Bayesian analysis of a new, large class of problems that were hitherto difficult, or not possible, to tackle from the Bayesian viewpoint. This research we believe should have numerous positive ramifications for the growth and practice and teaching of Bayesian statistics.

Supplementary Materials

The supplementary materials in the online Appendix contain the technical proofs of the results in the paper, additional applications, and R code, packages and data for replicating the examples. This replication content can also be accessed from \url{https://apps.olin.wustl.edu/faculty/chib/chibshinsimoni2021replication/}. {19.0pt}

{21.0pt}

{1.5pt}