EconBase
← Back to paper

Semiparametric Bayesian Inference for a Conditional Moment Equality Model

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

77,703 characters · 21 sections · 122 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Semiparametric Bayesian Inference for a Conditional Moment Equality Model

abstractI propose a semiparametric Bayesian inference framework for conditional moment equalities. The core idea is that these models deterministically map a conditional distribution of data to a structural parameter via the restriction that a conditional expectation equals zero. Consequently, a posterior for the conditional distribution leads to a posterior for the structural parameter by minimizing the distance of the conditional moments to zero. The method has similar flexibility to frequentist semiparametric estimators and does not require converting the conditional moments into unconditional moments. I also establish frequentist asymptotic optimality of my proposal via a semiparametric Bernstein-von Mises theorem (BvM), which establishes that the posterior for the structural parameter is asymptotically normal and matches the chamberlain1987asymptotic semiparametric efficiency bound. The BvM conditions are verified for Gaussian process priors and complement the numerical aspects of the paper in which these priors are used to estimate welfare effects. \\ Keywords: Bayesian Semiparametrics, Bernstein-von Mises, Optimal Instruments

Introduction

Overview

Many economic parameters are identified via conditional moment equalities. These restrictions state that a finite-dimensional parameter of interest $\gamma $ uniquely solves $E_{P_{W|Z}}[g(W,Z,\gamma)|Z] = 0$, where $W $ and $Z $ are observed random vectors, $P_{W|Z}$ is the conditional distribution of $W $ given $Z$, and $g$ is a known vector-valued function. A canonical example is linear instrumental variables (IV) in which the regression model $Y=\gamma_{0}+\gamma_{1}D+\gamma_{2}X+U$ with $E[U|X,A] = 0$ leads to $E_{P_{W|Z}}[g(W,Z,\gamma)|Z]=0$ with $W=(Y,D)'$, $Z = (X,A)' $, and $g(W,Z,\gamma) = Y-\gamma_{0}-\gamma_{1}D-\gamma_{2}X$. Other examples include Euler equations hansen1982generalized, demand and supply models berry1995automobile, and peer effects models graham2008identifying, to name a few. Conditional moment equalities are also theoretically interesting because they impose overidentifying restrictions on $P_{W|Z}$. This makes efficient estimation nontrivial and has motivated important work on estimation using the chamberlain1987asymptotic optimal instruments newey1990efficient,NEWEY1993419,donald2001choosing,donald2003empirical,kitamura2004empirical,donald2009choosing,chen2021harmless,chib2022bayesian.

This paper makes three contributions. The first is introducing a semiparametric Bayesian inference framework for conditional moment equalities. Consistent with some approaches to conditional moment equalities (e.g., ai2003efficient), I leave $P_{W|Z}$ unrestricted and redefine $\gamma$ as a minimum distance estimand $\operatorname*{arg\,min}_{\gamma \in \Gamma}||E_{P_{W|Z}}[g(W,Z,\gamma)|Z]||^{2}$, where $||\cdot||$ is a norm. Observing that this defines a deterministic relationship between $P_{W|Z}$ and $\gamma$, a Bayesian framework arises in which a posterior for $P_{W|Z}$ implies a posterior for $\gamma$. The prior for $P_{W|Z}$ is supported over a nonparametric family, a feature ensuring that this Bayesian approach has similar flexibility to frequentist semiparametric estimators that do not rely on a tightly parametrized likelihood model. Moreover, by targeting $P_{W|Z}$, the posterior for $\gamma$ does not require converting the conditional moments into unconditional moments. This is appealing because ad hoc selections of unconditional moments can result in efficiency loss, and, more severely, some (efficient) choices can result in $\gamma$ not being identifiable.\footnote{See Example 2 of dominguez2004consistent for an instance where the unconditional moment equalities based on the optimal IVs have multiple roots, despite $\gamma$ uniquely minimizing the distance of the conditional moments to zero.} These benefits are compounded by the fact that, as a fully Bayesian procedure, it inherits all the basic benefits of Bayesian analysis (e.g., likelihood principle, expected utility decisions, simultaneous inference).

The second contribution is a semiparametric Bernstein-von Mises theorem (BvM) for the posterior for $\gamma$. BvMs provide conditions under which a posterior has the same limiting distribution as an efficient frequentist estimator, and are often invoked as a claim of large-sample robustness of the posterior to the choice of prior. Although BvMs hold under mild conditions in correctly specified regular parametric models, several papers find that these theorems can be delicate in semiparametric applications where the parameter of interest is a functional of a smooth nonparametric object, such as a probability density function, a regression function, etc. (see bickel2012semiparametric,castillo2012gaussian,castillo2012semiparametric,rivoirard2012bernstein,castillo2015bernstein,ray2020semiparametric). Since $\gamma$ is a functional of a nonparametric $P_{W|Z}$, these findings raise the possibility that the posterior for $\gamma$ may be sensitive to the choice of prior for $P_{W|Z}$ in large samples. To resolve these concerns, I prove a general semiparametric BvM for the posterior for $\gamma$ in Section (ref), providing conditions under which it converges to a normal distribution, centered at an efficient estimator, and with variance equal to the chamberlain1987asymptotic efficiency bound. As far as I am aware, this is the first fully Bayesian semiparametric BvM for the conditional moment equality model that does not require explicit conversion into unconditional moments.\footnote{See kato2013quasi and kankanala2025generalized for quasi-Bayesian results of a similar nature.} A byproduct of the BvM is that some Bayes estimators are asymptotically efficient frequentist estimators.

The third contribution is providing implementation details and verifying the BvM conditions for a flexible class of priors based on Gaussian processes (GP). GPs induce priors for conditional densities supported on the smoothness classes typically imposed on first-stage nuisance parameters, making concrete the flexibility of this Bayesian approach. Section (ref) demonstrates implementation for these priors with an empirical illustration focused on estimating welfare effects, an exercise capturing a common use of conditional moments (i.e., estimating model primitives and evaluating counterfactuals).\footnote{A complementary data-calibrated simulation comparing my method with alternatives is in Appendix (ref).} Section (ref) complements these numerical aspects by formally verifying the high-level BvM conditions for these GP priors. The key requirement is that the chamberlain1987asymptotic optimal residual is well-approximated by a reproducing kernel Hilbert space (RKHS), a `no-bias' condition that restricts the smoothness of the optimal residual relative to the GP.

Literature

This paper connects to several literatures in econometrics and statistics. The first is the Bayesian analysis of moment conditions. chamberlain2003nonparametric perform inference for structural parameters defined by unconditional moment equalities by solving the moment restrictions using draws from a Dirichlet process posterior for the data distribution. The key conceptual distinction between their approach and mine is that the former cannot be applied to conditional moment equalities with continuous $Z$ unless the conditional moments are converted into unconditional moments.\footnote{In an unpublished manuscript, chamberlain1995semiparametric consider conditional moment equalities, however their Dirichlet priors explicitly require the support of $Z$ to be small relative to the sample size, ruling out continuous $Z$.} chib2022bayesian is also relevant because they propose a Bayesian inference procedure for conditional moment equalities in which the conditional moments are converted into a sieve of unconditional moments, and inference is performed using the Bayesian exponentially tilted empirical likelihood (ETEL) framework of schennach2005bayesian. This conversion (and the use of ETEL) makes chib2022bayesian fundamentally different to my proposal. Other papers on Bayesian analysis of moments include lancaster2010bayesian, kitamura2011bayesian, pelenis2014bayesian, shin2015bayesian, bornn2018moment, and chib2018bayesian.

Another related literature is on semiparametric BvMs. castillo2015bernstein prove a general BvM for smooth functionals of nonparametric models, with substantial analysis devoted to those of unconditional probability density functions (building on rivoirard2012bernstein). I characterize $\gamma$ as a smooth functional of a conditional probability density function, and, via conditional asymptotics, I am able to modify their proof to show that the $\gamma$ posterior achieves the chamberlain1987asymptotic limit. chib2022bayesian also prove a BvM compatible with chamberlain1987asymptotic, however their proof differs from mine because, under appropriate conditions, the sieve of unconditional moment equalities makes parametric BvM arguments (cf. Theorem 10.1 of vaart_1998) applicable to the ETEL. Other contributions to the semiparametric BvM literature include shen2002asymptotic, bickel2012semiparametric, castillo2012gaussian,castillo2012semiparametric, norets2015bayesian, ray2020semiparametric, monard2021bernstein, breunig2025double,breunig2025semiparametricbayesiandifferenceindifferences, yiu2025semiparametric, and walker2024parametrization.

A third literature is efficient frequentist estimation of conditional moment equalities. The posterior for $\gamma$ is formalized as that of a conditional estimand that solves an efficiently weighted minimum distance problem. For this reason, my approach offers a Bayesian analog to the ai2003efficient framework (except I focus on finite-dimensional parameters). Moreover, the first-order conditions of the minimum distance estimand relates my framework to method of moments estimators that plug in estimators of the chamberlain1987asymptotic optimal instruments (e.g., those studied in robinson1987heteroskedasticity, newey1990efficient, NEWEY1993419, and chen2021harmless). Other frequentist approaches to optimal instruments estimation include Generalized Method of Moments (GMM) and Generalized Empirical Likelihood (EL) estimators based on sieves of unconditional moments donald2001choosing,donald2003empirical,donald2009choosing, and localized EL estimators ,kitamura2004empirical.

A final related literature is quasi-Bayes. Popularized in chernozhukov2003mcmc, quasi-Bayes is a sampling-based frequentist estimation framework in which an extremum objective forms a log quasi-likelihood, a prior is assumed for the parameter, and Bayes rule is used to obtain a quasi-posterior. kato2013quasi and kankanala2025generalized develop quasi-Bayesian frameworks for conditional moment equalities, and prove quasi-Bayesian BvMs for a structural function defined by conditional moment equalities and linear functionals of it, respectively. A key conceptual distinction between my approach and quasi-Bayes is that mine has a finite-sample Bayesian interpretation. Other papers on quasi-Bayesian estimation include liao2011posterior, florens2012nonparametric,florens2021gaussian, chen2018monte, andrews2022optimal, and kankanala2023quasi.

Outline

The paper is organized as follows. Section (ref) sets up the general Bayesian inference framework and provides a flexible class of priors, Section (ref) presents an empirical illustration, Section (ref) states the semiparametric BvM and related results, Section (ref) verifies the BvM conditions for the GP priors, and Section (ref) concludes with some extensions. The Appendix contains proofs of theorems and corollaries, technical details about implementation, a simulation study calibrated to the empirical application, and further discussion. Proofs of propositions, and lemmas, as well as some additional discussion are in the Online Appendix walker2026supplement. Notation is introduced when appropriate.

Bayesian Inference Framework

This section formalizes semiparametric Bayesian inference and describes implementation for a class of priors based on Gaussian processes.

Semiparametric Bayesian Inference

I start with Bayesian inference for the conditional distribution $P_{W|Z}$ of $W$ given $Z$. Let $\{(W_{i}',Z_{i}')'\}_{i \geq 1}$, $W_{i} \in \mathcal{W} \subseteq \mathbb{R}^{d_{w}}$ and $Z_{i} \in \mathcal{Z} \subseteq \mathbb{R}^{d_{z}}$, be a sequence of random vectors for which the first $n \geq 1$ elements $\{(W_{i}',Z_{i}')'\}_{i=1}^{n}$ forms the observed data. The sampling model conditions on the realization $\{z_{i}\}_{i \geq 1}$ of $\{Z_{i}\}_{i \geq 1}$ and assumes that the elements of $W^{(n)}= (W_{1}',...,W_{n}')'$ satisfy $W_{i}|p_{W|Z} \overset{ind}{\sim} P_{W|z_{i}}$ for $i=1,...,n$, where $p_{W|Z}$ is a conditional probability density function, and $P_{W|z_{i}}$ is the probability distribution attached to $p_{W|Z}(\cdot|z_{i})$. Since the model parameters are conditional densities $p_{W|Z}$, the prior $\Pi$ is a (conditional) probability measure defined over the space $\mathcal{P}_{W|Z}$ of conditional densities (equipped with a $\sigma$-algebra $\mathscr{P}_{W|Z}$).\footnote{The qualifier `conditional' is used because implicitly $\Pi$ also defined conditional on $\{Z_{i}\}_{i \geq 1}$.} There are no parametric restrictions on $\mathcal{P}_{W|Z}$, so $\Pi$ should be thought of as a probability measure over an infinite-dimensional space of conditional densities.

assumptionFor each $n \geq 1$ and almost every fixed realization $\{z_{i}\}_{i \geq 1}$ of $\{Z_{i}\}_{i \geq 1}$, the mapping $(w^{(n)},p_{W|Z}) \mapsto \prod_{i=1}^{n}p_{W|Z}(w_{i}|z_{i})$ is measurable on $(\mathcal{W}^{n} \times \mathcal{P}_{W|Z}, \mathscr{W}^{\otimes n} \otimes \mathscr{P}_{W|Z})$, where $\mathscr{W}$ is a $\sigma$-algebra over $\mathcal{W}$, and $\mathscr{W}^{\otimes n}$ is the corresponding product $\sigma$-algebra.
assumptionLet $p_{0,W|Z}$ be the true conditional density. For each $ n \geq 1$ and almost every fixed realization $\{z_{i}\}_{i \geq 1}$ of $\{Z_{i}\}_{i \geq 1}$, $P_{0,W|Z}^{(n)}(\int \prod_{i=1}^{n}p_{W|Z}(W_{i}|z_{i})d\Pi(p_{W|Z}) \in (0,\infty)) = 1$, where $P_{0,W|Z}^{(n)} = \bigotimes_{i=1}^{n}P_{0,W|z_{i}}$ (and $P_{0,W|z_{i}}$ is the distribution attached to $p_{0,W|Z}(\cdot|z_{i})$).

Assumptions (ref) and (ref) are standard regularity conditions that imply that the conditional distribution $\Pi(p_{W|Z}\in \cdot |W^{(n)})$ of $p_{W|Z}$ given $W^{(n)}$ (known as the posterior) has a version satisfying Bayes theorem,\footnote{See, for example, Section 1.3 of ghosal2017fundamentals for similar regularity conditions.}

align*[align* omitted — 229 chars of source]

Since $\Pi(p_{W|Z} \in \cdot | W^{(n)})$ contains an inference theory for (transformations of) $p_{W|Z}$, inference for $\gamma$ is obtained upon formalizing the estimand $\operatorname*{arg\,min}_{\gamma \in \Gamma}||E_{P_{W|Z}}[g(W,Z,\gamma)|Z]||^{2}$.

The estimand is an iterated minimum distance estimand. For a given conditional density $p_{W|Z}$, let $m(z_{i},\gamma) = E_{P_{W|Z}}[g(W,Z,\gamma)|Z =z_{i}]$ be the $d_{g}\times 1$ vector of conditional moments at $z_{i}$, let $\Sigma(z_{i},\gamma) = E_{P_{W|Z}}[g(W,Z,\gamma)g(W,Z,\gamma)'|Z=z_{i}]$ be the $d_{g}\times d_{g}$ matrix of conditional second moments at $z_{i}$, and, for $\tilde{\gamma} \in \Gamma$ fixed, define a conditional estimand $$q_{n}(\tilde{\gamma},p_{W|Z}) = \operatorname*{arg\,min}_{\gamma \in \Gamma}Q_{n}(\gamma,\tilde{\gamma},p_{W|Z}), $$ where\footnote{The terminology `conditional estimand' follows abadie2014inference.}

align[align omitted — 178 chars of source]

To eliminate dependence on $\tilde{\gamma}$, I adopt a similar technique to hansen2021inference and consider an iterated estimand defined by the recursive relation

align[align omitted — 151 chars of source]

where $\gamma_{n,0}$ is some element of $\Gamma$ (e.g., the identity-weighted estimand). Notice that $\gamma_{n}$ does not require converting the conditional moments into unconditional moments. Further, since $\gamma_{n}$ satisfies the first-order conditions $n^{-1}\sum_{i=1}^{n}M(z_{i},\gamma_{n})'\Sigma^{-1}(z_{i},\gamma_{n})m(z_{i},\gamma_{n})=0$, it implicitly plugs in an estimate $M(z_{i},\gamma_{n})' \Sigma^{-1}(z_{i},\gamma_{n})$ of the chamberlain1987asymptotic optimal instruments, where $M(z_{i},\gamma)$ is the $d_{g} \times d_{\gamma}$ Jacobian of $m(z_{i},\gamma)$ in $\gamma$. This is important for the efficiency guarantees in Section (ref).

assumptionFor almost every fixed realization $\{z_{i}\}_{i \geq 1}$ of $\{Z_{i}\}_{i \geq 1}$ and for each $n \geq 1$, the limit $\gamma_{n}$ exists for each $p_{W|Z} \in \mathcal{P}_{W|Z}$ and $p_{W|Z} \mapsto \gamma_{n}$ is $\mathscr{P}_{W|Z}$-measurable.

Assumption (ref) guarantees that $\Pi(p_{W|Z} \in \cdot |W^{(n)})$ leads to a well-defined marginal posterior for $\gamma_{n}$ given by the pushforward measure $\Pi(\gamma_{n} \in \cdot |W^{(n)}) = \Pi(p_{W|Z} \in \cdot|W^{(n)}) \circ \gamma_{n}^{-1}$.\footnote{Two points. First, the pushforward satisfies $(\Pi(p_{W|Z} \in \cdot|W^{(n)}) \circ \gamma_{n}^{-1})(A) = \Pi(p_{W|Z}: \gamma_{n} \in A|W^{(n)})$ for events $A$. Second, Appendix (ref) contains sufficient conditions for Assumption (ref) that may be of independent interest.} This formalizes the posterior for $\gamma$. Like $\Pi(p_{W|Z} \in \cdot | W^{(n)})$, it leads to inference about $\gamma_{n}$ and any function $f(\gamma_{n})$. For example, in the leading case where $f(\gamma)$ is scalar, a point estimate can be obtained using the posterior median $c_{n,f}(0.5)$, while uncertainty can be quantified using a $(1-\alpha)$-equitailed probability interval $CS_{n}(1-\alpha) = [c_{n,f}(\alpha/2),c_{n,f}(1-\alpha/2)]$ (a type of credible set), where $\alpha \in (0,1/2)$, and $c_{n,f}(q)$, $q \in (0,1)$, is the $q$-quantile of $\Pi(f(\gamma_{n}) \in \cdot | W^{(n)})$. Section (ref) establishes the asymptotic equivalence of these Bayes estimators and frequentist estimators based on the chamberlain1987asymptotic optimal instruments.

exampleRecall that linear IV sets $W = (Y,D')'$, $Z = (X',A')'$, and $g(W,Z,\gamma) = Y-\gamma_{0}-\gamma_{1}'D-\gamma_{2}'X$. In such case, $\gamma = (\gamma_{0},\gamma_{1}',\gamma_{2}')'$, $m(z,\gamma) = m_{y}(z)-M(z)'\gamma$, where $m_{y}(z) = E_{P_{W|Z}}[Y|Z=z]$ and $M(z) = (1,m_{d}(z),x)'$ with $m_{d}(z) = E_{P_{W|Z}}[D|Z=z]$, and $\Sigma(z,\gamma) = E_{P_{W|Z}}[g^{2}(W,Z,\gamma)|Z=z]$. Substituting $m(z,\gamma)$ and $\Sigma(z,\tilde{\gamma})$ into ((ref)), \begin{align} q_{n}(\tilde{\gamma},p_{W|Z}) = \left(\frac{1}{n}\sum_{i=1}^{n}\Sigma^{-1}(z_{i},\tilde{\gamma})M(z_{i})M(z_{i})' \right)^{-1}\frac{1}{n}\sum_{i=1}^{n}\Sigma^{-1}(z_{i},\tilde{\gamma})M(z_{i})m_{y}(z_{i}) \end{align} Since $p_{W|Z}$ entirely determines $\{q_{n}(\tilde{\gamma},p_{W|Z}): \tilde{\gamma} \in \Gamma\}$, a posterior for $p_{W|Z}$ implies a marginal posterior for $\{q_{n}(\tilde{\gamma},p_{W|Z}):\tilde{\gamma} \in \Gamma\}$, and, under Assumption (ref), leads to a posterior for $\gamma_{n}$. An example of $f$ is $f(\gamma_{n}) = \gamma_{n,1}$, the subvector of coefficients on $D$.
remark[Other Minimum Distance Estimands] Although this paper focuses on the iterated estimand $\gamma_{n}$, the `two-step' estimand $\gamma_{n,1} = q_{n}(\gamma_{n,0},p_{W|Z})$ and continuous updating estimand $\gamma_{n,CU} = \operatorname*{arg\,min}_{\gamma \in \Gamma} Q_{n}(\gamma,\gamma,p_{W|Z})$ can also be viewed as deterministic transformations of $p_{W|Z}$. Consequently, $\Pi(p_{W|Z} \in \cdot | W^{(n)})$ leads to a joint posterior for $(\gamma_{n},\gamma_{n,1},\gamma_{n,CU})$. I leave analysis of this joint posterior as future work.
remark[Prior on the Structural Parameter] The marginal prior $\Pi_{\Gamma}$ for $\gamma$, given by $\Pi_{\Gamma} = \Pi \circ \gamma_{n}^{-1}$, may not agree with a researcher's prior $\tilde{\Pi}_{\Gamma}$ about $\gamma$. Under regularity conditions, a marginal prior $\tilde{\Pi}_{\Gamma}$ can be incorporated using techniques from dunson2015marginally.
remark[Extensions] The posterior $\Pi(p_{W|Z} \in \cdot | W^{(n)})$ can be used for other inference problems related to the conditional moment equalities (i.e., beyond inference for $\gamma_{n}$). See Section (ref) for several examples that could each form the basis for future research.

A Flexible Class of Priors

I present a flexible class of priors for $p_{W|Z}$. The priors are based on logistic transformations of Gaussian processes (GPs), and are a leading class of priors in the Bayesian nonparametrics literature tokdar2007towards,TOKDAR200734,vaart2008rates,tokdar2010bayesian,vehtari2014laplace. Specifically, the conditional density of $W$ given $Z$ is modeled as an infinite-dimensional exponential family

align*[align* omitted — 132 chars of source]

where $f_{\theta}(\cdot|z)$, $\theta \in \Theta \subseteq \mathbb{R}^{d_{\theta}}$ with $d_{\theta} < \infty$, is a conditional density, $F_{\theta,z}(\cdot)$ is the corresponding $d_{w}\times 1$ vector of conditional cumulative distribution functions, $F_{0}(\cdot)$ is a known invertible mapping taking values in $[0,1]^{d_{z}}$, and $B: [0,1]^{d} \rightarrow \mathbb{R}$ is a function with $d = d_{w}+ d_{z}$.\footnote{Let $f_{\theta,j}(w_{j}|w_{1},...,w_{j-1},z)$ denote the density of $W_{j}$ given $\{W_{j'}\}_{j'=1}^{j-1}$ and $Z$ under $f_{\theta}(w|z)$. The elements of $F_{\theta,z}(w)=(F_{\theta,1,z}(w),...,F_{\theta,d_{w},z}(w))$ satisfy $F_{\theta,j,z}(w) = \int_{-\infty}^{w_{j}} f_{\theta,j}(t|w_{1},...,w_{j-1},z)dt$. Using this and the change of variables $u=F_{\theta,z}(w)$, one can show $\int_{\mathcal{W}}p_{\theta,B}(w|z)dw = 1$ for each $z$.} Since $(\theta,B)$ determines $p_{\theta,B}$, a prior for $p_{\theta,B}$ is implied by a prior $\Pi$ for $(\theta,B)$. I set $\Pi = \Pi_{\Theta} \otimes \Pi_{\mathcal{B}}$, where $\Pi_{\Theta}$ is a probability distribution over $\Theta$, and $\Pi_{\mathcal{B}}$ is the probability law of a centered GP. To define the latter, let $(C([0,1]^{d},\mathbb{R}),||\cdot||_{\infty})$ be the space of continuous $h: [0,1]^{d} \rightarrow \mathbb{R}$ with $||h||_{\infty} = \sup_{t \in [0,1]^{d}}|h(t)|$.

definitionA centered GP in $(C([0,1]^{d},\mathbb{R}),||\cdot||_{\infty})$ is a continuous stochastic process $\{B(t):t \in [0,1]^{d}\}$ such that, for every $t_{1},...,t_{k} \in [0,1]^{d}$ and $k \geq 1$, $(B(t_{1}),...,B(t_{k}))'$ follows a mean-zero Gaussian distribution with covariance matrix $(\kappa(t_{j},t_{l}))_{j,l=1}^{k}$. The symmetric, nonnegative-definite $\kappa: [0,1]^{d} \times [0,1]^{d} \rightarrow \mathbb{R}$ is called the covariance function.\footnote{The function $\kappa$ being symmetric and nonnegative-definite means that, for any $t_{1},...,t_{k} \in [0,1]^{d}$, the matrix $(\kappa(t_{j},t_{l}))_{j,l=1}^{k}$ is symmetric and positive-semidefinite, thereby making it suitable for covariance matrices.}

Logistically transformed GPs are flexible. Fixing $\theta$ and imposing $z \in [0,1]^{d_{z}}$ for simplicity, $p_{0,W|Z}$ is in the support of the prior if $\Pi_{\mathcal{B}}(||B-b_{0,\theta}||_{\infty} < \delta) > 0$ for each $\delta > 0$, where $b_{0,\theta}(t) = \log p_{0,W|Z}(F_{\theta,z}^{-1}(u)|z) - \log f_{\theta}(F_{\theta,z}^{-1}(u)|z)$ for $t = (u',z')' \in [0,1]^{d}$ is $\log(p_{0,W|Z}/f_{\theta})$ with $w$ transformed to be defined on $[0,1]^{d_{w}}$.\footnote{Two comments. First, the support claim holds because the Kullback-Leibler support of the prior for $p_{\theta,B}$, which is the standard definition of the support of a nonparametric prior, is determined by the uniform support of $\Pi_{\mathcal{B}}$ (Lemma 3.1 in vaart2008rates). Second, to allow for $z \in \mathcal{Z}\nsubseteq [0,1]^{d_{z}}$, replace $z$ with $F_{0}^{-1}(v)$ where $v \in [0,1]^{d_{z}}$.} For common GPs, this condition is typically satisfied if $b_{0,\theta}$ is continuous; however, to ensure certain frequentist properties for statistical functionals (i.e., $\gamma_{n}$), continuity of $b_{0,\theta}$ is often strengthened to H\"{o}lder or Sobolev-type smoothness restriction on $b_{0,\theta}$.\footnote{The H\"{o}lder space $C^{\alpha}([0,1]^{d},\mathbb{R})$, $\alpha > 0$, is the set of functions $f: [0,1]^{d} \rightarrow \mathbb{R}$ that are $\lfloor \alpha \rfloor$-times differentiable with bounded derivatives, and, additionally, the $\lfloor \alpha \rfloor$th derivative is $(\alpha-\lfloor \alpha \rfloor)$-H\"{o}lder continuous. It is equipped with norm $||f||_{\alpha} = \max_{|k| \leq \lfloor \alpha \rfloor}||D^{k}f||_{\infty}+ \max_{|k| = \lfloor \alpha \rfloor}\sup_{t,s \in [0,1]^{d}: t \neq s}\frac{|(D^{k}f)(t)-(D^{k}f)(s)|}{||t-s||_{2}^{\alpha-\lfloor\alpha \rfloor}}$, where $k=(k_{1},...,k_{d})'$ is a vector of nonnegative integers, $|k| = \sum_{j=1}^{d}k_{j}$, $D^{k}$ is the differential operator, and $||\cdot||_{2}$ is the Euclidean norm. The Sobolev space $S^{\alpha}([0,1]^{d},\mathbb{R})$, $\alpha > 0$, comprises functions $f:[0,1]^{d} \rightarrow \mathbb{R}$ that are restrictions of functions $f: \mathbb{R}^{d} \rightarrow \mathbb{R}$ with Fourier transforms $\hat{f}$ such that $||f||_{2,2,\alpha}^{2} = \int_{\mathbb{R}^{d}} (1+||\lambda||_{2}^{2})^{\alpha}|\hat{f}(\lambda)|^{2} < \infty$. The Sobolev space norm is $||\cdot||_{2,2,\alpha}$. } Similar smoothness conditions are often imposed on first-stage nuisance parameters in frequentist estimation of conditional moment equalities (e.g., newey1990efficient,NEWEY1993419, ai2003efficient, kankanala2025generalized).

exampleA centered GP belongs to the Mat\'{e}rn class if the covariance function satisfies $\kappa(t,s) = \int_{\mathbb{R}^{d}}\exp(i\lambda'(t-s))\mu(\lambda)d\lambda$ for $t,s \in [0,1]^{d}$, where $\mu(\lambda) = (1+||\lambda||_{2}^{2})^{-\alpha-d/2}$, $\alpha > 0$, and, here, $i$ is the imaginary unit. The hyperparameter $\alpha$ indexes smoothness because the sample paths $B$ take values in the H\"{o}lder spaces $ C^{a}([0,1]^{d},\mathbb{R})$ for any $a< \alpha$ van2011information. In Section (ref), I show that if $b_{0,\theta}$ is in the intersection of a H\"{o}lder space $C^{\alpha_{0}}([0,1]^{d},\mathbb{R})$ and Sobolev space $S^{\alpha_{0}}([0,1]^{d},\mathbb{R})$ for some $\alpha_{0} > 0$ (and other conditions hold), then the general BvM for $\gamma_{n}$ from Section (ref) can be verified for Mat\'{e}rn GPs. In practice, special functions are used to compute $\kappa$ williams2006gaussian.

Posterior computation for logistic GP priors proceeds as follows. The MCMC algorithms proposed in tokdar2007towards and tokdar2010bayesian can be used to obtain $S \geq 1$ draws $\{(\theta^{[s]},B^{[s]})\}_{s=1}^{S}$ from the posterior for $(\theta,B)$ (see Appendix (ref)). Given a posterior draw $(\theta^{[s]},B^{[s]})$, the conditional moments $m^{[s]}(z_{i},\gamma)$ and $\Sigma^{[s]}(z_{i},\gamma)$ at $z_{i}$, $i=1,...,n$, can be estimated using importance sampling: for each $i$, 1. generate $\{W_{i,j}^{[s]}\}_{j=1}^{J}\overset{iid}{\sim} f_{\theta^{[s]}}(\cdot|z_{i})$, 2. compute importance weights $\{\omega_{i,j}^{[s]}\}_{j=1}^{J}$,

align*[align* omitted — 207 chars of source]

and 3. compute importance-weighted averages

align*[align* omitted — 236 chars of source]

A draw $\gamma_{n}^{[s]}$ from $\Pi(\gamma_{n} \in \cdot | W^{(n)})$ is then obtained by solving the iterated minimum distance problem, replacing $m(z_{i},\gamma)$ and $\Sigma(z_{i},\tilde{\gamma})$ in ((ref)) with $m_{J}^{[s]}(z_{i},\gamma)$ and $\Sigma_{J}^{[s]}(z_{i},\tilde{\gamma})$, respectively. Performing these steps over $\{(\theta^{[s]},B^{[s]})\}_{s=1}^{S}$ leads to a sample $\{\gamma_{n}^{[s]}\}_{s=1}^{S}$ from $\Pi(\gamma_{n} \in \cdot | W^{(n)})$. A sample from $\Pi(f(\gamma_{n}) \in \cdot | W^{(n)})$ is obtained by computing $\{f(\gamma_{n}^{[s]})\}_{s=1}^{S}$ using $\{\gamma_{n}^{[s]}\}_{s=1}^{S}$. Bayes estimators (e.g., posterior medians, credible sets) are computed using the empirical distribution of $\{f(\gamma_{n}^{[s]})\}_{s=1}^{S}$.

remark[Other Priors] There are many other priors for conditional densities, such as covariate-dependent mixtures of normals norets2010approximation,norets2014posterior,norets2017adaptive, and covariate-dependent stick-breaking priors dunson2008kernel,ren2011logistic,rodriguez2011nonparametric. The modularity of my approach (due to the unrestricted $p_{W|Z}$) means that, in principle, it could be applied to all of these priors (with precise implementation details depending on the class), and verifying the high-level BvM conditions in Section (ref) may offer some interesting new theoretical insights. I leave this for future research.

Empirical Illustration: Estimating Welfare Effects

This section uses my proposal to estimate welfare effects of price changes. A data-calibrated simulation that compares my approach with alternatives is in Appendix (ref).

Dataset and Target Parameter

I use the 2001 National Household Travel Survey gasoline demand dataset from blundell2012measuring,blundell2017nonparametric.\footnote{The dataset was obtained from the publicly available replication files of chen2018optimal. They can be found using the link \url{https://github.com/timothymchristensen/NPIV}.} This dataset contains household gasoline consumption $Q$ (in gallons), average price $P$ of gasoline (in dollars per gallon) in the household's county, household income $Y$ (in dollars), and distance $A$ (in 1,000 kilometers) of a household's state capital to a major oil platform in the Gulf of Mexico. There are $n=4,812$ households in the dataset.

The parameter of interest is the welfare effect of a gasoline price change (e.g., due to taxation). Household gasoline demand is modeled using a constant elasticity specification

align*[align* omitted — 77 chars of source]

where the error term $U$ satisfies $E[U|\log Y, \log A] = 0$.\footnote{Distance as an excluded IV is proposed in blundell2012measuring and is implemented in chen2018optimal too. A justification is that distance is a cost-shifter in that it affects prices only through its impact on firm transportation costs.} The welfare effect of a price change from $p^{0}$ to $p^{1}$ at income $y$ is measured using deadweight loss (DWL),

align*[align* omitted — 91 chars of source]

where $S(p^{0},y,\gamma)$ is consumer surplus hausman1981exact,hausman1995nonparametric, and $q(p,y,\gamma) = \exp(\gamma_{0}+\gamma_{1} \log p + \gamma_{2} \log y)$. walker2026supplement details DWL calculation.

Semiparametric Bayesian Inference

Let $W = (\log Q ,\log P)'$, $Z = (\log Y, \log A)'$, and $g(W,Z,\gamma) = \log Q - \gamma_{0} - \gamma_{1} \log P -\gamma_{2} \log Y$. Since $E[U|\log Y, \log A] = 0$ implies $E_{P_{W|Z}}[g(W,Z,\gamma)|Z] = 0$, Bayesian inference for $\gamma$ can be obtained via the approach in Section (ref) (in fact, it is an instance of Example (ref)). Furthermore, since $DWL(p^{0},p^{1},y,\gamma)$ is a deterministic function of $\gamma$ (as the researcher sets $(p^{0},p^{1},y)$), one automatically obtains a posterior for $DWL(p^{0},p^{1},\gamma)$. This highlights that my framework offers simultaneous inference for model primitives ($\gamma$) and counterfactuals ($DWL(p^{0},p^{1},y,\gamma)$).

Figure (ref) and Table (ref) report the posterior densities and Bayes estimates, respectively, of the DWL for $p^{0} = \$1.22$, $p^{1} = \$1.44$, and $y \in \{\$42500,\$57500,\$72500\}$.\footnote{These price-income combinations are the same as those reported in blundell2012measuring.} The posteriors are based on a logistic GP prior for $p_{W|Z}$ (from Section (ref)), with a homoskedastic Gaussian linear regression of $W$ on $Z$ for $f_{\theta}$, diffuse priors for the location-scale parameters $\theta$, and a Mat\'{e}rn GP for $B$ with $\alpha = 5/2$.\footnote{Two comments. First, Appendix (ref) contains more details about the prior. Second, setting $\alpha = 5/2$ follows recommendations in williams2006gaussian; walker2026supplement shows changing $\alpha$ has a modest effect on the estimates.} Figure (ref) indicates that the DWL posteriors are approximately symmetric and bell-shaped across all income groups. The estimates in Table (ref) reveal that DWL as a percentage of tax paid is almost identical across the income groups, however, as a proportion of income, deadweight loss decreases monotonically with income level. These findings are qualitatively similar to the constant elasticity results in blundell2012measuring, though my results relax price exogeneity.

figure[figure omitted — 160 chars of source]
table[table omitted — 1,168 chars of source]

Semiparametric Bernstein-von Mises Theorem

This section proves a general BvM, proving that $\Pi(\gamma_{n} \in \cdot|W^{(n)})$ is asymptotically Gaussian and achieves the chamberlain1987asymptotic semiparametric efficiency bound.

Assumptions about the Data-Generating Process (DGP)

I start with assumptions about the DGP. For notation, let $||\cdot||_{2}$ be the Euclidean norm, let $||\cdot||_{op}$ be the operator norm (i.e., the maximum singular value), let $\lambda_{min}(A)$ and $\lambda_{max}(A)$ be the minimum and maximum eigenvalues, respectively, of matrix $A$, let $\text{int}(B)$ be the interior of a set $B$, and let $vec(A)$ be the vectorized matrix $A$. A function class is Glivenko-Cantelli if it obeys a uniform strong law of large numbers (USLLN).

assumption1. $\{(W_{i}',Z_{i}')'\}_{i \geq 1}$ is an i.i.d sequence with common distribution $P_{0,WZ} = P_{0,W|Z} \otimes P_{0,Z}$, where $P_{0,W|Z}$ is the conditional distribution of $W$ given $Z$, and $P_{0,Z}$ is the marginal distribution of $Z$; 2. $P_{0,W|Z}$ has a density $p_{0,W|Z}$.
assumption1. $\Gamma$ is a compact subset of $\mathbb{R}^{d_{\gamma}}$, $d_{\gamma} < \infty$; 2. there is a unique $\gamma_{0} \in \text{int}(\Gamma)$ such that $m_{0}(z,\gamma_{0})= 0$ for every $z \in \mathcal{Z}$, where $m_{0}(z,\gamma)$ denotes $m(z,\gamma)$ at $P_{0,W|Z}$.
assumption1. $\sup_{(z,\gamma) \in \mathcal{Z}\times \Gamma}||m_{0}(z,\gamma)||_{2} < \infty$; 2. $m_{0}(z,\gamma)$ is twice continuously differentiable in $\gamma$ in a neighborhood $\Gamma_{0}$ of $\gamma_{0}$ for each $z \in \mathcal{Z}$, $\sup_{(z,\gamma) \in \mathcal{Z}\times \Gamma_{0}}||M_{0}(z,\gamma)||_{op}< \infty $, and $\sup_{(z,\gamma) \in \mathcal{Z}\times \Gamma_{0}}||\partial vec(M_{0}(z,\gamma))/\partial \gamma'||_{op} < \infty$, where $M_{0}(z,\gamma)$ is the Jacobian of $m_{0}(z,\gamma)$ in $\gamma$.
assumptionLet $\Sigma_{0}(z,\gamma)$ denote $\Sigma(z,\gamma)$ at $P_{0,W|Z}$. 1. there exists $0 < \underline{\lambda} < \overline{\lambda}< \infty$ such that $\underline{\lambda} \leq \inf_{(z,\gamma) \in \mathcal{Z}\times \Gamma}\lambda_{min}(\Sigma_{0}(z,\gamma))\leq \sup_{(z,\gamma) \in \mathcal{Z} \times \Gamma} \lambda_{max}(\Sigma_{0}(z,\gamma)) \leq \overline{\lambda}$; 2. $\Sigma_{0}(z,\gamma)$ is continuously differentiable in $\gamma$ in a neighborhood $\Gamma_{0}$ of $\gamma_{0}$ for each $z \in \mathcal{Z}$ with Jacobian satisfying $\sup_{(z,\gamma) \in \mathcal{Z}\times \Gamma_{0}}||\partial vec(\Sigma_{0}(z,\gamma))/\partial \gamma'||_{op} < \infty$.
assumption1. the class $\{m_{0}(\cdot,\gamma)'\Sigma_{0}^{-1}(\cdot,\tilde{\gamma})m_{0}(\cdot,\gamma):(\gamma,\tilde{\gamma}) \in \Gamma \times \Gamma\}$ is $P_{0,Z}$-Glivenko-Cantelli; 2. for every $\xi>0$, $\inf_{\gamma \in \mathbb{R}^{d_{\gamma}}:||\gamma-\gamma_{0}||_{2} \geq \xi} E_{P_{0,Z}}[||m_{0}(Z,\gamma)||_{2}^{2}]> 0$.
assumptionLet $V_{0}(Z,\tilde{\gamma}) = M_{0}(Z,\gamma_{0})'\Sigma_{0}^{-1}(Z,\tilde{\gamma})M_{0}(Z,\gamma_{0})$. 1. $\{V_{0}(\cdot,\tilde{\gamma}): \tilde{\gamma} \in \Gamma\}$ is $P_{0,Z}$-Glivenko-Cantelli; 2. $E_{P_{0,Z}}[M_{0}(Z,\gamma_{0})'M_{0}(Z,\gamma_{0})]$ is positive definite.

Assumptions (ref)--(ref) are similar restrictions to those encountered in the optimal instruments literature. Assumption (ref) imposes the same sampling restriction as chamberlain1987asymptotic, and, relative to Section (ref), is a restriction on $\{Z_{i}\}_{i \geq 1}$. Assumption (ref) imposes correct specification, requires that $\gamma_{0}$ be point identified, and that $\gamma_{0}$ belongs to the interior of a compact $\Gamma \subseteq \mathbb{R}^{d_{\gamma}}$. Misspecification is discussed in Section (ref). Assumption (ref) imposes some smoothness restrictions on $\gamma \mapsto m_{0}(z,\gamma)$. It accommodates the differentiability conditions on $g(W,Z,\gamma)$ that are routinely imposed in the optimal instruments literature (e.g., chamberlain1987asymptotic, newey1990efficient,NEWEY1993419, donald2003empirical, and chib2022bayesian), while allowing settings where $g(W,Z,\gamma)$ is nondifferentiable but $m_{0}(Z,\gamma)$ is differentiable (e.g., quantile IV chernozhukov2005iv,chernozhukov2006instrumental). Assumption (ref) imposes a compactness restriction on $\{\Sigma_{0}(z,\gamma): (z,\gamma) \in \mathcal{Z} \times \Gamma\}$ and a mild smoothness requirement for $\gamma \mapsto \Sigma_{0}(z,\gamma)$. Assumption (ref).1 requires $Q_{n}(\gamma,\tilde{\gamma},p_{0,W|Z})$ obeys a USLLN over $\Gamma \times \Gamma$, while Assumption (ref).2, combined with Assumption (ref).1, implies that the criterion $Q_{P_{0,Z}}(\gamma,\tilde{\gamma},p_{0,W|Z}) := E_{P_{0,Z}}[m_{0}(Z,\gamma)'\Sigma_{0}^{-1}(Z,\tilde{\gamma})m_{0}(Z,\gamma)]$ has a uniformly well-separated minimum at $\gamma_{0}$, an important condition for consistent estimation of $\gamma_{0}$. Assumption (ref).1 is a USLLN condition for the class $\{V_{0}(\cdot,\tilde{\gamma}): \tilde{\gamma} \in \Gamma\}$, and Assumption (ref).2 is a strong identification condition that, combined with Assumption (ref).1, implies $V_{0} :=E_{P_{0,Z}}[V_{0}(Z,\gamma_{0})]$ is positive definite, so the chamberlain1987asymptotic efficiency bound, $V_{0}^{-1}$, is finite.

Assumptions about the Prior

The next assumption concerns $\Pi(p_{W|Z} \in \cdot |W^{(n)})$. For notation, let $||f||_{n,2}^{2} = n^{-1}\sum_{i=1}^{n}||f(z_{i})||^{2}$ and $||f||_{n,\infty} = \max_{1 \leq i \leq n }||f(z_{i})||$, where $||\cdot||$ is $||\cdot||_{2}$ if $f(z)$ is a vector and is $||\cdot||_{op}$ if $f(z)$ is a matrix, let $$\tilde{\chi}_{0}(W,Z) = -V_{0}^{-1}M_{0}(Z,\gamma_{0})'\Sigma_{0}^{-1}(Z,\gamma_{0})g(W,Z,\gamma_{0}) $$ be the chamberlain1987asymptotic efficient influence function (EIF) and note that $V_{0}^{-1}$ also satisfies $V_{0}^{-1} = E_{P_{0,WZ}}[\tilde{\chi}_{0}(W,Z)\tilde{\chi}_{0}(W,Z)']$, let $\overset{P_{0,W|Z}^{(n)}}{\longrightarrow}$ denote convergence in probability under $P_{0,W|Z}^{(n)}=\bigotimes_{i=1}^{n}P_{0,W|z_{i}}$, let $P_{0,Z}^{\infty}$ be the probability law of the i.i.d sequence $\{Z_{i}\}_{i \geq 1}$ (i.e., $P_{0,Z}^{\infty} = \bigotimes_{i \geq 1}P_{0,Z}$), and sequences $\{x_{n}\}_{n \geq 1}$ and $\{y_{n}\}_{n \geq 1}$ satisfy $x_{n} = o(y_{n})$ if $x_{n}/y_{n} \rightarrow 0$ as $n\rightarrow \infty$ and $x_{n} = O(y_{n})$ if there is a constant $C > 0$ such that $|x_{n}/y_{n}| \leq C$ for large $n$.

assumptionFor $P_{0,Z}^{\infty}$-almost every fixed realization $\{z_{i}\}_{i \geq 1}$ of $\{Z_{i}\}_{i \geq 1}$, there exists a sequence of sets $\{\tilde{\mathcal{P}}_{n,W|Z}\}_{n\geq 1}$ such that \begin{enumerate} • $\Pi(p_{W|Z} \in \tilde{\mathcal{P}}_{n,W|Z} |W^{(n)}) \longrightarrow 1$ in $P_{0,W|Z}^{(n)}$-probability as $n\rightarrow \infty$ • For each $p_{W|Z} \in \tilde{\mathcal{P}}_{n,W|Z}$, \begin{enumerate} • $m(z_{i},\gamma)$, $i=1,...,n$, is twice continuously differentiable in $\gamma$ over $\Gamma_{0}$$\Sigma(z_{i},\gamma)$, $i=1,...,n$, is once continuously differentiable in $\gamma$ over $\Gamma_{0}$, \end{enumerate} where $\Gamma_{0}$ is the neighborhood of $\gamma_{0}$ from Assumption (ref). • The following holds: \begin{enumerate} • $\sup_{(\gamma,p_{W|Z}) \in \Gamma \times \tilde{\mathcal{P}}_{n,W|Z}}||m(\cdot,\gamma)-m_{0}(\cdot,\gamma)||_{n,2} = o(n^{-\frac{1}{4}})$$\sup_{(\gamma,p_{W|Z}) \in \Gamma_{0} \times \tilde{\mathcal{P}}_{n,W|Z}}||M(\cdot,\gamma)-M_{0}(\cdot,\gamma)||_{n,2} = o(n^{-\frac{1}{4}})$$\sup_{(\gamma,p_{W|Z}) \in \Gamma \times \tilde{\mathcal{P}}_{n,W|Z}}||\Sigma(\cdot,\gamma)-\Sigma_{0}(\cdot,\gamma)||_{n,\infty} = o(n^{-\frac{1}{4}})$$\sup_{(\gamma,p_{W|Z}) \in \Gamma_{0} \times \tilde{\mathcal{P}}_{n,W|Z}}||\partial \text{vec}(M(\cdot,\gamma)')/\partial \gamma'||_{n,2} = O(1)$$\sup_{(\gamma,p_{W|Z}) \in \Gamma_{0} \times \tilde{\mathcal{P}}_{n,W|Z}}||\partial\text{vec}(\Sigma(\cdot,\gamma))/\partial \gamma'||_{n,\infty}= O(1)$$\sup_{p_{W|Z} \in \tilde{\mathcal{P}}_{n,W|Z}}\max_{1 \leq i \leq n}E_{P_{W|Z}}[\exp(|t'\tilde{\chi}_{0}(W,Z)|)|Z=z_{i}] = O(1)$ \end{enumerate} for any $t$ in a neighborhood of zero.\footnote{Use of $o(\cdot)$ and $O(\cdot)$ is appropriate because, conditional on $\{Z_{i}\}_{i \geq 1}$, these objects are nonrandom.} \end{enumerate}

Assumption (ref) states that, conditional on $\{Z_{i}\}_{i \geq 1}$, there are sets $\{\tilde{\mathcal{P}}_{n,W|Z}\}_{n \geq 1}$ on which the posterior concentrates as $n\rightarrow \infty$, and these sets are structured enough to ensure certain smoothness (in $\gamma$) and convergence guarantees for $m$, $M$, and $\Sigma$. The assumption is general in that it is stated for an arbitrary prior $\Pi$ for $p_{W|Z}$; I derive examples of $\{\tilde{\mathcal{P}}_{n,W|Z}\}_{n \geq 1}$ for certain GPs in Section (ref). As for the contents, Parts 3(a) and 3(b) state that $m(\cdot,\gamma)$ and $M(\cdot,\gamma)$ converge to their true counterparts $m_{0}(\cdot,\gamma)$ and $M_{0}(\cdot,\gamma)$, respectively, at a rate faster than $n^{-1/4}$ under the empirical $L^{2}$ norm. Part 3(c) states that $\Sigma(\cdot,\gamma)$ converges to $\Sigma_{0}(\cdot,\gamma)$ at a rate faster than $n^{-1/4}$ under the empirical supremum norm, a condition analogous to Assumption 3.4(iii) in ai2003efficient. Requiring $o(n^{-1/4})$ nuisance convergence rates is generally considered a mild condition in semiparametric applications. Part 3(d) and 3(e) are boundedness conditions that help show that $\tilde{\gamma} \mapsto q_{n}(\tilde{\gamma},p_{W|Z})$ is a uniform contraction mapping in large samples, a property that helps establish that $\gamma_{n}$ uniformly converges to $\gamma_{0}$ along $\{\tilde{\mathcal{P}}_{n,W|Z}\}_{n \geq 1}$. Part 3(f) requires that the EIF has sufficiently thin tails over $\tilde{\mathcal{P}}_{n,W|Z}$, a condition sufficient to control a remainder in an expansion of $\log L_{n}(p_{W|Z})$.

Bernstein-von Mises Theorem

Main Result

Assumptions (ref)--(ref) lead to a Bernstein-von Mises theorem (BvM). Informally, the BvM states that $\gamma_{n}|W^{(n)}\overset{a}{\sim} \mathcal{N}(\hat{\gamma}_{n},\frac{1}{n}V_{0}^{-1})$ for large $n$, where $\hat{\gamma}_{n} = \gamma_{0}+n^{-1}\sum_{i=1}^{n}\tilde{\chi}_{0}(W_{i},z_{i})$ is a hypothetical efficient estimator of $\gamma_{0}$. Consequently, the posterior for $\gamma_{n}$ behaves like a best regular estimator $\hat{\gamma}_{n}$ for large $n$, thereby establishing its frequentist asymptotic optimality.

The key to the BvM is an asymptotically linear representation of $\gamma_{n}$. Stated as a theorem below, it shows that $\gamma_{n}$ is asymptotically linear in $p_{W|Z}$ with Riesz representer given by $\tilde{\chi}_{0}$. It can be thought of as a Bayesian analogue to the asymptotically linear representation of an efficient estimator, and validates that $\Pi(\gamma_{n} \in \cdot | W^{(n)})$ accounts for the optimal instruments as $n\rightarrow \infty$.

theoremIf Assumptions (ref), (ref), (ref), (ref), (ref), (ref), (ref), and (ref).1--(ref).5 hold, then, for $P_{0,Z}^{\infty}$-almost every fixed realization $\{z_{i}\}_{i \geq 1}$ of $\{Z_{i}\}_{i \geq 1}$, \begin{align*} \sup_{p_{W|Z} \in \tilde{\mathcal{P}}_{n,W|Z}}\left | \left | \sqrt{n}(\gamma_{n}-\gamma_{0}) - \frac{1}{\sqrt{n}}\sum_{i=1}^{n}E_{P_{W|Z}}[\tilde{\chi}_{0}(W,Z)|Z=z_{i}] \right | \right |_{2} =o(1), \end{align*} as $n\rightarrow \infty$, where $\{\tilde{\mathcal{P}}_{n,W|Z}\}_{n \geq 1}$ are the sets defined in Assumption (ref).

Theorem (ref) establishes that, asymptotically, $\gamma_{n}$ is an (efficient) linear functional of $p_{W|Z}$. Consequently, a semiparametric BvM for $\gamma_{n}$ can be established by building on castillo2015bernstein, who prove BvMs for unconditional density functions with i.i.d data.

theoremSuppose that Assumptions (ref)--(ref) hold and that, for $P_{0,Z}^{\infty}$-almost every fixed realization $\{z_{i}\}_{i \geq 1}$ of $\{Z_{i}\}_{i \geq 1}$, the prior invariance condition \begin{align} \frac{\int_{\tilde{\mathcal{P}}_{n,W|Z}}L_{n}(p_{t,n,W|Z})d\Pi(p_{W|Z})}{\int_{\tilde{\mathcal{P}}_{n,W|Z}}L_{n}(p_{W|Z}) d\Pi(p_{W|Z})} = 1 + o_{P_{0,W|Z}^{(n)}}(1) \end{align} holds for any $t$ in a neighborhood of zero, where \begin{align*} p_{t,n,W|Z}(w|z)= \frac{p_{W|Z}(w|z)\exp\left(-\frac{1}{\sqrt{n}}t'\tilde{\chi}_{0}(w,z)\right)}{\int_{\mathcal{W}}p_{W|Z}(\tilde{w}|z)\exp\left(-\frac{1}{\sqrt{n}}t'\tilde{\chi}_{0}(\tilde{w},z)\right)d\tilde{w}}. \end{align*} Then, for almost every fixed realization $\{z_{i}\}_{i \geq 1}$ of $\{Z_{i}\}_{i \geq 1}$, the structural posterior satisfies \begin{align*} d_{BL}\left(\Pi\left(\sqrt{n}(\gamma_{n}-\hat{\gamma}_{n}) \in \cdot\middle |W^{(n)}\right), \mathcal{N}(0,V_{0}^{-1})\right) \overset{P_{0,W|Z}^{(n)}}{\longrightarrow} 0 \end{align*} as $n\rightarrow \infty$, where $d_{BL}$ is the bounded Lipschitz metric.\footnote{For any two probability measures $P$ and $Q$ on a measurable space $(\mathcal{X},\mathscr{X})$, the bounded Lipschitz distance is given by $d_{BL}(P,Q) = \sup_{f \in Lip_{1}}|\int f d P - \int f dQ|$, where $Lip_{1}$ is the set of functions $f: \mathcal{X} \rightarrow [-1,1]$ with $\sup_{x \in \mathcal{X}}|f(x)| \leq 1$ and $|f(x)-f(y)| \leq d(x,y)$ for all $x,y \in \mathcal{X}$ with $x \neq y$, where $d$ is a metric over $\mathcal{X}$. Convergence in $d_{BL}$ is equivalent to convergence in distribution.}

Comments about the BvM

Several aspects of Theorem (ref) should be highlighted. First, Theorem (ref) implies that some Bayesian point estimators and credible sets are asymptotically efficient. Let $f: \mathbb{R}^{d_{\gamma}} \rightarrow \mathbb{R}$ be continuously differentiable at $\gamma_{0}$ and recall that $c_{n,f}(q)$, $q \in (0,1)$, is the $q$-quantile of $\Pi(f(\gamma_{n}) \in \cdot|W^{(n)})$. Corollary (ref) states that, conditional on $\{Z_{i}\}_{i \geq 1}$, $c_{n,f}(q)=f(\hat{\gamma}_{n}) + \Phi^{-1}(q)n^{-1/2}\Omega_{0,f}^{1/2}+o_{P_{0,W|Z}^{(n)}}(n^{-1/2})$, where $\Phi(\cdot)$ is the $\mathcal{N}(0,1)$ cumulative distribution function and $\Omega_{0,f} =\frac{\partial f(\gamma_{0})}{\partial \gamma'}V_{0}^{-1}\frac{\partial f(\gamma_{0})'}{\partial \gamma}$ is the asymptotic variance of $f(\hat{\gamma}_{n})$. Consequently, by setting $q= 0.5$, the posterior median $c_{n,f}(0.5)$ is first-order asymptotically equivalent to the efficient estimator $f(\hat{\gamma}_{n})$.\footnote{The restrictions on $f$ imply the delta method preserves efficiency (see Section 25.7 of vaart_1998).} Moreover, the equitailed probability interval $CS_{n,f}(1-\alpha)= [c_{n,f}(\alpha/2), c_{n,f}(1-\alpha/2)]$, where $\alpha \in (0,1/2)$, is first-order asymptotically equivalent to an efficient Wald confidence interval $[f(\hat{\gamma}_{n})+\Phi^{-1}(\alpha/2)\sqrt{\Omega_{0,f}/n},f(\hat{\gamma}_{n})-\Phi^{-1}(\alpha/2)\sqrt{\Omega_{0,f}/n}]$, thereby leading to best asymptotic frequentist uncertainty quantification.

corollaryIf the assumptions of Theorem (ref) hold and $f:\mathbb{R}^{d_{\gamma}}\rightarrow \mathbb{R}$ is continuously differentiable at $\gamma_{0}$, then, for any $q \in (0,1)$ and $P_{0,Z}^{\infty}$-almost every $\{z_{i}\}_{i \geq 1}$, $c_{n,f}(q)$ satisfies \begin{align*} c_{n,f}(q) = f(\hat{\gamma}_{n}) + \Phi^{-1}(q)\sqrt{\frac{\Omega_{0,f}}{n}} + o_{P_{0,W|Z}^{(n)}}\left(\frac{1}{\sqrt{n}}\right), \quad \Omega_{0,f} = \frac{\partial f(\gamma_{0})}{\partial \gamma'}V_{0}^{-1}\frac{\partial f(\gamma_{0})'}{\partial \gamma}. \end{align*}

A second aspect is that Theorem (ref) also holds unconditionally, where stochastic convergence is defined with respect to $P_{0,WZ}$. Corollary (ref) formalizes this, and, by extension, implies an unconditional version of Corollary (ref) (see Remark (ref)).

corollaryIf the assumptions of Theorem (ref) holds, then \begin{align*} d_{BL}\left(\Pi\left(\sqrt{n}(\gamma_{n}-\hat{\gamma}_{n}) \in \cdot \middle |W^{(n)}\right), \mathcal{N}\left(0,V_{0}^{-1}\right)\right)\overset{P_{0,WZ}^{n}}{\longrightarrow} 0 \end{align*} as $n\rightarrow \infty$, where $P_{0,WZ}^{n}$ is $n$-fold product of $P_{0,WZ}$.

A final aspect worth discussing is condition ((ref)). These conditions are common in semiparametric BvMs (see castillo2012semiparametric,castillo2012gaussian, rivoirard2012bernstein, castillo2015bernstein, ray2020semiparametric, and monard2021bernstein, to name a few), and arise from expanding the log-likelihood ($\log L_{n}(p_{W|Z})$) around the least favorable parametric submodel ($p_{t,n,W|Z}$). The name `prior invariance' reflects that its verification typically requires the prior be roughly unchanged under shifts in the direction of the EIF. Indeed, Section (ref) verifies ((ref)) for GPs by finding a sequence $\{\tilde{g}_{n}\}_{n \geq 1}$ that uniformly approximates the chamberlain1987asymptotic optimal residual $\tilde{g}_{0} = -V_{0}\tilde{\chi}_{0}$, with $o(n^{-1/4})$ approximation error, and for which the law $\Pi_{n,t}$ of $p_{n,t,W|Z}\propto p_{W|Z}\exp(-t'V_{0}^{-1}\tilde{g}_{n}/\sqrt{n})$ under $\Pi$ satisfies $\Pi_{n,t} \ll \Pi$ and $d \Pi_{n,t}/d\Pi \rightarrow 1$ uniformly along $\{\tilde{\mathcal{P}}_{n,W|Z}\}_{n \geq 1}$ as $n\rightarrow \infty$. For the Mat\'{e}rn process (i.e., Example (ref)), this restricts the GP sample path smoothness relative to $\tilde{g}_{0}$.

Verifying Assumption (ref) and Prior Invariance with GPs

I provide sufficient conditions for Assumption (ref) and ((ref)) for the priors from Section (ref).

Setup and Assumptions

Recall $p_{\theta,B}(w|z) \propto f_{\theta}(w|z)\exp(B(F_{\theta,z}(w),F_{0}(z)))$, where $f_{\theta}(\cdot|z)$, $\theta \in \Theta \subseteq \mathbb{R}^{d_{\theta}}$ with $d_{\theta} < \infty$, is a conditional density, $F_{\theta,z}(\cdot)$ is the $d_{w}\times 1$ vector of conditional CDFs attached to $f_{\theta}(\cdot|z)$, $F_{0}: \mathcal{Z} \rightarrow [0,1]^{d_{z}}$ is a known invertible transformation, and $B: T \rightarrow \mathbb{R}$ is a function. A prior for $p_{\theta,B}$ is obtained from a prior $\Pi$ for $(\theta,B)$. Also, recall that $(C^{\alpha}([0,1]^{d},\mathbb{R}),||\cdot||_{\alpha})$ and $(S^{\alpha}([0,1]^{d},\mathbb{R}),||\cdot||_{2,2,\alpha})$ denote H\"{o}lder and Sobolev spaces, respectively, of order $\alpha >0$ (see Footnote (ref) for precise details).

assumption1. $\mathcal{W}\subseteq \mathbb{R}^{d_{w}}$ is compact; 2. $\mathcal{Z} = [0,1]^{d_{z}}$; 3. $p_{0,W|Z}$ is continuous and positive on $\mathcal{W} \times \mathcal{Z}$; 4. $P_{0,Z}$ admits a Lebesgue density $p_{0,Z}$ that is bounded away from zero and infinity.
assumption1. $(w,z,\gamma) \mapsto g_{j}(w,z,\gamma)$ is continuous in all arguments over $\mathcal{W}\times \mathcal{Z} \times \Gamma$ for $j=1,...,d_{g}$; 2. $g_{j}(w,z,\gamma)$, $j=1,...,d_{g}$, is differentiable in $\gamma$ over a neighborhood $\Gamma_{0}$ of $\gamma_{0}$ with first and second derivatives continuous in all arguments over $\mathcal{W} \times \mathcal{Z} \times \Gamma_{0}$.
assumption1. $\Theta \subseteq \mathbb{R}^{d_{\theta}}$ is compact; 2. $(w,z,\theta) \mapsto f_{\theta,j}(w_{j}|w_{1},...,w_{j-1},z)$ is continuous in all arguments for all $j=1,...,d_{w}$; 3. $f_{\theta,j}(w_{j}|w_{1},...,w_{j-1},z) > 0$ for all $(w,z,\theta) \in \mathcal{W} \times \mathcal{Z} \times \Theta$ and all $j=1,...,d_{w}$; 4. there is a constant $L>0$ such that $\sup_{(w,z) \in \mathcal{W} \times \mathcal{Z}}|\log f_{\theta_{1}}(w|z) - \log f_{\theta_{2}}(w|z)| \leq L ||\theta_{1}-\theta_{2}||_{2}$ for all $\theta_{1},\theta_{2} \in \Theta$.\footnote{Recall that $f_{\theta,j}(w_{j}|w_{1},...,w_{j-1},z)$ is the density of $W_{j}$ given $\{W_{j'}\}_{j'=1}^{j-1}$ and $Z$ under $f_{\theta}(w|z)$.}
assumption$\Pi = \Pi_{\Theta} \otimes \Pi_{\mathcal{B}}$, where 1. $\Pi_{\Theta}$ is a probability distribution over $\Theta$ with a continuous Lebesgue density $\pi_{\Theta}$, and 2. $\Pi_{\mathcal{B}}$ is the law $\tilde{\Pi}_{\mathcal{B}}$ of a centered Gaussian process $B$ in $(C([0,1]^{d},\mathbb{R}),||\cdot||_{\infty})$ conditioned on $\{||B||_{\infty}\leq \bar{B}\}$ for some large $\bar{B} \in (0,\infty)$ with $B \sim \tilde{\Pi}_{\mathcal{B}}$ taking values in $(C^{a}([0,1]^{d},\mathbb{R}),||\cdot||_{a})$ for every $0<a<\alpha$ and $\alpha > d_{z}/2$.

Assumptions (ref)--(ref) comprise sufficient conditions on the DGP and prior that will be used to verify Assumption (ref). The main limitation of Assumption (ref) is compactness of $\mathcal{W}$, however it is not overly restrictive relative to the literature on semiparametric BvMs for density functionals rivoirard2012bernstein,castillo2015bernstein, and, more broadly, posterior consistency with infinite-dimensional exponential families scricciolo2006convergence,rivoirard2012posterior. I expect compact $\mathcal{W}$ can be relaxed (see Footnote (ref)). Compactness of $\mathcal{Z}$ is a standard assumption in nonparametric conditional density estimation, and, given such an assumption, setting $\mathcal{Z} = [0,1]^{d_{z}}$ is without loss of generality (and enables setting $F_{0}(z)=z$). Assumption (ref) is also stronger than the high-level smoothness conditions imposed in Section (ref), as it rules out models with nondifferentiable $g(W,Z,\gamma)$. Assumption (ref) imposes regularity conditions on $\{f_{\theta}: \theta \in \Theta\}$ that are satisfied, for example, when $f_{\theta}$ is a Gaussian linear regression (truncated to $\mathcal{W}$), $\mathcal{Z}$ is compact (i.e. Assumption (ref).2), and $\Theta$ constrains the slope and variance to compact sets, with the variance bounded away from zero. Assumption (ref) is a sample path smoothness restriction that many Gaussian processes satisfy (e.g., the Mat\'{e}rn process from Example (ref)). Uniformly bounding $B$ is a technical device used to control random approximation errors, and, since the magnitude of $\bar{B}$ has no role, it can be viewed as an arbitrarily large but finite number.\footnote{Lemma 5.1 in van2008reproducing shows $\tilde{\Pi}_{\mathcal{B}}(||B||_{\infty} \leq \bar{B}) > 0$ for any $\bar{B} > 0$.}

Verifying Assumption (ref)

Under Assumptions (ref) and (ref), only Assumptions (ref).3(a)--10.3(c) need verification. Let $h(p_{\theta,B}(\cdot|z),p_{0,W|Z}(\cdot|z))$ be the Hellinger distance between $p_{\theta,B}$ and $p_{0,W|Z}$ at $z$, and let $h_{n,2}(p_{\theta,B},p_{0,W|Z}) = (n^{-1}\sum_{i=1}^{n}h^{2}(p_{\theta,B}(\cdot|z_{i}),p_{0,W|Z}(\cdot|z_{i})))^{1/2}$ be the root mean square Hellinger (RMSH) distance.\footnote{$h(p_{\theta,B}(\cdot|z),p_{0,W|Z}(\cdot|z))$ satisfies $h^{2}(p_{\theta,B}(\cdot|z),p_{0,W|Z}(\cdot|z)) = \int_{\mathcal{W}} (p_{\theta,B}^{1/2}(w|z)-p_{0,W|Z}^{1/2}(w|z))^{2}dw$.} Under Assumptions (ref).1, (ref), (ref), the uniform Lipschitz conditions hold: $\sup_{\gamma \in \Gamma}||m(\cdot,\gamma)-m_{0}(\cdot,\gamma)||_{n,2} \leq C_{1} h_{n,2}(p_{\theta,B},p_{0,W|Z})$ and $\sup_{\gamma \in \Gamma_{0}}||M(\cdot,\gamma)-M_{0}(\cdot,\gamma)||_{n,2} \leq C_{2} h_{n,2}(p_{\theta,B},p_{0,W|Z})$ for constants $C_{1},C_{2}>0$. Consequently, a RMSH convergence rate is an upper bound on the rate of convergence for $m(\cdot,\gamma)$ and $M(\cdot,\gamma)$ to $m_{0}(\cdot,\gamma)$ and $M_{0}(\cdot,\gamma)$, respectively, in the empirical $L^{2}$ norm.\footnote{This Lipschitz property (and the resulting convergence implications) seems to be where compactness of $\mathcal{W}$ is most crucial. Subject to modified restrictions on $\{f_{\theta}: \theta \in \Theta\}$, I conjecture that compactness could be relaxed the following ways: (i) boundedness of $g_{j}(W,Z,\gamma)$ and its derivatives in $\gamma$, (ii) uniform integrability of $\{g_{j}(W,z,\gamma): (z,\gamma) \in \mathcal{Z} \times \Gamma\}$ (and similarly the derivatives in $\gamma$, except replacing $\Gamma$ with $\Gamma_{0}$) with respect to $p_{0,W|Z}$ and $\{p_{\theta,B}: \theta \in \Theta, \ ||B||_{\infty} \leq \bar{B}\}$ (or a subset on which the posterior concentrates as $n\rightarrow \infty$), or (iii) convergence around $p_{0,W|Z}$ under a stronger norm (e.g., $L^{2}$). The intuition is that (i) and (ii) preserves RMSH convergence implying $||\cdot||_{n,2}$-convergence of conditional expectations, while (iii) could relax boundedness via, in the case of $L^{2}$, the Cauchy-Schwarz inequality. } This means that Assumptions (ref).3(a) and (ref).3(b) are verified if the RMSH rate is $o(n^{-1/4})$ (by restricting $\tilde{\mathcal{P}}_{n,W|Z}$ to be contained in an appropriately shrinking RMSH ball around $p_{0,W|Z}$).

Let $\mathcal{H}$ be the Reproducing Kernel Hilbert Space (RKHS) of the unrestricted centered GP (i.e., corresponding to $\tilde{\Pi}_{\mathcal{B}}$), and let $||\cdot||_{\mathcal{H}}$ denote the RKHS norm (i.e., $||\cdot||_{\mathcal{H}}^{2} = \langle \cdot,\cdot \rangle_{\mathcal{H}}$, where $\langle \cdot,\cdot \rangle_{\mathcal{H}}$ is the RKHS inner product).\footnote{See van2008reproducing for a formal definition of the RKHS of a centered GP.} The RMSH rate is determined by the RKHS. To formalize this, recall that $b_{0,\theta}(u,z) = \log p_{0,W|Z}(F_{\theta,z}^{-1}(u)|z)-\log f_{\theta}(F_{\theta,z}^{-1}(u)|z)$ for each $(u',z')' \in [0,1]^{d}$, and define the supremum norm concentration function at $b_{0,\theta}$ as

align*[align* omitted — 196 chars of source]

for $\delta > 0$. The concentration function uses the RKHS to measure the prior mass around $b_{0,\theta}$ (see Lemma 5.3 in van2008reproducing). Building on ghosal2007convergence and vaart2008rates, Proposition (ref) states that the RMSH rate $\delta_{n}$ satisfies $\sup_{\theta \in \Theta}\varphi_{b_{0,\theta}}(\delta_{n}) \leq n \delta_{n}^{2}$. For notation, let $\overline{\mathcal{H}}$ denote the $||\cdot||_{\infty}$-closure of $\mathcal{H}$.

propositionSuppose Assumptions (ref), (ref), (ref), and (ref) holds. If 1. $\sup_{\theta \in \Theta}||b_{0,\theta}||_{\infty} \leq \bar{B}$, 2. $\{b_{0,\theta}: \theta \in \Theta\} \subseteq \overline{\mathcal{H}}$, and 3. $\sup_{\theta \in \Theta}\varphi_{b_{0,\theta}}(\delta_{n}) \leq n \delta_{n}^{2}$ for some $\{\delta_{n}\}_{n \geq 1}$ satisfying $\delta_{n} \rightarrow 0$ and $n\delta_{n}^{2} \rightarrow \infty$ as $n\rightarrow \infty$, then, for $P_{0,Z}^{\infty}$-almost every fixed realization $\{z_{i}\}_{i \geq 1}$ of $\{Z_{i}\}_{i \geq 1}$, $\Pi(h_{n,2}(p_{\theta,B},p_{0,W|Z}) < D \delta_{n}| W^{(n)}) \rightarrow 1$ in $P_{0,W|Z}^{(n)}$-probability as $n\rightarrow \infty$ for some $D>0$.

Proposition (ref) is not sufficient for Assumption (ref).3(c). Proposition (ref) below proves that, under smoothness conditions, an RMSH rate leads to a convergence rate in the empirical supremum Hellinger distance $h_{n,\infty}(p_{\theta,B},p_{0,W|Z})=\max_{1 \leq i \leq n} h(p_{\theta,B}(\cdot|z_{i}),p_{0,W|Z}(\cdot|z_{i}))$.\footnote{The smoothness conditions on $f_{\theta}$ hold, for example, if $f_{\theta}(w|z)$ is a homoskedastic Gaussian linear regression (truncated to $\mathcal{W}$), with bounded $z$, and with compactly supported slopes and variances (and variances bounded away from zero).} By Assumptions (ref), (ref), and (ref), this leads to conditions for Assumption (ref).3(c) by a similar Lipschitz property (i.e., $\sup_{\gamma \in \Gamma}||\Sigma(\cdot,\gamma)-\Sigma_{0}(\cdot,\gamma)||_{n,\infty} \leq C_{3} h_{n,\infty}(p_{\theta,B},p_{0,W|Z})$ for some $C_{3} > 0$).\footnote{A similar conjecture to Footnote (ref) applies.} It is important to note the $h_{n,\infty}$ rate is only an upper bound and sharper convergence rates may be possible; I leave full treatment of optimal $h_{n,\infty}$ rates for future research because, to the best of my knowledge, posterior contraction rates in the supremum distance for conditional densities is an open question (and is of general statistical interest).

propositionSuppose Assumptions (ref), (ref), (ref), and (ref) hold, the conditions of Proposition (ref) hold, $p_{0,W|Z}$ satisfies $\int_{\mathcal{W}} ||\log p_{0,W|Z}(w|\cdot)||_{\alpha_{0}}dw < \infty$ for some $\alpha_{0} > d_{z}/2$, and $\{f_{\theta}: \theta \in \Theta\}$ satisfies $\sup_{\theta \in \Theta} \int_{\mathcal{W}}|| \log f_{\theta}(w|\cdot) ||_{a}dw < \infty$ and $\sup_{\theta \in \Theta}\int_{\mathcal{W}}||F_{\theta,\cdot}(w)||_{a}^{\lfloor a \rfloor}dw < \infty$ for each $a \in (0,\alpha)$. Then, for $P_{0,Z}^{\infty}$-almost every fixed realization $\{z_{i}\}_{i \geq 1}$ of $\{Z_{i}\}_{i \geq 1}$, there exists $\{\tilde{\mathcal{P}}_{1,n,W|Z}\}_{n \geq 1}$ and $\{\tilde{\delta}_{n}\}_{n \geq 1}$ such that $\Pi(\tilde{\mathcal{P}}_{n,1,W|Z}|W^{(n)}) \rightarrow 1$ in $P_{0,W|Z}^{(n)}$-probability as $n\rightarrow \infty$ and $\sup_{p_{\theta,B} \in \tilde{\mathcal{P}}_{1,n,W|Z}}h_{n,\infty}(p_{\theta,B},p_{0,W|Z})\leq \tilde{D}\tilde{\delta}_{n}$ for a constant $\tilde{D}>0$, where $\tilde{\delta}_{n}$ depends on $\alpha,\alpha_{0}$, $d_{z}$, and the RMSH rate $\delta_{n}$.

\addtocounter{example}{-1}

example[Continued] Suppose $\{b_{0,\theta}: \theta \in \Theta\} \subseteq C^{\alpha_{0}}([0,1]^{d},\mathbb{R}) \cap S^{\alpha_{0}}([0,1]^{d},\mathbb{R})$ for some $\alpha_{0} > 0$, $\sup_{\theta \in \Theta}||b_{0,\theta}||_{\alpha_{0}} < \infty$, and $\sup_{\theta \in \Theta}||b_{0,\theta}||_{2,2,\alpha_{0}} < \infty$. Then, Lemma 3 of van2011information and Lemma 13 in walker2026supplement imply that the $h_{n,2}$ rate is $\delta_{n} = n^{-\min\{\alpha,\alpha_{0}\}/(2 \alpha + d)}$. Consequently, $\delta_{n} = o(n^{-1/4})$ iff $\alpha > d/2$ and $\alpha_{0} > \alpha/2 + d/4$. The second inequality restricts oversmoothing because, if $\alpha \leq \alpha_{0}$ (i.e., undersmoothing), the contraction rate is $\delta_{n} = n^{-\alpha/(2\alpha+d)}$ and $\delta_{n} = o(n^{-1/4})$ if $\alpha > d/2$. If $\alpha = \alpha_{0}$, then $\delta_{n}$ is the minimax rate $n^{-\alpha_{0}/(2\alpha_{0}+d)}$ and $\delta_{n} = o(n^{-1/4})$ if $\alpha_{0} > d/2$. Substituting $\delta_{n} = n^{-\min\{\alpha,\alpha_{0}\}/(2\alpha + d)}$ into the $h_{n,\infty}$ rate $\tilde{\delta}_{n}$, one can show $\tilde{\delta}_{n} = o(n^{-1/4})$ if $\alpha/2 + 3d_{z}/4<\alpha_{0}$ and $1/4 + (d_{z}+(\alpha-a))/(2\alpha+d_{z}) < \min\{\alpha,\alpha_{0}\}/(2\alpha + d)$.\footnote{See the discussion following Lemma 13 in walker2026supplement for more details on the rate calculations.}

Verifying Prior Invariance

The next proposition states that ((ref)) holds if the chamberlain1987asymptotic optimal residual, $\tilde{g}_{0}(w,z,\gamma_{0}) = M_{0}(z,\gamma_{0})'\Sigma_{0}^{-1}(z,\gamma_{0})g(w,z,\gamma_{0})$, is well-approximated by sequences in $\mathcal{H}$. Let $\tilde{g}_{0,\theta}(u,z,\gamma_{0}) = \tilde{g}_{0}(F_{\theta,z}^{-1}(u),z,\gamma_{0}),z,\gamma_{0})$ for $(u,z,\theta) \in [0,1]^{d_{w}} \times [0,1]^{d_{z}} \times \Theta$ be the optimal residual transformed so that is defined on $[0,1]^{d}$.

propositionSuppose Assumptions (ref), (ref), (ref), (ref), (ref), (ref), and the conditions of Proposition (ref) hold. Further, suppose there are sequences $\{\tilde{g}_{n,\theta,j}^{*}: \theta \in \Theta\}_{n \geq 1}$, $j=1,...,d_{\gamma},$ in $ \mathcal{H}$ and $\{\zeta_{n}\}_{n\geq 1}$ in $\mathbb{R}_{++}$ with $\zeta_{n}\rightarrow 0$ such that 1. $\sup_{\theta \in \Theta}\max_{1 \leq j \leq d_{\gamma}}||\tilde{g}_{n,\theta,j}^{*}-\tilde{g}_{0,\theta,j}||_{\infty} \lesssim \zeta_{n}$, 2. $\sup_{\theta \in \Theta}\max_{1 \leq j \leq d_{\gamma}}||\tilde{g}_{n,\theta,j}^{*}||_{\mathcal{H}} \lesssim \sqrt{n}\zeta_{n}$, 3. $\sqrt{n}\zeta_{n}\delta_{n}\rightarrow 0$ as $n\rightarrow \infty$, where $\delta_{n}$ is the RMSH rate, and 4. $\{\Delta_{n,\theta}^{*}: \theta \in \Theta\}$ with $\Delta_{n,\theta}^{*}(w,z)=\tilde{g}_{n,\theta}^{*}(F_{\theta,z}(w),z)-\tilde{g}_{0,\theta}(F_{\theta,z}(w),z)$ for $(w,z) \in \mathcal{W} \times [0,1]^{d_{z}}$ satisfies, for $P_{0,Z}^{\infty}$-almost every fixed realization $\{z_{i}\}_{i\geq 1}$ of $\{Z_{i}\}_{i \geq 1}$, \begin{align*} \sup_{\theta \in \Theta}\left|\left|\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\{\Delta_{n,\theta}^{*}(W_{i},z_{i})-E_{P_{0,W|Z}}[\Delta_{n,\theta}^{*}(W,Z)|Z=z_{i}]\}\right|\right|_{2} \overset{P_{0,W|Z}^{(n)}}{\longrightarrow} 0 \end{align*} as $n\rightarrow \infty$. Then, for $P_{0,Z}^{\infty}$-almost every fixed realization $\{z_{i}\}_{i\geq 1}$ of $\{Z_{i}\}_{i \geq 1}$, there is $\{\tilde{\mathcal{P}}_{n,W|Z}\}_{n \geq 1}$ such that $\Pi(p_{W|Z} \in \tilde{\mathcal{P}}_{n,W|Z}|W^{(n)}) \rightarrow 1$ in $P_{0,W|Z}^{(n)}$-probability and ((ref)) holds.

Some comments on Proposition (ref). Sequences satisfying 1. and 2. can be found by solving $\sup_{\theta \in \Theta}\inf_{f \in \mathcal{H}: ||f-\tilde{g}_{0,\theta,j}||_{\infty} \leq \zeta_{n,j}}\frac{1}{2}||f||_{\mathcal{H}}^{2} \leq n \zeta_{n,j}^{2}$ for $j=1,...,d_{\gamma}$ and setting $\zeta_{n} = \max_{1 \leq j \leq d_{\gamma}}\zeta_{n,j}$. Since LHS is the component of the concentration function that describes the approximation of $\tilde{g}_{0,\theta,j}$ by $\mathcal{H}$, this can be checked for many GPs vaart2008rates,van2008reproducing,van2011information. The condition $\sqrt{n}\zeta_{n}\delta_{n}\rightarrow 0$ as $n\rightarrow \infty$ in 3. is known as a `no-bias' condition, and, since $\delta_{n} = o(n^{-1/4})$, it holds if $\zeta_{n} = o(n^{-1/4})$.\footnote{This label follows castillo2012gaussian,castillo2012semiparametric, rivoirard2012bernstein, and castillo2015bernstein, where similar conditions are likened to frequentist `no-bias' conditions (e.g., Section 25.8 of vaart_1998).} Example (ref) below demonstrates that this restricts the smoothness of $B$ relative to $\tilde{g}_{0,\theta}$. Part 4. amounts to a complexity constraint on $\{\Delta_{n,\theta}^{*}: \theta \in \Theta\}$ because $\sup_{\theta \in \Theta}||\Delta_{n,\theta}^{*}||_{\infty} = o(1)$ implies that the convergence holds pointwise. Finally, if $\{g_{0,\theta,j}: \theta \in \Theta\} \subseteq \mathcal{H}$ for each $j$, then Proposition (ref) holds trivially (i.e., $g_{n,\theta,j}^{*} = \tilde{g}_{0,\theta,j}$ and $\zeta_{n,j} = 0$ for each $j$), however, for many GPs, the RKHS comprises a relatively small class of functions. \addtocounter{example}{-1}

example[Continued] Suppose that $\{\tilde{g}_{0,\theta,j}: \theta \in \Theta\} \subseteq C^{\beta_{0,j}}(T,\mathbb{R}) \cap H^{\beta_{0,j}}([0,1]^{d},\mathbb{R})$ for some $\beta_{0,1},...,\beta_{0,d_{\gamma}} >0$, and $\sup_{\theta \in \Theta}||\tilde{g}_{0,\theta,j}||_{\beta_{0,j}} < \infty$ and $\sup_{\theta \in \Theta}||\tilde{g}_{0,\theta,j}||_{2,2,\beta_{0,j}} < \infty$ for $j=1,...,d_{\gamma}$. If $\beta_{0,j} < \alpha$ for some $j$, then Lemma 13 in walker2026supplement yields that $\zeta_{n} = n^{-\underline{\beta}/(2\alpha + d)}$, where $\underline{\beta} = \min_{1 \leq j \leq d_{\gamma}}\beta_{0,j}$, and $ \zeta_{n} = o(n^{-1/4})$ if $\underline{\beta} > \alpha/2+ d/4$. If $\alpha + d/2 >\underline{\beta} > \alpha$, then $\zeta_{n} = o(n^{-1/4})$ holds as long as $\alpha > d/2$. Finally, if $\underline{\beta} > \alpha + d/2$, then $\{\tilde{g}_{0,\theta,j}: \theta \in \Theta\} \subseteq \mathcal{H}$ for all $j$ and Proposition (ref) holds trivially.

Conclusion and Extensions

This paper proposes semiparametric Bayesian inference for conditional moment equalities. The central idea is that a posterior for a conditional distribution of data implies a posterior for a minimum distance estimand based on the conditional moments. The framework has similar flexibility to frequentist semiparametric estimators, and does not require converting the conditional moments to unconditional moments. I also establish the method's frequentist optimality via a BvM, providing conditions under which the posterior is asymptotically equivalent to a chamberlain1987asymptotic efficient estimator.

My paper offers several directions for future research. First, there are important settings in which $\gamma$ is infinite-dimensional. An example is nonparametric IV in which $W=(Y,D)'$, $Z = (X,A)'$, and $g(W,Z,\gamma) = Y-\gamma(D,X)$, where $\gamma(\cdot,\cdot)$ is an unknown function ai2003efficient,newey2003instrumental. Conceptually, the same ideas apply: a posterior for $P_{W|Z}$ implies a posterior for a function-valued $\gamma$. However, theoretical issues, such as non-compact function spaces, ill-posedness of identifying conditions, etc., may warrant different minimum distance estimands (e.g., penalized estimands like in chen2012estimation).

Second, model misspecification is an important concern for conditional moment equality models (i.e., when $E_{P_{W|Z}}[g(W,Z,\gamma)|Z] \neq 0$ for all $\gamma \in \Gamma$, with positive probability). The posterior for the value function $Q_{n}(\gamma_{n},\gamma_{n},p_{W|Z})$, a `$J$-statistic'-like estimand, contains information about the compatibility of the conditional moments with the data. Developing a Bayesian specification assessment framework using this posterior and comparing it with classical overidentifying restrictions tests would be interesting. Moreover, reporting the posterior distribution of $Q_{n}(\gamma_{n},\gamma_{n},p_{W|Z})$ may also connect with the recommendations in andrews2025purpose of reporting $J$-statistics (not $J$-tests) in overidentified models.

Finally, competing structural models often lead to nonnested conditional moment equalities $E_{P_{W|Z}}[g_{1}(W,Z,\gamma_{1})|Z] = 0$ and $E_{P_{W|Z}}[g_{2}(W,Z,\gamma_{2})|Z] = 0$. A posterior for $p_{W|Z}$ leads to a joint posterior for $(\gamma_{n,1},Q_{n,1},\gamma_{n,2},Q_{n,2})$, where $Q_{n,j}$, $j \in \{1,2\}$, is shorthand for $Q_{n,j}(\gamma_{n,j},\gamma_{n,j},p_{W|Z})$, with $Q_{n,j}(\gamma,\tilde{\gamma},p_{W|Z})$ being ((ref)) for $E_{P_{W|Z}}[g_{j}(W,Z,\gamma_{j})|Z]$. The posterior for $Q_{n,1}-Q_{n,2}$ can be used to assess which conditional moments are most plausible.