Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
77,703 characters · 21 sections · 122 citation commands
Semiparametric Bayesian Inference for a Conditional Moment Equality Model
Many economic parameters are identified via conditional moment equalities. These restrictions state that a finite-dimensional parameter of interest $\gamma $ uniquely solves $E_{P_{W|Z}}[g(W,Z,\gamma)|Z] = 0$, where $W $ and $Z $ are observed random vectors, $P_{W|Z}$ is the conditional distribution of $W $ given $Z$, and $g$ is a known vector-valued function. A canonical example is linear instrumental variables (IV) in which the regression model $Y=\gamma_{0}+\gamma_{1}D+\gamma_{2}X+U$ with $E[U|X,A] = 0$ leads to $E_{P_{W|Z}}[g(W,Z,\gamma)|Z]=0$ with $W=(Y,D)'$, $Z = (X,A)' $, and $g(W,Z,\gamma) = Y-\gamma_{0}-\gamma_{1}D-\gamma_{2}X$. Other examples include Euler equations hansen1982generalized, demand and supply models berry1995automobile, and peer effects models graham2008identifying, to name a few. Conditional moment equalities are also theoretically interesting because they impose overidentifying restrictions on $P_{W|Z}$. This makes efficient estimation nontrivial and has motivated important work on estimation using the chamberlain1987asymptotic optimal instruments newey1990efficient,NEWEY1993419,donald2001choosing,donald2003empirical,kitamura2004empirical,donald2009choosing,chen2021harmless,chib2022bayesian.
This paper makes three contributions. The first is introducing a semiparametric Bayesian inference framework for conditional moment equalities. Consistent with some approaches to conditional moment equalities (e.g., ai2003efficient), I leave $P_{W|Z}$ unrestricted and redefine $\gamma$ as a minimum distance estimand $\operatorname*{arg\,min}_{\gamma \in \Gamma}||E_{P_{W|Z}}[g(W,Z,\gamma)|Z]||^{2}$, where $||\cdot||$ is a norm. Observing that this defines a deterministic relationship between $P_{W|Z}$ and $\gamma$, a Bayesian framework arises in which a posterior for $P_{W|Z}$ implies a posterior for $\gamma$. The prior for $P_{W|Z}$ is supported over a nonparametric family, a feature ensuring that this Bayesian approach has similar flexibility to frequentist semiparametric estimators that do not rely on a tightly parametrized likelihood model. Moreover, by targeting $P_{W|Z}$, the posterior for $\gamma$ does not require converting the conditional moments into unconditional moments. This is appealing because ad hoc selections of unconditional moments can result in efficiency loss, and, more severely, some (efficient) choices can result in $\gamma$ not being identifiable.\footnote{See Example 2 of dominguez2004consistent for an instance where the unconditional moment equalities based on the optimal IVs have multiple roots, despite $\gamma$ uniquely minimizing the distance of the conditional moments to zero.} These benefits are compounded by the fact that, as a fully Bayesian procedure, it inherits all the basic benefits of Bayesian analysis (e.g., likelihood principle, expected utility decisions, simultaneous inference).
The second contribution is a semiparametric Bernstein-von Mises theorem (BvM) for the posterior for $\gamma$. BvMs provide conditions under which a posterior has the same limiting distribution as an efficient frequentist estimator, and are often invoked as a claim of large-sample robustness of the posterior to the choice of prior. Although BvMs hold under mild conditions in correctly specified regular parametric models, several papers find that these theorems can be delicate in semiparametric applications where the parameter of interest is a functional of a smooth nonparametric object, such as a probability density function, a regression function, etc. (see bickel2012semiparametric,castillo2012gaussian,castillo2012semiparametric,rivoirard2012bernstein,castillo2015bernstein,ray2020semiparametric). Since $\gamma$ is a functional of a nonparametric $P_{W|Z}$, these findings raise the possibility that the posterior for $\gamma$ may be sensitive to the choice of prior for $P_{W|Z}$ in large samples. To resolve these concerns, I prove a general semiparametric BvM for the posterior for $\gamma$ in Section (ref), providing conditions under which it converges to a normal distribution, centered at an efficient estimator, and with variance equal to the chamberlain1987asymptotic efficiency bound. As far as I am aware, this is the first fully Bayesian semiparametric BvM for the conditional moment equality model that does not require explicit conversion into unconditional moments.\footnote{See kato2013quasi and kankanala2025generalized for quasi-Bayesian results of a similar nature.} A byproduct of the BvM is that some Bayes estimators are asymptotically efficient frequentist estimators.
The third contribution is providing implementation details and verifying the BvM conditions for a flexible class of priors based on Gaussian processes (GP). GPs induce priors for conditional densities supported on the smoothness classes typically imposed on first-stage nuisance parameters, making concrete the flexibility of this Bayesian approach. Section (ref) demonstrates implementation for these priors with an empirical illustration focused on estimating welfare effects, an exercise capturing a common use of conditional moments (i.e., estimating model primitives and evaluating counterfactuals).\footnote{A complementary data-calibrated simulation comparing my method with alternatives is in Appendix (ref).} Section (ref) complements these numerical aspects by formally verifying the high-level BvM conditions for these GP priors. The key requirement is that the chamberlain1987asymptotic optimal residual is well-approximated by a reproducing kernel Hilbert space (RKHS), a `no-bias' condition that restricts the smoothness of the optimal residual relative to the GP.
This paper connects to several literatures in econometrics and statistics. The first is the Bayesian analysis of moment conditions. chamberlain2003nonparametric perform inference for structural parameters defined by unconditional moment equalities by solving the moment restrictions using draws from a Dirichlet process posterior for the data distribution. The key conceptual distinction between their approach and mine is that the former cannot be applied to conditional moment equalities with continuous $Z$ unless the conditional moments are converted into unconditional moments.\footnote{In an unpublished manuscript, chamberlain1995semiparametric consider conditional moment equalities, however their Dirichlet priors explicitly require the support of $Z$ to be small relative to the sample size, ruling out continuous $Z$.} chib2022bayesian is also relevant because they propose a Bayesian inference procedure for conditional moment equalities in which the conditional moments are converted into a sieve of unconditional moments, and inference is performed using the Bayesian exponentially tilted empirical likelihood (ETEL) framework of schennach2005bayesian. This conversion (and the use of ETEL) makes chib2022bayesian fundamentally different to my proposal. Other papers on Bayesian analysis of moments include lancaster2010bayesian, kitamura2011bayesian, pelenis2014bayesian, shin2015bayesian, bornn2018moment, and chib2018bayesian.
Another related literature is on semiparametric BvMs. castillo2015bernstein prove a general BvM for smooth functionals of nonparametric models, with substantial analysis devoted to those of unconditional probability density functions (building on rivoirard2012bernstein). I characterize $\gamma$ as a smooth functional of a conditional probability density function, and, via conditional asymptotics, I am able to modify their proof to show that the $\gamma$ posterior achieves the chamberlain1987asymptotic limit. chib2022bayesian also prove a BvM compatible with chamberlain1987asymptotic, however their proof differs from mine because, under appropriate conditions, the sieve of unconditional moment equalities makes parametric BvM arguments (cf. Theorem 10.1 of vaart_1998) applicable to the ETEL. Other contributions to the semiparametric BvM literature include shen2002asymptotic, bickel2012semiparametric, castillo2012gaussian,castillo2012semiparametric, norets2015bayesian, ray2020semiparametric, monard2021bernstein, breunig2025double,breunig2025semiparametricbayesiandifferenceindifferences, yiu2025semiparametric, and walker2024parametrization.
A third literature is efficient frequentist estimation of conditional moment equalities. The posterior for $\gamma$ is formalized as that of a conditional estimand that solves an efficiently weighted minimum distance problem. For this reason, my approach offers a Bayesian analog to the ai2003efficient framework (except I focus on finite-dimensional parameters). Moreover, the first-order conditions of the minimum distance estimand relates my framework to method of moments estimators that plug in estimators of the chamberlain1987asymptotic optimal instruments (e.g., those studied in robinson1987heteroskedasticity, newey1990efficient, NEWEY1993419, and chen2021harmless). Other frequentist approaches to optimal instruments estimation include Generalized Method of Moments (GMM) and Generalized Empirical Likelihood (EL) estimators based on sieves of unconditional moments donald2001choosing,donald2003empirical,donald2009choosing, and localized EL estimators ,kitamura2004empirical.
A final related literature is quasi-Bayes. Popularized in chernozhukov2003mcmc, quasi-Bayes is a sampling-based frequentist estimation framework in which an extremum objective forms a log quasi-likelihood, a prior is assumed for the parameter, and Bayes rule is used to obtain a quasi-posterior. kato2013quasi and kankanala2025generalized develop quasi-Bayesian frameworks for conditional moment equalities, and prove quasi-Bayesian BvMs for a structural function defined by conditional moment equalities and linear functionals of it, respectively. A key conceptual distinction between my approach and quasi-Bayes is that mine has a finite-sample Bayesian interpretation. Other papers on quasi-Bayesian estimation include liao2011posterior, florens2012nonparametric,florens2021gaussian, chen2018monte, andrews2022optimal, and kankanala2023quasi.
The paper is organized as follows. Section (ref) sets up the general Bayesian inference framework and provides a flexible class of priors, Section (ref) presents an empirical illustration, Section (ref) states the semiparametric BvM and related results, Section (ref) verifies the BvM conditions for the GP priors, and Section (ref) concludes with some extensions. The Appendix contains proofs of theorems and corollaries, technical details about implementation, a simulation study calibrated to the empirical application, and further discussion. Proofs of propositions, and lemmas, as well as some additional discussion are in the Online Appendix walker2026supplement. Notation is introduced when appropriate.
This section formalizes semiparametric Bayesian inference and describes implementation for a class of priors based on Gaussian processes.
I start with Bayesian inference for the conditional distribution $P_{W|Z}$ of $W$ given $Z$. Let $\{(W_{i}',Z_{i}')'\}_{i \geq 1}$, $W_{i} \in \mathcal{W} \subseteq \mathbb{R}^{d_{w}}$ and $Z_{i} \in \mathcal{Z} \subseteq \mathbb{R}^{d_{z}}$, be a sequence of random vectors for which the first $n \geq 1$ elements $\{(W_{i}',Z_{i}')'\}_{i=1}^{n}$ forms the observed data. The sampling model conditions on the realization $\{z_{i}\}_{i \geq 1}$ of $\{Z_{i}\}_{i \geq 1}$ and assumes that the elements of $W^{(n)}= (W_{1}',...,W_{n}')'$ satisfy $W_{i}|p_{W|Z} \overset{ind}{\sim} P_{W|z_{i}}$ for $i=1,...,n$, where $p_{W|Z}$ is a conditional probability density function, and $P_{W|z_{i}}$ is the probability distribution attached to $p_{W|Z}(\cdot|z_{i})$. Since the model parameters are conditional densities $p_{W|Z}$, the prior $\Pi$ is a (conditional) probability measure defined over the space $\mathcal{P}_{W|Z}$ of conditional densities (equipped with a $\sigma$-algebra $\mathscr{P}_{W|Z}$).\footnote{The qualifier `conditional' is used because implicitly $\Pi$ also defined conditional on $\{Z_{i}\}_{i \geq 1}$.} There are no parametric restrictions on $\mathcal{P}_{W|Z}$, so $\Pi$ should be thought of as a probability measure over an infinite-dimensional space of conditional densities.
Assumptions (ref) and (ref) are standard regularity conditions that imply that the conditional distribution $\Pi(p_{W|Z}\in \cdot |W^{(n)})$ of $p_{W|Z}$ given $W^{(n)}$ (known as the posterior) has a version satisfying Bayes theorem,\footnote{See, for example, Section 1.3 of ghosal2017fundamentals for similar regularity conditions.}
Since $\Pi(p_{W|Z} \in \cdot | W^{(n)})$ contains an inference theory for (transformations of) $p_{W|Z}$, inference for $\gamma$ is obtained upon formalizing the estimand $\operatorname*{arg\,min}_{\gamma \in \Gamma}||E_{P_{W|Z}}[g(W,Z,\gamma)|Z]||^{2}$.
The estimand is an iterated minimum distance estimand. For a given conditional density $p_{W|Z}$, let $m(z_{i},\gamma) = E_{P_{W|Z}}[g(W,Z,\gamma)|Z =z_{i}]$ be the $d_{g}\times 1$ vector of conditional moments at $z_{i}$, let $\Sigma(z_{i},\gamma) = E_{P_{W|Z}}[g(W,Z,\gamma)g(W,Z,\gamma)'|Z=z_{i}]$ be the $d_{g}\times d_{g}$ matrix of conditional second moments at $z_{i}$, and, for $\tilde{\gamma} \in \Gamma$ fixed, define a conditional estimand $$q_{n}(\tilde{\gamma},p_{W|Z}) = \operatorname*{arg\,min}_{\gamma \in \Gamma}Q_{n}(\gamma,\tilde{\gamma},p_{W|Z}), $$ where\footnote{The terminology `conditional estimand' follows abadie2014inference.}
To eliminate dependence on $\tilde{\gamma}$, I adopt a similar technique to hansen2021inference and consider an iterated estimand defined by the recursive relation
where $\gamma_{n,0}$ is some element of $\Gamma$ (e.g., the identity-weighted estimand). Notice that $\gamma_{n}$ does not require converting the conditional moments into unconditional moments. Further, since $\gamma_{n}$ satisfies the first-order conditions $n^{-1}\sum_{i=1}^{n}M(z_{i},\gamma_{n})'\Sigma^{-1}(z_{i},\gamma_{n})m(z_{i},\gamma_{n})=0$, it implicitly plugs in an estimate $M(z_{i},\gamma_{n})' \Sigma^{-1}(z_{i},\gamma_{n})$ of the chamberlain1987asymptotic optimal instruments, where $M(z_{i},\gamma)$ is the $d_{g} \times d_{\gamma}$ Jacobian of $m(z_{i},\gamma)$ in $\gamma$. This is important for the efficiency guarantees in Section (ref).
Assumption (ref) guarantees that $\Pi(p_{W|Z} \in \cdot |W^{(n)})$ leads to a well-defined marginal posterior for $\gamma_{n}$ given by the pushforward measure $\Pi(\gamma_{n} \in \cdot |W^{(n)}) = \Pi(p_{W|Z} \in \cdot|W^{(n)}) \circ \gamma_{n}^{-1}$.\footnote{Two points. First, the pushforward satisfies $(\Pi(p_{W|Z} \in \cdot|W^{(n)}) \circ \gamma_{n}^{-1})(A) = \Pi(p_{W|Z}: \gamma_{n} \in A|W^{(n)})$ for events $A$. Second, Appendix (ref) contains sufficient conditions for Assumption (ref) that may be of independent interest.} This formalizes the posterior for $\gamma$. Like $\Pi(p_{W|Z} \in \cdot | W^{(n)})$, it leads to inference about $\gamma_{n}$ and any function $f(\gamma_{n})$. For example, in the leading case where $f(\gamma)$ is scalar, a point estimate can be obtained using the posterior median $c_{n,f}(0.5)$, while uncertainty can be quantified using a $(1-\alpha)$-equitailed probability interval $CS_{n}(1-\alpha) = [c_{n,f}(\alpha/2),c_{n,f}(1-\alpha/2)]$ (a type of credible set), where $\alpha \in (0,1/2)$, and $c_{n,f}(q)$, $q \in (0,1)$, is the $q$-quantile of $\Pi(f(\gamma_{n}) \in \cdot | W^{(n)})$. Section (ref) establishes the asymptotic equivalence of these Bayes estimators and frequentist estimators based on the chamberlain1987asymptotic optimal instruments.
I present a flexible class of priors for $p_{W|Z}$. The priors are based on logistic transformations of Gaussian processes (GPs), and are a leading class of priors in the Bayesian nonparametrics literature tokdar2007towards,TOKDAR200734,vaart2008rates,tokdar2010bayesian,vehtari2014laplace. Specifically, the conditional density of $W$ given $Z$ is modeled as an infinite-dimensional exponential family
where $f_{\theta}(\cdot|z)$, $\theta \in \Theta \subseteq \mathbb{R}^{d_{\theta}}$ with $d_{\theta} < \infty$, is a conditional density, $F_{\theta,z}(\cdot)$ is the corresponding $d_{w}\times 1$ vector of conditional cumulative distribution functions, $F_{0}(\cdot)$ is a known invertible mapping taking values in $[0,1]^{d_{z}}$, and $B: [0,1]^{d} \rightarrow \mathbb{R}$ is a function with $d = d_{w}+ d_{z}$.\footnote{Let $f_{\theta,j}(w_{j}|w_{1},...,w_{j-1},z)$ denote the density of $W_{j}$ given $\{W_{j'}\}_{j'=1}^{j-1}$ and $Z$ under $f_{\theta}(w|z)$. The elements of $F_{\theta,z}(w)=(F_{\theta,1,z}(w),...,F_{\theta,d_{w},z}(w))$ satisfy $F_{\theta,j,z}(w) = \int_{-\infty}^{w_{j}} f_{\theta,j}(t|w_{1},...,w_{j-1},z)dt$. Using this and the change of variables $u=F_{\theta,z}(w)$, one can show $\int_{\mathcal{W}}p_{\theta,B}(w|z)dw = 1$ for each $z$.} Since $(\theta,B)$ determines $p_{\theta,B}$, a prior for $p_{\theta,B}$ is implied by a prior $\Pi$ for $(\theta,B)$. I set $\Pi = \Pi_{\Theta} \otimes \Pi_{\mathcal{B}}$, where $\Pi_{\Theta}$ is a probability distribution over $\Theta$, and $\Pi_{\mathcal{B}}$ is the probability law of a centered GP. To define the latter, let $(C([0,1]^{d},\mathbb{R}),||\cdot||_{\infty})$ be the space of continuous $h: [0,1]^{d} \rightarrow \mathbb{R}$ with $||h||_{\infty} = \sup_{t \in [0,1]^{d}}|h(t)|$.
Logistically transformed GPs are flexible. Fixing $\theta$ and imposing $z \in [0,1]^{d_{z}}$ for simplicity, $p_{0,W|Z}$ is in the support of the prior if $\Pi_{\mathcal{B}}(||B-b_{0,\theta}||_{\infty} < \delta) > 0$ for each $\delta > 0$, where $b_{0,\theta}(t) = \log p_{0,W|Z}(F_{\theta,z}^{-1}(u)|z) - \log f_{\theta}(F_{\theta,z}^{-1}(u)|z)$ for $t = (u',z')' \in [0,1]^{d}$ is $\log(p_{0,W|Z}/f_{\theta})$ with $w$ transformed to be defined on $[0,1]^{d_{w}}$.\footnote{Two comments. First, the support claim holds because the Kullback-Leibler support of the prior for $p_{\theta,B}$, which is the standard definition of the support of a nonparametric prior, is determined by the uniform support of $\Pi_{\mathcal{B}}$ (Lemma 3.1 in vaart2008rates). Second, to allow for $z \in \mathcal{Z}\nsubseteq [0,1]^{d_{z}}$, replace $z$ with $F_{0}^{-1}(v)$ where $v \in [0,1]^{d_{z}}$.} For common GPs, this condition is typically satisfied if $b_{0,\theta}$ is continuous; however, to ensure certain frequentist properties for statistical functionals (i.e., $\gamma_{n}$), continuity of $b_{0,\theta}$ is often strengthened to H\"{o}lder or Sobolev-type smoothness restriction on $b_{0,\theta}$.\footnote{The H\"{o}lder space $C^{\alpha}([0,1]^{d},\mathbb{R})$, $\alpha > 0$, is the set of functions $f: [0,1]^{d} \rightarrow \mathbb{R}$ that are $\lfloor \alpha \rfloor$-times differentiable with bounded derivatives, and, additionally, the $\lfloor \alpha \rfloor$th derivative is $(\alpha-\lfloor \alpha \rfloor)$-H\"{o}lder continuous. It is equipped with norm $||f||_{\alpha} = \max_{|k| \leq \lfloor \alpha \rfloor}||D^{k}f||_{\infty}+ \max_{|k| = \lfloor \alpha \rfloor}\sup_{t,s \in [0,1]^{d}: t \neq s}\frac{|(D^{k}f)(t)-(D^{k}f)(s)|}{||t-s||_{2}^{\alpha-\lfloor\alpha \rfloor}}$, where $k=(k_{1},...,k_{d})'$ is a vector of nonnegative integers, $|k| = \sum_{j=1}^{d}k_{j}$, $D^{k}$ is the differential operator, and $||\cdot||_{2}$ is the Euclidean norm. The Sobolev space $S^{\alpha}([0,1]^{d},\mathbb{R})$, $\alpha > 0$, comprises functions $f:[0,1]^{d} \rightarrow \mathbb{R}$ that are restrictions of functions $f: \mathbb{R}^{d} \rightarrow \mathbb{R}$ with Fourier transforms $\hat{f}$ such that $||f||_{2,2,\alpha}^{2} = \int_{\mathbb{R}^{d}} (1+||\lambda||_{2}^{2})^{\alpha}|\hat{f}(\lambda)|^{2} < \infty$. The Sobolev space norm is $||\cdot||_{2,2,\alpha}$. } Similar smoothness conditions are often imposed on first-stage nuisance parameters in frequentist estimation of conditional moment equalities (e.g., newey1990efficient,NEWEY1993419, ai2003efficient, kankanala2025generalized).
Posterior computation for logistic GP priors proceeds as follows. The MCMC algorithms proposed in tokdar2007towards and tokdar2010bayesian can be used to obtain $S \geq 1$ draws $\{(\theta^{[s]},B^{[s]})\}_{s=1}^{S}$ from the posterior for $(\theta,B)$ (see Appendix (ref)). Given a posterior draw $(\theta^{[s]},B^{[s]})$, the conditional moments $m^{[s]}(z_{i},\gamma)$ and $\Sigma^{[s]}(z_{i},\gamma)$ at $z_{i}$, $i=1,...,n$, can be estimated using importance sampling: for each $i$, 1. generate $\{W_{i,j}^{[s]}\}_{j=1}^{J}\overset{iid}{\sim} f_{\theta^{[s]}}(\cdot|z_{i})$, 2. compute importance weights $\{\omega_{i,j}^{[s]}\}_{j=1}^{J}$,
and 3. compute importance-weighted averages
A draw $\gamma_{n}^{[s]}$ from $\Pi(\gamma_{n} \in \cdot | W^{(n)})$ is then obtained by solving the iterated minimum distance problem, replacing $m(z_{i},\gamma)$ and $\Sigma(z_{i},\tilde{\gamma})$ in ((ref)) with $m_{J}^{[s]}(z_{i},\gamma)$ and $\Sigma_{J}^{[s]}(z_{i},\tilde{\gamma})$, respectively. Performing these steps over $\{(\theta^{[s]},B^{[s]})\}_{s=1}^{S}$ leads to a sample $\{\gamma_{n}^{[s]}\}_{s=1}^{S}$ from $\Pi(\gamma_{n} \in \cdot | W^{(n)})$. A sample from $\Pi(f(\gamma_{n}) \in \cdot | W^{(n)})$ is obtained by computing $\{f(\gamma_{n}^{[s]})\}_{s=1}^{S}$ using $\{\gamma_{n}^{[s]}\}_{s=1}^{S}$. Bayes estimators (e.g., posterior medians, credible sets) are computed using the empirical distribution of $\{f(\gamma_{n}^{[s]})\}_{s=1}^{S}$.
This section uses my proposal to estimate welfare effects of price changes. A data-calibrated simulation that compares my approach with alternatives is in Appendix (ref).
I use the 2001 National Household Travel Survey gasoline demand dataset from blundell2012measuring,blundell2017nonparametric.\footnote{The dataset was obtained from the publicly available replication files of chen2018optimal. They can be found using the link \url{https://github.com/timothymchristensen/NPIV}.} This dataset contains household gasoline consumption $Q$ (in gallons), average price $P$ of gasoline (in dollars per gallon) in the household's county, household income $Y$ (in dollars), and distance $A$ (in 1,000 kilometers) of a household's state capital to a major oil platform in the Gulf of Mexico. There are $n=4,812$ households in the dataset.
The parameter of interest is the welfare effect of a gasoline price change (e.g., due to taxation). Household gasoline demand is modeled using a constant elasticity specification
where the error term $U$ satisfies $E[U|\log Y, \log A] = 0$.\footnote{Distance as an excluded IV is proposed in blundell2012measuring and is implemented in chen2018optimal too. A justification is that distance is a cost-shifter in that it affects prices only through its impact on firm transportation costs.} The welfare effect of a price change from $p^{0}$ to $p^{1}$ at income $y$ is measured using deadweight loss (DWL),
where $S(p^{0},y,\gamma)$ is consumer surplus hausman1981exact,hausman1995nonparametric, and $q(p,y,\gamma) = \exp(\gamma_{0}+\gamma_{1} \log p + \gamma_{2} \log y)$. walker2026supplement details DWL calculation.
Let $W = (\log Q ,\log P)'$, $Z = (\log Y, \log A)'$, and $g(W,Z,\gamma) = \log Q - \gamma_{0} - \gamma_{1} \log P -\gamma_{2} \log Y$. Since $E[U|\log Y, \log A] = 0$ implies $E_{P_{W|Z}}[g(W,Z,\gamma)|Z] = 0$, Bayesian inference for $\gamma$ can be obtained via the approach in Section (ref) (in fact, it is an instance of Example (ref)). Furthermore, since $DWL(p^{0},p^{1},y,\gamma)$ is a deterministic function of $\gamma$ (as the researcher sets $(p^{0},p^{1},y)$), one automatically obtains a posterior for $DWL(p^{0},p^{1},\gamma)$. This highlights that my framework offers simultaneous inference for model primitives ($\gamma$) and counterfactuals ($DWL(p^{0},p^{1},y,\gamma)$).
Figure (ref) and Table (ref) report the posterior densities and Bayes estimates, respectively, of the DWL for $p^{0} = \$1.22$, $p^{1} = \$1.44$, and $y \in \{\$42500,\$57500,\$72500\}$.\footnote{These price-income combinations are the same as those reported in blundell2012measuring.} The posteriors are based on a logistic GP prior for $p_{W|Z}$ (from Section (ref)), with a homoskedastic Gaussian linear regression of $W$ on $Z$ for $f_{\theta}$, diffuse priors for the location-scale parameters $\theta$, and a Mat\'{e}rn GP for $B$ with $\alpha = 5/2$.\footnote{Two comments. First, Appendix (ref) contains more details about the prior. Second, setting $\alpha = 5/2$ follows recommendations in williams2006gaussian; walker2026supplement shows changing $\alpha$ has a modest effect on the estimates.} Figure (ref) indicates that the DWL posteriors are approximately symmetric and bell-shaped across all income groups. The estimates in Table (ref) reveal that DWL as a percentage of tax paid is almost identical across the income groups, however, as a proportion of income, deadweight loss decreases monotonically with income level. These findings are qualitatively similar to the constant elasticity results in blundell2012measuring, though my results relax price exogeneity.
This section proves a general BvM, proving that $\Pi(\gamma_{n} \in \cdot|W^{(n)})$ is asymptotically Gaussian and achieves the chamberlain1987asymptotic semiparametric efficiency bound.
I start with assumptions about the DGP. For notation, let $||\cdot||_{2}$ be the Euclidean norm, let $||\cdot||_{op}$ be the operator norm (i.e., the maximum singular value), let $\lambda_{min}(A)$ and $\lambda_{max}(A)$ be the minimum and maximum eigenvalues, respectively, of matrix $A$, let $\text{int}(B)$ be the interior of a set $B$, and let $vec(A)$ be the vectorized matrix $A$. A function class is Glivenko-Cantelli if it obeys a uniform strong law of large numbers (USLLN).
Assumptions (ref)--(ref) are similar restrictions to those encountered in the optimal instruments literature. Assumption (ref) imposes the same sampling restriction as chamberlain1987asymptotic, and, relative to Section (ref), is a restriction on $\{Z_{i}\}_{i \geq 1}$. Assumption (ref) imposes correct specification, requires that $\gamma_{0}$ be point identified, and that $\gamma_{0}$ belongs to the interior of a compact $\Gamma \subseteq \mathbb{R}^{d_{\gamma}}$. Misspecification is discussed in Section (ref). Assumption (ref) imposes some smoothness restrictions on $\gamma \mapsto m_{0}(z,\gamma)$. It accommodates the differentiability conditions on $g(W,Z,\gamma)$ that are routinely imposed in the optimal instruments literature (e.g., chamberlain1987asymptotic, newey1990efficient,NEWEY1993419, donald2003empirical, and chib2022bayesian), while allowing settings where $g(W,Z,\gamma)$ is nondifferentiable but $m_{0}(Z,\gamma)$ is differentiable (e.g., quantile IV chernozhukov2005iv,chernozhukov2006instrumental). Assumption (ref) imposes a compactness restriction on $\{\Sigma_{0}(z,\gamma): (z,\gamma) \in \mathcal{Z} \times \Gamma\}$ and a mild smoothness requirement for $\gamma \mapsto \Sigma_{0}(z,\gamma)$. Assumption (ref).1 requires $Q_{n}(\gamma,\tilde{\gamma},p_{0,W|Z})$ obeys a USLLN over $\Gamma \times \Gamma$, while Assumption (ref).2, combined with Assumption (ref).1, implies that the criterion $Q_{P_{0,Z}}(\gamma,\tilde{\gamma},p_{0,W|Z}) := E_{P_{0,Z}}[m_{0}(Z,\gamma)'\Sigma_{0}^{-1}(Z,\tilde{\gamma})m_{0}(Z,\gamma)]$ has a uniformly well-separated minimum at $\gamma_{0}$, an important condition for consistent estimation of $\gamma_{0}$. Assumption (ref).1 is a USLLN condition for the class $\{V_{0}(\cdot,\tilde{\gamma}): \tilde{\gamma} \in \Gamma\}$, and Assumption (ref).2 is a strong identification condition that, combined with Assumption (ref).1, implies $V_{0} :=E_{P_{0,Z}}[V_{0}(Z,\gamma_{0})]$ is positive definite, so the chamberlain1987asymptotic efficiency bound, $V_{0}^{-1}$, is finite.
The next assumption concerns $\Pi(p_{W|Z} \in \cdot |W^{(n)})$. For notation, let $||f||_{n,2}^{2} = n^{-1}\sum_{i=1}^{n}||f(z_{i})||^{2}$ and $||f||_{n,\infty} = \max_{1 \leq i \leq n }||f(z_{i})||$, where $||\cdot||$ is $||\cdot||_{2}$ if $f(z)$ is a vector and is $||\cdot||_{op}$ if $f(z)$ is a matrix, let $$\tilde{\chi}_{0}(W,Z) = -V_{0}^{-1}M_{0}(Z,\gamma_{0})'\Sigma_{0}^{-1}(Z,\gamma_{0})g(W,Z,\gamma_{0}) $$ be the chamberlain1987asymptotic efficient influence function (EIF) and note that $V_{0}^{-1}$ also satisfies $V_{0}^{-1} = E_{P_{0,WZ}}[\tilde{\chi}_{0}(W,Z)\tilde{\chi}_{0}(W,Z)']$, let $\overset{P_{0,W|Z}^{(n)}}{\longrightarrow}$ denote convergence in probability under $P_{0,W|Z}^{(n)}=\bigotimes_{i=1}^{n}P_{0,W|z_{i}}$, let $P_{0,Z}^{\infty}$ be the probability law of the i.i.d sequence $\{Z_{i}\}_{i \geq 1}$ (i.e., $P_{0,Z}^{\infty} = \bigotimes_{i \geq 1}P_{0,Z}$), and sequences $\{x_{n}\}_{n \geq 1}$ and $\{y_{n}\}_{n \geq 1}$ satisfy $x_{n} = o(y_{n})$ if $x_{n}/y_{n} \rightarrow 0$ as $n\rightarrow \infty$ and $x_{n} = O(y_{n})$ if there is a constant $C > 0$ such that $|x_{n}/y_{n}| \leq C$ for large $n$.
Assumption (ref) states that, conditional on $\{Z_{i}\}_{i \geq 1}$, there are sets $\{\tilde{\mathcal{P}}_{n,W|Z}\}_{n \geq 1}$ on which the posterior concentrates as $n\rightarrow \infty$, and these sets are structured enough to ensure certain smoothness (in $\gamma$) and convergence guarantees for $m$, $M$, and $\Sigma$. The assumption is general in that it is stated for an arbitrary prior $\Pi$ for $p_{W|Z}$; I derive examples of $\{\tilde{\mathcal{P}}_{n,W|Z}\}_{n \geq 1}$ for certain GPs in Section (ref). As for the contents, Parts 3(a) and 3(b) state that $m(\cdot,\gamma)$ and $M(\cdot,\gamma)$ converge to their true counterparts $m_{0}(\cdot,\gamma)$ and $M_{0}(\cdot,\gamma)$, respectively, at a rate faster than $n^{-1/4}$ under the empirical $L^{2}$ norm. Part 3(c) states that $\Sigma(\cdot,\gamma)$ converges to $\Sigma_{0}(\cdot,\gamma)$ at a rate faster than $n^{-1/4}$ under the empirical supremum norm, a condition analogous to Assumption 3.4(iii) in ai2003efficient. Requiring $o(n^{-1/4})$ nuisance convergence rates is generally considered a mild condition in semiparametric applications. Part 3(d) and 3(e) are boundedness conditions that help show that $\tilde{\gamma} \mapsto q_{n}(\tilde{\gamma},p_{W|Z})$ is a uniform contraction mapping in large samples, a property that helps establish that $\gamma_{n}$ uniformly converges to $\gamma_{0}$ along $\{\tilde{\mathcal{P}}_{n,W|Z}\}_{n \geq 1}$. Part 3(f) requires that the EIF has sufficiently thin tails over $\tilde{\mathcal{P}}_{n,W|Z}$, a condition sufficient to control a remainder in an expansion of $\log L_{n}(p_{W|Z})$.
Assumptions (ref)--(ref) lead to a Bernstein-von Mises theorem (BvM). Informally, the BvM states that $\gamma_{n}|W^{(n)}\overset{a}{\sim} \mathcal{N}(\hat{\gamma}_{n},\frac{1}{n}V_{0}^{-1})$ for large $n$, where $\hat{\gamma}_{n} = \gamma_{0}+n^{-1}\sum_{i=1}^{n}\tilde{\chi}_{0}(W_{i},z_{i})$ is a hypothetical efficient estimator of $\gamma_{0}$. Consequently, the posterior for $\gamma_{n}$ behaves like a best regular estimator $\hat{\gamma}_{n}$ for large $n$, thereby establishing its frequentist asymptotic optimality.
The key to the BvM is an asymptotically linear representation of $\gamma_{n}$. Stated as a theorem below, it shows that $\gamma_{n}$ is asymptotically linear in $p_{W|Z}$ with Riesz representer given by $\tilde{\chi}_{0}$. It can be thought of as a Bayesian analogue to the asymptotically linear representation of an efficient estimator, and validates that $\Pi(\gamma_{n} \in \cdot | W^{(n)})$ accounts for the optimal instruments as $n\rightarrow \infty$.
Theorem (ref) establishes that, asymptotically, $\gamma_{n}$ is an (efficient) linear functional of $p_{W|Z}$. Consequently, a semiparametric BvM for $\gamma_{n}$ can be established by building on castillo2015bernstein, who prove BvMs for unconditional density functions with i.i.d data.
Several aspects of Theorem (ref) should be highlighted. First, Theorem (ref) implies that some Bayesian point estimators and credible sets are asymptotically efficient. Let $f: \mathbb{R}^{d_{\gamma}} \rightarrow \mathbb{R}$ be continuously differentiable at $\gamma_{0}$ and recall that $c_{n,f}(q)$, $q \in (0,1)$, is the $q$-quantile of $\Pi(f(\gamma_{n}) \in \cdot|W^{(n)})$. Corollary (ref) states that, conditional on $\{Z_{i}\}_{i \geq 1}$, $c_{n,f}(q)=f(\hat{\gamma}_{n}) + \Phi^{-1}(q)n^{-1/2}\Omega_{0,f}^{1/2}+o_{P_{0,W|Z}^{(n)}}(n^{-1/2})$, where $\Phi(\cdot)$ is the $\mathcal{N}(0,1)$ cumulative distribution function and $\Omega_{0,f} =\frac{\partial f(\gamma_{0})}{\partial \gamma'}V_{0}^{-1}\frac{\partial f(\gamma_{0})'}{\partial \gamma}$ is the asymptotic variance of $f(\hat{\gamma}_{n})$. Consequently, by setting $q= 0.5$, the posterior median $c_{n,f}(0.5)$ is first-order asymptotically equivalent to the efficient estimator $f(\hat{\gamma}_{n})$.\footnote{The restrictions on $f$ imply the delta method preserves efficiency (see Section 25.7 of vaart_1998).} Moreover, the equitailed probability interval $CS_{n,f}(1-\alpha)= [c_{n,f}(\alpha/2), c_{n,f}(1-\alpha/2)]$, where $\alpha \in (0,1/2)$, is first-order asymptotically equivalent to an efficient Wald confidence interval $[f(\hat{\gamma}_{n})+\Phi^{-1}(\alpha/2)\sqrt{\Omega_{0,f}/n},f(\hat{\gamma}_{n})-\Phi^{-1}(\alpha/2)\sqrt{\Omega_{0,f}/n}]$, thereby leading to best asymptotic frequentist uncertainty quantification.
A second aspect is that Theorem (ref) also holds unconditionally, where stochastic convergence is defined with respect to $P_{0,WZ}$. Corollary (ref) formalizes this, and, by extension, implies an unconditional version of Corollary (ref) (see Remark (ref)).
A final aspect worth discussing is condition ((ref)). These conditions are common in semiparametric BvMs (see castillo2012semiparametric,castillo2012gaussian, rivoirard2012bernstein, castillo2015bernstein, ray2020semiparametric, and monard2021bernstein, to name a few), and arise from expanding the log-likelihood ($\log L_{n}(p_{W|Z})$) around the least favorable parametric submodel ($p_{t,n,W|Z}$). The name `prior invariance' reflects that its verification typically requires the prior be roughly unchanged under shifts in the direction of the EIF. Indeed, Section (ref) verifies ((ref)) for GPs by finding a sequence $\{\tilde{g}_{n}\}_{n \geq 1}$ that uniformly approximates the chamberlain1987asymptotic optimal residual $\tilde{g}_{0} = -V_{0}\tilde{\chi}_{0}$, with $o(n^{-1/4})$ approximation error, and for which the law $\Pi_{n,t}$ of $p_{n,t,W|Z}\propto p_{W|Z}\exp(-t'V_{0}^{-1}\tilde{g}_{n}/\sqrt{n})$ under $\Pi$ satisfies $\Pi_{n,t} \ll \Pi$ and $d \Pi_{n,t}/d\Pi \rightarrow 1$ uniformly along $\{\tilde{\mathcal{P}}_{n,W|Z}\}_{n \geq 1}$ as $n\rightarrow \infty$. For the Mat\'{e}rn process (i.e., Example (ref)), this restricts the GP sample path smoothness relative to $\tilde{g}_{0}$.
I provide sufficient conditions for Assumption (ref) and ((ref)) for the priors from Section (ref).
Recall $p_{\theta,B}(w|z) \propto f_{\theta}(w|z)\exp(B(F_{\theta,z}(w),F_{0}(z)))$, where $f_{\theta}(\cdot|z)$, $\theta \in \Theta \subseteq \mathbb{R}^{d_{\theta}}$ with $d_{\theta} < \infty$, is a conditional density, $F_{\theta,z}(\cdot)$ is the $d_{w}\times 1$ vector of conditional CDFs attached to $f_{\theta}(\cdot|z)$, $F_{0}: \mathcal{Z} \rightarrow [0,1]^{d_{z}}$ is a known invertible transformation, and $B: T \rightarrow \mathbb{R}$ is a function. A prior for $p_{\theta,B}$ is obtained from a prior $\Pi$ for $(\theta,B)$. Also, recall that $(C^{\alpha}([0,1]^{d},\mathbb{R}),||\cdot||_{\alpha})$ and $(S^{\alpha}([0,1]^{d},\mathbb{R}),||\cdot||_{2,2,\alpha})$ denote H\"{o}lder and Sobolev spaces, respectively, of order $\alpha >0$ (see Footnote (ref) for precise details).
Assumptions (ref)--(ref) comprise sufficient conditions on the DGP and prior that will be used to verify Assumption (ref). The main limitation of Assumption (ref) is compactness of $\mathcal{W}$, however it is not overly restrictive relative to the literature on semiparametric BvMs for density functionals rivoirard2012bernstein,castillo2015bernstein, and, more broadly, posterior consistency with infinite-dimensional exponential families scricciolo2006convergence,rivoirard2012posterior. I expect compact $\mathcal{W}$ can be relaxed (see Footnote (ref)). Compactness of $\mathcal{Z}$ is a standard assumption in nonparametric conditional density estimation, and, given such an assumption, setting $\mathcal{Z} = [0,1]^{d_{z}}$ is without loss of generality (and enables setting $F_{0}(z)=z$). Assumption (ref) is also stronger than the high-level smoothness conditions imposed in Section (ref), as it rules out models with nondifferentiable $g(W,Z,\gamma)$. Assumption (ref) imposes regularity conditions on $\{f_{\theta}: \theta \in \Theta\}$ that are satisfied, for example, when $f_{\theta}$ is a Gaussian linear regression (truncated to $\mathcal{W}$), $\mathcal{Z}$ is compact (i.e. Assumption (ref).2), and $\Theta$ constrains the slope and variance to compact sets, with the variance bounded away from zero. Assumption (ref) is a sample path smoothness restriction that many Gaussian processes satisfy (e.g., the Mat\'{e}rn process from Example (ref)). Uniformly bounding $B$ is a technical device used to control random approximation errors, and, since the magnitude of $\bar{B}$ has no role, it can be viewed as an arbitrarily large but finite number.\footnote{Lemma 5.1 in van2008reproducing shows $\tilde{\Pi}_{\mathcal{B}}(||B||_{\infty} \leq \bar{B}) > 0$ for any $\bar{B} > 0$.}
Under Assumptions (ref) and (ref), only Assumptions (ref).3(a)--10.3(c) need verification. Let $h(p_{\theta,B}(\cdot|z),p_{0,W|Z}(\cdot|z))$ be the Hellinger distance between $p_{\theta,B}$ and $p_{0,W|Z}$ at $z$, and let $h_{n,2}(p_{\theta,B},p_{0,W|Z}) = (n^{-1}\sum_{i=1}^{n}h^{2}(p_{\theta,B}(\cdot|z_{i}),p_{0,W|Z}(\cdot|z_{i})))^{1/2}$ be the root mean square Hellinger (RMSH) distance.\footnote{$h(p_{\theta,B}(\cdot|z),p_{0,W|Z}(\cdot|z))$ satisfies $h^{2}(p_{\theta,B}(\cdot|z),p_{0,W|Z}(\cdot|z)) = \int_{\mathcal{W}} (p_{\theta,B}^{1/2}(w|z)-p_{0,W|Z}^{1/2}(w|z))^{2}dw$.} Under Assumptions (ref).1, (ref), (ref), the uniform Lipschitz conditions hold: $\sup_{\gamma \in \Gamma}||m(\cdot,\gamma)-m_{0}(\cdot,\gamma)||_{n,2} \leq C_{1} h_{n,2}(p_{\theta,B},p_{0,W|Z})$ and $\sup_{\gamma \in \Gamma_{0}}||M(\cdot,\gamma)-M_{0}(\cdot,\gamma)||_{n,2} \leq C_{2} h_{n,2}(p_{\theta,B},p_{0,W|Z})$ for constants $C_{1},C_{2}>0$. Consequently, a RMSH convergence rate is an upper bound on the rate of convergence for $m(\cdot,\gamma)$ and $M(\cdot,\gamma)$ to $m_{0}(\cdot,\gamma)$ and $M_{0}(\cdot,\gamma)$, respectively, in the empirical $L^{2}$ norm.\footnote{This Lipschitz property (and the resulting convergence implications) seems to be where compactness of $\mathcal{W}$ is most crucial. Subject to modified restrictions on $\{f_{\theta}: \theta \in \Theta\}$, I conjecture that compactness could be relaxed the following ways: (i) boundedness of $g_{j}(W,Z,\gamma)$ and its derivatives in $\gamma$, (ii) uniform integrability of $\{g_{j}(W,z,\gamma): (z,\gamma) \in \mathcal{Z} \times \Gamma\}$ (and similarly the derivatives in $\gamma$, except replacing $\Gamma$ with $\Gamma_{0}$) with respect to $p_{0,W|Z}$ and $\{p_{\theta,B}: \theta \in \Theta, \ ||B||_{\infty} \leq \bar{B}\}$ (or a subset on which the posterior concentrates as $n\rightarrow \infty$), or (iii) convergence around $p_{0,W|Z}$ under a stronger norm (e.g., $L^{2}$). The intuition is that (i) and (ii) preserves RMSH convergence implying $||\cdot||_{n,2}$-convergence of conditional expectations, while (iii) could relax boundedness via, in the case of $L^{2}$, the Cauchy-Schwarz inequality. } This means that Assumptions (ref).3(a) and (ref).3(b) are verified if the RMSH rate is $o(n^{-1/4})$ (by restricting $\tilde{\mathcal{P}}_{n,W|Z}$ to be contained in an appropriately shrinking RMSH ball around $p_{0,W|Z}$).
Let $\mathcal{H}$ be the Reproducing Kernel Hilbert Space (RKHS) of the unrestricted centered GP (i.e., corresponding to $\tilde{\Pi}_{\mathcal{B}}$), and let $||\cdot||_{\mathcal{H}}$ denote the RKHS norm (i.e., $||\cdot||_{\mathcal{H}}^{2} = \langle \cdot,\cdot \rangle_{\mathcal{H}}$, where $\langle \cdot,\cdot \rangle_{\mathcal{H}}$ is the RKHS inner product).\footnote{See van2008reproducing for a formal definition of the RKHS of a centered GP.} The RMSH rate is determined by the RKHS. To formalize this, recall that $b_{0,\theta}(u,z) = \log p_{0,W|Z}(F_{\theta,z}^{-1}(u)|z)-\log f_{\theta}(F_{\theta,z}^{-1}(u)|z)$ for each $(u',z')' \in [0,1]^{d}$, and define the supremum norm concentration function at $b_{0,\theta}$ as
for $\delta > 0$. The concentration function uses the RKHS to measure the prior mass around $b_{0,\theta}$ (see Lemma 5.3 in van2008reproducing). Building on ghosal2007convergence and vaart2008rates, Proposition (ref) states that the RMSH rate $\delta_{n}$ satisfies $\sup_{\theta \in \Theta}\varphi_{b_{0,\theta}}(\delta_{n}) \leq n \delta_{n}^{2}$. For notation, let $\overline{\mathcal{H}}$ denote the $||\cdot||_{\infty}$-closure of $\mathcal{H}$.
Proposition (ref) is not sufficient for Assumption (ref).3(c). Proposition (ref) below proves that, under smoothness conditions, an RMSH rate leads to a convergence rate in the empirical supremum Hellinger distance $h_{n,\infty}(p_{\theta,B},p_{0,W|Z})=\max_{1 \leq i \leq n} h(p_{\theta,B}(\cdot|z_{i}),p_{0,W|Z}(\cdot|z_{i}))$.\footnote{The smoothness conditions on $f_{\theta}$ hold, for example, if $f_{\theta}(w|z)$ is a homoskedastic Gaussian linear regression (truncated to $\mathcal{W}$), with bounded $z$, and with compactly supported slopes and variances (and variances bounded away from zero).} By Assumptions (ref), (ref), and (ref), this leads to conditions for Assumption (ref).3(c) by a similar Lipschitz property (i.e., $\sup_{\gamma \in \Gamma}||\Sigma(\cdot,\gamma)-\Sigma_{0}(\cdot,\gamma)||_{n,\infty} \leq C_{3} h_{n,\infty}(p_{\theta,B},p_{0,W|Z})$ for some $C_{3} > 0$).\footnote{A similar conjecture to Footnote (ref) applies.} It is important to note the $h_{n,\infty}$ rate is only an upper bound and sharper convergence rates may be possible; I leave full treatment of optimal $h_{n,\infty}$ rates for future research because, to the best of my knowledge, posterior contraction rates in the supremum distance for conditional densities is an open question (and is of general statistical interest).
\addtocounter{example}{-1}
The next proposition states that ((ref)) holds if the chamberlain1987asymptotic optimal residual, $\tilde{g}_{0}(w,z,\gamma_{0}) = M_{0}(z,\gamma_{0})'\Sigma_{0}^{-1}(z,\gamma_{0})g(w,z,\gamma_{0})$, is well-approximated by sequences in $\mathcal{H}$. Let $\tilde{g}_{0,\theta}(u,z,\gamma_{0}) = \tilde{g}_{0}(F_{\theta,z}^{-1}(u),z,\gamma_{0}),z,\gamma_{0})$ for $(u,z,\theta) \in [0,1]^{d_{w}} \times [0,1]^{d_{z}} \times \Theta$ be the optimal residual transformed so that is defined on $[0,1]^{d}$.
Some comments on Proposition (ref). Sequences satisfying 1. and 2. can be found by solving $\sup_{\theta \in \Theta}\inf_{f \in \mathcal{H}: ||f-\tilde{g}_{0,\theta,j}||_{\infty} \leq \zeta_{n,j}}\frac{1}{2}||f||_{\mathcal{H}}^{2} \leq n \zeta_{n,j}^{2}$ for $j=1,...,d_{\gamma}$ and setting $\zeta_{n} = \max_{1 \leq j \leq d_{\gamma}}\zeta_{n,j}$. Since LHS is the component of the concentration function that describes the approximation of $\tilde{g}_{0,\theta,j}$ by $\mathcal{H}$, this can be checked for many GPs vaart2008rates,van2008reproducing,van2011information. The condition $\sqrt{n}\zeta_{n}\delta_{n}\rightarrow 0$ as $n\rightarrow \infty$ in 3. is known as a `no-bias' condition, and, since $\delta_{n} = o(n^{-1/4})$, it holds if $\zeta_{n} = o(n^{-1/4})$.\footnote{This label follows castillo2012gaussian,castillo2012semiparametric, rivoirard2012bernstein, and castillo2015bernstein, where similar conditions are likened to frequentist `no-bias' conditions (e.g., Section 25.8 of vaart_1998).} Example (ref) below demonstrates that this restricts the smoothness of $B$ relative to $\tilde{g}_{0,\theta}$. Part 4. amounts to a complexity constraint on $\{\Delta_{n,\theta}^{*}: \theta \in \Theta\}$ because $\sup_{\theta \in \Theta}||\Delta_{n,\theta}^{*}||_{\infty} = o(1)$ implies that the convergence holds pointwise. Finally, if $\{g_{0,\theta,j}: \theta \in \Theta\} \subseteq \mathcal{H}$ for each $j$, then Proposition (ref) holds trivially (i.e., $g_{n,\theta,j}^{*} = \tilde{g}_{0,\theta,j}$ and $\zeta_{n,j} = 0$ for each $j$), however, for many GPs, the RKHS comprises a relatively small class of functions. \addtocounter{example}{-1}
This paper proposes semiparametric Bayesian inference for conditional moment equalities. The central idea is that a posterior for a conditional distribution of data implies a posterior for a minimum distance estimand based on the conditional moments. The framework has similar flexibility to frequentist semiparametric estimators, and does not require converting the conditional moments to unconditional moments. I also establish the method's frequentist optimality via a BvM, providing conditions under which the posterior is asymptotically equivalent to a chamberlain1987asymptotic efficient estimator.
My paper offers several directions for future research. First, there are important settings in which $\gamma$ is infinite-dimensional. An example is nonparametric IV in which $W=(Y,D)'$, $Z = (X,A)'$, and $g(W,Z,\gamma) = Y-\gamma(D,X)$, where $\gamma(\cdot,\cdot)$ is an unknown function ai2003efficient,newey2003instrumental. Conceptually, the same ideas apply: a posterior for $P_{W|Z}$ implies a posterior for a function-valued $\gamma$. However, theoretical issues, such as non-compact function spaces, ill-posedness of identifying conditions, etc., may warrant different minimum distance estimands (e.g., penalized estimands like in chen2012estimation).
Second, model misspecification is an important concern for conditional moment equality models (i.e., when $E_{P_{W|Z}}[g(W,Z,\gamma)|Z] \neq 0$ for all $\gamma \in \Gamma$, with positive probability). The posterior for the value function $Q_{n}(\gamma_{n},\gamma_{n},p_{W|Z})$, a `$J$-statistic'-like estimand, contains information about the compatibility of the conditional moments with the data. Developing a Bayesian specification assessment framework using this posterior and comparing it with classical overidentifying restrictions tests would be interesting. Moreover, reporting the posterior distribution of $Q_{n}(\gamma_{n},\gamma_{n},p_{W|Z})$ may also connect with the recommendations in andrews2025purpose of reporting $J$-statistics (not $J$-tests) in overidentified models.
Finally, competing structural models often lead to nonnested conditional moment equalities $E_{P_{W|Z}}[g_{1}(W,Z,\gamma_{1})|Z] = 0$ and $E_{P_{W|Z}}[g_{2}(W,Z,\gamma_{2})|Z] = 0$. A posterior for $p_{W|Z}$ leads to a joint posterior for $(\gamma_{n,1},Q_{n,1},\gamma_{n,2},Q_{n,2})$, where $Q_{n,j}$, $j \in \{1,2\}$, is shorthand for $Q_{n,j}(\gamma_{n,j},\gamma_{n,j},p_{W|Z})$, with $Q_{n,j}(\gamma,\tilde{\gamma},p_{W|Z})$ being ((ref)) for $E_{P_{W|Z}}[g_{j}(W,Z,\gamma_{j})|Z]$. The posterior for $Q_{n,1}-Q_{n,2}$ can be used to assess which conditional moments are most plausible.