EconBase
← Back to paper

Parametrization, Prior Independence, and the Semiparametric Bernstein-von Mises Theorem for the Partially Linear Model

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

73,498 characters · 17 sections · 84 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Parametrization, Prior Independence, and the Semiparametric Bernstein-von Mises Theorem for the Partially Linear Model

abstractI prove a semiparametric Bernstein-von Mises theorem for a partially linear regression model with independent priors for the low-dimensional parameter of interest and the infinite-dimensional nuisance parameters. My result avoids a challenging prior invariance condition that arises from a loss of information associated with not knowing the nuisance parameter. The key idea is to employ a feasible reparametrization of the partially linear regression model that reflects the semiparametric structure of the model. This allows a researcher to assume independent priors for the model parameters while automatically accounting for the loss of information associated with not knowing the nuisance parameters. The theorem is verified for uniform wavelet series priors and Mat\'{e}rn Gaussian process priors.

Introduction

Overview

This paper is concerned with Bayesian inference for a partially linear regression model. The partially linear model states that the outcome $Y \in \mathbb{R}$ is related to covariates $X \in \mathbb{R}$ and $W \in \mathbb{R}^{d_{w}}$ via the regression model

align[align omitted — 69 chars of source]

where $\beta_{0}$ is a scalar parameter, $\eta_{0}$ is a function-valued parameter, and $U$ is a scalar unobservable that satisfies $E[U|X,W] = 0$ and $Var[U|X,W] = \sigma_{01}^{2}$ for some scalar $\sigma_{01}^{2}\in (0,\infty)$. Partially linear models appear in a variety of applications including modeling the relationship between temperature and electricity demand engle1986semiparametric, economic models of production with endogenous input choices olley1996dynamics,levinsohn2003estimating,wooldridge2009estimating,ackerberg2015identification, and sample selection models ahn1993semiparametric,das2003nonparametric. Moreover, the model has also gained recent theoretical attention as a leading example for which debiased machine learning methods are applied (see, for example, chernozhukov2018double).

My analysis views $\beta_{0}$ as the parameter of interest while treating $\eta_{0}$ as an infinite-dimensional nuisance parameter. The goal is to derive conditions under which the Bernstein-von Mises theorem holds (i.e., the marginal posterior for $\beta$ is asymptotically normal and matches an efficient frequentist estimator). A key challenge with Bernstein-von Mises theory for the partially linear model is that there is loss of information associated with not knowing the nuisance parameter $\eta_{0}$, and, as a result, the natural assumption that $\beta$ and $\eta$ are independent a priori leads to prior invariance. Prior invariance refers to constructing a change of variables that respects the model's semiparametric structure and leaves the prior for the nuisance function $\eta$ roughly unchanged. It was first emphasized as a general condition for semiparametric Bernstein-von Mises theorems in the important papers of castillo2012semiparametric,castillo2012semiparametricBVM, and it introduces both practical and technical issues.\footnote{Prior invariance is also encountered in statistical functional estimation rivoirard2012bernstein,castillo2015bernstein,ray2020semiparametric. Moreover, the condition has been shown to be an almost necessary condition for the Bernstein-von Mises theorem in some models (e.g., the curve alignment model in castillo2012semiparametricBVM).} From a practical standpoint, prior invariance stipulates that the prior for $\eta$ must be adequate for estimating of the true control function $\eta_{0}$ and the conditional expectation $m_{02}$ of $X$ given $W$, a condition that restricts the smoothness $m_{02}$ relative to $\eta_{0}$, and, as a result, limits the data-generating processes for which the Bernstein-von Mises theorem can be applied. From a technical perspective, showing the prior for $\eta$ is roughly unchanged after the change of variables can be difficult to verify for some priors because it involves performing an infinite-dimensional change of measure (e.g., this is a key reason that castillo2012semiparametric,castillo2012semiparametricBVM focuses on Gaussian priors).

The main contribution of this paper is a Bernstein-von Mises theorem for the partially linear model that bypasses prior invariance. Theorem (ref) in Section (ref) establishes this result, and the theorem is verified for two examples in Section (ref). A common theme in the examples is that there are no cross-restrictions on the smoothness of nuisance functions, which demonstrates that my approach can eliminate the aforementioned practical challenge of requiring a prior for a single nuisance parameter (i.e., $\eta$) be suitable for estimation of multiple quantities (i.e., $\eta_{0}$ and $m_{02}$). To make things concrete, my proposal is to embed an interest-respecting reparametrization of ((ref)) given by

align[align omitted — 72 chars of source]

where $m_{01}(W)=E_{P_{0}}[Y|W]$ and $m_{02}(W)=E_{P_{0}}[Y|W]$, within a (quasi-)likelihood model that is suitable for identification of $(\beta_{0},m_{01},m_{02})$, and then perform Bayesian inference by placing independent priors on $\beta$ and $m=(m_{1},m_{2})$. The transformed regression model ((ref)) is commonly referred to as the robinson1988root transformation, and is observationally equivalent to ((ref)) because $\eta_{0}$ is identified as $m_{01}-\beta_{0}m_{02}$. A key property of the (quasi-)likelihood model is that there is no loss of information associated with not knowing the nuisance parameter $m_{0}=(m_{01},m_{02})$ (i.e., it is adaptive in the sense of bickel1982adaptive). Consequently, there is no need to change variables because ordinary/parametric local asymptotic normality expansions respect the semiparametric structure of the model. For terminology, I refer to ((ref)) as the $(\beta,\eta)$-parametrization and ((ref)) as the $(\beta,m)$-parametrization throughout.

To my knowledge, Theorem (ref) is the first Bernstein-von Mises theorem for the $(\beta,m)$-parametrization of the partially linear model. It has several important takeaways. The first takeaway is that the findings demonstrate that the choice of parametrization can influence the conditions under which the semiparametric Bernstein-von Mises theorem holds. This is demonstrated explicitly in Section (ref) in which the $(\beta,\eta)$- and $(\beta,m)$-parametrizations are compared for Mat\'{e}rn Gaussian process priors, a class of Gaussian process priors popular in applications (see, for example, williams2006gaussian). Since the $(\beta,m)$-parametrization is orthogonal (and the $(\beta,\eta)$-parametrization is not), this takeaway is similar to the recent frequentist debiased machine learning literature chernozhukov2018double and is reminiscent of classical results on improving parametric likelihood estimators via orthogonal parametrizations cox1973parameter. The second takeaway is that the use of reparametrization to eliminate prior invariance conditions avoids data-driven prior/posterior adjustments. For this reason, my proposal offers a conceptually distinct approach to obtaining debiased Bernstein-von Mises theorems relative to those encountered in a growing literature on data-driven prior/posterior adjustments yang2015semiparametric,ray2019debiased,ray2020semiparametric,breunig2022double,yiu2025semiparametric. Determining whether the approach of this paper is more generally applicable is an important area for future research. A final takeaway is that the Bernstein-von Mises theorem can be valid even if the (quasi-)likelihood model is misspecified. This reflects classical results on quasi-maximum likelihood estimation gourieroux1984pseudo,white1994estimation, and Section (ref) makes the claim precise by verifying Theorem (ref) for uniform wavelet series priors without requiring the sampling model be correctly specified.

Related Literature

This paper connects to several active literatures in statistics and econometrics. First, it is related to the literature on semiparametric Bernstein-von Mises theorems. A Bernstein-von Mises theorem for the $(\beta,\eta)$-parametrization with independent priors for $\beta$ and $\eta$ is studied in shen2002asymptotic, bickel2012semiparametric, yang2015semiparametric, and xie2020adaptive. castillo2012semiparametric,castillo2012semiparametricBVM are also relevant to my paper because these important papers emphasize the role of the prior invariance condition in separated semiparametric models with information loss. Relatedly, rivoirard2012bernstein and castillo2015bernstein demonstrated the importance of prior invariance for general smooth functionals of nonparametric models (with the latter treating separated semiparametric models as a special case). yang2015semiparametric,ray2020semiparametric,breunig2022double, and yiu2025semiparametric propose data-driven prior/posterior corrections to prove Bernstein-von Mises theorems under weaker conditions than their unmodified counterparts. A key difference between these papers and my proposal is that I do not require any data-driven adjustment because the lack of prior invariance follows from the adaptivity of the $(\beta,m)$-parametrization. yang2019posterior derives Bernstein-von Mises theorem for a single coordinate in a high-dimensional linear regression model under sparsity using a different type of reparametrization and a prior correction. Although I do not explicitly study high-dimensional linear regression, Section (ref) outlines how my approach can be extended to this setting and an interesting subject for future research would be determining whether my approach leads to results similar to yang2019posterior in this setting. Other papers on semiparametric Bernstein-von Mises theorems include kim2006bernstein, dejonge2013semiparametric, norets2015bayesian, chae2019semi, nickl2019bernstein, nickl2020bernstein, monard2021statistical, and l2023semiparametric.

Second, some of the auxiliary results in Section (ref) connect this paper to the literature on posterior contraction rates in nonparametric models. Specifically, Propositions (ref) and (ref) provide a general approach to deriving posterior contraction rates for nuisance functions $m_{01}$ and $m_{02}$. These propositions are based on empirical process conditions that arise from direct expansions of the (quasi-)likelihood ratio process, and, for this reason, are conceptually similar to the convergence rate framework of shen2001rates in which empirical process-type conditions for likelihood ratios (e.g., bracketing integrals) are used to derive posterior contraction rates. Since Propositions (ref) and (ref) concern posterior contraction rates in (possibly) misspecified Gaussian nonparametric regression, the auxiliary results in Section (ref) are also related to the theory presented in Section 4 of kleijn2006misspecification. Other papers on posterior contraction rates in general nonparametric models include ghosal2000convergence,ghosal2007convergence,castillo2008lower,vaart2008rates,10.1214/08-AOS678,van2011information,gine2011rates,dejonge2012adaptive,castillo2014supremum, and shen2015adaptive.

Finally, since the Bernstein-von Mises theorem establishes asymptotic equivalence between Bayesian and frequentist estimators, my paper also dovetails with the literature on frequentist semiparametric estimation. The $(\beta,m)$-parametrization was first utilized in robinson1988root to derive asymptotically normal least squares estimators of $\beta_{0}$ based on kernel estimators of $m_{01}$ and $m_{02}$. DONALD199430 also apply the partialling out technique of robinson1988root to show that asymptotic normality of the least squares estimators of $\beta_{0}$ can be achieved under very mild conditions when series regression methods are used to estimate the nuisance functions. My paper provides a Bayesian analogue to these papers because it shows the $(\beta,m)$-parametrization can be used to derive a Bernstein-von Mises theorem that bypasses prior invariance. The proposed Bayesian inference framework is also related to the literature on profile likelihood estimation severini1992profile, murphy2000profile because the transformed regression model ((ref)) mimics the Gaussian profile maximum likelihood problem for ((ref)). Since $(\beta,m)$-parametrization is orthogonal, my paper is also related to the literature on two-step semiparametric estimation based on Neyman orthogonal moment conditions (e.g., andrews1994asymptotics, newey1994asymptotic,chernozhukov2018double). Using a (quasi-)likelihood to perform Bayesian inference a low-dimensional parameter of interest defined by conditional moment restrictions also connects my proposal to the quasi-maximum likelihood estimation literature gourieroux1984pseudo,white1994estimation,KOMUNJER2005137.

Outline of Paper

The rest of the paper is organized as follows. Section (ref) formalizes the data-generating process and puts forward the Bayesian inference framework. Section (ref) derives conditions under which the marginal posterior for $\beta$ satisfies a Bernstein-von Mises theorem and verifies the conditions for two important classes of priors. Section (ref) provides an extensive discussion the findings of Section (ref). Section (ref) concludes and proposes some directions for future research. All proofs are in the Appendix and notation is introduced when appropriate.

Data Generating Process and Bayesian Inference

Data Generating Process

Let $Y \in \mathcal{Y} \subseteq \mathbb{R}$ be an outcome variable, let $X \in \mathcal{X} \subseteq \mathbb{R}$ be a covariate of interest, let $W \in \mathcal{W} \subseteq \mathbb{R}^{d_{w}}$, $d_{w} < \infty$, be some control variables, and let $P_{0}$ be the joint distribution of $(Y,X,W')'$. The observed data $\{(Y_{i},X_{i},W_{i}')'\}_{i=1}^{n}$ consists of the first $n$ elements of a sequence $\{(Y_{i},X_{i},W_{i}')'\}_{i \geq 1}$ of independent and identically distributed (i.i.d) random vectors drawn from $P_{0}$. Assumption (ref) states the restrictions on the data distribution $P_{0}$.

assumptionThe distribution $P_{0}$ of $(Y,X,W')'$ satisfies the following restrictions: \begin{enumerate} • The distribution of $Y$ given $X$ and $W$ satisfies $E_{P_{0}}[Y|X,W]= X\beta_{0}+\eta_{0}(W)$ and $Var_{P_{0}}[Y|X,W] = \sigma_{01}^{2}$ for some $(\beta_{0},\eta_{0}, \sigma_{01}^{2}) \in \mathcal{B} \times \mathcal{H} \times \mathbb{R}_{++}$ with $\mathcal{B} \subseteq \mathbb{R}$ and $\mathcal{H} \subseteq L^{2}(\mathcal{W})$, where $f \in L^{2}(\mathcal{W})$ iff $E_{P_{0}}[f^{2}(W)] < \infty$. • There exists constants $0<\underline{c}< \overline{c}<\infty$ such that $E_{P_{0}}[(X-E_{P_{0}}[X|W])^{2}]\in [\underline{c},\overline{c}]$. • The conditional distribution $P_{0,YX|W}$ of $(Y,X)'$ given $W$ has a density $p_{0,YX|W}$ with respect to the Lebesgue measure and $E_{P_{0}}|\log p_{0,YX|W}(Y,X|W)| \in (0,\infty)$ • The projection errors $Y- E_{P_{0}}[Y|W]$ and $X-E_{P_{0}}[X|W]$ satisfy $E_{P_{0}}[(Y-E_{P_{0}}[Y|W])^{2}] \in (0,\infty)$ and $E_{P_{0}}[(X-E_{P_{0}}[X|W])^{4}]\in (0,\infty)$. \end{enumerate}

I briefly discuss Assumption (ref). Part 1 of Assumption (ref) imposes that the conditional moment restriction ((ref)) is satisfied at $P_{0}$ and that the projection error $Y - X\beta_{0}-\eta_{0}(W)$ is homoskedastic. I treat $\sigma_{01}^{2}$ as known for exposition, however, Appendix (ref) extends the main results to accommodate unknown $\sigma_{01}^{2}$ and multivariate $X$. Part 2 of Assumption (ref) is a strong identification condition that is necessary and sufficient for regular estimation of $\beta_{0}$ robinson1988root. The remaining parts are technical conditions that enable the application of stochastic limit theorems and ensure certain population objective functions are well-defined.

Bayesian Inference

Bayesian inference starts with a conditional model for the data (sampling model) and a distribution over the model parameters (prior). I describe each of these in turn. The sampling model is

align[align omitted — 304 chars of source]

for $i=1,...,n$, where $\beta \in \mathcal{B}$, $m = (m_{1},m_{2}) \in \mathcal{M}$ with $\mathcal{M} = \mathcal{M}_{1} \times \mathcal{M}_{2}$ and $\mathcal{M}_{j} \subseteq L^{2}(\mathcal{W})$ for $j \in \{1,2\}$, and $\sigma_{01}^{2},\sigma_{02}^{2} \in (0,\infty)$ are known scalars. Display ((ref)) is a model for the conditional distribution of $Y$ given $X$ and $W$ that imposes the robinson1988root transformation of the partially linear model, a result that establishes the observational equivalence between ((ref)) and the ‘partialed out' regression model $Y = m_{01}(W) + (X-m_{02}(W))\beta_{0} + U$, where $m_{01}(W) = E_{P_{0}}[Y|W]$, $m_{02}(W) = E_{P_{0}}[X|W]$, and $U$ satisfies $E[U|X,W] = 0$ and $Var[U|X,W] = \sigma_{01}^{2}$. Display ((ref)) is a model for the conditional distribution of $X$ given $W$ that is included to account for the fact that $m_{01}$ and $m_{02}$ are not separately identified from ((ref)).\footnote{Since ((ref)) is for identification of $m_{02}$ only, $\sigma_{02}^{2}$ has no intrinsic meaning and can be set at an arbitrary fixed value (e.g., $\sigma_{02}^{2}=1$).} Together, ((ref)) and ((ref)) imply a (quasi-)likelihood $L_{n}(\beta,m)$ given by

align*[align* omitted — 77 chars of source]

where $p_{\beta,m}(\cdot|w)$ is the probability density function of a bivariate Gaussian distribution with mean vector $m(w)=(m_{1}(w),m_{2}(w))'$ and covariance matrix

align*[align* omitted — 176 chars of source]

The term `(quasi-)likelihood' is used because the set of conditional distributions for $(Y,X)$ given $W$ compatible with Assumption (ref) is much larger than those that are compatible with ((ref))--((ref)). Sections (ref) and (ref) elaborate on the sampling model's suitability for estimating $\beta_{0}$ despite it being much more restrictive than Assumption (ref).

Since the model parameters are $(\beta,m)$, the prior $\Pi$ is a probability distribution over $\mathcal{B} \times \mathcal{M}$. Throughout, I maintain that $\beta \sim \Pi_{\mathcal{B}}$, $m \sim \Pi_{\mathcal{M}}$, and $\beta \protect\mathpalette{\protect\independenT}{\perp} m$, where $\Pi_{\mathcal{B}}$ and $\Pi_{\mathcal{M}}$ are probability measures over $\mathcal{B}$ and $\mathcal{M}$, respectively. This means that $\Pi = \Pi_{\mathcal{B}} \otimes \Pi_{\mathcal{M}}$. Prior independence between the low-dimensional target parameter and the infinite-dimensional nuisance parameters is standard in semiparametric Bayesian inference bickel2012semiparametric,castillo2012semiparametric,castillo2012semiparametricBVM,ghosal2017fundamentals. Assumption (ref) states some regularity conditions for $\Pi$.

assumptionThe following conditions hold: \begin{enumerate} • The prior is such that $P_{0}^{\infty}(\int_{\mathcal{B}\times \mathcal{M}}L_{n}(\beta,m)d \Pi(\beta,m) \in (0,\infty))=1$, where $P_{0}^{\infty}$ is the law of $\{(Y_{i},X_{i},W_{i}')'\}_{i\geq 1}$. • The prior $\Pi_{\mathcal{B}}$ has a Lebesgue density $\pi_{\mathcal{B}}$ that is continuous and positive over a neighborhood $\mathcal{B}_{0}$ of $\beta_{0}$. \end{enumerate}

The sampling model ((ref))--((ref)) and prior $\Pi$ implies a conditional distribution $\Pi((\beta,m) \in \cdot | \{(Y_{i},X_{i},W_{i}')'\}_{i=1}^{n})$ for the model parameters $(\beta,m)$ given the data $\{(Y_{i},X_{i},W_{i}')'\}_{i=1}^{n}$. This is known as the posterior, and, under Assumption (ref), posterior probabilities are computed via Bayes rule. That is, for any event $A$,

align*[align* omitted — 184 chars of source]

Bayesian inference revolves around the posterior distribution. As examples, researchers interested in a point estimator for $\beta$ can report the posterior median or posterior mean, while researchers interested in quantifying uncertainty about $\beta$ report a credible set, a set estimator $CS_{n}(1-\alpha)$ that satisfies $\Pi(\beta \in CS_{n}(1-\alpha)|\{(Y_{i},X_{i},W_{i}')'\}_{i=1}^{n}) \geq 1-\alpha$.

remark[Posterior Computation] Although this paper focuses on theoretical properties of the posterior, I provide some brief remarks on posterior computation. An approach to sampling from the posterior is Gibbs sampling. That is, the elements of $(\beta,m_{1},m_{2})$ are updated component-wise while fixing the remaining parameters. Holding $\beta$ and $m_{2}$ fixed, ((ref)) reveals that Bayesian updating of $m_{1}$ is that of a Gaussian nonparametric regresion of $Y-(X-m_{2}(W))\beta$ on $W$. Similarly, factorizing $p_{\beta,m}$ into the product of the conditional distribution of $X|Y,W$ and $Y|W$ indicates that Bayesian updating of $m_{2}$ (while holding $\beta$ and $m_{1}$ fixed) is that of a Gaussian nonparametric regression of $X-(\beta\sigma_{2}^{2}/(\sigma_{01}^{2}+\beta^{2}\sigma_{02}^{2}))(Y-m_{1}(W))$ on $W$. Finally, holding $m_{1}$ and $m_{2}$ fixed, the first part of ((ref)) reveals that Bayesian updating of $\beta$ conforms to that of a Gaussian linear regression of $Y-m_{1}(W)$ on $X-m_{2}(W)$. Section (ref) provides some theoretical guarantees for the case where $m_{1} \protect\mathpalette{\protect\independenT}{\perp} m_{2}$ and $m_{j}$, $j \in \{1,2\}$, follows a Gaussian process. This is relevant for computation because these priors are conjugate for the first two steps.

Bernstein-von Mises Theorem

This section provides Bernstein-von Mises theory for the marginal posterior for $\beta$. It has two subsections. The first establishes a general Bernstein-von Mises theorem based on high-level posterior consistency and empirical process conditions, and the second verifies the assumptions for two classes of priors.

General Theorem

I prove a general Bernstein-von Mises theorem for the marginal posterior of $\beta$. The asymptotic analysis is conducted conditional on the realizations of $\{W_{i}\}_{i \geq 1}$, meaning that $\{(Y_{i},X_{i})\}_{i=1}^{n}$ should be viewed as a sequence of independent but not identically distributed random vectors (i.e., $(Y_{i},X_{i}) \protect\mathpalette{\protect\independenT}{\perp} (Y_{j},X_{j})|\{W_{i}\}_{i \geq 1}$ for $i \neq j$). For notation, $P_{0,W}^{\infty}$ is the joint law of the i.i.d sequence $\{W_{i}\}_{i \geq 1}$, $\langle f, g \rangle_{n,2} = n^{-1}\sum_{i=1}^{n}f(w_{i})g(w_{i})$ and $||f||_{n,2}= \sqrt{\langle f,f \rangle_{n,2}}$ are the empirical $L^{2}$ inner product and induced norm, respectively, for functions $f,g: \mathcal{W} \rightarrow \mathbb{R}$ and for a given realization $\{w_{i}\}_{i \geq 1}$ of $\{W_{i}\}_{i \geq 1}$, $B_{n,\mathcal{M}}(m_{0},\delta)$ is a ball centered at $m_{0}$ with radius $\delta > 0$ in the product empirical $L^{2}$ norm $\max\{||m_{1}||_{n,2},||m_{2}||_{n,2}\}$, and $P_{0,YX|W}^{(n)} = \bigotimes_{i=1}^{n}P_{0,YX|w_{i}}$ with $P_{0,YX|w_{i}}$ denoting the probability distribution associated with the probability density function $p_{0,YX|W}(\cdot|w_{i})$ for $i=1,...,n$. The next two assumptions concern the limiting behavior of the posterior.

assumptionThere exists a sequence $\{\delta_{n}\}_{n \geq 1}$ such that \begin{align*} \delta_{n} = o(n^{-1/4}) \end{align*} and \begin{align*} \Pi(B_{n,\mathcal{M}}(m_{0},D\delta_{n})|\{(Y_{i},X_{i},w_{i}')'\}_{i=1}^{n}) \overset{P_{0,YX|W}^{(n)}}{\longrightarrow} 1 \end{align*} for $P_{0,W}^{\infty}$-almost every fixed realization $\{w_{i}\}_{i \geq 1}$ of $\{W_{i}\}_{i \geq 1}$ as $n\rightarrow \infty$ for some constant $D>0$ large.
assumptionThere are sets $\{\mathcal{M}_{n}\}_{n \geq 1} \subseteq \mathcal{M}$ such that \begin{align*} \Pi(\mathcal{M}_{n}|\{(Y_{i},X_{i},w_{i}')'\}_{i=1}^{n}) &\overset{P_{0,YX|W}^{(n)}}{\longrightarrow} 1 \\ \sup_{m \in \mathcal{M}_{n}\cap B_{n,\mathcal{M}}(m_{0},D\delta_{n})}\left| G_{n}^{(1)}(m) \right|&\overset{P_{0,YX|W}^{(n)}}{\longrightarrow} 0 \\ \sup_{m \in \mathcal{M}_{n}\cap B_{n,\mathcal{M}}(m_{0},D\delta_{n})}\left| G_{n}^{(2)}(m) \right|&\overset{P_{0,YX|W}^{(n)}}{\longrightarrow} 0 \end{align*} for $P_{0,W}^{\infty}$-almost every fixed sequence $\{w_{i}\}_{i \geq 1}$ of $\{W_{i}\}_{i \geq 1}$ as $n\rightarrow \infty$, where $G_{n}^{(1)},G_{n}^{(2)}: \mathcal{M} \rightarrow \mathbb{R}$ satistify \begin{align*} G_{n}^{(1)}(m) = \frac{1}{\sqrt{n}}\sum_{i=1}^{n}\varepsilon_{i2}(m_{1}(w_{i})-m_{01}(w_{i})) \end{align*} and \begin{align*} G_{n}^{(2)}(m) = \frac{1}{\sqrt{n}}\sum_{i=1}^{n}(U_{i}-\beta_{0}\varepsilon_{i2})(m_{2}(w_{i})-m_{02}(w_{i})), \end{align*} with $\varepsilon_{i2} = X_{i}-m_{02}(w_{i})$ and $U_{i} = Y_{i}-m_{01}(w_{i})-(X_{i}-m_{02}(w_{i}))\beta_{0}$.

Assumptions (ref) and (ref) are high-level conditions about the the limiting behavior of the posterior. Assumption (ref) enforces that the marginal posterior for $m$ concentrates around $m_{0}$ in the empirical $L^{2}$-norm and the rate of convergence $\delta_{n}$ is faster than $n^{-1/4}$. Assumption (ref) imposes that there are sets $\{\mathcal{M}_{n}\}_{n \geq 1}$ that the posterior concentrates on such that, when intersected with $B_{n,\mathcal{M}}(m_{0},\delta_{n})$, are structured enough to ensure that multiplier empirical processes $\{G_{n}^{(1)}\}_{n\geq 1}$ and $\{G_{n}^{(2)}\}_{n \geq 1}$ satisfy stochastic equicontinuity-type conditions. These requirements are conceptually similar to those encountered in the literature on asymptotic normality of plug-in method of moments estimators based on orthogonal moment conditions andrews1994asymptotics,chernozhukov2018double and semiparametric maximum likelihood estimators murphy2000profile. Combined with Assumptions (ref) and (ref), Assumptions (ref) and (ref) imply that the marginal posterior for $\beta$ satisfies a Bernstein-von Mises theorem.

theoremLet $\tilde{I}_{n}(m_{0})$ be given by \begin{align*} \tilde{I}_{n}(m_{0}) = \frac{1}{n}\sum_{i=1}^{n}\left(\frac{X_{i}-m_{02}(w_{i})}{\sigma_{01}}\right)^{2} \end{align*} and let $\tilde{\Delta}_{n,0}$ be given by \begin{align*} \tilde{\Delta}_{n,0} = \tilde{I}_{n}(m_{0})^{-1}\frac{1}{\sqrt{n}}\tilde{\ell}_{n}(\beta_{0},m_{0}), \end{align*} where \begin{align*} \tilde{\ell}_{n}(\beta_{0},m_{0}) = \frac{1}{\sigma_{01}^{2}}\sum_{i=1}^{n}(Y_{i}-m_{01}(w_{i})-(X_{i}-m_{02}(w_{i}))\beta_{0})(X_{i}-m_{02}(w_{i})). \end{align*} If Assumptions (ref), (ref), (ref), and (ref) hold and $\beta_{0} \in \text{int}(\mathcal{B})$, then \begin{align*} \left | \left |\Pi\left(\beta \in \cdot |\{(Y_{i},X_{i},w_{i}')'\}_{i=1}^{n}\right) - \mathcal{N}\left(\beta_{0}+\frac{\tilde{\Delta}_{n,0}}{\sqrt{n}},\frac{1}{n}\tilde{I}_{n}(m_{0})^{-1} \right) \right | \right |_{TV} \overset{P_{0,YX|W}^{(n)}}{\longrightarrow} 0 \end{align*} for $P_{0,W}^{\infty}$-almost every fixed realization $\{w_{i}\}_{i \geq 1}$ of $\{W_{i}\}_{i \geq 1}$ as $n\rightarrow \infty$, where $||\cdot||_{TV}$ is total variation distance.\footnote{The total variation distance between probability measures $P$ and $Q$ is $||P-Q||_{TV} = \sup_{A}|P(A)-Q(A)|$.}
remark[Unconditional Asymptotics] Theorem (ref) conditions on the sequence of control variables $\{W_{i}\}_{i \geq 1}$. The same conclusion holds in the unconditional asymptotic thought experiment in which $P_{0,YX|W}^{(n)}$ is replaced by the $n$-fold product measure $P_{0}^{n}$ (i.e., the sequence $\{W_{i}\}_{i \geq 1}$ is also treated as random). This follows from an application of the Bounded Convergence Theorem and is stated as a corollary below.
corollaryIf Assumptions (ref), (ref), (ref), and (ref) hold and $\beta_{0} \in \text{int}(\mathcal{B})$, then \begin{align*} \left | \left |\Pi\left(\beta \in \cdot |\{(Y_{i},X_{i},W_{i}')'\}_{i=1}^{n}\right) - \mathcal{N}\left(\beta_{0}+\frac{\tilde{\Delta}_{n,0}}{\sqrt{n}},\frac{1}{n}\tilde{I}_{n}(m_{0})^{-1} \right) \right | \right |_{TV} \overset{P_{0}^{n}}{\longrightarrow} 0 \end{align*} as $n\rightarrow \infty$.
remark[Efficiency] Under i.i.d sampling, the probability limit of $\tilde{I}_{n}^{-1}(m_{0})$ is the chamberlain1992efficiency semiparametric efficiency bound under homoskedasticity and $\hat{\beta}_{n} =\beta_{0}+\tilde{\Delta}_{n,0}/\sqrt{n}$ is an estimator that achieves Chamberlain's bound. Consequently, the above results provide conditions under which the marginal posterior for $\beta$ is first-order asymptotically equivalent to a (locally) efficient frequentist estimator of the regression coefficient $\beta_{0}$. Efficiency is local because chamberlain1992efficiency does not constrain the conditional variance of $U$ given $X$ and $W$ to be constant when deriving the information bound (see Section 4.3 of newey1990semiparametric for more on local efficiency).
remark[Uncertainty Quantification] Theorem (ref) implies Bayesian credible sets are asymptotically (locally) efficient frequentist confidence sets for data-generating processes compatible with Assumption (ref). To see why, let $c_{n}(q)$, $q \in (0,1)$, be the $q$-quantile of the marginal posterior $\Pi(\beta \in \cdot |\{(Y_{i},X_{i},W_{i}')'\}_{i=1}^{n})$ and consider the equitailed probability interval $[c_{n}(\alpha/2),c_{n}(1-\alpha/2)]$, where $\alpha \in (0,1/2)$. Corollary (ref) establishes that the endpoints of these intervals are asymptotically equivalent to the endpoints of Wald confidence intervals based on the estimator $\hat{\beta}_{n}=\beta_{0}+\tilde{\Delta}_{n,0}/\sqrt{n}$ that achieves the chamberlain1992efficiency efficiency bound under homoskedasticity. Consequently, the credible interval $[c_{n}(\alpha/2),c_{n}(1-\alpha/2)]$ achieves asymptotic frequentist coverage of $1-\alpha$. Note that the result is stated conditionally (i.e., $P_{0,YX|W}^{(n)}$), however, the same argument can be applied using Corollary (ref) to obtain it unconditionally (i.e., $P_{0}^{n}$). \begin{corollary} If Assumptions (ref), (ref), (ref), and (ref) hold and $\beta_{0} \in \text{int}(\mathcal{B})$, then the $q$-quantile $c_{n}(q)$, $q \in (0,1)$, of the marginal posterior $\Pi(\beta \in \cdot| \{(Y_{i},X_{i},w_{i}')'\}_{i=1}^{n})$ satisfies \begin{align*} c_{n}(q) = \hat{\beta}_{n} + \Phi^{-1}(q)\sqrt{\frac{\tilde{I}_{n}^{-1}(m_{0})}{n}} + o_{P_{0,YX|W}^{(n)}}\left(\frac{1}{\sqrt{n}}\right) \end{align*} for $P_{0,W}^{\infty}$-almost every fixed realization $\{w_{i}\}_{i \geq 1}$ of $\{W_{i}\}_{i \geq 1}$ as $n\rightarrow \infty$, where $\Phi$ is the cumulative distribution function of $\mathcal{N}(0,1)$, $\hat{\beta}_{n} = \beta_{0}+\tilde{\Delta}_{n,0}/\sqrt{n}$, and $\tilde{I}_{n}(m_{0})=\sigma_{01}^{-2}n^{-1}\sum_{i=1}^{n}(X_{i}-m_{02}(w_{i}))^{2}$. \end{corollary}

Sufficient Conditions and Examples

This section has two subsections. Section (ref) presents sufficient conditions for the assumptions in Section (ref). Sections (ref) and (ref) verify Theorem (ref) for two classes of priors using these sufficient conditions.

Useful Preliminary Results

I start with a proposition that enables finding sequences $\{\delta_{n}\}_{ n\geq 1}$ for which Assumption (ref) holds. It has similarities with Theorem 2 of shen2001rates in that it establishes posterior consistency by directly analyzing the (quasi-)likelihood ratio process. For notation, let $\varepsilon = (Y-m_{01}(W),X-m_{02}(W))'$, and let $\lambda_{min}(A)$ and $\lambda_{max}(A)$ denote the minimum and maximum eigenvalues of a matrix $A$, respectively.

propositionSuppose that $\mathcal{B}$ is compact, $\sup_{w \in \mathcal{W}}\lambda_{max}(E(\varepsilon\varepsilon'|W=w))< \infty$, and there is a sequence $\{\delta_{n}\}_{n \geq 1}$ such that the following conditions hold: \begin{enumerate} • There exists a constant $C>0$ such that $\Pi_{\mathcal{M}}(m \in B_{n,\mathcal{M}}(m_{0},\delta_{n})) \gtrsim \exp(-Cn\delta_{n}^{2})$ conditionally given $P_{0,W}^{\infty}$-almost every realization $\{w_{i}\}_{i \geq 1}$ of $\{W_{i}\}_{i \geq 1}$. • There exists increasing function $\delta \mapsto \omega_{n}(\delta)$ for which $\omega_{n}(\delta_{n}) \leq \sqrt{n}\delta_{n}^{2}$, $\delta \mapsto \omega_{n}(\delta)/\delta^{\upsilon}$ is decreasing for some $\upsilon \in (0, 2)$, and \begin{align} E_{P_{0,YX|W}^{(n)}}\left[\sup_{(\beta,m) \in \mathcal{B} \times B_{n,\mathcal{M}}(m_{0},\delta)}\left|\frac{1}{\sqrt{n}}\sum_{i=1}^{n}(m(w_{i})-m_{0}(w_{i}))'V^{-1}(\beta)\varepsilon_{i}\right|\right] \lesssim \omega_{n}(\delta) \end{align} for $P_{0,W}^{\infty}$-almost every fixed realization $\{w_{i}\}_{i \geq 1}$ of $\{W_{i}\}_{i \geq 1}$. • The sequence $\{\delta_{n}\}_{n \geq 1}$ satisfies $n\delta_{n}^{2} \rightarrow \infty$ as $n\rightarrow \infty$. \end{enumerate}Then the marginal posterior for $m$ satisfies \begin{align*} \Pi\left(m \in B_{n,\mathcal{M}}(m_{0},D\delta_{n})^{c}\middle |\{(Y_{i},X_{i},w_{i}')'\}_{i =1}^{n}\right) \overset{P_{0,YX|W}^{(n)}}{\longrightarrow} 0 \end{align*} for $P_{0,W}^{\infty}$-almost every fixed realization $\{w_{i}\}_{i \geq 1}$ of $\{W_{i}\}_{i \geq 1}$ as $n\rightarrow \infty$ for some large constant $D>0$.

Proposition (ref) states that $\delta_{n}$ is determined by the 1. mass the nuisance prior $\Pi_{\mathcal{M}}$ assigns to empirical $L^{2}$ neighborhoods of $m_{0}$, and 2. the rate of convergence of the least squares estimator of $m_{0}$. Since $||\cdot||_{n,2} \leq ||\cdot||_{\infty}$ with $||f||_{\infty}= \sup_{w \in \mathcal{W}}|f(w)|$ denoting the supremum norm, the prior mass condition can be checked by analyzing the probability mass that $\Pi_{\mathcal{M}}$ assigns to uniform balls centered at $m_{0}$, a quantity that has been studied for many different priors ghosal2017fundamentals. The least squares connection is through the second condition because $\delta \mapsto \omega_{n}(\delta)$ can be identified as the `continuity modulus' of the multiplier empirical process that determines the least squares rate of convergence (see, for example, Chapter 9 of geer2000empirical and Section 3.2 of vaartwellner96book). Given a valid $\delta_{n}$, conditions that ensure the rate restriction $\delta_{n} = o(n^{-1/4})$ can be established.

remark[Relaxing Part 2 of Proposition (ref)] The supremum in ((ref)) may be infinite for some choices of $\mathcal{M}$ (e.g., if $\mathcal{M}$ is not compact).\footnote{I do not relax compactness of $\mathcal{B}$ because, unlike infinite-dimensional parameter spaces, a compactness condition for a finite-dimensional parameter space is not overly restrictive.} The following result modifies Proposition (ref) to hold along sieves $\{\mathcal{M}_{n}\}_{n \geq 1}$ such that the posterior $\Pi(m \in \mathcal{M}_{n}|\{(Y_{i},X_{i},w_{i}')'\}_{i=1}^{n}) \rightarrow 1$ in $P_{0,YX|W}^{(n)}$-probability for $P_{0,W}^{\infty}$-almost every fixed realization $\{w_{i}\}_{i \geq 1}$ of $\{W_{i}\}_{i \geq 1}$ as $n\rightarrow \infty$. This relaxation is relevant for the Gaussian process example in Section (ref), and, like Proposition (ref), has some similarities with shen2001rates (i.e., Theorem 4 of their paper).
propositionSuppose that $\mathcal{B}$ is compact, $\sup_{w \in \mathcal{W}}\lambda_{max}(E(\varepsilon\varepsilon'|W=w))< \infty$, and there is a sequences $\{\delta_{n}\}_{n \geq 1}$ and $\{\mathcal{M}_{n}\}_{n \geq 1}$ such that the following conditions hold: \begin{enumerate} • There exists a constant $C>0$ such that $\Pi_{\mathcal{M}}(m \in B_{n,\mathcal{M}}(m_{0},\delta_{n})) \gtrsim \exp(-Cn\delta_{n}^{2})$ conditionally given $P_{0,W}^{\infty}$-almost every realization $\{w_{i}\}_{i \geq 1}$ of $\{W_{i}\}_{i \geq 1}$. • The posterior satisfies $\Pi(\mathcal{M}_{n}^{c}|\{(Y_{i},X_{i},w_{i}')'\}_{i=1}^{n}) \rightarrow 0$ in $P_{0,YX|W}^{(n)}$-probability for $P_{0,W}^{\infty}$-almost every fixed realization $\{w_{i}\}_{i \geq 1}$ of $\{W_{i}\}_{i \geq 1}$ as $n\rightarrow \infty$. • There exists an increasing function $\delta \mapsto \omega_{n}(\delta)$ for which $\omega_{n}(\delta_{n}) \leq \sqrt{n}\delta_{n}^{2}$, $\delta \mapsto \omega_{n}(\delta)/\delta^{\upsilon}$ is decreasing for some $\upsilon \in (0, 2)$, and \begin{align*} E_{P_{0,YX|W}^{(n)}}\left[\sup_{(\beta,m) \in \mathcal{B} \times (\mathcal{M}_{n}\cap B_{n,\mathcal{M}}(m_{0},\delta))}\left|\frac{1}{\sqrt{n}}\sum_{i=1}^{n}(m(w_{i})-m_{0}(w_{i}))'V^{-1}(\beta)\varepsilon_{i}\right|\right] \lesssim \omega_{n}(\delta) \end{align*} for $P_{0,W}^{\infty}$ almost-every fixed realization $\{w_{i}\}_{i \geq 1}$ of $\{W_{i}\}_{i \geq 1}$. • The sequence satisfies $n\delta_{n}^{2} \rightarrow \infty$ as $n\rightarrow \infty$. \end{enumerate} Then the marginal posterior for $m$ satisfies \begin{align*} \Pi\left(m \in B_{n,\mathcal{M}}(m_{0},D\delta_{n})^{c}\middle |\{(Y_{i},X_{i},w_{i}')'\}_{i =1}^{n}\right) \overset{P_{0,YX|W}^{(n)}}{\longrightarrow} 0 \end{align*} for $P_{0,W}^{\infty}$-almost every fixed realization $\{w_{i}\}_{i \geq 1}$ of $\{W_{i}\}_{i \geq 1}$ as $n\rightarrow \infty$ for some large constant $D>0$.
remark[Propositions (ref) and (ref) Can Help Verify Assumption (ref)] Verifying Proposition (ref) can also lead to the verification of Assumption (ref). Since $\beta \mapsto V^{-1}(\beta)$ is continuous and $\mathcal{B}$ is compact, the verification of ((ref)) can be reduced to a collection of univariate problems of finding increasing functions $\delta \mapsto \omega_{n,j_{1},j_{2}}(\delta)$, $j_{1},j_{2} \in \{1,2\}$, such that $\delta \mapsto \omega_{n,j_{1},j_{2}}(\delta)/\delta^{\upsilon_{j_{1},j_{2}}}$ is decreasing for some $\upsilon_{j_{1},j_{2}} \in (0,2)$ and \begin{align} E_{P_{0,YX|W}^{(n)}}\left[\sup_{m_{j_{1}} \in \mathcal{M}_{j_{1}}: ||m_{j_{1}}-m_{0,j_{1}}||_{n,2} < D \delta_{n}}\left|\frac{1}{\sqrt{n}}\sum_{i=1}^{n}(m_{j_{1}}(w_{i})-m_{0,j_{1}}(w_{i}))'\varepsilon_{i,j_{2}}\right|\right] \lesssim \omega_{n,j_{1},j_{2}}(\delta). \end{align} Indeed, given $\{\omega_{n,j_{1},j_{2}}: j_{1},j_{2} \in \{1,2\}\}$, I can check the inequality $\max_{j_{1},j_{2} \in \{1,2\}}\omega_{n,j_{1},j_{2}}(\delta_{n}) \leq \sqrt{n}\delta_{n}^{2}$ to verify that ((ref)) holds (ignoring multiplicative constants). Since it is required that $\delta_{n} = o(n^{-1/4})$ and the error $U$ satisfies $U = \varepsilon_{1} - \beta_{0}\varepsilon_{2}$, the inequality $$\max_{j_{1},j_{2} \in \{1,2\}}\omega_{n,j_{1},j_{2}}(\delta_{n}) \leq \sqrt{n}\delta_{n}^{2}$$ immediately implies Assumption (ref) holds for $\mathcal{M}_{n} = \mathcal{M}$. For the case where the suprema in ((ref)) may be unbounded, a symmetric argument may be applied except that ((ref)) is modified to incorporate $\mathcal{M}_{n,j_{1}}$, $j_{1} \in \{1,2\}$. That is, \begin{align*} E_{P_{0,YX|W}^{(n)}}\left[\sup_{m_{j_{1}} \in \mathcal{M}_{n,j_{1}}: ||m_{j_{1}}-m_{0,j_{1}}||_{n,2} < D \delta_{n}}\left|\frac{1}{\sqrt{n}}\sum_{i=1}^{n}(m_{j_{1}}(w_{i})-m_{0,j_{1}}(w_{i}))'\varepsilon_{i,j_{2}}\right|\right] \lesssim \omega_{n,j_{1},j_{2}}(\delta). \end{align*} In such case, Assumption (ref) holds for a candidate $\mathcal{M}_{n}$ used to verify Proposition (ref). I use the technique outlined in this remark for the examples in Sections (ref) and (ref), so that Assumptions (ref) and (ref) are simultaneously verified.

Example 1 -- Uniform Wavelet Series Priors

I verify Theorem (ref) for a class of priors based on untruncated uniform wavelet series. These priors have feature regularly in the Bayesian nonparametrics literature gine2011rates,10.1214/13-AOS1133,castillo2014supremum,l2023semiparametric, however, as far as I am aware, the Bernstein-von Mises theorem for the partially linear model has not been verified using these priors. To make things formal, suppose that $\mathcal{W}= [0,1]$, the marginal distribution $P_{0,W}$ of $W$ has a Lebesgue density $p_{0,W}$ on $[0,1]$ that is bounded away from zero and infinity, and that there are known constants $M>0$ and $\alpha_{0,1},\alpha_{0,2}>1/2$ such that $||m_{0,j}||_{\infty,\infty,\alpha_{0,j}} \leq M$ for all $j \in \{1,2\}$, where $||\cdot||_{\infty,\infty,\alpha}$ denotes the norm of a H\"{o}lder-Zygmund space $B_{\infty,\infty}^{\alpha}([0,1])$, $\alpha > 0$.\footnote{For $\alpha \notin \mathbb{N}$, H\"{o}lder-Zygmund spaces coincide with the usual H\"{o}lder space $C^{\alpha}([0,1])$ of functions that are $\lfloor \alpha \rfloor$-times continuously differentiable and $\lfloor \alpha \rfloor$th derivative is $(\alpha-\lfloor \alpha \rfloor)$-Lipschitz and the norm $||\cdot||_{\infty,\infty,\alpha}$ is equivalent to the usual H\"{o}lder norm $||\cdot||_{\alpha}$ (i.e., given by (4.111) in Gine_Nickl_2015). However, for $\alpha \in \mathbb{N}$, these spaces are no longer the same. See Section 4.3.3 of Gine_Nickl_2015 for more details.} Given these assumptions, I set the nuisance parameter space to be $\mathcal{M}_{j}=\{f \in L^{2}([0,1]): ||f||_{\infty,\infty,\alpha_{0,j}} \leq M\}$ for each $j \in \{1,2\}$ and a prior $\Pi_{\mathcal{M}}$ over $\mathcal{M}_{1} \times \mathcal{M}_{2}$ is constructed as follows. I take a boundary adapted wavelet basis $\{\psi_{lk}: l \geq 0, 0 \leq k \leq 2^{l}-1\}$ for $L^{2}([0,1])$ that is sufficiently regular to characterize the H\"{o}lder-Zygmund spaces $(B_{\infty,\infty}^{\alpha}([0,1]),||\cdot||_{\infty,\infty,\alpha})$, $\alpha \in \{\alpha_{0,1},\alpha_{0,2}\}$, and set the marginal prior $\Pi_{\mathcal{M}_{j}}$ for $m_{j}$ to be the law of the random function

align*[align* omitted — 109 chars of source]

where $\nu_{lk,j}$ are independent and identically distributed (across $l$, $k$, and $j$) uniform random variables supported on $[-M,M]$.\footnote{The wavelet basis being sufficiently regular to characterize $(B^{\alpha}_{\infty,\infty}([0,1]),||\cdot||_{\alpha,\infty,\infty})$ means that the norm $||\cdot||_{\infty,\infty,\alpha}$ satisfies $||f||_{\infty,\infty,\alpha}=\sup_{l \geq 0}\max_{0 \leq k \leq 2^{l}-1}2^{l(\alpha+1/2)}|\langle f, \psi_{lk} \rangle_{2} |$. An example is the wavelet basis proposed in cohen1993wavelets.} Since the wavelet coefficients are i.i.d across $j$, it follows that the joint law of $m=(m_{1},m_{2})$ satisfies $\Pi_{\mathcal{M}}= \Pi_{\mathcal{M}_{1}}\otimes \Pi_{\mathcal{M}_{2}}$. One can also check that the support of $\Pi_{\mathcal{M}}$ is $\mathcal{M}_{1}\times \mathcal{M}_{2}$. The next proposition verifies Theorem (ref) for these priors under a subgaussian restriction on the projection errors $\varepsilon$.

propositionSuppose that Assumption (ref) holds, $\varepsilon$ is subgaussian conditional on $W$ and the conditional distribution is continuous in $W$, the marginal distribution $P_{0,W}$ of the control variable $W$ has a probability density function $p_{0,W}$ with respect to the Lebesgue measure on $[0,1]$ that is bounded away from zero and infinity, $m_{0} \in \mathcal{M}_{1} \times \mathcal{M}_{2}$ with $\mathcal{M}_{j} = \{f \in L^{2}([0,1]): ||f||_{\infty,\infty,\alpha_{0,j}} \leq M\}$ for each $j \in \{1,2\}$, where $\alpha_{0,1},\alpha_{0,2} > 1/2$ and $M < \infty$, and $\beta_{0} \in (-B,B)$ for some $B \in (0,\infty)$. Then Theorem (ref) holds for the prior $\Pi = \Pi_{\mathcal{B}} \otimes \Pi_{\mathcal{M}}$, where $\Pi_{\mathcal{B}}$ is the uniform distribution over $\mathcal{B} = [-B,B]$ and $\Pi_{\mathcal{M}}=\Pi_{\mathcal{M}_{1}}\otimes \Pi_{\mathcal{M}_{2}}$ with $\Pi_{\mathcal{M}_{j}}$ being the law of the uniform wavelet series.
remark[Model Specification and Role of Subgaussianity] Proposition (ref) indicates that the uniform wavelet series priors are a class of priors for which the Bernstein-von Mises theorem holds even when the Gaussian model ((ref))--((ref)) is misspecified (i.e., it only requires $\varepsilon$ be subgaussian conditional on $W$). The subgaussian restriction is only imposed to relate $\omega_{n}(\delta)$ to the metric entropy integrals $\int_{0}^{\delta} \sqrt{\log N(\tau,\mathcal{M}_{j},||\cdot||_{n,2})}d\tau$, $j \in \{1,2\}$, so that well-known metric entropy bounds for Besov balls (e.g., Theorem 4.3.36 of Gine_Nickl_2015) can be applied to find sequences that satisfy Part 2 of Proposition (ref).\footnote{The notation $N(\tau,\mathcal{F},||\cdot||)$ refers to the $\tau$-covering number of the set $\mathcal{F}$ under the (semi)norm $||\cdot||$. That is, the smallest number of $||\cdot||$-balls with radius $\tau$ required to cover $\mathcal{F}$. The logarithm of the covering number is known as the metric entropy.} For this reason, I consider subgaussianity to be a technical restriction and it may be that heavier tailed distributions for $\varepsilon$ can be accommodated using a more refined complexity measure for $\mathcal{M}_{j}$.
remark[No Cross Restrictions on Regularity of $m_{01}$ and $m_{02}$] There are no restrictions on the relative magnitude of $\alpha_{0,1}$ and $\alpha_{0,2}$ because it is only required that $\alpha_{0,1} > 1/2$ and $\alpha_{0,2}>1/2$. In other words, the uniform wavelet series example demonstrates that the $(\beta,m)$-parametrization can allow for $m_{01}$ to be very smooth relative to $m_{02}$ (and vice versa), so long as this baseline smoothness condition is met (i.e., regularity greater than $1/2$)

Example 2 -- Mat\'{e}rn Gaussian Process Priors

This section verifies Theorem (ref) for a popular class of Gaussian process priors. Suppose that $\mathcal{W} = [0,1]^{d_{w}}$ for some $1 \leq d_{w} < \infty$. The prior for $m=(m_{1},m_{2})$ is $\Pi_{\mathcal{M}} = \Pi_{\mathcal{M}_{1}} \otimes \Pi_{\mathcal{M}_{2}}$, where $\Pi_{\mathcal{M}_{j}}$ is the law of a centered Mat\'{e}rn Gaussian process on $[0,1]^{d_{w}}$ with regularity parameter $\alpha_{j} > 0$. Specifically, $m_{j} \sim \Pi_{\mathcal{M}_{j}}$ if and only if, for any set of indices $w_{1},...,w_{k} \in [0,1]^{d_{w}}$, the vector of function values $(m_{j}(w_{1}),...,m_{j}(w_{k}))'$ follows a centered multivariate normal distribution with covariance matrix elements given by

align*[align* omitted — 192 chars of source]

The regularity parameter $\alpha_{j}$ determines the smoothness of the sample paths because it can be shown that $m_{j}$ takes values in the H\"{o}lder space $(C^{a}([0,1]^{d_{w}}),||\cdot||_{a})$ for any $a < \alpha_{j}$ (see Section 3.1 of van2011information).\footnote{The H\"{o}lder space $C^{\alpha}([0,1])$ is the space of functions on $[0,1]^{d_{w}}$ that are $\lfloor \alpha \rfloor$-times continuously differentiable and $\lfloor \alpha \rfloor$th derivative is $(\alpha-\lfloor \alpha \rfloor)$-Lipschitz. The H\"{o}lder norm $||\cdot||_{\alpha}$ is given by (4.111) in Gine_Nickl_2015, however, its specific form is not important for Proposition (ref) (nor Proposition (ref) in Section (ref) that establishes the analogous result for the $(\beta,\eta)$-parametrization).} The next proposition verifies Theorem (ref) for the Mat\'{e}rn prior in the case where ((ref))--((ref)) is correctly specified (see Remark (ref)). For notation, $H^{\alpha}([0,1]^{d_{w}})$ denotes the Sobolev space of functions $f$ on $[0,1]^{d_{w}}$ that can be extended to a function on $\mathbb{R}^{d}$ with Fourier transform $\hat{f}$ that satisfies $\int_{\mathbb{R}^{d_{w}}}|\hat{f}(\lambda)|^{2}(1+||\lambda||^{2}_{2})^{\alpha}d \lambda < \infty$.

propositionSuppose that $\Pi = \Pi_{\mathcal{B}} \otimes \Pi_{\mathcal{M}}$, where $\Pi_{\mathcal{B}}$ is the uniform distribution over $\mathcal{B} =[-B,B]$ and $\Pi_{\mathcal{M}} = \Pi_{\mathcal{M}_{1}} \otimes \Pi_{\mathcal{M}_{2}}$ with $\Pi_{\mathcal{M}_{j}}$ being the law of a Mat\'{e}rn process with regularity $\alpha_{j} > 0$. Further, suppose that the true conditional distribution of $(Y,X)'$ given $W$ satisfies $(Y,X)'|W \sim \mathcal{N}(m_{0}(W),V(\beta_{0}))$ for $\beta_{0} \in (-B,B)$ and $m_{0}= (m_{01},m_{02})$ with $m_{0,j} \in C^{\alpha_{0,j}}([0,1]^{d_{w}}) \cap H^{\alpha_{0,j}}([0,1]^{d_{w}})$ for each $j \in \{1,2\}$. Theorem (ref) holds if the regularity parameters $\alpha_{1},\alpha_{0,1},\alpha_{2},\alpha_{0,2}>0$ satisfy the following inequalities: \begin{align} \alpha_{1} > \frac{d_{w}}{2}, \quad \alpha_{0,1} > \frac{\alpha_{1}}{2}+ \frac{d_{w}}{4} \quad \alpha_{2} > \frac{d_{w}}{2}, \quad \alpha_{0,2} > \frac{\alpha_{2}}{2}+ \frac{d_{w}}{4} . \end{align}
remark[Role of Correct Specification] Proposition (ref) assumes that the sampling model ((ref))--((ref)) is correctly specified. This is restrictive relative to Theorem (ref) (and Proposition (ref)). Correct specification is used at only one point of the proof of Proposition (ref) and that is to verify Part 2 of Proposition (ref). Specifically, I use the property $E_{P_{0,YX|W}^{(n)}}[L_{n}(\beta,m)/L_{n}(\beta_{0},m_{0})] = 1$ to show the expected value of the numerator of the posterior satisfies \begin{align*} E_{P_{0,YX|W}^{(n)}}\int_{[-B,B] \times \mathcal{M}_{n}^{c}}\frac{L_{n}(\beta,m)}{L_{n}(\beta_{0},m_{0})}d \Pi(\beta,m) \leq \Pi_{\mathcal{M}}(\mathcal{M}_{n}^{c}). \end{align*}This allows $\Pi_{\mathcal{M}}(\mathcal{M}_{c}^{c}) \leq \exp(-Mn\delta_{n}^{2})$ for $M>0$ large to verify $\Pi(\mathcal{M}_{n}^{c}|\{(Y_{i},X_{i},w_{i}')'\}_{i=1}^{n})\rightarrow 0$ in $P_{0,YX|W}^{(n)}$-probability, an inequality that the sieves $\{\mathcal{M}_{n}\}_{n \geq 1}$ introduced in the Step 2 of the proof satisfy. If it could be shown that $\Pi(\mathcal{M}_{n}^{c}|\{(Y_{i},X_{i},w_{i}')'\}_{i=1}^{n})\rightarrow 0$ in $P_{0,YX|W}^{(n)}$-probability without correct specification of ((ref))--((ref)) for the sieves $\{\mathcal{M}_{n}\}_{n \geq 1}$ (or for some modification of $\{\mathcal{M}_{n}\}_{n \geq 1}$ that does not change the metric entropy in a meaningful way), then Theorem (ref) could be verified only under the restriction that $\varepsilon$ is subgaussian conditional on $W$ because the maximal inequalities used to verify Part 3 of Proposition (ref) (and Assumption (ref)) only require subgaussianity (i.e., they apply Corollary 2.2.9 of vaartwellner96book).
remark[No Cross Restrictions on Regularity of $m_{01}$ and $m_{02}$] Proposition (ref) is similar to the uniform wavelet series example from Section (ref) in the sense that the result does not impose cross restrictions on the nuisance function. The inequalities of the proposition only compare the regularity of the prior draws $m_{j}$ to the function $m_{0,j}$ they are targeting. Specifically, the restriction $\alpha_{0,j} > \alpha_{j}/2 + d_{w}/4$ means draws $m_{j}$ from the marginal $\Pi_{\mathcal{M}_{j}}$ cannot be too smooth relative to the true nuisance parameter $m_{0,j}$, however, this does not restrict the smoothness of the draws from $\Pi_{\mathcal{M}_{1}}$ and $m_{02}$ (and vice versa). I will return to this point in Section (ref).

Discussion

This section presents a discussion of Theorem (ref) and compares the Bernstein-von Mises theorem for the $(\beta,m)$-parametrization and the $(\beta,\eta)$-parametrization of the partially linear model.

Why Theorem (ref) Works

Theorem (ref) proves that the marginal posterior for the coefficient of interest $\beta$ is asymptotically normal. The main requirements are that the marginal posterior for $m$ concentrates around $m_{0}$ at a rate faster than $n^{-1/4}$ and certain multiplier empirical processes satisfy asymptotic equicontinuity conditions. These conditions are mild and mirror those required for the asymptotic normality of method of moments estimators based on orthogonal estimating equations andrews1994asymptotics,chernozhukov2018double. Moreover, Theorem (ref) can be valid when the sampling model ((ref))--((ref)) is misspecified. Specifically, the uniform wavelet series prior example in Section (ref) satisfies Theorem (ref) without requiring that the true conditional distribution $P_{0,YX|W}$ of $(Y,X)$ given $W$ be compatible with ((ref))--((ref)). The next result is important for explaining these features.

theoremIf Assumption (ref) holds, then \begin{enumerate} • For any $\beta \in \mathcal{B}$ fixed, $m_{0}=(m_{01},m_{02})$ is the unique solution to the minimization problem \begin{align} \min_{m \in \mathcal{M}}E_{P_{0}}\left[KL\left(p_{0,YX|W}(\cdot|W),p_{\beta,m}(\cdot|W)\right)\right]. \end{align} • The true parameter $(\beta_{0},m_{0})$ is the unique solution to the minimization problem \begin{align} \min_{(\beta,m) \in \mathcal{B} \times \mathcal{M}}E_{P_{0}}\left[KL\left(p_{0,YX|W}(\cdot|W),p_{\beta,m}(\cdot|W)\right)\right], \end{align} \end{enumerate} where $KL(p,q)$ denotes the Kullback-Leibler divergence between two densities $p$ and $q$.

The connection between Theorem (ref) and Theorem (ref) is as follows. Part 1 shows that the true nuisance function $m_{0}$ maximizes the population log-likelihood function when $\beta$ is fixed, and, since the Gaussian distributions with unknown means and known variances form linear exponential families, it may be viewed as a nonparametric application of Theorem 1 of gourieroux1984pseudo (and Theorem 5.4 of white1994estimation).\footnote{Since $E_{P_{0}}[\log p_{\beta,m}(Y,X|W)] = -E_{P_{0}}[KL(p_{0,YX|W}(\cdot|W),p_{\beta,m}(\cdot|W))] + E_{P_{0}}[\log p_{0,YX|W}(Y,X|W)]$, the solutions to ((ref)) and ((ref)) coincide with that obtained when maximimizing $E_{P_{0}}[\log p_{\beta,m}(Y,X|W)]$ with $m$ and $(\beta,m)$, respectively.} The result is important for Theorem (ref) because the independence of the solution $m_{0}$ from $\beta$ reveals that there is no loss of information associated with not knowing the nuisance parameter $m_{0}$. This means the $(\beta,m)$-parametrization is an example of an adaptive semiparametric model bickel1982adaptive. Adaptive semiparametric models have the property that the ordinary score (i.e., the derivative of $\log p_{\beta,m}$ with respect to $\beta$ at $(\beta_{0},m_{0})$) coincides with the efficient score (i.e., the ordinary score of least favorable parametric submodel $\beta \mapsto p_{\beta,m_{0}}$ at $\beta_{0}$). Consequently, the asymptotic analysis of the semiparametric model is essentially that of a regular parametric model because ordinary local asymptotic normality (LAN) expansions along parametric submodels $\beta \mapsto p_{\beta,m}$ respect the semiparametric structure of the model. This makes the primary concern whether $(\beta_{0},m_{0})$ is identifiable from the sampling model ((ref))--((ref)). Part 2 confirms identifiability because it establishes that the true parameter $(\beta_{0},m_{0})$ uniquely maximizes the population log-likelihood even though the sampling model may be misspecified, and, similar to Part 1, may be viewed as a nonparametric extension of Theorem 6 of gourieroux1984pseudo since Gaussian distributions with unknown means and variances form quadratic exponential families. The sole purpose of Assumptions (ref) and (ref) is to ensure that remainders of these ordinary LAN expansions (i.e., due to estimation of $m_{0}$) vanish appropriately as the sample size grows.

remark[Efficiency] The achievement of the chamberlain1992efficiency efficiency bound under homoskedasticity despite ((ref))--((ref)) being possibly misspecified reflects the fact that maximum likelihood estimation in the least favorable parametric submodel $\{p_{\beta,m_{0}}: \beta \in \mathcal{B}\}$ is that of a linear exponential family, and, as a result, the maximum likelihood estimator has the same asymptotic distribution as a weighted least squares estimator based on the conditional moment restrictions $E_{P_{0}}[Y|X,W] = m_{01}(W)+(X-m_{02}(W))\beta_{0}$ with weights $1/\sigma_{01}^{2}$ (see, for example, Theorem 4 of gourieroux1984pseudo and Theorem 6.13 of white1994estimation).
remark[Connection to castillo2012semiparametric] Theorem (ref) establishes that the $(\beta,m)$-parametrization is an example of an adaptive semiparametric model, and, as a result, it falls within the class of semiparametric models castillo2012semiparametric refers to as `the case without loss of information' (see Section 1.3 of his paper). Consequently, it is foreseeable that Theorem 1 of his paper could be verified to establish the Bernstein-von Mises theorem (at least in the correctly specified case). The structure of the Gaussian (quasi-)likelihood makes it convenient to apply direct arguments rather than employ his more abstract, yet more generally applicable, LAN conditions.

Comparison with the Original Parametrization

I compare my results with the Bernstein-von Mises theory for the $(\beta,\eta)$-parametrization. Bayesian inference in the $(\beta,\eta)$-parametrization is typically based on the following model for the conditional distribution of $Y$ given $X$ and $W$ (for example, Section 7 of bickel2012semiparametric):

align[align omitted — 299 chars of source]

where $\beta \in \mathcal{B} \subseteq \mathbb{R}$, $\eta \in \mathcal{H} \subseteq L^{2}(\mathcal{W})$, $\Pi_{\mathcal{B}}$ is a probability distribution over $\mathcal{B}$, $\Pi_{\mathcal{H}}$ is a probability distribution over $\mathcal{H}$, and $\sigma_{01}^{2} \in (0,\infty)$ is a known constant. Since the parameters $(\beta_{0},\eta_{0})$ are identifiable from the conditional distribution of $Y$ given $X$ and $W$, there is no need to include a model for the conditional distribution of $X$ given $W$.

The key distinction between the sampling model ((ref)) and the sampling model ((ref))--((ref)) is that the former is not adaptive. Indeed, for $\beta \in \mathcal{B}$ fixed, the maximizer of the population log-likelihood function with respect to $\eta$ is $\eta_{\beta} = \eta_{0} - (\beta-\beta_{0})m_{02}$, which means there is loss of information associated with not knowing the nuisance parameter $\eta$ (unless $m_{02} = 0$). This makes finding priors/data generating processes that satisfy the Bernstein-von Mises theorem more challenging because the prior $\Pi_{\mathcal{H}}$ must be chosen so it is suitable for both estimation of $\eta_{0}$ and $m_{02}$. Proposition (ref) demonstrates this for the case where $\Pi_{\mathcal{H}}$ is the law of a centered Mat\'{e}rn Gaussian process.\footnote{The proposition also imposes the same restrictions on the data-generating process as Proposition (ref) to ensure fair comparison.} A discussion follows.

propositionSuppose that $Y|X,W \sim \mathcal{N}(X\beta_{0}+\eta_{0}(W),\sigma_{01}^{2})$ and $X|W \sim \mathcal{N}(m_{02}(W),\sigma_{02}^{2})$, where $\beta_{0} \in (-B,B)$, $\eta_{0} \in C^{\alpha_{0,\eta}}([0,1]^{d_{w}}) \cap H^{\alpha_{0,\eta}}([0,1]^{d_{w}})$, $\alpha_{0,\eta} > 0$, $m_{02} \in C^{\alpha_{0,2}}([0,1]^{d_{w}}) \cap H^{\alpha_{0,2}}([0,1]^{d_{w}})$, $\alpha_{0,2}>0$, and $\sigma_{01}^{2},\sigma_{02}^{2} \in (0,\infty)$ are known. Suppose further that Bayesian inference is based on ((ref)) with $\Pi_{\mathcal{B}}$ equal to the uniform distribution over $\mathcal{B} =[-B,B]$ and $\Pi_{\mathcal{H}}$ equal to the probability law of a centered Mat\'{e}rn Gaussian process with regularity parameter $\alpha_{\eta} > 0$. If $\alpha_{\eta},\alpha_{0,\eta}$, and $\alpha_{0,2}$ satisfy the following inequalities \begin{align} \alpha_{\eta} > \frac{d_{w}}{2}, \quad \alpha_{0,\eta} > \frac{\alpha_{\eta}}{2} + \frac{d_{w}}{4}, \quad \alpha_{0,2} > \frac{\alpha_{\eta}}{2} + \frac{d_{w}}{4}, \end{align} then the marginal posterior $\Pi\left(\beta \in \cdot |\{(Y_{i},X_{i},W_{i}')'\}_{i=1}^{n}\right)$ satisfies \begin{align*} \left | \left |\Pi\left(\beta \in \cdot |\{(Y_{i},X_{i},w_{i}')'\}_{i=1}^{n}\right) - \mathcal{N}\left(\beta_{0}+\frac{\tilde{\Delta}_{n,0}}{\sqrt{n}},\frac{1}{n}\tilde{I}_{n}(m_{0})^{-1} \right) \right | \right |_{TV} \overset{P_{0,YX|W}^{(n)}}{\longrightarrow} 0 \end{align*} for $P_{0,W}^{\infty}$-almost every fixed realization $\{w_{i}\}_{i \geq 1}$ of $\{W_{i}\}_{i \geq 1}$ as $n\rightarrow \infty$.

I compare Proposition (ref) and (ref) to elaborate on the aforementioned challenges associated with Bayesian inference for the $(\beta,\eta)$-parametrization. Since both propositions assume correct specification and are based on the same class of Gaussian process priors, the inequalities for the regularity parameters is the natural point of comparison. The first two inequalities of ((ref)) are comparable with the inequalities ((ref)) in Proposition (ref) in that they require the draws from the prior to not be too smooth relative to the function they are targeting (i.e., draws from $\Pi_{\mathcal{H}},\Pi_{\mathcal{M}_{1}}$, and $\Pi_{\mathcal{M}_{2}}$ cannot be too smooth relative to $\eta_{0}$, $m_{01}$, and $m_{02}$ respectively). However, as mentioned earlier, Proposition (ref) also restricts the regularity of the draws from the prior $\Pi_{\mathcal{H}}$ relative to $m_{02}$ because the third part of ((ref)) states that $\eta\sim \Pi_{\mathcal{H}}$ cannot be too smooth relative to $m_{02}$. This relates to the nonadaptivity of the $(\beta,\eta)$-parametrization. Indeed, verifying the appropriate LAN expansions requires a change of parametrization $(\beta,\eta)\mapsto (\beta_{0},\tilde{\eta}_{n}(\beta,\eta))$, where $\tilde{\eta}_{n}(\beta,\eta) = \eta + (\beta-\beta_{0})m_{n,2}$ for some sequence $\{m_{n,2}\}_{n \geq 1}$ in the reproducing kernel Hilbert space $(\mathbb{H}_{\eta},||\cdot||_{\mathbb{H}_{\eta}})$ of the Gaussian process $\eta$ for which $||m_{n,2}||_{\mathbb{H}_{\eta}} \leq 2\sqrt{n}\rho_{n}$ and $||m_{n,2}-m_{02}||_{\infty} \leq \rho_{n}$ for $\rho_{n} = n^{-\min\{\alpha_{\eta},\alpha_{0,2}\}/(2\alpha_{\eta}+d_{w})}$. Control of the LAN remainder requires $\rho_{n} = o(n^{-1/4})$, a condition that holds if and only if $\alpha_{\eta} > d_{w}/2$ and $\alpha_{0,2} > \alpha_{\eta}/2 + d_{w}/4$ (i.e., the first and third parts of ((ref))). This indicates that the Bernstein-von Mises theorem may fail for the $(\beta,\eta)$-parametrization in situations where $m_{02}$ is not smooth relative to $\eta \sim \Pi_{\mathcal{H}}$. In contrast, Remark (ref) points out that there are no cross-restrictions between the regularity of the nuisance functions and the draws from the prior $\Pi_{\mathcal{M}}$ in Proposition (ref) (i.e., $m_{02}$ could potentially be much less smooth than $m_{1} \sim \Pi_{\mathcal{M}_{1}}$ without violating Proposition (ref)). Put differently, the $(\beta,m)$-parametrization enables the researcher to direct nuisance priors towards the specific functions they are estimating without concern that this will adversely affect estimation the regression coefficient $\beta_{0}$.

A more subtle point of comparison relates to the requirement that the sequence $\{m_{n,2}\}_{n \geq 1}$ be in the reproducing kernel Hilbert space $(\mathbb{H}_{\eta},||\cdot||_{\mathbb{H}_{\eta}})$ of the Gaussian process $\eta$. In applying the change of parametrization $(\beta,m)\mapsto (\beta_{0},\tilde{\eta}_{n}(\beta,\eta))$, it must also be shown that the prior $\Pi_{\mathcal{H}}$ is roughly unchanged under the mapping $\eta \mapsto \tilde{\eta}_{n}(\beta,\eta)$ for $\beta$ fixed. Since, for $\beta$ fixed, $\tilde{\eta}_{n}(\beta,\eta)$ concerns a shift of a Gaussian process, it is necessary and sufficient that $\{m_{n,2}\}_{ n\geq 1}$ be in $(\mathbb{H}_{\eta},||\cdot||_{\mathbb{H}_{\eta}})$ so that the Cameron-Martin theorem (Proposition I.20 of ghosal2017fundamentals) can be applied to establish the required stability of $\Pi_{\mathcal{H}}$ via an infinite-dimensional change of measure. Since there is no need to change variables in the $(\beta,m)$-parametrization (i.e., ordinary LAN expansions respect the semiparametric structure), this prior stability condition does not arise in my proposed framework. This is important because it can be technically difficult to check this condition for non-Gaussian priors supported on infinite-dimensional spaces (i.e., Cameron-Martin is a tool specific to Gaussian priors), whereas Propositions (ref) and (ref) indicate that the key assumptions of Theorem (ref) can be verified using standard small ball probability conditions for the prior and maximal inequalities for multiplier empirical processes.\footnote{Readers should not be concerned about whether the sets $\{\mathcal{M}_{n}\}_{n \geq 1}$ in Proposition (ref) pose additional challenges for the $(\beta,m)$-parametrization relative to the $(\beta,\eta)$-parametrization. The proof of Proposition (ref) illustrates that exact analogs of the sets $\{\mathcal{M}_{n}\}_{n \geq 1}$ constructed in the proof of Proposition (ref) are utilized to verify corresponding asymptotic equicontinuity conditions in the LAN expansions for the $(\beta,\eta)$-parametrization (specifically, see $\{\tilde{\mathcal{H}}_{n}\}_{n \geq 1}$ in Step 1 of the proof of Proposition (ref)).} This allows me to check Theorem (ref) for other priors (e.g., untruncated uniform wavelet series prior) using similar technical devices to the Gaussian case. The points raised in this paragraph and the previous paragraph illustrate that there may be benefits from the $(\beta,m)$-parametrization.

Conclusion

This paper proves a Bernstein-von Mises theorem for the partially linear model. The theorem is based on a feasible adaptive parametrization of the model that is based on robinson1988root transformation. An important feature of the result is that it avoids the prior invariance condition. The examples in Sections (ref) and (ref) indicate that this is beneficial because they show that my approach does not impose cross-restrictions on the regularity of unknown functions. In contrast, the original parametrization with independent priors requires that the prior for the nuisance function $\eta$ be simultaneously appropriate for estimating $\eta_{0}$ and $m_{02}$, a condition that leads to restrictions on smoothness of the conditional expectation $m_{02}$ relative to the prior draws $\eta \sim \Pi_{\mathcal{M}}$. Moreover, the avoidance of a change of variables in the $(\beta,m)$-parametrization means that my approach can bypass an infinite-dimensional change of measure, a technical condition that may be hard to verify for complicated priors.

There are a number of extensions that would be interesting to pursue. First, it would be interesting to consider rate-adaptive priors for $m_{1}$ and $m_{2}$. I conjecture Theorem (ref) could be verified for rate adaptive priors (at least in the case of correct specification). The Mat\'{e}rn prior example in Section (ref) follows a similar argument as Theorem 3.1 of dejonge2013semiparametric, and, as a result, it is reasonable that the argument of Theorem 4.1 in their paper could also be modified to verify Theorem (ref) in the case of hierarchical spline priors, an example of a prior that can yield adaptive, rate-optimal posterior rates of contraction dejonge2012adaptive. However, the extent to which adaptive priors can be accommodated should be formally investigated. Second, it would be interesting to consider the case where $\eta(W) = W'\eta$ but the dimension of $W$ is large relative to the sample size $n$. In this case, the sampling model ((ref))--((ref)) coincides with the setup of hahn2018regularization and their simulations indicate favorable performance when using shrinkage priors for the nuisance parameters. hahn2018regularization do not provide any formal asymptotic analysis, and, for this reason, it would be interesting to examine whether an asymptotic normality result could be established in this setting. Finally, the general idea in this paper is to use knowledge of the profile likelihood severini1992profile,murphy2000profile to derive an adaptive parametrization of the partially linear model. An important extension is determining a more general class of models for which this technique could be applied.