Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
73,498 characters · 17 sections · 84 citation commands
Parametrization, Prior Independence, and the Semiparametric Bernstein-von Mises Theorem for the Partially Linear Model
This paper is concerned with Bayesian inference for a partially linear regression model. The partially linear model states that the outcome $Y \in \mathbb{R}$ is related to covariates $X \in \mathbb{R}$ and $W \in \mathbb{R}^{d_{w}}$ via the regression model
where $\beta_{0}$ is a scalar parameter, $\eta_{0}$ is a function-valued parameter, and $U$ is a scalar unobservable that satisfies $E[U|X,W] = 0$ and $Var[U|X,W] = \sigma_{01}^{2}$ for some scalar $\sigma_{01}^{2}\in (0,\infty)$. Partially linear models appear in a variety of applications including modeling the relationship between temperature and electricity demand engle1986semiparametric, economic models of production with endogenous input choices olley1996dynamics,levinsohn2003estimating,wooldridge2009estimating,ackerberg2015identification, and sample selection models ahn1993semiparametric,das2003nonparametric. Moreover, the model has also gained recent theoretical attention as a leading example for which debiased machine learning methods are applied (see, for example, chernozhukov2018double).
My analysis views $\beta_{0}$ as the parameter of interest while treating $\eta_{0}$ as an infinite-dimensional nuisance parameter. The goal is to derive conditions under which the Bernstein-von Mises theorem holds (i.e., the marginal posterior for $\beta$ is asymptotically normal and matches an efficient frequentist estimator). A key challenge with Bernstein-von Mises theory for the partially linear model is that there is loss of information associated with not knowing the nuisance parameter $\eta_{0}$, and, as a result, the natural assumption that $\beta$ and $\eta$ are independent a priori leads to prior invariance. Prior invariance refers to constructing a change of variables that respects the model's semiparametric structure and leaves the prior for the nuisance function $\eta$ roughly unchanged. It was first emphasized as a general condition for semiparametric Bernstein-von Mises theorems in the important papers of castillo2012semiparametric,castillo2012semiparametricBVM, and it introduces both practical and technical issues.\footnote{Prior invariance is also encountered in statistical functional estimation rivoirard2012bernstein,castillo2015bernstein,ray2020semiparametric. Moreover, the condition has been shown to be an almost necessary condition for the Bernstein-von Mises theorem in some models (e.g., the curve alignment model in castillo2012semiparametricBVM).} From a practical standpoint, prior invariance stipulates that the prior for $\eta$ must be adequate for estimating of the true control function $\eta_{0}$ and the conditional expectation $m_{02}$ of $X$ given $W$, a condition that restricts the smoothness $m_{02}$ relative to $\eta_{0}$, and, as a result, limits the data-generating processes for which the Bernstein-von Mises theorem can be applied. From a technical perspective, showing the prior for $\eta$ is roughly unchanged after the change of variables can be difficult to verify for some priors because it involves performing an infinite-dimensional change of measure (e.g., this is a key reason that castillo2012semiparametric,castillo2012semiparametricBVM focuses on Gaussian priors).
The main contribution of this paper is a Bernstein-von Mises theorem for the partially linear model that bypasses prior invariance. Theorem (ref) in Section (ref) establishes this result, and the theorem is verified for two examples in Section (ref). A common theme in the examples is that there are no cross-restrictions on the smoothness of nuisance functions, which demonstrates that my approach can eliminate the aforementioned practical challenge of requiring a prior for a single nuisance parameter (i.e., $\eta$) be suitable for estimation of multiple quantities (i.e., $\eta_{0}$ and $m_{02}$). To make things concrete, my proposal is to embed an interest-respecting reparametrization of ((ref)) given by
where $m_{01}(W)=E_{P_{0}}[Y|W]$ and $m_{02}(W)=E_{P_{0}}[Y|W]$, within a (quasi-)likelihood model that is suitable for identification of $(\beta_{0},m_{01},m_{02})$, and then perform Bayesian inference by placing independent priors on $\beta$ and $m=(m_{1},m_{2})$. The transformed regression model ((ref)) is commonly referred to as the robinson1988root transformation, and is observationally equivalent to ((ref)) because $\eta_{0}$ is identified as $m_{01}-\beta_{0}m_{02}$. A key property of the (quasi-)likelihood model is that there is no loss of information associated with not knowing the nuisance parameter $m_{0}=(m_{01},m_{02})$ (i.e., it is adaptive in the sense of bickel1982adaptive). Consequently, there is no need to change variables because ordinary/parametric local asymptotic normality expansions respect the semiparametric structure of the model. For terminology, I refer to ((ref)) as the $(\beta,\eta)$-parametrization and ((ref)) as the $(\beta,m)$-parametrization throughout.
To my knowledge, Theorem (ref) is the first Bernstein-von Mises theorem for the $(\beta,m)$-parametrization of the partially linear model. It has several important takeaways. The first takeaway is that the findings demonstrate that the choice of parametrization can influence the conditions under which the semiparametric Bernstein-von Mises theorem holds. This is demonstrated explicitly in Section (ref) in which the $(\beta,\eta)$- and $(\beta,m)$-parametrizations are compared for Mat\'{e}rn Gaussian process priors, a class of Gaussian process priors popular in applications (see, for example, williams2006gaussian). Since the $(\beta,m)$-parametrization is orthogonal (and the $(\beta,\eta)$-parametrization is not), this takeaway is similar to the recent frequentist debiased machine learning literature chernozhukov2018double and is reminiscent of classical results on improving parametric likelihood estimators via orthogonal parametrizations cox1973parameter. The second takeaway is that the use of reparametrization to eliminate prior invariance conditions avoids data-driven prior/posterior adjustments. For this reason, my proposal offers a conceptually distinct approach to obtaining debiased Bernstein-von Mises theorems relative to those encountered in a growing literature on data-driven prior/posterior adjustments yang2015semiparametric,ray2019debiased,ray2020semiparametric,breunig2022double,yiu2025semiparametric. Determining whether the approach of this paper is more generally applicable is an important area for future research. A final takeaway is that the Bernstein-von Mises theorem can be valid even if the (quasi-)likelihood model is misspecified. This reflects classical results on quasi-maximum likelihood estimation gourieroux1984pseudo,white1994estimation, and Section (ref) makes the claim precise by verifying Theorem (ref) for uniform wavelet series priors without requiring the sampling model be correctly specified.
This paper connects to several active literatures in statistics and econometrics. First, it is related to the literature on semiparametric Bernstein-von Mises theorems. A Bernstein-von Mises theorem for the $(\beta,\eta)$-parametrization with independent priors for $\beta$ and $\eta$ is studied in shen2002asymptotic, bickel2012semiparametric, yang2015semiparametric, and xie2020adaptive. castillo2012semiparametric,castillo2012semiparametricBVM are also relevant to my paper because these important papers emphasize the role of the prior invariance condition in separated semiparametric models with information loss. Relatedly, rivoirard2012bernstein and castillo2015bernstein demonstrated the importance of prior invariance for general smooth functionals of nonparametric models (with the latter treating separated semiparametric models as a special case). yang2015semiparametric,ray2020semiparametric,breunig2022double, and yiu2025semiparametric propose data-driven prior/posterior corrections to prove Bernstein-von Mises theorems under weaker conditions than their unmodified counterparts. A key difference between these papers and my proposal is that I do not require any data-driven adjustment because the lack of prior invariance follows from the adaptivity of the $(\beta,m)$-parametrization. yang2019posterior derives Bernstein-von Mises theorem for a single coordinate in a high-dimensional linear regression model under sparsity using a different type of reparametrization and a prior correction. Although I do not explicitly study high-dimensional linear regression, Section (ref) outlines how my approach can be extended to this setting and an interesting subject for future research would be determining whether my approach leads to results similar to yang2019posterior in this setting. Other papers on semiparametric Bernstein-von Mises theorems include kim2006bernstein, dejonge2013semiparametric, norets2015bayesian, chae2019semi, nickl2019bernstein, nickl2020bernstein, monard2021statistical, and l2023semiparametric.
Second, some of the auxiliary results in Section (ref) connect this paper to the literature on posterior contraction rates in nonparametric models. Specifically, Propositions (ref) and (ref) provide a general approach to deriving posterior contraction rates for nuisance functions $m_{01}$ and $m_{02}$. These propositions are based on empirical process conditions that arise from direct expansions of the (quasi-)likelihood ratio process, and, for this reason, are conceptually similar to the convergence rate framework of shen2001rates in which empirical process-type conditions for likelihood ratios (e.g., bracketing integrals) are used to derive posterior contraction rates. Since Propositions (ref) and (ref) concern posterior contraction rates in (possibly) misspecified Gaussian nonparametric regression, the auxiliary results in Section (ref) are also related to the theory presented in Section 4 of kleijn2006misspecification. Other papers on posterior contraction rates in general nonparametric models include ghosal2000convergence,ghosal2007convergence,castillo2008lower,vaart2008rates,10.1214/08-AOS678,van2011information,gine2011rates,dejonge2012adaptive,castillo2014supremum, and shen2015adaptive.
Finally, since the Bernstein-von Mises theorem establishes asymptotic equivalence between Bayesian and frequentist estimators, my paper also dovetails with the literature on frequentist semiparametric estimation. The $(\beta,m)$-parametrization was first utilized in robinson1988root to derive asymptotically normal least squares estimators of $\beta_{0}$ based on kernel estimators of $m_{01}$ and $m_{02}$. DONALD199430 also apply the partialling out technique of robinson1988root to show that asymptotic normality of the least squares estimators of $\beta_{0}$ can be achieved under very mild conditions when series regression methods are used to estimate the nuisance functions. My paper provides a Bayesian analogue to these papers because it shows the $(\beta,m)$-parametrization can be used to derive a Bernstein-von Mises theorem that bypasses prior invariance. The proposed Bayesian inference framework is also related to the literature on profile likelihood estimation severini1992profile, murphy2000profile because the transformed regression model ((ref)) mimics the Gaussian profile maximum likelihood problem for ((ref)). Since $(\beta,m)$-parametrization is orthogonal, my paper is also related to the literature on two-step semiparametric estimation based on Neyman orthogonal moment conditions (e.g., andrews1994asymptotics, newey1994asymptotic,chernozhukov2018double). Using a (quasi-)likelihood to perform Bayesian inference a low-dimensional parameter of interest defined by conditional moment restrictions also connects my proposal to the quasi-maximum likelihood estimation literature gourieroux1984pseudo,white1994estimation,KOMUNJER2005137.
The rest of the paper is organized as follows. Section (ref) formalizes the data-generating process and puts forward the Bayesian inference framework. Section (ref) derives conditions under which the marginal posterior for $\beta$ satisfies a Bernstein-von Mises theorem and verifies the conditions for two important classes of priors. Section (ref) provides an extensive discussion the findings of Section (ref). Section (ref) concludes and proposes some directions for future research. All proofs are in the Appendix and notation is introduced when appropriate.
Let $Y \in \mathcal{Y} \subseteq \mathbb{R}$ be an outcome variable, let $X \in \mathcal{X} \subseteq \mathbb{R}$ be a covariate of interest, let $W \in \mathcal{W} \subseteq \mathbb{R}^{d_{w}}$, $d_{w} < \infty$, be some control variables, and let $P_{0}$ be the joint distribution of $(Y,X,W')'$. The observed data $\{(Y_{i},X_{i},W_{i}')'\}_{i=1}^{n}$ consists of the first $n$ elements of a sequence $\{(Y_{i},X_{i},W_{i}')'\}_{i \geq 1}$ of independent and identically distributed (i.i.d) random vectors drawn from $P_{0}$. Assumption (ref) states the restrictions on the data distribution $P_{0}$.
I briefly discuss Assumption (ref). Part 1 of Assumption (ref) imposes that the conditional moment restriction ((ref)) is satisfied at $P_{0}$ and that the projection error $Y - X\beta_{0}-\eta_{0}(W)$ is homoskedastic. I treat $\sigma_{01}^{2}$ as known for exposition, however, Appendix (ref) extends the main results to accommodate unknown $\sigma_{01}^{2}$ and multivariate $X$. Part 2 of Assumption (ref) is a strong identification condition that is necessary and sufficient for regular estimation of $\beta_{0}$ robinson1988root. The remaining parts are technical conditions that enable the application of stochastic limit theorems and ensure certain population objective functions are well-defined.
Bayesian inference starts with a conditional model for the data (sampling model) and a distribution over the model parameters (prior). I describe each of these in turn. The sampling model is
for $i=1,...,n$, where $\beta \in \mathcal{B}$, $m = (m_{1},m_{2}) \in \mathcal{M}$ with $\mathcal{M} = \mathcal{M}_{1} \times \mathcal{M}_{2}$ and $\mathcal{M}_{j} \subseteq L^{2}(\mathcal{W})$ for $j \in \{1,2\}$, and $\sigma_{01}^{2},\sigma_{02}^{2} \in (0,\infty)$ are known scalars. Display ((ref)) is a model for the conditional distribution of $Y$ given $X$ and $W$ that imposes the robinson1988root transformation of the partially linear model, a result that establishes the observational equivalence between ((ref)) and the ‘partialed out' regression model $Y = m_{01}(W) + (X-m_{02}(W))\beta_{0} + U$, where $m_{01}(W) = E_{P_{0}}[Y|W]$, $m_{02}(W) = E_{P_{0}}[X|W]$, and $U$ satisfies $E[U|X,W] = 0$ and $Var[U|X,W] = \sigma_{01}^{2}$. Display ((ref)) is a model for the conditional distribution of $X$ given $W$ that is included to account for the fact that $m_{01}$ and $m_{02}$ are not separately identified from ((ref)).\footnote{Since ((ref)) is for identification of $m_{02}$ only, $\sigma_{02}^{2}$ has no intrinsic meaning and can be set at an arbitrary fixed value (e.g., $\sigma_{02}^{2}=1$).} Together, ((ref)) and ((ref)) imply a (quasi-)likelihood $L_{n}(\beta,m)$ given by
where $p_{\beta,m}(\cdot|w)$ is the probability density function of a bivariate Gaussian distribution with mean vector $m(w)=(m_{1}(w),m_{2}(w))'$ and covariance matrix
The term `(quasi-)likelihood' is used because the set of conditional distributions for $(Y,X)$ given $W$ compatible with Assumption (ref) is much larger than those that are compatible with ((ref))--((ref)). Sections (ref) and (ref) elaborate on the sampling model's suitability for estimating $\beta_{0}$ despite it being much more restrictive than Assumption (ref).
Since the model parameters are $(\beta,m)$, the prior $\Pi$ is a probability distribution over $\mathcal{B} \times \mathcal{M}$. Throughout, I maintain that $\beta \sim \Pi_{\mathcal{B}}$, $m \sim \Pi_{\mathcal{M}}$, and $\beta \protect\mathpalette{\protect\independenT}{\perp} m$, where $\Pi_{\mathcal{B}}$ and $\Pi_{\mathcal{M}}$ are probability measures over $\mathcal{B}$ and $\mathcal{M}$, respectively. This means that $\Pi = \Pi_{\mathcal{B}} \otimes \Pi_{\mathcal{M}}$. Prior independence between the low-dimensional target parameter and the infinite-dimensional nuisance parameters is standard in semiparametric Bayesian inference bickel2012semiparametric,castillo2012semiparametric,castillo2012semiparametricBVM,ghosal2017fundamentals. Assumption (ref) states some regularity conditions for $\Pi$.
The sampling model ((ref))--((ref)) and prior $\Pi$ implies a conditional distribution $\Pi((\beta,m) \in \cdot | \{(Y_{i},X_{i},W_{i}')'\}_{i=1}^{n})$ for the model parameters $(\beta,m)$ given the data $\{(Y_{i},X_{i},W_{i}')'\}_{i=1}^{n}$. This is known as the posterior, and, under Assumption (ref), posterior probabilities are computed via Bayes rule. That is, for any event $A$,
Bayesian inference revolves around the posterior distribution. As examples, researchers interested in a point estimator for $\beta$ can report the posterior median or posterior mean, while researchers interested in quantifying uncertainty about $\beta$ report a credible set, a set estimator $CS_{n}(1-\alpha)$ that satisfies $\Pi(\beta \in CS_{n}(1-\alpha)|\{(Y_{i},X_{i},W_{i}')'\}_{i=1}^{n}) \geq 1-\alpha$.
This section provides Bernstein-von Mises theory for the marginal posterior for $\beta$. It has two subsections. The first establishes a general Bernstein-von Mises theorem based on high-level posterior consistency and empirical process conditions, and the second verifies the assumptions for two classes of priors.
I prove a general Bernstein-von Mises theorem for the marginal posterior of $\beta$. The asymptotic analysis is conducted conditional on the realizations of $\{W_{i}\}_{i \geq 1}$, meaning that $\{(Y_{i},X_{i})\}_{i=1}^{n}$ should be viewed as a sequence of independent but not identically distributed random vectors (i.e., $(Y_{i},X_{i}) \protect\mathpalette{\protect\independenT}{\perp} (Y_{j},X_{j})|\{W_{i}\}_{i \geq 1}$ for $i \neq j$). For notation, $P_{0,W}^{\infty}$ is the joint law of the i.i.d sequence $\{W_{i}\}_{i \geq 1}$, $\langle f, g \rangle_{n,2} = n^{-1}\sum_{i=1}^{n}f(w_{i})g(w_{i})$ and $||f||_{n,2}= \sqrt{\langle f,f \rangle_{n,2}}$ are the empirical $L^{2}$ inner product and induced norm, respectively, for functions $f,g: \mathcal{W} \rightarrow \mathbb{R}$ and for a given realization $\{w_{i}\}_{i \geq 1}$ of $\{W_{i}\}_{i \geq 1}$, $B_{n,\mathcal{M}}(m_{0},\delta)$ is a ball centered at $m_{0}$ with radius $\delta > 0$ in the product empirical $L^{2}$ norm $\max\{||m_{1}||_{n,2},||m_{2}||_{n,2}\}$, and $P_{0,YX|W}^{(n)} = \bigotimes_{i=1}^{n}P_{0,YX|w_{i}}$ with $P_{0,YX|w_{i}}$ denoting the probability distribution associated with the probability density function $p_{0,YX|W}(\cdot|w_{i})$ for $i=1,...,n$. The next two assumptions concern the limiting behavior of the posterior.
Assumptions (ref) and (ref) are high-level conditions about the the limiting behavior of the posterior. Assumption (ref) enforces that the marginal posterior for $m$ concentrates around $m_{0}$ in the empirical $L^{2}$-norm and the rate of convergence $\delta_{n}$ is faster than $n^{-1/4}$. Assumption (ref) imposes that there are sets $\{\mathcal{M}_{n}\}_{n \geq 1}$ that the posterior concentrates on such that, when intersected with $B_{n,\mathcal{M}}(m_{0},\delta_{n})$, are structured enough to ensure that multiplier empirical processes $\{G_{n}^{(1)}\}_{n\geq 1}$ and $\{G_{n}^{(2)}\}_{n \geq 1}$ satisfy stochastic equicontinuity-type conditions. These requirements are conceptually similar to those encountered in the literature on asymptotic normality of plug-in method of moments estimators based on orthogonal moment conditions andrews1994asymptotics,chernozhukov2018double and semiparametric maximum likelihood estimators murphy2000profile. Combined with Assumptions (ref) and (ref), Assumptions (ref) and (ref) imply that the marginal posterior for $\beta$ satisfies a Bernstein-von Mises theorem.
This section has two subsections. Section (ref) presents sufficient conditions for the assumptions in Section (ref). Sections (ref) and (ref) verify Theorem (ref) for two classes of priors using these sufficient conditions.
I start with a proposition that enables finding sequences $\{\delta_{n}\}_{ n\geq 1}$ for which Assumption (ref) holds. It has similarities with Theorem 2 of shen2001rates in that it establishes posterior consistency by directly analyzing the (quasi-)likelihood ratio process. For notation, let $\varepsilon = (Y-m_{01}(W),X-m_{02}(W))'$, and let $\lambda_{min}(A)$ and $\lambda_{max}(A)$ denote the minimum and maximum eigenvalues of a matrix $A$, respectively.
Proposition (ref) states that $\delta_{n}$ is determined by the 1. mass the nuisance prior $\Pi_{\mathcal{M}}$ assigns to empirical $L^{2}$ neighborhoods of $m_{0}$, and 2. the rate of convergence of the least squares estimator of $m_{0}$. Since $||\cdot||_{n,2} \leq ||\cdot||_{\infty}$ with $||f||_{\infty}= \sup_{w \in \mathcal{W}}|f(w)|$ denoting the supremum norm, the prior mass condition can be checked by analyzing the probability mass that $\Pi_{\mathcal{M}}$ assigns to uniform balls centered at $m_{0}$, a quantity that has been studied for many different priors ghosal2017fundamentals. The least squares connection is through the second condition because $\delta \mapsto \omega_{n}(\delta)$ can be identified as the `continuity modulus' of the multiplier empirical process that determines the least squares rate of convergence (see, for example, Chapter 9 of geer2000empirical and Section 3.2 of vaartwellner96book). Given a valid $\delta_{n}$, conditions that ensure the rate restriction $\delta_{n} = o(n^{-1/4})$ can be established.
I verify Theorem (ref) for a class of priors based on untruncated uniform wavelet series. These priors have feature regularly in the Bayesian nonparametrics literature gine2011rates,10.1214/13-AOS1133,castillo2014supremum,l2023semiparametric, however, as far as I am aware, the Bernstein-von Mises theorem for the partially linear model has not been verified using these priors. To make things formal, suppose that $\mathcal{W}= [0,1]$, the marginal distribution $P_{0,W}$ of $W$ has a Lebesgue density $p_{0,W}$ on $[0,1]$ that is bounded away from zero and infinity, and that there are known constants $M>0$ and $\alpha_{0,1},\alpha_{0,2}>1/2$ such that $||m_{0,j}||_{\infty,\infty,\alpha_{0,j}} \leq M$ for all $j \in \{1,2\}$, where $||\cdot||_{\infty,\infty,\alpha}$ denotes the norm of a H\"{o}lder-Zygmund space $B_{\infty,\infty}^{\alpha}([0,1])$, $\alpha > 0$.\footnote{For $\alpha \notin \mathbb{N}$, H\"{o}lder-Zygmund spaces coincide with the usual H\"{o}lder space $C^{\alpha}([0,1])$ of functions that are $\lfloor \alpha \rfloor$-times continuously differentiable and $\lfloor \alpha \rfloor$th derivative is $(\alpha-\lfloor \alpha \rfloor)$-Lipschitz and the norm $||\cdot||_{\infty,\infty,\alpha}$ is equivalent to the usual H\"{o}lder norm $||\cdot||_{\alpha}$ (i.e., given by (4.111) in Gine_Nickl_2015). However, for $\alpha \in \mathbb{N}$, these spaces are no longer the same. See Section 4.3.3 of Gine_Nickl_2015 for more details.} Given these assumptions, I set the nuisance parameter space to be $\mathcal{M}_{j}=\{f \in L^{2}([0,1]): ||f||_{\infty,\infty,\alpha_{0,j}} \leq M\}$ for each $j \in \{1,2\}$ and a prior $\Pi_{\mathcal{M}}$ over $\mathcal{M}_{1} \times \mathcal{M}_{2}$ is constructed as follows. I take a boundary adapted wavelet basis $\{\psi_{lk}: l \geq 0, 0 \leq k \leq 2^{l}-1\}$ for $L^{2}([0,1])$ that is sufficiently regular to characterize the H\"{o}lder-Zygmund spaces $(B_{\infty,\infty}^{\alpha}([0,1]),||\cdot||_{\infty,\infty,\alpha})$, $\alpha \in \{\alpha_{0,1},\alpha_{0,2}\}$, and set the marginal prior $\Pi_{\mathcal{M}_{j}}$ for $m_{j}$ to be the law of the random function
where $\nu_{lk,j}$ are independent and identically distributed (across $l$, $k$, and $j$) uniform random variables supported on $[-M,M]$.\footnote{The wavelet basis being sufficiently regular to characterize $(B^{\alpha}_{\infty,\infty}([0,1]),||\cdot||_{\alpha,\infty,\infty})$ means that the norm $||\cdot||_{\infty,\infty,\alpha}$ satisfies $||f||_{\infty,\infty,\alpha}=\sup_{l \geq 0}\max_{0 \leq k \leq 2^{l}-1}2^{l(\alpha+1/2)}|\langle f, \psi_{lk} \rangle_{2} |$. An example is the wavelet basis proposed in cohen1993wavelets.} Since the wavelet coefficients are i.i.d across $j$, it follows that the joint law of $m=(m_{1},m_{2})$ satisfies $\Pi_{\mathcal{M}}= \Pi_{\mathcal{M}_{1}}\otimes \Pi_{\mathcal{M}_{2}}$. One can also check that the support of $\Pi_{\mathcal{M}}$ is $\mathcal{M}_{1}\times \mathcal{M}_{2}$. The next proposition verifies Theorem (ref) for these priors under a subgaussian restriction on the projection errors $\varepsilon$.
This section verifies Theorem (ref) for a popular class of Gaussian process priors. Suppose that $\mathcal{W} = [0,1]^{d_{w}}$ for some $1 \leq d_{w} < \infty$. The prior for $m=(m_{1},m_{2})$ is $\Pi_{\mathcal{M}} = \Pi_{\mathcal{M}_{1}} \otimes \Pi_{\mathcal{M}_{2}}$, where $\Pi_{\mathcal{M}_{j}}$ is the law of a centered Mat\'{e}rn Gaussian process on $[0,1]^{d_{w}}$ with regularity parameter $\alpha_{j} > 0$. Specifically, $m_{j} \sim \Pi_{\mathcal{M}_{j}}$ if and only if, for any set of indices $w_{1},...,w_{k} \in [0,1]^{d_{w}}$, the vector of function values $(m_{j}(w_{1}),...,m_{j}(w_{k}))'$ follows a centered multivariate normal distribution with covariance matrix elements given by
The regularity parameter $\alpha_{j}$ determines the smoothness of the sample paths because it can be shown that $m_{j}$ takes values in the H\"{o}lder space $(C^{a}([0,1]^{d_{w}}),||\cdot||_{a})$ for any $a < \alpha_{j}$ (see Section 3.1 of van2011information).\footnote{The H\"{o}lder space $C^{\alpha}([0,1])$ is the space of functions on $[0,1]^{d_{w}}$ that are $\lfloor \alpha \rfloor$-times continuously differentiable and $\lfloor \alpha \rfloor$th derivative is $(\alpha-\lfloor \alpha \rfloor)$-Lipschitz. The H\"{o}lder norm $||\cdot||_{\alpha}$ is given by (4.111) in Gine_Nickl_2015, however, its specific form is not important for Proposition (ref) (nor Proposition (ref) in Section (ref) that establishes the analogous result for the $(\beta,\eta)$-parametrization).} The next proposition verifies Theorem (ref) for the Mat\'{e}rn prior in the case where ((ref))--((ref)) is correctly specified (see Remark (ref)). For notation, $H^{\alpha}([0,1]^{d_{w}})$ denotes the Sobolev space of functions $f$ on $[0,1]^{d_{w}}$ that can be extended to a function on $\mathbb{R}^{d}$ with Fourier transform $\hat{f}$ that satisfies $\int_{\mathbb{R}^{d_{w}}}|\hat{f}(\lambda)|^{2}(1+||\lambda||^{2}_{2})^{\alpha}d \lambda < \infty$.
This section presents a discussion of Theorem (ref) and compares the Bernstein-von Mises theorem for the $(\beta,m)$-parametrization and the $(\beta,\eta)$-parametrization of the partially linear model.
Theorem (ref) proves that the marginal posterior for the coefficient of interest $\beta$ is asymptotically normal. The main requirements are that the marginal posterior for $m$ concentrates around $m_{0}$ at a rate faster than $n^{-1/4}$ and certain multiplier empirical processes satisfy asymptotic equicontinuity conditions. These conditions are mild and mirror those required for the asymptotic normality of method of moments estimators based on orthogonal estimating equations andrews1994asymptotics,chernozhukov2018double. Moreover, Theorem (ref) can be valid when the sampling model ((ref))--((ref)) is misspecified. Specifically, the uniform wavelet series prior example in Section (ref) satisfies Theorem (ref) without requiring that the true conditional distribution $P_{0,YX|W}$ of $(Y,X)$ given $W$ be compatible with ((ref))--((ref)). The next result is important for explaining these features.
The connection between Theorem (ref) and Theorem (ref) is as follows. Part 1 shows that the true nuisance function $m_{0}$ maximizes the population log-likelihood function when $\beta$ is fixed, and, since the Gaussian distributions with unknown means and known variances form linear exponential families, it may be viewed as a nonparametric application of Theorem 1 of gourieroux1984pseudo (and Theorem 5.4 of white1994estimation).\footnote{Since $E_{P_{0}}[\log p_{\beta,m}(Y,X|W)] = -E_{P_{0}}[KL(p_{0,YX|W}(\cdot|W),p_{\beta,m}(\cdot|W))] + E_{P_{0}}[\log p_{0,YX|W}(Y,X|W)]$, the solutions to ((ref)) and ((ref)) coincide with that obtained when maximimizing $E_{P_{0}}[\log p_{\beta,m}(Y,X|W)]$ with $m$ and $(\beta,m)$, respectively.} The result is important for Theorem (ref) because the independence of the solution $m_{0}$ from $\beta$ reveals that there is no loss of information associated with not knowing the nuisance parameter $m_{0}$. This means the $(\beta,m)$-parametrization is an example of an adaptive semiparametric model bickel1982adaptive. Adaptive semiparametric models have the property that the ordinary score (i.e., the derivative of $\log p_{\beta,m}$ with respect to $\beta$ at $(\beta_{0},m_{0})$) coincides with the efficient score (i.e., the ordinary score of least favorable parametric submodel $\beta \mapsto p_{\beta,m_{0}}$ at $\beta_{0}$). Consequently, the asymptotic analysis of the semiparametric model is essentially that of a regular parametric model because ordinary local asymptotic normality (LAN) expansions along parametric submodels $\beta \mapsto p_{\beta,m}$ respect the semiparametric structure of the model. This makes the primary concern whether $(\beta_{0},m_{0})$ is identifiable from the sampling model ((ref))--((ref)). Part 2 confirms identifiability because it establishes that the true parameter $(\beta_{0},m_{0})$ uniquely maximizes the population log-likelihood even though the sampling model may be misspecified, and, similar to Part 1, may be viewed as a nonparametric extension of Theorem 6 of gourieroux1984pseudo since Gaussian distributions with unknown means and variances form quadratic exponential families. The sole purpose of Assumptions (ref) and (ref) is to ensure that remainders of these ordinary LAN expansions (i.e., due to estimation of $m_{0}$) vanish appropriately as the sample size grows.
I compare my results with the Bernstein-von Mises theory for the $(\beta,\eta)$-parametrization. Bayesian inference in the $(\beta,\eta)$-parametrization is typically based on the following model for the conditional distribution of $Y$ given $X$ and $W$ (for example, Section 7 of bickel2012semiparametric):
where $\beta \in \mathcal{B} \subseteq \mathbb{R}$, $\eta \in \mathcal{H} \subseteq L^{2}(\mathcal{W})$, $\Pi_{\mathcal{B}}$ is a probability distribution over $\mathcal{B}$, $\Pi_{\mathcal{H}}$ is a probability distribution over $\mathcal{H}$, and $\sigma_{01}^{2} \in (0,\infty)$ is a known constant. Since the parameters $(\beta_{0},\eta_{0})$ are identifiable from the conditional distribution of $Y$ given $X$ and $W$, there is no need to include a model for the conditional distribution of $X$ given $W$.
The key distinction between the sampling model ((ref)) and the sampling model ((ref))--((ref)) is that the former is not adaptive. Indeed, for $\beta \in \mathcal{B}$ fixed, the maximizer of the population log-likelihood function with respect to $\eta$ is $\eta_{\beta} = \eta_{0} - (\beta-\beta_{0})m_{02}$, which means there is loss of information associated with not knowing the nuisance parameter $\eta$ (unless $m_{02} = 0$). This makes finding priors/data generating processes that satisfy the Bernstein-von Mises theorem more challenging because the prior $\Pi_{\mathcal{H}}$ must be chosen so it is suitable for both estimation of $\eta_{0}$ and $m_{02}$. Proposition (ref) demonstrates this for the case where $\Pi_{\mathcal{H}}$ is the law of a centered Mat\'{e}rn Gaussian process.\footnote{The proposition also imposes the same restrictions on the data-generating process as Proposition (ref) to ensure fair comparison.} A discussion follows.
I compare Proposition (ref) and (ref) to elaborate on the aforementioned challenges associated with Bayesian inference for the $(\beta,\eta)$-parametrization. Since both propositions assume correct specification and are based on the same class of Gaussian process priors, the inequalities for the regularity parameters is the natural point of comparison. The first two inequalities of ((ref)) are comparable with the inequalities ((ref)) in Proposition (ref) in that they require the draws from the prior to not be too smooth relative to the function they are targeting (i.e., draws from $\Pi_{\mathcal{H}},\Pi_{\mathcal{M}_{1}}$, and $\Pi_{\mathcal{M}_{2}}$ cannot be too smooth relative to $\eta_{0}$, $m_{01}$, and $m_{02}$ respectively). However, as mentioned earlier, Proposition (ref) also restricts the regularity of the draws from the prior $\Pi_{\mathcal{H}}$ relative to $m_{02}$ because the third part of ((ref)) states that $\eta\sim \Pi_{\mathcal{H}}$ cannot be too smooth relative to $m_{02}$. This relates to the nonadaptivity of the $(\beta,\eta)$-parametrization. Indeed, verifying the appropriate LAN expansions requires a change of parametrization $(\beta,\eta)\mapsto (\beta_{0},\tilde{\eta}_{n}(\beta,\eta))$, where $\tilde{\eta}_{n}(\beta,\eta) = \eta + (\beta-\beta_{0})m_{n,2}$ for some sequence $\{m_{n,2}\}_{n \geq 1}$ in the reproducing kernel Hilbert space $(\mathbb{H}_{\eta},||\cdot||_{\mathbb{H}_{\eta}})$ of the Gaussian process $\eta$ for which $||m_{n,2}||_{\mathbb{H}_{\eta}} \leq 2\sqrt{n}\rho_{n}$ and $||m_{n,2}-m_{02}||_{\infty} \leq \rho_{n}$ for $\rho_{n} = n^{-\min\{\alpha_{\eta},\alpha_{0,2}\}/(2\alpha_{\eta}+d_{w})}$. Control of the LAN remainder requires $\rho_{n} = o(n^{-1/4})$, a condition that holds if and only if $\alpha_{\eta} > d_{w}/2$ and $\alpha_{0,2} > \alpha_{\eta}/2 + d_{w}/4$ (i.e., the first and third parts of ((ref))). This indicates that the Bernstein-von Mises theorem may fail for the $(\beta,\eta)$-parametrization in situations where $m_{02}$ is not smooth relative to $\eta \sim \Pi_{\mathcal{H}}$. In contrast, Remark (ref) points out that there are no cross-restrictions between the regularity of the nuisance functions and the draws from the prior $\Pi_{\mathcal{M}}$ in Proposition (ref) (i.e., $m_{02}$ could potentially be much less smooth than $m_{1} \sim \Pi_{\mathcal{M}_{1}}$ without violating Proposition (ref)). Put differently, the $(\beta,m)$-parametrization enables the researcher to direct nuisance priors towards the specific functions they are estimating without concern that this will adversely affect estimation the regression coefficient $\beta_{0}$.
A more subtle point of comparison relates to the requirement that the sequence $\{m_{n,2}\}_{n \geq 1}$ be in the reproducing kernel Hilbert space $(\mathbb{H}_{\eta},||\cdot||_{\mathbb{H}_{\eta}})$ of the Gaussian process $\eta$. In applying the change of parametrization $(\beta,m)\mapsto (\beta_{0},\tilde{\eta}_{n}(\beta,\eta))$, it must also be shown that the prior $\Pi_{\mathcal{H}}$ is roughly unchanged under the mapping $\eta \mapsto \tilde{\eta}_{n}(\beta,\eta)$ for $\beta$ fixed. Since, for $\beta$ fixed, $\tilde{\eta}_{n}(\beta,\eta)$ concerns a shift of a Gaussian process, it is necessary and sufficient that $\{m_{n,2}\}_{ n\geq 1}$ be in $(\mathbb{H}_{\eta},||\cdot||_{\mathbb{H}_{\eta}})$ so that the Cameron-Martin theorem (Proposition I.20 of ghosal2017fundamentals) can be applied to establish the required stability of $\Pi_{\mathcal{H}}$ via an infinite-dimensional change of measure. Since there is no need to change variables in the $(\beta,m)$-parametrization (i.e., ordinary LAN expansions respect the semiparametric structure), this prior stability condition does not arise in my proposed framework. This is important because it can be technically difficult to check this condition for non-Gaussian priors supported on infinite-dimensional spaces (i.e., Cameron-Martin is a tool specific to Gaussian priors), whereas Propositions (ref) and (ref) indicate that the key assumptions of Theorem (ref) can be verified using standard small ball probability conditions for the prior and maximal inequalities for multiplier empirical processes.\footnote{Readers should not be concerned about whether the sets $\{\mathcal{M}_{n}\}_{n \geq 1}$ in Proposition (ref) pose additional challenges for the $(\beta,m)$-parametrization relative to the $(\beta,\eta)$-parametrization. The proof of Proposition (ref) illustrates that exact analogs of the sets $\{\mathcal{M}_{n}\}_{n \geq 1}$ constructed in the proof of Proposition (ref) are utilized to verify corresponding asymptotic equicontinuity conditions in the LAN expansions for the $(\beta,\eta)$-parametrization (specifically, see $\{\tilde{\mathcal{H}}_{n}\}_{n \geq 1}$ in Step 1 of the proof of Proposition (ref)).} This allows me to check Theorem (ref) for other priors (e.g., untruncated uniform wavelet series prior) using similar technical devices to the Gaussian case. The points raised in this paragraph and the previous paragraph illustrate that there may be benefits from the $(\beta,m)$-parametrization.
This paper proves a Bernstein-von Mises theorem for the partially linear model. The theorem is based on a feasible adaptive parametrization of the model that is based on robinson1988root transformation. An important feature of the result is that it avoids the prior invariance condition. The examples in Sections (ref) and (ref) indicate that this is beneficial because they show that my approach does not impose cross-restrictions on the regularity of unknown functions. In contrast, the original parametrization with independent priors requires that the prior for the nuisance function $\eta$ be simultaneously appropriate for estimating $\eta_{0}$ and $m_{02}$, a condition that leads to restrictions on smoothness of the conditional expectation $m_{02}$ relative to the prior draws $\eta \sim \Pi_{\mathcal{M}}$. Moreover, the avoidance of a change of variables in the $(\beta,m)$-parametrization means that my approach can bypass an infinite-dimensional change of measure, a technical condition that may be hard to verify for complicated priors.
There are a number of extensions that would be interesting to pursue. First, it would be interesting to consider rate-adaptive priors for $m_{1}$ and $m_{2}$. I conjecture Theorem (ref) could be verified for rate adaptive priors (at least in the case of correct specification). The Mat\'{e}rn prior example in Section (ref) follows a similar argument as Theorem 3.1 of dejonge2013semiparametric, and, as a result, it is reasonable that the argument of Theorem 4.1 in their paper could also be modified to verify Theorem (ref) in the case of hierarchical spline priors, an example of a prior that can yield adaptive, rate-optimal posterior rates of contraction dejonge2012adaptive. However, the extent to which adaptive priors can be accommodated should be formally investigated. Second, it would be interesting to consider the case where $\eta(W) = W'\eta$ but the dimension of $W$ is large relative to the sample size $n$. In this case, the sampling model ((ref))--((ref)) coincides with the setup of hahn2018regularization and their simulations indicate favorable performance when using shrinkage priors for the nuisance parameters. hahn2018regularization do not provide any formal asymptotic analysis, and, for this reason, it would be interesting to examine whether an asymptotic normality result could be established in this setting. Finally, the general idea in this paper is to use knowledge of the profile likelihood severini1992profile,murphy2000profile to derive an adaptive parametrization of the partially linear model. An important extension is determining a more general class of models for which this technique could be applied.