Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
124,365 characters · 18 sections · 51 citation commands
Revisiting identification concepts in Bayesian analysis
}
\fontsize{10.95}{14pt plus.8pt minus .6pt}\selectfont
This paper investigates Bayesian analysis in models that lack identification. We first revisit theoretical concepts related to identification. We highlight how lack of identification has a different impact for statistical inference depending on whether one develops Bayesian or frequentist inference. Specifically, Bayesian analysis can be carried out without imposing any identifying restrictions. Then, we distinguish between nonidentified and partially identified models. For partially identified models we propose to construct the prior and posterior distributions of a set parameter by using the capacity functional and we introduce the concepts of prior (resp. posterior) capacity functional and prior (resp. posterior) coverage function.\\ As Lindley1971 remarked, the problem of non identification causes no real difficulty in the Bayesian approach. Indeed, if a proper prior distribution is specified, then the posterior distribution is well-defined. Kadane1974 observed that identification is a property of the likelihood function which is the same irrespective of whether it is considered from Bayesian or frequentist perspectives. It is however necessary to distinguish between three concepts of identification depending on the level of specification of the model: \emph{sampling} (or \emph{frequentist}) \emph{identification}, \emph{measurable identification} and \emph{Bayesian identification}, see \textit{e.g.} FlorensMouchartRolin1985,FlorensMouchartRolin1990 and FlorensMouchart1986. Sampling identification is defined without the introduction of a ${\Greekmath 011B}$-field associated with the parameter space. On the other hand, such a ${\Greekmath 011B}$-field is necessary for the other two concepts of identification. In addition, the notion of Bayesian identification requires the introduction of a unique joint probability measure over the sample and the parameters.\\ When a ${\Greekmath 011B}$-field associated with the parameter space is introduced, the concept of identification is related to the \emph{minimal sufficient} parameter. Namely, the observed sample brings information only on the \emph{minimal sufficient} parameter and hence, the parameter of the model is identified if it is equal to the minimal sufficient parameter. In other words, the \emph{minimal sufficient} parameter is the smallest ${\Greekmath 011B}$-field on the parameter space that makes the sampling probability measurable. So, it is the identified parameter. Conditionally on this parameter, the Bayesian experiment is completely non informative: the prior distribution of a nonidentified parameter is not revised through the information brought by the data so that the conditional posterior and conditional prior distributions (conditioned on the identified parameter) are the same.\\ Equivalently, we say that a model is nonidentified if the parametrization is redundant. It is then natural to wonder why one should introduce a ${\Greekmath 011B}$-field larger than the minimal sufficient parameter. In nonexperimental fields, redundant parametrization is usually introduced either as an early stage of model building or as a support for relevant prior information or because the parameter of interest (making \textit{e.g.} the loss function measurable) is larger than the minimal sufficient parameter (see \textbf{Example} (ref)). In experimental fields, it may be the case that the experimental design will not provide information on all the parameters of a theoretically relevant model, see FlorensMouchartRolin1990.\\ This paper makes the following contributions. (I) For nonidentified models, we show that there are situations where the introduction of a non-degenerate prior distribution can make a parameter that is unidentified in frequentist theory identified in Bayesian theory. Specifically, we demonstrate that this is true for nonparametric models with heterogeneity modeled either as a Gaussian process or as a Dirichlet process where the parameter of interest is the (hyper)parameter of the heterogeneity distribution. We show that in these models it is possible to obtain Bayesian identification since the hyperparameter can be expressed as a known function of the identified parameter. We stress that this is not a property of the prior of the nonidentified parameter, but instead it is a property of the conditional prior of the identified parameter, given the unidentified parameter. Such a result is no longer true in parametric models where a degenerate prior for the nonidentified parameter is required to get Bayesian identification, which then is completely artificial. (II) We provide the example of latent variable models to illustrate that it is preferable to conduct Bayesian inference and develop Markov Chain Monte Carlo (MCMC) algorithms for the nonidentified model instead of introducing identifying restrictions. Working with the nonidentified model grants better mixing properties of the MCMC.
(III) We propose a procedure to make Bayesian analysis for partially identified models where the identified parameter is a set, called the identified set. For these models we build up a new Bayesian nonparametric approach, based on the Dirichlet process prior, and construct prior and posterior distributions for set parameters. We propose to define these distributions in terms of prior and posterior capacity functionals. The posterior capacity functional is an appealing tool to build estimators and credible sets for the identified set and, in addition, it can be easily computed either by simulations or in closed-form. (IV) We show that, when the model contains an identified parameter and a parameter of interest that lacks identification, the latter can be identified in the marginal model where the identified parameter has been marginalized out from the likelihood function with respect to its conditional prior given the unidentified parameter. Therefore, the integrated likelihood depends on the unidentified parameter and the prior of the latter is revised by the information brought by the data in the marginal model. To complement our theoretical approach we develop many examples and simulations that show how our method can be implemented.
Our results contribute to show that the Bayesian approach is appealing in models that lack identification for several reasons. First, if the prior distribution on all the parameters is proper, Bayesian analysis of nonidentified and partially identified models is always possible since the posterior distribution always exists. Second, when the parameter in the model is multidimensional, with some components that are identified and others that are nonidentified, data can be marginally informative about the nonidentified parameter. That is, if the parameters are a priori dependent, then after we marginalise out the identified parameter from the likelihood with respect to its conditional prior given the nonidentified parameter, the integrated likelihood will depend on the nonidentified parameter. Third, the issue of non- and/or partial-identification can be reduced (or even eliminated) by introducing an informative prior. Lastly, even when the model is nonidentified or partially identified, Bayesian procedures have computational advantages over frequentist ones. In particular, MCMC algorithms have better mixing properties if they are specified for the nonidentified model instead of imposing identifying restrictions.\\ The paper is organized as follows. In section (ref) we discuss the three concepts of identification given above and the concept of partial identification. Section (ref) studies models with heterogeneity and models with latent variables. Bayesian analysis of partially identified models is developed in section (ref). Finally, in section (ref) we discuss identification by marginalisation for both unidentified and partially identified models. Section (ref) concludes. Examples and all the proofs are in Appendix E in the Supplement. In the paper we abbreviate “almost surely” by “a.s.”, $P$ will denote the data distribution and $p$ the associated Lebesgue density. For an event $A$, $\mathbbm{1}\{A\}$ denotes the indicator function which takes the value $1$ if the event $A$ is satisfied and $0$ otherwise. \paragraph{Literature Review.} Our paper is related to two strands of the Bayesian literature, the one focusing on nonidentified models and the Bayesian literature on partially identified models. We provide here a concise review of the previous contributions that are the most relevant for our paper.\\ Initial discussions of nonidentified models which lay the foundations of nonidentification in a Bayesian experiment in a measure-theoretic framework can be found in Lindley1971, Kadane1974 and Picci1977, among others. FlorensMouchartRolin1985,FlorensMouchartRolin1990 and FlorensMouchart1986 resume these works and provide further developments. In particular, they provide a rigorous discussion on the difference among sampling, measurable and Bayesian identification. Compared to these contributions, in section (ref) we provide a unified framework that gather together the concepts and results related to nonidentification that are the most relevant for applied Bayesian analysis in econometrics. We adopt a measure-theoretic framework and present the results of identification in terms of ${\Greekmath 011B}$-fields of sets of the parameter space. An aspect that we do not consider in our paper is the prior elicitation process for nonidentified parameters which is investigated for instance in SanMartinGonzales2010.\\ As in Poirier1998, we emphasise the difference between marginal and conditional uninformativeness of the data for Bayesian analysis of nonidentified models. The main contribution of Poirier1998 consists in analysing the diverse effect of nonidentification for Bayesian analysis in the two following situations: the case with proper priors, and the case with improper priors. Our paper does not emphasizes this difference between proper and improper priors. Instead, one of our main contributions is to demonstrate the different role played by the prior in nonparametric and parametric models that are nonidentified. The question of nonidentification in nonparametric models has not been explicitly considered in the past Bayesian literature to the best of our knowledge. For these models we prove that Bayesian identification, i.e. through the prior, arises in a non-artificial way in some cases.\\ Gustafson2005 analyses nonidentifiability by taking a nonconventional approach. Instead of contracting the model, which consists in reducing the redundant parametrisation to get identification, he proposes to expand the model to a supermodel that can at best yield identification by the prior. This is an alternative approach to ours.\\ The second strand of literature that relates to our paper is the Bayesian literature on partial identification. It includes relatively few contributions in comparison to the vast frequentist literature on partially identified models. For excellent overviews of frequentist and Bayesian inference for partially identified econometric models see BontempsMagnac2017, MOLINARI2020 and references therein. Most of the previous Bayesian literature is interested in constructing Bayesian procedures that provide asymptotically valid frequentist inferences for partially identified models, see for instance LiaoJiang2010, Gustafson2010,Gustafson2011, MS12, NoretsTang2013, KlineTamer2015, chen2017monte, LiaoSimoni2019 and GiacominiKitagawa2018. They have proposed Bayesian or quasi-Bayesian approaches for constructing (asymptotically) valid frequentist confidence sets of the identified set. Their methods are mostly based either on the limited information likelihood of CH03 or on the support function. Compared to this literature, we take a different approach that is based on Dirichlet process prior placed on the identified reduced-form parameter. We construct the prior and posterior probabilities for the identified set as prior and posterior capacity functionals, which is new in the literature. In addition, we are not interested in frequentist asymptotic properties of our procedure and conduct a fully Bayesian analysis.\\ Marginal analysis in partially identified models, which we consider in section 5, has been considered in LiaoJiang2010 but in a way different from ours. They marginalise out the slackness parameters in the partially identified models characterized by moment inequalities while we marginalise out the identified (reduced-form) parameter.
In this section we recall basic definitions in the non-Bayesian framework, called in the paper sampling theory approach. Let us consider a statistical model defined by a sampling space $X$ provided with a ${\Greekmath 011B}$-field $\mathcal{X}$ and by a collection of probabilities on this space. The collection of sampling probability measures defined on $(X,\mathcal{X})$ is denoted by $(P^{{\Greekmath 0112}})_{{\Greekmath 0112}\in\Theta}$ and is in general indexed by a parameter ${\Greekmath 0112} \in \Theta$, which can be finite dimensional or a functional parameter. Hence, the sampling statistical model $\mathcal{E}_{s}$ is defined as $\mathcal{E}_{s}:=\{\Theta,(X,\mathcal{X}),(P^{{\Greekmath 0112}})_{{\Greekmath 0112}\in\Theta}\}$.\\ In a small sample approach we observe a finite number of realizations of a random variable taking values in a measurable space -- for example an iid sample $x:=(x_{1},\ldots,x_{n})\in (X,\mathcal{X})$ -- where the sample size is not made explicit and is kept fixed. In an asymptotic approach, $\mathcal{X}$ is provided with a filtration $(\mathcal{X}_{n})_{n\geq 1}$, where $\mathcal{X}_{n}$ represents the information contained in a sample of size $n$. We denote by $\mathcal{X}_{\infty} := \bigvee_{n\geq 1}\mathcal{X}_{n}$ the ${\Greekmath 011B}$-field generated by $\bigcup_{n\geq 1}\mathcal{X}_{n}$ and write $\mathcal{X}_{n}\rightarrow \mathcal{X}_{\infty}$ as $n\rightarrow \infty$.\\ Two parameters ${\Greekmath 0112}_{1}$ and ${\Greekmath 0112}_{2}$ are said to be observationally equivalent, in symbol ${\Greekmath 0112}_{1}\sim{\Greekmath 0112}_{2}$, if $P^{{\Greekmath 0112}_{1}} = P^{{\Greekmath 0112}_{2}}$. This relation defines an equivalence relation and we denote by $\widetilde{\Theta}$ the quotient space $\Theta/\sim$, i.e. the elements $\widetilde{\Greekmath 0112}$ of $\widetilde{\Theta}$ are the equivalence classes over $\Theta$ by $\sim$. Notice that $\widetilde{{\Greekmath 0112}}$ is a set and the elements of $\widetilde{\Theta}$ are sets of parameters. We now define the concept of sampling identification.
This definition is equivalent to say that ${\Greekmath 0112}_{1}\sim{\Greekmath 0112}_{2}$ implies ${\Greekmath 0112}_{1} = {\Greekmath 0112}_{2}$ or to say that any equivalence class is reduced to a singleton. In turn, this is equivalent to say that the sampling model $\mathcal{E}_{s}$ is identified if the mapping ${\Greekmath 0112}\rightarrow P^{{\Greekmath 0112}}$ is injective.\\ For any sampling statistical model $\mathcal{E}_{s}$ there exists a canonical identified model $\widetilde{\mathcal{E}}_{s}$ defined as the sampling statistical model with parameter space $\widetilde{\Theta}:\widetilde{\mathcal{E}}_{s}=\{\widetilde{\Theta},(X,\mathcal{X}),(P^{\widetilde{{\Greekmath 0112}}})_{\widetilde{{\Greekmath 0112}}\in\widetilde{\Theta}}\}$ where $P^{\widetilde{{\Greekmath 0112}}} = P^{{\Greekmath 0112}}$ for any ${\Greekmath 0112}\sim\widetilde{{\Greekmath 0112}}$. Equivalently, $\widetilde{\mathcal{E}}_{s}$ is the set identified statistical model associated with $\mathcal{E}_{s}$. Thus, one may construct an identified sampling model by selecting a single element in each equivalence class. Let us call section a function ${\Greekmath 011B}:\widetilde{\Theta}\rightarrow\Theta$ such that ${\Greekmath 011B}(\widetilde{{\Greekmath 0112}})\in \widetilde{{\Greekmath 0112}}$ and ${\Greekmath 011B}(\widetilde{{\Greekmath 0112}})\in\Theta$. By using such a function one may define the identified sampling model
where $\Theta_{{\Greekmath 011B}}$ is the image of $\widetilde{\Theta}$ by ${\Greekmath 011B}$ and $P^{{\Greekmath 0112}} = P^{{\Greekmath 011B}(\widetilde{{\Greekmath 0112}})}$. In this context it is natural to look for continuous sections ${\Greekmath 011B}(\cdot)$ or to bicontinuous bijections between $\widetilde{\Theta}$ and $\Theta_{{\Greekmath 011B}}$. In general, it is not possible to devise constructive rules for selecting elements in $\Theta$ for each $\widetilde{{\Greekmath 0112}}\in\widetilde{\Theta}$. The existence of $\Theta_{{\Greekmath 011B}}$ cannot be proved by using axioms from set theory if we do not have some structure on the parameter space but it must be asserted as an additional axiom called axiom of choice, see e.g. KolmogorovFomin1975.\\ In the sampling theory statistics, a topological structure is needed in the parameter space in order to define statistical decision rules based on convergence, risk or loss function. A canonical topological structure is defined on $\Theta$ and then may be carried on $\widetilde{\Theta}$. Let ${\Greekmath 011A}:\Theta\rightarrow \widetilde{\Theta}$ be the canonical application ${\Greekmath 0112} \rightarrow {\Greekmath 011A}({\Greekmath 0112}) = \widetilde{{\Greekmath 0112}}$. The natural topological structure is the smallest one for which ${\Greekmath 011A}$ is continuous, i.e. ${\Greekmath 011A}^{-1}(\widetilde{O})$ is open in $\Theta$ whenever $\widetilde{O} \subset \widetilde{\Theta}$ is open. We will not detail the topological aspects of set identification, which is beyond the scope of this paper, and refer to Husmoller1994 and appendix to chapter 3 in DellacherieMeyer1975.
Let us define a measurable statistical model $\mathcal{E}_{m}$ as $\mathcal{E}_{m} := \{(\Theta,\mathcal{A}), (X,\mathcal{X}), (P^{{\Greekmath 0112}})_{{\Greekmath 0112}\in\Theta}\}$, where we use the notation previously introduced and in addition we define $\mathcal{A}$ to be a ${\Greekmath 011B}$-field on $\Theta$ such that $P^{{\Greekmath 0112}}$ is a transition probability. Recall that given two measurable spaces $(\Theta,\mathcal{A})$ and $(X,\mathcal{X})$, the mapping $P^{(\cdot)}(\cdot):\Theta\times\mathcal{X}\rightarrow [0,1]$ is a transition probability if: (i) $\forall {\Greekmath 0112}\in\Theta$, $P^{{\Greekmath 0112}}(\cdot)$ is a probability measure on $(X,\mathcal{X})$, and (ii) $\forall E\in\mathcal{X}$, $P^{(\cdot)}(E)$ is a measurable function on $(\Theta,\mathcal{A})$. The introduction of the ${\Greekmath 011B}$-field $\mathcal{A}$ of subsets of $\Theta$ is necessary in order to introduce a joint probability measure on the product space $\Theta\times X$. This probability will be introduced in section (ref). In this sense, the introduction of a measurable statistical model is a preliminary step for the construction of a Bayesian model.\\ A ${\Greekmath 011B}$-field represents an information structure and the parameter of interest is naturally introduced as a ${\Greekmath 011B}$-field $\mathcal{A}$ that makes both the transition probability $P^{(\cdot)}(E)$, $\forall E\in\mathcal{X}$, and the loss function of the underlying decision model, $\mathcal{A}$-measurable. We now introduce the concept of sufficient ${\Greekmath 011B}$-field which is essential in order to discuss identification.
The structure of model $\mathcal{E}_{m}$ allows for the possibility of introducing a given distribution on $\Theta$. Here, we discuss identification in model $\mathcal{E}_{m}$ without the specification of a (prior) distribution on $(\Theta,\mathcal{A})$ which will be discussed in the next section. The following proposition introduces the concept of identification in model $\mathcal{E}_{m}$, called measurable identification.
The notion of sampling and measurable identification are identical if $\Theta$ is a measurable subset of $\mathbb{R}^{k}$ provided with the Borelian ${\Greekmath 011B}$-field and if $X$ is included in $\mathbb{R}^{n}$ and also provided with the Borelian ${\Greekmath 011B}$-field. This equivalence is true more generally and requires that $\Theta$ “looks like” a Borelian of $\mathbb{R}$ and that the ${\Greekmath 011B}$-field on $X$ is separable or equivalently generated by a countable family of subsets. Such a link between sampling and measurable identification is established in the next theorem where we assume that $(\Theta,\mathcal{A})$ is a Souslin space. We recall that a measurable space $(\Theta,\mathcal{A})$ is a Souslin space if there exists an analytic set $B\subset \mathbb{R}$ and a bimeasurable bijection between $(\Theta,\mathcal{A})$ and $(B,B\cap\mathcal{B})$, where $B\cap\mathcal{B}$ denotes the restriction to $B$ of the Borel ${\Greekmath 011B}$-fields $\mathcal{B}$ of $\mathbb{R}$. An analytic set on $\mathbb{R}$ is the projection on $\mathbb{R}$ of a Borelian set in $\mathbb{R}^{2}$. In particular, all Borelian sets are analytic. The property that $(\Theta,\mathcal{A})$ is a Souslin space is clearly true for finite dimensional parameter spaces or more generally for Polish spaces, which include the $L^{2}$ spaces on a real space. The majority of the functional spaces usually considered in statistical and econometric applications are Polish, so the requirement that $(\Theta,\mathcal{A})$ is a Souslin space is almost always satisfied.
The first point of the theorem provides an additional caracterisation of measurable identification. The second part establishes that sampling and measurable identification are equivalent when $\mathcal{X}$ is separable.
A Bayesian model consists of a measurable statistical model and a measure on $(\Theta,\mathcal{A})$ which can be either proper (if it is a probability measure) or improper, called prior distribution and denoted by ${\Greekmath 0116}$. For simplicity we will consider a probability measure. Then, ${\Greekmath 0116}$ and $P^{{\Greekmath 0112}}$ generate a unique measure on $(\Theta\times X,\mathcal{A}\otimes\mathcal{X})$ denoted by $\Pi$:
or, equivalently $\Pi = {\Greekmath 0116}\otimes P^{{\Greekmath 0112}} = P\otimes {\Greekmath 0116}^{x}$, where $P$ is the marginal probability measure on $(X,\mathcal{X})$ called the predictive distribution and ${\Greekmath 0116}^{x}$ is the posterior distribution. We assume that there exists a regular version of the conditional probability on $\Theta$ given $X$, that is, ${\Greekmath 0116}^{x}(\cdot)$ is a transition probability (see definition in section (ref)). The Bayesian model is then defined by the following probability space
In the following, we may use the notation ${\Greekmath 0116}(\cdot|x)$ instead of ${\Greekmath 0116}^{x}(\cdot)$ when it is more appropriate. Moreover, by abuse of notation we use ${\Greekmath 0116}({\Greekmath 0112}|x)$ (resp. ${\Greekmath 0116}({\Greekmath 0112})$) to denote both the posterior (resp. the prior) distribution and its Lebesgue density function. The sampling (resp. predictive) density function with respect to the Lebesgue measure is denoted by $p(x|{\Greekmath 0112})$ (resp. $p(x)$). Before defining the concept of Bayesian identification, we recall the definition of sufficiency and minimal sufficiency in the Bayesian model.
The conditional independence of the previous definition has two equivalent characterizations: (i) for every positive function $t : \mathcal{X}\rightarrow\mathbb{R}_{+}$, $\mathbf{E}(t(x)|\mathcal{A}) = \mathbf{E}(t(x)|\mathcal{B})$, $\Pi-a.s.$ provided that the conditional expectations exist; (ii) for every positive function $a : \mathcal{A} \rightarrow \mathbb{R}_{+}$, $\mathbf{E}(a({\Greekmath 0112})|\mathcal{X}\vee\mathcal{B})=\mathbf{E}(a({\Greekmath 0112})|\mathcal{B})$, $\Pi-a.s.$ The first characterization (i) weakens the concept of sufficient ${\Greekmath 011B}$-field because the property is only required almost surely with respect to the prior probability. If $\mathcal{A}$ is generated by a function $a(\cdot)$ and $\mathcal{B}$ by a function $b(\cdot)$ then (i) means that the likelihood functions $p(x|a)$ and $p(x|b)$ are a.s. equal. The second characterization (ii) says that the conditional prior and posterior distributions are a.s. equal given a sufficient ${\Greekmath 011B}$-field. Equivalently, \textit{(ii)} means that the posterior distribution ${\Greekmath 0116}(a|x,b)$ of $a$ is \textit{a.s.} equal to the prior ${\Greekmath 0116}(a|b)$.\\ It may be proven that there exists a \emph{minimal Bayesian sufficient ${\Greekmath 011B}$-field} in $\mathcal{A}$, denoted by $\mathcal{A}_{*}^{{\Greekmath 0116}}$ and defined as follows.
The minimal Bayesian sufficient ${\Greekmath 011B}$-field $\mathcal{A}_{*}^{{\Greekmath 0116}}$ is a ${\Greekmath 011B}$-field on the parameter space and depends on the prior ${\Greekmath 0116}$. An equivalent definition of $\mathcal{A}_{*}^{{\Greekmath 0116}}$ is that $\mathcal{A}_{*}^{{\Greekmath 0116}}$ is the smallest ${\Greekmath 011B}$-field which makes the sampling probabilities $P^{(\cdot)}(E)$ measurable, $\forall E\in\mathcal{X}$, completed by the null sets of $\mathcal{A}$ with respect to ${\Greekmath 0116}$. It may be easily verified that $\overline{\mathcal{A}_{*}}\cap\mathcal{A} = \mathcal{AX} \equiv \mathcal{A}_{*}^{{\Greekmath 0116}}$ where $\overline{\mathcal{A}_{*}}$ is the ${\Greekmath 011B}$-field generated by $\mathcal{A}_{*}$ and all the null sets of $\mathcal{A}\vee\mathcal{X}$, and $\overline{\mathcal{A}_{*}}\cap\mathcal{A}$ is the ${\Greekmath 011B}$-field generated by $\mathcal{A}_{*}$ and all the null sets of $\mathcal{A}$ with respect to ${\Greekmath 0116}$.\\ We are now ready to give the definition of Bayesian identification.
More generally, the concept of identification may be introduced for a sub-${\Greekmath 011B}$-field (i.e. a parameter) $\mathcal{B}\subset\mathcal{A}$:
We recall that $\mathcal{BS}$ denotes the projection of $\mathcal{S}$ on $\mathcal{B}$, that is, $$\mathcal{BS}:={\Greekmath 011B}(\{\mathbf{E}[s|\mathcal{B}]; s \textrm{ belongs to the set of positive random variables defined on }(X,\mathcal{X})\}),$$ where ${\Greekmath 011B}(\{G\})$ denotes the ${\Greekmath 011B}$-field generated by the set $G$.\\ Definition (ref) means that in a Bayesian identified model, $\mathcal{A}_{*}^{{\Greekmath 0116}}$ is almost surely equal to $\mathcal{A}$, that is, $\overline{\mathcal{A}_{*}}\cap\mathcal{A} = \mathcal{A}$ or: $\forall A\in\mathcal{A}$, $\exists B\in\mathcal{A}_{*}$ such that ${\Greekmath 0116}(A\bigtriangleup B) = 0$, where $\bigtriangleup$ denotes the symmetric difference. This property shows that Bayesian identification is an almost sure measurable identification with respect to the prior. We also have the following characterization of Bayesian identification, see FlorensMouchartRolin1990.
It is clear from the theorem that sampling identification (and equivalently measurable identification) implies Bayesian identification but the reverse is not true. For instance, a degenerate prior that puts mass one on specific values of the parameter space makes Bayesian identified a model that is not sampling identified as the following trivial example shows. Suppose that we observe a random variable $x$ from the sampling model $x|{\Greekmath 0112}_1,{\Greekmath 0112}_2 \sim \mathcal{N}({\Greekmath 0112}_1 + {\Greekmath 0112}_2,1)$, ${\Greekmath 0112} = ({\Greekmath 0112}_1,{\Greekmath 0112}_2)' \in \Theta = \mathbb{R}^2$, and that we endow the parameter ${\Greekmath 0112}$ with the degenerate prior ${\Greekmath 0112}\sim\mathcal{N}(({\Greekmath 0112}_{10},{\Greekmath 0112}_{20})',{\Greekmath 0113} {\Greekmath 0113}')$, where ${\Greekmath 0113} = (1,1)'$. This prior gives probability zero to all the value in $\mathbb{R}^2$ but the values on the line ${\Greekmath 0112}_2 = {\Greekmath 0112}_1$. Therefore, $\Theta_0 = \mathbb{R}^2\setminus \{{\Greekmath 0112}\in\mathbb{R}^2; {\Greekmath 0112}_2 = {\Greekmath 0112}_1\}$ and the sampling distribution is injective on $\Theta - \Theta_0 = \{{\Greekmath 0112}\in\mathbb{R}^2; {\Greekmath 0112}_2 = {\Greekmath 0112}_1\}$. This model is not sampling identified nor measurable identified. However, this specification of the prior makes the model Bayesian identified.\\
Before concluding this section we point out the following relationships existing between the three concepts of identification that we have seen. (i) Measurable identification implies Bayesian identification for any ${\Greekmath 0116}$. (ii) Measurable identification implies sampling identification if and only if $\mathcal{A}$ is separating, that is, all the atoms of $\mathcal{A}$ are singletons.
In this section we present the important connection that exists between the concepts of identification and of exact estimability in the Bayesian models. We start by defining exact estimability.
The ${\Greekmath 011B}$-field $\overline{\mathcal{X}}$ is a ${\Greekmath 011B}$-field on the product space $\Theta \times X$ and not a ${\Greekmath 011B}$-field on the sampling space $X$. So, it is possible that $\mathcal{B}\subset \overline{\mathcal{X}}$ but that $\mathcal{B} \not\subset \mathcal{X}$. For instance, consider $x_i|{\Greekmath 0112} \sim^{iid}\mathcal{N}({\Greekmath 0112},1)$ and $\mathcal{B} = {\Greekmath 011B}(\{{\Greekmath 0112}\})$. Consider a prior for ${\Greekmath 0112}$ that puts all its mass on the sample mean $\bar{x}$, then we clearly have that $\mathcal{B}\subset \overline{\mathcal{X}}$ but $\mathcal{B} \not\subset \mathcal{X}$. This situation is artificial in small sample, but it describes well what it happens asymptotically if $\bar{x}$ is a consistent estimator since $\bar{x}\rightarrow {\Greekmath 0112}$, $\Pi$-a.s. without requiring a degenerate prior, see also Definition (ref) and the discussion below it.\\ The inclusion $\mathcal{B}\subset \overline{\mathcal{X}}$ means that for any positive random variable $a$ defined on $(\Theta,\mathcal{B})$, the posterior expectation $\mathbf{E}(a|\mathcal{X}) = a$ $\Pi$-a.s., provided that the conditional expectation exists. This means that a posteriori -- that is, after observing the sample -- we know $a$ $\Pi$-a.s. Thus, any sub-${\Greekmath 011B}$-field of $\mathcal{A}\cap\overline{\mathcal{X}}$ is exactly estimable. The next proposition states a link between exact estimability and Bayesian identification.
The result in the proposition, together with the following result, allows to show that in an i.i.d. experiment, the minimal sufficient ${\Greekmath 011B}$-field is Bayesian identified.
Moreover, if $\mathcal{AX}$ is exactly estimable, then the following theorem shows that the reverse implication of Proposition (ref) holds.
The concept of exact estimability has a particular interest in asymptotic models. Let $x^{(n)}:=(x_1,\ldots,x_n)$ denote a sequence of $n$ observations and $x^{(\infty)}$ denote the sample of infinite size. The ${\Greekmath 011B}$-field generated by $x^{(n)}$ (resp. $x^{(\infty)}$) is denoted by $\mathcal{X}_{n}$ (resp. $\mathcal{X}_{\infty}$). For any function $a(\cdot)$ of the parameter ${\Greekmath 0112}$, the martingale convergence theorem implies that $\mathbf{E}(a|\mathcal{X}_{n})\rightarrow \mathbf{E}(a|\mathcal{X}_{\infty})$ a.s. This convergence is a.s. with respect to the joint distribution $\Pi$ and also with respect to the predictive probability $P$. This is one concept of Bayesian consistency for which the convergence must be taken with respect to the joint probability distribution $\Pi$. Another concept of Bayesian consistency is convergence with respect to the sampling measure $P^{{\Greekmath 0112}}$, that is, $\mathbf{E}(a|\mathcal{X}_{n})\rightarrow a({\Greekmath 0112})$, $P^{{\Greekmath 0112}}$- a.s. This convergence does not follow from the previous argument and there are cases in which it is not be verified, see e.g. DiaconisFreedman1986 and FlorensSimoni2010.\\ From the above concept of Bayesian consistency, if $\mathcal{B}\subset \overline{\mathcal{X}}_{\infty}$, we have $\mathbf{E}(a|\mathcal{X}_{n})\rightarrow a$ $\Pi$-a.s. for any $\mathcal{B}$-measurable positive function $a$ of ${\Greekmath 0112}$, that is, the posterior mean of $a$ converges a.s. to the true model. This is a consequence of the asymptotically exact estimability which we now define.
This definition and the martingale convergence theorem imply that if $\mathcal{B}\subset \mathcal{A}$ is asymptotically exactly estimable then, for every integrable real random variable $a$ defined on $\mathcal{B}$, $\mathbf{E}(a|\mathcal{X}_{n})\rightarrow a$ $\Pi$-a.s.\\ In a Souslin space, asymptotic exact estimability means that the posterior distribution is asymptotically a Dirac measure on a function of the sample, see FlorensMouchartRolin1990. The intuition is that, if $\mathbf{E}(a|\mathcal{X}_{n})\rightarrow a$ $\Pi$-a.s., then also $\mathbf{E}(a^{2}|\mathcal{X}_{n})\rightarrow a^{2}$ $\Pi$-a.s. which implies that the posterior variance of $a$ converges to $0$. This clearly explains that the posterior distribution concentrates around the $\Pi$-a.s. limit of its mean. More generally, an integrable real random variable $a(\cdot)$ defined on $\mathcal{B}$ is asymptotically exactly estimable if there exists a strongly consistent sequence of estimators of $a({\Greekmath 0112})$, i.e. if there exists a set of random variables $(t_{n})_{n\in\mathbb{N}}$, each defined on $\mathcal{X}$, such that $t_{n}\rightarrow a({\Greekmath 0112})$, $\Pi$-a.s. In sampling theory, a necessary condition for the existence of such a sequence is the identifiability of the parameter.
Between the two extreme situations of identification and non-identification there are models that are partially identified. Partial identification arises when the combination of available data and assumptions that are plausible for the model only allows to place the population parameter of interest ${\Greekmath 010D}$ within a proper subset $\Gamma_{I}$ of the parameter space $\Gamma$ called identified set. If one has very precise prior information, identification can be restored by specifying a prior distribution degenerated at some point in $\Gamma_I$ as we explain in section (ref) - Model 5. However, this strategy can be unsuitable if one is concerned with robustness of the Bayesian procedure. Therefore, a more appealing approach consists in working directly with the parameter $\Gamma_I$ which is a set and: first, define the prior and posterior distribution of $\Gamma_I$ (see section (ref)); second, define a prior for ${\Greekmath 010D}$ inside this set (see section (ref)).\\ In a partially identified model, usually, there is an identified parameter, say ${\Greekmath 0112}$, and a partially identified parameter which is the parameter of interest, denoted by ${\Greekmath 010D}$. The latter enters the model as a supplementary parameter and is linked to ${\Greekmath 0112}$ through a relation of the form $A({\Greekmath 0112},{\Greekmath 010D})\in A_{0} \subset \Phi$ where $\Phi$ is a suitably defined space and $A$ is a given function of $({\Greekmath 0112}, {\Greekmath 010D})$. Such relation characterizes the identified set as $\Gamma_{I}:=\{{\Greekmath 010D};\:A({\Greekmath 0112},{\Greekmath 010D})\in A_{0} \in \Phi\}$ which may not be a singleton. The function $({\Greekmath 0112}, {\Greekmath 010D})\mapsto A({\Greekmath 0112}, {\Greekmath 010D})$ may be a parametric likelihood function, flat over a region of ${\Greekmath 010D}$s, and $A_0$ is its maximum value. Alternatively, $\Gamma_{I}$ may be a set of moment restrictions as in Example (ref) below. See chen2017monte for more details on Bayesian analysis of these examples.
To develop Bayesian analysis it is important to understand that the specification of the prior distribution differs in nonidentified models and in partially identified models. In partially identified models, the prior distribution is naturally decomposed in the marginal prior on ${\Greekmath 0112}$ and the conditional prior on ${\Greekmath 010D}$ given ${\Greekmath 0112}$ which incorporates the link between ${\Greekmath 0112}$ and ${\Greekmath 010D}$: $x|{\Greekmath 0112},{\Greekmath 010D} \sim P^{{\Greekmath 0112}}$, ${\Greekmath 0112}\sim {\Greekmath 0116}_{{\Greekmath 0112}}$ and ${\Greekmath 010D}|{\Greekmath 0112}\sim {\Greekmath 0116}_{{\Greekmath 010D}}(\cdot|{\Greekmath 0112})$. \\ On the other hand, models with nonidentified parameters arise when ${\Greekmath 0112}$ is partitioned in ${\Greekmath 0112} = ({\Greekmath 010C},{\Greekmath 010D})$ where ${\Greekmath 010C}$ is identified and ${\Greekmath 010D}$ is not. For instance, ${\Greekmath 010D}$ can be the parameter of the distribution of the heterogeneity parameter ${\Greekmath 010C}$. Therefore, the prior is naturally decomposed in the prior on ${\Greekmath 010C}$ conditional on ${\Greekmath 010D}$, ${\Greekmath 0116}_{{\Greekmath 010C}}(\cdot|{\Greekmath 010D})$, and the prior ${\Greekmath 0116}_{{\Greekmath 010D}}$ on ${\Greekmath 010D}$. This decomposition of the prior distribution is the reverse of the decomposition made for partially identified models.
This section is made of two subsections that describe two different options one has to make Bayesian inference for nonidentified models. We first discuss and construct models that may satisfy Bayesian identification even if they are not sampling identified nor measurable identified. This happens when the function ${\Greekmath 0112} \mapsto P^{{\Greekmath 0112}}$ is not injective but the prior assigns a mass zero on the nonidentified component of the model, as it has been shown in Theorem (ref). A second option, discussed in section (ref), consists in working with the nonidentified model without introducing any identifying priors or identifying assumptions. The beauty of the Bayesian approach is that, as long as the prior is proper, one can make Bayesian inference because the posterior distribution is well defined and then MCMC algorithms can be designed to simulate from the posterior even in a nonidentified model.
This section considers models where the parameter of interest ${\Greekmath 010D}$ is a parameter of the distribution of a heterogeneity parameter ${\Greekmath 010C}$. More precisely, the parameter space $(\Theta,\mathcal{A})$ contains two subparameter spaces $B$ and $\Gamma$ with the associated sub-${\Greekmath 011B}$-fields $\mathcal{B}$ and $\mathcal{G}$, where $\mathcal{B}$ is the ${\Greekmath 0116}$-a.s. minimal sufficient ${\Greekmath 011B}$-field in the parameter space. The ${\Greekmath 011B}$-field $\mathcal{G}$ is not identified in the sense that $\mathcal{G}\nsubseteq\mathcal{B}$. The sampling distribution only depends on $\mathcal{B}$ and the Lebesgue data density, if it exists, satisfies $p(x|{\Greekmath 010C},{\Greekmath 010D}) = p(x|{\Greekmath 010C})$ a.s., where ${\Greekmath 010C}\in(B,\mathcal{B})$ and ${\Greekmath 010D}\in (\Gamma,\mathcal{G})$. The posterior satisfies ${\Greekmath 0116}({\Greekmath 010D}|x,{\Greekmath 010C}) = {\Greekmath 0116}({\Greekmath 010D}|{\Greekmath 010C})$ a.s. An interesting situation, known as local identification, arises when the ${\Greekmath 011B}$-field $\mathcal{G}$ is not identified only for particular values taken on by the parameter $\mathcal{B}$, that is, $p(x|{\Greekmath 010C}=\bar{{\Greekmath 010C}},{\Greekmath 010D}) = p(x|{\Greekmath 010C}=\bar{{\Greekmath 010C}})$ for some value $\bar{{\Greekmath 010C}}$, see e.g. DREZE1983, KleibergenVanDijk1998, HoogerheideEtAL2007 and KleibergenMavroeidis2011. We do not analyze this situation in this paper.\\ The prior is naturally specified as the product of the conditional distribution of $\mathcal{B}$ given $\mathcal{G}$ and the marginal on $\mathcal{G}$. In this case, $\mathcal{G}$ is interpreted as a parameter of the prior on $\mathcal{B}$, usually called an hyperparameter, see e.g. Berger1985, or latent variable. Even if this model is not sampling (nor measurable) identified, it may satisfy Bayesian identification if the prior is conveniently specified as it has been shown in Theorem (ref). This identification by the prior is fully artificial in the finite-dimensional parameter case where a degenerate prior is required to get identification and we do not consider this case. Instead, we consider nonparametric and asymptotic settings where identification can be obtained without a degenerate prior. The main argument that we use to get Bayesian identification in this way is the following.
Notice that the a.s. in the proposition are related to the prior and that $\mathcal{B}$ is the minimal sufficient ${\Greekmath 011B}$-field. The proposition states that any sub-${\Greekmath 011B}$-field $\mathcal{G}\subset \mathcal{A}$ is identified if and only if it is almost surely included in $\mathcal{B}$. Intuitively, the result of the proposition holds if the conditional prior distribution on $\mathcal{G}$ given $\mathcal{B}$ is degenerate into a Dirac measure on a $\mathcal{B}$-measurable function. No requirement are made on the marginal prior on $\mathcal{G}$. We now discuss four models where Bayesian identification can be obtained without a degenerate marginal prior on ${\Greekmath 010D}$ and a fifth example where the model is partially identified, and contrarily to the previous cases, Bayesian identification can be only obtained artificially with a degenerate marginal prior. \paragraph{Model 1: unobserved heterogeneity (or incidental parameter).} The incidental parameter problem typically arises with panel data models when a regression model includes an agent specific intercept ${\Greekmath 010C}_{i}$ which is a latent variable. The general conditional model can be simplified as
where ${\Greekmath 011B}^{2}$ is known and ${\Greekmath 010D}$ is a hyperparameter characterizing the distribution of ${\Greekmath 010C}_{i}$. If ${\Greekmath 010C}_{i}$ is i.i.d. across individuals then ${\Greekmath 010D}$ is the common parameter in the population. Let ${\Greekmath 010C}:=({\Greekmath 010C}_{1},{\Greekmath 010C}_{2},\ldots)$. The ${\Greekmath 0116}$-a.s. minimal sufficient ${\Greekmath 011B}$-field $\mathcal{B}$ is generated by ${\Greekmath 010C}$, that is, ${\Greekmath 010C}$ is the identified parameter. Instead of estimating the distribution of the heterogeneity parameter ${\Greekmath 010C}_{i}$ one could be satisfied with the estimation of the common parameter ${\Greekmath 010D}$. Unfortunately, ${\Greekmath 010D}$ is not measurably identified even if it may be Bayesian identified. To illustrate this, let us specify a prior probability for ${\Greekmath 010C}$ and ${\Greekmath 010D}$ as follows:
where ${\Greekmath 011B}_{0}^{2}$ is known and ${\Greekmath 0116}_{{\Greekmath 010D}}$ is a probability measure. By the strong law of large numbers, ${\Greekmath 010D} = \lim_{n\rightarrow\infty}\frac{1}{n}\sum_{i=1}^{n}{\Greekmath 010C}_{i}$ a.s. with respect to both the conditional prior distribution of ${\Greekmath 010C}$ given ${\Greekmath 010D}$ and the joint prior distribution of $({\Greekmath 010C},{\Greekmath 010D})$. Therefore, the conditional distribution of ${\Greekmath 010D}|\{{\Greekmath 010C}_i\}_{i=1}^{\infty}$ puts all its mass on $\lim_{n\rightarrow\infty}\frac{1}{n}\sum_{i=1}^{n}{\Greekmath 010C}_{i}$ and the model is fully Bayesian identified. We stress that this property does not depend on the marginal prior on ${\Greekmath 010D}$, which does not have to be degenerate on some value, but crucially depends on the prior of ${\Greekmath 010C}$ given ${\Greekmath 010D}$. \paragraph{Model 2: Gaussian process and hyperparameter in the mean.} Let $L^{2}[0,1]$ denote the space of square integrable functions defined on $[0,1]$ with respect to the uniform distribution, endowed with the $L^{2}$-inner product $\langle\cdot,\cdot\rangle$ and the induced norm $\|\cdot\|$. We consider a sample space $X=L^{2}[0,1]$. The sample is made of a single observation which is a trajectory $x$ from a Gaussian process in $X$: $x|{\Greekmath 010C},{\Greekmath 010D} \sim \mathcal{GP}({\Greekmath 010C},\Sigma)$ where ${\Greekmath 010C}\in X$ and $\Sigma : X\rightarrow X$ is a covariance operator which is bounded, linear, positive definite, self-adjoint and trace-class. It follows that $\mathbf{E}||x||^{2}<\infty$. The parameter ${\Greekmath 010C}$ is measurably identified but the parameter ${\Greekmath 010D}$ is not and it may be interpreted as a low dimensional parameter that controls the distribution of ${\Greekmath 010C}$.\\ Suppose that ${\Greekmath 010C}$ and ${\Greekmath 010D}$ are two random variables with values in $(B,\mathcal{B})$ and $(\mathbb{R}_{+},\mathcal{R})$, respectively, with $B = L^{2}[0,1]$ and $\mathcal{B}$ and $\mathcal{R}$ are the associated Borel ${\Greekmath 011B}$-field. The joint prior distribution on $(B\times\mathbb{R}_{+},\mathcal{B}\otimes\mathcal{R})$ is specified as:
where ${\Greekmath 010C}_{0}$, ${\Greekmath 010D}_{0}$, ${\Greekmath 011B}_{0}^{2}$ and $\Omega$ are known hyperparameters. The covariance operator $\Omega$ is bounded, linear, positive definite, self-adjoint and trace-class and its eigensystem is denoted by $({\Greekmath 0115}_{j},{\Greekmath 0127}_{j})_{j\geq 1}$. By definition of Gaussian processes in Hilbert space, ${\Greekmath 010C}|{\Greekmath 010D} \sim \mathcal{GP}({\Greekmath 010D}{\Greekmath 010C}_{0}, {\Greekmath 011B}_{0}^{2}\Omega)$ if and only if $\langle {\Greekmath 010C},{\Greekmath 0127}_{j}\rangle |{\Greekmath 010D} \sim \mathcal{N}(\langle {\Greekmath 010C}_{0},{\Greekmath 0127}_{j}\rangle {\Greekmath 010D}, {\Greekmath 011B}_{0}^{2}{\Greekmath 0115}_{j})$ for all $j\geq 1$. Hence, it follows that $$\left.\frac{\langle {\Greekmath 010C},{\Greekmath 0127}_{j}\rangle}{{\Greekmath 011B}_{0}\sqrt{{\Greekmath 0115}_{j}}}\right|{\Greekmath 010D} \sim \mathcal{N}\Big(\langle {\Greekmath 010C}_{0},{\Greekmath 0127}_{j}\rangle/({\Greekmath 011B}_{0}\sqrt{{\Greekmath 0115}_{j}}){\Greekmath 010D},1\Big).$$ Let us write ${\Greekmath 010C} = {\Greekmath 010C}_0{\Greekmath 010D} + U$, where $U\sim\mathcal{GP}(0,{\Greekmath 011B}_0^2 \Omega)$. Let ${\Greekmath 010D}^*({\Greekmath 010C})$ denote the minimizer of the least squares criterion: ${\Greekmath 010D}^*({\Greekmath 010C}):=\arg\min_{{\Greekmath 010D}}\sum_{j=1}^{\infty} (\langle{\Greekmath 010C},{\Greekmath 0127}_{j}\rangle - \langle{\Greekmath 010C}_0,{\Greekmath 0127}_{j}\rangle{\Greekmath 010D})^2/({\Greekmath 011B}_{0}^{2}{\Greekmath 0115}_{j})$ which takes the form
where ${\Greekmath 0118}_{j} := \langle U,{\Greekmath 0127}_{j}\rangle / ({\Greekmath 011B}_0\sqrt{{\Greekmath 0115}_j}) \sim i.i.d. \; \mathcal{N}(0,1)$, $j\geq 1$. If $\langle\Omega^{-1/2}U,\Omega^{-1/2}{\Greekmath 010C}_{0}\rangle =0$ then the second term on the right hand side of (ref) is zero and so ${\Greekmath 010D}^*({\Greekmath 010C}) = {\Greekmath 010D}$. It follows that we can write ${\Greekmath 010D}$ as a function of ${\Greekmath 010C}$ and we have exact estimability and Bayesian identification of ${\Greekmath 010D}$, that is, ${\Greekmath 010D}\in\overline{\mathcal{B}}$. As in Model 1, this property does not depend on the marginal prior on ${\Greekmath 010D}$ but crucially depends on the prior of ${\Greekmath 010C}$ given ${\Greekmath 010D}.$ The condition $\langle\Omega^{-1/2}U,\Omega^{-1/2}{\Greekmath 010C}_{0}\rangle =0$ means that the scaled error term and the prior mean functions have to be orthogonal in $L^2[0,1]$.
\paragraph{Model 3: Gaussian process and hyperparameter in the variance.} Let us consider the same setting as in Model 2 but suppose that now we are interested in the common variance parameter ${\Greekmath 011B}^2$. That is, the prior distribution for $({\Greekmath 010C},{\Greekmath 011B}^2)\in(B\times\mathbb{R}_{+},\mathcal{B}\otimes\mathcal{R})$ is specified as:
where $\mathcal{I}\Gamma$ denotes an inverse-gamma distribution and $\Omega:X\rightarrow X$ is a known covariance operator which is bounded, linear, positive definite, self-adjoint and trace-class. The hyperparameters ${\Greekmath 0117}_{0}$ and ${\Greekmath 011B}_{0}^{2}$ are known. Knowledge of $\Omega$ implies knowledge of its eigensystem $({\Greekmath 0115}_{j},{\Greekmath 0127}_{j})_{j\geq 1}$.\\ By definition, ${\Greekmath 010C}|{\Greekmath 011B}^2\sim \mathcal{GP}(0,{\Greekmath 011B}^{2}\Omega)$ if and only if $\left\{\langle{\Greekmath 010C},{\Greekmath 0127}_{j}\rangle/\sqrt{{\Greekmath 0115}_{j}}\right\}_{j\geq 1}|{\Greekmath 011B}^2\sim^{i.i.d.}\mathcal{N}(0,{\Greekmath 011B}^{2})$. By the strong law of large numbers we have
This convergence is valid with respect to both the conditional prior ${\Greekmath 0116}({\Greekmath 010C}|{\Greekmath 011B}^{2})$ of ${\Greekmath 010C}$ given ${\Greekmath 011B}^{2}$ and the joint prior ${\Greekmath 0116}({\Greekmath 010C},{\Greekmath 011B}^{2})$ of $({\Greekmath 010C},{\Greekmath 011B}^{2})$ and shows that ${\Greekmath 011B}^{2}$ is a ${\Greekmath 0116}-a.s.$ function of ${\Greekmath 010C}$. Therefore, ${\Greekmath 011B}^{2}$ is Bayesian identified. \paragraph{Model 4: Bayesian nonparametric and Dirichlet process.} We observe an $n$-sample from the following sampling model $x_{i}|F,G\sim\:i.i.d.\: F$, $i=1,\ldots,n$, where $F$ and $G$ are two probability distributions -- for instance on $\mathbb{R}$. We can interpret $F$ as a nuisance parameter and $G$ is the only parameter of interest which characterizes the distribution of $F$. The parameter $F$ generates a ${\Greekmath 011B}$-field $\mathcal{B}$ that is measurable identified while $G$ generates a ${\Greekmath 011B}$-field $\mathcal{G}$ that is not measurable identified. We specify a nonparametric Dirichlet process prior for $F$ with parameters $n_{0}$ and $G$: $F|G \sim \mathcal{D}ir(n_{0},G)$ and $G \sim {\Greekmath 0116}$, where ${\Greekmath 0116}$ denotes a probability measure on the space of distributions that generate almost surely a diffuse distribution, i.e. $G(x) = 0$, $\forall x$. The $F$ generated from the above Dirichlet process can be equivalently generated by using the stick-breaking representation as $F=\sum_{j}{\Greekmath 010B}_{j}{\Greekmath 010E}_{{\Greekmath 0118}_{j}}$, where $\{{\Greekmath 0118}_{j}\}_{j\geq 1}$ are independent draws from $G$, ${\Greekmath 010B}_{j} = v_{j}\prod_{k=1}^{j}(1 - v_{k})$ with $\{v_{j}\}_{j\geq 1}$ independent draws from a Beta distribution $\mathcal{B}e(1,n_{0})$ and $\{v_{j}\}_{j\geq 1}\perp\{{\Greekmath 0118}_{j}\}_{j\geq 1}$, see Appendix D in the Supplement. Then, $\lim_{J\rightarrow \infty}\frac{1}{J}\sum_{j=1}^{J}{\Greekmath 010E}_{{\Greekmath 0118}_{j}} = G$,\quad ${\Greekmath 0116}-a.s.$ which proves that $G$ is a ${\Greekmath 0116}$-a.s. function of $F$ and so, it is Bayesian identified. \paragraph{Model 5: Moment conditions and partially identified models.} Consider the setting of Example (ref) where the parameter of interest ${\Greekmath 010D}\in(\Gamma,\mathcal{G})$ is characterized through a moment condition of the type $\mathbf{E}_{F}(h(x,{\Greekmath 010D}))\geq 0$ for a known function $h$ and a data generating process $F$. We observe $n$ realisations $(x_{1},\ldots,x_{n})$ of $n$ i.i.d. random variables from $F$ and specify for $F$ a Dirichlet process prior with parameters $n_{0}$ and $F_{0}$:
The parameter $F$ generates the ${\Greekmath 011B}$-field $\mathcal{B}$ and is measurable identified. Suppose that ${\Greekmath 010D}$ is partially identified. Despite of this, ${\Greekmath 010D}$ can be Bayesian identified by specifying a conditional prior ${\Greekmath 0116}({\Greekmath 010D}|F)$ for ${\Greekmath 010D}$, given $F$, degenerate on a given functional of $F$. For example, if the restriction $\mathbf{E}_{F}(h(x,{\Greekmath 010D})) \geq 0$ writes ${\Greekmath 010D}\in[{\Greekmath 011E}_{1}(F),{\Greekmath 011E}_{2}(F)]$ for two functionals ${\Greekmath 011E}_{1}$, ${\Greekmath 011E}_{2}$ of $F$, then ${\Greekmath 0116}({\Greekmath 010D}|F)$ could be specified as a Dirac on $({\Greekmath 011E}_{1}(F) + {\Greekmath 011E}_{2}(F))/2$. We stress that Bayesian identification in this model is obtained artificially and in a different way than in Models 1-4 above because of the construction of a conditional degenerate prior for the partially identified parameter ${\Greekmath 010D}$, given the identified one. Instead, in Models 1-4 we have specified a marginal prior for the unidentified parameter that was not degenerate and Bayesian identification arose because of the specification of the conditional prior for the identified parameter.\\ This artificial way to obtain Bayesian identification through a degenerate prior can be reprehensible as it can be seen against the logic of partial identification. In section (ref) we will proceed without imposing Bayesian identification.
In this section we take the example of latent variable models to illustrate that it is possible to make Bayesian inference directly for the nonidentified model without introducing an identifying prior or identifying assumptions. Because these models are well known in the literature we discuss them briefly.\\ Consider discrete observable outcomes $y_i$ arising from the model $y_i = g(z_i)$, where $g(\cdot)$ is a given function and $z_i$ is a latent variable satisfying: $z_i = x_i'{\Greekmath 010E} + u_i$, where $x_i$ is an observable real-valued random vector, ${\Greekmath 010E}$ is the parameter vector of interest, and $u_i$ is an unobservable component with unrestricted variance. Identification problems appear if the $g(\cdot)$ function is invariant to location or scale transformations of $z$. Differently from frequentist analysis, imposing identifying restrictions is not necessary if one wants to conduct Bayesian analysis. With proper priors, one obtains the posterior for the nonidentified model, constructs an MCMC algorithm to simulate from this posterior and then marginalises to obtain the posterior of the identified parameters. It is often the case that it is easier to construct MCMC algorithms for unrestricted parameters and they have better mixing properties than MCMC algorithms for the corresponding identified model, see e.g. vanDykMeng2001.\\ The simpler example of a latent variable model is the binary logit / probit, where the observable outcome $y_i$ is binary and modeled as: $y_i = \mathbbm{1}\{z_i>0\}$ with $z_i$ a real-valued random variable. If $u_i$ follows a Gaussian (resp. Logistic) distribution we have the probit (resp. logit) model. The identification problem arises because the indicator function is invariant to scale transformation of $z_i$: multiplication of $z_i$ by a positive constant does not change the likelihood. Identification can be restored by imposing for instance the restriction $Var(u_i) = 1$. However, Bayesian analysis does not require any identifying restrictions. Other example of latent variable models are provided in Appendix A.
This section considers partially identified models as described in section (ref). Let us consider a model with an identified parameter ${\Greekmath 0112}$ and another parameter ${\Greekmath 010D}$ characterized by the condition
where ${\Greekmath 010D}\in\Gamma\subset\mathbb{R}^{d_{{\Greekmath 010D}}}$, $A: (\Theta,\mathcal{A},{\Greekmath 0116}) \times\Gamma\rightarrow \Phi$, for some finite or infinite dimensional space $\Phi$ and $A_{0}$ is a subset of $\Phi$. The function $A({\Greekmath 0112},{\Greekmath 010D})$ is usually either a likelihood function or a moment function and condition (ref) can contain both inequalities and equalities, see chen2017monte for interesting examples. The parameter ${\Greekmath 010D}$ is the parameter of interest and, depending on the relation (ref), it may be only partially identified which means that for a given ${\Greekmath 0112}$ there may exist more than one value of ${\Greekmath 010D}$ in $\Gamma$ satisfying (ref) but not all the values in $\Gamma$ satisfy the condition.\\ Let $\Gamma_{I} := \Gamma_{I}({\Greekmath 0112}) = \{{\Greekmath 010D}\in\Gamma; A({\Greekmath 0112},{\Greekmath 010D}) \in A_{0}\}$ be the identified set which is a proper subset of $\Gamma$ and depends on ${\Greekmath 0112}$. The model is point identified if $\Gamma_{I}$ is a singleton and is partially identified otherwise. Inference of partially identified models can focus either on ${\Greekmath 010D}$, or on the set $\Gamma_I$ or on both. In this section, we focus on $\Gamma_I$ and so do not endow ${\Greekmath 010D}$ with a prior. A prior on ${\Greekmath 010D}$ will be introduced in section (ref).\\ In the Bayesian analysis, the quantity $A({\Greekmath 0112},{\Greekmath 010D})$ in (ref) is a random element, where the randomness comes from ${\Greekmath 0112}$: for each ${\Greekmath 010D}\in\Gamma$, $A(\cdot,{\Greekmath 010D})$ is a function on $(\Theta,\mathcal{A})$ which takes values in the state space $(\Phi,\mathcal{B}(\Phi))$, where $\mathcal{B}(\Phi)$ denotes the Borel ${\Greekmath 011B}$-field associated with $\Phi$. Therefore, condition (ref) has to hold a.s. with respect to the prior distribution of ${\Greekmath 0112}$. For a realisation of ${\Greekmath 0112}$ in $\Theta$, the set $\Gamma_{I}$ is constructed on the basis of the trajectory ${\Greekmath 010D} \mapsto A({\Greekmath 0112},{\Greekmath 010D})$, ${\Greekmath 010D}\in\Gamma$. Therefore, $\Gamma_{I}({\Greekmath 0112})$ has to be interpreted as a random set where the randomness comes from ${\Greekmath 0112}$. This differs from the frequentist analysis of partially identified models where $\Gamma_I$ is non random.\\ The next definition describes a random closed set. For this, let $\mathcal{C}$ (resp. $\mathcal{O}$) denote a family of closed (resp. open) subsets of $\Gamma$.
The multivalued function $\Gamma_{I}(\cdot):\Theta \rightarrow \mathcal{O}$ is called a random open set if its complement is a random closed set. The following proposition gives a condition that guarantees that $\Gamma_{I}$ is a random closed set. A regular closed set $C$ is such that $C$ coincides with the closure of its interior, see Molchanov2005.
We assume in the following that $\Gamma_{I}({\Greekmath 0112})$ is a closed random set. Our analysis can be extended to the case where $\Gamma_{I}({\Greekmath 0112})$ is an open random set. It is convenient to introduce the stochastic process $\{g_{{\Greekmath 0112}}({\Greekmath 010D})\}_{{\Greekmath 010D}\in \Gamma}$ associated with condition (ref), where for every ${\Greekmath 010D}\in\Gamma$
So, $g_{{\Greekmath 0112}}({\Greekmath 010D})$ is the indicator of the random set $\Gamma_{I}({\Greekmath 0112})\subset\Gamma$ and ${\Greekmath 010D}$ is the index of the stochastic process $g_{(\cdot)}(\cdot):(\Theta,\mathcal{A},{\Greekmath 0116})\times\Gamma \rightarrow \{0,1\}$. When $\Gamma_{I}({\Greekmath 0112})$ is a separable random set, see Molchanov2005, then it can be completely described through its indicator function.
A Bayesian analysis requires the specification of a prior distribution for the random set $\Gamma_{I}$. Since $\Gamma_{I}$ depends on ${\Greekmath 0112}$, we propose to construct such a prior by first specifying a prior distribution for ${\Greekmath 0112}$ and then recovering from it the prior for $\Gamma_{I}$. Suppose $X \subseteq \mathbb{R}$ and let $F$ denote the data distribution. We suppose that ${\Greekmath 0112}$ may be written as a measurable functional of the distribution $F$, that is, there exists a measurable functional ${\Greekmath 011E}(\cdot)$ such that ${\Greekmath 0112} = {\Greekmath 011E}(F)$. For instance, ${\Greekmath 0112}=\mathbf{E}_{F}(x) = \int x dF(x)$.\\ In our analysis $F$ is not restricted to belong to some parametric class. Thus, we specify a Dirichlet process prior for $F$ and write $F\sim\mathcal{D}ir(n_{0},F_{0})$ where $n_{0}\in\mathbb{R}_{+}$ and $F_{0}$ is a diffuse probability measure on $\mathbb{R}$, i.e. $F_{0}(y) = 0$, $\forall y\in\mathbb{R}$, see e.g. Ferguson1973 and Florens2002. The prior distribution for $\Gamma_{I}({\Greekmath 0112})$ is obtained from this prior for $F$ through the prior capacity functional. Define $\mathcal{K}$ as the family of compact subsets of $\Gamma$. The prior capacity functional $T_{\Gamma_{I}}:\mathcal{K} \mapsto[0,1]$ is given by
where the probability $P$ is determined by the prior distribution of ${\Greekmath 0112}$ which in turns is determined by the Dirichlet process prior on $F$: $F\sim \mathcal{D}ir(n_{0},F_{0})$ since ${\Greekmath 0112}={\Greekmath 011E}(F)$.\\ When the prior capacity functional is defined on singletons instead of on $\mathcal{K}$, i.e. $K = \{{\Greekmath 010D}\}$ for some ${\Greekmath 010D}\in\Gamma$, then it is called prior coverage function of $\Gamma_{I}$ and denoted by $p_{\scriptscriptstyle{\Gamma_{I}}}({\Greekmath 010D})$. In particular, $\forall {\Greekmath 010D} \in\Gamma$
where $g_{{\Greekmath 0112}}(\cdot)$ is the stochastic process defined in (ref) and $P$ and $\mathbf{E}$ are the probability and expectation, respectively, taken with respect to the prior of ${\Greekmath 0112}$. The appealing fact of the prior coverage function with respect to the prior capacity functional is that $p_{\scriptscriptstyle{\Gamma_{I}}}({\Greekmath 010D})$ can be represented graphically in an easier way than $T_{\Gamma_{I}}(K)$ (at least if $\Gamma\subset\mathbb{R}$), see for instance Figures (ref), (ref) and 9 where we represent $p_{\scriptscriptstyle{\Gamma_{I}}}(\cdot)$ and the posterior coverage function $p_{\scriptscriptstyle{\Gamma_{I}}}(\cdot|x)$. We remark that the prior coverage function $p_{\scriptscriptstyle{\Gamma_{I}}}({\Greekmath 010D})$ does not characterize the distribution of the stochastic process $g_{{\Greekmath 0112}}({\Greekmath 010D})$ which is instead characterized by its finite-dimensional distributions.\\ We detail now the construction of the prior and posterior capacity functionals with the help of a generic example. Suppose ${\Greekmath 0112} := ({\Greekmath 0112}_{1},{\Greekmath 0112}_{2})' \in\mathbb{R}^2$, $A({\Greekmath 0112}, {\Greekmath 010D}) = ({\Greekmath 0112}_{1} - {\Greekmath 010D}, {\Greekmath 010D} - {\Greekmath 0112}_{2})'$ for some ${\Greekmath 010D}\in\mathbb{R}$ and $A_{0} = (-\infty, 0]\times (-\infty, 0]$. Hence, the condition $A({\Greekmath 0112},{\Greekmath 010D})\in A_{0}$ writes as ${\Greekmath 010D} \in [{\Greekmath 0112}_{1},{\Greekmath 0112}_{2}]$ and $\Gamma_{I} = [{\Greekmath 0112}_{1},{\Greekmath 0112}_{2}]$. We recall that the aim is to make inference on the identified set $\Gamma_I$ and not on the partially identified parameter. To start with, suppose that we specify a parametric prior for ${\Greekmath 0112}$. For instance, ${\Greekmath 0112}_{1}\sim \mathcal{U}[0,1]$ and ${\Greekmath 0112}_{2}\sim\mathcal{U}[1,2]$. For any $K\in\mathcal{K}$, write $K = [K_{1},K_{2}]$, $\bar{K} = [\bar{K}_{1},\bar{K}_{2}] := K \cap[0,1]$ and $\bar{\bar{K}} = [\bar{\bar{K}}_{1},\bar{\bar{K}}_{2}] := K \cap[1,2] $. The prior capacity functional then is:
while the prior coverage function is given by $p_{\scriptscriptstyle{\Gamma_{I}}}({\Greekmath 010D}) = P({\Greekmath 010D} \in [{\Greekmath 0112}_{1},{\Greekmath 0112}_{2}]) = {\Greekmath 010D} \mathbbm{1}\{{\Greekmath 010D}\in[0,1]\} + (2 - {\Greekmath 010D}) \mathbbm{1}\{{\Greekmath 010D}\in[1,2]\}$, where $P$ is the distribution with respect to the prior of ${\Greekmath 0112}$.\\ With this intuition in mind, let us move to the nonparametric Bayesian approach. This approach is based on a Dirichlet process prior and requires to write ${\Greekmath 0112} =({\Greekmath 0112}_{1},{\Greekmath 0112}_{2})'$ as ${\Greekmath 0112}_{i} = {\Greekmath 011E}_{i}(F)$, for a measurable functional ${\Greekmath 011E}_{i}$, $i=1,2$. If we observe realizations of a random vector $Y$ from a distribution $F$ then the Bayesian model writes $Y|F \sim F$, $F \sim \mathcal{D}ir(n_{0},F_{0})$. The probability measure $F_{0}$ should be chosen such that ${\Greekmath 011E}_{2}(F) > {\Greekmath 011E}_{1}(F)$, ${\Greekmath 0116}$-a.s. By using the stick-breaking representation of the Dirichlet process, see Appendix D in the Supplement, the prior capacity functional of $\Gamma_{I} = [{\Greekmath 011E}_{1}(F),{\Greekmath 011E}_{2}(F)]$ is given by
where $\{{\Greekmath 0118}_{j}\}_{j\geq 1}$ are independent draws from $F_{0}$, ${\Greekmath 010E}_{{\Greekmath 0118}_{j}}$ denotes the Dirac mass in ${\Greekmath 0118}_{j}$, ${\Greekmath 010B}_{j} = v_{j}\prod_{l=1}^{j-1}(1 - v_{l})$ with $\{v_{l}\}_{l\geq 1}$ independent draws from a Beta distribution $\mathcal{B}e(1,n_{0})$ and $\{v_{j}\}_{j \geq 1}$ are independent of $\{{\Greekmath 0118}_{j}\}_{j\geq 1}$. In Appendix D we recall how to simulate ${\Greekmath 011E}_{i}(F)$, $i=1,2$, from the prior and posterior distribution by using this representation. The prior coverage function of $\Gamma_{I}$ is: for every ${\Greekmath 010D}\in\Gamma$,
where we have taken $K = \{{\Greekmath 010D}\}$ and $P$ is the prior distribution of ${\Greekmath 0112}$, $\{{\Greekmath 010B}_j\}$ and $\{{\Greekmath 0118}_j\}$. In general we do not have an analytic form for $T_{\Gamma_{I}}$ and $p_{\scriptscriptstyle{\Gamma_{I}}}$ but we have a perfect knowledge of them since we can easily simulate from $T_{\Gamma_{I}}$ and $p_{\scriptscriptstyle{\Gamma_{I}}}$ by using the stick-breaking representation of the Dirichlet process.\\ After observing an $n$-sample of $Y$, $(y_{1}, \ldots,y_{n})$, one computes the posterior distribution of $F$ as
where $F_{n}(\cdot) := \frac{1}{n}\sum_{j}{\Greekmath 010E}_{y_{j}}(\cdot)$ denotes the empirical cumulative distribution. If the true data distribution $F$ is such that ${\Greekmath 011E}_{2}(F) > {\Greekmath 011E}_{1}(F)$ then the same is true for the distribution $F$ generated by the posterior. The posterior capacity functional is denoted by $T_{\Gamma_{I}}(K|\{y_{i}\}_{i=1}^{n})$ and given by: $\forall K\in\mathcal{K}$,
where $\{{\Greekmath 010B}_{j}\}$, $\{{\Greekmath 010E}_{{\Greekmath 0118}_{j}}\}$, $\{{\Greekmath 010E}_{y_{j}}\}$ and $\{{\Greekmath 0118}_{j}\}$, are as above, ${\Greekmath 011A}$ is drawn form a Beta distribution $\mathcal{B}e(n,n_{0})$ independently of the other quantities and $({\Greekmath 010C}_{1}, \ldots, {\Greekmath 010C}_{n})$ are drawn from a Dirichlet distribution with parameters $(1,\ldots, 1)$ on the simplex $S_{n-1}$ of dimension $(n-1)$. For every ${\Greekmath 010D}\in\Gamma$, the posterior coverage function $p_{\scriptscriptstyle{\Gamma_{I}}}({\Greekmath 010D}|\{y_{i}\}_{i=1}^{n})=P\Big({\Greekmath 010D}\in[{\Greekmath 011E}_{1}(F),{\Greekmath 011E}_{2}(F)]\Big|\{y_{i}\}_{i=1}^{n}\Big)$ is: $\forall {\Greekmath 010D}\in\Gamma$,
For simplicity we have presented only the case where the condition $A({\Greekmath 0112},{\Greekmath 010D})\in A_{0}$ writes as ${\Greekmath 010D}\in[{\Greekmath 011E}_{1}(F),{\Greekmath 011E}_{2}(F)]$ but our nonparametric method can be generalized to the case where $\Gamma_{I}$ is not an interval. In that case, if a Dirichlet process prior is specified for $F$, the prior capacity functional of $\Gamma_{I}$ is given by: $\forall K\in\mathcal{K}$,
and the posterior capacity functional is: $\forall K\in\mathcal{K}$,
where $\{{\Greekmath 010B}_{j}\}$, $\{{\Greekmath 010E}_{{\Greekmath 0118}_{j}}\}$, $\{{\Greekmath 010E}_{y_{j}}\}$, $\{{\Greekmath 0118}_{j}\}$, ${\Greekmath 011A}$ and $({\Greekmath 010C}_{1},\ldots,{\Greekmath 010C}_{n})$ are as above.\\ Once the posterior capacity functional is available, an estimator for $\Gamma_{I}({\Greekmath 0112})$ can be easily constructed. One possibility is to fix ${\Greekmath 0112}$ equal to its posterior mean or median, denote it by $\widehat{\Greekmath 0112}$, and take the corresponding $\Gamma_{I}(\widehat{\Greekmath 0112})$ as an estimator for $\Gamma_{I}$. In alternative, one could construct a closed set $C_{{\Greekmath 010B}}$ that satisfies the following condition
for some ${\Greekmath 010B}\in[0,1]$. This is the usual posterior credible region. A similar estimator is proposed e.g. by LiaoSimoni2019 and chen2017monte. We remark that the probability in ((ref)) is determined by the posterior Dirichlet process and is the posterior containment functional evaluated at $C_{{\Greekmath 010B}}$.\\
In this section we provide two examples where we use our proposed Bayesian nonparametric method described in section (ref). Other two examples are developed in Appendix B in the Supplement. For simplicity, we only focus on the prior and posterior coverage functions, which are easy to represent graphically.
In sections (ref) and (ref) we have considered statistical models where no parameter is marginalized out. We refer to these models as full models. When the parameter of interest is a sub-parameter ${\Greekmath 010D}$ of the whole model parameter one might want to perform the analysis by getting rid of the parameters of the model that are not of interest. Therefore, it is natural to examine the marginal model in this sub-parameter.
Let ${\Greekmath 0112} := ({\Greekmath 010C},{\Greekmath 010D})$ denote the whole parameter of the model where ${\Greekmath 010C}$ is identified and ${\Greekmath 010D}$ is the parameter of interest that is nonidentified. For instance, ${\Greekmath 010D}$ is the parameter related to a latent variable as in section (ref). One can specify the prior for ${\Greekmath 0112}$ as ${\Greekmath 0116}({\Greekmath 0112}) = {\Greekmath 0116}({\Greekmath 010C}|{\Greekmath 010D}){\Greekmath 0116}({\Greekmath 010D})$. Hence, the marginal model is obtained by integrating out the parameter ${\Greekmath 010C}$ in the original model with respect to the prior ${\Greekmath 0116}({\Greekmath 010C}|{\Greekmath 010D})$:
where $P^{{\Greekmath 010D}}(\cdot)$ denotes the integrated sampling distribution which depends on ${\Greekmath 010D}$. The corresponding Lebesgue density function (or marginal likelihood) writes $p(x|{\Greekmath 010D}) = \int p(x|{\Greekmath 010C},{\Greekmath 010D}){\Greekmath 0116}({\Greekmath 010C}|{\Greekmath 010D}) d{\Greekmath 010C}$. The predictive density $p(x)$, obtained by integrating out ${\Greekmath 010D}$ from $p(x|{\Greekmath 010D})$ with respect to the prior, is the same as in the full model and the marginal posterior ${\Greekmath 0116}({\Greekmath 010D}|x)$ of ${\Greekmath 010D}$ is obtained by marginalizing the joint posterior ${\Greekmath 0116}({\Greekmath 010C},{\Greekmath 010D}|x)$ with respect to ${\Greekmath 010C}$.
Note that, in the previous example, even if ${\Greekmath 010D}$ is identified in the marginal model it is not exactly estimable. In fact, Theorem (ref) does not apply because the marginal model is not i.i.d. Exact estimability would hold only if the conditional distribution of ${\Greekmath 010D}$ given ${\Greekmath 010C}$ was a degenerated Dirac measure on a deterministic function of ${\Greekmath 010C}$.\\ In the setting of Example (ref), ${\Greekmath 010C}$ can be interpreted as an heterogeneity parameter whose distribution depends on an unidentified parameter ${\Greekmath 010D}$, see e.g. HeckamnSigner1984. In the frequentist setting, the conditional distribution of ${\Greekmath 010C}|{\Greekmath 010D}$ is part of the data generating process while in the Bayesian setting the conditional distribution of ${\Greekmath 010C}|{\Greekmath 010D}$ is the prior.\\ The next theorem considers the asymptotic behavior of a sub-parameter which might be nonidentified in the full model. It states that the posterior mean of a sub-parameter $c$ converges $\Pi$-a.s. to the conditional prior mean given the identified parameter. Remark that the a.s. in the theorem is with respect to the joint distribution.
Consider now the partially identified model of section (ref). Let $Y$ be an observable random variable with distribution $F$. The parameter $F$ is identified and suppose that there is another parameter ${\Greekmath 0112}$ of the model that is identified and that can be written as ${\Greekmath 0112} = {\Greekmath 011E}(F)$ for some functional ${\Greekmath 011E}$. The parameter of interest is denoted by ${\Greekmath 010D}$ and is related to ${\Greekmath 0112}$ by relation ((ref)): $A({\Greekmath 0112},{\Greekmath 010D}) \in A_{0}\subset\Phi$. Let $\Gamma$ be provided with a ${\Greekmath 011B}$-field $\mathcal{G}$. Hence, the parameter space is $(\Theta \times \Gamma, \mathcal{A}\otimes\mathcal{G})$. By using the structural relation (ref) we now specify a restricted prior ${\Greekmath 0116}({\Greekmath 010D}|{\Greekmath 0112})$ for ${\Greekmath 010D}$ conditional on ${\Greekmath 0112}$. In particular, ${\Greekmath 0116}({\Greekmath 010D}|{\Greekmath 0112})$ has support equal to the set of the ${\Greekmath 010D}$s that satisfy the constraint $A({\Greekmath 0112},{\Greekmath 010D})\in A_{0}$ in (ref) for a given ${\Greekmath 0112}\in \Theta$.\footnote{In alternative, we may relax this constraint on the support of ${\Greekmath 0116}({\Greekmath 010D}|{\Greekmath 0112})$ into a constraint on the hyperparameter of the distribution of ${\Greekmath 010D}$, as illustrated by the examples below.} The marginal posterior of ${\Greekmath 010D}$ is: $\forall \Gamma_{1}\in\mathcal{G},$
where ${\Greekmath 0116}({\Greekmath 0112}|y)$ denotes the posterior distribution of ${\Greekmath 0112}$ and ${\Greekmath 0116}(dF|y)$ denotes the posterior distribution of $F$. In the second equality we have written the integral in terms of ${\Greekmath 0116}(dF|y)$ to stress the fact that the prior distribution of ${\Greekmath 0112}$ is recovered from the Dirichlet process prior for $F$ as described in section (ref). It is clear that ${\Greekmath 010D}$ is identified in the marginal model since its marginal prior distribution is updated by the data. In addition, Theorem (ref) above applies also to the case of partially identified models.\\ While the marginal posterior density ${\Greekmath 0116}({\Greekmath 010D}|y)$ of ${\Greekmath 010D}$ is not usually known in closed-form -- in particular if ${\Greekmath 0112}$ is infinite dimensional -- one can easily simulate from it. For this, one first simulates ${\Greekmath 0112}$ given $y$ from ${\Greekmath 0116}({\Greekmath 0112}|y)$ and then, for each draw of ${\Greekmath 0112}$, one simulates ${\Greekmath 010D}$ given ${\Greekmath 0112}$ from ${\Greekmath 0116}({\Greekmath 010D}|{\Greekmath 0112})$. This simulation scheme produces draws from ${\Greekmath 0116}({\Greekmath 010D}|y)$ and this is due to the lack of identification of ${\Greekmath 010D}$ which implies that ${\Greekmath 010D} \perp y|{\Greekmath 0112}$, see section (ref). Having a marginal posterior distribution ${\Greekmath 0116}({\Greekmath 010D}|y)$ of the parameter of interest ${\Greekmath 010D}$ is important for instance in a decision problem setting. In fact, knowledge of ${\Greekmath 0116}({\Greekmath 010D}|y)$ allows to select the most likely value of ${\Greekmath 010D}$ or the region inside $\Gamma_{I}$ with the highest posterior probability. Such a selection is clearly affected by the choice of the prior on ${\Greekmath 010D}$.
In this section we develop further Example (ref) of section (ref) by endowing the parameter ${\Greekmath 010D}$ with a conditional prior distribution given ${\Greekmath 0112}={\Greekmath 011E}(F)$ that we denote by ${\Greekmath 0116}_{{\Greekmath 010D}}^{F}$: ${\Greekmath 010D}|{\Greekmath 0112}\sim {\Greekmath 0116}_{{\Greekmath 010D}}^{F} := {\Greekmath 0116}({\Greekmath 010D}|{\Greekmath 011E}(F))$. Additional examples are developed in Appendix C in the Supplement. In our simulations we consider four different specifications for ${\Greekmath 0116}_{{\Greekmath 010D}}^{F}$, where the hyperparameters $a_0$ and $b_0$ are specified in each specific example:
\setcounter{example}{2}
\setcounter{section}{5}
This paper studies theoretical properties and implementation of the Bayesian approach in various models that lack identification. As examples of unidentified models, we analyse nonparametric models with heterogeneity modeled either as a Gaussian process or as a Dirichlet process where the parameter of interest is the (hyper)parameter of the heterogeneity distribution which is unidentified. We also analyse unidentified latent variable and partially identified models.\\ In partially identified models we propose to construct the prior and posterior of the identified set through the prior and posterior capacity functionals. The prior capacity functional is obtained as the transformation of a Dirichlet process, so that our approach is completely nonparametric. The proposed procedure is appealing since, even if the posterior capacity functional has a complicated expression, simulating from it is simple.\\ Finally, we discuss models that have some parameters that are identified and others that are not. For these models, we show that the parameter that is unidentified or partially identified in the full model is identified in the marginal model.