The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
111,051 characters
Nonparametric Bayesian Inference for Partially Identified Discrete Response Models
\pagestyle{plain}
\maketitle
\begin{abstract}
This paper proposes a nonparametric Bayesian inference framework for partially identified discrete response models. The key observation is that these models map a reduced-form conditional choice probability to an identified set. Consequently, nonparametric Bayesian inference for the conditional probability mass function leads to Bayesian inference for the identified set. The inference framework nests conditional moment inequalities and linear systems with unknown coefficients as special cases. Importantly, our proposal does not require converting conditional moments into unconditional moments or discretizing covariates. We show that the posterior is consistent for the true identified set when the model is correctly specified, show that the posterior can consistently detect model misspecification, and show posterior consistency for a pseudo-identified set that is valid under misspecification. We also verify the assumptions for a class of priors based on Gaussian processes that we use to implement our proposal. These priors offer similar flexibility to frequentist partial identification methods, and are computationally attractive because posterior sampling can be performed in closed-form. We also show that many of the ideas in this paper extend to continuous responses and aggregated discrete responses (e.g., market shares).
\end{abstract}
\newpage
\section{Introduction}
\subsection{Overview.} Many economic models with discrete responses lead to set-identified structural parameters. These include games with multiple equilibria \citep{tamer2003incomplete,ciliberto2009market}, semiparametric static binary/multinomial choice models with fixed effects \citep{manski1987semiparametric,shi2018estimating,gao2024identification,pakes2024moment}, dynamic models with state dependence \citep{Heckman1981a,honore2006bounds,torgovitsky2019nonparametric,khan2021inference,pakes2021unobserved,khan2023identification}, limited consideration discrete choice models \citep{barseghyan2021heterogeneous,LU2022368}, ordered choice models \citep{CHESHER201233,pakes2015moment}, network formation models \citep{sheng2020structural}, and binary regression models with interval-censored covariates \citep{manski2002inference}, to name a few. A common statistical theme across these examples is that the identified set is known up to an unrestricted reduced-form conditional probability mass function (PMF). Consequently, uncertainty about the identified set entirely reflects that of the PMF, and, as a result, statistical inference for the PMF should translate to the identified set.
Using this observation, our paper proposes a nonparametric Bayesian inference framework for identified sets defined by conditional PMFs. Concretely, let the observed data be $(Y,X')'$, where $Y$ is a discrete random variable and $X$ is a (possibly) continuous random vector, and let the target parameter be an identified set $\Gamma_{I} = \{\gamma: f(X,p(X),\gamma) \geq 0 \ a.s.\}$, where $\gamma$ is a finite-dimensional economic parameter (e.g., payoff parameters), $p(X)$ is a reduced-form conditional PMF for $Y$ given $X$, and $f(X,p(X),\gamma)$ is a known vector of possibly nonseparable functions (with inequalities evaluated elementwise). The central idea of our proposal is that $\Gamma_{I}$ can be viewed as the output of a correspondence that takes $p$ as an input. Consequently, a Bayesian inference framework for $\Gamma_{I}$ arises in which a marginal posterior $\Gamma_{I}|Y,X$ is obtained from a posterior $p|Y,X$ for $p$. Beyond this, we provide implementation details for Gaussian process (GP) priors, prove a general posterior consistency theorem for $\Gamma_{I}|Y,X$ (and for a pseudo-identified set that accommodates misspecification), and verify the conditions of the theorem for GPs. The examples below highlight the breadth of our framework; Section \ref{sec:extensions} also demonstrates how many of our results extend to continuous $Y$ and aggregated discrete choice (e.g., $Y$ is a market share).
\begin{example}[Conditional Moment Inequalities]\label{ex:momentsineq}
Let $g(Y,X,\gamma)$ be a vector of known functions. The conditional moment inequality (CMI) model says $\gamma$ is in the identified set iff $E[g(Y,X,\gamma)|X] \geq 0$ a.s. CMIs arise in discrete choice, for instance, due to model incompleteness \citep{tamer2003incomplete,ciliberto2009market,sheng2020structural,barseghyan2021heterogeneous,LU2022368} or weak assumptions about unobserved heterogeneity \citep{manski2002inference,shi2018estimating,pakes2021unobserved,pakes2024moment}. Defining $f(X,p(X),\gamma) = \sum_{y}g(y,X,\gamma)p_{y}(X),$ where $p_{y}(X)$ is the conditional probability that $Y=y$ under $p(X)$, CMIs fall within our framework.
\end{example}
\begin{example}[Linear Systems]
Some discrete choice models lead to identifying restrictions of the form $\mathbf{1}\{p(X) \in R\}c_{R}(X)\gamma \geq 0 \ \forall \ R \in \mathcal{R}$ a.s., where $R$ is some restriction on $p(X)$, $c_{R}(X)$ is a known transformation that depends on $R$, $\mathcal{R}$ is a set of restrictions, and $\gamma$ is a structural parameter.\footnote{For instance, in the dynamic binary panel data model of \cite{khan2023identification}, inequalities like $\mathbf{1}\{p(Y_{t-1}=1,Y_{t}=1|X) + P(Y_{s-1}=1,Y_{s} = 0|X) \geq 1\}\{(X_{t}-X_{s})'\beta + \theta\} \geq 0$, where $s,t \in \{1,...,T\}$ are time periods, and $X=(X_{t})_{t=1}^{T}$ is a vector of covariates, characterize the sharp identified set.} These are a special case of a class of linear systems $A(X,p(X))\gamma \geq b(X,p(X))$ a.s., where $A(X,p(X))$ is a matrix and $b(X,p(X))$ is a vector, which is cast within our framework by defining $f(X,p(X),\gamma) = A(X,p(X))\gamma-b(X,p(X))$.
\end{example}
Our proposal has several appealing features. First, our approach accommodates continuous $X$ without requiring the conversion of conditional moments into unconditional moments or explicitly discretizing the support of $X$ into bins. In CMI models, the one-sided nature of inequality restrictions means that making these choices in a way that preserves identifying information is challenging \citep{khan2009inference, andrews2013inference}. This has resulted in a widespread empirical practice of basing inference on ad hoc choices. For instance, in many empirical implementations of CMIs, the conditioning variables are discretized into finite cells or groups. The resulting cell-level inequalities can then be treated as a finite set of unconditional moment inequalities, typically by multiplying the conditional moment by cell indicators or other nonnegative `instruments' and averaging (see, for example, \cite{kline2016bayesian,chen2018monte,ciliberto2021market, pakes2021unobserved}, among many others). Ad hoc selections of unconditional moments may result in substantial loss of sharpness under correct specification, while conclusions drawn from outer sets based on unconditional moments may be sensitive to instrument choice when the underyling CMIs are misspecified \citep{li2024discordant}. Similarly, empirical applications of linear systems typically involve coarsened covariates. For instance, the linear programming formulation \cite{honore2006bounds} presumes that the covariates have finite support (so as to obtain a finite set of inequalities), while the empirical implementation of the identified set in \cite{khan2023identification} binarizes the covariates, despite the finding that rich covariate support can lead to much sharper identified sets (and even point identification in some cases). Similar discretizing steps were taken in \cite{ciliberto2009market} and \cite{kline2016bayesian}. In both CMIs and linear systems, our procedure obviates the need for such discretization and/or selecting unconditional moments by targeting the conditional PMF directly.
Second, our approach is flexible and amenable to computation. We provide implementation details for a class of priors that combine stick-breaking and Gaussian process (GP) methods.\footnote{Stick-breaking is a technique that dates back to \cite{halmos1944random} and can be thought of as representing a discrete distribution in terms of its hazard functions. It was famously employed by \cite{sethuraman1994constructive} as a computational device for the \cite{ferguson1973bayesian} Dirichlet process, however these techniques are increasingly used for Bayesian conditional distribution estimation \citep{dunson2008kernel,chung2009nonparametric,rodriguez2011nonparametric,ren2011logistic}.} GP priors allow a researcher to flexibly incorporate nonparametric properties, such as smoothness, and produce nonparametric Bayesian procedures with similar flexibility to frequentist nonparametric estimators \citep{vaart2008rates}. Consequently, our implementation shares the advantageous flexibility of frequentist inference for identified sets, while enjoying all of the niceties of Bayesian inference. These priors are also computationally attractive because the stick-breaking characterization enables the use of the P\'{o}lya-Gamma data augmentation technique of \cite{polson2013bayesian} to derive Gibbs sampling algorithms in which GPs form conditionally conjugate priors. As a result, sampling from the posterior for the reduced-form PMFs amounts to generating draws from standard parametric distributions. Given draws from $p|Y,X$, the researcher can use \textit{any} appropriate method for computing $\Gamma_{I}$ to obtain the corresponding draws from $\Gamma_{I}|Y,X$, while the two-step nature of our proposal means that computing the $\Gamma_{I}$ draws is parallelizable.
The main theoretical results in the paper are posterior consistency guarantees for the identified set. A possible concern about Bayesian inference for partially identified models is that, even with infinite data, the prior does not fully revise \citep{poirier1998revising,moon2012bayesian,giacomini2021robust}. In our setting, the prior is placed directly on an identifiable reduced-form parameter $p$ and the identified set $\Gamma_{I}$ is viewed as a point identified, set-valued parameter. Consequently, under suitable continuity restrictions on the mapping from reduced form PMFs to identified sets, fully revised beliefs for the PMF should translate to fully revised beliefs about the identified set. Our core posterior consistency theorem confirms this intuition, providing high-level conditions under which posterior consistency for the conditional choice probabilities leads to the posterior for the identified set to concentrate around the true identified set. In this sense, the prior for $p$ has an asymptotically negligible effect on inferences about $\Gamma_{I}$. We also prove posterior consistency for a pseudo-identified set $\tilde{\Gamma}_{I}$ that is always nonempty (and coincides with $\Gamma_{I}$ when $\Gamma_{I}$ is not empty), and show that the posterior probability that the identified set is empty is a consistent diagnostic for model misspecification. This is important because several of the examples that fit within our framework are prone to model misspecification, for instance, due to parametric restrictions on unobserved heterogeneity \citep{ciliberto2009market,barseghyan2021heterogeneous,ciliberto2021market}.
Our posterior consistency theorems are based on assumptions that the posterior concentrates on sets of $p$ for which certain uniform convergence and lower hemicontinuity conditions are satisfied. While these conditions are straightforward to verify in finite-dimensional models, posterior consistency in infinite-dimensional models can be a delicate issue, evidenced by canonical inconsistency results, such as those in \cite{freedman1963asymptotic, freedman1965asymptotic} and \cite{diaconis1986inconsistent,diaconis1986consistency,diaconis1993nonparametric}, and more recent work highlighting the intricacies of posterior consistency in norms relevant to statistical functionals \citep{gine2011rates,castillo2014bayesian,ho2024bayesian}. Since we cast the identified set as a statistical functional, this second set of concerns is relevant for our framework. Conscious of this, we enhance the credibility of our general theory with a set low-level sufficient conditions for an important class of identified sets (i.e., they apply to linear systems and some conditional moment inequality models) under both correct specification (i.e., $\Gamma_{I}$) and misspecification (i.e., $\tilde{\Gamma}_{I}$). Then, noticing a common underlying supremum norm posterior consistency requirement for $p$, we derive conditions under which the stick-breaking GP prior used for the implementation meets this condition, and thereby provides a complete verification of posterior consistency. As an intermediate technical result, we find conditions under which the stick-breaking GP posterior for $p$ converges to the truth $p_{0}$ in empirical mean-square at the minimax optimal rate, which may be of independent interest.
\subsection{Literature.}
This paper is related to several literatures in econometrics. The first is Bayesian analysis of partially identified models. We build on \cite{kline2016bayesian} by performing Bayesian inference for an identified set by evaluating it at draws from the posterior of an identifiable reduced-form parameter (i.e., the PMF). A key difference, however, is that \cite{kline2016bayesian} explicitly restrict the reduced-form parameter to be finite-dimensional, thereby ruling out partially identified discrete response models with continuous covariates. \cite{LIAO2019338} and \cite{florens2021revisiting} also derive reduced-form Bayesian inference frameworks for identified sets using support functions and Dirichlet processes, respectively. These papers similarly restrict attention to reduced-form parameters that are estimable at the parametric rate, precluding settings where the identified set is indexed by a reduced-form conditional PMF with continuous conditioning variables. \cite{norets2014semiparametric} propose Bayesian inference for \cite{rust1987optimal}-type structural dynamic discrete choice models by assuming a joint prior for the reduced-form conditional choice probabilities and observable state transition distributions, but restrict attention to finite observable state spaces. More recently, taking a posterior for an identified set as given, \cite{kline2024counterfactual} propose a Bayesian inference framework for counterfactual prediction sets, and establish continuous mapping theorems under which posterior consistency for the identified set implies predictive consistency for the counterfactuals. Since our posterior consistency results for $\Gamma_{I}$ match those required by \cite{kline2024counterfactual}, an implication of our theory is posterior consistency for counterfactual prediction sets without requiring discretization of covariates. Other papers on Bayesian inference under partial identification include \cite{poirier1998revising}, \cite{liao2010bayesian}, \cite{moon2012bayesian}, \cite{giacomini2021robust}, and \cite{christensen2026optimal}.\footnote{There is also a related literature in statistics on nonparametric Bayesian level set estimation. See for example \cite{gayraud2005rates,gayraud2007consistency}'s work on density level sets and \cite{li2021posterior} who focus on level sets of a general unknown function.}
More broadly, our paper fits within the literature on reduced-form Bayesian inference for econometric models, which refers to situations in which a parameter of interest is viewed as a functional of an identifiable reduced-form parameter (for which a prior is assumed).\footnote{The reduced-form Bayesian literature can be connected to a large literature in statistics on Bayesian estimation of statistical functionals. A nonexhaustive list of contributions includes \cite{ferguson1973bayesian}, \cite{rubin1981bayesian}, \cite{newton1994approximate}, \cite{rivoirard2012bernstein}, \cite{castillo2015bernstein}, \cite{lyddon2019general}, \cite{nickl2020nonparametric}, \cite{monard2021statistical}, \cite{ray2021bernstein}, and \cite{li2026empiricallikelihoodgenerativeai}.} \cite{walker2026semiparametric} considers Bayesian inference for parameters point identified by conditional moment equalities via a nonparametric prior for the conditional density of the endogenous variables given the exogenous variables. Our paper follows a similar thought experiment, except that it assumes a nonparametric prior for a conditional PMF and focuses on identified sets determined by inequality restrictions on the conditional distribution. In this sense, the conceptual relationship between our paper and \cite{walker2026semiparametric} parallels how \cite{kline2016bayesian} connects the seminal inference framework of \cite{chamberlain2003nonparametric} to identified sets. \cite{norets2022adaptiveconditional,norets2022adaptive} develop a nonparametric Bayesian frameworks for reduced-form mixed continuous-discrete distributions, with a key empirical motivation being that firm entry parameters (e.g., \cite{pakes2007simple}) are functionals of these distributions. \cite{ray2018semiparametric}, \cite{breunig2025double,breunig2026semiparametricbayesiandifferenceindifferences}, \cite{ditraglia2025bayesian}, \cite{yiu2025semiparametric}, and \cite{ye2026nonparametric} propose reduced-form Bayesian approaches for treatment effect parameters.
Finally, our paper contributes to a large literature on inference under partial identification. Since we characterize identified sets as level sets of criterion functions, our approach offers a Bayesian analogue to frequentist inference procedures proposed in \cite{manski2002inference}, \cite{chernozhukov2007estimation}, \cite{bugni2010bootstrap}, \cite{romano2010inference}, \cite{menzel2014consistent}, \cite{chernozhukov2015inference}, and \cite{chen2018monte}. Our two leading examples, conditional moment inequalities and linear systems, connects our Bayesian framework to frequentist inference frameworks for these models, such as \cite{khan2009inference}, \cite{andrews2013inference,andrews2014nonparametric,andrews2017inference}, \cite{chernozhukov2013intersection}, \cite{armstrong2014weighted,armstrong2015asymptotically}, \cite{armstrong2016multiscale}, \cite{chetverikov2018adaptive}, \cite{andrews2023inference}, and \cite{chernozhukov2023constrained} for conditional moment inequalities, and \cite{bai2022testing}, \cite{fang2023inference}, \cite{goff2025inference}, and \cite{bai2026inference} for linear systems. There are many other frequentist approaches to inference under partial identification; see \cite{canay2017practical}, \cite{Ho_Rosen_2017}, \cite{MOLINARI2020355}, \cite{KLINE2021345}, and \cite{CANAY2023105558} and the references therein.
\subsection{Outline.} The paper is organized as follows. Section \ref{sec:generalsetup} formalizes the Bayesian inference framework, Section \ref{sec:GPimplementation} puts forward a flexible class of priors for $p$ and illustrates their use in an application to state dependence, Section \ref{sec:posteriorconsistency} presents our posterior consistency results, Section \ref{sec:extensions} offers several extensions of the framework, and Section \ref{sec:conclusion} concludes. Notation is introduced when appropriate and proofs are in the Appendix.
\section{Bayesian Inference Framework}\label{sec:generalsetup}
This section formalizes how Bayesian inferences about identified sets of the form $\Gamma_{I} = \{\gamma: f(X,p(X),\gamma) \geq 0 \ a.s.\}$ can be obtained via nonparametric Bayesian inference for the unrestricted reduced-form conditional PMF $p$ of $Y$ given $X$.
\subsection{Nonparametric Bayesian Inference for PMFs.} Let $\{(Y_{i},X_{i}')'\}_{i\geq 1}$, where $Y_{i} \in \mathcal{Y}$, $\mathcal{Y} = \{y_{1},...,y_{K}\}$, and $X_{i} \in \mathcal{X} \subseteq \mathbb{R}^{d_{x}}$ for each $i$, be a sequence of random vectors for which $\{(Y_{i},X_{i}')'\}_{i=1}^{n}$, $n \geq 1$, forms the observed data. There are no restrictions on $\mathcal{X}$ (i.e., $X_{i}$ can have continuous components). Our sampling model conditions on the realization $\{x_{i}\}_{i \geq 1}$ of $\{X_{i}\}_{i \geq 1}$ and imposes that $Y^{(n)} = (Y_{1},...,Y_{n})$ satisfies
\begin{align*}
Y^{(n)}|p \sim P^{(n)}, \quad P^{(n)} = \bigotimes_{i=1}^{n}Categorical(p(x_{i})),
\end{align*}
where $Categorical(p(x))$ denotes a discrete distribution over $\{y_{1},...,y_{K}\}$ with probability mass function $p(x) = (p_{1}(x),...,p_{K}(x))$ for $x \in \mathcal{X}$. For Bayesian inference, we endow the infinite-dimensional space $\mathcal{P}$ of PMFs $p$ with a prior $\Pi$, a conditional (on $\{X_{i}\}_{i \geq 1}$) probability measure over $(\mathcal{P},\mathscr{P})$, where $\mathscr{P}$ is a $\sigma$-algebra over $\mathcal{P}$. Section \ref{sec:GPimplementation} presents an example.
\begin{assumption}\label{as:samplingmodel}
For each $n \geq 1$ and almost every fixed realization $\{x_{i}\}_{i \geq 1}$ of $\{X_{i}\}_{i \geq 1}$, $(y^{(n)},p) \mapsto L_{n}(p) :=\prod_{i=1}^{n}\prod_{k=1}^{K}p_{k}(x_{i})^{\mathbf{1}\{y_{i}=y_{k}\}}$ is measurable on $(\mathcal{Y}^{n} \times \mathcal{P}, 2^{\mathcal{Y}^{n}}\otimes\mathscr{P})$.
\end{assumption}
\begin{assumption}\label{as:marginallikelihood}
The following conditions hold:
\begin{enumerate}
\item For each $n \geq 1$ and almost every fixed realization $\{x_{i}\}_{i\geq 1}$ of $\{X_{i}\}_{i \geq 1}$, the marginal distribution $\Pi_{n}$ of $(p(x_{1}),...,p(x_{n}))$ under $\Pi$ admits a density $\pi_{n}$.
\item Let $p_{0}$ denote the true conditional PMF. For each $n \geq 1$ and almost every fixed realization $\{x_{i}\}_{i \geq 1}$ of $\{X_{i}\}_{i \geq 1}$, $\int_{\mathcal{P}}L_{n}(p)\pi_{n}(p)dp \in (0,\infty)$ a.e. $[P_{0}^{(n)}]$.
\end{enumerate}
\end{assumption}
Assumptions \ref{as:samplingmodel} and \ref{as:marginallikelihood} are standard regularity conditions \citep{ghosal2017fundamentals}, and imply that the vector of in-sample PMFs $(p(x_{1}),...,p(x_{n}))$ has a well-defined posterior distribution $\Pi_{n}((p(x_{1}),...,p(x_{n})) \in \cdot | Y^{(n)})$ with density that satisfies
\begin{align*}
\pi_{n}(p(x_{1}),...,p(x_{n})|Y^{(n)}) \propto L_{n}(p)\pi_{n}(p(x_{1}),...,p(x_{n})).
\end{align*}
for each $p \in \mathcal{P}$. The posterior contains a full inference theory for the vector $(p(x_{1}),...,p(x_{n}))$ and any transformation of it. Consequently, by treating the identified set as a function of this vector, we obtain a Bayesian inference theory for the identified set.
\subsection{Bayesian Inference for Identified Sets.} We now show how $\pi_{n}(\cdot|Y^{(n)})$ leads to a Bayesian inference framework for an identified set. Let $\Gamma \subseteq \mathbb{R}^{d_{\gamma}}$, $d_{\gamma} < \infty$, be the structural parameter space, and, for a fixed realization $\{x_{i}\}_{i \geq 1}$ of $\{X_{i}\}_{i \geq 1}$, define the \textit{conditional identified set} $\Gamma_{n,I}(p)$ at $p$ as
\begin{align*}
\Gamma_{n,I}(p) = \{\gamma \in \Gamma: f(x_{i},p(x_{i}),\gamma) \in \mathbb{R}_{+}^{d_{f}} \ \forall \ i=1,...,n\},
\end{align*}
where $f: \mathcal{X} \times \Delta^{K-1} \times \Gamma \rightarrow \mathbb{R}^{d_{f}}$, $d_{f} < \infty$, is a known function.\footnote{Theorem \ref{thm:fullsupport} resolves the apparent discrepancy between $\Gamma_{n,I}$ and $\Gamma_{I}$.} The conditional identified set can be equivalently expressed as the level set $\Gamma_{n,I}(p) = \{\gamma \in \Gamma: Q_{n}(\gamma,p) = 0\}$, where $Q_{n}: \Gamma \times \mathcal{P} \rightarrow \mathbb{R}_{+}$ is any criterion function that only depends on $p$ through $(p(x_{1}),...,p(x_{n}))$, and, importantly, satisfies the property\footnote{Section \ref{sec:examples} demonstrates $Q_{n}$ may depend on $\{x_{i}\}_{i=1}^{n}$, but we suppress for ease of exposition.}
\begin{align*}
Q_{n}(\gamma,p) = 0 \iff f(x_{i},p(x_{i}),\gamma)\in \mathbb{R}_{+}^{d_{f}} \ \forall \ i=1,...,n .
\end{align*}
Crucially, since $Q_{n}$ only depends on the in-sample response probabilities $(p(x_{1}),...,p(x_{n}))$, posterior probability statements about $(p(x_{1}),...,p(x_{n}))$ translate to statements about $\Gamma_{n,I}(p)$, provided that the correspondence $\Gamma_{n,I}: \mathcal{P} \rightarrow 2^{\Gamma}$ is suitably measurable.
For this purpose, we maintain that $\Gamma_{n,I}: \mathcal{P} \rightarrow 2^{\Gamma}$ forms a random closed set \citep{molchanov2017theory,molchanov2018random}. That is, $\Gamma_{n,I}(p)$ is a closed set for each $p$ and $\Gamma_{n,I}^{-}(A) := \{p: \Gamma_{n,I}(p) \cap A \neq \emptyset\} \in \mathscr{P}$ for each compact subset $A$ of $\Gamma$. In such case, the probability law of $\Gamma_{n,I}|Y^{(n)}$ is uniquely determined by the \textit{posterior capacity functional} $T_{\Gamma_{n,I}|Y^{(n)}}$, which satisfies $T_{\Gamma_{n,I}|Y^{(n)}}(A) = \Pi_{n}(\Gamma_{n,I}(p) \cap A \neq \emptyset | Y^{(n)})$ for $A \subseteq \Gamma$ compact. The posterior capacity functional formalizes $\Gamma_{I}|Y,X$. Assumptions \ref{as:parameterspace} and \ref{as:criterion}, both standard, are sufficient for this characterization (with Theorem \ref{thm:randomset} serving as validation).\footnote{Two comments. First, compactness of $\Gamma$ is not used in the proof of Theorem \ref{thm:randomset}, however, for the application of random set theory, the ambient space (i.e., $\Gamma$) is typically restricted to be locally compact second countable Hausdorff (LCSCH) (see \cite{molchanov2017theory,molchanov2018random}). Assumption \ref{as:parameterspace} is sufficient for this and will be used in the subsequent results of the paper, so we impose it outright but explicitly acknowledge that it is not necessary for Theorem \ref{thm:randomset} (i.e, we can impose that $\Gamma$ is LCSCH only). Second, if Assumption \ref{as:criterion}.1 is strengthened to continuity and Assumption \ref{as:parameterspace} holds, then Assumption \ref{as:criterion}.2 holds if $Q_{n}(\gamma,\cdot): \mathcal{P} \rightarrow \mathbb{R}_{+}$ is $\mathscr{P}$-measurable for each $\gamma \in \Gamma$ (see, for example, Theorem 17.18 in \cite{aliprantis1999infinite}). For sufficient conditions based on lower semicontinuity, see \cite{stinchcombe1992some}.}
\begin{assumption}\label{as:parameterspace}
$\Gamma$ is a compact subset of $(\mathbb{R}^{d_{\gamma}},||\cdot||_{2})$, where $||\cdot||_{2}$ is the Euclidean norm.
\end{assumption}
\begin{assumption}\label{as:criterion}
For each $n \geq 1$ and almost every fixed realization $\{x_{i}\}_{i \geq 1}$ of $\{X_{i}\}_{i \geq 1}$,
\begin{enumerate}
\item The mapping $\gamma \mapsto Q_{n}(\gamma,p)$ is lower semicontinuous for each $p \in \mathcal{P}$.
\item The mapping $p \mapsto \inf_{\gamma \in A}Q_{n}(\gamma,p)$ is measurable function from $(\mathcal{P},\mathscr{P})$ to $(\mathbb{R}_{+},\mathcal{B}(\mathbb{R}_{+}))$ for each compact $A \subseteq \Gamma$.\footnote{The notation $\mathcal{B}(A)$ denotes the Borel $\sigma$-algebra generated by the open subsets of $A$.}
\end{enumerate}
\end{assumption}
\begin{theorem}\label{thm:randomset}
Suppose that Assumptions \ref{as:parameterspace} and \ref{as:criterion} hold. Then, for each $n \geq 1$ and almost every fixed realization $\{x_{i}\}_{i \geq 1}$ of $\{X_{i}\}_{i \geq 1}$, $\Gamma_{n,I}$ defines a random closed set over $(\mathcal{P},\mathscr{P})$.
\end{theorem}
There are several aspects of our framework that should be emphasized. First, the posterior for the identified set implies a marginal posterior for the identified set of functions of the structural parameter. Let $g: \mathbb{R}^{d_{\gamma}} \rightarrow \mathbb{R}^{d_{h}}$, $d_{g} < \infty$, be a known continuous transformation. The conditional identified set for $g(\gamma)$ is $G_{n,I}(p) = \{g(\gamma): \gamma \in \Gamma_{n,I}(p)\}$. Corollary \ref{cor:functionrandomset} shows that $G_{n,I}: \mathcal{P} \rightarrow 2^{g(\Gamma)}$ forms a random closed set. Consequently, $\pi_{n}(\cdot|Y^{(n)})$ leads to a posterior for $G_{n,I}(p)$ that is similarly described its capacity functional. Defining $g$ as the coordinate projection, this nests the identified set for subvectors of $\gamma$.
\begin{corollary}\label{cor:functionrandomset}
Suppose Assumptions \ref{as:parameterspace} and \ref{as:criterion} hold, and $g: \Gamma \rightarrow \mathbb{R}^{d_{g}}$ is continuous. Then, for each $n \geq 1$ and almost every fixed realization $\{x_{i}\}_{i \geq 1}$ of $\{X_{i}\}_{i \geq 1}$, $G_{n,I}$ defines a random closed set over $(\mathcal{P},\mathscr{P})$.
\end{corollary}
Second, our framework also accommodates model misspecification. A partially identified model is misspecified at $p \in \mathcal{P}$ if $\Gamma_{n,I}(p) = \emptyset$. Setting $A = \Gamma$, $\{p: \Gamma_{n,I}(p)= \emptyset\} $ is $\mathscr{P}$-measurable and $1-T_{\Gamma_{n,I}|Y^{(n)}}(\Gamma)$ is a misspecification diagnostic because a high posterior probability of an empty identified set provides strong evidence that the identifying restrictions are not compatible with the data. Alternatively, one can report the posterior for a pseudo-identified set $\tilde{\Gamma}_{n,I}(p) = \{\gamma \in \Gamma: \tilde{Q}_{n}(\gamma,p) = 0\}$, where $\tilde{Q}_{n}(\gamma,p) = Q_{n}(\gamma,p) - \inf_{\tilde{\gamma} \in \Gamma}Q_{n}(\tilde{\gamma},p)$ (i.e., $\tilde{\Gamma}_{n,I}(p)$ is the set of minimizers of $Q_{n}(\cdot,p)$). Like $\Gamma_{n,I}(p)$, $\tilde{\Gamma}_{n,I}(p)$ is a set-valued functional of $(p(x_{1}),...,p(x_{n}))$ so it has a posterior under $\pi_{n}(\cdot|Y^{(n)})$ as long as it is a random closed set.\footnote{There are, of course, conceptual issues associated with reporting pseudo-true parameters \citep{white1982maximum,muller2013risk,hansen2021inference,andrews2026true}. Our paper does not take a stance on whether one \textit{should} report $\tilde{\Gamma}_{n,I}(p)$, but simply provides it as an option for those concerned about misspecification at $p$.} Similarly, a valid inference framework for functions of $\gamma$ is obtained via the posterior for $\tilde{G}_{n,I}(p) = \{g(\gamma): \gamma \in \tilde{\Gamma}_{n,I}(p)\}$.
\begin{theorem}\label{thm:pseudoset}
Suppose that Assumptions \ref{as:parameterspace} and \ref{as:criterion} hold. Then, for each $n \geq 1$ and almost every fixed realization $\{x_{i}\}_{i \geq 1}$ of $\{X_{i}\}_{i \geq 1}$, $\tilde{\Gamma}_{n,I}$ defines a random closed set over $(\mathcal{P},\mathscr{P})$.
\end{theorem}
Finally, most literature on partial identification with covariates assumes that $\{(Y_{i},X_{i}')'\}_{i \geq 1}$ forms an independent and identically distributed (i.i.d) sequence, and defines the identified set at $p$ as $\Gamma_{I}(p) = \{\gamma \in \Gamma: f(x,p(x),\gamma) \geq 0 \ \forall \ x \in \mathcal{S}_{X}\}$ to define the identified set at $p$, where $\mathcal{S}_{X}$ is the support of the marginal distribution $P_{0,X}$ of $X$ (see, for example, \cite{Ho_Rosen_2017}, \cite{MOLINARI2020355}, \cite{KLINE2021345}, and \cite{CANAY2023105558}). This creates an apparent discrepancy between our target parameter $\Gamma_{n,I}(p)$ and the focus of much of the partial identification literature (an exception being \cite{rosen2025finite}, who define the identified set for the \cite{manski1975maximum,manski1985semiparametric} maximum score model conditionally). Fortunately, the next result establishes that, under a well-separatedness condition for $\Gamma_{I}(p)$ and other standard regularity conditions, there is virtually no difference between $\Gamma_{n,I}(p)$ and $\Gamma_{I}(p)$ for large $n$. For notation, $f_{j}(x,p(x),\gamma)$ is the $j$th element of $f(x,p(x),\gamma)$, $\mathscr{X}$ is the $\sigma$-algebra over $\mathcal{X}$ for which $P_{0,X}$ is defined, $\mathscr{X}^{\infty}$ is the associated product $\sigma$-algebra, and $d_{\mathcal{H}}$ is the Hausdorff distance.\footnote{Given two sets $\Gamma_{1},\Gamma_{2}$ in $(\mathbb{R}^{d_{\gamma}},||\cdot||_{2})$, the Hausdorff distance is defined as $d_{\mathcal{H}}(\Gamma_{1},\Gamma_{2}) = \max\{\sup_{\gamma \in \Gamma_{1}}\inf_{\tilde{\gamma} \in \Gamma_{2}}||\gamma-\tilde{\gamma}||_{2},\sup_{\gamma \in \Gamma_{2}}\inf_{\tilde{\gamma} \in \Gamma_{1}}||\gamma-\tilde{\gamma}||_{2}\}$.}
\begin{theorem}\label{thm:fullsupport}
Suppose that the following hold for each $p \in \mathcal{P}$: 1. $\gamma \mapsto Q_{n}(\gamma,p)$ is lower semicontinuous for each $\{x_{i}\}_{i \geq 1} \in \mathcal{X}^{\infty}$ and each $n \geq 1$, 2. $\{x_{i}\}_{i \geq 1} \mapsto \inf_{\gamma \in A}Q_{n}(\gamma,p)$ is a measurable function from $(\mathcal{X}^{\infty},\mathscr{X}^{\infty})$ to $(\mathbb{R}_{+},\mathcal{B}(\mathbb{R}_{+}))$ for each compact $A \subseteq \Gamma$ and each $n \geq 1$, 3. $\Gamma$ is a compact subset of $(\mathbb{R}^{d_{\gamma}},||\cdot||_{2})$, 4. the sequence $\{X_{i}\}_{i\geq 1}$ is i.i.d with common distribution $P_{0,X}$, 5. the mapping $x \mapsto \sup_{\gamma \in A}\min_{1 \leq j \leq d_{f}}f_{j}(x,p(x),\gamma) $ is a measurable function from $(\mathcal{X},\mathscr{X})$ to $(\mathbb{R},\mathcal{B}(\mathbb{R}))$ for every compact $A \subseteq \Gamma$, 6. $\Gamma_{I}(p) \neq \emptyset$, and 7. for each compact $A\subseteq \Gamma \setminus \Gamma_{I}(p)$, $P_{0,X}(\sup_{\gamma \in A}\min_{1 \leq j \leq d_{f}}f_{j}(X,p(X),\gamma) < 0)>0$. Then for each $p \in \mathcal{P}$: $\lim_{n\rightarrow \infty}d_{\mathcal{H}}\left(\Gamma_{n,I}(p),\Gamma_{I}(p)\right) = 0$ a.e. $[P_{0,X}^{(\infty)}]$, where $P_{0,X}^{(\infty)}= \bigotimes_{i \geq 1}P_{0,X}$.
\end{theorem}
\subsection{Examples.}\label{sec:examples}
This section revisits Examples \ref{ex:momentsineq} and \ref{ex:linsyst}. Let $||\cdot||_{n,r}$ satisfy $||h||_{n,r}= (n^{-1}\sum_{i=1}^{n}|h(x_{i})|^{r})^{1/r}$ for $r \in [1,\infty)$ and $||h||_{n,\infty} = \max_{1 \leq i \leq n}|h(x_{i})|$ for real-valued $h$ (replacing $|h(x)|$ with $||h(x)||_{2}$ and $||h(x)||_{op}$ for vector valued and matrix valued $h$, respectively, where $||\cdot||_{op}$ is the operator norm). Let $\text{dist}_{W}(\cdot,A)$ satisfy $\text{dist}_{W}(t,A) = \inf_{\tilde{t} \in A}||t-\tilde{t}||_{W}$ for each $t$, where $||t||_{W}^{2} = t'Wt$ for some symmetric, positive-definite matrix $W$.
\setcounter{example}{0}
\begin{example}[Continued]
Recall that, for CMIs, $f(x_{i},p(x_{i}),\gamma) = E[g(Y,X,\gamma)|X=x_{i}]$, where $g(Y,X,\gamma)$ be a $d_{g} \times 1$ vector of known functions and the expectation is taken with respect to $p(x_{i})$. A class of criterion functions for CMIs are of the mimimum distance form,
\begin{align*}
Q_{n,r}(\gamma,p) = \left|\left|\text{dist}_{W(\cdot,p(\cdot),\gamma)}(f(\cdot,p(\cdot),\gamma),\mathbb{R}_{+}^{d_{g}})\right| \right|_{n,r},
\end{align*} where $W(x_{i},p(x_{i}),\gamma)$ is some symmetric, positive-definite matrix and $r \in [1,\infty]$. Since $\{Q_{n}(\gamma,p): \gamma \in \Gamma\}$ is fully determined by $(p(x_{1}),...,p(x_{n}))$, the posterior $\pi_{n}(\cdot|Y^{(n)})$ leads to a posterior over the values of $Q_{n}(\gamma,p)$, $\gamma \in \Gamma$, and, by extension, the set of minimizers. These criterions are common in the moment inequality literature (see \cite{canay2017practical} and the references therein). Indeed, setting $r = 2$ and $W(x_{i},p(x_{i}),\gamma) = \Sigma^{-1}(x_{i},p(x_{i}),\gamma)$, where $\Sigma(x_{i},p(x_{i}),\gamma) = Var(g(Y,X,\gamma)|X=x_{i})$, one obtains a quasi-likelihood ratio (QLR) criterion,
\begin{align*}
&Q_{n,QLR}(\gamma,p)
=\sqrt{\frac{1}{n}\sum_{i=1}^{n}\min_{t \in \mathbb{R}^{d_{g}}_{+}}\left\{(f(x_{i},p(x_{i}),\gamma)-t)'\Sigma^{-1}(x_{i},\gamma)(f(x_{i},p(x_{i}),\gamma)-t)\right\}}
\end{align*}while setting $r = 2$ and $W(x_{i},p(x_{i}),\gamma) = D^{-1}(x_{i},p(x_{i}),\gamma)$, with $D(x_{i},p(x_{i}),\gamma) = \text{diag}\Sigma(x_{i},p(x_{i}),\gamma)$, yields a modified method of moments (MMM) criterion
\begin{align*}
Q_{n,MMM}(\gamma,p) = \sqrt{\frac{1}{n}\sum_{i=1}^{n}||(D^{-1/2}(x_{i},p(x_{i})\gamma)f(x_{i},p(x_{i}),\gamma))_{-}||_{2}^{2}},
\end{align*}
where $(t)_{-} = (\min\{t_{1},0\},...,\min\{t_{d},0\})'$ for $t \in \mathbb{R}^{d}$. In specific relation to the frequentist conditional moment inequality literature (e.g., \cite{andrews2013inference} and \cite{armstrong2014weighted}), $Q_{n,r}(\gamma,p)$ with $r \in [1,\infty)$ and $r= \infty$ can be viewed as Cr\'{a}mer-von Mises-type (CvM) and Kolmogorov-Smirnov-type (KS) statistics, respectively. A key distinction is that we aggregate over the observed covariates $\{x_{i}\}_{i=1}^{n}$, whereas the papers listed above aggregate over a user-specified class of implied unconditional moment restrictions. Consequently, applied researchers willing to adopt a Bayesian perspective can avoid selecting unconditional moments.
\end{example}
\begin{example}[Continued]\label{ex:linsyst}
Recall that, for linear systems, $f(x_{i},p(x_{i}),\gamma) = A(x_{i},p(x_{i}))\gamma - b(x_{i},p(x_{i}))$, where $A(x_{i},p(x_{i})) \in \mathbb{R}^{d_{a}\times d_{\gamma}}$ and $b(x_{i},p(x_{i})) \in \mathbb{R}^{d_{a}}$. Like Example \ref{ex:momentsineq}, a possible criterion function is the minimum distance objective
\begin{align*}
Q_{n,r}(\gamma,p) = ||\text{dist}_{W(\cdot,p(\cdot),\gamma)}(A(\cdot,p(\cdot))\gamma-b(\cdot,p(\cdot)),\mathbb{R}_{+}^{d_{a}})||_{n,r},
\end{align*}
where $W(x_{i},p(x_{i}),\gamma)$ is a symmetric, positive-definite weight matrix and $r \in [1,\infty]$. Alternatively, one could report
\begin{align*}
Q_{n}^{LP}(\gamma,p) = \frac{1}{n}\sum_{i=1}^{n}||(A(x_{i},p(x_{i}))\gamma-b(x_{i},p(x_{i})))_{-}||_{1},
\end{align*}
where $||\cdot||_{1}$ is the taxicab norm.\footnote{The taxicab norm is $||z||_{1} = \sum_{j=1}^{d_{z}}|z_{j}|$.}$^{,}$\footnote{This statistic is not specific to linear systems because, for any candidate $f$, one can define $Q_{n}^{LP}(\gamma,p) = n^{-1}\sum_{i=1}^{n}||(f(x_{i},p(x_{i}),\gamma))_{-}||_{1}$.} This criterion function is particularly compelling for linear systems because if $\Gamma$ is a convex polytope (i.e., $\Gamma=\{\gamma\in\mathbb R^{d_\gamma}:A_{\Gamma}\gamma\leq b_{\Gamma}\}$ for some matrix $A_{\Gamma}$ and vector $b_{\Gamma}$, and is bounded), then the set of minimizers of $Q_n^{LP}(\cdot,p)$ is the projection onto the $\gamma$-coordinates of the optimal-solution set of the following linear program:\footnote{If $\Gamma$ is merely compact (with a relevant distinguishing case being the spherical normalization $\Gamma =\{\gamma \in \mathbb{R}^{d_{\gamma}}: ||\gamma||_{2}= 1\}$), then the same optimization problem can be written with the constraint $\gamma\in\Gamma$, although it is not necessarily a linear program.}
\begin{align*}
\min_{\gamma,s_1,\cdots,s_n} & \frac{1}{n}\sum_{i=1}^n\mathbf 1_{d_a}'s_i\\ \text{s.t.}\quad& A(x_i,p(x_i))\gamma+s_i-b(x_i,p(x_i))\geq 0, \ i=1,\ldots,n,\\
&A_{\Gamma}\gamma\leq b_{\Gamma},\quad s_i\geq 0,\ i=1,\ldots,n . \end{align*}
A similar story to Example \ref{ex:momentsineq} applies: $(p(x_{1}),...,p(x_{n}))$ determines $Q_{n}^{LP}(\cdot,p)$ and $Q_{n,r}(\cdot,p)$, meaning that a posterior $\pi_{n}(\cdot|Y^{(n)})$ leads to a posterior for the set of minimizers. Importantly, our approach does not require discretizing the covariates.
\end{example}
\section{Implementation with Gaussian Processes.}\label{sec:GPimplementation}
This section describes implementation for priors based on Gaussian processes (GPs), and presents a Monte Carlo simulation.
\subsection{A Flexible Class of Priors.}\label{sec:GPimplementationdetails}
Our implementation is based on the `stick-breaking' parametrization of a categorical distribution from \cite*{linderman2015dependent}. Define the stick-breaking parametrization of the PMF $p$ as follows,
\begin{align}\label{eq:stickbreak}
p_{1} = q_{1}, \quad p_{k} = q_{k}\prod_{j = 1}^{k-1}(1-q_{j}), \ j=2,...,K-1, \quad p_{K} = \prod_{j=1}^{K-1}(1-q_{j}),
\end{align}
where $q_{k}: \mathcal{X} \rightarrow [0,1]$ for each $k=1,...,K-1$.\footnote{For any discrete distribution $p$, there exists a well-defined stick-breaking representation because $q_{k}= p_{k}/\sum_{j \geq k}p_{j}$ for all $k=1,...,K-1$, so this parametrization does not entail any loss of generality.} By substituting the stick-breaking weights, the likelihood function
\begin{align*}
L_{n}(p) = \prod_{i=1}^{n}\prod_{k=1}^{K}p_{k}(x_{i})^{\mathbf{1}\{Y_{i}=y_{k}\}}
\end{align*}
satisfies $L_{n}(p) = L_{n}^{*}(q)$, where\footnote{$L_{n}^{*}(q)$ presumes the support of $Y$ has been enumerated (with $y_{k}$ being an enumerated support point) so that the event $Y > y_{k}$ has meaning.}
\begin{align*}
L_{n}^{*}(q) = \prod_{i=1}^{n}\prod_{k=1}^{K-1}q_{k}(x_{i})^{\mathbf{1}\{Y_{i}=y_{k}\}}(1-q_{k}(x_{i}))^{\mathbf{1}\{Y_{i} > y_{k}\}}.
\end{align*}
Consequently, by specifying a prior for $q := (q_{1},...,q_{K-1})$ and computing the posterior for $(q(x_{1}),...,q(x_{n}))$ using $L_{n}^{*}(q)$, Bayesian inference for $(p(x_{1}),...,p(x_{n}))$ is obtained via the induced probability distribution for $p$ under the stick-breaking parametrization (\ref{eq:stickbreak}).
Since the stick-breaking functions $q$ \textit{only} need to satisfy $0 \leq q_{k} \leq 1$ for each $k=1,...,K-1$, we specify the prior for $q$ as the probability law of $(\Lambda(B_{1}),...,\Lambda(B_{K-1}))$, where $\Lambda(z) = \exp(z)/(1+\exp(z))$ for each $z \in \mathbb{R}$, and
\begin{align*}
B_{k}\overset{ind}{\sim} GP(\mu_{k},\kappa_{k}), \quad k=1,...,K-1,
\end{align*}
with $GP(\mu_{k},\kappa_{k})$ denoting a Gaussian process with mean function $\mu_{k}: \mathcal{X} \rightarrow \mathbb{R}$ and covariance function $\kappa_{k}: \mathcal{X} \times \mathcal{X} \rightarrow \mathbb{R}$.\footnote{The parametrization $\Lambda(B_{k})$ is fully nonparametric because $B_{k} = \Lambda^{-1}(q_{k})$.} Following similar constructions in the conditional density estimation literature (e.g., \cite{ren2011logistic}), we call this the logistic stick-breaking Gaussian process prior, and use $LSBGP(\{\mu_{k}\}_{k=1}^{K-1},\{\kappa_{k}\}_{k=1}^{K-1})$ to denote this class of priors. The specification of $\kappa_{k}$ allows for flexible incorporation of nonparametric properties like smoothness, while the choice of $\mu_{k}$ enables shrinkage towards a particular parametric model for $q$, such as a linear index model \citep{williams2006gaussian}. Moreover, the prior is compatible with our general assumptions. Indeed, under the GP, the vector $(B_{n,1},...,B_{n,K-1})$, with $B_{n,k} :=(B_{k}(x_{i}))_{i=1}^{n}$, satisfies
\begin{align*}
B_{n,k} \overset{ind}{\sim}\mathcal{N}(\mu_{n,k},\kappa_{n,k}), \quad k=1,...,K-1
\end{align*}
where $\mu_{n,k}:=(\mu_{k}(x_{i}))_{i=1}^{n}$ and $\kappa_{n,k}:=(\kappa_{k}(x_{i},x_{j}))_{i,j=1}^{n}$ for $k=1,...,K-1$. This implies that Assumption \ref{as:marginallikelihood}.1 holds because $B \mapsto p$ has a differentiable inverse. In practice, we set $\mu_{k} = 0$ for each $k=1,...,K-1$, so prior elicitation amounts to specifying $\{\kappa_{k}\}_{k=1}^{K-1}$.
\subsection{Posterior Sampling}\label{sec:samplerdescribe} We describe sampling from the posterior for $(B_{n,1},...,B_{n,K-1})$, the quantities relevant for computing $(p(x_{1}),...,p(x_{n}))$. Under the independent GP prior, the posterior density for $(B_{n,1},...,B_{n,K-1})$ satisifes
\begin{align*}
\pi_{n}(B_{n,1},...,B_{n,K-1}|Y^{(n)}) &\propto \prod_{k=1}^{K-1}\prod_{i \in \mathcal{I}_{k}}\Lambda(B_{k}(x_{i}))^{\mathbf{1}\{Y_{i}=y_{k}\}}(1-\Lambda(B_{k}(x_{i})))^{\mathbf{1}\{Y_{i} > y_{k}\}} \\
&\quad \times \prod_{k=1}^{K-1}\pi_{n,k}(B_{n,k}) \\
&\propto \prod_{k=1}^{K-1}\pi_{n,k}(B_{n,k}|Y^{(n)}),
\end{align*}
where $\pi_{n,k}$ is the density of $\mathcal{N}(\mu_{n,k},\kappa_{n,k})$, $\mathcal{I}_{k} = \{i:Y_{i} \geq y_{k}\}$, and
\begin{align}\label{eq:binarylogit}
\pi_{n,k}(B_{n,k}|Y^{(n)}) \propto \prod_{i \in \mathcal{I}_{k}}\Lambda(B_{k}(x_{i}))^{\mathbf{1}\{Y_{i}=y_{k}\}}(1-\Lambda(B_{k}(x_{i})))^{\mathbf{1}\{Y_{i} > y_{k}\}}\pi_{n,k}(B_{n,k}).
\end{align}
This means that the blocks $B_{n,1}$,...,$B_{n,K-1}$ are independent under the posterior. Consequently, a draw from the posterior for $(p(x_{1}),...,p(x_{n}))$ is obtained as follows:
\begin{enumerate}
\item Draw $B_{n,k} \overset{ind}{\sim} \pi_{n,k}(\cdot|Y^{(n)})$ for $k=1,...,K-1$.
\item Compute $p(x_{i})$ for $i=1,...,n$ using (\ref{eq:stickbreak}).
\end{enumerate}
Given a draw $(p(x_{1}),...,p(x_{n}))$, a draw from $\Gamma_{n,I}|Y^{(n)}$ is acquired by computing $\Gamma_{n,I}(p)$ using \textit{any} appropriate method (and, as a post-processing exercise, can be made parallel). Step 1 can also be parallelized across $k$, which is an attractive feature for settings with large $K$.
We briefly outline sampling from $\pi_{n,k}(\cdot|Y^{(n)})$; Appendix \ref{ap:posteriordraws} contains the precise details. For a given $k$, let $B_{\mathcal{I}_{k},k}=(B_{k}(x_{i}))_{i \in \mathcal{I}_{k}}$ and $B_{\mathcal{I}_{k}^{c},k} = (B_{k}(x_{i}))_{i \in \mathcal{I}_{k}^{c}}$. Partitioning $B_{n,k} $ into $(B_{\mathcal{I}_{k},k},B_{\mathcal{I}_{k}^{c},k})$ and factorizing $\pi_{n,k}(B_{n,k}|Y^{(n)})$ as
\begin{align*}
\pi_{n,k}\left(B_{n,k}\middle |Y^{(n)}\right) = \pi_{n,k}\left(B_{\mathcal{I}_{k}^{c},k}\middle |Y^{(n)},B_{\mathcal{I}_{k},k}\right) \pi_{n,k}\left(B_{\mathcal{I}_{k},k}\middle | Y^{(n)}\right),
\end{align*}
we sample from $B_{n,k}|Y^{(n)}$ by first sampling from $\pi_{n,k}(B_{\mathcal{I}_{k},k}|Y^{(n)})$ and then sampling from $\pi_{n,k}(B_{\mathcal{I}_{k}^{c},k}|Y^{(n)},B_{\mathcal{I}_{k},k})$. Expression (\ref{eq:binarylogit}) indicates that $\pi_{n,k}(B_{\mathcal{I}_{k},k}|Y^{(n)})$ is the posterior of a nonparametric binary logit model for the subsample $(Y_{i})_{i \in \mathcal{I}_{k}}$. Consequently, P\'{o}lya-Gamma data augmentation \citep{polson2013bayesian} leads to a Gibbs sampler with \textit{closed-form} sweeps. For $\pi_{n}(B_{\mathcal{I}_{k}^{c},k}|Y^{(n)},B_{\mathcal{I}_{k},k})$, the conditional likelihood in (\ref{eq:binarylogit}) is a constant function of $B_{\mathcal{I}_{k}^{c},k}$, which implies that $\pi_{n,k}(B_{\mathcal{I}_{k}^{c},k}|Y^{(n)},B_{\mathcal{I}_{k},k})=\pi_{n}(B_{\mathcal{I}_{k}^{c},k}|B_{\mathcal{I}_{k},k})$. Since $B_{n,k} \sim \mathcal{N}(\mu_{n,k},\kappa_{n,k})$, properties of multivariate normal distributions implies that sampling from $\pi_{n,k}(B_{\mathcal{I}_{k}^{c},k}|Y^{(n)},B_{\mathcal{I}_{k},k})$ simply amounts to generating draws from a Gaussian distribution.
\begin{remark}[Prior Hyperparameters]
The mean and covariance functions often depend on hyperparameters. That is, $\mu_{k}(x)=\mu_{k,\tau_{k,1}}(x)$ and $\kappa_{k}(x,\tilde{x}) = \kappa_{k,\tau_{k,2}}(x,\tilde{x})$ for some $\tau_{k}:=(\tau_{k,1}',\tau_{k,2}')' \in \mathbb{R}^{d_{\tau_{k}}}$. Factorizing the posterior into the product
\begin{align*}
\pi_{n,k}(B_{n,k},\tau_{k}|Y^{(n)}) = \pi_{n,k}(B_{\mathcal{I}_{k}^{c},k}|Y^{(n)},B_{\mathcal{I}_{k},k},\tau_{k})\pi_{n,k}(B_{\mathcal{I}_{k},k},\tau_{k}|Y^{(n)}),
\end{align*}
accommodating hyperparameter selection requires sampling from $\pi_{n,k}(B_{\mathcal{I}_{k},k},\tau_{k}|Y^{(n)})$ and then $\pi_{n,k}(B_{\mathcal{I}_{k}^{c},k}|Y^{(n)},B_{\mathcal{I}_{k},k},\tau_{k})$. The former requires a minor modification to the Gibbs sampler that, importantly, preserves the tractability of P\'{o}lya-Gamma data augmentation (see Appendix \ref{ap:hyperparameters}). Generating draws from $\pi_{n,k}(B_{\mathcal{I}_{k}^{c},k}|Y^{(n)},B_{\mathcal{I}_{k},k},\tau_{k})$ is identical to $\pi_{n,k}(B_{\mathcal{I}_{k}^{c},k}|Y^{(n)},B_{\mathcal{I}_{k},k})$, except that $\mu_{n,k}$ and $\kappa_{n,k}$ are adjusted to reflect the sampled $\tau_{k}$.
\end{remark}
\begin{remark}[Discrete Covariates]\label{remark:discrete}
Suppose that $x_{i} = (x_{i,c}',x_{i,d}')'$, where $x_{i,d} \in \{x_{1},...,x_{G}\}$ with $G < \infty$. In such case, $q_{k}(x_{i}) = \prod_{g=1}^{G}q_{k,g}(x_{i,c})^{\mathbf{1}\{x_{i,d}=x_{g}\}}$ for each $k=1,...,K-1$, and, $L_{n}^{*}(q) =\prod_{i=1}^{n}\prod_{k=1}^{K-1}\prod_{g=1}^{G}q_{k,g}(x_{i,c})^{\mathbf{1}\{x_{i,d}=g,Y_{i}= k\}}(1-q_{k,g}(x_{i,c})^{\mathbf{1}\{x_{i,d}=g,Y_{i} > k\}}$. Consequently, by writing $q_{k,g}(x_{i,c}) = \Lambda(B_{k,g}(x_{i,c}))$ and assuming that $B_{k,g} \overset{ind}{\sim} GP(\mu_{k,g},\kappa_{k,g})$ for each $k,g$, we can apply the same algorithm as above, except now with an additional factorization across discrete covariate values $\{x_{g}\}_{g=1}^{G}$.
\end{remark}
\subsection{Simulation.}
We demonstrate the logistic stick-breaking process in a simulated dynamic panel binary choice model.
\subsubsection{Data-Generating Process and Identified Sets.} The data-generating process imposes that a binary outcome $Y_{t}$ at time $t$ is related to a contemporaneous covariate $X_{t}$ and the previous period's outcome $Y_{t-1}$ via the threshold-crossing model,
\begin{align*}
Y_{t} = \mathbf{1}\{\beta t + X_{t} + \theta Y_{t-1} \geq \alpha + U_{t}\}, \quad t=1,2
\end{align*}
where $\beta = 1$, $\theta = 0.5$, $\alpha | X_{1},X_{2} \sim \mathcal{N}(0,1)$, $(U_{1},U_{2})|\alpha,X_{1},X_{2} \sim \mathcal{N}(0_{2},\Sigma_{U})$ with correlation matrix $\Sigma_{U}$ having off-diagonal equal to $0.5$, and $Y_{0}|X_{1},X_{2}\sim Bernoulli(0.5)$. We generate covariates according to the rule $X_{1},X_{2} \overset{iid}{\sim}U[-2,2]$. Our simulations focus on the posterior for the conditional identified set of $\gamma := (\beta,\theta)'$ under the stationarity assumption
\begin{align*}
U_{1}|X_{1},X_{2},\alpha \sim U_{2}|X_{1},X_{2},\alpha.
\end{align*}
Conditional stationarity is a common approach to index coefficient identification in semiparametric panel discrete choice models \citep{manski1987semiparametric,shi2018estimating,khan2021inference,pakes2021unobserved,khan2023identification,pakes2024moment}.
The identified set for $\gamma$ is determined by the conditional probability mass function $p$ of $(Y_{0},Y_{1},Y_{2})'$ given $X = (X_{1},X_{2})$. Specifically, Theorem 1 of \cite{khan2023identification} show that the sharp identified set is characterized by inequalities of the form $p(x) \in R \implies c_{R}(x,\gamma) \geq 0$, where $R$ is some restriction and $c_{R}(x,\gamma)$ is a known transformation of $X$ that depends on $R$. One illustrative inequality is
\begin{align*}
p(Y_{2}=1|X=x) \geq p(Y_{1}=1|X=x) \implies \beta + (x_{2}-x_{1}) + |\theta| \geq 0
\end{align*}
Hence, the dynamic panel binary choice model falls within our framework by defining $f(x,p(x),\gamma) = (f_{1}(x,p(x),\gamma),...,f_{18}(x,p(x),\gamma))'$, where $f_{j}(x,p(x),\gamma) = \mathbf{1}\{p(x) \in R_{j}\}c_{R_{j}}(x,\gamma)$, the $R_{j}$s are the restrictions that define the identified set, and $J=18$ is the number of restrictions for each $x$.\footnote{Appendix \ref{ap:simulation} lists the full set of inequalities. Importantly, by partitioning the parameter space based on the sign of $\theta$, the identified set can be viewed as the union of convex polytopes (of which fall into Example \ref{ex:linsyst}).} Figure \ref{fig:conditionalidset} presents the conditional identified set for $n=100,500,1000,2000$ based on a sequence of covariates $\{x_{i}\}_{i=1}^{n}$ that we keep fixed throughout the simulations.\footnote{Specifically, we first generate $\{(X_{i1},X_{i2})\}_{i=1}^{2000}$ i.i.d $U([-2,2])^{\otimes 2} $ and then compute $\Gamma_{n,I}(p_{0})$ using the truncated sequence $\{(X_{i1},X_{i2})\}_{i=1}^{n}$. Since the sequence of covariates are i.i.d, Theorem \ref{thm:fullsupport} implies that the conditional identified sets converge to the identified set based on the full support of the uniform distribution. In Appendix \ref{ap:comparisonwithothersets}, we compare the conditional identified set at $n=2000$ with the full support identified set. The latter is modestly smaller.}
\begin{figure}
\centering
\includegraphics[width=\linewidth]{ConditionalSets.pdf}
\caption{Conditional Identified Sets for Different Sample Sizes}
\label{fig:conditionalidset}
\end{figure}
\subsubsection{Simulation Results} Conditional on the fixed set of covariates $\{x_{i}\}_{i=1}^{n}$, we simulate $200$ independent panels comprised of $n$ independent cross-sectional units and implement the logistic Gaussian process Gibbs sampler to generate draws from the posterior for the reduced form PMF $p$. Since there are eight possible outcome paths (i.e., $K = |\{0,1\}^{3}| = 8$), we need to specify seven independent GPs to implement the logistic Gaussian process stick-breaking Gibbs sampler. For each $k$, we set $\mu_{k} = 0$ and specify that $\kappa_{k}$ belongs to the Mat\'{e}rn class,
\begin{align}\label{eq:materncovariance}
\kappa_{k}(x,\tilde{x}) = \frac{2^{1-\alpha_{k}}}{\Gamma(\alpha)}\left(\sqrt{2\alpha_{k}}\frac{||x-\tilde{x}||_{2}}{l_{k}}\right)^{\alpha_{k}}C_{\alpha_{k}}\left(\sqrt{2\alpha_{k}}\frac{||x-\tilde{x}||_{2}}{l_{k}}\right),
\end{align}
where $x = (x_{1},x_{2})'$, $\alpha_{k} > 0$ is a smoothness parameter, $l_{k}>0$ is a length-scale parameter, and $C_{\alpha_{k}}$ is the modified Bessel function of the second kind.\footnote{We standardize the covariates when evaluating $\kappa_{k}$.} The hyperparameter $\alpha_{k}$ indexes smoothness because the sample paths $B$ are $\lfloor\alpha_{k} \rfloor$-times differentiable \citep{JMLR:v12:vandervaart11a}, while the length-scale hyperparameter $l_{k}$ describes the rate at which the correlation between function values decays across the covariate space. We follow the common practice of setting $\alpha_{k} = 1.5$ for each $k=1,...,7$ \citep{williams2006gaussian}. To select the length scale, we update $l_{k}$ during the first $10,000$ iterations of the sampler (assuming $\log l_{k} \overset{iid}{\sim} \log N(0,1)$), and then fix $(l_{1},...,l_{7})$ at the medians of the draws obtained during the second half of this exploration phase.\footnote{A similar approach to length-scale selection is employed in \cite{kankanala2025generalized}.} We perform $20,000$ iterations of the Gibbs sampling algorithm, discarding the first $16,000$ iterations as burn-in and then thinning every four draws. This results in $1,000$ total draws from the posterior.\footnote{Even if updating $l_{k}$ for the entire chain, our prior is specified so that the product posterior from Section \ref{sec:samplerdescribe} arises, which means that if a user's computer has at least $7$ cores, then the posteriors $B_{n,k}|Y^{(n)}$, $k=1,...,7$, can be sampled completely in parallel, effectively only \textit{one} Gibbs sampling chain is performed.}
Using the output from the Gibbs sampler, we compute choice probabilities $(p(x_{1}),...,p(x_{n}))$, and then $\operatorname*{arg\,min}_{\gamma \in [-2,2] \times [-1,1]}Q_{n}^{LP}(\gamma,p)$ to obtain identified set posterior draws, where
\begin{align*}
Q_{n}^{LP}(\gamma,p) = \frac{1}{n}\sum_{i=1}^{n}\sum_{j=1}^{18}\left|\min\{\mathbf{1}\{p(x_{i}) \in R_{j}\}c_{R_{j}}(x_{i},\gamma),0\}\right|,
\end{align*}
In other words, we report the posterior for $\tilde{\Gamma}_{n,I}(p)=\{\gamma \in \Gamma: \tilde{Q}_{n}^{L}(\gamma,p) = 0\}$. The set of minimizers is found by evaluating $Q_{n}^{LP}(\gamma,p)$ over a grid in $[-2,2]\times [-1,1]$ with step size $10^{-2}$. Figure \ref{fig:inclusionprobabilities} plot the posterior probabilities that $\beta \in \tilde{B}_{n,I}(p)$ and $\theta \in \tilde{\Theta}_{n,I}(p)$ (averaged over the $200$ simulated datasets), where $\tilde{B}_{n,I}(p)$ and $\tilde{\Theta}_{n,I}(p)$ are the coordinate projections of $\tilde{\Gamma}_{n,I}(p)$ in the $\beta$ and $\theta$ directions, respectively. Comparing this with Figure \ref{fig:conditionalidset}, the identified set places increasing probability on the elements of the conditional identified set as $n$ grows. Table \ref{tab:mad} presents the mean absolute deviation (MAD) of $E[d_{\mathcal{H}}(\tilde{\Gamma}_{n,I}(p),\Gamma_{n,I}(p_{0}))|Y^{(n)}]$ from zero, which, by Markov's inequality, summarizes the amount of mass the posterior for $\tilde{\Gamma}_{n,I}(p)$ places on $d_{\mathcal{H}}$ neighborhoods of $\Gamma_{n,I}(p_{0})$. The results indicate that the posterior becomes more concentrated around $\Gamma_{n,I}(p_{0})$ as $n$ grows.
\begin{figure}[h!]
\centering
\includegraphics[width=\linewidth]{Prob_Plots.pdf}
\caption{Posterior Inclusion Probabilities}
\label{fig:inclusionprobabilities}
\end{figure}
\begin{table}[h!]
\centering
\label{tab:hausdorff}
\begin{tabular}{lllll}
& $n=100$ & $n=500$ & $n=1000$ & $n=2000$ \\ \cline{2-5}
\multicolumn{1}{l|}{MAD} & $1.733$ & $0.995$ & $0.705$ & $0.548$\\
& & & &
\end{tabular}
\caption{MAD for Different Sample Sizes}\label{tab:mad}
\end{table}
\section{Posterior Consistency for the Identified Set}\label{sec:posteriorconsistency}
This section derives conditions under which $\Gamma_{n,I}|Y^{(n)}$ concentrates around $\Gamma_{n,I}(p_{0})$ in large samples, where $p_{0}$ is the true conditional distribution of $Y$ given $X$.
\subsection{General Theorem.} This subsection presents a general posterior consistency theorem for $\Gamma_{n,I}|Y^{(n)}$ that applies to any prior $\Pi$ and criterion $Q_{n}$ satisfying some conditions. We start with some assumptions. For notation, let $P_{0,X}^{(\infty)}$ denote the probability law of $\{X_{i}\}_{i \geq 1}$, let $\text{dist}(x,A)$ denote the Euclidean distance between a point $x$ and a set $A$, and let $\Gamma_{n,I}^{\varepsilon} (p)= \{\gamma \in \Gamma: \text{dist}(\gamma,\Gamma_{n,I}(p)) < \varepsilon\}$ for each $\varepsilon > 0$ and $p \in \mathcal{P}$.
\begin{assumption}[DGP]\label{as:dgp}
Conditional on $P_{0,X}^{(\infty)}$-almost every fixed realization $\{x_{i}\}_{i\geq 1}$ of $\{X_{i}\}_{i \geq 1}$, the following hold:
\begin{enumerate}
\item Independent Outcomes: the sequence $\{Y_{i}\}_{i \geq 1}$ satisfies $Y_{i} \overset{ind}{\sim} Categorical(p_{0}(x_{i}))$ for every $i \geq 1$, where $p_{0} \in \mathcal{P}$.
\item Correct Specification and Well-Separated Criterion:
\begin{enumerate}
\item The true conditional identified set $\Gamma_{n,I}(p_{0})$ satisfies $\Gamma_{n,I}(p_{0}) \neq \emptyset$ for all $n \geq 1$.
\item For every $\varepsilon >0$, there exists $\xi > 0$ and $N \geq 1$ such that $Q_{n}(\gamma,p_{0}) \geq \xi$ for all $\gamma \in \Gamma \setminus \Gamma_{n,I}^{\varepsilon}(p_{0})$ and all $n \geq N$ and $\{\gamma \in \Gamma: Q_{n}(\gamma,p_{0}) \geq \xi\}$ is compact subset of $\Gamma$ for all $n \geq N$.
\end{enumerate}
\end{enumerate}
\end{assumption}
\begin{assumption}[Prior]\label{as:prior}
For $P_{0,X}^{(\infty)}$-almost every fixed realization $\{x_{i}\}_{i \geq 1}$ of $\{X_{i}\}_{ i\geq 1}$, there is a sequence of semimetrics $\{d_{n}\}_{n \geq 1}$ over $\mathcal{P}$, sequences $\{\delta_{n}\}_{n \geq 1}$ with $\delta_{n} \downarrow 0$ as $n\rightarrow \infty$, and sieves $\{\mathcal{P}_{n}\}_{n \geq 1} \subseteq \mathcal{P}$ such that
\begin{enumerate}
\item \textit{Posterior Concentration:}
\begin{enumerate}
\item $\Pi_{n}(p \in \mathcal{P}_{n}|Y^{(n)}) \rightarrow 1$ in $P^{(n)}_{0}$-probability as $n\rightarrow \infty$.
\item $\Pi_{n}(d_{n}(p,p_{0})< \delta_{n}|Y^{(n)}) \rightarrow 1$ in $P^{(n)}_{0}$-probability as $n\rightarrow \infty$.
\end{enumerate}
\item \textit{Uniform Convergence:} the sequence of criterion function $\{Q_{n}\}_{n \geq 1}$ is such that
\begin{align*}
\limsup_{n\rightarrow \infty}\sup_{(\gamma,p) \in \Gamma \times V_{n,\delta_{n}}}|Q_{n}(\gamma,p)-Q_{n}(\gamma,p_{0})| =0,
\end{align*}
where $V_{n,\delta_{n}} = \mathcal{P}_{n} \cap \{p: d_{n}(p,p_{0}) < \delta_{n}\}$.
\item \textit{Lower Hemicontinuity:} for every $\varepsilon >0$, there exists $N \geq 1$ such that $p \in V_{n,\delta_{n}}$ and $n \geq N$ implies $\Gamma_{n,I}(p) \cap B(\gamma,\varepsilon) \neq \emptyset$ for all $\gamma \in \Gamma_{n,I}(p_{0})$.
\end{enumerate}
\end{assumption}
We discuss the assumptions. Assumption \ref{as:dgp}.1 requires that the conditional sampling model be correctly specified. Assumption \ref{as:dgp}.2 imposes that the identified set $\Gamma_{n,I}(p_{0})$ is nonempty and, for large $n$, forms a well-separated set of minimizers of the criterion function $Q_{n}(\cdot,p_{0})$.\footnote{The compactness of $\{\gamma \in \Gamma: Q_{n}(\gamma,p_{0}) \geq \xi\}$ is a technical measurability condition. It holds, for instance, if lower semicontinuity of $Q_{n}(\gamma,p_{0})$ is strengthened to continuity in $\gamma$.} Separation conditions are often imposed for consistency of estimators under both point identification \citep{newey1994large} and partial identification \citep{chernozhukov2007estimation}. We treat misspecification (i.e., $\Gamma_{n,I}(p_{0}) = \emptyset$) momentarily. Assumption \ref{as:prior}.1 requires that the posterior concentrates on (possibly) infinite-dimensional sieves $\mathcal{P}_{n}$ and neighborhoods of $p_{0}$ with respect to some semimetric $d_{n}$ at rate $\delta_{n}$. Assumption \ref{as:prior}.2 states the criterion functions $Q_{n}(\gamma,p)$ converges uniformly to $Q_{n}(\gamma,p_{0})$ over $\Gamma \times V_{n,\delta_{n}}$, while Assumption \ref{as:prior}.3 is a lower hemicontinuity-type condition on the correspondence $p \mapsto \Gamma_{n,I}(p)$. Section \ref{sec:suff1} provides low-level sufficient conditions for Assumptions \ref{as:prior}.2 and \ref{as:prior}.3, while Section \ref{sec:GPverification} verifies Assumption \ref{as:prior}.1 for the logistic stick-breaking Gaussian process prior.
\begin{theorem}\label{thm:consistency}
Suppose Assumptions \ref{as:samplingmodel}--\ref{as:prior} hold. Then, for $P_{0,X}^{(\infty)}$-almost every fixed realization $\{x_{i}\}_{i \geq 1}$ of $\{X_{i}\}_{i \geq 1}$ and for every $\varepsilon > 0$,
\begin{align*}
\Pi_{n}\left(d_{\mathcal{H}}(\Gamma_{n,I}(p),\Gamma_{n,I}(p_{0})) \geq \varepsilon \middle|Y^{(n)}\right) \overset{P_{0}^{(n)}}{\longrightarrow} 0,
\end{align*}
as $n\rightarrow \infty$, where $d_{\mathcal{H}}$ is the Hausdorff distance and $P_{0}^{(n)} = \bigotimes_{i=1}^{n}Categorical(p_{0}(x_{i}))$.
\end{theorem}
There are two immediate corollaries of Theorem \ref{thm:consistency}. First, the posterior is consistent for the identified set $G_{n,I}(p_{0})$ of continuous transformations of $\gamma$ (of which a special case is the identified set for a subvector of $\gamma$). Second, the posterior is consistent for the full support identified set $\Gamma_{I}(p_{0})$ if, additionally, the conditions of Theorem \ref{thm:fullsupport} hold.
\begin{corollary}\label{cor:subvectors}
Suppose Assumptions \ref{as:samplingmodel} -- \ref{as:prior} and $g: \Gamma \rightarrow \mathbb{R}^{d_{g}}$ is continuous. Then, for $P_{0,X}^{(\infty)}$-almost every fixed realization $\{x_{i}\}_{i \geq 1}$ of $\{X_{i}\}_{i \geq 1}$ and every $\varepsilon > 0$, the following is true:
\begin{align*}
\Pi_{n}\left(d_{\mathcal{H}}(G_{n,I}(p),G_{n,I}(p_{0})) \geq \varepsilon \middle|Y^{(n)}\right) \overset{P_{0}^{(n)}}{\longrightarrow} 0,
\end{align*}
as $n\rightarrow \infty$, where $d_{\mathcal{H}}$ denotes the Hausdorff distance.
\end{corollary}
\begin{corollary}\label{cor:fullsupportidset}
Suppose Assumptions \ref{as:samplingmodel}--\ref{as:prior} hold and the assumptions of Theorem \ref{thm:fullsupport} hold too. Then, for $P_{0,X}^{(\infty)}$-almost every fixed realization $\{x_{i}\}_{i \geq 1}$ of $\{X_{i}\}_{i \geq 1}$ and every $\varepsilon >0$,
\begin{align*}
\Pi_{n}\left(d_{\mathcal{H}}\left(\Gamma_{n,I}(p),\Gamma_{I}(p_{0})\right) \geq \varepsilon \middle | Y^{(n)}\right) \overset{P_{0}^{(n)}}{\longrightarrow} 0
\end{align*}
as $n\rightarrow \infty$. If, additionally, $g: \mathbb{R}^{d_{\gamma}}\rightarrow \mathbb{R}$ is continuous, then, for $P_{0,X}^{(\infty)}$-almost almost every fixed realization $\{x_{i}\}_{i \geq 1}$ of $\{X_{i}\}_{i \geq 1}$ and every $\varepsilon > 0$,
\begin{align*}
\Pi_{n}\left(d_{\mathcal{H}}(G_{n,I}(p),G_{I}(p_{0})) \geq \varepsilon \middle|Y^{(n)}\right) \overset{P_{0}^{(n)}}{\longrightarrow} 0,
\end{align*}
as $n\rightarrow \infty$, where $d_{\mathcal{H}}$ denotes Hausdorff distance.
\end{corollary}
Theorem \ref{thm:consistency}'s assumptions (or slight modifications of them) also imply some asymptotic properties under misspecification. Specifically, the posterior probability that $\Gamma_{n,I}(p) \neq \emptyset$ is a consistent diagnostic for model misspecification, and the posterior for the pseudo-identified set $\tilde{\Gamma}_{n,I}(p) = \{\gamma \in \Gamma: \tilde{Q}_{n}(\gamma,p)=0\}$ is consistent at $\tilde{\Gamma}_{n,I}(p_{0})$.
\begin{theorem}\label{thm:test}
Suppose Assumptions \ref{as:samplingmodel} -- \ref{as:prior} hold. Then, the following holds for $P_{0,X}^{(\infty)}$-almost every fixed realization $\{x_{i}\}_{i \geq 1}$ of $\{X_{i}\}_{i \geq 1}$,
\begin{enumerate}
\item If $\liminf_{n\rightarrow \infty}\inf_{\gamma \in \Gamma}Q_{n}(\gamma,p_{0}) > 0$, then $\Pi_{n}(\Gamma_{n,I}(p) = \emptyset | Y^{(n)}) = 1 + o_{P_{0}^{(n)}}(1)$.
\item If $\Gamma_{n,I}(p_{0}) \neq \emptyset$ for all $n \geq 1$, then $\Pi_{n}(\Gamma_{n,I}(p) \neq \emptyset | Y^{(n)}) = 1 + o_{P_{0}^{(n)}}(1)$.
\end{enumerate}
\end{theorem}
\begin{theorem}\label{thm:misspec_consistency}
Suppose Assumptions \ref{as:samplingmodel}--\ref{as:prior} hold, except with the following modifications:
\begin{enumerate}
\item Modification of Assumption \ref{as:dgp}.2: for every $\varepsilon >0$, there exists $\xi > 0$ and $N \geq 1$ such that $\tilde{Q}_{n}(\gamma,p_{0}) \geq \xi$ for all $\gamma \in \Gamma \setminus \tilde{\Gamma}_{n,I}^{\varepsilon}(p_{0})$ and all $n \geq N$, and $\{\gamma \in \Gamma: \tilde{Q}_{n}(\gamma,p_{0}) \geq \xi\}$ is a compact subset of $\Gamma$ for all $n \geq N$.
\item Modification of Assumption \ref{as:prior}.3: for every $\varepsilon >0$, there exists $N \geq 1$ such that $p \in V_{n,\delta_{n}}$ and $n \geq N$ implies $\tilde{\Gamma}_{n,I}(p) \cap B(\gamma,\varepsilon) \neq \emptyset$ for all $\gamma \in \tilde{\Gamma}_{n,I}(p_{0})$..
\end{enumerate}
Then, for $P_{0,X}^{(\infty)}$-almost every fixed realization $\{x_{i}\}_{i \geq 1}$ of $\{X_{i}\}_{i \geq 1}$ and every $\varepsilon > 0$, the following is true:
\begin{align*}
\Pi_{n}\left(d_{\mathcal{H}}(\tilde{\Gamma}_{n,I}(p),\tilde{\Gamma}_{n,I}(p_{0})) \geq \varepsilon \middle|Y^{(n)}\right) \overset{P_{0}^{(n)}}{\longrightarrow} 0,
\end{align*}
as $n\rightarrow \infty$, where $d_{\mathcal{H}}$ denotes the Hausdorff distance.
\end{theorem}
\subsection{Verifying Assumption \ref{as:prior}.}\label{sec:verifyingconditions}
This section offers sufficient conditions for Assumption \ref{as:prior}. There are three subsections: the first two present low-level sufficient conditions for Assumptions \ref{as:prior}.2 and \ref{as:prior}.3 (and its misspecified counterpart). Then, noticing a common uniform convergence requirement, we verify uniform convergence for the GP priors from Section \ref{sec:GPimplementation} in the third subsection to find a sufficient condition for the existence of sieved neighborhoods $\{V_{n,\delta_{n}}\}_{n \geq 1}$.
\subsubsection{Sufficient Conditions for Assumption \ref{as:prior}.2}\label{sec:suff1} The first proposition gives sufficient conditions for uniform convergence of $Q_{n}(\gamma,p)$ (Assumption \ref{as:prior}.2) in the case where
\begin{align*}
Q_{n}(\gamma,p) = ||\text{dist}_{W(\cdot,p(\cdot),\gamma)}(f(\cdot,p(\cdot),\gamma),\mathbb{R}_{+}^{d_{f}})||_{n,r}
\end{align*}
for some $r \in [1,\infty]$. For notation, let $\lambda_{min}(A)$ and $\lambda_{max}(A)$ denote the minimum and maximum eigenvalues, respectively, of a matrix $A$.
\begin{proposition}\label{prop:ucon}
Suppose that, for $P_{0,X}^{(\infty)}$-almost every fixed realization $\{x_{i}\}_{i \geq 1}$ of $\{X_{i}\}_{i \geq 1}$, the following conditions hold:
\begin{enumerate}
\item Bounded criterion:
\begin{align*}
\sup_{\gamma \in \Gamma}||\text{dist}_{W(\cdot,p_{0}(\cdot),\gamma)}(f(\cdot,p_{0}(\cdot),\gamma),\mathbb{R}_{+}^{d_{f}})||_{n,r} < \infty.
\end{align*}
\item Bounded weighting matrix: there exists constants $0 < \underline{\lambda}_{W} \leq \overline{\lambda}_{W} < \infty$ such that
\begin{align*}
\underline{\lambda}_{W} \leq \inf_{\gamma \in \Gamma} \inf_{n \geq 1}\min_{1 \leq i \leq n} \lambda_{min}(W(x_{i},p_{0}(x_{i}),\gamma))
\end{align*}
and
\begin{align*}
\sup_{\gamma \in \Gamma} \sup_{n \geq 1}\max_{1 \leq i \leq n} \lambda_{max}(W(x_{i},p_{0}(x_{i}),\gamma)) \leq \overline{\lambda}_{W}.
\end{align*}
\item Uniform convergence: $\{\mathcal{P}_{n}\}_{n \geq 1}$, $\{d_{n}\}_{n \geq 1}$, and $\{\delta_{n}\}_{n \geq 1}$ from Assumption \ref{as:prior} are such that
\begin{align}\label{eq:supf}
\sup_{(\gamma,p) \in \Gamma \times V_{n,\delta_{n}}}\left| \left| f(\cdot,p(\cdot),\gamma)- f(\cdot,p_{0}(\cdot),\gamma) \right| \right|_{n,\infty} = o(1)
\end{align}
and
\begin{align*}
\sup_{(\gamma,p) \in \Gamma \times V_{n,\delta_{n}}}\left| \left| W(\cdot,p(\cdot),\gamma)-W(\cdot,p_{0}(\cdot),\gamma)\right| \right|_{n,\infty} = o(1)
\end{align*}
as $n\rightarrow \infty$.
\end{enumerate}
Then Assumption \ref{as:prior}.2 holds for $Q_{n}(\gamma,p) = ||\text{dist}_{W(\cdot,p(\cdot),\gamma)}(f(\cdot,p(\cdot),\gamma),\mathbb{R}_{+}^{d_{f}})||_{n,r}$.
\end{proposition}
Proposition \ref{prop:ucon} states that Assumption \ref{as:prior}.2 holds if $Q_{n}(\cdot,p_{0}): \Gamma \rightarrow \mathbb{R}_{+}$ is bounded, the class of weight matrices $\{W(x_{i},p_{0}(x_{i}),\gamma): 1 \leq i \leq n, \ n \geq 1, \ \gamma \in \Gamma\}$ forms a compact set of positive-definite matrices, and the sieve neighborhoods $V_{n,\delta_{n}}$ are structured enough to ensure convergence of $W(\cdot,p(\cdot),\gamma)$ and $f(\cdot,p(\cdot),\gamma)$ to $W(\cdot,p_{0}(\cdot),\gamma)$ and $f(\cdot,p_{0}(\cdot),\gamma)$, respectively, in the empirical supremum norm, and uniformly over $\Gamma \times V_{n,\delta_{n}}$.\footnote{It should be noted convergence of $f$ in $||\cdot||_{\infty}$ is sufficient for simultaneous control over all $r \in [1,\infty]$, however one can show that Assumption \ref{as:prior}.2 still holds pointwise across $r$ if (\ref{eq:supf}) is weakened to $||\cdot||_{n,r}$; we prefer to state the proposition in terms of $||\cdot||_{n,\infty}$ as this convergence is important for verifying Assumption \ref{as:prior}.3.} The first two conditions hold if $x$ takes values in a compact set, $(x,\gamma) \mapsto f_{0}(x,p_{0}(x),\gamma)$ is continuous, and $(x,\gamma) \mapsto W(x,p_{0}(x),\gamma)$ is continuous and $W(x,p_{0}(x),\gamma)$ is positive-definite for each $(x,\gamma) \in \mathcal{X} \times \Gamma$. Condition 3 is typically satisfied if $V_{n,\delta_{n}}$ is contained in a $||\cdot||_{n,\infty}$-ball centered at $p_{0}$ with shrinking radius.
\setcounter{example}{0}
\begin{example}[Continued]
Suppose that $\mathcal{X}$ is compact and $g(y,x,\gamma)$ is continuous in $(x,\gamma)$ for each $y$ (alternatively, if $g(y,x,\gamma)$ is bounded in $(x,\gamma)$), then, coupled with Assumption \ref{as:parameterspace}, there is a constant $C \in (0,\infty)$ that is independent of $\{x_{i}\}_{i=1}^{n}$ such that
\begin{align*}
||f(\cdot,p(\cdot),\gamma) - f(\cdot,p_{0}(\cdot),\gamma)||_{n,\infty} \leq C||p-p_{0}||_{n,\infty}.
\end{align*}
Consequently, the first uniform convergence requirement is met if $V_{n,\delta_{n}}$ is contained in an empirical supremum norm ball centered at $p_{0}$ with slowly shrinking radius. Similarly, under positive definiteness of $\{W(\cdot,p_{0}(\cdot),\gamma): \gamma \in \Gamma\}$, the second uniform convergence condition is met for $W(x,p(x),\gamma) = \Sigma^{-1}(x,\gamma)$ and $W(x,p(x),\gamma) = D^{-1}(x,\gamma)$ if $V_{n,\delta_{n}}$ is contained in an empirical supremum norm ball centered at $p_{0}$ with slowly shrinking radius.
\end{example}
\begin{example}[Continued]
Since $\Gamma$ is compact, $\sup_{\gamma \in \Gamma}||f(\cdot,p(\cdot),\gamma) - f(\cdot,p_{0}(\cdot),\gamma)||_{n,\infty}$ is bounded from above by a constant multiple of the maximum of $||A(\cdot,p(\cdot))-A(\cdot,p_{0}(\cdot))||_{n,\infty}$ and $||b(\cdot,p(\cdot))-b(\cdot,p_{0}(\cdot))||_{n,\infty}$. Hence, uniform convergence is satisfied if $A(\cdot,p(\cdot))$ and $b(\cdot,p(\cdot))$ is continuous in $p$ with respect to $||\cdot||_{n,\infty}$ and $V_{n,\delta_{n}}$ is contained in an empirical supremum norm ball centered at $p_{0}$ with slowly shrinking radius. For concreteness, consider \cite{khan2023identification}-type inequalities $\mathbf{1}\{p(x_{i}) \in R\}c_{R}(x_{i})'\gamma \geq 0$ for all $i=1,...,n$ and for all $ R \in \mathcal{R}$, where $R$ is a polyhedral restriction on $p(x)$, $\mathcal{R}$ is a finite set of restrictions $R$, and $c_{R}(x)$ is a known vector. Let $r_{n}\downarrow 0$ as $n\rightarrow \infty$ be such that $V_{n,\delta_{n}} \subseteq \{p: ||p-p_{0}||_{n,\infty} < r_{n}\}$ and let $\rho_{n}:=\min_{1 \leq i \leq n}\min_{R \in \mathcal{R}}\text{dist}(p_{0}(x_{i}),\partial R)$. If there is an $N \geq 1$ such that $r_{n} < \rho_{n}$, then $||A(\cdot,p(\cdot))-A(\cdot,p_{0}(\cdot))||_{n,\infty} = 0$ for all $n \geq N$, which guarantees the desired uniform convergence of $f$ (we can ignore $b(x,p(x))$ as it is zero for all $x$ and $p$).
\end{example}
Since Proposition \ref{prop:ucon} applies to minimum distance criterion functions based on weighted Euclidean norms, it does not immediately apply the criterion $Q_{n}^{LP}(\gamma,p)$ introduced in Example \ref{ex:linsyst}. Fortunately, $Q_{n}^{LP}(\gamma,p)$ satisfies Assumption \ref{as:prior}.2 under empirical $L^{1}$ convergence of $f$ (implied by the same uniform convergence assumption on $f$ as Proposition \ref{prop:ucon}).
\begin{proposition}\label{prop:taxicab}
Suppose that, for $P_{0,X}^{(\infty)}$-almost every fixed realization $\{x_{i}\}_{i \geq 1}$ of $\{X_{i}\}_{i \geq 1}$, $\{\mathcal{P}_{n}\}_{n \geq 1}$, $\{d_{n}\}_{n \geq 1}$, and $\{\delta_{n}\}_{n \geq 1}$ from Assumption \ref{as:prior} are such that
\begin{align*}
\sup_{(\gamma,p) \in \Gamma \times V_{n,\delta_{n}}}\left| \left| f(\cdot,p(\cdot),\gamma)- f(\cdot,p_{0}(\cdot),\gamma) \right| \right|_{n,1} = o(1)
\end{align*}
as $n\rightarrow \infty$. Then Assumption \ref{as:prior}.2 holds for $Q_{n}^{LP}(\gamma,p) = ||(f(\cdot,p(\cdot),\gamma))_{-}||_{n,1}$.
\end{proposition}
\subsubsection{Verifying Assumption \ref{as:prior}.3 and its Misspecified Counterpart}\label{sec:suffhemi} We provide low-level sufficient conditions for Assumption \ref{as:prior}.3. Proposition \ref{prop:hemi} states Assumption \ref{as:prior}.3 holds if, additionally, $\Gamma_{n,I}(p)$ is `strongly identified' for all $p \in V_{n,\delta_{n}}$.
\begin{proposition}\label{prop:hemi}
If, for $P_{0,X}^{(\infty)}$-almost every fixed realization $\{x_{i}\}_{i \geq 1}$ of $\{X_{i}\}_{i \geq 1}$, the sequences $\{\mathcal{P}_{n}\}_{n \geq 1}$, $\{d_{n}\}_{n \geq 1}$, and $\{\delta_{n}\}_{n \geq 1}$ from Assumption \ref{as:prior} are such that
\begin{align*}
\sup_{(\gamma,p) \in \Gamma \times V_{n,\delta_{n}}}\left| \left| f(\cdot,p(\cdot),\gamma)- f(\cdot,p_{0}(\cdot),\gamma) \right| \right|_{n,\infty} = o(1),
\end{align*}
and there is a constant $c \in (0,\infty)$ and $\tilde{N} \geq 1$ such that, for each $\gamma \in \Gamma$,
\begin{align}\label{eq:errorbound}
\max_{1 \leq i \leq n} ||(f(x_{i},p(x_{i}),\gamma))_{-}||_{2} \geq c \text{dist}(\gamma,\Gamma_{n,I}(p))
\end{align}
holds for all $p \in V_{n,\delta_{n}}$ and all $n \geq \tilde{N}$, then Assumption \ref{as:prior}.3 holds.
\end{proposition}
Conditions like (\ref{eq:errorbound}) are regularly encountered in the literature on inference for identified sets. For example, it is similar to a linear version of the polynomial minorant condition (4.5) in \cite{chernozhukov2007estimation} and an unstudentized version of Assumption 3 in \cite{kaido2022constraint}. We call it a strong identification condition on $\Gamma_{n,I}(p)$ because it states that, as $\gamma$ moves away from $\Gamma_{n,I}(p)$, the maximal violation of the inequality restrictions grows sufficiently fast. The result readily extends to `union bounds' in which $\Gamma_{n,I}(p) = \bigcup_{j=1}^{J}\Gamma_{n,I}^{(j)}(p)$ with $\Gamma_{n,I}^{(j)} = \{\gamma \in \Gamma^{(j)}: f^{(j)}(x_{i},p(x_{i}),\gamma) \in \mathbb{R}_{+}^{d_{f^{(j)}}} \ \forall \ i=1,...,n\}$ satisfying the conditions of Proposition \ref{prop:hemi}. This result, combined with Proposition \ref{prop:convexhemi} below, is useful for the dynamic binary panel model studied in Section \ref{sec:GPimplementation} (we elaborate on this after stating Proposition \ref{prop:convexhemi}).
\begin{corollary}\label{cor:unions}
Suppose that
\begin{align*}
\Gamma_{n,I}(p) = \bigcup_{j=1}^{J}\Gamma_{n,I}^{(j)}(p),
\end{align*}
where, for each $j=1,...,J$,
\begin{align*}
\Gamma_{n,I}^{(j)} = \left\{\gamma \in \Gamma^{(j)}: f^{(j)}(x_{i},p(x_{i}),\gamma) \in \mathbb{R}_{+}^{d_{f^{(j)}}} \ \forall \ i=1,...,n\right\}
\end{align*} satisfies the conditions of Proposition \ref{prop:hemi}. Then Assumption \ref{as:prior}.3 holds.
\end{corollary}
Condition (\ref{eq:errorbound}) (and Corollary \ref{cor:unions}) is still quite high-level, and results from the mathematical optimization literature indicate that its satisfaction may depend subtly on the geometry of $\Gamma_{n,I}$ (see \cite{IOFFE_2016} for a recent survey). For this reason, we provide an interpretable set of conditions for (\ref{eq:errorbound}) for an important class of identified sets.
\begin{proposition}\label{prop:convexhemi}
Suppose that the following conditions hold:
\begin{enumerate}
\item Parameter Space: $\Gamma$ is a compact and convex subset of $\mathbb{R}^{d_{\gamma}}$.
\item Smooth concavity: $f(x,p(x),\gamma)$ is differentiable and concave in $\gamma$ for all $(x,p)$.
\end{enumerate}
Further suppose that, for $P_{0,X}^{(\infty)}$-almost every fixed realization $\{x_{i}\}_{i\geq 1}$ of $\{X_{i}\}_{i \geq 1}$, the following conditions hold:
\begin{enumerate}
\setcounter{enumi}{2}
\item Interior true identified set: there exists $\bar{N}_{1} \geq 1$ and $\xi>0$ such that
\begin{align*}
\Gamma_{n,I}(p_{0}) \subseteq \text{int}(\Gamma)
\end{align*}
and
\begin{align*}
\sup_{\gamma \in \partial \Gamma}\max_{1\leq i \leq n}\max_{1 \leq j \leq d_{f}}f_{j}(x_{i},p_{0}(x_{i}),\gamma) \leq -\xi
\end{align*}
for all $n \geq \bar{N}_{1}$.
\item Slater condition: there exists $\gamma^{\circ} \in \Gamma$, $\eta > 0$, and $\bar{N}_{2} \geq 1$ such that
\begin{align*}
\min_{1\leq i \leq n}\min_{1 \leq j \leq d_{f}}f_{j}(x_{i},p_{0}(x_{i}),\gamma^{\circ}) \geq \eta
\end{align*}
for all $n \geq \bar{N}_{2}$.
\item Uniform convergence: the sequences $\{\mathcal{P}_{n}\}_{n \geq 1}$, $\{d_{n}\}_{n \geq 1}$, and $\{\delta_{n}\}_{n \geq 1}$ from Assumption \ref{as:prior} are such that
\begin{align*}
\sup_{(\gamma,p) \in \Gamma \times V_{n,\delta_{n}}}\left| \left| f(\cdot,p(\cdot),\gamma)- f(\cdot,p_{0}(\cdot),\gamma) \right| \right|_{n,\infty} = o(1).
\end{align*}
\end{enumerate}
Then Assumption \ref{as:prior}.3 holds.
\end{proposition}
Proposition \ref{prop:convexhemi} gives conditions under which Assumption \ref{as:prior}.3 holds when $\Gamma_{n,I}(p)$ is convex. Convex $\Gamma_{n,I}(p)$ is satisfied by several discrete response models, such as interval censored binary choice models \citep{manski2002inference}, ordered choice models \citep{pakes2015moment}, and panel multinomial choice models \citep{shi2018estimating}. Provided the components $\Gamma_{n,I}^{(j)}(p)$ of $\Gamma_{n,I}(p) = \bigcup_{j=1}^{J}\Gamma_{n,I}^{(j)}(p)$ satisfy Conditions 1--5 of Proposition \ref{prop:convexhemi}, Corollary \ref{cor:unions} implies that it readily extends to identified sets that are unions of convex sets. Related to Section \ref{sec:GPimplementation}, this case conforms to the identification results from \cite{khan2023identification}, where the identified sets are unions of convex polytopes determined by the sign of a state dependence parameter. In terms of the conditions themselves, the key condition on the data-generating process $p_{0}$ in Proposition \ref{prop:convexhemi} is the existence of a uniform slater point $\gamma^{\circ}$ because, combined with the other conditions, it allows us to find a constant $c > 0$ through an asymptotic uniform boundedness (over $\Gamma \times V_{n,\delta_{n}}$) property for the Lagrange multipliers associated with the convex program $\inf_{\tilde{\gamma} \in \Gamma_{n,I}(p)}||\gamma-\tilde{\gamma}||_{2}$.\footnote{The requirements on $p_{0}$ that $\Gamma_{n,I}(p_{0})$ is in the interior of $\Gamma$ is a standard condition (e.g., it is similar to Assumption 1(b) in \cite{kaido2022constraint}), while the separation from the boundary of $\Gamma$ is merely a technical condition that provides enough uniformity to account for $p$ being random.} The uniform convergence for $f(\cdot,p(\cdot),\gamma)$ is the same as Propositions \ref{prop:ucon}-\ref{prop:hemi}, and it is the \textit{only} role for the prior $\Pi$.
Propositions \ref{prop:hemi} and \ref{prop:convexhemi} are catered towards the correctly specified case where $\Gamma_{n,I}(p_{0}) \neq \emptyset$. The next result provides a sufficient condition for the modified Assumption \ref{as:prior}.3 that appears in Theorem \ref{thm:misspec_consistency} (i.e., the posterior consistency for $\tilde{\Gamma}_{n,I}$); it is a counterpart to Proposition \ref{prop:hemi}. \begin{proposition}\label{prop:minorantmisspec}
Suppose that for $P_{0,X}^{(\infty)}$-almost every fixed realization $\{x_{i}\}_{i \geq 1}$ of $\{X_{i}\}_{i \geq 1}$, the following holds:
\begin{align}\label{eq:hemi_misspec_ucon}
u_{n}:=\sup_{(\gamma,p) \in \Gamma \times V_{n,\delta_{n}}}\left|\tilde{Q}_{n}(\gamma,p)-\tilde{Q}_{n}(\gamma,p_{0})\right| \longrightarrow 0
\end{align}
as $n\rightarrow \infty$. Further, suppose that, for $P_{0,X}^{(\infty)}$-almost every fixed realization $\{x_{i}\}_{i \geq 1}$ of $\{X_{i}\}_{i \geq 1}$, there exists a sequence $\{r_{n}\}_{n \geq 1}$, constants $c,\zeta >0$, and an integer $\tilde{N} \geq 1$ such that $r_{n} \geq 0$ for all $n \geq 1$ and $r_{n}\rightarrow 0$ as $n\rightarrow \infty$, and, for each $\gamma \in \Gamma$,
\begin{align}\label{eq:hemi_misspec_minorant}
\tilde{Q}_{n}(\gamma,p) \geq c \cdot \text{dist}(\gamma,\tilde{\Gamma}_{n,I}(p))^{\zeta}- r_{n}
\end{align}
for all $p \in V_{n,\delta_{n}}$ and $n \geq \tilde{N}$. Then, the following holds for $P_{0,X}^{(\infty)}$-almost every fixed realization $\{x_{i}\}_{i \geq 1}$ of $\{X_{i}\}_{i \geq 1}$: for every $\varepsilon >0$, there exists $N \geq 1$ such that $p \in V_{n,\delta_{n}}$ and $n \geq N$ implies $\tilde{\Gamma}_{n,I}(p) \cap B(\gamma,\varepsilon) \neq \emptyset$ for all $\gamma \in \tilde{\Gamma}_{n,I}(p_{0})$.
\end{proposition}
Proposition \ref{prop:minorantmisspec} states that the analogue of Assumption \ref{as:prior}.3 for $\tilde{\Gamma}_{n,I}(p)$ holds under a uniform convergence and asymptotic \textit{weak sharp minimum-type property} for $\tilde{Q}_{n}(\gamma,p)$. We label the latter this way because if $\zeta = 1$ and $r_{n} = 0$ for all $n \geq 1$, then it collapses to a stochastic version of the idea introduced in \cite{doi:10.1137/0331063}. We also note that, by allowing $\zeta >0$, the $r_{n} = 0$ case closely relates to the polynomial minorant condition from \cite{chernozhukov2007estimation}. Like Proposition \ref{prop:hemi}, we also provide a set of interpretable low-level sufficient conditions for Proposition \ref{prop:minorantmisspec}.
\begin{proposition}\label{prop:misspec_hemi_convex}
Suppose that the following conditions hold:
\begin{enumerate}
\item Parameter space: $\Gamma \subseteq \mathbb{R}^{d_{\gamma}}$ is a compact and convex subset of $\mathbb{R}^{d_{\gamma}}$.
\end{enumerate}
Further, suppose that, for $P_{0,X}^{(\infty)}$-almost every fixed realization $\{x_{i}\}_{i \geq 1}$ of $\{X_{i}\}_{i \geq 1}$, the following conditions hold:
\begin{enumerate}
\setcounter{enumi}{1}
\item Piecewise Maximum of Smooth Functions: for every $n \geq 1$, there exists a finite index set $\mathcal{A}_{n}$ and functions $\{q_{n,a}: a \in \mathcal{A}_{n}\}$ defined on $\Gamma^{o} \times \mathcal{P}$, where $\Gamma^{o}$ is an open set that contains $\Gamma$, such that $$Q_{n}(\gamma,p) = \max_{a \in \mathcal{A}_{n}}q_{n,a}(\gamma,p)$$ for each $n \geq 1$, and $q_{n,a}(\gamma,p)$ is differentiable and convex in $\gamma$ for each $p$, $a$, and $n$.
\item Uniform convergence: the sequences $\{\mathcal{P}_{n}\}_{n \geq 1}$, $\{d_{n}\}_{n \geq 1}$, and $\{\delta_{n}\}_{n \geq 1}$ from Assumption \ref{as:prior} are such that
\begin{align*}
\Delta_{n}&:=\sup_{(\gamma,p) \in \Gamma \times V_{n,\delta_{n}}}\max_{a \in \mathcal{A}_{n}}|q_{n,a}(\gamma,p)-q_{n,a}(\gamma,p_{0})| \longrightarrow 0 \\
\Delta_{n}'&:=\sup_{(\gamma,p) \in \Gamma \times V_{n,\delta_{n}}}\max_{a \in \mathcal{A}_{n}}\left|\left|\nabla_{\gamma} q_{n,a}(\gamma,p)-\nabla_{\gamma}q_{n,a}(\gamma,p_{0})\right| \right|_{2} \longrightarrow 0
\end{align*}
as $n\rightarrow \infty$.
\item Gradient: $\nabla_{\gamma}q_{n,a}(\gamma,p_{0})$ is $L_{0}$-Lipschitz on $\Gamma$, where $L_{0} \in (0,\infty)$ is independent of $n$ and $a$, and there exists $\Gamma_{nbd} \subseteq \Gamma$ and $\breve{N} \geq 1$ such that $\tilde{\Gamma}_{n,I}(p_{0}) \subseteq \Gamma_{nbd}$ for all $n \geq \breve{N}$ and $\sup_{n \geq \breve{N}}\sup_{\gamma \in \Gamma_{nbd}}\max_{a \in \mathcal{A}_{n}}||\nabla_{\gamma}q_{n,a}(\gamma,p_{0})||_{2} < \infty$.
\item Uniform Nondegeneracy: there exists $\eta > 0$ and $\bar{N} \geq 1$ such that, for all $n \geq \bar{N}$ and all $\tilde{\gamma} \in \tilde{\Gamma}_{n,I}(p_{0})$,
\begin{align*}
\{t \in \mathbb{R}^{d_{\gamma}}: ||t||_{2} \leq \eta\} \subseteq \operatorname{conv} \left\{\nabla_{\gamma}q_{n,a}(\tilde{\gamma},p_{0}): a \in \mathcal{A}_{n}^{*}(\tilde{\gamma},p_{0})\right\},
\end{align*}
where $\operatorname{conv}$ denotes convex hull and $\mathcal{A}_{n}^{*}(\gamma,p) = \{a \in \mathcal{A}_{n}: q_{n,a}(\gamma,p) = Q_{n}(\gamma,p)\}$.
\end{enumerate}
Then Proposition \ref{prop:minorantmisspec} holds.
\end{proposition}
Some brief remarks on the conditions of Proposition \ref{prop:misspec_hemi_convex}. First, the prior only enters in Condition 3, where it is assumed that $\{V_{n,\delta_{n}}\}_{n \geq 1}$ ensure uniform convergence of the component functions and their gradients. Appendix \ref{ap:uconmisspec} argues that, for some $Q_{n}(\gamma,p)$, $\Delta_{n}$ and $\Delta_{n}'$ are bounded by the empirical uniform norm of $f$ and $\nabla_{\gamma}f$, so a similar takeaway to the correctly specified case applies: Condition 3 can be satisfied if $V_{n,\delta_{n}}$ is contained in a $||\cdot||_{n,\infty}$-ball centered at $p_{0}$ with shrinking radius. Second, the key restriction on $p_{0}$ is Condition 5, which is interpreted as the misspecification Slater condition in Proposition \ref{prop:convexhemi}. Appendix \ref{ap:complementarity} argues these properties are complementary: Condition 5 can fail under correct specification when the slater condition holds, while Condition 5 deals with $\Gamma_{n,I}(p_{0}) = \emptyset$, a setting where, mechanically, Slater fails.
Finally, Condition 5 implies that $\tilde{\Gamma}_{n,I}(p_{0})$ is a singleton for large $n$. This is not an artifact of the proof: with a flat bottom, lower hemicontinuity of the exact pseudo-identified set can fail under Conditions 1-4. Indeed, in such case, posterior draws of $p$ act as small tilts of the criterion that collapse the exact set of minimizers onto draw-dependent faces of $\tilde{\Gamma}_{n,I}(p_{0})$, so the posterior for the exact set concentrates on strict subsets of $\tilde{\Gamma}_{n,I}(p_{0})$ and is spuriously precise -- a Bayesian analogue of the spurious precision under misspecification analyzed by \cite{andrews2024misspecified}.\footnote{Appendix \ref{ap:failurehemicontinuity} presents an example of this.} When a non-singleton pseudo-identified set is anticipated, we recommend reporting the posterior of a slightly relaxed pseudo-identified set $\tilde{\Gamma}_{n,I}^{t}:=\{\gamma \in \Gamma: \tilde{Q}_{n}(\gamma,p) \leq t\}$, $t > 0$, rather than the exact set of minimizers. Proposition \ref{prop:relaxed} in Appendix \ref{ap:failurehemicontinuity} shows the relaxation remedies this: if $\tilde{\Gamma}_{n,I}(p_{0})$ is a set of weak sharp minima of $\tilde{Q}_{n}(\cdot,p_{0})$, which, unlike Condition 5, is a condition compatible with flat bottoms, then, for large $n$ and uniformly over $p \in V_{n,\delta_{n}}$, $\tilde{\Gamma}_{n,I}^{t}(p)$ contains $\tilde{\Gamma}_{n,I}(p_{0})$ and lies within Hausdorff distance $2t/\eta$ of it. Furthermore, taking $t:=t_{n} \downarrow 0$ with $u_{n}/t_{n}\rightarrow 0$ as $n\rightarrow \infty$ restores Hausdorff consistency without any singleton requirement, which mirrors the slack sequences of \cite{chernozhukov2007estimation}.
\subsubsection{Uniform Convergence of Choice Probabilities Under Gaussian Priors.}\label{sec:GPverification}
A key takeaway from Section \ref{sec:suff1} is that, for certain classes of $\Gamma_{n,I}(p_{0})$ (and $\tilde{\Gamma}_{n,I}(p_{0})$), empirical uniform convergence of the conditional choice probabilities is a sufficient condition for Assumption \ref{as:prior}. This section derives conditions under which $\Pi_{n}((p(x_{1}),...,p(x_{n})) \in \cdot | Y^{(n)})$ concentrates around $(p_{0}(x_{1}),...,p_{0}(x_{n}))$ in the empirical supremum norm for the logistic stick-breaking GP prior. For notation, let $C([0,1]^{d_{x}},\mathbb{R})$ be the space of real-valued continuous functions defined on $[0,1]^{d_{x}}$, let $||\cdot||_{\infty}$ be the supremum norm (i.e., $||b||_{\infty} = \sup_{x \in [0,1]^{d_{x}}}|b(x)|$), and let $C^{a}([0,1]^{d_{x}},\mathbb{R})$, $a > 0$, be the H\"{o}lder space of functions that are $\lfloor a \rfloor$ times differentiable with $\lfloor a \rfloor$th derivative that is H\"{o}lder continuous with index $(a- \lfloor a \rfloor)$.
\begin{assumption}[Covariates]\label{as:covariates}
The following conditions hold:
1. $\mathcal{X} = [0,1]^{d_{x}}$, 2. $\{X_{i}\}_{i \geq 1}$ is i.i.d (i.e., $P_{0,X}^{(\infty)} = \bigotimes_{i \geq 1}P_{0,X}$), and 3. $P_{0,X}$ admits a density $p_{0,X}$ that is uniformly bounded away from zero and infinity on $[0,1]^{d_{x}}$.
\end{assumption}
\begin{assumption}[Gaussian Process]\label{as:GP}
The prior $\Pi$ for $p $ is $LSBGP(\{0\}_{k=1}^{K-1},\{\kappa_{k}\}_{k=1}^{K-1})$ with the Gaussian processes $B_{1},...,B_{K-1}$ additionally satisfying 1. $B_{k} \in C^{a_{k}}([0,1]^{d_{x}},\mathbb{R})$ for all $a_{k} < \alpha_{k}$, and 2. are truncated to $||B_{k}||_{\infty} \leq \overline{B}_{k}$ for some $\overline{B}_{k} < \infty$ for each $k=1,...,K-1$.
\end{assumption}
Assumptions \ref{as:covariates} and \ref{as:GP} are standard restrictions on $\{X_{i}\}_{i \geq 1}$ and the prior. Assumption \ref{as:covariates}.1 imposes that the covariates take values in $[0,1]^{d_{x}}$. This condition is without loss of generality if $\mathcal{X}$ is a compact subset of $\mathbb{R}^{d_{x}}$, and is routinely imposed in nonparametric econometrics. Assumption \ref{as:covariates}.2 imposes that the covariates are i.i.d, a standard assumption in the partial identification literature. Assumption \ref{as:covariates}.3 requires that the marginal distribution of $X_{i}$ has a density which is uniformly bounded away from zero and infinity on $[0,1]^{d_{x}}$, a condition that is satisfied, for instance, if $\mathcal{X}$ is compact, and $p_{0,X}$ is continous and positive everywhere. Assumption \ref{as:GP}.1 requires that the sample paths of the Gaussian process $B_{k}$ are almost $\alpha_{k}$ regular in the sense that they take values in the H\"{o}lder space $C^{a_{k}}([0,1]^{d_{x}})$. Common Gaussian processes (in particular, the Mat\'{e}rn processes from Section \ref{sec:GPimplementation}) satisfy this condition \citep{vaart2008rates,JMLR:v12:vandervaart11a}. Assumption \ref{as:GP}.2 imposes that the sample paths are uniformly bounded. It is a technical condition because $\overline{B}_{k}$ permitted to be arbitrarily large, but finite, and, for this reason, can be ignored in practice. Uniform boundedness restrictions on GP priors are often imposed when deriving strong norm posterior concentration guarantees \citep{gine2011rates,norets2015bayesian,walker2026semiparametric}.\footnote{Lemma 5.1 of \cite{van2008reproducing} guarantees that the event $\{||B_{k}||_{\infty} \leq \bar{B}_{k}\}$ receives positive measure under the probability law of the Gaussian process.}
We start with an intermediate empirical mean square posterior concentration result. For notation, let $(\mathcal{H}_{k},||\cdot||_{\mathcal{H}_{k}})$ be the reproducing kernel Hilbert space (RKHS) attached to $B_{k}$, let $\overline{\mathcal{H}}_{k}$ be the closure of $\mathcal{H}_{k}$ in $(C([0,1]^{d_{x}},\mathbb{R}),||\cdot||_{\infty})$, and let $$\varphi_{k,b_{0,k}}(\delta) = \inf_{f_{k} \in \mathcal{H}_{k}: ||f_{k}-b_{0,k}||_{\infty}< \delta}\frac{1}{2}||f_{k}||_{\mathcal{H}_{k}} - \log P_{GP(0,\kappa_{k})}(||B||_{k}< \delta)$$ for $\delta > 0$ be the concentration function of $GP(0,\kappa_{k})$ at $b_{0,k}$, where $b_{0,k} = \Lambda^{-1}(q_{0,k})$ for each $k=1,...,K-1$.
\begin{proposition}\label{prop:L2}
Suppose Assumptions \ref{as:dgp}.1, \ref{as:covariates}.1, and \ref{as:GP}.2 hold, $b_{0,k} \in \overline{\mathcal{H}}_{k}$ and $||b_{0,k}||_{\infty} \leq \bar{B}_{k}$ for each $k \in \{1,...,K-1\}$, and, for each $k \in \{1,...,K-1\}$, there exists a sequence $\{\tilde{\delta}_{n,k}\}_{n \geq 1}$ with $\tilde{\delta}_{n,k} \rightarrow 0$ as $n\rightarrow \infty$, $n\tilde{\delta}_{n,k}^{2} \rightarrow \infty$ as $n\rightarrow \infty$, and $\varphi_{k,b_{0,k}}(\tilde{\delta}_{n,k}) \leq n \tilde{\delta}_{n,k}^{2}$. Then for $P_{0,X}^{(\infty)}$-almost every fixed realization $\{x_{i}\}_{i \geq 1}$ of $\{X_{i}\}_{i \geq 1}$,
\begin{align*}
\Pi_{n}\left(||p-p_{0}||_{n,2}\geq M\tilde{\delta}_{n}\middle | Y^{(n)}\right) \overset{P_{0}^{(n)}}{\longrightarrow} 0
\end{align*}
as $n\rightarrow \infty$, where $M>0$ is a large constant and $\tilde{\delta}_{n} = \max_{1 \leq k \leq K-1}\tilde{\delta}_{n,k}$.
\end{proposition}
Proposition \ref{prop:L2} states that if $b_{0,k}$ can be uniformly approximated by sequences in the RKHS $\mathcal{H}_{k}$, then the posterior concentrates around empirical $L^{2}$ balls centered at $p_{0,k}$ (and the rate of convergence is quantified by $\varphi_{b_{0,k}}$). It can be viewed as an extension of the binary regression results in \cite{vaart2008rates} (i.e., Theorem 3.2) to categorical regression, and, since it is based on the logit stick breaking Gaussian process prior, it may be of independent general interest, as we are not aware of formal posterior contraction rate guarantees for these priors.\footnote{\cite{pati2013posterior} and \cite{norets2014posterior} derive posterior consistency results (without rates) in the related context of covariate dependent stick breaking models for conditional densities.} In some cases, these posteriors contract at the optimal rate. For example, if $\kappa_{k}$ belongs to the Mat\'{e}rn class with regularity parameter $\alpha_{k} > 0$ and $b_{0,k}$ is in the intersection of a H\"{o}lder and a Sobolev space of order $\alpha_{0,k}> 0$, then Theorem 5 of \cite{JMLR:v12:vandervaart11a} implies that $\tilde{\delta}_{n,k} = n^{-\min\{\alpha_{k},\alpha_{0,k}\}/(2\alpha_{k}+d_{x})}$ for each $k=1,...,K-1$. Consequently, when the prior and true smoothness match (i.e., $\alpha_{k} = \alpha_{0,k}$ for each $k=1,...,K-1$), we conclude that $\tilde{\delta}_{n} = n^{-\alpha_{0}/(2\alpha_{0}+d_{x})}$, where $\alpha_{0} = \min\{\alpha_{0,1},...,\alpha_{0,K-1}\}$, thereby achieving the optimal rate from \cite{stone1982optimal}.
Proposition \ref{prop:post_sup_norm} establishes that, when combined with smoothness restrictions on $b_{0,k}$ and $B_{k}$, posterior consistency in the empirical $L^{2}$ distance (Proposition \ref{prop:L2}) implies posterior consistency in the empirical supremum norm. The argument is based on wavelet interpolation and follows a similar logic to Proposition 2 of \cite{walker2026semiparametric}, except that it applies to a different class of priors (i.e., the logistic stick-breaking GP prior rather than an infinite-dimensional exponential family). The only additional condition, Assumption \ref{as:wavelet}, is a condition on a wavelet basis, and, since its exposition is somewhat cumbersome, we defer it to the Appendix. For Mat\'{e}rn GPs with $\alpha_{k} = \alpha_{0,k}$ for each $k=1,...,K-1$, the condition holds if $\min_{1 \leq k \leq K-1}\alpha_{k}>(1+\sqrt{5})d_{x}/4$, which is a slight strengthening of the baseline restriction $\min_{1 \leq k \leq K-1}\alpha_{k} > d_{x}/2$ in Assumption \ref{as:GP} (see Appendix \ref{ap:verifyingwavelet} for more discussion).
\begin{proposition}\label{prop:post_sup_norm}
Suppose that Assumptions \ref{as:dgp}.1, \ref{as:covariates}, \ref{as:GP}, and \ref{as:wavelet} hold, the conditions of Proposition \ref{prop:L2} hold, and for each $k \in \{1,...,K-1\}$, there exists $\alpha_{0,k} > d_{x}/2$ such that $b_{0,k} \in C^{\alpha_{0,k}}([0,1]^{d_{x}},\mathbb{R})$. Then, for $P_{0,X}^{(\infty)}$-almost every fixed realization $\{x_{i}\}_{i \geq 1}$ of $\{X_{i}\}_{i \geq 1}$, the following holds: for every $\varepsilon > 0$,
\begin{align*}
\Pi_{n}\left(||p-p_{0}||_{n,\infty} \geq \varepsilon \middle | Y^{(n)}\right) \overset{P_{0}^{(n)}}{\longrightarrow} 0
\end{align*}
as $n\rightarrow \infty$.
\end{proposition}
\section{Extensions}\label{sec:extensions}
This section extends our framework to settings with aggregated discrete responses and continuous endogenous variables.
\subsection{Aggregated Discrete Responses.} In some economic settings (e.g., empirical industrial organization), the researcher may only have access to aggregated discrete outcome over some units (e.g., market shares). Our framework readily extends to this setting. For concreteness, suppose that the observed data is $(Y',X')'$, where $X = (\tilde{X}',M)$ is a vector of observed covariates $\tilde{X}$ and an exogenous number of units $M$ over which the aggregation is performed (e.g., market size), $Y = M\tilde{Y} \in \mathbb{N}^{K}$ is the vector of counts, and $\tilde{Y}$ is a $K$-dimensional random vector that takes values in the probability simplex (i.e., shares). Assuming the aggregated units are independent, the reduced-form model is
\begin{align*}
Y_{i}|p \overset{ind}{\sim} Multinomial(m_{i},p(\tilde{x}_{i})), \quad i=1,...,n, \quad p \sim \Pi,
\end{align*}
where $Multinomial(m,p(\tilde{x}))$ is the multinomial distribution based on $m$ trials and event probabilities $p(\tilde{x})$, and $\Pi$ is the same prior over $p$ as before. Consequently, by defining $\Gamma_{n,I}(p) = \{\gamma \in \Gamma: f(x_{i},p(\tilde{x}_{i}),\gamma) \in \mathbb{R}^{d_{f}}_{+} \ \forall \ i=1,...,n\}$ (i.e., the same as before except acknowledging the partition $x_{i}=(\tilde{x}_{i},m_{i})$), a posterior for $(p(\tilde{x}_{1}),...,p(\tilde{x}_{n}))$ implies a posterior for $\Gamma_{n,I}(p)$. This demonstrates that, conceptually, there is virtually no difference between our main setup and the aggregated discrete choice setting.
Similarly, the stick-breaking technique used for our implementation readily extends to this setting. Using the stick-breaking weights, $p_{k} = q_{k}\prod_{j < k}(1-q_{j})$ for $k=1,...,K-1$ and $p_{K} = \prod_{j=1}^{K-1}(1-q_{j})$, the multinomial likelihood $L_{n}(p)$ satisfies $L_{n}(p) = L_{n}^{*}(q)$ with
\begin{align*}
L_{n}^{*}(q) = \prod_{k=1}^{K-1}\prod_{i \in \mathcal{I}_{k}}{m_{i,k} \choose Y_{i,k}}q_{k}(\tilde{x}_{i})^{Y_{i,k}}(1-q_{k}(\tilde{x}_{i}))^{m_{i,k}-Y_{i,k}},
\end{align*}
where $m_{i,k} = m_{i} - \sum_{j < k}Y_{i,j}$, $Y_{i,k}$ is the $k$th element of $Y_{i}$, and $\mathcal{I}_{k} = \{i: m_{i,k} > 0\}$. Notice that, for each $k$, the factor is the density of $|\mathcal{I}_{k}|$ independent Binomial distributions with $m_{i,k}$ trials. Consequently, by specifying the law of $(q_{1},...,q_{K-1})$ to be that of $(\Lambda(B_{1}),...,\Lambda(B_{K-1}))$, where $B_{k} \overset{ind}{\sim} GP(\mu_{k},\kappa_{k})$ for each $k=1,...,K-1$, the binomial version of \cite{polson2013bayesian}'s P\'{o}lya-Gamma data augmentation can be used to sample from $\pi_{n,k}(B_{\mathcal{I}_{k},k}|Y^{(n)})$ (see Appendix \ref{ap:posteriordraws}). The other implementation aspects (e.g., post-processing to sample from $\pi_{n,k}(B_{\mathcal{I}_{k}^{c},k}|Y^{(n)},B_{\mathcal{I}_{k},k})$ and computing $\Gamma_{n,I}(p)$) do not change because they do not depend on the likelihood.
Since the reduced-form parameter $p$ is unchanged in the aggregated setting, the only possible concern for the asymptotic theory is whether this extension affects posterior concentration of $p$ around $p_{0}$. That is, whether the results from Section \ref{sec:GPverification} extends. The next result states that if $\{m_{i}\}_{i \geq 1}$ forms a bounded sequence, then Propositions \ref{prop:L2} and \ref{prop:post_sup_norm} are unaffected by the aggregation. Consequently, under bounded $M$, aggregated discrete choice settings are both practically and theoretically accommodated.
\begin{proposition}\label{prop:GPmultinomial}
Suppose that the true data-generating process satisfies $Y_{i}|\{X_{j}\}_{j \geq 1} \overset{ind}{\sim} Multinomial(M_{i},p_{0}(\tilde{X}_{i}))$ for every $i \geq 1$, $\{\tilde{X}_{i}\}_{i \geq 1}$ satisfies Assumption \ref{as:covariates}, and $\{M_{i}\}_{i \geq 1}$ is bounded with $P_{0,X}^{(\infty)}$-probability equal to one. Further, suppose that Assumptions \ref{as:GP} and \ref{as:wavelet} hold, $b_{0,k} \in \overline{\mathcal{H}}_{k} \cap C^{\alpha_{0,k}}([0,1]^{d_{x}},\mathbb{R})$ for some $\alpha_{0,k} > d_{x}/2$ and $||b_{0,k}||_{\infty} \leq \bar{B}_{k}$ for each $k \in \{1,...,K-1\}$, and, for each $k \in \{1,...,K-1\}$, there exists a sequence $\{\tilde{\delta}_{n,k}\}_{n \geq 1}$ with $\tilde{\delta}_{n,k} \rightarrow 0$ as $n\rightarrow \infty$, $n\tilde{\delta}_{n,k}^{2} \rightarrow \infty$ as $n\rightarrow \infty$, and $\varphi_{k,b_{0,k}}(\tilde{\delta}_{n,k}) \leq n \tilde{\delta}_{n,k}^{2}$. Then, for $P_{0,X}^{(\infty)}$-almost every fixed realization $\{x_{i}\}_{i \geq 1}$ of $\{X_{i}\}_{i \geq 1}$, the following is true: for every $\varepsilon > 0$,
\begin{align*}
\Pi_{n}\left(||p-p_{0}||_{n,\infty} \geq \varepsilon \middle | Y^{(n)}\right) \overset{P_{0}^{(n)}}{\longrightarrow} 0
\end{align*}
as $n\rightarrow \infty$, where $P_{0}^{(n)} = \bigotimes_{i=1}^{n}Multinomial(m_{i},p_{0}(\tilde{x}_{i}))$
\end{proposition}
\subsection{Continuous Response Variables.} Some important partially identified models involve continuously distributed $Y$, such as linear regression models with interval censoring \citep{manski2002inference}, incomplete auction models \citep{haile2003inference}, and models of exporter choice \citep{dickstein2018exporters}, to name a few. Our framework extends to this setting. First, we redefine the reduced-form model as
\begin{align*}
Y_{i}|p_{Y|X}\overset{ind}{\sim} p_{Y|X}(\cdot|x_{i}), \quad i=1,...,n, \quad p_{Y|X} \sim \Pi,
\end{align*}
where $p_{Y|X}$ is the conditional probability density function of $Y|X$ and $\Pi$ is a prior for $p_{Y|X}$. Then, we redefine the conditional identified set as
\begin{align*}
\Gamma_{n,I}(p_{Y|X}) = \left\{\gamma \in \Gamma: f(x_{i},p_{Y|X}(\cdot|x_{i}),\gamma) \in \mathbb{R}_{+}^{d_{f}} \ \forall \ i=1,...,n\right\}.
\end{align*}
Subject to regularity conditions, our Bayesian inference framework extends: the posterior for $(p_{Y|X}(\cdot|x_{1}),...,p_{Y|X}(\cdot|x_{n}))$ implies a posterior for $\Gamma_{n,I}(p_{Y|X})$ via the mapping $(p_{Y|X}(\cdot|x_{1}),...,p_{Y|X}(\cdot|x_{n})) \mapsto \Gamma_{n,I}(p_{Y|X})$. Importantly, by defining $f(x_{i},p_{Y|X}(\cdot|x_{i}),\gamma) = \int g(y,x_{i},\gamma)p_{Y|X}(y|x_{i})dy$, this nests conditional moment inequalities with continuous $Y$.
Since the consistency theory only uses the finite support of $Y$ in Section \ref{sec:GPverification}, some of our posterior consistency results readily extend to continuous $Y$. Suppose that the posterior for $p_{Y|X}$ concentrates on $V_{n,\delta_{n}}^{dens} :=\mathcal{P}_{n,Y|X} \cap \{p_{Y|X}: e_{n}(p_{Y|X},p_{0,Y|X}) < \delta_{n}\}$, where $e_{n}$ is a semimetric over conditional densities and $\mathcal{P}_{n,Y|X}$ is a sieve, and that these sets are structured enough to ensure that
\begin{align}
\sup_{(\gamma,p) \in \Gamma \times V_{n,\delta_{n}}^{dens}}\max_{1 \leq i \leq n}||f(x_{i},p_{Y|X}(\cdot|x_{i}),\gamma)-f(x_{i},p_{0,Y|X}(\cdot|x_{i}),\gamma)||_{2}\longrightarrow 0 \label{eq:continuous_ucon1}\\
\sup_{(\gamma,p) \in \Gamma \times V_{n,\delta_{n}}^{dens}}\max_{1 \leq i \leq n}||W(x_{i},p_{Y|X}(\cdot|x_{i}),\gamma)-W(x_{i},p_{0,Y|X}(\cdot|x_{i}),\gamma)||_{2}\longrightarrow 0 \label{eq:continuous_ucon2}
\end{align}
as $n\rightarrow \infty$. Then, a continuous $Y$ extension of Assumption \ref{as:prior} for convex $\Gamma_{n,I}(p_{Y|X})$ and $Q_{n,r}(\gamma,p_{Y|X}) = ||\text{dist}_{W(\cdot,p_{Y|X}(\cdot|\cdot),\gamma)}(f(\cdot,p_{Y|X}(\cdot|\cdot),\gamma),\mathbb{R}_{+}^{d_{f}})||_{n,r}$, $r \in [1,\infty]$, holds so long as the true density $p_{0,Y|X}$ satisfies analogous conditions to $p_{0}$ in Propositions \ref{prop:ucon} and \ref{prop:convexhemi}.\footnote{Continuous $Y$ extensions of Propositions \ref{prop:taxicab}, \ref{prop:minorantmisspec}, and \ref{prop:misspec_hemi_convex} can be achieved. For conciseness, we omit them.} Proposition \ref{prop:continuousY} below formalizes this (with Assumption \ref{as:continuousDGP} in Appendix \ref{ap:extensionproofs} stating conditions on $p_{0,Y|X}$). Consequently, if the DGP satisfies $Y_{i}|\{X_{j}\}_{j \geq 1} \overset{ind}{\sim}p_{0,Y|X}$ and $Q_{n,r}(\gamma,p_{0,Y|X})$ is well-separated at $\Gamma_{n,I}(p_{0,Y|X}) \neq \emptyset$ (i.e., a modification of Assumption \ref{as:dgp} holds), then a virtually identical argument to Theorem \ref{thm:consistency} establishes posterior consistency for $\Gamma_{n,I}(p_{Y|X})$ at $\Gamma_{n,I}(p_{0,Y|X})$ in the Hausdorff distance. For conditional moment inequalities with bounded $g(Y,X,,\gamma)$ and $W(x,p_{Y|X}(\cdot|x),\gamma)$ based on $Var(g(Y,X,\gamma)|X=x)$, Proposition 2 of \cite{walker2026semiparametric} provides sufficient conditions for (\ref{eq:continuous_ucon1}) and (\ref{eq:continuous_ucon2}) under GP priors for $p_{Y|X}$.
\begin{proposition}\label{prop:continuousY}
Suppose that, for $P_{0,X}^{(\infty)}$-almost every fixed realization $\{x_{i}\}_{i \geq 1}$ of $\{X_{i}\}_{i \geq1}$, there exists sets $\{\mathcal{P}_{n,Y|X}\}_{n \geq 1}$, metrics $\{e_{n}\}_{n \geq 1}$, and sequences $\{\delta_{n}\}_{n \geq 1}$ such that $\delta_{n} \rightarrow 0$ as $n\rightarrow \infty$, and
\begin{align*}
\Pi_{n}(p_{Y|X} \in \mathcal{P}_{n,Y|X}|Y^{(n)}) &= 1+o_{P_{0,Y|X}^{(n)}}(1) \\
\Pi_{n}(e_{n}(p_{Y|X},p_{0,Y|X})<\delta_{n}|Y^{(n)}) &= 1+o_{P_{0,Y|X}^{(n)}}(1)
\end{align*}
Moreover, suppose that the criterion is $Q_{n,r}(\gamma,p) = ||\text{dist}_{W(\cdot,p_{Y|X}(\cdot|\cdot),\gamma)}(f(\cdot,p_{Y|X}(\cdot|\cdot),\gamma),\mathbb{R}_{+}^{d_{f}})||_{n,r}$ for some $r \in [1,\infty]$, and, for $P_{0,X}^{(\infty)}$-almost every fixed realization $\{x_{i}\}_{i \geq 1}$ of $\{X_{i}\}_{i \geq1}$,
\begin{align*}
\sup_{(\gamma,p) \in \Gamma \times V_{n,\delta_{n}}^{dens}}\max_{1 \leq i \leq n}||f(x_{i},p_{Y|X}(\cdot|x_{i}),\gamma)-f(x_{i},p_{0,Y|X}(\cdot|x_{i}),\gamma)||_{2}\longrightarrow 0 \\
\sup_{(\gamma,p) \in \Gamma \times V_{n,\delta_{n}}^{dens}}\max_{1 \leq i \leq n}||W(x_{i},p_{Y|X}(\cdot|x_{i}),\gamma)-W(x_{i},p_{0,Y|X}(\cdot|x_{i}),\gamma)||_{2}\longrightarrow 0
\end{align*}
as $n\rightarrow \infty$. If Assumption \ref{as:continuousDGP} holds, then Assumption \ref{as:prior} holds (with $p_{Y|X}$ in place of $p$).
\end{proposition}
\section{Conclusion}\label{sec:conclusion}
This paper develops a nonparametric Bayesian approach to inference in partially identified discrete response models. The central observation is simple: in a large class of models, the identified set is a functional of an unrestricted reduced-form conditional choice probability. Rather than placing prior structure on unidentified structural parameters, we place a flexible prior on this identifiable reduced form parameter and map its posterior into a posterior for the identified set. This delivers a fully Bayesian procedure for set-valued parameters while keeping the source of statistical uncertainty transparent. The extensions demonstrate that our insights paper also apply to aggregated discrete choice and continuous outcomes.
A main advantage of this approach relative to existing methods is that it works directly with the conditional restrictions that define the model. In conditional moment inequality models, there is no need to replace conditional moments with a user-chosen collection of unconditional moments or instruments. More generally, the procedure does not require discretizing continuously distributed covariates simply to make the problem finite dimensional. These steps are common in empirical implementations, but they can discard identifying information and make inference sensitive to choices that are external to the economic model. By learning the conditional probability mass function directly, our framework accommodates rich covariate variation while preserving the original identifying restrictions. In this sense, it combines the flexibility typically associated with nonparametric frequentist methods with a genuinely Bayesian treatment of uncertainty about the identified set.
Implementation is also straightforward. With the logistic stick-breaking Gaussian process priors considered here, P\'{o}lya--Gamma data augmentation yields conditionally Gaussian posterior updates, and the stick-breaking components can be sampled independently and in parallel. Given a conditional choice probability, draws of the identified set are obtained by solving the same optimization problem that would be used if those probabilities were known, making identified set computation a parallelizable post-processing step.
Our general asymptotic theory shows that our proposal has a clear large-sample interpretation: under correct specification, the posterior concentrates on the true identified set; under misspecification, the posterior probability of an empty identified set provides a diagnostic, while inference can instead be conducted on a pseudo-identified set. The low-level analysis in Section \ref{sec:verifyingconditions} highlights that our general asymptotic theory applies to the logistc stick-breaking Gaussian process priors. Taken together, these results provide a practical route to fully Bayesian inference in partially identified models without requiring researchers to coarsen covariates or alter the conditional identifying content of the model.
\bibliographystyle{ecca}
\bibliography{multinomial}