EconBase
← Back to paper

An Adversarial Approach to Identification

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

118,872 characters · 19 sections · 85 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

An Adversarial Approach to Identification

frontmatter\runtitle{An Adversarial Approach to Identification} \begin{aug} \address[id=add1]{ \orgdiv{Department of Economics}, \orgname{McMaster University}} \address[id=add2]{ \orgdiv{Department of Economics}, \orgname{University of North Carolina Wilmington}} \end{aug} \support{We thank Bo Honor\'e, Jiaying Gu, Hide Ichimura, and Jim Powell for discussions and suggestions. Botosaru gratefully acknowledges financial support from the Canada Research Chairs Program. This paper was presented at the University of Arizona in March 2024 and the Southern Economic Association in November 2024.} \begin{abstract} We introduce a new framework for characterizing identified sets of structural and counterfactual parameters in econometric models. By reformulating the identification problem as a set membership question, we leverage the separating hyperplane theorem in the space of observed probability measures to characterize the identified set through the zeros of a discrepancy function with an adversarial game interpretation. The set can be a singleton, resulting in point identification. A feature of many econometric models, with or without distributional assumptions on the error terms, is that the probability measure of observed variables can be expressed as a linear transformation of the probability measure of latent variables. This structure provides a unifying framework and facilitates computation and inference via linear programming. We demonstrate the versatility of our approach by applying it to nonlinear panel models with fixed effects, with parametric and nonparametric error distributions, and across various exogeneity restrictions, including strict and sequential. \end{abstract} \begin{keyword} \kwd{partial identification} \kwd{nonlinear panel models} \kwd{counterfactual parameters} \kwd{linear programming} \end{keyword}

Introduction

Identification of structural and counterfactual parameters is a central challenge in econometric models with unobserved heterogeneity. In many cases, the distribution of unobserved heterogeneity is not point identified, leading to partial identification of structural and/or counterfactual parameters. This issue is pervasive in nonlinear panel models with fixed effects. Fixed effects obstruct the point identification of both structural parameters and partial effects in all but a narrow class of models ArellanoBonhomme2012. This highlights the need for methods that achieve sharp identification for both types of parameters, while remaining computationally feasible and enabling valid inference.

Addressing these issues in nonlinear panel models is challenging. Many existing approaches are tailored to specific model features, relying on parametric assumptions about error distributions or support and exogeneity restrictions on the covariates. Additionally, existing methods focus on either structural or counterfactual parameters. This has led to a fragmented literature with various methods addressing only isolated aspects of the broader identification problem.

We introduce a novel and unifying framework for characterizing identified sets of structural and counterfactual parameters in econometric models with unobserved heterogeneity. Departing from existing approaches, we reformulate the identification problem as a set membership question in the space of observed probability measures (i.e., probability measures of observed random variables). This reformulation allows us to leverage the separating hyperplane theorem in the infinite-dimensional space of observed probability measures, providing a characterization of the identified set via the zeros of a new discrepancy function. The discrepancy function checks whether there exists at least one hyperplane that separates the observed probability measure from the set of model probabilities, defined as the collection of all probability measures of observed random variables consistent with a given parameter value. By aggregating over all hyperplanes and all measures in the set of model probabilities, the discrepancy function reveals whether the observed measure belongs to the set of model probabilities. The discrepancy function has a maximin representation or adversarial game interpretation, inspiring the name of our approach.\footnote{Our approach is distinct from the simulation-based adversarial estimation method of Kaji2023. Ours is an identification approach.}

We establish new sufficient and necessary conditions for sharp identification: when the set of model probabilities is convex, our approach obtains a sharp characterization of the identified set; when convexity fails, our approach characterizes an outer set. When the identified set is a singleton, we obtain point identification. Remarkably, many econometric models naturally feature a convex set of model probabilities.

A key innovation of our framework is its ability to accommodate both parametric and nonparametric error distributions, along with a wide range of exogeneity restrictions and conditioning variables --- whether continuous or discrete, strictly exogenous or predetermined. To demonstrate the power of our approach, we characterize the identified set for structural and counterfactual parameters in the semiparametric binary choice panel model under sequential exogeneity (or “predeterminedness”) and without parametric assumptions on the distribution of the error terms, addressing a long-standing gap in the literature on nonlinear panel models. We further highlight the versatility of our method by applying it to the binary choice panel model with error terms that follow a fixed but arbitrary distribution. For the nested case of logistic errors, we recover established results for the structural parameter and recently derived results for counterfactual parameters. These applications illustrate the versatility of our framework and its potential to advance partial identification in nonlinear econometric models.

Let $Z$ denote the observed random variables, and \(\mu^*_Z\) the observed probability measure characterizing the distribution of \(Z\). Let $\theta\in\Theta$ denote the parameter of interest, which can include both structural parameters and arbitrary functionals of the latent probability measure, and $\overline{\mathcal{M}}_\theta$ denote the closure of the set of model probabilities.\footnote{ Explicitly taking the closure ensures the inclusion of limit measures, such as those arising from degenerate distributions when latent random variables reach extreme values (e.g., fixed effects approaching \(\pm \infty\)). } We reformulate the identification problem as a set membership question: For a candidate parameter value \(\theta\), does \(\mu^*_Z\) belong to \(\overline{\mathcal{M}}_\theta\)? Accordingly, the identified set for \(\theta\) is given by:

align[align omitted — 132 chars of source]

This set contains all parameter values compatible with \(\mu^*_Z\). If \(\mu^*_Z \in \overline{\mathcal{M}}_\theta\), the observed probability measure and the model probabilities are indistinguishable at \(\theta\); otherwise, \(\theta\) is incompatible with \(\mu^*_Z\).

To determine membership of $\mu^*_Z$ in $\overline{\mathcal{M}}_\theta$, we leverage the separating hyperplane theorem. Specifically, when $\mu_Z^*\notin\overline{\mathcal{M}}_\theta$ and $\mathcal{M}_\theta$ is convex, there exists a hyperplane $\phi$ that separates $\mu_Z^*$ from $\mathcal{M}_\theta$. Aggregating across all hyperplanes $\phi$ leads to the following discrepancy function:

align[align omitted — 222 chars of source]

where $\mathcal{Z}$ denotes the support of $Z$ and $\Phi_b(\mathcal{Z})$ a set of bounded functions supported on $\mathcal{Z}$, defined in Section (ref).

The discrepancy function $T(\theta)$ evaluates to zero if and only if no separating hyperplane $\phi$ exists. Consequently, the zeros of $T(\theta)$ can be used to characterize $\Theta_\mathrm{I}$. This characterization is sharp provided that $\mathcal{M}_\theta$ is convex --- a property shared by many econometric models. When $\mathcal{M}_\theta$ is not convex, the zeros of $T(\theta)$ describe an outer set. When $\Theta_\mathrm{I}$ is a singleton, our framework obtains point identification of $\theta$.

In many econometric models, the discrepancy function admits a low-dimensional representation through an extreme point characterization. This facilitates the computation of zeros of $T(\theta)$ using linear programming. Two features of many models are central to this result: (i) the observed probability measure can be expressed as a linear transformation of the latent probability measure (i.e., the measure of latent random variables),\footnote{In semiparametric models with unrestricted error distributions, this transformation corresponds to a pushforward measure, whereas in models with parametric error distributions, it takes the form of a linear operator.} and (ii) the latent probability measures lie within a convex set.

For example, the semiparametric binary choice panel model under strict exogeneity exhibits both features, allowing for efficient computation of the identified set for the structural and counterfactual parameters using linear programming. Conversely, the semiparametric binary choice panel model under sequential exogeneity satisfies the linearity condition but not the convexity of the set of latent probabilities; instead, this model features a convex set of model probabilities $\mathcal M_\theta$, which yields sharp identification. However, computing the identified set in this case requires an extension of our linear programming approach. Beyond these computational advantages, the two features enable the sample analog of the discrepancy function to serve as a test statistic, facilitating valid inference. We establish the asymptotic distribution of this test statistic and construct critical values for hypothesis tests that are uniformly valid across the underlying probability distributions and parameter values.

A few key features distinguish our proposed method. The first distinguishing feature underscores the simplicity of our method in establishing sharp identification. A sufficient condition for sharpness is the convexity of $\mathcal M_\theta$. This can be established directly, or, given the linearity of the transformation, via the convexity of the set of latent probability measures. The latter holds in two key cases: (i) when no assumptions are imposed on the latent probability measure, such as in panel models with no distributional assumptions on the error terms, and (ii) when the latent probability measure is required to satisfy a finite number of linear restrictions. Such linear restrictions typically arise in two contexts: (i) when analyzing counterfactual parameters, many of which can be expressed as linear functionals of the latent probability measure,\footnote{See, e.g., ChristensenConnault2023, torgovitsky2019.} and (ii) when imposing exogeneity conditions, such as zero-mean or zero-median constraints on the error terms. We show this in our examples.

The second distinguishing feature is the versatility and broad applicability of our approach. Our method applies to many econometric models with unobserved heterogeneity, regardless of whether the models involve error terms that follow parametric distributions, or outcomes and covariates that are discrete or continuous. Our method accommodates various exogeneity restrictions on the covariates, and, if desired, restrictions on the latent probability measures.

The third distinguishing feature of our method is its computational ease. We implement our procedure and examine the size of the identified set in the canonical semiparametric binary choice model without parametric restrictions on the error terms, both in the cross-sectional case and in the panel case with strictly exogenous regressors, time effects, and fixed effects. With predetermined regressors, although the set of latent measures is not convex, the set of model probabilities is, so the identified set can be computed via an extension of our linear programming method. These illustrations to the semiparametric binary choice panel model with strictly exogenous or predetermined covariates contribute new results to the nonlinear panel literature.

The fourth distinguishing feature is that the sample analog of the discrepancy function can be used as a test statistic for valid inference.

Related literature

This paper contributes to the literature on sharp identification in general classes of models and to the literature on nonlinear panel models.

A wide range of methods has been developed to characterize identified sets across different models.\footnote{For reviews, see BontempsMagnac2017,CanayShaikh2017,Molinari2020,ChesherRosen2020,KlineTamer2023.} Approaches include the use of random set theory beresteanuSharpIdentificationRegions2011, ChesherRosen2017, optimal transport GalichonHenry2009,EkelandGalichonHenry2010,GalichonHenry2011, and information-theoretic methods schennachEntropicLatentVariable2014. Other contributions, such as torgovitsky2019, focus on extending subdistributions, while more recent work explores minimum relevant partition and latent space enumeration Tebaldi2023,GuRussellStringham2022. Some of these methods focus exclusively on complete models, while others are explicitly designed to also address incomplete models.\footnote{In particular, models where the relationship between the observed random variables and the latent random variables is a correspondence, e.g., Tamer2003.} Many of these methods, like ours, leverage convex analysis for sharp identification or low-dimensional representations.

Our framework differs by reformulating the identification problem as a set membership question in the space of observed probability measures, treating \(\mathcal{M}_\theta\) as primitive. This allows us to (i) apply the separating hyperplane theorem in the space of observed probability measures, (ii) link sharpness to the convexity of \(\mathcal{M}_\theta\) (with outer sets characterized when convexity fails), and (iii) exploit two features common to many econometric models: the linearity of the transformation between observed and latent probability measures and the convexity of the set of latent probability measures. These properties facilitate the computation of the identified set via linear programming. This structure enables a unified approach for structural and counterfactual parameters, accommodating both parametric and nonparametric error distributions. This versatility is noteworthy. For example, methods designed for nonparametric error distributions rarely handle parametric restrictions (e.g., guDualApproachWassersteinRobust2023, ChesherRosenZhang), while parametric approaches often rely on specific distributional assumptions (e.g., Bonhomme2012,daveziesFixedEffectsBinary2023,GuRussellStringham2022).

While we focus on models where the relationship between observed and latent random variables is a mapping rather than a correspondence, our results can be extended to models involving correspondences. This is possible since our approach operates in the space of observed probability measures, and a sufficient assumption for all our results is that the observed probability measure is a linear map of the latent probability measure. Correspondences between random variables can induce such linear maps. However, as we show in the paper, such linear maps often arise in models commonly used in the literature on nonlinear panel models. Since our methodology is tailored to address a long-standing gap in the literature on nonlinear panel models, we leave a detailed investigation of correspondence-defined models for future research.

Identification challenges have long been a hallmark of nonlinear panel models due to the presence of fixed effects, see, e.g., arellanoPanelDataModels2001, honoreBoundsParametersPanel2006. As ArellanoBonhomme2012 note, point identification of structural parameters is rare, and even then, functionals of the fixed effects distribution, such as partial effects, remain only partially identified.\footnote{For exceptions of point-identified counterfactual parameters in specific models, see honoreMarginalEffectsSemiparametric2008, aguirregabiriaIdentificationAverageMarginal2024, and danoTransitionProbabilitiesIdentifying2023.} Addressing dynamic exogeneity in nonlinear models remains an open challenge, especially when both structural and counterfactual parameters are of interest, see HonoredePaula2021, ArkhangelskyImbens.

We showcase adversarial identification by applying it to the semiparametric binary choice panel model (see Manski1987). We obtain new results for the identified set for both structural parameters and the average structural function (ASF), without imposing parametric restrictions on the distribution of the error terms and under various dynamic exogeneity assumptions, such as strict and sequential exogeneity.

Most work on the semiparametric binary choice panel model focuses on the identification of structural parameters; for a non-exhaustive list, see Manski1987, aristodemouSemiparametricIdentificationPanel2021, khanIdentificationDynamicBinary2023, botosaruIdentificationTimevaryingTransformation2021, mbakopIdentificationDiscreteChoice2023, gaoIdentificationNonlinearDynamic2024, ChesherRosenZhang. The identification of partial effects has received less attention: ChernValHahnNewey2013 derive results under time-homogeneity and restrictions on the support of the covariates; BotosaruMuris2024 relax the latter assumptions while assuming that the structural parameters are a priori either point- or partially-identified.

When the error distribution is fixed but arbitrary, the structural parameters can be point-identified in a narrow class of models, see Bonhomme2012. Recent work focuses on partial identification of counterfactual parameters starting from point-identified or $\sqrt{n}$-estimable structural parameters, see dobronyiIdentificationDynamicPanel2021,daveziesFixedEffectsBinary2023,pakelBoundsAverageEffects2024. Seminal contributions by honoreBoundsParametersPanel2006 and ChernValHahnNewey2013 treat partial identification of structural parameters and average treatment effects, relying on discrete outcomes and covariates, and support conditions on the fixed effects.

Even with parametric restrictions on the error terms, identification of both structural and counterfactual parameters under sequential exogeneity is challenging, cf. arellanoPanelDataModels2001, ArellanoBonhomme2012, and HonoredePaula2021. Results for the identification of structural parameters under sequential exogeneity and error terms that follow parametric distributions can be found in arellanoBinaryChoicePanel2003, piginiConditionalInferenceBinary2022, chamberlainIdentificationDynamicBinary2023, and for partial effects in bonhommeIdentificationBinaryChoice2023.

In contrast, our approach obtains the identified set of structural and counterfactual parameters for models with known or unknown error distributions, does not require functional form restrictions such as an index structure or additivity in the unobservables, time-homogeneity, or support restrictions on the covariates or on the fixed effects. This flexibility extends to models with various forms of dynamic exogeneity --- strict, sequential, or lagged, offering new insights for nonlinear panel models. To the best of our knowledge, ours are the only identification results in nonlinear panel models with sequential exogeneity and without parametric restrictions on the error terms.

Organization.

We introduce our main result in Section (ref). This result applies to a very general class of models. We specialize the result to semiparametric models with unobserved heterogeneity in Section (ref) and with observed covariates in Section (ref). We also illustrate our approach on the semiparametric binary choice model with a discrete regressor in Section (ref). Section (ref) shows that the discrepancy function and identified set can be computed via linear programming. Section (ref) extends all results to models with parametric error terms: Section (ref) focuses on the identification of common parameters, Section (ref) on the identification of both structural and counterfactual parameters, and Section (ref) discusses computation via linear programming. Section (ref) applies the results to panel models. Appendix (ref) contains additional remarks, including inference in Section (ref), while Appendix (ref) contains proofs not contained in the main text.

Notation.

For a Polish space $\mathcal{S}$ endowed with its Borel $\sigma$-algebra, $\mathfrak{B}(\mathcal{S})$ denotes the set of all Borel measures on the set $\mathcal{S}$, $\mathcal{P}(\mathcal{S}) \subseteq \mathfrak{B}(\mathcal{S})$ denotes the set of all Borel probability measures supported on $\mathcal{S}$, and $\delta_{s}$ denotes the Dirac measure at $s\in\mathcal{S}$. For an index $\theta$, $ \Gamma_\theta(\mathcal{S})\subseteq \mathcal{P}(\mathcal{S})$ denotes a generic set of Borel probability measures supported on $\mathcal{S}$. The set $C_c(\mathcal{S})$ denotes the space of compactly supported continuous functions defined on $\mathcal{S}$. The product of two or more Polish spaces is endowed with the product topology.

For an arbitrary convex set $\mathcal{C}$ in a linear space, $\mathrm{ext}( \mathcal{C})$ denotes the set of all extreme points of $\mathcal{C}$, $\mathrm{co}(\mathcal{C})$ denotes the convex hull of $\mathcal{C}$, and $\overline{\text{co}}(\mathcal{C})$ denotes the smallest closed convex set containing the set $\mathcal{C}$.

We use $\subset$ to denote a strict subset, i.e. $A \subset B$ means that $x\in A \Rightarrow x \in B$ and $\exists b \in B: b \not \in A$. We use $A \subseteq B$ to denote that $A \subset B$ or $A = B$.

For arbitrary Borel measures $\mu, \mu'$, the total variation norm between $\mu$ and $\mu'$ is denoted by $\norm{\mu -\mu'}_{\mathrm{TV}} \equiv \sup_{B \text{ Borel}} |\mu(B) - \mu'(B)|$. For a given arbitrary measure $\mu$ and an arbitrary vector-valued function $f$, $f \in L^1(\mu)$ if each component function of $f$ is integrable with respect to $\mu$. We let $\EE{\mu}{\cdot}$ denote integration against a general probability measure $\mu$.

We denote by $\Phi(\cdot)$ the cumulative distribution function of the standard normal distribution, and by $\Lambda(\cdot)$ the cumulative distribution function of the logistic distribution.

Main result

Let $Z$ denote an observable Borel measurable random variable supported on a space $\mathcal{Z}$, and let $\mu^*_{Z}$ denote the true probability measure of $Z$. Denote by $\Theta$ the set of possible values of the parameter of interest $\theta$, and by \(\Gamma_\theta\) a set of auxiliary parameters $\gamma$ that may vary with $\theta$. The parameter of interest may include functions of $\gamma\in\Gamma_\theta$. For example, in the binary choice model of Section (ref), $\gamma$ is the unknown distribution of an unobservable error term and $\theta$ includes counterfactual choice probabilities, which are functionals of $\gamma$.

For each $\theta \in \Theta$ and $\gamma \in \Gamma_\theta$, the econometric model for $Z$ specifies a probability measure $\mu_{Z,(\theta,\gamma)}$, which we call the model probability. For fixed $\theta$, define the set of model probabilities as:

align[align omitted — 117 chars of source]

This set is the collection of all probability measures of $Z$ that are consistent with the econometric model under parameter value $\theta$. The set is a fundamental object for our analysis, and its geometric properties are essential for our main result in Theorem (ref) below.

The identified set for $\theta$ is defined as the set of parameter values compatible with the probability measure $\mu_Z^*$, formally defined in (ref), where $\overline{\mathcal{M}}_{\theta}$ is the closure of $\mathcal{M}_\theta$ with respect to the topology induced by the total variation (TV) norm:

align[align omitted — 251 chars of source]

Defining the identified set through the closure \(\overline{\mathcal{M}}_{\theta}\) offers two benefits. First, it explicitly includes limit measures, which may arise when, e.g., $\gamma$ converges to a degenerate measure.\footnote{Alternatively, such measures could be directly incorporated into $\mathcal M_\theta$.} Second, it ensures that the identified set includes all parameter values $\theta$ for which the model measures are indistinguishable from the observed measures in the total variation norm, thereby avoiding issues related to impossible inference. Further details are provided in Section (ref).

Computing the identified set based on (ref) involves a search over $\overline{\mathcal M}_\theta$, which can be challenging when $\Gamma_\theta$ is an infinite-dimensional space. This motivates the alternative characterization of the identified set in Theorem (ref) below, which uses the following assumptions.

asm$\mathcal{Z}$ is a Polish space.
asmFor all $\theta \in \Theta$, there exists some $\sigma$-finite positive measure $\lambda_\theta \in \mathfrak{B}(\mathcal{Z})$ with respect to which every $\mu \in \M_\theta$ is continuous.

Assumption (ref) accommodates random variables $Z$ supported on various separable spaces, and excludes spaces that are non-metrizable, which rarely arise in econometric applications. Assumption (ref) requires that measures in $\mathcal{M}_\theta$ be well-behaved with respect to a $\sigma$-finite Borel measure $\lambda_{\theta}$ that is allowed to vary with $\theta$. The assumption is mild, allowing for a broad class of random variables, including continuous, discrete, and mixed types, and excluding singular measures. We discuss the necessity of this assumption in Remark (ref) in Section (ref), while in Section (ref) we relax it.

Consider the discrepancy function defined in (ref) with

align[align omitted — 124 chars of source]

and define the set of parameter values that set (ref) to zero as:

align[align omitted — 82 chars of source]

Our main result below clarifies the connection between $\Theta_\mathrm{I}$ in (ref) and $\Thetam$ in (ref), and serves as a building block for the subsequent analysis.

thm[Main result] Let Assumptions (ref) and (ref) hold. For any $\mu^*_Z \in \mathcal{P}(\mathcal{Z})$, \begin{align} \Theta_{\mathrm{I}} \subseteq \Thetam. \end{align} Additionally, let $\overline{\mathcal{M}}_\theta$ be convex for all $\theta$. Then $\Theta_{\mathrm{I}} = \Theta_{\mathrm{MI}}$.
proofTheorem (ref) is a direct implication of Proposition (ref), see Section (ref).

Convexity of \(\mathcal{M}_\theta\) plays a central role in our approach.\footnote{While this convexity may influence the geometry of the identified set, convexity of the identified set itself is neither implied by nor required for our results.} First, it serves as a sufficient condition for sharp identification by ensuring that \(\overline{\mathcal{M}}_\theta\) is convex. When convexity of \(\mathcal{M}_\theta\) fails, the characterization via (ref) provides an outer set. Second, if extreme points of \(\mathcal{M}_\theta\) exist, convexity of \(\mathcal{M}_\theta\) reduces the search in (ref) over \(\mu \in \mathcal{M}_\theta\) to a search over these extreme points. While characterizing these extreme points can be a complex, model-specific task, we show that many econometric models have a specific feature that renders this step unnecessary. This feature refers to the fact that the model probability can be expressed as a linear map on the space \(\Gamma_\theta\). That is,

equation[equation omitted — 121 chars of source]

where \(\mathcal{T}_\theta: \Gamma_\theta \to \mathcal{P}(\mathcal{Z})\) is a linear map. We illustrate this structure with concrete examples in subsequent sections: semiparametric models where the error distribution is either unrestricted or linearly restricted (see Section (ref)) or parametrically specified (see Section (ref)). The representation in (ref) implies that convexity of \(\Gamma_\theta\) is sufficient for Theorem (ref) to apply,\footnote{This assumption is not necessary. We show in Section (ref) that under sequential exogeneity, \(\mathcal{M}_\theta\) is convex despite \(\Gamma_\theta\) failing to be convex.}enabling dimensionality-reduction via a search over the extreme points of the convex set \(\Gamma_\theta\), which are easier to characterize.

Both linearity of the map and convexity of \(\Gamma_\theta\) enable the computation of the zeros of \(T(\theta)\) via a linear program (see Sections (ref) and (ref)), making computation of \(\Theta_{\mathrm{I}}\) through \(\Theta_{\mathrm{MI}}\) a feasible task, even when direct computation using (ref) may be impractical.

remarkThe discrepancy function in (ref) follows directly from our definition of the identified set in (ref) and is new to the literature. We provide a comparison to other criterion functions used in the literature in Remark (ref) in Section (ref).

Intuition

We provide a heuristic interpretation for the discrepancy function $T(\theta)$ and Theorem (ref). The maintained assumption for the discussion here is that $\mathcal{M}_\theta$ is convex.

Given a \(\theta \in \Theta\) and a probability measure \(\mu^*_Z \in \mathcal{P}(\mathcal{Z})\), there may exist several \(\gamma \in \Gamma_\theta\) such that the corresponding model probabilities \(\mu_{Z,(\theta,\gamma)} \in \mathcal{M}_\theta\) are indistinguishable from the true \(\mu^*_Z\) in the sense of (ref). If there exists at least one such \(\gamma\),\footnote{Or if a sequence of \(\gamma\)'s can be constructed such that \(\mu_{Z, (\theta, \gamma)}\) converges to \(\mu^*_Z\).} then \(\mu^*_Z \in \overline{\mathcal{M}}_\theta\), and consequently, \(\theta \in \Theta_{\mathrm{I}}\). Theorem (ref) uses insights from convex analysis to solve this existence problem. In particular, the proof of Theorem (ref) shows that $\mu^*_Z \in \overline{\mathcal{M}}_\theta$ if and only if

align[align omitted — 167 chars of source]

The discrepancy function \(T(\theta)\) in (ref) is then obtained after rearranging and taking the supremum over \(\phi\). The decision rule is based on the sign of \(T(\theta)\): If \(T(\theta) > 0\), the parameter \(\theta\) is excluded from the identified set, while if \(T(\theta) \leq 0\), \(\theta\) is included in the identified set. When $\mathcal M_\theta$ is convex, this characterization is sharp.

The discrepancy function $T(\theta)$ has an adversarial formulation where two opposing players, a critic and a defender, interact strategically. The critic selects a feature $\phi\in\Phi_b(\mathcal Z)$ and the defender selects a measure $\gamma\in\Gamma_\theta$. The critic seeks to maximize the discrepancy between the feature observed in the data, i.e. $\mathbb{E}_{\mu^*_Z}[\phi]$, and the corresponding feature predicted by the model, i.e. $\mathbb{E}_{\mu}[\phi]$, for a given parameter value $\theta$ and taking into account the defender's action. In response, the defender adjusts $\gamma$ to minimize this discrepancy. The resulting discrepancy function captures the maximum discrepancy that the critic can enforce, even after the defender optimally adjusts the probability measure of the unobserved heterogeneity. The sign of the discrepancy function determines a decision rule: A positive value means that the critic has identified a feature where the model, under a specified parameter value, fails to replicate the observed data; the parameter value is then excluded from the identified set. A non-positive value means that the defender can always find a measure that aligns the model's prediction with the observed data; the parameter value is then included in the identified set. Note that if \(T(\theta) \leq 0\) for all \(\phi\), the critic selects \(\phi = 0\) to ensure \(T(\theta) = 0\).

This interpretation, together with the central role of the discrepancy function in our identification, computation, and inference results, forms the basis of our adversarial approach.

Models without parametric restrictions

Models in this section are described as follows. Let the input variables be denoted by \(W \in \mathcal{W}\) (which may include latent variables) and continue to denote the observed random variables by \(Z \in \mathcal{Z}\). \footnote{For example, \(Z\) may consist of observed outcomes \(Y\) and conditioning covariates \(X\), such as regressors and instrumental variables, whereas \(W\) may contain \(X\) along with stochastic error terms or other latent random variables. See Section (ref) for further discussion.} In this setting, \(\Gamma_\theta \subseteq \mathcal{P}(\mathcal{W})\) is the set of probability measures supported on \(\mathcal{W}\) that are allowed by the model under the parameter value \(\theta\); we denote this set by \(\Gamma_\theta(\mathcal{W})\) to emphasize its dependence on \(\mathcal{W}\).

For any $\theta\in\Theta$, there exists a measurable map $\psi_\theta:\mathcal{W} \mapsto \mathcal{Z}$ known up to $\theta$, such that

equation[equation omitted — 74 chars of source]

This specification includes semiparametric models with, e.g., outcome equations such as $Z=h(\beta,W)$, $\theta=(h,\beta)$, with $h$ an unknown function and $\beta$ an unknown finite-dimensional parameter, and models with outcome equations such as $Y=m(W)$, where $\theta=m$ is an unknown function.

For fixed $\theta$, $\psi_\theta$ induces a pushforward measure $(\psi_\theta)_*: \Gamma_\theta(\mathcal W) \ra \mathcal{P}(\mathcal{Z})$ defined as:

align[align omitted — 172 chars of source]

where \(\psi_\theta^{-1}(S) = \{ w \in \mathcal{W} : \psi_\theta(w) \in S \}\) is the preimage of \(S\) under \(\psi_\theta\). Then,

equation[equation omitted — 176 chars of source]
asmFor any $\theta$, there exists a measurable map $\psi_\theta:\mathcal W \mapsto \mathcal Z$ such that for any $\gamma\in\Gamma_\theta$, the model probability is given by: $$\mu_{Z,(\theta,\gamma)}(S) = (\psi_\theta)_*\gamma (S), \text { for any Borel } S \subseteq \mathcal{Z}.$$

Assumption (ref) is naturally satisfied in the class of models with outcome equations described by (ref). Here, the linear map in (ref) is the pushforward measure $(\psi_\theta)_*$, so that measures on \(\mathcal{Z}\) are obtained by “pushing forward” \(\gamma \in \Gamma_\theta(\mathcal{W})\) via \((\psi_\theta)_*\). \footnote{While $\psi_\theta$ here is not a correspondence, our result in Corollary (ref) may still apply to such models. This is because a sufficient condition for that result is the existence of $\mathcal T_\theta$ in (ref), which can be induced by a correspondence between $\mathcal W$ and $\mathcal Z$. In Section (ref), we extend our results to cover linear operators from \(\Gamma_\theta(\mathcal{W})\) to \(\mathcal{P}(\mathcal{Z})\). Notably, a correspondence between \(\mathcal{W}\) and \(\mathcal{Z}\) can induce a map or an operator between spaces of probability measures on \(\mathcal{W}\) and \(\mathcal{Z}\). Deriving sufficient conditions for when such correspondences lead to well-defined linear maps or operators is left for future work.}

corollaryLet Assumptions (ref), (ref), and (ref) hold, and assume that $\Gamma_\theta(\mathcal W)$ is convex. Then, $\mathcal{M}_\theta$ is convex and $\Theta_{\mathrm{I}} = \Theta_{\mathrm{MI}}$.
proofThe set $\mathcal{M}_\theta$ is convex since it is the image of a convex set under a linear map. This follows from Assumption (ref), which defines elements of $\mathcal{M}_\theta$ via the application of the linear pushforward measure $(\psi_\theta)_*$ to elements of the convex set $\Gamma_\theta(\mathcal{W})$. That $\Theta_{\mathrm{I}} = \Theta_{\mathrm{MI}}$ follows from Theorem (ref), and from the convexity of $\mathcal M_\theta$.

The set $\Gamma_\theta(\mathcal W)$ plays a central role in our analysis. In some semiparametric models, it is unrestricted, i.e. $\Gamma_\theta(\mathcal{W}) = \mathcal{P}(\mathcal{W})$ is the set of all probability measures on $\mathcal W$, while in other models, $\gamma$ is known to satisfy certain linear restrictions, i.e. $\Gamma_\theta(\mathcal{W}) = \mathcal{P}(\mathcal{W})^g$ which we define below. Models with parametric restrictions on $W$ are discussed in Section (ref).

For a known vector of functions $g:\Theta\times\mathcal{W}\to\mathbb{R}^{d_g}$, define

align[align omitted — 184 chars of source]

This is the set of all probability measures on $\mathcal W$ that satisfy a set of $d_g<\infty$ linear restrictions that may depend on $\theta$. As will become clear from examples throughout this paper, such restrictions are important in many econometric models. First, many restrictions commonly made on the distribution of latent variables, such as mean- or median-independence, can be expressed as in (ref). Second, many counterfactuals of interest can be cast in the form $\EE{\gamma}{g(\theta, w)}=0$, implicitly imposing linear restrictions on \(\gamma\). Section (ref) illustrates this for a semiparametric binary choice model.

Corollary (ref) shows that convexity of $\Gamma_{\theta}(\mathcal W)$ is sufficient for sharp identification. Under Assumption (ref) below, each of $\mathcal{P}(\mathcal{W})$ and $\mathcal{P}(\mathcal{W})^g$ is convex, and Corollary (ref) applies.

asm$\mathcal{W}$ is a Polish space.

We are now ready to establish an extremal point characterization of the result in Theorem (ref). Characterizing the identified set using (ref) involves a search over the space $\mathcal{M}_\theta$. Given Assumption (ref), the search can instead be conducted over $\Gamma_\theta(\mathcal{W})$, i.e. for any $\phi\in\Phi_b(\mathcal{Z})$,

align[align omitted — 172 chars of source]

Proposition (ref) below shows that this search can further be confined to a smaller space.

proposition[Extremal point characterization of the identified set] Let Assumptions (ref), (ref), (ref) and (ref) hold. Additionally, \begin{itemize} • if $\Gamma_\theta(\mathcal{W}) = \mathcal{P}(\mathcal{W})$, then $\theta \in \Theta_\mathrm{I}$ if and only if \begin{align} \EE{\mu^*_Z}{\phi} \le \sup_{w \in \mathcal{W}} (\phi \circ \psi_\theta)(w) for all \phi\in \Phi_b(\mathcal{Z}). \end{align} • if $\Gamma_\theta(\mathcal{W}) = \mathcal{P}(\mathcal{W})^g$, then $\theta \in \Theta_\mathrm{I}$ if and only if \begin{align} \EE{\mu^*_Z}{\phi} \le &\sup_{\{c_j, w_j\}_{j=1}^{d_g+1}} \sum_{j=1}^{d_g+1} c_j \left(\phi \circ \psi_\theta(w_j)\right) for all \phi\in \Phi_b(\mathcal{Z}) \nonumber \\ &subject to \sum_{j=1}^{d_g+1} c_j g(\theta, w_j) = 0, \, \sum_{j=1}^{d_g+1} c_j = 1, \, c_j \geq 0 . \end{align} \end{itemize}

This result has important implications. For instance, when \(\gamma\) is unrestricted, so that \(\Gamma_\theta(\mathcal{W}) = \mathcal{P}(\mathcal{W})\), Proposition (ref)(i) implies that the identified set can be characterized by searching over the set of extreme points of \(\Gamma_\theta(\mathcal{W})\), all of which correspond to Dirac measures on \(\mathcal{W}\). Thus, the search is restricted to \(\mathcal{W}\) rather than the much larger set of probability measures on \(\mathcal{W}\). This refinement is particularly useful in nonlinear panel models, as it means that determining whether \(\theta \in \Theta_I\) requires only a search over the values of the fixed effects, rather than over all possible conditional distributions of the fixed effects.

remarkThe discrepancy function $T(\theta)$ in (ref) simplifies. For example, under the conditions for case (i) in Proposition (ref), \begin{equation} T(\theta) = \sup_{\phi \in \Phi_b(\mathcal Z)} \inf_{w \in \mathcal W} \left( \EE{\mu_Z^*}{\phi} - (\phi \circ \psi_\theta)(w) \right). \end{equation} Section (ref) uses this insight to show that $T(\theta)$ can be recast as a linear program.
remarkUnder additional regularity conditions, the dimensionality of the search can be further reduced by considering all continuous functions $\phi:\mathcal{Z}\to [0,1]$, see Section (ref).

Semiparametric regression models

In this section, we focus on semiparametric models with both observed and unobserved heterogeneity. Our primary goal is to clarify how conditioning variables (including regressors and instrumental variables) are treated within our framework. We also provide a blueprint for verifying the assumptions of Corollary (ref) for models with discrete and continuous conditioning variables. Section (ref) provides an example.

Many semiparametric models have an outcome equation of the form:

equation[equation omitted — 59 chars of source]

where $Y \in \mathcal{Y}$, $X \in \mathcal{X}$, and $U \in \mathcal{U}$ are random variables, $\theta \in \Theta$ is an unknown parameter, and $h:\mathcal{X} \times \mathcal {U} \times \Theta \to \mathcal{Y}$ is a (structural) function known up to $\theta$. $Y$ denotes outcome variables, $X$ denotes observed variables, and $U$ unobserved variables.

Using the notation from Section (ref), let \(Z = (Y, X)\), \(W = (X, U)\), \(\mathcal{Z} = \mathcal{Y} \times \mathcal{X}\), and \(\mathcal{W} = \mathcal{X} \times \mathcal{U}\). Denote the distribution of $W$ supported on $\mathcal W$ by \(\gamma \in \Gamma_\theta(\mathcal{W})\subseteq\mathcal{P}(\mathcal{W})\), the marginal distribution of \(X\) under \(\gamma\) by \(\gamma_\mathcal{X}\in\Gamma_{\theta, \mathcal{X}}(\mathcal{X})\subseteq \mathcal{P}(\mathcal{X})\), and the conditional distribution of \(U \mid X = x\) for each \(x \in \mathcal{X}\) by \(\gamma_{U \mid x}\in\Gamma_{\theta, x}(\mathcal{U})\subseteq \mathcal{P}(\mathcal{U})\). By the disintegration theorem, measures \(\gamma \in \Gamma_\theta(\mathcal{W})\) have differential \(\mathrm{d}\gamma = \mathrm{d}\gamma_{U|x} \, \mathrm{d} \gamma_X\) for all \(x \in \mathcal{X}\), allowing thus for correlation between $X$ and $U$.

Finally, for any \(\theta \in \Theta\), the mapping

equation[equation omitted — 78 chars of source]

induces, for each \(\gamma \in \Gamma_\theta(\mathcal{W})\), the pushforward measure \((\psi_\theta)_*\) on \(\mathcal{P}(\mathcal{Z})\), as defined in (ref). Consequently, the set \(\Gamma_\theta(\mathcal{W})\) induces the set \(\mathcal{M}_\theta = (\psi_\theta)_* \Gamma_\theta(\mathcal{W})\) as in (ref).

The assumption below allows us to specialize the assumptions of Corollary (ref) to models with outcome equation as in (ref).

asm(i) $\mathcal{Y}$ and $\mathcal{X}$ are Polish spaces; (ii) The set $\Gamma_{\theta, \mathcal{X}}$ is convex, and the set $\Gamma_{\theta, x}$ is convex for every $x\in\mathcal X$; (iii) There exists a $\sigma$-finite $\lambda_{\theta, \mathcal{X}} \in \mathfrak{B}(\mathcal{X})$ with respect to which every $\gamma_X \in \Gamma_{\theta, \mathcal{X}}(\mathcal{X})$ is continuous; (iv) There is a collection of $\sigma$-finite $\mathcal{X}$-measurable measures $\{\lambda_{\theta, x}: x \in \mathcal{X}\} \subseteq \mathcal{P}(\mathcal{Y})$ such that, for all $x \in \mathcal{X}$ and $\gamma_{U|x} \in \Gamma_{\theta, x}(\mathcal{U})$, the pushforward of $\gamma_{U|x}$ under $h(x,\cdot; \theta): \mathcal{U} \mapsto \mathcal{Y}$ is continuous with respect to $\lambda_{\theta, x}$.

Assumption (ref)(i) is a regularity condition which fulfills the requirements of Assumption (ref). Assumption (ref)(ii) specifies sufficient conditions for the convexity of $\Gamma_\theta(\mathcal{W})$, and consequently the convexity of $\mathcal{M}_\theta$. Typically, $X$ are treated as conditioning variables, in the sense that the model specifies assumptions on features of the conditional distribution $\gamma_{U|x}, \; x\in\mathcal X$. By including $X$ in $W$, typical conditions on $\gamma_{U|x}$ can be formulated as linear restrictions on \(\gamma\) as in (ref). If, additionally, $\mathcal U$ is a Polish space so that Assumption (ref) holds, convexity of $\Gamma_\theta(\mathcal W)$ is preserved and straightforward to verify by determining whether assumptions on $U \mid X$ can be expressed as in (ref). Alternatively, convexity of $\Gamma_\theta(\mathcal W)$ can be established by verifying Assumption (ref)(ii). Note that here $\gamma_\mathcal X$ is treated as a latent distribution. Since $X$ enters $Z$, it is possible to treat $\gamma_X$ as known and equal to the marginal distribution of $X$ under the observed distribution of $Z$. In this case, $\Gamma_{\theta,\mathcal X}$ is a singleton and trivially convex. Convexity of $\Gamma_{\theta, x}$ is enforced via assumptions on $\gamma_{U \mid x}$ that guarantee convexity of the set for every $x\in\mathcal X$.

Assumptions (ref)(iii) and (iv) are mild continuity conditions that guarantee that all measures in $\mathcal{M}_\theta$ are continuous with respect to the $\sigma$-finite measure $\lambda_\theta \in \mathfrak{B}(\mathcal{Y} \times \mathcal{X})$ with differential $\d \lambda_\theta = \d \lambda_{\theta, x} \, \d \lambda_{\theta, \mathcal{X}}$ for all $x \in \mathcal{X}$, thereby satisfying Assumption (ref). These assumptions require, respectively, that the marginal and conditional distributions of measures in $\mathcal{M}_\theta$ have density with respect to $\sigma$-finite measures, which may depend upon $\theta$. These dominating measures may be continuous, discrete, or a mixture of both. For example, when $\mathcal{Y}$ is discrete, $\lambda_{\theta, x}$ is the counting measure on $\mathcal{Y}$ for all $x$, in which case Assumption (ref) is fulfilled. In contrast to much of the existing relevant literature, Assumptions (ref)(iii) and (iv) allow for $(Y,X)$ to be continuous, discrete, or a mixture. Importantly, Corollary (ref) pertains to semiparametric models with outcome equations of the form (ref), where, for example, $X$ is continuous with respect to Lebesgue measure $\lambda_{\theta,\mathcal{X}}$ and either $Y$ is discrete with $\lambda_{\theta,x}$ the counting measure or $Y$ is continuous with $\lambda_{\theta,x}$ the Lebesgue measure (or some mixture of these cases).

corollaryConsider the outcome equation in (ref). Let Assumption (ref) hold for all $\theta \in \Theta$. Then $\Thetaw = \Thetam$.

Corollary (ref) provides a blueprint for checking the assumptions of Corollary (ref), and, consequently, for establishing sharp identification in semiparametric models with an outcome equation as in (ref).

remark[Extremal point representation] Consider now the extremal point representation of Proposition (ref). Restrictions on the probability measure of $X$ in $Z$, such as Assumption (ref)(iii), impose restrictions on the measure of $W$; these restrictions may constrain $\mathcal P(\mathcal W)$ and $\mathcal P(\mathcal W)^g$ in measure theoretic ways that are not allowed by Proposition (ref). When $X$ is discrete and $\lambda_{\theta,\mathcal X}$ in Assumption (ref)(iii) is the counting measure, this phenomenon does not occur. In this case, the assumptions of Proposition (ref) hold, and $\Thetam(=\Thetaw)$ is characterized by the extremal point representation in either (ref) or (ref). When $X$ is continuous and $\lambda_{\theta,\mathcal X}$ in Assumption (ref)(iii) is the Lebesgue measure, the extremal point representation in (ref) or (ref) describes an outer set for $\Thetam$. It is still the case that $\Thetaw=\Thetam$, but (ref) or (ref) may give an outer set for $\Thetam$.\footnote{In Section (ref), we recover an extremal point representation result for $\Thetam$ when $X$ is continuous. In that setting, some components of $U$ satisfy parametric restrictions. The additional structure on the conditional distribution of $Y$ allows us to further relax the already mild requirements on the distribution of $X$ in Corollary (ref).}

Example: semiparametric binary response

We illustrate our approach using the model in manskiMaximumScoreEstimation1975 with a discrete regressor. We characterize the identified set for both regression coefficients and counterfactual choice probabilities. \footnote{The purpose of this section is to illustrate our approach rather than to obtain new results. For existing results, see komarovaBinaryChoiceModels2013, blevinsNonStandardRates2015, and torgovitsky2019. }

Consider the binary choice model with outcome equation:

equation[equation omitted — 105 chars of source]

where $Y\in\{0,1\} = \mathcal Y$, $X\in\{x_1,\cdots,x_K\} = \mathcal X$ has $K$ points of support, $U\in\mathcal U = \mathbb{R}$ is a scalar error term that satisfies the following conditional median-zero assumption:

align[align omitted — 99 chars of source]

and $\beta = (\beta_0,\beta_1)$ are unknown structural parameters. The counterfactual choice probabilities are given by:

align[align omitted — 146 chars of source]

which correspond to the average counterfactual outcome obtained by exogenously setting the values of the observed regressors to counterfactual values $x^*$ for the subpopulation given by $X=x_k$.

To characterize the identified set of

align[align omitted — 106 chars of source]

we verify the assumptions of Corollary (ref). Using the notation of the previous sections, $Z=(Y,X),\; W=(X,U)$, $h$ is given by (ref), $\gamma_x$ denotes the marginal distribution of $X$, $\gamma_{U|x}$ denotes the conditional distribution of $U \mid X=x, \; x\in\mathcal X$ satisfying (ref), and $\gamma \in \Gamma_\theta(\mathcal{W})$ denotes the distribution of $W$ with differential \(\mathrm{d}\gamma = \mathrm{d}\gamma_{U|x} \, \mathrm{d} \gamma_X\) for all \(x \in \mathcal{X}\).

Assumption (ref)(i) is trivially satisfied because $\mathcal Y$ and $\mathcal X$ are finite. The same is true for (ref)(iii) and (ref)(iv) with $\lambda_{\theta,\mathcal X}$ and $\lambda_{\theta,x}$ counting measures. (ref)(ii) is also satisfied: (a) $\Gamma_{\theta, \mathcal X}$ is a point, so convexity is trivially satisfied; (b) for each $x$, the set $\Gamma_{\theta,x}$ is the set of all probability measures $\gamma_{U|x}$ that satisfy the linear restrictions in (ref) and (ref), so $\Gamma_{\theta,x}$ is convex for all $x$ (see proof of Proposition (ref)). To see that (ref) and (ref) impose linear restrictions on $\gamma_{U|x}$ for each $x\in\mathcal X$ consider that these restrictions can be written as $$E_{\gamma_{U|x}}\left[ \widetilde g(U,\theta) \mid X = x\right] = 0$$ where

align[align omitted — 176 chars of source]

Because Assumption (ref) is satisfied, Corollary (ref) guarantees sharp identification of $\theta$ via $\Thetam$. Moreover, the results of Proposition (ref) also apply.

It follows that the identified set can be computed using the linear programming approach developed in Section (ref). Section (ref) describes in detail how to apply it to the semiparametric binary choice model, and presents numerical results. The computation time for the identified set of $\theta$ is trivial, see Section (ref).

Linear programming

The computation of $\Thetam$ requires computation of the discrepancy function (ref). This may seem challenging, as it involves a search over measures in $\mathcal{M}_\theta$ and over functions $\phi \in \Phi_b(\mathcal{Z})$. However, computing $T(\theta)$ only requires solving a linear program (LP). Below, we provide a sketch, leaving the details to Section (ref). The approach extends to the models in Section (ref).

Our theoretical results allow for $\mathcal Z$ and $\mathcal W$ to be Polish spaces, cf. Assumptions (ref) and (ref). For the purpose of computation, we may restrict attention to the case of finite supports, \( \mathcal Z = \{z_1,\cdots,z_L\}, \; \mathcal W = \{w_1,\cdots,w_M\}. \) Note that the linear program presented here does not rely on discreteness of $\mathcal Z$ and $\mathcal W$. Rather, it relies on the linearity of the map between probability measures as in (ref). When $\mathcal Z$ and $\mathcal W$ are discrete, the map takes the form of a matrix, which is the case here.

The probability measure $\mu_Z^*$ can be represented by a probability mass function (pmf) $p_Z^* = \left(p_{Z,l}^*\right)$, an $L \times 1$ probability vector, so that the first term in (ref) is \[ \EE{\mu_Z^*}{\phi} = \sum_{l = 1}^L \phi(z_l) p_{Z,l}^* = \phi^\prime p_Z^*. \]

Similarly, every model probability $\mu_{Z,(\theta,\gamma)}$ corresponds to a pmf $p_{Z}^{(\theta,\gamma)} = \left(p_{Z,l}^{(\theta,\gamma)}\right) $, and every probability measure $\gamma$ for $W$ to an $M\times 1$ pmf $p_W = \left(p_{W,m}\right)$. By the pushforward representation in Section (ref), cf. (ref) and Assumption (ref), there exists a matrix $\widetilde C_\theta \in \mathbb R^{L \times M}$ such that

equation[equation omitted — 101 chars of source]

For a given $\theta$, this pushfoward matrix maps the pmf of $W$ to a model pmf of $Z$ under parameter value $\theta$. For examples of $\widetilde C_\theta$, see Sections (ref) and (ref).

The second term of (ref) can be written as

equation[equation omitted — 81 chars of source]

so the infimum over $\mu$ in (ref) can be replaced by a minimum over $p_W$ subject to

equation[equation omitted — 87 chars of source]

These constraints enforce that $p_W$ is a probability vector and allow for additional constraints corresponding to $\Gamma_\theta$. Without additional constraints, $A_\theta = \iota_M^\prime$ and $b_\theta = 1$. For examples of other constraints, see Sections (ref) and (ref).

Taken together, we have

align[align omitted — 238 chars of source]

where

equation[equation omitted — 98 chars of source]

In Section (ref), we show that the bilinear program in (ref) is equivalent to the LP

equation[equation omitted — 307 chars of source]

where $\lambda$ is the Lagrange multiplier associated with the equality constraints in (ref). We conclude that determining whether $\theta \in \Theta_I$ amounts to solving the LP (ref) and checking that $T(\theta) \leq 0$. This task has negligible computation time even for very large $(L,M)$, see for example the computation times reported in Section (ref). It is also trivial to write the code for a specific model. The user specifies (i) supports $\mathcal Z, \, \mathcal W$; (ii) true parameter values $(\theta^*,p_W^*)$; (iii) the pushforward matrix $\widetilde C_\theta$; (iv) restrictions $(A_\theta,b_\theta)$; then hands off the LP (ref) to a solver.

remark[Relationship of LP to Proposition (ref)] Proposition (ref) and Remark (ref) establish the extremal point problem, i.e. when optimizing a linear functional over a convex set of probability measures, the search can be restricted to the set of extreme points. Expression (ref) reformulates the discrepancy function as a bilinear program, where the inner minimization is over the convex hull of extreme points. This is a relaxation of the extremal point problem that leaves the value of $T(\theta)$ unchanged, but facilitates computation via the dual LP in (ref). There, the problem is expressed in terms of Lagrange multipliers associated with the constraints on the extreme points, eliminating the need to compute them explicitly.
remarkThe resulting LP is different from the LP in honoreBoundsParametersPanel2006 and bonhommeIdentificationBinaryChoice2023, and from the quadratic program in ChernValHahnNewey2013. In contrast to those approaches, ours avoids a search over the space of all conditional distributions of unobserved heterogeneity, with the inner minimization restricted instead to the support of the unobserved heterogeneity, cf. Proposition (ref). This is expressed by the dual in the LP discussed above. In the resulting LP, $\phi$ ranges over $\Phi_b$, which is typically of low dimension. Searching over a subset of $\phi$ yields an outer set. In contrast, an incomplete search over fixed effects distributions may lead to an inner set.

Models with parametric restrictions

In this section, the latent component $U$ in (ref) consists of two components: one that follows a known distribution and another whose distribution is unrestricted. The unrestricted component is labeled $\alpha$ to evoke the literature on parametric nonlinear panel models, see, e.g., Bonhomme2012.

We consider the following semiparametric regression model:

align[align omitted — 170 chars of source]

where $\beta \in \mathrm{B}$ is the parameter of interest, the observed random variables are $Z=(Y,X)$, the latent random variables $U=(\alpha,V)$ consist of a component $\alpha\in\mathcal{A}$ and a component $V\in\mathcal{V}$, where $\gamma_{X,\alpha} \in \mathcal{P}(\mathcal{X} \times \mathcal{A})$ denotes the joint distribution of $(X,\alpha)$ and $F_{V|X,\alpha;\beta}$ denotes the distribution of $V$ that is allowed to depend on $(X,\alpha,\beta)$. Here, $\gamma_{X,\alpha}$ is unknown and unrestricted, while $F_{V|X,\alpha;\beta}$ is known up to $\beta$.

In Section (ref), we characterize the identified set for $\beta$. In Section (ref), we characterize the identified set of counterfactual parameters. In Section (ref), we extend the LP approach in Section (ref). In Section (ref), we apply the tools developed here to a parametric binary choice model with fixed effects.

Identification of common parameters

For the models in this section, the pushforward representation for the model probability in Assumption (ref) does not hold. Instead, we show that the model probability can be represented as a linear integral operator similar to the one in Bonhomme2012.

In the following assumption, $f_{Y|X,\alpha}$ is the conditional density of $Y$ given $(X,\alpha)$ determined by the outcome equation and the parametric distribution of $V$, see (ref).

asmFor $\beta \in \mathrm{B}$, the densities $(y,x,a) \mapsto f_{Y|X,\alpha}(y| x, a;\beta)$ are defined relative to a $\sigma$-finite measure $\lambda_\mathcal{Y} \in \mathfrak{B}(\mathcal{Y})$, and Borel measurable for all $(x,a)\in\mathcal X \times \mathcal A$.

The parametric nature of the distribution of $V$ allows us to drop Assumption (ref), with Assumption (ref) imposing a measure continuity condition only on the conditional distributions of $Y$ given $X = x, \alpha = a$ for $\beta \in \mathrm{B}$.

Denote by $\gamma_X$ the $\mathcal{X}$-marginal of $\gamma_{X,\alpha}$. Under (ref) and Assumption (ref), for fixed $\beta$, the conditional density $f_{Y|X,\alpha}$ and the joint measure $\gamma_{X, \alpha}$ induce the following linear integral operator:

equation[equation omitted — 275 chars of source]

where $S\subseteq \mathcal Y \times \mathcal X$ is a measurable set and $\mathbf{1}_S(y, x)$ is the indicator function for $S$. That is, for any joint probability measure $\gamma_{X, \alpha}\in \mathcal{P}(\mathcal{X} \times\mathcal{A})$, $\mathcal{L}_\beta[\gamma_{X,\alpha}]$ is the probability measure of $Z$ under $\gamma_{X, \alpha}$. In particular,

equation[equation omitted — 162 chars of source]

That is, when a latent component follows a parametric distribution, the linear map $\mathcal T_\theta$ in (ref) is the integral operator $\mathcal{L}_\beta$. The set of all model probabilities compatible with $\beta$ can then be defined as \[ \mathcal M_\beta = \left\{ \mathcal{L}_\beta[\gamma_{X,\alpha}]: \gamma_{X,\alpha} \in \mathcal P(\mathcal X \times \mathcal A) \right\}. \]

propositionLet $\mathcal Y, \mathcal X, \mathcal A$ be Polish spaces and let Assumption (ref) hold for the model described by (ref). Let $ \mu_{Z}^*$ be the true probability measure of $Z=(Y,X)$. Then $\beta$ is in the identified set if and only if \begin{align} \EE{\mu^*_{Z}}{\phi(Z)} & \le \sup_{\substack{ x \in \mathcal{X} \\ a \in \mathcal{A} }} \int_{\mathcal{Y}} \phi(y,x) f_{Y|X,\alpha} (y|x, a; \beta) \, \d \lambda_{\mathcal{Y}} for all \phi\in\Phi_b(\mathcal Y \times \mathcal X). \end{align}

Identification of counterfactual parameters

We now turn to the identification of counterfactual parameters \(\tau\). Here, \(\theta = (\beta, \tau)\) denotes the joint parameter of interest. The counterfactual parameters $\tau$ are defined as solutions to the integral equation:

align[align omitted — 157 chars of source]

where \(g(x, \alpha, \beta, \tau): \mathcal{X} \times \mathcal{A} \times \Theta \to \mathbb{R}^{d_g}\) is a known, vector-valued function. Typically, $\tau$ is a function from $\mathcal{X}$ to $\R^{d_g}$ and $g(x,\alpha, \beta, \tau) = g_0(x,\alpha, \beta) - \tau(x)$ for some $g_0: \mathcal{X} \times \mathcal{A} \times \mathrm{B} \ra \R^{d_g}$, the linear restriction in (ref) becomes

align[align omitted — 95 chars of source]

Various specifications of $g_0$ in (ref) allow a researcher to restrict the conditional moments of $\gamma_{X,\alpha}$ and determine if specific conditional moments, parameterized by $\tau$, belong to the identified set.

Using the operator \(\mathcal{L}_\beta\) defined in (ref), the set of model probabilities is then:

align[align omitted — 334 chars of source]

In order to identify $(\beta, \tau)$, we impose a light regularity constraint on $g$. In the assumption that follows, $\R^0$ is identified with $\{0\}$, and a function that maps to $\R^0$ is the zero-map.

asm$g: \mathcal{X} \times \mathcal{A} \times \Theta \ra \R^{d_g}$ is Borel and satisfies \begin{align} 0 \in \mathrm{co}\{g(x,a, \beta, \tau): a \in \mathcal{A}\}, for all x \in \mathcal{X}. \end{align}
propositionLet $\mathcal Y, \mathcal X, \mathcal A$ be Polish spaces and let Assumptions (ref) and (ref) hold for the model described by (ref). Let $ \mu_{Z}^*$ be the true probability measure of $Z=(Y,X)$. Then, \(\theta = (\beta, \tau)\) is in the identified set if and only if, for all $\phi\in\Phi_b(\mathcal Y \times \mathcal X)$: \begin{align} \mathbb{E}_{\mu_Z^*}[\phi(Z)] \le &\sup_{\substack{x \in \mathcal{X} \\ (c_j, a_j)_{j = 1}^{d_g + 1} }} \sum_{j = 1}^{d_g + 1} c_j \int_{\mathcal{Y}} \phi(y,x) f_{Y|X,\alpha} (y|x,a_j; \beta) \, \d \lambda_{\mathcal{Y}} \nonumber \\ & subject to \sum_{j = 1}^{d_g + 1} c_j g(x, a_j, \beta, \tau) = 0,\, \sum_{j = 1}^{d_g +1} c_j = 1, \, c_j \ge 0. \end{align}

Linear programming

We extend the LP approach developed in Section (ref) to the present setting in Appendix (ref). Here, we note that in the parametric error case, the optimization problem is reformulated using a linear integral operator, reflecting the parametric structure imposed on \( V \). Specifically, the model probabilities are expressed as an integral transform of the joint measure of the unrestricted components of the latent variables through a conditional density function. This representation mirrors the structure observed in the pushforward map from Section (ref) but adapts to the parametric setting.

Crucially, the dimensionality of the resulting LP depends only on the support of \( X \) and \( \alpha \), and not on that of \( V \). This invariance ensures that the computational complexity remains manageable even as the dimension of the latent space increases.

Binary choice with fixed effects

We study the two-period binary choice model with outcome equation

align[align omitted — 113 chars of source]

which has a fixed effect $\alpha \in \mathcal A \subseteq \mathbb R$, error terms $V_t \in \mathcal{V}_t \subseteq \mathbb R$, and time-varying regressors $X_1 \in \mathcal X_1 = \{x_{11},\cdots,x_{1K_1}\}$ and $X_2 \in \mathcal X_2 = \{x_{21},\cdots,x_{2K_2}\}$.

We consider the particular specification in (ref) because of its canonical status within the nonlinear panel literature. Our method applies to a much larger class of models, and does not require the index structure or additive separability.

In Sections (ref) and (ref), we do not impose parametric assumptions on the distribution of the error terms. In Section (ref), we provide new results for partial effects under strict exogeneity, and discuss the extension to correlated random coefficients.

In Section (ref), we provide new results for both structural parameters and partial effects under sequential exogeneity (predetermined covariates). To the best of our knowledge, these are the first such results for the binary choice panel model without parametric restrictions on the distribution of the error terms. Section (ref) presents the results from a numerical experiment and visualizes the identified sets.

In Section (ref), we use our results in Section (ref) to analyze parametric binary choice models with fixed effects (logit, probit, and other distributions). We recover the results in the numerical experiment of ChernValHahnNewey2013 and explore some variations of their design. The proofs of all results can be found in Appendix (ref).

Partial effects under strict exogeneity

We are interested in the average structural functions (ASF):

align[align omitted — 120 chars of source]

This is the period $t$ choice probability obtained by exogenously setting regressors $X_t$ to a fixed value $x$. We characterize the identified set of $\theta = (\beta,\tau_1,\tau_2)$ under the following assumption: \footnote{Section (ref) provides results under a weaker, sequential exogeneity condition.}

asmThe random variables $(\alpha,V_1,V_2,X_1,X_2)$ satisfy \[V_1 | \alpha, X_1, X_2 \stackrel{d}{=} V_2 | \alpha, X_1, X_2.\]

This is the standard strict stationarity or exogeneity assumption for nonlinear panel models, see Manski1987 and ChernValHahnNewey2013.

To map this model to the framework of Section (ref), define $Y = (Y_1,Y_2)$ and $X = (X_1,X_2)$ so that $Z = (Y,X)$. The input variables are $W = (\alpha, V_1, V_2, X_1, X_2) \in \mathcal W$. The set of probability distributions of $W$ consists of all $\gamma\in\Gamma_\theta(\mathcal W)$ that satisfy Assumption (ref). Finally, the mapping \(\psi_\theta\) is given by

align[align omitted — 147 chars of source]

Using this mapping and Corollary (ref), the following result establishes sharp identification of partial effects.

thmConsider the model described by outcome equation (ref) and Assumption (ref). The set $\Theta_{\mathrm{MI}}$ is the sharp identified set for $\theta = (\beta,\tau_1,\tau_2)$.

Computation of $\Thetam$ is straightforward via the LP approach that we developed in Section (ref). In Section (ref), we use this approach to visualize $\Thetam$.

remarkConsider a version of this model with the outcome equation $$Y_t = 1\{X_{1t}^\prime \beta + X_{2t}^\prime \alpha + V_t \geq 0\},$$ where $\alpha$ is a random vector of individual-specific coefficients that may be correlated with $X_1, X_2$. Without restrictions on the distribution of $(\alpha,X)$, the proof of Theorem (ref) carries through without modification and the identified set of $\theta$, which can include moments of the random coefficients, is characterized by $\Thetam$.
remarkaristodemouSemiparametricIdentificationPanel2021 and ChesherRosenZhang consider the identification of regression coefficients in nonlinear panel models under (conditional) independence restrictions on the errors and regressors. For example, under $V \perp X$ instead of Assumption (ref), $\Gamma_\theta$ is convex, and our results deliver partial effects.

Predeterminedness

Assumption (ref) does not allow for correlation between current covariates and past shocks. A sequential exogeneity assumption that allows for such feedback is:

asmThe random variables $(\alpha,V_1,V_2,X_1,X_2)$ satisfy \begin{equation} V_2 | \alpha, X_1 , X_2 \stackrel{d}{=} V_1 | \alpha, X_1. \end{equation}

This is Assumption 3 in ChernValHahnNewey2013, which is less restrictive than the sequential exogeneity assumption in parametric models.\footnote{It does not specify the marginal distribution of the error terms, nor does it require serial independence.} Under Assumption (ref), $\Gamma_\theta(\mathcal W)$ is not convex.\footnote{The restriction that $V_2$ and $X_2$ are independent conditional on $(A,X_1)$ is nonlinear.} Nonetheless, the following result establishes convexity of the set of model probabilities $\mathcal M_\theta$ and uses Theorem (ref) to characterize the identified set.

thmConsider the model described by outcome equation (ref) and Assumption (ref). The set $\Thetam$ is the sharp identified set for $\theta=(\beta,\tau_1,\tau_2)$.

$\Thetam$ may be computed using Theorem (ref), the pushforward representation of Section (ref), and the extremal point representation of Proposition (ref). Setting $\psi_\theta$ as in (ref) and $W$ as in Section (ref), Theorem (ref) implies that $\Thetam$ is the set of $\theta$ satisfying

align[align omitted — 189 chars of source]

where $\Gamma_\theta(\mathcal{W})$ is the set of all distributions satisfying (ref). Let $\Gamma^\mathrm{seq}$ be the set of all distributions of $(V_1, V_2, X_2)$ which satisfy $V_2|X_2 \overset{d}{=} V_1$. Then, because both sides of (ref) are conditional on $(\alpha, X_1)$, the extremal points of $\Gamma_\theta(\mathcal{W})$ all take the form $\gamma \times \delta_{(a, x_1)}$, $\gamma \in \Gamma^\mathrm{seq}$. As in the proof of Proposition (ref), we may rewrite the right hand side of (ref) as the supremum over $a, x_1$ of

align[align omitted — 265 chars of source]

where $\gamma_{X_2}$ is a marginal distribution of $X_2$ and $\Gamma^\mathrm{seq}(\gamma_{X_2})$ is the subset of $\Gamma^\mathrm{seq}$ having $X_2$-marginal $\gamma_{X_2}$. The inner supremum in (ref) is an LP and the outer supremum is only over the space of distributions of $X_2$ and points $a, x_1$. Section (ref) below applies this extremal point characterization to obtain $\Thetam$ for various $\theta_0$.

Numerical experiment

figure[figure omitted — 1,158 chars of source]

We conduct a numerical experiment for two different DGPs. Using DGP1, we explore the size of the identified sets of the regression coefficient and the partial effects under strict exogeneity. Using DGP2, we explore the relative widths of the identified set of the regression coefficient under strict and sequential exogeneity.

DGP1. This DGP includes a time dummy $X_{1t} = t-1$ and a regressor $X_{2t}$ whose support we vary across designs as follows. In design $p$, $X_{21} = 0$ and $X_{22} \in \{-p,-p+1,\cdots,p\}$, so that higher values of $p$ imply more variation in the second regressor. All designs have $X$ discrete uniform, and $(\alpha,V_1,V_2)$ independent of $X$, with $P((V_1,V_2,\alpha) = (u_1,u_2,a) \propto \exp(-u_1^2/2 - u_2^2 /2 - a^2/2)$, with support for $V_1,V_2$ as $\{-3,-2.9,\cdots,2.9,3\}$ and the support of $\alpha$ as $\{-2,-1,\cdots,2\}$. The identified sets are computed via linear programming, see (ref) for details on the implementation.

Figure (ref) presents results for $\beta_0 = (1,-0.5)$ and designs $p \in \{4,5\}$. We plot the discrepancy function $T(\theta)$ against candidate values $\beta_2$. The identified set is smaller for Design 5 (green line) as expected, because of increased variation in $X$. Figure (ref) presents the identified set for $\beta_2$ as we vary the true value of the regression coefficient. As expected, additional variation in the regressors tightens the bounds on the regression coefficient, see Design 10 (purple line).\footnote{Manski1987 shows that point identification obtains under continuous variation in (one of the) regressors over the entire real line.} Figure (ref) presents results for the (joint) identified set for the regression coefficient and the ASF, at $\beta_2 = 0.5$. The ASF is for time period $t=1$, counterfactual value $x = 1$, and a subpopulation with $X_{21}=0,X_{22}=1$. The identified set for the ASF is the height of the box; the identified set for the regression coefficient is its width (coinciding with that in Figure (ref)). The ASF is not point identified because of the time dummy (see BotosaruMuris2024). The identified sets for the ASF are informative across all designs. The information about the regression coefficient does not translate to information about counterfactual parameters.

DGP2. We consider a worst case of the preceding design, with $X_{2t}$ independently and uniformly distributed on $\{0,1\}$. The time effect is $\beta_1 = 2$. The distribution of $\alpha + V_1$ is uniformly distributed on 11 equidistant points in the interval $[-3,3]$, and we set $V_1 = V_2$ (so strict exogeneity holds). We compute the identified set for $\beta_2$ under strict exogeneity as before, and under sequential exogeneity by applying (ref) as described in Section (ref).\footnote{A genetic optimizer from the deap Python package handles the outer optimization over $\phi \in \Phi_b$.}

Figure (ref) depicts identified sets for $\beta_2$ under the assumptions of sequential and strict exogeneity as $\beta_2$ is varied. The identified sets are larger than the ones under DGP1. This is as expected, as DGP2 has little variation in $X_{2t}$. The results clearly show that the identified set is larger under sequential exogeneity.

Parametric errors

Let $H(\cdot)$ denote an arbitrary cumulative distribution function. In this section, we assume that the conditional distribution of the error terms is known, and given by \[ P(V_t \leq v | \alpha = a, X = x) = H(v). \] Furthermore, we assume that the $V_t$ are independent across $t$ conditional on $(\alpha,X)$. This obtains the static binary choice model with fixed effects,

equation[equation omitted — 155 chars of source]

with $\alpha \in \mathcal A = \mathbb R$, and the sequence of regressors $X = (X_{1},\cdots,X_{T}) \in \mathcal X$. If $H(\cdot) = \Lambda(\cdot)$ is the standard logistic distribution, this is the static panel logit model raschStudiesMathematicalPsychology1960, rasch1961general. Setting $H(\cdot)=\Phi(\cdot)$ obtains the probit version, etc.

Conditional serial independence of $V_t$ implies that $(Y_{1},\cdots,Y_{T})$ are independent conditional on $(\alpha,X)$. For any $y = (y_1,\cdots,y_T) \in \{0,1\}^T$, we therefore have

equation[equation omitted — 198 chars of source]

so that the model fits the framework of Section (ref). Assumptions (ref)(i) and (ii) are trivially satisfied for any $H(\cdot)$. It follows that equation (ref) in Proposition (ref) characterizes the identified set for $\beta$.

We now describe results from a numerical experiment, starting from the special case of $T=2$, $X_1 = 0,\; X_2 = 1$.\footnote{A full description of the computation is in Appendix (ref), with details about this specific model in (ref).}

Figure (ref) plots $T(\theta)$ for the probit model with $\beta_0 = 1$, with $\alpha$ on a grid of $101$ equally spaced points from -5 to 5, and with $P(\alpha = a) \propto \exp(-a^2/2)$. We consider a grid for $\beta$ of 601 equally spaced points from 0.7 to 1.3. On a single core i7-11370H at 3.30GHz, it takes 0.0036 seconds to compute $T(\theta)$ per value of $\theta$. For more details on computation times, see Section (ref). For the probit model, $\beta_0$ is not point-identified, and we find $\Theta_I = [0.968,1.065]$.

Figure (ref) presents $\Theta_I$ for five different choices of $H(\cdot)$. “Logistic” corresponds to the standard logit model, for which we recover the well-known result that $\beta_0$ is point-identified. “Normal” refers to the result in Figure (ref). “Uniform” sets $H(\cdot)$ equal to the cumulative distribution function of the continuous uniform distribution on $[-2,2]$. “Cauchy” sets it equal to the standard Cauchy distribution, and “Truncated normal” sets it equal to a standard normal distribution truncated to $[-2,3]$.

The width of these identified sets is not directly comparable across distributions, due to their different scales. Nonetheless, the identified set is rather small across all distributions, especially given the limited variation in $X$. This is in line with previous findings on the size of identified sets in parametric nonlinear panels, see honoreBoundsParametersPanel2006,ChernValHahnNewey2013.

figure[figure omitted — 669 chars of source]

Our second result is for the regression coefficient and average treatment effects for the probit model with $T \in \{2,3\}$ and binary regressors $\mathcal X = \{0,1\}^T$. Similarly to the numerical experiments in ChernValHahnNewey2013, \S 8, we set $P(X = x) = 2^{-T}$ for all $x$, and the support and distribution of the fixed effects, independent of $X$. Finally, $\beta_0 \in \left\{0, 0.1, \cdots, 1.9, 2\right\}$ and $\beta \in \{\beta_0 - 0.3, \beta_0 - 0.299, \cdots, \beta_0 + 0.3\}$. Appendix (ref) describes how to modify the computation from the baseline case above to this more general setup.

Figure (ref) presents the identified sets as a function of $\beta_0$, for $T = 2$ and $T=3$. \footnote{These results recover those in the top left panel in Figures 2 and 3 in ChernValHahnNewey2013.} Unless $\beta_0 = 0$, the regression coefficient is partially identified. The identified set is small for $T=2$, much smaller for $T=3$, and indistinguishable from the 45-degree line when $T=4$ (not reported). The figures are based on 12621 values of $(\beta, \beta_0)$. The computation time per value is 0.013 seconds for $T=2$, and 0.066 seconds for $T=3$ on a single core i7-11370H at 3.30GHz.

figure[figure omitted — 541 chars of source]

Next, we compute the identified set for the average treatment effect of moving a randomly selected individual's $x_t$ from $0$ to $1$, i.e. \[ \text{ATE}(0, 1;\beta) = E[H(\alpha + \beta) - H(\alpha)]. \] Proposition (ref) applies, and (ref) gives an expression for the identified set. Figure (ref) presents the identified sets as a function of $\beta_0$, for $T = 2$ and $T=3$. \footnote{These results recover those in the bottom right panel in Figures 2 and 3 in ChernValHahnNewey2013.} The size of the identified sets is greatly reduced when moving from $T=2$ to $T=3$.

figure[figure omitted — 512 chars of source]
appendix\section{Additional remarks} Throughout the Appendices, we let $\Theta^0_\mathrm{I}$ denote the set: \begin{align} \Theta^0_{\mathrm{I}} &\equiv \{\theta \in \Theta: \exists \gamma\in\Gamma such that \mu^*_{Z} = \mu_{Z,(\theta,\gamma)} a.e. Z\}, \end{align} which is closely related to the definition of the identified set for $\theta$ in honoreBoundsParametersPanel2006. This set can be equivalently expressed as: \begin{align} \Theta^0_\mathrm{I} = \{\theta \in \Theta: \mu_Z^* \in \mathcal{M}_{\theta}\}. \end{align} \begin{remark}[The identified set and the closure of the set of model probabilities] The identified set \(\Theta_{\mathrm{I}}\) in (ref) is defined using the closure \(\overline{\mathcal{M}}_\theta\). We consider \(\Theta_{\mathrm{I}}\) instead of \(\Thetai\) in (ref) or (ref) for two reasons. First, we avoid issues related to infeasible inference. In Remark (ref), we show that elements of \(\Theta_{\mathrm{I}}\) are indistinguishable from those of \(\Thetai\), making it impossible to construct a hypothesis test that can differentiate between \(\Theta_{\mathrm{I}}\) and \(\Theta^0_{\mathrm{I}}\) with power exceeding size. Second, the need to define the identified set in terms of \(\overline{\mathcal{M}}_\theta\) also arises in nonlinear panel models. For instance, allowing the support of the fixed effects to include $\{-\infty,\infty\}$ is crucial, as discussed in Chamberlain2010. Without finiteness assumptions on the support, the model probability lies in \(\overline{\mathcal{M}}_\theta\) but not in \(\mathcal{M}_\theta\). Thus, the identified set must include values of \(\theta\) consistent with \(\mu^*_Z \in \overline{\mathcal{M}}_\theta\). Under finiteness assumptions, such as in honoreBoundsParametersPanel2006, there is no distinction between \(\mathcal{M}_\theta\) and \(\overline{\mathcal{M}}_\theta\), implying \(\Theta^0_{\mathrm{I}} = \Theta_{\mathrm{I}}\). schennachEntropicLatentVariable2014 also considers a refinement of the conventional notion of the identified set within her framework by defining it in terms of the closure of the set of all feasible values for the moment function. This allows for the support of the unobserved heterogeneity in her framework to be unbounded, see also li2019identification. \end{remark} \begin{remark}[Norm choice for defining the closure] The choice of norm defining the closure is important. In the literature on impossible inference, the total variation (TV) and the L\'evy-Prokhorov (LPv) norms play important roles, see e.g., BERTANHA2020247. Naturally, the choice of norm impacts the resulting identified set. For example, our identified set $\Thetaw$ involves the TV norm, and includes values of $\theta$ such that $\mu^*_Z \in \overline{\mathcal{M}}_\theta$. Our main result in Theorem (ref) establishes that $\Theta_\mathrm{I}$ includes values of $\theta$ that cannot be distinguished from those in $\Thetai$ with any bounded function $\phi:\mathcal{Z}\to [0,1]$. Instead, consider the closure of $\mathcal{M}_\theta$ with respect to the topology induced by the LPv norm, denoted by $\overline{\mathcal{M}}^\mathrm{LP}_\theta$, and define $\Theta^\mathrm{LP}_\mathrm{I} \equiv \{\theta \in \Theta: \mu_Z^* \in \overline{\mathcal{M}}^\mathrm{LP}_\theta\}$. It is possible to show that $\Theta^\mathrm{LP}_\mathrm{I}$ includes values of $\theta$ that cannot be distinguished from those in $\Thetai$ with any continuous and bounded function $\phi_c:\mathcal{Z}\to [0,1]$. Since $\overline{\mathcal{M}}_\theta \subseteq \overline{\mathcal{M}}^\mathrm{LP}_\theta$, it follows that $\Theta_\mathrm{I} \subseteq \Theta^\mathrm{LP}_\mathrm{I}$. Under additional assumptions, these sets can be made equal. For example, when $\mathcal{Z}$ is discrete, $\Theta_\mathrm{I} = \Theta^\mathrm{LP}_\mathrm{I}$ indicating robustness to the choice of norm. See BERTANHA2020247 for alternative assumptions on $\mu_{Z,(\theta,\gamma)}$ that lead to this equality result. \end{remark} \begin{remark}[Necessity of Assumption (ref)] Assumption (ref) is essential for Theorem (ref). Consider, for example, the case where \(\mu^*_Z\) is the Lebesgue measure on \([0,1]\) and, for some \(\theta \in \Theta\), \(\mathcal{M}_\theta = \mathrm{co}(\{\delta_z: z \in [0,1]\})\), which does not satisfy Assumption (ref). For all bounded Borel maps \(\phi: [0,1] \to [0,1]\), we have \(\mathbb{E}_{\mu^*_Z}[\phi] \leq \sup_{z \in [0,1]} \phi(z) = \sup_{\mu \in \mathcal{M}_\theta} \mathbb{E}_{\mu}[\phi]\), implying \(\theta \in \Theta^0_\mathrm{I}\). However, for every \(\mu \in \sum_{i=1}^k \alpha_i \delta_{z_i} \in \mathcal{M}_\theta\), \(d_{\mathrm{TV}}(\mu^*_Z, \mu) = 1\) since \(\mu^*_Z(\{z_1, \ldots, z_k\}) = 0\) and \(\mu(\{z_1, \ldots, z_k\}) = 1\). Therefore, \(\mu^*_Z \notin \overline{\mathcal{M}}_\theta\) and \(\theta \notin \Theta_\mathrm{I}\). \end{remark} \begin{remark}[The discrepancy function] The discrepancy function $T(\theta)$ in (ref) is new to the literature. Computing the zeros of this function involves a search over two infinite-dimensional spaces: the space of features $\Phi_b(\mathcal Z)$ and the space of model probabilities $\mathcal M_\theta$: \begin{align} T(\theta) \equiv \sup_{\phi \in \Phi_b(\mathcal Z)} \left( \EE{\mu_Z^*}{\phi} - \sup_{\mu \in \mathcal{M}_\theta}\EE{\mu}{\phi} \right). \end{align} The function $T(\theta)$ can be seen as a generalization of the Maximum Mean Discrepancy (MMD) measure defined in (ref) below. The MMD is an integral probability metric that has been used in the machine learning literature to discriminate between two observed probability measures, see, e.g., Gretton2007,Gretton2012. To see this, suppose that $\mathcal M_{\theta} = \{\overline{\mu}\}$ is a singleton (be it because the critic knows or has identified $\overline{\mu}$, the defender plays a fixed strategy $\overline{\mu}$, etc). Then the critic selects $\phi$ that solves: \begin{align} MMD(\mu^*_Z,\overline{\mu}) & \equiv \sup_{\phi\in\Phi_b(\mathcal{Z})}\left(\EE{\mu_Z^*}{\phi} - \EE{\overline{\mu}}{\phi}\right) \text{ for known } \mu^*_Z \text{ and } \overline{\mu}. \end{align} In our framework $\mathcal{M}_\theta$ is not a singleton. Consequently, \(T(\theta)\) can be seen as a generalization of the MMD to situations where $\overline{\mu}$ in (ref) is not known. In schennachEntropicLatentVariable2014 and li2019identification, see also guDualApproachWassersteinRobust2023 -- and more generally, approaches based on support-functions, e.g., beresteanuSharpIdentificationRegions2011 and ChesherRosenZhang, the criterion function that characterizes the identified set depends on a finite dimensional moment function $g\in\mathbb{R}^{d_g},\; d_g < \infty$, and it is given by: \begin{align} T_{\mathrm{SF}}(\theta) &\equiv \sup_{\eta\in\mathbb{R}^{d_g}:\vert\vert\eta\vert\vert=1}\mathbb{E}_{\mu^*_Z}[\inf_{u\in\mathcal{U}}\eta^\prime g\left(u,Z,\theta\right)], \end{align} where $\mathcal U$ is the support of the latent random variables. Supposing that $\mu$ is absolutely continuous with respect to $\mu^*_Z$, our criterion function $T(\theta)$ can be written as \begin{equation} T(\theta) = \sup_{\phi\in\Phi_b(\mathcal Z): \phi\in [0,1]}\inf_{\mu\in\mathcal M_\theta}\mathbb{E}_{\mu^*_Z}\left[\phi\left(\frac{d\mu}{d\mu^*_Z}-1\right)\right]. \end{equation} A comparison of (ref) and (ref) reveals several differences: (i) the supremum in (ref) is over a finite-dimensional vector, whereas in (ref), it is over a space of functions; (ii) the infimum in (ref) is taken inside the expectation, while in (ref), it is taken outside; and (iii) the infimum in (ref) is over the support of unobserved random variables, whereas in (ref), it is over a set of probability measures. Additionally, note that $T_{\mathrm{SF}}(\theta) = 0$ at $\theta = \theta_0$ (the true parameter value), whereas $T(\theta) = 0$ when $\mu^*_Z = \mu$ (the observed measure matches the model probability measure). These distinctions reflect fundamental differences in the approaches: an approach based on support functions versus one based on set membership for observed probability measures. Essentially, the difference in approaches is driven by the space in which they operate. Existing approaches operate in the space of random variables defined by (a finite number of) moment conditions, using hyperplanes to test specific linear combinations of these moments via support functions. In contrast, our framework operates in the infinite-dimensional space of probability measures over observables, where hyperplanes separate the true probability measure $\mu^*_Z$ from the set $\overline{\mathcal{M}}_\theta$. The hyperplanes in our framework correspond to bounded functions in $\Phi_b(\mathcal{Z})$. All approaches are grounded in convex analysis and leverage the fundamental principle of separating hyperplanes. For all approaches, convexity of the feasible set - whether in the moment space or the probability measure space - is essential for ensuring the applicability of separation results. Finally, note that the dual formulation of optimal transport resembles the MMD problem, despite arising from different principles. In the dual formulation, the separating hyperplane can be thought of as operating in the space of potential functions. These potential functions encode the cost structure and enforce the feasibility of the transport plan, often subject to smoothness assumptions and additional constraints dictated by the cost function. Moreover, as in the MMD problem, the dual formulation of optimal transport does not involve minimization over the model probability measure. \end{remark} \section{Proofs} Our main result, Theorem (ref), is an implication of the more general result in Proposition (ref) stated below. We first state and prove that more general result. The proofs of other results in the main text follow. Some supporting lemmas, and the proofs for the results in Section (ref), are in the supplementary material. The result here uses $\Thetai$ defined in (ref), $\Theta_\mathrm{I}$ defined in (ref), and $\Thetam$ defined in (ref). We denote by $\mu^n$ the product measure on $\mathcal{D}^n$ with the corresponding product topology. This is the joint distribution of $n$ points drawn independently from distribution $\mu$. \begin{proposition} For $n \in \N$ let $\M_\theta^{(n)} = \{\mu^n: \mu \in \M_\theta\}$. Consider the following statements: \begin{enumerate} • $\theta \in \Theta_{\mathrm{I}}$. • For all $n \in \N$ and all $f: \mathfrak{B}(\mathcal{Z})^n \ra \R$ continuous with respect to the total variation norm, $f((\mu_Z^*)^n) \in \overline{f(\M_\theta^{(n)})}$. \setcounter{enumi}{2} • $\theta \in \Thetam$. • For all compactly supported Borel $\phi: \mathcal{Z} \ra [0,1]$, \begin{align} \EE{\mu^*_Z}{\phi} \le \sup_{\mu \in \M_\theta} \EE{\mu}{\phi}. \end{align} \end{enumerate} Then the following are true: (a) Statements I and II are equivalent. (b) Statements III and IV are equivalent, and implied by statements I and II. (c) Let Assumptions (ref) and (ref) hold, and suppose that $\overline{\mathcal{M}}_\theta$ is convex. Then, statements I, II, III, and IV are equivalent. This equivalence holds for every $\mu^*_Z \in \mathcal{P}(\mathcal{Z})$ only if $\overline{\mathcal{M}}_\theta$ is convex. \end{proposition} \begin{proof} See page (ref). \end{proof} \begin{remark} The equivalence of I.\ and II.\ in Proposition (ref) implies that points \(\theta \in \Thetaw\) are statistically indistinguishable from those in \(\Thetai\). Specifically, if \(\theta \in \Thetaw\), any value attained by a function \(f((\mu^*_Z)^n)\), where \((\mu^*_Z)^n\) is the distribution of \(n\) i.i.d. observations from \(\mu^*_Z\) and \(f\) is continuous in the TV norm, must be a limit point of the image of \(\mathcal{M}_\theta^{(n)}\) under the same function. This makes it impossible to test the hypotheses: \begin{align*} &H_0: \theta \in \Thetai, \\ &H_1: \theta \notin \Thetai, \end{align*} with nontrivial power. To illustrate, suppose such a test \(\phi: \mathcal{Z} \to [0,1]\) exists with uniform size \(\alpha\) over \(\mathcal{M}_\theta\), the set of distributions of \(Z\) consistent with \(\theta\): \begin{align*} \sup_{\mu^n \in (\mathcal{M}_\theta)^n} \int \phi(z_1, \ldots, z_n) \, d\mu^n \le \alpha. \end{align*} Since the map \(\mu^n \mapsto \int \phi(z_1, \ldots, z_n) \, d\mu^n\) is continuous in the total variation norm, Proposition (ref) implies that the test's power to reject \(H_0\) when \(\mu^*_Z\) is the distribution of observables for \(\theta \in \Theta^0_\mathrm{I}\) cannot exceed \(\alpha\): \begin{align*} \int \phi \, d(\mu^*_Z)^n \le \alpha. \end{align*} \end{remark} \begin{proof}[Proof of Proposition (ref)] The proof of statements (a) and (b) is routine: \begin{enumerate} • If $\theta \in \Thetaw$, then for all $\ve > 0$ there is some $\mu \in \mathcal{M}_\theta$ which has $d_{\mathrm{TV}}(\mu^*_Z , \mu) < \ve$. For each such $\ve$ and $\mu$, $\mu^*_Z$ and $\mu$ have density with respect to the measure $\lambda \equiv \mu^*_Z + \mu$. Hence, by the arguments below, one has $d_{\mathrm{TV}}((\mu^*_Z)^n, \mu) < n \ve$ for $\ve $ arbitrary, and thus $(\mu^*_Z)^n \in \overline{\M_\theta^{(n)}}$. By a straightforward topological argument, $f(\overline{\M_\theta^{(n)}}) \subseteq \overline{f(\M_\theta^{(n)})}$, so II is implied by I. • It suffices to consider the case $n = 1$ and the function $f: \mu' \mapsto \inf_{\mu \in \M_\theta} \norm{\mu' - \mu}_{\mathrm{TV}}$, which is continuous by the triangle inequality. • This is straightforward from the definition of $\Thetam$, and by invariance of the upper bound in (ref) with respect to translations of $\phi$ and rescalings by positive constants. • We prove the contrapositive. If $\theta \not\in \Thetam$, there is some bounded Borel $\phi$ for which $\EE{\mu^*_Z}{\phi} > \sup_{\mu \in \mathcal{M}_\theta} \EE{\mu}{\phi}$. By taking a linear transformation, it may be assumed without loss of generality that $\phi$ is positive and bounded above by $1$. By the tightness of Polish Borel measures (VW1996, Lemma 1.3.2), there is some compact $K \subseteq \mathcal{Z}$ for which $\EE{\mu^*_Z}{\phi \cdot \one_K} > \sup_{\mu \in \mathcal{M}_\theta} \EE{\mu}{\phi} \ge \sup_{\mu \in \mathcal{M}_\theta} \EE{\mu}{\phi \cdot \one_{K}}$, so that IV.\ cannot be true. • For any bounded and Borel $\phi$, the map $\mu \mapsto \EE{\mu}{\phi}$ is continuous with respect to the total variation norm. Thus, (ref) follows from II. \end{enumerate} Now, make Assumptions (ref) and (ref), so that the measures $\mu \in \M_\theta$ and $\mu^*_Z$ are continuous (and have densities with respect to) the $\sigma$-finite measure $\lambda \equiv \mu^*_Z + \lambda_\theta$. Let $f[\mu]$ denote the mapping sending $\mu$ to its density with respect to $\lambda$. By Scheff\'{e}'s Lemma (tsybakov2008introduction, Lemma 2.1), $\norm{\mu^n-\phi^n}_{\mathrm{TV}} = \frac{1}{2} \norm{f[\mu^n] - f[\phi^n]}_{L^1(\lambda^n)}$ (so that $f$ is a linear isometry with respect to the total variation norm and $L^1$ topology), and by Lemma B.8 of Gh2017, this quantity is bounded above by $n\norm{\mu-\phi}_{\mathrm{TV}}$. Moreover, it follows that all measures $\mu$ and $\mu^*_Z$ invoked in this proof have densities in the Banach space $L^1(\lambda)$, which has as its dual $L^\infty(\lambda)$ (fremlin2000measure, Theorem 243G). Suppose now that $\overline{\M}_\theta$ is convex. Then, a straightforward argument shows that $\overline{\M}_\theta = \overline{\mathrm{co}}(\M_\theta)$. The Hahn-Banach theorem (C1994) and the observations above imply that \begin{align*} \overline{\M}_\theta &= \overline{\mathrm{co}}(\M_\theta) \\ & = \{\mu'\text{ continuous wrt }\lambda_\theta: \EE{\mu'}{\phi} \le \sup_{\mu \in \overline{\M}_\theta} \EE{\mu}{\phi} \,\forall \text{ bounded, Borel }\phi\}. \end{align*} Suppose that $\theta \in \Thetam$. Because we can take $\phi$ to be the indicator function of any Borel set in (ref), $\mu^*_Z$ must be continuous with respect to $\lambda_\theta$, and it follows from the previous display that $\mu^*_Z \in \overline{\M}_\theta$ and $\theta$ is in the identified set. On the other hand, suppose that for all Borel $\mu^*_Z$, $\theta \in \Thetam$ if and only if $\theta \in \Thetaw$. By the Hahn-Banach theorem, $\theta \in \Thetam$ if and only if $\mu^*_Z \in \overline{\mathrm{co}}(\M_\theta)$, and by definition, $\theta \in \Thetaw$ if and only if $\mu^*_Z \in \overline{\M}_\theta$. Thus, $\overline{\M}_\theta = \overline{\mathrm{co}}(\M_\theta)$, and the former set must be convex. \end{proof} \begin{proof}[Proof of Proposition (ref)] First, Assumption (ref) ensures that $\Gamma_\theta(\mathcal W)$ is convex. If there are no restrictions on $\gamma$, then $\Gamma_\theta(\mathcal{W}) = \mathcal{P}(\mathcal{W})$, which is inherently convex. When $\Gamma_\theta(\mathcal{W}) = \mathcal{P}(\mathcal{W})^g$, convexity is maintained since $\mathcal{P}(\mathcal{W})^g$ is the intersection of convex sets, each corresponding to a linear constraint imposed by the components of $g$. Then, by Corollary (ref), $\mathcal{M}_\theta$ is convex. By Theorem (ref), $\theta$ is in $\Thetaw$ if and only if \begin{align} \EE{\mu^*_Z}{\phi} \le \sup_{\gamma \in \Gamma_\theta} \EE{\gamma}{\phi \circ \psi_\theta}, \text{ for all } \phi \in \Phi_b(\mathcal{Z}), \end{align} (i) Suppose that $\Gamma_\theta = \mathcal{P}(\mathcal{W})$. Then the right hand side of (ref) is equal to the right hand side of (ref), because $\mathcal{P}(\mathcal{W})$ contains all of the Dirac measures $\delta_w, w \in \mathcal{W}$. (ii) Suppose that $\Gamma_\theta = \mathcal{P}(\mathcal{W})^g$. Let $\Delta^*(d_g)$ be the subset of $\Gamma_\theta$ consisting of measures supported on at most $d_g+1$ points: \begin{align*} \Delta^*(d_g) = \{\gamma \in \mathcal{P}(\mathcal{W})^g : \gamma = \sum_{j = 1}^{d_g + 1} c_j \delta_{w_j} : c_j \ge 0, w_i \in \mathcal{W}\}. \end{align*} Let $\gamma \in \Gamma_\theta$, so that $\EE{\gamma}{g} =0$. By the equality constraint case of Theorems 2.1 and 3.1 of winkler88 (c.f.\ 1907.07934, Theorem 2.1), there exists some probability measure $\nu$ supported on $\Delta^*(d_g)$ with the property that $\gamma$ is in the barycenter of $\nu$ and \begin{align*} \EE{\gamma}{\phi \circ \psi_\theta} = \int_{\Delta^*(d_g)}\EE{\gamma'}{\phi \circ \psi_\theta} \, \d \nu(\gamma') \end{align*} for all $\phi \in \Phi_b (\mathcal{Z})$. Hence, $\EE{\gamma}{\phi \circ \psi_\theta} \le \sup_{\gamma' \in \Delta^*(d_g)} \EE{\gamma'}{\phi \circ \psi_\theta}$. Taking the supremum over $\gamma \in \Gamma_\theta$ implies that $\sup_{\gamma \in \Gamma_\theta} \EE{\gamma}{\phi \circ \psi_\theta} = \sup_{\gamma' \in \Delta^*(d_g)} \EE{\gamma'}{\phi \circ \psi_\theta}$. This identity implies the equivalence of (ref) and (ref). \end{proof} \begin{proof}[Proof of Corollary (ref)] As Assumption (ref) ensures that $\mathcal{Z} = \mathcal{Y} \times \mathcal{X}$ is a Polish space, we verify that Assumption (ref) holds, and that $\mathcal{M}_\theta$ is convex, allowing us to invoke Theorem (ref). Assumption (ref) implies that every measure in $\mathcal{M}_\theta = (\psi_\theta)_* \Gamma_\theta (\mathcal{W})$ has a density with respect to the measure $\lambda_\theta$, which has the marginal distribution $\lambda_{\theta, \mathcal{X}}$ on $\mathcal{X}$ and conditional distributions $\lambda_{\theta, x}$ for all $x \in \mathcal{X}$. By the assumed measurability of the collection $\{\lambda_{\theta, x}: x \in \mathcal{X}\}$, this defines a measure on $\mathcal{Y} \times \mathcal{X}$ (see Section 10.4 in bogachev2007measure2). Indeed, if $A \subseteq \mathcal{Y} \times \mathcal{X}$ is a $\lambda_\theta$-null set, then for any $\gamma \in \Gamma_\theta$ and $\mu = (\psi_\theta)_* \gamma$ in $\mathcal{M}_\theta$, one has \begin{align*} \EE{\mu}{\one_{(Y,X) \in A}} & = \EE{\mu}{ \EE{\mu}{\one_{(Y,X) \in A}|X} } = \int_{\mathcal{X}} \int_{\mathcal{Y}} \one_{(y,x) \in A} \, \d h(x,\cdot;\theta)_*\gamma_{U|x} \, \d \gamma_X \\ & = \int_\mathcal{X} \int_\mathcal{Y} \one_{(y,x) \in A} \frac{\d h(x,\cdot;\theta)_*\gamma_{U|x}}{\d \lambda_{\theta, x}} \frac{\d \gamma_X}{\d \lambda_{\theta, \mathcal{X}}} \, \d \lambda_{\theta, x}\, \d \lambda_{\theta, \mathcal{X}} = 0. \end{align*} Finally, the convexity of $\mathcal{M}_\theta$ follows from the convexity assumptions of Assumption (ref), by Lemma (ref) in Section (ref), and by the linearity of the pushforward map $\psi_\theta$. \end{proof} \begin{proof}[Proof of Proposition (ref)] This proof makes use of Lemma (ref) and Lemma (ref) that can be found in the supplementary material. Convexity of $\mathcal{M}_\beta$ is a consequence of Assumption (ref)(ii) and Lemma (ref). An application of Lemma (ref) with $g = 0$ implies that $\beta$ is in the identified set if and only if $\EE{\mu^*_{Y,X}}{\phi(Y,X)} \le \sup_{\mu \in \mathcal{M}_\beta}\EE{\mu}{\phi(Y,X)}$ for all $\phi \in \Phi_b(\mathcal{Y} \times \mathcal{X})$. Because the parameter $\gamma_{X,\alpha}$ is not constrained in any way in the definition of $\mathcal{M}_\beta$, one has \begin{align*} \sup_{\mu \in \mathcal{M}_\beta}\EE{\mu}{\phi(Y,X)} = \sup_{\substack{ x \in \mathcal{X} \\ a \in \mathcal{A} }} \int_{\mathcal{Y}} \phi(y,x) f_{Y|X,\alpha} (y|x,a ; \beta) \, \d \lambda_{\mathcal{Y}}, \end{align*} which implies (ref). \end{proof} \begin{proof}[Proof of Proposition (ref)] This proof makes use of Lemma (ref) and Lemma (ref) that can be found in the supplementary material. Let $\mu_{Y,X} \in \mathcal{M}_\theta$ correspond to some $\gamma_{X,\alpha} \in \mathcal{P}(\mathcal{X} \times \mathcal{A})$ and let $\phi \in \Phi_b(\mathcal{Y} \times \mathcal{X})$. Then, \begin{align} \EE{\mu_{Y,X}}{\phi(Y,X)} = &\int_{\mathcal{X}} \int_{\mathcal{A}} \int_{\mathcal{Y}} \phi(y,x) f_{Y|X,\alpha}(y|x,a; \beta) \d \lambda_{\mathcal{Y}}\, \d \gamma_{\alpha|x} \, \d \gamma_X \text{, where} \nonumber\\ &\int_{\mathcal{A}} g(x,a, \beta, \tau) \, \d \gamma_{\alpha|x} = 0\text{, } \gamma_X\text{-almost surely}. \end{align} By modifying $\gamma$ on a set of measure $0$, we may assume that the second line of (ref) holds for all $x$. It follows that \begin{align} \EE{\mu_{Y,X}}{\phi(Y,X)} \le \sup_{\substack{x \in \mathcal{X}, \gamma_{\alpha|x} \\ \int_{\mathcal{A}} g(x,a, \beta, \tau) \, \d \gamma_{\alpha|x} = 0 }} \int_{\mathcal{A}} \int_{\mathcal{Y}} \phi(y,x) f_{Y|X,\alpha}(y|x,a; \beta) \d \lambda_{\mathcal{Y}}\, \d \gamma_{\alpha|x}. \end{align} By Theorems 2.1 and 3.1 of winkler88, the right hand side of (ref) is bounded above by the right hand side of (ref). On the other hand, for any particular $x \in \mathcal{X}$ and sequence of $c_j \in [0,1]$, $a_j \in \mathcal{A}$ meeting the constraints of (ref), the measure $\gamma_0 = \sum_{j = 1}^{d_g + 1} c_j \delta_{(x, a_j)}$ is in $\Gamma_{\tau}$, and letting $\mu_{Y,X}$ be the distribution associated with $\gamma_0$ has the property that \[ \EE{\mu_{Y,X}}{\phi(Y,X)} = \sum_{j = 1}^{d_g + 1} c_j \int_{\mathcal{Y}} \phi(y,x) f_{Y|X,\alpha} (y|x,a_j; \beta) \, \d \lambda_{\mathcal{Y}}. \] Hence, $\sup_{\mu \in \mathcal{M}_{\theta}} \EE{\mu}{\phi(Y,X)}$ is precisely the right hand side of (ref), and Lemma (ref) concludes the proof. \end{proof}

\setcounter{page}{1}

\setcounter{section}{0} {S\arabic{section}} \setcounter{subsection}{0} {S\arabic{section}.\arabic{subsection}} \setcounter{table}{0} {S\arabic{table}} \setcounter{figure}{0} {S\arabic{figure}} \setcounter{equation}{0} {S\arabic{equation}}