Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
118,872 characters · 19 sections · 85 citation commands
An Adversarial Approach to Identification
Identification of structural and counterfactual parameters is a central challenge in econometric models with unobserved heterogeneity. In many cases, the distribution of unobserved heterogeneity is not point identified, leading to partial identification of structural and/or counterfactual parameters. This issue is pervasive in nonlinear panel models with fixed effects. Fixed effects obstruct the point identification of both structural parameters and partial effects in all but a narrow class of models ArellanoBonhomme2012. This highlights the need for methods that achieve sharp identification for both types of parameters, while remaining computationally feasible and enabling valid inference.
Addressing these issues in nonlinear panel models is challenging. Many existing approaches are tailored to specific model features, relying on parametric assumptions about error distributions or support and exogeneity restrictions on the covariates. Additionally, existing methods focus on either structural or counterfactual parameters. This has led to a fragmented literature with various methods addressing only isolated aspects of the broader identification problem.
We introduce a novel and unifying framework for characterizing identified sets of structural and counterfactual parameters in econometric models with unobserved heterogeneity. Departing from existing approaches, we reformulate the identification problem as a set membership question in the space of observed probability measures (i.e., probability measures of observed random variables). This reformulation allows us to leverage the separating hyperplane theorem in the infinite-dimensional space of observed probability measures, providing a characterization of the identified set via the zeros of a new discrepancy function. The discrepancy function checks whether there exists at least one hyperplane that separates the observed probability measure from the set of model probabilities, defined as the collection of all probability measures of observed random variables consistent with a given parameter value. By aggregating over all hyperplanes and all measures in the set of model probabilities, the discrepancy function reveals whether the observed measure belongs to the set of model probabilities. The discrepancy function has a maximin representation or adversarial game interpretation, inspiring the name of our approach.\footnote{Our approach is distinct from the simulation-based adversarial estimation method of Kaji2023. Ours is an identification approach.}
We establish new sufficient and necessary conditions for sharp identification: when the set of model probabilities is convex, our approach obtains a sharp characterization of the identified set; when convexity fails, our approach characterizes an outer set. When the identified set is a singleton, we obtain point identification. Remarkably, many econometric models naturally feature a convex set of model probabilities.
A key innovation of our framework is its ability to accommodate both parametric and nonparametric error distributions, along with a wide range of exogeneity restrictions and conditioning variables --- whether continuous or discrete, strictly exogenous or predetermined. To demonstrate the power of our approach, we characterize the identified set for structural and counterfactual parameters in the semiparametric binary choice panel model under sequential exogeneity (or “predeterminedness”) and without parametric assumptions on the distribution of the error terms, addressing a long-standing gap in the literature on nonlinear panel models. We further highlight the versatility of our method by applying it to the binary choice panel model with error terms that follow a fixed but arbitrary distribution. For the nested case of logistic errors, we recover established results for the structural parameter and recently derived results for counterfactual parameters. These applications illustrate the versatility of our framework and its potential to advance partial identification in nonlinear econometric models.
Let $Z$ denote the observed random variables, and \(\mu^*_Z\) the observed probability measure characterizing the distribution of \(Z\). Let $\theta\in\Theta$ denote the parameter of interest, which can include both structural parameters and arbitrary functionals of the latent probability measure, and $\overline{\mathcal{M}}_\theta$ denote the closure of the set of model probabilities.\footnote{ Explicitly taking the closure ensures the inclusion of limit measures, such as those arising from degenerate distributions when latent random variables reach extreme values (e.g., fixed effects approaching \(\pm \infty\)). } We reformulate the identification problem as a set membership question: For a candidate parameter value \(\theta\), does \(\mu^*_Z\) belong to \(\overline{\mathcal{M}}_\theta\)? Accordingly, the identified set for \(\theta\) is given by:
This set contains all parameter values compatible with \(\mu^*_Z\). If \(\mu^*_Z \in \overline{\mathcal{M}}_\theta\), the observed probability measure and the model probabilities are indistinguishable at \(\theta\); otherwise, \(\theta\) is incompatible with \(\mu^*_Z\).
To determine membership of $\mu^*_Z$ in $\overline{\mathcal{M}}_\theta$, we leverage the separating hyperplane theorem. Specifically, when $\mu_Z^*\notin\overline{\mathcal{M}}_\theta$ and $\mathcal{M}_\theta$ is convex, there exists a hyperplane $\phi$ that separates $\mu_Z^*$ from $\mathcal{M}_\theta$. Aggregating across all hyperplanes $\phi$ leads to the following discrepancy function:
where $\mathcal{Z}$ denotes the support of $Z$ and $\Phi_b(\mathcal{Z})$ a set of bounded functions supported on $\mathcal{Z}$, defined in Section (ref).
The discrepancy function $T(\theta)$ evaluates to zero if and only if no separating hyperplane $\phi$ exists. Consequently, the zeros of $T(\theta)$ can be used to characterize $\Theta_\mathrm{I}$. This characterization is sharp provided that $\mathcal{M}_\theta$ is convex --- a property shared by many econometric models. When $\mathcal{M}_\theta$ is not convex, the zeros of $T(\theta)$ describe an outer set. When $\Theta_\mathrm{I}$ is a singleton, our framework obtains point identification of $\theta$.
In many econometric models, the discrepancy function admits a low-dimensional representation through an extreme point characterization. This facilitates the computation of zeros of $T(\theta)$ using linear programming. Two features of many models are central to this result: (i) the observed probability measure can be expressed as a linear transformation of the latent probability measure (i.e., the measure of latent random variables),\footnote{In semiparametric models with unrestricted error distributions, this transformation corresponds to a pushforward measure, whereas in models with parametric error distributions, it takes the form of a linear operator.} and (ii) the latent probability measures lie within a convex set.
For example, the semiparametric binary choice panel model under strict exogeneity exhibits both features, allowing for efficient computation of the identified set for the structural and counterfactual parameters using linear programming. Conversely, the semiparametric binary choice panel model under sequential exogeneity satisfies the linearity condition but not the convexity of the set of latent probabilities; instead, this model features a convex set of model probabilities $\mathcal M_\theta$, which yields sharp identification. However, computing the identified set in this case requires an extension of our linear programming approach. Beyond these computational advantages, the two features enable the sample analog of the discrepancy function to serve as a test statistic, facilitating valid inference. We establish the asymptotic distribution of this test statistic and construct critical values for hypothesis tests that are uniformly valid across the underlying probability distributions and parameter values.
A few key features distinguish our proposed method. The first distinguishing feature underscores the simplicity of our method in establishing sharp identification. A sufficient condition for sharpness is the convexity of $\mathcal M_\theta$. This can be established directly, or, given the linearity of the transformation, via the convexity of the set of latent probability measures. The latter holds in two key cases: (i) when no assumptions are imposed on the latent probability measure, such as in panel models with no distributional assumptions on the error terms, and (ii) when the latent probability measure is required to satisfy a finite number of linear restrictions. Such linear restrictions typically arise in two contexts: (i) when analyzing counterfactual parameters, many of which can be expressed as linear functionals of the latent probability measure,\footnote{See, e.g., ChristensenConnault2023, torgovitsky2019.} and (ii) when imposing exogeneity conditions, such as zero-mean or zero-median constraints on the error terms. We show this in our examples.
The second distinguishing feature is the versatility and broad applicability of our approach. Our method applies to many econometric models with unobserved heterogeneity, regardless of whether the models involve error terms that follow parametric distributions, or outcomes and covariates that are discrete or continuous. Our method accommodates various exogeneity restrictions on the covariates, and, if desired, restrictions on the latent probability measures.
The third distinguishing feature of our method is its computational ease. We implement our procedure and examine the size of the identified set in the canonical semiparametric binary choice model without parametric restrictions on the error terms, both in the cross-sectional case and in the panel case with strictly exogenous regressors, time effects, and fixed effects. With predetermined regressors, although the set of latent measures is not convex, the set of model probabilities is, so the identified set can be computed via an extension of our linear programming method. These illustrations to the semiparametric binary choice panel model with strictly exogenous or predetermined covariates contribute new results to the nonlinear panel literature.
The fourth distinguishing feature is that the sample analog of the discrepancy function can be used as a test statistic for valid inference.
This paper contributes to the literature on sharp identification in general classes of models and to the literature on nonlinear panel models.
A wide range of methods has been developed to characterize identified sets across different models.\footnote{For reviews, see BontempsMagnac2017,CanayShaikh2017,Molinari2020,ChesherRosen2020,KlineTamer2023.} Approaches include the use of random set theory beresteanuSharpIdentificationRegions2011, ChesherRosen2017, optimal transport GalichonHenry2009,EkelandGalichonHenry2010,GalichonHenry2011, and information-theoretic methods schennachEntropicLatentVariable2014. Other contributions, such as torgovitsky2019, focus on extending subdistributions, while more recent work explores minimum relevant partition and latent space enumeration Tebaldi2023,GuRussellStringham2022. Some of these methods focus exclusively on complete models, while others are explicitly designed to also address incomplete models.\footnote{In particular, models where the relationship between the observed random variables and the latent random variables is a correspondence, e.g., Tamer2003.} Many of these methods, like ours, leverage convex analysis for sharp identification or low-dimensional representations.
Our framework differs by reformulating the identification problem as a set membership question in the space of observed probability measures, treating \(\mathcal{M}_\theta\) as primitive. This allows us to (i) apply the separating hyperplane theorem in the space of observed probability measures, (ii) link sharpness to the convexity of \(\mathcal{M}_\theta\) (with outer sets characterized when convexity fails), and (iii) exploit two features common to many econometric models: the linearity of the transformation between observed and latent probability measures and the convexity of the set of latent probability measures. These properties facilitate the computation of the identified set via linear programming. This structure enables a unified approach for structural and counterfactual parameters, accommodating both parametric and nonparametric error distributions. This versatility is noteworthy. For example, methods designed for nonparametric error distributions rarely handle parametric restrictions (e.g., guDualApproachWassersteinRobust2023, ChesherRosenZhang), while parametric approaches often rely on specific distributional assumptions (e.g., Bonhomme2012,daveziesFixedEffectsBinary2023,GuRussellStringham2022).
While we focus on models where the relationship between observed and latent random variables is a mapping rather than a correspondence, our results can be extended to models involving correspondences. This is possible since our approach operates in the space of observed probability measures, and a sufficient assumption for all our results is that the observed probability measure is a linear map of the latent probability measure. Correspondences between random variables can induce such linear maps. However, as we show in the paper, such linear maps often arise in models commonly used in the literature on nonlinear panel models. Since our methodology is tailored to address a long-standing gap in the literature on nonlinear panel models, we leave a detailed investigation of correspondence-defined models for future research.
Identification challenges have long been a hallmark of nonlinear panel models due to the presence of fixed effects, see, e.g., arellanoPanelDataModels2001, honoreBoundsParametersPanel2006. As ArellanoBonhomme2012 note, point identification of structural parameters is rare, and even then, functionals of the fixed effects distribution, such as partial effects, remain only partially identified.\footnote{For exceptions of point-identified counterfactual parameters in specific models, see honoreMarginalEffectsSemiparametric2008, aguirregabiriaIdentificationAverageMarginal2024, and danoTransitionProbabilitiesIdentifying2023.} Addressing dynamic exogeneity in nonlinear models remains an open challenge, especially when both structural and counterfactual parameters are of interest, see HonoredePaula2021, ArkhangelskyImbens.
We showcase adversarial identification by applying it to the semiparametric binary choice panel model (see Manski1987). We obtain new results for the identified set for both structural parameters and the average structural function (ASF), without imposing parametric restrictions on the distribution of the error terms and under various dynamic exogeneity assumptions, such as strict and sequential exogeneity.
Most work on the semiparametric binary choice panel model focuses on the identification of structural parameters; for a non-exhaustive list, see Manski1987, aristodemouSemiparametricIdentificationPanel2021, khanIdentificationDynamicBinary2023, botosaruIdentificationTimevaryingTransformation2021, mbakopIdentificationDiscreteChoice2023, gaoIdentificationNonlinearDynamic2024, ChesherRosenZhang. The identification of partial effects has received less attention: ChernValHahnNewey2013 derive results under time-homogeneity and restrictions on the support of the covariates; BotosaruMuris2024 relax the latter assumptions while assuming that the structural parameters are a priori either point- or partially-identified.
When the error distribution is fixed but arbitrary, the structural parameters can be point-identified in a narrow class of models, see Bonhomme2012. Recent work focuses on partial identification of counterfactual parameters starting from point-identified or $\sqrt{n}$-estimable structural parameters, see dobronyiIdentificationDynamicPanel2021,daveziesFixedEffectsBinary2023,pakelBoundsAverageEffects2024. Seminal contributions by honoreBoundsParametersPanel2006 and ChernValHahnNewey2013 treat partial identification of structural parameters and average treatment effects, relying on discrete outcomes and covariates, and support conditions on the fixed effects.
Even with parametric restrictions on the error terms, identification of both structural and counterfactual parameters under sequential exogeneity is challenging, cf. arellanoPanelDataModels2001, ArellanoBonhomme2012, and HonoredePaula2021. Results for the identification of structural parameters under sequential exogeneity and error terms that follow parametric distributions can be found in arellanoBinaryChoicePanel2003, piginiConditionalInferenceBinary2022, chamberlainIdentificationDynamicBinary2023, and for partial effects in bonhommeIdentificationBinaryChoice2023.
In contrast, our approach obtains the identified set of structural and counterfactual parameters for models with known or unknown error distributions, does not require functional form restrictions such as an index structure or additivity in the unobservables, time-homogeneity, or support restrictions on the covariates or on the fixed effects. This flexibility extends to models with various forms of dynamic exogeneity --- strict, sequential, or lagged, offering new insights for nonlinear panel models. To the best of our knowledge, ours are the only identification results in nonlinear panel models with sequential exogeneity and without parametric restrictions on the error terms.
We introduce our main result in Section (ref). This result applies to a very general class of models. We specialize the result to semiparametric models with unobserved heterogeneity in Section (ref) and with observed covariates in Section (ref). We also illustrate our approach on the semiparametric binary choice model with a discrete regressor in Section (ref). Section (ref) shows that the discrepancy function and identified set can be computed via linear programming. Section (ref) extends all results to models with parametric error terms: Section (ref) focuses on the identification of common parameters, Section (ref) on the identification of both structural and counterfactual parameters, and Section (ref) discusses computation via linear programming. Section (ref) applies the results to panel models. Appendix (ref) contains additional remarks, including inference in Section (ref), while Appendix (ref) contains proofs not contained in the main text.
For a Polish space $\mathcal{S}$ endowed with its Borel $\sigma$-algebra, $\mathfrak{B}(\mathcal{S})$ denotes the set of all Borel measures on the set $\mathcal{S}$, $\mathcal{P}(\mathcal{S}) \subseteq \mathfrak{B}(\mathcal{S})$ denotes the set of all Borel probability measures supported on $\mathcal{S}$, and $\delta_{s}$ denotes the Dirac measure at $s\in\mathcal{S}$. For an index $\theta$, $ \Gamma_\theta(\mathcal{S})\subseteq \mathcal{P}(\mathcal{S})$ denotes a generic set of Borel probability measures supported on $\mathcal{S}$. The set $C_c(\mathcal{S})$ denotes the space of compactly supported continuous functions defined on $\mathcal{S}$. The product of two or more Polish spaces is endowed with the product topology.
For an arbitrary convex set $\mathcal{C}$ in a linear space, $\mathrm{ext}( \mathcal{C})$ denotes the set of all extreme points of $\mathcal{C}$, $\mathrm{co}(\mathcal{C})$ denotes the convex hull of $\mathcal{C}$, and $\overline{\text{co}}(\mathcal{C})$ denotes the smallest closed convex set containing the set $\mathcal{C}$.
We use $\subset$ to denote a strict subset, i.e. $A \subset B$ means that $x\in A \Rightarrow x \in B$ and $\exists b \in B: b \not \in A$. We use $A \subseteq B$ to denote that $A \subset B$ or $A = B$.
For arbitrary Borel measures $\mu, \mu'$, the total variation norm between $\mu$ and $\mu'$ is denoted by $\norm{\mu -\mu'}_{\mathrm{TV}} \equiv \sup_{B \text{ Borel}} |\mu(B) - \mu'(B)|$. For a given arbitrary measure $\mu$ and an arbitrary vector-valued function $f$, $f \in L^1(\mu)$ if each component function of $f$ is integrable with respect to $\mu$. We let $\EE{\mu}{\cdot}$ denote integration against a general probability measure $\mu$.
We denote by $\Phi(\cdot)$ the cumulative distribution function of the standard normal distribution, and by $\Lambda(\cdot)$ the cumulative distribution function of the logistic distribution.
Let $Z$ denote an observable Borel measurable random variable supported on a space $\mathcal{Z}$, and let $\mu^*_{Z}$ denote the true probability measure of $Z$. Denote by $\Theta$ the set of possible values of the parameter of interest $\theta$, and by \(\Gamma_\theta\) a set of auxiliary parameters $\gamma$ that may vary with $\theta$. The parameter of interest may include functions of $\gamma\in\Gamma_\theta$. For example, in the binary choice model of Section (ref), $\gamma$ is the unknown distribution of an unobservable error term and $\theta$ includes counterfactual choice probabilities, which are functionals of $\gamma$.
For each $\theta \in \Theta$ and $\gamma \in \Gamma_\theta$, the econometric model for $Z$ specifies a probability measure $\mu_{Z,(\theta,\gamma)}$, which we call the model probability. For fixed $\theta$, define the set of model probabilities as:
This set is the collection of all probability measures of $Z$ that are consistent with the econometric model under parameter value $\theta$. The set is a fundamental object for our analysis, and its geometric properties are essential for our main result in Theorem (ref) below.
The identified set for $\theta$ is defined as the set of parameter values compatible with the probability measure $\mu_Z^*$, formally defined in (ref), where $\overline{\mathcal{M}}_{\theta}$ is the closure of $\mathcal{M}_\theta$ with respect to the topology induced by the total variation (TV) norm:
Defining the identified set through the closure \(\overline{\mathcal{M}}_{\theta}\) offers two benefits. First, it explicitly includes limit measures, which may arise when, e.g., $\gamma$ converges to a degenerate measure.\footnote{Alternatively, such measures could be directly incorporated into $\mathcal M_\theta$.} Second, it ensures that the identified set includes all parameter values $\theta$ for which the model measures are indistinguishable from the observed measures in the total variation norm, thereby avoiding issues related to impossible inference. Further details are provided in Section (ref).
Computing the identified set based on (ref) involves a search over $\overline{\mathcal M}_\theta$, which can be challenging when $\Gamma_\theta$ is an infinite-dimensional space. This motivates the alternative characterization of the identified set in Theorem (ref) below, which uses the following assumptions.
Assumption (ref) accommodates random variables $Z$ supported on various separable spaces, and excludes spaces that are non-metrizable, which rarely arise in econometric applications. Assumption (ref) requires that measures in $\mathcal{M}_\theta$ be well-behaved with respect to a $\sigma$-finite Borel measure $\lambda_{\theta}$ that is allowed to vary with $\theta$. The assumption is mild, allowing for a broad class of random variables, including continuous, discrete, and mixed types, and excluding singular measures. We discuss the necessity of this assumption in Remark (ref) in Section (ref), while in Section (ref) we relax it.
Consider the discrepancy function defined in (ref) with
and define the set of parameter values that set (ref) to zero as:
Our main result below clarifies the connection between $\Theta_\mathrm{I}$ in (ref) and $\Thetam$ in (ref), and serves as a building block for the subsequent analysis.
Convexity of \(\mathcal{M}_\theta\) plays a central role in our approach.\footnote{While this convexity may influence the geometry of the identified set, convexity of the identified set itself is neither implied by nor required for our results.} First, it serves as a sufficient condition for sharp identification by ensuring that \(\overline{\mathcal{M}}_\theta\) is convex. When convexity of \(\mathcal{M}_\theta\) fails, the characterization via (ref) provides an outer set. Second, if extreme points of \(\mathcal{M}_\theta\) exist, convexity of \(\mathcal{M}_\theta\) reduces the search in (ref) over \(\mu \in \mathcal{M}_\theta\) to a search over these extreme points. While characterizing these extreme points can be a complex, model-specific task, we show that many econometric models have a specific feature that renders this step unnecessary. This feature refers to the fact that the model probability can be expressed as a linear map on the space \(\Gamma_\theta\). That is,
where \(\mathcal{T}_\theta: \Gamma_\theta \to \mathcal{P}(\mathcal{Z})\) is a linear map. We illustrate this structure with concrete examples in subsequent sections: semiparametric models where the error distribution is either unrestricted or linearly restricted (see Section (ref)) or parametrically specified (see Section (ref)). The representation in (ref) implies that convexity of \(\Gamma_\theta\) is sufficient for Theorem (ref) to apply,\footnote{This assumption is not necessary. We show in Section (ref) that under sequential exogeneity, \(\mathcal{M}_\theta\) is convex despite \(\Gamma_\theta\) failing to be convex.}enabling dimensionality-reduction via a search over the extreme points of the convex set \(\Gamma_\theta\), which are easier to characterize.
Both linearity of the map and convexity of \(\Gamma_\theta\) enable the computation of the zeros of \(T(\theta)\) via a linear program (see Sections (ref) and (ref)), making computation of \(\Theta_{\mathrm{I}}\) through \(\Theta_{\mathrm{MI}}\) a feasible task, even when direct computation using (ref) may be impractical.
We provide a heuristic interpretation for the discrepancy function $T(\theta)$ and Theorem (ref). The maintained assumption for the discussion here is that $\mathcal{M}_\theta$ is convex.
Given a \(\theta \in \Theta\) and a probability measure \(\mu^*_Z \in \mathcal{P}(\mathcal{Z})\), there may exist several \(\gamma \in \Gamma_\theta\) such that the corresponding model probabilities \(\mu_{Z,(\theta,\gamma)} \in \mathcal{M}_\theta\) are indistinguishable from the true \(\mu^*_Z\) in the sense of (ref). If there exists at least one such \(\gamma\),\footnote{Or if a sequence of \(\gamma\)'s can be constructed such that \(\mu_{Z, (\theta, \gamma)}\) converges to \(\mu^*_Z\).} then \(\mu^*_Z \in \overline{\mathcal{M}}_\theta\), and consequently, \(\theta \in \Theta_{\mathrm{I}}\). Theorem (ref) uses insights from convex analysis to solve this existence problem. In particular, the proof of Theorem (ref) shows that $\mu^*_Z \in \overline{\mathcal{M}}_\theta$ if and only if
The discrepancy function \(T(\theta)\) in (ref) is then obtained after rearranging and taking the supremum over \(\phi\). The decision rule is based on the sign of \(T(\theta)\): If \(T(\theta) > 0\), the parameter \(\theta\) is excluded from the identified set, while if \(T(\theta) \leq 0\), \(\theta\) is included in the identified set. When $\mathcal M_\theta$ is convex, this characterization is sharp.
The discrepancy function $T(\theta)$ has an adversarial formulation where two opposing players, a critic and a defender, interact strategically. The critic selects a feature $\phi\in\Phi_b(\mathcal Z)$ and the defender selects a measure $\gamma\in\Gamma_\theta$. The critic seeks to maximize the discrepancy between the feature observed in the data, i.e. $\mathbb{E}_{\mu^*_Z}[\phi]$, and the corresponding feature predicted by the model, i.e. $\mathbb{E}_{\mu}[\phi]$, for a given parameter value $\theta$ and taking into account the defender's action. In response, the defender adjusts $\gamma$ to minimize this discrepancy. The resulting discrepancy function captures the maximum discrepancy that the critic can enforce, even after the defender optimally adjusts the probability measure of the unobserved heterogeneity. The sign of the discrepancy function determines a decision rule: A positive value means that the critic has identified a feature where the model, under a specified parameter value, fails to replicate the observed data; the parameter value is then excluded from the identified set. A non-positive value means that the defender can always find a measure that aligns the model's prediction with the observed data; the parameter value is then included in the identified set. Note that if \(T(\theta) \leq 0\) for all \(\phi\), the critic selects \(\phi = 0\) to ensure \(T(\theta) = 0\).
This interpretation, together with the central role of the discrepancy function in our identification, computation, and inference results, forms the basis of our adversarial approach.
Models in this section are described as follows. Let the input variables be denoted by \(W \in \mathcal{W}\) (which may include latent variables) and continue to denote the observed random variables by \(Z \in \mathcal{Z}\). \footnote{For example, \(Z\) may consist of observed outcomes \(Y\) and conditioning covariates \(X\), such as regressors and instrumental variables, whereas \(W\) may contain \(X\) along with stochastic error terms or other latent random variables. See Section (ref) for further discussion.} In this setting, \(\Gamma_\theta \subseteq \mathcal{P}(\mathcal{W})\) is the set of probability measures supported on \(\mathcal{W}\) that are allowed by the model under the parameter value \(\theta\); we denote this set by \(\Gamma_\theta(\mathcal{W})\) to emphasize its dependence on \(\mathcal{W}\).
For any $\theta\in\Theta$, there exists a measurable map $\psi_\theta:\mathcal{W} \mapsto \mathcal{Z}$ known up to $\theta$, such that
This specification includes semiparametric models with, e.g., outcome equations such as $Z=h(\beta,W)$, $\theta=(h,\beta)$, with $h$ an unknown function and $\beta$ an unknown finite-dimensional parameter, and models with outcome equations such as $Y=m(W)$, where $\theta=m$ is an unknown function.
For fixed $\theta$, $\psi_\theta$ induces a pushforward measure $(\psi_\theta)_*: \Gamma_\theta(\mathcal W) \ra \mathcal{P}(\mathcal{Z})$ defined as:
where \(\psi_\theta^{-1}(S) = \{ w \in \mathcal{W} : \psi_\theta(w) \in S \}\) is the preimage of \(S\) under \(\psi_\theta\). Then,
Assumption (ref) is naturally satisfied in the class of models with outcome equations described by (ref). Here, the linear map in (ref) is the pushforward measure $(\psi_\theta)_*$, so that measures on \(\mathcal{Z}\) are obtained by “pushing forward” \(\gamma \in \Gamma_\theta(\mathcal{W})\) via \((\psi_\theta)_*\). \footnote{While $\psi_\theta$ here is not a correspondence, our result in Corollary (ref) may still apply to such models. This is because a sufficient condition for that result is the existence of $\mathcal T_\theta$ in (ref), which can be induced by a correspondence between $\mathcal W$ and $\mathcal Z$. In Section (ref), we extend our results to cover linear operators from \(\Gamma_\theta(\mathcal{W})\) to \(\mathcal{P}(\mathcal{Z})\). Notably, a correspondence between \(\mathcal{W}\) and \(\mathcal{Z}\) can induce a map or an operator between spaces of probability measures on \(\mathcal{W}\) and \(\mathcal{Z}\). Deriving sufficient conditions for when such correspondences lead to well-defined linear maps or operators is left for future work.}
The set $\Gamma_\theta(\mathcal W)$ plays a central role in our analysis. In some semiparametric models, it is unrestricted, i.e. $\Gamma_\theta(\mathcal{W}) = \mathcal{P}(\mathcal{W})$ is the set of all probability measures on $\mathcal W$, while in other models, $\gamma$ is known to satisfy certain linear restrictions, i.e. $\Gamma_\theta(\mathcal{W}) = \mathcal{P}(\mathcal{W})^g$ which we define below. Models with parametric restrictions on $W$ are discussed in Section (ref).
For a known vector of functions $g:\Theta\times\mathcal{W}\to\mathbb{R}^{d_g}$, define
This is the set of all probability measures on $\mathcal W$ that satisfy a set of $d_g<\infty$ linear restrictions that may depend on $\theta$. As will become clear from examples throughout this paper, such restrictions are important in many econometric models. First, many restrictions commonly made on the distribution of latent variables, such as mean- or median-independence, can be expressed as in (ref). Second, many counterfactuals of interest can be cast in the form $\EE{\gamma}{g(\theta, w)}=0$, implicitly imposing linear restrictions on \(\gamma\). Section (ref) illustrates this for a semiparametric binary choice model.
Corollary (ref) shows that convexity of $\Gamma_{\theta}(\mathcal W)$ is sufficient for sharp identification. Under Assumption (ref) below, each of $\mathcal{P}(\mathcal{W})$ and $\mathcal{P}(\mathcal{W})^g$ is convex, and Corollary (ref) applies.
We are now ready to establish an extremal point characterization of the result in Theorem (ref). Characterizing the identified set using (ref) involves a search over the space $\mathcal{M}_\theta$. Given Assumption (ref), the search can instead be conducted over $\Gamma_\theta(\mathcal{W})$, i.e. for any $\phi\in\Phi_b(\mathcal{Z})$,
Proposition (ref) below shows that this search can further be confined to a smaller space.
This result has important implications. For instance, when \(\gamma\) is unrestricted, so that \(\Gamma_\theta(\mathcal{W}) = \mathcal{P}(\mathcal{W})\), Proposition (ref)(i) implies that the identified set can be characterized by searching over the set of extreme points of \(\Gamma_\theta(\mathcal{W})\), all of which correspond to Dirac measures on \(\mathcal{W}\). Thus, the search is restricted to \(\mathcal{W}\) rather than the much larger set of probability measures on \(\mathcal{W}\). This refinement is particularly useful in nonlinear panel models, as it means that determining whether \(\theta \in \Theta_I\) requires only a search over the values of the fixed effects, rather than over all possible conditional distributions of the fixed effects.
In this section, we focus on semiparametric models with both observed and unobserved heterogeneity. Our primary goal is to clarify how conditioning variables (including regressors and instrumental variables) are treated within our framework. We also provide a blueprint for verifying the assumptions of Corollary (ref) for models with discrete and continuous conditioning variables. Section (ref) provides an example.
Many semiparametric models have an outcome equation of the form:
where $Y \in \mathcal{Y}$, $X \in \mathcal{X}$, and $U \in \mathcal{U}$ are random variables, $\theta \in \Theta$ is an unknown parameter, and $h:\mathcal{X} \times \mathcal {U} \times \Theta \to \mathcal{Y}$ is a (structural) function known up to $\theta$. $Y$ denotes outcome variables, $X$ denotes observed variables, and $U$ unobserved variables.
Using the notation from Section (ref), let \(Z = (Y, X)\), \(W = (X, U)\), \(\mathcal{Z} = \mathcal{Y} \times \mathcal{X}\), and \(\mathcal{W} = \mathcal{X} \times \mathcal{U}\). Denote the distribution of $W$ supported on $\mathcal W$ by \(\gamma \in \Gamma_\theta(\mathcal{W})\subseteq\mathcal{P}(\mathcal{W})\), the marginal distribution of \(X\) under \(\gamma\) by \(\gamma_\mathcal{X}\in\Gamma_{\theta, \mathcal{X}}(\mathcal{X})\subseteq \mathcal{P}(\mathcal{X})\), and the conditional distribution of \(U \mid X = x\) for each \(x \in \mathcal{X}\) by \(\gamma_{U \mid x}\in\Gamma_{\theta, x}(\mathcal{U})\subseteq \mathcal{P}(\mathcal{U})\). By the disintegration theorem, measures \(\gamma \in \Gamma_\theta(\mathcal{W})\) have differential \(\mathrm{d}\gamma = \mathrm{d}\gamma_{U|x} \, \mathrm{d} \gamma_X\) for all \(x \in \mathcal{X}\), allowing thus for correlation between $X$ and $U$.
Finally, for any \(\theta \in \Theta\), the mapping
induces, for each \(\gamma \in \Gamma_\theta(\mathcal{W})\), the pushforward measure \((\psi_\theta)_*\) on \(\mathcal{P}(\mathcal{Z})\), as defined in (ref). Consequently, the set \(\Gamma_\theta(\mathcal{W})\) induces the set \(\mathcal{M}_\theta = (\psi_\theta)_* \Gamma_\theta(\mathcal{W})\) as in (ref).
The assumption below allows us to specialize the assumptions of Corollary (ref) to models with outcome equation as in (ref).
Assumption (ref)(i) is a regularity condition which fulfills the requirements of Assumption (ref). Assumption (ref)(ii) specifies sufficient conditions for the convexity of $\Gamma_\theta(\mathcal{W})$, and consequently the convexity of $\mathcal{M}_\theta$. Typically, $X$ are treated as conditioning variables, in the sense that the model specifies assumptions on features of the conditional distribution $\gamma_{U|x}, \; x\in\mathcal X$. By including $X$ in $W$, typical conditions on $\gamma_{U|x}$ can be formulated as linear restrictions on \(\gamma\) as in (ref). If, additionally, $\mathcal U$ is a Polish space so that Assumption (ref) holds, convexity of $\Gamma_\theta(\mathcal W)$ is preserved and straightforward to verify by determining whether assumptions on $U \mid X$ can be expressed as in (ref). Alternatively, convexity of $\Gamma_\theta(\mathcal W)$ can be established by verifying Assumption (ref)(ii). Note that here $\gamma_\mathcal X$ is treated as a latent distribution. Since $X$ enters $Z$, it is possible to treat $\gamma_X$ as known and equal to the marginal distribution of $X$ under the observed distribution of $Z$. In this case, $\Gamma_{\theta,\mathcal X}$ is a singleton and trivially convex. Convexity of $\Gamma_{\theta, x}$ is enforced via assumptions on $\gamma_{U \mid x}$ that guarantee convexity of the set for every $x\in\mathcal X$.
Assumptions (ref)(iii) and (iv) are mild continuity conditions that guarantee that all measures in $\mathcal{M}_\theta$ are continuous with respect to the $\sigma$-finite measure $\lambda_\theta \in \mathfrak{B}(\mathcal{Y} \times \mathcal{X})$ with differential $\d \lambda_\theta = \d \lambda_{\theta, x} \, \d \lambda_{\theta, \mathcal{X}}$ for all $x \in \mathcal{X}$, thereby satisfying Assumption (ref). These assumptions require, respectively, that the marginal and conditional distributions of measures in $\mathcal{M}_\theta$ have density with respect to $\sigma$-finite measures, which may depend upon $\theta$. These dominating measures may be continuous, discrete, or a mixture of both. For example, when $\mathcal{Y}$ is discrete, $\lambda_{\theta, x}$ is the counting measure on $\mathcal{Y}$ for all $x$, in which case Assumption (ref) is fulfilled. In contrast to much of the existing relevant literature, Assumptions (ref)(iii) and (iv) allow for $(Y,X)$ to be continuous, discrete, or a mixture. Importantly, Corollary (ref) pertains to semiparametric models with outcome equations of the form (ref), where, for example, $X$ is continuous with respect to Lebesgue measure $\lambda_{\theta,\mathcal{X}}$ and either $Y$ is discrete with $\lambda_{\theta,x}$ the counting measure or $Y$ is continuous with $\lambda_{\theta,x}$ the Lebesgue measure (or some mixture of these cases).
Corollary (ref) provides a blueprint for checking the assumptions of Corollary (ref), and, consequently, for establishing sharp identification in semiparametric models with an outcome equation as in (ref).
We illustrate our approach using the model in manskiMaximumScoreEstimation1975 with a discrete regressor. We characterize the identified set for both regression coefficients and counterfactual choice probabilities. \footnote{The purpose of this section is to illustrate our approach rather than to obtain new results. For existing results, see komarovaBinaryChoiceModels2013, blevinsNonStandardRates2015, and torgovitsky2019. }
Consider the binary choice model with outcome equation:
where $Y\in\{0,1\} = \mathcal Y$, $X\in\{x_1,\cdots,x_K\} = \mathcal X$ has $K$ points of support, $U\in\mathcal U = \mathbb{R}$ is a scalar error term that satisfies the following conditional median-zero assumption:
and $\beta = (\beta_0,\beta_1)$ are unknown structural parameters. The counterfactual choice probabilities are given by:
which correspond to the average counterfactual outcome obtained by exogenously setting the values of the observed regressors to counterfactual values $x^*$ for the subpopulation given by $X=x_k$.
To characterize the identified set of
we verify the assumptions of Corollary (ref). Using the notation of the previous sections, $Z=(Y,X),\; W=(X,U)$, $h$ is given by (ref), $\gamma_x$ denotes the marginal distribution of $X$, $\gamma_{U|x}$ denotes the conditional distribution of $U \mid X=x, \; x\in\mathcal X$ satisfying (ref), and $\gamma \in \Gamma_\theta(\mathcal{W})$ denotes the distribution of $W$ with differential \(\mathrm{d}\gamma = \mathrm{d}\gamma_{U|x} \, \mathrm{d} \gamma_X\) for all \(x \in \mathcal{X}\).
Assumption (ref)(i) is trivially satisfied because $\mathcal Y$ and $\mathcal X$ are finite. The same is true for (ref)(iii) and (ref)(iv) with $\lambda_{\theta,\mathcal X}$ and $\lambda_{\theta,x}$ counting measures. (ref)(ii) is also satisfied: (a) $\Gamma_{\theta, \mathcal X}$ is a point, so convexity is trivially satisfied; (b) for each $x$, the set $\Gamma_{\theta,x}$ is the set of all probability measures $\gamma_{U|x}$ that satisfy the linear restrictions in (ref) and (ref), so $\Gamma_{\theta,x}$ is convex for all $x$ (see proof of Proposition (ref)). To see that (ref) and (ref) impose linear restrictions on $\gamma_{U|x}$ for each $x\in\mathcal X$ consider that these restrictions can be written as $$E_{\gamma_{U|x}}\left[ \widetilde g(U,\theta) \mid X = x\right] = 0$$ where
Because Assumption (ref) is satisfied, Corollary (ref) guarantees sharp identification of $\theta$ via $\Thetam$. Moreover, the results of Proposition (ref) also apply.
It follows that the identified set can be computed using the linear programming approach developed in Section (ref). Section (ref) describes in detail how to apply it to the semiparametric binary choice model, and presents numerical results. The computation time for the identified set of $\theta$ is trivial, see Section (ref).
The computation of $\Thetam$ requires computation of the discrepancy function (ref). This may seem challenging, as it involves a search over measures in $\mathcal{M}_\theta$ and over functions $\phi \in \Phi_b(\mathcal{Z})$. However, computing $T(\theta)$ only requires solving a linear program (LP). Below, we provide a sketch, leaving the details to Section (ref). The approach extends to the models in Section (ref).
Our theoretical results allow for $\mathcal Z$ and $\mathcal W$ to be Polish spaces, cf. Assumptions (ref) and (ref). For the purpose of computation, we may restrict attention to the case of finite supports, \( \mathcal Z = \{z_1,\cdots,z_L\}, \; \mathcal W = \{w_1,\cdots,w_M\}. \) Note that the linear program presented here does not rely on discreteness of $\mathcal Z$ and $\mathcal W$. Rather, it relies on the linearity of the map between probability measures as in (ref). When $\mathcal Z$ and $\mathcal W$ are discrete, the map takes the form of a matrix, which is the case here.
The probability measure $\mu_Z^*$ can be represented by a probability mass function (pmf) $p_Z^* = \left(p_{Z,l}^*\right)$, an $L \times 1$ probability vector, so that the first term in (ref) is \[ \EE{\mu_Z^*}{\phi} = \sum_{l = 1}^L \phi(z_l) p_{Z,l}^* = \phi^\prime p_Z^*. \]
Similarly, every model probability $\mu_{Z,(\theta,\gamma)}$ corresponds to a pmf $p_{Z}^{(\theta,\gamma)} = \left(p_{Z,l}^{(\theta,\gamma)}\right) $, and every probability measure $\gamma$ for $W$ to an $M\times 1$ pmf $p_W = \left(p_{W,m}\right)$. By the pushforward representation in Section (ref), cf. (ref) and Assumption (ref), there exists a matrix $\widetilde C_\theta \in \mathbb R^{L \times M}$ such that
For a given $\theta$, this pushfoward matrix maps the pmf of $W$ to a model pmf of $Z$ under parameter value $\theta$. For examples of $\widetilde C_\theta$, see Sections (ref) and (ref).
The second term of (ref) can be written as
so the infimum over $\mu$ in (ref) can be replaced by a minimum over $p_W$ subject to
These constraints enforce that $p_W$ is a probability vector and allow for additional constraints corresponding to $\Gamma_\theta$. Without additional constraints, $A_\theta = \iota_M^\prime$ and $b_\theta = 1$. For examples of other constraints, see Sections (ref) and (ref).
Taken together, we have
where
In Section (ref), we show that the bilinear program in (ref) is equivalent to the LP
where $\lambda$ is the Lagrange multiplier associated with the equality constraints in (ref). We conclude that determining whether $\theta \in \Theta_I$ amounts to solving the LP (ref) and checking that $T(\theta) \leq 0$. This task has negligible computation time even for very large $(L,M)$, see for example the computation times reported in Section (ref). It is also trivial to write the code for a specific model. The user specifies (i) supports $\mathcal Z, \, \mathcal W$; (ii) true parameter values $(\theta^*,p_W^*)$; (iii) the pushforward matrix $\widetilde C_\theta$; (iv) restrictions $(A_\theta,b_\theta)$; then hands off the LP (ref) to a solver.
In this section, the latent component $U$ in (ref) consists of two components: one that follows a known distribution and another whose distribution is unrestricted. The unrestricted component is labeled $\alpha$ to evoke the literature on parametric nonlinear panel models, see, e.g., Bonhomme2012.
We consider the following semiparametric regression model:
where $\beta \in \mathrm{B}$ is the parameter of interest, the observed random variables are $Z=(Y,X)$, the latent random variables $U=(\alpha,V)$ consist of a component $\alpha\in\mathcal{A}$ and a component $V\in\mathcal{V}$, where $\gamma_{X,\alpha} \in \mathcal{P}(\mathcal{X} \times \mathcal{A})$ denotes the joint distribution of $(X,\alpha)$ and $F_{V|X,\alpha;\beta}$ denotes the distribution of $V$ that is allowed to depend on $(X,\alpha,\beta)$. Here, $\gamma_{X,\alpha}$ is unknown and unrestricted, while $F_{V|X,\alpha;\beta}$ is known up to $\beta$.
In Section (ref), we characterize the identified set for $\beta$. In Section (ref), we characterize the identified set of counterfactual parameters. In Section (ref), we extend the LP approach in Section (ref). In Section (ref), we apply the tools developed here to a parametric binary choice model with fixed effects.
For the models in this section, the pushforward representation for the model probability in Assumption (ref) does not hold. Instead, we show that the model probability can be represented as a linear integral operator similar to the one in Bonhomme2012.
In the following assumption, $f_{Y|X,\alpha}$ is the conditional density of $Y$ given $(X,\alpha)$ determined by the outcome equation and the parametric distribution of $V$, see (ref).
The parametric nature of the distribution of $V$ allows us to drop Assumption (ref), with Assumption (ref) imposing a measure continuity condition only on the conditional distributions of $Y$ given $X = x, \alpha = a$ for $\beta \in \mathrm{B}$.
Denote by $\gamma_X$ the $\mathcal{X}$-marginal of $\gamma_{X,\alpha}$. Under (ref) and Assumption (ref), for fixed $\beta$, the conditional density $f_{Y|X,\alpha}$ and the joint measure $\gamma_{X, \alpha}$ induce the following linear integral operator:
where $S\subseteq \mathcal Y \times \mathcal X$ is a measurable set and $\mathbf{1}_S(y, x)$ is the indicator function for $S$. That is, for any joint probability measure $\gamma_{X, \alpha}\in \mathcal{P}(\mathcal{X} \times\mathcal{A})$, $\mathcal{L}_\beta[\gamma_{X,\alpha}]$ is the probability measure of $Z$ under $\gamma_{X, \alpha}$. In particular,
That is, when a latent component follows a parametric distribution, the linear map $\mathcal T_\theta$ in (ref) is the integral operator $\mathcal{L}_\beta$. The set of all model probabilities compatible with $\beta$ can then be defined as \[ \mathcal M_\beta = \left\{ \mathcal{L}_\beta[\gamma_{X,\alpha}]: \gamma_{X,\alpha} \in \mathcal P(\mathcal X \times \mathcal A) \right\}. \]
We now turn to the identification of counterfactual parameters \(\tau\). Here, \(\theta = (\beta, \tau)\) denotes the joint parameter of interest. The counterfactual parameters $\tau$ are defined as solutions to the integral equation:
where \(g(x, \alpha, \beta, \tau): \mathcal{X} \times \mathcal{A} \times \Theta \to \mathbb{R}^{d_g}\) is a known, vector-valued function. Typically, $\tau$ is a function from $\mathcal{X}$ to $\R^{d_g}$ and $g(x,\alpha, \beta, \tau) = g_0(x,\alpha, \beta) - \tau(x)$ for some $g_0: \mathcal{X} \times \mathcal{A} \times \mathrm{B} \ra \R^{d_g}$, the linear restriction in (ref) becomes
Various specifications of $g_0$ in (ref) allow a researcher to restrict the conditional moments of $\gamma_{X,\alpha}$ and determine if specific conditional moments, parameterized by $\tau$, belong to the identified set.
Using the operator \(\mathcal{L}_\beta\) defined in (ref), the set of model probabilities is then:
In order to identify $(\beta, \tau)$, we impose a light regularity constraint on $g$. In the assumption that follows, $\R^0$ is identified with $\{0\}$, and a function that maps to $\R^0$ is the zero-map.
We extend the LP approach developed in Section (ref) to the present setting in Appendix (ref). Here, we note that in the parametric error case, the optimization problem is reformulated using a linear integral operator, reflecting the parametric structure imposed on \( V \). Specifically, the model probabilities are expressed as an integral transform of the joint measure of the unrestricted components of the latent variables through a conditional density function. This representation mirrors the structure observed in the pushforward map from Section (ref) but adapts to the parametric setting.
Crucially, the dimensionality of the resulting LP depends only on the support of \( X \) and \( \alpha \), and not on that of \( V \). This invariance ensures that the computational complexity remains manageable even as the dimension of the latent space increases.
We study the two-period binary choice model with outcome equation
which has a fixed effect $\alpha \in \mathcal A \subseteq \mathbb R$, error terms $V_t \in \mathcal{V}_t \subseteq \mathbb R$, and time-varying regressors $X_1 \in \mathcal X_1 = \{x_{11},\cdots,x_{1K_1}\}$ and $X_2 \in \mathcal X_2 = \{x_{21},\cdots,x_{2K_2}\}$.
We consider the particular specification in (ref) because of its canonical status within the nonlinear panel literature. Our method applies to a much larger class of models, and does not require the index structure or additive separability.
In Sections (ref) and (ref), we do not impose parametric assumptions on the distribution of the error terms. In Section (ref), we provide new results for partial effects under strict exogeneity, and discuss the extension to correlated random coefficients.
In Section (ref), we provide new results for both structural parameters and partial effects under sequential exogeneity (predetermined covariates). To the best of our knowledge, these are the first such results for the binary choice panel model without parametric restrictions on the distribution of the error terms. Section (ref) presents the results from a numerical experiment and visualizes the identified sets.
In Section (ref), we use our results in Section (ref) to analyze parametric binary choice models with fixed effects (logit, probit, and other distributions). We recover the results in the numerical experiment of ChernValHahnNewey2013 and explore some variations of their design. The proofs of all results can be found in Appendix (ref).
We are interested in the average structural functions (ASF):
This is the period $t$ choice probability obtained by exogenously setting regressors $X_t$ to a fixed value $x$. We characterize the identified set of $\theta = (\beta,\tau_1,\tau_2)$ under the following assumption: \footnote{Section (ref) provides results under a weaker, sequential exogeneity condition.}
This is the standard strict stationarity or exogeneity assumption for nonlinear panel models, see Manski1987 and ChernValHahnNewey2013.
To map this model to the framework of Section (ref), define $Y = (Y_1,Y_2)$ and $X = (X_1,X_2)$ so that $Z = (Y,X)$. The input variables are $W = (\alpha, V_1, V_2, X_1, X_2) \in \mathcal W$. The set of probability distributions of $W$ consists of all $\gamma\in\Gamma_\theta(\mathcal W)$ that satisfy Assumption (ref). Finally, the mapping \(\psi_\theta\) is given by
Using this mapping and Corollary (ref), the following result establishes sharp identification of partial effects.
Computation of $\Thetam$ is straightforward via the LP approach that we developed in Section (ref). In Section (ref), we use this approach to visualize $\Thetam$.
Assumption (ref) does not allow for correlation between current covariates and past shocks. A sequential exogeneity assumption that allows for such feedback is:
This is Assumption 3 in ChernValHahnNewey2013, which is less restrictive than the sequential exogeneity assumption in parametric models.\footnote{It does not specify the marginal distribution of the error terms, nor does it require serial independence.} Under Assumption (ref), $\Gamma_\theta(\mathcal W)$ is not convex.\footnote{The restriction that $V_2$ and $X_2$ are independent conditional on $(A,X_1)$ is nonlinear.} Nonetheless, the following result establishes convexity of the set of model probabilities $\mathcal M_\theta$ and uses Theorem (ref) to characterize the identified set.
$\Thetam$ may be computed using Theorem (ref), the pushforward representation of Section (ref), and the extremal point representation of Proposition (ref). Setting $\psi_\theta$ as in (ref) and $W$ as in Section (ref), Theorem (ref) implies that $\Thetam$ is the set of $\theta$ satisfying
where $\Gamma_\theta(\mathcal{W})$ is the set of all distributions satisfying (ref). Let $\Gamma^\mathrm{seq}$ be the set of all distributions of $(V_1, V_2, X_2)$ which satisfy $V_2|X_2 \overset{d}{=} V_1$. Then, because both sides of (ref) are conditional on $(\alpha, X_1)$, the extremal points of $\Gamma_\theta(\mathcal{W})$ all take the form $\gamma \times \delta_{(a, x_1)}$, $\gamma \in \Gamma^\mathrm{seq}$. As in the proof of Proposition (ref), we may rewrite the right hand side of (ref) as the supremum over $a, x_1$ of
where $\gamma_{X_2}$ is a marginal distribution of $X_2$ and $\Gamma^\mathrm{seq}(\gamma_{X_2})$ is the subset of $\Gamma^\mathrm{seq}$ having $X_2$-marginal $\gamma_{X_2}$. The inner supremum in (ref) is an LP and the outer supremum is only over the space of distributions of $X_2$ and points $a, x_1$. Section (ref) below applies this extremal point characterization to obtain $\Thetam$ for various $\theta_0$.
We conduct a numerical experiment for two different DGPs. Using DGP1, we explore the size of the identified sets of the regression coefficient and the partial effects under strict exogeneity. Using DGP2, we explore the relative widths of the identified set of the regression coefficient under strict and sequential exogeneity.
DGP1. This DGP includes a time dummy $X_{1t} = t-1$ and a regressor $X_{2t}$ whose support we vary across designs as follows. In design $p$, $X_{21} = 0$ and $X_{22} \in \{-p,-p+1,\cdots,p\}$, so that higher values of $p$ imply more variation in the second regressor. All designs have $X$ discrete uniform, and $(\alpha,V_1,V_2)$ independent of $X$, with $P((V_1,V_2,\alpha) = (u_1,u_2,a) \propto \exp(-u_1^2/2 - u_2^2 /2 - a^2/2)$, with support for $V_1,V_2$ as $\{-3,-2.9,\cdots,2.9,3\}$ and the support of $\alpha$ as $\{-2,-1,\cdots,2\}$. The identified sets are computed via linear programming, see (ref) for details on the implementation.
Figure (ref) presents results for $\beta_0 = (1,-0.5)$ and designs $p \in \{4,5\}$. We plot the discrepancy function $T(\theta)$ against candidate values $\beta_2$. The identified set is smaller for Design 5 (green line) as expected, because of increased variation in $X$. Figure (ref) presents the identified set for $\beta_2$ as we vary the true value of the regression coefficient. As expected, additional variation in the regressors tightens the bounds on the regression coefficient, see Design 10 (purple line).\footnote{Manski1987 shows that point identification obtains under continuous variation in (one of the) regressors over the entire real line.} Figure (ref) presents results for the (joint) identified set for the regression coefficient and the ASF, at $\beta_2 = 0.5$. The ASF is for time period $t=1$, counterfactual value $x = 1$, and a subpopulation with $X_{21}=0,X_{22}=1$. The identified set for the ASF is the height of the box; the identified set for the regression coefficient is its width (coinciding with that in Figure (ref)). The ASF is not point identified because of the time dummy (see BotosaruMuris2024). The identified sets for the ASF are informative across all designs. The information about the regression coefficient does not translate to information about counterfactual parameters.
DGP2. We consider a worst case of the preceding design, with $X_{2t}$ independently and uniformly distributed on $\{0,1\}$. The time effect is $\beta_1 = 2$. The distribution of $\alpha + V_1$ is uniformly distributed on 11 equidistant points in the interval $[-3,3]$, and we set $V_1 = V_2$ (so strict exogeneity holds). We compute the identified set for $\beta_2$ under strict exogeneity as before, and under sequential exogeneity by applying (ref) as described in Section (ref).\footnote{A genetic optimizer from the deap Python package handles the outer optimization over $\phi \in \Phi_b$.}
Figure (ref) depicts identified sets for $\beta_2$ under the assumptions of sequential and strict exogeneity as $\beta_2$ is varied. The identified sets are larger than the ones under DGP1. This is as expected, as DGP2 has little variation in $X_{2t}$. The results clearly show that the identified set is larger under sequential exogeneity.
Let $H(\cdot)$ denote an arbitrary cumulative distribution function. In this section, we assume that the conditional distribution of the error terms is known, and given by \[ P(V_t \leq v | \alpha = a, X = x) = H(v). \] Furthermore, we assume that the $V_t$ are independent across $t$ conditional on $(\alpha,X)$. This obtains the static binary choice model with fixed effects,
with $\alpha \in \mathcal A = \mathbb R$, and the sequence of regressors $X = (X_{1},\cdots,X_{T}) \in \mathcal X$. If $H(\cdot) = \Lambda(\cdot)$ is the standard logistic distribution, this is the static panel logit model raschStudiesMathematicalPsychology1960, rasch1961general. Setting $H(\cdot)=\Phi(\cdot)$ obtains the probit version, etc.
Conditional serial independence of $V_t$ implies that $(Y_{1},\cdots,Y_{T})$ are independent conditional on $(\alpha,X)$. For any $y = (y_1,\cdots,y_T) \in \{0,1\}^T$, we therefore have
so that the model fits the framework of Section (ref). Assumptions (ref)(i) and (ii) are trivially satisfied for any $H(\cdot)$. It follows that equation (ref) in Proposition (ref) characterizes the identified set for $\beta$.
We now describe results from a numerical experiment, starting from the special case of $T=2$, $X_1 = 0,\; X_2 = 1$.\footnote{A full description of the computation is in Appendix (ref), with details about this specific model in (ref).}
Figure (ref) plots $T(\theta)$ for the probit model with $\beta_0 = 1$, with $\alpha$ on a grid of $101$ equally spaced points from -5 to 5, and with $P(\alpha = a) \propto \exp(-a^2/2)$. We consider a grid for $\beta$ of 601 equally spaced points from 0.7 to 1.3. On a single core i7-11370H at 3.30GHz, it takes 0.0036 seconds to compute $T(\theta)$ per value of $\theta$. For more details on computation times, see Section (ref). For the probit model, $\beta_0$ is not point-identified, and we find $\Theta_I = [0.968,1.065]$.
Figure (ref) presents $\Theta_I$ for five different choices of $H(\cdot)$. “Logistic” corresponds to the standard logit model, for which we recover the well-known result that $\beta_0$ is point-identified. “Normal” refers to the result in Figure (ref). “Uniform” sets $H(\cdot)$ equal to the cumulative distribution function of the continuous uniform distribution on $[-2,2]$. “Cauchy” sets it equal to the standard Cauchy distribution, and “Truncated normal” sets it equal to a standard normal distribution truncated to $[-2,3]$.
The width of these identified sets is not directly comparable across distributions, due to their different scales. Nonetheless, the identified set is rather small across all distributions, especially given the limited variation in $X$. This is in line with previous findings on the size of identified sets in parametric nonlinear panels, see honoreBoundsParametersPanel2006,ChernValHahnNewey2013.
Our second result is for the regression coefficient and average treatment effects for the probit model with $T \in \{2,3\}$ and binary regressors $\mathcal X = \{0,1\}^T$. Similarly to the numerical experiments in ChernValHahnNewey2013, \S 8, we set $P(X = x) = 2^{-T}$ for all $x$, and the support and distribution of the fixed effects, independent of $X$. Finally, $\beta_0 \in \left\{0, 0.1, \cdots, 1.9, 2\right\}$ and $\beta \in \{\beta_0 - 0.3, \beta_0 - 0.299, \cdots, \beta_0 + 0.3\}$. Appendix (ref) describes how to modify the computation from the baseline case above to this more general setup.
Figure (ref) presents the identified sets as a function of $\beta_0$, for $T = 2$ and $T=3$. \footnote{These results recover those in the top left panel in Figures 2 and 3 in ChernValHahnNewey2013.} Unless $\beta_0 = 0$, the regression coefficient is partially identified. The identified set is small for $T=2$, much smaller for $T=3$, and indistinguishable from the 45-degree line when $T=4$ (not reported). The figures are based on 12621 values of $(\beta, \beta_0)$. The computation time per value is 0.013 seconds for $T=2$, and 0.066 seconds for $T=3$ on a single core i7-11370H at 3.30GHz.
Next, we compute the identified set for the average treatment effect of moving a randomly selected individual's $x_t$ from $0$ to $1$, i.e. \[ \text{ATE}(0, 1;\beta) = E[H(\alpha + \beta) - H(\alpha)]. \] Proposition (ref) applies, and (ref) gives an expression for the identified set. Figure (ref) presents the identified sets as a function of $\beta_0$, for $T = 2$ and $T=3$. \footnote{These results recover those in the bottom right panel in Figures 2 and 3 in ChernValHahnNewey2013.} The size of the identified sets is greatly reduced when moving from $T=2$ to $T=3$.
\setcounter{page}{1}
\setcounter{section}{0} {S\arabic{section}} \setcounter{subsection}{0} {S\arabic{section}.\arabic{subsection}} \setcounter{table}{0} {S\arabic{table}} \setcounter{figure}{0} {S\arabic{figure}} \setcounter{equation}{0} {S\arabic{equation}}