EconBase
← Back to paper

Random Set Quantile Estimation of Partially Identified Discrete Response Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

230,531 characters · 24 sections · 87 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Random Set Quantile Estimation of Partially Identified Discrete Response Models

\spacing{1}

abstractSemiparametric discrete choice models are widely applied in economics, yet a fundamental tension arises when covariates are discrete as regression coefficients that are point identified under continuous regressors may become only partially identified. We show that this is not merely an identification problem but creates serious estimation pathologies. Classical estimators, including the maximum score estimator of manski1975, not only have population maximizers that are {\em outer regions} of the identified set (komarova2013) but also converge to a random set drawn from a finite collection of deterministic regions that partition that outer region. To resolve this failure, we introduce the Random Set Quantile (RSQ) estimator which extracts the ${\Greekmath 011C}$-quantile of the classical estimator for ${\Greekmath 011C} \in (1/2,1)$. We prove this result for a class of widely used models, which includes binary/multinomial choice and discrete outcome panel data models. This construction is consistent and locally robust across the full parameter space, including precisely those configurations where classical estimators break down. A feasible implementation based on the $m$-out-of-$n$ bootstrap inherits both properties. We apply the methodology to the 2019 UK General Election, where the discrete support of Brexit-related covariates generates the partial identification our theory analyzes. \begin{description} • Maximum score estimation, Partial identification, Identified set, Robustness, Random Set Quantile, Panel data discrete choice, Multinomial choice. • C14, C31, C25. \end{description}

\spacing{1.2}

Introduction

Early work on modeling discrete choices of economic agents with random utility over those choices led to the creation of an important class of semiparametric econometric models. Here a discrete outcome variable is a parametric function of observed regressors, but also depends on the random noise whose distribution is unknown and is not specified parametrically. Pioneering this line of work, manski1975, manski1985, manski2 became foundational papers, elucidating conditions under which the parameters of such models attain point identification. Specifically, in a binary choice model, by stipulating the conditional quantile of the random noise to be zero, coupled with, most notably, a continuity assumption regarding the distribution of at least one regressor, these papers established the precise conditions sufficient for parameter point identification.

In practice, however, discrete-only covariates are common, particularly in a range of policy-relevant applications. Income categories, education levels, geographic indicators, treatment assignments are all discrete by nature. When all regressors are discrete, the continuity condition is violated which may lead to the failure of point identification, the case analyzed in komarova2013. In this paper we take discrete regressors seriously as an important modeling feature rather than an inconvenience.

We demonstrate that depending on the true parameter value of the data generating process, the identified set may be either a singleton or a set of strictly positive dimension. We study not only identification but also what happens to estimation when the identified set changes its structure across the parameter space. Our analysis delivers three interconnected contributions.

Our first contribution is formalizing the concepts of both an exhaustive characterization and a coarse characterization of a semiparametric model. An exhaustive characterization fully exploits all observable implications of the model and maps every observable functional of the conditional distribution of the outcome to a constraint on the underlying parameter of interest. A coarse characterization uses a strict subset of these implications. We show that three leading classical estimators, which are the maximum score estimator of manski1975, manski1985 for cross-sectional binary choice, the conditional maximum score estimator of manski1987 for panel binary choice, and the multinomial maximum score estimator of manski1975, each induce a specific coarse structure. Using these coarse structures is essentially equivalent to using the respective exhaustive ones when at least one regressor is continuous, but this changes when all regressors are discrete.

Our second contribution is to show, to our knowledge for the first time in full generality, that a broad class of classical estimators that induce coarse structures (including those above) behave fundamentally differently with discrete regressors than in settings with continuous covariates and point identification. We establish four general theorems that characterize the weak limits of such estimators when coarse and exhaustive characterizations differ with positive probability. These results show that the estimators may be inconsistent: instead of converging in probability to the identified set, they may converge weakly to a random set that is a draw from a finite collection of deterministic regions partitioning a strict superset of the identified set, with the identified set being on the boundary of each region and its closure being equal to the intersection of any strict majority of the corresponding closed regions. These results apply to all three leading examples and, more broadly, to any model within our general framework.

In our leading examples, configurations in which inconsistencies of classical estimators occur correspond to economically meaningful situations, such as indifference or threshold behavior. As a result, the failure of standard estimators is not a theoretical curiosity but a feature that can be expected in a wide range of empirical applications. Thus, what appears as a knife-edge phenomenon in models with a continuous regressor, becomes a first-order feature of the data once covariates are discrete, and therefore should not be ignored in practice.

Our third contribution builds on the structure described for classical estimators in four general theorems. We propose a new class of estimators based on the quantile of a random set molchanov2006book. Our Random Set Quantile (RSQ) estimator maps the random set output of a classical estimator into its ${\Greekmath 011C}$-quantile, ${\Greekmath 011C} \in (1/2,1)$, taken over its closure. This is the set of parameter values that belong to the estimator’s output with probability at least ${\Greekmath 011C}$. Leveraging the majority intersection property, and the fact that all deterministic regions receive equal asymptotic probability in our leading examples, the RSQ estimator recovers the closure of the identified set exactly in the limit, for any ${\Greekmath 011C} \in (1/2,1)$.

The RSQ estimator satisfies both criteria we use for evaluation. It is consistent at every parameter value in the data generating process, including those where the classical estimators are inconsistent. It is also locally robust in the sense that its limiting distribution varies continuously under $\sqrt{n}$-local perturbations of the true parameter, in all directions that resolve indifference relations driving the coarseness. Additionally, we show that maximum score estimators are also locally robust in those same directions, hence the two estimators differ on exactly the consistency criterion. The RSQ estimator dominates. Consistency is a natural and desirable property for any set estimator. Robustness is equally important in our setting as a lack of robustness would imply that small perturbations in the data generating process associated e.g. with minor changes in the distribution of discrete covariates can lead to large and discontinuous changes in inference. Both criteria are introduced and discussed in Section (ref).

A feasible version of the RSQ estimator applied to maximum score estimator uses an $m$-out-of-$n$ bootstrap with $m = o(n)$. This inherits the theoretical properties of the infeasible version.

We illustrate our methodology using the 2019 UK General Election, where Labour secured only 202 of 650 constituencies, a historic low. The central question is whether the Leave vote in the 2016 Brexit referendum reduced Labour's probability of retaining a constituency. The discrete support of these covariates in this application generates exactly the partial identification structure our theory analyzes. The RSQ estimate delivers a set that is a strict subset of the maximum score region, has a strictly smaller affine dimension and yields an economically important conclusion about the differential effect of the Leave vote on Labour-held constituencies that is missed by the maximum score estimate.

Our paper draws on rich prior literature on identification and inferences for point and partially identified semiparametric discrete choice models. This includes classic work on the maximum score estimator in manski1975, manski1985, manski2 with the distribution theory developed in kimpollard and the smoothed version of the estimator horowitz-sms yielding asymptotically normally distributed parameters. This model was further studied in settings with heteroskedastic errors in khan:13, and in settings with discrete regressors\footnote{Important other work in the discrete regressor setting includes torgovitskyqe, where the parameter of interest was not the regression coefficients as is the case here and other referenced papers, but the the average treatment effect, for which he derived new methods to characterize the identified set. In other related work to the paper considered here, rosenura2022 considered finite sample properties of a related estimator that is based on moment inequalities. } in komarova2013.

The paper is also related to the partial identification literature, particularly the one using random set theory which was introduced to econometrics by beresteanumolinari2008 with the focus on using the Aumann expectation. Our approach is distinct in that we exploit the quantile of a random set rather than its expectation.

The paper also connects to the literature on nonregular inference, particularly andrews-boundary on boundary parameters and andrews-guggenberger-ema on hybrid inference in moment inequality models. An important example which they include is the set of moment inequality models, where discontinuity arises when inequalities are binding.\footnote{Recent important work on inference in these models, particularly for subvector inference can be found in bugnicanayshi and kaidomolinaristoye.} Our robustness analysis mirrors their local-alternatives framework but applies it to set-valued estimators. The structural parallel between the discontinuity of identified sets in moment inequality models and in our discrete-regressor setting, together with the important differences in the underlying mechanism, is discussed in detail in Section (ref).

The rest of the paper is organized as follows. Section (ref) introduces the general model, the concepts of exhaustive and coarse characterizations, and the three theoretical examples. Section (ref) formalizes the discrete setup and states the four general theorems on the asymptotic behavior of coarse-structure-based estimators, with applications to the three examples. Section (ref) introduces consistency and local robustness as our two evaluation criteria for set-estimators. Section (ref) develops the RSQ estimator and establishes its consistency and robustness. Section (ref) compares the RSQ estimator to the maximum score, establishing their shared robustness and highlighting differing consistency. Section (ref) illustrates the relative performance of the RSQ and Maximum Score Estimators through a simulation study that considers both point and partially identified designs. It also gives a guide to feasible RSQ estimation in practice. Section (ref) presents the empirical application. Section (ref) concludes. All proofs are in the Appendix.

commentOur first finding is diagnostic and, to our knowledge, previously undocumented in its full generality. Namely, we show that a wide class of classical estimators, including the maximum score estimators of manski1975, manski1985, manski1987 behave fundamentally differently when regressors are discrete.\footnote{komarova2013 considers cross-sectional binary choice model unde rthe median condition of manski1985 and shows when exactly in the discrete regressors case the maximizer of the population maximum score objective function results in a superset of the identified set, and then suggests a slack variables estimation approach instead of the maximum score to deal with this issue. In contrast, we focus on the original raw maximum score estimation not just in this model but other models we consider as well.} We provide a general framework that links classical estimators to what we call coarse characterizations of the model and formally explains why they may fail in discrete designs. In a nutshell, coarse characterizations do not fully exploits all observable implications but instead rely on weaker implications. When at least one covariate is continuous, the use of coarse characterizations does not have a detrimental impact on the consistency of these maximum score estimators, but this changes when all covariates are discrete. The consequence is twofold. First, these estimators may become inconsistent and rather than converging in probability to the identified set, they converge weakly to a random set which can be described as a random draw from a finite collection of deterministic regions that partition a strict superset of the identified set. Second, this inconsistency, when it occurs, is not pathological and has precise structure. The number and geometry of these regions, and the probabilities with which they are selected, are fully characterized by our theory. Our second finding is constructive in that we propose a new class of estimators based on the concept of a quantile of a random set from the random set theory (e.g., see molchanov2006book). We call it a Random Set Quantile (RSQ estimator). It takes the random set output of any classical estimator (such as the maximum score) and then considers its quantile to extract a consistent estimator of the identified set's closure. This finding is driven by our structural characterization of the classical estimators' weak limits when they are not consistent for the identified set. The RSQ estimator successfully addresses the changing structure of the identified set over the parameter space from singleton to non-singleton sets. We compare our new estimator with existing methods in proposed in the literature, examples of which include maximum score estimators. It can be implemented in practice using a simple $m$-out-of-$n$ bootstrap procedure. We use two criteria for evaluation and comparison of estimation methodologies. First, {\em consistency} is the property of an estimator to approximate the identified set\footnote{Some strands of the partial identification literature employ the term {\em sharp} when referring to estimation procedures with this property. } with high probability for a given value of parameter of the data generating process. The rationale behind employing this criterion is clear-cut; our objective is to ascertain the truth in the probability limit. Second, {\em robustness} is the property that the distribution of an estimator varies continuously in the small neighborhood of a particular parameter value of the data generating process. This property is important as the lack of the local continuity in distribution required for robustness can also result in failure of bootstrap and other resampling-based methods for inference. This resembles the case of the parameter on the boundary of the parameter space leading to discontinuity in the distribution limit and the failure of bootstrap, e.g., characterized by andrews-boundary\footnote{Recent contributions on inference in these models is considered in andrews-guggenberger-ema, who propose a novel hybrid procedure for a class of {\em nonregular} models. An important example which they include is the set of moment inequality models, where discontinuity arises when inequalities are binding. Recent important work on inference in these models, particularly for subvector inference can be found in bugnicanayshi and kaidomolinaristoye.} We find that our new quantile random set estimator dominates existing classical maximum score estimators for semiparametric cross-sectional binary choice, semiparametric panel data binary choice and semiparametric multinomial choice models based on these criteria. The aforementioned estimators are not consistent over the whole parameter space. E.g., the maximum score estimator can converge in probability to a singleton or a set, but it can also weakly converge to a random set. In contrast, our new RSQ estimator proves to be consistent at all parameter values in the data generating process (in partciular, at those where maximum score approaches are inconsistent) establishing its superiority over all previously discussed estimation methodologies in our discrete-only setting with regard to consistency. We show that both the maximum score estimators and the RSQ estimator are locally robust with respect to directions that resolve the boundary ties driving inconsistency of the maximum score estimators, and we characterize precisely the set of directions for which robustness can and cannot hold. Since both estimators are locally robust in the same directions, their main difference is in consistency. We derive the results of the RSQ estimator properties for three leading models. In the binary choice cross-sectional model under the ${\Greekmath 010D}$-quantile restriction, the coarseness of the maximum score estimator manifests when $P(Y=1|X=x)={\Greekmath 010D}$ with a positive probability. In the static panel binary choice model of manski1987, the same mechanism operates through the sign of probability differences across time periods. In the multinomial choice model of Manski manski1975, it operates through probability ties across alternatives. In all three cases, we verify our general conditions explicitly, characterize the geometry of the limiting random set, and establish consistency and robustness of the RSQ estimator. Our general framework that links classical estimators to coarse characterizations may apply to other settings and traditional estimators as well and, hence, the RSQ estimation approach will be valuable in those settings too. For expositional purposes, we choose to focus on maximum score type of classical estimators. We illustrate the empirical relevance of these findings by studying the impact of the 2016 Brexit referendum vote on constituency-level outcomes in the 2019 UK General Election trying to explain poor performance of the Labour's party. The discrete support of these covariates generates exactly the partial identification structure our theory analyzes. The RSQ estimate delivers a set that is a strict subset of the maximum score region, has a strictly smaller affine dimension and yields an economically important conclusion about the differential effect of the Leave vote on Labour-held constituencies that is missed by the maximum score estimate. Our paper draws on rich prior literature on identification and inferences for point and partially identified semiparametric discrete choice models. This includes classic work on the maximum score estimator in manski1975 and manski1985 with the distribution theory developed in kimpollard and the smoothed version of the estimator horowitz-sms yielding asymptotically normally distributed parameters. This model was further studied in setting with heteroskedastic errors in khan:13, and in setting with discrete regressors\footnote{Important other work in the discrete regressor setting includes torgovitskyqe, where the parameter of interest was not the regression coefficients as is the case here and other referenced papers, but the the average treatment effect, for which he derived new methods to characterize the identified set. In other related work to the paper considered here, rosenura2022 considered finite sample properties of a related estimator that is based on moment inequalities. } in komarova2013. The paper is also related to the partial identification literature, particularly the one using random set theory which was introduced to eonometrics by beresteanumolinari2008 with the focus on using the Aumann expectation. Our approach is distinct in that we exploit the quantile of a randpom set rather than its expectation. The paper is also related to the literature on nonregular inference, particularly andrews:00 on boundary parameters and andrews-Guggenberger-et on hybrid inference in moment inequality models. Our robustness analysis mirrors their local-alternatives framework but applies it to set-valued estimators. More broadly, our results highlight the importance of accounting for set-valued asymptotics in semiparametric models and suggest that random set–based methods, in particular the ones based on the quantile notion, may be useful in a wide class of partially identified problems. The rest of the paper is organized as follows. Section (ref) develops a general framework that distinguishes between exhaustive (using all information) and coarse (using a subset of information) characterizations of a semiparametric model. It presents three theoretical examples (cross section binary choice under conditional quantile independence as in manski1985, panel data binary choice model as in manski87, multinomial choice model under conditional i.i.d. of unobservables as in manski75) that are used to illustrate the coarse structures induced by the maximum score estimation approaches in these models and then later used throughout that paper to analyze consequences of this and properties of our new suggested RSQ estimation approach. Section (ref) formalizes the discrete setup and characterizes estimators based on coarse structures. It states our four general theorems that deliver a structured form of the weak limits of the coarse-structure-based estimators in cases when the difference between the coarse and exhaustive characerizations is occurs with a positive probability. It illustrates the applications of the general theorems to our three theoretical examples. In Section (ref) we create a toolkit for a general analysis of estimators and introduces our criteria for their evaluation -- {\it consistency} and {\it robustness}. As will be formally defined, consistency of an estimator is its property where it can approximate the identified set with high probability. {\it Robustness} as the property of local continuity of the limiting distribution of such an estimator and, as we show later in the paper, it may be true for an estimator which is not consistent. In Section (ref) we develop our main novel estimation methodology based on the concept of a quantile of a random set to construct a new class of estimators. We show them to be both consistent and robust (with respect to direction which make the distinction between coarse and exhaustive structures negligible) in our theoretical examples. In Section (ref) we demonstrate the advantages of our quantile random set estimator by comparing it to existing maximum score methods proposed in the literature for our theoretical examples, with our comparison and establish their properties from the perspectives of our consistency and robustness criteria. Even though maximum score estimation approaches are robust with respect to the same directions as the RSQ estimators, based on their inconsistency for some parameters values in the data generating process, we conclude our random set quantile estimator is preferable. In Section (ref) we provide an empirical illustration highlighting results from all the estimation approaches discussed in the paper. Section (ref) concludes by summarizing results and discussing areas for future research, and the appendix collects all proofs of the main theorems.

General Setting

The three theoretical examples used throughout this paper (cross-sectional and panel binary choice, multinomial choice), as will be detailed by our discussion later, all share a structure that may not be immediately apparent from their individual formulations. In each case, the observable data impose restrictions on the unknown parameter ${\Greekmath 010B}$ through a mapping from observable conditional probabilities to constraints on the linear index $x'{\Greekmath 010B}$. But a classical maximum score approach in each case turns out to silently drop some of the restrictions. To formalize this insight and unify the analysis across the three models (as well as make it applicable in other semiparametric settings), we introduce a general framework that captures these features.

We consider a general semiparametric model indexed by an outcome $Y \in \mathbb{R}^p$, covariates $X \in \mathbb{R}^k$, and an unknown parameter ${\Greekmath 010B} \in \mathcal{A} \subseteq \mathbb{R}^k$, where $\mathcal{A}$ is a compact set. It is characterized by the following equivalence:

equation[equation omitted — 164 chars of source]

where ${\Greekmath 011E}: \mathcal{F} \to \mathbb{R}^{q_1}$ is a $q_1$-dimensional functional of the conditional cumulative distribution function (c.d.f.) $F(Y|X=x)$, and ${\Greekmath 0120}: \mathbb{R}^k \times \mathbb{R}^k \to \mathbb{R}^{q_2}$ is a $q_2$-dimensional function of the linear index $x'{\Greekmath 010B}$. The sets $C_m$ belong to a collection $\mathcal{C} = \{C_m\}_{m \in \mathcal{M}}$ of subsets of $\mathbb{R}^{q_1}$, and the sets $D_m$ belong to a collection $\mathcal{D} = \{D_m\}_{m \in \mathcal{M}}$ of subsets of $\mathbb{R}^{q_2}$.

We impose the following conditions on the model structure:

conditionThe collections $\mathcal{C} = \{C_m\}_{m \in \mathcal{M}}$ and $\mathcal{D} = \{D_m\}_{m \in \mathcal{M}}$ consist of sets that are either identical or disjoint. That is, for any $C_m, C_{m'} \in \mathcal{C}$, either $C_m = C_{m'}$ or $C_m \cap C_{m'} = \emptyset$, and similarly for $D_m, D_{m'} \in \mathcal{D}$.

Condition 1 formalizes the requirement that distinct configurations of ${\Greekmath 011E}(F(Y|X=x))$ or ${\Greekmath 0120}(x,{\Greekmath 010B}) $ are distinguishable and do not overlap.

The index set $\mathcal{M}$ is unrestricted in its cardinality; it may be finite, countably infinite, or uncountable. However, in our applications, $\mathcal{M}$ typically represents a finite collection of indices.

conditionFor every $C_m \in \mathcal{C}$, $m \in \mathcal{M}$, there exists a distribution of $X$ and a conditional c.d.f. $F(Y|X=x)$ such that $P({\Greekmath 011E}(F(Y|X=x)) \in C_m) > 0$. Consequently, for the corresponding $D_m \in \mathcal{D}$, it follows that $P({\Greekmath 0120}(x,{\Greekmath 010B}) \in D_m) > 0$.

This condition ensures that each set $C_m$ in the collection $\mathcal{C}$ is empirically relevant, meaning there exists at least one distribution of the covariates $X$ and a corresponding conditional distribution $F(Y|X=x)$ for which the functional ${\Greekmath 011E}(F(Y|X=x))$ has a positive probability of lying in $C_m$. This guarantees that no set $C_m$ is redundant or unattainable under the model. Similarly, through the equivalence in equation (ref), each corresponding set $D_m$ in $\mathcal{D}$ is also attainable with positive probability under the index ${\Greekmath 0120}(x,{\Greekmath 010B})$.

definitionThe characterization in equation (ref) is exhaustive, and the collections $\mathcal{C} = \{C_m\}_{m \in \mathcal{M}}$ and $\mathcal{D} = \{D_m\}_{m \in \mathcal{M}}$ are irreducible, if there does not exist another valid characterization of the model with collections $\widetilde{\mathcal{C}} = \{\widetilde{C}_{\tilde{m}}\}_{\tilde{m} \in \widetilde{\mathcal{M}}}$ and $\widetilde{\mathcal{D}} = \{\widetilde{D}_{\tilde{m}}\}_{\tilde{m} \in \widetilde{\mathcal{M}}}$ that satisfy the following: \begin{itemize} • Conditions (ref) and (ref). • Each $\widetilde{C}_{\tilde{m}} \in \widetilde{\mathcal{C}}$ (respectively, $\widetilde{D}_{\tilde{m}} \in \widetilde{\mathcal{D}}$), $\tilde{m} \in \widetilde{\mathcal{M}}$, is a subset of some $C_m \in \mathcal{C}$ (respectively, $D_m \in \mathcal{D}$), $m \in \mathcal{M}$. • For some $\tilde{m} \in \widetilde{\mathcal{M}}$ and $m \in \mathcal{M}$, (i) $\widetilde{C}_{\tilde{m}}$ is a proper subset of $C_m$, (ii) there exists a distribution of $X$ and a conditional c.d.f. $F(Y|X=x)$ such that $P({\Greekmath 011E}(F(Y|X=x)) \in C_m \setminus \widetilde{C}_{\tilde{m}}) > 0$. \end{itemize}

In essence, the characterization in equation (ref) is exhaustive if it cannot be refined further by constructing strictly finer collections $\widetilde{\mathcal{C}}$ and $\widetilde{\mathcal{D}}$ that still satisfy the model’s conditions. Intuitively, it fully exploits all observable implications of the data generating process and is, therefore, closely tied to the concept of the identified set for the parameter ${\Greekmath 010B}$.

Notation: For convenience of our subsequent discussion, for every $x$ we let $m(x) \in \mathcal M$ denote the unique index such that ${\Greekmath 011E}(F(Y| X=x)) \in C_{m(x)}$.

definition[Identified Set] Consider an {exhaustive} characterization of the model in equation (ref). The identified set $\mathcal{A}_0$ for the parameter ${\Greekmath 010B}$ is defined as: \begin{equation} \mathcal{A}_0 = \left\{ a \in \mathbb{R}^k : {\Greekmath 0120}(x,a) \in D_{m(x)}, \; \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{for a.e. } x \right\}. \end{equation} In semiparametric models, it is standard to normalize the norm of ${\Greekmath 010B}$ (or one of its components) prior to identification analysis. We assume that such a normalization is already incorporated into the definition of $\mathcal{A}_0$ as well as the original parameter set $\mathcal{A}$.

We now formalize the notion of a coarse characterization. A coarse characterization replaces the tight constraint sets $D_m$ from the exhaustive characterization with larger sets, thereby potentially admitting additional parameter values as consistent with the observable data.

definitionSuppose the characterization in equation (ref) is exhaustive. A one-directional characterization of the model, given by \begin{equation} {\Greekmath 011E}(F(Y|X=x)) \in C_m \quad \implies \quad {\Greekmath 0120}(x,{\Greekmath 010B}) \in D^*_m, \quad m \in \mathcal{M}, \end{equation} is called coarse if the collection $\mathcal{D}^* = \{D^*_m\}_{m \in \mathcal{M}}$ satisfies the following: \begin{itemize} • Condition (ref). • For each $m \in \mathcal{M}$, the set $D^*_m$ is a superset of $D_m$, i.e., $D_m \subseteq D^*_m$. • For some $m \in \mathcal{M}$ such that $D^*_m$ is a proper superset of $D_m$ ($D^*_m \supset D_m$), there exists a distribution of $X$ and a value of ${\Greekmath 010B}$ such that $P({\Greekmath 0120}(x,{\Greekmath 010B}) \in D^*_m \setminus D_m) > 0$. \end{itemize}

This coarse characterization is one-directional, mapping knowledge of the observable functional ${\Greekmath 011E}(F(Y|X=x))$ to information about the parameter ${\Greekmath 010B}$ via the sets $D^*_m$. Unlike the exhaustive characterization, the collection $\mathcal{D}^* = \{D^*_m\}_{m \in \mathcal{M}}$ is not required to satisfy Condition (ref), allowing the sets $D^*_m$ to overlap in non-trivial ways. Furthermore, for some $m \in \mathcal{M}$, the sets $D^*_m$ are strictly larger than their counterparts $D_m$ in the exhaustive characterization, with positive probability under certain data generating processes.

This relaxation implies that identification analysis based on a coarse characterization may not yield the true identified set $\mathcal{A}_0$. Instead, it often produces a strict superset of $\mathcal{A}_0$ for some DGPs, as will be demonstrated in subsequent examples. For now, we define the set of all parameter vectors consistent with a given coarse characterization as:

equation*[equation* omitted — 192 chars of source]

with the normalization already incorporated. Specific examples of such sets and their relations to $\mathcal{A}_0$ will be provided later.

The interest in coarse characterizations stems from their prevalence in the literature. As argued in this paper, many significant estimators are built on coarse structures rather than exhaustive ones. This is primarily because these estimation approaches were designed with continuous covariates (or at least one continuous covariate) in mind, where the index $x'a$ falls into $D^*_m \setminus D_m$ with only zero probability. However, when discrete covariates are allowed, $D^*_m \setminus D_m$ may contain index values with positive probability, leading to significant consequences for the properties of these estimation approaches, as demonstrated in this paper.

\paragraph*{Objective functions and estimators} Estimators based on coarse characterizations may take various forms, typically being sample versions of population criteria that reward compliance with the coarse restriction ${\Greekmath 0120}(x,a)\in D^*_{m(x)}$. One example of such objective functions would be based on score function $s(\cdot, \cdot)$ that depends on ${\Greekmath 011E}(F(Y|X=x))$, drawn from the suitable space of conditional probability distributions, and on $t \in \mathbb{R}^q$ such that

align[align omitted — 468 chars of source]

With such a score function, we can consider, for instance, the population objective function

equation[equation omitted — 138 chars of source]

with weights $ w(x)\ge 0$, ${\mathbb P}(w(x)>0)=1$. Under correct specification, (ref)-(ref) imply that for each $x$ the pointwise maximizers are those $a$ such that ${\Greekmath 0120}(x,a)\in D_{m(x)}^*$, and therefore $\arg\max_a Q(a)=\mathcal A^*$.

Given data $\{(y_i,x_i)\}_{i=1}^n$ and a plug-in estimator of ${\Greekmath 011E}(\widehat{F}(\cdot| X=x))$), the direct sample criterion is

equation[equation omitted — 191 chars of source]

where $\mathcal X_n$ is the sample support of $X$.

Two natural choices of the score function are the following.

enumerate• Indicator scores:\quad $s({\Greekmath 011E}(F(Y|X=x)), t)=s({\Greekmath 011E}(F(Y|X=x)),\mathbf 1\!\{t\in D_{m(x)}^*\})$. • Soft (distance) scores:\quad $s({\Greekmath 011E}(F(Y|X=x)),t)=s({\Greekmath 011E}(F(Y|X=x)),-{\Greekmath 011A}\!\big(\mathrm{dist}(t,D_{m(x)}^*)\big))$ with ${\Greekmath 011A}$ increasing and ${\Greekmath 011A}(0)=0$.

Generally, we can interpret the objectives (ref) as general monotone-score objectives due to conditions (ref)-(ref). These objectives include max-score, rank-based, inequality-penalty, and smoothed versions thereof.

In the exposition above, we introduced a coarse structure and then considered population objectives maximized over the corresponding coarse set $\mathcal{A}^*$. This ordering is purely expository. In practice, researchers do not choose a coarse structure first but rather use established estimators that implicitly induce one. As we show below, maximum score estimators endogenously induce coarse characterizations because they are tailored to settings with continuously distributed covariates, where the exhaustive and coarse structures differ only on sets of measure zero. When applied to environments with discrete regressors and partial identification, however, this induced coarseness becomes consequential and can lead to different identification and estimation properties.

We next consider our three leading examples and show that some well-known estimators in econometrics can be interpreted within this exhaustive/coarse framework.

Theoretical Example 1: Binary choice cross-sectional model manski1975,manski1985,manski-jasa

Consider the model

align[align omitted — 300 chars of source]

where $\mathcal{Q}(\cdot|\cdot)$ denotes the conditional quantile operator, ${\Greekmath 010D}\in(0,1)$, $X$ is a $k$-dimensional random vector and ${\Greekmath 010B}$ is a $k$-dimensional unknown parameter vector of interest. This model is widely used in applications such as program participation or treatment choice, where the outcome reflects whether an individual crosses a latent utility threshold.

Exhaustive characterization. To show that this model fits our setting of exhaustive characterization, denote $${\Greekmath 011E}(F(Y|X=x)) = P(Y=1|X=x)-{\Greekmath 010D}, \quad {\Greekmath 0120}(x,{\Greekmath 010B})=x'{\Greekmath 010B},$$ and let the collections $\mathcal{C}$ and $\mathcal{D}$ in Definition (ref) consist of the same three sets:

align[align omitted — 194 chars of source]

With these definitions of $\mathcal{C}$ and $\mathcal{D}$ the model can be equivalently written as in ((ref)).

Coarse characterization induced by the maximum score estimator. We focus on the coarse characterization induced by the maximum score estimator commonly used in this setting. Given a sample $\{(y_i,x_i\}_{i=1}^n$, the maximum score estimator maximizes $ \sum_{i=1}^n (y_i-{\Greekmath 010D}) \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{sgn}(x_i'a)$ over $a$ in the parameter space, which can be rewritten as the maximization over $a$ of $$\sum_{x \in \mathcal{X}_n} (\widehat{P}(Y=1|x)-{\Greekmath 010D}) \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{sgn}(x'a) \widehat{P}(X=x), $$ with $\mathcal{X}_n$ being the sample support , $\widehat{P}(Y=1|x)=\frac{1/n\sum_{i=1}^n y_i 1(x_i=x)}{1/n\sum_{i=1}^n 1(x_i=x)}$, $\widehat{P}(X=x)=1/n\sum_{i=1}^n 1(x_i=x)$, and $\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{sgn}(z)=1(z\geq 0)-1(z<0))$. The population version of this objective function is $$E \left[({P}(Y=1|X)-{\Greekmath 010D}) \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{sgn}(X'a \geq 0)\right].$$ The set of maximizers of this population objective is

equation[equation omitted — 363 chars of source]

This immediately implies that the coarse characterization induced by the MS estimation uses the following collection $\mathcal{D}^*$:

equation[equation omitted — 117 chars of source]

Thus, ((ref)) describes the coarse set $\mathcal{A}^*$ given by this induced coarse characterization.

Note that both population and sample MS objective functions here can be written in the form of (ref) and (ref), respectively, where

equation[equation omitted — 170 chars of source]

From the perspective of the MS estimator performance, as shown in komarova2013, the problematic aspect of the coarse characterization is the implication ${\Greekmath 011E}(F(Y|X=x)) \in C_3 \implies {\Greekmath 0120}(x'a) \in D^*_3 = \mathbb{R}$. When $P(Y=1|X=x) = {\Greekmath 010D}$, the corresponding $x$ does not constrain the estimation of $a$, as $x'a$ can take any value in $\mathbb{R}$. For continuous covariates, the event $P(Y=1|X=x) = {\Greekmath 010D}$ occurs with probability zero, but for discrete covariates, this event may have positive probability, impacting the estimator's performance. The difference between $D_1$ (an open half-line) and $D_1^*$ (a closed half-line) affects only the boundary of the maximizer set and does not change its affine dimension. Hence, this coarsening is relatively innocuous. In contrast, the discrepancy between $D_3$ (a singleton) and $D_3^*$ (the entire real line) eliminates binding restrictions on the index and is thus substantially more consequential.

Economically, observations with $m(x)=3$ are agents sitting exactly at the decision margin $x'a=0$. In a voting model, this is someone whose latent preference difference between candidates is zero (with ${\Greekmath 010D}=1/2$ these are indifferent voters). In a program participation model, this is someone whose utility from enrolling exactly meets the threshold. With continuous covariates such cases are negligible by assumption. With discrete covariates that may involve income brackets, education categories, treatment indicators, a whole mass of observations can sit at this margin. These observations are arguably the most interesting ones for policy, since they are the first to respond to any perturbation. Yet they are exactly where the maximum score estimator goes silent. As we show in Section (ref), this has real consequences for consistency.

Parameter vector normalization. Before we proceed with a simple example, let us touch upon a normalization issue. As discussed in multiple works by Manski (e.g. manski1985), one can only hope to identify parameters up to scale. For that reason, a common tradition in the literature is to normalize one of the components of ${\Greekmath 010B}$ to 1 (alternatively, $-1$ depending on the perceived direction of the effect). The parameter space $\mathcal{A} \subseteq \mathbb{R}^k$, the identified set $\mathcal{A}_0$ and the coarse set $\mathcal{A}^*$ have to respect such normalization thus making their affine dimensions at most $k-1$. We will keep such a normalization in mind without unnecessarily overemphasizing it.

Simple example. Let $X$ be discrete with two values, $x_1$ and $x_2$, such that $P(Y=1\mid X=x_1)>{\Greekmath 010D}$ and $P(Y=1\mid X=x_2)={\Greekmath 010D}$. The exhaustive characterization implies that the identified set $\mathcal{A}_0$ (with normalization embedded in $\mathcal{A}$) consists of all $a\in\mathcal{A}$ satisfying $x_1'a>0$, $x_2'a=0.$ By contrast, the coarse set $\mathcal{A}^*$ induced by the maximum score estimator consists of all $a\in\mathcal{A}$ such that $x_1'a\ge 0.$ If $\mathcal{A}$ has nonempty relative interior in the $(k-1)$-dimensional normalized space and is large enough, then both $\mathcal{A}$ and $\mathcal{A}^*$ have affine dimension $k-1$, whereas $\mathcal{A}_0$ has affine dimension $k-2$.\footnote{This discussion requires that $\mathcal{A}$ is not fully contained in $\{a: x_2'a=0\}$.} Thus, in such discrete designs, where some $P(Y=1| X=x)-{\Greekmath 010D}$ lies in $C_3$, the coarse set $\mathcal{A}^*$ is typically a strict superset of $\mathcal{A}_0$.

Theoretical Example 2: Static panel data manski1987

manski1987 proposed a semiparametric extension of the panel binary choice model of andersen70:

equation[equation omitted — 130 chars of source]

where $i=1,2,...n$ are the cross-sectional units and $ s=1,2$ are the time periods. The binary variable $Y_{is}$ and the $k$-dimensional regressor vector $X_{is}$ are each observed and the parameter of interest is the $k \times 1$ vector ${{\Greekmath 010B}}$. The variables not observed in the data are $c_i$, and ${\Greekmath 010F}_{is}$, the former not varying with $s$ and often referred to as the “fixed effect" or the individual specific effect. manski1987 imposed no specific distributions on unobservables.

Denote $\mathbf{X}_i=(X_{i1},X_{i2})$ and in the spirit of manski1987, suppose that

itemize$F_{{\Greekmath 010F}_{i1}|\mathbf{X}_i,c_i}(\cdot|\cdot)=F_{{\Greekmath 010F}_{i2}|\mathbf{X}_i,c_i}(\cdot|\cdot)$ for all $(c_i,\mathbf{X}_i)$ in the joint support. • The support of ${\Greekmath 010F}_{it}\,\big|\,\mathbf{X}_i,c_i$ is $\mathbb{R}$.

Note that these conditions, in particular, imply that $P(Y_{is}=1|\mathbf X_i=\mathbf x) \in (0,1)$, $s=1,2$.

Exhaustive characterization. To obtain an exhaustive characterization ((ref)) of this model under the stated assumptions, define $${\Greekmath 011E}(F(Y_{i1},Y_{i2})|\mathbf X_i=\mathbf x)) = P(Y_{i2}=1|\mathbf X_i=\mathbf x)- P(Y_{i1}=1|\mathbf X_i=\mathbf x), \quad {\Greekmath 0120}(\mathbf X,a)=X_{i2}'a-X_{i1}'a.$$ The exhaustive characterization is then obtained using these functionals and collections $\mathcal{C}$ and $\mathcal{D}$ in ((ref)) and ((ref)), respectively.

Conditional maximum score estimator and the induced coarse structure. Given a sample $\{(y_{i1},y_{i2},x_{i1},x_{i2})\}_{i=1}^n$, the classical estimation approach in this literature is the maximization over index parameter $a$ of the {\it conditional maximum score} objective function proposed by manski1987:

equation[equation omitted — 184 chars of source]

This is equivalent to the maximization over $a$ of $$\sum_{\mathbf x \in \mathcal{X}_n} (\widehat{P}(y_{2}=1|\mathbf x) -\widehat{P}(y_{1}=1| \mathbf x))\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{sgn}\left((x_{i2}-x_{i1})'a\right) \widehat{P}(\mathbf X=\mathbf x),$$ where $\mathcal{X}_n$ is the sample support of $\mathbf X$, $\mathbf x=(x_1,x_2)$, $\widehat{P}(y_s=1|\mathbf x)$ and $\widehat{P}(\mathbf X=\mathbf x)$ are the plug-in estimators for ${P}(y_s=1|\mathbf x)$ and ${P}(\mathbf X=\mathbf x)$, respectively, analogous to those in Theoretical Example 1, The corresponding population objective function takes the form $$E\left( ({P}(y_{2}=1|\mathbf x) -{P}(y_{1}=1| \mathbf x))\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{sgn}\left((x_{2}-x_{1})'a\right)\right).$$ Analogously to the cross-sectional binary choice in Theoretical Example 1, we can conclude that the optimization of such an objective induces the coarse structure with the same collection $\{D_m^*\}$ of coarse sets as in Theoretical Example 1. Both the population and the sample objective function in this case can be rewritten in the form (ref) and (ref) with $$s({\Greekmath 011E}(F(Y=1|X=x)),u)=|{P}(y_{2}=1|x) -{P}(y_{1}=1|x))|\cdot \mathbf{1}(u \in D_{m(x)}^* ), \qquad w(x)=1.$$

Analogous to Theoretical Example 1, the coarseness has a particularly drastic effect when we have the case of discretely distributed $\mathbf X$ with some points in this support having $m(\mathbf x)=3$ (that is, satisfying ${\Greekmath 011E}(F(y|\mathbf x)) \in C_3$) as then $D_{3}=\{0\}$ is replaced with $D_3^*=\mathbb{R}$. Individuals with such $\mathbf x$ are those whose behavior, on average, did not change between periods. These stable individuals are often the primary target of the policy being studied. A subsidy meant to induce take-up, a change in eligibility rules, a tax incentive may be designed to move people who were not moving before. When discrete covariates create a positive mass of such observations, ignoring in coarse structure the restriction they impose removes from the estimation precisely the variation that the policy question is about.

Just as in Theoretical Example 1, we normalize one of coefficients in ${\Greekmath 010B}$ and assume that the parameter space $\mathcal{A}$ has a relative interior in the $(k-1)$-dimensional normalized space and is large enough to ensure that $\mathcal{A}^*$ also has a relative interior in that space. Case $m(x)=3$ happening with a positive probability then guarantees that the identified set $\mathcal{A}_0$ built on the exhaustive characterization has a strictly smaller affine dimension that the coarse set $\mathcal{A}^*$.

Simple example. Let $\mathbf X$ be discrete with two values, $\widetilde{\mathbf x}$ and $\mathbf x^{\diamond}$, such that $P(Y_2=1|\mathbf X=\widetilde{\mathbf x})=P(Y_1=1| \mathbf X=\widetilde{\mathbf x})$ and $P(Y_2=1|\mathbf X=\mathbf x^{\diamond})>P(Y_1=1|\mathbf X=\mathbf x^{\diamond})$. In other words, ${\Greekmath 011E}(F(Y_{1},Y_{2}|\mathbf X=\widetilde{\mathbf x}) \in C_3$, ${\Greekmath 011E}(F(Y_{1},Y_{2}|\mathbf X=\mathbf x^{\diamond}) \in C_1.$

The exhaustive characterization implies that the identified set $\mathcal{A}_0$ consists of all $a \in \mathcal{A}$ (subject to normalization) satisfying ${\Greekmath 0120}(\widetilde{ \mathbf x}',a) \in D_3$, ${\Greekmath 0120}(\mathbf x^{{\diamond}},a) \in D_1,$ or equivalently, $(\widetilde{x}_{2}-\widetilde{x}_{1})'a=0, \quad (x^{\diamond}_{2}-x^{\diamond}_{1})'a>0$ (affine dimension of $\mathcal{A}_0$ is $k-2$). In contrast, $\mathcal{A}^*$ consists of all $a$ (subject to normalization) satisfying only $( x^{\diamond}_{2}- x^{\diamond}_{1})'a \geq 0$ (affine dimension of $\mathcal{A}^*$ is $k-1$). In particular, $\mathcal{A}^*$ can be a strict superset of $\mathcal{A}_0$ with a strictly positive Hausdorff distance between them.

commentthe configurations that drive coarseness are boundary cases $C_3/D_3$ in which $P(Y_{i2}=1 \mid X_i=x) - P(Y_{i1}=1 \mid X_i=x)=0$. At such support points, the exhaustive characterization imposes $(X_{i2}-X_{i1})'a=0$, whereas the coarse characterization places no restriction on the index. Economically, these cases correspond to individuals whose behavior does not systematically change across periods, reflecting a balance in latent utilities over time. These individuals may precisely be the targets of policies designed to induce behavioral change over time, such as subsidies, tax incentives, or changes in eligibility rules. When such boundary configurations occur with positive probability, the data contain mass points where no observed change occurs. At these support points, the coarse characterization discards the restriction linking changes in covariates to changes in behavior, effectively treating these individuals as uninformative. As a result, the response of marginal individuals is not properly captured, which can lead to misleading conclusions about the effectiveness of policies aimed at inducing transitions over time. In this sense, ignoring these cases obscures the very margin along which policy operates.

Theoretical Example 3: Multinomial choice manski1975

We consider the multinomial choice model studied by manski1975. For each individual $i$, a latent utility for object $j$ is given by $U^*_{ij}=x_{ij}'{\Greekmath 010B}_j+{\Greekmath 0122}_{ij},$ $j=1,\ldots,J,$ where $({\Greekmath 0122}_{i1},\ldots,{\Greekmath 0122}_{iJ})$ are i.i.d.\ conditional on $\mathbf x_i:=(x_{i1}',\ldots,x_{iJ}')'$. Agents choose the alternative with the highest latent utility: $Y_i=j \quad\iff\quad U^*_{ij}>\max_{k\neq j}U^*_{ik}.$ Let $Y$ denote the choice variable, which takes values in $\{1,\ldots,J\}$.

In the spirit on manski1975 assume that the support of each ${\Greekmath 0122}_j|\mathbf x$. $j=1, \ldots, J$, is $\mathbb{R}$ and the c.d.f. is continuous. Under assumed conditions, events $U^*_{ij}=U^*_{ik}$ occur with probability zero and $P(Y=j|\mathbf{x}) \in (0,1)$ for all $j=1,\ldots, J$, and a.e. $\mathbf x$.

To obtain an exhaustive characterization, define the functionals ${\Greekmath 011E}(F(Y| \mathbf X=\mathbf x)) = \bigl(P(Y=1|\mathbf x),\ldots,P(Y=J|\mathbf x)\bigr)'$, $ {\Greekmath 0120}(\mathbf x,a) = \bigl(x_1'a_1,\ldots,x_J'a_J\bigr)'$, and for each pair $(j,k)$ with $j<k$, let $$\widetilde C_{jk;>}=\{c\in\mathbb R^J:\ c_j>c_k\},\quad \widetilde C_{jk;<}=\{c\in\mathbb R^J:\ c_j<c_k\},\quad \widetilde C_{jk;=}=\{c\in\mathbb R^J:\ c_j=c_k\},$$ and analogously $\widetilde D_{jk;>}$, $\widetilde D_{jk;<}$, and $\widetilde D_{jk;=}$ in $\mathbb R^J$. Let ${\Greekmath 0114}_{jk}\in\{>,<,=\}$ denote one of the three relations. Collections $\mathcal C$ and $\mathcal D$ consist of all nonempty intersections $C_m=\bigcap_{j<k}\widetilde C_{jk;{\Greekmath 0114}_{jk}}$, $ D_m=\bigcap_{j<k}\widetilde D_{jk;{\Greekmath 0114}_{jk}},$ taken over all assignments $\{{\Greekmath 0114}_{jk}\}_{j<k}$. These sets correspond to complete (possibly weak) orderings of the choice probabilities/utility indices. As follows from manski1975, the model admits the exhaustive characterization (ref) with the aforementioned ${\Greekmath 011E}$, ${\Greekmath 0120}$, $\mathcal{C}$, $\mathcal{D}$.

\vskip 0.05in

Original multinomial maximum score. The multinomial maximum score estimator proposed by manski1975 maximizes, over $a$, the function $\frac{1}{n}\sum_{i=1}^n\sum_{j=1}^J \mathbf 1(Y_i=j)\, W (\sum_{k\neq j}\mathbf 1(x_{ij}'a_j>x_{ik}'a_k)),$ where $W(\cdot)$ is a strictly increasing sequence of real numbers. That is, $W(J-1) > W(J-2) > \cdots > W(1) > W(0)$. Equivalently the sample objective function can be written as $ \sum_{x \mathcal{X_n}} \sum_{j=1}^{J} \widehat{P}(y={j}|\mathbf{x}) \, W\left( \sum_{k \neq j} \mathbf{1}\left( x_{j}' a > x_{k}' a \right) \right) \widehat{P}(\mathbf{X}=\mathbf{x})$ and its population version is

equation[equation omitted — 166 chars of source]

When conditional choice probabilities are strictly ordered, ((ref)) is maximized by aligning deterministic utilities in the same way, so the population maximizers coincide with $\mathcal A_0$.

When $P(Y=j|\mathbf x)=P(Y=k|\mathbf x)$ for some $j\neq k$, the objective breaks ties via $x_{ij}'a_j-x_{ik}'a_k$, favoring strict inequalities over equality. Thus the maximizers need not include $\mathcal A_0$ (only its closure does), unlike the binary case where ties impose no restriction.

\vskip 0.05in

A tie-indifferent (block) maximum score objective. We modify the objective to coincide with the original without ties, but to impose no ordering restrictions when ties occur, yielding a superset of $\mathcal{A}_0$. This mirrors the resolution that we saw in the binary choice cases and ensures that the coarse $\mathcal{A}^*$ contains the identfied set $\mathcal{A}_0$.

Let $p$ denote a generic probability vector from the $(J-1)$-dimensional probability simplex. Let $\mathfrak B(p)=(B_1(p),\ldots,B_{R(p)}(p))$ denote the ordered partition of $\{1,\ldots,J\}$ into probability tie blocks, ordered from highest to lowest probability. Define the score function \[ s^{\mathrm{blk}}(p,u) = \sum_{r=1}^{R(p)} \left( \sum_{j \in B_r(p)} \frac{p_j}{B_r(p)}\right) \sum_{d=1}^{|B_r(p)|}W \left( \sum_{r'\neq r} \mathbf 1\Big\{\min_{j\in B_r(p)} u_j>\max_{k\in B_{r'}(p)} u_k\Big\} +d-1 \right), \qquad u\in\mathbb R^J, \] where $W(\cdot)$ is strictly increasing sequence (as before) which is defined at points $0, \ldots, J-1$, with all those used in the estimator. The corresponding population objective function is then of the form (ref) with $w(x)=1$:

equation[equation omitted — 185 chars of source]

This objective coincides with the original multinomial maximum score when probabilities are strictly ordered, but becomes flat in directions corresponding to within-block index differences when probabilities tie.The sample analogue is obtained by replacing conditional probabilities with their plug-in estimates:

equation[equation omitted — 230 chars of source]

\vskip 0.05in

Coarse characterization induced by the tie-indifferent maximum score objective. As in the previous two theoretical examples, we now take a closer look at which coarse structure is induced by our tie-indifferent objective. For a given $\mathbf{x}$, we have its exhaustive regime $m(\mathbf{x})$ (thus, with ${\Greekmath 011E}( (F(Y|\mathbf{x}))\in C_{m(\mathbf{x})}$) and the block partition $\mathfrak B({\Greekmath 011E}( (F(Y|\mathbf{x})))$ of the alternatives. The score $s^{\mathrm{blk}}\!\bigl({\Greekmath 011E}(F(Y|\mathbf{x})),{\Greekmath 0120}(\mathbf{x},a)\bigr)$ is maximized (pointwise in $\mathbf{x}$) whenever indices respect the strict between-block ordering, while imposing no restriction within blocks. This induces the following coarse set for the regime $m(x)$:

equation[equation omitted — 231 chars of source]

where $\mathfrak B({\Greekmath 011E}( (F(Y|\mathbf{x}))))=(B_1({\Greekmath 011E}( (F(Y|\mathbf{x})))),\ldots,B_{R(\mathbf{x})}({\Greekmath 011E}( (F(Y|\mathbf{x})))))$. Equivalently, (ref) keeps all strict inequalities implied by strict probability orderings across blocks, and drops all equality restrictions implied by ties within blocks. In other words, formally, for a regime $m$ with relations ${\Greekmath 0114}_{jk}$, $$ D_m^* = \bigcap_{j<k}

cases\widetilde D_{jk;>}, & {\Greekmath 0114}_{jk} \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ is } >,\\ \widetilde D_{jk;<}, & {\Greekmath 0114}_{jk} \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ is } <,\\ \mathbb R, & {\Greekmath 0114}_{jk}\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ is } =.

\quad \supseteq \quad D_m= \bigcap_{j<k}

cases\widetilde D_{jk;>}, & {\Greekmath 0114}_{jk} \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ is } >,\\ \widetilde D_{jk;<}, & {\Greekmath 0114}_{jk} \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ is } <,\\ \widetilde D_{jk;=}, & {\Greekmath 0114}_{jk}\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ is } =.

. $$ Thus, whenever $P(Y=j|\mathbf x)=P(Y=k|\mathbf x)$, the coarse structure imposes no restrictions on $x_{ij}'a_j-x_{ik}'a_k$. Thus, the block estimator is effectively built on the coarse characterization $ {\Greekmath 0120}(\mathbf x,{\Greekmath 010B})\in D^*_{m(\mathbf x)},$ where $D^*_{m(\mathbf x)}\supseteq D_{m(\mathbf x)}$ whenever $C_{m(\mathbf x)}$ contains at least one equality and $D^*_{m(\mathbf x)}=D_{m(\mathbf x)}$ if ${\Greekmath 011E}(F(Y|\mathbf X=\mathbf x))$ is strictly ordered (no ties). The corresponding coarse set is $\mathcal A^* = \bigl\{a\in\mathcal A:\ {\Greekmath 0120}(\mathbf x,a)\in D^*_{m(\mathbf x)}\ \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{for a.e.\ }\mathbf x\bigr\},$ which always satisfies $\mathcal A_0\subseteq \mathcal A^*$ and is strict superset of $\mathcal A_0$ when a tie occurs with a positive probability (e.g. under discrete regressors). $\mathcal A^*$ maximizes (ref).

Observations with $\mathbf{x}$ that tie some options are central to understanding substitution. E.g. in a model of product demand, they determine cross-price elasticities. In occupational sorting, they are the workers who would switch given a small wage change. When discrete covariates make such ties occur with positive probability, the coarse characterization effectively removes the identifying variation that comes from observing agents at these margins.

\vskip 0.05in

Simple example To illustrate the difference between the identified set and the coarse set, take $J=3$ and suppose $\mathbf{x}$ is discrete and its support consists of two points $\widetilde{\mathbf{x}}$ and $\mathbf{x}^{\diamond}$ for which we have ${P}(Y=1|\widetilde{\mathbf{x}}) = {P}(Y=2|\widetilde{\mathbf{x}})< {P}(Y=3|\widetilde{\mathbf{x}})$, ${P}(Y=1|\mathbf{x}^{\diamond}) < {P}(Y=2|\mathbf{x}^{\diamond})= {P}(Y=3|\mathbf{x}^{\diamond}), $ In other words, ${\Greekmath 011E}(F(y|\widetilde{\mathbf{x}})) \in C_{m_1}=\widetilde{C}_{12; =} \cap \widetilde{C}_{13; < }\cap \widetilde{C}_{23; < }$, ${\Greekmath 011E}(F(y|\mathbf{x}^{\diamond})) \in C_{m_2}=\widetilde{C}_{12; <} \cap \widetilde{C}_{13; < }\cap \widetilde{C}_{23; = }.$ The exhaustive characterization implies $\widetilde{x}_1'a_1=\widetilde{x}_2'a_2<\widetilde{x}_3'a_3$, $ x^{\diamond '}_1 a_1<x^{\diamond '}_2 a_2=x^{\diamond '}_3 a_3.$

Under the block maximum score objective, alternatives $1$ and $2$ for $\widetilde{\mathbf{x}}$ and alternatives $2$ and $3$ for $\mathbf{x}^{\diamond}$ form single blocks, and the induced coarse restriction is $\max\{\widetilde{x}_1' a_1,\widetilde{x}_2'a_2\}<\widetilde{x}_3'a_3,$ $ x^{\diamond '}_1 a_1<\min\{x^{\diamond '}_2 a_2,x^{\diamond '}_3 a_3\}$ with no restrictions on the relative ordering of $\widetilde{x}_1'a_1$, $\widetilde{x}_2'a_2$ and of $x^{\diamond '}_2 a_2$, $x^{\diamond '}_3 a_3$.

Thus, as long as the parameter space $\mathcal{A}$ has a nonempty relative interior in the normalized space (suitable normalizations are applied to $a_1$, $a_2$, $a_3$), the coarse set $\mathcal{A}^*$ therefore has strictly higher affine dimension than the exhaustive identified set $\mathcal{A}_0$.

commentThus, in the multinomial choice setting, the configurations that drive coarseness are boundary cases in which the latent utilities of two or more alternatives coincide at some support points of the covariates. Economically, these cases correspond to individuals who are indifferent between competing alternatives. In applications such as product choice, occupational sorting, or location decisions, these are agents for whom small changes in covariates can shift the ranking of options and lead to different choices. When such boundary configurations occur with positive probability, the data contain regions where observed choice frequencies do not clearly reveal a strict ordering across alternatives. At these support points, the coarse characterization discards the restrictions that determine alternatives being equally ranked, effectively weakening information about the structure of substitution. Ignoring these cases can therefore lead qualitatively incorrect predictions of substitution patterns under price intervention or tax policies shifting incentives.

Illustrative designs for theoretical examples

For the discussion that follows, it will sometimes be helpful to refer to specific cases from the theoretical examples outlined above. To support this, we present illustrative designs for the first two examples in this section, and include an additional design for the third in the Appendix.

Illustrative Design 1 (Cross-sectional binary choice under median condition) Let the vector of covariates be $(1,X_1,X_2)$ and, thus, have a fixed first component. Let $X=(1,X_1,X_2)' \in \{(1,0,1),(1,1,0)\}$. Suppose ${\mathbb P}(X=(1,0,1)')=q>\frac12$ (without loss of generality). We take ${{\Greekmath 010B}} \equiv ({\Greekmath 010B}_{1},1,{\Greekmath 010B}_{2})$. Let $\mbox{Med}(u|X)=0.$

In this design the linear index ${X}'a$ takes two possible forms: $a_1+a_2$ (for the first support point) and $a_1+1$ (for the second support point). If the parameter vector ${\Greekmath 010B}$ in the DGP makes at least one of these indices zero, then the corresponding probability ${\mathbb P}(Y=1\,|\,{X})=\frac12.$

We can express the identified set $\mathcal{A}_0$ using the exhaustive characterization in terms of observed choice probabilities $P(Y=1|X=(1,0,1)')$ and $P(Y=1|X=(1,1,0)').$ Since $P(Y=1|X=(1,0,1)') > (<)[=] \iff {\Greekmath 010B}_{1}+{\Greekmath 010B}_{2}> (<)[=] 0$. $P(Y=1|X=(1,1,0)') > (<)[=] \iff {\Greekmath 010B}_{1}+1> (<)[=] 0$, then equivalently, it can be done in terms of ${\Greekmath 010B}$ in the DGP.

Using the $\operatorname*{sign}$ function such that $\operatorname*{sign}(0)=0$,\footnote{Note that the function $\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{sgn}$ used by us above has the property $\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{sgn}(0)=1$ which aligns with its properties in manski1985,manski-jasa} the identified set is $$ {\mathcal A}_0=\left\{ (a_1,1,a_2)\,:\, \operatorname*{sign}(a_1+a_2)= \operatorname*{sign}({\Greekmath 010B}_{1}+{\Greekmath 010B}_{2}),\, \operatorname*{sign}(a_1+1)=\operatorname*{sign}({\Greekmath 010B}_{1}+1) \right\} \cap {\mathcal A}. $$ Suppose ${\mathcal A}$ has a relative interior in the normalized 2-dimensional space and is large enough to contain $(-1,1,1)$. Then we have the following four cases for the identified set ${\mathcal A}_0:$

enumerate• If ${\Greekmath 010B}=(-1,1,1)$, then ${\mathcal A}_0=\{(-1,1,1)\}$ (the only point identification case). • If ${\Greekmath 010B}_{1}=-1$, ${\Greekmath 010B}_{2}\neq 1$, then ${\mathcal A}_0= \left\{(-1, 1,a_{2}): \operatorname*{sign}(-1+a_{2})=\operatorname*{sign}(-1+{\Greekmath 010B}_{2})\}\right\} \cap {\mathcal A}.$ ${\mathcal A}_0$ is fully contained in one-dimensional hyperplane ${\Greekmath 010B}_{1}=-1$ (partial identification case; identified set has affine dimension 1). • If ${\Greekmath 010B}_{1}\neq -1$, ${\Greekmath 010B}_{1}+{\Greekmath 010B}_{2}=0$, then ${\mathcal A}_0= \left\{(a_{1}, 1,-a_{1}): \operatorname*{sign}(a_{1}+1)=\operatorname*{sign}({\Greekmath 010B}_{1}+1)\}\right\} \cap {\mathcal A}.$ ${\mathcal A}_0$ is fully contained in the one-dimensional hyperplane $a_{1}+a_{2}=0$ (partial identification case; identified set has affine dimension 1). • If ${\Greekmath 010B}_{1}\neq -1$ and ${\Greekmath 010B}_{1}+{\Greekmath 010B}_{2}\neq 0$, then only the general expression provided above applies (partial identification case; identified set has a non-empty 2-dimensional interior in the normalized parameter space and, hence, has affine dimension 2).

This design exhibits some interesting complexity where even in partially identified cases, certain conditional probabilities of choice may remain equal to $1/2$, in which case identified sets have a smaller affine dimension than the dimension of the non-normalized parameters.

Illustrative design 2 (Two-period static binary choice panel data model under identical distribution of errors)

Suppose the support of each ${X}_{it}$ is $\widetilde{\mathcal{X}}=\{(0,0)', (0,1)', (1,0)', (1,1)'\}$, $t=1,2$, and the support $\mathcal{X}$ of $(X_{i1},X_{i2})'$ will be the Cartesian product $\widetilde{\mathcal{X}} \times \widetilde{\mathcal{X}}$. Suppose the second parameter is normalized to 1, thus leaving the unknown parameter in the normalized space to take the form $({\Greekmath 010B}_1,1)$.

Suppose ${\Greekmath 010B}_1=1$ in the DGP. When $x_{i1}=(0,1)'$ and $x_{i2}=(1,0)'$ (or the other way around), we are in situation when $P(Y_{i1}=1|{X}_{i1}=(0,1)', {X}_{i2}=(1,0)')=P(Y_{i2}=1|{X}_{i1}=(0,1)', {X}_{i2}=(1,0)')$, $P(Y_{i1}=1|{X}_{i1}=(1,0)', {X}_{i2}=(0,1)')=P(Y_{i2}=1|{X}_{i1}=(1,0)', {X}_{i2}=(0,1)')$. The exhaustive characterization leads to the conclusion that the identified set is $\mathcal{A}_0=\{(1,1)\}$. Similarly, for the case ${\Greekmath 010B}_1=0$ is the case of point identification with $\mathcal{A}_0=\{(0,1)\}$. If ${\Greekmath 010B}_1 \notin \{0,1\}$ in the DGP, then the identified set will be an intersection of an open interval in the normalized space with $\mathcal{A}$.

Properties of estimators based on coarse structures

The three theoretical examples share a common feature with some support points \(x\) have \({\Greekmath 011E}(F(Y| X = x))\) lying exactly on the boundary between regions \(C_m\), \(m \neq m(x)\). With discrete covariates, such boundary events can occur with positive probability, creating a set of ambiguous support points central to what follows. In this section, we formalize this and develop a unified theory of how classical estimators, which were designed for continuous settings and treating boundary events as measure-zero, behave when such events have positive probability.

We first characterize empirical settings in which estimators inducing coarse structures exhibit potentially undesirable properties (formally defined later in Section (ref)), establish their asymptotic properties, and relate these findings to our theoretical examples (Sections (ref)--(ref)) and the illustrative designs in Section (ref).

For an estimator that induces a coarse structure, the definitions of exhaustiveness and coarseness imply that the coarse set $\mathcal{A}^*$ contains the identified set $\mathcal{A}_0$. In our illustrative designs with discrete regressors, several possibilities arise. For some values of ${\Greekmath 010B}$ in the DGP, $\mathcal{A}^*$ coincides with $\mathcal{A}_0$. For others, $\mathcal{A}^*$ differs from $\mathcal{A}_0$ only at the boundary. In still other cases, $\mathcal{A}^*$ differs from $\mathcal{A}_0$ more substantially, with a positive Hausdorff distance $d_H(\mathcal{A}^*, \mathcal{A}_0)$. In our theoretical examples and illustrative designs of this type, $\mathcal{A}^*$ has strictly larger affine dimension than $\mathcal{A}_0$. Our attention will ultimately be on such cases.

Discrete setup: finite support and regime classification

Before we proceed to our formal results, we describe a discrete setting underlying them.

Assume $X$ is discrete with finite support $\mathcal X=\{x_1,\ldots,x_{|\mathcal X|}\}$ and an estimator $\widehat F(\cdot\mid x)$ which is used in a plug-in sample objective function ((ref)) is consistent on $\mathcal X$. Throughout, we assume that function ${\Greekmath 011E}$ and ${\Greekmath 0120}$ in the description of exhaustive and coarse structures are continuous.

For each $x\in\mathcal X$, define the finite-sample feasible regime set $$ \mathcal M_{fs}(x) :=\Bigl\{m\in\mathcal M:\ {\mathbb P}\bigl({\Greekmath 011E}(\widehat F(Y\mid X=x))\in C_m\bigr)>0\Bigr\}, $$ and the asymptotically relevant regime set $$ \mathcal M_{as}(x) :=\Bigl\{m\in\mathcal M:\ \liminf_{n\to\infty} {\mathbb P}\bigl({\Greekmath 011E}(\widehat F(Y\mid X=x))\in C_m\bigr)>0\Bigr\}. $$ Denote the set of persistently ambiguous covariate points $$\mathcal X_0:=\{x\in\mathcal X:\ |\mathcal M_{as}(x)|\geq 2\}=\{x_1,\ldots,x_T\}.$$ In our formal results later $\mathcal{X}_0 \neq \varnothing$ will be exactly the reasons why the Hausdorff distance between $\mathcal{A}^*$ and $\mathcal{A}_0$ will be zero.

Consequently, for $x\notin\mathcal X_0$ we will impose asymptotic uniqueness of regime classification:

align[align omitted — 103 chars of source]

\paragraph{Feasible systems and their solution sets.} For each regime of indices $m=(m_1,\ldots,m_T)$, $m_t\in\mathcal M_{as}(x_t)$, $t=1,\ldots,T$, we define the associated constraint system:

align[align omitted — 352 chars of source]

Let $A(m)$ be the solution set to (ref)--(ref) in $\mathcal A$. Some $A(m)$ may be empty. Let $A_1,\ldots,A_L$ denote the distinct nonempty sets among $\{A(m)\}$. Thus, each region $A_{\ell}$ corresponds to some resolution of the boundary regime induced by sampling noise at a finite number of ambiguous covariate points in $\mathcal{X}_0$.

For each $\ell= 1,\ldots,L$, and each $t\in\{1,\ldots,T\}$, fix an index $m_{\ell,t}\in\mathcal M_{as}(x_t)$ such that $$A_\ell=A(m_{\ell,1},\ldots,m_{\ell,T}).\footnote{Hypothetically, at this stage we can agree that if multiple index vectors potentially generate the same $A_\ell$, then we pick any one as our results do not depend on which representative is chosen. As we will see later, such cases will be eliminated by our conditions in the general theorems.}$$

For each $\ell= 1,\ldots,L$, define the set of nonredundant ambiguous constraints for $A_\ell$ as any subset $\mathcal X_{0,\ell}\subseteq\mathcal X_0$ such that removing a constraint ${\Greekmath 0120}(x_t'a)\in D_{m_{\ell,t}}$ for $x_t\in\mathcal X_{0,\ell}$ changes $A_\ell$.

First, note that for ${\Greekmath 010B}$ in the DGP that results in $\mathcal{X}_0=\varnothing$, we have then $L=1$ and $A_1=\mathcal{A}^*$ as only ((ref)) apply. Second, if ${\Greekmath 010B}$ in the DGP results in $\mathcal{X}_0 \neq \varnothing$, then little can be said about the nature of $A_1$, \ldots, $A_L$ without further describing the properties of how $D^*_{m_t}$ across different $m_t \in \mathcal{M}_{as}(x_t)$. For our theoretical examples we will show that in such cases $L \geq 2$ given tha the parameter space $\mathcal{A}$ is large enough and has non-empty relative interior in the normalized space.

These considerations motivate us to impose the following Assumption (ref).

assumptionWe will assume that $L\geq 1$ when $\mathcal{X}_0 \neq \varnothing$.

This assumption is weak requirement on the existence of the combination of indices $m_1,\ldots,m_T$, $m_t \in \mathcal{M}_{as}(x_t)$ for which the sample objective function of optimized when each term takes its highest possible value. Note that our conditions in general theorems below will ultimately require more from $L$ and $A_1$, \ldots, $A_L$ (e.g. conditions (C9)-(C11)), but these requirements will be aligned what happens in our theoretical examples 1-3.

Estimators and general theorems

Given our discrete setup we consider a population objective in the form of (ref) and its sample version (ref) with a continuous score function $s$ that satisfies (ref)-(ref).

As discussed earlier, the Argmax set for (ref) coincides with $\mathcal{A}^*$. We will denote the set of maximizers of the sample objective function (ref) as $\widehat{A}$: $$\widehat{A} := \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{Argmax}_{a \in \mathcal{A}}\sum_{x\in\mathcal X_n} w(x) s\!\big({\Greekmath 011E}(\widehat{F}(Y\mid X=x)),{\Greekmath 0120}(x,a)\big) \, \widehat P(X=x).$$ We present our results related to the properties of $\widehat{A}$ in four different theorems.

Our first general Theorem (ref) shows, among other things, that in our discrete setup the estimator $\widehat{A}$ under a coarse structure does not converge to a single deterministic set if $L>1$. Instead, it converges in distribution to a {random choice} among deterministic regions $A_1,\ldots,A_L$.

theoremAssume a discrete setup in Section (ref). Let Assumption (ref) hold. Consider the following three conditions: \begin{enumerate} • For each $t=1,\ldots,T$, it holds that $D^*_{m(x_t)}\supseteq D^*_{m_t} \quad \forall m_t \in \mathcal{M}_{as}(x_t)$. • For each $t=1,\ldots,T$, it holds that $ D^*_{m_t} \cap D^*_{m'_t} =\emptyset \quad \forall m_t, m_t' \in \mathcal{M}_{as}(x_t)$. • For each $t$ and each $m_t \in \mathcal{M}_{as}(x_t)$, $t=1,\ldots,T$, \begin{align}{\Greekmath 011E}(\widehat{F}(Y|X=x_t)) \in \mathcal{C}_{m_t} \quad & \Rightarrow \quad s({\Greekmath 011E}(\widehat{F}(Y|X=x_t)), u) > s({\Greekmath 011E}(\widehat{F}(Y|X=x_t)), u') \notag \\ \forall \; u \in \cup_{m \in \mathcal{M}_{as}(x_t)} D^*_{m}, & \quad \forall \; u' \notin \cup_{m \in \mathcal{M}_{as}(x_t)} D^*_{m}. \qquad \end{align} Also, for any $\ell=1,\ldots, L$, \begin{equation} \lim \inf_{n \to \infty} {\mathbb P}\left( \left(\cap_{x \in \mathcal{X}\setminus \mathcal{X}_0} ({\Greekmath 011E}(\widehat{F}(y=1|x)) \in C_{m(x)})\right) \cap \left(\cap_{t=1}^T ({\Greekmath 011E}(\widehat{F}(y=1|x_t)) \in C_{m_{\ell,t}}\right) \right)>0. \end{equation} Also, for each $t$ and every pair $\ell \neq \ell'$, $\ell,\ell' \in \{1,\ldots, L\}$, {\begin{align} \lim_{n\to\infty} {\mathbb P}\Bigg( \sum_{t=1}^T w(x_t)\,\widehat P(X=x_t) \Big[ s\!\big({\Greekmath 011E}(\widehat F(Y|X=x_t)),u_{\ell,t}\big) - s\!\big({\Greekmath 011E}(\widehat F(Y|X=x_t)),u_{\ell',t}\big) \Big] =0 \Bigg) =0, \end{align}} where $u_{\ell,t}\in D^*_{m_{\ell,t}}$ and $u_{\ell',t}\in D^*_{m_{\ell',t}}$. (The expression does not depend on the particular choice of $u_{\ell,t}$ and $u_{\ell',t}$ because $s({\Greekmath 011E}(\widehat{F}(Y|X=x_t)),\cdot)$ is constant on each $D_m^*$ by (ref).) \end{enumerate} Then: \begin{enumerate} • Under (C1), $A_{\ell} \subseteq \mathcal{A}^*$, $\ell =1,\ldots, L$. • Under (C2), $A_{\ell} \cap A_{\ell}'=\varnothing$ for $\ell \neq \ell'$. • Condition (ref) in (C3) implies that $${\mathbb P}\left(\widehat{A} \in \left\{ \cup_{\ell \in \mathcal{I}} A_{\ell} : \mathcal{I} \subseteq \{1,\ldots, L\}\right\} \right) \to 1 \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ as } n \to \infty.$$ Condition (ref) in (C3) when considered together with (ref) additionally implies that for each $\ell=1, \ldots, L$, $$\inf \lim_{n \rightarrow} {\mathbb P}\left(A_{\ell} \subseteq \widehat{A} \right) >0 \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ as } n \to \infty.$$ Condition (ref) in (C3) when considered together with (ref), (ref) additionally implies that for $\ell \neq \ell'$. $$\inf \lim_{n \rightarrow} {\mathbb P}\left(A_{\ell} \subseteq \widehat{A}, A_{\ell'} \subseteq \widehat{A} \right) \to 0 \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ as } n \to \infty.$$ Overall, the implications of all (ref)-(ref) can be summarized as \begin{equation} \widehat{A} \stackrel{d}{\to} \sum_{\ell=1}^{L} d_{\ell} A_{\ell}, \end{equation} where $d_{\ell}$ are binary $0/1$ variables such that $\sum_{\ell=1}^{L} d_{\ell} =1$, $d_{\ell} d_{\ell'}=0$ for $\ell \neq \ell'$. \end{enumerate}

Conditions (C1) and (C2) are straightforward. Condition (ref) in (C3) ensures that, with probability approaching 1, the sample objective function attains its maximum over one of the sets $A_1,\dots,A_L$. Condition (ref) ensures that each of these sets remains asymptotically relevant. This is the first point at which we require a statement about the joint behavior of ${\Greekmath 011E}(\widehat F(y\mid x))$ across all $x\in\mathcal X$. Earlier, when classifying $\mathcal X$ into $\mathcal X_0$ and $\mathcal X\setminus\mathcal X_0$ on the basis of $\mathcal M_{as}(x)$, it was enough to study the behavior at each $x$ separately, whereas now we require control of how these objects behave simultaneously across the support. Condition (ref) guarantees asymptotic uniqueness of the maximizing region. We have stated this condition deliberately in a generic form. In many applications, one can replace it with more interpretable sufficient conditions tailored to the structure of the objective function. Our theoretical examples illustrate how such conditions can be verified in settings where the choice probabilities enter linearly, and they make clear the kinds of arguments that can be used more generally to establish this property.

Note that in formulation (ref) we do not need to invoke the notion of weak convergence from random set theory. In that literature, weak convergence is typically defined for random closed sets (see molchanov2006book) and is based on the Fell topology, which in particular requires closedness of the sets. The sets $A_1,\ldots,A_\ell$ need not be closed, and taking closures at this stage would unnecessarily complicate the analysis. Instead, we exploit that $\widehat A$ takes values in a finite collection and can therefore be treated as a discrete random element. This renders the use of random set convergence machinery unnecessary at this stage.

The next theorem provides conditions under which the identified set $\mathcal A_0$ lies on the boundary of each $A_\ell$, reflecting that the information lost under the coarse characterization consists of equality-type restrictions (e.g.\ $x'{\Greekmath 010B}=0$), which generically define lower-dimensional manifolds.

theoremAssume a discrete setup in Section (ref). Let Assumption (ref) hold. Consider the following three conditions: \begin{enumerate} • For every $\ell$ and every $a\in\mathcal A$ such that \[ {\Greekmath 0120}(x'a)\in \overline D^*_{m(x)} \quad \forall x\notin\mathcal X_0, \qquad {\Greekmath 0120}(x_t'a)\in \overline D^*_{m_{\ell,t}} \quad \forall x_t\in\mathcal X_{0,\ell}, \] and for every ${\Greekmath 0122}>0$, there exists $a^+_{\Greekmath 0122}\in\mathcal A$ with $\|a^+_{\Greekmath 0122}-a\|<{\Greekmath 0122}$ such that \[ {\Greekmath 0120}(x'a^+_{\Greekmath 0122})\in D^*_{m(x)} \quad \forall x\notin\mathcal X_0, \qquad {\Greekmath 0120}(x_t'a^+_{\Greekmath 0122})\in D^*_{m_{\ell,t}} \quad \forall x_t\in\mathcal X_{0,\ell}. \] • For every ${\Greekmath 010B}\in\mathcal A_0$, every $\ell$, for every ${\Greekmath 0122}>0$ there exist $a^+_{\Greekmath 0122},a^-_{\Greekmath 0122}\in\mathcal A$ with $\|a^\pm_{\Greekmath 0122}-{\Greekmath 010B}\|<{\Greekmath 0122}$ such that \[ {\Greekmath 0120}(x,a^\pm_{\Greekmath 0122})\in D^*_{m(x)} \quad \forall x\notin\mathcal X_0; \] \[ {\Greekmath 0120}(x_s,a^+_{\Greekmath 0122}) \in D^*_{m_{\ell,s}}, \quad \forall x_s\in\mathcal X_{0,\ell}, \qquad {\Greekmath 0120}(x_s,a^-_{\Greekmath 0122})\notin D^*_{m_{\ell,s}}, \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ for some } x_s\in\mathcal X_{0,\ell}. \] • For each $x_t\in\mathcal X_0$ and each $m\in\mathcal M_{as}(x_t)$, we have $D_{m(x_t)} \subseteq \partial D^*_{m_{\ell,t}} \quad \forall \ell \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ with } x_t\in\mathcal X_{0,\ell}$. \end{enumerate} Then: \begin{enumerate} • Under (C4), for each $\ell=1,\ldots,L$. \[ \overline A_\ell = \Bigl\{a\in\mathcal A:\ {\Greekmath 0120}(x'a)\in \overline D^*_{m(x)} \ \forall x\notin\mathcal X_0,\ {\Greekmath 0120}(x_t'a)\in \overline D^*_{m_{\ell,t}} \ \forall x_t\in\mathcal X_{0,\ell} \Bigr\}. \] • When $\mathcal{X}_0 \neq \varnothing$, (C5) and (C6) imply that $\mathcal A_0 \subseteq \partial A_\ell$ for all $\ell=1,\ldots,L.$ \end{enumerate}

Condition (C4) says that any point satisfying the closed versions of the constraints defining $A_\ell$ (i.e.\ with $\overline{D^*_m}$ in place of $D^*_m$) can be approximated arbitrarily closely by a point satisfying the open constraints. In other words, no point on the boundary of $A_\ell$ is isolated from the interior of $A_\ell$. This is a standard regularity condition ensuring that the closure $\overline{A_\ell}$ is the closure of the interior of $A_\ell$, rather than containing isolated boundary points.

Condition (C5) is a local crossing condition at the identified set. It says that within any neighborhood of ${\Greekmath 010B} \in \mathcal{A}_0$, one can find both a point satisfying all constraints of $A_\ell$ and a point violating at least one constraint (hence, not being in $A_{\ell}$). This crossing property captures the fact that $\mathcal{A}_0$ sits precisely on the boundary between feasibility and infeasibility for the constraints at $\mathcal{X}_0$.

Condition (C6) says that the exhaustive constraint set $D_{m(x_t)}$ at each ambiguous point $x_t\in\mathcal{X}_0$ lies on boundaries of competing $D_{m_t}$ for $m_t \in \mathcal{M}_{as}(x_t)$. In particular, this condition captures the property that exhaustive sets $D_{m(x_t)}$ corresponding to ambiguous points are “thin”.

Theorem (ref) below establishes under which conditions the closure of the coarse identified set $\mathcal A^*$ is the union of the closures of regions $A_{\ell}$, which means that coarsening expands the identified set by filling in neighborhoods around the true boundary constraints.

theoremAssume a discrete setup in Section (ref) and Let Assumption (ref) hold. Define the closed coarse-feasible set \[ \mathcal A^{*,\bullet} := \{a\in\mathcal A:\ {\Greekmath 0120}(x,a)\in \overline D^*_{m(x)}\ \forall x\in\mathcal X\}. \] Consider the following conditions: \begin{enumerate} • for every $a\in\mathcal A^{*,\bullet}$ with ${\Greekmath 0120}(x,a)\in \overline D^*_{m(x)}\setminus D^*_{m(x)}$ for at least one $x$, there exists a neighborhood $U\ni a$ such that \[ \{{\Greekmath 0120}(x,a): a \in U\}\cap \prod_{x\in\mathcal X} D^*_{m(x)}\neq\emptyset, \] • for each $t=1,\ldots,T$, \[ \overline D^*_{m(x_t)} \subseteq \overline{\bigcup_{m\in\mathcal M_{as}(x_t)} D^*_m }, \] \end{enumerate} Then: \begin{enumerate} • Under (C7), $\overline{\mathcal A^*}=\mathcal A^{*,\bullet}$. • (C1), (C4) (C7), (C8) imply that $\overline{\mathcal A^*}= \bigcup_{\ell=1}^L \overline A_\ell$. \end{enumerate}

Condition (C7) is an analogue of (C4) but for the coarse set $\mathcal{A}^*$ and rules out isolated boundary points of $\mathcal{A}^*$ that cannot be approximated from the interior. Condition (C8) says that any point that is feasible under the closed coarse constraint at an ambiguous point $x_t$ is also feasible under the closed coarse constraint of at least one of the asymptotically relevant regimes, ensuring no points are “lost” when passing from $\overline{\mathcal{A}}$ to $\bigcup_\ell \overline{A_\ell}$.

The next theorem will give us a result on how $\overline{\mathcal{A}}_0$ can be recovered from a sub-collection of $\{A_{\ell}\}_{\ell=1}^L$. It will also imply a sample mechanism through which this can be attained. To formulate this theorem, let us bring back the fact that each $A_{\ell}$ is obtained through a choice of indices \[ \mathbf{m}=(m_1,\ldots,m_T),\qquad m_t\in\mathcal M_{as}(x_t), \quad t=1, \ldots, T. \] For convenience we may represent such profiles as $\mathbf{m} =(m_t,m_{-t})$ when dealing with a particular $t=1, \ldots, T$.

For a given $A_{\ell}$ let $\mathcal{F}_{\ell}$ collect all profiles $\mathbf{m}$ that result in $A_{\ell}$. Since $A_{\ell}$ is non-empty, then $\mathcal{F}_{\ell} \neq \emptyset$.

theoremAssume a discrete setup in Section (ref) with $\mathcal{X}_0 \neq \varnothing$ and let Assumption (ref) hold. Suppose all conditions of Theorems (ref)- (ref) hold. Consider the following conditions: \begin{itemize} • Each $\mathcal{F}_{\ell}$ is a singleton. For each $x_t$ and each subset $\mathcal{J}_t \subseteq \mathcal{M}_{as}(x_t)$ such that $|\mathcal{J}_t| \geq \lfloor \frac{|\mathcal{M}_{as}(x_t)|}{2}\rfloor +1$, it holds that $\overline D_{m(x_t)} = \bigcap_{m\in\mathcal J_t} \overline D_m^*.$ • For each $\ell$ and each $t$ in the decomposition $\mathbf{m}=(m_t,m_{-t}) \in \mathcal{F}_{\ell}$, all $(|\mathcal{M}_{as}(x_t)|-1)$ profiles $\mathbf{m}=(\widetilde{m}_t,m_{-t})$ with $\widetilde{m}_t \in \mathcal{M}_{as}(x_t)$, $\widetilde{m}_t \neq m_t$, belong to other $\mathcal{F}_{\ell'}$. • For each $t$ every $m_t \in \mathcal{M}_{as}(x_t)$ is part of some $\mathbf{m}=(m_t,m_{-t}) \in \cup_{\ell=1}^L \mathcal{F}_{\ell}$. Moreover, for each $\ell$, $t$ and each $m_t$ in the decomposition $\mathbf{m}=(m_t,m_{-t}) \in \mathcal{F}_{\ell}$, the number of profiles $\mathbf{m}=({m}_t,\widetilde{m}_{-t}) \in \cup_{\ell=1}^L \mathcal{F}_{\ell}$ with $\widetilde{m}_{-t} \neq m_{-t}$, is the same number ${N}(x_t)$ \end{itemize} If (C9), (C10) hold or (C9), (C11) hold, then the intersection of any $\lfloor L/2\rfloor+1$ sets $\overline A_\ell$ equals the closure of the identified set: $$\bigcap_{\ell\in S}\overline A_\ell = \overline{\mathcal A}_0 \qquad \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{for all } S\subset\{1,\ldots,L\},\ |S|\ge \lfloor L/2\rfloor+1.$$

Theorem (ref) tells us that the identified set can be recovered (up to the closure) from the intersection of any $[L/2]+1$ sets $\overline{A}_{\ell}$. This will be a major motivation for us to consider in the discrete setup of Section (ref) a Random Set Quantile estimator in a sample.

Condition (C9) consists of two parts. The first part requires each family $\mathcal{F}_\ell$ to be a singleton, meaning that each nonempty region $A_\ell$ is generated by exactly one profile $\mathbf{m}=(m_1,\ldots,m_T)$. The second part of (C9) requires that the majority intersection of the closed coarse sets $\overline{D^*_m}$ recovers the closed exhaustive constraint set $\overline{D_{m(x_t)}}$ at each $x_t\in\mathcal{X}_0$.

From a substantive perspective, (C10) is a local injectivity requirement. For a fixed $t$ and a fixed configuration $m_{-t}$, the mapping from $m_t$ to $\overline A_\ell$ is injective as changing the local component $m_t$ necessarily produces a new and previously unseen set $\overline A_\ell$. Hence, incompatibility at index $t$ is translated into exclusion from many distinct $\overline A_\ell$'s. The majority property is then obtained pointwise, profile by profile. Condition (C11), on the other hand, enforces a symmetry condition by requiring that each local index $m_t \in \mathcal M_{as}(x_t)$ appears exactly $N(x_t)$ times across the family $\{\mathcal F_\ell\}$. This uniform replication ensures that any incompatibility associated with a given $m_t$ is neither over- nor under-represented. In other words, while violation may recur across multiple sets, it does so evenly, so that incorrect parameter values still appear in at most half of the $\overline A_\ell$'s.

When we verify the conditions of this theorem later for our theoretical examples, we will see that e.g. in the cross-sectional binary choice models under quantile independence, (C10) will be applicable under linear independence of elements in $\mathcal{X}_0$, whereas (C11) will be applied more generally, even in the case of linear dependence.\footnote{Thus, in that Theoretical example 1 condition (C11) will be more general than (C10) but this does not mean this is true generally.}

We can summarize the results of general theorems in the following result.

theoremAssume a discrete setup in Section (ref). \begin{enumerate} • Suppose parameter value ${\Greekmath 010B}$ in the data generating process yields $\mathcal{X}_0 \neq \varnothing$. Let Assumption (ref) hold, and let $A_1$, \ldots, $A_L$, $L \geq 1$, be the deterministic nonempty distinct sets obtained as all possible distinct solutions to ((ref))-((ref)). If (C1)-(C9) hold, and either (C10) or more general (C11) holds, then $L \geq 2$ and $$\widehat{A} \stackrel{d}{\rightarrow} d_1 A_1 + \ldots +d_L A_L,$$ where $\cup_{\ell=1}^L \overline{A}_{\ell} = \overline{{\mathcal A}^*}$. In addition, \begin{itemize} • $d_1$, \ldots, $d_L$ are dummy variables such that $d_{\ell} d_m =0$ for $\ell \neq m$ (mutually exclusive), and $d_1+\ldots+d_L=1$ (collectively exhaustive), and ${\mathbb P}(d_{\ell}=1)\in (0,1)$ for each $\ell$, $\sum_{\ell=1}^L {\mathbb P}(d_{\ell}=1)=1$. • The boundary of each $A_{\ell}$ contains the identified set ${\mathcal A}_0$. • The intersection of any $[L/2]+1$ closed sets $\overline{A}_{\ell}$ coincides with the closure $\overline{{\mathcal A}}_{0}$. \end{itemize} Also, $\overline{\widehat{A}} \stackrel{W}{\rightarrow } \mathbf{A}({\Greekmath 010B})$, where $\stackrel{W}{\rightarrow }$ is the weak convergence from the random set theory, and random set $\mathbf{A}({\Greekmath 010B})$ has the following distribution: $${\mathbb P}(\mathbf{A}({\Greekmath 010B})=\overline{A}_{\ell})={\mathbb P}(d_{\ell}=1), \quad \sum_{\ell=1}^L {\mathbb P}(\mathbf{A}({\Greekmath 010B})=\overline{A}_{\ell}) =1.$$ • Suppose parameter value ${\Greekmath 010B}$ in the data generating process yields $\mathcal{X}_0 = \varnothing$. Then ${\mathbb P}(\widehat{A}\neq \mathcal{A}^*) \to 0 \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ as } n \to \infty,$ which implies $\widehat{A} \stackrel{d}{\rightarrow} \mathcal{A}^*$, $\overline{\widehat{A}} \stackrel{W}{\rightarrow} \overline{\mathcal{A}^*}$. \end{enumerate}

Note that Case 2 of Theorem (ref) follows directly from the discrete setup of Section (ref) and the finite support of $X$, independently of Theorems (ref)-(ref).

Convergence $\stackrel{d}{\rightarrow}$ of $\widehat{A}$, which is not necessarily a closed set. was discussed after Theorem (ref). Properties (a)-(c) in the first statement will be very important for us later on when we motivate a new estimator in our theoretical examples. Since the support $\overline{\widehat{A}}$ is finite, weak convergence $\stackrel{W}{\rightarrow}$ of closed $\overline{\widehat{A}}$ is equivalent to convergence of the probabilities. E.g., in statement 1 this is equivalent to ${\mathbb P}(\overline{\widehat{A}}=\overline{A}_{\ell}) \to {\mathbb P}(\overline{\widehat{A}}=\overline{A}_{\ell})$, $\ell=1, \ldots, L$. The Fell topology is the standard hit-or-miss topology on the hyperspace of closed subsets of $\mathcal{A}$.

To summarize, Theorems (ref)-(ref) show that estimators based on coarse structures may converge to a random set rather than a fixed target, with the identified set on the boundary of every realization. Section (ref) formalizes what we should ask of an estimator in this setting; Section (ref) proposes one that delivers it.

Application of general theorems to our examples

We now illustrate how our general theorems apply to the theoretical examples in Sections (ref)–(ref). In each case, the same mechanism as in our general discrete setup arises, with discrete covariates generating boundary regimes that classical estimators ignore or weaken, resulting in multiple admissible regions.

Binary choice cross-sectional model

We continue to consider the model discussed in Section (ref), where we gave exhaustive and coarse characterizations. Since the coarse characetrization of interest is the one induced by the maximum score estimator, it means that application of Theorems (ref)-(ref) will produce the result for the asymptotic behavior of the maximum score estimator $\widehat{A}$ obtained as a maximizer of the sample objetive function (ref) with the score function (ref). Recall that the coarse $\mathcal{A}^*$ in this case is the maximizer of the population objective (ref) with the same score function.

In line with the discrete design of interest, suppose $X$ has finite support $\mathcal X=\{x_1,\ldots,x_{|\mathcal X|}\}$. Before we engage with the terminology of Section (ref) for this model, let us formulate the following result about the joint and individual asymptotic behaviors of plug-in estimators $\widehat P(Y=1|{x})$.

propositionSuppose $X$ has finite support $\mathcal{X}$ as described above and we have a random sample $\{(x_i,y_i)\}_{i=1}^n$. Then, for a given $x \in \mathcal{X}$: \begin{itemize} • If $P(Y=1|{x})-{\Greekmath 010D} \in C_j$ for $j \in \{1,2\}$, then ${\mathbb P}(\widehat P(Y=1|{x}) -{\Greekmath 010D} \in C_j) \to 1$ as $n \to \infty$. • If $P(Y=1|{x})-{\Greekmath 010D} \in C_3$, then ${\mathbb P}(\widehat P(Y=1|{x}) -{\Greekmath 010D} \in C_m) \to \frac{1}{2}$ for $m \in \{1,2\}$. \end{itemize}

Parts (a) and (b) in Proposition (ref) imply that in the terminology and description of Section (ref) the subset $\mathcal{X}_0 \subseteq \mathcal{X}$ of ambiguous points with multiple possible asymptotic regimes will be composed of all $x \in \mathcal{X}$ for which $m(x)=3$:

equation[equation omitted — 112 chars of source]

For each $x_t \in \mathcal{X}_0$ we have $\mathcal M_{as}(x_t)=\{1,2\}$ whereas $\mathcal M_{as}(x)=\{m(x)\}$ for each $x \in \mathcal{X}\setminus \mathcal{X}_0$,

If $\mathcal X_0=\varnothing$, then things are pretty straightforward as all asymptotic regimes are aligned with the true regimes 1 or 2. $\mathcal A^*$ in this case differs from $\mathcal A_0$ at most at the boundary (due to the weak inequality in $D_1^*$), we have $L=1$ and $A_1=\mathcal A^*$.

The challenging case is that of $\mathcal X_0\neq\varnothing$, in which identification uses boundary equalities $x'{\Greekmath 010B}=0$ that are weakened or dropped by the coarse structure. For this case, we suppose, in line with out discussion earlier, that the compact parameter set $\mathcal{A}$ has a nonempty relative interior in the $(k-1)$-dimensional normalized space. This guarantees that the set $\mathcal{A}_{out} = \{a \in \mathcal{A}: x'a \in D_{m(x)} \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ for all } x \in \mathcal{X}\setminus \mathcal{X}_0\}$ described by only strict inequalities in the exhaustive characterization has a nonempty relative interior in the $(k-1)$-dimensional normalized space. This feature of $\mathcal{A}_{out}$ is important in verifying conditions of Theorem (ref) for this model. The coarse set $\mathcal{A}^*$ only differs from $\mathcal{A}_{out}$ at the boundary as $D_1^*\setminus D_1 =\{0\}$. This condition on $\mathcal{A}$ also guarantees that with $\mathcal X_0\neq\varnothing$ the affine dimension of $\mathcal{A}_0$ is strictly smaller than that of $\mathcal{A}^*$ and that those points in $\mathcal{A}_0$ that are not at the boundary at $\mathcal{A}$ are in the relative interior of $\mathcal{A}^*$ in the $(k-1)$-dimensional normalized space.

The mechanics of constructing sets $A_1,..,A_L$ when $\mathcal X_0\neq\varnothing$ is straightforward. E.g., if $|\mathcal{X}_0|=1$, then our condition on the non-empty relative interior of the parameters space guarantees that we can construct two distinct regions $A_1$ and $A_2$, defined by replacing the equality constraint $x_{t}'a=0$ with one of the two inequalities. When $|\mathcal{X}_0|>1$, then analogously each equality is replaced with one of the two strict inequalities for the construction of $A_{\ell}$ going all possible combinatorial ways of picking inequalities. Some of these systems may give an empty set (when this occurs is discussed in the verification of (C11) ) but there is always at least two distinct non-empty $A_{\ell}$, as again follows from our verification of (C10) and (C11) and which is quite intuitive given the logic for $|\mathcal{X}_0|=1$.

To sum up, a non-empty $A_\ell$ in the construction of Section (ref) applied to this model, satisfies $x'a\geq 0$ if $m(x)=1$, $x'a<0$ if $ m(x)=2$, and for ambiguous points $x_t\in\mathcal X_{0,\ell}\subseteq\mathcal X_0$, imposes one of the two coarse constraints: $x_t'a\ge 0$ if $m_{\ell,t}=1$, or $x_t'a<0$ if $m_{\ell,t}=2$, where, recall, $m_{\ell,t}$ denotes the relevant regime from $\mathcal{M}_{as}(x_t)$ that participates in the construction of $A_\ell$.

Proposition (ref) below helps us to verify (C3).

propositionSuppose $X$ has finite support $\mathcal{X}$ as described above and we have a random sample $\{(x_i,y_i)\}_{i=1}^n$. Let $\mathcal{X}_0=\{x_1,\ldots,x_T\}$ for $\mathcal{X}_0$ defined in (ref). Then \begin{itemize} • $\sqrt{n}(\widehat P(Y=1|{x}_1)-P(Y=1|{x}_1), \ldots, \widehat P(Y=1|{x}_T)-P(Y=1|x_T))' \stackrel{d}{\to} \mathcal{N}(0, \Sigma)$ for a positive definite $\Sigma$. • for any $(m_1,\ldots, m_T) \in \{1,2\}^T$. $${\mathbb P}\left(\left(\cap_{x \in \mathcal{X}\setminus \mathcal{X}_0} (\widehat P(Y=1|{x}) -{\Greekmath 010D} \in C_{m(x)}) \right) \cap \left(\cap_{t=1}^T (\widehat P(Y=1|{x}_t) -{\Greekmath 010D} \in C_{m_t}) \right) \right) \to \frac{1}{2^T} $$ \end{itemize}

Verification of conditions (C1)-(C9) and then (C10) or (C11) is given in the Appendix. From that verification it is clear that (C10) implies (C11) via coordinate-wise pairing of feasible profiles. (C10) holds when elements in $\mathcal{X}_0$ are linearly independent, whereas (C11) holds generally, in particular under linear dependence in $\mathcal{X}_0$.

The following general result, presented in Theoretical Example 1, builds on Theorems (ref)–(ref) as they apply to this model, while also establishing a new result concerning the equal probabilities of deterministic sets arising in the distributional limit of the maximum score estimator.

theoremConsider a cross-sectional binary choice model under assumptions discussed in Section (ref) and within the discrete setup given in Section (ref). Suppose that the parameter space has a relative interior in the $(k-1)$-dimensional normalized space. 1. If the parameter value ${\Greekmath 010B}$ in the DGP results in $\mathcal{X}_0 \neq \varnothing$, then there are $L \geq 2$ distinct nonempty deterministic sets $A_1, \ldots, A_L$ constructed in the way described earlier in this section (aligned with considering all possible nonempty solutions sets to (ref)-(ref)) and the maximum score estimator $\widehat{A}$ satisfies part 1 of Theorem (ref) with these sets $A_1$, ..., $A_L$ appearing in the weak limits. Moreover, ${\mathbb P}(\widehat{A}=A_{\ell}) \to \frac{1}{L}$, $\ell=1,\ldots, L$. 2. If the parameter value ${\Greekmath 010B}$ in the DGP results in $\mathcal{X}_0 = \varnothing$, then $d_H(\mathcal{A}_0, \mathcal{A}^*)=0$, $L=1$ with $A_1=\mathcal{A}^*$ and ${\widehat{A}} \stackrel{d}{\rightarrow} \mathcal{A}^*$, $\overline{\widehat{A}} \stackrel{W}{\rightarrow} \overline{\mathcal{A}}_0.$

The majority of the first statement in this theorem follows immediately from Theorem (ref) and our verification of all their conditions ( whether (C10) or more general (C11) is applied depends on linear dependence of elements in $\mathcal{X}_0$). The new statement here is on equal asymptotic probabilities of sets $A_{\ell}$. Hence, the proof of this part in the appendix is focused on showing this fact.

The second statement is quite straightforward as part (b) of Proposition (ref) implies that ${\mathbb P}(\widehat{A}=\mathcal{A}^*) = {\mathbb P}( \cap_{x \in \mathcal{X}} (\widehat{P}(Y=1|x) - {\Greekmath 010D} \in C_{m(x)}) ) \to 1$. Additionally, the only difference between $\mathcal{A}^*$ and $\mathcal{A}_0$ in the case $\mathcal{X}_0=\varnothing$ is for $x$ with $m(x)=1$ appearing in $\mathcal{A}_0$ with the strict inequality $x'{\Greekmath 010B}>0$ and appearing in $\mathcal{A}^*$ with the weak inequality $x'{\Greekmath 010B} \geq 0$ (thus, $\mathcal{A}^*$ and $\mathcal{A}_0$ may only differ at the boundary).

We now give a concrete application of our theorems in Illustrative Design 1.

\vskip 0.05in

Illustrative Design 1 (continued) In line with our discussion previously, we continue to assume that the parameter set $\mathcal{A}$ is the normalized 3-dimensional space (effectively making it 2-dimensional) has a 2-dimensional non-empty interior and is covex and large. Denote the points in the covariate support as $x^a=(1,0,1)'$, $x^b=(1,1,0)'$.

(1A): ${\Greekmath 010B}_{1}+{\Greekmath 010B}_{2} = 0$, ${\Greekmath 010B}_{1}+1 =0$ in the DGP. Then $\mathcal{A}_0=\{(-1,1,1)\}$, $\mathcal X_0=\{x^a,x^b\}$, $\mathcal M_{as}(x^a)=\mathcal M_{as}(x^b)=\{1,2\}$. There are four possible inequality selections $m=(m_1,m_2)\in\{1,2\}^2$. The corresponding regions that partition the coarse $\mathcal{A}^*$ are

align*[align* omitted — 253 chars of source]

Since $\mathcal A$ can be assumed to be large enough to intersect all four wedges, then $L=4$ and $\{A_\ell\}_{\ell=1}^4$ are exactly these sets. With $A_{(k,j)}$, $k,j=1,2$, defined in this way, the maximum score estimator $\widehat{A}$ asymptotically behaves as $$\widehat{A} \stackrel{d}{\rightarrow} \sum_{k,j=1}^2 d_{k,j} A_{(k,j)},$$ where $d_{(k,j)}$, $k,j=1,2$, are binary variables such that $\sum_{k,j=1}^2 d_{k,j}=1$, $d_{k,j} \cdot d_{k_1,j_1} =0$ for any two different dummies. This is a direct implication of Theorem (ref). In addition $P(d_{(k,j)}=1)=0.25$ for each $(k,j)$, as is implied by ${\mathbb P}(\widehat{P}(Y=1|x^b) -0.5 \in D^*_{m_1}, \widehat{P}(Y=1|x^b) -0.5 \in D^*_{m_2}) \to 0.25$. Finally, note that the intersection of any three $\overline{A}_{k,j}$ out of these four sets gives us $\overline{A}_0.$

(1B): ${\Greekmath 010B}_{1}+{\Greekmath 010B}_{2} \neq 0$, ${\Greekmath 010B}_{1}+1 =0$ in the DGP. In this case, $\mathcal{X}_0=\{x^b\}$, $\mathcal M_{as}(x^b)=\{1,2\}$. As for the other support point, $\mathcal M_{as}(x^a)=\{m(x^a)\}$ with $m(x^a)=1$ if $ {\Greekmath 010B}_{0,2}>1$ and $m(x^a)=2$ if $ {\Greekmath 010B}_{0,2}<1$. We can write $\mathcal A_0 =\{(-1,1,a_2)\in\mathcal A: \, a_2-1\in D_{m(x^a)}\}$. We obtain that $L=2$ and the two sets sets that partition the coarse $\mathcal{A}^*=\{(a_1,1,a_2)\in\mathcal A: \, a_2-1\in D^*_{m(x^a)}\}$ by \[ A_1^{(B)}= \{a\in\mathcal A:\ a_1+a_2\in D_{m(x^a)},\ a_1+1\ge 0\}, \quad A_2^{(B)}= \{a\in\mathcal A:\ a_1+a_2\in D_{m(x^a)},\ a_1+1< 0\} \] (again, both are non-empty as long as $\mathcal A$ is not too small which we have assumed).

With $A_1^{(B)}$, $A_2^{(B)}$ defined in this way, the MS estimator $\widehat{A}$ asymptotically behaves as $$\widehat{A}_{ms} \stackrel{d}{\rightarrow} d_1A_1^{(B)}+d_2A_2^{(B)},$$ where $d_i$, $i=1,2$, are binary variables such that $d_1+d_2=1$, $d_1 \cdot d_2=0$. In addition, $P(d_1=1)=P(d_2=1)=0.5$, as is implied by ${\mathbb P}(\widehat{P}(Y=1|x^b) -0.5 \in D^*_1) =1- {\mathbb P}(\widehat{P}(Y=1|x^b) -0.5 \in D^*_2)\to 0.5$. Finally, note that $\overline{A}_1^{(B)} \cap \overline{A}_2^{(B)}= \overline{\mathcal A_0}$.

(1C): ${\Greekmath 010B}_{1}+{\Greekmath 010B}_{2} = 0$, ${\Greekmath 010B}_{1}+1 \neq 0$ in the DGP. Here ${\Greekmath 010B}_{2}=-{\Greekmath 010B}_{1}$ and ${\Greekmath 010B}_{1}\neq -1$. Then $\mathcal X_0=\{x^a\}$, $\mathcal M_{as}(x^a)=\{1,2\}$. As for $x^b$, $\mathcal M_{as}(x^b)=\{m(x^b)\}$, where $m(x^b)=1$ if ${\Greekmath 010B}_{1}>-1$ and $m(x^b)=2$ if ${\Greekmath 010B}_{1}<-1$.

The identified set can be written as $\mathcal A_0=\{(a_1,1,-a_1)\in\mathcal A:\ a_1+1\in D_{m(x^b)}\}$ (it is one-dimensional). We have $L=2$ as different selection regimes result in partitioning the coarse $\mathcal{A}^*=\{(a_1,1,a_2)\in\mathcal A:\ a_1+1\in D_{m(x^b)}\}$ by the following two regions in our collection: $$A_1^{(C)}= \{a\in\mathcal A:\ a_1+1\in D_{m(x^b)},\ a_1+a_2\ge 0\}, \quad A_2^{(C)}= \{a\in\mathcal A:\ a_1+1\in D_{m(x^b)},\ a_1+a_2< 0\},$$ which are non-empty if $\mathcal{A}$ is not too small. The MS estimator $\widehat{A}$ asymptotically behaves as $$\widehat{A} \stackrel{d}{\rightarrow} d_1A_1^{(C)}+d_2A_2^{(C)},$$ where $d_i$, $i=1,2$, are binary variables such that $d_1+d_2=1$, $d_1d_2=0$. Additionally, $P(d_1=1)=P(d_2=1)=0.5$ as in this case ${\mathbb P}(\widehat{P}(Y=1|x^a) -0.5 \in D^*_1) =1- {\mathbb P}(\widehat{P}(Y=1|x^a) -0.5 \in D^*_2)\to 0.5$. Finally, note that $\overline{A}_1^{(C)}\cap\overline{A}_2^{(C)}= \{(a_1,1,-a_1)\in\mathcal A: a_1+1 \in \overline{D}_{m(x^b)}\}=\overline{\mathcal{A}}_0.$

(1D): ${\Greekmath 010B}_{1}+{\Greekmath 010B}_{2} \neq 0$, ${\Greekmath 010B}_{1}+1 \neq 0$ in the DGP. In this case, $\mathcal{X}_0=\varnothing$, $\mathcal M_{as}(x^j)=\{m(x^j)\}$, $j \in \{a,b\}$. Hence, $L=1$ and $A_1=\mathcal{A}^*=\{a\in\mathcal A:\ a_1+a_2 \in D_{m(x^a)},\ a_1+1 \in D_{m(x^b)}\}.$ It also holds in this case $\overline{A}_1=\overline{\mathcal{A}}_0 =\{a\in\mathcal A:\ a_1+a_2 \in \overline{D}_{m(x^a)},\ a_1+1 \in \overline{D}_{m(x^b)}\}.$

Geometric interpretations of this design for all cases (1A)-(1D) are given in Figure (ref), where we plot the sets described here in their projection on the 2-dimensional space of non-normalized components.

figure[figure omitted — 4,013 chars of source]

Application to Theoretical Example 2 (static panel binary choice)

Denote $\Delta x_i = x_{i2}-x_{i1}$, $\Delta P(Y=1|x_i) = P(Y_{i2}=1|x_i)-P(Y_{i1}=1|x_i)$, $\Delta \widehat{P}(Y=1|x_i) = \widehat{P}(Y_{i2}=1|x_i)-\widehat{P}(Y_{i1}=1|x_i)$. In line with the discrete design of interest, suppose $X$ has finite support $\mathcal X=\{x_1,\ldots,x_{|\mathcal X|}\}$. Analogously to cross-sectional binary choice case, we assume that the parameter space $\mathcal{A}$ has relative interior in the normalized space. In Proposition (ref) we formulate the individual asymptotic behaviors of plug-in estimators $\widehat P(Y=1|{x})$.

propositionSuppose $X$ has finite support $\mathcal{X}$ as described above and we have a random sample $\{(x_i,y_i)\}_{i=1}^n$. Then, for a given $x \in \mathcal{X}$: \begin{itemize} • If $\Delta P(Y=1|{x}) \in C_j$ for $j \in \{1,2\}$, then ${\mathbb P}(\Delta \widehat P(Y=1|{x}) \in C_j) \to 1$ as $n \to \infty$. • If $\Delta P(Y=1|{x}) \in C_3$, then ${\mathbb P}(\Delta \widehat P(Y=1|{x}) \in C_m) \to \frac{1}{2}$ for $m \in \{1,2\}$. \end{itemize}

Analogously to Theoretical Example 1, on the basis of results of Proposition (ref) we define

equation[equation omitted — 96 chars of source]

and construct sets $A_1$, .., $A_L$ in the same way as described there and aligned with approach given in Section (ref) as all distinct non-empty solutions to systems (ref)-(ref) . Analogous arguments can also be used to establish that when $\mathcal{X}_0 \neq \varnothing$, then we have at least two such setsimplying $L \geq 2$. Proposition (ref) below is needed to verify (C3).

propositionSuppose $X$ has finite support $\mathcal{X}$ as described above and we have a random sample $\{(x_i,y_i)\}_{i=1}^n$. Let $\mathcal{X}_0=\{x_1,\ldots,x_T\}$ for $\mathcal{X}_0$ defined in (ref). Then \begin{itemize} • $\sqrt{n}(\Delta \widehat P(Y=1|{x}_1)-\Delta P(Y=1|{x}_1), \ldots, \Delta \widehat P(Y=1|{x}_T)-\Delta P(Y=1|x_T))' \stackrel{d}{\to} \mathcal{N}(0, \Sigma)$ for some positive definite $\Sigma$. • For any $(m_1,\ldots, m_T) \in \{1,2\}^T$. $${\mathbb P}\left(\left(\cap_{x \in \mathcal{X}\setminus \mathcal{X}_0} (\Delta \widehat P(Y=1|{x}) \in C_{m(x)}) \right) \cap \left(\cap_{t=1}^T (\Delta \widehat P(Y=1|{x}_t) \in C_{m_t}) \right) \right) \to \frac{1}{2^T} $$ \end{itemize}

We are not going to verify conditions (C1)-(C11) of general Theorems (ref)-(ref) as such verification is completely analogous to the cross-sectional case. (C10) holds when all first differences $\Delta x_{1}, \ldots, \Delta x_T$ for elements $x_1,...,x_T$ in $\mathcal{X}_0$ are linearly independent. Condition (C11) , once again, is applicable generally (and more general than (C10)), in particular, in the case of linearly dependent $\Delta x_{1}, \ldots, \Delta x_T$. In fact, such linear dependence occurs in our Illustrative Design 2, as shown below. Thus, in the case of this design we apply (C11) in Theorem (ref).

We can also show for this model results analogous to those in Theorem (ref), including the asymptotic equivalence of probabilities assigned to distinct deterministic sets in the limiting distribution of the conditional maximum score estimator. Since this equal probabilities property is subsequently used to establish the consistency of our new Random Set Quantile estimator for this model, a more detailed discussion of how it is obtained is provided in the Appendix, following the proof of Theorem (ref).

\vskip 0.05in

Illustrative Design 2 (continued) Let us start by considering the following case: ${\Greekmath 010B}_0 \in \{0,1\}$. then $\mathcal{A}_0=\{{\Greekmath 010B}_0\}$ and $\mathcal{A}^*=(0,+\infty)\cap \mathcal{A}$ when ${\Greekmath 010B}_0=1$, and $\mathcal{A}^*=(-\infty,1)\cap \mathcal{A}$ when ${\Greekmath 010B}_0=0$. Suppose the parameter space $\mathcal{A}$ is a large interval containing both 0 and 1.

For concreteness, take ${\Greekmath 010B}_0=1$, Then $\mathcal{A}^*=[0,+\infty) \cap \mathcal{A}$, $\mathcal{X}_0=\{x^a, x^b\}$, where $x^a=((0,1)',(1,0)')'$, $x^b=((0,1)',(1,0)')'$, and $\mathcal M_{as}(x^a)=\mathcal M_{as}(x^b)=\{1,2\}$. Thus, there are four possible selections of inequalities with $m=(m_1,m_2)\in\{1,2\}^2$. However, selections $m_1=1$, $m_2=2$ and $m_1=2$, $m_2=1$ produce empty sets due to contradictory inequalities $a-1 \geq 0$, $a-1<0$. Thus, despite having four possible selections of inequalities there are only two non-empty sets $A_1$, $A_2$ partitioning $\mathcal{A}^*$: $A_1=[0,1)$, $A_2=[1,+\infty) \cap \mathcal{A}$. This also explains why we cannot apply (C10) in this setting and have to apply (C11). Note that $\overline{A}_1 \cap \overline{A}_2 = \overline{A}_0=\{1\}$.

The case of ${\Greekmath 010B}_0=0$ is analogous to the case ${\Greekmath 010B}_0=1$.

If e.g. ${\Greekmath 010B}_0 \in (0,1)$ in the DGP, then $\mathcal{A}^*=\mathcal{A}_0=(0,1)$, $L=1$, and $A_1$ coincides with $\mathcal{A}^*$.

Multinomial choice and block maximum score estimation

In line with our discussion of this case before, we take it as given that $0<P(Y=k| \mathbf{x})<1$ a.e. Let $K=k_1+...+k_J$ be the total number of covariates across all the indices. As one parameter value in each index is conventionally normalized in the semiparametric literature, the effective dimension of $\mathbf{X}$ is $K-J$. The support $\mathcal{X}$ is taken to be $\{\mathbf{x}_1,\ldots,\mathbf{x}_{|\mathcal X|}\}$. Analogously to binary choice cases, we assume that the parameter space $\mathcal{A}$ has relative interior in the normalized space. We start by formulating the individual asymptotic behaviors of plug-in estimators $\widehat{\mathcal{P}}(\mathbf{x}) := {\Greekmath 011E}(\widehat{F}(Y|\mathbf{x}))= (\widehat P(Y=1|\mathbf{x}), \ldots, \widehat{P}(Y=J|\mathbf{x}))^{\top}$ in Proposition (ref).

propositionSuppose $\mathbf{X}$ has finite support $\mathcal{X}$ as described above and we have a random sample $\{(\mathbf{x}_i,y_i)\}_{i=1}^n$. Then, for a given $\mathbf{x} \in \mathcal{X}$: \begin{itemize} • If $P(Y=k_1|\mathbf{x}) >\ldots > P(Y=k_J|\mathbf{x})$ for a permutation $(k_1,\ldots,k_J)$ of $(1,\ldots, J)$, then ${\mathbb P}(\widehat{P}(Y=k_1|\mathbf{x}) >\ldots > \widehat{P}(Y=k_J|\mathbf{x})) \to 1$ as $n \to \infty$. • If $P(Y=k_1|\mathbf{x}) {\Greekmath 0114}_1 \ldots {\Greekmath 0114}_{j-1} P(Y=k_J|\mathbf{x})$ for a permutation $(k_1,\ldots,k_J)$ of $(1,\ldots, J)$, where ${\Greekmath 0114}_j \in \{>,=\}$, then {\begin{multline*}{\mathbb P}\left( \bigcap_{{\Greekmath 0114}_j \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ is } ">"} \left(\widehat{P}(Y=k_jj\mathbf{x}) >\widehat{P}(Y=k_{j+1}|\mathbf{x})\right), \bigcap_{{\Greekmath 0114}_j \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ is } "="} \left(\widehat{P}(Y=k_j|\mathbf{x}) \, \mathrm{str}({\Greekmath 0114}_j) \, \widehat{P}(Y=k_{j+1}|\mathbf{x})\right)\right) \\ \to \frac{1}{2^{\Upsilon(\mathbf{x})}}, \end{multline*}} where $\mathrm{str}({\Greekmath 0114}_j) \in \{>,<\}$, $\Upsilon(\mathbf{x})=\sum_j \mathbf{1}({\Greekmath 0114}_j \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ is } "=") $. \end{itemize}

On the basis of results of Proposition (ref) we define

equation[equation omitted — 212 chars of source]

Using the ordered partitioned block terminology of Section (ref), we can equivalently define it as $\mathcal{X}_0 = \{\mathbf{x} \in \mathcal{X}: R(\mathbf{x})<J\},$ where $R(\mathbf{x})$ was the number of ordered tie blocks for $\mathbf{x}$.

For $\mathbf{x}_t\in\mathcal X_0$ , $\mathfrak B({\Greekmath 011E}({F}(y|\mathbf{x}_t)))$ has at least one block of size $\ge 2$ and $m(\mathbf{x}_t)$ is such that $$ D_{m(\mathbf{x}_t)}= \bigcap_{j<k}

cases\widetilde D_{jk;>}, & \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ if } P(Y=j|\mathbf{x}_t)>P(Y=k|\mathbf{x}_t),\\ \widetilde D_{jk;<}, & \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ if } P(Y=j|\mathbf{x}_t)<P(Y=k|\mathbf{x}_t),\\ \widetilde D_{jk;=}, & \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ if } P(Y=j|\mathbf{x}_t)=P(Y=k|\mathbf{x}_t)

. $$ The collection $ M_{as}(\mathbf{x}_t)$ consists of all $m_t$ such that $D_{m_t}$ replaces each $\widetilde D_{jk;=}$ in this definition with either $\widetilde D_{jk;>}$ or $\widetilde D_{jk;<}$ in such a way that the resulting intersection remains non-empty (i.e., the induced pairwise orderings are mutually consistent). Equivalently, each regime $m_t \in\mathcal M_{as}(\mathbf{x}_t)$ corresponds to a refinement of $\mathfrak B({\Greekmath 011E}({F}(y|\mathbf{x}_t)))$ that breaks all the within-block ties into a strict complete ordering within that block, while preserving the between-block ordering.

We can construct sets $A_1$, .., $A_L$ by solving (ref)-(ref) with various profiles $(m_1,\ldots, m_T)$ for $m_t \in \mathcal{M}_{as}(\mathbf{x}_t)$. Arguments analogous to the binary choice case can be used to establish that when $\mathcal{X}_0 \neq \varnothing$ for $\mathcal{X}_0$ defined in (ref), then we have at least two such sets. Hence, $L \geq 2$. We next have Proposition (ref) which helps us to verify (C3).

propositionSuppose $\mathbf{x}$ has finite support $\mathcal{X}$ as described above and we have a random sample $\{(x_i,y_i)\}_{i=1}^n$. Then \begin{itemize} • $\sqrt{n}(\widehat{\mathcal{P}}(\mathbf{x}_1)-\mathcal{P}(\mathbf{x}_1), \ldots, \widehat{\mathcal{P}}(\mathbf{x}_T)-\mathcal{P}(\mathbf{x}_T))' \stackrel{d}{\to} \mathcal{N}(0, \Sigma)$ for some positive definite $\Sigma$. • For any $(m_1,\ldots, m_T) \in \times_{t=1}^T \mathcal{M}_{as}(\mathbf{x}_t)$. $${\mathbb P}\left(\left(\cap_{x \in \mathcal{X}\setminus \mathcal{X}_0} (\widehat{\mathcal{P}}(\mathbf{x}) \in D_{m(\mathbf{x})}) \right) \cap \left(\cap_{t=1}^T (\widehat{\mathcal{P}}(\mathbf{x}_t) \in D_{m_t}) \right) \right) \to \frac{1}{2^{\sum_{t=1}^T \Upsilon(\mathbf{x}_t)}}.$$ \end{itemize}

Verification of conditions of our theorems (ref)-(ref) for this case is in the Appendix. If $\mathcal{X}_0$ defined in (ref) is empty, then things are straightforward as $L=1$ and our objective function is asymptotically equivalent to the original Manski (1975) objective function, and $A_1=\mathcal{A}_0$. Thus, the main case we focus on in verifying conditions is when $\mathcal{X}_0 \neq \varnothing$ and consequently $L \geq 2$.

Recall that in the binary choice case, (C11) was more general than (C10). Namely, while (C11) was true generally, (C10) required linear independence of the elements of $\mathcal{X}_0$. An analogous relationship holds in the present setting, as discussed in the Appendix. Specifically, (C11) continues to hold generally, whereas (C10) holds under a condition that parallels linear independence in the binary choice case. Linear independence itself is no longer the appropriate criterion, since ties may involve vectors associated with multiple alternatives. In its place, we impose a dimensionality reduction condition tied to each tie event. In the binary choice case this condition would have been equivalent to linear independence of elements in $\mathcal{X}_0$, Appendix contains the details of this dimensionality reduction condition as well as the application of general theorems to Illustrative Design 3.

For this model we can show results analogous to those in Theorem (ref), including the asymptotic equivalence of probabilities assigned to distinct deterministic sets in the limiting distribution of the block maximum score estimator. Jus as for the panel data binary choice model, a more detailed discussion of how the asymptotically equal probabilities are obtained is provided in the Appendix, following the proof of Theorem (ref).

Criteria used to compare estimators

We now turn to what we consider desirable properties for estimators of the identified set.

definition[Consistency] We say that $\widehat{{A}}$ is consistent at a given parameter value ${\Greekmath 010B}$ in the DGP if $d_H\!\left(\overline{\widehat{{A}}},\, \overline{\mathcal{A}}_0\right) \stackrel{p}{\rightarrow} 0$, where $d_H(\cdot,\cdot)$ denotes the Hausdorff distance on the normalized parameter space $\mathcal{A}$.

As seen in the theoretical examples above, neither $\mathcal{A}_0$ nor $\widehat{A}$ need be closed, and hence not necessarily compact. Since the Hausdorff metric is a proper metric only on the space of nonempty compact sets in finite-dimensional Euclidean spaces, we take closures of both sets in Definition (ref). Importantly, this entails no loss of content as two sets with identical closures are indistinguishable under $d_H$, so consistency is inherently defined only up to boundary points. This formulation also aligns with the notion of convergence in probability for random sets in the random set literature (Definition 6.19 in molchanov2006book).\footnote{Note that beresteanumolinari2008 used an analogous definition of consistency, though in that setting all sets were compact by assumption, rendering the closure operation unnecessary.}

As is clear from both the Theoretical examples and the discussion in Section (ref), the same set estimator (say, maximum score in the Theoretical example 1) may fail to be consistent for some values of the DGP parameter ${\Greekmath 010B}$ making consistency a pointwise rather than uniform property. Namely, as follows from Theorem (ref) in Theoretical Example 1, the maximum score estimator will be consistent at ${\Greekmath 010B}$ in DGP that give $\mathcal{X}_0=\varnothing$ and inconsistent at ${\Greekmath 010B}$ in DGP that give $\mathcal{X}_0 \neq \varnothing$. Analogous conclusions apply to theoretical examples 2 and 3.

When a set estimator $\widehat{{A}}$ fails to be consistent at a given DGP parameter ${\Greekmath 010B}$, we will maintain the weaker assumption that $\overline{\widehat{{A}}}$ converges weakly to a limiting random set $\mathbf{A}({\Greekmath 010B})$ in the sense of the weak convergence in the random set theory. Weak convergence of random sets, as discussed in molchanov2006book, is characterized by the convergence of Choquet capacities: the sequence $T^{\overline{\widehat{A}}}(\mathcal{K})$ converges to $T^{\mathbf{A}({\Greekmath 010B})}(\mathcal{K})$ for all compact sets $\mathcal{K}$ in the Fell topology on the compact ${\mathcal{A}}$,\footnote{The Choquet capacity of a random set $\mathcal{B}$ is defined as $T^{\mathcal{B}}(\mathcal{K}) = \mathbb{P}(\mathcal{B} \cap \mathcal{K} \neq \emptyset)$ for all compact $K \subseteq \mathcal{A}$.} The assumption of such weak convergence is reasonable as this is exactly what we have generically under conditions in our general theorems (ref)-(ref), as summarized in Theorem (ref).

To motivate our second criterion we will be working with, recall that the original Hodges estimator (e.g., see hodges)\footnote{For important work on properties of the original Hodges estimator, see for example LeebPots2005,LeebPots2006,LeebPots2008.} illustrated the role of continuity of the estimator's limit when comparing its properties to MLE. We aim to use a similar principle here, though now within the partial identification paradigm, to evaluate another aspect of the behavior of the set estimators in our settings. Namely, we are concerned with potential dependence of weak limit of set estimators $\overline{\widehat{A}}$ on the underlying parameter ${\Greekmath 010B}$ of the DGP as the limit may change discontinuously analogously to the behavior of the identified set in our illustrative designs above.\footnote{E.g. in Illustrative Design 2, the identified set ${\mathcal A}_0$ depends on the DGP parameter ${\Greekmath 010B}$ as follows: $${\mathcal A}_0= \left( (-\infty,-1) \cdot \mathbf{1}({\Greekmath 010B}<-1) + \{-1\}\cdot \mathbf{1}({\Greekmath 010B}=-1) + (-1,0) \cdot \mathbf{1}({\Greekmath 010B}\in (-1,0)) + \{0\} \cdot \mathbf{1}({\Greekmath 010B}=0) + (0, +\infty) \cdot \mathbf{1}({\Greekmath 010B}>0)\right) \cap \mathcal{A}.$$ The set-values map from ${\Greekmath 010B}$ to $\mathcal{A}_0$ is lower semicontinuous but not upper semicontinuous.}

The discontinuity of the map ${\Greekmath 010B} \mapsto \mathcal{A}_0({\Greekmath 010B})$ that motivates our local robustness criterion below has a structural parallel in the moment inequality literature, where the identified set can also vary discontinuously with the underlying DGP. This is a central challenge in that literature. E.g., kaidomolinaristoye construct confidence sets that are valid uniformly over DGPs precisely to accommodate such abrupt changes in the identified set. The phenomenon is also present in our Illustrative Designs as in Design 1, case (1A) yields a singleton identified set, whereas cases (1B)--(1D) yield sets of positive dimension, with discontinuous transitions across configurations of ${\Greekmath 010B}$.

The mechanism driving this discontinuity is, however, different in the two settings. In moment inequality models, discontinuities arise when inequality restrictions move from slack to binding. In our setting, the discontinuity instead arises from particular realizations of $X$ which lead to the identified set having a dimensional collapse.

A further distinction concerns the implications for estimation. In moment inequality models, standard criterion-based estimators remain well behaved in the sense that the argmax set converges to the identified set (cht), so the main challenge is inference. In our setting, by contrast, the estimator itself can fail with the argmax of the sample maximum score objective converging to a random set rather than to the identified set, with the identified set lying on the boundary of each realization. This estimation feature, driven by the coarse characterization induced by the maximum score objective, motivates the RSQ estimator in Section (ref).

Returning to the notion of robustness in our setiing, drawing on the classical local alternatives framework of ibragimov, we study weak convergence of $\overline{\widehat{A}}$ under sequences of locally perturbed parameters ${\Greekmath 010B}_n(h) \to {\Greekmath 010B}$ (e.g.${\Greekmath 010B}_n(h) ={\Greekmath 010B}+h/\sqrt{n}$). In that setting, local robustness requires that the limiting distribution under such perturbations varies continuously in $h$. Extending this idea to partial identification requires care. First, $\overline{\widehat{A}}$ is set-valued, so its asymptotics are described by weak convergence of random closed sets in the sense of molchanov2006book. Second, the identified set $\mathcal{A}_0({\Greekmath 010B})$ itself may change discontinuously in ${\Greekmath 010B}$, so requiring the drifting limit to coincide with $\overline{\mathcal{A}_0({\Greekmath 010B})}$ would be too strong. Instead, we compare the asymptotic law under local perturbations to that obtained under nearby fixed parameters approached along the same direction.

Suppose that for each fixed $h$, the estimator $\overline{\widehat A}$ converges weakly under $P_{{\Greekmath 010B}_n(h)}$ to a random closed set $\mathbf A_h$, and that for each fixed parameter value $a$ near ${\Greekmath 010B}$, $\overline{\widehat A}$ converges weakly under $P_a$ to a random closed set $\mathbf A(a)$. These assumptions are supported by the estimators' behavior in Theorem (ref) and Theoretical examples 1-3.

definition[Local robustness] The estimator $\widehat A$ is locally robust at ${\Greekmath 010B}$ with respect to the set of directions $\mathcal{H}({\Greekmath 010B})$ if for every ${\Greekmath 0122}>0$, \[ \sup_{h \in \mathcal{H}({\Greekmath 010B}), \|h\|>0} \;\limsup_{a \in \mathcal{E}(h)} {\mathbb P}\!\left(d_H(\mathbf A_h,\mathbf A(a))>{\Greekmath 0122}\right)=0, \] where $\mathcal{E}(h):=\left\{a: a \to {\Greekmath 010B}, \; \; (a-{\Greekmath 010B})/\|a-{\Greekmath 010B}\| \to h/\|h\|\right\}.$

Since these limits may be random, the Hausdorff distance is interpreted as a distance between coupled realizations of random closed sets defined on a common probability space. This is natural in our setting because the local limits can be constructed from a common Gaussian vector.

Thus, local robustness requires that the asymptotic law under $n$- and $h$-dependent local perturbations agrees with the law obtained under nearby fixed parameters approached along the same direction. It can be easily shown that Definition (ref) implies the continuity of the law of the local limit: \[ \mathcal{L}(\mathbf A_h)\Rightarrow \mathcal{L}(\mathbf A({\Greekmath 010B})) \quad \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{as } \|h\|\downarrow 0, \; \; h \in \mathcal{H}({\Greekmath 010B}). \]

This notion of local robustness is closely related to the literature emphasizing the role of continuity of asymptotic distributions for robust inference, particularly the work of andrews-boundary and andrews-guggenberger-ema. A key insight of that literature is that discontinuities in limiting distributions can lead to failures of standard asymptotic approximations. Our definition extends this idea to set-valued estimators but instead of requiring continuity of a scalar or vector distribution, we require stability of the limiting random set.

The combined concepts of consistency and robustness allow us to analyze properties of set estimators. In particular, if an estimator is consistent at a given ${\Greekmath 010B} \in {\mathcal A},$ it may not be necessarily robust with respect to a certain range of sequences ${\Greekmath 010B}_n(h)$. At the same time, an estimator which is robust at a given point of the parameter space may not necessarily be consistent because robustness is the property of continuity of the limiting distribution. In cases where an estimator only converges weakly and not in probability, it cannot converge to a fixed (e.g. closure of the identified) set.

Equipped with these two criteria, we are now able to evaluate the existing classical estimators and propose a new one. As is clear from our discussion in Section (ref), the maximum score estimators in Theoretical examples 1-3 fail to be consistent at ${\Greekmath 010B}$ the result in $ \mathcal{X}_0 \neq \varnothing $, even though in Section (ref) it will be shown that they satisfy local robustness. No estimator in the current literature satisfies both conditions simultaneously. Section (ref) fills this gap.

Random Set Quantile Estimator

Our earlier analysis in Sections (ref) and (ref) left us with the findings that in discrete setting estimators may have fluctuating asymptotic behavior whenever the parameter ${\Greekmath 010B}$ in DGP results in $\mathcal{X}_0 \neq \varnothing$ rendering such estimators inconsistent at those ${\Greekmath 010B}$ in the DGP under which the Hausdorff distance between $\mathcal{A}_0$ and $\mathcal{A}^*$ is strictly positive (which we saw to be the case in our theoretical examples). Theorems (ref)-(ref) summarized in Theorem (ref) also gave us quite a clear structure to the weak limit of these estimators in cases of inconsistency. That structure is behind this section's proposal of a novel approach to constructing consistent and robust estimators,

Our approach introduces a new class of estimators grounded in the concept of random set theory, a framework that has been integral to Econometrics since the pioneering work of beresteanumolinari2008. Unlike previous applications of random set theory in econometrics, which centered on the Aumann expectation (beresteanumolinari2008, beresteanumolchanovmolinari2011, BERESTEANU201217), our approach exploits the quantile of a random set. The motivation is obvious from Theorem (ref) which shows that the maximum score estimator converges to a random set whose majority intersection recovers $\overline{A}_0$ and whose deterministic components in the limit have the same asymptotic probability. The quantile then is precisely the tool that extracts this majority intersection from data. This quantile approach can serve as a basis for constructing consistent estimators for identified sets not just in our discrete setups, but potentially in other partially identified models as well.

First, to give some insights on what our estimation approach will deliver, let's look at Illustrative Design 2 with compact $\mathcal{A}$ large enough to contain $[0,1]$ in its interior. When ${\Greekmath 010B}=1$ is the parameter value in the DGP, then as illustrated in Section (ref), $$\widehat{A} \stackrel{d}{\to} d \cdot [0,1) + (1-d) \cdot [1,+\infty) \cap \mathcal{A},$$ where $d$ is a Bernoulli random variable with parameter $\frac12$. For the closure we have $P(\overline{\widehat{A}} = [1,+\infty) \cap \mathcal{A} ) \to \frac{1}{2}$, $P(\overline{\widehat{A}} = \cdot [0,1] \cap \mathcal{A}) \to \frac{1}{2}$. If we look at the weak limit of $\overline{\widehat{A}}$, which is a random set, we can see that $1$ is the only point in the parameter space that happens to be in the closure of realizations of this random set in at least $100(1/2+\Delta) \%$ cases for $\Delta>0$. Indeed, $1$ belongs to the boundary of both $ [1,+\infty) \cap \mathcal{A} $ and $[0,1] $ and no other point simultaneously belongs to the boundary of both sets. More generally, the idea of our new estimation method relies on the properties of the distribution limit summarized in Theorems (ref) and Theorem (ref) for the maximum score specifically.

An elegant way to utilize that property is through the notion of the random set quantile. Namely, for estimator $\widehat{A}$ whose asymptotioc behavior is described by Theorem (ref) (either in (a) or (b)) we define our new estimator as the random set ${\Greekmath 011C}$-th quantile of the closure of ${\widehat{A}}$:

equation[equation omitted — 137 chars of source]

Following the definition of the ${\Greekmath 011C}$-th quantile of a random set in molchanov2006book (p.176),

equation[equation omitted — 177 chars of source]

for {\it coverage function } $p(u\,;\,\overline{\widehat{A}}) \equiv {\mathbb P}( u \in \overline{\widehat{A}}).$ Note that this notion from molchanov2006book also applies to vector settings.\footnote{Molchanov2017, p.36 defines quantiles of a random closed set via a family $\mathfrak{M}$ of compact subsets, allowing for a range of notions depending on this choice. In particular, different selections of $\mathfrak{M}$ (e.g., balls or more general sets) yield quantiles that reflect geometric or neighborhood-based features of the random set. In this paper, we adopt the canonical specialization $\mathfrak{M}$ as the family of singletons. This choice yields the finest possible resolution and is particularly well suited to our setting.}

An important consideration lies in the selection of suitable quantile indices ${\Greekmath 011C}$ that would ensure consistency of this estimator. To foreshadow our formal result, Theorem (ref) plus equal asymptotic probabilities of $A_1, \ldots, A_L$ will imply the choice of ${\Greekmath 011C}$-th quantile of $\overline{\widehat{A}}$ to be for ${\Greekmath 011C} \in (1/2, 1)$. To demonstrate why it works, take once again Illustrative Design 2 with the parameter value ${\Greekmath 010B}=1$ in the DGP. In this scenario, the ${\Greekmath 011C}$-th quantile of the closure $\overline{\widehat{A}}$ with ${\Greekmath 011C} \in (1/2,1)$, reduces to only $\{1\}$ for sufficiently large $n$, thereby converging in probability to the identified set.\footnote{It's worth noting that for ${\Greekmath 011C}<1/2$, the ${\Greekmath 011C}$-th quantile of $\overline{\widehat{A}}$ is $[0,+\infty) \cap \mathcal{A}$ for sufficiently large $n$. This set if $\overline{A^*}$. On the other hand, the $1/2$-th quantile of $\overline{\widehat{A}}$ for large enough $n$ can, with additional effort, be demonstrated to be the two-element set $\{0,1\}$, failing to asymptotically recover the identified set.} Case 2 in Theorem (ref) recovers $\overline{\mathcal{A}}_0$ for any ${\Greekmath 011C} \in (0,1)$. In particular, it does so for ${\Greekmath 011C} \in (1/2,1)$.

Given that the concept of a random set quantile may be unfamiliar to econometricians, it's helpful to draw parallels to voting rules for additional clarity. Let's explore this by examining the median of a random set within the framework of a simple discrete model. Suppose that we can generate infinitely many random samples $\{(x_i,y_i\}_{i=1}^n$ of a fixed size $n$. Each sample casts votes for any number of elements of $\mathcal{A}$ which maximize the objective (ref). Essentially, a given sample votes for its respective estimate $\widehat{A}$. After the completion of set $\widehat{A}$ to its closure $\overline{\widehat{A}}$, majority winners are selected -- namely, those elements in $\mathcal{A}$ that are voted for by at least 50% of the samples. The collection of those majority winners would give us $q_{.5} ( \overline{\widehat{A}})$. For any arbitrary index ${\Greekmath 011C} \in (0,1)$, the quantile $q_{{\Greekmath 011C}}(\overline{\widehat{A}})$ comprises those elements from $\overline{\mathcal{A}}$ that adhere to the "quota rule" with a threshold of ${\Greekmath 011C}$.

Consistency of the random set quantile estimator in Theoretical examples 1-3

We now formulate our result regarding the asymptotic behavior of the random quantile set estimator in the settings of interest. Note that Theorem (ref) immediately implies that $\overline{\mathcal{A}_0}\subseteq q_{{\Greekmath 011C}}(\overline{\widehat{A}})$ but to show that exact equality we will need additional properties in line with those stated in Theorem (ref).

theoremConsider an estimator $\widehat{A}$ that satisfies the conditions of Theorem (ref). In addition, suppose that (i) in Case 1 of Theorem (ref), ${\mathbb P}(d_\ell)=\frac{1}{L}$, $\ell=1, \ldots, L$, (ii) in Case 2 of Theorem (ref), $d_H(\overline{\mathcal{A}^*},\overline{\mathcal{A}_0})=0$. Let $\widehat{A}_{RSQ,{\Greekmath 011C}}$ be the ${\Greekmath 011C}$-th random set quantile of $\overline{\widehat{A}}$ as defined in ((ref)). Then for ${\Greekmath 011C}>1/2$, $$d_H(\widehat{A}_{RSQ,{\Greekmath 011C}},\overline{{\mathcal A}_0}) \stackrel{p}{\rightarrow} 0.$$ Hence, $\widehat{A}_{RSQ,{\Greekmath 011C}}$ is consistent for $\mathcal{A}_0$ at any parameter ${\Greekmath 010B} \in \mathcal{A}$ in DGP.

The key additional conditions in Theorem (ref) are, first, equal asymptotic probabilities of $A_{\ell}$ in the distribution limit of $\widehat{A}$ in Case 1, and (ii) the coarse $D^*_m$ resulting in $\mathcal{A}^*$ and $\mathcal{A}_0$ being different at most at the boundary. In Section (ref) these additional conditions were shown to be true hold for the maximum score estimator in the cross-sectional binary choice model (see Theorem (ref)). Even though formal results were not formulated for the other two theoretical examples, analogous results hold for them too with model-specific propositions replacing Proposition (ref) to establish those results, as discussed in more detail in the Appendix after the proof of Theorem (ref). Thus, we conclude that Theorem (ref) is applicable to all our theoretical examples.

The result of Theorem (ref) of $\widehat{A}_{RSQ,{\Greekmath 011C}}$ being consistent on the entire parameter space implies that there is no longer a need to differentiate between various “parameter regimes" ${\Greekmath 010B}$ in DGP when considering maximum score estimators in Theoretical examples 1-3.

Even though Theorem (ref) shows that consistency holds for any ${\Greekmath 011C} \in (1/2, 1)$, the choice of ${\Greekmath 011C}$ can be important in finite samples. as larger ${\Greekmath 011C}$ may produces a smaller tighter set, whereas a smaller ${\Greekmath 011C}$ close to $1/2$ may produce a larger set. We do not develop a theory for an optimal finite-sample choice of ${\Greekmath 011C}$. Instead, we recommend that researchers report RSQ estimates over a range of ${\Greekmath 011C}$ values, as this can provide insight into the estimator's behavior across the admissible range.

comment\begin{theorem} Consider a cross-sectional binary choice model under assumptions discussed in Section (ref) and within the discrete setup given in Section (ref). Suppose that the parameter space has a relative interior in the $(k-1)$-dimensional normalized space. Consequently, the maximum score estimator $\widehat{A}$ has asymptotic behavior as described in Theorem (ref) (case 1 or case 2). Let $\widehat{A}_{RSQ,{\Greekmath 011C}}$ be the ${\Greekmath 011C}$-th random set quantile of $\overline{\widehat{A}}$ as defined in ((ref)). Then for ${\Greekmath 011C}>1/2$, $$d_H(\widehat{A}_{RSQ,{\Greekmath 011C}},\overline{{\mathcal A}_0}) \stackrel{p}{\rightarrow} 0.$$ Hence, $\widehat{A}_{RSQ,{\Greekmath 011C}}$ is consistent for $\mathcal{A}_0$ at any parameter ${\Greekmath 010B} \in \mathcal{A}$ in DGP. \end{theorem}

Robustness of the random set quantile estimator in Theoretical examples 1-3

Our result of the consistency of the random set quantile estimator was general for any estimation approach that has asymptotic behavior described by Theorem (ref) with some additional conditions. For the analysis of robustness, we will focus on specific cases of our theoretical examples.

We start with the cross-sectional binary choice model and the random set quantile estimator built on the maximum score estimator $\widehat{A}$. The case of ${\Greekmath 010B}$ in DGP that results in $\mathcal{X}_0=\varnothing$ is quite straightforward from the robustness perspective. A more challenging case is when $\mathcal{X}_0 \neq \emptyset$, as this is when the identified set will change discontinuously.

We will take rate $1/\sqrt{n}$ in ${\Greekmath 010B}_n(h)={\Greekmath 010B}+h/\sqrt{n}$ as thia is the natural choice given that the estimated conditional probabilities $\widehat{P}(Y=1|x_t)$ at ambiguous points $x_t\in\mathcal{X}_0$ converge at exactly this rate. The drift therefore shifts the signal $P_{{\Greekmath 010B}_n(h)}(Y=1|x_t)-{\Greekmath 011C}$ by an amount of the same order as the estimation noise, placing the local perturbation precisely at the boundary between tie-resolving and tie-preserving. Slower drifts would overwhelm the sampling noise and trivially resolve all ties whereas faster drifts would be swamped by the sampling noise. This rate also connects to andrews-guggenberger-ema, whose drifting sequences are calibrated to the convergence rate of the relevant estimator. In our setting tha convergence rate for the sample proportion $\widehat{P}(Y=1|x_t)$ is $\sqrt{n}$.

Before formulating the local robustness result, we establish the following lemma.

lemmaConsider a cross-sectional binary choice model under assumptions discussed in Section (ref) and within the discrete setup given in Section (ref). Suppose that the parameter space has a relative interior in the $(k-1)$-dimensional normalized space. Additionally, assume that for any $x \in \mathcal{X}$ the conditional c.d.f. $F_{u\mid X}(\cdot\mid x)$ is differentiable at $0$, with $\left.\frac{\partial}{\partial v}F_{u|X}(v|x)\right|_{v=0} = f_{u|X}(0|x)\in(0,\infty)$. Then under ${\Greekmath 010B}_n(h)={\Greekmath 010B}+h/\sqrt{n}$, $$P_{{\Greekmath 010B}_n(h)}(Y=1|x_t)-{\Greekmath 010D} =f_{u| X}(0\mid x_t)\frac{x_t'h}{\sqrt{n}}+o(n^{-1/2})$$ for any $x_t \in \mathcal{X}_0$ with $\mathcal{X}_0=\{x \in \mathcal{X}: P(Y=1|x)={\Greekmath 010D}\}$ induced by this ${\Greekmath 010B}$. \begin{comment} Denoting $\mathcal{X}_0 = \{x_,\ldots, x_T\}$, we get \begin{equation}\sqrt{n}(\widehat{P}_{{\Greekmath 010B}_n(h)}(Y=1\mid x_1)-{\Greekmath 011C}, \ldots, \widehat{P}_{{\Greekmath 010B}_n(h)}(Y=1\mid x_T)-{\Greekmath 011C})' \stackrel{d}{\to} \mathcal{N}\left(\Pi({\Greekmath 010B})\,h, \Sigma\right).\end{equation} where $\Pi({\Greekmath 010B})$ is the $T \times K$ matrix $$\Pi({\Greekmath 010B})=\left(f_{{\Greekmath 010F}|\widetilde{X}}(0|\bar{x}_1)\bar{x}_1, \ldots,f_{{\Greekmath 010F}|\widetilde{X}}(0\,|\,\bar{x}_M)\bar{x}_M \right)^{\top}.$$ \end{comment}
theorem[Local robustness of RSQ estimator for Theoretical example 1] Consider a cross-sectional binary choice model under assumptions discussed in Section (ref) and within the discrete setup given in Section (ref). Suppose that the parameter space has a relative interior in the $(k-1)$-dimensional normalized space and suppose the additional condition on the c.d.f. differentiability from Lemma (ref). Let ${\Greekmath 010B}$ be the parameter value in the DGP. Consider sequence of the DGP parameters ${\Greekmath 010B}_n(h) = {\Greekmath 010B} + h/\sqrt{n}$, where $h \in \mathbb{R}^k$ is chosen in such a way that ${\Greekmath 010B}_n(h) \in \mathcal{A}$. Then the random set quantile $\widehat A_{RSQ,{\Greekmath 011C}}$ of the maximum score estimator is locally robust at ${\Greekmath 010B}$ in the sense of Definition (ref) with respect to $$\mathcal H({\Greekmath 010B})=\{h: x_t'h \neq 0 \quad \forall x_t \in \mathcal{X}_0\}$$ (and , of course, $h$ need to comply with normalization constraints to ensure that ${\Greekmath 010B}_n(h)$ belong to the normalized space, which we take as being true).

Theorem (ref) excludes directions $h\notin\mathcal{H}({\Greekmath 010B})$, i.e.\ those for which $x_t'h=0$ for some $x_t\in\mathcal{X}_0$. Let us explain why. For $h\notin\mathcal{H}({\Greekmath 010B})$, the tie at $x_t$ persists at first order along ${\Greekmath 010B}_n(h)$, and the sample maximum score objective fluctuates among the regions $\{\overline{A}_\ell:\ell\in\mathcal{L}(h)\}$, where $\mathcal{L}(h)\subseteq\{1,\ldots,L\}$ collects those indices consistent with the partially resolved signs at $\{x_t\in\mathcal{X}_0: x_t'h\neq 0\}$. Te RSQ weak limit $\mathbf{A}_h^{RSQ}$ is still deterministic as the quantile operation for ${\Greekmath 011C} \in (1/2,1)$ extracts the majority intersection of $ \{\overline{A_\ell}: \ell\in\mathcal{L}(h)\}$. However, for a nearby fixed parameter $a \in \mathcal{E}(h)$, the directional condition $(a-{\Greekmath 010B})/\|a-{\Greekmath 010B}\|\to h/\|h\|$ only requires the direction of approach to converge to $h/\|h\|$ asymptotically, and does not force $a$ to lie exactly on the ray through $h$. Consequently, one may construct sequences of such $a$ that include a smaller-order perturbation breaking the tie at $x_t\in \{x_s\in\mathcal{X}_0: x_s'h= 0\}$, while still satisfying the directional condition. Along such sequences, $\mathcal{X}_0(a)=\varnothing$ and $\mathbf{A}^{RSQ}(a)= \overline{A_{\ell(a)}}$ for some $\ell(a)$ determined by the sign of $x_t'(a-{\Greekmath 010B})$ at $x_t \in \{x_s\in\mathcal{X}_0: x_s'h= 0\}$, which differs from $\mathbf{A}_h^{RSQ}$ as the majority intersection will be a a subset of $\overline{A_{\ell(a)}}$. Hence local robustness fails for $h\notin\mathcal{H}({\Greekmath 010B})$, and the restriction to $\mathcal{H}({\Greekmath 010B})$ in Theorem (ref) is tight.

The local robustness argument extends to Theoretical examples 2 and 3 with only model-specific adaptations. In each case the key ingredient is a first-order expansion of the relevant probability differential under ${\Greekmath 010B}_n(h)$, which pins down the sign of the differential and hence the resolved regime, provided $h$ lies in the appropriate set $\mathcal{H}({\Greekmath 010B})$.

For the static panel binary choice model of Section (ref), the relevant differential is $\Delta P_{{\Greekmath 010B}_n(h)}(Y=1| \mathbf x_t)=P_{{\Greekmath 010B}_n(h)}(Y_2=1| \mathbf x_t) -P_{{\Greekmath 010B}_n(h)}(Y_1=1| \mathbf x_t)$. Under the assumption that $F_{{\Greekmath 010F}|\mathbf X}(\cdot| \mathbf x)$ is differentiable at $0$ with $\frac{dF_{{\Greekmath 010F}|\mathbf X}(v|x)}{dv}\big\vert_{v=0} =:f_{{\Greekmath 010F}|\mathbf X}(0|\mathbf x)\in(0,\infty)$, the same Taylor expansion as in Lemma (ref) gives $$\Delta P_{{\Greekmath 010B}_n(h)}(Y=1| \mathbf x_t) =f_{{\Greekmath 010F}|\mathbf X}(0|\mathbf x_t)\, \frac{(x_{t2}-x_{t1})'h}{\sqrt{n}}+o(n^{-1/2}) $$ for each $\mathbf x_t\in\mathcal{X}_0$. The sign is therefore $\mathrm{sgn}((x_{t2}-x_{t1})'h)$, and the natural set of tie-resolving directions is $\mathcal{H}({\Greekmath 010B})=\{h:(x_{t2}-x_{t1})'h\neq 0\ \forall \mathbf x_t\in\mathcal{X}_0\}$. With this adaptation, the proof of Theorem (ref) applies straightforwardly, replacing $x_t'h$ with $(x_{t2}-x_{t1})'h$ throughout, and the RSQ estimator built on the conditional maximum score is locally robust at every ${\Greekmath 010B}\in\mathcal{A}$ with respect to $\mathcal{H}({\Greekmath 010B})$.

For the multinomial choice model of Section (ref), the relevant differentials are $P_{{\Greekmath 010B}_n(h)}(Y=j|\mathbf{x}_t)- P_{{\Greekmath 010B}_n(h)}(Y=k|\mathbf{x}_t)$ for each tied pair $(j,k)$ at each $\mathbf{x}_t\in\mathcal{X}_0$. Under the assumption that the c.d.f.\ of ${\Greekmath 0122}_{ij}-{\Greekmath 0122}_{ik}$ conditional on $\mathbf{x}$ is differentiable at $0$ with $\frac{dF_{{\Greekmath 0122}_j-{\Greekmath 0122}_k|\mathbf{x}}(v|\mathbf{x})}{dv}\big\vert_{v=0} =:f_{jk}(0|\mathbf{x})\in(0,\infty)$, an analogous expansion gives: $$P_{{\Greekmath 010B}_n(h)}(Y=j|\mathbf{x}_t)- P_{{\Greekmath 010B}_n(h)}(Y=k|\mathbf{x}_t) =f_{jk}(0|\mathbf{x}_t)\, \frac{(x_{t,j}-x_{t,k})'h}{\sqrt{n}}+o(n^{-1/2}).$$ The tie-resolving set is therefore: $$\mathcal{H}({\Greekmath 010B})=\{h\in\mathbb{R}^K: (x_{t,j}-x_{t,k})'h\neq 0\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ for all }\mathbf{x}_t \in\mathcal{X}_0\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ and all tied pairs }(j,k)\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ at } \mathbf{x}_t\},$$ and the RSQ estimator built on the block maximum score is locally robust at every ${\Greekmath 010B}\in\mathcal{A}$ with respect to $\mathcal{H}({\Greekmath 010B})$. The proof again follows Theorem (ref) rather straightforwardly, replacing $x_t'h$ with $(x_{t,j}-x_{t,k})'h$ for each tied pair and invoking conditions (ref)--(ref) as verified for the block maximum score in the Appendix.

Feasible version of the random set quantile estimator

The estimator $\widehat{A}_{RSQ,{\Greekmath 011C}}$ as defined in (ref) is infeasible because the coverage function $p(u;\overline{\widehat{A}})={\mathbb P}(u\in\overline{\widehat{A}})$ is a population object depending on the unknown distribution of $(Y,X)$. To construct a feasible version, we approximate the coverage function using an $m$-out-of-$n$ bootstrap.

We describe this procedure for an RSQ estimator based on a general classical estimator in Section (ref). Given the original data sample $\{(y_i,x_i)\}_{i=1}^n$, draw with replacement $m$ observations to obtain bootstrap samples $(y^*_1,x^*_1),\ldots,(y^*_m,x^*_m)$ with $m\to\infty$ and $m=o(n)$. For each bootstrap replication $b=1,\ldots,B$, compute the estimate on the $b$-th bootstrap sample as $$\widehat{A}^{(b)} := \mathrm{Argmax}_{a\in\mathcal{A}} \sum_{x\in\mathcal{X}_{m,b}} w(x)\, s\!\big({\Greekmath 011E}(\widehat{F}^{(b)}(Y| X=x)),{\Greekmath 0120}(x,a)\big) \,\widehat{P}^{(b)}(X=x),$$ where $\mathcal{X}_{m,b}$ is the sample support of the $b$-th bootstrap sample, $\widehat{F}^{(b)}(\cdot| X=x)$ and $\widehat{P}^{(b)}(X=x)$ are the corresponding plug-in estimators of $F(Y| X=x)$ and $P(X=x)$.

Using the closures $\overline{\widehat{A}^{(1)}},\ldots, \overline{\widehat{A}^{(B)}}$ of the bootstrap maximum score estimates, construct the bootstrap approximation to the coverage function:

equation[equation omitted — 147 chars of source]

The feasible random set quantile estimator is then defined as

equation[equation omitted — 211 chars of source]

We now establish the consistency and robustness of the feasible RSQ estimator in Theoretical example 1.

theorem[Consistency of the feasible RSQ estimator] Consider an i.i.d. sample $\{(y_i,x_i)\}_{i=1}^n$ from the cross-sectional binary choice model under the assumptions of Section (ref) and the discrete setup of Section (ref). Suppose that the parameter space has a relative interior in the $(k-1)$-dimensional normalized space. Let $\{(y^*_i,x^*_i)\}_{i=1}^m$ be i.i.d.\ bootstrap replicates drawn with replacement from $\{(y_i,x_i)\}_{i=1}^n$. Suppose that $m/n+1/m=o(1)$ as $n\to\infty$. Then for ${\Greekmath 011C}\in(1/2,1)$: $$d_H\!\left(\widehat{A}^{boot}_{RSQ,{\Greekmath 011C}},\, \overline{\mathcal{A}_0}\right)\stackrel{p}{\longrightarrow}0.$$
theorem[Local robustness of the feasible RSQ estimator] Under the same conditions as Theorem (ref) and the additional c.d.f. differentiability assumption of Lemma (ref), when ${\Greekmath 010B}_n(h)={\Greekmath 010B}+h/\sqrt{n}$, we have that $\widehat{A}^{boot}_{RSQ,{\Greekmath 011C}}$ is locally robust at every ${\Greekmath 010B}\in\mathcal{A}$ with respect to the same $\mathcal{H}({\Greekmath 010B})$ as defined in Theorem (ref).

The key mechanism driving robustness of the feasible estimator is the interplay between the rates $n^{-1/2}$ and $m^{-1/2}$. Under the drifting sequence ${\Greekmath 010B}_n(h)$, the probability gap $P_{{\Greekmath 010B}_n(h)}(Y=1|x_t)-{\Greekmath 010D}=O(n^{-1/2})$ at $x_t\in\mathcal{X}_0$ is of the same order as the original sample fluctuation, so the bootstrap with rate $m^{-1/2}\gg n^{-1/2}$ randomizes over the sign at $x_t$ and reproduces the correct coverage function. Under a fixed nearby parameter $a\in\mathcal{E}(h)$, the gap $P_a(Y=1|x_t)-{\Greekmath 010D}$ is a fixed positive constant, so the bootstrap fluctuation $O_p(m^{-1/2})$ is asymptotically negligible relative to it, and the bootstrap concentrates on the single region $\overline{A_{\ell(h)}}$ determined by the sign of $x_t'h$. In both cases the feasible RSQ estimator recovers $\overline{A_{\ell(h)}}$, establishing robustness.

Thus, condition $m=o(n)$ ensures the rate separation in the sense that $m^{-1/2}\gg n^{-1/2}$ under ${\mathbb P}_{{\Greekmath 010B}_n(h)}$ but $m^{-1/2}\ll |P_a(Y=1|x_t)-{\Greekmath 010D}|$ under ${\mathbb P}_a$. It also illustrates why the $m$-out-of-$n$ bootstrap is the natural tool for this problem rather than the standard $n$-out-of-$n$ bootstrap, which would not achieve the correct rate separation.

For Theoretical examples 2 and 3 consistency and robustness of the feasible RSQ estimators based on their respective conditional maximum score and block maximum score estimators can be shown to be consistent and robust in an analogous way.

The RSQ estimator is now fully characterized, It is consistent and locally robust in both its infeasible and feasible forms, for all three classes of models. What remains is to understand how it compares to the maximum score estimator it is built upon. The comparison turns out to be clean as the two estimators share the same local robustness properties, and differ only in consistency. The next section establishes this formally.

Comparison of RSQ estimator to Maximum Score

Section (ref) implies that the RSQ estimator is consistent and locally robust across all three theoretical examples. We now establish that the maximum score estimator shares the robustness property (but, of course, not consistency).

Below we establish that the maximum score estimator is locally robust with respect to the tie-breaking directions with respect to the same $\mathcal{H}({\Greekmath 010B})$ as the RSQ estimator at any ${\Greekmath 010B}$ in the DGP. A powerful takeaway from this is the result that it is locally robust even at those ${\Greekmath 010B}$ in the DGP where it is inconsistent.

For simplicity, we start with the maximum score estimator in Theoretical example 1.

theorem[Local robustness of maximum score estimator] Consider a cross-sectional binary choice model under assumptions discussed in Section (ref) and within the discrete setup given in Section (ref). Suppose that the parameter space has a relative interior in the $(k-1)$-dimensional normalized space. Also suppose the additional condition on the density from Lemma (ref). Let ${\Greekmath 010B}$ be the parameter value in the DGP. Consider sequence of the DGP parameters ${\Greekmath 010B}_n(h) = {\Greekmath 010B} + h/\sqrt{n}$, where $h \in \mathbb{R}^k$ is chosen in such a way that ${\Greekmath 010B}_n(h) \in \mathcal{A}$. Then the maximum score estimator $\widehat A$ is locally robust at ${\Greekmath 010B}$ in the sense of Definition (ref) applied to the closure $\overline{\widehat{A}}$ and with respect to $$\mathcal H({\Greekmath 010B})=\{h: x_t'h \neq 0 \quad \forall x_t \in \mathcal{X}_0\}$$ (and, of course, $h$ need to comply with normalization constraints to ensure that ${\Greekmath 010B}_n(h)$ belong to the normalized space -- we take this property as given).

The maximum score estimator $\widehat{A}$ and the RSQ estimator $\widehat{A}_{RSQ,{\Greekmath 011C}}$ are both locally robust for the same reason: any $h\in\mathcal{H}({\Greekmath 010B})$ resolves all ties in $\mathcal{X}_0$, making both the drifting limit $\mathbf{A}_h$ and the fixed-parameter limit $\mathbf{A}(a)$ concentrate on the same region $\overline{A_{\ell(h)}}$.

Both theorems (ref) and (ref) restrict attention to $h\in\mathcal{H}({\Greekmath 010B})$, i.e.\ directions that resolve all ties in $\mathcal{X}_0$. For the RSQ estimator we have already discussed why for directions $h\notin\mathcal{H}({\Greekmath 010B})$ (meaning $x_t'h=0$ for some $x_t\in\mathcal{X}_0$), local robustness fails. For $h\notin\mathcal{H}({\Greekmath 010B})$ local robustness will also fail for the maximum score estimator, but for a different reason. When $x_t'h=0$, the tie at $x_t$ persists along ${\Greekmath 010B}_n(h)$ for all $n$ since $P_{{\Greekmath 010B}_n(h)}(Y=1|x_t)-{\Greekmath 011C}=f_{u|X}(0|x_t)x_t'h/ \sqrt{n}=0$. Hence $\mathcal{X}_0({\Greekmath 010B}_n(h))\ni x_t$ for all $n$, and by Theorem (ref) part 1 the drifting limit $\mathbf{A}_h$ of the maximum score estimator is a genuinely random set, supported on the subcollection $\{A_\ell:\ell\in \mathcal{L}(h)\}$ of regions consistent with the partially resolved signs at $\{x_s\in\mathcal{X}_0:x_s'h\neq 0\}$. By contrast, for any fixed $a\in\mathcal{E}(h)$ sufficiently close to ${\Greekmath 010B}$, the directional condition $(a-{\Greekmath 010B})/\|a-{\Greekmath 010B}\|\to h/\|h\|$ only requires the direction to converge to $h/\|h\|$ asymptotically and does not force $x_t'(a-{\Greekmath 010B})=0$. For any fixed $a$ with all $x_t'(a-{\Greekmath 010B})$ nonzero, we have $\mathcal{X}_0(a)=\varnothing$ and $\mathbf{A}(a)= \overline{A_{\ell(a)}}$ is deterministic. Since $\mathbf{A}_h$ is random while $\mathbf{A}(a)$ is deterministic, ${\mathbb P}(d_H(\mathbf{A}_h,\mathbf{A}(a))=0) <1$ and robustness fails. Thus, for maximum score robustness fails for $h \notin \mathcal{H}({\Greekmath 010B})$ because its drifting limit $\mathbf{A}_h$ is random while $\mathbf{A}(a)$ is deterministic. This is qualitatively different from the RSQ estimator, where robustness failed for $h \notin \mathcal{H}({\Greekmath 010B})$ because its drifting limit $\mathbf{A}_h^{RSQ}$ was strictly smaller than $\mathbf{A}^{RSQ}(a)$, with the gap reflecting residual ambiguity at those covariate values in $\mathcal{X}_0$ that the drift direction $h$ leaves unresolved.

In a nutshell, the distinction between the the maximum score and RSQ estimators is not local robustness but consistency. The RSQ estimator $\widehat{A}_{RSQ,{\Greekmath 011C}}$ is always consistent whereas the maximum score $\widehat{A}$ fails to be consistent at ${\Greekmath 010B}$ for which $\mathcal{X}_0\neq\varnothing$.

The local robustness of maximum score estimators in Theoretical examples 2 and 3 can be established in an analogous way with a suitable choice of $\mathcal{H}({\Greekmath 010B})$, as discussed previously in the context of the random set quantile estimator.

Let us look at the implications of Theorem (ref) for Illustrative Design 2. When ${\Greekmath 010B} \notin \{0,1\}$, the conditional maximum score is consistent and drifting does not impact its limit. When ${\Greekmath 010B} \in \{0,1\}$, the conditional maximum score is inconsistent for $\mathcal{A}_0$ at those ${\Greekmath 010B}$. If, for concreteness, ${\Greekmath 010B}=1$, then the maximum score estimator converges in distribution to $d \cdot [0,1) +(1-d)[1,+\infty)\cap \mathcal{A} $ with dummy variable $d$ taking values 0 and 1 with equal probabilities.

When $h > 0$ in the drifting ${\Greekmath 010B}_n(h)={\Greekmath 010B}+\frac{h}{\sqrt{n}}$, then the maximum score estimator puts a point mass of 1 on the set $[1,+\infty) \cap \mathcal{A}$. When $h<0$, then the maximum score estimator puts a point mass of 1 on the set $[0,1)$. Thus, as $h$ varies in $\mathcal{H}({\Greekmath 010B})$, the regions visit all possible deterministic sets in the distribution limit of the maximum score estimator.

Thus, as $h$ varies from $0$ to all the directions in $\mathcal{H}({\Greekmath 010B})$, the distribution of the limit random set varies from equal randomization between $[0,1)$ and $[1,+\infty)\cap \mathcal{A}$ (when $h=0$) to selecting a fixed set $[1,+\infty)\cap \mathcal{A}$ or $[0,1)$. Thus, this choice of the drifting sequence bridges these cases. As a result, even though the maximum score estimator is not consistent at that point, it is locally robust.

The theoretical analysis is now complete. The RSQ estimator dominates the maximum score estimator on our criteria as it adds consistency at no cost to robustness. The following section further illustrates these ersults in finite samples in a small scale simulation study and Section (ref) brings this conclusion to data.

Monte Carlo simulations and practical guide

In this section, we first conduct a Monte Carlo simulation study to evaluate the finite-sample performance of the feasible Random Set Quantile (RSQ) estimator and compare it against the standard maximum score estimator. The experimental setup adapts is built in Illustrative Design 1(A) and 1(B) introduced in Section (ref).

We then present a guide to practitioners on the choice of quantile indices in the finite sample and what to expect in practice.

Monte Carlo

Setting. We consider a binary choice model in the Illustrative Design 1: $Y_i = \mathbf{1}\{{\Greekmath 010B}_0 + X_{1i} + {\Greekmath 010B}_2 X_{2i} - u_i \geq 0\}, $ where $Med(u_i|X_{1i}.X_{2i})=0$, and the covariates $(X_{1i}, X_{2i})$ are drawn from a discrete support $\{(0,1), (1,0)\}$ with probabilities: $P((X_1,X_2) = (0,1)) = q > 0.5$, $P((X_1,X_2) = (1,0)) = 1-q$.

In our simulations, we set $q = 0.7$ and take $u_i \sim \mathcal{N}(0,1)$ (we will assume that the researcher only knows/is confident about the median independence restriction).

We analyze the following two settings: (i) Illustrative Design 1(A): ${\Greekmath 010B}_0 = -1$, ${\Greekmath 010B}_2 = 1$ (case of point identification); (ii) special case within Illustrative Design 1(B): ${\Greekmath 010B}_0 = -1$, ${\Greekmath 010B}_2 = 0.5$ (identified set has dimension 1).

\vskip 0.1in

Estimation procedure We evaluate the estimators over four sample sizes: $n \in \{10, 100, 1000, 10000\}$. The utility indices $a_0+1$ (obtained when $(X_1,X_2)=(1,0)$) and $a_0+a_2$ (obtained when $(X_1,X_2)=(0,1)$ naturally split the parameter space into 9 cells, which are depicted as $S_1$, \ldots, $S_9$ in Figure (ref): $S_1$ is obtained when $<;<$ (both indices are negative), $S_2$ is when $=;<$ (first index is zero, the second one is negative), $S_3$ is when $>;<$, $S_4$ is when $>;=$, $S_5$ is when $>;>$, $S_6$ is when $=;>$, $S_7$ is when $<;>$, $S_8$ is when $<;=$, $S_9$ is when $=;=$.

In Illustrative Design 1A we have $\mathcal{A}_0$ associated with $S_9$, $\mathcal{A}^*$ being the parameter space (associated with the union of all these 9 cells), The maximum score estimator asymptotically fluctuates among four sets $A_{(k,j)}$, $k,j=1,2,$ introduced in an earlier discussion of this design in Section (ref). Set $A_{(1,1)}$ is associated with $S_4 \cup S_5 \cup S_6 \cup S_9$, Set $A_{(1,1)}$ is associated with $S_4 \cup S_5 \cup S_6 \cup S_9$ and is closed. Set $A_{(2,1)}$ is associated with $S_2 \cup S_3$ with its closure being associated with $S_2 \cup S_3 \cup S_4 \cup S_9$. Set $A_{(1,2)}$ is associated with $S_7 \cup S_8$ with its closure being associated with $S_6 \cup S_7 \cup S_8 \cup S_9$. Set $A_{(2,2)}$ is associated with $S_1$ with its closure being associated with $S_1 \cup S_2 \cup S_8 \cup S_9$.

In Illustrative Design 1B with our choice of simulation parameters we have $\mathcal{A}_0$ associated with $S_2$ and $\mathcal{A}^*$ associated with $S_1\cup S_2 \cup S_3$. The maximum score estimator asymptotically fluctuates among two sets $A_1^{(B)}$ and $A_2^{(B)}$ (notation used in our earlier discussion of this design in Section (ref). Set $A_1^{(B)}$ is associated with $S_2\cup S_3$ and its closure is associated with $S_2\cup S_3 \cup \cup S_4 \cup S_9$. Set $A_2^{(B)}$ is associated with $S_1$ and its closure is associated with $S_1 \cup S_2\cup S_8 \cup S_9$ . We compare the following estimators.

enumerate• Standard Maximum Score estimator. We apply Manski's maximum score estimator to the full Monte Carlo sample. In both our designs in each sample the maximum score objective function is optimized at one of $A_{(k,j)}$, $k,j=1,2$, or their various unions but these four sets are the smallest sets which can be maximizers in the sample. In our case of Illustrative Design 1B, as mentioned earlier, according to our theory only $A_1^{(B)}=A_{(2,1)}$ and $A_2^{(B)}=A_{(2,2)}$ will survive asymptotically. • Random Set Quantile (RSQ) estimator: We construct the empirical coverage function using an $m$-out-of-$n$ bootstrap procedure. The RSQ estimator evaluates the bootstrap probability that the maximum-score closure contains each set $j$. The estimated set is formed by the cells whose bootstrap probability exceeds a strict majority (i.e., $0.5 + {\Greekmath 010E}$). Thus, the feasible RSQ estimator is represented as a subset of the finite 9-cell partition.To test robustness, we evaluate the estimator using four subsample size rules: $m \sim n^{0.1}, n^{0.25}, n^{0.5}$, and $n^{0.75}$. According to our theory, the RSQ estimator asymptotically recovers $S_9$ in Illustrative Design 1A and asymptotically recovers $S_2 \cup S_9$ (associated with the closure of the identified set) in our case of Illustrative Design 1B.

One of estimators' performance metrics is the identified-set containment frequency.

Simulation results.

figure[figure omitted — 495 chars of source]
figure[figure omitted — 478 chars of source]

Figures (ref) and (ref) display the empirical coverage function over the nine-cell partition for Illustrative Designs 1A and 1B, respectively. The figures are computed using a very large Monte Carlo sample ($n=10{,}000{,}000$), so the empirical conditional probabilities are already very close to their population counterparts. Despite the large sample size, the behavior of the feasible RSQ estimator is driven primarily by the bootstrap subsample size $m$. In the figures, we deliberately take $m$ relatively small compared to $n$ in order to illustrate the type of finite-sample behavior that may realistically arise in practice.

In Illustrative Design 1A, the true identified set corresponds to the singleton cell $S_9$, which receives empirical coverage essentially equal to one. Several neighboring cells receive coverage slightly above $0.5$, with the largest nuisance coverage approximately equal to $0.52$. Consequently, for any quantile level ${\Greekmath 011C}>0.52$, the feasible RSQ estimator selects only the true cell $S_9$. Thus, the estimator successfully recovers the point-identified set over a broad range of quantile levels.

The reason why neighboring nuisance cells still receive nontrivial coverage probabilities, and therefore why the feasible RSQ estimator produces strict supersets of $S_9$ for ${\Greekmath 011C}\in(0.5,0.52]$, is the variability induced by the $m$-out-of-$n$ bootstrap itself. Although the original Monte Carlo sample is extremely large, the bootstrap resamples involve substantially smaller effective sample sizes, so neighboring maximum-score regions continue to appear with positive probability in bootstrap draws.

In Illustrative Design 1B, the target set for the RSQ estimator is $S_2\cup S_9$, corresponding to the closure of the identified set. The empirical coverage of $S_2 \cup S_9$ is equal to one. All nuisance cells receive substantially smaller coverage probabilities, approximately up to $0.51$. Hence, for any \( {\Greekmath 011C}\in(0.51,1), \) the feasible RSQ estimator correctly recovers the closure of the identified set.

Tables (ref) and (ref) report empirical identified-set containment frequencies for Illustrative Designs 1A and 1B, respectively. For each estimator, the reported containment frequency is the proportion of Monte Carlo replications in which the estimated set contains the true identified set $\mathcal A_0$ (equivalently, contains all cells associated with $\mathcal A_0$ in the nine-cell partition).

The results reveal substantial differences between the standard maximum score estimator and the proposed RSQ procedure. In both designs, the standard maximum score estimator behaves according to predictions from our asymptotic theory, Namely. As the sample size increases, the containment probability of the standard maximum score estimator converges to $1/4$ in Design 1A\footnote{The maximum score estimator asymptotically fluctuates among four sets $A_{j,k}$, $j,k=1,2$, with equal probabilities $1/4$ and it is only one of those sets that contains the identified set corresponding to $S_9$.} and converges to $1/2$ in Design 1B\footnote{The maximum score estimator asymptotically fluctuates between two sets with equal probabilities $1/2$ and it is only one of those sets that contains the identified set corresponding to $S_2$.}

By contrast, the RSQ estimator exhibits strong containment performance across both designs. In both the point-identified and partially identified settings, the containment frequency approaches $1$ as the sample size increases. Moreover, the results are remarkably stable across all considered bootstrap subsample rules, ranging from $m\sim n^{0.1}$ to $m\sim n^{0.75}$.

Overall, the simulations demonstrate that the standard maximum score estimator does not reliably recover the identified set or its closure either in finite samples or asymptotically, whereas the proposed RSQ procedure delivers highly accurate containment performance and appears relatively insensitive to the precise choice of the bootstrap subsample scaling parameter $m$ within the range considered here.

table[table omitted — 622 chars of source]
table[table omitted — 620 chars of source]

Practical guide

The Monte Carlo results provide a useful benchmark for practical implementation of the feasible RSQ estimator. In our designs with a very small discrete support, the bootstrap random set stabilized sharply and the informative quantile region is easy to identify. In empirical applications, however, the support is typically richer and the bootstrap random set correspondingly more complex. In practice, we recommend trying two or three values of $m$ spanning the admissible range and checking that conclusions are stable across them. One practical lower bound is that each bootstrap resample should contain enough observations to estimate $\widehat{P}(Y=1|x)$ at every support point.

The theory guarantees consistency for any ${\Greekmath 011C}\in(1/2,1)$, but the right approach in a finite sample is not to fix ${\Greekmath 011C}$ in advance. Instead, compute $\widehat{A}_{RSQ,{\Greekmath 011C}}$ over a fine grid of ${\Greekmath 011C}$ values and examine the resulting sequence. Because the family is nested ($\widehat{A}_{RSQ,{\Greekmath 011C}_2}\subseteq \widehat{A}_{RSQ,{\Greekmath 011C}_1}$ for ${\Greekmath 011C}_2>{\Greekmath 011C}_1$) it contracts monotonically as ${\Greekmath 011C}$ increases, doing so in discrete jumps separated by flat intervals over which the estimated set is unchanged. We call these flat intervals plateaus. Each plateau corresponds to a set of parameter restrictions that survive in at least a ${\Greekmath 011C}$-fraction of bootstrap draws, and the width of the plateau interval measures how robustly those restrictions are present in the bootstrap distribution. Lower plateaus yield conservative sets close to the full-sample maximum score estimate; higher plateaus provide sharper nested refinements. We recommend reporting the full plateau sequence rather than a single estimate, as this gives a transparent picture of which restrictions are robust across what fraction of bootstrap draws.

It is also natural that the $m$-out-of-$n$ bootstrap RSQ estimator becomes empty for ${\Greekmath 011C}$ near 1. The emptiness for ${\Greekmath 011C}$ near 1 is driven by the fact that each bootstrap draw resamples all observations and, as a result, even support points whose full-sample functional values ${\Greekmath 011E}(\widehat{F}(y|x))$ are comfortably in $D_{m(x)}$ (in the Theoretical Example 1 with the median restriction this would mean their choice probabilities are comfortably away from $1/2$) may occasionally be resampled in a way that puts it in a different $D_{\widetilde{m}}$ (in the Theoretical Example 1 this means reversing the corresponding sample inequality). These occasional perturbations accumulate across bootstrap draws. This should not be interpreted as evidence of point identification.

commentThe Monte Carlo results provide a useful benchmark for practical implementation of the feasible RSQ estimator. In our designs with a very small discrete support, the bootstrap random set stabilized sharply and the informative quantile region is easy to identify. In empirical applications, however, the support is typically richer and the bootstrap random set correspondingly more complex. In practice, we recommend trying two or three values of $m$ and checking that conclusions are stable. One practical lower bound is that each bootstrap resample should contain enough observations to estimate $\widehat{P}(Y=1|x)$ at every support point. The theory guarantees consistency for any ${\Greekmath 011C}\in(1/2,1)$, but the right approach in a finite sample is not to fix ${\Greekmath 011C}$ in advance. Instead, compute $\widehat{A}_{RSQ,{\Greekmath 011C}}$ over a fine grid of ${\Greekmath 011C}$ values and examine the resulting sequence. Because the family is nested, it contracts in discrete jumps as ${\Greekmath 011C}$ increases, with plateaus (flat intervals) between jumps. Each plateau represents restrictions that are stable across a range of bootstrap proportions, and plateau width measures that stability. We recommend reporting the full plateau sequence rather than a single estimate. Lower plateaus yield conservative sets close to the full-sample maximum score estimate; higher plateaus provide sharper nested refinements. Reporting all stable plateaus gives a transparent picture of which restrictions are robust across what fraction of bootstrap draws. It is also natural that the $m$-out-of-$n$ bootstrap RSQ estimator becomes empty for ${\Greekmath 011C}$ near 1. The emptiness for ${\Greekmath 011C}$ near 1 is driven by the fact that each bootstrap draw resamples all observations and, as a result, even support points whose full-sample functional values ${\Greekmath 011E}(\widehat{F}(y|x))$ are comfortably in $D_{m(x)}$ (in the Theoretical Example 1 with the median restriction this would mean their choice probabilities are comfortably away from $1/2$) may occasionally be resampled in a way that puts it in a different $D_{\widetilde{m}}$ (in the Theoretical Example 1 this means reversing the corresponding sample inequality). These occasional perturbations accumulate across bootstrap draws. This should not be interpreted as evidence of point identification. In empirical work, the informative object is typically not a single quantile but the pattern of stabilization across quantiles. We therefore recommend a researcher finds the RSQ estimates for a range of quantile levels ${\Greekmath 011C} \in (1/2,1)$. By construction, the family is nested in the sense that $\widehat{\mathcal A}^{RSQ}({\Greekmath 011C}_2) \subseteq \widehat{\mathcal A}^{RSQ}({\Greekmath 011C}_1)$ for ${\Greekmath 011C}_2>{\Greekmath 011C}_1$. Lower quantiles close to $1/2$ yield conservative estimates likely close to the full-sample maximum score estimate. When we increase ${\Greekmath 011C}$ gradually we will observe sharpening of the estimate. We can identify intervals of ${\Greekmath 011C}$ over which the estimated set is unchanged.These intervals define stable quantile plateaus. Higher plateaus, when present, provide sharper nested refinements. This gives the researcher an idea of robust empirical cores obtained for different proportions of bootstrap samples. Reporting the sequence of stable plateaus gives a transparent ranking of restrictions by empirical stability.

Empirical application

The UK General Election in 2019 marked a significant development for the Labour Party as it faced a decline in its constituency victories. With a total of 650 constituencies, Labour secured only 202 seats during this electoral contest, a historic low both in terms of numerical count and proportion since the year 1935. Various media analyses pointed towards the aftermath of the Brexit referendum as a contributing factor to Labour's electoral setbacks. A telling example is a headline from The Guardian that succinctly captured the sentiment: "It was Brexit, not left-wing policies, that lost Labour this election."\footnote{\tiny \url{https://www.theguardian.com/politics/2020/jun/18/key-points-from-review-of-2019-labour-election-defeat}}

The central question for our analysis is whether the 2016 Brexit referendum Leave vote did indeed reduce Labour's probability of retaining constituencies it had held, and whether this effect operated differently in Labour-held versus non-Labour-held areas. This question is well suited to our methodology. First, the semiparametric binary choice framework imposes minimal distributional assumptions on voter preferences. Second, the discrete support of Brexit-related covariates generates exactly the partial identification structure our theory analyzes, making the RSQ estimator the appropriate tool.

Our units of observation are UK constituencies and the outcome variable $Y_i$ is an indicator for the same party winning a given constituency $i$ in both the 2015 and 2019 General Elections. This formulation focuses on seat retention rather than vote share, which is the economically relevant margin for a party's parliamentary presence. We construct five covariates: (1) $X_{i1}$ is the indicator for a constituency voting “Leave” in the Brexit referendum; (2) $X_{i2}$ is the indicator for the 2015 General Election result being within a 5% margin, capturing fiercely contested constituencies prior to Brexit; (3) $X_{i3}$ is an ordered variable (0, 1, 2) for mean constituency income growth from 2015 to 2019, taking value 0 for negative growth (8.15% of constituencies), 1 for positive but below-median growth (41.84%), and 2 for above-median growth (50%); (4) $X_{i4}$ is the indicator for Labour winning the constituency in 2015; (5) $X_{i5}$ is the interaction between $X_{i1}$ and $X_{i4}$, capturing the differential effect of the Leave vote on constituencies held by Labour in 2015.

Table (ref) summarizes these variables, and Table (ref) shows the joint distribution of Labour victories in 2015 and Leave votes. Of 650 constituencies, 519 (79.8%) retained the same winning party across elections. Of the 232 constituencies Labour held in 2015, 149 (64.2%) also voted Leave. This configuration that sits at the heart of Labour's 2019 difficulties and at the heart of our analysis.

table[table omitted — 738 chars of source]
table[table omitted — 330 chars of source]

We estimate the semiparametric binary choice model as in Theoretical Example 1, with the median restriction $Q_{1/2}(u_i|X_i)=0$ and covariate vector ${X}_i=(1,X_{i1},\ldots,X_{i5})'$ and normalize the coefficient on $X_{i2}$ to $-1$, reflecting the natural prior that a close 2015 election is negatively associated with retaining a seat in 2019. The normalized parameter vector is ${{\Greekmath 010B}}=({\Greekmath 010B}_0,{\Greekmath 010B}_1,-1,{\Greekmath 010B}_3, {\Greekmath 010B}_4,{\Greekmath 010B}_5)'$.

commentIn our data, there are 23 unique discrete realizations of ${X}$. For 4 elements in the sample support of ${X}$, we have $\widehat{P}(Y=1|{X})<0.5$, and for the remaining 17 elements, $\widehat{P}(Y=1|{X})>0.5$, and For 2 elements in the sample support of ${X}$, we have $\widehat{P}(Y=1|{X})=0.5$. The coarse characterization induced by the maximum score estimator replaces $D_3=\{0\}$ with $D_3^*=\mathbb{R}$ for the last two support points in $C_3$, effectively being indifferent about the sign of ${x}'{{\Greekmath 010B}}$ at those points. This is precisely the coarsening identified in Section (ref) but at the sample level: There are also $\widehat{P}(Y=1|{X})$ that are not too far from 0.5 and, thus, these points may potentially correspond to the point with $m(x)=3$ in the population. This will be dealt with in our $m$-out-of-$n$ bootstrap approach. The histogram of the estimated choice probabilities, along with their joint display against the empirical frequencies of the support points, is given in the Appendix.

The sample support of $X_i$ contains 23 distinct values. Of these, 4 have $\widehat{P}(Y=1|X=x)<0.5$, 17 have $\widehat{P}(Y=1|X=x)>0.5$, and 2 have $\widehat{P}(Y=1|X=x)=0.5$ exactly. These two support points are unambiguously in $\mathcal{X}_0$ at the sample level and the maximum score estimator places no restriction on $x'{\Greekmath 010B}$ at these points. However, the population $\mathcal{X}_0$ may be larger. Support points whose estimated probabilities are close to but not exactly $0.5$ in the full sample may belong to the population $\mathcal{X}_0$ but fail to land exactly on $0.5$ due to sampling variation. The histogram of estimated choice probabilities in Figure (ref) shows that several support points have $\widehat{P}$ in the range close to 0.6 and these are plausible candidates for population membership in $\mathcal{X}_0$. The $m$-out-of-$n$ bootstrap handles this naturally as by resampling observations of size $m$, support points near the $0.5$ boundary cross it in many bootstrap draws even when they do not cross it in the full sample, allowing the bootstrap coverage function to capture the identifying information these near-boundary points carry. The two support points where $\widehat{P}=0.5$ exactly generate the dominant randomization across $A_1,\ldots,A_L$ in the bootstrap, but the near-boundary points contribute to the progressive sharpening visible across the plateau sequence, and their influence is precisely why the RSQ estimate delivers restrictions that go beyond what the full-sample maximum score system implies.

Maximum score estimate. The maximum score estimate $\widehat{A}$ is the set of all ${\Greekmath 010B}=({\Greekmath 010B}_0,{\Greekmath 010B}_1,-1,{\Greekmath 010B}_3,{\Greekmath 010B}_4,{\Greekmath 010B}_5)'$ satisfying 17 weak inequality constraints (from support points where $\widehat{P}(Y=1|x)>0.5$) and 4 strict inequality constraints (from those where $\widehat{P}(Y=1|x)<0.5$). This system is feasible and the resulting set has a nonempty interior in the normalized five-dimensional space, confirming partial identification. For the full set of inequalities see the Appendix.

From the structure of $\widehat{A}$ we can already establish several directional conclusions. The constraints imply ${\Greekmath 010B}_1\geq 0$, that is, the Leave vote is weakly positive for non-Labour constituencies, consistent with Leave areas outside Labour's base tending to consolidate behind the Conservatives. The constraints also imply ${\Greekmath 010B}_1+{\Greekmath 010B}_5<0$, that is, the combined effect of the Leave vote in Labour-held constituencies is strictly negative, so the interaction term ${\Greekmath 010B}_5$ is sufficiently negative to overwhelm ${\Greekmath 010B}_1$. This confirms that the Leave vote damaged Labour's prospects specifically where Labour was defending seats. We can make even more of the interpretation of $\widehat{A}$ if we work with the latent utility indices: $U^*_{00}={\Greekmath 010B}_0-X_2+{\Greekmath 010B}_3 X_3$ (Remain, non-Labour 2015); $U^*_{01}={\Greekmath 010B}_0-X_2+{\Greekmath 010B}_3 X_3+{\Greekmath 010B}_4$ (Remain, Labour 2015); $U^*_{10}={\Greekmath 010B}_0+{\Greekmath 010B}_1-X_2+{\Greekmath 010B}_3 X_3$ (Leave, non-Labour 2015); $U^*_{11}={\Greekmath 010B}_0+{\Greekmath 010B}_1-X_2+{\Greekmath 010B}_3 X_3 +{\Greekmath 010B}_4+{\Greekmath 010B}_5$ (Leave, Labour 2015). The fact that $\widehat{{A}}\subseteq [0,1.5]\times[0,+\infty)\times[-0.5,0.5]\times [0,+\infty)\times(-\infty,0].$ implies $U^*_{01} \geq U^*_{00}, \quad U^*_{10} \geq U^*_{00}, \quad U^*_{01} \geq U^*_{11}.$ Since the system of inequalities for $\widehat{A}$ additionally implies that ${\Greekmath 010B}_4+{\Greekmath 010B}_5 <0$ and ${\Greekmath 010B}_0+{\Greekmath 010B}_1+{\Greekmath 010B}_4+{\Greekmath 010B}_5 \geq 0$, we also get $U^*_{11} < U^*_{10}, \quad U^*_{11} \geq U^*_{00}.$ The sign of ${\Greekmath 010B}_4-{\Greekmath 010B}_1$ is undetermined in maximum score meaning we cannot rank $U^*_{01}$ and $U^*_{10}$. The estimator thus documents the direction of the main effect but cannot resolve the full electoral ordering.

commentThe sample maximum score objective excludes the two support points in $\mathcal{X}_0$. Its possible largest values are attained at ${\Greekmath 010B}$ that satisfy 17 weak inequality constraints (${x}'{{\Greekmath 010B}} \geq 0$) and 4 strict inequality constraints (${x}'{{\Greekmath 010B}}<0$) from the remaining 21 support points, if the set of such ${\Greekmath 010B}$ is, of course, non-empty (if it is empty, then the sample maximum score objective will look for a set with violations of some of these inequalities). This system in our application does have a nonempty solution, and the resulting maximum score estimate $\widehat{{A}}$ has a nonempty interior in the normalized space (effectively, $\mathbb{R}^5$), confirming partial identification. Concretely, the maximum score estimate collects all ${{\Greekmath 010B}}=({\Greekmath 010B}_0,{\Greekmath 010B}_1,-1,{\Greekmath 010B}_3,{\Greekmath 010B}_4,{\Greekmath 010B}_5)'$ that satisfy $$({\Greekmath 010B}_0,{\Greekmath 010B}_1,{\Greekmath 010B}_3,{\Greekmath 010B}_4,{\Greekmath 010B}_5) \left( \begin{array}{ccccccccccccccccc} 1 & 1 & 1 & 1 & 1 & 1 & 1 & 1 & 1 & 1 & 1 & 1 & 1 & 1 & 1 & 1 & 1\\ 0 & 0 & 0 & 0 & 0 & 0 & 0 & 0 & 1 & 1 & 1 & 1 & 1 & 1 & 1 & 1 & 1\\ 0 & 0 & 1 & 1 & 2 & 2 & 1 & 2 & 0 & 0 & 1 & 1 & 2 & 2 & 0 & 1 & 2\\ 0 & 1 & 0 & 1 & 0 & 1 & 1 & 1 & 0 & 1 & 0 & 1 & 0 & 1 & 0 & 0 & 0\\ 0 & 0 & 0 & 0 & 0 & 0 & 0 & 0 & 0 & 1 & 0 & 1 & 0 & 1 & 0 & 0 & 0\\ \end{array} \right) \geq c_1 $$ with $c_1=(0,0,0,0,0,0,1,1,0,0,0,0,0,0,1,1,1)$ and $$ ({\Greekmath 010B}_0,{\Greekmath 010B}_1,{\Greekmath 010B}_3,{\Greekmath 010B}_4,{\Greekmath 010B}_5) \left( \begin{array}{cccc} 1 & 1 & 1 & 1 \\ 0 & 0 & 1 & 1 \\ 1 & 2 & 0 & 1 \\ 0 & 0 & 1 & 1 \\ 0 & 0 & 1 & 1 \\ \end{array} \right) < c_2, \qquad c_2=(1,1,1,1).$$ We can establish that $\widehat{{A}}\subseteq [0,1.5]\times[0,+\infty)\times[-0.5,0.5]\times [0,+\infty)\times(-\infty,0].$ The maximum score estimate finds ${\Greekmath 010B}_1$, which is the effect of Leave vote on non-Labour constituencies in 2015, to be non-negative. Working directly with the maximum-score system of inequalities, we can derive ${\Greekmath 010B}_1 +{\Greekmath 010B}_5 < 0$. Since ${\Greekmath 010B}_1 +{\Greekmath 010B}_5 $ is the effect of the Leave vote on Labour constituencies. this confirms that the Leave vote hurt chances of retaining seat in Labour-held constituencies. Thus, maximum score estimation partially supports the Guardian narrative: the Leave vote reduced the utility of retaining Labour constituencies (since $U^*_{11}<U^*_{01}$ and $U^*_{11}<U^*_{10}$), but the ranking between Leave & non-Labour and Remain&Labour constituencies ($U^*_{01}$ vs.\ $U^*_{10}$) remains unresolved. Thus, maximum score estimation does not allow us to create a full electoral ordering across constituency types and, thus, does not allow us to conclude that Labour does worse in Leave areas across all relevant comparisons. However, not being able to compare $U^*_{01}$ vs.\ $U^*_{10}$ does not erase the main message that Leave was especially bad for Labour-held Leave constituencies.

We next examine whether random set quantile estimation gives sharper empirical conclusions.

Random set quantile estimate. We implement the feasible RSQ estimator using $m=550$ and $B=5,000$ bootstrap draws. We implement the feasible RSQ estimator using $B=5{,}000$ bootstrap draws and $m=550$. The choice of $m$ is guided by, first, $n=650$ being small by the standards of applications where asymptotic rate separation between $m^{-1/2}$ and $n^{-1/2}$ is easily achieved, and, second, by the identifying variation in our application coming substantially from support points whose estimated probabilities are close to but not exactly $0.5$. As discussed above, these near-boundary points are plausible members of the population $\mathcal{X}_0$ and their contribution to the RSQ estimate depends on the bootstrap randomizing their sign reliably across draws. This choice of $m$ ensures that near-boundary points cross the $0.5$ threshold in a meaningful fraction of bootstrap draws.

As ${\Greekmath 011C}$ increases from $1/2$ toward 1, the RSQ estimate contracts monotonically by construction. This contraction occurs in discrete steps, producing a sequence of stable plateaus (intervals of ${\Greekmath 011C}$ over which the estimated set is constant). These plateaus are empirically informative with each representing a set of parameter restrictions that survive at least a ${\Greekmath 011C}$-fraction of bootstrap draws. Higher plateaus corresponding to more stringent restrictions. In our application the plateau structure is quite robust to the choice of $m$ with other $m$ resulting only in minor changes in the exact starting and end points of the corresponding ${\Greekmath 011C}$-intervals. This indicates that the plateau structure is not driven by a particular bootstrap tuning choice.

Table (ref) summarizes the three main plateaus. The RSQ estimate is empty for ${\Greekmath 011C} \geq 0.87$, reflecting the finite-sample bootstrap variability discussed in Section (ref).

table[table omitted — 830 chars of source]

The first plateau for ${\Greekmath 011C}_1\in(0.5,0.62)$ adds two inequality restrictions beyond the maximum score region. Both restrictions arise from the two support points in $\mathcal{X}_0$, which are the same points the maximum score estimator ignores.

commentThe first plateau occurs for ${\Greekmath 011C}_1\in(0.5,0.62)$, where $\widehat A_{RSQ,{\Greekmath 011C}_1} = \widehat A \cap \left\{ {\Greekmath 010B}_0\leq 1,\; {\Greekmath 010B}_0+{\Greekmath 010B}_1+2{\Greekmath 010B}_3+{\Greekmath 010B}_4+{\Greekmath 010B}_5\leq 1 \right\}$. This already sharpens the maximum score estimate. Both restrictions arise from support points that contribute zero to the sample maximum score criterion and are therefore ignored by the classical estimator. The RSQ estimator nevertheless recovers stable information from these observations, since the corresponding inequalities appear systematically across bootstrap resamples. The second plateau appears for ${\Greekmath 011C}\in[0.62,0.69)$, where $\widehat A_{RSQ,{\Greekmath 011C}_2} = \widehat A \cap \left\{ {\Greekmath 010B}_0\leq 1,\; {\Greekmath 010B}_0+2{\Greekmath 010B}_3=1, \; {\Greekmath 010B}_1+{\Greekmath 010B}_4+{\Greekmath 010B}_5\leq 0 \right\}$. Relative to the first plateau, the estimator now selects a lower-dimensional subset of the parameter space, strengthening the earlier restrictions through an equality that persists more uniformly under resampling.

The second plateau for ${\Greekmath 011C}\in[0.62,0.69)$ strengthens the first in two ways. It adds the equality ${\Greekmath 010B}_0+2{\Greekmath 010B}_3=1$ and the inequality ${\Greekmath 010B}_1+{\Greekmath 010B}_4+{\Greekmath 010B}_5\leq 0$. The inequality is already substantively important as it says that the net effect of the Leave vote on Labour-held constituencies is weakly negative, meaning that the strongly negative interaction term ${\Greekmath 010B}_5$ is large enough in absolute value to at least cancel the combined positive contributions of the general Leave effect ${\Greekmath 010B}_1$ and the Labour incumbency advantage ${\Greekmath 010B}_4$. Labour-Leave constituencies were therefore no better placed than the baseline in terms of seat retention, and possibly worse.

The third plateau for ${\Greekmath 011C}\in(0.71,0.839]$ sharpens the second plateau's inequality to an equality, simultaneously pinning down ${\Greekmath 010B}_0 = 1$ ${\Greekmath 010B}_3 = 0$, and ${\Greekmath 010B}_1 + {\Greekmath 010B}_4 + {\Greekmath 010B}_5 = 0$ in the maximum score set. The progression from the second to the third plateau is the key inferential step since what was established as a weakly negative net effect in the second plateau is now identified as exactly zero. The Labour incumbency advantage and the general Leave-area effect together were precisely canceled by the negative interaction. Labour-Leave constituencies were pushed exactly to the decision boundary, a 50% probability of changing hands, rather than being merely disadvantaged relative to their prior safe-seat status. This distinction between a negative effect and an exact cancellation is economically meaningful as it says that the Brexit backlash erased Labour's defensive advantage without additional margin, leaving these seats as knife-edge contests rather than inevitable losses.

We report the third plateau as our robust empirical core for the following reasons. Its ${\Greekmath 011C}$ -interval is the widest of the three plateaus and it fully incorporates the randomization of the signs of inequalities corresponding to the two support points where $\widehat{P}(Y=1|x)=0.5$ (that is, the radnomization of ${\Greekmath 010B}_0\geq (\leq ) 1$,\; ${\Greekmath 010B}_0+{\Greekmath 010B}_1+2{\Greekmath 010B}_3+{\Greekmath 010B}_4+{\Greekmath 010B}_5\geq (\leq) 1$; inequalities $\leq $ are weak due to us considering the closure). The restriction ${\Greekmath 010B}_0 = 1$ pins down the baseline utility of retaining a constituency for a non-Leave, non-Labour, non-close-margin seat with negative income growth. The restriction ${\Greekmath 010B}_3 = 0$ indicates that income growth over the 2015--2019 period had no differential effect on seat retention conditional on the other covariates, consistent with the view that the Brexit signal dominated economic conditions as an electoral driver during this period. The restriction ${\Greekmath 010B}_1 + {\Greekmath 010B}_4 + {\Greekmath 010B}_5 = 0$ has been discussed above.

commentThe most informative plateau occurs for ${\Greekmath 011C}_3\in(0.71,0.839]$, where $\widehat A_{RSQ,{\Greekmath 011C}_3} = \widehat A \cap \left\{ {\Greekmath 010B}_0=1,\; {\Greekmath 010B}_3=0,\; {\Greekmath 010B}_1+{\Greekmath 010B}_4+{\Greekmath 010B}_5=0 \right\}.$ We interpret this plateau as the robust empirical core of the estimator. While the RSQ estimate sharpens further over a narrow range of larger values of ${\Greekmath 011C}$ before becoming empty, the plateau $\widehat A_{RSQ,{\Greekmath 011C}_3}$ persists over a wide interval. For this reason, it provides the clearest summary of the high-coverage empirical content of the bootstrap distribution.\footnote{ Geometrically, earlier plateaus retain sign cells that appear only intermittently throughout the resampling distribution, whereas the plateau for ${\Greekmath 011C}_3$ removes these unstable directions and isolates the sign pattern reproduced very consistently across bootstrap draws.} For larger values of ${\Greekmath 011C}$, the RSQ estimate sharpens briefly and then becomes empty for ${\Greekmath 011C}\geq 0.87$. We use $m=500$ and $B=5,000$ bootstrap draws. As ${\Greekmath 011C} \in (1/2, 1)$ progresses, we notice the changes in the RSQ estimate. There are two plateaus in the RSQ estimates that stand out. The first occurs for ${\Greekmath 011C}_1 \in (0.5, 0.62)$ where $\widehat{A}_{RSQ,{\Greekmath 011C}_1}=\widehat{A}\cap \{{\Greekmath 010B}_0\leq 1, {\Greekmath 010B}_0 +{\Greekmath 010B}_1 +2{\Greekmath 010B}_3 +{\Greekmath 010B}_4+{\Greekmath 010B}_5 \leq 1\}$. This already sharpens the maximum score estimate as the two support points that has zero input to the maximum score objective function (thus, not providing any constraints then). Then next plateau for ${\Greekmath 011C}_2 \in [0.62, 0.69)$ gives $\widehat{A}_{RSQ,{\Greekmath 011C}_2}=\widehat{A}_{RSQ,{\Greekmath 011C}_1} \cap \{{\Greekmath 010B}_0 +2{\Greekmath 010B}_3 =1\}$. The plateau for ${\Greekmath 011C}_3 \in (0.71,0.839]$ gives $\widehat{A}_{RSQ,{\Greekmath 011C}_3}=\widehat{A} \cap \{{\Greekmath 010B}_0 = 1, {\Greekmath 010B}_3=0, {\Greekmath 010B}_1 +{\Greekmath 010B}_4+{\Greekmath 010B}_5 = 0\}$. The RSQ estimate is empty set when ${\Greekmath 011C}\geq 0.87$. (ref). \begin{table}[!ht] \begin{center} \caption{RSQ estimates. $\widehat{A}$ denotes the full sample maximum score estimator. } \vskip 0.05in \begin{tabular}{l|l} \hline\hline ${\Greekmath 011C}$ range & RSQ estimate\\ \hline $(0.5, 0.7)$ & $\widehat{A}_{RSQ,{\Greekmath 011C}_1}=\widehat{A}\cap \{{\Greekmath 010B}_0\leq 1, {\Greekmath 010B}_0 +{\Greekmath 010B}_1 +2{\Greekmath 010B}_3 +{\Greekmath 010B}_4+{\Greekmath 010B}_5 \leq 1\}$ \\ $[0.7, 0.75)$ & $\widehat{A}_{RSQ,{\Greekmath 011C}_2}=\widehat{A}\cap \{{\Greekmath 010B}_0 = 1, {\Greekmath 010B}_1 +2{\Greekmath 010B}_3 +{\Greekmath 010B}_4+{\Greekmath 010B}_5 \leq 0\}$ \\ $[0.75, 0.79)$ & $\widehat{A}_{RSQ,{\Greekmath 011C}_3}=\widehat{A}\cap \{{\Greekmath 010B}_0 = 1, {\Greekmath 010B}_3=0, {\Greekmath 010B}_1 +{\Greekmath 010B}_4+{\Greekmath 010B}_5 \leq 0\}$ \\ $[0.79, 0.915)$ & $\widehat{A}_{RSQ,{\Greekmath 011C}_4}=\widehat{A}\cap \{{\Greekmath 010B}_0 = 1, {\Greekmath 010B}_3=0, {\Greekmath 010B}_1 +{\Greekmath 010B}_4+{\Greekmath 010B}_5 = 0\}$ \\ $[0.915, 0.935)$ & $\widehat{A}_{RSQ,{\Greekmath 011C}_4}=\widehat{A}\cap \{{\Greekmath 010B}_0 = 1, {\Greekmath 010B}_3=0, {\Greekmath 010B}_1 +{\Greekmath 010B}_4+{\Greekmath 010B}_5 = 0\}$ \\ \hline \end{tabular} \end{center} \end{table} \begin{align*} \widehat{{A}}_1&=\widehat{{A}} \cap\{{\Greekmath 010B}_0\geq 1,\; {\Greekmath 010B}_0+{\Greekmath 010B}_1+2{\Greekmath 010B}_3+{\Greekmath 010B}_4+{\Greekmath 010B}_5\geq 1\}\\ \widehat{{A}}_2&=\widehat{{A}} \cap\{{\Greekmath 010B}_0\geq 1,\; {\Greekmath 010B}_0+{\Greekmath 010B}_1+2{\Greekmath 010B}_3+{\Greekmath 010B}_4+{\Greekmath 010B}_5<1\}\\ \widehat{{A}}_3&=\widehat{{A}} \cap\{{\Greekmath 010B}_0<1,\; {\Greekmath 010B}_0+{\Greekmath 010B}_1+2{\Greekmath 010B}_3+{\Greekmath 010B}_4+{\Greekmath 010B}_5\geq 1\}\\ \widehat{{A}}_4&=\widehat{{A}} \cap\{{\Greekmath 010B}_0<1,\; {\Greekmath 010B}_0+{\Greekmath 010B}_1+2{\Greekmath 010B}_3+{\Greekmath 010B}_4+{\Greekmath 010B}_5<1\}. \end{align*} Not surprisingly, these regions correspond to the four sign combinations at the two sample ambiguous support points. As a consequence, the RSQ estimate for ${\Greekmath 011C}>1/2$ is the intersection of any three of the four closed sets $\overline{\widehat{{A}}_\ell}$, $\ell=1,\ldots,4$. All such intersections coincide, yielding the RSQ estimate: $$\widehat{{A}}_{RSQ,{\Greekmath 011C}}=\left\{ ({\Greekmath 010B}_0,{\Greekmath 010B}_1,-1,{\Greekmath 010B}_3,{\Greekmath 010B}_4,{\Greekmath 010B}_5)': {\Greekmath 010B}_0=1,\;{\Greekmath 010B}_3=0,\; {\Greekmath 010B}_1+{\Greekmath 010B}_4+{\Greekmath 010B}_5=0\right\}.$$
commentOverall, we see a striking sharpening of the maximum score estimate as we go from a five-dimensional region to a two-dimensional affine subspace ib $\widehat A_{RSQ,{\Greekmath 011C}_3}$. The three binding constraints in $\widehat A_{RSQ,{\Greekmath 011C}_3}$ have clear interpretations. First, ${\Greekmath 010B}_0=1$ pins down the baseline utility of retaining a constituency. Second, ${\Greekmath 010B}_3=0$ indicates that income growth had no significant differential effect on seat retention in our model. Third, and most substantively, ${\Greekmath 010B}_1+{\Greekmath 010B}_4+ {\Greekmath 010B}_5=0$ says imply that while Labour’s incumbency (${\Greekmath 010B}_4$) and possibly the general effect of Leave (${\Greekmath 010B}_1$) would on their own give Labour an advantage in retaining seats, the negative interaction term (${\Greekmath 010B}_5$) captures a strong backlash specifically in Labour-held Leave constituencies that exactly wipes out this advantage and makes those seats knife-edge contests right at the tipping point (50% probability of winning). With these seats being pushed to the decision boundary (that is, a 50–50 outcome), we conclude that the disproportionate erosion of the Labour incumbency effect by the Leave vote turned previously relatively safe seats into highly competitive ones. In terms of the utility indices, on $\widehat A_{RSQ,{\Greekmath 011C}_3}$ we have a refinement relative to the maximum score findings by being able to conclude now that $U_{11}^*=U_{00}^*$, but we still cannot rank $U^*_{01}$ and $U^*_{10}$.

To make the progression concrete, recall the latent utility indices. At the second plateau, the inequality ${\Greekmath 010B}_1 + {\Greekmath 010B}_4 + {\Greekmath 010B}_5 \leq 0$ implies $U^*_{11} \leq U^*_{00}$ which means that Labour-Leave constituencies were no better placed than the non-Labour Remain baseline, and possibly worse. The maximum score estimate had already established some orderings, so by the second plateau we know Labour-Leave constituencies were at the bottom of the retention ordering relative to all other types. What remained open was whether they sat strictly below the baseline or exactly at it. The third plateau resolves this since ${\Greekmath 010B}_1 + {\Greekmath 010B}_4 + {\Greekmath 010B}_5 = 0$ implies $U^*_{11} = U^*_{00}$ meaning that Labour-Leave constituencies are at the same retention probability as the non-Labour Remain baseline. Thus, Labour-Leave seats were not driven below the baseline into near-certain loss, but rather stripped of their incumbency protection and left as genuine $50-50$ contests.

This finding sharpens and qualifies the Guardian narrative that Brexit lost Labour the election. The RSQ estimate does not support the view that the Leave vote uniformly damaged Labour. Rather, the damage was targeted and it cancelled Labour's incumbency advantage in the specific constituencies where Labour held seats and the local majority had voted Leave, a configuration affecting 149 of Labour's 232 held seats. For non-Labour Leave constituencies, the Leave vote if anything supported continuity. The mechanism was not broad electoral punishment of Labour but a precise erosion of the defensive advantage that should have protected Labour's existing seats in Leave areas.

What the RSQ estimator cannot resolve, even at the third plateau, is the ranking between $U^*_{01}$ (Remain Labour) and $U^*_{10}$ as the sign of ${\Greekmath 010B}_1 - {\Greekmath 010B}_4$ remains undetermined. In others words, we cannot say semiparametrically whether a Remain-voting Labour seat was more or less likely to be retained than a Leave-voting non-Labour seat. The data do not contain sufficient variation to resolve this comparison without distributional assumptions and the RSQ estimator correctly reports it as the boundary of what is identified.

Comparison with Probit. Table (ref) reports Probit estimates as a benchmark. Since the ideal maximum score inequality system is feasible, standard results imply Probit estimates lie within $\widehat{A}$ (after normalisation), which is confirmed. The Probit point estimate gives $\widehat{{\Greekmath 010B}}_1 + \widehat{{\Greekmath 010B}}_4 + \widehat{{\Greekmath 010B}}_5 = -0.231$, and a Wald test of $H_0: {\Greekmath 010B}_1 + {\Greekmath 010B}_4 + {\Greekmath 010B}_5 = 0$ fails to reject at conventional levels ( p-value is $0.132$). This is consistent with our RSQ finding. The Probit coefficient on income growth $\widehat{{\Greekmath 010B}}_3 = 0.015$ with the standard error $0.098$ is similarly consistent with the RSQ restriction ${\Greekmath 010B}_3=0$ at the third plateau.

{

table[table omitted — 691 chars of source]

}

commentAs a benchmark, Table (ref) reports Probit estimates. Since the maximum score system of inequalities is feasible and, thus, implying perfect separation of constituencies with $\widehat{P}(Y=1|{X}) >1/2$ from those with $\widehat{P}(Y=1|{X})<1/2$, then the standard results imply that Probit estimates (after re-normalization to enforce ${\Greekmath 010B}_2=-1$) must lie within the maximum score estimate $\widehat{{A}}$ (subject to normalisation), which is confirmed in Table (ref). \begin{table}[!ht] \caption{Estimates from the Probit model} \begin{center} \begin{tabular}{lc} \hline Variable & Probit \\ \hline Indicator for “Leave" vote & 0.6997 \\ \quad & (0.1516) \\ Indicator for GE 2015 being within 5% margin & -0.8948 \\ \quad & (0.1820) \\ 2015-to-2019 mean income growth category & 0.0145 \\ \quad & (0.0981) \\ Indicator if Labour won in 2015 & 1.3352 \\ \quad & (0.3007) \\ Indicator if Labour won in 2015 $\times$ Indicator for “Leave" vote & -2.2655 \\ \quad & (0.0911) \\ constant & 0.6486 \\ \quad & (0.1608) \\ \hline \end{tabular} \end{center} { Robust standard errors in parentheses.} \end{table} The Probit estimates give $\widehat{{\Greekmath 010B}}_1+\widehat{{\Greekmath 010B}}_4+\widehat{{\Greekmath 010B}}_5 =0.700+1.335-2.266=-0.231$, suggesting a negative net effect of the Leave vote on Labour retention. This is in contrast with the Random Set Quantile estimation, which found a zero net effect (or LAbour-Leave constituencies pushed to the decision boundary). Here, in probit, we find Labour-Leave constituencies being pushed on the other side of the decision boundary, indicating at the first glance a stronger negative impact than the one given by the RSQ estimation. However, A Wald test of $H_0:{\Greekmath 010B}_1+{\Greekmath 010B}_4+{\Greekmath 010B}_5=0$ fails to reject at conventional levels ($p$-value $=0.132$), which is consistent with the RSQ point estimate of zero for this combination. However, the RSQ estimator delivers this conclusion directly as a point estimate of the identified set, without requiring a separate hypothesis test or parametric distributional assumptions. This illustrates the practical advantage of the RSQ approach as it recovers sharp information about the parameter that the maximum score estimator leaves partially identified, and does so semiparametrically, while the Probit model can only approximate this conclusion through a hypothesis test that relies on parametric assumptions that may be misspecified.

Conclusion

In this paper we study semiparametric discrete choice models when covariates are discrete, violating the assumption of the continuity of distribution of regressors required to establish point identification. While the parameters of the model are generally only partially identified, depending on the true value of the parameters of the data generating process, the identified set can be a singleton or a non-singleton set.

commentWe focus on the question if this behavior of the identified set is accurately captured by a given estimator. We propose two criteria for evaluation of a given estimator. Sharpness of an estimator for a given true parameter of the data generating is the property where it converges in probability to the identified set corresponding to the true parameter value. Robustness of an estimator is the property of continuity of the distribution limit of an estimator with respect to local changes in the parameter of the data generating process. We explore existing estimators for discrete choice models including the maximum score estimator and the closed-form two-step estimator. We find that the maximum score estimator is not sharp everywhere on the parameter space, however, it is robust everywhere. In contrast, the closed form estimator is sharp only in parts of the parameter space where the model is point identified and is not robust at those points. We also consider direct combinations of these existing estimators using a “switching” device as well a combination of their objective functions. The latter combination idea allows us to construct estimators which are sharp, and this is not true for the latter combination idea. Both combination estimators fail to be robust.

As the main contribution in the paper, we propose a novel class of estimators based on the concept of a quantile of a random set. It uses the random set output by existing estimators, such as the maximum score estimator, and produces an estimator which is both consistent and robust on the entire parameter space. We illustrate the performance of our estimator and compare it with existing estimators first through a small scale simulation study and then by analyzing the impact of the Brexit referendum vote on outcomes of the 2019 UK General Elections. Our estimator both performs better and provides more meaningful results than the alternatives. We also show that our framework extends to other important settings including the semiparametric multinomial choice model as well as static panel data models with discrete outcomes.

Our results suggest that care is required when applying standard semiparametric estimators in discrete environments, and that methods explicitly designed for partial identification can yield substantial gains. More broadly, the paper highlights the usefulness of random set methods in econometrics and suggests that similar approaches may be fruitful in other models with set-valued limits.

Several directions for future research remain. First, extending the framework to allow for high-dimensional covariates may broaden its empirical applicability. Second, developing inference procedures that fully exploit the structure of random set quantiles is an important next step. Finally, applying these methods in substantive empirical contexts may shed further light on the practical importance of the issues we document.