Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
168,562 characters · 24 sections · 7 citation commands
{ Semiparametric discrete choice models are widely used in a variety of practical applications. While these models are point identified in the presence of continuous covariates, they can become partially identified when covariates are discrete. In this paper we find that classical estimators, including the maximum score estimator, (\citeasnoun{manski1975}), loose their attractive statistical properties without point identification. First of all, they are not {\it sharp} with the estimator converging to an {\em outer region} of the identified set, (\citeasnoun{komarova2013}), and in many discrete designs it weakly converges to a {\em random set}. Second, they are not {\it robust,} with their distribution limit discontinuously changing with respect to the parameters of the model. We propose a novel class of estimators based on the concept of a quantile of a random set, which we show to be both sharp and robust. We demonstrate that our approach extends from cross-sectional settings to classical static and dynamic discrete panel data models. }
{\bf Key Words:} Maximum Score, Identified Set, Robustness, Random Set Quantile, Panel Data Discrete Choice.
{\bf JEL Codes:} C14, C21, C25.
Early work on modeling discrete choices of economic agents with random utility over those choices led to the creation of an important class of semiparametric econometric models where discrete outcome variable is a parametric function of observed regressors, but also depends on the random noise whose distribution is unknown and is not specified parametrically. Pioneering this line of work, \citeasnoun{manski1975} advanced a series of foundational papers, elucidating conditions under which the parameters of such models attain point identification. Specifically, by stipulating the conditional median of the random noise to be zero, coupled with, most notably, a continuity assumption regarding the distribution of at least one regressor, he established the precise conditions sufficient for parameter point identification. Discrete-only data, however, is commonplace in a range of policy-relevant applications and in this paper, we focus on the settings of the discrete choice model where regressors have only a discrete support. This means that the continuity assumption used in the prior literature is violated which may lead to the failure of point identification, the case analyzed in \citeasnoun{komarova2013}. We demonstrate that depending on the data generating process parameters, our model can be either point- or partially-identified, resulting in the identified set being either a singleton or a non-singleton.
Within the discrete-only regressors framework, we propose a new class of estimators for discrete choice models that can be applied in either partially or point identified scenarios. These estimators are based on the concept of a quantile of a random set from the random set theory (e.g., see \citeasnoun{molchanov2006book}). We show that they successfully address the changing structure of the identified set over the parameter space from singleton to non-singleton sets. We compare our new estimator with the maximum score estimator and the closed-form estimator based on ideas of \citeasnoun{ichimura1994}, which were both originally intended for point identification scenarios. Additionally, we contrast our new estimator with two novel variants of the maximum score estimator and the closed-form estimator, designed to combine the strengths of both approaches.
We use two criteria for evaluation and comparison of estimation methodologies. First, {\em sharpness} is the property of an estimator to approximate the identified set with high probability for a given value of parameter of the data generating process. The rationale behind employing this criterion is clear-cut; our objective is to ascertain the truth in the probability limit. Second, {\em robustness} is the property that the distribution of an estimator varies continuously in the small neighborhood of a particular parameter value of the data generating process. This property is important as the lack of the local continuity in distribution required for robustness can also result in failure of bootstrap and other resampling-based methods for inference. This resembles the case of the parameter on the boundary of the parameter space leading to discontinuity in the distribution limit and the failure of bootstrap, e.g., characterized by \citeasnoun{andrews-boundary}.
We find that the maximum score and closed-form estimators are not sharp over the whole parameter space. E.g., the maximum score estimator can converge in probability to a singleton or a set, but it can also weakly converge to a random set. However, the maximum score estimator is robust whereas the closed-form estimator is not. Our two novel variants of the maximum score estimator and the closed-form estimator, as mentioned earlier, aim to amalgamate the strengths of both approaches, potentially yielding improved properties. One variant combines the objective functions of the the maximum score the closed-form estimator. The other variant combines the set estimates produced by those estimators. The former combination idea leads to sharp estimation, whereas the second one does not. At the same time, both of these combination estimators fail to be robust.
In contrast, our new estimator rooted in the concept of a quantile of a random set proves to be both sharp and robust, establishing its superiority over all previously discussed estimation methodologies in our discrete-only setting. To elaborate, this estimator incorporates the classic maximum score estimator and defines a new estimator based on the quantile of the random set produced by the maximum score method. It can be implemented in practice using a simple re-sampling procedure. We posit that the estimators of this kind may be appealing in other partially identified models with discrete variables.
Our paper draws on rich prior literature on identification and inferences for point and partially identified semiparametric discrete choice models. This includes classic work on the maximum score estimator in \citeasnoun{manski1975} and \citeasnoun{manski1985} with the distribution theory developed in \citeasnoun{kimpollard} and the smoothed version of the estimator \citeasnoun{horowitz-sms} yielding asymptotically normally distributed parameters. This model was further studied in setting with heteroskedastic errors in \citeasnoun{khan:13}, and in setting with discrete regressors in \citeasnoun{komarova2013}. \citeasnoun{ichimura1994} and \citeasnoun{ahnetal2018} consider an alternative set of procedures focusing on the implication for the choice probability to be close to 0.5 for particular values of covariates. On the side of inference, important relevant results are discussed in \citeasnoun{abrevayahuang2005} showing the failure of the standard bootstrap for the classic maximum score estimator, and later \citeasnoun{cattaneojansson2020} developed a valid bootstrap procedure for the maximum score. \citeasnoun{rosenura2022} considered finite sample properties of a related estimator that is based on moment inequalities.
The rest of the paper is organized as follows. In Section (ref) we present the general model, underlying technical assumptions and its simple version which we showcase the main technical points throughout the paper. We demonstrate that, depending on the values of parameter of the data generating process, the model can either be point- or partially identified. We create a toolkit for analysis of estimators which can capture this behavior of the identified set and discuss in detail two main criteria which we use for evaluation in the paper -- {\it sharpness} and {\it robustness} introduced earlier.
In Section (ref) we consider the maximum score estimator \citeasnoun{manski1975} and the closed-form estimator which stems from the ideas in \citeasnoun{ichimura1994} and establish their properties from the perspectives of our sharpness and robustness criteria. As mentioned earlier, neither of these is sharp over the whole parameter space but the patterns of their non-sharpness are very different. To elaborate, the maximum score estimator is sharp on the set of full Lebesgue measure in the parameter space whereas the closed-form estimator is sharp on the set of a zero Lebesgue measure in the parameter space. On the second criterion, the maximum score estimator is robust over the whole parameter space whereas the closed-form estimator is not robust everywhere except for the points in that set of Lebesgue measure zero where it is sharp. This section also introduces two novel combination estimation approaches designed to build on the respective strengths of the maximum score and closed-form estimators and establishes that one of them sharp over the whole parameter space, the other one is not, and neither of them is robust.
In Section (ref) we develop our main novel estimation methodology based on the concept of a quantile of a random set to construct a new class of estimators which are shown to be both sharp and robust everywhere on the parameter space.
In Section (ref) we (a) provide an empirical illustration highlighting results from all the estimation approaches discussed in the paper; and (b) propose an idea for a feasible implementation of a random set quantile and demonstrate the performance of our random set quantile estimator in a series of Monte Carlo simulations.
Section (ref) demonstrates how ideas from our simple cross-sectional model extend to single-index models as well as static and dynamic panel data models enabling inference when those models are partially identified.
Finally, Section (ref) concludes by summarizing results and discussing areas for future research, and the appendix collects all proofs of the main theorems.
In this section, we illustrate our central ideas by reviewing the fundamental concept of identification and introducing set estimation relevant to the binary choice models under consideration. This will serve as a foundation for formally defining the concepts of sharpness and robustness in the context of our estimators.
To introduce the formal concepts we turn to the simplest cross-sectional form of the semiparametric discrete choice model
where the outcome random variable $Y$ with range $\{0,\,1\}$ is generated from the vector of covariates $\widetilde{X}$ and an unobserved disturbance ${\Greekmath 010F}.$ To avoid confusion, we use capital letters throughout the paper for random variables and small letters to denote their specific realizations. We impose the following assumption on the structure of the model and the data generating process:
Assumptions (ref) (i), (ii) and (iv) are classic assumptions of manski1975 and following literature while Assumption (ref) (iii) extends the original framework to the case where regressors $\widetilde{X}$ are discrete under additional smoothness assumption on the distribution of the unobserved shock ${\Greekmath 010F}$. This, in particular, guarantees that the conditional c.d.f. $F_{{\Greekmath 010F}|\widetilde{X}}(\cdot\,|\,\widetilde{X}=\widetilde{x})$ is strictly increasing around 0, which has implication on the characterization of the identified set.
Under Assumption (ref), parameter $\widetilde{{\Greekmath 010B}}_0$ characterizing the data generating process can, in general, be no longer identified. In fact, as can be deduced from \citeasnoun{manski1975} and \citeasnoun{manski1985}, the identified set for $\widetilde{{\Greekmath 010B}}_0$ -- which we will denote as ${\mathcal A}_0 $ -- is characterized as set of the values of $\widetilde{{\Greekmath 010B}}$ such that
and that satisfy the normalization restriction in Assumption (ref)(i). Let ${\mathcal A}$ and ${\mathcal A}_{0}$ denote the projections of the parameter space $\widetilde{{\mathcal A}}$ and the identified set $\widetilde{{\mathcal A}}_0$, respectively, onto the ${\mathbb R}^{K-1}$ excluding the normalized $k$-th component in the original sets.
To explain our arguments, we will use two illustrative designs.
Normality of the error distribution on these designs is not important and is simply used for convenience to fully characterize the heteroskedastic distribution of the unobserved term.
It is easy to see from our characterization of the identified set\footnote{See Appendix (ref) for technical characterization of the identified set.} above that in our Illustrative Design 1 the following holds (suppose the parameter space ${\mathcal A}$ is large enough to contain values $0$ and $-1$):
For simplicity, the case of ${\Greekmath 010B}_0>0$ will be our main case for partial identification in the context of this design (case ${\Greekmath 010B}_0<-1$ results in the identified set ${\mathcal A}_0=(-\infty,-1) \cap {\mathcal A}$, and the case ${\Greekmath 010B}_0 \in (-1,0)$ gives the identified set ${\mathcal A}_0=(-1,0)$.
In Illustrative Design 2 the vector of regressors has only 2 support points and the linear index $\widetilde{X}'\widetilde{{\Greekmath 010B}}$ forms two hyperplanes for those support points: ${\Greekmath 010B}_1+{\Greekmath 010B}_2=0$ and ${\Greekmath 010B}_1+1=0.$ If the parameter vector takes a value on a given hyperplane, then the corresponding probability ${\mathbb P}(Y=1\,|\,\widetilde{X})=\frac12.$ If the parameter vector is not on particular hyperplane, we can only identify which side of the hyperplane it is on, but not its value. Using the convention for the sign function that $\operatorname*{sign}(0)=0,$ we can express the identified set as $$ {\mathcal A}_0=\left\{ ({\Greekmath 010B}_1,{\Greekmath 010B}_2)\,:\, \operatorname*{sign}({\Greekmath 010B}_1+{\Greekmath 010B}_2)= \operatorname*{sign}({\Greekmath 010B}_{1,0}+{\Greekmath 010B}_{2,0}),\, \operatorname*{sign}({\Greekmath 010B}_1+1)=\operatorname*{sign}({\Greekmath 010B}_{1,0}+1) \right\} \cap {\mathcal A}. $$ Suppose the parameter space ${\mathcal A}$ is large enough to contain $(-1,1)$, then there are 4 following cases for the identified set ${\mathcal A}_0:$
As evident, Illustrative Design 2 exhibits greater complexity, wherein even in partially identified cases, certain conditional probabilities of choice may remain equal to $1/2$. Consequently, this leads to identified sets characterized by an empty interior in $\mathbb{R}^{K-1}$.
For a substantial portion of our findings, drawing upon the insights derived from Illustrative Design 1 will prove sufficient. Thus, this design will serve as our primary illustrative tool throughout the paper. However, a select few results will be more effectively explained by employing Illustrative Design 2.
One a more general note, note that ((ref)) defines a convex polyhedron. Its dimension $K-1-r_0$ depends on the rank $r_0$ of the system of equations $\widetilde{x}'\widetilde{{\Greekmath 010B}}=0$ in (((ref))) (here we use the normalization of the parameter $\widetilde{{\Greekmath 010B}}$ to obtain $K-1$ as the largest dimension). When $K-1-r_0>0$, the boundary of this polyhedron when considered in ${\mathbb R}^{K-1-r_0}$ is excluded due to strict inequalities. The identified set ${\mathcal A}_0$ is obtained as the intersection of this polyhedron with ${\mathcal A}$.
Having demonstrated that in general our model is partially identified, we now turn to the characterization of estimators for the identified set ${\mathcal A}_0.$ Let $\widehat{{\mathcal A}}$ denote an estimator for the identified set obtained from sample $(y_i,\,\widetilde{x}_i)^n_{i=1}$.
In this definition ${\mathcal A}^*$ is a probability limit of estimator $\widehat{{\mathcal A}}$ treated as a random set and coincides with the definition of convergence in probability from the random set theory (see molchanov2006book, Definition 6.19). This definition allows us to introduce the concept of {\it sharpness} of the set estimator $\widehat{{\mathcal A}}.$
Definition (ref) characterizes sharpness by two essential properties of a set estimator. First, a sharp estimator needs to converge in probability to a deterministic set. Second, the limit set has to be within Hausdorff distance 0 from the identified set. As stated in the definition, sharpness is a pointwise property and the same set estimator may not necessarily remain sharp for different values of parameter ${\Greekmath 010B}_0$ of the DGP.
The next step in characterizing behavior of set estimators $\widehat{{\mathcal A}}$ involves examining scenarios where a particular estimator lacks sharpness. Since sharpness is linked to both the convergence of an estimator in probability and the probability limit being the identified set, the absence of sharpness typically implies a lack of convergence in probability. In the absence of sharpness, we maintain the assumption that for parameter ${\Greekmath 010B}_0$ of the DGP the estimator $\widehat{{\mathcal A}}$ still converges weakly to a random set ${\bf A}({\Greekmath 010B}_0).$ Weak convergence of random sets as discussed in \citeasnoun{molchanov2006book} can be characterized as the convergence of the sequence of Choquet capacities $T^{\widehat{{\mathcal A}}}_{{\Greekmath 010B}_0}(\mathcal{K})$ of random sets $\widehat{{\mathcal A}}$ to the capacity of the limit random set $T^{{\bf A}({\Greekmath 010B}_0)}_{{\Greekmath 010B}_0}(\mathcal{K})$ for the parameter ${\Greekmath 010B}_0 \in {\mathcal A}$ of the DGP for all sets $\mathcal{K}$ in the Fell topology on ${\mathcal A}.$\footnote{Recall that Choquet capacity of the random set ${\mathcal B}$ is defined as $T^{\mathcal B}(\mathcal{K})={\mathbb P} ({\mathcal B} \cap \mathcal{K}).$}
To motivate our second criterion, recall that the original Hodges estimator (e.g., see hodges) \footnote{For important work on properties of the original Hodges estimator, see for example \citeasnoun{LeebPots2005},\citeasnoun{LeebPots2006},\citeasnoun{LeebPots2008}.} illustrated the role of continuity of the estimator's limit when comparing its properties to MLE. We aim to use a similar principle here, though now in the partially identified setting, to evaluate another aspect of the behavior of the set estimators in our settings. Namely, we are concerned with potential dependence of weak limit of set estimators $\widehat{{\mathcal A}}$ on the underlying parameter ${\Greekmath 010B}_0$ of the DGP as the limit may change discontinuously analogously to the behavior of the identified set in our illustrative designs above.
Drawing from the analysis for the point identified case in \citeasnoun{ibragimov}, we consider weak convergence of set estimators $\widehat{{\mathcal A}}$ under locally varying parameters of the DGP ((ref)). Let ${\Greekmath 010B}(n,t;{\Greekmath 010B}_0) \in {\mathcal A}$ be a sequence of parameter values indexed by constant $t \in {\mathbb R}$ and sample size $n$ converging to ${\Greekmath 010B}_0$ as $n \rightarrow \infty.$ Parameter ${\Greekmath 010B}(n,t;{\Greekmath 010B}_0)$ for each fixed $t \in {\mathbb R}$ and $n$ determines the probability distribution ${\mathbb P}_{{\Greekmath 010B}(n,t;{\Greekmath 010B}_0)}$ for the random vector of observable variables $(Y,\widetilde{X}).$ We then construct the Choquet capacity of the set estimator $\widehat{{\mathcal A}}$ induced by this probability distribution as $T^{\widehat{{\mathcal A}}}_{{\Greekmath 010B}(n,t;{\Greekmath 010B}_0)}(\mathcal{K})={\mathbb P}_{{\Greekmath 010B}(n,t;{\Greekmath 010B}_0)}\left(\widehat{{\mathcal A}} \cap \mathcal{K} \right).$
Thus, robustness is the property of the continuity of the distribution of the estimator's limit at a given point ${\Greekmath 010B}_0$ with respect to some sequences of parameters of the DGP converging to that point.
The combined concepts of sharpness and robustness allow us to analyze properties of set estimators. In particular, if an estimator is sharp at a given ${\Greekmath 010B}_0 \in {\mathcal A},$ then it is not necessarily robust with respect to a certain range of sequences ${\Greekmath 010B}(n,t;{\Greekmath 010B}_0),$ as we show in the next section for specific instances of set estimators. At the same time, an estimator which is robust at a given point of the parameter space is not necessarily sharp because robusness is the property of continuity of the limiting distribution. In case where an estimator only converges in distribution and not in probability, it cannot converge to a fixed (identified) set.
An important question in our analysis will be the selection of the drifting rates for the sequences ${\Greekmath 010B}(n,t;{\Greekmath 010B}_0)$. Since the the shape of the identified set is determined by conditional choice probabilities $p(\widetilde{x})=P(Y=1\,|\,\widetilde{X}=\widetilde{x})$ (through the threshold crossing condition ((ref))), the question of robustness of the estimators can be approached by analyzing the limits of estimators along the sequences ${\Greekmath 010B}(n,t;{\Greekmath 010B}_0)$ that induce the sequences of conditional probabilities $p^{(n)}(\widetilde{x})$ varying with the sample size $n$ and approach $p(\widetilde{x})$ in the limit, with cases when some $p(\widetilde{x})$ are equal to $\frac12$ being particularly interesting as they result either in the case of point identification or that of partial identification but with the identified set having an empty interior in the parameter space.
Due to the discrete nature of our regressors, the conditional probability $P(Y=1\,|\,\widetilde{X}=\widetilde{x})$ can be estimated as the ratio of two sample means, which has $\sqrt{n}$-rate of convergence. This can be translated into the parameter sequences with the behaviour $\|{\Greekmath 010B}(n,t;{\Greekmath 010B}_0)- {\Greekmath 010B}_0\|=O(1/\sqrt{n})$. Such sequences will play the central role in our subsequent robustness analysis.
In this section we analyze two classic estimators -- maximum score and the closed-form estimator -- for the semiparametric discrete choice models originally developed for the point identified models. We study their sharpness and robutsness properties. We also propose and analyze two novel estimators that build on the bases of those two classic estimators and are designed to combine their respective strengths.
In order to emphasize the key technical aspects, we will provide a detailed proof of the results in this section specifically for one of our illustrative designs. Results for generic discrete settings will also be presented, accompanied by an overview of how their proofs align with the cases in illustrative designs and the additional algebraic considerations they entail.
Maximum score estimator was among the first and most influential estimation methods proposed in \citeasnoun{manski1975} for semiparametric discrete choice models. The key technical assumption behind the maximum score estimator is the median condition ((ref)) (iv) leading to the point identification of the model coefficients in case where a regressor has continuous distribution and has a non-trivial impact on the index. We show that in the discrete regressors setting (thus, with point identification generally failing), the maximum score estimator is robust everywhere on the parameter space, but not uniformly sharp.
The objective function for the maximum score estimator constructed from the sample $(y_i,\widetilde{x}_i)^n_{i=1}$ takes the form
The estimator yields the maximum to this objective function: $ \widehat{\cal A}_{ms}= \arg \max_{{\Greekmath 010B} \in {\cal A}} MS_n({\Greekmath 010B}).$ An alternative way to define the objective function proposed in \citeasnoun{komarova2013} is:
where $\hat p(\widetilde{x})$ and $ \widehat{P}(\widetilde{X}=\widetilde{x})$ denote, respectively, a conditional choice probability estimator (estimated as a group average) and the sample frequency of $\widetilde{X}$ at a given support point $\widetilde{x}$.
In general discrete regressors settings, $\widehat{\cal A}_{ms}$ will be a non-singleton set in the parameter space ${\mathcal A}$. The question we are interested in what this set converges to as the sample size, denote by $n$, gets arbitrarily large.
We start our analysis by considering the infeasible version of the maximum score estimator as the maximizer of $ MS_{n,INF}({\Greekmath 010B}) =\sum_{\widetilde{x} \in \mathcal{X}} \left(p(\widetilde{x})-1/2 \right) \operatorname*{sign}(\widetilde{x}' \widetilde{{\Greekmath 010B}}) \widehat{P}(\widetilde{X}=\widetilde{x}), $ which takes conditional choice probabilities $p(\widetilde{x})$ as known. This infeasible maximum score estimator is denoted as $\widehat{\cal A}_{ms, INF}$.
By arguments completely analogous to \citeasnoun{komarova2013}, we can show that $\widehat{\cal A}_{ms, INF}$ can be a strict superset of ${\mathcal A}_{0}$ -- namely, this may happen whenever there are $\widetilde{x}$ in the discrete support for which $p(\widetilde{x})=1/2$ as these cases are simply ignored by $MS_{n,INF}({\Greekmath 010B})$, as evident from the representation above. As shown in the Appendix, for all values of parameter $\widetilde{{\Greekmath 010B}}_0$ of the DGP, the infeasible estimator $\widehat{\cal A}_{ms, INF}$ converges in probability to the maximizer ${\cal A}_{ms}$ of the population objective function
To give more details, ${\mathcal A}_{ms}$ is defined as the set of $\widetilde{{\Greekmath 010B}} \in {\mathcal A}$ that satisfies the following:
and, thus, ignores the cases $p(\widetilde{x})=1/2$ at the decision making boundary. ${\mathcal A}_{ms}$ always has a non-empty interior in ${\mathbb R}^{K-1}$ (recall that ${\mathcal A}_0$ can have smaller dimension $K-1-r_0$ which depends on equality constraints). Description ((ref)) by itself presents a convex polyhedron with its boundary excluded. This polyhedron is then intersected with ${\mathcal A}$. This description of ${\mathcal A}_{ms}$ is sufficient to conclude that ${{\mathcal A}}_{ms}$ will coincide with ${{\mathcal A}}_{0}$ when $p(\widetilde{x}) \neq 1/2$ for all $\widetilde{x} \in \mathcal{X}$, and will be a superset of ${{\mathcal A}}_{0}$ otherwise.
In Illustrative Design 1, as we discussed in Section (ref), we distinguish between the case of point identification (whenever ${\Greekmath 010B}_0 \in \{-1,0\}$) and the case of partial identification of the parameter of interest. While we defer formal derivations to the appendix, we can characterize the maximizer of the population maximum score objective as follows. When ${\Greekmath 010B}_0=0$, then ${\mathcal A}_{ms}=[-1,+\infty) \cap {\mathcal A}$. If ${\Greekmath 010B}_0=-1$, then ${\mathcal A}_{ms}=[-\infty,1)$. If ${\Greekmath 010B}_0 \not\in \{0,\,-1\}$ and, without loss of generality, ${\Greekmath 010B}_0>0$ then ${\mathcal A}_{ms}=\overline{{\mathcal A}}_{0} =[0,\,+\infty) \cap {\mathcal A}.$ Note that cases ${\Greekmath 010B}_0 \in \{0,\,-1\}$ in Illustrative Design 1, of course, correspond to situations when there is $x$ in the support such that $p({x})=\frac12$. Such $x$ drives point identification of ${\Greekmath 010B}_0$ in the the formal identification analysis but at the same time is fully disregarded by the population maximum score objective function $MS({\Greekmath 010B})$ (as well as by $MS_{n,INF}({\Greekmath 010B})$). Thus, in these cases the infeasible maximum score estimator is not sharp. In cases ${\Greekmath 010B}_0 \not\in \{0,\,-1\}$, the maximizer of the infeasible maximum score objective function is the closure of the identified set with probability approaching 1. Consequently, the unfeasible maximum score estimator can be sharp for those values of parameters of the DGP.
Now, let's shift our focus to the original (feasible) maximum score estimator, which maximizes the (feasible) sample objective function $MS_{n}({\Greekmath 010B})$. In this context, situations where $p(\widetilde{x})=1/2$ once again present challenges, albeit in a distinct manner compared to the issues encountered in $MS({\Greekmath 010B})$. Here, these cases are not disregarded in the objective function but lack consistent consideration. This is due to the fact that, for arbitrarily large samples, the estimator $\hat p(\widetilde{x})$ fluctuates on different sides of $1/2$, with probabilities approaching $1/2$. As turns out, this will translate in the fluctuating behavior of the maximum score estimator $\widehat{{\mathcal A}}_{ms}$ itself. We first describe this fluctuating behavior occurs in the context of Illustrative Design 1. While the technical details of this derivation are deferred to the appendix, when can can show that while ${\Greekmath 010B}_0 \not \in \{0,-1\},$ then the probability limit of $\widehat{{\mathcal A}}_{ms}$ is the set $[0,\,+\infty),$ which is the closure of the identified set. However, whenever ${\Greekmath 010B}_0=0$ (or ${\Greekmath 010B}_0=-1$) , then whenever ${\Greekmath 010B}_0 \in \{0,\,-1\},$ then $\widehat{{\mathcal A}}_{ms}$ neither converges to ${\mathcal A}_{ms}$ nor to ${\mathcal A}_0.$ Using the terms of the random set theory in molchanov2006book, we can characterize the limit as a random set
where $B$ is a Bernoulli random variable with parameter $\frac12$.
This means that the maximum score estimator is {\it not sharp} at points ${\Greekmath 010B}_0 \in \{0,\,-1\} $ and is sharp elsewhere in the parameter space.
To give a deeper statistical interpretation of the the limit of the maximum score estimator, we rely on our discussion from Appendix (ref), analyzing uniform convergence of $MS_n({\Greekmath 010B})$ to $MS({\Greekmath 010B})$. One useful observation from there is that for the normalized empirical process $\sqrt{n}\left( MS_n({\Greekmath 010B})-MS({\Greekmath 010B}) \right)$ we can establish the following weak convergence in $\ell_{\infty}(A)$ (the space of bounded functions in $\infty$-norm on $A$) for all $A \subset {\mathcal A}:$
where $Z^0 $ and $Z^1 $ and are independent mean zero Gaussian random variables.
The convergence result ((ref)) sheds the light on the mechanics of the formal characterization of the maximum score estimator for Illustrative Design 1. The right-hand side of ((ref)) is often referred to as “stochastic residual." In fact, whenever ${\Greekmath 010B}_0 \not\in \{0,\,-1\},$ then we can simply rely on the finite variance of the Gaussian variables and boundedness of the sign function to conclude that $\sup\limits_{{\Greekmath 010B} \in {\mathcal A}}\|MS_n({\Greekmath 010B})-MS({\Greekmath 010B})\|=o_p(1)$ leading to the conclusion that the limit of the maximizer of the sample maximum score objective function is the same as the maximizer of the population objective function.
However, whenever ${\Greekmath 010B}_0=0$ (or ${\Greekmath 010B}_0=-1$) the right-hand side no longer plays the role of the “residual." The population objective function $MS({\Greekmath 010B})$ in that case is equal to zero whenever ${\Greekmath 010B}<-1$ and is equal to the strictly positive constant for ${\Greekmath 010B} \geq -1,$ thus, attaining its maximum on the set $[-1,+\infty).$ Due to the presence of Gaussian variables $Z^0$ and $Z^1$ symmetrically distributed around zero, the term $\mbox{sign}({\Greekmath 010B})\,Z^0+ \mbox{sign}(1+{\Greekmath 010B})\,Z^1$ adds a positive weight on either set ${\Greekmath 010B}<0$ or ${\Greekmath 010B} \geq 0$ with equal probabilities each. This means that, even though, the “stochastic residual" term is infinitesimal, it randomly selects one of two subsets in the partitioning of the set ${\mathcal A}_{ms}$ (which, as a reminder, maximizes the population $MS({\Greekmath 010B})$), designating that subset as the maximizer of the sample objective function $MS_n({\Greekmath 010B})$. It is important to note that the partitioning of ${\mathcal A}_{ms}$ occurs at the parameter value ${\Greekmath 010B}_0$ in the DGP, carrying significant implications for our search for a superior estimator compared to the maximum score at a later stage.
Our findings here extend to general discrete-only scenarios. For instance, in Illustrative Design 2 we can partition the parameter space into the following sets: ${\mathcal C}_1=\{{\Greekmath 010B}_1+{\Greekmath 010B}_2<0,\,{\Greekmath 010B}_1+1<0\},$ ${\mathcal C}_2=\{{\Greekmath 010B}_1+{\Greekmath 010B}_2<0,\,{\Greekmath 010B}_1+1>0\},$ ${\mathcal C}_3=\{{\Greekmath 010B}_1+{\Greekmath 010B}_2>0,\,{\Greekmath 010B}_1+1>0\},$ ${\mathcal C}_4=\{{\Greekmath 010B}_1+{\Greekmath 010B}_2>0,\,{\Greekmath 010B}_1+1<0\},$ as well as hyperplanes ${\Greekmath 010B}_1+{\Greekmath 010B}_2=0,$ ${\Greekmath 010B}_1+1=0$ with and without the point $(-1,1).$
In Case 2A the maximizer of the population objective is the entire parameter space ${\mathcal A}$ and the maximum score estimator has the asymptotic distribution outputting sets $\overline{{\mathcal C}}_1$ -- $\overline{{\mathcal C}}_4$ with probabilities $\frac14$ each.
In case 2B, the maximizer of the population objective is a half-space $\{({\Greekmath 010B}_1,{\Greekmath 010B}_2)\,:\,\operatorname*{sign}({\Greekmath 010B}_1+{\Greekmath 010B}_2)= \operatorname*{sign}({\Greekmath 010B}_{0,1}+{\Greekmath 010B}_{0,2})\}$ and the maximum score estimator has asymptotic distribution outputting sets $\overline{{\mathcal C}}_3$ and $\overline{{\mathcal C}}_4$ (if ${\Greekmath 010B}_{0,1}+{\Greekmath 010B}_{0,2}>0$) or $\overline{{\mathcal C}}_1$ and $\overline{{\mathcal C}}_3$ (if ${\Greekmath 010B}_{0,1}+{\Greekmath 010B}_{0,2}<0$) with probabilities $\frac12$ each.
In case 2C, the maximizer of the population objective function is a half-space $\{({\Greekmath 010B}_1,{\Greekmath 010B}_2)\,:\,\operatorname*{sign}({\Greekmath 010B}_1+1)= \operatorname*{sign}({\Greekmath 010B}_{0,1}+1)\}$ and the maximum score estimator has asymptotic distribution outputting sets $\overline{{\mathcal C}}_2$ and $\overline{{\mathcal C}}_3$ (if ${\Greekmath 010B}_{0,1}+1>0$) or $\overline{{\mathcal C}}_1$ and $\overline{{\mathcal C}}_4$ (if ${\Greekmath 010B}_{0,1}+1<0$) with probabilities $\frac12$ each.
In case 2D the maximum score estimator converges in probability to one of the sets $\overline{{\mathcal C}}_1$ -- $\overline{{\mathcal C}}_4,$ coinciding with the closure of the identified set.
Theorem (ref) formulates a general result.
Or next theorem gives a more elaborate result regarding the asymptotic behaviour of the maximum score estimator when $p(\widetilde{x})= 1/2$ for some $\widetilde{x} \in \mathcal{X}$, and gives a form of its distribution limit.
Our final discussion here relates to the nature of the sets ${\mathcal C}_{\ell}$. Without a loss of generality, let the first $M$, $M\geq 1$, points in $\cal{X}$ be those at the decision making boundary $p(\widetilde{x})= 1/2$ while the rest are not. Let us denote the collection of these first $M$ support points as $\mathcal{X}_{db}$. As explained earlier, the maximizer $\mathcal{A}_{ms}$ of the population maximum score objective function is purely defined by $\widetilde{x} \notin \mathcal{X}_{db}$ through inequalities ((ref)) for all $\widetilde{x} \notin \mathcal{X}_{db}$.
Consider any combination of inequalities $\widetilde{x}'\widetilde{{\Greekmath 010B}} \geq 0$ or $\widetilde{x}'\widetilde{{\Greekmath 010B}}<0$ for $\widetilde{x} \in \mathcal{X}_{db}$ -- a combination has to include an inequality for each $\widetilde{x} \in \mathcal{X}_{db}$. Each this combination results in a maximum score estimate ${\mathcal C}_{\ell}$ which is a subset of ${\mathcal A}_{ms}$ (since ((ref)) hold asymptotically for all $\widetilde{x} \notin \mathcal{X}_{db}$ and, thus, are taken as given). Different combinations of signs of $\widetilde{x}'\widetilde{{\Greekmath 010B}}$ for $\widetilde{x} \in \mathcal{X}_{db}$ result either in disjoint or identical maximum score estimates ${\mathcal C}_{\ell}$. We denote the collection of all these unique maximum score estimates as ${\mathcal C}_{1}$, \ldots, ${\mathcal C}_{L}$. For more details see the proof of Theorem (ref).
In the previous subsection, we established that the maximum estimator is not sharp whenever there exist points $\widetilde{x}$ in the support of covariate $\widetilde{X}$ where $p(\widetilde{x})= \frac12.$ If such points do not exist, it is sharp and, moreover, it converges in probability to the identified set ${\mathcal A}_0.$ However, at those values the maximum score estimator only converges weakly to a random set. This means, that the weak limit of the maximum score estimator in our partially identified is {\it discontinuously} changes with respect to the parameter of the underlying data generating process.
In this subsection we study {\it robustness} of the maximum score estimator ((ref)) for our simple discrete design at ${\Greekmath 010B}_0 \in \{0,-1\}$ as the property of continuity of the weak limit of the estimator with respect to specific class of sequences of parameters of the data generating process. As we discussed in Section (ref), parameter sequences of particular interest to us take the form ${\Greekmath 010B}(n,t;{\Greekmath 010B}_0)={\Greekmath 010B}_0+t/\sqrt{n}.$
The choice of local data generating processes indexed by these sequences reflects the foundational property of the maximum score estimator allowing us to interpret it as an ensemble of weak learners (e.g., see section 10.1 in shalev:14). To see this, we take our Illustrative Design 1. Objective function ((ref)) can be viwed as the aggregator of “votes" where each observation casts a vote for one of the sets $(-\infty,-1),$ $[-1,0)$ and $[0,+\infty).$ To see this, consider a single element of the maximum score objective function $ y_i\,I[{\Greekmath 010B}+x_i \geq 0]+(1-y_i)\,I[ {\Greekmath 010B}+x_i<0 ]. $ For each possible combination $(y_i,x_i)$ the element of the objective function selects the intervals $[-1,+\infty),$ $[0,+\infty),$ $(-\infty,-1)$ and $(-\infty,0).$
This element is a classifier\footnote{This classifier is referred to as a “weak learner" in the terminology of the Machine learning literature. The term “weak learner" is used because the probability that it classifies the object of interest correctly is not diminishing in larger samples.} which can select one of the sets $(-\infty,-1),$ $[-1,0)$ and $[0,+\infty)$ or their pairwise unions. For the pairwise union, the classifier selects each set in the union.
The classification outcome for observation $i$ is $(-\infty,-1)$ with probability $(1-p(1))q,$ $(-\infty,-1) \cup [-1,0)$ with probability $(1-p(0))(1-q),$ $[0,+\infty)$ with probability $p(0)(1-q)$ and $[-1,0)\cup (0,+\infty)$ with probability $p(1)q$ (where $p(x)=P(Y=1\,|\,X=x)$ and $q=P(X=1)$). We consider the concept of the weak learner in context of partial identification, novel to the machine learning literature.
Let $v_i$ be a three-dimensional vector with elements $1/0$ depending on whether the corresponding set $(-\infty,-1),$ $[-1,0)$ or $[0,+\infty)$ (for each of the three dimensions) was selected by observation $i.$ Then $\bar{v}=\frac1n \sum^n_{i=1}v_i$ produces a vector of collective “votes" of $n$ classifiers for each observation $i$ and the maximum score estimator can be written as $$ \widehat{{\mathcal A}}_{ms}= \mbox{(}-\infty,-1\mbox{)}\cdot {\bf 1}\{\operatorname*{arg\,max} \bar{v}=1\}+ \mbox{[}-1,0 \mbox{)}\cdot {\bf 1}\{\operatorname*{arg\,max} \bar{v}=2\}+ \mbox{[}0,+\infty\mbox{)}\cdot {\bf 1}\{\operatorname*{arg\,max} \bar{v}=3\}. $$
The estimator is a function of the sample mean $\bar{v}$ converging at the standard parametric rate $\sqrt{n}$ to a normal random variable. For the analysis of continuity of distribution limit of such parameters the literature (e.g., ibragimov, Lecam1953) suggest considering sequences of the population distributions of the underlying random variable with its expectation drifting at the same parametric rate. Since the expectation of the random vector $v_i$ is linear in the probability $p(0),$ sequences of probabilities $p^{(n)}(0)$ drifting at $1/\sqrt{n}$ rate will result in an equivalent drifting of the expectation of $v_i$ at that rate.
We now establish the general result for the maximizer of ((ref)).
Theorem (ref) has the following two major implications for Illustrative Design 1. First, when the sequence of parameters of the data generating process converge to parameter values where the maximum score estimator is sharp for the Illustrative Design 1, drifting does not impact its limit and it converges to the identified set ${\mathcal A}_0.$ Second, when the sequence of parameters of the data generating process converges to the value where the maximum score estimator is not sharp, the maximum score estimator converges to a random set whose distribution depends on the constant indexing a particular parameter sequence. In particular, it has the property that $\sup\limits_{t>0}\lim\limits_{n \rightarrow \infty}{\mathbb P}\left( d_H(\widehat{{\mathcal A}}_{ms},\,[0,+\infty)\cap {\mathcal A})=0\right) =1,$ i.e., it converges to limit of the estimator at parameter values of the data generating process not equal to zero (but, possibly, arbitrarily close to zero). This means that the maximum score estimator is {\it locally robust} respect to parameter sequences ${\Greekmath 010B}(t,n;{\Greekmath 010B}_0)={\Greekmath 010B}_0+t/\sqrt{n}$ at points where it is not sharp, according to Definition (ref).
Parameter drifting to the parameter values where the maximum score estimator is not sharp ranges between two regimes. The first regime corresponds to the distribution limit of the maximum score estimator which is a random set taking sets $[-1,0)$ and $[0,+\infty)$ with equal probabilities (corresponds to $t=0$). In the second regime is where the maximum score estimator puts a point mass of 1 on the set $[0,+\infty),$ which is the maximizer of the population maximum score objective function and the closure of the identified set whenever parameter of the data generating process is fixed at values outside of $-1$ and $0$ (corresponding to $t=+\infty$).
To summarize, for parameter of the data generating process drfiting towards ${\Greekmath 010B}_0=0$ as $t$ varies from $0$ to $+\infty,$ the distribution of the limit random set varies from equal randomization between $[-1,0)$ and $[0,+\infty)$ to selecting a fixed set $[0,+\infty).$ Thus, this choice of the drifting sequence bridges the two cases. As a result, even though the maximum score estimator is not sharp at that point, it is locally robust.
For Illustrative Design 2, Theorem (ref) leads to a similar behavior of the limit as for Illustrative Design 1, albeit, with more complex structure of limit. When ${\Greekmath 010B}_0$ corresponds to the value where $p(\bar{x})=\frac12$ for one or more points in the support of $\widetilde{X},$ the distribution limit of the maximizer of ((ref)) takes value on two or more sets ${\mathcal C}_1, {\mathcal C}_2,{\mathcal C}_3,{\mathcal C}_4$ formed by intersections of half-spaces with boundary hyperplanes ${\Greekmath 010B}_1+{\Greekmath 010B}_2=0$ and ${\Greekmath 010B}_1+1=0.$
Whenever ${\Greekmath 010B}(t,n;{\Greekmath 010B}_0)= {\Greekmath 010B}_0+t/\sqrt{n} \in {\mathcal C}_k$ and $\|t\|$ increases, the probability that random set has a realization ${\mathcal C}_k$ increases. Consequently, $\sup\limits_{t,\,{\Greekmath 010B}(t,n;{\Greekmath 010B}_0) \in {\mathcal C}_k}\lim\limits_{n \rightarrow \infty}{\mathbb P}\left( d_H(\widehat{{\mathcal A}}_{ms},\,{\mathcal C}_k\cap {\mathcal A})=0\right) =1.$ Thus, the maximum score estimator is robust.
We next consider the second classic estimator in our semiparametric binary choice model with discrete regressors. This estimator relates directly to procedures introduced in \citeasnoun{ichimura1994} and \citeasnoun{ahnetal2018}. They are based on the following implication of Assumption (ref) (including strict monotonicity of $F_{{\Greekmath 010F}|X}$ at 0):
Based on the above implication, \citeasnoun{ichimura1994} proposed the estimator to be minimizer of the $\frac{1}{n}\sum_{i=1}^n \hat w_i(\widetilde{x}_i'\,\widetilde{{\Greekmath 010B}})^2,$ subject to normalization on $\widetilde{{\Greekmath 010B}}$ imposed by Assumption (ref) (ii). In this setup $\hat w_i$ is a smoothing function which uses a nonparametrically estimated conditional probability $p(\widetilde{x})$ for each observation and puts higher weight on the observations for which this predicted probability is close to $1/2$.
This estimator is easy to implement, as it is of closed form in each stage. Under stated conditions ensuring point identification, this estimator was shown to have desirable asymptotic properties, specifically being asymptotically equivalent to maximum score or smoothed maximum score (\citeasnoun{horowitz-sms}). Its disadvantage compared to maximum score is that because it involves nonparametric procedures, it requires one to make the choice of kernel function and bandwidth sequence.
In our discrete regressors setup, we can e.g. take $w_i={\bf 1}\left\{|\widehat{p}(\widetilde{x}_i)-\frac12|<h_n\right\}$ where sequence $h_n \rightarrow 0$ is the tuning parameter of the estimator.
With $w_i$ chosen in this way, we denote the minimizer of $\frac{1}{n}\sum_{i=1}^n w_i(\widetilde{x}_i'\,\widetilde{{\Greekmath 010B}})^2,$ subject to the required normalization as $\widehat{{\mathcal A}}_{CF}.$ If $|\widehat{p}(\widetilde{x}_i)-\frac12|<h_n$ is not satisfied for any observations we set $\widehat{{\mathcal A}}_{CF}\equiv {\mathcal A}$ to ensure the estimation procedure does not halt.
Note that the closed form estimator can be considered as maximizing the sample objective function
with $w_i$ chosen as above. The respective population objective function is then
whose maximizer in ${\mathbb R}^{K-1}$ we denote as ${\mathcal A}_{CF}$. We take ${\mathcal A}_{CF}$ to the entire parameter space ${\mathcal A}$ where ${\mathbb P}(p(\widetilde{X})=1/2)=0.$
Let us turn to Illustrative designs to exemplify closed form estimation.
In Illustrative Design 1, $$ {\mathcal A}_{CF}=\left\{
\right. $$
This description indicates that the population objective function attains its optimum, equivalent to the identified set, solely when the parameter ${\Greekmath 010B}_0$ in the DGP takes on values from the set ${0, -1}$ and, thus, when we have the case of point identification. When ${\Greekmath 010B}_0$ assumes any other value, thus, we are in the situation of partial identification, then the population objective function produces the default of the entire parameter space, which is, of course, a superset of the identified set.
In Illustrative Design 2, $$ {\mathcal A}_{CF}=\left\{
\right. $$ In line with our conclusions for Illustrative Design 1, ${\mathcal A}_{CF}$ coincides with the identified set ${\mathcal A}_0$ only in the point identification case (Case 2A). In situations of partial identification (Cases 2B-2D), ${\mathcal A}_{CF}$ is a strict superset of ${\mathcal A}_0$.
These conclusions are intuitive. Maximizer of ((ref)) disregards values $\widetilde{X}$ that do not lie on the decision-making boundary characterized by $p(\widetilde{X})=1/2$. In scenarios involving point identification (like 2A in Illustrative Design 2 or 1A in Illustrative Design 1), such point identification is obtained because there are enough of $\widetilde{X}$ instances at the decision boundary. The closed-form method then efficiently utilizes all such $\widetilde{X}$ instances. However, in cases where there are insufficient $\widetilde{X}$ values at the decision boundary to achieve point identification (as observed in Cases 2B and 2C in Illustrative Design 2), or none at all (as seen in Case 2D in Illustrative Design 2 or Case 1B in Illustrative Design 1), the closed-form method relies solely on limited information (or none at all in Cases 2D and 1B). It overlooks information from all other $\widetilde{X}$ values not on the decision boundary, even though they may contribute to the structure and shape of the identified set.
A result for a generic discrete regressors setting is given in Theorem (ref).
Our next focus is on the sampling behavior of the closed-form estimator formally established in Theorem
Results of Theorems (ref) and (ref) imply that the closed-form estimator is not sharp when ${\Greekmath 010B}_0$ in the DGP does not result in point identification.
Some further insights can be obtained from considering illustrative designs. The maximum score estimator in Illustrative Design 1 was sharp only when ${\Greekmath 010B}_0 \notin \{0,-1\}$ (1B) in the DGP and was not sharp otherwise (1A). Conversely, the closed-form estimator in the same design exhibits contrasting behavior -- it is sharp solely when ${\Greekmath 010B}_0 \in {0,-1}$ (1A) within the DGP, and loses sharpness otherwise (1B). This disparity is unsurprising, given that these two estimation methods rely on distinct principles -- the closed-form estimator exclusively utilizes regressor values at the decision-making boundary, while the maximum score estimator is sharp in the absence of such regressor values.
In Illustrative Design 2, maximum score was sharp only in Case 2D whereas the closed-form estimator is sharp only in Case 2A, thus leaving 2B and 2C as cases where neither of these two estimators is sharp.
To analyze robustness of the closed-form estimator we use the sequences of parameters of the data generating process with the property $\|{\Greekmath 010B}(t,n; {\Greekmath 010B}_0)-{\Greekmath 010B}_0\|=O(1/\sqrt{n}).$ We do so by the same reasoning as in case of the maximum score estimator, given that the core object of the estimator ((ref)), the conditional probability $P(Y=1\,|\,X=x),$ is represented by a sample mean. The following theorem characterizes the distribution limit of $\widehat{{\mathcal A}}_{CF}$ under those drifting parameter sequences.
This theorem demonstrates that the convergence of the maximizer of ((ref)) is not affected by the choice of the drifting sequence as long as it is not entirely contained in the subsets of the parameter space for which ${\mathbb P}(Y=1\,|\,X=x) \equiv \frac12$ for some $x$ in the support of $X$.
To summarize, the maximizer of ((ref)), $\widehat{{\mathcal A}}_{CF}$, is neither {\it sharp}, nor is it{\it robust.}
Our previous discussion highlighted the sharpness, or lack thereof, exhibited by both maximum score and closed-form estimators. It became evident that merging the underlying concepts of these approaches is necessary to formulate an estimation procedure surpassing either individual method. Essentially, this entails leveraging all regressor values—both those at the decision boundary and those away from it -- in a judicious manner. The primary challenge lies in devising such an approach.
Our first idea for a new estimator is, in some sense, an obvious one. It is to directly combine the maximum score and the closed-form estimators through a feasible data-driven “switching device”. This is how it goes. First, construct estimators for the choice probabilities $\hat p(\cdot)$ from the sample $\{(y_i,\widetilde{x}_i)\}^n_{i=1}$. Next, define the minimum deviation of estimated choice probabilities from $1/2$: $ \widehat{P}=\min_{i=1,\ldots,n}|\hat{p}(\widetilde{x}_i)-\frac{1}{2}|$.
Our proposed “switching" estimator is then explicitly defined as
with ${\Greekmath 0117}_n \rightarrow 0$ being the tuning parameter sequence. Even tough we use the same notation for a tuning parameter here as we used in the definition of the closed-form estimator, these two sequences are not necessarily the same.
We refer to it as a hybrid estimator. This estimator builds on the idea that $\widehat{P}$ being far enough from zero is evidence against having choice probabilities being $1/2$, and $\widehat{P}$ being close enough from zero is supportive of the opposite case. Naturally, in the former case we can rely on the maximum score estimator and, in the latter case, on the closed form estimator. $\widehat{\cal A}_H$ may both output a set or a singleton in finite samples, depending on the data generating process.
In Illustrative Design 1 with the right choice of the rate of ${\Greekmath 0117}_n \to 0$ (e.g., $\sqrt{n}{\Greekmath 0117}_n \to \infty$) this estimator is going to be sharp. When ${\Greekmath 010B}_0 \in \{0,-1\}$ (1A), we expect $\widehat{P}$ in large enough sample to be less than ${\Greekmath 0117}_n$ and, thus, picking the closed-form estimator which is sharp in Case 1A. When ${\Greekmath 010B}_0 \notin \{0,-1\}$ (1B), we expect $\widehat{P}$ in large enough sample to be greater than ${\Greekmath 0117}_n$ and, thus, picking the maximum score estimator which is sharp in Case 1B.
The situation in Illustrative Design 2, however, is very different. Even with the right choice of the rate of ${\Greekmath 0117}_n \to 0$, this estimator is not sharp. Namely, it will be sharp in Case 2A as the closed-form estimator is then sharp and it will be selected by the “switching device” (there will be two sample conditional choice probabilities close to 1/2), and it will also be sharp in Case 2D as the maximum score estimator is then sharp and it will be selected by the “switching device” (no sample conditional choice probabilities close to 1/2). In Cases 2B and 2C a the “switching device” will, once again, select the closed-form estimator (as there will be one sampling conditional choice probability close to 1/2) but the that estimator is not sharp in these cases.
Our results in Illustrative Design 2 also make clear when the hybrid will be sharp in general discrete regressors settings. This is formulated in Theorem (ref).
It is evident that the issues regarding the sharpness of the hybrid estimator, as seen in Illustrative Design 2, arise from the inability of the switching device to distinguish between cases where only one sample conditional choice probability is close to 1/2 and cases where there are multiple such probabilities. In light of this observation, it is prudent to explore more sophisticated switching devices for $K>2$, such as those based on the $(|\mathcal{X}_n|-K+2)$-th order statistic\footnote{This would require $|\mathcal{X}_n|-K+2\geq 1$.} of the $|\mathcal{X}_n|$-dimensional vector of ${|\hat{p}(\widetilde{x}_i)-1/2|}$ for unique $\widetilde{x}_i$ in finite-sample support $\mathcal{X}_n$ ($|\mathcal{X}_n|$ denotes the cardinality of $\mathcal{X}_n$). This adjustment would help in Illustrative Design 2 and would result in the sharpness of the modified hybrid estimator. However, in general discrete regressors scenarios this would not guarantee sharpness as having $K-1$ conditional choice probabilities close to 1/2 does not necessarily guarantee sufficient variation in the respective $\widetilde{x}_i$ to uniquely identify the parameter (the only case when the closed-form estimator is sharp). Consequently, we opt not to pursue this extension formally.
The next estimator we introduce that integrates idea from both maximum score and closed-form estimation relies on modifying the maximum score objective function ((ref)) which uses estimated choice probabilities in the first step. Our proposed modified objective function is
where ${\Greekmath 0117}_n$ is a tuning parameter akin to that used in the hybrid estimator discussed earlier. We denote the set of maximizers of ((ref)) as $\widehat{\mathcal{A}}_{OFC}$.\footnote{A similar 2-step rank estimator for heteroskedastic monotone index models was proposed in \citeasnoun{khan2001} for a point identified model. An estimator involving a Maximum score objective function with nonparametrically estimated regressors in the first stage was proposed in \citeasnoun{chenetal2014}.}
The objective function ((ref)) blends principles of the objective functions ((ref)) and ((ref)). Analogous to the maximum score estimator, it incorporates inequalities of the index function. Notably, by utilizing estimated choice probabilities and formulating inequalities of the index function as weak in both directions, it assigns importance to observations where the index equals zero. A “slackness" tuning parameter ${\Greekmath 0117}_n$, as proven later, ensures sharpness of the estimator $\widehat{\cal A}_{OFC}$ in both point and set identified cases.\footnote{A related slack variables idea was introduced in \citeasnoun{komarova2013}.}
To provide further insight into the functioning of this estimation method, let's examine Illustrative Design 1.
First, consider ${\Greekmath 010B}_0$ as outlined in Case 1A. In this scenario, the estimated choice probability $\hat{p}(x_i)$ for $x_i=-{\Greekmath 010B}_0$ is expected to be distributed around 0.5, with a positive probability (provided the rate of ${\Greekmath 0117}_n$ is suitably chosen). In the objective function ((ref)), regardless of whether $\hat{p}(-{\Greekmath 010B}_0)$ falls below or above 0.5, as long as it lies within the interval $(0.5-{\Greekmath 0117}_n, 0.5+{\Greekmath 0117}_n)$, both inequalities ${\Greekmath 010B}-{\Greekmath 010B}_0 \geq 0$ and ${\Greekmath 010B}-{\Greekmath 010B}_0 \leq 0$ are equally likely to be satisfied. Consequently, in the limit, this yields the identified set $\{{\Greekmath 010B}_0\}$.
When $x_i=1+{\Greekmath 010B}_0$, $\hat{p}(1+{\Greekmath 010B}_0)$ is expected to deviate from 0.5 by at least ${\Greekmath 0117}_n$, but this deviation does not influence the shape of the identified set. To elaborate further, when ${\Greekmath 010B}_0=0$, $\hat{p}(1)$ exceeds $0.5+{\Greekmath 0117}_n$ with probability approaching 1, resulting in the inequality ${\Greekmath 010B} \geq -1$. When combined with ${\Greekmath 010B}=0$, this ultimately converges to 0 in probability. Conversely, when ${\Greekmath 010B}_0=-1$, $\hat{p}(0)$ falls below $0.5-{\Greekmath 0117}_n$ with probability approaching 1, leading to the inequality ${\Greekmath 010B} \leq 0$. When combined with ${\Greekmath 010B}=1$, this converges to $-1$ in probability.
Now let us consider the partially identified Case 1B in Illustrative Design 1. When ${\Greekmath 010B}_0>0$, both $\hat{p}(0)$ and $\hat{p}(1)$ exceed $0.5+{\Greekmath 0117}_n$ with probability approaching 1, resulting in the limit set ${\mathcal A}_{OFC}$ that includes all the points from parameter set ${\mathcal A}$ that satisfy the inequality ${\Greekmath 010B} \geq 0$. This differs from ${\mathcal A}_0 =(0,+\infty) \cap {\mathcal A}$ only in the inclusion of the boundary point 0, thus giving $d_H({\mathcal A}_{OFC}, {\mathcal A}_0)=0$. Other sub-cases ${\Greekmath 010B}_0 \in (-1,0)$ and ${\Greekmath 010B}_0 \in (-\infty,-1)\cap {\mathcal A}$ within Case 1B will give analogous results.
We now provide a general result on the asymptotic behavior of our second proposed combination estimator.
In particular, Theorem (ref) means that when ${\mathcal A}_0$ is a singleton, then $\widehat{\cal A}_{OFC}$ converges in probability to the singleton (as the boundary of ${\mathcal A}_0$ is empty). When ${\mathcal A}_0$ is a non-singleton, then even for arbitrarily large samples $\widehat{\cal A}_{OFC}$ may differ from ${\mathcal A}_0$ in including some boundary points of ${\mathcal A}_0$. This, however, does not affect the Hausdorff distance.
In summary, Theorem (ref) establishes sharpness of our second combination estimator.
Similarly to the analysis of the classic estimators for semiparametric binary choice model we consider robustness property of the new estimators constructed in this section by exploring continuity of the limit with respect to sequences of parameters of the data generating process ${\Greekmath 010B}(t,n;{\Greekmath 010B}_0)$ with the property $\|{\Greekmath 010B}(t,n;{\Greekmath 010B}_0)-{\Greekmath 010B}_0\|=O(1/\sqrt{n}).$ As observed in the analysis of robustness of the maximum score estimator, its objective function is constructed from the estimated choice probability which in the discrete design is a sample average which determined the choice of the rate of drifting of the parameter sequence towards ${\Greekmath 010B}_0.$
We then provide an analogous result for the maximizer of ((ref)):
The main conclusion is that neither estimator has continuity properties under drifting asymptotics. Estimator $\widehat{{\mathcal A}}_{H}$ converges to the population maximizer of the closed-form estimator for any sequence of parameters converging to the point ${\Greekmath 010B}_0$ where $p(\widetilde{x})=\frac12$ for some $\widetilde{x} \in {\mathcal X}.$ Estimator $\widehat{{\mathcal A}}_{OFC}$ converges to the identified set corresponding to the parameter in the limit of the parameter sequence ${\Greekmath 010B}(n,t;{\Greekmath 010B}_0).$ As a result, arbitrarily small change in ${\Greekmath 010B}_0$ resulting in equalities $p(\widetilde{x})=\frac12$ not hold, results in a discontinuous change in the limit of both estimators.
In other words, neither of the constructed estimators are robust.
In this section we offer a novel approach to construction of sharp and robust estimators which can serve as a basis for constructing consistent estimators for identified sets in partially identified discrete choice models. Our approach introduces a new class of estimators grounded in the concept of random set theory, a framework that has been integral to Econometrics since the pioneering work of \citeasnoun{beresteanumolinari2008}. While previous econometric research has primarily concentrated on the concept of selection (or Aumann) expectation within this theory and its associated estimators (see also subsequent studies by \citeasnoun{beresteanumolchanovmolinari2011}) and \citeasnoun{BERESTEANU201217}), our focus diverges due to the distinctive characteristics of our model and its classical estimators within the discrete-only regressors framework. Notably, such feature as the fluctuating behavior of the maximum score estimator in some scenarios prompts us to adopt a fundamentally different approach from the random set theory.
First, to give some insights on what our estimation approach will deliver, we focus on Illustrative Design 1. For this design Theorem (ref) implied in Case 1A -- for concreteness, let us take ${\Greekmath 010B}_0=0$, -- $$\widehat{{\mathcal A}}_{ms} \stackrel{d}{\to} {\bf A} \equiv t \cdot \underbrace{[0,+\infty) \cap \mathcal{A}}_{\equiv B_1} + (1-t) \cdot \underbrace{[-1,0) \cap \mathcal{A}}_{\equiv B_2},$$ where $t$ is a Bernoulli random variable with parameter $\frac12$.
If we look at the distribution limit, which is a random set, we can see that ${\Greekmath 010B}_0=0$ is the only point in the parameter space that happens to be in the closure of realizations of random set ${\bf A}$ in at least $100(1/2+\Delta) \%$ cases for $\Delta>0$. Indeed, $0$ belong in the boundary of both $B_1$ and $B_2$ and no other point simultaneously belongs to the boundary of both sets. Moreover, the proof of Theorem (ref) in the Appendix will make it clear that the same result will hold for $\widehat{{\mathcal A}}_{ms}$ in a finite-sample for large enough $n$.
This naturally brings us to the estimation approach based on the notion of the random set quantile. Namely we define our estimator as the random set ${\Greekmath 011C}$-th quantile of the closure of the maximum score estimator:
Following the definition of the ${\Greekmath 011C}$-th quantile of a random set in \citeasnoun{molchanov2006book} (p.176),
for {\it coverage function } $p\left(u\,;\,\overline{\widehat{{\mathcal A}}_{ms}}\right) \equiv {\mathbb P}\left( u \in \overline{\widehat{{\mathcal A}}_{ms}}\right).$ Note that this notion from \citeasnoun{molchanov2006book} also applies to vector settings.
An important consideration lies in the selection of suitable quantile indices ${\Greekmath 011C}$ for our estimation. To foreshadow our formal result, we propose utilizing a ${\Greekmath 011C}=1/2+\Delta$-th quantile of the closure of $\widehat{\mathcal{A}}_{ms}$, where $\Delta \in (0, 1/2)$.
Now, let's revisit Illustrative Design 1 for the parameter value ${\Greekmath 010B}_0=0$ in the DGP. In this scenario, the ${\Greekmath 011C}$-th quantile of the closure $\overline{\widehat{\mathcal{A}}_{ms}}$ with ${\Greekmath 011C}=1/2+\Delta$, $\Delta>0$, reduces to only $\{0\}$ for sufficiently large $n$, thereby converging in probability to the identified set.\footnote{It's worth noting that for indices below $1/2$, such as $(1/2-\Delta)$-th quantile of $\overline{\widehat{\mathcal{A}}{ms}}$ with $\Delta \in (0,1/2)$, the resulting quantile set is $[-1,+\infty) \cap \mathcal{A}$ for sufficiently large $n$. This set is a superset of the identified set and aligns with the maximization of the population maximum score objective function in this scenario. On the other hand, the $1/2$-th quantile of $\overline{\widehat{A}_{ms}}$ for large enough $n$ can, with additional effort, be demonstrated to be the two-element set ${-1,0}$, failing to asymptotically recover the identified set.}
Given that the concept of a random set quantile may be unfamiliar to econometricians, it's helpful to draw parallels to voting rules for additional clarity. Let's explore this by examining the median of a random set within the framework of a simple discrete model.
Suppose that we can generate infinitely many random samples $\{(x_i,y_i\}_{i=1}^n$ of a fixed size $n$. Each sample casts votes for any number of elements of $\mathcal{A}$ which maximize the objective ((ref)). Essentially, a given sample votes for its respective maximum score estimate $\widehat{{\mathcal A}}_{ms}$. After the completion of set $\widehat{{\mathcal A}}_{ms}$ to its closure $\overline{\widehat{A}_{ms}}$, majority winners are selected -- namely, those elements in $\mathcal{A}$ that are voted for by at least 50% of the samples. The collection of those majority winners would give us $q_{.5} \left( \overline{\widehat{{\mathcal A}}_{ms}}\right)$. For any arbitrary index ${\Greekmath 011C} \in (0,1)$, the quantile $q_{{\Greekmath 011C}}(\overline{\widehat{{\mathcal A}}_{ms}})$ comprises those elements from $\mathcal{A}$ that adhere to the 'quota rule' with a threshold of ${\Greekmath 011C}$. In simpler terms, these elements within $\mathcal{A}$ must garner votes from at least $100{\Greekmath 011C} \%$ of the samples to be included.
The estimator $\widehat{{\mathcal A}}_{RSQ,{\Greekmath 011C}}$ is infeasible due to the distribution of the random set $\overline{\widehat{{\mathcal A}}_{ms}}$ (of maximizers of ((ref))) being a population object. Consequently, we have rely on the finite sample to approximate the population quantile of $\overline{\widehat{{\mathcal A}}_{ms}}$.
To illustrate our ideas on how we can proceed with this, let us focus on Illustrative Design 1 and Case 1A -- for concreteness, we take ${\Greekmath 010B}_0=0$ as the parameter value in the DGP. Despite the weak limit of the estimator in this case being an essentially a random choice between $[-1,0)$ or $[0,+\infty)$, in a {\it concrete sample} we only observe just one set -- either $[-1,0)$ or $[0,+\infty)$ (whichever of them maximizes ((ref))) with high probability (other sets can be realized as maximizers with a decreasingly small probability). Consequently, to simulate the uncertainty across samples, we need to employ suitable sampling techniques. Simultaneously, these methods must be constructed in a manner that their impact on cases of $x$ where $p(x) \neq 1/2$ remains inconsequential.
To illustrate our proposed approach, we rely on the representation of the maximum score objective function as in ((ref)). We propose a sampling technique that will draw the conditional choice probability $\hat p_s(x)$, $s =1, \ldots, S$, for each value of discrete regressor in its sample support. Since by the standard Moivre-Laplace theorem for each $x $ in the sample support, $$ \sqrt{n}\frac{\hat p(x)-p(x)}{\sqrt{p(x)(1-p(x)}} \stackrel{a}{\longrightarrow} \mathcal{N}\left(0,\,1\right),$$ the for a given sample of size $n$ it may seem natural to draw $\hat p_s(x)$ from $ \mathcal{N}(\hat p(x), \frac{\hat p(x)(1-\hat p(x))}{n})$. However, if $p(x)=1/2$, then this symmetric way of drawing on both sides of $\hat p(x)$ will give a probabilistic advantage to one of the sets from $[-1,0)$ or $[0,+\infty)$ -- whichever of them was the maximum score estimate in our sample. Indeed, if in Case 1A for parameter value ${\Greekmath 010B}_0$ in DGP we have $\hat p(-{\Greekmath 010B}_0)>1/2$, then $\hat p_s(-{\Greekmath 010B}_0)$ will be primarily located above $1/2$ as well while we need them $\hat p_s(-{\Greekmath 010B}_0)$ to appear in approximately similar proportions on both sides of $1/2$. This leads us to proposing an asymmetric sampling technique.
Namely, if $\hat p(x)>1/2$ for $x$ in the sample support, then $$ \textstyle \hat p_s(x) \sim \mathcal{N}\left( \hat p(x), \frac{\hat p(x)(1-\hat p(x))}{n}\right) \cdot {\bf 1}\left\{\hat p_s(x) > \hat p(x) \right\} +\mathcal{N}\left( \hat p(x), {\Greekmath 011C} \frac{\hat p(x)(1-\hat p(x))}{n}\right) \cdot {\bf 1}\left\{\hat p_s(x) \leq \hat p(x) \right\}, $$ and if $\hat p(x)<1/2$, then $$ \textstyle \hat p_s(x) \sim \mathcal{N}\left( \hat p(x), {\Greekmath 011C}\frac{p(x)\hat p(x)(1-\hat p(x))}{n}\right) \cdot {\bf 1}\left\{\hat p_s(x) > \hat p(x) \right\} +\mathcal{N}\left( \hat p(x), \frac{\hat p(x)(1-\hat p(x))}{n}\right) \cdot {\bf 1}\left\{\hat p_s(x) \leq \hat p(x) \right\}. $$ The value of ${\Greekmath 011C}>1$ may be data driven and may depend on the desired quantile index as well as probabilities of values close to the decision making boundary (defined as $p(x)=0.5)$.
Finally, for each $s=1,\ldots,S$ we produce a set
The feasible version of the quantile estimator $\widehat{{\mathcal A}}_{RSQ,{\Greekmath 011C}}$ is approximated from the simulation sample by taking a sample $(0.5+\Delta)$-quantile based on the sample of $\{\widehat{\cal A}_{ms,s}\}_{s=1}^S$.
We now formulate our result regarding the asymptotic behavior of the random quantile set estimator. We consider a general discrete regressors setting and, for technical convenience, focus on $\widehat{{\mathcal A}}_{RSQ,{\Greekmath 011C}}$ defined as in ((ref)).
To formulate the sharpness result, we first use a refined result on the asymptotic behavior of the maximum score estimator in Theorem (ref).
The result of Theorem (ref) immediately implies that the estimator is sharp on the entire parameter space.
We note that there is no difference between various “parameter regimes" ${\Greekmath 010B}_0.$ Regardless of whether the model is point or partially identified, the $(\frac12+\Delta)$-quantile of the random set of closure of the maximum score estimator converges to the identified set.
In our analysis of robustness of the maximum score estimator in Section (ref) we characterize the maximizer of ((ref)) as a random set. Just as in our analysis of sharpness of the estimator, in this section for simplicity of exposition we focus on the infeasible quantile estimator under drifting sequences of the data generating process ${\Greekmath 010B}(t,n;{\Greekmath 010B}_0)$ with the property $\|{\Greekmath 010B}( t,n;{\Greekmath 010B}_0 )-{\Greekmath 010B}_0\|=O(1/\sqrt{n})$ where ${\Greekmath 010B}_0$ is the where $p(\widetilde{x})=\frac12$ at least for one $\widetilde{x} \in {\mathcal X}.$ In other words, the maximum score estimator converges to a random set when the data generating process is indexed by parameter ${\Greekmath 010B}_0.$ The following theorem characterizes the structure of the probability along the such parameter sequences.
Thus, estimator $\widehat{{\mathcal A}}_{RSQ,{\Greekmath 011C}}$ is robust: its limit ranges from the closure of the identified set corresponding to the data generating process indexed by the limit of the parameter sequence to the closure of the identified set corresponding to the data generating process indexed by the elements of the parameter sequence. As we established earlier,$\widehat{{\mathcal A}}_{RSQ,{\Greekmath 011C}}$ is also sharp and, unlike the maximum score estimator, it converges to the identified set for all parameter values of the data generating process.\footnote{ While this estimator is infeasible, a consistent estimator of this quantile (see our proposed approach earlier ((ref)) would also have this property.}
\setcounter{equation}{0}
The UK General Election in 2019 marked a significant development for the Labour Party as it faced a decline in its constituency victories. With a total of 650 constituencies, Labour secured only 202 seats during this electoral contest, a historic low both in terms of numerical count and proportion since the year 1935. Various media analyses pointed towards the aftermath of the Brexit referendum as a contributing factor to Labour's electoral setbacks. A telling example is a headline from The Guardian that succinctly captured the sentiment: "It was Brexit, not left-wing policies, that lost Labour this election."\footnote{\tiny \url{https://www.theguardian.com/politics/2020/jun/18/key-points-from-review-of-2019-labour-election-defeat}}
We consider the outcome representing the indicator for a political party winning a given constituency in 2015 and retaining its seat in the 2019 elections. We focus on the question whether the “Leave" vote in the Brexit referendum and its consequences has impacted preferences of the UK voters.
To address our question, we construct 5 factors and measure their impact on the outcome of 2019 election for the Labour. The first is an indicator variable reflecting the “Leave" vote. The second is an indicator denoting a winning margin in the 2015 General election of less than 5%, capturing fiercely contested constituencies before the Brexit referendum. The third is an ordered variable with values 0, 1, 2. It takes on 0 if the mean income growth from 2015 to 2019 was negative (8.15% in our data), 1 if positive but below the median growth across all constituencies (41.84% of constituencies in the data), and 2 if above the median growth (naturally, 50% of constituencies). The forth is an indicator that the winning party in a constituency in 2015 was Labour. We also include an interaction between the indicators of the Leave vote and Labour's 2015 victory, capturing the potential differential effects of the Brexit referendum on constituencies held by Labour in 2015.
Table (ref) provides a summary of these variables. It shows that 79.8% of constituencies had the same party win the elections in 2015 and 2019 while elections in 2015 were close within 5% in 8.7% of constituencies. 22.9% of constituencies voted to “Leave" and had Labour won in 2015 while overall, Labor won 35.7% of constituencies.
We estimate the semiparametric model in ((ref)) under Assumption (ref), notably using the median restriction in Assumption (ref) (iv). We set the vector of covariates $\widetilde{X}=(1,X_1,\ldots,X_5)'$ with the first element corresponding to the intercept and 5 non-constant covariates we outlined above. We normalize the vector of estimated coefficients $\widetilde{{\Greekmath 010B}}=({\Greekmath 010B}_0,{\Greekmath 010B}_1,-1, {\Greekmath 010B}_3,{\Greekmath 010B}_4,{\Greekmath 010B}_5),$ using a natural normalization for the coefficient of the indicator for the 2015 elections being close in a given constituency.
In our data, there are 23 unique discrete realizations of $\widetilde{X}$. For 2 elements in the sample support of $\widetilde{X}$, we have ${\mathbb P}(Y=1|\widetilde{X})=0.5$. For 4 elements, ${\mathbb P}(Y=1|\widetilde{X})<0.5$, and for the remaining 17 elements, ${\mathbb P}(Y=1|\widetilde{X})>0.5$.
We now use the estimators considered in Sections (ref) and (ref).
\paragraph{Maximum Score estimator}. Objective function ((ref)) of the maximum score estimator excludes the points in the support of covariates where ${\mathbb P}(Y=1| \widetilde{X})=0.5$. For the remaining 21 points in the support of the covariates we construct the following system of inequalities:
The set of solutions to this system contains all maximands of the objective function ((ref)). Notably, in our model and data, this system does indeed have a solution. Consequently, the maximum score estimator can be succinctly characterized by the following system of inequalities (with normalization $\widetilde{{\Greekmath 010B}}=({\Greekmath 010B}_0,{\Greekmath 010B}_1,-1,{\Greekmath 010B}_3,{\Greekmath 010B}_4,{\Greekmath 010B}_5)'$): $$({\Greekmath 010B}_0,{\Greekmath 010B}_1,{\Greekmath 010B}_3,{\Greekmath 010B}_4,{\Greekmath 010B}_5) \left(
\right) \geq c_1 $$ with $c_1=(0,0,0,0,0,0,1,1,0,0,0,0,0,0,1,1,1)$ and $$ ({\Greekmath 010B}_0,{\Greekmath 010B}_1,{\Greekmath 010B}_3,{\Greekmath 010B}_4,{\Greekmath 010B}_5) \left(
\right) < c_2$$ with $c_2=(1,1,1,1)$.
Estimate $\widehat{{\mathcal A}}_{ms}$ obtained from this system of inequalities has a non-empty interior in $\mathbb{R}^5$ and is contained in $[0, 1.5]\times [0, +\infty]\times [-0.5, 0.5] \times [0, +\infty) \times (-\infty,0] $.
We can provide a meaningful interpretation for the estimate $\widehat{{\mathcal A}}_{ms}$ in light of the fact that $\widetilde{X}=(1,X_1,\ldots,X_5)'$. In our model specification, the base group is non-Labour constituencies in 2015 that voted Remain. In Table (ref) we present raw joint counts of the outcome of the vote for Labour party in 2015 and the Brexit vote.
Given $X_2$ and $X_3$, the utility indices of the latent utilities $U^*={\Greekmath 010B}_0+{\Greekmath 010B}_1 X_1 -X_2 +{\Greekmath 010B}_3 X_3 +{\Greekmath 010B}_4 X_4 + {\Greekmath 010B}_5 X_5 + {\Greekmath 0122}$ manifest as follows:
Following from the directions of the effects (signs of ${\Greekmath 010B}_1$, ${\Greekmath 010B}_4$, ${\Greekmath 010B}_5$) that $U^*_{01} \geq U^*_{00}, \quad U^*_{10} \geq U^*_{00}, \quad U^*_{01} \geq U^*_{11}.$ System of inequalities for maximizers of ((ref)) additionally implies that ${\Greekmath 010B}_4+{\Greekmath 010B}_5 <0$ and ${\Greekmath 010B}_0+{\Greekmath 010B}_1+{\Greekmath 010B}_4+{\Greekmath 010B}_5 \geq 0$ thus giving $U^*_{11} < U^*_{10}, \quad U^*_{11} \geq U^*_{00}.$ The inequalities allow for either ${\Greekmath 010B}_1\geq {\Greekmath 010B}_4$ or ${\Greekmath 010B}_1\geq {\Greekmath 010B}_4$ thus leaving the relationship between $U^*_{01}$ and $U^*_{10}$ ambiguous.
To summarize, we conclude that given $X_2$, $X_3$, $U^*_{01} \geq U^*_{11} \geq U^*_{00}$ and $U^*_{10} > U^*_{11} \geq U^*_{00}$ with the inequality between $U^*_{01}$ and $U^*_{10}$ left undetermined.
At the level of utility indices, one can infer that our findings from the estimate $\widehat{{\mathcal A}}_{ms}$ partially align with the perspectives presented in the Guardian article. This partial alignment applies to comparisons between Labour 2015 & Leave vs. Labour 2015 & Remain, as well as Labour 2015 & Leave vs. non-Labour 2015 & Leave. However, it does not extend to the comparison between Labour 2015 & Leave and non-Labour 2015 & Remain, introducing nuances to the interpretation of the underlying relations.
\paragraph{Closed form estimator.} Next we consider the maximizer of ((ref)) and note that if the tuning parameter ${\Greekmath 0117}_n$ used to construct the weights is smaller than 0.1, then we can use only 2 elements with ${\mathbb P}(Y=1|\widetilde{X})\in [0.5-h_n, 0.5+h_n[$ for the estimator. This leads to two equalities (under the same normalization on $\widetilde{{\Greekmath 010B}}$ as above): $$ 0={\Greekmath 010B}_0+{\Greekmath 010B}_1-1+2{\Greekmath 010B}_3+{\Greekmath 010B}_4+{\Greekmath 010B}_5=0\;\;\mbox{and}\;\; 0={\Greekmath 010B}_0-1 $$ leading to estimate $\widehat{{\mathcal A}}_{CF} = \{({\Greekmath 010B}_0, {\Greekmath 010B}_1, {\Greekmath 010B}_3, {\Greekmath 010B}_4, {\Greekmath 010B}_5): {\Greekmath 010B}_0=1, \; {\Greekmath 010B}_1+2{\Greekmath 010B}_3 +{\Greekmath 010B}_4 +{\Greekmath 010B}_5=0\},$ which is a three-dimensional hyperplane in ${\mathbb R}^5$.
Thus, the closed-form estimator leaves ambiguous the directions and relative (to normalized ${\Greekmath 010B}_2=-1$) effects of $X_1$, $X_3$, $X_4$ and $X_5$ just indicating that their effects combined in a certain way sum up to 0.
\paragraph{Combination estimators.} We start with the estimator ((ref)) based on combination of the estimates $\widehat{{\mathcal A}}_{ms}$ and $\widehat{{\mathcal A}}_{CF}.$ With the tuning parameter ${\Greekmath 0117}_n$ less than 0.1, we get the combination estimate that coincides with $\widehat{{\mathcal A}}_{CF}$ and, thus, has properties as discussed above.
Next we consider estimator $\widehat{{\mathcal A}}_{OFC}$ maximizing ((ref)). While $\widehat{{\mathcal A}}_H$ opts between all equality constraints provided by maximizers of ((ref)) and all inequality constraints provided by maximizers of ((ref)), $\widehat{{\mathcal A}}_{OFC}$ combines both equality constraints and inequality information. In our application, for the nuisance parameter ${\Greekmath 0117}_n \in (0, 0.1) $ obtained estimate is expressed as: $\widehat{{\mathcal A}}_{OFC} = \left\{({\Greekmath 010B}_0, {\Greekmath 010B}_1, {\Greekmath 010B}_3, {\Greekmath 010B}_4, {\Greekmath 010B}_5): {\Greekmath 010B}_0=1, {\Greekmath 010B}_3=0, {\Greekmath 010B}_1+{\Greekmath 010B}_4+{\Greekmath 010B}_5=0 \right\}. $ This estimate reveals that there is no discernible impact (${\Greekmath 010B}_3=0$) of the change in economic well-being on the utility index for the re-election of the incumbent party, accounting for other factors.
Building on our earlier findings: $U^*_{01} \geq U^*_{11} $ and $U^*_{10} > U^*_{11}. $ These conclusions persist, but now the additional constraint ${\Greekmath 010B}_1+{\Greekmath 010B}_4+{\Greekmath 010B}_5=0$ also allows us to assert $U^*_{11} = U^*_{00}.$ This finding lends support, at least at the level of utility indices, to the proposition that Labour lost many constituencies due to the Leave vote within those constituencies. Given $X_2$ and $X_3$, the Leave vote in other constituencies resulted in a greater utility index, while the Remain vote across all constituencies had either the same or a greater utility index.
\paragraph{Random set quantile estimation.} We rely on our approach to a feasible sample quantile of random set outlines in the text. With this approach, there are four sets each occurring with a probability close to 0.25 (but strictly less than 0.25).
Each of these sets is a subset of $\widehat{{\mathcal A}}_{ms}$:
Taking closures of these sets, as well as closures of any other realization of $\widehat{{\mathcal A}}_{ms}$ in our sample (although other outcomes occur with negligible probabilities), the feasible ${\Greekmath 011C}$-th random set quantile for ${\Greekmath 011C} > 0.5$ must be in the intersection of at least three out of four sets $\overline{\widehat{{\mathcal A}}_1}$, \ldots, $\overline{\widehat{{\mathcal A}}_4}$. This leads to the set $\left\{({\Greekmath 010B}_0, {\Greekmath 010B}_1, {\Greekmath 010B}_3, {\Greekmath 010B}_4, {\Greekmath 010B}_5)': {\Greekmath 010B}_0=1, {\Greekmath 010B}_3=0, {\Greekmath 010B}_1+ {\Greekmath 010B}_4+{\Greekmath 010B}_5=0 \right\}$ as our ${\Greekmath 011C}$-random set quantile estimate.
Consequently, in this case, a feasible quantile random set estimate aligns with $\widehat{{\mathcal A}}_{OFC}$.
\paragraph{Probit/Logit.} Even though we consider model ((ref)) without parametric specification for the unobserved shock ${\Greekmath 010F}$, one might wonder about the implications of employing parametric estimation methods such as Probit or Logit.
Given that the system of inequalities corresponding to maximizers of ((ref)) has a solution, delineating a set of hyperplanes that perfectly separate two classes of points (those with ${\mathbb P}(Y=1| \widetilde{X})>1/2$ and those with ${\mathbb P}(Y=1| \widetilde{X})<1/2$), it can be demonstrated that Probit and Logit estimates must fall within the set $\widehat{{\mathcal A}}_{ms}$ after re-normalization to enforce ${\Greekmath 010B}_2=-1$. In our model and data, this alignment is evident, as illustrated in the Probit output in Table (ref). (The conclusions from the results from the Logit model are analogous; hence, Logit estimation is omitted for brevity.)
If we take Probit estimates, $\widehat{{\Greekmath 010B}}_1+\widehat{{\Greekmath 010B}}_4+\widehat{{\Greekmath 010B}}_5=-0.2307,$ giving a stronger support to the Guardian article statement regarding the impact of Leave vote on the 2019 Labour electoral performance. The test for the null hypothesis $H_0: {\Greekmath 010B}_1+{\Greekmath 010B}_4+{\Greekmath 010B}_5=0,$ based on Probit suggests does not reject the null ($p-value=0.1315$) which is is consistent with estimators $\widehat{{\mathcal A}}_{OFC}$ and $\widehat{{\mathcal A}}_{RSQ,{\Greekmath 011C}}$. However, $\widehat{{\mathcal A}}_{OFC}$ and random set quantile estimates are more explicit about the relationship ${\Greekmath 010B}_1+{\Greekmath 010B}_4+{\Greekmath 010B}_5=0$ whereas this message appears convoluted in the output of the Probit model.
\setcounter{equation}{0}
In this section we demonstrate that our analysis naturally extends to several classes of models including the binary choice model with independent noise as well as static and dynamic binary choice panel data models with fixed effects and discrete regressors.
Identification argument in \citeasnoun{manski1987} applies in the simplified setting where the error term ${\Greekmath 010F}$ in model ((ref)) is statistically independent. However, in case where regressors are continuous, parameters of the model can be estimated by exploiting additional information of independence, opening the possibility to construct estimators with properties superior to those of the maximum score estimator of \citeasnoun{manski1987}.
In this section we consider the performance of this estimator when regressors are discrete. We focus on model ((ref)) under the following stronger version of Assumption (ref) (iii): \\ {\bf Assumption (ref).} {\it (iii”) ${\Greekmath 0122} \perp \widetilde{X}$ and distribution density $f_{{\Greekmath 010F}}(\cdot)$ is above zero and strictly above $L>0$ in some fixed neighborhood of $0.$ } \\ Also, just as in Section (ref), we use parameter normalization ${\Greekmath 010B}^k=1$ for one $k=1, \ldots, K$, and denote ${\Greekmath 010B}$ as the non-normalized part of $\widetilde{{\Greekmath 010B}}$.
From the results in \citeasnoun{bierenshartog}, model ((ref)) under Assumption (ref) (i), (ii) and (iii”) is not point identified. The identified set ${\mathcal A}_0$ is fully characterized as a collection of parameter values ${\Greekmath 010B}$ satisfying
From this definition of ${\mathcal A}_0$ it is clear that depending on the value of the parameter of the DGP, the identified set can be a singleton or a non-singleton set. Namely, it will be a singleton if there are enough pairs of $\widetilde{x}$ and $\widetilde{x}^*$, $ \widetilde{x} \neq \widetilde{x}^* $, such that $ p(\widetilde{x}) = p(\widetilde{x}^*)$. to point identify parameter from the respective system of equations $(\widetilde{x}-\widetilde{x}^*)'\widetilde{{\Greekmath 010B}}=0$.
\citeasnoun{MRC} introduced the maximum rank correlation estimator for single index models which include our binary choice models ((ref)) under Assumption (ref) (iii') of independent random shock ${\Greekmath 0122}$. This estimator maximizes
\citeasnoun{MRC} established the properties of this estimator in single index model under a continuity of a regressor with non-trivial impact (WLOS, this is a regressor with a normalized component of $\widetilde{{\Greekmath 010B}}$). In this case the parameter in the DGP is identified up to a location and the MRC estimator of the parameter subvector excluding intercept is consistent.
In our binary choice model we can alternatively represent ((ref)) as
Let $\widehat{{\mathcal A}}_{MRC}$ denote the set of maximizers of ((ref)). For the convenience of the technical discussion, we will suppose that the index specification in our model does not contain intercept.
We now analyze the {\it sharpness} of $\widehat{{\mathcal A}}_{MRC}$ maximizing ((ref)). To do that we construct the population analog of the objective function ((ref)) using the same structure as in representation ((ref)):
We denote the maximizer of ((ref)) by ${\mathcal A}_{MRC},$
We note the similarity of objective ((ref)) to ((ref)) in the way both objectives treat the cases of exact equality. While the maximum score objective ((ref)) drops the terms where the choicen probability $p(\widetilde{x})$ is exactly $1/2$, objective ((ref)) drops the terms where choice probabilities $p(\widetilde{x})$ are the same in two different points in the support of $\widetilde{X}.$
To help clarify our discussion in this section, we introduce a discrete-only illustrative design and use it throughout this section.
In Illustrative Design 3, ${\mathcal A}_0=\{{\Greekmath 010B}_0\}$ whenever the parameter of the data generating process ${\Greekmath 010B}_0 \in \{-1,0, 1\}.$ Each of the three cases is based on the equalities $p((1,1)')=p((0,0)')$ or $p((1,0)')=p((0,0)')$ or $p((1,0)')=p((0,1)')$, respectively. For all other values of parameter ${\Greekmath 010B}_0$ in the DGP the model is partially identified. Namely, (a) ${\mathcal A}_0=(-\infty,-1)$ when ${\Greekmath 010B}_0<-1$; (b) ${\mathcal A}_0=(-1,0)$ when $-1<{\Greekmath 010B}_0<0$; (c) ${\mathcal A}_0=(0,1)$ when $0<{\Greekmath 010B}_0<1$; (d) and ${\mathcal A}_0=(1,+\infty)$ when ${\Greekmath 010B}_0>1$.
If ${\Greekmath 010B}_0 \not\in \{-1,0,1\}$ in the DGP, then ${{\mathcal A}}_{MRC}={\mathcal A}_0$. When ${\Greekmath 010B}_0 \in \{-1,0,1\}$, then ${{\mathcal A}}_{MRC}$ is a superset of ${\mathcal A}_0$ as ((ref)) ignores the cases $p(\widetilde{x})=p(\widetilde{x}^*)$, $\widetilde{x} \neq \widetilde{x}^*$, which are certainly relevant in the definition and formation of the identified set (as explained above, in our design in these cases the identified sets are singletons).
Namely, ${\mathcal A}_{MRC}=(-\infty,0) \cap {\mathcal A}$ when ${\Greekmath 010B}_0=-1$ (because $p((1,1)')=p((0,0)')$ is ignored by ((ref))), ${\mathcal A}_{MRC}=(-1,1) \cap {\mathcal A}$ when ${\Greekmath 010B}_0=0$ (because $p((1,0)')=p((0,0)')$ is ignored by ((ref))), and ${\mathcal A}_{MRC}=(0,+\infty)\cap {\mathcal A}$ when ${\Greekmath 010B}_0=1$ (because $p((1,0)')=p((0,1)')$ is ignored by ((ref))). Thus, analogously to the maximum score case, the infeasible population MRC maximizer is not sharp when ${\Greekmath 010B}_0$ in the DGP results in cases $p(\widetilde{x})=p(\widetilde{x}^*)$, $\widetilde{x} \neq \widetilde{x}^*$, and is sharp otherwise.
Consequently, for the feasible MRC estimator $\widehat{{\mathcal A}}_{MRC}$ in Illustrative Design 3 we expect a fluctuating behavior whenever ${\Greekmath 010B}_0 \in \{-1,0,1\}$ in the DGP, which is completely analogous to what we had in the maximum score estimation. Specifically, we can establish that
where $B$ is a Bernoulli random variable with parameter $1/2$.
The similarity of the structure of the distribution limit ((ref)) to that of the maximum score estimator, allows us to establish {\it robustness} of the estimator $\widehat{{\mathcal A}}_{MRC}.$ In particular, the selection of sequences ${\Greekmath 010B}(t,n;{\Greekmath 010B}_0)={\Greekmath 010B}_0+t/\sqrt{n}$ for ${\Greekmath 010B}_0 \in \{-1,0,1\}$ allows us to establish that $$ \sup\limits_{t>0}\lim\limits_{n \rightarrow \infty}{\mathbb P} \left( d_H(\widehat{{\mathcal A}}_{MRC},{\mathcal A}^L)=0 \right)= 1\;\;\mbox{and}\;\; \sup\limits_{t<0}\lim\limits_{n \rightarrow \infty}{\mathbb P} \left( d_H(\widehat{{\mathcal A}}_{MRC},{\mathcal A}^R)=0 \right) = 1, $$ where ${\mathcal A}^L$ is the identified set whenever parameter of the data generating process takes values above ${\Greekmath 010B}_0$ and ${\mathcal A}^R$ is the identified set whenever parameter of the data generating process takes values below ${\Greekmath 010B}_0$ in some bounded neighborhood of ${\Greekmath 010B}_0.$
Taking our discussion from Illustrative Design 3 to general cases, we will be able to conclude that $\widehat{A}_{MRC}$ will not have a probability limit but will have a weak limit representing a mixture of sets whenever there are cases $p(\widetilde{x})=p(\widetilde{x}^*)$, $\widetilde{x} \neq \widetilde{x}^*$, and will be sharp otherwise.
We next take ideas of the closed form estimation (inspired by \citeasnoun{ahnetal2018}) to our case of independent errors. It is based on similar principles as the maximizer of ((ref)) and relies on equality restrictions in identification condition ((ref)) as they are most important from the perspective of identifying power. Specifically, the estimator is defined as the set of maximizers of
with $w_{ij}={\bf 1}\left\{ |\widehat{p}(\widetilde{x}_i)-\widehat{p}(\widetilde{x}_j)|<{\Greekmath 0117}_n \right\}$ for the tuning parameter ${\Greekmath 0117}_n \rightarrow 0.$ If $w_{ij}$ for any pair $\widetilde{x}_i$ and $\widetilde{x}_j$, $\widetilde{x}_i \neq \widetilde{x}_j$, then the estimator is taken to be the entire parameter space ${\mathcal A}$.
We can establish that the maximizer of ((ref)) is sharp only when ${\Greekmath 010B}_0$ in the DGP results in cases of point identification (thus, there have to be “enough” cases ((ref))), as is not sharp otherwise. As for robustness, we can establish that it is robust in the absence of equalities ((ref))) for distinct $\widetilde{x}, \widetilde{x}^*$, and is not robust otherwise. E.g., in Illustrative Design 3, the CF estimator is only when ${\Greekmath 010B}_0 \in \{-1,0,1\}$, and is only robust when ${\Greekmath 010B}_0 \notin \{-1,0,1\}$. In general, one can have cases when the closed form estimator is neither sharp nor robust. These results are completely analogous to those in our main model under the median independence.
Similarly to Section (ref), we can construct estimators based on combination ideas. We can combine the MRC and the closed form estimator directly based on whether $p(\widetilde{x})$ is close for two points in the support of the distribution of covariates $\widetilde{X}$, which will serve as a “switching device”. This estimator will have same drawbacks as the analogous estimator in Section (ref) -- it will not be universally sharp (it will only be sharp for ${\Greekmath 010B}_0$ in the DGP when either the MRC or closed form estimators are sharp) and will not be generally robust either. The second approach combines objective functions ((ref)) and ((ref)) and considers
for ${\Greekmath 0117}_n \to 0$, $n{\Greekmath 0117}_n^2 \to \infty$. We can establish that this estimator is sharp. However, they it it not robust for ${\Greekmath 010B}_0$ in the DGP that results in cases ${p}(\widetilde{x})={p}(\widetilde{x}^*)$ for $\widetilde{x} \neq \widetilde{x}^*$.
Finally, we also can construct the estimator based on the idea of the quantile of a random set in \citeasnoun{molchanov2006book}. This estimator exploits the distribution limit ((ref)) of the estimator $\widehat{{\mathcal A}}_{MRC}$. E.g. in Illustrative design 3, when ${\Greekmath 010B}_0 \in \{-1,0,1\}$, the MRC selects the sets above and below the true parameter of the data generating process with equal probabilities (see details in ((ref))). As a result, selecting an estimator which for $\Delta>0$ outputs ${\Greekmath 011C}=1/2+\Delta$-quantile of the closure of the random set $\widehat{{\mathcal A}}_{MRC},$ expressed as $q_{{\Greekmath 011C}}\left(\overline{\widehat{{\mathcal A}}_{MRC}}\right),$ produces an estimator which is sharp over the entire parameter space. Analogously to Section (ref), we can establish that universal sharpness extends to general cases, and that the estimator is universally robust too.
Panel setting extends the cross-sectional model ((ref)) by assuming the presence of the fixed effects with unknown distribution impacting the random utility while also allowing multiple observations over time available for each cross-sectional unit. The version of this model where the error terms follow a logistic distribution with time-varying covariates have been analyzed in \citeasnoun{anderson70},\footnote{Interestingly, recent work demonstrates how crucial the logistic distribution assumption is for point identification and consistent estimation. \citeasnoun{chamberlain2010} shows that when the observable covariates have bounded support, the logistic assumption on an unobserved component is {\em necessary } for point identification.} which proved that a conditional maximum likelihood estimator consistently estimates the model parameters up to scale without additional assumptions about the fixed effects.
\citeasnoun{manski1987} considered the following semiparametric version of the model considered in \citeasnoun{anderson70}:
where $i=1,2,...n$ are the cross-sectional units and $ t=1,2$ are the time periods. The binary variable $Y_{it}$ and the $K$-dimensional regressor vector $X_{it}$ are each observed and the parameter of interest is the $K$ dimensional vector $\widetilde{{\Greekmath 010B}}$. The variables not observed in the data are $c_i$, and ${\Greekmath 010F}_{it}$, the former not varying with $t$ and often referred to as the “fixed effect" or the individual specific effect. \citeasnoun{manski1987} imposed no specific distributions on unobservables. Under scale normalization for $\widetilde{{\Greekmath 010B}},$ unbounded support and continuity of distribution of at least one of the covariates, and only additionally maintaining the assumption that error terms for each cross-sectional unit are known only to be time-stationary with unbounded support, \citeasnoun{manski1987} established point-identification of coefficients $\widetilde{{\Greekmath 010B}}.$
Our framework for model ((ref)) directly extends to model ((ref)) in which we consider the setting of all discrete regressors. We maintain Assumption (ref) regarding the structure of the data generating process.
Assumption (ref) (iii) deviates from the continuity assumption of \citeasnoun{manski1987} and, ultimately, leads to a loss of point identification of parameter $\widetilde{{\Greekmath 010B}}.$ We illustrate below that, closely following the case of the cross-sectional model ((ref)) the identified set can either be a singleton or a non-singleton convex set.
The identified set for parameter $\widetilde{{\Greekmath 010B}}_0$ is characterized as the set of solutions in $\mathcal{A}$ to
Suppose in Illustrative Design 4 ${\Greekmath 010B}_0=1$ in the DGP. When $x_{i1}=(0,1)'$ and $x_{i2}=(1,0)'$ (or the other way around), we are in situation ((ref)) which leads us to the equalities identifying ${\Greekmath 010B}$ (we would have the same equalities if we additionally conditioned on $c_i$) as they lead to $({x}_{i2}-{x}_{i1})'(1,{\Greekmath 010B})'=1-{\Greekmath 010B}=0,$ thus meaning that the identified set is $\mathcal{A}_0=\{1\}$ (it is also consistent with inequalities ((ref)). If e.g. ${\Greekmath 010B}_0=1/2$ in the DGP, then the only conditions defining the identified set are inequalities in the form of ((ref)) and we get ${\mathcal A}_0=(0,1)$.
Similarly to our prior analysis we can evaluate the performance of the classic estimators for model ((ref)) under Assumption (ref). We start with the classic maximum score estimator and express it for model ((ref)) as the maximizer of the objective function of the {\it conditional maximum score} proposed by \citeasnoun{manski1987}:
The corresponding population objective function takes the form
We can see that the sum ((ref)) completely ignores (zeros out) cases when $x_{i2} \neq x_{i1}$ and $P\left(Y_{i2}=1 |{X}_i={x}_i\right) = P\left(Y_{i1}=1 | {X}_i={x}_i\right)$, which, as seen from ((ref))-((ref)), are essential in shaping the identified set. This already indicates an issue similar to that of the population maximum score objective function, which in its turn ignored or zeroed out cases of conditional choice probabilities equal to $1/2$. Thus, we conclude that the maximizer of ((ref)) will in general be a superset of the identified set ${\mathcal A}_0$.
In Illustrative Design 4, let us once again take ${\Greekmath 010B}_0=1$ in the DGP. We note that ((ref)) will have non-zero entries in the following points in the support of regressors: (i) ${x}_{i{\Greekmath 011C}}=(1,1)'$, ${x}_{i,{\Greekmath 011C}'}=(0,0)';$ (ii) ${x}_{i{\Greekmath 011C}}=(1,1)'$, ${x}_{i,{\Greekmath 011C}'}=(1,0)';$ (iii) ${x}_{i{\Greekmath 011C}}=(1,1)'$, ${x}_{i,{\Greekmath 011C}'}=(0,1)';$ (iv) ${x}_{i{\Greekmath 011C}}=(0,0)$, ${x}_{i,{\Greekmath 011C}'}=(1,0);$ (v) ${x}_{i{\Greekmath 011C}}=(0,0)$, ${x}_{i,{\Greekmath 011C}'}=(0,1),$ for $({\Greekmath 011C},{\Greekmath 011C}')=(1,2)$ or $(2,1).$ This objective function is maximized for $\widetilde{{\Greekmath 010B}}=(1,{\Greekmath 010B})$ which satisfies inequalities: ${\Greekmath 010B}>0,$ and $1+{\Greekmath 010B}>0.$ As a result, the set of maximizers of ((ref)) (over ${\mathcal A}$) is $(0,+\infty) \cap {\mathcal A}$ which is a superset of the identified set ${\mathcal A}_0=\{1\}$. In Appendix (ref) we show that the maximizer of the sample objective function ((ref)) for the simple design considered here, just like in case of the cross-sectional semiparametric discrete choice model, converges to a random set when model ((ref)) under Assumption (ref) has cases $P\left(Y_{i2}=1 |{X}_i={x}_i\right) = P\left(Y_{i1}=1 | {X}_i={x}_i\right)$ for $x_{i2} \neq x_{i1}$. This means that the maximum score estimator maximizing ((ref)) is not universally sharp.
The two-step closed form estimator, similar to that constructed in \citeasnoun{ichimura1994}, relies exclusively on conditions ((ref)) for $x_{i2} \neq x_{i1}$, unlike the \citeasnoun{manski1987} maximum score estimator. To write its formal objective function, it is convenient to denote $\Delta p_{12}(x_i)={\mathbb E}\left[ Y_{i2}-Y_{i1} \,|\,X_i=x_i\right]$, and $\widehat{\Delta p_{12}}(x_i)$ denote its sample analogue. The objective function for the estimator is constructed by selecting observations with small values of $\widehat{\Delta p_{12}}(x_i):$
where $w_i={\bf 1}\{|\widehat{\Delta p_{12}}(x_i)| <{\Greekmath 0117}_n\}$. Sequence ${\Greekmath 0117}_n \rightarrow 0$ is the tuning parameter for this estimator. The only cases informative in ((ref)) for the parameter value $\widetilde{{\Greekmath 010B}}$ will be those with $|\widehat{\Delta p_{12}}(x_i)| <{\Greekmath 0117}_n$ and $x_{i2}-x_{i1} \neq 0$. The default value of the estimator is the whole parameter space $\mathcal{a}$ if $|\widehat{\Delta p_{12}}(x_i)| <{\Greekmath 0117}_n$ either does not hold or only holds when $x_{i1}=x_{i2}$. The reliance of the maximizer of ((ref)) on the near zero value of $|\widehat{\Delta p_{12}}(x_i)|$ for some $x_i$ in $\mathcal{X}$ demonstrates that for the semiparametric binary choice model in the panel data settings we observe the same behavior of the closed form estimator as we did for the closed form estimator in the cross-sectional model in Section (ref). Namely, this estimator is not sharp universally on the parameter space and is not generally robust.
We can proceed to construct new estimators based on combination ideas and analogous to those constructed for the cross-sectional setting in Section (ref). We can also construct the random set quantile estimator similar to that in Section (ref).
The behavior of the maximum score estimator and the closed form estimator is completely analogous to that in the cross sectional settings. Namely, neither the conditional maximum score nor the closed form estimator estimators is sharp, though the conditional maximum score estimator is robust. The estimator that combines the conditional maximum score and the closed form estimators directly using a switching device would be neither universally sharp nor robust. The second combination estimator would modify the maximum score objective function ((ref)) and consider
where is the summation is over $({x}_{i1},{x}_{i2})$ in the sample support, and ${\Greekmath 0117}_n \to 0$. Analogously to the cross-section case in Section (ref) this combination estimator is universally sharp but not universally robust.
Finally, the random set quantile estimator is both sharp and robust following from the convergence of the maximizer of the conditional maximum score objective to a random set which is a mixture of deterministic sets each of whom contains the identified set ${\mathcal A}_0$ on its boundary.
An important extension of model ((ref)) adds the dependence of the random utility on the lagged realizations of the discrete outcome. A simple form of that model has only dependence on one lag and takes the form
with $t=1,\ldots,T$ and $i=1,\ldots,n.$ Note that model ((ref)) differs from model ((ref)) by one additional term and, respectively, one additional parameter ${\Greekmath 010D}$ which needs to be recovered from the data. Just as in model ((ref)), the unobserved components are represented by (a) a time invariant fixed effect $c_i$ which captures the systematic correlation of the unobservables over time; (b) an idiosyncratic error term ${\Greekmath 010F}_{it}$ which is randomly sampled both over time and the cross-sectional units. The parameter ${\Greekmath 010D}$ is of special interest as it measures the effect of state dependence in the model.
There is a rich literature in econometric theory which studies ((ref)) and develops estimators for its parameters under different conditions on the distribution of covariates $X_{it},$ fixed effect $c_i$ and the random shock ${\Greekmath 010F}_{it}.$ In this section we focus on the setting considered in \citeasnoun{honorekyriazidou}, which proposed an estimator based on a conditional maximum score objective function\footnote{To our knowledge, this was the first paper to consider semiparametric identification and estimation of a dynamic binary choice panel data model. While they focused exclusively on point identification, other recent work also considers partial identification as we do here. See for example \citeasnoun{kpt2022} and references therein, and subsequently \citeasnoun{crz2023}, and \citeasnoun{gaowang2023}. These other approaches do discuss sharpness but not robustness as we do here. There is also a very recent literature under {\em parametric} and serially independent assumptions on ${\Greekmath 010F}_{it}$, notably an i.i.d {\em logit} assumption - see \citeasnoun{kitazawa2022}, \citeasnoun{dobronyietal2023}, \citeasnoun{honoreweidner2023}, and references therein. These approaches attain moment conditions that rely on the functional form of the logistic distribution.} The identification argument there relies on the presence of at least 3 time periods (with 4 periods of outcome observations) and the overlapping support of regressors such that one can find observations where those values are close in periods 2 and 3. Then one can verify the sign condition similar to ((ref))-((ref)), conditional on those close observations across periods.
In this section we consider the setting where assumption that regressors have continuous distribution used in \citeasnoun{honorekyriazidou} does not hold and instead use Assumption (ref) with the following modification.
{\bf Assumption (ref).} {\it (i') $\left\{\left((y_i \equiv (y_{i0},y_{i1},y_{i2},y_{i3}),x_i \equiv (x_{i1,x_{i2}},x_{i3}) \right) \right\}^{n}_{i=1}$ is an i.i.d. random sample from the joint distribution $(Y_{i}\equiv(Y_{i0},Y_{i1},Y_{i2},Y_{i3}),X_i \equiv (X_{i1},X_{i2},X_{i3}))$ induced by ((ref)) for some $(\widetilde{{\Greekmath 010B}}_0,{\Greekmath 010D}_0) \in {\mathcal A} \subset {\mathbb R}^{K+1}$ and some distribution of initial values $Y_{i0}.$ \\ (iii') Distribution of regressor vector ${X}_i$ is discrete with the support $\mathcal{X} \subset {\mathbb R}^{3K}$. The support of ${X}_{i2}-{X}_{i1}$ does not lie in any proper linear subspace of ${\mathbb R}^K$ and ${\mathbb P}(X_{i2}=X_{i3})>0.$ }
The objective function in \citeasnoun{honorekyriazidou} is
The corresponding population objective function can be written as
We expositional convenience we will consider an simplified design given in Definition (ref).
In Illustrative Design 5 let ${\Greekmath 010B}_0=-1$ and ${\Greekmath 010D}_0=-1$ in the DGP. We show that in this case the identified set is a singleton.
To establish this, consider the following events defined by values $d_0,\,d_3 \in \{0,1\}$:
Then
If we consider values $x_{i1}= (0,0)'$ and $x_{i2}=x_{i3}=(0,1)'$, respectively, then $x_{i1}'\widetilde{{\Greekmath 010B}}=0,$ $x_{i2}'\widetilde{{\Greekmath 010B}}=1,$ and $x_{i3}'\widetilde{{\Greekmath 010B}}=1.$ As a result, $$
$$
This implies immediately the identifiability restriction $(x_{i2}-x_{i1})'\widetilde{{\Greekmath 010B}} + {\Greekmath 010D}\,(y_{i3}-y_{i0})=0$ which takes the form $ 1+{\Greekmath 010D}=0.$ Analogously, $$
$$ This implies the identifiability restriction $ -{\Greekmath 010B}+{\Greekmath 010D}=0.$ The combination of equations $1+{\Greekmath 010D}=1$ and $-{\Greekmath 010B}+{\Greekmath 010D}=0$ allows us to conclude that the identified set is a singleton $\{((-1,1,-1)'\}$.
We now construct the set of maximizers of the {\it population} objective function ((ref) under the normalization $\widetilde{{\Greekmath 010B}}=({\Greekmath 010B},1)'$. The elements in the sum in ((ref)) will be zero for $x_i$ such that $x_{i2}=x_{i3}$ if $${\mathbb P}(A(d_0, d_3) | c_i, X_i=x_i) = {\mathbb P}(B(d_0, d_3) | c_i, X_i=x_i). $$ This is because for those values of covariates $$
$$ thus resulting in $${\mathbb E}\left[Y_{i2}-Y_{i1} \big| X_i=x_i, Y_{i0}=d_0, Y_{i3}=d_3\right]=0.$$
This is analogous to the behavior of the maximum score objective function both in the cross sectional case ((ref)) and in the case of static panel data model ((ref)) where population objective function drops the terms that are most important in shaping the identified set and will give point identification in some cases (such as in case ${\Greekmath 010B}_0=1$, ${\Greekmath 010D}_0=-1$ in Illustrative Design 5).
To further simplify, we assume that fixed effect $c_i$ are binary. Then we have the following system of inequalities defining the maxmimzer of the population objective function corresponding to : $ {\Greekmath 010B}+{\Greekmath 010D} <-1$, ${\Greekmath 010B}+{\Greekmath 010D} <1,$ $ {\Greekmath 010B}-{\Greekmath 010D} <1,$ $ -{\Greekmath 010B}+{\Greekmath 010D}<1.$ Removing redundant inequalities leaves three relevant inequalities ${\Greekmath 010B}+{\Greekmath 010D} <-1,$ ${\Greekmath 010B}-{\Greekmath 010D} <1,$ $-{\Greekmath 010B}+{\Greekmath 010D}<1.$
The solution set to these inequalities can be described as
It includes the point $({\Greekmath 010B}_0=-1,{\Greekmath 010D}_0=-1)$ as an interior point (take $t=-1/2$, ${\Greekmath 0115}=2$). Thus, the maximizer of the population Honore-Kyriazidou objective function is a (unbounded) superset of the identified set. This maximizer is displayed in Figure (ref).
The maximizer of the sample objective function ((ref)), analogously to the cross sectional case, will fluctuate among a finite number of disjoint sets with probabilities bounded away from zero and one. Each of such sets will have the true parameters ${\Greekmath 010B}_0=-1$ and ${\Greekmath 010D}_0=-1$ of the DGP at the boundary and the union of all such sets will give the described above maximizer of the population objective function (again, analogous to our illustrative example in the cross sectional case).
We can also construct the closed form estimator by identifying points in the support of covariates which satisfy $x_{i2}=x_{i3}$ and for some fixed values of the outcomes in the initial and the last period gives equal conditional probabilities of choice in periods 1 and 2: $P\left(Y_{i2}=1 \big| X_i=x_i, Y_{i0}=y_{i0}, Y_{i3}=y_{i3}\right) =P\left(Y_{i1}=1 \big| X_i=x_i, Y_{i0}=y_{i0}, Y_{i3}=y_{i3}\right)$. With discrete covariates, both of those probabilities can be estimated as corresponding sample averages. The parameters can be estimated by minimizing with respect to $(\widetilde{{\Greekmath 010B}},{\Greekmath 010D})$ the quadratic objective that consists of terms $ \left((x_{i2}-x_{i1})'\widetilde{{\Greekmath 010B}}+{\Greekmath 010D}\,(y_{i3}-y_{i0})\right)^2 $ multiplied by weights that are non-zero only if $x_{i2}=x_{i3}$ and $\widehat{P}\left(Y_{it}=1 \big| X_i=x_i, Y_{i0}=y_{i0}, Y_{i3}=y_{i3}\right), t=1,2$, are close to each other. As before, the estimator would output the entire parameter space if no points with properties described above are found.
Similar to previous models, we can construct new estimators based on combination of the output and combination of objectives of the conditional maximum score estimator and the closed-form two-step estimator.
We can once again show that the random set quantile of the feasible estimator for quantile indices strictly above $1/2$ will give the true parameter value ${\Greekmath 010B}_0=-1$ and ${\Greekmath 010D}_0=-1$ for large enough sample sizes.
All our estimation techniques extend from Illustrative Design 5 to a general case of discrete regressors. Similar to our previous observation in Section (ref) these observations allow us to conclude that the maximizer of the conditional maximum score objective function ((ref)) will produce an estimator which is not sharp on the entire parameter space but robust. In contrast, the closed form estimator and the combination estimator that directly conditional maximum score and closed form estimator are neither sharp nor robust. The estimator that combines the objective functions of the conditional maximum score and closed form estimator is sharp everywhere on the parameter space but not robust. Finally, the random set quantile estimator will be both sharp and robust on the entire parameter space.
In this paper we study semiparametric discrete choice models when covariates are discrete, violating the assumption of the continuity of distribution of regressors required to establish point identification. The parameters of the model are generally only partially identified. However, depending on the true value of the parameters of the data generating process the identified set can be a singleton or a non-singleton set. We focus on the question if this behavior of the identified set is accurately captured by a given estimator. We propose two criteria for evaluation of a given estimator. Sharpness of an estimator for a given true parameter of the data generating is the property where it converges in probability to the identified set corresponding to the true parameter value. Robustness of an estimator is the property of continuity of the distribution limit of an estimator with respect to local changes in the parameter of the data generating process. We explore classic estimators for discrete choice models including the maximum score estimator and the closed-form two-step estimator. We find that the maximum score estimator is not sharp everywhere on the parameter space, however, it is robust everywhere. In contrast, the closed form estimator is sharp only in parts of the parameter space where the model is point identified and is not robust at those points.
To construct estimators with better properties, we consider a direct combination of the two classic estimators using a “switching” device as well a combination of their objective functions. The latter combination idea allows us to construct estimators which are sharp, and this is not true for the latter combination idea. Both combination estimators fail to be robust.
We then propose a novel class of estimators based on the concept of a quantile of a random set. It uses the random set output by the maximum score estimator and produces an estimator which is both sharp and robust on the entire parameter space. We illustrate the performance of our novel estimator and compare it with classic estimators by analyzing the impact of the Brexit referendum vote on outcomes of the 2019 UK General Elections. Our new estimator both performs better and provides more meaningful results than the alternatives. We also show that our framework extends to other important settings including the maximum rank correlation estimator for the discrete choice model under independence as well as static and dynamic panel data models with discrete outcomes.