EconBase
← Back to paper

New possibilities in identification of binary choice models with fixed effects

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

80,040 characters

New possibilities in identification of binary choice models with fixed effects


\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\global\long
\title{New possibilities in identification of binary choice models with fixed
effects}
\author{Yinchu Zhu\thanks{Email: [email removed]. I am extremely grateful to Whitney
Newey for introducing me to the literature on fixed effects in nonlinear
models and for having many inspiring discussions with me. I thank
Roy Allen, Whitney Newey and Martin Mugnier for pointing out mistakes
in previous versions and for providing helpful comments on the paper.
I am also grateful for comments from Xavier D'Haultf{\oe}uille, Shakeeb
Khan, Elie Tamer and participants of various seminars and conferences.
All the errors are my own. }\\
\\
Department of Economics, \\
Brandeis University}
\maketitle
\begin{abstract}
We study the identification of binary choice models with fixed effects.
We propose a condition called sign saturation and show that this condition
is sufficient for identifying the model. In particular, this condition
can guarantee identification even when all the regressors are bounded,
including multiple discrete regressors. We also establish that without
this condition, the model is not identified unless the error distribution
belongs to a special class. Moreover, we show that sign saturation
is also essential for identifying the sign of treatment effects. Finally,
we introduce a measure for sign saturation and develop tools for its
estimation and inference.
\end{abstract}
Key words: identification, panel model, binary choice, fixed effects

\section{Introduction}

This paper considers panel models with binary outcomes in the presence
of fixed effects. These models are convenient in economic analysis
as they allow for fairly general unobserved individual heterogeneity.
The nonlinear nature of the binary choice models makes it difficult
to eliminate the fixed effects by differencing. Since we view the
fixed effects or their conditional distribution as a nuisance parameter,
the identification of the parameter of interest becomes tricky. In
this paper, we provide new results and insights on what drives the
identification of the coefficients on the regressors and what this
means for applied work.

Consider independent and identically distributed (i.i.d) observations
$\{(Y_{i},X_{i})\}_{i=1}^{n}$ with $Y_{i}=(Y_{i,0},Y_{i,1})$ and
$X_{i}=(X_{i,0},X_{i,1})$ from the following model
\[
Y_{i,t}=\mathbf{1}\{X_{i,t}'\beta+\alpha_{i}\geq u_{i,t}\}\qquad t\in\{0,1\},
\]
where the fixed effects are represented by the scalar variable $\alpha_{i}$
and $\beta$ is a non-random vector of coefficients. This is a semiparametric
model as the distribution of $\alpha_{i}$ given $X_{i}$ is unrestricted.
For notational simplicity, we drop the $i$ subscript in the rest
of the paper. Therefore, we write $Y=(Y_{0},Y_{1})$ and $X=(X_{0},X_{1})\in\mathcal{X}_{0}\times\mathcal{X}_{1}$
with
\begin{equation}
Y_{t}=\mathbf{1}\{X_{t}'\beta+\alpha\geq u_{t}\}\qquad t\in\{0,1\}.\label{eq: FE model}
\end{equation}

In this paper, we mainly focus on the identification of $\beta$ but
we will also discuss quantities related to treatment effects. There
are roughly two approaches to identifying $\beta$, depending on whether
or not we impose a parametric model on the distribution of $u_{t}$
given $(X,\alpha)$. Methods without parametric assumptions on the
error distribution are typically based on \citet{manski1987semiparametric}
and maximum-score-type estimation. This approach only assumes that
the distribution $u_{t}\mid(X,\alpha)$ does not depend on $t$. In
contrast, the more parametric approach relies on additional assumptions
on the functional form of the distribution $u_{t}\mid(X,\alpha)$.
For example, perhaps the most popular parametric assumption is that
$u_{t}$'s are i.i.d logistic errors across $t$ and are independent
of $(X,\alpha)$. This approach typically adopts an estimation scheme
based on the (conditional) likelihood.

Despite the strong restrictions on the functional form, the approach
relying on parametric restrictions has received considerable attention
arguably due to identification reasons. The literature has pointed
out the widespread identification failure, such as \citet{arellano2011nonlinear}.
In particular, identification seems infeasible outside the logistic
case unless the support of $X$ is unbounded. For example, Assumption
2 in \citet{manski1987semiparametric} requires the unboundedness
of at least one component of $W=X_{1}-X_{0}$: for $W=(W_{1},...,W_{K})'$
and $\beta=(\beta_{1},...,\beta_{K})'$, there exists $k$ such that
$\beta_{k}\neq0$ and the conditional distribution $W_{k}\mid(W_{1},...,W_{k-1},W_{k+1},...,W_{K})$
has support equal to $\mathbb{R}$ almost surely. Theorem 1 of \citet{chamberlain2010binary}
goes even further: for bounded $X$, the identification fails in certain
regions of the parameter space if no further restrictions are imposed
on the distribution of the error term $u_{t}$.
\begin{comment}
The literature has pointed out the widespread identification failure,
such as \citet{arellano2011nonlinear}.
\end{comment}

In this paper, we provide a more precise picture than \citet{chamberlain2010binary}
by showing that identification without parametric assumptions on the
error distribution is quite possible even for bounded $X$. The main
motivation of this paper is to explore new possibilities without functional-form
assumptions. This raises concerns about stability. If identification
crucially hinges on the imposed functional form, the reliability of
the analysis could be a concern. After all, if the model parameter
is unidentified under every distribution function other than the logistic
one, how much should we trust this model? Hence, it is helpful to
explore a more robust setting. In this direction, this paper considers
the identification issue and leaves the estimation problem to future
research. We make the following contributions.

First, we provide tight and simple identification conditions without
parametric restrictions on the error distribution. We show that a
sufficient condition for identification is what we refer to as sign
saturation (Assumption \ref{assu: sign saturation}), which states
that $P\left(E(Y_{1}-Y_{0}\mid X)>0\right)$ and $P\left(E(Y_{1}-Y_{0}\mid X)<0\right)$
are both strictly positive. This condition can hold with bounded regressors
$X$ and allow for multiple discrete regressors and interaction terms.
Moreover, we show that this is also a necessary condition for identification
unless the distribution of $u_{t}$ is in a special class. Therefore,
although we know (from \citet{chamberlain2010binary}) that for bounded
regressors the identification fails at some values of $\beta$, our
results pinpoint these values: the identification fails exactly at
points where the sign saturation fails. Therefore, whether the regressors
are bounded is not what really drives the identification. The identifiability
is more closely related to the sign saturation condition, which is
simple and intuitive.

A key advantage of the proposed sign saturation condition is that
it is stated in terms of observed variables only and does not involve
the parameter that we are trying to identify. Hence, it can be used
to answer the question of whether identification holds under the distribution
of observed data, thereby making identification directly testable.
This is different from the \citet{manski1987semiparametric}-type
condition. For example, suppose that $K=2$, $W_{1}$ is bounded and
the conditional distribution $W_{2}\mid W_{1}$ has support $\mathbb{R}$
almost surely. The identification condition in \citet{manski1987semiparametric}
becomes $\beta_{2}\neq0$. However, how can we check $\beta_{2}\neq0$
when the identifiability of $\beta$ is not yet established? It does
not seem obvious how to do so with the data. In contrast, under the
sign saturation condition, we check statements \textit{only} on observed
data: $P\left(E(Y_{1}-Y_{0}\mid X)>0\right)$ and $P\left(E(Y_{1}-Y_{0}\mid X)<0\right)$.

Second, we show that the sign saturation condition is also sufficient
and necessary for learning the sign of the marginal effects. Even
though the magnitude of different notions of the marginal effects
is usually not point identified, the sign of the effects is typically
the same as the sign of a component of $\beta$. It turns out that
outside a special class of distributions, if the sign saturation condition
fails, we cannot guarantee to distinguish between zero effects and
strictly positive effects. In many empirical studies, one important
task is to check whether the identified set for the effects includes
zero. Therefore, it is worthwhile to check the sign saturation condition
before exploring additional assumptions that give bounds to treatment
effects.

Third, we introduce a tool that can be used to assess the sign saturation
condition empirically. We show that $E\max\{P(Y_{1}-Y_{0}\mid X),0\}$
and $E\min\{P(Y_{1}-Y_{0}\mid X),0\}$ can be directly estimated from
the data and confidence intervals can be constructed by a simple bootstrap
without any tuning parameters such as bandwidth or number of basis
functions. For bounded regressors, the existing literature highlights
the risk of identification failure: the identification possibly (not
definitely) fails. This generic warning may have discouraged the application
of fixed effect models in practice. The proposed tool serves as a
more accurate diagnosis so that the identifiability can be assessed
from the data.
\begin{comment}
It turns out that the computation is already familiar to applied researchers
in this area as it only involves computing the usual maximum score
estimator and its bootstrapping. We are not aware of any existing
tools for checking identification in this model.
\end{comment}
\begin{comment}
One important objective of our work is to provide robust tools for
empirical research. Although conditional maximum likelihood or any
other approach based on logistic errors is popular in applied work,
we argue that this assumption can be problematic in applications.
If so, it might be prudent to leave the error distribution nonparametric.
Since estimation and inference based on maximum score estimates are
already available (e.g., \citet{kim1990cube}, \citet{seo2018local}
and \citet{cattaneo2020bootstrap}), we only need to make sure that
we can reasonably assume the identification. This paper fills this
very need: tight and testable conditions for identification.
\end{comment}

Our work is closely related to the literature of binary response models
with logistic errors. \citet{rasch1960probabilistic}, \citet{andersen1970asymptotic}
and \citet{chamberlain1980analysis} consider the estimation and inference
in this case. The logistic link function is also singled out in \citet{chamberlain2010binary}
as the ``nice'' link function: the identification does not require
unbounded regressors. Recently, the work of \citet{Mugnier2009.08108}
points out that if there are more than 2 time periods, then the nice
link function can be extended to a generalized version of the logistic
distribution. This suggests that imposing a special class of parametric
structures is a useful strategy of achieving identification and this
special class seems to be the logistic functions. We provide new insight
on this. We show that this special class also includes those such
that $\dot{G}(\cdot)$ is periodic, where $\dot{G}(\cdot)$ is the
derivative of $\ln\frac{F(\cdot)}{1-F(\cdot)}$ and $F(\cdot)$ is
the distribution function of $u_{t}$. If $F(\cdot)$ is logistic,
then $\dot{G}(\cdot)$ is a constant function, which is clearly periodic.
However, we find that identification is possible for any periodic
$\dot{G}(\cdot)$. Although \citet{chamberlain2010binary} tells us
that for non-logistic distributions, identification fails in an open
neighborhood, we show that for any periodic $\dot{G}(\cdot)$, identification
must hold in another open neighborhood. Therefore, it seems to us
that the truly problematic distributions are those with non-periodic
$\dot{G}(\cdot)$. Indeed, outside this special class of periodic
$\dot{G}(\cdot)$, we show that identification is impossible at \textit{every}
point if the sign saturation condition is not satisfied. Thus, for
generic distribution functions, sign saturation is a necessary condition
for identification. We summarize the relation between identifiability
and error distribution in Table \ref{tab: ID and sign sat}.
\begin{comment}
Logistic errors are also assumed in the study of identification in
dynamic models, including recent works of \citet{Honore2005.05942},
\citet{khan2020identification} among many others.
\end{comment}

Our work is also closely related to works that do not assume logistic
errors. First, some results in the literature allow for bounded regressors
but require all regressors to be continuous. For example, Assumption
3.3 of \citet{shi2018estimating} gives a sufficient condition for
identification under bounded support. As commented in the paper, their
assumption essentially requires all regressors to be continuous. Continuous
distribution on all the regressors are also required by Assumption
6 of \citet{toth2017} and Assumption 4' of \citet{gao2023logical}.
Examples on nonseparable models include \citet{hoderlein2012nonparametric},
\citet{chernozhukov2015nonparametric} and \citet{chernozhukov2019nonseparable};
their arguments rely on the derivatives with respect to $X$ and thus
$X$ needs to be continuous. Here, we allow for one or multiple discrete
regressors, which are important in empirical studies as many treatment
variables are binary. Second, Corollary 4.1 of \citet{horowitz2009semiparametric}
considers a related model in the cross-sectional setting. It is possible
to translate this to the panel data in terms of conditional densities,
but these are still not the most general conditions; for example,
some strictly increasing continuous distribution functions have zero
density almost everywhere, e.g., \citet{salem1943some} and \citet{takacs1978increasing}.
Most importantly, none of the aforementioned works show whether their
conditions are necessary. We contribute to the literature by finding
what must be assumed and developing empirical tools to assess it.
\begin{comment}
For this model, proposes a condition for the identification of $\beta$.
Under the normalization of $\beta=(1,\beta_{-1}')'$ and the corresponding
partition $X=(X_{1},X_{-1}')'$, the key requirement is that there
exist some small $\delta>0$ and some set $N_{\delta}$ with $P(X_{-1}\in N_{\delta})>0$
such that for almost every $x_{-1}\in N_{\delta}$, the density of
$X_{1}+X_{-1}'\beta_{-1}$ conditional on $X_{-1}=x_{-1}$ is everywhere
positive on $[-\delta,\delta]$. One might wonder if a similar condition
can be stated for the panel setting. Like \citet{manski1987semiparametric},
this condition involves the unknown parameter $\beta$: to determine
the identifiability of $\beta$ from the observed data, we need to
check the density of $X'\beta$ condition on $X_{k}$ for some $k$.
As discussed above, the sign saturation condition is stated only in
terms of observed data and is thus directly testable. The proposed
test also has no tricky tuning parameters or curse of dimensionality
that plague many nonparametric estimators, especially in multiple
dimensions. Another technical issue is that statements of densities
might not give the most general conditions; for example, some strictly
increasing distribution functions have zero density almost everywhere,
e.g., \citet{takacs1978increasing}.
\end{comment}
\begin{comment}
\begin{rem}
We note that this condition does not cover all the identified cases.
For example, consider $X=(1,Z)'$ and $\beta=(0,3)'$, i.e., $X'\beta=1+3Z$.
If we set $X_{1}=1$ and $X_{-1}=Z$, then $X'\beta$ conditional
on $X_{-1}$ has a point mass; if we set $X_{1}=Z$ and $X_{-1}=1$,
then

Due to the normalization, we can only condition on $Z$ because the
coefficient on $1$ is zero; however, $X'\beta=1+3Z$ conditional
on $Z$ has no density. Moreover,
\end{rem}
\end{comment}
{}
\begin{comment}
Notice that this condition rules out a discrete $X_{1}$; if $X_{1}$
is discrete, then the density of $X_{1}+X_{-1}'\beta_{-1}$ conditional
on $X_{-1}=x_{-1}$ does not exist. This can be viewed as a \textbf{conditional}
sign saturation, namely, conditional on $X_{-1}$, $X'\beta$ has
sign saturation. Our sign saturation is an \textbf{unconditional}
statement. The conditional sign saturation is stronger than necessary
even for continuous $X$; in Lemma \ref{lem: example horowitz} in
the appendix, we give an example with $X=(X_{1},X_{2})'$ in which
almost surely, one of $P(X'\beta>0\mid X_{1})$ and $P(X'\beta<0\mid X_{1})$
is exactly zero and one of $P(X'\beta>0\mid X_{2})$ and $P(X'\beta<0\mid X_{2})$
is exactly zero.
\end{comment}
{}
\begin{comment}
\begin{lem}
\label{lem: example horowitz}Let $X=(X_{1},X_{2})'$ with $X_{2}=X_{1}\xi$,
where ${\rm supp}(X_{1})=[-1,1]$, ${\rm supp}(\xi)=[1/2,1]$ and $X_{1}$ and
$\xi$ are independent. Assume that $\beta=(\beta_{1},\beta_{2})'$
with $\beta_{1},\beta_{2}>0$. Then \\
(1) for any $x_{1}\in(-1,1)$, one of $P(X'\beta>0\mid X_{1}=x_{1})$
and $P(X'\beta<0\mid X_{1}=x_{1})$ is exactly zero. \\
(2) for any $x_{2}\in(-1,1)$, one of $P(X'\beta>0\mid X_{2}=x_{2})$
and $P(X'\beta<0\mid X_{2}=x_{2})$ is exactly zero.
\end{lem}
\begin{proof}
We first show the result for $x_{1}$. Notice that $X'\beta=X_{1}(\beta_{1}+\beta_{2}\xi)$.
If $x_{1}>0$, then $P(X'\beta<0\mid X_{1}=x_{1})=P(\beta_{1}+\beta_{2}\xi<0\mid X_{1}=x_{1})=0$
since ${\rm supp}(\xi)=[1/2,1]$ and $\beta_{1},\beta_{2}>0$. If $x_{1}<0$,
then $P(X'\beta>0\mid X_{1}=x_{1})=P(\beta_{1}+\beta_{2}\xi<0\mid X_{1}=x_{1})=0$.
If $x_{1}=0$, then both $P(X'\beta>0\mid X_{1}=x_{1})$ and $P(X'\beta<0\mid X_{1}=x_{1})$
are zero.

To see the result for $x_{2}$, notice that $X_{1}=X_{2}/\xi$. If
$x_{2}>0$, then
\[
P(X'\beta<0\mid X_{2}=x_{2})=P(\beta_{1}X_{2}/\xi+\beta_{2}X_{2}<0\mid X_{2}=x_{2})=P(\beta_{1}/\xi+\beta_{2}<0\mid X_{2}=x_{2})=0.
\]

If $x_{2}<0$, then
\[
P(X'\beta>0\mid X_{2}=x_{2})=P(\beta_{1}X_{2}/\xi+\beta_{2}X_{2}>0\mid X_{2}=x_{2})=P(\beta_{1}/\xi+\beta_{2}<0\mid X_{2}=x_{2})=0.
\]

If $x_{2}=0$, then both $P(X'\beta>0\mid X_{1}=x_{1})$ and $P(X'\beta<0\mid X_{1}=x_{1})$
are zero.
\end{proof}
\end{comment}
\begin{comment}
Moreover, checking this condition on density in practice seems to
require (1) finding a correct normalization (which component of $\beta$
is positive or at least non-zero) and (2) multiple-dimensional nonparametric
estimation, which might involve tricky tuning parameters. In contrast,
the sign saturation condition proposed here is stated in terms of
observed variables without any unknown parameters and can be tested
in a straight-forward way.
\end{comment}

Our necessity results complement those in \citet{pakes2024moment}.
By Proposition 3 therein, $\arg\max_{\beta}E(Y_{1}-Y_{0})\cdot\mathbf{1}\{W'\beta\geq0\}$,
the set of maximizers of the maximum score criterion function, is
the sharp identification set when we only assume the stationarity
of the distribution $u_{t}\mid(X,\alpha)$. Our results show that
this sharp identified reduces to a singleton up to scale under the
sign saturation condition. A natural question is the necessity of
this condition, especially if we are willing to impose additional
assumptions, namely i.i.d errors that are independent of $(X,\alpha)$.
By classical results, we know that under logistic errors, $\arg\max_{\beta}E(Y_{1}-Y_{0})\cdot\mathbf{1}\{W'\beta\geq0\}$
is not the sharp identified set. Our results in Section \ref{subsec: necessary cond}
further complete the puzzle by characterizing point identification
as the periodicity of $\dot{G}(\cdot)$.
\begin{comment}
However, the proof of their Theorem 2 shows that every point in the
identified set can be rationalized by a stationary error $u_{t}\mid X$
without any fixed effects $P(\alpha=0)=1$. Therefore, the class of
models with arbitrary stationary errors is so large that the fixed
effects are not important any more. In this sense, our sign saturation
condition is necessary for point identification up to scale. \footnote{One can show that the maximum score criterion function is maximized
at a unique point (up to scale) under the sign saturation condition
if the support of $Z$ is convex with non-empty interior.}
\end{comment}
{}

\section{\label{sec: id result}Identification of coefficients}

In the setting of (\ref{eq: FE model}), \citet{manski1987semiparametric}
identifies $\beta$ with the distributional stationarity condition:
\begin{assumption}
\label{assu: stationary error}The distribution of $u_{t}\mid(X,\alpha)$
does not depend on $t\in\{0,1\}$. Let the cumulative distribution
function (c.d.f) of this common distribution be $F(\cdot|x,\alpha)$.
Assume that for any $(x,\alpha)$, $F(\cdot\mid x,\alpha)$ is a strictly
increasing function on $\mathbb{R}$.
\end{assumption}
Assumption \ref{assu: stationary error} rules out dynamic models.
In this paper, we focus on static models with two time periods. Without
any restriction on the magnitude of $\alpha$ and $u_{t}$ in (\ref{eq: FE model}),
it is impossible to identify the magnitude of $\beta$ but the following
result by \citet{manski1987semiparametric} gives the identification
of the ``direction'' of $\beta$; in other words, $\beta$ is identified
up to scaling.
\begin{lem}[\citet{manski1987semiparametric}]
\label{lem: manski lem 1}Let Assumption \ref{assu: stationary error}
hold. Then with probability one,
\begin{equation}
{\rm sgn}\left(E(Y_{1}-Y_{0}\mid X)\right)={\rm sgn}\left(W'\beta\right),\label{eq: manski condition}
\end{equation}
where $W=X_{1}-X_{0}$ and ${\rm sgn}(\cdot)$ is defined as ${\rm sgn}(t)=1$
for $t>0$, ${\rm sgn}(t)=-1$ for $t<0$ and ${\rm sgn}(0)=0$.
\end{lem}
The key idea of our result is based on the condition in (\ref{eq: manski condition}).
Clearly, the conditional mean function $E(Y_{1}-Y_{0}\mid X)$ is
identified. For simplicity, suppose that $P(E(Y_{1}-Y_{0}\mid X)=0)=0$.
Then (\ref{eq: manski condition}) implies that there is a hyperplane
of $W$ that perfectly classifies ${\rm sgn}(E(Y_{1}-Y_{0}\mid X))$.
This is an ideal support vector machine (SVM) as illustrated in Figure
\ref{fig:SVM}, where the red and blue dots represent $-1$ and $1$,
respectively for ${\rm sgn}(E(Y_{1}-Y_{0}\mid X))$.\footnote{In \citet{komarova2013binary}, the geometry based on SVM is also
to study binary response models with a median restriction; the model
there does not have a panel structure. We thank Christopher Walker
for this reference. } If we want to identify the classification boundary in the SVM, we
obviously need to have both classes (red and blue) in the data. Since
these two classes correspond to $E(Y_{1}-Y_{0}\mid X)>0$ and $E(Y_{1}-Y_{0}\mid X)<0$,
this simple requirement is formalized as the following sign saturation
condition.

\begin{figure}
\caption{\label{fig:SVM}Lemma \ref{lem: manski lem 1} as a support vector
machine}

\begin{centering}
\includegraphics[clip,scale=0.6]{SVM4}
\par\end{centering}
{\small{}This demonstrates the case with $W\in\mathbb{R}^{2}$. The red dots
correspond to ${\rm sgn}(W'\beta)=1$ and the blue ones correspond
to ${\rm sgn}(W'\beta)=-1$.}{\small\par}
\end{figure}

\begin{assumption}
\label{assu: sign saturation}Both $P(E(Y_{1}-Y_{0}\mid X)>0)$ and
$P(E(Y_{1}-Y_{0}\mid X)<0)$ are strictly positive.
\end{assumption}
From Figure \ref{fig:SVM}, it is also clear that in addition to having
both classes, we also need the data points to be ``dense'' around
the hyperplane. For simplicity, we will require that the support of
$W$ is convex and has no-empty interior so the support of $W$ is
dense. This is all we need for uniquely determining the hyperplane
and there is nothing about unbounded supports. We will formalize this
intuition in in the next subsection and extend the results to the
case of multiple discrete regressors in Section \ref{sec: multiple discrete}.
\begin{rem}
Phrasing the problem in terms of support vector machines makes it
easy to see that when all the regressors are discrete, we cannot in
general expect to point identify $\beta$ up to scale. To see this,
notice that by \citet{pakes2024moment}, the sharp identified set
for $\beta$ is $\arg\max_{b}E(Y_{1}-Y_{0}){\rm sgn}(W'b)$, which
can be written as
\[
\arg\max_{b}E\left[\left|E(Y_{1}-Y_{0}\mid X)\right|\cdot{\rm sgn}(W'\beta)\cdot{\rm sgn}(W'b)\right].
\]
Therefore, if ${\rm sgn}(W'\tilde{\beta})={\rm sgn}(W'\beta)$ with
probability one, then $\tilde{\beta}$ is in the identified set. When
the regressors are discrete, we can in general move $\beta$ a bit
and retain the same ${\rm sgn}(W'\beta)$. As illustrated in Figure
\ref{fig:SVM 2}, the black and gray lines both perfectly classify
the red and blue dots and thus represent two points in the identified
set that are not scalar multiple of each other.
\end{rem}
\begin{figure}
\caption{\label{fig:SVM 2}Discrete regressors as a support vector machine }

\begin{centering}
\includegraphics[clip,scale=0.6]{discrete_covariates1}
\par\end{centering}
{\small{}This demonstrates the case in which $W\in\mathbb{R}^{2}$ is discrete.
The red dots correspond to ${\rm sgn}(W'\beta)=1$ and the blue ones
correspond to ${\rm sgn}(W'\beta)=-1$.}{\small\par}
\end{figure}


\subsection{The sign saturation condition is sufficient for identification}

Following \citet{chamberlain2010binary}, we assume that one of the
regressors is binary in that it is equal to zero at $t=0$ and is
equal to one at $t=1$; in other words, one component of $W=X_{1}-X_{0}$
is one. Thus, without loss of generality, we can partition $W=(Z',1)'\in\mathbb{R}^{K}$.
If all the regressors are continuous, then the result can still be
applied because the coefficient for the binary regressor is allowed
to be zero, see Corollary \ref{cor: only cont X}. The first main
result of this paper is to show that the sign saturation condition
(together with additional weak assumptions) is enough to guarantee
identification up to scaling. To state the formal result, we first
recall the definition of the support of a random variable (or its
corresponding probability measure): the support of probability measure
$\lambda$ is the smallest closed set $A$ such that $\lambda(A)=1$,
e.g., page 227 of \citet{dudley_2002}.
\begin{thm}
\label{thm: mixed reg}Let Assumption \ref{assu: stationary error}
hold. Partition $W=(Z',1)'\in\mathbb{R}^{K}$. Suppose that the support of
$Z$ is convex and has non-empty interior. Let Assumption \ref{assu: sign saturation}
hold.

Then $b=\mu\beta$ for some $\mu>0$ if and only if $R(b)=0$, where
\begin{equation}
R(b)=P\left({\rm sgn}\left(E(Y_{1}-Y_{0}\mid X)\right)\neq{\rm sgn}(W'b)\right).\label{eq: R function}
\end{equation}

Therefore, $\beta$ is identified up to scale.
\end{thm}
The requirement on the distribution of $Z$ is mild. Let $\mathcal{Z}$
be the support of $Z$. First, we do not require $\mathcal{Z}$ to be an
unbounded set. Second, the distribution of $Z$ does not have to admit
a density and ``atoms'' (or point masses) are allowed. Theorem \ref{thm: mixed reg}
imposes this via the convexity and non-empty interior of $\mathcal{Z}$.
When $\mathcal{Z}$ is not convex, the result is still useful: as long as
$\mathcal{Z}$ contains a convex subset with non-empty interior, we can
apply the result on this subset by restricting the sample, i.e., the
sign saturation condition holds on the restricted sample. If there
are multiple discrete variables, then we set $Z$ to be the difference
in continuous variables, see Section \ref{sec: multiple discrete}.

By Theorem \ref{thm: mixed reg}, to check whether $b$ is a rescaled
version of the true $\beta$, we only need to check $R(b)$ from the
distribution of the observed data. Notice that the identification
does not assume that $u_{t}$'s are independent across $t$; this
is because Assumption \ref{assu: stationary error} only requires
them to have the same marginal distribution given $(X,\alpha)$. To
give some intuition on the identifying power of sign saturation under
bounded $\mathcal{Z}$, we use the following simple example to illustrate
why the criterion function $R(\cdot)$ fails to deliver identification
when Assumption \ref{assu: sign saturation} is not satisfied.
\begin{example}
\label{exa: time dummy}Suppose that $W=(Z,1)$, where $Z$ is a continuous
random variable with support $[-1,1]$. For simplicity, we consider
$\beta=(\beta_{1},\beta_{2})'$ and $b=(b_{1},b_{2})'$ with $\beta_{1},b_{1}>0$.
The question is whether we can derive $b=\mu\beta$ for some $\mu>0$
from $R(b)=0$. Here, identifying $\beta$ up to scale is equivalent
to identifying $\beta_{2}/\beta_{1}$. After straight-forward calculations,
we can see that
\[
R(b)=P\left(Z\in\left(-\frac{\beta_{2}}{\beta_{1}},-\frac{b_{2}}{b_{1}}\right)\bigcup\left(-\frac{b_{2}}{b_{1}},-\frac{\beta_{2}}{\beta_{1}}\right)\right).
\]

Suppose that the sign saturation condition fails, say $P(W'\beta>0)=0$,
which implies that $-\beta_{2}/\beta_{1}\geq1$. In this case, the
set $\left(-\frac{\beta_{2}}{\beta_{1}},-\frac{b_{2}}{b_{1}}\right)\bigcup\left(-\frac{b_{2}}{b_{1}},-\frac{\beta_{2}}{\beta_{1}}\right)$
has empty intersection with $[-1,1]$ whenever $-b_{2}/b_{1}>1$.
In other words, if $P(W'\beta>0)=0$, then $R(b)=0$ whenever $-b_{2}/b_{1}>1$;
as a result, we cannot always verify $b_{2}/b_{1}=\beta_{2}/\beta_{1}$
through $R(b)$.

On the other hand, if the sign saturation condition holds, then we
have $-1<\beta_{2}/\beta_{1}<1$. In this case, the intersection of
$\left(-\frac{\beta_{2}}{\beta_{1}},-\frac{b_{2}}{b_{1}}\right)\bigcup\left(-\frac{b_{2}}{b_{1}},-\frac{\beta_{2}}{\beta_{1}}\right)$
and $[-1,1]$ always has non-empty interior whenever $b_{2}/b_{1}\neq\beta_{2}/\beta_{1}$.
Consequently, $R(b)=0$ if and only $b_{2}/b_{1}=\beta_{2}/\beta_{1}$.
\qed
\end{example}
The following example demonstrates how Theorem 1 gives identification
conditions for $\beta$ with interaction terms.
\begin{example}[Interaction terms]
Let $D_{t}=\mathbf{1}\{t=1\}$ and consider a continuous variable $H_{t}$.
Suppose that we are interested in the treatment effect of $D_{t}$
controlling for $H_{t}$. For a more flexible specification, we include
the interaction $D_{t}H_{t}$. Then $X_{t}=(D_{t},H_{t},D_{t}H_{t})'\in\mathbb{R}^{3}$
with the corresponding partition $\beta=(\beta_{1},\beta_{2},\beta_{3})'\in\mathbb{R}^{3}$.
For simplicity, suppose that the support for $(H_{0},H_{1})'$ is
$[-1,1]\times[-1,1]$. Then $Z=(H_{1},D_{1}H_{1})'-(H_{0},D_{0}H_{0})'=(H_{1}-H_{0},H_{1})'$.
One can easily verify that $\mathcal{Z}$ is convex and has non-empty interior.\footnote{To see the convexity, simply notice that $Z=\begin{pmatrix}-1 & 1\\
0 & 1
\end{pmatrix}\begin{pmatrix}H_{0}\\
H_{1}
\end{pmatrix}$ and $(H_{0},H_{1})'$ has a convex support.} By Lemma \ref{lem: manski lem 1}, straight-forward calculations
show that the sign saturation condition becomes $|\beta_{2}+\beta_{3}|+|\beta_{3}|>|\beta_{1}|$.
For example, this condition holds when $\beta_{1}=\beta_{3}=0$ and
$\beta_{2}\neq0$. Therefore, to test the null hypothesis of ``zero
effects whatsoever'', it suffices to have a relevant control variable.
\qed
\end{example}
We now discuss an important special case with only continuous covariates.
\begin{cor}
\label{cor: only cont X}Let Assumption \ref{assu: stationary error}
hold. If the interior of the support of $X_{1}-X_{0}$ contains $\boldsymbol{0}_{\dim(X_{t})}$
(zero in $\mathbb{R}^{\dim(X_{t})})$ and $\beta\neq\boldsymbol{0}_{\dim(X_{t})}$,
then $\beta$ is identified up to scale.
\end{cor}
A comparison between Example \ref{exa: time dummy} and Corollary
\ref{cor: only cont X} highlights the role of discrete variables
in identification. If all the covariates are continuous and the change
in covariates contains zero in the support, then we can generally
expect identification (up to scale) of $\beta$, regardless of whether
the covariates are bounded. In contrast, if the covariates include
discrete variables (such as the time fixed effect in \citet{chamberlain2010binary}),
we no longer have such a generic guarantee of identification.

\subsection{\textcolor{blue}{\label{sec: multiple discrete}}Extension to multiple
discrete regressors}

In this subsection, we consider the case of multiple discrete regressors.\footnote{I thank Xavier D'Haultf{\oe}uille for a discussion of this case.}
We partition $X_{t}=(X_{(1),t}',X_{(2),t}')'$, where $X_{(1),t}\in\mathbb{R}^{K_{1}}$
are discrete regressors and $X_{(2),t}\in\mathbb{R}^{K_{2}}$ are continuous
regressors. We can partition the corresponding $W=X_{1}-X_{0}=(D',Z')'$
with $D=X_{(1),1}-X_{(1),0}$ and $Z=X_{(2),1}-X_{(2),0}$. Let $\mathcal{D}$
denote the support of $D$. We show that identification still holds
under a conditional version of sign saturation.
\begin{thm}
\label{thm: discrete cov ID}Let Assumption \ref{assu: stationary error}
hold. Suppose that there exist linearly independent $d_{1},...,d_{K_{1}}\in\mathcal{D}$
with such that for $j\in\{1,...,K_{1}\}$,\\
(1) both $P(E(Y_{1}-Y_{0}\mid X,D=d_{j})>0)$ and $P(E(Y_{1}-Y_{0}\mid X,D=d_{j})<0)$
are strictly positive\\
(2) the support of $Z$ conditional on $D=d_{j}$ is convex and has
non-empty interior.

Then $\beta$ is identified up to scale.
\end{thm}
We note that $|\mathcal{D}|$ can be much larger than $K_{1}$. For example,
if $X_{(1),t}$ represents $K_{1}$ binary variables, then the support
of $X_{(1),t}$ is $\{0,1\}^{K_{1}}$ and $\mathcal{D}=\{-1,0,1\}^{K_{1}}$,
which means $|\mathcal{D}|=3^{K_{1}}$. Theorem \ref{thm: discrete cov ID}
is convenient in that it does not require the sign saturation to hold
conditional on $D=d$ for \textbf{every} $d\in\mathcal{D}$. It suffices
to require this for $K_{1}$ linearly independent elements of $\mathcal{D}$.
If $K_{1}=1$ ($X_{(1),t}$ is a time dummy), then this requirement
reduces to Assumption \ref{assu: sign saturation}. Theorem \ref{thm: discrete cov ID}
also allows for more complicated cases as demonstrated in the following
example.
\begin{example}[Time dummy and categorical variables]
Suppose that $R_{t}$ is a categorical variable taking values in
$\{1,...,K_{1}\}$. Consider $X_{(1),t}=(\mathbf{1}\{t=1\},\mathbf{1}\{R_{t}=1\},...,\mathbf{1}\{R_{t}=K_{1}-1\})'\in\mathbb{R}^{K_{1}}$,
i.e., a time dummy and $K_{1}-1$ dummy variables for $R_{t}$. Then
we can write the support of $X_{(1),t}\in\mathbb{R}^{K_{1}}$ as $\{\mathbf{1}\{t=1\}\}\times\{e_{1},...,e_{K_{1}-1},\boldsymbol{0}_{K_{1}-1}\}$,
where $e_{j}$ is the $j$-th column of the $(K_{1}-1)\times(K_{1}-1)$
identity matrix and $\boldsymbol{0}_{K_{1}-1}=(0,...,0)'\in\mathbb{R}^{K_{1}-1}$.
Let us assume that the support of $(R_{0},R_{1})'$ contains $(K_{1},j)$
for any $j\in\{1,...,K_{1}\}$. Then $\mathcal{D}$ contains $\{1\}\times\{e_{1},...,e_{K_{1}-1},\boldsymbol{0}_{K_{1}-1}\}$.
In other words, $\mathcal{D}$ contains all the columns of the following
matrix:
\[
\begin{pmatrix}1 & 1 & 1 & \cdots & 1\\
0 & 1 & 0 & \cdots & 0\\
0 & 0 & 1 & \cdots & 0\\
\vdots & \vdots & \vdots & \ddots & \vdots\\
0 & 0 & 0 & \cdots & 1
\end{pmatrix}\in\mathbb{R}^{K_{1}\times K_{1}}.
\]

We can easily see that the determinant of this upper triangular matrix
is equal to one. Therefore, the above matrix has full rank, which
means that $\mathcal{D}$ contains $K_{1}$ linearly independent elements.
Then $\beta$ is identified up to scale if for every $j\in\{1,...,K_{1}\}$,
$P(E(Y_{1}-Y_{0}\mid Z,R_{0}=K_{1},R_{1}=j)>0)$ and $P(E(Y_{1}-Y_{0}\mid Z,R_{0}=K_{1},R_{1}=j)<0)$
are strictly positive and the support of $Z$ conditional on $(R_{0},R_{1})=(K_{1},j)$
is convex and has non-empty interior.\qed
\end{example}

\subsection{\label{subsec: necessary cond}Is the sign saturation condition is
also necessary?}

It turns out that in the many situations, sign saturation characterizes
identifiability. To illustrate this point, we consider the setting
studied by \citet{chamberlain2010binary} and assume that $u_{t}$
is independent of $(X,\alpha)$ and is from a known distribution.\footnote{To establish necessity, it is enough to consider the case in which
$u_{t}$ is independent of $(X,\alpha)$. This is similar to showing
minimax lower bounds. If identification cannot be guaranteed even
with the additional assumption of independence between $u$ and $(X,\alpha)$,
then it is definitely not guaranteed without this assumption.} We show that without sign saturation, identification fails at every
point unless the distribution of $u_{t}$ is from a special class.

Suppose that in time period $t\in\{0,1\}$, we observe $(Y_{t},X_{t})$,
where $X_{t}=(X_{t,1}',X_{t,2})'$ with $X_{t,1}\in\mathbb{R}^{K-1}$ and
$X_{t,2}=\mathbf{1}\{t=1\}$. Suppose that $u_{t}$ is independent of $(X,\alpha)$
and has c.d.f $F(\cdot)$. Then from the data, the distribution of
$Y$ given $(X,\alpha)$ is determined by the vector
\begin{align}
L(X;\beta,\alpha) & =\begin{pmatrix}F(X_{0,1}'\beta_{1}+\alpha)\\
F(X_{1,1}'\beta_{1}+\beta_{2}+\alpha)\\
F(X_{0,1}'\beta_{1}+\alpha)\cdot F(X_{1,1}'\beta_{1}+\beta_{2}+\alpha)
\end{pmatrix}\nonumber \\
 & =\begin{pmatrix}F(X_{0,1}'\beta_{1}+\alpha)\\
F(X_{0,1}'\beta_{1}+\alpha+W'\beta)\\
F(X_{0,1}'\beta_{1}+\alpha)\cdot F(X_{0,1}'\beta_{1}+\alpha+W'\beta)
\end{pmatrix},\label{eq: L fun}
\end{align}
where $\beta=(\beta_{1}',\beta_{2})$ is partitioned as $\beta_{1}\in\mathbb{R}^{K-1}$
and $\beta_{2}\in\mathbb{R}$, and $W:=X_{1}-X_{0}=(Z',1)'$ with $Z=X_{1,1}-X_{0,1}$.
As pointed out in \citet{chamberlain2010binary}, here we should aim
for identification of $\beta$, not just up to scaling, because ``our
scale normalization is built in to the given specification for the
$u_{t}$ distribution''. Let $\mathcal{Z}$ be the support of $Z$, i.e.,
again the smallest closed set with the full probability mass. Since
$W=(Z',1)'$, the support of $W$ is $\mathcal{W}=\mathcal{Z}\times\{1\}$.

The special class of functions is best described in terms of a transformation
of $F(\cdot)$. Define the function $G(\cdot)=\ln\frac{F(\cdot)}{1-F(\cdot)}$.
Let $\dot{G}(\cdot)$ denote the derivative of $G(\cdot)$. If $F(\cdot)$
is the logistic function, then $G(\cdot)$ is an affine function and
$\dot{G}(\cdot)$ is a constant function. It turns out that if we
are willing to rule out cases with $\dot{G}(\cdot)$ being a periodic
function,\footnote{A function $h(\cdot)$ on $\mathbb{R}$ is a periodic function if there exists
$c\neq0$ such that $h(c+t)=h(t)$ for any $t\in\mathbb{R}$. In this case,
$c$ is said to be a period of $h(\cdot)$. Notice that we assume
that a period has to be non-zero; otherwise, every function would
be a periodic function with a period being zero.} then sign saturation is a necessary condition for identification.
This special class includes the logistic function, whose $\dot{G}(\cdot)$
is a constant function and is clearly periodic. (Any real number is
a period of a constant function.)

We now introduce the notations for describing identification. Let
$\Pi$ denote the set of probability measures on $\mathbb{R}$. The distribution
of $Y\mid X=x$ is determined by $\int L(x;\beta,\alpha)d\pi_{x}(\alpha)$
for some $\pi_{x}\in\Pi$. Here, $\pi_{x}$ denotes the distribution
$\alpha\mid X=x$. We now recall the notation of identification failure
discussed in \citet{chamberlain2010binary}.
\begin{defn}[Identification failure at $\beta$]
\label{def: chamberlain def}We say that the identification fails
at $\beta$ if there exists $b\neq\beta$ such that for any $x$ in
the support, there exist $\pi_{x},\tilde{\pi}_{x}\in\Pi$ depending
on $x$ such that $\int L(x;\beta,\alpha)d\pi_{x}(\alpha)=\int L(x;b,\alpha)d\tilde{\pi}_{x}(\alpha)$.
\end{defn}
This definition says that there exist two ``mixing distributions''
$\pi_{x}$ and $\tilde{\pi}_{x}$ in $\mathcal{G}$ such that the two mixtures
(representing the distribution $Y\mid X=x$) $\int L(x;\beta,\alpha)d\pi_{x}(\alpha)$
and $\int L(x;b,\alpha)d\tilde{\pi}_{x}(\alpha)$ are identical. Following
\citet{chamberlain2010binary}, we can state identification in terms
of convex hulls of $L(x;\beta,\alpha)$. The identification fails
at $\beta$ if there exists $b\neq\beta$ such that ${\rm conv}\{L(x;\beta,\alpha):\alpha\in\mathbb{R}\}\bigcap{\rm conv}\{L(x;b,\alpha):\alpha\in\mathbb{R}\}\neq\emptyset$
for any $x$, where ${\rm conv}$ denotes the convex hull.

Notice that $X_{0,1}'\beta_{1}$ can be absorbed by $\alpha$ since
the distribution of $\alpha$ is allowed to have arbitrary dependence
on $x$. Without loss of generality, we view the entire $X_{0,1}'\beta_{1}+\alpha$
as the fixed effects to simplify notations. Then we can restate Definition
\ref{def: chamberlain def} as follows.
\begin{defn}
\label{def: chamb alter}We say that the identification fails at $\beta$
if there exists $b\neq\beta$ such that $\mathcal{A}(w'\beta)\bigcap\mathcal{A}(w'b)\neq\emptyset$
for any $w\in\mathcal{W}$, where for any $t\in\mathbb{R}$, $\mathcal{A}(t)={\rm conv}\{p(t,\alpha):\alpha\in\mathbb{R}\}$
and
\[
p(t,\alpha):=\begin{pmatrix}F(\alpha)\\
F(t+\alpha)\\
F(\alpha)\cdot F(t+\alpha)
\end{pmatrix}.
\]
\end{defn}
Perhaps the simplest way to see the equivalence between the two definitions
is to notice that ${\rm conv}\{L(x;\beta,\alpha):\alpha\in\mathbb{R}\}=\mathcal{A}(w'\beta)$,
where $w=x_{1}-x_{0}$. We now define the parameter space in which
the sign saturation fails: $\mathcal{B}_{+}=\{\beta\in\mathbb{R}^{K}:\ w'\beta>0\ \forall w\in\mathcal{W}\}$.
The main result for the necessity is the following.
\begin{thm}
\label{thm: neccessity part 1}Suppose that $\mathcal{Z}$ is bounded. Suppose
that $F(\cdot)$ is continuously differentiable with support on $\mathbb{R}$.
If $\dot{G}(\cdot)$ is not a periodic function, then the identification
fails at every point in $\mathcal{B}_{+}$.
\end{thm}
We compare this with Theorem 1 of \citet{chamberlain2010binary},
which states that if $F(\cdot)$ is outside a special class, identification
fails in an open neighborhood. Theorem \ref{thm: neccessity part 1}
complements this result by showing that if $F(\cdot)$ is outside
a special class, identification fails at every point, not just in
an open neighborhood. This is also why the proof of Theorem \ref{thm: neccessity part 1}
follows a very different strategy from the proof of Theorem 1 of \citet{chamberlain2010binary};
the latter essentially only needs to find one point at which the identification
fails whereas we need to show that identification fails universally
in $\mathcal{B}_{+}$.

The special class in \citet{chamberlain2010binary} only includes
logistic functions, but here our special class is larger and includes
all functions with periodic $\dot{G}(\cdot)$. This enlargement of
the special class cannot be avoided because we show that there are
instances of non-logistic $F(\cdot)$ for which identification holds,
e.g., $F(t)=[1+\exp(-2t-\sin(t))]^{-1}$. We can push our arguments
further and show that the identification under periodic $\dot{G}(\cdot)$
is ``robust'' and is based on a set with strictly positive probability
mass.
\begin{thm}
\label{thm: ID period fun}Suppose that $\dot{G}(\cdot)$ is a continuous
periodic function that is non-constant. Let $\beta=(\beta_{1}',\beta_{2})'\in\mathcal{B}_{+}$.
Then\\
(1) there exists a minimal positive period $\eta>0$ for $\dot{G}(\cdot)$,
i.e., the set $\{a>0:\ \dot{G}(a+t)=\dot{G}(t)\ \forall t\in\mathbb{R}\}$
has a smallest element. \\
(2) If $(z'\beta_{1}+\beta_{2})/\eta$ is an integer for some $z$
in the interior of $\mathcal{Z}$, then $\beta$ is identified; moreover,
for any $b\in\mathcal{B}_{+}$ with $b\neq\beta$, $P\left(\mathcal{A}(W'\beta)\bigcap\mathcal{A}(W'b)=\emptyset\right)>0$,
where $W=(Z',1)'$. \\
(3) If $\mathcal{Z}$ is compact and has non-empty interior, then there
exists an open ball $D\subset\mathcal{B}_{+}$ such that every point in
$D$ is identified.
\end{thm}
By Theorem \ref{thm: ID period fun}, as long as $\dot{G}(\cdot)$
is periodic and $z'\beta_{1}+\beta_{2}$ is in $\eta\cdot\mathbb{N}$ for
some interior point $z$ ($\mathbb{N}$ denotes the set of positive integers),
we still have identification. This complements Theorem 1 of \citet{chamberlain2010binary}
in an interesting way. For non-logistic $F(\cdot)$, it is true that
the identification fails at every point in an open set. On the other
hand, we show that the identification also holds at every point in
another open set for periodic $\dot{G}(\cdot)$. It is worth noting
that in Theorem \ref{thm: ID period fun}, the identification of $\beta$
is robust in the sense that it is not based on a small number of ``unimportant
points'' in $\mathcal{W}$; the points in $\mathcal{W}$ that allow us to identify
$\beta$ have strictly positive probability mass. Therefore, the truly
hopeless cases for identification are those with non-periodic $\dot{G}(\cdot)$.
Finally, for the case of compact and convex $\mathcal{Z}$ with non-empty
interior, we summarize our results and the existing literature in
Table \ref{tab: ID and sign sat}.

\begin{table}[h]
\caption{\label{tab: ID and sign sat}Identification and sign saturation}

\bigskip{}

For bounded and convex $\mathcal{Z}$ with non-empty interior:

{\centering

\begin{tabular}{c|cc}
\multicolumn{3}{c}{}\tabularnewline
 & Sign saturation holds & Sign saturation does not hold\tabularnewline
\hline
 &  & \tabularnewline
$\dot{G}(\cdot)$ non-periodic & \begin{tabular}{@{}c@{}}ID holds at every point\\ (Theorem \ref{thm: mixed reg})\end{tabular} & \begin{tabular}{@{}c@{}}ID fails at every point\\ (Theorem \ref{thm: neccessity part 1})\end{tabular}\tabularnewline
 &  & \tabularnewline
 &  & \tabularnewline
\begin{tabular}{@{}c@{}}$\dot{G}(\cdot)$ periodic \\ but non-constant\end{tabular} & \begin{tabular}{@{}c@{}}ID holds at every point\\ (Theorem \ref{thm: mixed reg})\end{tabular} & \begin{tabular}{@{}c@{}}ID holds in an open set \\(Theorem \ref{thm: ID period fun})\\ but fails in another open set \\(\cite{chamberlain2010binary}) \end{tabular}\tabularnewline
 &  & \tabularnewline
 &  & \tabularnewline
\begin{tabular}{@{}c@{}}$\dot{G}(\cdot)$ constant \\(logistic $F$)\end{tabular} & \begin{tabular}{@{}c@{}}ID holds at every point\\ (\cite{chamberlain1980analysis}\\ \cite{DaveziesID2021})\end{tabular} & \begin{tabular}{@{}c@{}}ID holds at every point\\ (\cite{chamberlain1980analysis}\\ \cite{DaveziesID2021})\end{tabular}\tabularnewline
 &  & \tabularnewline
\hline
\multicolumn{3}{c}{}\tabularnewline
\end{tabular}

}
\end{table}

Exploiting the periodicity of $\dot{G}(\cdot)$ can be seen as a generalization
of the conditional maximum likelihood estimator for logistic distributions.
To see this, we notice that
\[
\frac{P(Y_{1}=1,Y_{0}=0\mid X,\alpha)}{P(Y_{1}=0,Y_{1}=1\mid X,\alpha)}=\exp\left[G(W'\beta+\alpha)-G(\alpha)\right].
\]
Under logistic errors, $\dot{G}(\cdot)$ is a constant and thus $G(W'\beta+\alpha)-G(\alpha)$
does not depend on $\alpha$ for any $W$. If $\dot{G}(\cdot)$ is
a periodic function with the smallest positive period $\eta>0$, then
$G(W'\beta+\alpha)-G(\alpha)$ also does not depend on $\alpha$ whenever
$W'\beta/\eta$ is an integer. Since $W=(Z',1)'$, we only need to
have an interior point $z$ such that $(z',1)'\beta/\eta$ is an integer.
\begin{rem}
As we have seen, the sign saturation condition guarantees the identification
of $\beta$ up to scale. Does it also guarantee the parametric rate
for estimation? The answer is no because Theorem 2 of \citet{chamberlain2010binary}
shows that the information bound is always zero outside the logistic
case, regardless of the identification status. This highlights a key
difference between identification and the parametric rate in estimation.
The question of identification depends on the sign saturation condition,
whereas the root-$n$ estimability depends on the semiparametric efficiency
calculation or the functional projection argument developed by \citet{bonhomme2012functional}.
\begin{comment}
The above analysis shows that distributions in the special class of
periodic $\dot{G}(\cdot)$ can have sufficient identification power
without the sign saturation condition. One natural question is whether
this special class of periodic $\dot{G}(\cdot)$ also has special
properties for estimation. One thing we can say is that the parametric
rate is still not possible outside the logistic case. This is because
Theorem 2 of \citet{chamberlain2010binary} still applies. Of course,
periodicity of $\dot{G}(\cdot)$ might provide other benefits in terms
of estimation even though the root-$n$ rate is impossible. We will
leave the exploration of this direction to future research.
\end{comment}
\end{rem}

\section{Implications on marginal effects}

The identification of $\beta$ is directly related to the identification
of marginal effects in panel models with binary responses. The bounds
on the marginal effects often assume point identification of $\beta$,
see e.g., \citet*{DaveziesID2021}, \citet{liu2105.12891} and Theorem
6 of \citet*{chernozhukov2013average}.

Here, we explore an important link between $\beta$ and the treatment
effects through the sign of components of $\beta$. Let us consider
the setting of Section \ref{subsec: necessary cond}: (1) the errors
$u_{t}$'s are i.i.d across $t$ with distribution $F(\cdot)$ and
are independent of $(X,\alpha)$ and (2) $X_{2,t}$ is binary with
$X_{2,t}=\mathbf{1}\{t=1\}$. We can interpret $X_{2,t}$ as a binary treatment;
in period $t=0$, no one is treated and in period $t=1$, everyone
is treated. For $\beta=(\beta_{1}',\beta_{2})'$, the sign of $\beta_{2}$
corresponds to the sign of common measures of treatment effects. For
example, the average partial effect effect at $X_{1,t}=z$ is
\[
\Delta_{APE}(z)=P(z'\beta_{1}+\beta_{2}+\alpha\geq u_{t})-P(z'\beta_{1}+\alpha\geq u_{t}),
\]
see e.g., \citet{chernozhukov2013average}.

Although the magnitude of average partial effect is often not point
identified (see e.g., \citet{honore2006bounds} and \citet{chernozhukov2013average}),
there is hope that the sign of the effect is point identified, which
means that the identified set contains only positive numbers, only
negative numbers or only zero. Under Assumption \ref{assu: stationary error},
the support of $u_{t}$ is $\mathbb{R}$, which means that ${\rm sgn}(\Delta_{APE}(z))={\rm sgn}(\beta_{2})$
for any $z$. Hence, identifying ${\rm sgn}(\Delta_{APE}(z))$ is
equivalent to identifying ${\rm sgn}(\beta_{2})$. In the literature,
there are also terms such as average treatment effects or ceteris
paribus effects, e.g., \citet{hoderlein2012nonparametric} and \citet{chernozhukov2015nonparametric}.
Under the assumptions of a linear index, the sign of these other measures
of treatment effects is also the sign of $\beta_{2}$.
\begin{rem}
There are alternative ways of defining treatment effects but due to
the single-index nature of the model, the sign of other notions of
treatment effects is still related to ${\rm sgn}(\beta_{2})$. For
example, consider the individual effect $\mathbf{1}\{X_{1,t}\beta_{1}+\beta_{2}+\alpha\geq u_{t}\}-\mathbf{1}\{X_{1,t}\beta_{1}+\alpha\geq u_{t}\}$.
Clearly, this effect is negative with zero probability if and only
if $\beta_{2}>0$.
\end{rem}
We now show that the sign saturation condition is necessary for this
purpose. To formally state the result, we rephrase Definition \ref{def: chamberlain def}.
\begin{defn}
We say that $\beta$ and $b$ are observationally equivalent if for
any $x$ in the support, there exist $\pi_{x},\tilde{\pi}_{x}\in\Pi$
depending on $x$ such that $\int L(x;\beta,\alpha)d\pi_{x}(\alpha)=\int L(x;b,\alpha)d\tilde{\pi}_{x}(\alpha)$,
where $L$ is defined in (\ref{eq: L fun}).
\end{defn}
Recall that $\mathcal{B}_{+}=\{\beta\in\mathbb{R}^{K}:\ w'\beta>0\ \forall w\in\mathcal{W}\}$
is the set of parameter values that do not satisfy the sign saturation
condition ($E(Y_{1}-Y_{0}\mid X)$ always positive). Consider two
subsets $\mathcal{B}_{+,0}=\{\beta=(\beta_{1}',\beta_{2})'\in\mathcal{B}_{+}:\ \beta_{2}=0\}$
and $\mathcal{B}_{+,+}=\{\beta=(\beta_{1}',\beta_{2})'\in\mathcal{B}_{+}:\ \beta_{2}>0\}$,
which denote the set of parameter values that imply zero effects and
positive effects of $X_{2,t}$, respectively. The following result
states the identification failure of ${\rm sgn}(\beta_{2})$.
\begin{thm}
\label{thm: neccessity sign}Suppose that $\mathcal{Z}$ is bounded. Suppose
that $F(\cdot)$ is continuously differentiable such that $\dot{G}(\cdot)$
is not a periodic function. Then every point in $\mathcal{B}_{+,0}$ is
observationally equivalent to a point in $\mathcal{B}_{+,+}$.
\end{thm}
By Theorem \ref{thm: mixed reg}, the sign saturation condition guarantees
point identification of $\beta$ up to scale and thus point identification
of ${\rm sgn}(\beta_{2})$. By Theorem \ref{thm: neccessity sign},
when the sign saturation condition fails, we cannot distinguish between
$\Delta_{APE}(z)=0$ and $\Delta_{APE}(z)>0$ unless $\dot{G}(\cdot)$
is periodic. Therefore, unless $\dot{G}(\cdot)$ is periodic, the
sign saturation condition is sufficient and necessary to guarantee
point identification of the sign of the marginal effects of $X_{2,t}$.
\begin{comment}
By a continuity argument, one can also show that . For any $\varepsilon>0$,
we define $\mathcal{B}_{+}^{(\varepsilon)}:=\{\beta\in\mathbb{R}^{K}:\ w'\beta>\varepsilon\ \forall w\in\mathcal{W}\}$.
Clearly, $\mathcal{B}_{+}=\mathcal{B}_{+}^{(0)}$.

For any $\varepsilon>0$, $\mathcal{B}_{+,+,\varepsilon}=\{\beta=(\beta_{1}',\beta_{2})'\in\mathcal{B}_{+}:\ 0<\beta_{2}<\varepsilon\}$
and $\mathcal{B}_{+,+,-\varepsilon}=\{\beta=(\beta_{1}',\beta_{2})'\in\mathcal{B}_{+}:\ -\varepsilon<\beta_{2}<0\}$.
We show that there exists $\varepsilon>0$ such that every point in
$\mathcal{B}_{+,+,-\varepsilon}$ is observationally equivalent to a point
in $\mathcal{B}_{+,+}$.
\begin{proof}
We now show that there exists $\delta>0$ such that for any $z\in\mathcal{Z}$,
$\mathcal{A}(w'\beta)\bigcap\mathcal{A}(w'b)\neq\emptyset$, where $w=(z',1)'$.
We proceed by contradiction. Suppose that for any $\delta>0$, there
exists $z_{\delta}\in\mathcal{Z}$ such that $\mathcal{A}(w_{\delta}'b)\bigcap\mathcal{A}(w_{\delta}'\beta)=\emptyset$,
where $w_{\delta}=(z_{\delta}',1)'$. By Lemma \ref{lem: farkas},
there exists $v_{\delta}\neq0$ such that $\sup_{\alpha\in\mathbb{R}}v_{\delta}'p(w_{\delta}'b,\alpha)\leq\inf_{\alpha\in\mathbb{R}}v_{\delta}'p(w_{\delta}'\beta,\alpha)$.
By Lemma \ref{lem: key ID} and $w_{\delta}'b-w_{\delta}'\beta=\delta>0$,
we have that $\sup_{\alpha\in\mathbb{R}}v_{\delta}'p(w_{\delta}'b,\alpha)\geq\inf_{\alpha\in\mathbb{R}}v_{\delta}'p(w_{\delta}'\beta,\alpha)$.
Hence,
\[
\sup_{\alpha\in\mathbb{R}}v_{\delta}'p(w_{\delta}'\beta+\delta,\alpha)=\sup_{\alpha\in\mathbb{R}}v_{\delta}'p(w_{\delta}'b,\alpha)=\inf_{\alpha\in\mathbb{R}}v_{\delta}'p(w_{\delta}'\beta,\alpha).
\]

By Lemma \ref{lem: ID part 1} (with $t=w_{\delta}'\beta$),
\[
\inf_{\alpha\in\mathbb{R}}[G(w_{\delta}'\beta+\delta+\alpha)-G(\alpha)]>\sup_{\xi\in\mathbb{R}}[G(w_{\delta}'\beta+\xi)-G(\xi)].
\]

Notice that $w_{\delta}'\beta\in T$. Therefore, we have shown that
for any $\delta>0$, there exists $t_{\delta}\in T$ such that
\[
\inf_{\alpha\in\mathbb{R}}[G(t_{\delta}+\delta+\alpha)-G(\alpha)]>\sup_{\xi\in\mathbb{R}}[G(t_{\delta}+\xi)-G(\xi)].
\]

By Lemma \ref{lem: ID part 2}, $\dot{G}$ is a periodic function.
However, this contradicts the assumption that $\dot{G}$ is not a
periodic function.
\end{proof}
Define $\mathcal{B}_{+,-}=\{\beta=(\beta_{1}',\beta_{2})'\in\mathcal{B}_{+}:\ \beta_{2}<0\}$.
We fix an arbitrary $\beta_{*}=(\beta_{*,1}',\beta_{,*2})\in\mathcal{B}_{+,-}$.
Clearly, $(\beta_{*,1}',0)'\in\mathcal{B}_{+,0}$ since $\beta_{*,2}<0$.
In the proof of Theorem \ref{thm: neccessity part 1}, we proved that
$(\beta_{*,1}',0)'\in\mathcal{B}_{+,0}$ is observationally equivalent to
$(\beta_{*,1}',\delta)'$ for some $\delta_{*}>0$. Notice that we
can choose $\delta_{*}$ to depend only on $\beta_{1,*}$. We now
consider $\mathcal{B}_{+,+,-\varepsilon}$ with $\varepsilon=\delta_{*}/2$.

We also know that $z'\beta_{*,1}>-\beta_{*,2}$ for any $z$.

We fix an arbitrary $\beta=(\beta_{1}',\beta_{2})'\in\mathcal{B}_{+,+,-\varepsilon}$.
We define

Now we consider any $\beta=(\beta_{1}',\beta_{2})\in\mathcal{B}_{+}$ with
$\beta_{2}<0$.

Now we pick any $\beta\in\mathcal{B}_{+,0}$. Then $b=(\beta_{1}',\beta_{2}+\delta)'=(\beta_{1}',\delta)\in\mathcal{B}_{+,+}$.
The desired result follows.

Define $A=\{z'\beta_{1}:z\in\mathcal{Z}\}$. Then $T=A+\beta_{2}$ is contained
by $\overline{T}=[]$.
\end{comment}
{}

\section{\label{sec: test}Assessing the sign saturation condition}

We have seen that the sign saturation condition is sufficient and
necessary for the identification. The conditional mean function $\phi(X)=E(Y_{1}-Y_{0}\mid X)$
is identified. Therefore, ideally the sign saturation condition can
be checked in data. Here, we provide a simple check that does not
involve nonparametric estimation of $\phi$.

By Lemma \ref{lem: manski lem 1}, $P(\phi(X)\leq0)=1$ if and only
if $P(W'\beta\leq0)=1$. Hence, we define $\rho(q)=E\mathbf{1}\{W'q\geq0\}(Y_{1}-Y_{0})$
for $q\in\mathbb{R}^{K}$. Here, we notice that we can replace $\mathbb{R}^{K}$
with $[-1,1]^{K}$ or any set that contains an open neighborhood of
zero. It turns out that the sign saturation condition can be written
in terms of $\rho(\cdot)$ once we observe
\[
\rho(\beta)=E\left(\mathbf{1}\{\phi(X)\geq0\}\cdot\phi(X)\right)\quad{\rm and}\quad\rho(-\beta)=E\left(\mathbf{1}\{\phi(X)\leq0\}\cdot\phi(X)\right).
\]

We give the formal statement below.
\begin{lem}
\label{lem: test}Let Assumption \ref{assu: stationary error} hold.
Then
\[
\sup_{q\in\mathbb{R}^{K}}\rho(q)=E\max\left\{ \phi(X),0\right\} \ {\rm and}\ \inf_{q\in\mathbb{R}^{K}}\rho(q)=E\min\left\{ \phi(X),0\right\} .
\]

\end{lem}
Define $\tau_{*}=\min\{\tau_{1},-\tau_{2}\}$ with $\tau_{1}=\sup_{q\in\mathbb{R}^{K}}\rho(q)$
and $\tau_{2}=\inf_{q\in\mathbb{R}^{K}}\rho(q)$. By Lemma \ref{lem: test},
measuring $\tau_{*}$ can give us some indication of whether (or how
``well'') the sign saturation condition holds. In particular, the
sign saturation fails if and only if $\tau_{*}=0$. From the data,
we can construct a one-sided confidence interval for $\tau_{*}$.
The natural estimate for $\tau_{*}$ is $\hat{\tau}_{*}=\min\{\hat{\tau}_{1},-\hat{\tau}_{2}\}$,
where $\hat{\tau}_{1}=\sup_{q\in\mathbb{R}^{K}}\hat{\rho}_{n}(q)$, $\hat{\tau}_{2}=\inf_{q\in\mathbb{R}^{K}}\hat{\rho}_{n}(q)$
and
\[
\hat{\rho}_{n}(q)=n^{-1}\sum_{i=1}^{n}\mathbf{1}\{W_{i}'q\geq0\}(Y_{i,1}-Y_{i,0}).
\]

In the proof of Theorem \ref{thm: test} below, we show that
\[
\sqrt{n}\left(\hat{\tau}_{*}-\tau_{*}\right)\leq\left(\sup_{q\in\mathbb{R}^{K}}S_{n}(q)\right)\cdot\mathbf{1}\{\tau_{1}\leq-\tau_{2}\}+\left(\sup_{q\in\mathbb{R}^{K}}(-S_{n}(q))\right)\cdot\mathbf{1}\{\tau_{1}>-\tau_{2}\},
\]
where $S_{n}(q)=\sqrt{n}(\hat{\rho}_{n}(q)-\rho(q))$. By the standard
arguments of empirical processes, $S_{n}$ converges to a mean-zero
Gaussian process, which means that $\sup_{q\in\mathbb{R}^{K}}S_{n}(q)$ and
$\sup_{q\in\mathbb{R}^{K}}(-S_{n}(q))$ have the same distribution. Thus,
a simple $(1-\alpha)$-confidence interval for $\tau_{*}$ can be
obtained once we approximate the $(1-\alpha)$ quantile of $\sup_{q\in\mathbb{R}^{K}}S_{n}(q)$.
This motivates the following bootstrapping algorithm.
\begin{lyxalgorithm}
\label{alg: test }For a $(1-\alpha)$-confidence interval for $\tau_{*}$:
\begin{enumerate}
\item Collect data $\{(W_{i},Y_{i,1},Y_{i,0})\}_{i=1}^{n}$.
\item Compute the estimate $\hat{\tau}_{*}=\min\{\hat{\tau}_{1},-\hat{\tau}_{2}\}$.
\item Draw a random sample $\{(\tilde{W}_{i},\tilde{Y}_{i,1},\tilde{Y}_{i,0})\}_{i=1}^{n}$
with replacement from the data and compute $\sup_{q\in\mathbb{R}^{K}}\tilde{S}_{n}(q)$,
where $\tilde{S}_{n}(q)=\sqrt{n}\left(n^{-1}\sum_{i=1}^{n}\mathbf{1}\{\tilde{W}_{i}'q>0\}(\tilde{Y}_{i,1}-\tilde{Y}_{i,0})-\hat{\rho}_{n}(q)\right)$.
\item Repeat the previous step many times and compute $c_{1-\alpha}$, the
$(1-\alpha)$ quantile of $\sup_{q\in\mathbb{R}^{K}}\tilde{S}_{n}(q)$.\footnote{In practice, optimization over $q\in\mathbb{R}^{K}$ can be done by a grid
search over $N$ points on the unit sphere $\{v\in\mathbb{R}^{K}:\ \|v\|_{2}=1\}$.
The computation seems reasonably fast for a normal sample size. For
example, in the setting of Section \ref{sec: monte carlo} with $\beta_{2}=1$,
$n=5000$ and grid size $N=2000$, computing $\sup_{q}\tilde{S}_{n}(q)$
with 500 bootstrap samples take less than 17 seconds in Matlab on
a 2022 MacBook Pro laptop. Further speedups are possible depending
on $n$, $N$ and available memory.}
\item The $(1-\alpha)$-confidence interval for $\tau_{*}$ is $\left[\max(\hat{\tau}_{*}-c_{1-\alpha}n^{-1/2},0),\ 1\right]$.
\end{enumerate}
\end{lyxalgorithm}
Notice that computing $\hat{\tau}_{1}$ is equivalent to computing
the maximum score estimator\footnote{Notice that maximizing $\hat{\rho}_{n}(q)$ is equivalent to maximizing
$n^{-1}\sum_{i=1}^{n}{\rm sgn}(W_{i}'q)(Y_{i,1}-Y_{i,0})$ because
${\rm sgn}(W_{i}'q)=-1+2\cdot\mathbf{1}\{W_{i}'q\geq0\}$.} and all the existing computational algorithms and software for the
maximum score estimator can be used. Finding $\hat{\tau}_{2}$ also
reduces to computing the maximum score estimator once we swap $Y_{1}$
and $Y_{0}$. The above algorithm can be justified by the following
result.
\begin{thm}
\label{thm: test}Let Assumption \ref{lem: manski lem 1} hold. Assume
that $P(Y_{1}=Y_{0})<1$. Then
\[
\limsup_{n\rightarrow\infty}P\left(\sqrt{n}(\hat{\tau}_{*}-\tau_{*})>c_{1-\alpha}\right)\leq\alpha.
\]
\end{thm}
Note that the cube-root asymptotics typically associated with the
maximum score estimator (see e.g., \citet{kim1990cube} and \citet{seo2018local})
does not arise here. The reason is that $\hat{\tau}_{1}$ is the maximum
of $\hat{\rho}_{n}(\cdot)$, rather than the argmax. The cube-root
asymptotics of the maximum score estimator is due to the non-standard
rate of some terms in the expansion for analyzing the argmax. Fortunately,
we do not have to deal with such terms for our purpose.

It is worth pointing out that Algorithm \ref{alg: test } is asymptotically
exact in the sense that there exists a data-generating process under
which the asymptotic coverage probability is exactly $\alpha$. For
example, if $\rho(q)=0$ for any $q$, then $\sqrt{n}(\hat{\tau}_{*}-\tau_{*})=\sup_{q\in\mathbb{R}^{K}}S_{n}(q)$
and $P\left(\sqrt{n}(\hat{\tau}_{*}-\tau_{*})>c_{1-\alpha}\right)\rightarrow\alpha$.
\begin{comment}
If the support of $Z_{i}$ (with $W_{i}=(1,Z_{i}')'$) is not convex
but contains a convex open neighborhood of zero, we can restrict $Z_{i}$
to this neighborhood.
\end{comment}
\begin{comment}
We now revisit the empirical example in Section \ref{sec: test logistic}.
Since we have seen strong evidence against the logistic assumption,
we would like to check whether $\beta$ can be identified without
imposing parametric restrictions on the error terms. To accommodate
multiple time periods, we modify the test statistic by using
\[
\hat{\rho}_{n}(q)=n^{-1}\sum_{i=1}^{n}\sum_{s>t}\mathbf{1}\{(X_{i,s}-X_{i,t})'q\geq0\}(Y_{i,s}-Y_{i,t}).
\]
The bootstrapped version is adjusted accordingly. We now report the
p-value based on 2000 bootstrap samples.
\begin{center}
\begin{tabular}{cc}
Null hypotheses & p-value\tabularnewline
\hline
$H_{0}:\sup_{q\in\mathbb{R}^{K}}\rho(q)\leq0$ & 0.000\tabularnewline
$H_{0}:\inf_{q\in\mathbb{R}^{K}}\rho(q)\geq0$ & 0.027\tabularnewline
\hline
\hline
 & \tabularnewline
\end{tabular}
\par\end{center}
We see that at least at 5\% significance level, we reject the null
hypothesis of
\end{comment}
\begin{comment}
\begin{rem}
An alternative test is to construct a one-sided confidence interval
for $\tau_{1}$ and then for $\tau_{2}$. However, doing inference
simultaneously on both $\tau_{1}$ and $\tau_{2}$ is not straight-forward
and might invoke Bonferroni-type methods that sacrifice power. Instead,
Algorithm \ref{alg: test } circumvents this problem and gives an
asymptotically exact test.
\end{rem}
testing $E(Y_{1}-Y_{0}\mid X)\leq0$ almost surely can be done by
testing $\sup_{q\in\mathbb{R}^{K}}\rho(q)\leq0$. This test amounts to constructing
a one-sided confidence interval for $\sup_{q\in\mathbb{R}^{K}}\rho(q)$.
A similar strategy can be done for testing $E(Y_{1}-Y_{0}\mid X)\geq0$
almost surely.

Since the sign saturation condition is split into two hypotheses here,
one could consider a Bonferroni-type correction. However, Bonferroni-type
procedures can have low power. We now construct one statistic that
directly tests the sign saturation condition.

By Lemma \ref{lem: test}, $P\left(E(Y_{1}-Y_{0}\mid X)>0\right)=0$,
i.e., $P\left(E(Y_{1}-Y_{0}\mid X)\leq0\right)=1$, is equivalent
to $\tau_{1}\leq0$, where $\tau_{1}=\sup_{q\in\mathbb{R}^{K}}\rho(q)$.
Similarly, $P\left(E(Y_{1}-Y_{0}\mid X)<0\right)=0$, i.e., $P\left(E(Y_{1}-Y_{0}\mid X)\geq0\right)=1$,
is equivalent to $\tau_{2}\geq0$, where $\tau_{2}=\inf_{q\in\mathbb{R}^{K}}\rho(q)$.
Therefore, the testing problem of
\[
H_{0}:\ P\left(E(Y_{1}-Y_{0}\mid X)>0\right)=0\ \text{or}\ P\left(E(Y_{1}-Y_{0}\mid X)<0\right)=0
\]
versus
\[
H_{1}:\ P\left(E(Y_{1}-Y_{0}\mid X)>0\right),P\left(E(Y_{1}-Y_{0}\mid X)<0\right)>0
\]
is equivalent to the testing problem of
\begin{equation}
H_{0}:\ \min\{\tau_{1},-\tau_{2}\}\leq0\quad{\rm versus}\quad H_{1}:\ \min\{\tau_{1},-\tau_{2}\}>0.\label{eq: null hypo}
\end{equation}

To construct a test statistic, we consider the natural estimate $\min\{\hat{\tau}_{1},-\hat{\tau}_{2}\}$,
where $\hat{\tau}_{1}=\sup_{q\in\mathbb{R}^{K}}\hat{\rho}(q)$, $\hat{\tau}_{2}=\inf_{q\in\mathbb{R}^{K}}\hat{\rho}(q)$
and
\[
\hat{\rho}_{n}(q)=n^{-1}\sum_{i=1}^{n}\mathbf{1}\{W_{i}'q\geq0\}(Y_{i,1}-Y_{i,0}).
\]

The next question is how to construct the critical value. It turns
out that a one-sided confidence interval for $\min\{\tau_{1},-\tau_{2}\}$
is $[\min\{\hat{\tau}_{1},-\hat{\tau}_{2}\}-c_{1-\alpha}^{*},1]$,
where $c_{1-\alpha}^{*}$ satisfies $P(\sup_{q\in\mathbb{R}^{K}}S_{n}(q)>c_{1-\alpha}^{*})=\alpha$
with $S_{n}(q)=\sqrt{n}(\hat{\rho}_{n}(q)-\rho(q))$. Since $S_{n}(\cdot)$
is a simple empirical process, we can use a nonparametric bootstrap
to obtain $c_{1-\alpha}^{*}$. We summarize the test below.
\end{comment}

Our results characterize what drives the identification and introduce
$\tau_{*}$ as a measure for the identification strength. This theoretical
insight motivates a preliminary diagnostic step in applied work. Reporting
the confidence interval of $\tau_{*}$ in Algorithm \ref{alg: test }
can offer important insight on the identification strength. For example,
if this confidence interval contains zero, then one should be concerned
about identification failure.

However, testing $\tau_{*}=0$ might not detect all the problematic
cases. In practice, the identification might not be clear-cut and
additional caution is warranted if we worry about the ``weak'' identification
scenario, which can complicate econometric analysis due to the generic
issue of post-selection inference, see e.g., \citet{Leeb2005}.

Fortunately, since $\tau_{*}$ can be learned from the data, one option
is to proceed with the identification-based inference (i.e., cube-root
asymptotics) only when the identification is strong enough. For instance,
one might adopt a robust default inference method and switch to the
cube-root approach only when a test fails to reject $H_{0}:\ \tau_{*}\geq\tau_{0}$,
where $\tau_{0}>0$ is a pre-specified fixed threshold. This is uniformly
valid asymptotically as it rules out $\tau_{*}$ being ``local''
to zero.

This strategy is conceptually analogous to a common treatment of weak
instruments in the linear instrumental variable (IV) models: low correlation
between the instrument and the exogenous variable suggests potential
identification failure and one solution is to proceed with the classical
asymptotics only when this correlation is large enough, e.g., measured
by the first-stage $F$-statistic \citep{stock2002testing}.

This raises a natural question: what should this default identification-robust
method be? In the IV literature, approaches such as the Anderson-Rubin
test have been developed. In our model, one could consider inverting
the maximum score criterion function, in a manner similar to the Anderson--Rubin
approach. While more refined methods may be possible, they would require
substantial additional analysis on asymptotic theory, which lies beyond
the current focus on characterizing identification and are left for
future research.
\begin{comment}
More broadly, the results here still offer applied researchers a practical
diagnostic tool to rule out some (but not all) problematic situations.
If the confidence interval for $\tau_{*}$ in Algorithm \ref{alg: test }
contains zero, then we should definitely worry about identification
and avoid cube-root asymptotics.

includes zero, this would signal insufficient identification strength,
suggesting that cube-root asymptotics should be avoided.

On the other hand, the results here can still provide applied researchers
with helpful diagnostic. If the identification does not seem

Here, $\tau_{*}=0$ corresponds to the failure of the sign saturation
condition and thus implies lack of point-identification; similarly,
zero correlation between the IV and the exogenous variable implies
the identification failure in an IV model. To ensure uniform validity
of inference, one can choose to apply the classical asymptotic theory
only when the correlation between the IV and the exogenous variable
is above a fixed positive threshold. Of course, this approach does
not provide a concrete answer to how to proceed when the identification
is not strong enough. In the IV literature, various identification-robust
solutions, such as the Anderson-Rubin test, have been developed. In
our model, one could consider inverting the maximum score criterion
function, in a manner analogous to the Anderson--Rubin approach.
While more refined methods may be possible, they would require substantial
technical analysis on inference that deviates from the current focus
on characterizing identification---a direction we leave to future
research.
\end{comment}
{}
\begin{comment}
For the weak IV problem, one solution is to check whether the IV is
strong enough and proceed accordingly, such as Stock and Yogo (2005)??.
In both problems, the identification strength can be assessed from
the data. In our model, we can learn $\tau_{*}$; in the IV model,
we can learn the correlation between the IV and the exogenous variable.
We recommend the following approach for inference: start with an identification-robust
method for inference, check the identification strength and only switch
to the cubic-root asymptotics when we are satisfied with strong identification.
\end{comment}
\begin{comment}
\textbf{ADD weak identification discussion.} We should set the threshold
to be high enough.
\end{comment}
\begin{comment}
this approach does not provide a complete answer. What should we do
when the identification is not strong enough?
\end{comment}
\begin{comment}
For example, suppose that we have an identification-robust inference
method as a default method. We can set a threshold $\tau_{0}$ and
only switch from the default method to a method based on cube-root
asymptotics when we fail to reject
\end{comment}


\section{\label{sec: monte carlo}Numerical illustration}

We have seen that $\tau_{*}$ is a measure of the sign saturation
condition and thus serves a gauge for identification. Using Monte
Carlo simulations, we now illustrate how $\tau_{*}$ relates to the
identification strength and the accuracy of the cube-root asymptotics.

Consider the model in (\ref{eq: FE model}) with $X_{t}=(\mathbf{1}\{t=0\},X_{2,t})'$
for $t\in\{0,1\}$, where $X_{2,t}$ is from the uniform distribution
on $[-1,1]$, $u_{t}$ is from the standard normal distribution and
$\alpha_{t}=(X_{2,0}+X_{2,1})/2$. Here, $X_{2,0}$, $X_{2,1}$, $u_{0}$
and $u_{1}$ are mutually independent. We set $\beta=(1,\beta_{2})'$.
Therefore, $W'\beta=1+\beta_{2}(X_{2,1}-X_{2,0})$. Clearly, the support
of $W'\beta$ is $[-2|\beta_{2}|+1,2|\beta_{2}|+1]$, which means
that the sign saturation holds if and only if $|\beta_{2}|>0.5$.
In Figure \ref{fig: ID strength}, we plot $\tau_{*}$ (computed using
$\hat{\tau}_{*}$ with $n=10^{8}$) as a function of $\beta_{2}$.
It confirms the intuition that $\beta_{2}=0.5$ corresponds to $\tau_{*}=0$,
which means that the sign saturation condition fails. Higher value
of $\beta_{2}$ corresponds to a higher degree to which the sign saturation
condition holds.

\begin{figure}
\caption{\label{fig: ID strength}Coefficient and identification}

\centering{}\includegraphics[scale=0.6]{identification}
\end{figure}

When the identification holds (together with other regularity conditions),
one can expect the cube-root asymptotics. Given a particular value
of $\tau_{*}$, we compare the finite-sample distribution of the maximum
score estimator with its asymptotic distribution. Explicitly deriving
the cube-root asymptotic distribution requires quantities that are
difficult to compute. To circumvent this problem, we treat the distribution
of the estimator with a sample size of $n=1,000,000$ as the asymptotic
distribution.\footnote{Here, the goal is to assess the importance of identification by examining
how well the asymptotic theory applies. In practice, there is another
issue of approximating the asymptotic distribution either with sub-sampling
or with a modified bootstrap.} The finite-sample distribution is computed using $n=1,000$. To make
these two distributions comparable, we compare the maximum score estimator
$\hat{\beta}_{2}$ and consider the distribution of $n^{1/3}(\hat{\beta}_{2}-\beta_{2})$,
which should converge to a limiting distribution under the cube-root
asymptotics. We present the comparison in Figure \ref{fig: dist}
for $\beta_{2}\in\{0.6,\ 1\}$ based on 5,000 repetitions. We see
that when the identification is stronger ($\beta_{2}=1$) and the
finite-sample distribution of the maximum score estimator is better
approximated by the asymptotic distribution.
\begin{comment}

\section{Discussions}

This paper establishes the sign saturation condition as the key identification
condition and provides a simple way of checking this condition in
the data. Although the estimation and inference has been studied assuming
identification, we are not aware of any previous works that prescribe
a formal procedure for checking the identification. In practice, the
identification might not be a clear-cut issue and extra caution should
be exercised if we worry about the ``weak'' identification scenario
due to the post-selection inference problem by ???.

We draw a parallel to the identification problem with weak instrumental
variables (IV's). Notice that $\tau_{*}=0$ (defined in Section \ref{sec: test})
denotes the failure of the sign saturation condition and thus of point-identification.
Similarly, zero correlation between the IV and the exogenous variable
implies the identification failure in an IV model. For the weak IV
problem, one solution is to check whether the IV is strong enough
and proceed accordingly, such as Stock and Yogo (2005)??. In both
problems, the identification strength can be assessed from the data.
In our model, we can learn $\tau_{*}$; in the IV model, we can learn
the correlation between the IV and the exogenous variable. We recommend
the following approach for inference: start with an identification-robust
method for inference, check the identification strength and only switch
to the cubic-root asymptotics when we are satisfied with strong identification.

We now outline the details of this approach and establish its validity.
First, an identification-robust inference approach can be constructed
in a way similar to the Anderson-Rubin test in the linear IV model.
By \citet{pakes2024moment}, the identified set for $\beta$ is $\arg\max_{q}\rho(q)$.
\begin{thm}
Let Assumption \ref{lem: manski lem 1} hold. Assume that $P(Y_{1}=Y_{0})<1$.
Consider $c_{1-\alpha}$ defined in Algorithm \ref{alg: test }. Define
\begin{equation}
CS(1-\alpha)=\left\{ b:\ \hat{\rho}(b)\geq\sup_{q}\hat{\rho}(q)-n^{-1/2}c_{1-\alpha}\right\} .\label{eq: CS}
\end{equation}
Then
\[
\limsup_{n\rightarrow\infty}P\left(\beta\in CS(1-\alpha)\right)\geq1-\alpha.
\]
\end{thm}
Therefore, Algorithm \ref{alg: test } gives us the identification-robust
confidence set in (\ref{eq: CS}) as well as a confidence interval
for $\tau_{*}$, which is $[\hat{\tau}_{*}-n^{-1/2}c_{1-\alpha},1]$.
 Suppose that we give a

Alternatively, the literature has recommended an identification-robust
approach using methods such as inverting the Anderson-Rubin test???.

Our recommendation is to start with an identification-robust approach
here. If we have further evidence of strong idenfication, we can adopt

results in this paper suggest the following steps in practice. We
can start with blindly computing the maximum score estimator without
worrying about whether the identification fails. Then we check the
identification by verifying the sign saturation condition using Algorithm
\ref{alg: test }; in essence, the test statistic and the critical
values are essentially the objective function values of the maximum
score estimator and its bootstrapped version. Once we are confident
that the model is identified, the estimate we already computed is
econometrically justified and the bootstrap procedure in Algorithm
\ref{alg: test } can also be used for inference via a test inversion.

In this paper, we do not consider non-standard asymptotics under which
$P(E(Y_{1}-Y_{0}\mid X)>0)$ and $P(E(Y_{1}-Y_{0}\mid X)<0)$ are
local to zero in some sense. A similar issue arises in pre-tests for
the strength of instruments in instrumental variables regressions.
Perhaps a simple idea is to employ confidence intervals. For example,
the one-sided confidence interval with nominal coverage $1-\alpha$
for $\min\{\tau_{1},-\tau_{2}\}$ is simply $[\min\{\hat{\tau}_{1},-\hat{\tau}_{2}\}-n^{-1/2}c_{1-\alpha},1]$.
Even if we fail to reject $H_{0}$ in (\ref{eq: null hypo}), we would
still like to see a relatively high value of $\min\{\hat{\tau}_{1},-\hat{\tau}_{2}\}-n^{-1/2}c_{1-\alpha}$
to be comfortable with the identification strength. We will leave
a more complete solution to this issue as future research.
\end{comment}

\begin{figure}
\caption{\label{fig: dist}Distribution of $n^{1/3}(\hat{\beta}_{2}-\beta_{2})$
for $\beta_{2}=0.6$ (left) and $\beta_{2}=1$ (right)}

\centering{}\includegraphics[scale=0.4]{Plot_maxscore_beta_060}\includegraphics[scale=0.4]{Plot_maxscore_beta_100}
\end{figure}