EconBase
← Back to paper

Identification of Nonlinear Dynamic Panels under Partial Stationarity

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

108,505 characters

Identification in Nonlinear Dynamic Panel Models under Partial Stationarity


\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\def\mc#1{\mathscr{#1}}
\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\def\abs#1{\left|#1\right|}
\global\long\def\norm#1{\left\Vert #1\right\Vert }
\global\long\def\rest#1{\left.#1\right|}
\global\long\def\bracket#1#2{\left\langle #1\middle\vert#2\right\rangle }
\global\long\def\sandvich#1#2#3{\left\langle #1\middle\vert#2\middle\vert#3\right\rangle }
\global\long\def\turd#1{\frac{#1}{3}}
\global\long\global\long\def\sand#1{\left\lceil #1\right\vert }
\global\long\def\wich#1{\left\vert #1\right\rfloor }
\global\long\def\sandwich#1#2#3{\left\lceil #1\middle\vert#2\middle\vert#3\right\rfloor }
\global\long\def\abs#1{\left|#1\right|}
\global\long\def\norm#1{\left\Vert #1\right\Vert }
\global\long\def\rest#1{\left.#1\right|}
\global\long\def\inprod#1{\left\langle #1\right\rangle }
\global\long\def\ol#1{\overline{#1}}
\global\long\def\ul#1{\underline{#1}}
\global\long\def\td#1{\tilde{#1}}
\global\long\def\bs#1{\boldsymbol{#1}}
\global\long\global\long\global\long\global\long\global\long
\setlength{\abovedisplayskip}{6pt}
\setlength{\belowdisplayskip}{6pt}
\title{Identification in Nonlinear Dynamic Panel \\Models under Partial Stationarity\thanks{We are grateful to Jason Blevins, Xiaohong Chen, Bryan Graham, Jiaying
Gu, Robert de Jong, Shakeeb Khan, Kyoo il Kim, Louise Laage, Xiao
Lin, Eric Mbakop, Chris Muris, Adam Rosen, Frank Schorfheide, Liyang
Sun, Elie Tamer, Valentin Verdier, Jeffrey Wooldridge, as well as numerous
seminar and conference participants for helpful comments and suggestions.}}
\author{Wayne Yuan Gao\thanks{Department of Economics, University of Pennsylvania, 133 S 36th St,
Philadelphia, PA 19104, USA. Email: [email removed]}$\ \ $and Rui Wang\thanks{This work was completed while Rui Wang was at Ohio State University, prior to joining Amazon. Department of Economics, Ohio State University, 1945 N High St, Columbus, OH 43210, USA. Email: [email removed]} }


\maketitle
\begin{abstract}
This paper provides a general identification approach for a wide range of nonlinear panel data models, including binary choice, ordered response, and other types of limited dependent variable models. Our approach accommodates dynamic models with any number of lagged dependent variables as well as other types of endogenous covariates. Our identification strategy relies on a partial stationarity condition, which allows for not only an unknown distribution of errors, but also temporal dependencies in errors. We derive partial identification results under flexible model specifications and establish sharpness of our identified set in the binary choice setting. We demonstrate the robust finite-sample performance of our approach using Monte Carlo simulations, and apply the approach to the empirical analysis of income categories using various ordered choice models.


\noindent \textbf{~}\\
 \textbf{Keywords}: Panel Discrete Choice Models; Stationarity; Dynamic
Models; Partial Identification; Endogeneity
\end{abstract}
\newpage{}

\section{\label{sec:intr}Introduction}

This paper provides a general and unified identification approach
for a wide range of panel data models with limited dependent variables,
including various discrete (binary, multinomial, and ordered) choice
models and censored outcome models. In particular, our approach accommodates dynamic models with any number of lagged dependent variables and contemporaneously endogenous covariates. Moreover, the identification
approach does not impose parametric distributions on unobserved heterogeneity, nor on the exact form of endogeneity, thus allowing for more flexible model specifications.

To fix ideas, we start with the following dynamic binary choice model,
which is on its own of considerable theoretical and applied interest.
Section \ref{sec:gene} generalizes the approach to other limited
dependent variable models. Specifically, consider
\begin{equation}
Y_{it}=\mathbf{\mathbbm1}\left\{ W_{it}^{'}\theta_{0}+\alpha_{i}+\epsilon_{it}\geq0\right\} ,\label{eq:bin}
\end{equation}
where $Y_{it}\in\left\{ 0,1\right\} $ denotes a binary outcome variable
for individual $i=1,2,...$ and time $t=1,...,T$, $W_{it}\in\mathcal{R}^{d_{w}}$
denotes a vector of observed covariates, $\alpha_{i}\in\mathcal{R}$ denotes
the unobserved fixed effect for individual $i$, and $\epsilon_{it}$
denotes the unobserved time-varying error term for individual $i$
at time $t$. The objective is to identify the parameter $\theta_{0}$\footnote{We discuss in Appendix \ref{subsec:ID_CF} how our results can be used to derive bounds on certain counterfactual parameters.}
using a panel of observed variables $\left(Y_{i},W_{i}\right){}_{i=1}^{n}$,
where $W_{i}:=\left(W_{i1},...,W_{iT}\right)$, and similarly for
$Y_{i}$.  We focus on short panels, where the number of time periods
$T\geq2$ is fixed and finite.

The identification of model \eqref{eq:bin} has been explored in the
literature under various assumptions. For example, \citet{chamberlain1980}
examines identification under the logistic distribution of $\epsilon_{it}$
and the independence of $\epsilon_{it}$ with respect to $\left(\alpha_{i},W_{i}\right)$.
Subsequently, \citet{manski1987} relaxes the distributional assumption
and employs the following conditional stationarity of $\epsilon_{it}$
to achieve identification:
\begin{equation}
\epsilon_{is}\,\sim\,\epsilon_{it}\mid\alpha_{i},W_{i}\quad\forall s,t=1,...,T\label{eq:cond_sta}
\end{equation}
This condition is also referred to as ``group stationarity'' or
``group homogeneity'' and has been exploited in studies such
as \citet{chernozhukov2013}, \citet*{shi2018} and \citet{pakes2019}.\footnote{To be precise, condition \eqref{eq:cond_sta} is often stated in the
following weaker ``pairwise'' version in the literature,
\[
\epsilon_{is}\,\sim\,\epsilon_{it}\mid\alpha_{i},W_{is},W_{it},\ \quad\forall s,t=1,...,T,
\]
where only covariate realizations from the two periods $\left(s,t\right)$
are conditioned on. However, the difference between condition \eqref{eq:cond_sta}
and the pairwise version above usually only leads to a minor adaptation
of the results in the aforementioned papers (as well as in the current
one). See Remark \ref{rem:PairPS} for a follow-up discussion.} Condition \eqref{eq:cond_sta} does not impose parametric restrictions
on the distributions of $\epsilon_{it}$ and allows for dependence between the
fixed effect $\alpha_{i}$ and the covariates $W_{i}$. However, condition
\eqref{eq:cond_sta} does impose substantial restriction on the dependence
between $W_{i}$ and the time-varying error term $\epsilon_{it}$: it effectively
requires that all covariates in $W_{i}$ are exogenous with respect
to the time varying error $\epsilon_{it}$.\footnote{For instance, suppose $W_{it}=\left(Z_{it},X_{it}\right)$ and $\mathbb{E}\left[\rest{\epsilon_{it}}W_{i}\right]=X_{it}^{'}\eta$,
then the conditional distributions of $\epsilon_{it}$ and $\epsilon_{is}$ cannot
be the same as long as $X_{it}^{'}\eta\neq X_{is}^{'}\eta$. Hence
condition \eqref{eq:cond_sta} fails in general.}

In many economic applications, certain components of the observable
covariates $W_{i}$, namely $X_{i}$, may exhibit endogeneity issues in violation of condition \eqref{eq:cond_sta}. For
example, in a \emph{dynamic} setting where $X_{it}$ includes the
lagged outcome variable $Y_{i,t-1}$, then the dependence  of $Y_{i,t-1}$
on $\epsilon_{i,t-1}$
arises immediately, violating the condition \eqref{eq:cond_sta}.\footnote{We note that lagged outcome variable $Y_{it}$ are sometimes referred to as ``predetermined'' or ``weakly or sequentially exogenous'' covariates in panel data literature. We clarify that in this paper, we use the words  ``exogenous'' and ``endogenous'' in the stronger ``all-period'' sense, or more precisely, in the sense of whether condition \eqref{eq:cond_sta} holds when conditioned on the covariate in question.} For another example, if $X_{it}$ includes ``price''
or other variables that may be endogenously chosen by economic agents
after observing $\epsilon_{it}$, then $X_{it}$ would be correlated with the
contemporaneous $\epsilon_{it}$, so the exogeneity restriction imposed by
condition \eqref{eq:cond_sta} will again fail to hold.

In this paper, we instead impose and exploit a weaker version of
condition \eqref{eq:cond_sta}  by excluding all endogenous components
of $W_{i}$ from the conditioning set. To be precise,
from now on we suppose that we can decompose $W_{it}$ as:
\[
W_{it}\equiv\left(Z_{it},X_{it}\right),
\]
where $Z_{it}$ is of dimension $d_{z}$, and $X_{it}$ is of dimension
$d_{x}$ with $d_{w}=d_{z}+d_{x}$. Our ``partial stationarity''
assumption is then formulated as follows:
\begin{equation}
\epsilon_{is}\sim\epsilon_{it}\mid\alpha_{i},Z_{i},\quad\forall s,t=1,...,T.\label{eq:part_stat}
\end{equation}
Our partial stationarity condition \eqref{eq:part_stat}, as its name
suggests, only requires that the errors are stationary conditional
on the realizations of a subvector of the covariates (i.e., the exogenous
covariates denoted by $Z_{i}$) while allowing the remaining covariates
(denoted by $X_{i}$) to be endogenous in arbitrary manners.\footnote{Our identification strategy and results can be easily adapted under
the alternative ``pairwise partial stationarity'' condition $\epsilon_{is}\sim\epsilon_{it}\mid\alpha_{i},Z_{is},Z_{it}.$
See Remarks \ref{rem:PairPS} and \ref{rem:PairPS_prop1} for follow-up
discussions.} In short, condition \eqref{eq:part_stat} imposes exogeneity conditions
only on a subset of covariates, which we label as ``$Z$'' and will thereafter be referred to as the ``exogenous covariates''.\footnote{Condition \eqref{eq:part_stat} also accommodates the standard stationarity
assumption conditional on all covariates.}

We describe how to exploit the partial stationarity condition \eqref{eq:part_stat}
to derive the identified set on the model parameters $\theta_{0}$
through a class of conditional moment inequalities, which take the
form of lower and upper bounds for the conditional distribution $\epsilon_{it}+\alpha_{i}\mid Z_{i}$,
solely as functions of observed variables and the model parameters
$\theta_{0}$. We show that these bounds must have nonzero intersections
over time under the partial stationarity assumption, thereby forming
a class of identifying restrictions for the parameter $\theta_{0}$. Conditional
on the exogenous covariates $Z_{i}$, our class of inequalities is
indexed by a scalar $c\in\mathcal{R}$, which implicitly traces out all possible
values that the parametric index $Z_{it}^{'}\beta_{0}+X_{it}^{'}\gamma_{0}$
can take. That said, we show how the\emph{ effective }number of identifying
restrictions can be reduced to be finite when $X_{it}$ has finite
support, a condition naturally satisfied in the important special
case of ``$p$-th order autoregressive'' dynamic binary choice models,
where $X_{it}$ consists of lagged outcome variables $Y_{i,t-1},Y_{i,t-2},...,Y_{i,t-p}$
that are by construction discrete.

We demonstrate the sharpness of the identified set we derived for
binary choice models. More precisely, we show that for any $\theta$
that satisfies all the conditional moment inequalities we derived,
we can construct an observationally equivalent joint distribution
of the observed and unobserved variables in our model. Our proof of
sharpness consists of three main steps: we begin by demonstrating
``per-period'' sharpness, and then progressively generalize the
result from ``per-period'' to ``all-period'' sharpness, and from
discrete $X_{it}$ to general $X_{it}$. A key innovation in our proof
technique is using an explicit, simple, and general construction that
shows how marginal/aggregate stationarity restrictions and joint choice
probability restrictions can be satisfied simultaneously, which might
be of independent and wider use.

Our identification strategy based on partial stationarity applies
more broadly beyond the context of dynamic binary choice models. In
Section \ref{sec:gene}, we demonstrate its applicability in a general
nonseparable semiparametric model, and show how
it can be applied to a range of alternative limited dependent variable
models, such as ordered response models, multinomial choice models,
and censored outcome models. The results of our approach accommodate
both static and dynamic settings across all these models.

We characterize the identified set using a collection of conditional
moment inequalities, based on which estimation and inference can be
conducted using established econometric methods in the literature,
such as \citet*{chernozhukov2007}, \citet{andrews2013}, and \citet*{chernozhukov2013}.
Through Monte Carlo simulations, we demonstrate that our identification
method yields informative and robust finite-sample confidence intervals
for coefficients in both static and dynamic models.

\subsubsection*{Literature Review}

Our paper contributes directly to the line of econometric literature
on semiparametric panel discrete choice models. Dating back to \citet{manski1987},
a series of work exploits ``full'' stationarity conditions for identification,
such as \citet{abrevaya2000rank}, \citet*{chernozhukov2013}, \citet*{shi2018},
\citet{pakes2019}, \citet*{khan2021inference},
\citet{gao2020}, \citet{wang2022semiparametric}, and \citet*{botosaru2023identification}.
As discussed above, full stationarity conditions given all observable
covariates effectively require that all covariates are exogenous with
no dynamic effects (i.e., lagged dependent variables). In contrast,
we exploit the ``partial'' stationarity condition, allowing
for lagged dependent variables, as well as other endogenous covariates.

In the literature on dynamic discrete choice models, our paper is
most closely related to \citet*[KPT thereafter]{khan2023identification},
who studies the following dynamic panel binary choice model
\begin{equation}
Y_{it}=\mathbf{\mathbbm1}\left\{ Z_{it}^{'}\beta_{0}+Y_{i,t-1}\gamma_{0}+\alpha_{i}+\epsilon_{it}\geq0\right\} ,\label{eq:KPT}
\end{equation}
where the one-period lagged dependent variable $Y_{i,t-1}\in\{0,1\}$
serves as the endogenous covariate, and $Z_{it}$ are exogenous covariates.
KPT exactly imposes the ``partial stationarity'' condition \eqref{eq:part_stat}
in the specific context of \eqref{eq:KPT}, and derives the sharp
identified set for $\theta_{0}$ by explicitly enumerating the realizations
of the one-period lagged outcome variable $Y_{i,t-1}$ (across two
periods $t,s$). In contrast, our model \eqref{eq:bin}, along with
the ``partial stationarity'' condition, is stated with more general
specifications of the endogenous covariates $X_{it}$. The covariates
$X_{it}$ can include more than one lagged dependent variables (e.g.
$Y_{i,t-1},Y_{i,t-2},$ ...) and other endogenous variables (such
as ``price'' if $Y_{it}$ represents the purchase of a particular
product), which may be continuously valued. Consequently, our identification
strategy is substantially different from that of KPT, and can be applied
more broadly to various other dynamic nonlinear panel models. In the
specific model \eqref{eq:KPT}, we show that the identifying restrictions
we derived are equivalent to those derived in KPT and thus both approaches
lead to sharp identification. Relatedly, \citet{mbakop2023identification}
proposes a computational algorithm to derive conditional moment inequalities
in a general class of dynamic discrete choice models (potentially
with multiple lags). The focus of \citet{mbakop2023identification}
is on scenarios where the lagged discrete outcome variables are the
only endogenous covariates in the model, and the proposed algorithm
relies on the discreteness of these variables. Relative to these works,
our paper introduces an analytic approach that directly applies to
a more general class of dynamic binary choice models, as well as other
types of models with continuous limited dependent variables and any
number of endogenous covariates, regardless of whether they are discrete
or continuous and whether they take the form of lagged outcome variables
or not.

Our identification strategy relies on a type of stationarity condition,
while alternative approaches utilize other notions of exogeneity.
For example, \citet{honore2000panel} provides identification by exploiting
events where the exogenous covariates stay the same across two periods:
they consider both the logit case and a semiparametric case, but both
under the independence between time-changing errors and the lagged
dependent variable, as well as the intertemporal independence of errors.
Additionally, \citet{aristodemou2021semiparametric} exploits an alternative
type of ``full independence'' assumption for identification in dynamic binary
choice models. The ``full independence'' assumption essentially
requires that the time-varying errors from all time periods and the
exogenous variables from all time periods are independent (conditional
on initial conditions), but does not make intertemporal restrictions
on the errors (such as stationarity). Hence, such `full independence''
assumption and the partial stationarity assumption in our paper do
not nest each other as special cases. \citet*{chesher2023identification}
applies the framework of generalized instrumental variables \citep{chesher2017generalized}
to the context of various dynamic discrete choice models with fixed
effects, and utilizes a similar ``full independence'' assumption
\citep{aristodemou2021semiparametric} for identification.\footnote{Our identification strategy shares some conceptual similarity with
the idea of generalized instrumental variable (GIV) in \citet*{chesher2017generalized},
who proposes a general approach for representing the identified set
of structural models with endogeneity.
\citet*{chesher2017generalized}, \citet*{chesher2020generalized},
and \citet*{chesher2023identification} demonstrate how the GIV framework
can be applied to various settings, but focus mostly on the use of
exclusion restrictions and/or full independence assumptions. In this
paper, we neither impose exclusion restrictions nor independence assumptions
but instead explore identification under a partial stationarity condition.} More differently, some other papers work with sequential exogeneity
in various dynamic panel models and provide (non-)identification results
under different model restrictions. For example, \citet{shiu2013identification}
imposes a high-level invertibility condition along with a restriction
that rules out the dependence of covariates on past dependent variables.
More recently, \citet*{bonhomme2023identification} investigates panel
binary choice models with a single binary predetermined covariate
under \emph{sequential exogeneity}, whose evolution may depend on
the past history of outcome and covariate values. The sequential exogeneity
condition considered in these papers and the partial stationarity
condition in ours again do not nest each other as special cases: in
particular, our current paper accommodates contemporaneously endogenous
covariates that violate sequential exogeneity. In summary, the key
assumptions, identification strategy, and identification results of
these studies are substantially different from and thus not directly
comparable to those in our current paper.


\begin{comment}
That said, our identification strategy is conceptually related to
the idea of generalized instrumental variable (GIV) in \citet*{chesher2017generalized},
who proposes a general approach for representing the identified set
of a structural model with endogeneity as an infinite collection of
inequalities about conditional probabilities of the structural errors
taking values in an arbitary given set (conditional on the exogenous
variables). discuss how the approach can be used to obtain sharp identified
set under various model specifications and structural assumptions,
and their subsequent work applies the approach to specific settings.
\end{comment}

Our paper is also complementary to the literature that studies
dynamic logit models with fixed effects. This literature typically assumes that time-varying errors are conditionally independent across time, independent from all other variables, and follow the logistic distribution. The
logit assumption in panel settings has long been studied, such
as in \citet*{chamberlain1984panel} and \citet*{chamberlain2010binary}.
We do not impose the logit assumption, nor require conditional independence
across time, and our identification strategy is very different from
those based on the logit assumption.

Our paper also contributes to the general panel data literature on
linear and nonlinear models with and without endogeneity and dynamics.
Most relatedly, \citet{botosaru2017binarization} proposes a binarization
strategy for general panel data models with fixed effects without
requiring time homogeneity, but focuses on static settings. \citet{botosaru2022}
considers a model where the outcome variable is generated as a strictly
monotone (and thus invertible) transformation of a linear model, and
they exploit time homogeneity in conditional means (instead of the
whole distributions) for identification. Our current paper, with a
focus on discrete choice models, does not require strict monotonicity
and invertibility.

It should be acknowledged that our partial stationarity condition \eqref{eq:part_stat}, as well as the full stationarity condition \eqref{eq:cond_sta}, cannot handle random coefficients as in \cite*{berry1995automobile} and \cite{mcfadden2000mixed}. On the other hand, our approach can be naturally extended to handle nonadditive and functional fixed effects as discussed in Remark \ref{rem:scalaradd} and Section \ref{subsec:generic}. It would be interesting in future research to investigate whether our approach can be further adapted to incorporate heterogeneous coefficients.

~

The rest of the paper is organized as follows. Section \ref{sec:bin}
studies the sharp identification of panel binary choice models with
endogenous covariates. Section \ref{sec:gene} demonstrates how our
key identification strategy generalizes to a wide range of dynamic
nonlinear panel data models, such as ordered response models, multinomial
choice models, and censored outcome models. Section \ref{sec:simu}
presents simulation results about the finite-sample performances of
our approach, and Section \ref{sec:appl} explores the empirical application
of income categories using various ordered response models. We conclude
with Section \ref{sec:conc}.

\section{\label{sec:bin}Dynamic Binary Choice Model}

\subsection{Model}

To explain the partial stationarity and our key identification strategy,
we start with the canonical binary choice model, which is of wide
theoretical interest itself. In Section \ref{sec:gene}, we explain
how our identification strategy can be applied more generally.

Specifically, consider the same binary model as introduced in \eqref{eq:bin}:
$$
Y_{it}=\mathbf{\mathbbm1}\left\{ W_{it}^{'}\theta_{0}+\alpha_{i}+\epsilon_{it}\geq0\right\}.
$$
Recall that we decompose $W_{it}\equiv\left(Z_{it}^{'},X_{it}^{'}\right)^{'}$, and, throughout this paper, we will refer to $Z_{i}$ as ``exogenous covariates'',
and refer to $X_{i}$ as ``endogenous covariates''. The exact difference
between $Z_{i}$ and $X_{i}$ is formalized through the ``\emph{partial stationarity}'' condition, which we now state as a formal assumption:
\begin{assumption}[Partial Stationarity]
 \label{assu:PartStat} The conditional distribution of $\epsilon_{it}\mid Z_{i},\alpha_{i}$
is stationary over time, i.e.,
\[
\epsilon_{it}\mid Z_{i},\alpha_{i}\stackrel{d}{\sim}\epsilon_{is}\mid Z_{i},\alpha_{i}\quad\forall t,s=1,...,T.
\]
\end{assumption}
Assumption \ref{assu:PartStat} essentially requires that the (conditional)
distribution of $\epsilon_{it}$ stays the same across all time periods
$t=1,...,T$ even if $Z_{i}$ realize to different values. To illustrate,
suppose that there are only two periods $t=1,2$, and that $Z_{i1},Z_{i2}$
realize to two values $z_{1},z_{2}$, respectively, with $z_{1}<z_{2}$.
Then Assumption \ref{assu:PartStat} requires that $\epsilon_{i1}$ and
$\epsilon_{i2}$ still have the same (conditional) distributions: in particular,
$\epsilon_{i1}$ cannot be stochastically smaller (or larger) than $\epsilon_{i2}$
because of $z_{1}<z_{2}$. Hence, Assumption \ref{assu:PartStat}
can be thought as a definition of the ``exogeneity'' of the covariates
$Z_{it}$ in our context.

In contrast, Assumption \ref{assu:PartStat} imposes no such restrictions
on the (potentially) endogenous covariates $X_{i}$. In fact, since
$X_{i}$ does not appear in Assumption \ref{assu:PartStat} at all,
here we are completely agnostic about the dependence structure between
$\epsilon_{i}$ and $X_{i}$: in particular, the conditional distribution
of $\epsilon_{it}$ is allowed to vary across $t$ arbitrarily for any particular
realization of $X_{i}$. As a result, different forms of endogeneity
in $X_{i}$ can be incorporated under our framework in a unified manner,
as we illustrate in the examples below.
\begin{example}[Dynamic Effects via Lagged Outcomes]
\label{exa:dyn} Consider the following ``AR(1)'' dynamic binary
choice model studied in \citet*[KPT thereafter]{khan2023identification}:
\[
Y_{it}=\mathbf{\mathbbm1}\left\{ Z_{it}^{'}\beta_{0}+Y_{i,t-1}\gamma_{0}+\alpha_{i}+\epsilon_{it}\geq0\right\} ,
\]
which is a special case of our model with $X_{it}$ set to be the
one-period lagged binary outcome variable $Y_{i,t-1}$. Here, $X_{it}$
is endogenous since $X_{it}\equiv Y_{i,t-1}$ and $\epsilon_{i,t-1}$ is
by construction positively correlated with $Y_{i,t-1}$ for any $t$,
and thus the distribution of $\epsilon_{it}$ cannot be stationary across
$t$ when conditioned on the realizations of $Y_{i0},...,Y_{i,T-1}$.
For example, given $Y_{i0}=Y_{i1}=1,$ $Y_{i2}=0$ (and $Z_{i},\alpha_{i}$),
the conditional distribution of $\epsilon_{i1}$ will naturally be different
from that of $\epsilon_{i2}$. To obtain identification under the endogeneity
of $Y_{i,t-1}$, KPT imposes the stationarity of $\epsilon_{it}$ conditional
on the exogenous covariates $Z_{i}$ only, which coincides with our
``partial stationarity'' condition (Assumption \ref{assu:PartStat})
when specialized to their setting.

A natural generalization of the AR(1) model above in KPT is the following
``AR($p$)'' model, which is again a special case of our model with
$X_{it}$ taken to be the vector of $p$ lagged outcomes $Y_{i,t-1},...,Y_{i,t-p}$:
\[
Y_{it}=\mathbf{\mathbbm1}\left\{ Z_{it}^{'}\beta_{0}+\sum_{j=1}^{p}Y_{i,t-j}\gamma_{j}+\alpha_{i}+\epsilon_{it}\geq0\right\} .
\]
Similarly, $X_{it}$ is endogenous here due to dependence on $\epsilon_{i,t-1},...,\epsilon_{t-p}$,
which can again be handled in our framework under the ``partial stationarity''
assumption. While it is not clear how the identification results in
KPT can be easily generalized to the AR($p$) model above, we show
in the next subsection how our identification strategy provides a
simple and unified approach to derive moment inequalities regardless
of the exact specifications of $X_{it}$.
\end{example}
\begin{example}[Contemporaneously Endogenous Covariates]
\label{exa:ContempEndo} Alternatively, consider the following binary
choice model with contemporaneously endogenous covariates:
\begin{align*}
Y_{it} & =\mathbf{\mathbbm1}\left\{ Z_{it}^{'}\beta_{0}+X_{it}^{'}\gamma_{0}+\alpha_{i}+\epsilon_{it}\geq0\right\} ,\\
X_{it} & =\phi\left(Z_{it},u_{it}\right)
\end{align*}
where $\phi$ is an unknown ``first-stage'' function and $u_{it}$
is allowed to be arbitrarily correlated with $\epsilon_{it}$. For example,
$X_{it}$ may be a ``price '' variable that is strategically chosen
by a decision maker after observing the current-period error $\epsilon_{it}$,
which generates contemporaneous dependence between $X_{it}$ and $\epsilon_{it}$.
Even though contemporaneous endogeneity of this type is very different
in nature from the dynamic endogeneity discussed in the previous example,
it also induces non-stationarity of $\epsilon_{it}$ when conditioned on
$X_{i}$: for example, if $X_{it}$ and $\epsilon_{it}$ are positively
correlated, then, conditional on $X_{i1}<X_{i2}$, it is unreasonable
to assume the distribution of $\epsilon_{i1}$ is the same as $\epsilon_{i2}$.
That said, such type of contemporaneous endogeneity can also be handled
in our framework under the ``partial stationarity'' condition (Assumption
\ref{assu:PartStat}).
\end{example}
\begin{rem}[Combination of Dynamic and contemporaneous Endogeneity]
 We separately discussed two types of endogenous covariates, dynamic
covariates (lagged outcome variables) and contemporaneously endogenous
covariates, in the two examples above, but our identification strategy
also applies if both types of endogenous covariates are present together,
since our identification strategy works generally under ``partial
stationarity'', which does not impose or exploit any restrictions
on the form of endogeneity between $\epsilon_{it}$ and $X_{i}$.
\end{rem}
\begin{rem}[Full Stationarity as Special Case]
 Obviously, the standard ``full stationarity'' condition \eqref{eq:cond_sta}
is nested under ``partial stationarity'' condition (Assumption \ref{assu:PartStat})
as a special case, where the endogenous covariate $X_{it}$ contains
no variables. Hence, ``full stationarity'' is in general stronger
than ``partial stationarity''.
\end{rem}
\begin{rem}[Focus on Time-Varying Endogeneity]
 Technically, our partial stationarity condition also allows some
endogeneity between $\epsilon_{it}$ and $Z_{i}$, as long as such endogeneity
is time-invariant. This is because Assumption \ref{assu:PartStat}
is stated conditional on the full vector $Z_{i}=\left(Z_{i1},...,Z_{iT}\right)$
and the time-invariant fixed effect $\alpha_{i}$. Hence, as long as the
conditional distribution of $\epsilon_{it}$ depends on $Z_{i1},...,Z_{iT}$
and $\alpha_{i}$ in a time-invariant manner, the stationarity of $\epsilon_{it}$
can still hold. That said, since in empirical applications we are
mostly interested in ``time-varying endogeneity'', such as the dynamic
and contemporaneous endogeneity discussed in the examples above, in this
paper we refer to $Z_{i}$ as ``exogenous'' even though it may be
endogenous in a time-invariant manner, and only call $X_{i}$, which
features time-varying endogeneity, the ``endogenous'' covariates.
\end{rem}
\begin{rem}[Pairwise Version of Partial Stationarity]
\label{rem:PairPS} In Assumption \ref{assu:PartStat}, we impose
partial stationarity of $\epsilon_{it}$ conditional on $Z_{it}$ from all
periods $t=1,...,T$. Alternatively, we could impose partial stationarity
in a ``pairwise'' version, conditional on $\left(Z_{it},Z_{is}\right)$
from any pair of time periods $\left(t,s\right)$ only:
\begin{equation}
\text{\textbf{Pairwise Partial Stationarity}:\ensuremath{\quad}}\epsilon_{it}\mid Z_{it},Z_{is},\alpha_{i}\stackrel{d}{\sim}\epsilon_{is}\mid Z_{it},Z_{is},\alpha_{i},\quad\forall t,s=1,...,T.\label{eq:PartStat_ts}
\end{equation}
Clearly, the ``pairwise'' version is equivalent to the ``all-periods''
version when $T=2$, but is weaker when $T\geq3$. Our identification
strategy applies under both versions of partial stationarity, though
the identification results and the corresponding proofs have slightly
different representations. Essentially, conditioning on all-period
covariate realizations would be replaced with conditioning the realizations
in any specific pair of periods. See Remark \ref{rem:PairPS_prop1}
at the end of Section \ref{subsec:ind} for a follow-up discussion.
\end{rem}
\begin{rem}[Initial Conditions in Dynamic Settings]
 \label{rem:initial_cond} In dynamic settings where $X_{it}$ includes lagged
outcome variables such as $Y_{i,t-1}$, the treatment of the initial
condition $Y_{i0}$ warrants some additional discussion. Our current
setup \eqref{eq:bin} treats $X_{it}$ (and the lagged outcome
variables involved) as \emph{observed}\footnote{If only $\left(Y_{i1},...,Y_{iT}\right)$ are observed,
we can truncate the time periods to satisfy this requirement. For
example, in the AR(1) setting, we can treat $Y_{i1}$ as the initial
condition $Y_{i0}$ and relabel periods 2 as period 1.} and \emph{endogenous}. However, one may consider alternative setups
where $Y_{i0}$ is treated as unobserved and/or exogenous. In Appendix
\eqref{subsec:Yi0}, we explain how our approach can be adapted to
such settings.
\end{rem}
\begin{rem}[Scalar Additivity]\label{rem:scalaradd}
 We work with the binary choice model \eqref{eq:bin} with scalar-additive
fixed effects $\alpha_{i}$ and error $\epsilon_{it}$. This restriction is
unnecessary: We explain in Section \ref{sec:gene} that our identification
strategy does not rely at all on the scalar-additivity of $\alpha_{i}$
and $\epsilon_{it}$. However, in this section we stick with the scalar-additive
representation \eqref{eq:bin}, since it is the most standard
specification (or notation) that is adopted in a wide body of work on
binary choice models. It thus provides a context in which most clearly
we can explain our partial stationarity condition in relation to
previous work.
\end{rem}

\subsection{\label{subsec:ind} Key Identification Strategy}

We now explain our key identification strategy based on the partial
stationarity condition. Write $v_{it}:=-\left(\epsilon_{it}+\alpha_{i}\right)$
so that model \eqref{eq:bin} can be rewritten as
\[
Y_{it}=\mathbf{\mathbbm1}\left\{ v_{it}\leq W_{it}^{'}\theta_{0}\right\} .
\]
For any constant $c\in\mathcal{R}$, consider first the event
\[
Y_{it}=1\text{ and }W_{it}^{'}\theta_{0}\leq c.
\]
Whenever the event above happens, we must have $v_{it}\leq W_{it}^{'}\theta_{0}\leq c$,
implying that $v_{it}\leq c$. Formally, the above can be summarized
by the following inequality:
\begin{equation}
\begin{aligned}Y_{it}\mathbf{\mathbbm1}\left\{ W_{it}^{'}\theta_{0}\leq c\right\}  & =\mathbf{\mathbbm1}\left\{ v_{it}\leq W_{it}^{'}\theta_{0}\right\} \mathbf{\mathbbm1}\left\{ W_{it}^{'}\theta_{0}\leq c\right\} \leq\mathbf{\mathbbm1}\left\{ v_{it}\leq c\right\} \end{aligned}
\label{eq:vleqc_LB}
\end{equation}
Symmetrically, we can also consider the ``flipped'' event
\[
Y_{it}=0\text{ and }W_{it}^{'}\theta_{0}\geq c,
\]
which implies $v_{it}>c$:
\begin{align*}
\left(1-Y_{it}\right)\mathbf{\mathbbm1}\left\{ W_{it}^{'}\theta_{0}\geq c\right\} =\  & \mathbf{\mathbbm1}\left\{ v_{it}>W_{it}^{'}\theta_{0}\right\} \mathbf{\mathbbm1}\left\{ W_{it}^{'}\theta_{0}\geq c\right\} \\
\leq\  & \mathbf{\mathbbm1}\left\{ v_{it}>c\right\} \equiv1-\mathbf{\mathbbm1}\left\{ v_{it}\leq c\right\}
\end{align*}
Rearranging the above, we have
\begin{equation}
\mathbf{\mathbbm1}\left\{ v_{it}\leq c\right\} \leq1-\left(1-Y_{it}\right)\mathbf{\mathbbm1}\left\{ W_{it}^{'}\theta_{0}\geq c\right\} .\label{eq:vleqc_UB}
\end{equation}
Next, taking conditional expectations of \eqref{eq:vleqc_LB} and
\eqref{eq:vleqc_UB} given $Z_{i}=z$, we have
\begin{align}
\mathbb{P}\left(\rest{Y_{it}=1,\ W_{it}^{'}\theta_{0}\leq c}z\right)\leq\  & \mathbb{P}\left(\rest{v_{it}\leq c}z\right)\nonumber \\
=\  & \mathbb{P}\left(\rest{v_{is}\leq c}z\right)\nonumber \\
\leq\  & 1-\mathbb{P}\left(\rest{Y_{is}=0,\ W_{is}^{'}\theta_{0}\geq c}z\right)\label{eq:Bound_ts}
\end{align}
where ``$|z$'' is a shorthand for ``$Z_{i}=z$'' that we will
use throughout the paper. Note that the middle equality of \eqref{eq:Bound_ts}
follows from the partial stationarity condition (Assumption \ref{assu:PartStat}).\footnote{Specifically, observe that Assumption \ref{assu:PartStat} implies
the partial stationarity of $v_{it}$ given $Z_{i}$, since
\begin{align*}
\mathbb{P}\left(\rest{\alpha_{i}+\epsilon_{it}\leq c}z\right) & =\mathbb{E}\left[\rest{\mathbb{P}\left(\rest{\alpha_{i}+\epsilon_{it}\leq c}z,\alpha_{i}\right)}\right]\\
 & =\mathbb{E}\left[\rest{\mathbb{P}\left(\rest{\alpha_{i}+\epsilon_{is}\leq c}z,\alpha_{i}\right)}\right]=\mathbb{P}\left(\rest{\alpha_{i}+\epsilon_{is}\leq c}z\right)
\end{align*}
for any $c$, and hence $v_{it}\mid Z_{i}\stackrel{d}{\sim}v_{is}\mid Z_{i}.$} Essentially, in the above we exploit the joint occurrence of $v_{it}\leq W_{it}^{'}\theta_{0}$
and $W_{it}^{'}\theta_{0}\leq c$ to deduce an implication on the composite
error $v_{it}\leq c$ that is free of the endogenous covariates $X_{it}$,
and then leverage the partial stationarity of $v_{it}$ given $Z_{i}$
for intertemporal comparisons.

Since the lower and upper bounds in \eqref{eq:Bound_ts} hold for
any $t$ and $s$, we summarize the identifying restrictions \eqref{eq:Bound_ts}
across all time periods in the following proposition.
\begin{prop}[Identified Set]
\label{prop:bin} Write
\begin{align}
L_{t}\left(\rest cz,\theta\right) & :=\mathbb{P}\left(\rest{Y_{it}=1,\ W_{it}^{'}\theta\leq c}z\right)\nonumber \\
U_{t}\left(\rest cz,\theta\right) & :=1-\mathbb{P}\left(\rest{Y_{it}=0,\ W_{it}^{'}\theta\geq c}z\right)\label{eq:Lt_Ut}
\end{align}
and
\begin{equation}
\ol L\left(\rest cz;\theta\right):=\max_{t=1,...,T}L_{t}\left(c|z;\theta\right),\quad\ul U\left(\rest cz;\theta\right):=\min_{t=1,...,T}U_{t}\left(c|z;\theta\right),\label{eq:Lbar_Ubar}
\end{equation}
Define $\Theta_{I}$ as the set of $\theta\in\mathcal{R}^{d_{w}}$ such that
\begin{equation}
\ol L\left(\rest cz,\theta\right)\leq\ul U\left(\rest cz,\theta\right),\quad\forall c\in\mathcal{R},\ \forall z\in{\cal Z}:=\text{Supp\ensuremath{\left(Z_{i}\right)}},\label{eq:ID_Set}
\end{equation}
Then, under model \eqref{eq:bin} and Assumption \ref{assu:PartStat},
$\theta_{0}\in\Theta_{I}$.
\end{prop}
\begin{rem}
We note that, once conditioned on $z\equiv\left(z_{1},...,z_{T}\right)$,
the randomness in $W_{it}^{'}\theta=z_{t}^{'}\beta+X_{it}^{'}\gamma$ lies purely
in $X_{it}$ given $z$, and thus it is equivalent to write
\[
L_{t}\left(\rest cz,\theta\right):=\mathbb{P}\left(\rest{Y_{it}=1,\ z_{t}^{'}\beta+X_{it}^{'}\gamma\leq c}z\right)
\]
and similarly for $U_{t}$. We will continue to use the notation $W_{it}^{'}\theta_{0}$
for simplicity, but would like to emphasize this degeneracy of $Z_{it}^{'}\beta$
given $Z_{i}=z.$ In particular, this means that $z_{t}^{'}\beta$ can
be ``absorbed'' into the constant $c$, in a sense that will become
clearer below.
\end{rem}
Proposition \ref{prop:bin} characterizes the identified set $\Theta_{I}$
for $\theta_{0}$ as restrictions on the conditional joint distribution
of $Y_{it}$ and $X_{it}$ given $z$. More specifically, the restrictions
in \eqref{eq:ID_Set} can be regarded as a collection of conditional
moment inequalities that relate $\mathbf{\mathbbm1}\left\{ Y_{it}=1,\ W_{it}^{'}\theta\leq c\right\} $
and $\mathbf{\mathbbm1}\left\{ Y_{it}=0,\ W_{it}^{'}\theta\geq c\right\} $ conditional
on $z$.

Proposition \ref{prop:bin} holds regardless of whether the endogenous
covariates $X_{it}$ are discrete or continuous. When $X_{it}$ are
continuous (taking a continuum of values), then Proposition \ref{prop:bin}
requires that condition \eqref{eq:ID_Set} hold for a continuum of
constants $c\in\mathcal{R}$, so that (the information in) the whole joint
distribution of the binary variable $Y_{it}$ and the continuous variable
$W_{it}^{'}\theta=z_{t}^{'}\beta+X_{it}^{'}\gamma$ can be captured by the collection
of joint distributions of $\left(Y_{it},\mathbf{\mathbbm1}\left\{ W_{it}^{'}\theta\leq c\right\} \right)$
across all possible cutoff values $c$.

However, when $X_{it}$ are discrete, such as in the AR$\left(p\right)$
dynamic model where $X_{it}$ consists of $p$ lagged binary outcome
variables, there is no need to evaluate \eqref{eq:ID_Set} at all
possible values of $c\in\mathcal{R}$, since the inequalities in \eqref{eq:ID_Set}
can only bind at finitely many values of $c$. We formalize this observation
via the following Proposition.
\begin{prop}[Identified Set with Discrete Endogenous Covariates]
 \label{prop:disc} Suppose that the endogenous covariate $X_{it}$
can only take finite number of values in $\left\{ \ol x_{1},...,\ol x_{K}\right\} $
across all time periods $t=1,...,T$. Then $\Theta_{I}=\Theta_{I}^{disc},$
where $\Theta_{I}^{disc}$ consists of all $\theta=\left(\beta^{'},\gamma^{'}\right)^{'}\in\mathcal{R}^{d_{z}}\times\mathcal{R}^{d_{x}}$
that satisfy condition \eqref{eq:ID_Set} for any
\begin{equation}
c\in\left\{ z_{t}^{'}\beta+\ol x_{k}^{'}\gamma:k=1,...,K,t=1,...,T\right\} ,\label{eq:c_KTset}
\end{equation}
and for any $z\in{\cal Z}$.
\end{prop}
Proposition \ref{prop:disc} shows that the discreteness of the endogenous
covariates $X_{it}$ help reduce the infinite number of inequality
restrictions in Proposition \ref{prop:bin} to finitely many, or more
precisely, $KT$ ones (conditional on $z$).

The case of discrete $X_{it}$ is conceptually important, since it
nests the dynamic AR$\left(p\right)$ model widely studied in the
literature as a special case. Clearly, when $X_{it}$ consists of
$p$ (finitely many) lagged binary outcome variables $Y_{i,t-1},...,Y_{i,t-p}$,
then $X_{it}$ by construction can only take $K=2^{p}$ discrete values.
Specialized further to the AR(1) model in KPT, Proposition \ref{prop:disc}
shows that the identified set $\Theta_{I}$ is characterized by $2T$
conditional restrictions, which is drastically smaller than the $9T\left(T-1\right)$
conditional restrictions listed in KPT (even when $T$ is small).
\begin{rem}
\label{rem:PairPS_prop1} Following up on Remark \ref{rem:PairPS},
if pairwise partial stationarity is adopted, then Propositions \ref{prop:bin}
and \ref{prop:disc} continue to hold with \eqref{eq:ID_Set} adapted
to the following ``pairwise'' version:
\begin{equation}
\mathbb{P}\left(\rest{Y_{it}=1,\ W_{it}^{'}\theta\leq c}z_{ts}\right)\leq1-\mathbb{P}\left(\rest{Y_{is}=0,\ W_{is}^{'}\theta\geq c}z_{ts}\right),\label{eq:ID_Set_PS}
\end{equation}
for all $(t,s)$, where ``$|z_{ts}$'' denotes conditioning on the
event $\left(Z_{it},Z_{is}\right)=\left(z_{t},z_{s}\right)=:z_{ts}$.
Relative to \eqref{eq:ID_Set}, the statement in \eqref{eq:ID_Set_PS}
reflects the fact that pairwise partial stationarity is imposed on
all pairs of time periods separately instead of all $T$ time periods
jointly. It is straightforward to verify that the identification arguments
above, in particular \eqref{eq:vleqc_LB}-\eqref{eq:Bound_ts}, carry
over with all conditional probabilities/expectations taken conditional
on $z_{ts}$ instead of $z$.
\end{rem}

\subsection{\label{subsec:Sharp}Sharpness}

So far we have only shown that $\Theta_{I}$ is a valid identified set
for $\theta_{0}$. However, it is not yet clear whether it has incorporated
all the available information for $\theta_{0}$ under the current model
specification. We now proceed to establish the sharpness of $\Theta_{I}$
under appropriate conditions.

We start with the discrete case where the support of $X_{it}$ is
assumed to be finite. Remarkably, the sharpness of our identified
set can be established without any additional assumptions in this
case.
\begin{thm}[Sharpness: Discrete Case]
\label{thm:sharp_disc} Suppose that $X_{it}$ only takes finitely
many values for each $t$. Then, under model \eqref{eq:bin} and
Assumption \ref{assu:PartStat}, the identified set $\Theta_{I}^{disc}$
is sharp.
\end{thm}
The formal definition of sharpness, along with the complete proof
of Theorem \ref{thm:sharp_disc}, are available in Appendix \ref{subsec:pf_sharp_disc}.
In short, we show (by construction) that, for each $\theta\in\Theta_{I}\backslash\left\{ \theta_{0}\right\} $,
there exists a data generating process (DGP) that satisfies Assumption
\ref{assu:PartStat} and produces the same joint distribution of observable
data $\left(Y_{i},W_{i}\right)$ under model \eqref{eq:bin} with
parameter $\theta$. Theorem \ref{thm:sharp_disc} demonstrates that our
key identification strategy based on the bounding of (endogenous)
parametric index by arbitrary constants, as described in Section \ref{subsec:ind},
is able to extract all the available information for $\theta_{0}$ from
the model and the observable data, and thus it is impossible to further
differentiate $\theta_{0}$ from alternatives in the identified set $\Theta_{I}$
under model \eqref{eq:bin} and our assumption of partial stationarity
(without further restrictions).

Theorem \ref{thm:sharp_disc} immediately implies that, in the special
case of dynamic AR$\left(p\right)$ models where $X_{it}$ consists
of discrete lagged outcomes, our characterization of the identified
set $\Theta_{I}^{disc}$ in Proposition \ref{prop:disc} is sharp. In
particular, our result generalizes the corresponding result in KPT,
which focuses on the AR(1) model. Furthermore, KPT characterizes the
sharp identified set via $9T\left(T-1\right)$ conditional restrictions,
the derivation of which is based on an exhaustive enumeration of lagged
outcome realizations $Y_{i,t-1}$. In this paper we adopt an entirely
different (and much more general) identification strategy, and arrive
at a characterization of the identified set by $2T$ conditional restrictions,
which we also show to be sharp by Theorem \ref{thm:sharp_disc}. Since
our model and assumption specialize exactly to that in KPT under the
AR(1) specification, it follows that our $2T$ restrictions must be
able to reproduce all the $9T\left(T-1\right)$ restrictions in KPT.
This demonstrates that our identification strategy not only applies
more generally than the one in KPT, but also leads to a more elegant
characterization of the sharp identified set with much fewer restrictions.
We provide a more detailed explanation about this point in the next
subsection.

Another conceptually remarkable, or surprising, feature of Proposition
\ref{prop:disc} and Theorem \ref{thm:sharp_disc} is that they are
established without reference to the exact nature, or interpretation,
of the endogenous covariates $X_{it}$. The identified set $\Theta_{I}$
we characterized is valid and sharp regardless of whether $X_{it}$
are specified as lagged outcome variables, contemporaneously endogenous
covariates, or a combination of both.

Our proof of sharpness consists of two main steps. First, we show
for each $\theta\in\Theta_{I}\backslash\left\{ \theta_{0}\right\} $ how to construct
the per-period marginal distributions of errors that match the per-period
marginal choice probabilities. Second, we show how to combine the
$T$ per-period marginal distributions into an all-period joint distribution
that matches the all-period joint choice probabilities, so that observational
equivalence holds.

The proof techniques we exploited are also different from, and thus
novel relative to, those used in the related work that leverages stationarity-type
conditions for partial identification, such as \citet*{pakes2019}
for static multinomial choice model and KPT for dynamic AR(1) model.
Instead of showing existence only, we provide a more explicit construction
of the joint distribution of the latent variables, which is valid
regardless of the exact type of endogeneity in $X_{it}$. In particular,
a key challenge in proving sharpness based on stationarity-type conditions
lies in that stationarity imposes only aggregate restrictions (via
integrals/sums) of the joint distribution of errors, which is rather
implicit to work with. A key innovation in our proof technique is
to show how marginal/aggregate stationarity restrictions and joint
choice probability restrictions can be satisfied simultaneously by
an explicit, simple and general construction, which might be of independent
and wider use.

~

Next, we seek to establish the sharpness of our identification set
in the case where certain or all components of $X_{it}$ may be continuous.
Below we present an additional set of regularity conditions for the
continuous case and the corresponding sharpness result, followed by
a discussion of the conditions and the result.
\begin{assumption}[Regularity Conditions for the Continuous Case]
 \label{assu:Cts}Suppose that:
\begin{itemize}
\item[(a)] $\rest{W_{it}^{'}\theta_{0}}z$ is continuously distributed with strictly
positive density on a bounded interval support for each $t$.
\item[(b)] $\mathbb{P}\left(\rest{Y_{it}=1}W_{i}=w\right)\in\left(0,1\right)$ for each
$t$.
\item[(c)] $\ol L\left(\rest cz,\theta_{0}\right)=\ul U\left(\rest cz,\theta_{0}\right)$
only for $c$'s in a set of Lebesgue measure $0$.
\end{itemize}
\end{assumption}
\begin{thm}[Sharpness: Continuous Case]
\label{thm:sharp_cts} Let $\Theta_{I}^{cts}$ be the set of $\theta$ such
that model \eqref{eq:bin}, Assumptions \ref{assu:PartStat} and
\ref{assu:Cts} all hold with $\theta$ in lieu of $\theta_{0}$. Then $\Theta_{I}^{cts}$
is sharp.
\end{thm}
Theorem \ref{thm:sharp_cts} establishes the sharpness of our identification
set under the additional regularity conditions imposed in Assumption
\ref{assu:Cts}. The proof, presented in Appendix \ref{subsec:pf_sharp_cts},
follows the general construction strategy used in the discrete-case
proof, with some key adaptions to handle several continuity and measure-zero
issues arising in the continuous case. Such adaptions utilize the
conditions in Assumption \ref{assu:Cts}, which we now explain in
more details.

Assumption \ref{assu:Cts}(a) can be effectively regarded as a setup
of the continuous-case model. In our current context, given $z$,
the induced index $W_{it}^{'}\theta_{0}=z_{t}^{'}\beta_{0}+X_{it}^{'}\gamma_{0}$
is what enter most directly into our model, rather than $X_{it}$
per se. Part of Assumption \ref{assu:Cts}(a) states that $W_{it}^{'}\theta_{0}$
is continuously distributed on a bounded interval, which can be satisfied
with various lower-level conditions on $X_{it}$. For example, if
$X_{it}|z$ is continuously distributed on a bounded and connected
support with nonempty interior, and if $\gamma_{0}$ is restricted to
lie within a bounded set (which can be imposed as a scale normalization
without loss of generality), then $X_{it}^{'}\gamma_{0}|z$ is continuously
distributed on a bounded connected interval. If in addition $X_{it}|z$
is assumed to have a density that is strictly positive (almost) everywhere
on its support, then the induced density of $X_{it}^{'}\gamma_{0}|z$
will also be (almost) everywhere strictly positive. Note also Assumption
\ref{assu:Cts}(a) may also be satisfied if some (but not all) components
of $X_{it}$ are discrete, as long as some other component(s) of $X_{it}$
is continuously distributed with nonzero coefficient and a sufficiently
large support.

Assumption \ref{assu:Cts}(b), along with the assumptions of connectedness
(interval representation of the support) and strictly positive densities
(strictly increasing CDFs) for $X_{it}^{'}\gamma_{0}|z$ in Assumption
\ref{assu:Cts}(a), are imposed mainly as simplifying restrictions
that are not conceptually necessary but allow for a more convenient
notation. Essentially, they jointly imply that the per-period CCPs
on the left-hand and right-hand sides of \eqref{eq:ID_Set} are continuous
and strictly increasing in $c$ on connected intervals, leading to
simpler notation in the proof via the use of the inverse function
and the intermediate value theorem. Without these conditions, we would
need to handle ``flat regions'', ``jump points'', and ``continuously
increasing regions'' separately and then combine them together to
produce the final result, which  should be achievable using
a  combination of the proof techniques in the discrete case
and the continuous case.

Assumption \ref{assu:Cts}(c) is a key condition for the validity
of our adapted construction in the continuous case, but it is admittedly
the most nonstandard and implicit one, which warrants further explanation.
Effectively, Assumption 2(c) rules out certain ``knife-edge'' degenerate
DGPs that result in a ``flat region of contact'' between $\ol L$
and $\ul U$, though the exact form of such degeneracy can be rather
complicated given the nonlinear nature of the binary choice model
and the generality of the endogeneity we incorporate. Below we provide some intuition for why Assumption \ref{assu:Cts}(c) should be regarded as a relatively mild condition. For more detail, see Appendix \ref{subsec:cond_A2c} for two illustrative examples where Assumption \ref{assu:Cts}(c) (almost) trivially holds, as well as a general sufficient condition for Assumption \ref{assu:Cts}(c).

Note that under Assumption \ref{assu:Cts}(a)(b), we have $L_{t}\left(c|z,\theta_{0}\right)<U_{t}\left(c|z,\theta_{0}\right)$
with strict inequality, so $\ol L\left(\rest cz,\theta_{0}\right)\leq\ul U\left(\rest cz,\theta_{0}\right)$
can only hold with equality if there exist two different periods $t\neq s$
such that $L_{t}\left(c|z,\theta_{0}\right)=U_{s}\left(c|z,\theta_{0}\right)$.
If this holds for all $c$'s in a small open interval, i.e. with positive
Lebesgue measure in violation of Assumption 2(c), then we can deduce
that their derivatives in $c$ must also match, i.e.,
\begin{equation}
L_{t}^{'}\left(\rest cz,\theta_{0}\right)=U_{s}^{'}\left(\rest cz,\theta_{0}\right)\label{eq:L'=00003DU'}
\end{equation}
on an open interval, with $L_{t}^{'}$ and $U_{s}^{'}$ given by
\begin{align}
L_{t}^{'}\left(\rest cz,\theta_{0}\right) & =\mathbb{P}\left(Y_{it}=1|X_{it}^{'}\gamma_{0}=c-z_{t}^{'}\beta_{0},Z_{i}=z\right)\pi_{t}\left(\rest{c-z_{t}^{'}\beta_{0}}z\right),\nonumber \\
U_{s}^{'}\left(\rest cz,\theta_{0}\right) & =\mathbb{P}\left(Y_{is}=0|X_{is}^{'}\gamma_{0}=c-z_{s}^{'}\beta_{0},Z_{i}=z\right)\pi_{s}\left(\rest{c-z_{s}^{'}\beta_{0}}z\right),\label{eq:LU_deriv}
\end{align}
where $\pi_{t}\left(\rest{\cdot}z\right)$ denotes the conditional pdf
of $X_{it}^{'}\gamma_{0}$ given $Z_{i}=z$.

Consequently, $L_{t}^{'}=U_{s}^{'}$ on an open interval essentially
means that the density-weighted CCPs on the right-hand sides of \eqref{eq:LU_deriv}
must change continuously in $c$ in \emph{exactly the same functional
form on a continuum, despite all of the following}: (i) $L_{t}^{'}$
is defined on the event $Y_{it}=1$ while $U_{s}^{'}$ is defined
as on the event $Y_{is}=0$, which are ``flipped'' events that may
generally vary with $c$ in different manners, (ii) the values of
$z_{t}^{'}\beta_{0}$ and $z_{s}^{'}\beta_{0}$ can be different, so the
conditioning events are generally different for $L_{t}$ and $U_{s}$
as well, (iii) the conditional distribution of $X_{it}$ given $Z_{i}=z$
may be different (nonstationary) across periods $t$, so $\pi_{t}\left(\rest{c-z_{t}^{'}\beta_{0}}z\right)$
and $\pi_{s}\left(\rest{c-z_{s}^{'}\beta_{0}}z\right)$ may be different
even if $z_{t}^{'}\beta_{0}=z_{s}^{'}\beta_{0}$, (iv) the dependence structure
between $X_{it}$ and $\epsilon_{it}$ may vary across $t$. For all these
reasons, it appears unlikely that \eqref{eq:L'=00003DU'}
can hold for a continuum of $c$, except under carefully designed ``knife-edge'' DGPs. Hence, we consider Assumption 2(c) as a mild condition. See Appendix \ref{subsec:cond_A2c} for more examples and details.



\subsection{\label{subsec:Recon}Reconciliation with Related Work}

Our identifying restrictions in \eqref{eq:ID_Set} and \eqref{eq:c_KTset}
have a somewhat ``nonstandard'' representation in terms of (conditional)
joint probabilities of $Y_{it}$ and $X_{it}$ (given $Z_{i}$), instead
of conditional probabilities of $Y_{it}$ given $X_{it}$ (such as
lagged outcomes), which are more usually found in the related literature.
Hence, we provide a more detailed discussion about the content and
interpretation of our identifying restrictions, as well as a more
explicit explanation of how they relate to existing results in the
related literature.

\subsubsection*{Reconciliation with \citet{manski1987}}

Consider first the special case where there are \emph{no} endogenous
covariates $X_{it}$, or in other words, $X_{it}$ is degenerate.
In this case, our ``partial stationarity'' condition specializes
to the ``full stationarity'' condition \eqref{eq:cond_sta} as in
\citet{manski1987}. However, our identifying restriction \eqref{eq:ID_Set}
still has a different form than the identifying restriction in \citet{manski1987}.
To illustrate, focus on any two periods $\left(t,s\right)$, and observe
that our identifying restriction becomes:
\begin{equation}
\mathbb{P}\left(\rest{Y_{it}=1,\ z_{t}^{'}\beta_{0}\leq c}z\right)\leq1-\mathbb{P}\left(\rest{Y_{is}=0,\ z_{s}^{'}\beta_{0}\geq c}z\right),\ \forall c,\label{eq:ID_ts}
\end{equation}
while the ``maximum-score-type'' identifying restrictions in \citet{manski1987}
are of the form
\begin{equation}
z_{s}^{'}\beta_{0}\geq z_{t}^{'}\beta_{0}\ \Leftrightarrow\ \mathbb{P}\left(\rest{Y_{is}=1}z\right)\geq\mathbb{P}\left(\rest{Y_{it}=1}z\right).\label{eq:MaxScore}
\end{equation}
The ``maximum-score-type'' identifying restriction \eqref{eq:MaxScore}
has a quite intuitive and interpretable representation: across two
periods $\left(t,s\right)$ under full stationarity, the conditional
choice probability at period $s$ is larger if and only if the index
$z_{s}^{'}\beta_{0}$ is larger. In contrast, our restriction \eqref{eq:ID_ts}
has a somewhat twisted representation even in this simple setting.

However, a closer look reveals that our \eqref{eq:ID_ts} reproduces Manski's ``maximum-score-type'' identifying restrictions
in the current context. To see this, notice that, by setting $c=z_{t}^{'}\beta_{0}$
in \eqref{eq:ID_ts}, we obtain
\begin{align*}
\mathbb{P}\left(\rest{Y_{it}=1}z\right)=\, & \mathbb{P}\left(\rest{Y_{it}=1}z\right)\mathbf{\mathbbm1}\left\{ z_{t}^{'}\beta_{0}\leq z_{t}^{'}\beta_{0}\right\} \leq1-\mathbb{P}\left(\rest{Y_{is}=0}z\right)\mathbf{\mathbbm1}\left\{ z_{s}^{'}\beta_{0}\geq z_{t}^{'}\beta_{0}\right\}
\end{align*}
Hence, if $z_{s}^{'}\beta_{0}\geq z_{t}^{'}\beta_{0}$, i.e., the left-hand
side of \eqref{eq:MaxScore} holds, then the above implies that
\[
\mathbb{P}\left(\rest{Y_{it}=1}z\right)\leq1-\mathbb{P}\left(\rest{Y_{is}=0}z\right)=\mathbb{P}\left(\rest{Y_{is}=1}z\right),
\]
which becomes exactly the right-hand side of \eqref{eq:MaxScore}.
Switching $t$ with $s$ in the argument above produces the other
implication $z_{s}^{'}\beta_{0}\leq z_{t}^{'}\beta_{0}$ $\Rightarrow$ $\mathbb{P}\left(\rest{Y_{is}=1}z\right)\leq\mathbb{P}\left(\rest{Y_{it}=1}z\right)$.
Together these exactly constitute the ``if-and-only-if'' restriction
in \eqref{eq:MaxScore}. Hence, even though our inequality restriction
\eqref{eq:ID_ts} looks different from the more intuitive ``maximum-score-type''
restriction, they both incorporate the same information.

\subsubsection*{Reconciliation with KPT}

Now, consider the dynamic AR(1) model as studied in KPT, where the
only endogenous covariate is the one-period lagged outcome variable,
i.e., $X_{it}:=Y_{i,t-1}$.

To illustrate, first focus on any two periods $\left(t,s\right)$
only, and observe that in this case our identifying restriction becomes
\begin{equation}
\mathbb{P}\left(\rest{Y_{it}=1,\ z_{t}^{'}\beta_{0}+Y_{i,t-1}\gamma_{0}\leq c}z\right)\leq1-\mathbb{P}\left(\rest{Y_{is}=0,\ z_{s}^{'}\beta_{0}+Y_{i,s-1}\gamma_{0}\geq c}z\right),\ \forall c.\label{eq:ID_ts_KPT}
\end{equation}
Under the same model and assumption, KPT derives the following 9 inequality
implications for $\left(t,s\right)$:\footnote{We adapt the notation in KPT to our current notation, and state these
9 inequalities as strict inequalities, which lead to a simpler and
more focused explanation. The equivalence between our restriction
and the KPT restrictions still hold if their inequalities are stated
in the weak form.}

KPT(i): $\mathbb{P}\left(\rest{Y_{it}=1}z\right)>\mathbb{P}\left(\rest{Y_{is}=1}z\right)$
$\Rightarrow$ $\left(z_{t}-z_{s}\right)^{'}\beta_{0}+\left|\gamma_{0}\right|>0.$

KPT(ii): $\mathbb{P}\left(\rest{Y_{it}=1}z\right)>1-\mathbb{P}\left(\rest{Y_{i,s}=0,Y_{i,s-1}=1}z\right)$
$\Rightarrow$ $\left(z_{t}-z_{s}\right)^{'}\beta_{0}-\min\left\{ 0,\gamma_{0}\right\} >0.$

KPT(iii): $\mathbb{P}\left(\rest{Y_{it}=1}z\right)>1-\mathbb{P}\left(\rest{Y_{i,s}=0,Y_{i,s-1}=0}z\right)$
$\Rightarrow$ $\left(z_{t}-z_{s}\right)^{'}\beta_{0}+\max\left\{ 0,\gamma_{0}\right\} >0.$

KPT(iv): $\mathbb{P}\left(\rest{Y_{it}=1,Y_{it-1}=1}z\right)>\mathbb{P}\left(\rest{Y_{is}=1}z\right)$
$\Rightarrow$ $\left(z_{t}-z_{s}\right)^{'}\beta_{0}+\max\left\{ 0,\gamma_{0}\right\} >0.$

KPT(v): $\mathbb{P}\left(\rest{Y_{it}=1,Y_{it-1}=1}z\right)>1-\mathbb{P}\left(\rest{Y_{is}=0,Y_{i,s-1}=1}z\right)$
$\Rightarrow$ $\left(z_{t}-z_{s}\right)^{'}\beta_{0}>0.$

KPT(vi): $\mathbb{P}\left(\rest{Y_{it}=1,Y_{it-1}=1}z\right)>1-\mathbb{P}\left(\rest{Y_{is}=0,Y_{i,s-1}=0}z\right)$
$\Rightarrow$ $\left(z_{t}-z_{s}\right)^{'}\beta_{0}+\gamma_{0}>0.$

KPT(vii): $\mathbb{P}\left(\rest{Y_{it}=1,Y_{it-1}=0}z\right)>1-\mathbb{P}\left(\rest{Y_{is}=0}z\right)$
$\Rightarrow$ $\left(z_{t}-z_{s}\right)^{'}\beta_{0}-\min\left\{ 0,\gamma_{0}\right\} >0.$

KPT(viii): $\mathbb{P}\left(\rest{Y_{it}=1,Y_{it-1}=0}z\right)>1-\mathbb{P}\left(\rest{Y_{is}=0\text{\ensuremath{,}}Y_{i,s-1}=1}z\right)$
$\Rightarrow$ $\left(z_{t}-z_{s}\right)^{'}\beta_{0}-\gamma_{0}>0.$

KPT(ix): $\mathbb{P}\left(\rest{Y_{it}=1,Y_{it-1}=0}z\right)>1-\mathbb{P}\left(\rest{Y_{is}=0,Y_{i,s-1}=0}z\right)$
$\Rightarrow$ $\left(z_{t}-z_{s}\right)^{'}\beta_{0}>0.$

~

\noindent In a way, the 9 inequality restrictions in KPT above are
similar to the ``maximum-score restrictions'', in the sense that
all of them take the form of logical implications between intertemporal
comparisons of various conditional probabilities and intertemporal
differences of covariate indexes.

Using a very different identification strategy than the one in KPT,
we arrived at our inequality restriction \eqref{eq:ID_ts_KPT}, which
looks very different from the collection of 9 inequality restrictions
in KPT. At first sight it is not clear how \eqref{eq:ID_ts_KPT} relates
to and compares with the 9 KPT restrictions. However, a closer look
again reveals that our restriction \eqref{eq:ID_ts_KPT} can reproduce
all the 9 restrictions in KPT, and thus incorporate all the information
therein in a unified format.

Take KPT(v) as an illustration and suppose that the left-hand side
of KPT(v) holds, then it implies
\begin{equation}
\mathbb{P}\left(\rest{Y_{it}=1,Y_{it-1}=1}z\right)>1-\mathbb{P}\left(\rest{Y_{is}=0,Y_{i,s-1}=1}z\right).\label{eq:KPTv_LHS}
\end{equation}

With $X_{it}=Y_{i,t-1}$, our inequality restriction \eqref{eq:ID_ts_KPT}
can be equivalently rewritten as follows,
\begin{align}
 & \mathbb{P}\left(\rest{Y_{it}=1,\ Y_{i,t-1}=1}z\right)\mathbf{\mathbbm1}\left\{ z_{t}^{'}\beta_{0}+\gamma_{0}\leq c\right\} +\mathbb{P}\left(\rest{Y_{it}=1,\ Y_{i,t-1}=0}z\right)\mathbf{\mathbbm1}\left\{ z_{t}^{'}\beta_{0}\leq c\right\} \nonumber \\
\leq\  & 1-\mathbb{P}\left(\rest{Y_{is}=0,\ Y_{i,s-1}=1}z\right)\mathbf{\mathbbm1}\left\{ z_{s}^{'}\beta_{0}+\gamma_{0}\geq c\right\} -\mathbb{P}\left(\rest{Y_{is}=0,\ Y_{i,s-1}=0}z\right)\mathbf{\mathbbm1}\left\{ z_{s}^{'}\beta_{0}\geq c\right\} ,\label{eq:ID_KPT_expand}
\end{align}
where the realization of $Y_{i,t-1}$ is explicitly enumerated as
in KPT.

Note that we can further relax condition \eqref{eq:ID_KPT_expand}
by dropping the two probabilities $\mathbb{P}\left(\rest{Y_{it}=1,\ Y_{i,t-1}=0}z\right)\mathbf{\mathbbm1}\left\{ z_{t}^{'}\beta_{0}\leq c\right\} $
and $\mathbb{P}\left(\rest{Y_{is}=0,\ Y_{i,s-1}=0}z\right)\mathbf{\mathbbm1}\left\{ z_{s}^{'}\beta_{0}\geq c\right\} $
as it makes the lower bound smaller and the upper bound larger:
\begin{align*}
 & \mathbb{P}\left(\rest{Y_{it}=1,\ Y_{i,t-1}=1}z\right)\mathbf{\mathbbm1}\left\{ z_{t}^{'}\beta_{0}+\gamma_{0}\leq c\right\} \\
\leq\  & 1-\mathbb{P}\left(\rest{Y_{is}=0,\ Y_{i,s-1}=1}z\right)\mathbf{\mathbbm1}\left\{ z_{s}^{'}\beta_{0}+\gamma_{0}\geq c\right\} .
\end{align*}
Then, the statement that $\mathbf{\mathbbm1}\left\{ z_{t}^{'}\beta_{0}+\gamma_{0}\leq c\right\} $
and $\mathbf{\mathbbm1}\left\{ z_{s}^{'}\beta_{0}+\gamma_{0}\geq c\right\} $ both hold
is precisely equivalent to the following statement of
\[
z_{t}^{'}\beta_{0}\leq z_{s}^{'}\beta_{0}\ \Rightarrow\ \mathbb{P}\left(\rest{Y_{it}=1,\ Y_{i,t-1}=1}z\right)\leq\ 1-\mathbb{P}\left(\rest{Y_{is}=0,\ Y_{i,s-1}=1}z\right).
\]
By contraposition, it leads to exactly the same implication of KPT(v):
\[
\mathbb{P}\left(\rest{Y_{it}=1,\ Y_{i,t-1}=1}z\right)>\ 1-\mathbb{P}\left(\rest{Y_{is}=0,\ Y_{i,s-1}=1}z\right)\Longrightarrow z_{t}^{'}\beta_{0}>z_{s}^{'}\beta_{0}.
\]
Hence, we have shown that \eqref{eq:ID_KPT_expand} implies KPT(v).

Similarly, it is shown in Appendix \ref{appe:RelateKPT} that \eqref{eq:ID_KPT_expand}
implies all 9 restrictions in KPT. In fact, the representation \eqref{eq:ID_KPT_expand}
reveals why there are precisely 9 KPT-type restrictions. The two period-$t$
indicators $\mathbf{\mathbbm1}\left\{ z_{t}^{'}\beta_{0}+\gamma_{0}\leq c\right\} $ and
$\mathbf{\mathbbm1}\left\{ z_{t}^{'}\beta_{0}\leq c\right\} $ in the upper expression
of \eqref{eq:ID_KPT_expand} may take 3 ``useful''\footnote{The 4th combination, $\mathbf{\mathbbm1}\left\{ z_{t}^{'}\beta_{0}+\gamma_{0}\leq c\right\} =\mathbf{\mathbbm1}\left\{ z_{t}^{'}\beta_{0}\leq c\right\} =0$,
will make the upper expression of \eqref{eq:ID_KPT_expand} equal
to $0$, so that the inequality \eqref{eq:ID_KPT_expand} holds trivially.
Hence, this $\left(0,0\right)$ combination is not useful.} combinations $\left(1,0\right),\left(0,1\right)$ and $\left(1,1\right)$,
while the two period-$s$ indicators $\mathbf{\mathbbm1}\left\{ z_{s}^{'}\beta_{0}+\gamma_{0}\geq c\right\} $
and $\mathbf{\mathbbm1}\left\{ z_{s}^{'}\beta_{0}\geq c\right\} $ in the lower expression
of \eqref{eq:ID_KPT_expand} may also take 3 useful combinations.
Consequently, in total there are $3\times3=9$ useful combinations,
which exactly correspond to the $9$ left-hand-side suppositions in
the 9 KPT restrictions.

Hence, while our restriction \eqref{eq:ID_ts_KPT} appears very different
from the 9 KPT restrictions, it actually automatically incorporates
all the KPT restrictions. In particular, by treating the endogenous
covariate $X_{it}=Y_{i,t-1}$ as a random variable, our restriction
\eqref{eq:ID_ts_KPT} automatically aggregates the identifying information
across all possible realizations of $Y_{i,t-1}$, without the need
to explicitly consider each possibility separately.

Now, consider a general setting with $T\geq2$ periods. By our Proposition
\ref{prop:disc} and Theorem \ref{thm:sharp_disc}, the sharp identified
set can be characterized by $2T$ restrictions, which are generated
by evaluating \eqref{eq:ID_Set} at each $c$ of the $2T$ points
in $\left\{ z_{t}^{'}\beta,z_{t}^{'}\beta+\gamma:t=1,...,T\right\} $. In contrast,
across $T$ periods the KPT approach produces $9T\left(T-1\right)$
restrictions, which are generated by imposing the $9$ KPT restrictions
across all ordered time pairs $\left(t,s\right)$. Hence, our approach
provides a much simpler characterization of the sharp identified set,
using a significantly smaller number of restrictions. For example,
with $T=2$ periods, we have $4$ restrictions while KPT has $18$;
with $T=3$, we have $6$ restrictions while KPT has 54. Hence, the
reduction in the number of restrictions relative to KPT is quite remarkable.

~

In summary, while the representation of our identifying restrictions
in Propositions \ref{prop:bin} and \ref{prop:disc} may appear somewhat
unusual in the first place, it actually becomes equivalent to the more
familiar representations in the specialized settings of \citet{manski1987}
and KPT.

\section{\label{sec:gene}Extensions}

The key idea underlying our identification strategy extends beyond the binary choice model and can be applied to a wide range of nonlinear panel data models with dynamics and endogeneity. We first present our general identification strategy in a
generic semiparametric model (Section 3.1), and then demonstrate how this strategy
can be applied and adapted to ordered response (Section 3.2), multinomial
choice (Section 3.3) and censored outcome (Appendix B.3) settings.

\subsection{Identification Strategy in a Generic Model\label{subsec:generic}}

We start with a generic semiparametric model that helps convey the
generality of our key identification strategy
\begin{align}
Y_{it} & =G\left(W_{it}^{'}\theta_{0},\,\alpha_{i},\,\epsilon_{it}\right),\label{eq:model_gen}
\end{align}
where $Y_{it}\in\mathcal{Y}$ can be either a discrete or continuous
variable, $\alpha_{i}$ is the individual fixed effect of arbitrary dimension,
$\epsilon_{it}$ is the time-varying error of arbitrary dimension, $W_{it}$
is a vector of observable covariates, $\theta_{0}\in\mathcal{R}^{d_{w}}$
is a conformable vector of parameters, and the function $G$ is allowed
to be unknown, nonseparable but assumed to satisfy the following:
\begin{assumption}[Index Monotonicity]
 \label{assu:Mono} The mapping $\delta\longmapsto G\left(\delta,\alpha,\epsilon\right)$
is weakly increasing in $\delta\in\mathcal{R}$ for each realization of
$\left(\alpha,\epsilon\right)$.
\end{assumption}
Note that, we can obtain the binary choice model in Section \ref{sec:bin}
by setting $\alpha_{i},\epsilon_{it}$ to be scalar-valued, and $G\left(W_{it}^{'}\theta_{0},\,\alpha_{i},\,\epsilon_{it}\right)=\mathbf{\mathbbm1}\left\{ W_{it}^{'}\theta_{0}+\alpha_{i}+\epsilon_{it}\geq0\right\} $,
where $G$ is by construction weakly increasing in $W_{it}^{'}\theta_{0}$.

As before, we decompose $W_{it}$, and correspondingly
$\theta_{0}$, into two components, $W_{it}=\left(Z_{it}^{'},X_{it}^{'}\right)^{'}$
and $\theta_{0}=\left(\beta_{0}^{'},\gamma_{0}^{'}\right)^{'}$, and impose the
partial stationarity condition (Assumption \ref{assu:PartStat}).
We now show how partial stationarity can be exploited in conjunction
with weak monotonicity (Assumption \ref{assu:Mono}) to obtain identifying
restrictions in the presence of endogeneity.

Let $\mathcal{Y}$ denote the support of $Y_{it}$. For any $c\in\mathcal{R}$
and $y\in\mathcal{Y}$, observe that
\[
\begin{aligned}\mathbf{\mathbbm1}\left\{ Y_{it}\leq y,\ W_{it}^{'}\theta_{0}\geq c\right\}  & =\mathbf{\mathbbm1}\left\{ G\left(W_{it}^{'}\theta_{0},\,\alpha_{i},\,\epsilon_{it}\right)\leq y,\ W_{it}^{'}\theta_{0}\geq c\right\} \\
 & \leq\mathbf{\mathbbm1}\left\{ G\left(c,\,\alpha_{i},\,\epsilon_{it}\right)\leq y\right\} ,
\end{aligned}
\]
where the inequality holds by the monotonicity of the function $G$.
Symmetrically, we have
\[
\begin{aligned}\mathbf{\mathbbm1}\left\{ Y_{it}>y,\ W_{it}^{'}\theta_{0}\leq c\right\}  & =\mathbf{\mathbbm1}\left\{ G\left(W_{it}^{'}\theta_{0},\,\alpha_{i},\,\epsilon_{it}\right)>y,\ W_{it}^{'}\theta_{0}\leq c\right\} \\
 & \leq\mathbf{\mathbbm1}\left\{ G\left(c,\,\alpha_{i},\,\epsilon_{it}\right)>y\right\} \\
 & =1-\mathbf{\mathbbm1}\left\{ G\left(c,\,\alpha_{i},\,\epsilon_{it}\right)\leq y\right\} .
\end{aligned}
\]
which is equivalent to
\[
\mathbf{\mathbbm1}\left\{ G\left(c,\,\alpha_{i},\,\epsilon_{it}\right)\leq y\right\} \leq1-\mathbf{\mathbbm1}\left\{ Y_{it}>y,\ Z_{it}^{'}\beta_{0}+X_{it}^{'}\gamma_{0}\leq c\right\} .
\]
The partial stationarity assumption $\epsilon_{it}\mid Z_{i},\alpha_{i}\sim\epsilon_{is}\mid Z_{i},\alpha_{i}$
implies the stationarity of the transform function $G$: $G\left(c,\alpha_{i},\epsilon_{it}\right)\mid Z_{i},\alpha_{i}\sim G\left(c,\alpha_{i},\epsilon_{is}\right)\mid Z_{i},\alpha_{i}$.
After integrating out $\alpha_{i}$, the stationarity condition persists
conditioned on $Z_{i}$ alone:
\[
G\left(c,\alpha_{i},\epsilon_{it}\right)\mid Z_{i}\sim G\left(c,\alpha_{i},\epsilon_{is}\right)\mid Z_{i}.
\]
Combining the above derived bounds on $\mathbf{\mathbbm1}\{G\left(c,\,\alpha_{i},\,\epsilon_{it}\right)\leq y\}$,
we have

\begin{equation}
\begin{aligned} & \mathbb{P}\left(Y_{it}\leq y,\ Z_{it}^{'}\beta_{0}+X_{it}^{'}\gamma_{0}\geq c\mid z\right)\\
\leq \  & \mathbb{P}\left(G\left(c,\,\alpha_{i},\,\epsilon_{it}\right)\leq y\mid z\right)=\mathbb{P}\left(G\left(c,\,\alpha_{i},\,\epsilon_{is}\right)\leq y\mid z\right)\\
\leq\  & 1-\mathbb{P}\left(Y_{is}>y,\ Z_{is}^{'}\beta_{0}+X_{is}^{'}\gamma_{0}\leq c\mid z\right)=:U_{s}\left(\rest{c,y}z,\theta_{0}\right)
\end{aligned}
\label{eq:Bounds_z}
\end{equation}
The key difference of the above and the corresponding identifying
restrictions in Section 2 lies in that the ``middle term'' in \eqref{eq:Bounds_z}
is no longer the conditional CDF of $\alpha_{i}+\epsilon_{it}$, but the conditional
probability of $G\left(c,\,\alpha_{i},\,\epsilon_{is}\right)\leq y$, with the
latter representation not dependent on scalar-additivity of fixed
effect $\alpha_{i}$ and time-varying errors $\epsilon_{it}$.

We summarize the identifying restrictions derived above by the following
proposition:
\begin{prop}
\label{prop:gene} Define $\Theta_{I,gen}$ as the set of all $\theta\in\mathcal{R}^{d_{w}}$
such that
\begin{align}
\max_{t}\mathbb{P}\left(Y_{it}\leq y,\ Z_{it}^{'}\beta+X_{it}^{'}\gamma\geq c\mid z\right)\leq1-\max_{s}\mathbb{P}\left(Y_{is}>y,\ Z_{is}^{'}\beta+X_{is}^{'}\gamma\leq c\mid z\right),\label{eq:gen_id_ineq}
\end{align}
where for any $c\in\mathcal{R},$ $y\in\mathcal{Y}$, and any $z$.
Under model \eqref{eq:model_gen}, Assumptions \ref{assu:PartStat}
and \ref{assu:Mono}, $\theta_{0}\in\Theta_{I,gen}$.
\end{prop}
Note that in the binary choice setting of Section \ref{sec:bin},
it suffices to set $y=0$ in \eqref{eq:gen_id_ineq}, which then coincides
with the identifying results in Proposition \ref{prop:bin}. This also shows that the identified set does not change  at all, regardless of whether scalar-additivity of $\alpha_i$ and $\epsilon_{it}$ is imposed or not in the binary choice model.

The results in Proposition \ref{prop:gene} generally hold regardless
of whether the dependent variable and the endogenous covariate are
discrete or continuous. The next proposition shows that additional
discreteness in either the dependent variable or endogenous covariates
can further simplify and reduce the number of the identifying conditions
in \eqref{eq:gen_id_ineq}.
\begin{prop}
\label{prop:gene_dis} When $X_{it}\in\left\{ \ol x_{1},...,\ol x_{K}\right\} $
for any $t$, then $\Theta_{I,gen}=\Theta_{I,gen}^{disc_{x}},$ where $\Theta_{I,gen}^{disc_{x}}$
consists of all $\theta=\left(\beta^{'},\gamma^{'}\right)^{'}$ that satisfy
condition \eqref{eq:gen_id_ineq} for any $c\in\left\{ z_{t}^{'}\beta+\ol x_{k}^{'}\gamma:k=1,...,K,t=1,...,T\right\} $.
\end{prop}
Moreover, when $Y_{it}\in\left\{ \ol y_{1},...,\ol y_{K}\right\} $
with $\ol y_{j}\leq\ol y_{j+1}$ for any $t$, then $\Theta_{I,gen}=\Theta_{I,gen}^{disc_{y}},$
where $\Theta_{I,gen}^{disc_{y}}$ consists of all $\theta=\left(\beta^{'},\gamma^{'}\right)^{'}$
that satisfy condition \eqref{eq:gen_id_ineq} for any $y\in\left\{ \ol y_{1},...,\ol y_{K-1}\right\} $.

Proposition \ref{prop:gene_dis} shows that for the general model,
when both the outcome and the endogenous variable are discrete, it
is sufficient to focus on a finite number of identifying restrictions.
The number of these restrictions is determined by the support of the
outcome variable and the covariate index. The proof of Proposition \ref{prop:gene_dis} follows the same reasoning
as Proposition \ref{prop:disc}, so it is omitted here. The central
idea is that for any point $c$ or $y$ outside the range specified
in Proposition \ref{prop:gene_dis}, we can find a point within the
specified range that provides weakly more informative results. Therefore,
the inclusion of these outside points would not provide additional
information for the identified set.

\begin{rem} It is  natural to ask whether sharpness can be established in this general setup. While we do not present a formal result, we provide a discussion of this  in Appendix \ref{subsec:sharp_gen}.
\end{rem}


\subsection{Ordered Response Model\label{subsec:order}}

Consider that the outcome variable $Y_{it}$ takes $J$ ordered values:
$Y_{it}\in\left\{ y_{1},..,y_{J}\right\} $ with $y_{j}<y_{j+1}$.
Examples of such ordered outcomes include various income categories,
health outcomes, or levels of educational attainment. We study the
following panel ordered choice model:
\begin{equation}
\begin{aligned}Y_{it}^{*} & =W_{it}'\theta_{0}+v_{it},\\
Y_{it} & =\sum_{j=1}^{J}y_{j}\mathbf{\mathbbm1}\left\{ b_{j}<Y_{it}^{*}\leq b_{j+1}\right\} ,
\end{aligned}
\label{model:order}
\end{equation}
where $Y_{it}^{*}$ denotes the latent dependent variable, and $Y_{it}$
denotes the ordered outcome which takes value $y_{j}$ when $Y_{it}^{*}\in(b_{j},b_{j+1}]$.
The threshold parameters satisfy $b_{1}=-\infty,b_{J+1}=+\infty$,
and the remaining threshold parameters $b_{j}$ (where $b_{j}\leq b_{j+1}$)
can be either known or unknown for $2\leq j\leq J-1$. The binary
choice model in \eqref{eq:bin} is nested with $J=2$ and $b_{2}=0$.

While the ordered response model \eqref{model:order} here can be
regarded as a special case of the generic model \eqref{eq:model_gen},
the special ``ordered cutoffs'' structure in \eqref{model:order}
contains more information than an unknown generic $G$ function in
\eqref{eq:model_gen}. As a result, even though the general identification
strategy in Section \ref{subsec:generic} still applies, we can adapt
the identification argument to the special additional structure imposed
here, obtaining a sharper result than a direct application of Proposition
\ref{prop:gene}. In particular, we explain why the line of our identification
arguments help us find such an adaption that exploits the special
model structure.

We now explain this in more details. Following the arguments in Section
\ref{subsec:generic}, we have
\[
\mathbf{\mathbbm1}\left\{ Y_{it}\leq y_{j},\ b_{j+1}-W_{it}^{'}\theta_{0}\leq c\right\} \leq\mathbf{\mathbbm1}\left\{ v_{it}\leq c\right\}
\]
For a given $c$, the above inequality holds for any response index
$j$. This immediately implies that we can take the largest one to
get a tighter lower bound:
\[
\max_{j}\mathbf{\mathbbm1}\left\{ Y_{it}\leq y_{j},\ b_{j+1}-W_{it}^{'}\theta_{0}\leq c\right\} \leq\mathbf{\mathbbm1}\left\{ v_{it}\leq c\right\} .
\]
In addition, an inspection of the LHS reveals that the maximum is
attained at
\[
j=\ol j_{c}\left(W_{it}\right):=\max\left\{ j:b_{j+1}-W_{it}^{'}\theta_{0}\leq c\right\}
\]
since such a (random) $j$ would maximize $\mathbf{\mathbbm1}\left\{ Y_{it}\leq y_{j}\right\} $
subject to $b_{j+1}-W_{it}^{'}\theta_{0}\leq c$. Consequently, we obtain
\begin{align}
\mathbf{\mathbbm1}\left\{ v_{it}\leq c\right\}  & \geq\mathbf{\mathbbm1}\left\{ Y_{it}\leq y_{\ol j_{c}\left(W_{it}\right)},\ b_{\ol j_{c}\left(W_{it}\right)+1}-W_{it}^{'}\theta_{0}\leq c\right\} \nonumber \\
 & =\sum_{j=1}^{\ol j_{c}\left(W_{it}\right)}\mathbf{\mathbbm1}\left\{ Y_{it}=y_{j},\ b_{\ol j_{c}\left(W_{it}\right)+1}-W_{it}^{'}\theta_{0}\leq c\right\} \nonumber \\
 & =\sum_{j=1}^{J}\mathbf{\mathbbm1}\left\{ Y_{it}=y_{j},\ b_{j+1}-W_{it}^{'}\theta_{0}\leq c\right\} \label{eq:order_LB}
\end{align}
where the last equality holds since $\mathbf{\mathbbm1}\left\{ b_{j+1}-W_{it}^{'}\theta_{0}\leq c\right\} =0$
for any choice $j>\ol j_{c}\left(W_{it}\right)$.

The final expression \eqref{eq:order_LB} is particularly nice for
three reasons: First, it aggregates the information aggregated from
different $y_{j}$ together to produce a tighter lower bound. Second,
the expression circumvent the need to compute the maximizer cutoff
$\ol j_{c}$. Third, it is represented as a linear sum (instead of
a maximum) so that conditional expectation of \eqref{eq:order_LB}
remains a linear sum of conditional expectations.

To see the advantage of the third point above, we take conditional
expectation of \eqref{eq:order_LB} given $z$ as before, obtaining
\[
\mathbb{P}\left(\rest{v_{it}\leq c}z\right)\geq\sum_{j=1}^{J}\mathbb{P}\left(\rest{Y_{it}=y_{j},\ b_{j+1}-W_{it}^{'}\theta_{0}\leq c}z\right).
\]
where the RHS can be computed as a simple sum of CCPs about each ordered
value $y_{j}$.

Similarly, we can derive an upper bound
\[
\mathbb{P}\left(v_{is}\leq c\mid z\right)\leq1-\sum_{j=1}^{J}\mathbb{P}\left(\rest{Y_{is}=y_{j},b_{j}-W_{is}^{'}\theta_{0}\geq c}z\right),
\]
which can be combined with the lower bound to yield the following
result.
\begin{prop}
\label{prop:order} Define $\Theta_{I,order}$ as the set of $\theta=\left(\beta^{'},\gamma^{'}\right)^{'}$
such that
\begin{align}
 & \max_{t=1,...,T}\sum_{j=1}^{J}\mathbb{P}\left(Y_{it}=y_{j},b_{j+1}-z_{t}^{'}\beta-X_{it}^{'}\gamma\leq c\mid z\right)\nonumber \\
\leq\  & 1-\max_{s=1,...,T}\sum_{j=1}^{J}\mathbb{P}\left(Y_{is}=y_{j},b_{j}-z_{s}^{'}\beta-X_{is}^{'}\gamma\geq c\mid z\right),\label{eq:order}
\end{align}
for any $c\in\mathcal{R}$ and any realization $z$ in the support
of $Z_{i}$. Under Assumptions \ref{assu:PartStat}, $\theta_{0}\in\Theta_{I,order}$.
\end{prop}
We emphasize again that Proposition \ref{prop:order} is \emph{not
}a direct application of Proposition \ref{prop:gene}, since Proposition
\ref{prop:order} explicitly utilizes the special model structure
of the order response model to aggregate information from all response
index $j$ together to form tighter bounds for each $c$. In contrast,
a naive application of Proposition \ref{prop:gene} would yield
\begin{multline*}
\max_{t=1,...,T}\mathbb{P}\left(Y_{it}\leq y_{j},b_{j+1}-z_{t}^{'}\beta-X_{it}^{'}\gamma\leq c\mid z\right)\\
\leq1-\max_{s=1,...,T}\mathbb{P}\left(Y_{is}>y_{j},b_{j}-z_{s}^{'}\beta-X_{is}^{'}\gamma\geq c\mid z\right),\forall j,\ \forall\left(c,z\right)
\end{multline*}
which remains valid but is a collection of bounds imposed on each
$j$ separately, thus is generally not as tight as the bounds in \eqref{eq:order}.

\subsection{Multinomial Choice Model \label{subsec:mul}}

In this subsection, we apply our key identification strategy to panel
multinomial choice model with endogeneity. Specifically, consider
a set of unordered choice alternatives $\mathcal{J}=\left\{ 0,1,...,J\right\} $.
Let $u_{ijt}$ denote the latent utility for individual $i$ of selecting
choice $j$ at time $t$, which depends on the three components: observed
covariate $W_{ijt}=(Z_{ijt}',X_{ijt}')'$, unobserved fixed effects
$\alpha_{ij}$, and unobserved time-varying preference shock $\epsilon_{ijt}$.
Let $Y_{it}\in\mathcal{J}$ denote individual $i$'s choice at time
$t$. We study the following panel multinomial choice model:
\[
\begin{aligned}u_{ijt} & =W_{ijt}^{'}\theta_{0}+\alpha_{ij}+\epsilon_{ijt},\\
Y_{it} & =\arg\max_{j\in\mathcal{J}}u_{ijt},
\end{aligned}
\]
and impose the same partial stationarity assumption:
\[
\epsilon_{is}\mid Z_{i},\alpha_{i}\stackrel{d}{\sim}\epsilon_{it}\mid Z_{i},\alpha_{i}\quad\text{for any}\ s,t\leq T.
\]
with $Z_{it}:=\{Z_{ijt}\}_{j\in\mathcal{J}}$,$\alpha_{i}:=\{\alpha_{ij}\}_{j\in\mathcal{J}}$
and $\epsilon_{it}:=\{\epsilon_{ijt}\}_{j\in\mathcal{J}}$ defined to collect
terms across all $J$ choice alternatives.

We emphasize that this model is not a special case of the generic
model \eqref{subsec:generic} in Subsection \eqref{subsec:generic},
since in the current model the $J$ outcome values are unordered,
and the model involves multiple indexes and multivariate monotonicity.
Hence, we cannot directly apply Proposition \ref{prop:gene} to the
current setting. That said, we explain how the key idea from Subsection
\eqref{subsec:generic} can again be adapted to obtain identification
result in the panel multinomial choice setting.

We start by looking at the indicator variable $Y_{it}^{j}:=\mathbf{\mathbbm1}\{Y_{it}=j\}$
of choosing alternative $j$, which maintains a similar monotone structure
with Assumption \ref{assu:Mono}:
\begin{align*}
Y_{it}^{j}=1\  & \Leftrightarrow\ W_{ijt}'\theta_{0}+\alpha_{ij}+\epsilon_{ijt}\geq W_{ikt}'\theta_{0}+\alpha_{ik}+\epsilon_{ikt},\ \forall k\in\mathcal{J}\\
 & \Leftrightarrow\ W_{ijt}'\theta_{0}-W_{ikt}'\theta_{0}\geq\alpha_{ik}+\epsilon_{ikt}-\alpha_{ij}-\epsilon_{ijt},\ \forall k\in\mathcal{J}
\end{align*}
and the new variable $Y_{it}^{j}$ is increasing in $W_{ijt}'\theta_{0}-W_{ikt}'\theta_{0}$
$\forall k\in\mathcal{J}$

More generally, for any subset $K\subset\mathcal{J}$, the indicator
variable $Y_{it}^{K}:=\mathbf{\mathbbm1}\{Y_{it}\in K\}$ represents individual
$i$'s choice belonging to the subset $K$, given by
\begin{align*}
Y_{it}^{K}=1\  & \Leftrightarrow\ W_{ijt}'\theta_{0}+\alpha_{ij}+\epsilon_{ijt}\geq W_{ikt}'\theta_{0}+\alpha_{ik}+\epsilon_{ikt},\ \exists j\in K,\forall k\in\mathcal{J}\setminus K,\\
 & \Leftrightarrow\ W_{ijt}'\theta_{0}-W_{ikt}'\theta_{0}\geq\alpha_{ik}+\epsilon_{ikt}-\alpha_{ij}-\epsilon_{ijt}\ \exists j\in K,\forall k\in\mathcal{J}\setminus K
\end{align*}
and the variable $Y_{it}^{K}$ is increasing in $W_{ijt}'\theta_{0}-W_{ikt}'\theta_{0}$
for any $j\in K$ and $k\in\mathcal{J}\setminus K$.

Adapting the key identification idea to the current context, the identification results
for panel multinomial choice models are presented in the following
proposition.
\begin{prop}
\label{prop:mul} Define $\Theta_{I,mul}$ as the set that consists of all $\theta=\left(\beta^{'},\gamma^{'}\right)^{'}$
such that
\begin{align}
 & \max_{t=1,...,T}\mathbb{P}\left(\rest{Y_{it}^{K}=1,\left(W_{ijt}-W_{ikt}\right)^{'}\theta\leq c_{jk}\ \forall j\in K,k\in\mathcal{J}\setminus K}z\right)\nonumber \\
\leq & 1-\max_{t=1,...,T}\mathbb{P}\left(\rest{Y_{is}^{K}=0,\left(W_{ijs}-W_{iks}\right)^{'}\theta\geq c_{jk}\ \forall j\in K,k\in\mathcal{J}\setminus K}z\right),\label{eq:ind_mul}
\end{align}
for any subset $K\subset\mathcal{J}$, any $c_{jk}\in\mathcal{R}$, any $j\in K$
and $k\in\mathcal{J}\setminus K$, and any realization $z$ in the
support of $Z_{i}$. Then, under Assumption \ref{assu:PartStat},
$\theta_{0}\in\Theta_{I,mul}$.
\end{prop}

Below we show that Proposition \ref{prop:mul} specializes to the
corresponding result in \citet{pakes2019}, who focuses on the static
panel multinomial choice model without any endogeneity. Since \citet{pakes2019}
establishes the sharpness of their identification result under their
setup, our Proposition \ref{prop:mul} is also sharp (under their
static two-period setting).

However, a key improvement of our result relative to that in \citet{pakes2019}
is that Proposition \ref{prop:mul} allows for any type of endogeneity
including dynamic multinomial models with lagged dependent variable,
as well as the inclusion of contemporaneously endogenous variables
such as product prices. For example, consider the following dynamic
model:
\[
u_{ijt}=Z_{ijt}'\beta_{0}+\mathbf{\mathbbm1}\left\{ Y_{i,t-1}=j\right\} \gamma_{0,j}-P_{ijt}\lambda_{0}+\alpha_{ij}+\epsilon_{ijt}.
\]
where individual $i$'s utility at time $t$ can potentially depend
on their choices in the previous period $t-1$ and we allow the dynamic
effect $\gamma_{0,j}$ to vary across choices, and $P_{ijt}$ is the price
of product $j$ faced by consumer $i$ at time $t$. To the best of our knowledge,
no previous work has considered such generalization of \citet{pakes2019}
that can incorporate price endogeneity and past-choice dependence.
Even though Proposition \ref{prop:mul} is presented as a byproduct
of our general identification strategy, it nevertheless presents a
substantive progress in the related literature on panel multinomial
choice models.

\subsubsection*{Reconciliation with \citet{pakes2019}}

\label{subsec:mul_app} Next, we show that Proposition \ref{prop:mul}
specializes to those in \citet{pakes2019}, who focus on the static
panel multinomial choice model without any endogeneity. Since
\citet{pakes2019} establishes the sharpness of their identification
set in a two-period setting and our identification set reproduces theirs, the sharpness of our identification set follows immediately in this setting.

Formally, \citet{pakes2019} characterizes the sharp identified set
for $\theta_{0}$ under the full stationarity assumption given all covariates:
\[
\epsilon_{is}\mid W_{i},\alpha_{i}\stackrel{d}{\sim}\epsilon_{it}\mid W_{i},\alpha_{i}.
\]
Under this condition, for two periods $\left(t,s\right)$ our identifying
condition in \eqref{eq:ind_mul} is simplified to
\begin{align}
 & \mathbb{P}\left(Y_{is}^{K}=1,(w_{js}-w_{ks})'\theta_{0}\leq c_{jk}\ \forall j\in K,k\in\mathcal{J}\setminus K\mid w\right)\nonumber \\
\leq\  & 1-\mathbb{P}\left(Y_{it}^{K}=0,(w_{jt}-w_{kt})'\theta_{0}\geq c_{jk}\ \forall j\in K,k\in\mathcal{J}\setminus K\mid w\right)\label{eq:mul_sta}
\end{align}
The above equation is only informative when $(w_{js}-w_{ks})'\theta_{0}\leq c_{jk}\leq(w_{jt}-w_{kt})'\theta_{0}$
for any $j\in K,k\in k\in\mathcal{J}\setminus K$; otherwise either
the upper bound becomes one or the lower bound becomes zero so that
condition \eqref{eq:mul_sta} holds for any $\theta$. There exists one
value $c_{jk}$ satisfying the condition $(w_{js}-w_{ks})'\theta_{0}\leq c_{jk}\leq(w_{jt}-w_{kt})'\theta_{0}$
is equivalent to $(w_{js}-w_{ks})'\theta_{0}\leq(w_{jt}-w_{kt})'\theta_{0}$,
generating the following inequality: for any $K\subset\mathcal{J}$,
\begin{align*}
\text{If \ensuremath{\quad}} & (w_{js}-w_{ks})'\theta_{0}\leq(w_{jt}-w_{kt})'\theta_{0}\ \ \forall j\in K,k\in\mathcal{J}\setminus K\\
\text{then\ensuremath{\quad}} & \mathbb{P}\left(Y_{it}^{K}=0\mid w\right)\leq1-\mathbb{P}\left(Y_{is}^{K}=1\mid w\right)
\end{align*}
which becomes the same result in \citet{pakes2019} (Proposition 1,
P. 12):
\begin{align*}
\text{If \ensuremath{\quad}} & (w_{js}-w_{ks})'\theta_{0}\leq(w_{jt}-w_{kt})'\theta_{0}\ \ \forall j\in K,k\in\mathcal{J}\setminus K\\
\text{then\ensuremath{\quad}} & \mathbb{P}\left(Y_{is}\in K\mid w\right)\leq\mathbb{P}\left(Y_{it}\in K\mid w\right)
\end{align*}
since $Y_{it}^{K}=1$ is equivalent to $Y_{it}\in K$ by the definition.

\section{Simulation}

\label{sec:simu}

In this section, we focus on the static ordered response model Section
\ref{subsec:order} and implement the kernel-based CLR inference approach
proposed in the papers by \citet*{chernozhukov2013} and \citet{chen2019breaking},
which was developed to construct confidence interval based on general
conditional moment inequalities.

In Appendix \ref{subsec:Q_gamma}, we also conduct a simulation exercise of a different nature. We numerically compute and visualize the identified set under two DGP configurations in dynamic binary choice setting, but do not implement the finite-sample estimation and inference procedure.

\subsection{Static Ordered Response Model}

\label{subsec:static}

This section explores a static ordered choice model with three choices
$Y_{it}\in\{1,2,3\}$. We consider the following two-period model
with $T=2$, and the latent dependent variable $Y_{it}^{*}$ is generated
as:
\[
Y_{it}^{*}=Z_{it}^{1}\beta_{01}+Z_{it}^{2}\beta_{02}+\alpha_{i}+\epsilon_{it},
\]
where the covariate $Z_{it}^{k}$ satisfies $Z_{it}^{k}\sim\mathcal{N}(0,\sigma_{z})$
for $k\in\{1,2\}$; the fixed effects $\alpha_{i}$ are given as $\alpha_{i}=\sum_{t=1}^{T}(Z_{it}^{1}+Z_{it}^{2})/(4*\sigma_{z}*T)$,
so they are correlated with the covariates; the error term $(\epsilon_{i1},\epsilon_{i2})$
follows the normal distribution $\mathcal{N}(\mu,\Sigma)$ with $\mu=(0,0)$
and $\Sigma=(1\ \rho;\rho\ 1)$. The true parameter is $\beta_{0}:=(\beta_{0,1},\beta_{02})'=(1,1)'$,
the repetition number is $B=200$, and the sample size is $n=\{2000,8000\}$.
We consider three specifications for $\sigma_{z}\in\{1,1.5,2\}$ and $\rho\in\{0,0.25,0.5\}$.

The observed dependent variable $Y_{it}$ is given as
\[
Y_{it}=1*(Y_{it}^{*}\leq b_{2})+2*(b_{2}<Y_{it}^{*}\leq b_{3})+3*(Y_{it}^{*}>b_{3}),
\]
where $b_{2}=-1$ and $b_{3}=1$. We write $Y_{i}:=(Y_{i1},Y_{i2})$ and $Z_{i}:=(Z_{i1},Z_{i2})$,

Proposition \ref{prop:static_ordered} in Appendix \ref{subsec:order_stat}, as a corollary of Proposition \ref{prop:order} with pivotal choices of $c$'s, provides a characterization of the identified set for $\beta_{0}$ under the static ordered choice model via
the following conditional moment inequalities: for $s\neq t\leq2$,
\[
E[g(Z_{i},Y_{i};\beta_{0})\mid z]\geq0,
\]
where
\[
g(Z_{i},Y_{i};\beta)=\left\{ \begin{aligned} & \mathbf{\mathbbm1}\{b_{2}-Z_{is}'\beta\geq b_{2}-Z_{it}'\beta\}(\mathbf{\mathbbm1}\{Y_{is}=1\}-\mathbf{\mathbbm1}\{Y_{it}=1\});\\
 & \mathbf{\mathbbm1}\{b_{2}-Z_{is}'\beta\geq b_{3}-Z_{it}'\beta\}(\mathbf{\mathbbm1}\{Y_{is}=1\}-\mathbf{\mathbbm1}\{Y_{it}\in\{1,2\}\});\\
 & \mathbf{\mathbbm1}\{b_{3}-Z_{is}'\beta\geq b_{2}-Z_{it}'\beta\}(\mathbf{\mathbbm1}\{Y_{is}\in\{1,2\}\}-\mathbf{\mathbbm1}\{Y_{it}=1\});\\
 & \mathbf{\mathbbm1}\{b_{3}-Z_{is}'\beta\geq b_{3}-Z_{it}'\beta\}(\mathbf{\mathbbm1}\{Y_{is}\in\{1,2\}\}-\mathbf{\mathbbm1}\{Y_{it}\in\{1,2\}\}),
\end{aligned}
\right.
\]
upon which our estimation and inference exercise in this subsection will be based.

The first element $\beta_{01}$ of the parameter $\beta_{0}$ is normalized
to one, and we are interested in conducting inference for the parameter
$\beta_{02}$ using the CLR approach. Tables \ref{table:static_sigma}
and \ref{table:static_rho} report the average confidence interval
(CI) for $\beta_{02}$, the coverage probability (CP), the average length
of the CI (length), the power of the test at zero (power), and the
mean absolute deviation of the lower bound ($l_{MAD}$) and upper
bound ($u_{MAD}$) of the CI.

\begin{table}[!htbp]
\centering
\caption{Performance of $\protect\beta_{02}$ under different values of $\protect\sigma_{z}$
($\rho=0.25$)}
\label{table:static_sigma}
\begin{tabular}{c|cccccc}
\hline
$\sigma_{z}$  & CI  & CP  & length  & power  & $l_{MAD}$  & $u_{MAD}$ \tabularnewline
\hline
 & \multicolumn{6}{c}{ $N=2000$}\tabularnewline
\hline
$\sigma_{z}=1$  & {[}0.537, 1.760{]}  & 0.876  & 1.222  & 1.000 & 0.476  & 0.784 \tabularnewline
$\sigma_{z}=1.5$  & {[}0.556, 1.768{]}  & 0.934 & 1.212 & 1.000  & 0.454  & 0.773 \tabularnewline
$\sigma_{z}=2$  & {[}0.567, 1.791{]}  & 0.950  & 1.224  & 1.000  & 0.440  & 0.796 \tabularnewline
\hline
 & \multicolumn{6}{c}{$N=8000$}\tabularnewline
\hline
$\sigma_{z}=1$  & {[}0.570, 1.532{]}  & 0.939  & 0.962  & 1.000  & 0.439  & 0.548 \tabularnewline
$\sigma_{z}=1.5$  & {[}0.607, 1.561{]}  & 0.975  & 0.954  & 1.000  & 0.398  & 0.563 \tabularnewline
$\sigma_{z}=2$  & {[}0.618, 1.571{]}  & 0.985  & 0.953  & 1.000  & 0.383  & 0.573 \tabularnewline
\hline
\end{tabular}
\end{table}

\begin{table}[!htbp]
\centering
\caption{Performance of $\protect\beta_{02}$ under different values of $\rho$
($\protect\sigma_{z}=1$)}
\label{table:static_rho}
\begin{tabular}{c|cccccc}
\hline
$\rho$  & CI  & CP  & length  & power  & $l_{MAD}$  & $u_{MAD}$ \tabularnewline
\hline
 & \multicolumn{6}{c}{ $N=2000$}\tabularnewline
\hline
$\rho=0$  & {[}0.537, 1.755{]}  & 0.895  & 1.218  & 1.000 & 0.476  & 0.773 \tabularnewline
$\rho=0.25$  & {[}0.537, 1.760{]}  & 0.876  & 1.222  & 1.000 & 0.476  & 0.784 \tabularnewline
$\rho=0.5$  & {[}0.511, 1.765{]}  & 0.909  & 1.254  & 1.000  & 0.497 & 0.785 \tabularnewline
\hline
 & \multicolumn{6}{c}{ $N=8000$}\tabularnewline
\hline
$\rho=0$  & {[}0.584, 1.553{]}  & 0.933  & 0.969  & 1.000  & 0.436  & 0.568 \tabularnewline
$\rho=0.25$  & {[}0.570, 1.532{]}  & 0.939  & 0.962  & 1.000  & 0.439  & 0.548 \tabularnewline
$\rho=0.5$  & {[}0.573, 1.526{]}  & 0.934  & 0.954  & 1.000  & 0.442  & 0.541 \tabularnewline
\hline
\end{tabular}
\end{table}

As shown in Tables \ref{table:static_sigma} and \ref{table:static_rho},
our approach exhibits robust performance across various specifications
of standard deviation $\sigma$ and correlation coefficients $\rho$.
The coverage probabilities of the 95\% confidence interval (CI) for
$\beta_{02}$ are close to the nominal level, the length of the CI is
reasonably small, and the CI consistently excludes zero. When the
sample size increases, there is a significant decrease in CI length,
an improvement in coverage probability, and a reduction of the mean
absolute deviation (MAD) for the lower and upper bounds of the CI.
Overall, these results demonstrate the good performance of our approach
in different DGP designs.

\subsection{Dynamic Ordered Response Model}

In this section, we investigate a dynamic ordered choice model with
one lagged dependent variable $Y_{i,t-1}$. The latent dependent variable
$Y_{it}^{*}$ is generated as follows:
\[
Y_{it}^{*}=Z_{it}\beta_{0}+Y_{i,t-1}\gamma_{0}+\alpha_{i}+\epsilon_{it}.
\]
where the endogenous variable is the lagged dependent variable $Y_{i,t-1}$.
We study three periods $T=3$ to illustrate our approach with multiple
periods. The DGP is similar: the exogenous covariate $Z_{it}$ satisfies
$Z_{it}\sim\mathcal{N}(0,\sigma_{z})$; the fixed effects $\alpha_{i}$ are
given as $\alpha_{i}=\sum_{t=1}^{T}Z_{it}/(4*\sigma_{z}*T)$; the error term
$(\epsilon_{i1},\epsilon_{i2},\epsilon_{i3})$ follows the normal distribution $\mathcal{N}(\mu,\Sigma)$
with $\mu=(0,0,0)$ and $\Sigma=(0.5\ c\ c;c\ 0.5\ c;c\ c\ 0.5)$,
where $c=0.5*\rho$. The true parameter is $\theta_{0}:=(\beta_{0},\gamma_{0})'=(1,1)'$,
the repetition number is $B=200$, and the sample size is $n\in\{2000,8000\}$.
We consider three specifications for $\sigma_{z}\in\{1,1.5,2\}$ and $\rho\in\{0,0.25,0.5\}$.

The observed dependent variable $Y_{it}$ is given as
\[
Y_{it}=1*(Y_{it}^{*}\leq b_{2})+2*(b_{2}<Y_{it}^{*}\leq b_{3})+3*(Y_{it}^{*}>b_{3}),
\]
for $1\leq t\leq T$. The initial value $Y_{i0}\in\{1,2,3\}$ is generated
independently of all variables and follows the distribution $\mathbb{P}(Y_{i0}=1)=0.6,\mathbb{P}(Y_{i0}=2)=\mathbb{P}(Y_{i0}=3)=0.2$.

In this dynamic model, the covariates $Z_{i}:=(Z_{it})_{t=1}^{T}$
and the initial value $Y_{i0}$ are exogenous, while the lagged variable
$Y_{i,t-1}$ is endogenous. Proposition \ref{prop:order} characterizes
the identified set for $\theta_{0}$ with the following conditional moment
inequalities:
\begin{itemize}
\item[(1)] When $s\in\{2,3\}$.
\[
\begin{aligned} & \sum_{j=1}^{2}\mathbb{P}\left(Y_{is}=y_{j},b_{j+1}-z_{s}'\beta-Y_{is-1}\gamma\leq c\mid z,y_{0}\right),\\
\leq\  & 1-\sum_{j=2}^{3}\mathbb{P}\left(Y_{i1}=y_{j}\mid z,y_{0}\right)*\mathbf{\mathbbm1}\left\{ b_{j}-z_{1}'\beta-y_{0}\gamma\geq c\right\}
\end{aligned}
\]
\[
\begin{aligned} & \sum_{j=1}^{2}\mathbb{P}\left(Y_{i1}=y_{j}\mid z,y_{0}\right)*\mathbf{\mathbbm1}\{b_{j+1}-z_{1}'\beta-y_{0}\gamma\leq c\},\\
\leq\  & 1-\sum_{j=2}^{3}\mathbb{P}\left(Y_{is}=y_{j},b_{j}-z_{s}'\beta-Y_{is-1}\gamma\geq c\mid z,y_{0}\right)
\end{aligned}
\]
for any $c\in\{b_{j}-z_{1}'\beta-y_{0}\gamma,b_{j}-z_{s}'\beta-\gamma,b_{j}-z_{s}'\beta-2\gamma,b_{j}-z_{s}'\beta-3\gamma\}_{j=2}^{T}$;
\item[(2)] When $s,t\in\{2,3\}$,
\[
\begin{aligned} & \sum_{j=1}^{2}\mathbb{P}\left(Y_{it}=y_{j},b_{j+1}-z_{t}'\beta-Y_{it-1}\gamma\leq c\mid z,y_{0}\right),\\
\leq\  & 1-\sum_{j=2}^{3}\mathbb{P}\left(Y_{is}=y_{j},b_{j}-z_{s}'\beta-Y_{is-1}\gamma\geq c\mid z,y_{0}\right)
\end{aligned}
\]
for any $c\in\{b_{j}-z_{s}'\beta-\gamma,b_{j}-z_{s}'\beta-2\gamma,b_{j}-z_{s}'\beta-3\gamma,b_{j}-z_{t}'\beta-\gamma,b_{j}-z_{t}'\beta-2\gamma,b_{j}-z_{t}'\beta-3\gamma\}_{j=2}^{3}$.
\end{itemize}
We normalize the first parameter $\beta_{0}$ to one, and report the
performance of the coefficient $\gamma_{0}$ for the lagged dependent
variable. Tables \ref{table:dyn_sigma} and \ref{table:dyn_rho} illustrate
that our approach yields robust and informative results for the dynamic
ordered choice model across various DGP specifications. The coverage
probability of the CI nearly reaches 95\%, and the CI consistently
excludes zero, producing significant coefficients. These results remain
similar across different values of correlation coefficients. When
the standard deviation $\sigma_{z}$ increases, the length of the CI also
experiences a slight increase. This phenomenon occurs because, in
the dynamic model, only partial identification is achieved, and the
bound for $\gamma_{0}$ depends on the variation in $\Delta z'\beta_{0}$. A
larger variation in $\Delta z'\beta_{0}$ may result in a wider identified
set in this specification, but it still provides informative results.
As the sample size increases, the confidence interval shrinks, and
concurrently, the coverage probability improves in all specifications.

\begin{table}[!htbp]
\centering
 \caption{Performance of $\protect\gamma_{0}$ under different values of $\protect\sigma_{z}$
($\rho=0.25$)}
\label{table:dyn_sigma}
\begin{tabular}{c|cccccc}
\hline
$\sigma_{z}$  & CI  & CP  & length  & power  & $l_{MAD}$  & $u_{MAD}$ \tabularnewline
\hline
 & \multicolumn{6}{c}{ $N=2000$}\tabularnewline
\hline
$\sigma_{z}=1$  & {[}0.446, 1.606{]}  & 0.935  & 1.160 & 1.000 & 0.565  & 0.625 \tabularnewline
$\sigma_{z}=1.5$  & {[}0.375, 1.673{]}  & 0.959  & 1.298  & 1.000  & 0.629  & 0.693 \tabularnewline
$\sigma_{z}=2$  & {[}0.311, 1.730{]}  & 0.960  & 1.418  & 1.000 & 0.700 & 0.739 \tabularnewline
\hline
 & \multicolumn{6}{c}{ $N=8000$}\tabularnewline
\hline
$\sigma_{z}=1$  & {[}0.529, 1.495{]}  & 0.969  & 0.966  & 1.000  & 0.473  & 0.504 \tabularnewline
$\sigma_{z}=1.5$  & {[}0.460, 1.559{]}  & 0.965  & 1.100  & 1.000  & 0.548  & 0.564 \tabularnewline
$\sigma_{z}=2$  & {[}0.427, 1.585{]}  & 0.985  & 1.158  & 1.000 & 0.573 & 0.589 \tabularnewline
\hline
\end{tabular}
\end{table}

\begin{table}[!htbp]
\centering
 \caption{Performance of $\protect\gamma_{0}$ under different values of $\rho$
($\sigma_{z}=1$)}
\label{table:dyn_rho}
\begin{tabular}{c|cccccc}
\hline
$\rho$  & CI  & CP  & length  & power  & $l_{MAD}$  & $u_{MAD}$ \tabularnewline
\hline
 & \multicolumn{6}{c}{ $N=2000$}\tabularnewline
\hline
$\rho=0$  & {[}0.472, 1.593{]}  & 0.932  & 1.121  & 1.000  & 0.550  & 0.607 \tabularnewline
$\rho=0.25$  & {[}0.446, 1.606{]}  & 0.935  & 1.160 & 1.000 & 0.565  & 0.625 \tabularnewline
$\rho=0.5$  & {[}0.457, 1.631{]}  & 0.943  & 1.173  & 1.000 & 0.548  & 0.648 \tabularnewline
\hline
 & \multicolumn{6}{c}{ $N=8000$}\tabularnewline
\hline
$\rho=0$  & {[}0.528, 1.472{]}  & 0.958  & 0.945 & 1.000 & 0.475  & 0.487 \tabularnewline
$\rho=0.25$  & {[}0.529, 1.495{]}  & 0.969  & 0.966  & 1.000  & 0.473  & 0.504 \tabularnewline
$\rho=0.5$  & {[}0.535, 1.515{]}  & 0.975 & 0.980  & 1.000 & 0.467 & 0.519 \tabularnewline
\hline
\end{tabular}
\end{table}






















\section{Empirical Application}

\label{sec:appl}

In this section, we apply our proposed approach to explore the empirical
analysis of income categories using the NLSY79 dataset. The dependent
variable is three categories of (log) income, denoted by the three
values $\{1,2,3\}$, indicating whether an individual falls within
the top 33.3\% highest income bracket, the 33.3\%-66.6\% highest income
range, and the lowest 33.3\% income tier, respectively. We include
two covariates in this analysis: one is tenure, defined as the total
duration (in weeks) with the current employer, and the other is a
residence indicator for whether one lives in an urban or rural area.\footnote{This dataset also contains other crucial factors for income such as
gender and race. However, these variables are time-invariant and cannot
be included for panel models with fixed effects.} We use two periods of panel data from the years 1982 and 1983 as
well as the income data from 1981 as initial values, and there are
$n=5259$ individuals in each period. The following table presents
the summary statistics of these variables.

\begin{table}[!htbp]
\centering
 \caption{Application: Summary Statistics}
\label{table: sum}
\begin{tabular}{c|ccc}
\hline
 & income category  & residence  & tenure /100 \tabularnewline
mean  & 1.990  & 0.799  & 0.825 \tabularnewline
s.d. & 0.810  & 0.401  & 0.738 \tabularnewline
25\% quantile  & 1.000  & 1.000  & 0.220 \tabularnewline
median  & 2.000  & 1.000  & 0.605 \tabularnewline
75\% quantile  & 3.000  & 1.000  & 1.280 \tabularnewline
minimum  & 1.000  & 0.000  & 0.010 \tabularnewline
maximum  & 3.000  & 1.000  & 4.850 \tabularnewline
\hline
\end{tabular}
\end{table}

We adopt various ordered response models introduced in Section \ref{subsec:order}
to analyze the income category. The first model is the standard static
model without any endogeneity. The second is the static model, while
treating residence as an endogenous covariate. Residence is potentially
endogenous since the choice of living area is typically endogenously
determined and may be correlated with individuals' unobserved ability
or preference. The last model considers the dynamic model with one
lagged dependent variable, allowing people's income in current periods
to depend on their income in the last period. All three models allow
for individual fixed effects and do not impose any parametric distributions
on time-changing shocks. Proposition \ref{prop:order} characterizes
the identified set of the model coefficients for these three models
using conditional moment inequalities. Similar to Section \ref{sec:simu},
we exploit the kernel-based CLR inference method to construct confidence
intervals. The coefficient of the variable ``residence'' is normalized
to one. Table \ref{table: app} reports the confidence intervals for
the coefficients of the covariate ``tenure'' and the lagged dependent
variable (when applicable).

\begin{table}[!htbp]
\centering
 \caption{Application: Income Categories}
\label{table: app}
\begin{tabular}{c|ccc}
\hline
 & $\beta_{0,1}$ (residence)  & $\beta_{0,2}$ (tenure)  & $\gamma_{0}$ (lag) \tabularnewline
exogenous static model  & 1  & {[}0.612, 0.939{]}  & - \tabularnewline
endogenous static model  & 1  & {[}0.041, 0.939{]}  & - \tabularnewline
dynamic model  & 1  & {[}0.531, 0.694{]}  & {[}0.286, 0.612{]} \tabularnewline
\hline
\end{tabular}
\end{table}

As shown in Table \ref{table: app}, tenure exhibits a significantly
positive effect on the income category across all specifications.
When allowing for the endogeneity of residence, the confidence interval
for tenure becomes wider, as we need to account for all possible correlations
between residence and unobserved heterogeneity. The results from the
dynamic model show that the income category in the current period
is also positively affected by last period's income, and this effect
is significant. Furthermore, this analysis demonstrates the flexibility
of our approach, which can not only allow for endogeneity introduced
by dynamics but also account for contemporaneous endogeneity.

\section{Conclusion}

\label{sec:conc}

We introduce a general method to (partially) identify index parameters in nonlinear panel data models
based on a partial stationarity condition. This approach accommodates
dynamic models with an arbitrary finite number of lagged outcome variables
and other types of endogenous covariates. We demonstrate how our key
identification strategy can be applied to obtain informative identifying
restrictions in various limited dependent variable models, including
binary choice, ordered response, multinomial choice, as well as censored
outcome models. Finally, we further extend this approach to study
general nonseparable models.

There are some natural directions for follow-up research. In this
paper we focus on the identification of model parameters, but it would
also be interesting to investigate how our identification strategy
can be exploited to obtain informative bounds on average marginal
effects and other counterfactual parameters, say, following the approach
proposed in \citet{botosaru2022identification}.\footnote{\citet{botosaru2022identification} proposes an approach to obtain
bounds on counterfactual CCPs in semiparametric dynamic panel data
models, assuming that the index parameters are (partially) identified.} Also, our identification strategy should be adaptable to exploit
additional restrictions imposed by time-exchangeability assumptions
such as in \citet*{mbakop2023identification}, which not only impose
homogeneity on per-period marginals of errors but also on their intertemporal
dependence structures. Additionally, the idea of bounding an endogenous
object (parametric index in our case) by an arbitrary constant so
as to obtain an object free of endogeneity issues may have broader
applicability beyond the models studied in this work, and it remains
to see whether our key identification strategy can be further adapted
to other structures.

\bibliographystyle{ecta}
\bibliography{stationarity}

\newpage{}