Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
71,640 characters · 16 sections · 60 citation commands
Robust Analysis of Short Panels
This paper deals with models of processes delivering values of outcomes, $Y$ , given values of exogenous variables, $Z$, and latent, that is unobserved, variables $U$ and $V$. The models that are the focus of this paper all leave the distribution of $V$ on its known support and its covariation with all other variables completely unrestricted. By contrast, latent variable $U$ may be required to be, to some degree, independent of $Z$.
Leading examples of latent variables in structural econometric models employed in practice on whose distribution one may not want to impose restrictions are the individual-specific unobserved variables included in many panel data models, sometimes called \textquotedblleft fixed effects\textquotedblright\ and the historic values of outcomes dynamically determined by a process, commonly called \textquotedblleft initial conditions\textquotedblright .
The following example has both elements, a \textquotedblleft fixed effect\textquotedblright , $C$, and an initial condition, $Y_{10}$.
In many cases found in practice in which $T$ is large, the value $V$ takes for each observational unit is identified. In this case econometric analysis can proceed placing no restrictions at all on the distribution of $V$ and treating it as a parameter to be estimated. When $T$ is not large this is unattractive because there may be intolerable inaccuracy in the estimation of $V$ which may contaminate estimates of other parameters. Additionally there is the incidental parameters problem set out in Neyman/Scott:48 and reviewed in Lancaster:00.
Faced with this problem, for small $T$, most papers proceed to obtain information on structural features by placing distributional restrictions on $V$. Section (ref) lists many examples. It is good to know what knowledge of structural features can be obtained absent such restrictions. That allows the force of distributional restrictions on $V$ to be assessed and offers the possibility of detecting misspecification. This paper shows how that knowledge can be obtained. The results given here open the way to a relatively robust analysis of models like panel models with fixed effects in which there are latent variables on which one desires to place no distributional restrictions.
This paper presents characterizations of identified sets of structures and structural features in models admitting unobserved variables such as $V$ whose distribution is unrestricted. There can be endogenous explanatory variables as in Example (ref) when $\alpha \neq 0$. A model may be incomplete in the sense that, given values of all observed and all unobserved variables and a specification of parameter values and functional forms, the model can deliver a nonsingleton set of values of outcomes. Identified sets are characterized by systems of moment inequalities. Estimation and inference can proceed using established econometric methods.
The strategy employed here removes unrestricted latent variables, $V$, by projection.\footnote{Our eschewal of restrictions on the distribution of $V$ accords with the approach in Neyman/Scott:48 in which the elements of $V$ are treated as parameters, subject to no restrictions.} We derive, for each value of the observed variables, the set of values of unobserved $U$ compatible with that value. Values of $U$ in such a set are associated with alternative values of $V$. Typically a value of $U$ in such a set can deliver more than one value of $Y$. So, on removing latent variables $V$, there remains an incomplete model.\footnote{If the model is incomplete before projection then different values of $V$ can deliver different sets of values of $Y$.}
Identification analysis is conducted in the context of the Generalized Instrumental Variable (GIV) framework introduced in chesher2017generalized in which probability distributions of such sets of values of $U$ induced by the observed distributions of outcomes are essential elements.
Section (ref) considers the relationship of this work to some other results in the literature. Section (ref) presents characterizations of identified sets of structures. Sections (ref) to (ref) set out applications to linear panel models and to models of binary response panels, ordered choice panels, multiple discrete choice panels, simultaneous binary outcome panels, and models of panels with censored continuous outcomes.
rasch1960probabilistic, rasch1961general, andersen1970asymptotic, and chamberlain2010binary study point identifying static panel models (i.e. $\gamma =0$ in ((ref))) with restrictions requiring $U_{1},...,U_{T}$ to be independent over time and distributed independently of $Z$ and independently of the fixed effect and each with logistic marginal distributions. Like all the papers referred to in this section, except one paper which is noted, these models do not admit endogenous explanatory variables.
In the linear panel data model with fixed effects, differencing across time periods removes the fixed effect, delivering events whose probability of occurrence can be known and is invariant with respect to changes in the value of the fixed effect. Under suitable support restrictions this leads to point identification. In nonlinear panel data models with fixed effects, this simple differencing strategy does not apply. Nonetheless, in the Rasch-Andersen-Chamberlain set up, events whose probabilities of occurrence are invariant to changes in the value of the fixed effect are found. Under particular distributional restrictions point identification results. More recent papers on nonlinear panel data models have taken a similar approach.
honore2000panel study a dynamic model as in ((ref)) but with no endogenous explanatory variable ($\alpha =0$) with the $U_{t}$'s independent of the fixed effect, independent over time, distributed independently of $Z$ and with logistic distributions. That paper also studies a case in which the logistic distribution restriction is dropped and a case with multinomial logit panels with latent variables $U$ independent of the fixed effects and independent of $Z$. honore2019panel extends this work, studying multivariate dynamic panel data logit models with fixed effects. Many papers, like these, invoke restrictions requiring independence between $U_{t}$'s and the fixed effect conditional on some of the other observable variables including honore2006bounds, honore2021identification,\footnote{This paper considers models in which there are simultaneous equations in binary outcomes and so, endogenous explanatory variables.} Dobronyi/Gu/Kim:21, honore2022dynamic, Davezies/D'Haultfoeuille/Laage:22, Kitazawa:22, Bonhomme/Dano/Graham:23, Dano:23, Davezies/D'Haultfoeuille/Mugnier:23, and honore2021dynamic. Such independence restrictions are not imposed here.\footnote{ One approach in such settings, demonstrated by e.g. honore2022dynamic and honore2021dynamic, is the functional differencing approach developed in Bonhomme:12. This however requires knowledge of the distribution of $F_{Y|Z,C}$, which one does not have in models such as that of Example 1 without knowledge of the joint distribution of $U$ and $C$.} This permits for example the $U_t$'s to exhibit heteroskedastic variation with observational-unit-specific fixed effects.
There are many papers studying panel models of binary outcomes and multiple discrete choice under conditional stationarity restrictions on the distribution of the time varying latent variables introduced in manski1987semiparametric. These papers include Chernozhukov/Fernandez-Val/Hahn/Newey:09, shi2018estimating, Gao/Li:20, khan2021inference, pakes2021unobserved, pakes2022moment, Dobronyi/Ouyang/Yang:23, khan2023identification, and mbakop2023identification.
In all of these cases the stationarity restriction placed on time-varying unobservable heterogeneity is required to hold conditional on the value of the fixed effect and the observable exogenous variables, which restricts the covariation of the fixed effect and $U$.\footnote{ In the binary response specification ((ref)) conditional stationarity implies that for all $z$, $F_{U_1|Z=z,C=c} = F_{U_2|Z=z,C=c}$ and $ F_{U_1|Z=z,C=c^{\prime }} = F_{U_2|Z=z,C=c^{\prime }}$ for any $c,c^{\prime } $, which restricts how the conditional distribution of $U$ can change with values of the fixed effect $C$. As pointed out by Chernozhukov/Fernandez-Val/Hahn/Newey:09 the stationarity restriction $ U_t|C,Z \overset{d}{=}U_1|C,Z$ for all $t$ is equivalent to $(U_t,C)|Z \overset{d}{=}(U_1,C)|Z$ for all $t$.} In contrast, the models considered in this paper impose no restrictions on the covariation of the fixed effect with any variable.
The only previous paper of which we are aware that provides partial identification analysis for discrete outcome panel data models absent restrictions on the covariation of the fixed effect with any other variables is aristodemou2021semiparametric. That paper provides set-identifying moment inequalities in panel data models of binary response and ordered choice when the covariation of the fixed effects with other variables is unrestricted. The results developed in this paper provide a rule-directed procedure for enumerating all events whose probability is invariant with respect to the value of unrestricted latent variables thereby delivering sharp set identification for these and other nonlinear panel data models.
Application of sharp set identification analysis to panel data models with censored outcomes is demonstrated in Section (ref). Observable implications in the form of moment equalities are derived for such models with Tobit-type censoring at zero in both static and dynamic contexts in Honor{\'e} (1992, 1993),\nocite{Honore:92}\nocite{Honore:93} Honore/Hu:2002, and Hu:2002, all in models in which the $U_t$'s satisfy the conditional stationarity assumption that has also been used in discrete outcome panel models. The only previous paper of which we are aware that provides identification analysis for censored outcome panel models without restricting the covariation of the fixed effect with other variables is Khan/Ponomareva/Tamer:16 (KPT), which provides the sharp identified set for slope coefficient $\beta$ in a static two-period model in which $U_2 - U_1$ and $Z$ are independent. Extensions are provided to some specialized dynamic models with two periods of observations with an observed initial condition and inequality restrictions on parameters. The analysis here additionally accommodates more periods, unobserved initial conditions, and endogenous explanatory variables. Like KPT we allow the censoring value to vary and to be endogenous, nesting the classical case of fixed censoring found in Tobit models.
This paper presents a generally applicable approach to identification analysis in a wide class of nonlinear panel data models in which there are distributionally unrestricted latent variables and gives examples of the results it produces. Most of our examples feature discrete outcomes, but the application to the censored outcome model in Section (ref) demonstrates that the analysis applies more broadly.
First the notation employed in this paper is introduced.
Notation. Generically $\mathcal{R}_{A}$\ denotes the support of random variable $A$\ and $L_{A|Z=z}$ \ denotes a conditional probability distribution of random variable $A$ \ given $Z=z$\textit{. }$L_{A|Z=z}(\mathcal{S})$\textit{\ is the conditional probability }$A$\textit{\ takes a value in set }$\mathcal{S}$ \textit{\ given }$Z=z$\textit{. }$\mathcal{L}_{A|Z}\equiv \{L_{A|A=z};z\in R_{Z}\}$\textit{\ is the collection of conditional distributions delivered by a joint distribution }$L_{AZ}$\textit{\ when the support of }$Z$\textit{\ is }$\mathcal{R}_{Z}$\textit{. }$A \mathbin{\vbox{\baselineskip=0pt\lineskip=0pt \moveright2.5pt\hbox{$\|$} \hrule height 0.2pt width 10pt}} B$ \textit{denotes }$A$\textit{\ and }$B$\textit{\ are independently distributed.} \textit{Sets and set-valued random variables are expressed using calligraphic font. Collections of sets are expressed using sans serif font. }$ \mathbb{R} $ \textit{denotes the real line. The empty set is denoted $\emptyset $. For $T > 1$, notation $[T]$ denotes $\{1,...T\}$. For any random vectors $X_1,...,X_T$ notation $\Delta_{ts} X \equiv X_t - X_s$ is used throughout. }
Variables $Y$ are endogenous outcomes, variables $Z$ are exogenous\footnote{In the sense that their values are not affected by the evolution of the process.} and variables $U$ and $V$ are latent variables. Random vectors $(Y,Z,U,V)$ are defined on a probability space $(\Omega ,\mathsf{L},\mathbb{P})$ endowed with the Borel sets on $\Omega $. The support of $(Y,Z,U,V)$ is a subset of a finite dimensional Euclidean space. The sampling process identifies $F_{YZ}$, equivalently the collection of conditional distributions $\mathcal{F}_{Y|Z}$ and $F_Z$, as occurs for example under random sampling of observational units. It is assumed throughout for ease of exposition that each observational unit delivers the same number of observations, but unbalanced panels are easily accommodated with some added notation.
Models place restrictions on a structural function $h:\mathcal{R} _{YZUV}\rightarrow \mathbb{R} $ which specifies the combinations of these variables that can occur via the following restriction.\footnote{ In the case of ((ref)) a suitable $h$ function would be
with $V=(C,Y_{0})$.}
Models place restrictions on the conditional probability distributions of $U$ given $Z$ which are elements of a collection $\mathcal{G}_{U|Z}$. Coupled pairs $(h,\mathcal{G}_{U|Z})$ are called structures. A model $ \mathcal{M}$ is a collection of structures that obey the restrictions imposed a priori on the data generation process. This paper provides sharp identification analysis of structures $(h,\mathcal{G} _{U|Z}) \in \mathcal{M}$ and functionals thereof given knowledge of $ \mathcal{F}_{Y|Z}$.
The essential element of the models considered here is that they place no restrictions on the marginal distribution of $V$ and no restrictions on the covariation of $V$ with $(Z,U)$.
This paper shows how the framework set out in chesher2017generalized (CR) can be used to study cases with unobserved variables whose distribution and covariation with other variables is not subject to restrictions. The support of any initial condition components of $V$ is assumed known and the support of all “fixed effect” components of $V$ is assumed to be the entirety of the Euclidean space in which it resides. It is straightforward to generalize the analysis to cases in which the support of the fixed effect is restricted.
For all characterizations of identified sets of values of the pair $(h,\mathcal{G}_{U|Z}) \in \mathcal{M}$ it is assumed that a priori restrictions on $\mathcal{G}_{U|Z}$ are such that $U|Z$ is restricted absolutely continuous with respect to Lebesgue measure almost surely. This renders the boundary of sets $\mathcal{U}^{\ast}(y,z;h)$ to be measure zero with respect to any distribution $G_{U|Z=z}$. It is convenient to define the structural function $h$ such that sets $\mathcal{U}^{\ast }(Y,Z;h)$ are closed almost surely in the usual Euclidean topology, and we do so here, but this is of no substantive consequence and can be relaxed.\footnote{With some care equivalent results could be obtained allowing for random open sets and random closed sets, or by working with an alternative topology in which the sets under consideration are closed, such as the discrete topology when $\mathcal{R}_Y$ is discrete. One could also allow sets of values of unobservables that deliver “ties” in the optimal choice of discrete outcome with positive probability, and apply results of CR, with suitable care.}
Taken together the restrictions set out above ensure that Restrictions A1 - A6 of CR hold in the models considered, suitably modified to accommodate unobservable variables $(U,V)$ with the distribution of $V$ unrestricted. \footnote{The latent variables $U$ in restrictions A1-A6 of CR should be taken to include both the variables $U$ and $V$ of this paper. For completeness, these restrictions, adapted to the present context, are collected in Appendix (ref).}
Theorem (ref) provides a characterization of the identified set of structures, denoted $\mathcal{I}(\mathcal{M},\mathcal{F}_{Y|Z})$, delivered by a model $\mathcal{M}$ and a collection of distributions, $ \mathcal{F}_{Y|Z}$. This is the collection of distributions marginal with respect to $V$ obtained from some collection $\mathcal{F}_{YV|Z}$.
Formally the Theorem defines the identified set of structures $(h,\mathcal{G} _{U|Z})$ as
where, as in chesher2020generalized, for any random variable $A$ with distribution $F_{A}$ and random set $\mathcal{A}$, $F_{A}\preceq \mathcal{A}$ denotes that $F_{A}$ is selectionable with respect to the distribution of $ \mathcal{A}$.\footnote{ The probability distribution of random variable $A$ is selectionable with respect to the probabilty distribution of random set $\mathcal{A}$ when there exists (i) $\tilde{A}$ having the same distribution as $A$, and (ii) $ \widetilde{\mathcal{A}}$ having the same distribution as $\mathcal{A}$, both defined on the same probability space such that $\mathbb{P}[\tilde{A}\in \widetilde{\mathcal{A}}]=1$. See Definition 2 of Chesher and Rosen (2020).}
The proof relies on the following Lemma.
The proof of Theorem (ref) above proceeds as the proof of Theorem 2 in CR, replacing $U$ sets with $U^{\ast }$ sets.
The identified set of structures can be characterized as shown in Corollary (ref) using the characterization of selectionability given in artstein1983distributions, as in Corollary 2 of CR.
Some examples of the application of these results are now presented. \footnote{ The development of some of these results was done by exploiting the symbolic computational power of Mathematica, Mathematica.}
The approach set out in this paper delivers classical results when taken to the simple linear panel data model. Consider the simplest case with two periods of observation and the following model incorporating a conditional mean independence restriction
where $Z_{1}$ and $Z_{2}$ are scalar, $Z\equiv (Z_{1},Z_{2})$, and $z\equiv (z_{1},z_{2})$.\footnote{ The function
can serve as the function $h(Y,Z,U,V)$.}
The $Y^{\ast }$ and $U^{\ast }$ sets are as follows.
Theorem 5 of CR delivers the result that the values of $\beta _{1},$ say $ \beta _{1}^{+}$ in the identified set are all values such that zero is an element of the Aumann expectation of the set $\mathcal{U}^{\ast }(Y,Z;\beta _{1}^{+})$ conditional on $Z=z$ for all $z\in \mathcal{R}_{Z}$. The set $ \mathcal{U}^{\ast }(Y,Z;\beta _{1})$ is singleton in this example, so the Aumann expectation is simply the classical expectation of point-valued random variables and there is
which, set equal to zero, delivers the correspondence
which is point identifying as long as $z_{2}\neq z_{1}$.
Extension to $T>2$ and dynamic models is straightforward and need not be rehearsed here. The point is that the general approach proposed here delivers classical results.
However the approach will not deliver the well-known point identification result in binary response panel data models with logistic independently distributed time-varying latent variables because those models further impose $U \mathbin{\vbox{\baselineskip=0pt\lineskip=0pt \moveright2.5pt\hbox{$\|$} \hrule height 0.2pt width 10pt}} V$.\footnote{ See chamberlain2010binary.} In this paper the covariation of $V$ with all other variables is unrestricted.
This section studies the dynamic binary response model of Example (ref) under a variety of restrictions. Only in the final Section (ref) are models admitting endogenous explanatory variables considered. Section (ref) lists many papers that study binary response panel models with fixed effects. In all but one previous paper known to us there is a restriction on the joint distribution of the fixed effect and other variables such that the conditional distribution of other variables given the fixed effect is subject to restrictions. No such restrictions are imposed here. The one exception of which we are aware is aristodemou2021semiparametric, in which bounds are provided for binary response panel data models with an observed initial condition.
Section (ref) gives results for the two period dynamic binary response model when the initial condition ($Y_{0}$) is observed. This model is studied in aristodemou2021semiparametric. \ Three period dynamic models with unobserved initial condition are studied in Section (ref). Section (ref) gives results for a general case in which there may be endogenous explanatory variables. Extension to models with multiple lagged dependent variables is straightforward.
Define $Y=(Y_{1},\dots ,Y_{T})$ and $Z$ and $U$ similarly.
In the case considered in this section, $T=2$ and $Y_{0}$ is observed. Define $\Delta u\equiv u_{2}-u_{1}$, $\Delta z\equiv z_{2}-z_{1}$, and $ \theta =(\beta ^{\prime },\gamma )^{\prime }$.
The $U^{\ast }$ sets are as follows.
Unions of these $U^{\ast }$ sets do not deliver additional informative inequalities.\footnote{ Unions are either disjoint or equal to the support of $U$ depending on the sign of $\gamma$.}
Under the independence restriction $U \mathbin{\vbox{\baselineskip=0pt\lineskip=0pt \moveright2.5pt\hbox{$\|$} \hrule height 0.2pt width 10pt}} Z|Y_{0}$ the identified set of values of $(\theta ,G_{U|Y_{0}})$ comprises those values such that the following inequalities hold for $y_{0}\in \{0,1\}$ and a.e. $z\in \mathcal{R}_{Z}$.
These are the inequalities of Theorem 1 of aristodemou2021semiparametric. Setting $\gamma =0$ with $U \mathbin{\vbox{\baselineskip=0pt\lineskip=0pt \moveright2.5pt\hbox{$\|$} \hrule height 0.2pt width 10pt}} Z$, dropping conditioning on $Y_0$, delivers the inequalities defining the identified set in the two period static binary response panel model.
Table (ref) shows the $Y^{\ast }$ sets for the two period dynamic binary response panel model. The top half of the table shows the sets for the case in which $Y_{0}$ is observed. The bottom part shows the sets obtained when $Y_{0}$ is not observed.
Appendix (ref) derives sharp bounds on $\theta$ absent any specification of the distribution of $U$ using the method set out in Remark 8 of Section (ref).
For any $s,t \in [T]$ define $\Delta _{st}u\equiv u_{s}-u_{t}$, and $ \Delta _{st}z\equiv z_{s}-z_{t}$. With $T=3$, and treating both $V$ and $ Y_{0}$ as unobserved latent variables with unrestricted distributions the $ U^{\ast }$ sets are as shown in Table (ref).
For sets of values of $Y$, $\mathcal{T\subset R}_{Y}$, define functions
and\footnote{The set $\mathcal{T}$ can be a strict subset of $\mathcal{Y}(\mathcal{T},z;\theta )$. For example, this is the case when $\mathcal{T}$ contains two values of $Y$ and there is a third value of $Y$ such that its $U^{\ast }$ set is a subset of $\mathcal{S}(\mathcal{T},z;\theta )$ as in row 7 of Table (ref).}
The identified set of values of $\left( \beta ,\gamma ,\mathcal{G}_{U|Z=z}\right) $ comprises the values satisfying inequalities of the form
where the sets $\mathcal{Y}(\mathcal{T},z;\theta )$ and $\mathcal{T}$ are shown in the first and second columns of Tables (ref), (ref), and (ref), covering the cases in which $\gamma = 0$, $\gamma >0$, and $\gamma <0$, respectively.
Consider now the general specification of a dynamic panel data model from Example 1, allowing for endogeneity admitting $\alpha \neq 0$. This section illustrates application of our identification analysis to such cases, also allowing for arbitrary finite $T$ .\footnote{Here we impose $\mathcal{R}_C=\mathbb{R}$, as typically done in the literature. Extension to cases in which $\mathcal{R}_C$ is a subset of $ \mathbb{R}$ is straightforward.}
Define
denoting the sets of periods in which $Y_{1t}=0$ and $Y_{1t}=1$, respectively. Let $\mathcal{Y}_{0}$ denote the set of values in which the initial condition $Y_{10}$ is known to lie, with $\mathcal{Y} _{0}=\left\{ Y_{10}\right\} $ if the initial condition is observed and $ \mathcal{Y}_{0}=\{0,1\}$ if the initial condition is not observed.
The set $\mathcal{U}^{\ast }(Y,Z;h)$ defined in ((ref)) in this model can be written
This is so because the constituent inequalities may be equivalently expressed as
where
That $\underline{C}\leq \overline{C}$ for some $Y_{10}\in \mathcal{Y}_{0}$ guarantees there exist values $C\in \left[ \underline{C},\overline{C}\right] $ and $Y_{10}\in \mathcal{Y}_{0}$ such that ((ref)) holds.\footnote{ The $\max$ and $\min$ operators applied to the empty set are defined to be $ -\infty$ and $\infty$, respectively.}
Define $\theta =(\alpha ^{\prime },\beta ^{\prime },\gamma )^{\prime }$. For any panel data model for a binary outcome as in ((ref)) with $ U\sim G_{U}$ independent of $Z$, the identified set of values of $\left( \theta ,G_{U}\right) $ are those pairs satisfying, for an appropriately chosen collection\footnote{ The collection of all unions of $U^{\ast }$ sets, $\mathsf{U}^{\ast }(z;h)$, defined in ((ref)), will suffice. In practice there may be unions in this collection which need not be considered because they deliver redundant inequalities.} of sets $\mathcal{T}$, the inequalities
where the sets $\mathcal{S}(\mathcal{T},z;\theta )$ and $\mathcal{Y}( \mathcal{T},z;\theta )$ are as defined in ((ref)) and ((ref)).
This characterization applies for dynamic models and static models (for which $\gamma =0$ is imposed), models allowing endogenous explanatory variables (for which $\alpha \neq 0$ is permitted), and for arbitrary $T$.
In this section multiple discrete choice panel models are considered. The presence of the fixed effect renders the model incomplete as in the multiple discrete choice analysis of Chesher/Rosen/Smolinski:11MNLIV. In that analysis incompleteness arose due to the inclusion of potentially endogenous explanatory variables. Analysis of a static panel model with $T=2$ periods is considered.
In a three-choice model with two periods there is
where the $J_{dt}$ terms are random utilities with parameters $\theta \equiv (\beta _{1}^{\prime },\beta _{2}^{\prime })^{\prime }$ as follows.
The terms $V_{1}$ and $V_{2}$ are \textquotedblleft fixed effects\textquotedblright\ whose distribution and covariation with other variables is unrestricted.
Section (ref) lists many papers that study multiple discrete panel models with fixed effects. In all studies of multiple discrete choice panel data models known to us there are conditions imposed on the joint distribution of fixed effects and other variables such that the conditional distribution of other variables given the fixed effect is subject to restriction. No such restrictions are imposed here.
The $U^{\ast }$ sets are shown in Table (ref) using notation $\Delta U_d \equiv U_{d2} - U_{d1}$.
The identified set of values of $\left( \theta ,G_{U}\right) $ are those pairs satisfying, for all $z\in \mathcal{R}_{Z}$ the inequalities
where the sets $\mathcal{S}(\mathcal{T},z;\theta )$ and $\mathcal{Y}( \mathcal{T},z;\theta )$ are as defined in ((ref)) and ((ref)) and the sets $\mathcal{T}$ and $\mathcal{Y}(\mathcal{T} ,z;\theta )$ are shown in Table (ref). As we show for ordered choice panels in the following section, this characterization can be generalized to allow arbitrary periods $T$ and alternatives $\left\{1,...,K\right\}$, and can allow for dependence on lagged choices. Endogenous covariates can be permitted as is done for cross sectional multiple discrete choice in Chesher/Rosen/Smolinski:11MNLIV.
This section generalizes the binary response models of Section (ref) to models in which the outcome is an ordered response variable. Section (ref) gives results for a static two period model with three ordered outcomes. Section (ref) then gives results for a general ordered outcome model allowing an arbitrary finite number of ordered outcomes, arbitrary periods, and dynamics.
There are structural equations as follows.
Let $Y=(Y_{1},Y_{2})$, $Z=(Z_{1},Z_{2})$, $U=(U_{1},U_{2})$. Let $\theta =(\beta ^{\prime },c_{1},c_{2})^{\prime }$.\footnote{ In some applications $c_{1}$ and $c_{2}$ can have known values.} There is the restriction $U \mathbin{\vbox{\baselineskip=0pt\lineskip=0pt \moveright2.5pt\hbox{$\|$} \hrule height 0.2pt width 10pt}} Z$. This model is studied in aristodemou2021semiparametric where, as here, no restrictions are placed on the distribution of $V$ or on its covariation with other variables.
Define $\Delta u\equiv u_{2}-u_{1}$ and $\Delta z\equiv z_{2}-z_{1}$. The $ U^{\ast }$ sets are shown in Table (ref). The identified set of values of $(\theta ,G_{U})$ comprises the values satisfying, for $z\in \mathcal{R}_{Z}$, $7$ inequalities of the form
where $\mathcal{Y}$ and $\mathcal{S}$ are given in Table (ref).
Theorem 5 of aristodemou2021semiparametric delivers an outer set using the inequalities 1, 2 and 3 in Table (ref) and the inequalities:
and
which are implied by inequality 4, and
and
which are implied by inequality 5.
Consider now a general specification of an ordered response panel data model with $\mathcal{ R}_Y = \left\{0,...,J\right\}$ and allowing for dynamics as in e.g. honore2021dynamic in which for all $j \in \mathcal{R}_Y$:
where $c_0 \equiv -\infty$, $c_{J+1} \equiv \infty$, and $\imath_t \equiv \left(1\left[Y_{t-1}=0 \right],\ldots,1\left[Y_{t-1}=J \right]\right)$ with each component of $\gamma$ encoding the impact of lagged $Y$ on $Y_t$. \footnote{ It is straightforward to accommodate multiple lags.} Let $\tilde{Z}_t \equiv (Z_t,\imath_t)$, $\tilde{\beta} \equiv (\beta^{\prime}, \gamma^{\prime})^{\prime} $, $Y \equiv (Y_1,...,Y_T)$, $Z \equiv (Z_1,...,Z_T)$, $U\equiv(U_1,...,U_T)$. Let $\theta \equiv (\beta^{\prime},\gamma^{\prime},c_1,...,c_J)^{\prime }$ denote parameters of the structural function, restricted such that $c_1 < \dots < c_J$. The initial condition $Y_0$ is assumed observed, but it is straightforward to accommodate an unobserved initial condition as for the binary panel studied in Section (ref).
Sets $\mathcal{U}^{\ast}\left(Y,Z;h\right)$ are given by
This is verified by noting that for all $u \in \mathcal{U}^{\ast }(Y,Z;h)$ we have that
in turn implying the existence of $v$ such that
For all such $u,v$ it follows that ((ref)) holds for all $t$ with $U=u$ and $V=v$.
When the independence restriction $U \mathbin{\vbox{\baselineskip=0pt\lineskip=0pt \moveright2.5pt\hbox{$\|$} \hrule height 0.2pt width 10pt}} Z$ is imposed, the identified set for $ \left( \theta ,G_{U}\right) $ are those pairs satisfying
for an appropriately chosen collection of sets $\mathcal{T}$ where the sets $ \mathcal{S}(\mathcal{T},z;\theta )$ and $\mathcal{Y}(\mathcal{T},z;\theta )$ are as defined in ((ref)) and ((ref)).\footnote{ Once again the collection of all unions of $U^{\ast }$ sets, $\mathsf{U} ^{\ast }(z;h)$, defined in ((ref)), will suffice, but in practice some of these unions may not be necessary.} This characterization can be generalized to allow for endogenous variables on the right hand side of ((ref)) as done for cross section analysis of ordered choice models in Chesher/Smolinski:12 and Chesher/Rosen/Siddique:23. It is straightforward to allow $G_{U|Z=z}$ to vary with $z$ by replacing $G_U$ with $G_{U|Z=z}$ in the inequality above, which then delivers an identified set for pairs $\left( \theta ,\mathcal{G}_{U|Z}\right) $.
There is the model
with $t\in [T]$ and the independence restriction $(U_1,U_2) \mathbin{\vbox{\baselineskip=0pt\lineskip=0pt \moveright2.5pt\hbox{$\|$} \hrule height 0.2pt width 10pt}} Z\equiv (Z_{1},\dots ,Z_{T})$ where for each $j \in \{1,2\}$, $U_j \equiv (U_{j1},\dots ,U_{jT})$.\footnote{This strong exogeneity restriction can be relaxed.}
This is a simultaneous equations model with binary outcomes such as is found in simultaneous firm entry applications\footnote{ See for example tamer2003incomplete.} and models of social interactions, put into a panel context with \textquotedblleft fixed effects\textquotedblright , constant through time, one for each outcome.
honore2021identification study a restricted version of this model with $\beta _{1}=\beta _{2}$, $\alpha _{1}=\alpha _{2}$ and $U$ and $V$ restricted to be independently distributed. No such restrictions are imposed here.
Define $\theta \equiv (\alpha _{1},\alpha _{2},\beta _{1}^{\prime },\beta _{2}^{\prime })^{\prime }$. The distribution of $V\equiv (V_{1},V_{2})$ and the covariation of $V$ with other variables is unrestricted.
Consider the case with $T=2$ when $Y=(Y_{11},Y_{12},Y_{21},Y_{22})$. Extension to more time periods and outcomes is straightforward.
Define $\Delta u_{1}\equiv u_{12}-u_{11}$, $\Delta u_{2}\equiv u_{22}-u_{21}$ , $\Delta z\equiv z_{2}-z_{1}$. The $U^{\ast }$ sets, $\mathcal{U}^{\ast }(y,z;\theta )$, are as shown in Table (ref).
There are $12$ $U^{\ast }$ sets that are not equal to $\mathcal{R}_{U}$ and $ 4$ pairs of these $U^{\ast }$ sets are identical - for example $\mathcal{U} ^{\ast }(0,0,0,1),z;\theta )=\mathcal{U}^{\ast }(1,1,0,1),z;\theta )$, so there are unions of $8$ $U^{\ast }$ sets to be considered when calculating the identified set, that is $254$ unions in total. Only $24$ of these deliver inequalities that characterize the identified set of parameter values, the remaining unions delivering redundant inequalities.
The configuration of the unions of these $U^{\ast }$ sets depends on the signs of $\alpha _{1}$ and $\alpha _{2}$ and in practice there are likely to be restrictions on these. For example in a simultaneous firm entry application $\alpha _{1}\leq 0$ and $\alpha _{2}\leq 0$ would likely be imposed and in a model of couple's choices of activity (e.g. cinema attendance) $\alpha _{1}\geq 0$ and $\alpha _{2}\geq 0$.
Only the case with $\alpha _{1}\geq 0$ and $\alpha _{2}\geq 0$ is presented here. In this case, among the $U^{\ast }$ sets only the sets $\mathcal{U} ^{\ast }((0,1,0,1),z;\theta )$ and $\mathcal{U}^{\ast }((1,0,1,0),z;\theta )$ have a non-empty intersection.
The identified set of values of $\left( \theta ,G_{U}\right) $ are those pairs satisfying, for all $z\in \mathcal{R}_{Z}$ the inequalities
where the sets $\mathcal{S}(\mathcal{T},z;\theta )$ and $\mathcal{Y}( \mathcal{T},z;\theta )$ are as defined in ((ref)) and ((ref)) and the sets $\mathcal{T}$ and $\mathcal{Y}(\mathcal{T} ,z;\theta )$ are shown in Table (ref).
In a panel model with a censored outcome there is the following.
with $V \equiv (C, Y_{10})$ and $Y_{3t}$ denoting a censoring threshold, such that the outcome variable $Y_{1t}$ takes the value of the index $\alpha Y_{2t}+Z_{t}\beta +\gamma Y_{1t-1}+C+U_{t}$ when it exceeds the censoring threshold, and otherwise takes the value $Y_{3t}$. The censoring indicator $W_t \equiv 1\left[ Y_{1t}=Y_{3t} \right]$ is observed.
As in the models studied in KPT, the censoring threshold $Y_{3t}$ can be endogenous, and it may be correlated with elements of $U$ and $V$.\footnote{It is not necessary for $Y_{3t}$ to be observed in periods without censoring.} As before, endogenous $Y_{2t}$ is permitted in models with $\alpha \neq 0$, as in cross-sectional Tobit models studied in Chesher/Kim/Rosen:23. The set $\mathcal{Y}_{10}$ denotes the feasible set of values for the initial condition $Y_{10}$ given the observed variables.\footnote{So $\mathcal{Y}_{10}$ is the singleton $\{Y_{10}\}$ if the initial condition is observed, and would typically be its entire support if it is not observed. In a static model the analysis applies with $\gamma =0$ and $Y_{10}$ absent.}
Define
Adopting the strategy for obtaining $U^{\ast}$ sets described in remark 5 of Section (ref) we can define
The $U^{\ast }$ sets are then as follows.
where $\Delta_{ts}U \equiv U_{t} - U_{s}$, $\Delta_{ts}Z \equiv Z_{t} - Z_{s}$, $\Delta_{ts}Y_1 \equiv Y_{1t} - Y_{1s}$ and $\Delta_{ts}Y_2 \equiv Y_{2t} - Y_{2s}$.
Following the approach set out in Theorem (ref), the identified set of values of $\left( \theta ,\mathcal{G}_{U|Z}\right) $, where $\theta\equiv\left(\alpha,\beta,\gamma \right)$ are those satisfying the inequalities
for an appropriate selection of sets $\mathcal{T}$, where the sets $\mathcal{S}(\mathcal{T},z;\theta )$ and $\mathcal{Y}(\mathcal{T},z;\theta )$ are defined in ((ref)) and ((ref)). The required selection can be characterized following the same steps taken in the models studied in prior sections.
This paper delivers methods for producing identified sets when models admit unobserved, latent, variables on which no distributional restrictions are placed, opening the way to robust analysis of short panels. Examples found in econometric practice include models incorporating so-called fixed effects and initial conditions. Endogenous explanatory variables are easily accommodated.
The identified sets delivered by the models in this paper that place no restriction on the distribution of latent $V$ will contain the structures identified by more restrictive models if the restrictions of those models are satisfied by the process under study. The analysis set out here will show how sensitive the findings obtained using that more restrictive model are to those additional restrictions. In some cases it may be found that estimation employing a point-identifying model delivers a structure outside an estimator of the identified set obtained using a less restrictive model of the type studied in this paper. Such a finding would suggest the more restrictive model is misspecified. Formal development of such specification tests may be of interest for future research.