EconBase
← Back to paper

The Projection Solution to the Incidental Parameter Problem

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

120,199 characters · 25 sections · 103 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

The Projection Solution to the Incidental Parameter Problem

abstractThis paper introduces a new approach to econometric analysis of nonlinear panel data models when the number of observations per observational unit is small. In such models the presence of variables that are constant within, while varying across, units results in an incidental parameter problem. The approach taken in this paper removes these incidental parameters via projection, which produces a correspondence specifying all combinations of observed variables and within-unit-varying unobserved heterogeneity that are achievable by choice of some value of the unit-specific incidental parameters. With unit-specific variables removed, there is no need for assumptions concerning their joint distribution with other variables. The result is an incomplete model which is typically partially identifying. Identified sets are characterized via moment inequalities using tools of random set theory. Examples of application to static and dynamic models with discrete or continuous outcomes using distribution-free restrictions on within-unit-varying unobserved heterogeneity are presented.

Introduction

In econometric analysis using linear panel data models, unit-specific heterogeneity terms, so-called \textquotedblleft fixed effects\textquotedblright , can be removed by differencing. This enables identification and inference without the need for restrictions on the covariation of fixed effects with observed explanatory variables and other unobserved variables, and absent parametric distributional restrictions on unobservable heterogeneity.

By contrast, in almost all nonlinear panel data models, differencing of observed variables or functions thereof does not remove fixed effects. This may not be an issue when observational units deliver many realizations of outcomes, enabling “Large-$T$” analysis. In such cases well-behaved estimators of the values of fixed effects and the values of parameters common to observational units may be available.

However, when there are few realizations per unit, as is often the case in economic data, and the values of fixed effects are estimated, estimators of common parameters can be very poorly behaved. This is the incidental parameter problem set out in Neyman/Scott:48 and reviewed in Lancaster:00.

With very few exceptions, approaches to this problem in nonlinear panel models restrict the joint dependence of fixed effects and other heterogeneity terms. In many empirical settings such restrictions may not be appropriate. For example, in the context of production function estimation, firm-specific heterogeneity could be related to managerial ability affecting both the level and variability of output. It may thus be desirable to allow the covariation of firm-specific and within-unit-varying heterogeneity, upon which measures of total factor productivity depend, to be unrestricted. The moment restrictions of Arellano/Bond:91 used in the analysis of dynamic linear panel models have this feature, without imposing parametric restrictions on the distribution of unobservable heterogeneity. The approach in this paper enables analysis of nonlinear panel models that similarly leave the covariation of fixed effects with other heterogeneity unrestricted, and is applicable without imposing parametric distributional restrictions.

The approach taken here solves the incidental parameter problem by removing them from the model via projection. The approach is fundamentally different from others, including the functional differencing approach of Bonhomme:12. Functional differencing finds moments that are invariant to the conditional distribution of individual effects given covariates, effectively \textquotedblleft differencing out\textquotedblright\ the conditional distribution of individual effects given covariates. The functional differencing approach is only applicable in models with a parametric specification for the distribution of outcomes conditional on covariates and individual effects. The approach in this paper requires no such restrictions, because it instead projects out the individual effects themselves. With unit-specific heterogeneity removed, restrictions on its joint distributionß with other variables are irrelevant.

This paper's projection approach allows great flexibility in the restrictions on unobservable heterogeneity for which identification analysis can be conducted, including nonparametric specifications of the distribution of within-unit-varying heterogeneity. The paper demonstrates with examples that feature moment conditions as in conventional GMM analysis, independence restrictions, conditional quantile restrictions, and pairwise exchangeability restrictions. The projection approach is not tied to any specific type of distributional restriction and can admit many possibilities beyond the specific cases considered here.

To explain this it helps to bring some notation on board. Consider a panel data model for outcomes $Y\equiv (Y_{1},\dots ,Y_{T})$ with explanatory variables $X\equiv (X_{1},\dots ,X_{T})$.\footnote{It is straightforward to allow $T$ to vary across units. In many panel applications $t$ is an index for time, but $t$ could index a group, family, classroom, etc.} Let $U\equiv (U_{1},\dots ,U_{T})$ denote unobserved heterogeneity varying within units. Let $V$ denote unit-specific unobserved heterogeneity not varying within units. Each element, $Y_{t}$, $X_{t}$, $V$ and $U_{t}$ can be multidimensional.\footnote{Specific to each observational-unit $i$ there is $Y_{it}, X_{it}, V_i, U_{it}$. Observational-unit-specific indices $i$ are omitted to simplify notation.}

Panel data models restrict the functional relationship satisfied by $Y$, $X$, $V$ and $U$, defining sets of feasible values that these variables can simultaneously take. The approach proposed here works with the projection of these sets onto the space of $(Y,X,U)$. This is the set of values of $Y$, $X$, and $U$ that can be achieved by choice of one or more values of $V$. Various restrictions on the joint distribution of $Y$ , $X$, and $U$ can then be considered. With $V$ removed, econometric analysis can proceed with no restrictions on its joint distribution with other variables.

In the linear model the projection approach delivers the set of compatible values of $Y$, $X$ and $U$ as follows.

equation*[equation* omitted — 117 chars of source]

This set is defined by equalities, and the model is complete for differences in the outcome variables. In nonlinear models the set of compatible values of $Y$, $X$ and $U$ is typically defined by inequalities and there may be no nontrivial functions of outcomes for which the post-projection model is complete.

The paper shows how such projections can be used in identification analysis of nonlinear panel models using tools of random set theory, previously employed for identification analysis in Beresteanu/Molchanov/Molinari:09 and Chesher/Rosen:17.\footnote{Knowledge of that theory is not required to apply the results.} A novelty of the analysis here is the application of these tools to models that feature restrictions common in panel contexts but that are inherently absent in cross sectional settings, such as models with dynamics and models with weak exogeneity restrictions. Employing these different types of restrictions for identification analysis using random set theory is new to this paper. Identified sets for model parameters so-obtained are characterized by moment inequalities, enabling estimation and inference using approaches from the recent literature.\footnote{For example, approaches developed in Andrews/Shi:17, Chernozhukov/Chetverikov/Kato:19, Bai/Santos/Shaikh:22, and Marcoux/Russell/Wan:24 can be used for asymptotic inference with uncountably many conditional moment inequalities, see also the survey Shi:25. Characterizations based on a finite number of moment inequalities are amenable to even more approaches, see for example the recent guide Canay/Illanes/Velez:26.}

The paper's main contribution is the projection approach to nonlinear panel models, offering a general approach for removal of incidental parameters which is not tied to any one specific kind of model (e.g. binary response) or distributional restriction. Specific examples considered here show how the analysis delivers several contributions to the nonlinear panel literature, including the following.

enumerate• Strict and weak exogeneity restrictions allowing feedback can both be accommodated. This speaks to the emphasis in Chamberlain:22, Bonhomme/Dano/Graham:23, and Bonhomme:25 on the importance of relaxing strict exogeneity in panel models.\footnote{Chamberlain:22 is a posthumously published version of a 1993 working paper.} Recent developments in panel models with weak exogeneity include extension of the functional differencing approach of Bonhomme:12 to nonlinear models with weak exogeneity in Bonhomme/Dano/Graham:25, and partial identification of functionals of the distribution of heterogeneous individual-specific coefficients in linear panel models in Lee:26. The analysis in this paper contributes by allowing for weak exogeneity without placing any restrictions on the joint distribution of individual effects and $t$-varying heterogeneity, for example through the use of moment restrictions as in Section (ref). • There is flexible treatment of initial conditions in dynamic models, distinct from the random effects treatment in honore2006bounds. Distributional restrictions can be conditional or unconditional on the initial condition. Unlike previous treatments, initial conditions can be unobserved since they are then unit-specific heterogeneity terms that can be removed using this paper's projection approach, as demonstrated in Section (ref). • Dynamic models in which values of discrete outcomes depend upon lagged latent variables can be accommodated. This is useful in models in which continuous unobserved variables are coded into ordered categories as arises, for example, in studies of well-being or health status, as demonstrated in Section (ref). Previous papers on ordered response panel models such as honore2021dynamic have allowed dependence on observable lagged outcomes, but not on lagged latent variables that determine the discrete outcomes.\footnote{As honore2021dynamic note, whether it is more appropriate to model lagged dependence on the discrete outcome or on the continuous latent variable depends on the process being studied; see footnote 2 and Appendix D in that paper. We thank Bo Honor\'{e} for calling this to our attention.}

The paper proceeds as follows. Section (ref) sets out the class of panel models studied and defines a projection of the set of feasible values of $Y$, $X$, $V$ and $U$ onto the space of $Y$, $X$ and $U$. Identified sets of structures are characterized using random level sets of this projection.

Section (ref) presents two examples of panel data models for continuous outcomes and derives identified sets of structural features under restrictions on the correlations amongst $t$-varying heterogeneity and covariates. One example involves a linear model with censored outcomes, covariates, or both. The other example concerns a CES production function in which the elasticity of substitution which appears in the production function in a nonlinear fashion is firm-specific. The moment restrictions lead to characterizations of identified sets in terms of Aumann expectations of random sets, and the use of their support functions to obtain moment inequalities.

Section (ref) presents examples of dynamic panel data models for discrete outcomes and shows how to obtain identified sets for common parameters. This leads to characterizations of identified sets of structural features defined by Artstein's inequalities. There is a novel treatment of unobserved initial conditions and other missing values. A new method for accommodating autoregressive latent indexes in dynamic ordered outcome models is proposed, this by contrast to the commonly employed approaches in which outcomes follow an autoregressive process. Bounds on parameters of a dynamic binary outcome panel model are derived under quantile independence restrictions and under a conditional exchangeability restriction on the distribution of $U$ absent a parametric specification of that distribution. Section (ref) discusses the related literature on binary and ordered panel models, and the approaches used to deal with incidental parameters in those models.

Section (ref) presents numerical illustrations for the CES production function model introduced in Section (ref) under both strict and weak exogeneity restrictions. Section (ref) concludes.

The General Approach

This section first lays out the class of models covered, and then provides a general set identification characterization that will later be specialized to specific models to produce moment inequalities usable for estimation and inference.

Model, Notation, and Sampling Process

Consider a panel model specifying that

equation[equation omitted — 115 chars of source]

for some fixed and finite $T$, where $Y^{t-1}$ denotes the vector of values of $Y_s$ for $s<t$.\footnote{ Static models are accommodated by specifying $f$ invariant with respect to $ Y^{t-1}$. In dynamic models initial conditions, such as $Y_0$ in a model with a one period lag, may either be observable, in which case these are included in $Y^{t-1}$, or unobservable, in which case they are included in $V $.} Models allowing multi-valued functions $f$ are accommodated by (ref), thus allowing endogenous explanatory variables and models admitting multiple equilibria. The function $f$, all of whose arguments may be vectors, is restricted to belong to a set of functions $\mathcal{F}$, which may be parametrically or nonparametrically specified.

The panel models studied here additionally impose restrictions on the joint distribution of $U$ and $X$. To incorporate such restrictions, notation $G_{U|X}\left(\cdot|x \right)$ is used to denote a conditional distribution of $U$ conditional on $X = x$, where for any set $\mathcal{S} \subseteq \mathcal{R}_U$, $G_{U|X}\left(\mathcal{S}|x\right)$ denotes the probability of the event $U \in \mathcal{S}$ given $X = x$. Notation $G_{U|X}\left(\cdot|\cdot\right)$ denotes a collection of conditional distributions for $U$ given $X=x $ across all possible values of $x \in \mathcal{R}_{X}$. Throughout the paper notation $\mathcal{R}_A$ denotes the support of any random vector $A$.

When considering the restrictions imposed by a model, notation $\mathsf{G} _{U|X}$ is used to denote the family of collections $G_{U|X}\left(\cdot|\cdot\right)$ admitted by the model. For instance, if the components of $U$ are restricted to have zero mean conditional on certain components of $X$ then $\mathsf{G}_{U|X}$ contains all such $G_{U|X}\left(\cdot|\cdot\right)$.

A sampling process delivers realizations of $Y$ and $X$ such that their joint distribution, $F_{YX}$, is identified. Realizations of $V$ and $U \equiv \left(U_{1},...,U_{T} \right)$ are not observed. The former is invariant with respect to $t$. Variables $X \equiv \left(X_{1},...X_{T}\right)$ have components whose covariation with $t$-varying unobservable heterogeneity $U$ is restricted. The underlying probability space on which all variables are defined is assumed nonatomic throughout.\footnote{This is a mild technical requirement on the underlying probability space that ensures convexity of the Aumann expectation of the sets $\mathcal{Q}(Y,X;{\Greekmath 0112})$ defined in Section (ref). It is assured to hold whenever $U$ is absolutely continuously distributed with respect to Lebesgue measure. See Beresteanu/Molchanov/Molinari:09 for further discussion.}

Notation $\mathcal{M}$ will be used to denote a set of pairs $m$ of structural functions and collections of conditional distributions $m=(f,G_{U|X}(\cdot|\cdot))$ that satisfy a model's restrictions. The goal of our identification analysis is to determine which, if any, $m \in \mathcal{M}$ are capable of producing a distribution $F_{YX}$ of observable variables.

Identification Analysis

Using the notation laid out above, the restrictions so far described are formally collected in the following restriction.

Restriction Panel Model (PM): Euclidean random vectors $ \left( Y,X,V,U\right) $ are defined on a complete nonatomic probability space $\left( \Omega ,\mathsf{L},\mathbb{P}\right) $ endowed with the Borel sets on $\Omega$ such that ((ref)) holds, where $(f,G_{U|X}(\cdot|\cdot))\in \mathcal{M}$ and $V$ belongs to the set $\mathcal{R}_V$. The distribution of $(Y,X)$ is point identified. $\square $

Equivalent to (ref), $Y \in \mathcal{Y}(U,V,X;f)$ almost surely, where

equation[equation omitted — 255 chars of source]

The following defines the identified set of structures, denoted $\mathcal{I}(\mathcal{M},F_{YX})$, delivered by a model $\mathcal{M}$ and distribution $F_{YX}$.

definitionUnder Restriction PM the identified set of pairs of structural functions $f$ and conditional distributions of unobservable heterogeneity $G_{U|X}(\cdot|\cdot)$ is \begin{multline} \mathcal{I}(\mathcal{M},F_{YX})\equiv \left\{ \left(f,G_{U|X}(\cdot|\cdot)\right)\in \mathcal{M}:F_{Y|X}(\cdot|x)\preceq \mathcal{Y}(\tilde{U},\tilde{V},X;f)\right. \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ for some } (\tilde{U}, \tilde{V}) \\ \left. \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{where } \tilde{U} \sim G_{U|X}(\cdot|x) \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ conditional on } X=x\quad \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{a.e.}\quad x\in \mathcal{R}_{X}\right\}\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{.} \end{multline}

For any random vector $A$ with distribution $F_{A}$ and random set $\mathcal{A}$, $F_{A}\preceq \mathcal{A}$ denotes that $F_{A}$ is selectionable with respect to the distribution of $\mathcal{A}$.\footnote{The probability distribution of random variable $A$ is selectionable with respect to the probability distribution of random set $\mathcal{A}$ when there exists (i) $\tilde{A}$ having the same distribution as $A$, and (ii) $\widetilde{\mathcal{A}}$ having the same distribution as $\mathcal{A}$, both defined on the same probability space, such that $\mathbb{P}[\tilde{A}\in \widetilde{\mathcal{A}}]=1$. See Molchanov/Molinari:18 Chapter 2 or Definition 2 of Chesher/Rosen:Handbook2020. } The definition states that the identified set comprises those $\left(f,G_{U|X}(\cdot|\cdot)\right)$ pairs in $\mathcal{M}$ for which there exist random vectors $(\tilde{U},\tilde{V})$ with $\tilde{U}$ following the distributions of $G_{U|X}(\cdot|\cdot)$ such that the set of outcomes produced by the structural function $f$, namely $\mathcal{Y}(\tilde{U},\tilde{V},X;f)$ contains a random vector whose conditional distributions given $X$ match those of the observed conditional distributions $\{F_{Y|X}(\cdot|x):x\in \mathcal{R}_X\}$ almost surely. Then, and only then, the pair $\left(f,G_{U|X}(\cdot| \cdot)\right)$ are capable of producing the distribution $F_{YX}$.

Definition (ref) follows Chesher/Rosen:Handbook2020, here incorporating both types of unobservables $U$ and $V$, but it is not directly usable for estimation and inference. Making it usable requires two steps, as follows. First, it is shown that the unrestricted individual effects can be removed without loss of identifying power. Second, results from random set theory are used to yield characterizations that take the form of moment inequalities that can be used as a basis for estimation and inference.

Removing Individual Effects

Individual effects are removed from the model by use of the projection

equation[equation omitted — 402 chars of source]

which is the set of values of observable variables and $t$-varying unobservable heterogeneity that are mutually compatible when the structural function is $f$.

Level sets of the projection $\mathcal{R}_{YXU}(f)$ obtained by fixing a subset of the elements $(Y,X,U)$ are used for identification analysis. For any $(u,x,f) \in \mathcal{R}_{U} \times \mathcal{R}_{X} \times \mathcal{F}$

equation[equation omitted — 206 chars of source]

is the set of possible values for $y$ that can occur when $X=x$ and $U=u$ for some realization of $V$. For any $(y,x,f) \in \mathcal{R}_Y \times \mathcal{R}_X \times \mathcal{F}$

equation[equation omitted — 214 chars of source]

is the set of possible values for $u$ that can occur when $Y=y$ and $X=x$ for some realization of $V$. These sets are dual to each other in that for any $u,x,y$, and $f$:

equation[equation omitted — 182 chars of source]

There is the following Proposition.

propositionLet Restriction PM hold. Then \begin{multline} \mathcal{I}(\mathcal{M},F_{YX})= \left\{ \left(f,G_{U|X}(\cdot|\cdot)\right)\in \mathcal{M}:F_{Y|X}(\cdot|x)\preceq \mathcal{Y}(U,X;f)\right. \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{where } U \sim G_{U|X}(\cdot|x) \\ \left. \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ conditional on }X=x\quad \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{a.e.}\quad x\in \mathcal{R} _{X}\right\}\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{,} \end{multline} and equivalently, \begin{multline} \mathcal{I}(\mathcal{M},F_{YX})= \left\{ \left(f,G_{U|X}(\cdot|\cdot)\right)\in \mathcal{M}:G_{U|X}(\cdot|x)\preceq \mathcal{U}(Y,X;f)\right. \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{where } Y \sim F_{Y|X}(\cdot|x) \\ \left. \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ conditional on }X=x\quad \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{a.e.}\quad x\in \mathcal{R} _{X}\right\}\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{.} \end{multline}

Proposition (ref) provides high-level characterizations of the identified set of $\left(f,G_{U|X}(\cdot|\cdot)\right)$ from which unobservable individual effects $V$ are absent. The random sets $\mathcal{Y}(U,X;f)$ and $\mathcal{U}(Y,X;f)$ comprise, respectively, the set of random variables $Y$ compatible with structural function $f$ and $(U,X)$, and the set of random variables $U$ possible given knowledge of observable variables $(Y,X)$ under structural function $f$.\footnote{The selectionability statement requiring $G_{U|X}(\cdot|x)$ to be selectionable with respect to the distribution of $\mathcal{U}(Y,X;f)$ conditional on $X$ almost surely in (ref) in Proposition (ref) is equivalent to requiring that $(U,X)$ is selectionable with respect to the distribution of the random set $ \mathcal{U}(Y,X;f) \times \{X\}$ by Proposition 1 in Appendix B of Chesher/Rosen:15. An analagous statement holds regarding selectionability of $F_{Y|X}(\cdot|x)$ with respect to the distribution of $\mathcal{Y}(U,X;f)$ almost surely in (ref), and selectionability of $(Y,X)$ with respect to the distribution of $ \mathcal{Y}(U,X;f) \times \{X\}$. }

When the function $f$ is restricted to a parametric family indexed by a parameter, say ${\Greekmath 0112}$, such that $\mathcal{F} = \left\{f_{{\Greekmath 0112}}:{\Greekmath 0112} \in \Theta \right\}$, then the set-valued mappings defined in ((ref)) and ((ref)) will be indexed by ${\Greekmath 0112}$ rather than $f$, with the associated random sets denoted $\mathcal{Y}\left(U,X;{\Greekmath 0112} \right)$ and $\mathcal{U}\left(Y,X;{\Greekmath 0112} \right)$.

Characterization via Moment Inequalities

Proposition (ref) provides high-level, generally applicable characterizations of the identified set for $\left(f,G_{U|X}(\cdot|\cdot)\right)$ in which no distributional restrictions are placed on unobservable individual effects $V$.\footnote{Characterization of the identified set for $f$ or any functional of $\left(f,G_{U|X}(\cdot|\cdot)\right)$ then follows.} To make these characterizations practicable requires necessary and sufficient conditions for the stated selectionability properties usable for estimation and inference. The recent literature employing random set theory for identification analysis has made use of such conditions, which can be employed here as well.\footnote{For further details and alternative conditions for guaranteeing this selectionability property see Chapter 2 of Molchanov/Molinari:18.} One approach uses Artstein's inequality under the following mild restriction.

Restriction RCS: For all $\left(f,G_{U|X}(\cdot|\cdot)\right) \in \mathcal{M}$, $G_{U|X}\left( \mathsf{cl}\left(\mathcal{U}(y,x;f)\right) \setminus \mathcal{U}(y,x;f) |x\right) = 0$ $a.e.\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ } (y,x)$, where $\mathsf{cl}(\cdot)$ and $\cdot\setminus \cdot$ denote the closure of a set, and the set difference of two sets, respectively. $\square$

Restriction RCS holds automatically if $\mathcal{U}(Y,X;f)$ is closed almost surely. It also applies when that is not so, but the difference between $\mathcal{U}(Y,X;f)$ and its closure is measure zero almost surely.\footnote{This is a common occurrence in models in which inequalities determine the value of a limited dependent variable, for example when a discrete outcome is determined by whether a continuously distributed unobservable variable exceeds a threshold.} In this case the restriction enables characterization of identified sets by applying results from random set theory to the closure of $\mathcal{U}(Y,X;f)$, which is useful since selectionability criteria from random set theory are often stated for random closed sets. Lemma 1 in the Appendix provides the formal statement.

Additional Notation

For any random vector $A$ or realized value $a$ let $A^{\Delta}_{st} \equiv A_s - A_t $ and $a^{\Delta}_{st} \equiv a_s - a_t$. For any integer $k>0$, notation $\mathbf{0}_{k}$ denotes a zero vector of length $k$. In the definition of any set, such as $\mathcal{U}(y,x;f)$ in ((ref)) above, the support of a variable is omitted when it is clear from context. Notation expressing suprema and infima of conditional probabilities or expectations with respect to the conditioning variables are to be understood as essential suprema and infima, respectively. For any real number $c$, $c^{-} \equiv \lvert \min\{c,0\} \rvert $ and $c^{+}\equiv \max\{c,0\}$ denote the negative and positive part of $c$, respectively. For random vectors $A$ and $B$, $A \mathbin{\vbox{\baselineskip=0pt\lineskip=0pt \moveright2.5pt\hbox{$\|$} \hrule height 0.2pt width 10pt}} B$ signifies that $A$ and $B$ are stochastically independent. For any vectors $a$ and $b$, $a \cdot b$ denotes their dot product.

Models with continuous outcomes

This section considers two models with essential nonlinearity and continuous outcomes. It is shown how the projection approach leads to identification results and so to estimation and inference absent restrictions involving the distribution of unit-specific variables and absent a parametric specification of the distribution of within-unit-varying unobservables. Identification results are obtained under moment restrictions involving within-unit-varying unobservables and functions of covariates.

In Sections (ref) and (ref) the results of projection are derived. In Section (ref) identification analysis using these projections is provided.

CES Production Function Panel Models

This is an example of a nonlinear panel model with continuous outcomes, inspired by Example 3 of Bonhomme:12. Let log output $Y_{t}$ of an individual unit (e.g. a plant or firm) at time $t$ be generated by a constant elasticity of substitution (CES) production function such that

equation[equation omitted — 266 chars of source]

where $V = (S,C)$ is a pair of unit-specific unobservable variables, $C \in \mathbb{R}$, and

equation*[equation* omitted — 272 chars of source]

is the CES function with substitution parameter $s$.\footnote{So $\mathcal{R}_S$ is the subset of the extended real line on which all elements are no greater than one.}

Variables $X_{t}\equiv \left( L_{t},K_{t}\right) $ denote labor and capital inputs at time $t$, $U_{t}$ denotes $t$-varying unobservable variables, and ${\Greekmath 0112} =({\Greekmath 010C},{\Greekmath 010D})$ are common parameters with ${\Greekmath 010C} > 0$.

Thus (ref) holds for $f = f_{\Greekmath 0112}$ with

equation[equation omitted — 195 chars of source]

This specification is considered briefly in Bonhomme:12 as an example of nonlinear regression, with a unit-specific effect ($S$) entering nonlinearly.\footnote{Equation (ref) appears under Example 3 as equation (6) in Bonhomme:12, in the notation of that paper using the symbol ${\Greekmath 011B} $ where we use $S$. The discussion here uses a simplified version in which there is no low-skilled labor input.} In that paper $U\equiv (U_{1},\dots ,U_{T})$ is restricted Gaussian independent of $(X,V)$, where $V$ is denoted by ${\Greekmath 010B}$. It is stated on page 1344, \textquotedblleft Due to the nonlinearity, it does not seem possible to difference out ${\Greekmath 010B} $ in a straightforward way.\textquotedblright

The projection approach can be applied here, as is now illustrated. The Gaussian restriction proposed in Bonhomme:12 is not required.

The set of values of $(Y,X,U)$ delivered by the model for some value of $V$ is obtained as follows.

The function $g(l,k,{\Greekmath 010D} ,s)$ is a generalized mean, monotone increasing in $s\in [-\infty ,1]$, bounded with

equation*[equation* omitted — 227 chars of source]

With $x_t \equiv (x_{t1},x_{t2}) = (l_t,k_t)$ and $v = (s,c)$ define for each $t$:

equation[equation omitted — 208 chars of source]

which is monotone decreasing in each component of $v$.

The set of feasible combinations of $(y,x,u)$ obtained by projection across $v$ is

equation[equation omitted — 416 chars of source]

The $U$-level set of this projection for any realizations $y,x$ is

equation[equation omitted — 416 chars of source]

If $V$ were observable then realization of $V=v$ would reveal the realization of $U$ as the singleton value with components $m_t(y,x,{\Greekmath 0112},v)$, for each $t \in [T]$. The set $\mathcal{U}\left(y,x,{\Greekmath 0112}\right)$ is a manifold on $\mathbb{R}^T$ comprising the set of such vectors compatible with some value of unobservable $V$.

For ease of illustration consider the case in which $T=3$. Manifolds $\mathcal{U}(y,x;{\Greekmath 0112})$ are illustrated for an example in which $x=(l_t,k_t)$, $t=1,2,3$ are set according to:

equation*[equation* omitted — 164 chars of source]

and parameters are set at ${\Greekmath 010D}=0.6$ and ${\Greekmath 010C}=1$. The left panel of Figure (ref) depicts eight manifolds $\mathcal{U}(y,x;{\Greekmath 0112})$, one each for values of $y$ with $y^{\Delta}_{21}$ and $y^{\Delta}_{31}$ taking values shown in green in the right panel.\footnote{Because the individual effect $C$ enters additively, each manifold $\mathcal{U}\left(y,x,{\Greekmath 0112}\right)$ is represented in the space of differences $(u^{\Delta}_{21},u^{\Delta}_{31})$ and is fully determined by the values of $y^{\Delta}_{21}$ and $y^{\Delta}_{31}$.} Figure (ref) shows the set $A({\Greekmath 0112})$ in blue in the left hand panel comprising the union of the sets $\mathcal{U}(y,x;{\Greekmath 0112} )$ obtained as $y$ takes values such that $y^{\Delta}_{21}$ and $y^{\Delta}_{31}$ belong to the region $B$ shaded in green in the right hand panel. Because $(Y^{\Delta}_{21},Y^{\Delta}_{31})\in B$ implies that $(U^{\Delta}_{21},U^{\Delta}_{31})\in A({\Greekmath 0112})$, there is for all $x \in \mathcal{R}_X$ the inequality

equation*[equation* omitted — 218 chars of source]

which restricts the values of ${\Greekmath 0112} $ compatible with $F_{YX}$ because of the dependence of $A({\Greekmath 0112})$ on ${\Greekmath 0112}$. Implications of this sort also arise by use of Artstein's inequality. If $U$ and $X$ are stochastically independent the inequality above becomes

equation*[equation* omitted — 240 chars of source]

Section (ref) characterizes identified sets for the common parameters obtained in this model under moment restrictions on the product of elements of $U$ and functions or components of $X$. First, the following section presents a second example of a panel model with continuous outcomes.

figure[figure omitted — 1,173 chars of source]

Linear Panels with Censored Outcomes and Covariates

This section provides results for linear panel models when data are interval censored.\footnote{Analysis of panel models with censored outcomes has been studied in e.g. Honor\'{e} (1992, 1993),\nocite{Honore:92} \nocite{Honore:93} Hu:2002, Khan/Ponomareva/Tamer:16, and Abrevaya/Muris:20.}

The model specifies

equation[equation omitted — 131 chars of source]

where ${\Greekmath 0112}$ is a $k_x \times 1$ vector of common parameters, each $X_{t}^{\ast }$ is a $1\times k_{x}$ vector, and $V=C$. The unobserved variables are $Y^{\ast }$, $X^{\ast }$, $C$, and $U$. The observed variables are

equation[equation omitted — 236 chars of source]

where

equation[equation omitted — 229 chars of source]

and inequalities hold element-wise. There may be interval censoring of components of one or both of the outcome $Y^{\ast }$ and covariates $X^{\ast }$. Components that are not censored have $Y_{t}^{L}=Y_{t}^{H}$ for outcomes and $X_{tj}^{L}=X_{tj}^{H}$ for any uncensored components $X_{tj}^{\ast }$ of $X_{t}^{\ast }$. Missing data can be captured by having both lower and upper limits correspond to the end points of the support of the corresponding variables.

Define

align*[align* omitted — 375 chars of source]

There is, for all $s$ and $t$ in $[T]$:

align*[align* omitted — 331 chars of source]

Adding the two inequalities yields the projection of the model-admitted set of values of $(Y,X,C,U)$ onto the space of $(Y,X,U)$ as follows.

equation[equation omitted — 282 chars of source]

Recall that, absent censoring, the linear model by contrast delivers the projected set

equation[equation omitted — 172 chars of source]

in which there are equalities, whereas with censoring there are inequalities as in the CES production function case.

With censoring the level set of $U$-values that deliver $Y=y$ when $X=x$ is simply the slice through the projection defined in ((ref)) obtained fixing $(y,x)$ accordingly:

equation*[equation* omitted — 233 chars of source]

This characterization of $\mathcal{U}(Y,X;{\Greekmath 0112} )$ provides a starting point for identification analysis using various restrictions on the joint distribution of $(X,U)$. When covariates are censored, consideration of context may lead one to prefer restrictions on the joint distribution of $(X^{\ast },U)$, as considered in the following subsection.

Identified sets

Identification analysis for common parameters is now presented for the two models just considered under moment restrictions on within-unit-varying unobservables and covariates. A characterization for the CES model parameters under a conditional mean restriction is then provided. The methods employed can be applied to more general forms of moment conditions than the ones considered here.

Moment Restrictions

Consider the following restriction.

\noindentRestriction M. For all $t \in [T]$, $E\left[Z_t U_t \right]=\mathbf{0}_{J_t}$, where each $Z_t = Z_t(X)$ is a vector-valued function of $X$ of dimension $J_t$.

Weak exogeneity and strict exogeneity of components of $X$ can be accommodated by appropriate definition of each $Z_t$ as will be shown below. Flexibly defining $Z_t$ can further be used to specify that some covariates are strictly exogenous while others are only weakly exogenous.

Since Restriction M requires that $E\left[Z_t U_t \right]=\mathbf{0}_{J_t}$ for all $t$, it is useful to define the level set of possible values of $\left(\left(Z_1 U_1\right),...,\left(Z_T U_T\right)\right)$ obtained from the projection $\mathcal{R}_{YXU}({\Greekmath 0112})$ defined in (ref), making use of the $U$-level sets defined in (ref):

equation[equation omitted — 234 chars of source]

Thus $\mathcal{Q}(y,x;f)$ is a set of vectors, each element of which is a $J \equiv J_1 + \dots +J_T$ dimensional vector whose entries correspond to feasible values of components of $Z_t U_t$ under structural function $f$ across all $t \in[T]$, when $Y=y$ and $X=x$.

Replacing fixed arguments $(y,x)$ with $(Y,X)$ in (ref) yields $\mathcal{Q}(Y,X;f)$, a random set whose distribution is determined by that of $(Y,X)$. Under Restriction M, the panel model specification (ref) with structural function $f$ can produce the distribution of $(Y,X)$ if and only if there is a measurable selection of $\mathcal{Q}(Y,X;f)$, defined below, whose expected value is the zero vector $\mathbf{0}_J$.

definitionLet $\mathcal{Q}$ be a random closed set on $\left( \Omega ,\mathsf{L},\mathbb{P}\right)$ whose realizations are subsets of $\mathbb{R}^J$. A random vector $Q$ measurable on $\left( \Omega ,\mathsf{L},\mathbb{P}\right)$ is a measurable selection of $\mathcal{Q}$ if $Q({\Greekmath 0121})\in \mathcal{Q}({\Greekmath 0121}) $ for almost all ${\Greekmath 0121} \in \Omega$.

Let $\mathbf{L}^{1}(\mathcal{Q})$ denote the set of all integrable measurable selections of a random set $\mathcal{Q}$. The Aumann integral and Aumann expectation of $\mathcal{Q}$ are defined, respectively, as

equation*[equation* omitted — 339 chars of source]

Define the sets

equation[equation omitted — 374 chars of source]

Under the restrictions of Proposition (ref), the set $\mathcal{F}_I$ is the identified set for structural function $f$. The set $\overline{\mathcal{F}_I}$ is a superset of $\mathcal{F}_I$, and therefore provides bounds on $f$.

The set $\overline{\mathcal{F}_I}$ is the moment-closure of the identified set in the terminology of Li:26, and it is useful for estimation and inference for several reasons. First, the set $\overline{\mathcal{F}_I}$ can coincide with $ \mathcal{F}_I$ and hence be sharp, for example when the Aumann integral is closed, in which case the Aumann integral and Aumann expectation coincide. Such conditions hold under a mild restriction for the CES example of Section (ref), as shown in Proposition (ref) in the Appendix. Second, characterization of $\overline{\mathcal{F}_I}$ by way of the Aumann expectation is equivalent to a support function characterization which has the form of moment inequalities that only involve expectations of random variables rather than random sets. Third, there are conditions under which the identified set and its moment closure are statistically indistinguishable even if they do not coincide, as shown in Li:26.

The usefulness of the support function characterization for the moment closure follows from two consequences of the restrictions of Proposition (ref) below. First, there is the equivalence

equation[equation omitted — 273 chars of source]

where $\mathcal{B}^J \equiv \left\{r\in \mathbb{R}^J: \lVert r \rVert = 1 \right\}$ is the boundary of the unit ball in $\mathbb{R}^J$ and

equation*[equation* omitted — 150 chars of source]

denotes the support function of any set $\mathcal{Q}\subseteq \mathbb{R}^J$ evaluated at $r\in\mathbb{R}^J$.\footnote{This follows from Theorem 2.1.26 of Molchanov:17 because the underlying probability space is nonatomic under Restriction PM and $\mathcal{Q}(Y,X;{\Greekmath 0112})$ is integrable under Restriction M, so $\mathbb{E}\left[\mathcal{Q}(Y,X;{\Greekmath 0112})\right]$ is convex.} Second, the order of the expectation and support function can be swapped yielding

equation*[equation* omitted — 195 chars of source]

where on the right there is the expectation of a random variable.\footnote{This follows from Theorem 2.1.35 of Molchanov:17 because the underlying probability space is nonatomic.} Putting all this together, there is the following Proposition regarding $\mathcal{F}_I$ and $\overline{\mathcal{F}_I}$ defined in (ref).

propositionSuppose that Restrictions PM and M hold and that $\mathcal{Q}(Y,X;f)$ is closed almost surely. Then $\mathcal{F}_I$ is the identified set for the structural function $f$. Moreover, the moment closure $\overline{\mathcal{F}_I}$ of $\mathcal{F}_I$ comprises bounds on $f$ and admits the support function representation \begin{equation} \overline{\mathcal{F}_I} = \left\{f \in \mathcal{F}:\min_{r\in\mathcal{B}^J}E\left[h(\mathcal{Q}(Y,X;f),r) \right] \geq 0\right\}\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{.} \end{equation} If $\mathbb{E}_I\left[\mathcal{Q}(Y,X;f)\right]$ is closed then $\overline{\mathcal{F}_I} = \mathcal{F}_I$.

Proposition (ref) states that $\mathcal{F}_I$ is the identified set for $f$, and that its moment closure $\overline{\mathcal{F}_I}$ is characterized by the moment inequalities $E\left[h(\mathcal{Q}(Y,X;f),r) \right] \geq 0$, for all $r\in \mathbb{R}^J$. The reasoning is similar to that of Theorem 4.1 of Beresteanu/Molchanov/Molinari:09 which established sharp bounds on the best linear predictor with censored outcomes and covariates. The difference is that the analysis here applies to a nonlinear panel model with individual effects, rather than a model for cross section data, so the random set $\mathcal{Q}(Y,X;f)$ is constructed by taking the product of components of $Z$ as specified by Restriction M with the $U$-level set obtained from $ \mathcal{R}_{YXU}(f)$ after removing $V$ from the model by projection.\footnote{The projection step renders an additional integrably boundedness condition used in Theorem 4.1 of Beresteanu/Molchanov/Molinari:09 inapplicable without further restrictions in the present setting, necessitating the distinction between the identified set and its moment closure.}

A connection to the support function approach of Beresteanu/Molchanov/Molinari:09 is also made in Lee:26 which, in contrast to the projection approach, uses moments that restrict the joint distribution of individual effects with other heterogeneity terms. That paper uses duality theory for infinite dimensional programs to characterize the identified set of features of the distribution of random coefficients in linear panel models. It also shows that if the target parameter is a common parameter, the characterization is equivalent to that obtained by the support function approach used in Beresteanu/Molchanov/Molinari:09 Theorem 4.1. Thus, subject to regularity conditions, duality theory for infinite dimensional programs can also be used with the moment conditions of this paper for identification analysis. This relationship also highlights the possibility of using the formulation of Schennach:14 for developing estimation and inference approaches as a potential alternative to the moment inequality approach, an avenue which is left to future research.

Identification in the CES Model

In the CES model the class $\mathcal{F}$ is parameterized by ${\Greekmath 0112} = ({\Greekmath 010C},{\Greekmath 010D})$. Focus is given to the moment closure of the identified set for ${\Greekmath 0112}$, denoted $\Theta^{\ast}$, and the resulting moment inequalities. The notation replaces $f$ with ${\Greekmath 0112}$ accordingly. In the CES model

equation[equation omitted — 324 chars of source]

where $m_t(y,x,{\Greekmath 0112},v)$ defined in (ref) denotes for each $t$ the unique value of $u_t$ given fixed values of $(y,x,{\Greekmath 0112},v)$.

The characterization of Proposition (ref) is specialized under Restriction M with two different specifications for $Z$ as follows.

\noindentRestriction MS. Restriction M holds with $Z_t = (1,X_1,...,X_T)$ for all $t$.\\ \noindentRestriction MW. Restriction M holds with $Z_t = (1,X_1,...,X_t)$ for all $t$.

Consider first Restriction MS which requires that $E\left[X_{hk}U_t \right] = 0$ for all $h,t \in [T]$ and all $k=1,2$ where $X_{h1} = L_h$ and $X_{h2} = K_h$. This is a strict exogeneity restriction comprising $J=T(2T+1)$ moment restrictions. The expectation of the support function of $\mathcal{Q}(Y,X;{\Greekmath 0112})$ given by (ref) under the CES specification simplifies as

equation*[equation* omitted — 216 chars of source]

where for each $t\in[T]$,

equation*[equation* omitted — 149 chars of source]

and $r$ is a list of $J = T(2T+1)$ numbers $r_t,r_{thk}$ for all $t,h,k$.

Using the definition of $m_t$ in (ref) and simplifying, it follows from Proposition (ref) that under Restriction MS the set $\Theta^{\ast}$ is the set of ${\Greekmath 0112} = ({\Greekmath 010C},{\Greekmath 010D})$ that satisfy

equation[equation omitted — 280 chars of source]

Now suppose instead that only weak exogeneity is asserted, such that Restriction MW is imposed. Working through the same steps under this weaker restriction, Proposition (ref) delivers $\Theta^{\ast}$ as those parameter vectors satisfying the same inequalities (ref), but now with $r$ additionally restricted to satisfy $r_{thk}=0$ for all $t<h$. Minimization over this restricted set imposes the zero moment restrictions $E\left[ X_{hk} U_t \right]=0$ only for $t \geq h$.

Section (ref) demonstrates that $\Theta^{\ast}$ can produce informative sets by way of numerical illustrations under both weak and strict exogeneity restrictions. In that section further simplification of the inequalities (ref) is provided for that purpose.

Moment restrictions on other functions of covariates and components of $U$ may similarly be imposed through Restriction M beyond the two specific cases of Restriction MS and Restriction MW considered here. Moment conditions incorporating instrumental variables can be used by specifying components of $Z_t$ as functions of components of $X$ with respect to which the structural function is restricted to be invariant, for example by way of exclusion restrictions.

Identification in the Censored Linear Panel Model

Moment restrictions can also be used in the censored linear panel model described in Section (ref). Focus is again given to the moment closure of the identified set for ${\Greekmath 0112}$, denoted $\Theta^{\ast}$, and the resulting moment inequalities. In this model it may be desirable to invoke moment restrictions involving functions of censored covariate values $X^{\ast}$ rather than the observed endpoints of intervals on which $X^{\ast}$ is realized, which allows censoring to be endogenous.\footnote{If moment restrictions are made solely with respect to $X$, analysis following the steps of the previous section incorporating the linear panel specification applies directly, so is not repeated here.} Thus the following restriction is considered.\\ \noindentRestriction M$^{\ast}$: For all $t \in [T]$, $E\left[Z_t U_t \right]=\mathbf{0}_{J_t}$, where each $Z_t = Z_t(X^{\ast})$ is a vector-valued function of $X^{\ast}$ of dimension $J_t$.

The level set $\mathcal{Q}(y,x;{\Greekmath 0112})$ of possible values of $\bigl(\left(Z_1 U_1\right),...,\left(Z_T U_T\right)\bigr)$ obtained from the censored linear panel projection $\mathcal{R}_{YXU}({\Greekmath 0112})$ given in (ref) under Restriction $M^{\ast}$ is

equation*[equation* omitted — 298 chars of source]

where $\mathcal{D}(y,x)$ denotes the set of possible values of the censored variables given $(y,x)$:

equation*[equation* omitted — 319 chars of source]

Following the same reasoning used when considering Restriction M and the CES model the moment closure of the identified set, $\Theta^{\ast}$, comprises ${\Greekmath 0112}$ such that $\mathbf{0}_J \in \mathbb{E}\left[\mathcal{Q}(Y,X;{\Greekmath 0112})\right]$. Use of the support function yields an equivalent characterization via moment inequalities:

equation*[equation* omitted — 431 chars of source]

where each $\mathpalette\overrightarrow@{r_t}$ is a vector of length $J_t$ such that $r = (\mathpalette\overrightarrow@{r_1},...,\mathpalette\overrightarrow@{r_T})$.

As was the case for Restriction M, Restriction M$^{\ast}$ can accommodate different moment restrictions through specification of $Z_t$. For example, $Z_t = (1,X^{\ast}_1,...,X^{\ast}_T)$ and $Z_t= (1,X^{\ast}_1,...,X^{\ast}_t)$ for strict and weak exogeneity restrictions, respectively.

Conditional Moment Restrictions

Consider the following conditional moment restriction, which provides a strict exogeneity restriction stronger than that of Restriction MS.

\noindentRestriction CMS: For all $t \in [T]$, $E\left[U_t|X \right] = 0$.

Restriction CMS implies that $\mathbf{0}_T \in \mathbb{E}\left[\mathcal{U}(Y,X;{\Greekmath 0112}) |X\right]$ almost surely, i.e. that the zero vector is an element of the conditional Aumann Expectation of the $U$-level set, the form of which for the CES model is given in (ref).

Using the support function approach in the CES model this is equivalently that

equation*[equation* omitted — 290 chars of source]

where $r=(r_1,...,r_T)$. This can further be expressed

equation*[equation* omitted — 330 chars of source]

where ${\Greekmath 0116}_t(x) \equiv E\left[Y_t|X \right]$. From this it follows that for any ${\Greekmath 0112}$ only values of $r_1,...,r_T$ that sum to zero can provide the minimum over $r$, and the inequalities simplify to

equation*[equation* omitted — 320 chars of source]

So, as in the cases considered with unconditional moment restrictions, the bounds are characterized by an infinite collection of moment inequalities, in this case infinitely many conditional moment inequalities. As noted in the introduction estimation and inference methods from the recent literature are available, as discussed for example in the survey Shi:25.

Models with discrete outcomes

This section gives examples of application of the projection approach in static and dynamic panel models with discrete outcomes. A dynamic binary outcome model is studied in Section (ref); a dynamic ordered response model is studied in Section (ref).

It is shown how, in models in which ordered discrete outcomes encode the values of continuous latent variables, autoregressive dependence in the latent continuous variables can be accommodated. This approach can also be employed in other discrete response models, such as multinomial choice.

In dynamic models attention must be paid to initial values of outcomes if they are not observed. Following the approach of this paper, when they are not observed they are treated as unit-specific unobserved variables and are removed by projection. This is illustrated in the models considered in both Sections (ref) and (ref).

Section (ref) provides characterizations of identified sets for the binary and ordered response models of Sections (ref) and (ref). Moment-based restrictions such as those considered in continuous outcome models in Section (ref) can be uninformative in discrete outcome models; see Manski:88. So here attention is turned to the identifying power of stochastic independence restrictions. Section (ref) considers further options for conducting identification analysis absent a parametric specification of the distribution of $t$-varying unobservables, such as quantile and exchangeability restrictions.

Section (ref) provides discussion of the related literature on binary and ordered outcome panel models. In contrast to the approach here, nearly all such models in the literature restrict the joint distribution of unit-specific effects and within-unit-varying heterogeneity.\footnote{The analysis in aristodemou2021semiparametric is the sole exception of which we are aware.} To our knowledge, no previous models allow for unobserved initial conditions or for dependence on lagged latent variables.

A Dynamic Binary Outcome Panel Model

Consider the following binary outcome panel specification.

equation[equation omitted — 221 chars of source]

with parameter vector ${\Greekmath 0112} =({\Greekmath 010D} ,{\Greekmath 010C} ^{\prime })$, $\mathcal{R}_C = \mathbb{R}$, and $U$ continuously distributed with full support on $\mathbb{R}^T$ conditional on $X$.\footnote{This implies that Restriction RCS holds and there is no loss in using weak inequalities throughout in expressions for the sets $\mathcal{R}_{YXU}({\Greekmath 0112})$ and $\mathcal{U}(y,x;{\Greekmath 0112})$.} Higher order lags are easily accommodated.

In a dynamic model in which ${\Greekmath 010D} $ may be nonzero, there is an initial condition to be considered. If the value of $Y_{0}$ is observable, then the $t$-invariant individual unobservable is $V=C$, while if $Y_{0}$ is not observable $V=(C,Y_{0})$. Whichever is the case, define $\mathcal{Y}_{0}$ to be the set of possible values of the initial condition such that, if the initial condition is observable, then $\mathcal{Y}_{0}=\{y_{0}\}$ and if it is not then $\mathcal{Y}_{0}=\{0,1\}$.

Projecting away individual effects yields

equation*[equation* omitted — 368 chars of source]

It is convenient to define sets of indices

equation*[equation* omitted — 284 chars of source]

Observing that

equation*[equation* omitted — 476 chars of source]

the projection can be expressed as

equation*[equation* omitted — 431 chars of source]

The $U$-level set of this projection for any $(y,x)$ is

equation[equation omitted — 449 chars of source]

In static models with ${\Greekmath 010D} =0$ there is the further simplification

equation*[equation* omitted — 279 chars of source]

If observations in any periods $t$ are missing for a class of units, for example if there is an unbalanced panel, then such $t$ are in neither $ \mathcal{T}_{0}$ nor $\mathcal{T}_{1}$ and the set $\mathcal{U}(y,x;{\Greekmath 0112} )$ leaves the value of $U_{t}$ unrestricted. The identification analysis here still applies, and will result in inequalities that reflect the lack of restrictions on $U_{t}$ in such periods. In a dynamic model with the value of some $Y_{t}$ not observed that value becomes an additional unobserved unit-specific variable in the determination of $Y_{t+1}$ removed via projection along with other such variables. Unbalanced panels can be handled in the same manner.

To illustrate identification analysis in dynamic binary response panel models consider the following example.

Example 1: Two and three period binary response.

When $T=2$, if $Y_{0} = y_0$ is observed then $ \mathcal{Y}_{0}=\{y_{0}\}$ and ((ref)) simplifies as follows.

gather*[gather* omitted — 713 chars of source]

Identification regions for ${\Greekmath 0112} $ using the inequalities arising when $Y=(0,1)$ and when $Y=(1,0)$ are provided by aristodemou2021semiparametric for models in which $Y_{0}$ is observed and either one of $U\mathbin{\vbox{\baselineskip=0pt\lineskip=0pt \moveright2.5pt\hbox{$\|$} \hrule height 0.2pt width 10pt}}X\mid Y_{0}$ or $U\mathbin{\vbox{\baselineskip=0pt\lineskip=0pt \moveright2.5pt\hbox{$\|$} \hrule height 0.2pt width 10pt}}X$ hold. khan2023identification provide sharp identification regions for ${\Greekmath 0112} $ in dynamic binary response models for arbitrary finite $T$ under a conditional stationarity restriction, with $Y_{0}$ observed.

By contrast, taking the projection approach, it is not necessary to have $Y_{0}$ observed. When $Y_{0}$ is not observed the set $\mathcal{U} (y,x,{\Greekmath 0112} )$ is simply the union of the sets $\mathcal{U}(y,x;{\Greekmath 0112} )$ obtained on setting $y_{0}=0$ and then $y_{0}=1$. The sets corresponding to $Y\in \{(0,0),(1,1)\}$ are unchanged; the others are as follows.

equation*[equation* omitted — 370 chars of source]

where ${\Greekmath 010D}^- \equiv -\min\{{\Greekmath 010D},0\}$. Table (ref) shows the sets $\mathcal{U}(y,x;{\Greekmath 0112} )$ for the case in which $T=3$ and $Y_{0}$ is observable.\footnote{The inequalities that appear here and in similar models involving threshold crossing conditions and linear indexes are routine to derive using Fourier-Motzkin elimination.} Table (ref) shows the sets when $T=3$ and $Y_{0}$ is not observed.

For the case in which $T=3$ the sets $\mathcal{U}(y,x;{\Greekmath 0112} )$ can be expressed involving just $u_{31}^{\Delta }$ and $u_{32}^{\Delta }$, since $u_{21}^{\Delta }=u_{31}^{\Delta }-u_{32}^{\Delta }$, enabling visualization on the space of $(U_{31}^{\Delta },U_{32}^{\Delta })$. The six nontrivial sets $\mathcal{U}(y,x;{\Greekmath 0112} )$ for each $y\in \mathcal{R}_{Y}$ are depicted in Figure (ref) for the case in which ${\Greekmath 0112} =({\Greekmath 010D} ,{\Greekmath 010C} )$ with ${\Greekmath 010D} =1$ and $X=x$ such that $-x_{32}^{\Delta }{\Greekmath 010C} =-2$, and $-x_{31}^{\Delta }{\Greekmath 010C} =2$. $\square$

figure[figure omitted — 748 chars of source]
table[table omitted — 2,389 chars of source]
table[table omitted — 1,557 chars of source]

Ordered response models

Consider a dynamic ordered response panel model with

equation[equation omitted — 337 chars of source]

where ${\Greekmath 010B} _{J+1}=-{\Greekmath 010B} _{0}=\infty $, ${\Greekmath 0112} =\left( {\Greekmath 010B} _{1},...,{\Greekmath 010B} _{J},{\Greekmath 010C},{\Greekmath 010D}\right) $, and $U$ is continuously distributed with full support on $\mathbb{R}^T$ conditional on $X$.\footnote{Thus as in Section (ref) there is no loss in using weak inequalities in expressions for $\mathcal{U}(y,x;{\Greekmath 0112})$.} In the static case ${\Greekmath 010D} =0$.

The outcome $Y_{t}$ is ordered categorical, taking the value $j\in \{0,...,J\}$ if $Y_{t}^{\ast }\in \left[ {\Greekmath 010B} _{j},{\Greekmath 010B} _{j+1}\right) $, where $Y_{t}^{\ast }$ is a latent index. The normalization ${\Greekmath 010B} _{1}=0$ is imposed since one of the ${\Greekmath 010B} _{j}$ parameters can be absorbed by unit-specific variable $C$. The variable $L_{t}$ here denotes functions of lagged outcomes\ in a model in which $t$ indexes time. For example there could be $L_{t}={\Greekmath 0113} _{t-1}$ where ${\Greekmath 0113} _{t-1}\equiv \left( 1\left[ Y_{t-1}=1\right] ,...,1\left[ Y_{t-1}=J \right] \right)$ in a model with one-period lagged outcome dependence.\footnote{It is straightforward to accommodate multiple lags, for example two lags with $L_{t}=({\Greekmath 0113} _{t-1},{\Greekmath 0113} _{t-2})$. The indicator for one value of $Y_{t-1}$, here $1\{Y_{t-1}=0\}$, is omitted by normalization as in honore2021dynamic.} In this example ${\Greekmath 010D} =({\Greekmath 010D} _{1},...,{\Greekmath 010D} _{J})'$ so $L_{t}{\Greekmath 010D} ={\Greekmath 010D} _{j}$ if and only if $Y_{t-1}=j$.

Now expressions for the sets of values of $U$ that can occur given values of observed variables are derived. These are the $U$-level sets of the projection $\mathcal{R}_{YXU}({\Greekmath 0112})$ for the ordered response structural function (ref).

To deal with cases in which the lag variable is not observed, define $\mathcal{L}(y,x;{\Greekmath 0112} )$ as the set of possible values of the lag variables $(L_{1},...,L_{T})$ given the observability of lagged outcomes. In this exposition only one period lags are considered.

If $L_{t}={\Greekmath 0113} _{t-1}$, and the initial value $Y_{0}$ is not observed, then $\mathcal{L}(y,x;{\Greekmath 0112} )$ is the set of ${\Greekmath 0113} _{0},...,{\Greekmath 0113} _{T-1}$ produced by observed $y_{1},...,y_{t-1}$ with any value of ${\Greekmath 0113} _{0}$. If instead lag dependence manifests through unobservable realizations $y^{\ast }$ as studied below, then $\mathcal{L} (y,x;{\Greekmath 0112} )$ will restrict each $L_{t}$ to the interval implied by observed realizations $y$.

To obtain the $U$-level set $\mathcal{U}(y,x;{\Greekmath 0112} )$ note that there is for all $s$ and $t$

eqnarray*[eqnarray* omitted — 376 chars of source]

and upon adding

equation*[equation* omitted — 312 chars of source]

which leads to

multline*[multline* omitted — 566 chars of source]

In a static model with ${\Greekmath 010D} =0$ there is no lag dependence and the simplification

equation*[equation* omitted — 364 chars of source]

In contrast to other approaches to dynamic ordered response panel models, unobservable initial conditions can be accommodated using the projection approach. Moreover, period $t$ lags can include functions of both observable and latent variables, such as lagged values of the unobserved index $Y^{\ast }$. This is important because in many applications the ordered outcome $Y_{t}$ may depend not just on the value of $Y_{t-1}$ but on the location of $Y_{t-1}^{\ast }$ relative to the thresholds ${\Greekmath 010B} _{j}$. For example, if $Y_{t}$ is a categorical measure of health status, the effect of current health on future health may be different for two individuals in \textquotedblleft good\textquotedblright\ health, one of whom is close to the boundary for the \textquotedblleft fair\textquotedblright\ health category, and the other close to the boundary for the \textquotedblleft excellent\textquotedblright\ health category. Or, in application to letter grades obtained in a sequence of courses, dynamic impacts may be more effectively measured by a student's numerical score rather than, say, whether they achieved an \textquotedblleft A-\textquotedblright\ or \textquotedblleft B+\textquotedblright .

Lagged Outcome Dependence

With one period lagged outcome dependence $ \mathcal{L}(y,x;{\Greekmath 0112} )=\mathcal{L}_{1}(y,x;{\Greekmath 0112} )\times \dots \times \mathcal{L}_{T}(y,x;{\Greekmath 0112} )$ where $\mathcal{L} _{t}(y,x;{\Greekmath 0112} )=\{{\Greekmath 0113} _{t-1}\}$ for all $t>1$, and $\mathcal{L}_{1}(y,x;{\Greekmath 0112} )$ is the set of standard basis vectors in $\mathbb{R}^{J}$.

To illustrate consider such a model with two periods, three categories, $T=J=2$, and $Y_0$ not observed. When $Y=(0,0)$ or $Y=(2,2)$ any value of $U$ is possible, as there is always $C$ small enough or large enough, respectively, to produce either outcome. Table (ref) shows the other $U$-level sets in this model.

table[table omitted — 1,691 chars of source]

Lagged Latent Dependence

Now consider the case in which the period $t$ outcome depends on the lagged latent index $Y_{t-1}^{\ast }$ so $L_{t}=Y_{t-1}^{\ast }$. Repeated substitution for values of $Y_{s}^{\ast }$, $s<t$, in the equation for $Y_{t}^{\ast }$ in (ref) gives

equation*[equation* omitted — 218 chars of source]

So the set of values of $U$ that deliver $Y=y$ when $X=x$ is

equation*[equation* omitted — 516 chars of source]

The set $\mathcal{U}(y,x;{\Greekmath 0112} )$ differs from the case with lagged outcome dependence. The inequalities defining $\mathcal{U}(y,x;{\Greekmath 0112} )$ are linear in $c$ and $y_{0}^{\ast }$ so Fourier-Motzkin elimination can be used to remove these variables from the inequalities that define $\mathcal{U}(y,x;{\Greekmath 0112} )$.

Identified sets

With the $U$-level sets defined as in discrete outcome models such as those presented in Sections (ref) and (ref), moment inequality characterizations of identified sets for common parameters can be obtained using Artstein's inequality.\footnote{Artstein's inequality is established in artstein1983distributions, see also Molchanov:17 pages 83--84 and Molchanov/Molinari:18 Section 2.2.} The inequality provides the following corollary to Proposition (ref).

corollarySuppose that Restrictions PM and RCS hold. Then the identified set for $\left(f,G_{U|X}(\cdot|\cdot)\right)$ comprises those pairs $\left(f,G_{U|X}(\cdot|\cdot)\right)\in \mathcal{M}$ such that \begin{equation} \mathbb{P}\left[\mathcal{U}(Y,X;f) \subseteq \mathcal{S} |X=x \right] \leq G_{U|X}\left(\mathcal{S} |x\right), \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ a.e. } x\in\mathcal{R}_X\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{.} \end{equation} for all closed $\mathcal{S} \subseteq \mathcal{R}_U$. The identified set for $f$ is the set of $f$ such that ((ref)) holds for $f$ and some $G_{U|X}(\cdot|\cdot)\in \mathsf{G}_{U|X}$ with $\left(f,G_{U|X}(\cdot|\cdot)\right)\in \mathcal{M}$.

It will now be demonstrated how this corollary can be specialized to produce moment inequality characterizations of identified sets for common parameters. Prior applications of this inequality in the partial identification literature include Beresteanu/Molchanov/Molinari:10 and Chesher/Rosen:17, see Molinari:Handbook for further references. The novelty here is not in the use of Artstein's inequality for identification analysis, but rather its application to level sets of the projection $\mathcal{R}_{YXU}$ obtained by removal of incidental parameters, and under distributional restrictions commonly found in panel models having no counterpart in cross section models.

The characterization of the identified set provided by Corollary (ref) using Artstein's inequality comprises for each $x \in \mathcal{R}_X$ as many inequalities as the number of closed sets in $\mathcal{R}_U$. Previous papers such as Galichon/Henry:09, Chesher/Rosen:17, and Luo/Ponomarev/Wang:25 have characterized core determining collections that comprise a smaller collection of sets $\mathcal{S}$ such that if (ref) holds for all $\mathcal{S}$ in the collection, then it holds for all closed sets. To the best of our knowledge these results have not been previously employed in panel models but they are applicable here to simplify characterizations of identified sets delivered by Artstein's inequality.

With the characterizations of the $U$-level sets of Sections (ref) and (ref) the inequality (ref) delivers observable implications for the common parameters. This is now illustrated in the context of Example 1 in Section (ref). The same steps can be taken to characterize identified sets for the ordered response panel models of Section (ref).

\noindentExample 1, continued: Consider the dynamic binary panel model with unobservable initial condition and $T=3$. The $U$-level sets for a particular $x$ and ${\Greekmath 0112}$ are illustrated in Figure (ref). We can see immediately from the figure that with, for example, $\mathcal{S} = \left\{u: u^{\Delta}_{31} \geq u^{\Delta}_{32} \right\}$ the inequality (ref) becomes $\mathbb{P}\left[Y \in \{(0,1,0),(0,1,1) \} |X=x\right] \leq G_{U|X}\left(\mathcal{S}|x\right)$. This set is however not amongst the minimal core determining collection.

From Theorem 1 of Chesher/Rosen:17 it follows that only sets $\mathcal{S}$ that comprise unions of sets on the support of $\mathcal{U}(Y,X;f)$ need consideration. Theorem 3 of that paper establishes that among this collection, one need not consider those sets $\mathcal{S}$ that can be partitioned into two sets $\mathcal{S}_1$ and $\mathcal{S}_2$ such that either $\mathcal{U}(Y,X;f) \subseteq \mathcal{S}_1$ or $\mathcal{U}(Y,X;f) \subseteq \mathcal{S}_2$, but not both simultaneously. Such a set $\mathcal{S}$ is not self-connected in the terminology of Luo/Ponomarev/Wang:25.\footnote{That paper also shows that it is generally possible to achieve further refinement, establishing that among the class of sets comprising unions of sets on the support of $\mathcal{U}(Y,X;f)$, those that are both self-connected and complement-connected comprise a minimal core-determining collection. However, in the panel models studied here in which $\mathcal{U}(Y,X;f) = \mathcal{R}_U$ for some $y \in \mathcal{R}_Y$, the requirement that sets be complement-connected provides no reduction in the core-determining collection.} The minimal core determining collection of sets for the value of $x$ and ${\Greekmath 0112}$ that produce the $U$-level sets of Figure (ref) yields 32 moment inequalities of the form (ref). $\square$

If $U$ and $X$ are stochastically independent there is the following simplification of Artstein's inequality

equation[equation omitted — 256 chars of source]

This applies with both parametric and nonparametric restrictions on the class of functions $f$ and distributions $G_{U}(\cdot|\cdot)$ admitted by the model. For example, if the utility function and distribution of unobservable heterogeneity are parametrically specified up to ${\Greekmath 0112} \in \Theta \subseteq \mathbb{R}^d$ in (ref), $f$ may be replaced by ${\Greekmath 0112}$ and $G_U(\mathcal{S})$ by $G_U(\mathcal{S};{\Greekmath 0112})$.

Even when using only core determining collections of sets, the number of inequalities can be large. Nonetheless, approaches for asymptotic inference with infinitely many conditional moment inequalities can be used, see for instance Section 2.2 of Chernozhukov/Chetverikov/Kato:19 and Example 2 of Andrews/Shi:17.

The next section shows how identification analysis can proceed under nonparametric specifications of the distribution of within-unit-varying heterogeneity.

Nonparametric distributional specifications

An advantage of the projection approach developed in this paper is that it enables identification analysis when there are neither parametric nor stationarity restrictions on the distribution of within-unit-varying heterogeneity. In Section (ref) it was shown how this can be achieved using moment restrictions. Here are two alternative approaches more suited to models of discrete outcomes. Other such restrictions are possible.

Consider models such as the discrete outcome models considered in this section in which $U$ level sets are determined entirely by restrictions on differences $u_{st}^{\Delta }$ for a collection of values of $s$ and $t$.

First consider quantile independence restrictions. For chosen values of $s$ and $t$, first define an ascending sequence of quantile probabilities

equation*[equation* omitted — 84 chars of source]

which are specified values in $\left[ 0,1\right] $ with $p_{st}^{0}\equiv 0$ and $p_{st}^{K+1}\equiv 1$. Then define additional parameters, namely the unknown elements of

equation*[equation* omitted — 171 chars of source]

with ${\Greekmath 0115} _{st}^{0}\equiv -\infty $, ${\Greekmath 0115} _{st}^{K+1}\equiv +\infty $. These new parameters are values of the quantiles of the marginal distributions of the $U_{st}^{\Delta }$ at the chosen quantile probabilities, e.g. $(0.25,0.5,0.75)$, and there are the restrictions\footnote{It is easy to impose restrictions of symmetry and unimodality if that were desired.}

equation*[equation* omitted — 282 chars of source]

Moment inequalities characterizing the identified set of values of the parameters ${\Greekmath 0112} $ and the quantile values in ${\Greekmath 0115} _{st}$ are

equation*[equation* omitted — 256 chars of source]
equation*[equation* omitted — 208 chars of source]

for all pairs $(s,t)$ for which the quantile independence restrictions are maintained. Identified sets for ${\Greekmath 0112}$ are obtained as those values of ${\Greekmath 0112}$ for which there exist values of ${\Greekmath 0115}_{st}$ such that all such inequalities are satisfied. Chesher/Kim/Rosen:23 gives details and has an example of this approach in action in a different, non-panel, context.\footnote{Chesher/Kim/Rosen:23 studies an IV Tobit model in which explanatory variables may be endogenous. Proposition 4 in Section 4.2 deals with quantile independence restrictions and is easily extended to the models considered in this paper.}

Finally consider pairwise conditional exchangeability restrictions requiring that, for some chosen $s$ and $t$, $U_{s}$ and $U_{t}$ are exchangeable conditional on $X$. Under this restriction the median of $ U_{st}^{\Delta }$ conditional on $X=x$ is zero for all $x$. The inequalities above with $K=1$, ${\Greekmath 0115} _{st}^{1}=0$ and $p_{st}^{1}=0.5$ deliver bounds on ${\Greekmath 0112} $ absent parametric restrictions on the distribution of $U$. \footnote{Under the conditional exchangeability restriction the probability density function of $U_{st}^{\Delta }$ is symmetric around zero. This may deliver additional bounds in some cases.}

\noindentExample 1, continued: Consider again the dynamic binary outcome model as in (ref) with unobserved initial value $Y_{0}$ and $T=3$ with $U$-level sets shown in Table (ref).

The zero median independence restriction implied by pairwise conditional exchangeability of all elements of $U$ delivers the identified set of values of $({\Greekmath 010C} ,{\Greekmath 010D} )$ as those satisfying

equation[equation omitted — 172 chars of source]

where the expressions $w_{j}(x,{\Greekmath 010C} ,{\Greekmath 010D} )$ are shown in Table (ref). In this table, $p_{x}(y_{1},y_{2},y_{3})\equiv \mathbb{P} [Y=(y_{1},y_{2},y_{3})|X=x]$ and the column headed \textquotedblleft $ \mathcal{S}$\textquotedblright\ shows the set $\mathcal{S}$ in the zero median independence restriction

equation*[equation* omitted — 88 chars of source]

that delivers each row of the table. $\square$

table[table omitted — 1,743 chars of source]

Related Literature on Discrete Outcome Panel Models

Analysis of binary response panel models has a long history going back to Rasch (1960, 1961),\nocite{rasch1960probabilistic} \nocite{rasch1961general} andersen1970asymptotic, and see also chamberlain2010binary, in which a static model is studied with the elements of $U$ restricted to be i.i.d. logistic, independent of $(C,X)$. With these distributional restrictions ${\Greekmath 010C}$ is point-identified under a rank condition and consistently estimated by a conditional maximum likelihood estimator that conditions on $\sum\limits_{t=1}^T Y_t$.\footnote{A precise statement of the rank condition is provided as Assumption 2 in Davezies/D'Haultfoeuille/Laage:22.}

Extensions of panel logit models to dynamic models have been considered. honore2000panel provides results for identification and estimation of a dynamic binary panel model maintaining mutual independence of all elements of $U$ and independence of $U$ and $(C,X)$, and in most cases restricting the elements of $U$ to be logistically distributed. Kitazawa:22, Dano:23, and Honore/Weidner:25 provide moment equations in dynamic panel logit models, which can be used to study identification and estimation of common parameters. Dobronyi/Gu/Kim/Russell:25 analyze the full likelihood from the dynamic panel logit model and make a connection to the truncated moment problem to obtain all of the model's observable implications. That paper and Davezies/D'Haultfoeuille/Laage:22 also provide characterizations of certain average and marginal effects.

An alternative to these logit specifications in the binary outcome panel model is a conditional stationarity restriction introduced in manski1987semiparametric, requiring that conditional on $(X,C)$ the variables $U_1,...,U_T$ all have the same marginal distribution. This is implied by the panel logit distributional restriction, but is weaker. It does not require independence of $U$ and $(X,C)$, and it can allow for correlation in the components of $U$. Nonetheless it does restrict the joint distribution of $U$ and $C$.\footnote{As pointed out in Chernozhukov/Fernandez-Val/Hahn/Newey:09 the stationarity restriction $U_t|C,X \overset{d}{=}U_1|C,X$ for all $t$ is equivalent to $(U_t,C)|X \overset{d}{=}(U_1,C)|X$ for all $t$.}

Conditional stationarity restrictions have been used in several papers. Abrevaya:2000 studies a class of generalized regression models that nests binary response and censored outcome models under conditional stationarity and stronger restrictions. Chernozhukov/Fernandez-Val/Hahn/Newey:09 characterizes bounds for average and quantile effects in several nonseparable panel models, including binary response models. khan2023identification provides set identification results for common parameters in semiparametric dynamic binary response panel models. Conditional stationarity restrictions have also been used in multinomial response panel models, for example in shi2018estimating, khan2021inference, pakes2022moment, Pakes/Porter/Shepard/Calder-Wang:25, Gao/Wang:24, Gao/Li:20, and mbakop2023identification.

aristodemou2021semiparametric is the one paper of which we are aware that studies the binary response specification (ref) without restricting the covariation of $C$ with either $X$ or $U$. In that paper $U$ and $X$ are restricted to be independently distributed, in some cases conditional on an initial condition, but, importantly, not conditional on $C$. That paper provides bounds on parameters in the models studied but does not claim sharpness. The projection approach delivers characterizations of sharp identified sets and applies more broadly, for example allowing arbitrary $T$, unobserved initial conditions, and alternative restrictions on the joint distribution of $U$ and $X$.

The literature on fixed effects models of dynamic ordered response panels is recent and not extensive. honore2021dynamic employ functional differencing to provide moment conditions that can be used as a basis for estimation and inference in models in which the elements of $U$ are i.i.d. logistically distributed and independent of $X$ and $C$. In a model with these distributional restrictions on $U$ but an alternative lag dependence specification $ L_{t}=1[Y_{t-1}\geq k]$ for specified $k$, Muris/Raposo/Vandoros:25 develop a conditional maximum likelihood estimator building on insights from honore2000panel. References to the broader literature on ordered response panel models, including random effects approaches and static models, can be found in these papers. An unpublished chapter of Aristodemou:16 provides bounds on common parameters in some fixed effects ordered response panel models when $T=2$.\footnote{The models studied in Chapter 6 of Aristodemou:16, like those studied in this paper, impose no restrictions on the joint distribution of $C$ and $U $, although the analysis requires an observed initial condition, independence restrictions conditional on the initial condition, and does not allow dependence on lagged latent variables $Y^{\ast }$ as is allowed here. Moreover, non-sharp outer bounds are obtained. However, in contrast to honore2000panel and Muris/Raposo/Vandoros:25, in Aristodemou:16, as here, no logistic or other parametric distributional restriction on $U$ is required.}

Numerical Illustrations of Identified Sets

Illustrations of identified sets for $({\Greekmath 010C} ,{\Greekmath 010D} )$ are presented for the CES model of Section (ref). Both weak and strict exogeneity restrictions, MW and MS, are considered.

As in Section (ref), the model features firm-specific unobservables $C$ and $S$ in the CES production function. For the sake of illustration, a data generation process is considered in which $T=3$, and, although it is unknown to the econometrician, $C=S=0$ for all firms, so output follows a Cobb-Douglas specification \[ Y_{t}={\Greekmath 010C} _{0}\left( {\Greekmath 010D}_0 \log (X_{t1})+(1-{\Greekmath 010D} _{0})\log (X_{t2})\right) +U_{t}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{,}\qquad \quad t\in \{1,2,3\}. \] This is chosen for simplicity, but calculations are easily done for more complex cases.

Let $X$ denote the $3$x$2$ matrix with elements $X_{tj}$. To determine the support of $X$, values of its elements $X_{t1}$ and $X_{t2}$ were drawn i.i.d. with $\log X_{tj} \sim N(0,1/4)$ to produce $50$ support points, each of which was given equal probability.

The identified set is given by the inequalities (ref), equivalently for each $r$:

equation[equation omitted — 341 chars of source]

For $r$ such that $\operatorname{Pr}\left[ \sum_{t=1}^{T}w_{t}(r,X)=0\right] <1$, the last term can be made arbitrarily large and (ref) will be satisfied for any such $r$. Therefore we have the characterization

equation[equation omitted — 275 chars of source]

for which we need only consider values of $r$ such that $ \sum_{t=1}^{T}w_{t}(r,X)=0$ almost surely. If the support of $X$ is such that there exists no proper linear subspace of $\mathbb{R}^{2T+1}$ that contains $(1,X_1,...,X_T)$ almost surely, as is the case in this illustration, this is equivalent to imposing the restrictions

equation[equation omitted — 318 chars of source]

This has the effect of removing $c$ from the inequality.

On replacing $Y_{t}$ by its expectation conditional on $X$, denoted ${\Greekmath 0116}_t(X)$, there is:

equation[equation omitted — 318 chars of source]

where $\mathcal{R}$ is the set of $r$ such that $\Vert r \rVert = 1$ and (ref) holds.

Weak and strict exogeneity restrictions are distinguished by additionally imposing $r_{thk}=0$ for all $t<h$ under weak exogeneity, as described in Section (ref).

The maximisation with respect to $s$ is done using the modified golden section method provided by the optimise function of R.\footnote{R2025: \url{https://www.R-project.org/}.} Maximisation is done with respect to two alternative monotone transformations of $s\in (-\infty ,1]$ to the unit interval: \[ {\Greekmath 0115} _{1}(s)=\exp (s-1)\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{,}\quad \quad \quad {\Greekmath 0115} _{2}(s)=\frac{1}{{\Greekmath 0119} } \arctan (-\log (1-s))+\frac{1}{2}\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{.} \] Any internal maxima that are found are compared with the values obtained at $s=-\infty $ and $s=1$ and the largest value is chosen. The expectation in (ref) is obtained as the probability-weighted sum over the support of $X$.

Under the strict exogeneity Restriction MS stated in Section (ref), with $T=3$ there are $21$ elements in $r$ subject to $\lVert r \rVert =1$ and the seven restrictions (ref). Under the weak exogeneity Restriction MW there are an additional six restrictions $r_{thk} = 0$ for $t < h$.\footnote{Note that imposing $r_{thk}=0$ for $t<h$ in (ref) corresponds to a less restrictive model than when $r_{thk}=0$ for $t<h$ is not imposed.} Even in the case with $r$ more heavily restricted, minimization is hard to calculate precisely so an alternative calculation is done that delivers an outer region. The bounds obtained nonetheless demonstrate the informativeness of the CES model with multiple fixed effects, one entering nonlinearly.

The calculation proceeds by drawing $5000$ pseudo-random standard Gaussian values of $r$ which are subjected to the required restrictions. A value of $({\Greekmath 010C} ,{\Greekmath 010D} )$ is deemed out of the identified set if the inequality (ref) is violated at any of the values of $r$ considered. The same stream of $5000$ pseudo-random values of $r$ are employed as each value of $({\Greekmath 010C} ,{\Greekmath 010D} )$ is considered.

Figure (ref) shows the results obtained under the weak (dark blue) and strict (light blue) exogeneity restrictions when $({\Greekmath 010C} _{0},{\Greekmath 010D} _{0})=(1,0.5)$. These outer regions are quite informative even in this simple case in which $T=3$. Having larger $T$ or taking more than $5000$ pseudo-random draws would deliver tighter bounds.

figure[figure omitted — 545 chars of source]

Discussion and concluding remarks

In the econometrics and the statistics literature the incidental parameter problem arising with short panels has mostly been subject to analysis using particular parametric specifications of distributions of outcomes or unobservables. Notable examples are Neyman/Scott:48, honore2000panel, Lancaster:00, and Bonhomme:12. A problem for practicing researchers is: which distribution to choose - economic reasoning and context usually offers little guidance, and the literature says little about the consequences of an unsuitable choice.

A notable exception is the work based on the stationarity restrictions introduced in manski1987semiparametric. In that work no parametric restrictions are placed on probability distributions, but the approach is not universally applicable.

The situation in the year 2000 was summarized by Tony Lancaster as follows:\footnote{Lancaster:00, page 404.} \textquotedblleft The absence of a method guaranteed to work in a large class of econometric models means that any paper on [the incidental parameter problem] must be a catalogue of examples\textquotedblright . Little has changed in the years that followed. The projection approach introduced in this paper fills this gap, delivering a universally applicable solution to the incidental parameter problem.

Taking this projection approach, incidental parameters in any number are projected away from the space of observed and unobserved variables. The original model specification then delivers correspondences specifying the feasible combinations of the remaining variables. Classical \textquotedblleft fixed effects\textquotedblright , unobserved initial conditions and missing data are examples of variables that can be treated in this way.

Projection delivers an incomplete model whose identifying power can be determined by extension of available methods, for example as developed for the analysis of Generalized Instrumental Variable models in Chesher and Rosen (2017). Estimation and inference using the resulting characterizations of identified sets is off-the-shelf.

With the incidental parameters projected away, robust econometric analysis can proceed absent restrictions on their joint probability distribution with other variables. Importantly, progress can be made using nonparametric specifications of the distribution of the unobserved variables that vary within observational units. For example, mean and conditional mean restrictions can be employed as in the CES production function example of Section (ref) and quantile independence restrictions can be used as described in Section 4.

Four examples of application to econometric models have been set out in this paper. More can be found in the online working paper Chesher, Rosen and Zhang (2024), which includes applications to models admitting multiple indexes, for example, models of multiple discrete choice and simultaneous binary response.

Endogenous explanatory variables are easily accommodated following the GIV analysis of Chesher and Rosen (2017). All endogenous variables are placed in the list of outcomes, $Y$, and restrictions suitable for the context are imposed on the distribution of within-observation-unit-varying $U$ and explanatory variables, $X$. It is straightforward to impose weak exogeneity restrictions, for example, in dynamic panels requiring that for all $t$, $(X_{1},\dots ,X_{t})$ and $(U_{t},\dots ,U_{T})$ satisfy some suitable-for-context independence restriction while allowing feedback from historic shocks and outcomes to affect the determination of future $X$ values.

Finally, the results of this paper can be useful for conducting sensitivity analysis and specification testing. The identified sets delivered by this paper's models that place no restriction on the distribution of unobservable unit-specific effects will contain the structures identified by more restrictive models if their restrictions are satisfied by the process under study, as captured in the distribution of observable variables the process delivers. The analysis set out here can show how sensitive the findings obtained using those more restrictive models are to relaxation of their additional restrictions. It may be found that estimation employing a point-identifying model delivers a structure outside an estimator of the identified set obtained using a less restrictive model of the type studied in this paper. That will suggest the more restrictive model is misspecified. Formal development of such specification tests is a potentially fruitful topic for future research.