Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
152,485 characters · 13 sections · 122 citation commands
Identification of Dynamic Panel Logit Models with Fixed Effects We thank Victor Aguirregabiria, Roger Koenker, Ismael Mourifi\'e and Stanislav Volgushev for useful discussion. We are grateful to numerous seminar participants for their feedback, and are especially grateful to Francesca Molinari and three anonymous referees for their helpful comments. All errors are our own.
\thispagestyle{empty}
Keywords: Stieltjes Truncated Moment Problem, Dynamic Panel Logit Model, Fixed Effects, Semidefinite Programming
\pagenumbering{arabic}
Dynamic panel logit models are valuable empirical tools for modeling repeated choices made by households, firms and individual consumers. These models are favored in part because they can account for permanent unobserved heterogeneity, allowing the researcher to distinguish between true dynamics induced by lagged choice dependence, and spurious dynamics, which are a result of persistent individual heterogeneity (see heckman1981heterogeneity). Work on identification in these models has a long history, and can be roughly divided into two main areas: the sufficient statistics approach (e.g. rasch1960probabilistic, rasch1961general, andersen1970asymptotic, chamberlain1980analysis, chamberlain1985heterogeneity, honore2000panel, magnac2000subsidised, hahn2001information, gu2023information) and the functional differencing approach (e.g. johnson2004identification, bonhomme2012functional, honore2024moment).
In this paper we study a new and general approach to identification in a class of dynamic panel logit models with latent individual effects, providing an alternative to the sufficient statistics and functional differencing approaches. Using the structure of the logistic distribution, we show that the likelihood for these models can be written as a polynomial in certain generalized moments of the distribution of latent individual effects, revealing a connection to the truncated moment problem dating back to chebyshev1874valeurs and stieltjes1894recherches. Through this connection, we show that the identified set of structural parameters can be characterized by a set of conditional moment equalities subject to a certain set of shape restrictions on the model parameters. Estimation and inference is then based on repeatedly solving semidefinite programs, a special kind of convex program which can be solved quickly and reliably. We then show how to adapt the inference procedure in chernozhukov2023constrained to construct confidence sets for the model parameters. A key advantage of our method is its ability to handle partially identified models, and we show that the new characterization can deliver sharp bounds in cases where existing methods deliver no identifying restrictions. In addition, unlike many existing approaches, our approach can be used to construct the identified set of certain functionals of the distribution of latent individual effects, including average marginal effects and the average structural function.
There are two main challenges when studying dynamic panel logit models: the initial conditions problem, and the incidental parameters problem. The initial conditions problem arises because the joint distribution of the initial choices and the individual fixed effects is not nonparametrically point-identified (e.g. see heckman1981incidental and wooldridge2005fixed). The incidental parameters problem refers to the fact that, when the number of time periods is fixed, it is generally not possible to consistently estimate individual fixed effects, and attempting to do so can bias the estimates of the structural parameters (e.g. neyman1948consistent). This paper focuses on the incidental parameters problem, for which there are two common approaches: the random effects approach, and the fixed effects approach.\footnote{For a more complete survey of the literature, we refer the readers to arellano2001panel.} The (correlated) random effects approach places restrictions on the joint distribution of the initial conditions and the individual effects using a parametric distribution or a finite mixture (e.g. chamberlain1980CRC, wooldridge2005initial). When these assumptions are satisfied, the structural parameters and various functionals of the latent variable distribution are point-identified and can be consistently estimated. In contrast, the fixed effects approach treats the latent individual effects as random, but is entirely agnostic about their distribution and their dependence on the initial conditions.\footnote{Consistent with the existing literature, if the distribution of the time-invariant individual effects is not parametrically specified and is allowed to depend arbitrarily on covariates and initial conditions, then we refer to this as the “fixed effects” approach. See for instance honore2006bounds p. 612 for similar terminology. Throughout the paper we used “fixed effects” and “latent individual effects” interchangeably.} As a result, the fixed effects approach is more flexible, but presents a number of interesting identification and estimation challenges.
Under the fixed effects approach, in some cases the structural parameters are identified and can be consistently estimated using conditional maximum likelihood, pioneered by rasch1960probabilistic, rasch1961general, andersen1970asymptotic, chamberlain1980analysis, and chamberlain1985heterogeneity. This method involves finding a minimally sufficient statistic for the fixed effects, and constructing a partial likelihood that conditions on this statistic. By the definition of sufficiency, this partial likelihood no longer depends on the fixed effects. If this partial likelihood also depends on the structural parameters, then the first-order conditions to maximize the partial likelihood provide moment conditions that can be used for identification and estimation. honore2000panel extend this approach to dynamic logit models with time-varying covariates. Unfortunately, few models admit nontrivial sufficient statistics, and so the method does not always result in useful identifying restrictions. Even when it does, it can fail to exhaust all of the model's identifying content, and so can it can deliver nonidentification in cases when the structural parameters are point- or partially-identified.\footnote{There is one exception: if the likelihood of the sufficient statistics no longer depends on the structural parameters, then the conditional maximum likelihood method utilizes all relevant identifying information for the structural parameters. In many cases, including in the dynamic panel logit model, this condition is generally not satisfied.} In contrast, our approach always delivers the sharp set of model restrictions, and can always be used to construct the sharp identified set for the structural parameters. Later, we show specific examples where we are able to construct the sharp identified set for the structural parameters when conditional maximum likelihood delivers no identifying restrictions.
Our analysis also sheds light on the functional differencing approach proposed by johnson2004identification and bonhomme2012functional and used for a similar class of models by honore2024moment. At a high level, functional differencing searches for a collection of moment functions that do not depend on the latent variables, but that deliver some identifying information about the structural parameters. Functional differencing proceeds on a case-by-case basis, searching for moment conditions specific to each model. Finding the relevant collection of moment functions has historically been challenging. It is also difficult to determine whether all relevant moment functions have been found, and whether a collection of moment functions exhaust all the identifying restrictions of the model. Even given all relevant moment functions, it can be difficult to prove point identification of the structural parameters, a precondition for using standard estimation and inference methods. Despite these challenges, recent progress was made by honore2024moment, who found new moment conditions for the structural parameters in the AR(1) dynamic panel logit model with covariates, and proved point identification under certain conditions. They also found moment conditions in models for which the conditional maximum likelihood approach provides no identifying restrictions, such as the AR(2) dynamic panel logit model.
Relative to functional differencing, the advantage of our approach is its generality. In particular, it relies on a general structure of the logistic likelihood function that makes it relatively straightforward to apply to different models. It is also able to handle both point- and partially-identified models---such as short panels or models with limited covariate variation---and can be used to study functionals of the distribution of fixed effects. However, our approach also has an interesting connection to functional differencing: as a by-product of our analysis we show how the moment conditions from functional differencing can be constructed from the basis of the left null space of a certain matrix that arises in our approach. This allows us to provide a simple geometric explanation for why our approach sometimes provides more identifying restrictions than approaches based on functional differencing, and we provide a number of examples to illustrate when this is the case. This connection also suggests a new method for constructing moment conditions for functional differencing, which may be a promising avenue of future research.
As mentioned repeatedly, an important feature of our approach is its ability to study functionals of the distribution of fixed effects, including certain counterfactual parameters. This is done by linking the functional of interest to the generalized moments of the distribution of the latent individual effects. In contrast, both the conditional maximum likelihood approach and the functional differencing approach aim at removing the individual effects to derive moment conditions for the structural parameters. As a result, they cannot be used to study functionals of the distribution of latent individual effects. Our results on functionals relate to aguirregabiria2024identification, who were the first to show that the average marginal effect of the lagged choice in the AR(1) dynamic logit model is point-identified. While aguirregabiria2024identification restrict attention to models in which the structural parameters and the functional of interest are both point-identified, we generalize their setting to allow for partially-identified models, and cover a broader class of functionals. We also provide easily-checked sufficient conditions under which functionals are point-identified even when the latent variable distribution is not point-identified.
Outside of the sufficient statistics and functional differencing approaches, other approaches have been proposed that are based on discretizing the distribution of latent individual effects. This includes the linear programming approach in honore2006bounds and the quadratic programming approach in chernozhukov2013average. These approaches have similar advantages to our method, including the ability to handle point- or partially-identified models and the ability to study functionals of the latent variable distribution. However, both of these approaches require choosing a finite grid for the support of the latent distribution of individual effects. Furthermore, neither paper studies the impact of the approximation on rates of convergence or inference, and both papers focus on discrete covariates. Rather than construct a finite approximation to the infinite-dimensional latent variable distribution, we instead show that the latent variable distribution can be completely summarized by a finite vector of moments. Furthermore, our method maintains a similar computational cost.\footnote{For instance, in a simulation exercise with an AR(1) model with $T=3$ and $n=10^3$, across $100$ replications on average the quadratic program in chernozhukov2013average took about $0.0042$ seconds to solve with a grid size of $61$ points for the $\alpha$ distribution, and about $0.2374$ seconds to solve for a grid size of $601$ points. For the same model, our proposed semidefinite program took an average of $0.0082$ seconds to solve. }
The rest of the paper is organized as follows. Section (ref) introduces the identification problem and our main assumptions, and works through an example to illustrate our approach. General identification results and connections to the existing literature are presented in Section (ref). Estimation and inference using semidefinite programming is presented in Section (ref). An empirical application is presented in Section (ref), and Section (ref) concludes. The proofs of the main results, and additional material including a brief Monte Carlo study, can be found in the Online Supplementary Material.
We begin with some examples of models that fit into our framework.
We now present a general assumption that nests these examples as a special case. In the following, we let $\bm Y = (Y_{1}, \dots, Y_{T}) \in \mathcal{Y}^T$ denote a vector of observed choices, and we let $\bm X = (\bm X_{1}, \dots, \bm X_{T}) \in \mathcal{X}^T$ denote a vector of observed covariates. Throughout, we use $\bm W = (\bm W_{1}, \dots, \bm W_{T}) \in \mathcal{W}$ to denote a generic vector of conditioning variables, which includes any covariates $\bm X$ and may also include the initial conditions $(Y_{1-p}, Y_{2-p}, \ldots, Y_{0}) \in \mathcal{Y}^p$, depending on the model. Finally, the model also includes a latent individual effect $\alpha \in \mathbb{R}$ and a vector of structural parameters $\theta \in \Theta \subset \mathbb{R}^{d_{\theta}}$.
Assumption (ref) covers discrete choice models with idiosyncratic errors independent from the covariates and the fixed effects.\footnote{See aristodemou2021semiparametric and khan2023identification for results when this independence assumption is relaxed.} The assumption restricts attention to models whose conditional likelihood $f(\,\cdot\, \mid \bm w, \alpha; \theta)$ can be written as a polynomial in $\exp(\alpha)$, up to a common factor of $\kappa(\bm w, \alpha, \theta)$. Here $\kappa(\bm w, \alpha, \theta)^{-1}$ itself is a strictly positive polynomial of degree $S$ in $\exp(\alpha)$, which ensures that the function $f(\,\cdot\, \mid \bm w, \alpha; \theta)$ is bounded in $\alpha \in \mathbb{R}$. This will be important for our theoretical results.\footnote{Note this is actually implied by (ref) and the other positivity assumptions from Assumption (ref): summing over $\bm y \in \mathcal{Y}^T$, we have $1 = \kappa(\bm w, \alpha, \theta) \cdot \sum_{s=0}^S \exp(\alpha)^{s}\cdot \sum_{\bm y \in \mathcal{Y}^T} g_{s}(\bm y,\bm w,\theta)$, and rearranging for $\kappa(\bm w, \alpha, \theta)^{-1}$ shows it must be a strictly positive polynomial of degree $S$ in $\exp(\alpha)$.} The term $\kappa(\bm w, \alpha, \theta)$ changes depending on the model, but often its choice is obvious (e.g.\ see Example (ref) below). Assumption (ref) also fixes attention to the case where the support $\mathcal{Y}^T$ is finite, and emphasizes that $\alpha$ will be treated as a random variable with an unknown conditional distribution. Importantly, Assumption (ref) imposes no assumptions on the moments of $\alpha$, and no assumptions on the dependence between $\alpha$ and $\bm W$. Assumption (ref) also allows for interactions between the lagged outcomes and covariates, and known (up to a finite vector of parameters) nonlinear functions of the covariates, lagged outcomes, and model parameters to enter the index functions. The structure of the likelihood in (ref) in Assumption (ref) is essential to our approach, but is satisfied by a general class of logit models, including Examples (ref) - (ref). Throughout, let $\Lambda(u) := \frac{\exp(u)}{1+\exp(u)}$, and let $\bm g(\bm y, \bm w,\theta) = (g_{s}(\bm y, \bm w,\theta))_{s=0}^{S}$ denote an $(S+1)\times 1$ vector. We now show how Assumption (ref) applies to the examples introduced above. \setcounter{example}{0}
With these examples in hand, we now describe the general identification problem for models governed by Assumption (ref). Define $p(\bm y \mid \bm w):=P(\bm Y = \bm y \mid \bm W = \bm w)$, and fix a pair $(\theta, \bm w) \in \Theta\times\mathcal{W}$. Let $\mathcal{Q}$ denote the set of all Borel probability measures on $\mathbb{R}$, and consider a candidate conditional distribution $Q_{\alpha \mid \bm W}^\dagger \in \mathcal{Q}$ for the latent individual effect $\alpha$. We say that the conditional distribution $Q_{\alpha \mid \bm W}^\dagger$ can rationalize the observed conditional choice probabilities at $\theta\in \Theta$ if and only if:
a.s.\ for all $\bm y \in \mathcal{Y}^T$. The collection of all conditional probability measures $Q_{\alpha \mid \bm W}^\dagger$ that can rationalize the observed conditional choice probabilities for a fixed pair $(\theta, P)$ is given by:
Note that, depending on the value of $\theta \in \Theta$, this set may be empty. The set of all $\theta \in \Theta$ for which this set is nonempty is precisely the identified set of structural parameters.
To construct the identified set in practice, for each $\theta \in \Theta$ we must ask whether there exists a probability measure $Q_{\alpha \mid \bm W}^\dagger \in \mathcal{Q}$ that rationalizes the observed vector of conditional choice probabilities through (ref), $P_{\bm W}-$a.s. Since a probability measure is an infinite-dimensional object, verifying the existence of such a conditional probability measure is an infinite-dimensional existence problem.\footnote{This is a common feature of partially identified models. As far as we know, this terminology was first used by torgovitsky2019partial.} We now illustrate that the structure of the likelihood function $f(\,\cdot\mid \bm w, \alpha; \theta)$ in Assumption (ref) allows us to convert the infinite-dimensional existence problem to a tractable finite-dimensional problem.
Consider Example (ref) with $T = 2$ and $\gamma=0$ (i.e. without covariates).\footnote{The literature on the AR(1) model with $T=2$ is quite sparse. halliday2007testing studies testing for state dependence in an AR(1) model in which covariates are allowed to have time-varying coefficients, but focused on hypothesis testing. chamberlain2023identification proved the impossibility of point identification in the AR(1) model with bounded covariates when $T = 2$. Finally, in dobronyi2021identification, a previous working paper version of the current paper, we derived analytical bounds for the dynamic coefficient.} This simple example helps to illustrate a fundamental connection between identification in models governed by Assumption (ref) and the truncated moment problem in mathematics.\footnote{See schmudgen2017moment for a recent textbook treatment.} We use this simple example to provide the intuition for our approach before presenting our general identification results. This simple case is also interesting in itself: using functional differencing, honore2024moment show that there are no identifying restrictions for the parameter $\beta$. In contrast, we will show that the model still provides information about the structural parameters through a finite set of moment equalities and shape constraints. In particular, conditional on observing $Y_{0} = y_{0}$, the logistic distribution for $\epsilon_{t}$ implies that for any $\bm y \in \{0,1\}^2$: \[ f(\bm y \mid y_0 , \alpha; \theta) =\prod_{t=1}^2 \Lambda(\alpha + \beta y_{t-1})^{y_t} (1-\Lambda(\alpha + \beta y_{t-1}))^{1-y_t}. \] Now let $A := \exp(\alpha)$ and $B := \exp(\beta)$ and choose $\kappa(y_0,\alpha,\beta) =(1-\Lambda(\alpha + \beta y_0)) (1-\Lambda(\alpha)) (1-\Lambda(\alpha + \beta))$. Then we can write the likelihood as:
Relating to (ref) in Assumption (ref), in this example we have $S = 3$, and the entries in the rows of the matrix $\bm G(y_0,\beta)$ represent the coefficients $g_{s}(\bm y, y_0, \beta)$ of the polynomials of $A$ for the history $\bm y \in \mathcal{Y}^2$. Integrating the likelihood from (ref) with respect to any conditional distribution $Q_{\alpha\mid y_0}^\dagger(\alpha \mid y_0)$ for the individual effect yields: \[ \bm G(y_0,\beta)
= \bm G(y_0,\beta)
. \] To arrive at the second equality, we perform the change of measure:
and then let $\bar Q_{A\mid y_0}^\dagger(\,\cdot\,\mid y_0)$ denote the push-forward measure of $\bar{Q}_{\alpha\mid y_0}^\dagger(\,\cdot\,\mid y_0)$ under the map $\alpha\mapsto \exp(\alpha)$.\footnote{While this measure depends on the unknown parameters, the proof of our main identification results show that this fact has no identifying power for the structural parameters.} Since $\kappa(y_0,\alpha,\beta)$ is bounded and positive for all $\alpha \in \mathbb{R}$ by Assumption (ref), the measure $\bar Q_{A\mid y_0}^\dagger(\,\cdot\,\mid y_0)$ is a finite nonnegative Borel measure on $(\mathbb{R}_{+},\mathcal{B}(\mathbb{R}_{+}))$.\footnote{The fact that $\bar Q_{A\mid y_0}^\dagger(\emptyset \mid y_0)=0$ is obvious. Countable additivity follows by dominated convergence. } Now define the vector:
Then $\bm{r}(y_0)$ is a vector of moments of the variable $A$ up to order $3$ with respect to the measure $\bar Q_{A\mid y_0}^\dagger(A\mid y_0)$. We refer to $\bm{r}(y_0)$ as the vector of generalized moments of $\alpha$ throughout. Now let $\bm p(y_0)$ denote the vector of conditional probabilities $p(\bm y \mid y_{0})$ stacked across $\bm y \in \mathcal{Y}^2$.\footnote{The ordering of the choice sequence should match the order in (ref). We maintain a consistent ordering of choice sequences throughout: when the time period increases by one, we always append $0$ to all existing choice sequences, and then append $1$. } Then the question of whether a particular $\beta$ belongs to the identified set is equivalent to the question of whether, for each $y_0 \in \{0,1\}$, there exists a measure---specifically, a nonnegative Radon measure---whose moment vector $\bm{r}(y_0)$ satisfies $\bm p(y_0) = \bm G(y_0,\beta)\bm{r}(y_0)$.\footnote{When specialized to Euclidean space, a Radon measure is a nonnegative Borel measure that is finite on all compact sets. On Euclidean space, all finite nonnegative Borel measures are Radon, although not all Radon measures are finite measures; for example, the Lebesgue measure is a Radon measure. } This result reveals a fundamental connection between the identification of structural parameters in dynamic logit models and the moment problem from the mathematics literature.\footnote{See karlin1966tchebycheff, krein1977markov, and schmudgen2017moment for comprehensive treatments of this subject.} One of the main questions studied in the literature on the moment problem is whether there exists a Radon measure that rationalizes a sequence of real numbers as its moments. Given an infinite sequence of real numbers, this problem is referred to as the full moment problem. Given a finite sequence of real numbers, this problem is referred to as the truncated moment problem. When the Radon measure is restricted to have support on $\mathbb{R}_+$, as in our context, the truncated moment problem is known as the truncated Stieltjes moment problem, as it was first raised and analyzed by stieltjes1894recherches.
Let $\mathcal{P}_{+}$ denote the set of all nonnegative Radon measures on $(\mathbb{R}_{+},\mathcal{B}(\mathbb{R}_{+}))$, and define the moment space:
Referring back to Definition (ref), for the AR(1) model with $T=2$ we can write the identified set as:
This characterization of the identified set is not useful without a tractable means of verifying whether a vector $\bm{r}(y_0)$ belongs to the moment space $\mathcal{M}_S$ from (ref). However, the geometric structure of the moment space $\mathcal{M}_S$ has been studied extensively, and results from the literature on the moment problem lead to the following theorem.
Theorem (ref) uses Theorem 5.1 in curto1991recursiveness, with parts $(ii)$ and $(iii)$ providing a means of verifying whether the vectors $\bm{r}(0)$ and $\bm{r}(1)$ belong to the moment space $\mathcal{M}_{S}$ when $S=3$.\footnote{See also Theorems 9.35 and 9.36 in schmudgen2017moment.} To understand condition $(ii)$, the key insight is that the moment space $\mathcal{M}_{S}$ is a convex cone. As such, it has an associated dual cone given by:
Theorem II 9.1 in karlin1966tchebycheff derives the specific form of the dual cone, and shows that $\mathcal{M}_S^* = \mathcal{P}_{S}$, where:
In particular, $\mathcal{P}_{S}$ is the set of coefficients that produce a nonnegative polynomial on $\mathbb{R}_{+}$. By standard results in convex analysis, taking the dual of the dual cone $\mathcal{M}_{S}^*$ again recovers the closure of the moment space $\mathcal{M}_S$; that is, $(\mathcal{M}_{S}^*)^* = \text{cl}(\mathcal{M}_{S})$. Since $\mathcal{M}_{S}^* = \mathcal{P}_{S}$, the dual of the dual cone is:
Thus, $\bm{r}(0)$ and $\bm{r}(1)$ belong to $\text{cl}(\mathcal{M}_{S})$ if and only if they satisfy condition $(ii)$ in Theorem (ref).
To see how to check condition $(ii)$ from Theorem (ref) in practice, consider the case when $y_0 = 0$. Every nonnegative polynomial of $A$ with an odd degree $2m+1$ for some $m\in \mathbb{N}$ has a representation of the form:\footnote{For the even case, $\sum_{j = 0}^{2m} \eta_j A^j = f^2(A) + A q^2(A)$ where $f(A)$ are polynomials of A of at most order $m$ and $q(A)$ is a polynomial of $A$ of at most order $m-1$. See Corollary 8.1 in Chapter V of karlin1966tchebycheff and the further discussion in Section 10 of Chapter V. Also see Corollary 3.5 of schmudgen2017moment.} \[ \sum_{j=0}^{2m+1} \eta_{0,j} A^j= A f^2(A) + q^2(A)\geq 0, \] for all $A\in [0, \infty)$, where $f(A)$ and $q(A)$ are polynomials up to order $m$. In our AR(1) example with $T=2$, $S = 2m+1 = 3$, and thus $f(A)$ and $q(A)$ are polynomials of at most degree 1. Therefore, nonnegativity implies that we can write $f(A) = \xi_0 + \xi_1 A$ and $q(A) = \lambda_0 + \lambda_1 A$ for any coefficients $(\xi_0, \xi_1)$ and $(\lambda_0, \lambda_1)$ satisfying: \[ \sum_{j = 0}^{3} \eta_{0,j} A^j = A(\xi_0 + \xi_1A)^2 + (\lambda_0 + \lambda_1A)^2 \geq 0. \] Retrieving the corresponding coefficients $\eta_{0,j}$, the condition $\sum_{j=0}^3 \eta_{0,j} r_{j}(0) \geq 0$ requires:
which can be equivalently stated as:
for all coefficients $(\lambda_0, \lambda_1)$ and $(\xi_0, \xi_1)$. This condition is equivalent to checking that the two square matrices in (ref), defined using the elements of $\bm r(0)$, are positive semidefinite. These matrices are known as Hankel matrices in the truncated moment problem literature.\footnote{See Section 3.2 in schmudgen2017moment.}
Note that condition $(ii)$ ensures only that $\bm{r}(0)$ and $\bm{r}(1)$ belong to $\text{cl}(\mathcal{M}_{S})$, and not necessarily to $\mathcal{M}_{S}$.\footnote{To see what can go wrong, consider the vector $\bm{r}(0)^\top = [0,0,0,1]$. Then the matrices in (ref) are positive semidefinite, but clearly $\bm{r}(0)$ cannot be rationalized as a moment vector of a nonnegative Radon measure with support on $\mathbb{R}_{+}$, so that $\bm{r}(0)\notin\mathcal{M}_{S}$. This example is ruled out by condition $(iii)$: there are no coefficients satisfying $r_{3}(0) = a_{0,1} r_{1}(0) + a_{0,2} r_{2}(0)$, showing that $\bm{r}(0)$ cannot be rationalized as a moment vector.} Here condition $(iii)$ plays a role. When conditions $(ii)$ and $(iii)$ are combined, simple linear algebra combined with the discussion above shows that they are equivalent to checking if:
are positive semidefinite for some $\varsigma\geq 0$ (see Lemma 2.3 in curto1991recursiveness). The matrix $\bm H_{1}^*(\bm r, \varsigma)$ is called the Hankel extension of the corresponding Hankel matrix in (ref). Following a similar logic as above, positive semidefiniteness of these matrices is equivalent to:
Theorem V 3.1 in karlin1966tchebycheff then shows that $\text{cl}(\mathcal{M}_{S+1})$ can be expressed as:
so that the closure of the moment space is equal to the original moment space $\mathcal{M}_{S+1}$, but also includes a ray from the origin. Combining (ref) with (ref), we see that condition $(ii)$ and $(iii)$ are equivalent to checking if $(r_{0}(0),r_{1}(0),r_{2}(0),r_{3}(0)) \in \mathcal{M}_{S}$.\footnote{In particular, if $(r_{0}(0),r_{1}(0),r_{2}(0),r_{3}(0),\varsigma) \in \text{cl}(\mathcal{M}_{S+1})$ then there exists a $\lambda\geq 0$ such that $(r_{0}(0),r_{1}(0),r_{2}(0),r_{3}(0),\varsigma-\lambda) \in \mathcal{M}_{S+1}$. But then there exists a measure that supports this vector as its four moments, and this same measure must also support $(r_{0}(0),r_{1}(0),r_{2}(0),r_{3}(0))$ as its first three moments. This implies $(r_{0}(0),r_{1}(0),r_{2}(0),r_{3}(0)) \in \mathcal{M}_{S}$. }
Using Theorem (ref) we see that, in the specific case of the AR(1) model with $T=2$, the identified set can be constructed by checking two conditional moment equalities, and by checking if there exists a constant $\varsigma\in \mathbb{R}$ such that the matrices in (ref) are positive semidefinite.\footnote{For this specific model, it is possible to further derive analytical bounds on the parameter $\beta$ by converting matrix nonnegativity to inequalities on the determinants of all of its principal minors. See dobronyi2021identification.} By making a connection to the moment problem, our approach is able to obtain sharp restrictions on the structural parameters in examples like the AR(1) model with $T=2$ where competing approaches fail to deliver any nontrivial identifying restrictions.\footnote{See the discussion of this model in honore2024moment.} Intuitively, this is because our approach exploits two new facts that have not been considered by other methods: $(i)$ in a large class of models, the fixed effect distribution can be completely summarized by a finite vector of generalized moments, and $(ii)$ this vector of generalized moments must satisfy certain constraints which have nontrivial identifying content for the structural parameters.
While this section was meant to introduce the main assumptions and ideas through a simple example, in the next section we expand on the connection to the truncated moment problem and apply it to a larger class of models.
With the results from the dynamic panel logit model for $T = 2$ and $\gamma = 0$ in hand, we now generalize the identification analysis to all models governed by Assumption (ref). For the following, let $J:=|\mathcal{Y}|^T$, and define:
where $\bm y_1,\ldots,\bm y_{J},$ denotes an enumeration of the support $\mathcal{Y}^T$, and where $g_{s}(\bm y_j,\bm w,\theta)$ are the coefficients from Assumption (ref). The following is the main identification result of the paper.
Theorem (ref) shows that the identified set for the structural parameters $\theta \in \Theta$ for the class of models satisfying Assumption (ref) can be characterized by a set of moment equality conditions imposed on the conditional probabilities, as well as additional semidefinite shape restrictions on the parameter $\bm r(\bm w)$ coming from the moment space restrictions.
The following theorem shows the necessary and sufficient conditions to have $\bm r(\bm w) \in \mathcal{M}_S$, generalizing conditions $(ii)$ and $(iii)$ in Theorem (ref). The result follows from classic results in the moment literature (e.g. curto1991recursiveness). Here we use the notation $\bm A \succeq 0$ to represent the fact that the square matrix $\bm A$ is positive semidefinite.
Theorem (ref) shows that, to check that a vector $\bm r \in \mathbb{R}^{S+1}$ belongs to the moment space $\mathcal{M}_{S}$, it is both necessary and sufficient to check that two matrices are positive semidefinite. Checking if a matrix is positive semidefinite is equivalent to checking that all principle minors of the matrix are nonnegative.\footnote{See meyer2000matrix p.566. Recall that an $r\times r$ principle submatrix of an $n\times n$ matrix $\bm A$ is obtained by deleting the same set of $n-r$ rows and columns from the matrix $\bm A$. The principle minors of a matrix $\bm A$ are the determinants of the principle submatrices of $\bm A$. See meyer2000matrix p.494. } In this sense, the semidefinite restrictions on the matrices from Theorem (ref) can be viewed as nonlinear shape restrictions on the unknown vector of moments $\bm r(\bm w) \in \mathbb{R}^{S+1}$. Combining this idea with Theorem (ref), verifying whether a vector $\theta \in \Theta$ belongs to the identified set amounts to checking whether a certain set of conditional moment equalities hold subject to a set of shape restrictions on $\bm r(\bm w) \in \mathbb{R}^{S+1}$, $P_{\bm W}-$a.s. To formalize this, let $\mathcal{S}_{+}^d$ denote the space of symmetric $d \times d$ positive semidefinite matrices, and define the moment function:
Finally, let $\bm m(\bm y,\bm w,\theta,\bm r)$ denote the $J\times 1$ vector of moment functions of the form (ref) stacked across $j=1,\ldots,J$, and let $L^0(\mathcal{E}_1,\mathcal{E}_2)$ denote the set of all measurable functions from $\mathcal{E}_1$ to $\mathcal{E}_2$, where $\mathcal{E}_1$ and $\mathcal{E}_2$ are (subsets of) Euclidean space equipped with the Borel $\sigma-$algebra. The following is a simple corollary of Theorems (ref) and (ref).
In certain special cases, one can say more about the structure of the identified set. For instance, in the AR(1) model with $T=2$ and no covariates, the identified set for the dynamic coefficient is convex. For the AR(1) model with $T=3$ where the sole covariate is a time trend, the identified set for the dynamic coefficient can consist of at most two isolated points.\footnote{See dobronyi2021identification for these results.} In practice, we can check whether a given $\theta \in \Theta$ belongs to the identified set by solving a semidefinite program. To see this, for now consider the case when $\mathcal{W} = \{\bm w_1,\ldots, \bm w_{L} \}$ is finite, and $S=2m+1$ for some $m \in \mathbb{N}$ (i.e. $S$ is odd). Now consider the following optimization problem:
Both constraints $(i)$ and $(iii)$ in (ref) can be written as semidefinite constraints, which enforce the positive semidefiniteness of a matrix.\footnote{If $\bm \xi = (\xi_{11},\ldots,\xi_{JL})^\top$, then:
} The constraints in $(ii)$ are linear constraints. This makes the program (ref) a semidefinite program.\footnote{In general, semidefinite programs are programs that involve optimizing a linear objective function subject to linear constraints and semidefinite constraints. For an introduction see Section 4.6 in boyd2004convex, or Chapter 3 in ben2001lectures. } Semidefinite programs are convex optimization problems, are a special case of conic programs, and can be solved quickly and reliably with most commercially available solvers.\footnote{All computational results presented in this paper were obtained using the MOSEK interface in R. } For instance, for the AR(1) model with $T=2$ studied in the previous section, the average time to solve the corresponding semidefinite program is approximately $0.0029$ seconds. Average computational times for other models can be found in Section (ref) of the Online Supplementary Material. It is straightforward to see that, in the case when $\mathcal{W} = \{\bm w_1,\ldots, \bm w_{L} \}$, by Corollary (ref) we have $\theta \in \Theta_{I}(P)$ if and only if $\text{val}$((ref))$=0$. In Section (ref) we propose an estimator that replaces the population moment conditions in constraint $(ii)$ of the program (ref) with their sample analogs, and we study consistency and propose a method of inference. We also show how to extend the semidefinite programming approach introduced above to cases where $\bm W$ may be continuous or discrete.
In addition to providing a tractable representation of the identified set of structural parameters, our approach can be used when the researcher's parameter of interest is a functional of the distribution of latent individual effects. In particular, let $\psi: \mathcal{W} \times \mathbb{R} \times \Theta \to \mathbb{R}$ be a function of the form:
for some known sequence of coefficients $\bm{\eta}(\bm w, \theta) := (\eta_{0}(\bm w, \theta), \ldots, \eta_{S}(\bm w, \theta))^\top$, where $\kappa(\bm w,\alpha,\theta)$ is as in (ref). Then the function $\psi(\bm w,\alpha, \theta)$ is sum of polynomials with the same order and the same factor $\kappa(\bm w,\alpha,\theta)$ as the likelihood in (ref) from Assumption (ref). Now suppose the researcher's parameter of interest is:
for $\bm w \in \mathcal{W}$. Given this representation for $\Psi(\bm w,\theta_0)$, and given the form of $\psi(\bm w,\alpha, \theta_0)$ from (ref), for any given $\bm w \in \mathcal{W}$ we have:
for a known (up to $\theta_0$) vector $\bm{\eta}(\bm w, \theta_0)$. As we will show, a number of interesting functionals, including the average marginal effect of the lagged choice, can be written in this form. For the AR(1) model in Example (ref), the point identification of the average marginal effect of lagged choice was first discovered by aguirregabiria2024identification. Our results generalize to other functionals of the form (ref) for models satisfying Assumption (ref), and also allow for partial identification. The ability to bound functionals is also an advantage of our method over existing approaches like conditional maximum likelihood and functional differencing.
Note that if both $\theta\in \Theta$ and $\bm r(\bm w)$ are point-identified, then $\Psi( \bm w,\theta)$ is point-identified. Furthermore, point-identification of $\Psi( \bm w,\theta)$ can often be easily established using our framework.
Proposition (ref) provides a simple sufficient condition for point identification of the functional $\Psi(\bm w, \theta)$ that can be used even when the conditional distribution $Q_{\alpha \mid \bm W}$ is not point-identified. In particular, if $\theta\in\Theta$ is point-identified and $\bm G(\bm w, \theta_{0})$ has full column rank, then the generalized moments $\bm r(\bm w)$ are point-identified from the equation $\bm p(\bm w) = \bm G(\bm w, \theta_{0})\bm r(\bm w)$. Point identification of $\Psi(\bm w,\theta)$ then follows from (ref). We illustrate how to use this result in the examples ahead, which include functionals like the average marginal effect and the average structural function in the AR(1) model.
\setcounter{example}{0}
\setcounter{example}{0}
\setcounter{example}{0}
In the general case, the identified set for $\Psi(\bm w, \theta)$ can also be constructed using semidefinite programming. To see this, consider again the simplified case when $\mathcal{W} = \{\bm w_1,\ldots, \bm w_{L} \}$ is finite and $S=2m+1$ for some $m \in \mathbb{N}$ (i.e. $S$ is odd). Let $\bm w \in \mathcal{W}$ be some value, and consider the following optimization problem:
Compared to program (ref) introduced earlier, the program (ref) includes the additional constraint $(iv)$, and also adds an additional parameter $\xi_{\Psi}$ to constraint $(i)$. Since constraint $(iv)$ is linear in $\bm r(\bm w)$, the program (ref) remains a semidefinite program. It is straightforward to see that, in the case when $\mathcal{W} = \{\bm w_1,\ldots, \bm w_{L} \}$, the pair $(\theta,\Psi)$ belongs to the identified set if and only if $\text{val}$((ref))$=0$. The approach introduced in Section (ref) can also be used to extend the semidefinite program introduced here to cases where $\bm W$ may be continuous or discrete.
Functional differencing was proposed by johnson2004identification and bonhomme2012functional and recently used by honore2024moment, honore2025dynamic and davezies2023fixed. This method aims to find a vector of nonzero moment functions $\bm h(\,\cdot\,,\theta): \mathcal{Y}^{T} \times \mathcal{W} \to \mathbb{R}^{d_h}$ that satisfy:
$P_{\bm W}-$a.s.\ for all $\alpha$.\footnote{Note the number of moment functions $d_{h}$ is typically not known ahead of time. } Appealing to the discrete nature of $\bm Y$ under Assumption (ref), we can rewrite the moment conditions in (ref) as:
If (ref) holds for all $\alpha \in \R$, then it holds regardless of the true distribution of fixed effects. Provided the functions $\bm h(\,\cdot\,,\theta)$ are known, they can be used to obtain valid moment conditions to identify $\theta \in \Theta$. In particular, let $\bm f(\bm w, \alpha; \theta)$ denote the $J\times 1$ vector that stacks the likelihood function $f(\bm y\mid \bm w, \alpha; \theta)$ across all $\bm y\in \mathcal{Y}^T$. Then the set of moment functions that satisfy (ref) are given by:\footnote{Without loss of generality, we focus on finding moment functions that satisfy (ref) for all $(\bm w, \alpha)$, rather than $P_{\bm W}-$a.s.\ for all $\alpha$.}
Connecting with Assumption (ref), it is also clear that the collection of conditional moment functions satisfy $\bm h(\bm W, \theta)^\top \bm p(\bm W) = 0$ a.s. The challenge of using functional differencing lies in finding the functions $\bm h(\,\cdot\,, \theta)$. In some cases, these functions can be constructed numerically with the aid of a computer (see a detailed procedure in honore2024moment). However, these functions need to be found model-by-model and for each specific $T$.
In order to better compare our approach with functional differencing, we first provide a unified analytical method to find these functions for any model that has a likelihood function satisfying Assumption (ref).
Intuitively, Theorem (ref) suggests that the left null space of $\bm G(\bm w,\theta)$ provides a basis that spans the set $\bm D(\theta)$. Since $\bm G(\bm w,\theta)$ is a known matrix for fixed $\theta \in \Theta$ and $\bm w \in \mathcal{W}$, constructing a basis for the left null space can be done analytically, or by using symbolic computation with the aid of a computer.\footnote{Note that, as with the procedure of honore2024moment, there is no guarantee that all moment conditions are functions of $\theta$, and so some may be uninformative.} Checking whether (ref) holds at $\theta\in\Theta$ is then equivalent to checking if $\bm v(\bm w, \theta)^\top \bm p(\bm w)=0$ for all basis vectors $\bm v(\bm w, \theta)$ in the left null space of $\bm G(\bm w,\theta)$.
This connection provides additional insight into some results obtained earlier in the literature. For example, in the AR(1) model from Example (ref) with $T = 2$ and $\gamma = 0$, the $4\times 4$ matrix $\bm G(\bm w,\theta)$ has full rank for each $(\bm w,\theta)$, so that its left null space consists only of the zero vector. This explains why there are no moment conditions for $\beta$ using the functional differencing approach, a result reported by honore2024moment. Despite this, our approach still delivers identifying restrictions through the constraints $\bm p(\bm w) = \bm G(\bm w, \theta) \bm r(\bm w)$ and $\bm r(\bm w) \in\mathcal{M}_{S}$.
As another example of how Theorem (ref) can be helpful, consider the AR(1) model from Example (ref) with general $T$ and $\gamma=0$. For this model, honore2024moment find $2^T - 2T$ linearly independent moment conditions using a numerical search method, and they conjecture that these are all the moment conditions available. To use the approach suggested by Theorem (ref), first note that the matrix $\bm G(\bm w, \theta)$ is of dimension $2^T \times 2T$ and has full column rank.\footnote{For details, see the Additional Online Supplementary Material, which can be accessed \href{https://arxiv.org/abs/2104.04590}{here}.} Therefore, the left null space of $\bm G(\bm w, \theta)$ provides a basis with exactly $2^T - 2T$ linearly independent moment conditions, verifying the conjecture of honore2024moment. This result is also useful since ex ante it is not known how many linearly independent moment functions exist when using functional differencing. Our result suggests that honore2024moment have indeed found all the relevant moment functions.
As a final example, consider the AR(1) dynamic ordered logit model from Example (ref) with $M$ choice options and $T$ periods. The corresponding matrix $\bm G(\bm w, \theta)$ has dimension $M^T \times ((T-1)M^2 - (T-2)M)$ and is of full column rank.\footnote{For details, see the Additional Online Supplementary Material, which can be accessed \href{https://arxiv.org/abs/2104.04590}{here}.} Theorem (ref) thus confirms the conjecture made in honore2025dynamic that there are $M^T - (T-1)M^2 + (T-2)M$ linearly independent moment conditions available in this model.
Using Theorem (ref), the difference between functional differencing and our approach can be explained geometrically. For a fixed $(\bm w,\theta) \in \mathcal{W}\times\Theta$, let $\bm p_{\bm G}(\bm w)$ denote the projection of the choice probability vector $\bm p(\bm w)$ onto the column space of $\bm G(\bm w,\theta)$, and let $\bm r^*(\bm w,\theta)\in \mathcal{M}_{S}$ denote the vector that minimizes $|| \bm p(\bm w) - \bm G(\bm w, \theta) \bm r||$ over all $\bm r \in \mathcal{M}_{S}$. Note by Theorem (ref) we have $\theta\in \Theta_{I}(P)$ if and only if $||\bm p(\bm w) - \bm G(\bm w, \theta) \bm r^*(\bm w,\theta)||=0$, $P_{\bm W}-$a.s.\ It is straightforward to show that the vectors $\bm p(\bm w) - \bm p_{\bm G}(\bm w)$ and $\bm p_{\bm G}(\bm w) - \bm G(\bm w, \theta) \bm r^*(\bm w,\theta)$ are orthogonal, so that by Pythagoras' Theorem:\footnote{In particular, $\bm p(\bm w) - \bm p_{\bm G}(\bm w)$ is the least-squares residual, which lies in the null space of $\bm G(\bm w, \theta)^\top$, and so is orthogonal to the column space of $\bm G(\bm w, \theta)$. Thus, it is orthogonal to $\bm p_{\bm G}(\bm w) - \bm G(\bm w, \theta) \bm r^*(\bm w,\theta)$, which lies in the column space of $\bm G(\bm w,\theta)$.}
See Figure (ref) for an illustration. Now by Theorem (ref) and its following discussion, functional differencing searches for vectors $\bm v(\bm w, \theta)$ that form a basis for the left nullspace of $\bm G(\bm w,\theta)$, and that are orthogonal to $\bm p(\bm w)$. By the Fundamental Theorem of Linear Algebra, the condition $\bm v(\bm w, \theta)^\top \bm p(\bm w)=0$ holds for all basis vectors $\bm v(\bm w, \theta)$ in the left null space of $\bm G(\bm w,\theta)$ if and only if $\bm p(\bm w)$ lies in the column space of $\bm G(\bm w,\theta)$; that is, if and only if $\bm p(\bm w) = \bm p_{\bm G}(\bm w)$. By this reasoning, functional differencing is equivalent to checking whether term $(i)$ in (ref) is equal to zero, which is a necessary but not sufficient condition to have $\bm p(\bm w) = \bm G(\bm w, \theta) \bm r^*(\bm w,\theta)$ under the constraint $\bm r^*(\bm w,\theta) \in \mathcal{M}_S$. In contrast, our approach requires that both terms $(i)$ and $(ii)$ in (ref) are equal to zero. Seen in this way, functional differencing misses a piece of the orthogonal decomposition of $\bm p(\bm w) - \bm G(\bm w, \theta)\bm r^*(\bm w,\theta)$, and as a result it generally fails to pick up all relevant identifying restrictions.
In addition to providing a general approach to identification and allowing us to bound functionals of the distribution of the latent individual effects, our procedure delivers the sharp identified set even when there are no moment conditions available using functional differencing, it provides sharp bounds in cases where the functional differencing approach cannot, and it allows us to test for model misspecification.\footnote{Even when the structural parameters are point-identified from the functional differencing moment conditions, in some cases adding additional (binding) constraints on the model parameters can reduce asymptotic mean squared error. This was shown for the empirical likelihood estimator and the GMM estimator with an optimal weighting matrix by moon2009estimation in the specific case when the model parameters are point-identified by a set of moment equalities and the researcher has access to a single additional (drifting-to-)binding moment inequality. } We now illustrate these points using examples.
\setcounter{example}{0}
\setcounter{example}{0}
\setcounter{example}{0}
\setcounter{example}{2}
While our main results concern identification, in this section we propose a consistent estimator of the identified set that is applicable when the structural parameters are either point- or partially-identified, and we also propose an inference procedure. Our estimation and inference procedure allow for both discrete and continuous covariates, and our inference procedure is based on the procedure of chernozhukov2023constrained (CNS hereafter). The CNS procedure is designed for inference on (possibly infinite-dimensional) shape-constrained parameters in models defined by conditional moments, and allows for both point and partial identification. CNS also allows for a general class of shape constraints defined by equality and inequality restrictions. In our setting, the relevant shape constraints are on the moment vectors, since by Theorem (ref) any valid moment vector must be such that the Hankel matrix and its Hankel extension are positive semidefinite. Note that positive semidefiniteness of a matrix is equivalent to nonnegativity of the determinants of all of its principal minors. Thus, positive semidefiniteness of a matrix can be enforced by imposing certain nonlinear inequality constraints on the entries of the matrix, connecting our setting to the shape constraints allowed by CNS. However, in order to maintain the semidefinite programming structure discussed in the previous section, we use a conservative implementation of their procedure.\footnote{In the notation of CNS, we set $r_{n} = +\infty$, and take $\hat{V}_{n}(\theta,R \mid \ell_n)=\{\bm 0\}$. We also set the weighting matrix as the identity matrix. These are always feasible (but potentially conservative) choices. These choices greatly simplify computation by avoiding the need to optimize over the set $\hat{V}_{n}(\theta,R \mid \ell_n)$ in our bootstrap procedure, which would otherwise destroy the semidefinite programming structure of our bootstrap test statistic. Avoiding this minimization in the bootstrap test statistic leads to a larger-than-necessary critical value, but keeps the procedure tractable. } In addition to providing substantial computational gains, our simplified implementation also allows us to use a weaker set of assumptions than those provided in CNS. We outline this weaker set of assumptions in Appendix (ref). To keep notation simple, we focus on providing results for the identified set of structural parameters, although our approach extends to the functionals from Section (ref) under minimal additional assumptions.
Recall from Corollary (ref) and equation (ref) that the model constraints $\bm p(\bm w) = \bm G(\bm w, \theta) \bm r(\bm w)$ can be written as conditional moment equalities of the form:
where $\bm m(\bm Y_{i},\bm W_{i},\theta,\bm r)$ is a $J\times 1$ vector of moment functions with $j^{th}$ element:
While $\bm g(\bm y_j, \bm W_{i},\theta)^\top$ is a known function of the covariates and structural parameters, $\bm r(\bm W_{i})$ is an unknown vector-valued function that must be estimated. Furthermore, from Corollary (ref), we must also impose a number of shape constraints on these functions during estimation. Since the covariates may be continuous or discrete, it is desirable to allow for a flexible specification for the functions $\bm r:\mathcal{W}\to \mathbb{R}^{S+1}$, viewing the moments as a function of the covariates $\bm W_{i}$. Furthermore, the specification for these functions should be amenable to our implementation using semidefinite programming, even when the covariates are continuous. With these concerns in mind, we recommend a simple sieve approximation based on piecewise constant functions.\footnote{While this choice simplifies computation significantly, it is not necessary: researchers interested in other methods of approximation can consult Appendix (ref) for the minimal set of assumptions required by our procedure.}
Let $\mathcal{D}_{l_n}$ denote a growing partition of $\mathcal{W}$ into $l_n$ disjoint sets, and assume the sequence $\{\mathcal{D}_{l_n}\}_{n=1}^{\infty}$ is nested for all but finitely many $n\geq 1$. Now let $\mathcal{C}_{n}(\overline{\delta})$ denote the following set of functions:
Note that $\mathcal{C}_{n}(\overline{\delta})$ is the class of piecewise constant functions with uniformly bounded coefficients. Using this collection, we define a sieve for the functions $\bm r(\,\cdot\,):\mathcal{W}\to \mathbb{R}^{S+1}$ using all vector-valued functions whose elements are piecewise constant functions on the partition $\mathcal{D}_{l_n}$:
Note that $\mathcal{R}_{n}$ is the set of all piecewise constant vector-valued functions of the form $\bm r_{n}(\bm w) = \sum_{D \in \mathcal{D}_{l_n}} \bm \delta_{D}\cdot 1\{\bm w \in D\}$, where $\bm \delta_{D} \in [0,\overline{\delta}]^{S+1}$.\footnote{In many examples, $\overline{\delta}<\infty$ is guaranteed whenever $\mathcal{W}$ and $\Theta$ are compact.} Finally, let $\mathcal{R}$ denote the set of all functions that can be approximated as uniform limits of the sequences $\bm r_n \in \mathcal{R}_n$:\footnote{Other choices of the norm are possible.}
Then $\mathcal{R}$ is a subset of a Banach space, although the precise properties of $\mathcal{R}$ will depend on the sequence of partitions $\{\mathcal{D}_{l_n}\}_{n=1}^{\infty}$ chosen by the researcher.
Now since the model is characterized in terms of conditional moment equalities, we first convert the conditional moments into unconditional moments using instrument functions. In particular, given a collection $\mathcal{D}_{k_n}$ of $k_n$ Borel subsets of $\mathcal{W}$, define the $k_{n}\times 1$ vector of instrument functions:
For any such partition, the $J\times 1$ vector of conditional moment equalities of the form (ref) imply the following set of $J\cdot k_n \times 1$ vector of unconditional moment equalities:
With continuous covariates, we will generally require $k_n\uparrow \infty$ as $n\to\infty$. Now given an i.i.d.\ sample $\{(\bm Y_{i}, \bm W_{i})\}_{i=1}^{n}$, our estimator of the identified set is based on the minimizers of the following criterion function:
In particular, define the following set of shape restrictions:
Here $\varsigma^*(\bm w)$ is any choice that ensures either $\bm H_{m}^*(\bm r(\bm w),\varsigma^*(\bm w))\succeq 0$ (when $S$ is odd) or $\bm B_{m}^*(\bm r(\bm w),\varsigma^*(\bm w)) \succeq 0$ (when $S$ is even) whenever possible given a fixed $\bm r(\bm w)$.\footnote{Such a choice is always possible: see Lemma 2.3 in curto1991recursiveness. For theoretical purposes, it is convenient to view $\varsigma^*(\bm w)$ as a deterministic function of $\bm r(\bm w)$.} If $\bm r \in \mathcal{R}$, then the joint identified set for $(\theta,\bm r)$ is given by:
Note that $\Theta_{I}(P)=\text{Proj}_{\Theta}(\mathcal{I}^*(P))$ is exactly the projection of $\mathcal{I}^*(P)$ onto $\Theta$. Our estimator for the joint identified set for $(\theta,\bm r)$ is given by:
where $\tau_{n}\downarrow 0$ is a sequence of constants (see Remark (ref)). Now let $\hat{\Theta}_{I,n}=\text{Proj}_{\Theta}(\hat{\mathcal{I}}_{n})$ denote the corresponding projection of $\hat{\mathcal{I}}_{n}$ on $\Theta$. This set can be written as:
where $\Pi_{\mathcal{R}_n}(\mathcal{S}):=\{ \bm r \in \mathcal{R}_n : \exists \theta \in \Theta \text{ s.t. }(\theta,\bm r)\in \mathcal{S} \}$. The set $\hat{\Theta}_{I,n}$ represents our estimator for the identified set $\Theta_{I}(P)$.
Our next result shows that the set estimator $\hat{\Theta}_{I,n}$ is consistent for the identified set $\Theta_{I}(P)$ in the Hausdorff metric, uniformly over a certain class of data generating processes (DGPs).\footnote{Recall the Hausdorff distance between two sets $A$ and $B$ is given by:
} Before introducing our result, we require two additional assumptions. In the following, let $\mathcal{P}$ denote a subset of the set of all distributions on $\mathcal{Y}^T\times \mathcal{W}$, and for each element $(\theta,\bm r)\in \Theta\times \mathcal{R}$ let $\Pi_{n}(\theta,\bm r)$ denote its approximation on $\Theta\times \mathcal{R}_{n}$ and define $\mathcal{I}_{n}^*(P) := \left\{ \Pi_{n}(\theta,\bm r) : (\theta,\bm r)\in\mathcal{I}^*(P) \right\}$.
Assumption (ref)$(i)-(iv)$ are straightforward. Assumption (ref)$(v)$ implies the “asymptotic unbiasedness” condition required in CNS.\footnote{$\beta_a = 1/2$ is assumed in all of CNS's examples: see CNS Assumption 4.1$(iv)$ (heterogeneity and demand analysis), Assumption A.2.8$(iv)$ (consumer demand), and Assumption A.2.14$(iii)$ (quantile treatment effects).} It can be seen as a condition on the quality of the sieve space, imposing the restriction that the true (but unknown) vector of moment functions $\bm r \in \mathcal{R}$ is well-approximated by piecewise constant functions. It holds trivially if regressors are discrete, but otherwise depends on the chosen sequence $\{\mathcal{D}_{l_n}\}_{n=1}^{\infty}$ and the properties of $\bm r \in \mathcal{R}$.
For the next assumption, let $\vec{d}_{H}(A,B) = \sup_{a \in A} \inf_{b \in B}||a-b||$ denote the directed Hausdorff distance, and set:
That is, $Q_{n,P}(\theta,\bm r)$ is the analog of $Q_{n}(\theta,\bm r)$ when the sample moment conditions have been replaced by their population versions.
Assumption (ref) is similar to the standard polynomial minorant condition typically imposed in set estimation problems, going back to chernozhukov2007estimation (see their Condition C.2).\footnote{See also kaido2022constraint for an extensive discussion of this condition.} Intuitively, it requires that the criterion function (ref) “lifts off” sufficiently fast in a neighborhood of the identified set. However, Assumption (ref) is stronger than the typical polynomial minorant condition, since it imposes constraints on both the quality of the sieve $\mathcal{R}_{n}$ and the strength of identification associated with the instrument functions. In general the condition depends on the interaction between the instrument functions and piecewise constant functions at the population level, and rules out weak identification. In certain cases simple sufficient conditions can be developed.\footnote{ For instance, in the AR(1) model with $T=3$ and no covariates, this condition is satisfied if all the choices probabilities are bounded away from zero. In the AR(1) model with $T=2$ from Section (ref), the condition is satisfied if $\theta=\beta$ is bounded away from zero, and if certain degenerate distributions are ruled out for $\alpha_i$. For details, see the Additional Online Supplementary Material, which can be accessed \href{https://arxiv.org/abs/2104.04590}{here}.} For added flexibility, an alternative assumption, which can be used to replace Assumption (ref), is presented in Section (ref) of the Online Supplementary Material.
Under these additional assumptions, we have the following consistency result.
Theorem (ref) shows that our estimate of the identified set, given by (ref), converges to the true identified set in the Hausdorff distance uniformly over the class of DGPs $\mathcal{P}$ implicitly defined by Assumptions (ref), (ref) and (ref). Consistency requires that the sequence $\tau_{n}$ in (ref) tends to zero sufficiently slowly relative to the sample size and the number of instrument functions. We provide guidance on all tuning parameters at the end of this section.
As mentioned previously, our estimate of the identified set can be computed efficiently using semidefinite programming. In particular, let $\mathcal{D}_{l_n}:=\{D_1,\ldots,D_{l_n}\}$ and $\mathcal{D}_{k_n}:=\{D_1',\ldots,D_{k_n}'\}$. Since $\bm r \in \mathcal{R}_{n}$ implies that $\bm r(\bm w) = \sum_{\ell =1}^{l_n} \bm \delta_{\ell}\cdot 1\{\bm w \in D_{\ell}\}$ for some vector of coefficients $\{\bm \delta_{\ell}\}_{\ell=1}^{l_n}$, for each $j=1,\ldots,J,$ we have:
where the $\bm \delta_{k,\ell}$'s are auxiliary parameter satisfying the constraints $\bm \delta_{1,\ell} = \ldots = \bm \delta_{k_n,\ell}$ for $\ell = 1, \ldots, l_n$. Now the semidefinite constraints $\bm B_{m}(\bm r(\bm W_{i})) \in \mathcal{S}_+^{m+1}$ and $\bm H_{m}^*(\bm r(\bm W_{i}),\varsigma(\bm W_{i})) \in \mathcal{S}_+^{m+2}$ are equivalent to $\bm B_{m}(\bm \delta_{k,\ell}) \in \mathcal{S}_+^{m+1}$ for $\ell=1,\ldots,l_n$ and $\bm H_{m}^*(\bm \delta_{k,\ell},\varsigma_{k,\ell}) \in \mathcal{S}_+^{m+2}$ for $\ell=1,\ldots,l_n$ for some sequence of coefficients $\varsigma_{k,1},\ldots, \varsigma_{k,l_n}$. Let $\bm \zeta_{k} = (\zeta_{j,k})_{j=1}^{J}$ denote a vector for $k=1,\ldots, k_n$. Then for each $\theta \in \Theta$, for both continuous and discrete covariates minimizing $Q_{n}(\theta,\bm r)$ over $\bm r \in \Pi_{\mathcal{R}_n}(\mathcal{S})$ can be accomplished by solving the optimization problem:
The constraints in $(1)$ and $(3)$ are semidefinite constraints, and the constraints in $(2)$ and $(4)$ are linear constraints. This ensures that the program (ref) is a semidefinite program, which can be computed efficiently for each fixed $\theta \in \Theta$. Minimizing $Q_{n}(\theta,\bm r)$ over all $(\theta,\bm r) \in (\Theta \times \mathcal{R}_{n})\cap \mathcal{S}$ can then be accomplished by establishing a fine grid of evaluation points $\Theta^\dagger\subset \Theta$, solving (ref) at each $\theta\in \Theta^\dagger$, and then choosing the minimizing pair $(\theta,\bm r) \in (\Theta^\dagger \times \mathcal{R}_{n})\cap \mathcal{S}$. An estimate of the identified set can then be obtained by collecting all points $\Theta^\dagger$ satisfying the condition in (ref). This procedure is summarized in Algorithm (ref) at the end of the next subsection.
Building on the results of the previous subsection, in this section we propose a method of confidence set construction using hypothesis test inversion. In particular, define the following slightly revised set $\mathcal{S}(\vartheta)$ representing the shape restrictions:
Note that $\mathcal{S}(\vartheta)$ is the same as $\mathcal{S}$, but also has the additional restrictions that $\theta = \vartheta$ for some vector $\vartheta \in \Theta$. To construct a confidence set for $\theta$, we then invert the following hypothesis test:
where:
That is, the null hypothesis in (ref) tests whether there exists an $\bm r \in \mathcal{R}$ that satisfies all the moment conditions and semidefinite constraints when $\theta = \vartheta$. This will be the case if and only if $\vartheta \in \Theta_{I}(P)$, so that (ref) is equivalent to testing if $\vartheta \in \Theta_{I}(P)$. Due to the shape constraints on $\bm r \in \mathcal{R}$, we require an inference procedure that is valid under shape constraints, and we use a modified version of a procedure proposed by CNS. In particular, to test the null hypothesis from (ref), we propose the following test statistic:
where $Q_{n}(\theta,\bm r)$ is as in (ref). Our rejection decision is then based on comparing $T_{n}(\vartheta)$ to a critical value constructed using a multiplier bootstrap procedure. For i.i.d.\ $\{\xi_{i}^{b}\}_{i=1}^{n}$ with $\xi_{i}^{b}\sim N(0,1)$ independent of $\{(\bm Y_{i},\bm W_{i})\}_{i=1}^{n}$, define the multiplier bootstrap process:
Then our bootstrap test statistic is given by:
where:
for some sequence $\tilde{\tau}_n = o(1)$ satisfying $\tilde{\tau}_n\leq \tau_n$. At level $\alpha$, our rejection decision is based on whether $T_{n}(\vartheta)$ exceeds the $1-\alpha+\delta$ quantile of the bootstrap distribution of $T_{n}^{b}(\vartheta)$, where $\delta$ is some infinitesimal constant.\footnote{The inclusion of $\delta$ allows us to avoid high-level assumptions on the continuity of the asymptotic distribution of $T_{n}(\vartheta)$ under the null. andrews2013inference recommend $\delta=10^{-6}$.} Similar to estimation, the test statistic and bootstrap test statistic can be computed by solving a semidefinite program, which is demonstrated at the end of this section.
To introduce our next result, we require one final technical assumption to replace Assumption (ref). In the following, let $\{\tilde{D}_k\}_{k=1}^{\tilde{k}_n}$ denote an enumeration of all nonempty sets $D_{l}\cap D_{k}'$ where $D_{l} \in \mathcal{D}_{l_n}$ is any set used in the construction of the piecewise constant functions, and $D_{k}'\in \mathcal{D}_{k_n}$ is any set used in the construction the instruments. Now define:
which is a $(k_n+\tilde{k}_n)\times 1$ vector with components:
As illustrated at the end of Section (ref), each moment function $m_{j}(\bm y, \bm w, \theta, \bm r)$ can be written as a linear combination of the elements of the vector $\bm b_{n,j}(\bm y, \bm w,\theta)$ when $\bm r(\bm w)$ is a piecewise constant function. The properties of this vector, and the properties of the instrument vector $\bm q^{k_n}(\bm w)$, play an important role in determining the rate of the bootstrap coupling results in CNS which are crucial for our inference procedure. Define the matrices:
and let $\sigma_{max}(A)$ and $\sigma_{min}(A)$ denote the smallest and largest singular values of a matrix $A$, respectively.
Assumption (ref) replaces Assumption (ref) for our next result. Although it is not immediately obvious, Assumption (ref)$(i)$ is conceptually related to sieve ill-posedness. To see why Assumption (ref)$(i)$ is plausible, note that if $l_n$ and $k_n$ are of the same order and the singular values of $E_{P}[\bm G(\bm W_{i}, \theta) \mid \bm W_{i} \in D_{l}\cap D_{k}']$ are bounded away from zero uniformly in $P\in \mathcal{P}$, then the singular values of $M_{n,P}^{(1)}(\theta)$ and $M_{n,P}^{(2)}$ may be reasonably expected to decay at a rate of $O(k_n^{-1})$. Assumption (ref)$(i)$ comfortably allows for this kind of behaviour.\footnote{These sufficient conditions appear to rule out the AR(1) model with $T=2$ and no covariates from Section (ref) when $\beta=0$, since in this case the $4\times 4$ matrix $\bm G(\bm W_{i}, \theta)= \bm G(y_0, \theta)$ is rank deficient. However, when $\beta=0$ this model becomes the static panel logit model with fixed effects, and this model has a different $4 \times 3$ matrix $\bm G(y_0, \theta)$. Assumption (ref)$(i)$ applies when the researcher uses this alternative matrix when $\beta=0$. } Assumption (ref)$(ii)$ is a technical condition that places additional constraints on the class of DGPs $\mathcal{P}$. It requires that the nullspace of the matrix $M_{n,P}^{(1)}(\theta)$ does not change with $P$. It is trivially satisfied when $\mathcal{P}=\{P\}$ (“pointwise asymptotics”), and admits other possible classes, but it can fail, for instance, for classes $\mathcal{P}$ where the rank of the matrix $M_{n,P}^{(1)}(\theta)$ changes with $P$. Finally, Assumption (ref)$(ii)$ requires that the singular values of $M_{n,P}^{(3)}(\theta)$ are bounded away from infinity, and that they do not decay too fast. Again, if $l_n$ and $k_n$ are of the same order, the singular values of $M_{n,P}^{(3)}(\theta)$ may be reasonably expected to be of the order $O(k_{n}^{-1})$, which is allowed by Assumption (ref)$(ii)$. Note Assumption (ref)$(ii)$ is not required for our approach, but allows us to obtain a faster rate of convergence in the CNS bootstrap coupling result needed in the proofs of our main results, and allows us to maintain the same rate requirements on the sequences $a_{n}$ and $\tau_{n}$ as in Theorem (ref).\footnote{Similar assumptions are used in the leading application in CNS: see CNS Assumption 4.1 and 4.2.} With Assumption (ref) in hand, the following theorem provides the uniform validity of the testing procedure described above.
Theorem (ref) shows the validity of our proposed testing procedure, uniformly over the class of DGPs $\mathcal{P}$ implicitly determined by Assumptions (ref), (ref) and (ref). Using Theorem (ref), confidence sets for $\theta$ can be constructed via hypothesis test inversion by collecting the parameter vectors $\vartheta \in \Theta$ for which we fail to reject the null hypothesis in (ref). In particular, define:
where $\hat{q}_{1-\alpha+\delta}(\theta)$ is as in Theorem (ref). The following is a straightforward immediate consequence of the previous result.
Theorem (ref) and Corollary (ref) justify the testing and inference procedure described above. Combining our approximation based on piecewise constant functions with semidefinite programming provides a computationally efficient means of constructing confidence sets for structural parameters in the models we consider.
To use our inference procedure in practice, we require an efficient method of computing the test statistic $T_{n}(\vartheta)$ and the bootstrap test statistic $T_{n}^{b}(\vartheta)$. Note that computing the test statistic $T_{n}(\vartheta)$ from (ref) is equivalent to solving (ref) at $\theta=\vartheta$ (up to a rescaling by $\sqrt{n}$), so that our previous discussion of (ref) applies to $T_{n}(\vartheta)$. Computing $T_{n}^{b}(\vartheta)$ from (ref) requires only a few small modifications to this procedure. First, the objective function for $T_{n}^{b}(\vartheta)$ is different than $T_{n}(\vartheta)$. However, if $\bm r(\bm w) = \sum_{\ell =1}^{l_n} \bm \delta_{\ell}\cdot 1\{\bm w \in D_{\ell}\}$, some thought shows that (ref) is also linear in the coefficients $\{\bm \delta_{\ell}\}_{\ell=1}^{l_n}$. This makes the objective function for $T_{n}^{b}(\vartheta)$ the norm of a linear function, similar to the objective function for $T_{n}(\vartheta)$. Most of the constraints required to solve (ref) are also identical to those required to compute $T_{n}(\vartheta)$, with the exception that we must also impose the constraint:
The value of the infimum on the right is obtained as a by-product of estimating the identified set. As a result, this constraint can be added to the program as an additional semidefinite constraint. Summarizing, $T_{n}^{b}(\vartheta)$ can be computed by solving the following optimization problem at $\theta=\vartheta$:
Note that constraints $(4)$ and $(5)$ enforce the constraint (ref). Also note that the constraints in $(1)$, $(3)$ and $(4)$ are semidefinite constraints, and the constraints in $(2)$, $(5)$ and $(6)$ are linear constraints. This ensures that the program (ref) is a semidefinite program.
Finally, we note that our proposed bootstrap procedure can be simplified dramatically at the cost of a conservative distortion. In particular, optimization in (ref) can be avoided entirely by “recycling” the optimal vectors $\bm r_1,\ldots, \bm r_{l_n}$ obtained when computing the test statistic by substituting these optimal solutions into the bootstrap test statistic (ref) rather than re-optimizing. Inspecting (ref) and (ref), this makes our test more conservative, but can also dramatically improves computation time, allowing the researcher to trade-off between these two concerns. See marcoux2024simple for a similar procedure. The practical performance of our inference procedure is illustrated in a brief Monte Carlo exercise in Section (ref) of the Online Supplementary Material. In these simulation exercises, and in the application in the next section, we make use of the computational simplifications that come with “recycling” the optimal vectors $\bm r_1,\ldots, \bm r_{l_n}$ obtained when computing the test statistic in the bootstrap procedure.
Our entire estimation and inference procedure for the odd case is provided in Algorithm (ref). A similar algorithm works for the even case by replacing the semidefinite constraints $\bm B_{m}(\bm \delta_{\ell}) \in \mathcal{S}_+^{m+1}\text{ and } \bm H_{m}^*(\bm \delta_{\ell},\varsigma_{0\ell}) \in \mathcal{S}_+^{m+2}$ in (ref) and (ref) with $\bm H_{m}(\bm \delta_{\ell}) \in \mathcal{S}_+^{m+1}\text{ and } \bm B_{m}^*(\bm \delta_{\ell},\varsigma_{0\ell}) \in \mathcal{S}_+^{m+2}$. In terms of tuning parameters, for both estimation and inference the values $\tau_n=0.001 n^{-3/10}$, $\tilde{\tau}_n=0$, $k_n=1+\lceil n^{1/6} \rceil$, and $l_n = k_n-1$ meet all the theoretical requirements, and worked well in both the application in the next section and the Monte Carlo exercises in Section (ref) of the Online Supplementary Material. Setting $\tilde{\tau}_{n}=0$ and recycling the optimal $\bm r_1,\ldots, \bm r_{l_n}$ from the test statistic allows us to save substantial computational time by avoiding the need to re-optimize (ref) during the bootstrap. Researchers who are willing to trade increased computation time for increased testing power can instead set $\tilde{\tau}_n=\tau_n$ and repeatedly solve (ref) when computing the bootstrap test statistic. Finally, similar to the existing literature (e.g. andrews2013inference), the value of $\delta$ in our testing procedure does not play an important role in practice, and can be set arbitrarily small (e.g. $\delta=10^{-6}$).
In this section, we illustrate the proposed identification, estimation and inference procedure by applying it to data from the National Longitudinal Survey of Youth 1997 (NLSY97). The longitudinal surveys are sponsored by the United States Bureau of Labor Statistics with the aim of documenting labor market outcomes over a prolonged period of time. The first round of surveys began in 1997. Here, we use data from the years 2008 to 2010, which we label as periods $t=1,2,3,$ respectively. The outcome variable $Y_{it}$ is a binary variable representing an individual's employment status in a given year, and is equal to $1$ if the respondent worked more than $1000$ hours in year $t$.\footnote{Here we use the same variable definition as honore2024moment, who also use the NLSY97 data.} The value $Y_{i0}$ is defined similarly using data from the year 2007. Throughout we consider various cases of the following AR(1) model:
where $X_{it}$ is the respondent's spouse's income in hundreds of thousands of US dollars, $\epsilon_{it}$ is i.i.d.\ standard logistic, and $\alpha_{i}$ is the latent individual effect that can be arbitrarily dependent with all other random variables except $\epsilon_{it}$. In particular, in models of labor market outcomes it is especially important to distinguish between the true effect of state dependence, measured by $\beta$, and the effects of persistent unobserved heterogeneity, captured by the individual-specific effect $\alpha_{i}$ (see card1987measuring). We consider four specifications, labelled (S1) - (S4), which are based on the general model in (ref):
We drop all observations with missing data either on hours worked or spouse's income over the period we consider, which leaves $5097$ individuals for estimation. Since our procedure requires compactness of the support of the covariates, we winsorize spouse's income $X_{it}$ at one hundred thousand. Since spouses income is in hundreds of thousands, this ensures $X_{it} \in [0,1]$ for $t=1,2,3$. For the instrument functions, we interact indicators $1\{Y_{i0}=0\}$ and $1\{Y_{i0}=1\}$ with indicators of the form $1\{\max\{X_{i1},X_{i2},X_{i3}\} \in D_k\}$ where $D_{k} = \left(\frac{k-1}{k_n}, \frac{k}{k_n}\right]$ for $k=1,\ldots,k_{n} = 1+\lceil n^{1/6} \rceil = 6$. For the piecewise constant approximation to the moment vector, we use a similar partition, but with only $l_n = k_n-1 = 5$ subsets. Furthermore, since it is not known if the time trend model is point- or partially-identified, as per Remark (ref) we treat specifications (S3) and (S4) as if they were partially identified, and take $\tau_n = 0.001n^{-3/10}$. We then compare the results to estimates from functional differencing in models $(S1)$ and $(S2)$ using the procedure described in honore2024moment, which we refer to as “HW” in the results. Since it is not known whether the models $(S3)$ and $(S4)$ are point identified, we did not apply functional differencing. We also compare the results of our method to a model where $\alpha_{i}=\alpha$ for all $i=1,\ldots,n$, which is estimated using maximum likelihood. We refer to this comparison model as “Logit ML” in the results. Finally, we also include results from a model that estimates all the $\alpha_{i}$ as fixed effects using maximum likelihood, which we call “Logit ML FE.” Note that estimates from this model are inconsistent due to the incidental parameters problem (e.g. andersen1973conditional).
The results are displayed in Table (ref), which includes the (point and set) estimates of $\beta$ and $\gamma$, as well as $95\%$ confidence intervals displayed below the estimates. The results obtained using the methods developed in this paper are displayed under the heading “DGKR.” Across all specifications, we find that the effect of a lagged outcome is positive and significant at the $5\%$ level, indicating a strong and positive effect of the previous period's employment on future employment. We find the effect of the time trend to be negative and insignificant. Our estimates in models $(S1)$ and $(S2)$ are also similar to those obtained using the method of honore2024moment. Interestingly, the qualitative conclusions from our approach agree with the conclusions of the benchmark “Logit ML” that constrains $\alpha_{i}=\alpha$ for $i=1,\ldots,n$. However, without properly accounting for the effects of individual-specific permanent unobserved heterogeneity, the results of this model suggest a state-dependence effect that is approximately twice as large. Consistent with our results, the Logit ML model suggests the time effect is small in magnitude and insignificant. Finally, the table also displays the “Logit FE ML” estimates which come from estimating all fixed effects using maximum likelihood. Due to the incidental parameters problem, all estimates in this model are inconsistent. Unlike the previous models, this model delivers estimates of state dependence of employment that are negative and significant, contrary to intuition. Furthermore, unlike the previous methods, this method produces estimates of the effect of the time trend that is negative and significant. These unintuitive but highly significant results serve as a warning against this model, and motivation for using estimation methods that are consistent in the presence of latent individual effects like the one developed in this paper.
This paper presents a new characterization of the identified set for structural parameters and functionals of the latent variables in a large class of dynamic panel logit models. We do so by relating the problem of identification in these models to the truncated moment problem from the mathematics literature, which asks when a sequence of numbers can be rationalized as the moments of a Radon measure. In the case of structural parameters, we use this connection to show that the identified set can be characterized by a collection of conditional moment equalities subject to a certain set of shape restrictions on the model parameters. In addition to providing a general approach to identification, our procedure delivers the sharp identified set even in cases where previous methods fail. Building on the results of chernozhukov2023constrained, we present estimation and inference procedures that use semidefinite programming methods, are applicable with continuous or discrete covariates, and can be used if the model is point- or partially-identified. We also illustrate the usefulness of our results using a series of examples, and in an application to employment dynamics using data from the National Longitudinal Survey of Youth.
Although we did not pursue it here, our method might also be extended to accommodate environments where the initial outcome is unobserved, as in honore2006bounds. The connection to the truncated moment problem also clearly extends beyond logit models (e.g. heckman1990testing, d2017measuring), and there also exists a class of models with multidimensional fixed effects which we believe can also be connected to the truncated moment problem. These include multinomial panel logit models, and bivariate models involving choices made by multiple interacting individuals (e.g. honore2019panel, andaureo2021identification, and AGM2021). We therefore believe these tools will be useful in studying identification in a variety of other models.