Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
81,526 characters · 10 sections · 48 citation commands
Sufficient Statistics for Markovian Feedback Process and Unobserved Heterogeneity in Dynamic Panel Logit Models
JEL Codes: C23, C25
Keywords: Fixed Effect Dynamic Panel Logit, Sequential Exogeneity, Markov Process, Feedback, Sufficient Statistics, Conditional Maximum Likelihood Estimation.
Many economic problems are inherently dynamic over time in the sense that current choices depend on past choices. Behavioral persistence is a central feature of such settings, which shows the necessity of dynamic discrete choice models. Past choices may also shape subsequent covariates. Even after controlling for additive individual heterogeneity to eliminate potential spurious dynamics, two sources of dynamics remain. One is state dependence, or inertia, whereby individuals are likely to stick to past choices they have made. The other arises when covariates depend on past outcomes, which is referred to as a feedback process. Aside from the state dependence, past outcomes may also affect current outcomes via the feedback process of the covariates.
Despite its wide range of empirical applications, identifying the effects of state dependence and a feedback process in dynamic discrete choice models is difficult. This paper investigates identification and the minimum time periods required for identification in dynamic panel logit models with state dependence, first-order Markov feedback process, and individual unobserved heterogeneity by introducing sufficient statistics for the process and unobserved heterogeneity.
A main challenge in a binary panel logit model arises when there exists individual-specific heterogeneity unobserved from the data. Due to the unobserved heterogeneity, two types of dynamics arise: true dynamics from state dependence and spurious dynamics from unobserved heterogeneity. Distinguishing between these two sources of persistence is nontrivial. If heterogeneity is not factored into the model, the resulting estimates may reflect spurious dynamics driven by the correlation between unobserved heterogeneity and observed covariates, rather than genuine behavioral persistence heckman_1978, Heckman1981heterogeneity. However, the presence of the unobserved heterogeneity in nonlinear settings implies significant identification challenges, unlike in linear panel models, as unobserved heterogeneity is hard to be differenced out or eliminated due to the nonlinearity of the model. Treating individual heterogeneity as parameters to be estimated leads to the well-known incidental parameter problems where the number of incidental parameters to be estimated increases with the number of individual in the panel leading to inconsistent estimators neyman_1948.
Another important issue concerns whether identification is feasible with short panels and the minimum time periods required for identification. If the panel is long enough, the parameters and unobserved heterogeneity can be estimated consistently. For example, Bonhomme_discretizing_2022 propose a method to identify them by discretizing unobserved heterogeneity by grouping individuals using K-means clustering. However, with a short panel, identification becomes difficult due to the limited time periods. For example, in the case of a short panel with two periods, a necessary condition for point identification to be feasible under bounded and strictly exogenous covariates is that the error term follows a logistic distribution chamberlain2010binary.
Several methods have been introduced for identification in nonlinear panel models with strictly exogenous covariates, lagged dependent variables, and unobserved heterogeneity with a fixed time period. chamberlain1980 propose a conditional maximum likelihood estimation (CMLE) to address identification in a panel model under a static logit model; the corresponding estimator is consistent and asymptotically normal Anderson_1970. By conditioning on sufficient statistics for the unobserved heterogeneity, the conditional likelihood does not depend on the unobserved heterogeneity. Chamberlain_1985,magnac_2000 provide sufficient statistics under AR(p)-type dynamic logit model. Honore_sufficient_2000 further develop the method to accommodate the strictly exogenous variables which are continuous. Al-Sadoon21102017 suggest an exponential class specification and related estimation method, CMLE and generalized method of moment, to address identification with lagged variables. AGUIRREGABIRIA2021280 extend the approach to structural dynamic logit models with unobserved heterogeneity in the continuation value function accommodating various settings including lagged variables and forward-looking decision-making behavior. dano2025binarychoicelogitmodels study sufficient statistics-based identification strategies for general fixed effects under the binary panel logit model.
Instead of relying on sufficient statistics for the unobserved heterogeneity, which may not exist for some parametric specifications, a generalized method of moments based approach is also introduced. bonhomme_functional_2012 proposes a systematic approach to difference out the heterogeneity using the orthogonal projection of the mapping from heterogeneity to the dependent variables. KITAZAWA2022350 suggests moment conditions by transforming the dynamic binary panel logit model into linear form. dano2023transitionprobabilitiesmomentrestrictions presents a method to derive moment restrictions in AR(p)-type logit model with unobserved heterogeneity. Honore_Moment_functional_2024 extend the approach from bonhomme_functional_2012 to a more general setting with lagged variables, providing a numerical method to find explicit analytic expressions for the moment conditions under broader model specifications. Taking advantage of the polynomial structure of the logit-type model with respect to the individual unobserved heterogeneity, dobronyi2024identificationdynamicpanellogit establish the sharp identified set of the parameter, the heterogeneity, and the average treatment effect based on the truncated moment in the dynamic binary panel logit model.
Identified set for the parameter is often derived under weak assumptions. honore2006bounds provide the estimation methods to construct consistent estimates of the identified set. By exploiting the intersection of the inequalities satisfied by the parameters, ARISTODEMOU2021253,Khan2023 derive informative bounds and the sharp identified set in the presence of lagged variables and strictly exogenous covariates under minimal assumptions.
Aside from the state dependence, the explanatory variables may not be strictly exogenous, a common assumption in nonlinear panel models, even after controlling for the unobserved heterogeneity as pointed out Arellano_Honore_2001. As a remedy, several papers address issues of identification and non-identification while relaxing strict exogeneity. For example, gao2024identificationnonlineardynamicpanels provide a general identified set under partial stationarity where there are AR(p)-type of lagged dependent variables and contemporaneously endogenous variables.
Some papers address issues of non-identification and identification results with predetermined covariates. The covariates may depend on past dependent variables; this dependence is referred to as a feedback process. Contrary to the identification result of chamberlain2010binary, it is not possible to identify the parameters in a short panel with two periods if there exists a lagged dependent variable as a covariate even under logistic error Chamberlain_feedback_logit_nonidentification_2023. bonhomme_identification_2023 provide the conditions for identification under sequential exogeneity and show that even in the simple setup where all covariates are binary, point-identification often fails regardless of the time period. bonhomme2025momentrestrictionsnonlinearpanel characterize feedback and heterogeneity robust moment conditions under sequential exogeneity.
Identification of the parameters under sequential exogeneity is often possible under the additional assumptions. Under the assumption that there exists an explanatory variable independent of the unobserved heterogeneity and the errors conditional on the other covariates, Honore_Lewbel_2002 establish identification of the other predetermined covariates. ARELLANO2003125 present the identification of the parameters through the generalized method of moments when predetermined covariates depend on the conditional expectation of the unobserved heterogeneity. PIGINI202283 investigate the identification under the logit model with the assumption that the feedback process belongs to the exponential family via CMLE and pseudo CMLE.
This paper contributes to the literature by exploring identification and non-identification of the parameters in AR(1)-type dynamic logit models with a sequentially exogenous covariate. We show that, even in the presence of a first-order Markovian feedback process of the covariate and individual unobserved heterogeneity, it is possible to construct a conditional likelihood that is free of these components by treating them as nuisance parameters to be eliminated. Under the logit specification, which admits sufficient statistics for the individual heterogeneity, we derive sufficient statistics for the individual unobserved heterogeneity and the feedback process of discrete covariates. By conditioning on the sufficient statistics for both nuisance parameters, the conditional likelihood is free of any individual-specific unobserved parameters.
Analogous to the identification condition in the standard linear model requiring sufficient variation in the covariates, identification via conditional likelihood often fails because, once conditioning on the sufficient statistics, the conditional likelihood becomes a constant and contains no information about the parameter. This paper relies on this condition to investigate the identification of the parameter via conditional likelihood.
Even beyond identification via conditional likelihood, the identification can fail. This paper establishes a non-identification result under a first order Markovian feedback process by examining the non-existence of the moment function that is a non-constant function of the parameter bonhomme_identification_2023,bonhomme2025momentrestrictionsnonlinearpanel. Failure of point-identification implies that additional restrictions are required to identify $\beta$. This paper suggests the restrictions which are respectively imposed on the feedback process and on the initial condition to identify the parameter.
The rest of the paper is organized as follows. Section (ref) presents the AR(1)-type dynamic binary panel logit model, its assumptions, and establishes non-identification results in Subsections (ref) and (ref), which cover non-identification via conditional likelihood and non-identification beyond conditional likelihood. Section (ref) provides two assumptions to identify the parameters via conditional likelihood in Subsection (ref) and (ref). We conclude in Section (ref).
A sequence of the variable $X$ of individual $i$ from $s$ to $t$ period is denoted by $X_i^{s:t}=\left(X_{is},\hdots,X_{it}\right)$. $\left(G(x)\right)_{x\in\mathcal{X}}$ is the collection indexed by the support $\mathcal{X}$ with component $G(x)$ at each $x\in\mathcal{X}$. Cardinality of the set $\mathcal{X}$ is denoted by $|\mathcal{X}|$. The researcher observes panel data $\left(X_{i}^{1:T},Y_{i}^{0:T}\right)$ over $i=1,\hdots,N$ individuals where $Y_{it}\in\{0,1\}$ is a random binary dependent variable and $X_{it}\in\mathcal{X}$ is a random discrete covariate defined on the finite support $\mathcal{X}$. It is assumed that samples are randomly drawn.
Consider the following dynamic panel binary choice model with state dependence of $Y_{it}$ on $Y_{it-1}$, discrete sequentially exogenous $X_{it}$, individual unobserved heterogeneity $\alpha_i$, and individual time-varying unobserved error term $u_{it}$:
where $\mathbbm{1}\{\cdot\}$ is the indicator function. Let the parameters of interest be $\theta=\left(\rho,\beta\right)^\prime\in\Theta$ where $\Theta$ denotes the parameter space which is assumed to be a compact subset of $K$ dimensional Eucliden space, $\mathbbm{R}^K$. In the specification (ref), $K=2$.
In the model specification (ref), the error $u_{it}$ is assumed to follow a logistic distribution given $\left(X_{i}^{1:t},Y_{i}^{0:t-1},\alpha_i\right)$.
Individual unobserved heterogeneity $\alpha_i$ is independent and identically distributed given initial condition $I_i=\iota$. $\alpha_i$ is allowed to be arbitrarily correlated with $\left(X_{i}^{1:T},Y_{i}^{0:T}\right)$.
Initial condition varies according to the assumptions about $X_{it}$. Under Assumption (ref), initial conditions are $I_{i}=\left(Y_{i0},X_{i1}\right)$ and the support is $\mathcal{I}=\{0,1\}\times\mathcal{X}$. Under Assumption (ref), $I_{i}=Y_{i0}$ and the support is $\mathcal{I}=\{0,1\}$. Under Assumption (ref), $I_{i}=\left(X_{i0},Y_{i0}\right)$ and the support is $\mathcal{I}=\mathcal{X}\times\{0,1\}$.
The objective is to identify $\theta$ and determine the minimum period $T$\footnote{The total number of the time period required for identification is $T+1$ due to the initial condition period; in this paper, we do not count the period $T=0$ for identification.} for identification under Assumptions (ref) - (ref) while relaxing strict exogeneity on $X_{it}$, which is commonly assumed in the dynamic panel logit model. $X_{it}$ is sequentially exogenous, $$ E\left(u_{it}\:|\:X_{i}^{1:t}\right)=0,\quad E\left(u_{it}\:|\:X_{i}^{t+1:T}\right)\neq0, $$ because of the dependence of $X_{it}$ on the dependent variable in the past. Let the probability of $X_{it}$ at $x$ given $\left(X_{i}^{t-1},Y_{i}^{t-1}\right)$ be $$ P\left(X_{it}=x\:|\:X_{i}^{1:t-1},Y_{i}^{0:t-1}\right)=G_i\left(x\:|\:X_{i}^{1:t-1},Y_{i}^{0:t-1}\right). $$ Note that $X_{it}$ is a discrete random variable. Let this probability of $X_{it}$ be the feedback process of $X_{it}$, which is denoted by $G_i\left(x\:|\:X_{i}^{1:t-1},Y_{i}^{0:t-1}\right)$. The feedback process is assumed to follow the first-order Markovian feedback process in the sense that the process depends only on $\left(X_{it-1},Y_{it-1}\right)$, $$ G_i\left(x\:|\:X_{i}^{1:t-1},Y_{i}^{0:t-1}\right)=G_i\left(x\:|\:X_{it-1},Y_{it-1}\right),\quad \forall x\in\mathcal{X},\: \forall t\geq2. $$ This assumption does not require a homogeneous feedback process across individuals. Each individual may have a different feedback process.
Let $\mathbf{G}_i$ denote a Markov kernel with entries $G_i\left(x_2\:|\:x_1,y\right)$ for all $x_1,x_2\in\mathcal{X}$ and $y\in\{0,1\}$.
By Assumption (ref), all components of $\mathbf{G}_i$ are strictly positive. Assumptions (ref) - (ref) together with Assumptions about the feedback process of $X_{it}$, (ref) imply that the probability at any sample path $\left(x^{2:T},y^{1:T}\right)$ is strictly positive, $$ P\left(X_i^{2:T}=x^{2:T},Y_{i}^{1:T}=y^{1:T}\:|\:Y_{i0},X_{i1},\alpha_i,\mathbf{G}_i\:;\:\theta\right)\in\left(0,1\right),\quad \forall \left(x^{2:T},y^{1:T}\right)\in\mathcal{X}^{T-1}\times \{0,1\}^{T}. $$ As stated in bonhomme_identification_2023, these assumptions rule out cases of stayers in which $\alpha_i$ and the conditional distribution of $\left(X_{i}^{2:T},Y_{i}^{1:T}\right)$ induce identical values of $X_{it}$ and $Y_{it}$ over time.
Note that $\mathbf{G}_i$ is an individual-specific feedback process unobserved to the researchers and can be arbitrarily correlated with $\alpha_i$. Therefore, there are two sources of individual heterogeneity, $\alpha_i$ and $\mathbf{G}_i$, which are not observed in the data.
A key identification strategy is to exploit sufficient statistics, treating $\alpha_i$ and $\mathbf{G}_i$ as nuisance parameters to be eliminated. Sufficient statistics are derived using the factorization theorem. The logistic distribution is an example that admits sufficient statistics for $\alpha_i$ and $\mathbf{G}_i$. Under Assumption (ref) - (ref), the joint probability of $\left(X_i^{2:T},Y_{i}^{1:T}\right)$ at $\left(x^{2:T},y^{1:T}\right)$ given the initial condition $\left(Y_{i0},X_{i1}\right)=\left(y_0,x_1\right)$ in the model specification (ref) is
where the first equality holds due to the model specification from equation ((ref)) and Assumption (ref). The second equality follows from the fact that $Y_{it}$ is binary, and the third equality holds because the conditional distribution $F$ is the logistic distribution function.
The joint probability at $\left(x^{1:T},y^{0:T}\right)$ can be factorized into functions $f_{\theta}$, $f_{\alpha_i}$, and $f_{\mathbf{G}_i}$. $f_{\alpha_i}$ and $f_{\mathbf{G}_i}$ are functions of nuisance parameters $\alpha_i$ and $\mathbf{G}_i$, respectively, while $f_{\theta}$ is free of nuisance parameters. By the discrete support of $X_{it}$ and $Y_{it-1}$, the denominator in $f_{\alpha_i}$, $\prod_{t=1}^T \Big\{1+\exp{\left(\rho y_{t-1}+\beta x_{t}+\alpha_i\right)}\Big\}$ can be expressed in terms of the counts of possible values taken by $\Big\{1+\exp{\left(\rho y_{t-1}+\beta x_{t}+\alpha_i\right)}\Big\}$:
$f_{\alpha_i}$ can be represented as a function of the statistics related to $\alpha_i$. Let the vector of the statistics which indicates the number of transitions of $X_{it}$ given $Y_{it-1}$, $\sum_{t=1}^{T} \mathbbm{1}\{Y_{it-1}=y,X_{it}=x\}$, for all $x\in\mathcal{X},y\in\{0,1\}$ be $\mathbf{S}_1^{\alpha}$, $$ \mathbf{S}_1^{\alpha}\left(X_i^{1:T},Y_{i}^{0:T-1}\right)=\left(\sum_{t=1}^{T} \mathbbm{1}\{Y_{it-1}=y,X_{it}=x\}\right)_{x\in\mathcal{X},y\in\{0,1\}}. $$ The denominator of $f_{\alpha_i}$ is a function of $\left(X_{i}^{1:T},Y_{i}^{0:T-1}\right)$ via $\mathbf{S}_1^{\alpha}$. Based on the binary $Y_{it}$ and discrete $X_{it}$, if the number of all possible pairs $\left(Y_{it-1},X_{it}\right)$ in $(X_{i}^{1:T},Y_{i}^{0:T-1})$ is known given initial condition $\left(Y_{i0},X_{i1}\right)=\left(y_0,x_1\right)$, the denominator of $f_{\alpha_i}$ is known. Similarly, if $\sum_{t=1}^T Y_{it}$ is known, the numerator in $f_{\alpha_i}$ is known, which shows that the numerator of $f_{\alpha_i}$ depends on $\alpha_i$ through statistics $\sum_{t=1}^T Y_{it}$. Therefore, $f_{\alpha_i}$ depends on $\alpha_i$ through the statistics $\mathbf{S}^\alpha$, $$ \mathbf{S}^{\alpha}\left(X_i^{1:T},Y_{i}^{0:T}\right)=\left(\sum_{t=1}^T Y_{it},\mathbf{S}_1^{\alpha}\right). $$ $f_{\alpha_i}\left(x^{1:T},y^{0:T}\right)$ can be expressed as $$ f_{\alpha_i}\left(x^{1:T},y^{0:T}\right)=f_{\alpha_i}\left(s^\alpha\right) $$ where $s^\alpha$ denotes the value of $\mathbf{S}^\alpha$ at $\left(x^{1:T},y^{0:T}\right)$.
As an analogy of the representation of $f_{\alpha_i}$ using $\mathbf{S}^{\alpha}$, $f_{\mathbf{G}_i}$ can be represented in terms of the counts of possible cases in $G_i\left(x_t\:|\:x_{t-1},y_{t-1}\right)$:
Note that this representation holds regardless of the conditional distribution of $u_{it}$. $f_{\mathbf{G}_i}$ depends on $\mathbf{G}_i$ through a vector of statistics $\mathbf{S}^{\mathbf{G}}$ whose components are the number of all possible triples $\left(X_{it},Y_{it},X_{it+1}\right)$, $$ \mathbf{S}^{\mathbf{G}}\left(X_i^{1:T},Y_{i}^{1:T-1}\right)=\left(\sum_{t=1}^{T-1} \mathbbm{1}\{X_{it}=x_1,Y_{it}=y,X_{it+1}=x_2\}\right)_{x_1,x_2\in\mathcal{X},y\in\{0,1\}}. $$ Consequently, $$ f_{\mathbf{G}_i}\left(x^{1:T},y^{1:T-1}\right)=f_{\mathbf{G}_i}\left(s^\mathbf{G}\right). $$ where $s^\mathbf{G}$ denotes the value of $\mathbf{S}^\mathbf{G}$ at $\left(x^{1:T},y^{1:T-1}\right)$. First-order Markovian feedback Assumption (ref) allows the proposed approaches to be applied.
Therefore, the joint probability in ((ref)) can be factorized into functions of the statistics $\mathbf{S}^\alpha$ and $\mathbf{S}^{\mathbf{G}}$, and a function that is free of $\alpha_i$ and $\mathbf{G}_i$,
$f_\theta$ and $f_{\alpha_i}$ are strictly positive functions, and $f_{\mathbf{G}_i}$ is strictly positive according to Assumption (ref). Thus, by the factorization theorem,
are sufficient statistics for the nuisance parameters $\alpha_i$ and $\mathbf{G}_i$. Observe that the statistic $\mathbf{S}^\mathbf{G}$ imposes more constraints in $\left(X_{i}^{1:T},Y_{i}^{1:T-1}\right)$ than $\mathbf{S}^{\alpha}$ since $\mathbf{S}^{\mathbf{G}}$ counts the possible triples $\left(X_{it},Y_{it},X_{it+1}\right)$, while $\mathbf{S}^{\alpha}$ counts the possible pairs $\left(Y_{it-1},X_{it}\right)$. If $\mathbf{S}^\mathbf{G}$ is known, $\mathbf{S}^{\alpha}$ is determined. This restriction on possible $\left(X_i^{2:T},Y_{i}^{1:T}\right)$ is the source of non-identification of $\beta$ through the conditional likelihood under Assumption (ref).
Let $\mathcal{S}$ denote the support of $\mathbf{S}$. For each $s\in\mathcal{S}$, let $\mathcal{D}_s$ denote the set of possible values of $(x^{2:T},y^{1:T})$ such that $\mathbf{S}\left(x^{1:T},y^{0:T}\right)=s$:
From equation ((ref)) and ((ref)), the probability of $\left(X_i^{2:T},Y_{i}^{1:T}\right)$ at $\left(x^{2:T},y^{1:T}\right)$ conditional on $\mathbf{S}=s$ and the initial condition $\left(Y_{i0},X_{i1}\right)=\left(y_0,x_1\right)$ is
where $\tilde{x}_{t},\tilde{y}_{t}$ denotes $x_{t},y_{t}$ in $(\tilde{x}^{2:T},\tilde{y}^{1:T})$ and $\left(\tilde{y}_{0},\tilde{x}_{1}\right)=\left(y_0,x_1\right)$. Note that the probability ((ref)) does not depend on $\alpha_i$ and $\mathbf{G}_i$.
Identification of parameters of interest via conditional likelihood relies on the conditional probability ((ref)). Let $\theta_0$ denote the true value of the parameters. The identification is equivalent to verifying whether there exists $\mathbf{S}=s$ such that
is equivalent to $\theta_0=\theta_1$. The corresponding identification condition from equation ((ref)) is
To hold equality, for all possible $(\tilde{x}^{1:T},\tilde{y}^{0:T})$ conditional on $\mathbf{S}=s$,
holds. For all $s$ in $\mathcal{S}$, if there does not exist $(\tilde{x}^{1:T},\tilde{y}^{0:T})$ conditional on $\mathbf{S}=s$ such that
with positive probability, identification of the corresponding parameter via conditional likelihood fails because even if $\rho_0\neq\rho_1$ or $\beta_0\neq\beta_1$, equation ((ref)) can still hold. If every possible $(\tilde{x}^{2:T},\tilde{y}^{1:T})$ conditioning on $\mathbf{S}=s$ has identical value of both $\sum_{t=1}^T \tilde{y}_{t-1} \tilde{y}_{t}$ and $\sum_{t=1}^T \tilde{x}_{t} \tilde{y}_{t}$ or if $\mathcal{D}_s$ is a singleton, equation ((ref)) holds. In this case, the distributions of $\sum_{t=1}^T Y_{it-1} Y_{it}$ and $\sum_{t=1}^T X_{it} Y_{it}$ conditional on $\mathbf{S}=s$ are degenerate. The probability ((ref)) is a constant and holds for any parameter value. Therefore, identification of the parameters fails. Consequently, a key to the identification of parameters through conditional likelihood is to verify whether there exists $s$ such that condition ((ref)) holds once conditional on $\mathbf{S}=s$.
Proposition (ref) establishes the identification result via conditional likelihood in the model exploiting the condition ((ref)).
For the proof of Proposition (ref), see Appendix (ref). As an example, consider following two sample paths of $\left(X_i^{2:T},Y_{i}^{1:T}\right)$ given $\left(Y_{i0},X_{i1}\right)=\left(0,1\right)$ for the identification of $\rho$ under $T=3$: $$
$$ where $A$ and $B$ have identical values of the sufficient statistics $\mathbf{S}$. They have an identical number of occurrences of the feedback process $(X_{it},Y_{it},X_{it+1})$ over periods: $(1,0,1)$, $(1,1,1)$. In addition, $\sum_{t=1}^T Y^A_{it}=\sum_{t=1}^T Y^B_{it}=2.$ In terms of $\sum Y_{it-1} Y_{it}$, $A$ and $B$ have different values, $\sum_{t=1}^T Y_{it-1}^A Y_{it}^A=1$, $\sum_{t=1}^T Y_{it-1}^B Y_{it}^B=0$. Since there exists $\mathbf{S}=s$ and identification condition (\ref{eq:identifying_stat}) is satisfied, $\rho$ is identified.
Once conditional on the sufficient statistics (ref), there do not exist sufficient variations in the value of $\sum_{t=1}^TX_{it}Y_{it}$. There might exist multiple paths conditional on $\mathbf{S}=s$. However, all paths have an identical value of $\sum_{t=1}^TX_{it}Y_{it}$, violating the identification condition ((ref)). Conditioning on the sufficient statistics eliminates all variation in $\sum_{t=1}^TX_{it}Y_{it}$.
Non-identification result of $\beta$ through conditional likelihood does not imply that $\beta$ is not identified in general. CMLE only exploits the information about the parameters from the conditional likelihood ((ref)), not from the joint probability ((ref)). Moreover, there exist distributions $F$ that do not admit sufficient statistics for $\alpha_i$; hence, identification of the parameter conditioning on the sufficient statistics for $\alpha_i$ and $\mathbf{G}_i$ is not always plausible method for identification. Therefore, identification via conditional likelihood does not provide a generalized condition for identification. Section (ref) provides failure of point-identification under the Markovian process in panel logit models by examining non-existence of feedback and heterogeneity robust moment conditions described in bonhomme_identification_2023,bonhomme2025momentrestrictionsnonlinearpanel.
Analogous to the model specification in bonhomme_identification_2023, to illustrate the non-identification of $\beta$ via observed probabilities, consider the following simplified model with a binary sequentially exogenous $X_{it}$, individual unobserved heterogeneity $\alpha_i$, and individual time-varying unobserved error term $u_{it}$:
For simplicity, suppose the support of $\alpha_i$ is finite and contains more than $2^{2T-1}-1$ points\footnote{If the number of points in the support is less than $2^{2T-1}-1$, each point in the support would be pinned down as a solution to a nonlinear system of equations. In this case, the rank of the matrix $\mathbf{P}$ in the equation ((ref)) is $$ \operatorname*{rank}\left(\mathbf{P}\right)\leq \min\left(|\mathcal{A}|,2^{2T-1}-1\right)=|\mathcal{A}|. $$ To avoid this, it is assumed that the support $\mathcal{A}$ contains more than $2^{2T-1}-1$.}, $$ \mathcal{A}=\{\alpha_1,\hdots,\alpha_K\},\quad|\mathcal{A}|>2^{2T-1}-1. $$
Identical to the feedback process specification in bonhomme_identification_2023, for simplicity, assume that heterogeneity in $\mathbf{G}_i$ enters only through $\alpha_i$: $$ G_{i}\left(1\:|\:x,y\right)=G\left(1\:|\:x,y\:;\:\alpha_i\right)=G_{xy}\left(\alpha_i\right), \quad \forall x,y\in\{0,1\} $$ where $G_{xy}\left(\alpha_i\right)$ is shorthand for the feedback process evaluated at $\alpha_i$. Denote the value of the sufficient statistics $\mathbf{S}$ from ((ref)) of the probability of the path $j$ by $$ \left(n_y^j,\left(n_{x_1,x_2,y}^j\right)_{x_1,x_2,y\in\{0,1\}}\right). $$ and the exponential of $\alpha_i$ by $A_i$, $$ A_i=\exp\left(\alpha_i\right). $$ Using the notations, the integrated likelihood of $\left(X_i^{2:T},Y_{i}^{1:T}\right)$ at $j$ sample path $\left(x^{j,2:T},y^{j,1:T}\right)$ over $\alpha_i$ is
where $p_j\left(\alpha_k\right)$ is the shorthand notation of the conditional probability of $j$ sample paths given $\alpha_k$ and $\mathbf{G}=\left(G_{xy}\left(\alpha_k\right)\right)_{x,y\in\{0,1\},\alpha_k\in\mathcal{A}}$. There are $2^{2T-1}$ sample paths. Let $J+1=2^{2T-1}$ and $J$ paths need to be analyzed due to the fact that the sum of the probabilities of all sample paths is one.
Parameters in the model are structural parameter $\beta$, probability mass function of $\alpha_i$ at each point in the support $\mathcal{A}$, and feedback process of $X_{it}$, $\left(\beta,\mathbf{H}_{x_1},\mathbf{G}\right)$ where $\mathbf{H}_{x_1}=\left(H_{x_1}\left(\alpha_k\right)\right)_{\alpha_k\in\mathcal{A}}$. Let the vector of the integrated probabilities of all sample paths given $X_{i1}=x_1$ be $\psi_{x_1}\left(\beta,\mathbf{H}_{x_1},\mathbf{G}\right)$, $$ \psi_{x_1}\left(\beta,\mathbf{H}_{x_1},\mathbf{G}\right) =
$$ and let the corresponding Jacobian matrix be $$ \nabla \psi_{x_1}\left(\beta,\mathbf{H}_{x_1},\mathbf{G}\right)=
. $$
Identification requires the rank of the Jacobian matrix $\nabla \psi_{x_1}\left(\beta,\mathbf{H}_{x_1},\mathbf{G}\right)$ evaluated at the true parameter value, which is not attainable because the true value of the parameter is unknown. Regularity at the true parameter values ensures that the rank is locally constant in a neighborhood of the truth, so it is possible to compute the rank near the truth without knowing the exact true parameter.
As pointed out in bonhomme_identification_2023,bekker2001identification, Assumption (ref) is satisfied almost everywhere in the logit specification since the cumulative distribution of the logistic $F$ is an analytic function.
The parameter $\beta$ under the parametric models is identified if $$ \operatorname*{rank}\left(\nabla_{-\beta} \psi_{x_1}\left(\beta,\mathbf{H}_{x_1},\mathbf{G}\right)\right)<\operatorname*{rank}\left(\nabla \psi_{x_1}\left(\beta,\mathbf{H}_{x_1},\mathbf{G}\right)\right) $$ where $\nabla_{-\beta} \psi_{x_1}\left(\beta,\mathbf{H}_{x_1},\mathbf{G}\right)$ is the Jacobian matrix except $\beta$ column. Equivalently, if one can show $$ \operatorname*{rank}\left(\nabla_{-\beta} \psi_{x_1}\left(\beta,\mathbf{H}_{x_1},\mathbf{G}\right)\right)=\operatorname*{rank}\left(\nabla \psi_{x_1}\left(\beta,\mathbf{H}_{x_1},\mathbf{G}\right)\right), $$ the parameter $\beta$ is not identified rothenberg1971identification,bekker2001identification. Exploiting the conditions, bonhomme_identification_2023 investigates the necessary condition for identification of $\beta$ in dynamic panel logit models that if $\beta$ is identified, there exists a nonzero projection $m^{proj}$ of $\nabla_{\beta} \psi_{x_1}$ onto the orthogonal complement of the space spanned by $
$ such that $$ m^{proj}= \underbrace{\left(I-
^\dagger\right)}_{\mathbf{M}}\nabla_{\beta}\psi_{x_1} $$ where $\dagger$ denotes Moore-Penrose inverse matrix. By the properties of the projection which is nonzero, it follows that
Notice that since $m^{proj}$ is nonzero, $m^{proj}$ satisfies $\nabla_\beta \psi_{x_1}^\prime m^{proj}\neq0$.\footnote{Suppose $\nabla_\beta \psi_{x_1}^\prime m^{proj}=0$ holds. Since $m^{proj}=\mathbf{M}\nabla_{\beta}\psi_{x_1}$, $\nabla_{\beta} \psi_{x_1}^\prime m^{proj}=\nabla_{\beta}\psi_{x_1}^\prime \mathbf{M}\nabla_{\beta}\psi_{x_1}=\nabla_{\beta}\psi_{x_1}^\prime \mathbf{M}^2\nabla_{\beta}\psi_{x_1}= \left(m^{proj}\right)^\prime m^{proj}=0$. Therefore, $m^{proj}=0$.} $\nabla_{\mathbf{H}_{x_1}} \psi_{x_1}^\prime m^{proj}=0$ corresponds to the moment condition that
where $m^{proj}_j$ denotes the value of the moment function at the $j$ sample paths. Thus, identification of $\beta$ requires the existence of nonzero moment function $\phi$ satisfying the condition ((ref)) and ((ref)).
Consider an arbitrary moment function $m$. Under the model specification (ref) and the finite support of $\alpha_i$, the equation (ref) can be represented with matrix form,
If the rank of the matrix $\mathbf{P}$ satisfies $$ \operatorname*{rank}\left(\mathbf{P}\right)<\min\left(|\mathcal{A}|,J\right)=J, $$ there exists a nonzero moment function $m$. The existence of a moment function is equivalent to identifying the linear independence of the path probabilities among the columns in $\mathbf{P}$: $$ \sum_{j=1}^J m_j p_j\left(\alpha_k\right)=0,\quad\forall\alpha_k\in\mathcal{A}\quad \Longrightarrow\quad m_j =0,\quad\forall j=1,\hdots,J. $$ The dependence across the sample paths is the value of the moment function at $j$ sample paths, $m_j$. Consequently, if all path probabilities are linearly independent, there does not exist a nonzero moment function.
When $T\geq3$, under Markovian feedback process, there always exist linearly dependent probabilities of the sample path $j,w$ such that $$ p_j\left(\alpha_k\right)=p_w\left(\alpha_k\right), \quad\forall \alpha_k\in\mathcal{A}. $$ The following two sample paths are examples of the identical probabilities when $T=3$: $$
. $$ Probabilities of the path $A$ and $B$ are
Therefore, there exists a moment function such that for some $j$ and $w$, $m_j+m_w=0$. However, it does not mean that the function $m$ is a projection function satisfying the condition ((ref)). The moment function $m$ does not contain any information about $\beta$.
To distinguish the dependence which is a non-constant function of $\beta$ from the dependence due to the identical probabilities, which are not informative about the parameter, define the former as non-trivial dependence and the latter as trivial dependence.
In addition to the existence of the moment function satisfying the condition ((ref)), to identify $\beta$ under the Markovian feedback process, there should be a non-trivial moment function $m$ in the sense that the function $m$ is a non-constant function of $\beta$.
From the proposition (ref), if the sufficient statistics $\mathbf{S}$, $$ \mathbf{S}\left(X_i^{1:T},Y_{i}^{1:T}\right)=\left(\sum_{t=1}^T Y_{it},\left(\sum_{t=2}^T\mathbbm{1}\{X_{it-1}=x_1,Y_{it-1}=y,X_{it}=x_2\}\right)_{x_1,x_2,y\in\{0,1\}}\right) $$ are identical among the paths given the initial condition, the paths have an identical value of $\sum_{t=1}^T X_{it}Y_{it}$. Under the simple specification (ref), the paths have identical probability. Therefore, if there exist paths such that the probabilities of the paths are non-trivially linearly dependent, at least one of the statistics should be different.
Suppose there are $w_s$ sample paths in $\mathcal{D}_s$. By the proposition (ref), $p_w=p_{s}$ holds for the path $w\in \mathcal{D}_{s}$. A linear combination of the paths is
Consequently, the trivial dependence across probabilities of the paths in $\mathcal{D}_s$ can be reduced to one representative probability $p_{s}\left(\alpha_k\right)$ with $m_{s}$. In the subsequent lemmas and proposition, the representative probability $p_{s}\left(\alpha_k\right)$ will be used to avoid trivial dependent path probabilities.
The following lemmas restrict the sets of the paths which have non-trivially linearly dependent probabilities. The lemmas state that the paths should have the same polynomial degrees in terms of the nuisance parameters to have the dependence free of the nuisance parameters.
For the proof of Lemma (ref), see Appendix (ref). Lemma (ref) states that if there exists non-trivial dependence free of the nuisance parameters, the polynomial degrees of $A_i$ among the linearly dependent path probabilities should be identical, $n_y^j=n_y$.
Similarly, the linearly dependent probabilities of the sample paths have identical degrees in terms of $G_{xy}\left(\alpha_k\right)$ for all $x,y\in\{0,1\}$.
For the proof of Lemma (ref), see Appendix (ref). Lemma (ref) states that if the path probabilities have nontrivial linear dependence, the paths have identical $n_{x,y}$ such that $$ n_{x,y}=n_{x,y}^j=n_{x,1,y}^{j}+n_{x,0,y}^{j},\quad \forall x,y\in\{0,1\} $$ to match the polynomial degree in $G_{xy}\left(\alpha_k\right)$ across the path probabilities which are non-trivially linearly dependent.
Under Corollary (ref), Lemma (ref), and Lemma (ref), Lemma (ref) establishes the nonexistence of nontrivial linear dependence among path probabilities.
For the proof of Lemma (ref), see Appendix (ref). Consequently, according to Lemma (ref),(ref), and (ref), there is only trivial dependence among the paths and $m$ is the moment function reflecting trivial dependence. Since there is no non-trivial dependence, from the linear combination of the path probabilities (ref), all $m_s$ are zero. $$ \sum_{s\in \mathcal{S}}m _{s}p_{s}\left(\alpha_k\right)=0,\quad\forall \alpha_k\in\mathcal{A}\quad \Longrightarrow \quad m_{s}=\sum_{w=1}^{w_s} m_{w}=0,\quad \forall s\in \mathcal{S} $$ It implies that $\nabla_{\beta} \psi_{x_1}^\prime m=0$ since $$ \nabla_{\beta} \psi_{x_1}^\prime m=\sum_{\alpha_k \in \mathcal{A}}H_{x_1}\left(\alpha_k\right)\sum_{s\in \mathcal{S}}\sum_{w=1}^{w_s} m_{w} \frac{\partial p_{w}\left(\alpha_k\right)}{\partial \beta}=\sum_{\alpha_k \in \mathcal{A}}H_{x_1}\left(\alpha_k\right)\sum_{s\in \mathcal{S}}\underbrace{\left(\sum_{w=1}^{w_s} m_{w}\right)}_{=m_s=0}\frac{\partial p_{s}\left(\alpha_k\right)}{\partial \beta}=0, $$ where second equality holds because $p_w=p_s$ for all paths in $\mathcal{D}_s$ and third equality holds due to $m_s=0$ for all $s\in\mathcal{S}$. $m$ is not a projection function satisfying $\nabla_{\beta} \psi_{x_1}^\prime m\neq0$. Therefore, although there exist moment functions, the functions do not satisfy the conditions for the projection function. Since the projection function does not exist, $\beta$ is not identified regardless of the time period $T$ in the simple model specification (ref).
Proposition (ref) states that a vector of observed probabilities does not provide sufficient information to identify $\beta$. In addition to the observed probabilities, the model may imply additional valid moment conditions, such as those arising from the sequential exogeneity of $X_{it}$, although these are challenging to exploit in the nonlinear structure of the logit specification.
Consistent with the non-identification of $\beta$ under the logistic specification with an unrestricted feedback process in bonhomme_identification_2023, Proposition (ref) shows that $\beta$ is not identified even under a time-homogeneous first-order Markovian feedback process, which is more restrictive. This non-identification result necessitates additional restrictions to identify $\beta$ under the logistic specification, which will be investigated in Section (ref).
Non-identification result in Proposition (ref) implies that additional assumptions are required to identify $\beta$. This section considers two additional assumptions for the identification while holding heterogeneity of the feedback process across individuals. First assumption is imposed on the feedback process of the covariates. If the feedback process depends on $Y_{it-1}$ only, $\beta$ is identified via CMLE when $T\geq2$. Second assumption is imposed on the initial condition. If initial condition and the unobserved heterogeneity are independent, it provides source of the identification of $\beta$ via CMLE when $T\geq3$. This section also briefly investigates analogous identification result with multiple sequentially exogenous variable and discrete strict exogenous variables.
Consider logit model specification (ref) with a different feedback process where $\mathbf{G}_i$ depends only on $Y_{it-1}$.
Under Assumptions (ref) - (ref), and (ref), a joint probability of $\left(X_i^{1:T},Y_i^{1:T}\right)$ at the path $\left(x^{1:T},y^{1:T}\right)$ given the initial condition $Y_{i0}=y_0$ is
Analogous to the sufficient statistics (ref) for the first model specification (ref), associated sufficient statistics for $\alpha_i$ and $\mathbf{G}_i$ are, respectively,
Notice that sufficient statistics for transition of $X$ in $f_{\mathbf{G}_i}$ is different from those in $f_{\mathbf{G}_i}$ from equation ((ref)) since $\mathbf{G}_i$ does not depend on $X_{it-1}$. Therefore, sufficient statistics $\mathbf{S}$ for the nuisance parameters are given below,
The statistics impose weaker constraints on possible paths $(X_{i}^{1:T},Y_{i}^{1:T})$ conditional on $\mathbf{S}$ than the statistics (ref). These weaker restrictions allow for the identification of $\beta$ and $\rho$, which is shown in Proposition (ref).
Similar to the set (ref), let $\mathcal{D}_s$ denote the set of the paths of $(x^{2:T},y^{1:T})$ such that $\mathbf{S}=s$ be
The probability of $\left(X_i^{1:T},Y_{i}^{1:T}\right)$ at $\left(x^{1:T},y^{1:T}\right)$ conditional on $\mathbf{S}=s$ and the initial condition $Y_{i0}=y_0$ is
where $\tilde{x}_{t},\tilde{y}_{t}$ denotes $x_{t},y_{t}$ in $(\tilde{x}^{1:T},\tilde{y}^{1:T})$ and $\tilde{y}_{0}=y_0$.
Under the condition ((ref)), Proposition (ref) establishes identification result via conditional likelihood.
For the proof of Proposition (ref), see Appendix (ref). As an example for the identification of $\beta$ when $T=2$, consider next example: $$
. $$ Path $A$ and $B$ have the same value of the statistics $\mathbf{S}$: There is an identical number of occurrences of feedback $(Y_{it-1},X_{it})$ over periods, $(1,1)$, $(1,0)$, in $A$ and $B$ as well as $\sum_{t=1}^T Y_{it}^A=\sum_{t=1}^T Y_{it}^B=1$. However, $A$ and $B$ have different values in $\sum_{t=1}^T X_{it}Y_{it}$, $$ \sum_{t=1}^T X_{it}^A Y_{it}^A=1, \sum_{t=1}^T X_{it}^B Y_{it}^B=0. $$ Therefore, there exists $\mathbf{S}=s$ such that the distributions of $\sum_{t=1}^T X_{it} Y_{it}$ are not degenerate once conditional on $\mathbf{S}=s$. Condition (\ref{eq:identifying_stat}) holds and $\beta$ is identified when $T=2$.
To illustrate identification of $\rho$ when $T=3$, consider the following three paths $A,B$ and $C$ below $$
$$ where $A,B$ and $C$ have the same values of sufficient statistics: They have the same transitions of $(Y_{it-1},X_{it})$ over time: $(0,1)$, $(0,0)$, $(1,1)$. Also, $\sum_{t=1}^T Y^A_{it}=\sum_{t=1}^T Y^B_{it}=\sum_{t=1}^T Y^C_{it}=2.$ In terms of $\sum_{t=1}^T Y_{it-1} Y_{it}$, $A,B$ and $C$ have different values: $$ \sum_{t=1}^T Y_{it-1}^A Y_{it}^A=1,\sum_{t=1}^T Y_{it-1}^B Y_{it}^B=\sum_{t=1}^T Y_{it-1}^C Y_{it}^C=0. $$ Thus, the condition (\ref{eq:identifying_stat}) is satisfied and $\rho$ is identified when $T=3$.
Limited dependence of the feedback process of $X_{it}$ on $Y_{it-1}$ from the Assumption (ref) enables the identification of $\beta$ since the process in each period is independent of those in the other period. The process (ref) across periods is not independent because $\mathbf{G}_i$ depends on $X_{it-1}$ and $Y_{it-1}$. Therefore, feedback in the previous period directly limits the possible feedback in subsequent periods, which restricts possible paths and the corresponding value of $\sum_{t=1}^T X_{it}Y_{it}$ conditional on the sufficient statistics.
Assumption (ref) restricts the feedback process in that it depends only on $Y_{it-1}$, not $X_{it-1}$. An alternative way to identify $\beta$ while maintaining the first-order Markovian feedback process assumption (ref) is to impose a restriction on the initial condition. Consider panel data $\left(X_{i}^{0:T},Y_{i}^{0:T}\right)$ over $i=1,\hdots,N$ individuals where the initial condition is $I_i=\left(X_{i0},Y_{i0}\right)$ which is observed from the data. Assume that the distribution of the initial condition is independent of the nuisance parameters.
Let $p$ denote the vector of the initial condition probabilities, $p=\left(p_{xy}\right)_{x\in\mathcal{X},y\in\{0,1\}}^\prime$ where $$ p_{xy}=P\left(X_{i0}=x,Y_{i0}=y\right)$$ and let $\mathcal{P}$ be the admissible set of probability vectors. We also assume that the initial condition probabilities are uniformly bounded away from zero.
With Assumption (ref) and (ref), every path of $\left(X_{i}^{0:T},Y_{i}^{0:T}\right)$ is attainable with positive probability. The conditional likelihood of $\left(X_{i}^{0:T},Y_{i}^{0:T}\right)$ at $\left(x^{0:T},y^{0:T}\right)$ is
where last equality holds due to Assumption (ref). Analogous to the factorization (ref), the likelihood is decomposed,
Notice that $X_{i1}$ is not initial condition and has feedback process given $\left(X_{i0},Y_{i0}\right)$. Corresponding sufficient statistics are
The possible sample paths given sufficient statistics are
The probability of $\left(X_i^{0:T},Y_{i}^{0:T}\right)$ at $\left(x^{0:T},y^{0:T}\right)$ conditional on $\mathbf{S}=s$ is
Since the conditional probability does not condition on the initial condition, paths in $\mathcal{D}_s$ may have different initial conditions.
By Assumption (ref), the initial condition probabilities do not depend on $\alpha_i$ and $\mathbf{G}_i$, and they are estimated from the data. Once the initial condition probabilities are consistently estimated from the data, such as a frequency estimator, they can be used for identification of the parameters via pseudo CMLE. If the initial condition probabilities are instead jointly estimated via CMLE with $\beta$, $\beta$ is not identified because the variation in $\sum_{t=1}^T X_{it}Y_{it}$ can be absorbed by the initial condition probabilities. Thus, when identifying $\beta$, the initial condition probabilities should be estimated separately.
For the proof of Proposition (ref), see Appendix (ref). Consider the following two sample paths which have identical value of sufficient statistics when $T=2$: $$
. $$ Path $A$ and $B$ have identical $\sum_{t=1}^2 Y_{it}$ and occurrence of feedback: $\sum_{t=1}^2 Y_{it}^A=\sum_{t=1}^2 Y_{it}^B=1$, $G_i\left(1\:|\:0,1\right)$, and $G_i\left(0\:|\:1,1\right)$. On the other hand, $A$ and $B$ do not have identical values of $\sum_{t=1}^2 X_{it}Y_{it}$, $$ \sum_{t=1}^2 X_{it}^A Y_{it}^A=1, \sum_{t=1}^2 X_{it}^B Y_{it}^B=0. $$ Thus, $\beta$ is identified via conditional likelihood when $T=2$ since condition (\ref{eq:identifying_stat}) holds. Since $p_{01}$ and $p_{11}$ are identified from the data, the same identification condition ((ref)) can be applied.
For the identification of $\rho$ when $T=3$, consider the following two sample paths, $$
. $$ Each path has identical feedback processes, $G_i\left(0\:|\:0,0\right)$, $G_i\left(0\:|\:0,0\right)$, and $G_i\left(0\:|\:1,0\right)$ and identical $\sum_{t=1}^3 Y_{it}^A=\sum_{t=1}^3 Y_{it}^B=2$. $\sum_{t=1}^3 Y_{it-1} Y_{it}$ of path A and B are $$ \sum_{t=1}^3 Y_{it-1}^A Y_{it}^A=1, \sum_{t=1}^3 Y_{it-1}^B Y_{it}^B=0. $$ Consequently, there exist at least two sample paths satisfying the condition (\ref{eq:identifying_stat}) and $p_{00}$ is identified from the data, $\rho$ is identified when $T=3$, which can be generalized to the case with longer time periods.
Proposition (ref) and (ref) establish identification of $\beta$ under Assumptions (ref) and (ref), respectively, while maintaining sequential exogeneity of $X_{it}$ and without imposing any parametric restriction on the feedback process.
Identification in Propositions (ref) and Proposition (ref) extends to model specifications with multiple sequentially exogenous variables and strictly exogenous covariates defined on finite supports. Consider the following model specification:
where $X_{it}$ and $Z_{it}$ are vectors of discrete sequentially exogenous and strictly exogenous variables, respectively. The support of $Z_{it}$, denoted by $\mathcal{Z}$, is finite. We further allow that the law of motion of $X_{it}$ to depend on $Z_{it}$. The rest of the identification argument is unchanged. For example, if the conditional distribution of $X_{it}$ depends on $\left(Y_{it-1},Z_{it}\right)$, which is an analogous extension of Assumption (ref), sufficient statistics for $\alpha_i$ and $\mathbf{G}_i$ are
The resulting conditional probability is
where $\tilde{x}_{t},\tilde{y}_{t}$ denote $x_{t}$ and $y_{t}$ in $(\tilde{x}^{1:T},\tilde{y}^{1:T})$ and $\tilde{y}_{0}=y_0$. Analogously to Proposition (ref), the existence of strictly exogenous $Z_{it}$ and multiple sequentially exogenous covariates does not affect the identification result. Identification via conditional likelihood with Assumption (ref) augmented with the strictly exogenous covariates in Section (ref) is plausible in a similar manner. Aside from the specifications described above, there are many interesting extensions of the identification via conditional likelihood under sequential exogeneity once the covariates are discrete.
This paper investigates identification in dynamic panel logit models with the first-order Markovian feedback process, state dependence, and unobserved heterogeneity. We provide sufficient statistics for unobserved heterogeneity and the feedback process when the process of the covariate depends on the lagged dependent variable and the lagged covariate. These sufficient statistics can be generalized to settings that include discrete, strictly exogenous covariates, which may be useful for future research.
We also analyze the failure of identification, which necessitates additional identifying restrictions. Furthermore, we suggest two constraints for the identification of the sequentially exogenous variable without imposing parametric structure on the process: one is to impose that the covariate depends only on the lagged dependent variable and not on lagged values of the covariate, and the other is to impose a strictly exogenous initial condition, both of which allow the heterogeneous feedback processes.
In this paper, we derive identification by introducing sufficient statistics that rely on the discrete nature of variables. However, conditional maximum likelihood estimation is not constrained to discrete variables. The logic described in the paper may be extended to a continuous covariate. Furthermore, sufficient statistics for the feedback process do not depend on the logit specification, implying that there can be other model specifications which may allow identification under such feedback processes. We leave these for future research.
\printbibliography