Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
106,520 characters · 29 sections · 85 citation commands
Binary choice logit models with general fixed effects for panel and network data
\thispagestyle{empty} \setcounter{page}{0}
Binary outcome models with fixed effects have a long history in economics and related fields. They allow for flexible modeling of individual decisions while accounting for unobserved heterogeneity. These models arise naturally in two types of data structures: traditional panel data, where individuals or firms are observed over time, and network or dyadic data, where observations correspond to interactions between pairs of units. In both cases, researchers often wish to estimate the effect of observed covariates on a binary outcome, while allowing for unit-specific or pair-specific unobserved components. Examples include individual labor market choices, firm entry decisions, and the formation of social or economic networks.
A central challenge in these models is the treatment of fixed effects, which capture unobserved heterogeneity. When treated as parameters to be estimated, fixed effects can lead to the incidental parameter problem, first identified by neyman1948consistent. This issue arises because the number of fixed effects grows with the sample size. If the number of observations per parameter increases slowly, as in certain network models or panel data settings with increasing number of time periods, estimators may be asymptotically biased. In such cases, bias correction methods can be applied; see, for example, HahnNewey2004, and the surveys by ArellanoHahn2007 and FernandezValWeidner2018. On the other hand, when the number of parameters grows proportionally with the number of observations, as in standard panel data models with a fixed number of time periods, estimators are often inconsistent. In these settings, the most effective strategy is to eliminate the fixed effects.\footnote{In an alternative approach, HigginsJochmans2024 show that in parametric settings like the one considered in this paper, it is possible to use the bootstrap to do valid inference on the common parameters using the inconsistent and biased estimator that treats the fixed effects as parameters to be estimated.}
In this paper, we summarize and extend two existing methods for estimating binary choice logit models with fixed effects by eliminating them from the estimation problem. Our focus is on both static and dynamic models, allowing for very general fixed effect structures. We discuss models for standard panel data, as well as dyadic models that arise in the study of networks.
Two main general approaches have been developed to address the incidental parameter problem. The first is based on conditional likelihood and requires a parametric model conditional on the unobservable fixed effects. This method eliminates the fixed effects by conditioning on sufficient statistics. It was first introduced by rasch1960studies in the context of educational testing and later adapted to econometric applications by chamberlain_analysis_1980. In binary logit models, the fixed effects drop out of the likelihood when we condition on certain linear combinations of the outcomes. This approach has been extended to more complex fixed effect structures, including two-way fixed effects in dyadic data charbonneau2017multiple and network formation models graham2017econometric. Conditional likelihood methods are attractive because they provide a way to estimate structural parameters without modeling the distribution of the unobserved effects. However, they are only available in models where sufficient statistics can be found in closed form.
The second approach relies on moment conditions that do not depend on the fixed effects.\footnote{In this paper we focus on moment \textsl{equality} conditions. Moment \textsl{inequalities} form the basis for the semiparametric approach in Manski87, and have also been used to construct and estimate identified sets for the structural parameters. See, for example, PakesPorterShepardCalderWang2022.} This idea is sometimes called functional differencing. It was formalized by bonhomme2012functional and applied to discrete choice models by kitazawa2021transformations and honore2024dynamic. The basic idea is to construct functions of the data whose conditional expectation is zero for all values of the fixed effects. These functions then define valid moment conditions for the structural parameters. This approach has also been used to develop estimators in dynamic logit models with general fixed effects by dano2023transition. Moment-based methods can be used in a wider range of models than conditional likelihood. They are especially useful in dynamic models or in models with more complicated fixed effect structures.
This paper reviews both approaches in the context of binary choice logit models. We present the underlying identification ideas in a unified framework and show how they apply to standard examples. These include static panel models with individual fixed effects, models with heterogeneous time trends, two-way fixed effects, and network formation settings. For dynamic models, we discuss how moment conditions can be constructed under general assumptions, and how these conditions relate to the number of fixed effects in the model. We also give examples from dynamic panel and dynamic network settings. We focus on the logit model, rather than, say, the probit model or a semiparametric version, because it is known from chamberlain2010binary that in the simplest case of a static binary response panel data model with two time periods, regular root-$n$ consistent estimation of the common parameters is only possible in the logit model.\footnote{For a discussion of what can be learned about the common parameters in dynamic binary response models with panel data in a more semiparametric setting, see for example KHAN2023105515.}
The literature has long employed both conditional likelihood and moment condition approaches to estimate common parameters in fixed effects variants of other standard nonlinear models. For instance, HausmanHallGriliches84 applied conditional likelihood methods to panel data versions of Poisson regression models. Similarly, Honore92,HONORE199335 and Hu2002 developed estimators for static and dynamic censored and truncated regression models using moment functions. Additionally, Chamberlain_1992 and Wooldridge_1997 utilized moment conditions in a class of static and dynamic multiplicative models. A hybrid approach is taken by Kyriazidou97, Kyriazidou01, who combined conditional likelihood and moment conditions to construct estimators for static and dynamic sample selection models.
The fixed effects approach discussed earlier, and further developed in this paper, places no assumptions on the relationship between the individual-specific effects and the explanatory variables. A common alternative is to adopt a random effects framework, which assumes a specific distribution for the individual effects and estimates the model parameters via maximum likelihood. However, this approach faces two key challenges.
First, in many economic applications, explanatory variables are at least partially chosen by individuals. In such cases, it is often implausible to assume independence between these variables and the individual effects (see Mundlak1961). Second, dynamic models introduce the initial conditions problem: the distribution of the initial dependent variable typically depends on both the individual effects and the explanatory variables in a complex way (see Heckman81b).
A solution to these issues, originally proposed by Mundlak1978, is the correlated random effects model. This approach specifies a model for the individual effects conditional on the explanatory variables and, in dynamic settings, also on the initial conditions. For further discussion, see wooldridge2005.\footnote{For a recent overview of the historical distinction between fixed and random effects in panel data, see BellemareMillimet2025.}
The remainder of the paper is organized as follows. Section (ref) reviews and extends the conditional likelihood approach for estimating common parameters in static logit models. The key insight from this section is that under relatively simple conditions, one can identify a sufficient statistic for the fixed effects, which then serves as the basis for conditional likelihood estimation of the common parameters—specifically, the coefficients on strictly exogenous explanatory variables. Section (ref) turns to dynamic models in which the only explanatory variables are lagged outcomes. In this setting, there are again straightforward conditions under which a sufficient statistic for the fixed effects can be found. However, unlike the static case, the resulting conditional likelihood may fail to identify all the common parameters of the model. This limitation motivates Section (ref), which explores more general dynamic specifications. Within the logit framework, we demonstrate that it is often possible to construct moment conditions that allow for estimation of the parameters even when no non-trivial sufficient statistic exists, or when conditioning on a sufficient statistic yields a likelihood function that is uninformative about a subset of the parameters of interest. Section (ref) concludes.
We observe data $(Y_t, X_t, w_t)$ for $t = 1, \ldots, T$, where $Y_t \in \{0,1\}$ is a random binary outcome, $X_t \in \mathbb{R}^{d_x}$ is a random vector of covariates whose associated coefficient $\beta \in \mathbb{R}^{d_x}$ is the parameter of interest, and $w_t \in \mathbb{R}^{d_w}$ are non-random vectors associated with an unobserved effect $A \in \mathbb{R}^{d_w}$.\footnote{ Most results extend to random $w_t$, but since all our examples involve non-random $w_t$, we focus on that case. } We denote the outcome vector as $Y = (Y_1, \ldots, Y_T)' \in \mathcal{Y} = \{0,1\}^T$, and we write $X = (X_1, \ldots, X_T) \in \mathcal{X} \subset \mathbb{R}^{d_x \times T}$ and $\mathsf{W} = (w_1, \ldots, w_T) \in \mathbb{R}^{d_w \times T}$. We use capital letters to denote random variables (such as $Y_t$, $X_t$, and $A$) and upright font to denote non-random matrices (such as $\mathsf{W}$). The parameter $\beta$ is treated as non-random, while the fixed effect $A$ is allowed to be random.
This is a standard assumption for binary logit models with fixed effects. We could write all probability statements explicitly conditional on $\mathsf{W}$ (e.g.\ ${\rm Pr}(Y = y \mid \mathsf{W}, X, A)$ in Assumption (ref)), but since $\mathsf{W}$ is treated as non-random throughout, we omit this for notational simplicity.
Since we consider $A \in \mathbb{R}^{d_w}$, Assumption (ref) implies that
However, all results in this section extend to cases where (components of) $A$ may take values in $\pm \infty$, provided that the boundedness condition in (ref) continues to hold.
These are two classic examples of models satisfying the structure in Assumption (ref). Various generalizations are discussed later.
Our goal is to identify the parameter $\beta$ without imposing any assumptions on the distribution of the unobserved effect $A$. In the static binary logit model of Assumption (ref), it is well known that $\mathsf{W} Y$ is a sufficient statistic for $A$, implying that conditioning on this statistic removes dependence on $A$. Formally, for any two outcome vectors $y_1, y_2 \in \mathcal{Y} = \{0,1\}^T$ such that $\mathsf{W} y_1 = \mathsf{W} y_2$, one can show that
which implies that the distribution $Y$ conditional on $Y \in \{y_1, y_2\}$, $X$, and $A$ does not depend on $A$, as long as $\mathsf{W} y_1 = \mathsf{W} y_2$:
To construct such identifying pairs $y_1, y_2$, it is useful to reformulate the condition $\mathsf{W} y_1 = \mathsf{W} y_2$ as requiring the existence of a difference vector $w_{\perp} = y_1 - y_2$ satisfying $\mathsf{W} w_{\perp} = 0$. The following lemma formalizes this observation.
The condition $\mathsf{W} w_{\perp} = 0$ is familiar from linear models: if the $T$-vector $Y$ satisfies $Y = X' \beta + \mathsf{W}' A + \varepsilon$ and $\mathsf{W} w_{\perp} = 0$, then pre-multiplying by $w_{\perp}'$ eliminates the nuisance term $\mathsf{W} A$. Remarkably, the same condition allows us to eliminate $A$ in the binary logit model as well --- provided that $w_{\perp}$ satisfies $w_{\perp} \in \{-1,0,1\}^T$, which implies that it corresponds to a difference of two binary vectors. Ultimately, this leads to the following identification result.
Crucially, this identification result imposes no restrictions on the distribution of $A$ or its dependence with $X$. The condition $\mathsf{W} \, \mathsf{W}_{\perp} = 0$ ensures that the columns of $\mathsf{W}_{\perp}$ are orthogonal to the fixed effects. The rank condition guarantees that these orthogonal directions yield sufficient variation in $X$ to identify all components of $\beta$.
Note that the number of columns in $\mathsf{W}_{\perp}$, denoted $d_\perp$, can be smaller than $d_x$. What matters is that the matrix $X \mathsf{W}_{\perp}$ has enough variation to identify $\beta$. In particular, the vector $\beta$ can be identified with $d_\perp = 1$.
The presentation here is closely related to classical conditional likelihood approaches (e.g., rasch1960studies,andersen1970asymptotic,chamberlain_analysis_1980), but the formulation in Theorem (ref) highlights a useful structural parallel to linear models. While the sufficiency argument is well known, we are not aware of prior work that frames the construction of identifying pairs through the condition $W w_{\perp} = 0$ with $w_{\perp} \in \{-1,0,1\}^T$. This perspective unifies a number of existing results and provides a clean basis for constructing and analyzing conditional likelihood estimators in settings with general fixed effect structures, as illustrated in the examples that follow.
We now illustrate Theorem (ref) in several models. Some of the examples are well-known, while others are novel. In each case, we describe the fixed effect structure and provide explicit constructions of vectors $w_{\perp}$ (or matrices $\mathsf{W}_{\perp}$) that satisfy the orthogonality condition $W w_{\perp} = 0$. We aim to carry over as much intuition as possible from the linear model analogy in Remark (ref).
This example corresponds to the setup in Example (ref). Recall that $d_w = 1$ and $w_t = 1$ for all $t \in \{1,\ldots,T\}$, so that the index in the logit model is $X_t' \beta + A$, with a scalar fixed effect $A$ shared across all time periods, and $\mathsf{W} = (1, \ldots, 1) \in \mathbb{R}^{1 \times T}$. To understand the identification strategy, consider first the analogous linear model, making the cross-sectional index $i$ explicit: $$ Y_{it} = X_{it}' \beta + A_i + \varepsilon_{it}, \qquad t = 1, \ldots, T, \quad i=1,\ldots,n. $$ Here, we can eliminate $A_i$ by differencing, e.g. for $T = 2$: $$ Y_{i1} - Y_{i2} = (X_{i1} - X_{i2})' \beta + (\varepsilon_{i1} - \varepsilon_{i2}). $$ This corresponds to the linear combination $w_{\perp}' Y_i$ with $w_{\perp} = (1, -1)'$. Crucially, since $w_{\perp} \in \{-1,0,1\}^T$, the same vector works for the logit model:
For the static panel logit with $T=2$ we know that $ {\rm Pr} \left( Y_i = y \, \big| \, X_i, A_i, Y_{i1} + Y_{i2} = 1 \right) $ is free of $A_i$. The conditioning event $Y_{i1} + Y_{i2} = 1$ restricts us to the pair $y_1=(1,0)'$ and $y_2=(0,1)'$, for which $w_{\perp} = y_1 - y_2 = (1,-1)'$. This recovers the classical results of rasch1960studies and andersen1970asymptotic. For $T > 2$, any vector in $\{-1,0,1\}^T$ with components summing to zero provides a valid $w_{\perp}$. If $X_i w_{\perp}$ varies sufficiently across units, then $\beta$ is identified.
Let $d_w = 2$, and for each $t \in \{1, \ldots, T\}$, define the regressors associated with the fixed effect as $w_t = (1,t)'$, so that the index becomes $ X_t' \, \beta + w_t' \, A = X_t' \beta + A_1 + t A_2, $ where $A_1$ is a unit-specific intercept and $A_2$ a unit-specific time trend. The matrix $\mathsf{W} \in \mathbb{R}^{2 \times T}$ consists of a row of ones and a row of time indices. Again, we consider the linear model analog, making the cross-sectional index $i$ explicit: $$ Y_{it} = X_{it}' \beta + A_{i1} + t A_{i2} + \varepsilon_{it}, \qquad t = 1, \ldots, T, \quad i=1,\ldots,n. $$ For the linear model, $T=3$ is sufficient to eliminate both fixed effects: $$ Y_{i1} - 2 Y_{i2} + Y_{i3} = (X_{i1} - 2 X_{i2} + X_{i3})' \beta + (\varepsilon_{i1} - 2\varepsilon_{i2} + \varepsilon_{i3}). $$ This linear combination corresponds to $w_{\perp}=(1,-2,1)'$, which does not satisfy $w_{\perp} \in \{-1,0,1\}^T$, and is therefore not applicable to the logit model. Indeed, for the static panel logit model with heterogeneous time trends and $T=3$, the coefficient $\beta$ is generally not point-identified.
To eliminate both $A_{i1}$ and $A_{i2}$ in the logit model, we need $w_{\perp} \in \{-1,0,1\}^T$ such that $\sum_{t=1}^T w_{\perp,t} = 0$ and $\sum_{t=1}^T w_{\perp,t} \, t = 0$. For $T = 4$ this is satisfied for $w_{\perp} = (1, -1, -1, 1)'$, which implies that in the static logit model with heterogeneous time trends, ${\rm Pr} \left( Y=y_1 \, \big| \, X,A \right) / {\rm Pr} \left( Y=y_2 \, \big| \, X,A \right)$ is independent of $A=(A_1,A_2)$ for $y_1 = (1,0,0,1)'$ and $y_2=(0,1,1,0)'$, since $w_{\perp} = y_1 - y_2$.
More generally, we can consider $w_t=(1,t,t^2,\ldots,t^p)'$, with corresponding index $$ X_t' \, \beta + w_t' \, A =X_t' \, \beta + A_1 + t \, A_2 + t^2 \, A_3 + \ldots + t^p A_{p+1} . $$ Table (ref) shows the minimal number of time periods $T$ and corresponding weight vectors $w_{\perp} \in \{-1,0,1\}^T$ that satisfy $W w_{\perp} = 0$ for this model for $p \leq 5$ (for $p=6$ the minimal $T$ is $31$). For $p \geq 1$, the minimal weight vectors exhibit an alternating symmetry pattern (even $p$ antisymmetric, odd $p$ symmetric) and can be constructed recursively by placing the $(p-1)$-solution and its reversed copy (negated for antisymmetric cases) with zero-padding and optimal overlap to minimize length while maintaining the constraint $w_{\perp} \in \{-1,0,1\}^T$.
Consider a setting with overlapping fixed effects where $T=3$ and $d_w=2$. Define $$ \mathsf{W} =
, $$ which yields the index structure $$ X_t' \, \beta + w_t' \, A =
$$ Note that observation $t=2$ is affected by both fixed effects, while observations $t=1$ and $t=3$ are each affected by only one. To eliminate both $A_1$ and $A_2$, we can use $w_{\perp}=(1,-1,1)'$, which satisfies $\mathsf{W} w_{\perp} = 0$ since each fixed effect appears with net zero weight. This may be the simplest non-trivial extension of the standard panel fixed effects model to more general fixed effects.
Consider two-way fixed effects panel models, where each observation is indexed by a unit-time pair $t = (i, \tau)$, with $i \in \{1,\ldots,n\}$ denoting individuals and $\tau \in \{1,\ldots,\mathcal{T}\}$ denoting time periods. The total number of observations is $T = n \cdot \mathcal{T}$, and the logit index takes the form $ X_{i\tau}' \beta + A_i + B_\tau, $ where $A_i$ and $B_\tau$ are unit and time fixed effects, respectively. Again, consider the linear model analogue:
For $n = \mathcal{T} = 2$, the standard difference-in-differences strategy eliminates both fixed effects: $$ (Y_{11} - Y_{12}) - (Y_{21} - Y_{22}) = [(X_{11} - X_{12}) - (X_{21} - X_{22})]' \beta + [(\varepsilon_{11} - \varepsilon_{12}) - (\varepsilon_{21} - \varepsilon_{22})]. $$ With $Y=(Y_{11},Y_{12},Y_{21},Y_{22})'$, the linear combination in the last display corresponds to $w_{\perp} = (1, -1, -1, 1)'$, which satisfies $w_{\perp} \in \{-1,0,1\}^T$ and is therefore applicable to the logit model as well.
For such $2 \times 2$ subpanels with two-way fixed effects, identification strategies based on $w_{\perp} = (1, -1, -1, 1)'$ have been developed by charbonneau2017multiple for binary logit models, and by jochmans2017two for certain nonlinear models with multiplicative unobservables. Such model structures arise naturally in applications such as matched employer-employee data or international trade.
It is convenient to think of $w_{\perp}$ as the vectorization of an $n \times \mathcal{T}$ matrix. In the $n = \mathcal{T} = 2$ case: $$ w_{\perp} = \mathrm{vec}
=
. $$ Our general condition $\mathsf{W} w_{\perp} = 0$ is equivalent to requiring that this $n \times \mathcal{T}$ matrix has all row and column sums equal to zero. When $n = \mathcal{T} = 3$, many valid vectors $w_{\perp}$ exist, for instance: $$ w_{\perp} = \mathrm{vec}
. $$ Thus, the construction generalizes to larger $n \times \mathcal{T}$ subpanels.
This corresponds to the structure in Example (ref). Recall that $T = n(n-1)/2$ and $d_w = n$, where each observation corresponds to an unordered dyad $(i,j)$ with $i \neq j$ and index structure $ X_{ij}' \beta + A_i + A_j$. The corresponding linear model reads $$ Y_{ij} = X_{ij}' \beta + A_i + A_j + \varepsilon_{ij}, $$ which is essentially the same model as (ref), except for the symmetry $Y_{ij} = Y_{ji}$ and that $Y_{ii}$ is unobserved.
For $n=3$ we cannot eliminate the fixed effects in this model, because we have the same number of observations $(Y_{12},Y_{13},Y_{23})$ as fixed effects $(A_1,A_2,A_3)$. However, for $n=4$ we can simply reproduce the same different strategy as in model (ref) by considering the $2 \times 2$ subpanel given by $i \in \{1,2\}$ and $j \in \{3,4\}$, which gives $$ (Y_{13} - Y_{14}) - (Y_{23} - Y_{24}) = [(X_{13} - X_{14}) - (X_{23} - X_{24})]' \beta + [(\varepsilon_{13} - \varepsilon_{14}) - (\varepsilon_{23} - \varepsilon_{24})]. $$ Defining the appropriate vectorization operator by $$ (Y_{12},Y_{13},Y_{14},Y_{23},Y_{24},Y_{34})' = {\rm vech} \left(
\right), $$ we can write the differencing vector used here as $$ w_{\perp} = {\rm vech} \left(
\right), $$ which satisfies $w_{\perp} \in \{-1,0,1\}^T$, and is therefore equally applicable to the logit model. Indeed, the tetrad configuration in \citet{graham2017econometric} corresponds exactly to this vector $w_{\perp}$ and the corresponding outcome pairs satisfying $w_{\perp} = y_1 - y_2$.
This idea extends to larger subnetworks. For instance, with $n = 5$, one finds valid examples such as: $$ w_{\perp} = {\rm vech} \left(
\right), \quad w_{\perp} = {\rm vech} \left(
\right). $$ These $w_{\perp}$ again correspond to identifying configurations that satisfy the conditions of Theorem (ref).
This example generalizes the two-way panel model structure to three dimensions, following MurisPakel2025. Each observation is indexed by a triad $t = (i,j,k)$, with $i \in \{1,\ldots,n_1\}$, $j \in \{1,\ldots,n_2\}$, and $k \in \{1,\ldots,n_3\}$ denoting elements from three disjoint partite sets. The logit index takes the form $ X_{ijk}' \beta + A_{ij} + B_{jk} + C_{ik}, $ where $A_{ij}$, $B_{jk}$, and $C_{ik}$ are pairwise fixed effects. Consider the linear model analogue: $$ Y_{ijk} = X_{ijk}' \beta + A_{ij} + B_{jk} + C_{ik} + \varepsilon_{ijk}. $$ To eliminate all three sets of pairwise fixed effects, we need a triple-differencing strategy. Consider a hexad consisting of two nodes from each part: $i \in \{i_1, i_2\}$, $j \in \{j_1, j_2\}$, and $k \in \{k_1, k_2\}$. The triple difference $$[(Y_{i_1 j_1 k_1} - Y_{i_1 j_1 k_2}) - (Y_{i_1 j_2 k_1} - Y_{i_1 j_2 k_2})] - [(Y_{i_2 j_1 k_1} - Y_{i_2 j_1 k_2}) - (Y_{i_2 j_2 k_1} - Y_{i_2 j_2 k_2})] $$ eliminates all pairwise fixed effects. This corresponds to $w_{\perp} = (1, -1, 1, -1, 1, -1, 1, -1)'$, and since $w_{\perp} \in \{-1,0,1\}^T$, the same vector works for the logit model. Each pair appears exactly twice in the eight triads: once with weight $+1$ and once with weight $-1$, ensuring that the net effect for each $(i,j)$, $(j,k)$, and $(i,k)$ pair is zero, thus satisfying $\mathsf{W} w_{\perp} = 0$.
This hexad-based identification strategy for triadic models was developed by MurisPakel2025. Such models arise naturally in applications where a “team” is formed by selecting one node from each of three sets—for instance, firm-industry-time interactions, or trade triplets involving importer, exporter, and product. Unlike in dyadic models, however, MurisPakel2025 show that even under sparsity, identification and asymptotic normality require stronger conditions: the presence of a growing number of informative hexads. Their framework illustrates how richer forms of unobserved heterogeneity, while modeling more realistic structures, come at a cost in terms of data requirements.
We now turn to dynamic binary choice logit models with generalized fixed effects, extending the approach used for static models in Section (ref). In the current section, we focus on purely dynamic specifications --- generalized autoregressive models --- without additional covariates $X$. As in the static case, our identification strategy relies on finding sufficient statistics that eliminate the fixed effects. In contrast, Section (ref) below considers dynamic models that include covariates $X$, where identification will instead be based on moment conditions that are invariant to the fixed effects.
Most of the notation carries over from the static case. In particular, the definitions of $Y$, $W$, and $A$ remain unchanged. However, we require some additional structure to capture the temporal dependence in the outcomes. Let $Y^0$ denote a vector of initial conditions, and define $$ Y^{t-1} = (Y_{t-1}, Y_{t-2}, \ldots, Y_1, Y^0) $$ to be the history of outcomes up to time $t-1$, including initial conditions. The following assumption gives the class of generalized autoregressive models considered in this section.
Note that the specification of the unobserved fixed effects $A \in \mathbb{R}^{d_w}$ and the non-random regressor vectors $w_t \in \mathbb{R}^{d_w}$ is unchanged from the static case. Crucially, the model imposes no restrictions on the dependence between the fixed effects $A$ and the initial condition $Y^0$.
The following lemma is central to all identification results in this section.
The proof is given in the appendix.
Our first example of a model that satisfies Assumption (ref) and for which Lemma (ref) yields immediate and useful identification results is the AR(1) model. In this case, we take $\theta = \gamma \in \mathbb{R}$, and set $ \pi_t(Y^{t-1}, \theta) = Y_{t-1} \, \gamma, $ with initial condition $Y^0 = Y_0$ given by the observed outcome in period $t = 0$. Under this specification, the model in Assumption (ref) becomes:
We are going to discuss the AR(1) model in detail here, and afterwards briefly summarize the extension to AR($p$) model for $p>1$, with details given in the appendix.
In order to obtain identification results that mirror the results for the static model above as close as possible, we furthermore impose the following assumption on $w_t = (w_{t,1}, \ldots, w_{t,d_w})$ here:
{\color{red} }These constraints on $w_t$ may appear overly restrictive at first glance. However, as we will explain in Remark (ref), these constraints are in fact without loss of generality from the perspective of constructing sufficient statistics for the fixed effect $A$ in this model. To simplify notation, we define the lagged outcome $T$-vector as $Y_{\rm lag} = (Y_0, Y_1, \ldots, Y_{T-1})'.$ Similarly, we define the “lead” version of $\mathsf{W} = (w_1, \ldots, w_T)$ by $\mathsf{W}_{\rm lead} = (w_2, w_3,\ldots, w_T,0_{d_w \times 1}) \in \mathbb{R}^{d_w \times T}$.
The proof is provided in the appendix. Part (i) of the theorem establishes that the sufficient statistic for the fixed effect $A$ in this model is given by the pair $(\mathsf{W}Y, \mathsf{W}Y_{\rm lag})$, which can be explicitly written as $\left(\sum_{t=1}^T w_t y_t,\sum_{t=1}^T w_t y_{t-1}\right)$. The first term, $\sum_{t=1}^T w_t y_t$, is familiar from the static model in Section (ref), where it formed the sufficient statistic for $A$. The second term, $\sum_{t=1}^T w_t y_{t-1}$, is new and arises due to the autoregressive structure.
To understand the role of these statistics, consider two outcome paths $y, \tilde y \in \{0,1\}^T$. Using only $\sum_{t=1}^T w_t y_t = \sum_{t=1}^T w_t \tilde{y}_t$ and $y_{0} = \widetilde y_0$, the likelihood ratio simplifies as follows: $$ \frac{{\rm Pr}(Y=y | Y_0=y_0, A)}{{\rm Pr}(Y=\tilde y | Y_0=y_0, A)} = \exp\left[\gamma \sum_{t=1}^{T}(y_{t} y_{t-1} - \tilde y_{t} \tilde y_{t-1}) \right] \prod_{t=2}^T \frac{1+\exp( \tilde y_{t-1} \gamma+ w_t' A)}{1+\exp( y_{t-1} \gamma+ w_t' A)} . $$ Here, the second term still depends on $A$, which shows that, in contrast to the static model, matching only $\sum_{t=1}^T w_t y_t$ is not sufficient to eliminate the fixed effect from the likelihood ratio. To fully eliminate $A$ from this ratio, we also need to use that $\sum_{t=2}^T y_{t-1} = \sum_{t=2}^T \tilde y_{t-1}$ and $\sum_{t=2}^T w_t y_{t-1} = \sum_{t=2}^T w_t \tilde{y}_{t-1}$ and $w_{t,k} \in \{0,1\}$ with $\sum_{k=1}^{d_w} w_{t,k} =1$. Under those conditions, one can show that $\prod_{t=2}^T \frac{1+\exp(\gamma \tilde y_{t-1} + w_t' A)}{1+\exp(\gamma y_{t-1} + w_t' A)}=1$, and therefore $$ \frac{{\rm Pr}(Y=y | Y_0=y_0, A)}{{\rm Pr}(Y=\tilde y | Y_0=y_0, A)} = \exp\left[\gamma \sum_{t=1}^{T}(y_{t} y_{t-1} - \tilde y_{t} \tilde y_{t-1}) \right] $$ This shows that $\gamma$ can be point-identified from the likelihood ratio, as long as $\sum_{t=1}^{T} y_{t} y_{t-1} \neq \sum_{t=1}^T \tilde y_{t} \tilde y_{t-1}$, which is precisely the content of part (ii) of the theorem.
\newenvironment{examplecont}[1]{ \refstepcounter{example}
}
We now generalize the previous example by considering an AR(2) model with $\pi_t(Y^{t-1}, \theta) = Y_{t-1} \, \gamma_1+ Y_{t-2} \, \gamma_2$ and initial condition $Y^0 = (Y_0,Y_{-1})\in\{0,1\}^2$. Under this setup, the model specified in Assumption (ref) becomes:
We refer readers to Appendix Section (ref) for a general treatment of the AR($p$) models with $p>1$. Throughout, we maintain the assumption that for all $t \in \{1, \ldots, T\}$, the vector of weights $w_t = (w_{t,1}, \ldots, w_{t,d_w})$ satisfies the restrictions in (ref).
Applying Lemma (ref), we show in Appendix Section (ref) that the likelihood ratio for two outcome paths $y,\tilde y\in \{0,1\}^T$ \footnote{These generalize the sufficient statistics found in Chamberlain84 and magnac2000subsidised.} $$\frac{{\rm Pr}(Y=y | Y_0=y_0, A)}{{\rm Pr}(Y=\tilde y | Y_0=y_0, A)} $$ is invariant to $A$ if the following four conditions hold:
The first condition coincides with Assumption (i) in Lemma (ref) while the remaining three are equivalent statements of Assumption (ii) for the AR(2). That is, they ensure together that $\Big[ \big(w_t,y_{t-1}\, \gamma_1+y_{t-2}\, \gamma_2\big) \, : \, t=2,\ldots,T\Big] $ is a permutation of $\Big[ \big(w_t,\tilde y_{t-1}\, \gamma_1+\tilde y_{t-2}\, \gamma_2\big) \, : \, t=2,\ldots,T\Big]$.
We now consider the dyadic network formation model of graham2016homophily, introduced in Example (ref). This model falls within the class of generalized autoregressive logit models studied above, and we show how our identification strategy based on Lemma (ref) applies in this setting.
We model binary outcomes $Y_{ij\tau} = Y_{ji\tau} \in \{0,1\}$ for individuals $i,j \in \{1,\ldots,n\}$ with $i \neq j$, and time periods $\tau \in \{1,\ldots,{\cal T}\}$. The initial condition at $\tau=0$ is denoted by $Y^0$. Each dyad-time pair $(i,j,\tau)$ defines a single observation, and we set $T = \binom{n}{2} \cdot \mathcal{T}$ as the total number of observations. Each observation $t = (i,j,\tau)$ is treated as an element of the outcome vector $Y = (Y_1,\ldots,Y_T)' \in \{0,1\}^T$.
To express the model in the form of Assumption (ref), we impose a deterministic ordering of dyads within each time period. The overall observation index $t$ respects time ordering, i.e.\ if $\tau_1 < \tau_2$ then all observations from period $\tau_1$ precede those from $\tau_2$. For each observation $t = (i,j,\tau)$, we define the regressor $w_t \in \mathbb{R}^{\binom{n}{2}}$ as a unit vector selecting the dyad $(i,j)$, so that $w_t' A = A_{ij}$, where $A \in \mathbb{R}^{\binom{n}{2}}$ collects all dyad-specific fixed effects. The logit index for observation $t$ is then \[ \pi_t(Y^{t-1}, \theta) + w_t' A = Y_{ij,\tau-1} \, \gamma + R_{ij,\tau-1} \, \delta + A_{ij}, \] with parameters $\theta = (\gamma,\delta)'$ and $R_{ij,\tau-1} = \sum_{k \notin \{i,j\}} Y_{ik,\tau-1} Y_{jk,\tau-1}$ denoting the number of shared friends in the previous period.
Throughout, we write $Y_{\cdot \cdot ,\tau}$ for the collection of $Y_{ij,\tau}$ over all dyads, and use ${\cal D}_n$ to denote the set of all ${n \choose 2}$ dyads. Given $Y_{\cdot \cdot ,\tau}$ and $(i,j) \in {\cal D}_n$, the covariates entering the index at time $\tau+1$ are \[ Z_{ij}(Y_{\cdot \cdot,\tau}) := \left( Y_{ij,\tau}, \sum_{k \neq i,j} Y_{ik,\tau} Y_{jk,\tau} \right) \in \{0,1\} \times \mathbb{Z}_+. \] We now describe the sufficient statistic structure implied by Lemma (ref). To simplify the exposition, we focus on the minimal case ${\cal T} = 3$, which is also the specification considered in graham2016homophily. For each outcome vector $y \in \{0,1\}^T$, we define the set of alternative sequences
By Lemma (ref), the conditional likelihood ratio ${\rm Pr}(Y=y \mid Y^0, A) / {\rm Pr}(Y=\tilde y \mid Y^0, A)$ is invariant to $A$ for all $\tilde y \in {\cal Y}_{\rm cond}(y)$. This yields the conditional likelihood \[ \mathcal{L}_{\text{cond}}(\gamma, \delta) = \frac{\prod_{(i,j) \in \mathcal{D}_n} \exp\left( \sum_{\tau=2}^{3} Y_{ij,\tau} \, (\gamma Y_{ij,\tau-1} + \delta R_{ij,\tau-1}) \right) }{ \sum_{\tilde y \in {\cal Y}_{\rm cond}(Y)} \prod_{(i,j) \in \mathcal{D}_n} \exp\left( \sum_{\tau=2}^{3} \tilde{y}_{ij,\tau} \, (\gamma \tilde{y}_{ij,\tau-1} + \delta \tilde R_{ij,\tau-1}) \right) }, \] where $\tilde R_{ij,\tau-1} = \sum_{k \notin \{i,j\}} \tilde{y}_{ik,\tau-1} \cdot \tilde{y}_{jk,\tau-1}$. While this formulation is exact, the denominator is computationally infeasible to evaluate for large $n$, as the set ${\cal Y}_{\rm cond}(y)$ grows exponentially with the number of dyads.
To address this, graham2016homophily proposes a restricted subset of ${\cal Y}_{\rm cond}(y)$ based on the concept of “stable dyads”, which allows feasible computation in large-network settings. An alternative approach, suitable for small $n$ but many independent network observations (as in graham2013comment), is to compute the exact conditional likelihood using the full set ${\cal Y}_{\rm cond}(y)$ or a tractable subset.
A particularly convenient and computationally efficient subset is the two-element set
This set contains at most two elements: the original network $y$ and the version with periods $\tau=1$ and $2$ swapped. When $y_{\cdot\cdot,1} = y_{\cdot\cdot,2}$, the set is a singleton. For $n=3$, one can show that ${\cal Y}^*_{\rm cond}(y) = {\cal Y}_{\rm cond}(y)$ for all networks $y$. For $n=4$ and $n=5$, numerical checks confirm that ${\cal Y}^*_{\rm cond}(y) = {\cal Y}_{\rm cond}(y)$ in more than 95% of all configurations, making this approximation attractive in small-network applications.
In summary, this framework supports two estimation strategies: Firstly, in settings with large $n$ and few time periods, feasible estimation can be achieved using restricted conditioning sets such as those based on stable dyads (graham2016homophily). Secondly, in settings with small $n$ but many repeated network observations, exact conditional likelihood estimation is tractable and efficient using either ${\cal Y}_{\rm cond}(y)$ or its simplification ${\cal Y}^*_{\rm cond}(y)$.
The results extend naturally to longer time horizons ${\cal T} > 3$, though the structure of the conditioning sets becomes more complex. Likewise, the framework accommodates richer dynamic specifications, provided the conditions in Lemma (ref) continue to hold.
We now turn to an alternative identification strategy for non-linear panel models based on moment conditions that do not depend on the fixed effects. Unlike the conditional likelihood approach used in Sections (ref) and (ref), which relies on sufficient statistics, the moment-based approach developed here applies to a larger class of models and accommodates richer forms of unobserved heterogeneity.
While fixed-effect-free moment conditions had been constructed in specific models before, the general idea was formalized in the context of semiparametric panel models by bonhomme2012functional, who called it “functional differencing”. Its applicability to discrete choice models was recently emphasized by kitazawa2021transformations and honore2024dynamic.
The model and notation are essentially unchanged compared to Section (ref), the only difference is that we now also allow for additional strictly exogenous covariates $X$, as we had in Section (ref). Assumption (ref) then generalizes as follows.
Our goal is again to identify the parameter $\theta$ without imposing restrictions on the distribution of the fixed effect $A$. The model in Assumption (ref) nests many dynamic panel and network models. Our baseline Examples (ref) and (ref) remain essentially unchanged, except for the additional covariates $X_t$. Thus, in Example (ref) the specification becomes $$ \pi_t(Y^{t-1}, X_t, \theta) = Y_{t-1} \, \gamma + X_t' \, \beta, $$ where $\theta=(\gamma,\beta)'$ while generalizing Example (ref) gives $$ \pi_t(Y^{t-1}, X_t, \theta) = Y_{ij,\tau-1} \, \gamma + R_{ij,\tau-1} \, \delta + X_t' \, \beta $$ with $\theta = (\gamma, \delta)'$.
It is sometimes possible to extend the identification strategy based on sufficient statistics to models with additional covariates. For example, honore2000panel apply this approach to an AR(1) panel logit model with covariates by conditioning on subsets of the data where covariate values are identical across adjacent time periods. More generally, the results in Section (ref) can be applied to such models by treating $X$ as fixed throughout and absorbing the covariate dependence of $\pi_t(Y^{t-1}, X_t, \theta)$ into the definition of $\pi_t$. However, the sufficiency-based approach typically works only under restrictive support conditions on the covariates. In contrast, the moment condition strategy developed in this section applies more broadly and often accommodates arbitrary variation in $X$.
As discussed above, the sufficient statistics approach typically breaks down in dynamic settings with covariates that vary freely over time. This motivates an alternative strategy based on moment functions that depend on the parameters of interest but not on the fixed effects. Specifically, we aim to find functions $m(Y, Y^0, X, \theta)$ such that
Moment functions satisfying (ref) are valid for identification and estimation, as they are invariant to the fixed effect $A$ and therefore unaffected by the incidental parameter problem. To characterize when such functions exist, we use a combinatorial argument that provides a lower bound on the dimension of the space of fixed-effect-free moment functions,\footnote{See also dano2023transition for an alternative strategy tailored to standard AR($p$) logit models with $p \geq 1$.} leading to the following theorem.
To state the theorem, we introduce some additional notation. For each time period $t \in \{1, \ldots, T\}$, define the set of distinct values that the index $\pi_t(Y^{t-1}, X_t, \theta)$ can take across all possible outcome histories $(y_{t-1},y_{t-2},\ldots,y_1) \in \{0,1\}^{t-1}$ as $$ \Pi_t( Y^0,X,\theta) := \big\{ \pi_t(( y_{t-1},y_{t-2},\ldots,y_1,Y^{0}),X_t,\theta) \, : \, (y_{t-1},y_{t-2},\ldots,y_1) \in \{0,1\}^{t-1} \big\} , $$ Throughout this section, we treat $Y^0$, $X$, and $\theta$ as fixed. While we make these arguments explicit in the notation for mathematical precision, they are otherwise unimportant to the combinatorial structure we analyze and can be regarded as fixed and given (and thus ignored) for the remainder of the discussion. Let $$ Q_t(Y^0, X,\theta) := |\Pi_t(Y^0, X,\theta)| $$ denote the number of elements in the set $\Pi_t(Y^0, X, \theta)$, that is, the number of distinct values that the index $\pi_t((y_{t-1}, y_{t-2}, \ldots, y_1, Y^0), X_t, \theta)$ can take across different outcome histories. By construction, we have $Q_t(Y^0, X, \theta) \leq 2^{t-1}$, but in many practical models we have limited dependence on past outcomes and $Q_t(Y^0, X, \theta)$ is therefore often (much) smaller. Also, in most models, $Q_t(Y^0, X, \theta)$ remains constant across typical values of $Y^0$, $X$, and $\theta$, but may be different at specific points (e.g., when $\theta = 0$).
Next, define the set of all possible linear combinations of the $w_t$'s, where each coefficient $k_t$ is an integer between 0 and $Q_t(Y^0, X, \theta)$:
The remark and examples below will help clarify the intuition behind the definition of the set $\mathcal{D}(Y^0, X, \theta)$. As before, we write $|\mathcal{D}(Y^0, X, \theta)|$ to denote its cardinality. The following result provides a lower bound on the number of linearly independent fixed-effect-free moment conditions that exist in the class of dynamic models considered in this section.
Theorem (ref) provides a sufficient condition for the existence of non-zero moment functions $m(Y, Y^0, X, \theta)$ satisfying (ref). However, the result does not characterize the structure of these functions, nor does it guarantee that they depend on $\theta$ in a way that allows identification. In practice, constructing such moment functions, and verifying that they identify $\theta$, is typically model-specific and requires additional algebraic or numerical insight. Concrete examples are discussed below.
Even so, the existence result in Theorem (ref) is useful. For example, once existence is guaranteed, one can apply the general functional-analytic framework developed by bonhomme2012functional to compute the moment conditions numerically. In this “functional differencing” approach, moment functions can be approximated even when closed-form expressions are unavailable.
Whether obtained analytically or numerically, valid moment functions $m(Y, Y^0, X, \theta)$ can be used to construct GMM estimators that are root-$n$ consistent under standard regularity conditions.
We first consider the class of AR(\(p\)) panel logit models with covariates and scalar fixed effects. These models are covered by Assumption (ref), with the index and fixed effect specification given by
where $\theta = (\gamma_1, \ldots, \gamma_p,\beta')'$. The initial conditions $Y^0 = (Y_{-p+1}, \ldots, Y_0) \in \{0,1\}^p$ are treated as fixed and observed.
To apply Theorem (ref), we compute the number of distinct values that the index $\pi_t(Y^{t-1}, X_t, \theta)$ can take across binary outcome histories. Assuming $\gamma_r \neq 0$ for all $r \in \{1,\ldots,p\}$, this number is $$ Q_t = 2^{\min(p,\,t-1)}, $$ since the index depends on the most recent binary $p$ outcomes, or fewer in the initial periods. Because we are considering $w_t = 1$ for all $t$, the set $\mathcal{D} = \mathcal{D}(Y^0, X, \theta)$ consists of all integers between 0 and the maximum sum $d_{\max} := \sum_{t=1}^T Q_t$. Therefore, we have
By Theorem (ref), the number of linearly independent fixed-effect-free moment functions is therefore at least $$ 2^T - 2^p (T + 1 - p), $$ which is strictly positive whenever $T \geq 2+p$ (in addition to $p$ observed time periods as initial conditions). It turns out that this is not merely a lower bound: for generic values of the covariates $X$, this is in fact the exact number of linearly independent moment conditions in the AR(\(p\)) panel model. However, as we will see in the following example, additional moment conditions may become available for specific values of $X$, in particular when $X = 0$.
The result for the number of linearly independent moment conditions in the AR($p$) model was stated in honore2024dynamic, who also derive explicit analytical expressions for the corresponding moment functions in the cases $p = 1, 2, 3$, and provide a detailed analysis of identification and estimation for $p = 1$, building on the results of kitazawa2021transformations. The sharpness of the moment count has been established for $p = 1$ and $p = 2$ by kruiniger2020further, and further discussed in dobronyi2021identification. For general $p$, dano2023transition proves that the bound is sharp under free-varying covariates and provides analytical expressions for all the moment functions in closed form.
We now consider a special case of the previous example where $p = 2$ and no covariates are present. The model takes the form $$ \Pr(Y_t = 1 \mid Y_{t-1}, Y_{t-2}, A) = \frac{ \exp\left( Y_{t-1} \gamma_1 + Y_{t-2} \gamma_2 + A \right) } {1 + \exp\left( Y_{t-1} \gamma_1 + Y_{t-2} \gamma_2 + A \right)}, $$ which corresponds to the structure in Assumption (ref) with
In this model for $T = 3$, one can construct one valid fixed-effect-free moment function for each value of the inital condition $Y^0 = (Y_{-1}, Y_0) \in \{0,1\}^2$. For instance, when $y^0 = (0,0)$, the following function satisfies the moment condition (ref): $$ m(y, y^0, \theta) =
$$ and when $y^0 = (0,1)$, a valid moment function is given by $$ m(y, y^0, \theta) =
$$ see Section~3.3 of the first arXiv version of \citet{honore2024dynamic}, who provide closed-form moment functions for the AR($p$) model with $T = 3$ and $X_2=X_3$.
Firstly, this example shows that the lower bound in Theorem (ref) is not always sharp. For $p = 2$ and $T = 3$, we have $2^T - |\mathcal{D}| = 0$, so the theorem does not guarantee the existence of any fixed-effect-free moment functions. Nevertheless, such functions do exist, as demonstrated above.
Secondly, the first moment function implies that $\gamma_1$ is point-identified in this model, since $\exp(-\gamma_1)$ is strictly monotonic in $\gamma_1$, and thus the moment function is strictly monotonic as well. Once $\gamma_1$ is identified, the second moment function then identifies $\gamma_2$ through the same logic. This result provides analytical confirmation of earlier numerical findings in honore2019identification that suggested identification might be possible in this setup, even though conditional likelihood methods do not apply.
Thus, the moment condition approach point-identifies both $\gamma_1$ and $\gamma_2$ in this model as soon as $T \geq 3$. By contrast, as noted in Remark (ref), the sufficient statistics approach can identify only $\gamma_2$, and requires at least $T \geq 4$ to do so.
Next, consider the extension of Example (ref) with strictly exogenous covariates. For this specification, we have $\pi_t(Y^{t-1}, X_t, \theta) = Y_{t-1} \gamma + X_t' \beta$ and $w_t=(w_{t,1},w_{t,2},w_{t,3},w_{t,4})$ where $w_{t,q}$ is an indicator for quarter $q\in \{1,2,3,4\}$. It follows that
with cardinality $|\mathcal{D}(Y^0, X,\theta)| = \left(2 \lfloor \frac{T - 1}{4}\rfloor+2\right)\left(2 \lfloor \frac{T - 2}{4}\rfloor+3\right)\left(2 \lfloor \frac{T - 3}{4}\rfloor+3\right)\left(2 \lfloor \frac{T}{4}\rfloor+1\right)$. Hence, the bound $2^T-|\mathcal{D}(Y^0, X,\theta)|$ of Theorem (ref) becomes positive as soon as $T\geq 12$, ensuring the existence of moments. While informative, this bound is conservative in this instance since in the absence of regressors (which does not alter the bound here), Example (ref) already indicated the existence of identifying moments for $T=6$. Indeed, one can construct two linearly independent moment functions $m_{1}$ and $m_{2}$ with $T=6$ periods. Using the shorthand $x_{ts}=x_{t}-x_{s}$ for $t\neq s$, $m_1$ is given by
and $m_{2}(y, y_0,x,\theta)$ is obtained by substituting $y_{t}$ by $(1-y_{t})$ for $t=0,\ldots, 6$ and $x$ by $-x$ in the expression of $m_{1}(y, y_0,x,\theta)$. In other words, $m_{2}(y, y_0,x,\theta)= m_{1}(1-y, 1-y_0,-x,\theta)$. We refer readers to Appendix (ref) for detailed derivations of these expressions. Remark that the two moment functions depend on the common parameter $\theta=(\gamma,\beta')'$.
As an illustration of how the lower bound of Theorem (ref) can be applied to probe the existence of moment conditions in networks, consider the extension of graham2016homophily introduced earlier, incorporating strictly exogenous covariates $X$. In this setting, we specify $\pi_t(Y^{t-1}, X_t, \theta) = Y_{ij,\tau-1} \gamma + R_{ij,\tau-1} \delta + X_t' \beta$, where $t=(i,j,\tau)$ indexes dyad-time pairs, and $w_t$ denotes the basis vector in $\mathbb{R}^{\binom{n}{2}}$ with entry one for dyad $(i,j)$, and zeros elsewhere. Under this formulation, the model described in Assumption (ref) becomes
Since $R_{ij,\tau-1} := \sum_{k \neq i,j} Y_{ik,\tau-1} Y_{jk,\tau-1}$, for any fixed covariates $X$, the structure of $\pi_t$ implies that it can take at most $2(n-1)$ distinct values. Consequently, we have
and hence
A straightforward counting argument then gives $|\mathcal{D}(Y^0, X,\theta)| = \left[2(\mathcal{T}-1)(n-1)+2\right]^{\binom{n}{2}}$ and Theorem (ref) implies the existence of at least $2^T - \left|{\cal D}(Y_0,X) \right|=2^T-\left[2(\mathcal{T}-1)(n-1)+2\right]^{\binom{n}{2}}$ moment functions free from the fixed effects. Since $\log |\mathcal{D}(Y^0, X,\theta)|$ grows approximately proportionally to $\log \mathcal{T}$ - the number of time periods - while $\log(2^T)=\log(2)\binom{n}{2}\mathcal{T}$ grows linearly in $\mathcal{T}$, the existence of moment conditions is guaranteed for sufficiently large $\mathcal{T}$. For example, in the simplest case with $n=3$ agents, the lower bound is positive as soon as $T\geq 4$. In fact, dano2023transition shows that there is generally as much as $\binom{n}{2}$ fixed-effect-free moment conditions with only $T=3$ periods that are explicitly given by:
where $y$ denotes an undirected network, and $r_{ij}:= \sum_{k \neq i,j} y_{ik} y_{jk}$ denotes the number of friends that agents $i$ and $j$ share in common in $y$.
This paper has reviewed and extended two approaches for eliminating fixed effects in logit models: the conditional likelihood method and the construction of moment conditions. While the results are conceptually clean and methodologically promising, several important challenges remain, pointing to avenues for future research.
First, the moment-based framework often yields a large number of conditional moment conditions, linking it to the broader literature on optimal moment selection in panel data. Prior work, including bekker1994alternative, donald2001choosing, alvarez2003time, and okui2009optimal, has shown that using too many valid moments can introduce substantial finite-sample bias. This suggests that future research on nonlinear models should integrate identification strategies with principled approaches to moment selection.
Second, our focus has been on identifying and estimating structural parameters, rather than computing counterfactuals or marginal effects. In panel data settings, these quantities are typically not point-identified, even when the structural parameters are. Extending the methods discussed here to incorporate bounds on marginal effects, such as those proposed by PakelWeidner2024 and davezies2024identification, would enhance the empirical relevance of the fixed effects framework.
Finally, and perhaps most critically, the models considered here assume strict exogeneity of the explanatory variables (aside from lagged outcomes). This assumption is often unrealistic in economic applications. While recent work by ArellanoCarrasco2003, botosaru2024adversarial, bonhomme2023identification, and BonhommeDanoGraham2025 has made progress in relaxing this assumption, much remains to be done to develop robust methods that accommodate predetermined regressors in nonlinear panel models.