EconBase
← Back to paper

Binary choice logit models with general fixed effects for panel and network data

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

106,520 characters · 29 sections · 85 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Binary choice logit models with general fixed effects for panel and network data

\thispagestyle{empty} \setcounter{page}{0}

abstractThis paper systematically analyzes and reviews identification strategies for binary choice logit models with fixed effects in panel and network data settings. We examine both static and dynamic models with general fixed-effect structures, including individual effects, time trends, and two-way or dyadic effects. A key challenge is the incidental parameter problem, which arises from the increasing number of fixed effects as the sample size grows. We explore two main strategies for eliminating nuisance parameters: conditional likelihood methods, which remove fixed effects by conditioning on sufficient statistics, and moment-based methods, which derive fixed-effect-free moment conditions. We demonstrate how these approaches apply to a variety of models, summarizing key findings from the literature while also presenting new examples and new results.

Introduction

Binary outcome models with fixed effects have a long history in economics and related fields. They allow for flexible modeling of individual decisions while accounting for unobserved heterogeneity. These models arise naturally in two types of data structures: traditional panel data, where individuals or firms are observed over time, and network or dyadic data, where observations correspond to interactions between pairs of units. In both cases, researchers often wish to estimate the effect of observed covariates on a binary outcome, while allowing for unit-specific or pair-specific unobserved components. Examples include individual labor market choices, firm entry decisions, and the formation of social or economic networks.

A central challenge in these models is the treatment of fixed effects, which capture unobserved heterogeneity. When treated as parameters to be estimated, fixed effects can lead to the incidental parameter problem, first identified by neyman1948consistent. This issue arises because the number of fixed effects grows with the sample size. If the number of observations per parameter increases slowly, as in certain network models or panel data settings with increasing number of time periods, estimators may be asymptotically biased. In such cases, bias correction methods can be applied; see, for example, HahnNewey2004, and the surveys by ArellanoHahn2007 and FernandezValWeidner2018. On the other hand, when the number of parameters grows proportionally with the number of observations, as in standard panel data models with a fixed number of time periods, estimators are often inconsistent. In these settings, the most effective strategy is to eliminate the fixed effects.\footnote{In an alternative approach, HigginsJochmans2024 show that in parametric settings like the one considered in this paper, it is possible to use the bootstrap to do valid inference on the common parameters using the inconsistent and biased estimator that treats the fixed effects as parameters to be estimated.}

In this paper, we summarize and extend two existing methods for estimating binary choice logit models with fixed effects by eliminating them from the estimation problem. Our focus is on both static and dynamic models, allowing for very general fixed effect structures. We discuss models for standard panel data, as well as dyadic models that arise in the study of networks.

Two main general approaches have been developed to address the incidental parameter problem. The first is based on conditional likelihood and requires a parametric model conditional on the unobservable fixed effects. This method eliminates the fixed effects by conditioning on sufficient statistics. It was first introduced by rasch1960studies in the context of educational testing and later adapted to econometric applications by chamberlain_analysis_1980. In binary logit models, the fixed effects drop out of the likelihood when we condition on certain linear combinations of the outcomes. This approach has been extended to more complex fixed effect structures, including two-way fixed effects in dyadic data charbonneau2017multiple and network formation models graham2017econometric. Conditional likelihood methods are attractive because they provide a way to estimate structural parameters without modeling the distribution of the unobserved effects. However, they are only available in models where sufficient statistics can be found in closed form.

The second approach relies on moment conditions that do not depend on the fixed effects.\footnote{In this paper we focus on moment \textsl{equality} conditions. Moment \textsl{inequalities} form the basis for the semiparametric approach in Manski87, and have also been used to construct and estimate identified sets for the structural parameters. See, for example, PakesPorterShepardCalderWang2022.} This idea is sometimes called functional differencing. It was formalized by bonhomme2012functional and applied to discrete choice models by kitazawa2021transformations and honore2024dynamic. The basic idea is to construct functions of the data whose conditional expectation is zero for all values of the fixed effects. These functions then define valid moment conditions for the structural parameters. This approach has also been used to develop estimators in dynamic logit models with general fixed effects by dano2023transition. Moment-based methods can be used in a wider range of models than conditional likelihood. They are especially useful in dynamic models or in models with more complicated fixed effect structures.

This paper reviews both approaches in the context of binary choice logit models. We present the underlying identification ideas in a unified framework and show how they apply to standard examples. These include static panel models with individual fixed effects, models with heterogeneous time trends, two-way fixed effects, and network formation settings. For dynamic models, we discuss how moment conditions can be constructed under general assumptions, and how these conditions relate to the number of fixed effects in the model. We also give examples from dynamic panel and dynamic network settings. We focus on the logit model, rather than, say, the probit model or a semiparametric version, because it is known from chamberlain2010binary that in the simplest case of a static binary response panel data model with two time periods, regular root-$n$ consistent estimation of the common parameters is only possible in the logit model.\footnote{For a discussion of what can be learned about the common parameters in dynamic binary response models with panel data in a more semiparametric setting, see for example KHAN2023105515.}

The literature has long employed both conditional likelihood and moment condition approaches to estimate common parameters in fixed effects variants of other standard nonlinear models. For instance, HausmanHallGriliches84 applied conditional likelihood methods to panel data versions of Poisson regression models. Similarly, Honore92,HONORE199335 and Hu2002 developed estimators for static and dynamic censored and truncated regression models using moment functions. Additionally, Chamberlain_1992 and Wooldridge_1997 utilized moment conditions in a class of static and dynamic multiplicative models. A hybrid approach is taken by Kyriazidou97, Kyriazidou01, who combined conditional likelihood and moment conditions to construct estimators for static and dynamic sample selection models.

The fixed effects approach discussed earlier, and further developed in this paper, places no assumptions on the relationship between the individual-specific effects and the explanatory variables. A common alternative is to adopt a random effects framework, which assumes a specific distribution for the individual effects and estimates the model parameters via maximum likelihood. However, this approach faces two key challenges.

First, in many economic applications, explanatory variables are at least partially chosen by individuals. In such cases, it is often implausible to assume independence between these variables and the individual effects (see Mundlak1961). Second, dynamic models introduce the initial conditions problem: the distribution of the initial dependent variable typically depends on both the individual effects and the explanatory variables in a complex way (see Heckman81b).

A solution to these issues, originally proposed by Mundlak1978, is the correlated random effects model. This approach specifies a model for the individual effects conditional on the explanatory variables and, in dynamic settings, also on the initial conditions. For further discussion, see wooldridge2005.\footnote{For a recent overview of the historical distinction between fixed and random effects in panel data, see BellemareMillimet2025.}

The remainder of the paper is organized as follows. Section (ref) reviews and extends the conditional likelihood approach for estimating common parameters in static logit models. The key insight from this section is that under relatively simple conditions, one can identify a sufficient statistic for the fixed effects, which then serves as the basis for conditional likelihood estimation of the common parameters—specifically, the coefficients on strictly exogenous explanatory variables. Section (ref) turns to dynamic models in which the only explanatory variables are lagged outcomes. In this setting, there are again straightforward conditions under which a sufficient statistic for the fixed effects can be found. However, unlike the static case, the resulting conditional likelihood may fail to identify all the common parameters of the model. This limitation motivates Section (ref), which explores more general dynamic specifications. Within the logit framework, we demonstrate that it is often possible to construct moment conditions that allow for estimation of the parameters even when no non-trivial sufficient statistic exists, or when conditioning on a sufficient statistic yields a likelihood function that is uninformative about a subset of the parameters of interest. Section (ref) concludes.

Static logit models

Model setup

We observe data $(Y_t, X_t, w_t)$ for $t = 1, \ldots, T$, where $Y_t \in \{0,1\}$ is a random binary outcome, $X_t \in \mathbb{R}^{d_x}$ is a random vector of covariates whose associated coefficient $\beta \in \mathbb{R}^{d_x}$ is the parameter of interest, and $w_t \in \mathbb{R}^{d_w}$ are non-random vectors associated with an unobserved effect $A \in \mathbb{R}^{d_w}$.\footnote{ Most results extend to random $w_t$, but since all our examples involve non-random $w_t$, we focus on that case. } We denote the outcome vector as $Y = (Y_1, \ldots, Y_T)' \in \mathcal{Y} = \{0,1\}^T$, and we write $X = (X_1, \ldots, X_T) \in \mathcal{X} \subset \mathbb{R}^{d_x \times T}$ and $\mathsf{W} = (w_1, \ldots, w_T) \in \mathbb{R}^{d_w \times T}$. We use capital letters to denote random variables (such as $Y_t$, $X_t$, and $A$) and upright font to denote non-random matrices (such as $\mathsf{W}$). The parameter $\beta$ is treated as non-random, while the fixed effect $A$ is allowed to be random.

assumptionThe data-generating process is \begin{align*} {\rm Pr} \left( Y=y \, \big| \, X,A \right) &= \prod_{t=1}^T \, \left[ \frac {1} {1+\exp(X_t' \, \beta + w_t' \, A) } \right]^{1-y_t} \left[ \frac {\exp(X_t' \, \beta + w_t' \, A)} {1+\exp(X_t' \, \beta + w_t' \, A) } \right]^{y_t} . \end{align*}

This is a standard assumption for binary logit models with fixed effects. We could write all probability statements explicitly conditional on $\mathsf{W}$ (e.g.\ ${\rm Pr}(Y = y \mid \mathsf{W}, X, A)$ in Assumption (ref)), but since $\mathsf{W}$ is treated as non-random throughout, we omit this for notational simplicity.

Since we consider $A \in \mathbb{R}^{d_w}$, Assumption (ref) implies that

align[align omitted — 169 chars of source]

However, all results in this section extend to cases where (components of) $A$ may take values in $\pm \infty$, provided that the boundedness condition in (ref) continues to hold.

example[Standard fixed effects in panel data:] Consider $d_w = 1$ and $w_t = 1$ for all $t \in \{1, \ldots, T\}$. Then $$ X_t' \, \beta + w_t' \, A = X_t' \, \beta + A $$ corresponds to the standard panel logit model with an individual-specific intercept. We observe an i.i.d.\ sample of $(Y,X)$, denoted by $\{(Y_i, X_i) : i = 1, \ldots, n\}$, where the data for each unit $i$ satisfy Assumption (ref). The index in the logit specification becomes $X_{it}' \, \beta + A_i$. $\mathsf{W} = (1, \ldots, 1)$ and $\beta$ are non-random and constant across units $i$.
example[Dyadic fixed effects in network data:] Consider $d_w = n$ and $T = n(n - 1)/2$, where each observation $t = (i,j) = (j,i)$ corresponds to an unordered pair of distinct units $i, j \in \{1, \ldots, n\}$, with $i \neq j$. For each dyad $(i,j)$, we observe a binary outcome $Y_{ij} \in \{0,1\}$ indicating whether a link is present between units $i$ and $j$. Let $X_{ij} \in \mathbb{R}^{d_x}$ denote dyad-level covariates capturing observed characteristics of the pair, such as homophily measures. Define $w_{ij} \in \mathbb{R}^n$ as a selection vector such that\footnote{ For each unordered pair $t = (i,j)$ with $i \neq j$, define the vector $w_{ij} \in \mathbb{R}^n$ by $$ (w_{ij})_k = \begin{cases} 1, & \text{if } k = i \text{ or } k = j, \\ 0, & \text{otherwise}. \end{cases} $$ } $$ X_{ij}' \, \beta + w_{ij}' \, A = X_{ij}' \, \beta + A_i + A_j $$ corresponds to a network formation model with unit-specific fixed effects. This specification captures both homophily (through $X_{ij}$) and degree heterogeneity (through $A_i$ and $A_j$), as in graham2017econometric. The model can be interpreted as describing a small subnetwork within a larger graph, and by considering many such subnetworks, we obtain multiple independent or weakly dependent observations of $(Y, X)$. For example, graham2017econometric considers subnetworks of size $n=4$ to construct his “tetrad logit” estimator. Again, $\mathsf{W}$ and $\beta$ are non-random and constant across subnetworks.

These are two classic examples of models satisfying the structure in Assumption (ref). Various generalizations are discussed later.

Identification via conditional likelihood

Our goal is to identify the parameter $\beta$ without imposing any assumptions on the distribution of the unobserved effect $A$. In the static binary logit model of Assumption (ref), it is well known that $\mathsf{W} Y$ is a sufficient statistic for $A$, implying that conditioning on this statistic removes dependence on $A$. Formally, for any two outcome vectors $y_1, y_2 \in \mathcal{Y} = \{0,1\}^T$ such that $\mathsf{W} y_1 = \mathsf{W} y_2$, one can show that

align[align omitted — 211 chars of source]

which implies that the distribution $Y$ conditional on $Y \in \{y_1, y_2\}$, $X$, and $A$ does not depend on $A$, as long as $\mathsf{W} y_1 = \mathsf{W} y_2$:

align[align omitted — 222 chars of source]

To construct such identifying pairs $y_1, y_2$, it is useful to reformulate the condition $\mathsf{W} y_1 = \mathsf{W} y_2$ as requiring the existence of a difference vector $w_{\perp} = y_1 - y_2$ satisfying $\mathsf{W} w_{\perp} = 0$. The following lemma formalizes this observation.

lemmaThere exist $y_1,y_2 \in \{0,1\}^T$ with $y_1 \neq y_2$ and $\mathsf{W} y_1 = \mathsf{W} y_2$ if and only if there exists $w_{\perp} \in \{-1,0,1\}^T$ with $w_{\perp} \neq 0$ and $\mathsf{W} w_{\perp} = 0$.

The condition $\mathsf{W} w_{\perp} = 0$ is familiar from linear models: if the $T$-vector $Y$ satisfies $Y = X' \beta + \mathsf{W}' A + \varepsilon$ and $\mathsf{W} w_{\perp} = 0$, then pre-multiplying by $w_{\perp}'$ eliminates the nuisance term $\mathsf{W} A$. Remarkably, the same condition allows us to eliminate $A$ in the binary logit model as well --- provided that $w_{\perp}$ satisfies $w_{\perp} \in \{-1,0,1\}^T$, which implies that it corresponds to a difference of two binary vectors. Ultimately, this leads to the following identification result.

theoremSuppose Assumption (ref) holds. Assume further that for some integer $d_\perp \geq 1$, there exists a non-random matrix $\mathsf{W}_\perp \in \{-1, 0, 1\}^{T \times d_\perp}$ such that $\mathsf{W} \, \mathsf{W}_\perp = 0$, and there is no vector $b \in \mathbb{R}^{d_x}$ for which $b' X \mathsf{W}_\perp = 0$ almost surely. Then, the parameter $\beta$ is uniquely identified from the distribution of $(Y,X)$.

Crucially, this identification result imposes no restrictions on the distribution of $A$ or its dependence with $X$. The condition $\mathsf{W} \, \mathsf{W}_{\perp} = 0$ ensures that the columns of $\mathsf{W}_{\perp}$ are orthogonal to the fixed effects. The rank condition guarantees that these orthogonal directions yield sufficient variation in $X$ to identify all components of $\beta$.

Note that the number of columns in $\mathsf{W}_{\perp}$, denoted $d_\perp$, can be smaller than $d_x$. What matters is that the matrix $X \mathsf{W}_{\perp}$ has enough variation to identify $\beta$. In particular, the vector $\beta$ can be identified with $d_\perp = 1$.

remarkThe identification condition in Theorem (ref) closely parallels that of a standard linear fixed effects model. Suppose we observe $ Y = X' \beta + \mathsf{W}' A + \varepsilon$. In this setting, the parameter $\beta$ is identified if there exists a matrix $\mathsf{W}_{\perp} \in \mathbb{R}^{T \times d_\perp}$ such that $\mathsf{W} \mathsf{W}_{\perp} = 0$ and $X \mathsf{W}_{\perp}$ is non-collinear. Pre-multiplying the linear model equation by $\mathsf{W}_{\perp}'$ yields $ \mathsf{W}_{\perp}' Y = \mathsf{W}_{\perp}' X' \beta + \mathsf{W}_{\perp}' \varepsilon, $ in which $A$ no longer appears. The resulting equation can then be estimated by ordinary least squares (OLS). In the binary choice case, we are subject to the additional restriction that $\mathsf{W}_{\perp}$ must have integer entries in $\{-1,0,1\}$, since it must represent differences between binary outcome vectors. Nonetheless, the identification conditions for the logit model in Theorem (ref) are directly analogous to that of the linear model, modulo this constraint.
remarkAn alternative way to understand our identification result is via the population conditional maximum likelihood estimator (CMLE), as in davezies2024identification. Given any $s \in \mathbb{R}^{d_w}$, define $\mathcal{S}(s) = \{y \in \{0,1\}^T : \mathsf{W}y = s \}$ and consider the conditional log-likelihood \[ \ell_c(\beta \mid y,x,s) := \log \left[ {\rm Pr}(Y = y \mid X = x, \mathsf{W}Y = s) \right] , \] which is well-defined and free of $A$ due to sufficiency of $\mathsf{W}Y$. If the expected Hessian \[ H_{\rm CMLE} := \mathbb{E}\left[ - \frac{\partial^2 \ell_c(\beta \mid Y,X,\mathsf{W}Y)}{\partial \beta \, \partial \beta'} \right] \] is positive definite, then $\beta$ is the unique maximizer of the population CMLE objective. To connect this to Theorem (ref), observe that \[ H_{\rm CMLE} = \mathbb{E}\left[ \sum_{y_1,y_2 \in \mathcal{S}(\mathsf{W}Y)} \omega(\mathsf{W}Y, X, y_1, y_2) \, X (y_1 - y_2) (y_1 - y_2)' X' \right] \] for some positive weights $\omega(\mathsf{W}Y, X, y_1, y_2) > 0$. This shows that variation in $X(y_1 - y_2)$ along directions satisfying $\mathsf{W}(y_1 - y_2) = 0$ is essential for identification of $\beta$ under CMLE, just as in our rank condition.

The presentation here is closely related to classical conditional likelihood approaches (e.g., rasch1960studies,andersen1970asymptotic,chamberlain_analysis_1980), but the formulation in Theorem (ref) highlights a useful structural parallel to linear models. While the sufficiency argument is well known, we are not aware of prior work that frames the construction of identifying pairs through the condition $W w_{\perp} = 0$ with $w_{\perp} \in \{-1,0,1\}^T$. This perspective unifies a number of existing results and provides a clean basis for constructing and analyzing conditional likelihood estimators in settings with general fixed effect structures, as illustrated in the examples that follow.

Examples

We now illustrate Theorem (ref) in several models. Some of the examples are well-known, while others are novel. In each case, we describe the fixed effect structure and provide explicit constructions of vectors $w_{\perp}$ (or matrices $\mathsf{W}_{\perp}$) that satisfy the orthogonality condition $W w_{\perp} = 0$. We aim to carry over as much intuition as possible from the linear model analogy in Remark (ref).

Standard fixed effects in panel data

This example corresponds to the setup in Example (ref). Recall that $d_w = 1$ and $w_t = 1$ for all $t \in \{1,\ldots,T\}$, so that the index in the logit model is $X_t' \beta + A$, with a scalar fixed effect $A$ shared across all time periods, and $\mathsf{W} = (1, \ldots, 1) \in \mathbb{R}^{1 \times T}$. To understand the identification strategy, consider first the analogous linear model, making the cross-sectional index $i$ explicit: $$ Y_{it} = X_{it}' \beta + A_i + \varepsilon_{it}, \qquad t = 1, \ldots, T, \quad i=1,\ldots,n. $$ Here, we can eliminate $A_i$ by differencing, e.g. for $T = 2$: $$ Y_{i1} - Y_{i2} = (X_{i1} - X_{i2})' \beta + (\varepsilon_{i1} - \varepsilon_{i2}). $$ This corresponds to the linear combination $w_{\perp}' Y_i$ with $w_{\perp} = (1, -1)'$. Crucially, since $w_{\perp} \in \{-1,0,1\}^T$, the same vector works for the logit model:

For the static panel logit with $T=2$ we know that $ {\rm Pr} \left( Y_i = y \, \big| \, X_i, A_i, Y_{i1} + Y_{i2} = 1 \right) $ is free of $A_i$. The conditioning event $Y_{i1} + Y_{i2} = 1$ restricts us to the pair $y_1=(1,0)'$ and $y_2=(0,1)'$, for which $w_{\perp} = y_1 - y_2 = (1,-1)'$. This recovers the classical results of rasch1960studies and andersen1970asymptotic. For $T > 2$, any vector in $\{-1,0,1\}^T$ with components summing to zero provides a valid $w_{\perp}$. If $X_i w_{\perp}$ varies sufficiently across units, then $\beta$ is identified.

Heterogeneous time trends

Let $d_w = 2$, and for each $t \in \{1, \ldots, T\}$, define the regressors associated with the fixed effect as $w_t = (1,t)'$, so that the index becomes $ X_t' \, \beta + w_t' \, A = X_t' \beta + A_1 + t A_2, $ where $A_1$ is a unit-specific intercept and $A_2$ a unit-specific time trend. The matrix $\mathsf{W} \in \mathbb{R}^{2 \times T}$ consists of a row of ones and a row of time indices. Again, we consider the linear model analog, making the cross-sectional index $i$ explicit: $$ Y_{it} = X_{it}' \beta + A_{i1} + t A_{i2} + \varepsilon_{it}, \qquad t = 1, \ldots, T, \quad i=1,\ldots,n. $$ For the linear model, $T=3$ is sufficient to eliminate both fixed effects: $$ Y_{i1} - 2 Y_{i2} + Y_{i3} = (X_{i1} - 2 X_{i2} + X_{i3})' \beta + (\varepsilon_{i1} - 2\varepsilon_{i2} + \varepsilon_{i3}). $$ This linear combination corresponds to $w_{\perp}=(1,-2,1)'$, which does not satisfy $w_{\perp} \in \{-1,0,1\}^T$, and is therefore not applicable to the logit model. Indeed, for the static panel logit model with heterogeneous time trends and $T=3$, the coefficient $\beta$ is generally not point-identified.

To eliminate both $A_{i1}$ and $A_{i2}$ in the logit model, we need $w_{\perp} \in \{-1,0,1\}^T$ such that $\sum_{t=1}^T w_{\perp,t} = 0$ and $\sum_{t=1}^T w_{\perp,t} \, t = 0$. For $T = 4$ this is satisfied for $w_{\perp} = (1, -1, -1, 1)'$, which implies that in the static logit model with heterogeneous time trends, ${\rm Pr} \left( Y=y_1 \, \big| \, X,A \right) / {\rm Pr} \left( Y=y_2 \, \big| \, X,A \right)$ is independent of $A=(A_1,A_2)$ for $y_1 = (1,0,0,1)'$ and $y_2=(0,1,1,0)'$, since $w_{\perp} = y_1 - y_2$.

More generally, we can consider $w_t=(1,t,t^2,\ldots,t^p)'$, with corresponding index $$ X_t' \, \beta + w_t' \, A =X_t' \, \beta + A_1 + t \, A_2 + t^2 \, A_3 + \ldots + t^p A_{p+1} . $$ Table (ref) shows the minimal number of time periods $T$ and corresponding weight vectors $w_{\perp} \in \{-1,0,1\}^T$ that satisfy $W w_{\perp} = 0$ for this model for $p \leq 5$ (for $p=6$ the minimal $T$ is $31$). For $p \geq 1$, the minimal weight vectors exhibit an alternating symmetry pattern (even $p$ antisymmetric, odd $p$ symmetric) and can be constructed recursively by placing the $(p-1)$-solution and its reversed copy (negated for antisymmetric cases) with zero-padding and optimal overlap to minimize length while maintaining the constraint $w_{\perp} \in \{-1,0,1\}^T$.

table[table omitted — 669 chars of source]

Overlapping fixed effects

Consider a setting with overlapping fixed effects where $T=3$ and $d_w=2$. Define $$ \mathsf{W} =

pmatrix[pmatrix omitted — 37 chars of source]

, $$ which yields the index structure $$ X_t' \, \beta + w_t' \, A =

casesX_1' \, \beta + A_1 & if t=1, \\ X_2' \, \beta + A_1 + A_2 & if t=2, \\ X_3' \, \beta + A_2 & if t=3.

$$ Note that observation $t=2$ is affected by both fixed effects, while observations $t=1$ and $t=3$ are each affected by only one. To eliminate both $A_1$ and $A_2$, we can use $w_{\perp}=(1,-1,1)'$, which satisfies $\mathsf{W} w_{\perp} = 0$ since each fixed effect appears with net zero weight. This may be the simplest non-trivial extension of the standard panel fixed effects model to more general fixed effects.

Two-way fixed effects

Consider two-way fixed effects panel models, where each observation is indexed by a unit-time pair $t = (i, \tau)$, with $i \in \{1,\ldots,n\}$ denoting individuals and $\tau \in \{1,\ldots,\mathcal{T}\}$ denoting time periods. The total number of observations is $T = n \cdot \mathcal{T}$, and the logit index takes the form $ X_{i\tau}' \beta + A_i + B_\tau, $ where $A_i$ and $B_\tau$ are unit and time fixed effects, respectively. Again, consider the linear model analogue:

align[align omitted — 107 chars of source]

For $n = \mathcal{T} = 2$, the standard difference-in-differences strategy eliminates both fixed effects: $$ (Y_{11} - Y_{12}) - (Y_{21} - Y_{22}) = [(X_{11} - X_{12}) - (X_{21} - X_{22})]' \beta + [(\varepsilon_{11} - \varepsilon_{12}) - (\varepsilon_{21} - \varepsilon_{22})]. $$ With $Y=(Y_{11},Y_{12},Y_{21},Y_{22})'$, the linear combination in the last display corresponds to $w_{\perp} = (1, -1, -1, 1)'$, which satisfies $w_{\perp} \in \{-1,0,1\}^T$ and is therefore applicable to the logit model as well.

For such $2 \times 2$ subpanels with two-way fixed effects, identification strategies based on $w_{\perp} = (1, -1, -1, 1)'$ have been developed by charbonneau2017multiple for binary logit models, and by jochmans2017two for certain nonlinear models with multiplicative unobservables. Such model structures arise naturally in applications such as matched employer-employee data or international trade.

It is convenient to think of $w_{\perp}$ as the vectorization of an $n \times \mathcal{T}$ matrix. In the $n = \mathcal{T} = 2$ case: $$ w_{\perp} = \mathrm{vec}

pmatrix[pmatrix omitted — 31 chars of source]

=

pmatrix[pmatrix omitted — 33 chars of source]

. $$ Our general condition $\mathsf{W} w_{\perp} = 0$ is equivalent to requiring that this $n \times \mathcal{T}$ matrix has all row and column sums equal to zero. When $n = \mathcal{T} = 3$, many valid vectors $w_{\perp}$ exist, for instance: $$ w_{\perp} = \mathrm{vec}

pmatrix[pmatrix omitted — 53 chars of source]

. $$ Thus, the construction generalizes to larger $n \times \mathcal{T}$ subpanels.

Dyadic network formation

This corresponds to the structure in Example (ref). Recall that $T = n(n-1)/2$ and $d_w = n$, where each observation corresponds to an unordered dyad $(i,j)$ with $i \neq j$ and index structure $ X_{ij}' \beta + A_i + A_j$. The corresponding linear model reads $$ Y_{ij} = X_{ij}' \beta + A_i + A_j + \varepsilon_{ij}, $$ which is essentially the same model as (ref), except for the symmetry $Y_{ij} = Y_{ji}$ and that $Y_{ii}$ is unobserved.

For $n=3$ we cannot eliminate the fixed effects in this model, because we have the same number of observations $(Y_{12},Y_{13},Y_{23})$ as fixed effects $(A_1,A_2,A_3)$. However, for $n=4$ we can simply reproduce the same different strategy as in model (ref) by considering the $2 \times 2$ subpanel given by $i \in \{1,2\}$ and $j \in \{3,4\}$, which gives $$ (Y_{13} - Y_{14}) - (Y_{23} - Y_{24}) = [(X_{13} - X_{14}) - (X_{23} - X_{24})]' \beta + [(\varepsilon_{13} - \varepsilon_{14}) - (\varepsilon_{23} - \varepsilon_{24})]. $$ Defining the appropriate vectorization operator by $$ (Y_{12},Y_{13},Y_{14},Y_{23},Y_{24},Y_{34})' = {\rm vech} \left(

array[array omitted — 143 chars of source]

\right), $$ we can write the differencing vector used here as $$ w_{\perp} = {\rm vech} \left(

array[array omitted — 87 chars of source]

\right), $$ which satisfies $w_{\perp} \in \{-1,0,1\}^T$, and is therefore equally applicable to the logit model. Indeed, the tetrad configuration in \citet{graham2017econometric} corresponds exactly to this vector $w_{\perp}$ and the corresponding outcome pairs satisfying $w_{\perp} = y_1 - y_2$.

This idea extends to larger subnetworks. For instance, with $n = 5$, one finds valid examples such as: $$ w_{\perp} = {\rm vech} \left(

array[array omitted — 127 chars of source]

\right), \quad w_{\perp} = {\rm vech} \left(

array[array omitted — 131 chars of source]

\right). $$ These $w_{\perp}$ again correspond to identifying configurations that satisfy the conditions of Theorem (ref).

Triadic network models

This example generalizes the two-way panel model structure to three dimensions, following MurisPakel2025. Each observation is indexed by a triad $t = (i,j,k)$, with $i \in \{1,\ldots,n_1\}$, $j \in \{1,\ldots,n_2\}$, and $k \in \{1,\ldots,n_3\}$ denoting elements from three disjoint partite sets. The logit index takes the form $ X_{ijk}' \beta + A_{ij} + B_{jk} + C_{ik}, $ where $A_{ij}$, $B_{jk}$, and $C_{ik}$ are pairwise fixed effects. Consider the linear model analogue: $$ Y_{ijk} = X_{ijk}' \beta + A_{ij} + B_{jk} + C_{ik} + \varepsilon_{ijk}. $$ To eliminate all three sets of pairwise fixed effects, we need a triple-differencing strategy. Consider a hexad consisting of two nodes from each part: $i \in \{i_1, i_2\}$, $j \in \{j_1, j_2\}$, and $k \in \{k_1, k_2\}$. The triple difference $$[(Y_{i_1 j_1 k_1} - Y_{i_1 j_1 k_2}) - (Y_{i_1 j_2 k_1} - Y_{i_1 j_2 k_2})] - [(Y_{i_2 j_1 k_1} - Y_{i_2 j_1 k_2}) - (Y_{i_2 j_2 k_1} - Y_{i_2 j_2 k_2})] $$ eliminates all pairwise fixed effects. This corresponds to $w_{\perp} = (1, -1, 1, -1, 1, -1, 1, -1)'$, and since $w_{\perp} \in \{-1,0,1\}^T$, the same vector works for the logit model. Each pair appears exactly twice in the eight triads: once with weight $+1$ and once with weight $-1$, ensuring that the net effect for each $(i,j)$, $(j,k)$, and $(i,k)$ pair is zero, thus satisfying $\mathsf{W} w_{\perp} = 0$.

This hexad-based identification strategy for triadic models was developed by MurisPakel2025. Such models arise naturally in applications where a “team” is formed by selecting one node from each of three sets—for instance, firm-industry-time interactions, or trade triplets involving importer, exporter, and product. Unlike in dyadic models, however, MurisPakel2025 show that even under sparsity, identification and asymptotic normality require stronger conditions: the presence of a growing number of informative hexads. Their framework illustrates how richer forms of unobserved heterogeneity, while modeling more realistic structures, come at a cost in terms of data requirements.

Main Takeaways

itemize• For the static binary logit models with fixed effects in Assumption (ref), it is well known that conditioning on the sufficient statistic $\mathsf{W}Y$ removes dependence on the fixed effects $A$. However, identification of $\beta$ requires more than sufficiency: we need outcome pairs $(y_1, y_2)$ such that $\mathsf{W} y_1 = \mathsf{W} y_2$ and $X(y_1 - y_2)$ exhibits sufficient variation. This condition is central to Theorem (ref). • A key insight is the structural parallel to linear models (Remark (ref)): Differencing strategies that eliminate fixed effects in linear models also eliminate them in binary logit models --- provided the differencing vector satisfies $w_{\perp} \in \{-1,0,1\}^T$. This constraint reflects the binary nature of the data and leads to a unifying framework for constructing valid conditional likelihood estimators across a wide range of panel and network models.

Sufficient statistic for dynamic logit models

Model setup

We now turn to dynamic binary choice logit models with generalized fixed effects, extending the approach used for static models in Section (ref). In the current section, we focus on purely dynamic specifications --- generalized autoregressive models --- without additional covariates $X$. As in the static case, our identification strategy relies on finding sufficient statistics that eliminate the fixed effects. In contrast, Section (ref) below considers dynamic models that include covariates $X$, where identification will instead be based on moment conditions that are invariant to the fixed effects.

Most of the notation carries over from the static case. In particular, the definitions of $Y$, $W$, and $A$ remain unchanged. However, we require some additional structure to capture the temporal dependence in the outcomes. Let $Y^0$ denote a vector of initial conditions, and define $$ Y^{t-1} = (Y_{t-1}, Y_{t-2}, \ldots, Y_1, Y^0) $$ to be the history of outcomes up to time $t-1$, including initial conditions. The following assumption gives the class of generalized autoregressive models considered in this section.

assumptionThe data-generating process is \begin{align*} {\rm Pr} \left( Y=y \, \big| \, Y^0,A \right) &= \prod_{t=1}^T {\rm Pr} \left( Y_t=y_t \, \big| \, Y^{t-1},A \right) , \\ {\rm Pr} \left( Y_t=y_t \, \big| \, Y^{t-1},A \right) &= \frac {\left[\exp(\pi_t(Y^{t-1},\theta) + w_t' \, A)\right]^{y_t}} {1+\exp(\pi_t(Y^{t-1},\theta) + w_t' \, A) } , \end{align*} where $\theta$ is an unknown parameter and $\pi_t(\cdot,\cdot)$ is a known function for every $t \in \{1,\ldots,T\}$.

Note that the specification of the unobserved fixed effects $A \in \mathbb{R}^{d_w}$ and the non-random regressor vectors $w_t \in \mathbb{R}^{d_w}$ is unchanged from the static case. Crucially, the model imposes no restrictions on the dependence between the fixed effects $A$ and the initial condition $Y^0$.

example[Dynamic panel data] This example generalizes Example (ref) by allowing for state dependence through a lagged dependent variable. Specifically, suppose that $\pi_t(Y^{t-1}, \theta) = Y_{t-1} \, \gamma,$ so that the model becomes: $$ \Pr(Y_t = 1 \mid Y^{t-1}, A) = \frac{ \exp( Y_{t-1} \, \gamma + A) }{ 1 + \exp( Y_{t-1} \, \gamma + A) }, $$ where $\gamma \in \mathbb{R}$ captures state dependence while $A \in \mathbb{R}$ captures individual-specific heterogeneity as before. For this baseline model, there are well-known results due to Chamberlain78 and based on cox1958regression that show that how to identify and estimate $\gamma$ via conditional likelihood by conditioning on sufficient statistics based on transition counts.\footnote{Chamberlain84, magnac2004panel and d2010duration generalize those result further.} Our goal in this section is to generalize those existing results to more general fixed effect structures.
example[Dynamic dyadic network formation] This example extends the dyadic model of Example (ref) to a dynamic setting. Suppose we observe a sequence of undirected binary networks among $n$ agents over periods $\tau = 0, 1, \ldots, \mathcal{T}$. Each observation corresponds to a dyad $(i,j)$, and the link indicator $Y_{ij\tau} = Y_{ji\tau} \in \{0,1\}$ records whether a link between $i$ and $j$ exists in period $\tau$. The dynamic logit model considered in graham2016homophily takes the form: \[ \Pr(Y_{ij\tau} = 1 \mid Y_{ij,\tau-1}, R_{ij,\tau-1}, A_{ij}) = \frac{ \exp\left( Y_{ij,\tau-1} \, \gamma + R_{ij,\tau-1} \, \delta + A_{ij} \right) } { 1 + \exp\left( Y_{ij,\tau-1} \, \gamma + R_{ij,\tau-1} \, \delta + A_{ij} \right) }, \] where $ Y_{ij,\tau-1}$ is the lagged link indicator for the same dyad, $R_{ij,\tau-1} := \sum_{k \notin \{i,j\}} Y_{ik,\tau-1} Y_{jk,\tau-1}$ is the number of shared friends of $i$ and $j$ in the previous period, $A_{ij} \in \mathbb{R}$ is a dyad-specific fixed effect, and $\gamma$ and $\delta$ are unknown parameters. To map this into our general framework (Assumption (ref)), we define $T = \binom{n}{2} \cdot \mathcal{T}$, where $\mathcal{T}$ is the number of observed time periods in Graham’s model, and $\binom{n}{2}$ is the number of dyads. We then treat each dyad-time pair $(i,j,\tau)$ as a single observation indexed by $t = 1, \ldots, T$, see Section (ref) below for more details. In this formulation, the logit index for observation $t$ becomes $\pi_t(Y^{t-1}, X\theta) = Y_{ij,\tau-1} \, \gamma + R_{ij,\tau-1} \, \delta$.

Generalized sufficient statistics

The following lemma is central to all identification results in this section.

lemmaSuppose Assumption (ref) holds with fixed initial conditions $Y^0 = y^0$. Let $y, \tilde y \in \{0,1\}^T$ be any two outcome sequences, and define $y^{t-1} = (y_{t-1}, y_{t-2}, \ldots, y_1, y^0)$ and $\tilde y^{\, t-1} = (\tilde y_{t-1}, \tilde y_{t-2}, \ldots, \tilde y_1, y^0)$, for each $t \in \{2,\ldots,T\}$. Assume further that: \begin{itemize} • $\sum_{t=1}^T w_t y_t = \sum_{t=1}^T w_t \tilde{y}_t$. • $\Big[ \big(w_t,\pi_t(y^{t-1},\theta)\big) \, : \, t=2,\ldots,T\} \Big]$ is a permutation of $\Big[ \big(w_t,\pi_t(\tilde y^{t-1},\theta)\big) \, : \, t=2,\ldots,T\} \Big].$ \end{itemize} Then we have $$ \frac{{\rm Pr}(Y=y | Y^0=y^0, A)}{{\rm Pr}(Y=\tilde y | Y^0=y^0, A)} = \exp\left\{ \sum_{t=1}^{T}\left[ y_{t} \, \pi_t(y^{t-1},\theta) - \tilde y_{t} \, \pi_t(\tilde y^{t-1},\theta) \right] \right\} , $$ which does not depend on $A$.

The proof is given in the appendix.

AR($p$) panel data models

Our first example of a model that satisfies Assumption (ref) and for which Lemma (ref) yields immediate and useful identification results is the AR(1) model. In this case, we take $\theta = \gamma \in \mathbb{R}$, and set $ \pi_t(Y^{t-1}, \theta) = Y_{t-1} \, \gamma, $ with initial condition $Y^0 = Y_0$ given by the observed outcome in period $t = 0$. Under this specification, the model in Assumption (ref) becomes:

align[align omitted — 286 chars of source]

We are going to discuss the AR(1) model in detail here, and afterwards briefly summarize the extension to AR($p$) model for $p>1$, with details given in the appendix.

In order to obtain identification results that mirror the results for the static model above as close as possible, we furthermore impose the following assumption on $w_t = (w_{t,1}, \ldots, w_{t,d_w})$ here:

align[align omitted — 216 chars of source]

{\color{red} }These constraints on $w_t$ may appear overly restrictive at first glance. However, as we will explain in Remark (ref), these constraints are in fact without loss of generality from the perspective of constructing sufficient statistics for the fixed effect $A$ in this model. To simplify notation, we define the lagged outcome $T$-vector as $Y_{\rm lag} = (Y_0, Y_1, \ldots, Y_{T-1})'.$ Similarly, we define the “lead” version of $\mathsf{W} = (w_1, \ldots, w_T)$ by $\mathsf{W}_{\rm lead} = (w_2, w_3,\ldots, w_T,0_{d_w \times 1}) \in \mathbb{R}^{d_w \times T}$.

theoremWe assume the AR(1) model in (ref), with $w_t$ satisfying the restrictions in (ref). We treat the initial condition $Y_0 = y_0$ as fixed and known. \begin{enumerate}[(i)] • Then, $(\mathsf{W}Y, \mathsf{W}Y_{\rm lag})$ is a statistic that is sufficient for the nuisance parameter $A$ (conditional on $Y_0$), in the sense that the distribution of $Y$ given $(\mathsf{W}Y, \mathsf{W}Y_{\rm lag})$ and $Y_0$ does not depend on $A$. Since $Y_0 = y_0$ is treated as fixed, we can equivalently express this sufficient statistic as the linear transformation $\mathsf{V} Y$, where $\mathsf{V} := (\mathsf{W}' ,\mathsf{W}'_{\rm lead})'$. • Suppose further that there exist two binary vectors $y, \tilde y \in \{0,1\}^T$ such that $\mathsf{V} y = \mathsf{V} \tilde y $ and $\sum_{t=1}^{T} y_{t} y_{t-1} \neq \sum_{t=1}^T \tilde y_{t} \tilde y_{t-1}$, where we set $Y_0=y_0 = \tilde y_0$. Then, the autoregressive parameter $\gamma$ is uniquely identified from the distribution of $Y$ conditional on $Y_0 = y_0$. \end{enumerate}

The proof is provided in the appendix. Part (i) of the theorem establishes that the sufficient statistic for the fixed effect $A$ in this model is given by the pair $(\mathsf{W}Y, \mathsf{W}Y_{\rm lag})$, which can be explicitly written as $\left(\sum_{t=1}^T w_t y_t,\sum_{t=1}^T w_t y_{t-1}\right)$. The first term, $\sum_{t=1}^T w_t y_t$, is familiar from the static model in Section (ref), where it formed the sufficient statistic for $A$. The second term, $\sum_{t=1}^T w_t y_{t-1}$, is new and arises due to the autoregressive structure.

To understand the role of these statistics, consider two outcome paths $y, \tilde y \in \{0,1\}^T$. Using only $\sum_{t=1}^T w_t y_t = \sum_{t=1}^T w_t \tilde{y}_t$ and $y_{0} = \widetilde y_0$, the likelihood ratio simplifies as follows: $$ \frac{{\rm Pr}(Y=y | Y_0=y_0, A)}{{\rm Pr}(Y=\tilde y | Y_0=y_0, A)} = \exp\left[\gamma \sum_{t=1}^{T}(y_{t} y_{t-1} - \tilde y_{t} \tilde y_{t-1}) \right] \prod_{t=2}^T \frac{1+\exp( \tilde y_{t-1} \gamma+ w_t' A)}{1+\exp( y_{t-1} \gamma+ w_t' A)} . $$ Here, the second term still depends on $A$, which shows that, in contrast to the static model, matching only $\sum_{t=1}^T w_t y_t$ is not sufficient to eliminate the fixed effect from the likelihood ratio. To fully eliminate $A$ from this ratio, we also need to use that $\sum_{t=2}^T y_{t-1} = \sum_{t=2}^T \tilde y_{t-1}$ and $\sum_{t=2}^T w_t y_{t-1} = \sum_{t=2}^T w_t \tilde{y}_{t-1}$ and $w_{t,k} \in \{0,1\}$ with $\sum_{k=1}^{d_w} w_{t,k} =1$. Under those conditions, one can show that $\prod_{t=2}^T \frac{1+\exp(\gamma \tilde y_{t-1} + w_t' A)}{1+\exp(\gamma y_{t-1} + w_t' A)}=1$, and therefore $$ \frac{{\rm Pr}(Y=y | Y_0=y_0, A)}{{\rm Pr}(Y=\tilde y | Y_0=y_0, A)} = \exp\left[\gamma \sum_{t=1}^{T}(y_{t} y_{t-1} - \tilde y_{t} \tilde y_{t-1}) \right] $$ This shows that $\gamma$ can be point-identified from the likelihood ratio, as long as $\sum_{t=1}^{T} y_{t} y_{t-1} \neq \sum_{t=1}^T \tilde y_{t} \tilde y_{t-1}$, which is precisely the content of part (ii) of the theorem.

remarkThe restrictions $w_{t,k} \in \{0,1\}$ and $\sum_{k=1}^{d_w} w_{t,k} \in \{0,1\}$ in Assumption (ref) may appear overly restrictive. However, these conditions are without loss of generality for the purpose of constructing sufficient statistics and identifying $\gamma$ via conditional likelihood methods. To see this, consider the general autoregressive model where $w_t \in \mathbb{R}^{d_w}$ can take any values. Let $\Omega = \{\varphi_1, \varphi_2, \ldots, \varphi_{d_\omega}\} \subset \mathbb{R}^{d_w}$ denote the set of distinct values that $w_t$ assumes across $t = 1, \ldots, T$, where $d_\omega$ is the cardinality of $\Omega$. We can then define indicator variables $\omega_{t,k} = \mathbbm{1}(w_t = \varphi_k)$, for $k = 1, \ldots, d_\omega$, implying that $\omega_t = (\omega_{t,1}, \ldots, \omega_{t,d_\omega})'$ satisfies the binary restrictions on $w_t$ in (ref). We then have $w_t' A = \omega_t' A^*$, where $A^* = (\varphi_1' A, \ldots, \varphi_{d_\omega}' A)' \in \mathbb{R}^{d_\omega}$. One can show that the sufficient statistics for $A$ in the original model with $w_t' A$ are identical to those for $A^*$ in the transformed model with $\omega_t' A^*$, as characterized in Theorem (ref).
remarkApplying Lemma (ref) in the context of part (ii) of Theorem (ref) gives the following: There exist $y,\tilde y \in \{0,1\}^T$ with $y \neq \tilde y$ and $\mathsf{V} y_1 = \mathsf{V} y_2 $ if and only if there exists $w_{\perp} \in \{-1,0,1\}^T$ with $w_{\perp} \neq 0$ and $\mathsf{V} w_{\perp} = 0$. This can be computationally very useful to construct such pairs $y,\tilde y \in \{0,1\}^T$ that allow identification of $\gamma$. However, the additional identification condition $\sum_{t=1}^{T} y_{t} y_{t-1} \neq \sum_{t=1}^T \tilde y_{t} \tilde y_{t-1}$ is non-linear in the outcomes and can therefore not be expressed in terms of $w_{\perp}$ only.

\newenvironment{examplecont}[1]{ \refstepcounter{example}

example}{

}

examplecont{ex:DynamicPanel} The simplest example covered by Theorem (ref) is the case where $w_t = 1$ for all $t = 1, \ldots, T$, implying a standard individual-specific fixed effect $A$ that enters identically in every period. In this setting, the model reduces to a pure logit AR(1) panel model with a scalar fixed effect and no covariates. The sufficient statistics for $A$ in this model — namely, $(\sum_{t=1}^T Y_t, \sum_{t=1}^T Y_{t-1})$ — are well known since cox1958regression, and are often expressed in terms of $Y_0$ (which we condition on throughout), $\sum_{t=1}^{T-1} Y_t$, and $Y_T$. To identify the autoregressive parameter $\gamma$, one must observe at least $T = 3$ periods (in addition to the initial condition $Y_0$). For example, when $T=3$ and $Y_0 = 0$, the sequences $y = (1, 0, 1)$ and $\tilde{y} = (0, 1, 1)$ satisfy all conditions in part (ii) of Theorem (ref) to guarantee identification of $\gamma$.
exampleConsider quarterly panel data on binary outcomes $Y_0, Y_1, \ldots, Y_T$, and suppose we specify a logit AR(1) model with a single index of the form $$ Y_{t-1} \, \gamma + \sum_{q=1}^4 \, A_{q} \, w_{t,q} , $$ where $w_{t,q}$ is an indicator for quarter $q \in \{1,2,3,4\}$ and $A_{q}$ are quarter and unit-specific fixed effects. To identify $\gamma$ based on part (ii) of Theorem (ref) we require $y, \tilde y \in \{0,1\}^T$ that satisfy $$ \sum_{t=1}^T w_{qt}(y_{t} - \tilde y_{t}) = 0 , \quad \text{and} \quad \sum_{t=1}^T w_{qt}(y_{t-1} - \tilde y_{t-1}) = 0, \quad \text{for each } q = 1,2,3,4. $$ Finding a solution with $y \neq \tilde y$ requires $T = 6$ quarterly observations (plus the initial quarter $Y_0$), and the solution satisfies $y - \tilde y = \pm (1, 0, 0, 0, -1, 0)$. In addition, we need to satisfy the condition $\sum_{t=1}^{T} y_{t} y_{t-1} \neq \sum_{t=1}^T \tilde y_{t} \tilde y_{t-1}$, e.g.\ for $Y_0=0$, by choosing $y = (0,0,0,0,1,1)$ and $\tilde y = (1,0,0,0,0,1)$. This shows that $\gamma$ is identified for $T=6$.
exampleConsider the case $d_w = 2$ with $w_{t,1} = 1$ and $w_{t,2} = t$, corresponding to a model with heterogeneous linear time trends and single index \[ Y_{t-1} \, \gamma + A_1 + t \, A_2. \] By Remark (ref), this model can be reparameterized — for the purpose of constructing sufficient statistics — as a model with time-specific fixed effects: \[ Y_{t-1} \, \gamma + A_t, \] where $d_w = T$ and $A_t$ denotes a distinct fixed effect for each time period. In this formulation, the fixed effects vary freely over time, and no nontrivial sufficient statistics can be constructed to eliminate $A$. As a result, our sufficiency-based identification strategy cannot be applied, and Theorem (ref) yields a negative result in this case: the autoregressive parameter $\gamma$ is not identified via conditional likelihood. However, as discussed in Section 4.3 of honore2024dynamic, identification is still possible using an alternative approach. Specifically, in the autoregressive panel model with heterogeneous linear time trends (i.e., index $Y_{t-1} \, \gamma + A_1 + t \, A_2$), the parameter $\gamma$ can be identified via a moment condition strategy, as described in Section (ref) below, provided that $T \geq 8$. This example illustrates that for sufficiently general index structures — particularly those involving unrestricted time-varying fixed effects — the sufficient statistics approach may fail, while the moment condition approach can still succeed.

Generalization to AR($p$) models with $p>1$

We now generalize the previous example by considering an AR(2) model with $\pi_t(Y^{t-1}, \theta) = Y_{t-1} \, \gamma_1+ Y_{t-2} \, \gamma_2$ and initial condition $Y^0 = (Y_0,Y_{-1})\in\{0,1\}^2$. Under this setup, the model specified in Assumption (ref) becomes:

align[align omitted — 365 chars of source]

We refer readers to Appendix Section (ref) for a general treatment of the AR($p$) models with $p>1$. Throughout, we maintain the assumption that for all $t \in \{1, \ldots, T\}$, the vector of weights $w_t = (w_{t,1}, \ldots, w_{t,d_w})$ satisfies the restrictions in (ref).

Applying Lemma (ref), we show in Appendix Section (ref) that the likelihood ratio for two outcome paths $y,\tilde y\in \{0,1\}^T$ \footnote{These generalize the sufficient statistics found in Chamberlain84 and magnac2000subsidised.} $$\frac{{\rm Pr}(Y=y | Y_0=y_0, A)}{{\rm Pr}(Y=\tilde y | Y_0=y_0, A)} $$ is invariant to $A$ if the following four conditions hold:

align[align omitted — 385 chars of source]

The first condition coincides with Assumption (i) in Lemma (ref) while the remaining three are equivalent statements of Assumption (ii) for the AR(2). That is, they ensure together that $\Big[ \big(w_t,y_{t-1}\, \gamma_1+y_{t-2}\, \gamma_2\big) \, : \, t=2,\ldots,T\Big] $ is a permutation of $\Big[ \big(w_t,\tilde y_{t-1}\, \gamma_1+\tilde y_{t-2}\, \gamma_2\big) \, : \, t=2,\ldots,T\Big]$.

exampleConsider the standard AR(2) model with $d_w = 1$, $w_{t} = 1$ for all $t$. In this case, the conditions in (ref) can only be satisfied if $T\geq 4$. For instance, with $T=4$ and initial condition $Y^0=(0,1)$, $y=(1,0,0,0)$ and $\tilde y=(0,1,0,0)$ form a valid pair. See also Chamberlain_1985 and honore2000panel.
exampleIn the spirit of Example (ref), consider quarterly panel data on binary outcomes $Y_{-1},Y_0, Y_1, \ldots, Y_T$, and assume an AR(2) structure with a single index $$ Y_{t-1} \, \gamma_1+Y_{t-2} \, \gamma_2 + \sum_{q=1}^4 \, A_{q} \, w_{t,q} , $$ where $w_{t,q}$ are indicators for quarter $q \in \{1,2,3,4\}$. In this setting, the set of restrictions (ref) can only be satisfied for $T\geq 7$. For example, for the initial condition $Y^0=(0,0)$, the outcome sequences $y=(0,0,0,0,1,0,1)$ and $\tilde y = (1,0,0,0,0,0,1)$ form a pair that allows identification of $\gamma_2$.
remarkIn general, the conditional likelihood approach does not allow for the identification of parameters other than the final autoregressive coefficient $\gamma_p$ in the (pure) AR($p$) model. To build intuition, consider again the simple case $p=2$. Then, it is a straightforward exercise (see e.g Appendix Section (ref)) to show that the restrictions in (ref) imply $\sum_{t=1}^T y_{t}y_{t-1}=\sum_{t=1}^T \tilde{y}_{t}\tilde{y}_{t-1}$. In turn, this causes $\gamma_1$ to drop out from the likelihood ratio $\frac{{\rm Pr}(Y=y | Y_0=y_0, A)}{{\rm Pr}(Y=\tilde y | Y_0=y_0, A)}$. This example illustrates the property that the sufficient statistics resulting from Lemma (ref) absorb all variation relevant for identifying $\gamma_1,\ldots,\gamma_{p-1}$. This leaves $\gamma_{p}$ as the only parameter that can be identified via conditional likelihood in the AR($p$) setting. However, as was the case for the model with heterogeneous time trend in Example (ref), identification is possible using an alternative approach that relies on moment conditions.

Dynamic dyadic network formation

We now consider the dyadic network formation model of graham2016homophily, introduced in Example (ref). This model falls within the class of generalized autoregressive logit models studied above, and we show how our identification strategy based on Lemma (ref) applies in this setting.

We model binary outcomes $Y_{ij\tau} = Y_{ji\tau} \in \{0,1\}$ for individuals $i,j \in \{1,\ldots,n\}$ with $i \neq j$, and time periods $\tau \in \{1,\ldots,{\cal T}\}$. The initial condition at $\tau=0$ is denoted by $Y^0$. Each dyad-time pair $(i,j,\tau)$ defines a single observation, and we set $T = \binom{n}{2} \cdot \mathcal{T}$ as the total number of observations. Each observation $t = (i,j,\tau)$ is treated as an element of the outcome vector $Y = (Y_1,\ldots,Y_T)' \in \{0,1\}^T$.

To express the model in the form of Assumption (ref), we impose a deterministic ordering of dyads within each time period. The overall observation index $t$ respects time ordering, i.e.\ if $\tau_1 < \tau_2$ then all observations from period $\tau_1$ precede those from $\tau_2$. For each observation $t = (i,j,\tau)$, we define the regressor $w_t \in \mathbb{R}^{\binom{n}{2}}$ as a unit vector selecting the dyad $(i,j)$, so that $w_t' A = A_{ij}$, where $A \in \mathbb{R}^{\binom{n}{2}}$ collects all dyad-specific fixed effects. The logit index for observation $t$ is then \[ \pi_t(Y^{t-1}, \theta) + w_t' A = Y_{ij,\tau-1} \, \gamma + R_{ij,\tau-1} \, \delta + A_{ij}, \] with parameters $\theta = (\gamma,\delta)'$ and $R_{ij,\tau-1} = \sum_{k \notin \{i,j\}} Y_{ik,\tau-1} Y_{jk,\tau-1}$ denoting the number of shared friends in the previous period.

Throughout, we write $Y_{\cdot \cdot ,\tau}$ for the collection of $Y_{ij,\tau}$ over all dyads, and use ${\cal D}_n$ to denote the set of all ${n \choose 2}$ dyads. Given $Y_{\cdot \cdot ,\tau}$ and $(i,j) \in {\cal D}_n$, the covariates entering the index at time $\tau+1$ are \[ Z_{ij}(Y_{\cdot \cdot,\tau}) := \left( Y_{ij,\tau}, \sum_{k \neq i,j} Y_{ik,\tau} Y_{jk,\tau} \right) \in \{0,1\} \times \mathbb{Z}_+. \] We now describe the sufficient statistic structure implied by Lemma (ref). To simplify the exposition, we focus on the minimal case ${\cal T} = 3$, which is also the specification considered in graham2016homophily. For each outcome vector $y \in \{0,1\}^T$, we define the set of alternative sequences

align*[align* omitted — 425 chars of source]

By Lemma (ref), the conditional likelihood ratio ${\rm Pr}(Y=y \mid Y^0, A) / {\rm Pr}(Y=\tilde y \mid Y^0, A)$ is invariant to $A$ for all $\tilde y \in {\cal Y}_{\rm cond}(y)$. This yields the conditional likelihood \[ \mathcal{L}_{\text{cond}}(\gamma, \delta) = \frac{\prod_{(i,j) \in \mathcal{D}_n} \exp\left( \sum_{\tau=2}^{3} Y_{ij,\tau} \, (\gamma Y_{ij,\tau-1} + \delta R_{ij,\tau-1}) \right) }{ \sum_{\tilde y \in {\cal Y}_{\rm cond}(Y)} \prod_{(i,j) \in \mathcal{D}_n} \exp\left( \sum_{\tau=2}^{3} \tilde{y}_{ij,\tau} \, (\gamma \tilde{y}_{ij,\tau-1} + \delta \tilde R_{ij,\tau-1}) \right) }, \] where $\tilde R_{ij,\tau-1} = \sum_{k \notin \{i,j\}} \tilde{y}_{ik,\tau-1} \cdot \tilde{y}_{jk,\tau-1}$. While this formulation is exact, the denominator is computationally infeasible to evaluate for large $n$, as the set ${\cal Y}_{\rm cond}(y)$ grows exponentially with the number of dyads.

To address this, graham2016homophily proposes a restricted subset of ${\cal Y}_{\rm cond}(y)$ based on the concept of “stable dyads”, which allows feasible computation in large-network settings. An alternative approach, suitable for small $n$ but many independent network observations (as in graham2013comment), is to compute the exact conditional likelihood using the full set ${\cal Y}_{\rm cond}(y)$ or a tractable subset.

A particularly convenient and computationally efficient subset is the two-element set

align*[align* omitted — 339 chars of source]

This set contains at most two elements: the original network $y$ and the version with periods $\tau=1$ and $2$ swapped. When $y_{\cdot\cdot,1} = y_{\cdot\cdot,2}$, the set is a singleton. For $n=3$, one can show that ${\cal Y}^*_{\rm cond}(y) = {\cal Y}_{\rm cond}(y)$ for all networks $y$. For $n=4$ and $n=5$, numerical checks confirm that ${\cal Y}^*_{\rm cond}(y) = {\cal Y}_{\rm cond}(y)$ in more than 95% of all configurations, making this approximation attractive in small-network applications.

In summary, this framework supports two estimation strategies: Firstly, in settings with large $n$ and few time periods, feasible estimation can be achieved using restricted conditioning sets such as those based on stable dyads (graham2016homophily). Secondly, in settings with small $n$ but many repeated network observations, exact conditional likelihood estimation is tractable and efficient using either ${\cal Y}_{\rm cond}(y)$ or its simplification ${\cal Y}^*_{\rm cond}(y)$.

The results extend naturally to longer time horizons ${\cal T} > 3$, though the structure of the conditioning sets becomes more complex. Likewise, the framework accommodates richer dynamic specifications, provided the conditions in Lemma (ref) continue to hold.

Main Takeaways

itemize• The conditional likelihood approach based on sufficient statistics extends naturally to dynamic models. Lemma (ref) characterizes the sufficient statistics for generalized autoregressive logit models with fixed effects. We have shown how to apply this result in both dynamic panel and dynamic network settings. • In contrast to the static case, the sufficient statistics approach does not always identify all the common parameters in dynamic models. We illustrated this in two examples (i) in the AR(1) model with heterogeneous time trends, the sufficient statistics approach fails to identify the autoregressive parameter $\gamma$; (ii) in AR($p$) models with $p > 1$, the sufficient statistics approach can identify only the final lag coefficient $\gamma_p$. In the next section, we discuss an alternative to the sufficient conditions approach, which is based on moment conditions. This approach will lead to identification of all the common parameters in the two examples.

Moment conditions for dynamic logit models

We now turn to an alternative identification strategy for non-linear panel models based on moment conditions that do not depend on the fixed effects. Unlike the conditional likelihood approach used in Sections (ref) and (ref), which relies on sufficient statistics, the moment-based approach developed here applies to a larger class of models and accommodates richer forms of unobserved heterogeneity.

While fixed-effect-free moment conditions had been constructed in specific models before, the general idea was formalized in the context of semiparametric panel models by bonhomme2012functional, who called it “functional differencing”. Its applicability to discrete choice models was recently emphasized by kitazawa2021transformations and honore2024dynamic.

Model setup

The model and notation are essentially unchanged compared to Section (ref), the only difference is that we now also allow for additional strictly exogenous covariates $X$, as we had in Section (ref). Assumption (ref) then generalizes as follows.

assumptionThe data-generating process is \begin{align*} {\rm Pr} \left( Y=y \, \big| \, Y^0,X,A \right) &= \prod_{t=1}^T {\rm Pr} \left( Y_t=y_t \, \big| \, Y^{t-1},X,A \right) , \\ {\rm Pr} \left( Y_t=y_t \, \big| \, Y^{t-1},X,A \right) &= \frac {\left[\exp(\pi_t(Y^{t-1},X_t,\theta) + w_t' \, A)\right]^{y_t}} {1+\exp(\pi_t(Y^{t-1},X_t,\theta) + w_t' \, A) } , \end{align*} where $\theta$ is an unknown parameter and $\pi_t(\cdot,\cdot,\cdot)$ is a known function for every $t \in \{1,\ldots,T\}$.

Our goal is again to identify the parameter $\theta$ without imposing restrictions on the distribution of the fixed effect $A$. The model in Assumption (ref) nests many dynamic panel and network models. Our baseline Examples (ref) and (ref) remain essentially unchanged, except for the additional covariates $X_t$. Thus, in Example (ref) the specification becomes $$ \pi_t(Y^{t-1}, X_t, \theta) = Y_{t-1} \, \gamma + X_t' \, \beta, $$ where $\theta=(\gamma,\beta)'$ while generalizing Example (ref) gives $$ \pi_t(Y^{t-1}, X_t, \theta) = Y_{ij,\tau-1} \, \gamma + R_{ij,\tau-1} \, \delta + X_t' \, \beta $$ with $\theta = (\gamma, \delta)'$.

It is sometimes possible to extend the identification strategy based on sufficient statistics to models with additional covariates. For example, honore2000panel apply this approach to an AR(1) panel logit model with covariates by conditioning on subsets of the data where covariate values are identical across adjacent time periods. More generally, the results in Section (ref) can be applied to such models by treating $X$ as fixed throughout and absorbing the covariate dependence of $\pi_t(Y^{t-1}, X_t, \theta)$ into the definition of $\pi_t$. However, the sufficiency-based approach typically works only under restrictive support conditions on the covariates. In contrast, the moment condition strategy developed in this section applies more broadly and often accommodates arbitrary variation in $X$.

Identification via fixed-effect-free moment conditions

As discussed above, the sufficient statistics approach typically breaks down in dynamic settings with covariates that vary freely over time. This motivates an alternative strategy based on moment functions that depend on the parameters of interest but not on the fixed effects. Specifically, we aim to find functions $m(Y, Y^0, X, \theta)$ such that

align[align omitted — 118 chars of source]

Moment functions satisfying (ref) are valid for identification and estimation, as they are invariant to the fixed effect $A$ and therefore unaffected by the incidental parameter problem. To characterize when such functions exist, we use a combinatorial argument that provides a lower bound on the dimension of the space of fixed-effect-free moment functions,\footnote{See also dano2023transition for an alternative strategy tailored to standard AR($p$) logit models with $p \geq 1$.} leading to the following theorem.

To state the theorem, we introduce some additional notation. For each time period $t \in \{1, \ldots, T\}$, define the set of distinct values that the index $\pi_t(Y^{t-1}, X_t, \theta)$ can take across all possible outcome histories $(y_{t-1},y_{t-2},\ldots,y_1) \in \{0,1\}^{t-1}$ as $$ \Pi_t( Y^0,X,\theta) := \big\{ \pi_t(( y_{t-1},y_{t-2},\ldots,y_1,Y^{0}),X_t,\theta) \, : \, (y_{t-1},y_{t-2},\ldots,y_1) \in \{0,1\}^{t-1} \big\} , $$ Throughout this section, we treat $Y^0$, $X$, and $\theta$ as fixed. While we make these arguments explicit in the notation for mathematical precision, they are otherwise unimportant to the combinatorial structure we analyze and can be regarded as fixed and given (and thus ignored) for the remainder of the discussion. Let $$ Q_t(Y^0, X,\theta) := |\Pi_t(Y^0, X,\theta)| $$ denote the number of elements in the set $\Pi_t(Y^0, X, \theta)$, that is, the number of distinct values that the index $\pi_t((y_{t-1}, y_{t-2}, \ldots, y_1, Y^0), X_t, \theta)$ can take across different outcome histories. By construction, we have $Q_t(Y^0, X, \theta) \leq 2^{t-1}$, but in many practical models we have limited dependence on past outcomes and $Q_t(Y^0, X, \theta)$ is therefore often (much) smaller. Also, in most models, $Q_t(Y^0, X, \theta)$ remains constant across typical values of $Y^0$, $X$, and $\theta$, but may be different at specific points (e.g., when $\theta = 0$).

Next, define the set of all possible linear combinations of the $w_t$'s, where each coefficient $k_t$ is an integer between 0 and $Q_t(Y^0, X, \theta)$:

align*[align* omitted — 209 chars of source]

The remark and examples below will help clarify the intuition behind the definition of the set $\mathcal{D}(Y^0, X, \theta)$. As before, we write $|\mathcal{D}(Y^0, X, \theta)|$ to denote its cardinality. The following result provides a lower bound on the number of linearly independent fixed-effect-free moment conditions that exist in the class of dynamic models considered in this section.

theoremLet the model be given by Assumption (ref), and let $Y^0$, $X$, $\theta$ be given values of the initial conditions, covariates, and common parameters. Then, the number of linearly independent moment functions of the form (ref) is at least equal to $$ 2^T - \left|{\cal D}(Y^0,X,\theta) \right| . $$ In particular, if $\left|{\cal D}(Y^0,X,\theta) \right| < 2^T$, then non-zero moment functions of the form (ref) exist.

Theorem (ref) provides a sufficient condition for the existence of non-zero moment functions $m(Y, Y^0, X, \theta)$ satisfying (ref). However, the result does not characterize the structure of these functions, nor does it guarantee that they depend on $\theta$ in a way that allows identification. In practice, constructing such moment functions, and verifying that they identify $\theta$, is typically model-specific and requires additional algebraic or numerical insight. Concrete examples are discussed below.

Even so, the existence result in Theorem (ref) is useful. For example, once existence is guaranteed, one can apply the general functional-analytic framework developed by bonhomme2012functional to compute the moment conditions numerically. In this “functional differencing” approach, moment functions can be approximated even when closed-form expressions are unavailable.

Whether obtained analytically or numerically, valid moment functions $m(Y, Y^0, X, \theta)$ can be used to construct GMM estimators that are root-$n$ consistent under standard regularity conditions.

remarkThe static logit model provides a useful special case for understanding the structure and role of the set $\mathcal{D}(Y^0, X, \theta)$. In that model, we have $Q_t(Y^0, X, \theta) = 1$ for all $t$, so the set reduces to $$ \mathcal{D} = \left\{ \sum_{t=1}^T y_t w_t : y \in \{0,1\}^T \right\}, $$ which is simply the set of all possible values that the sufficient statistic $\mathsf{W} Y$ can take. As discussed in Section (ref), identification in the static model hinges on whether the mapping $Y \mapsto \mathsf{W} Y$ is injective. If $\mathsf{W} Y$ takes a distinct value for every realization $y \in \{0,1\}^T$, i.e., if $|\mathcal{D}| = 2^T$, then the existence of the sufficient statistic is not useful for identification or estimation of $\theta$, and no fixed-effect-free moment conditions exist. This explains the condition $|\mathcal{D}| < 2^T$ in Theorem (ref), which guarantees the existence of nontrivial moment functions precisely when the sufficient statistic does not uniquely index every outcome configuration --- that is, when different realizations of $Y$ map to the same value of $\mathsf{W} Y$. More generally, when sufficient statistics are not available, as in most dynamic models with general covariate values $X$, the set $\mathcal{D}(Y^0, X, \theta)$ still plays a closely analogous role. Although the linear combinations $\sum_{t=1}^T k_t w_t$ that define $\mathcal{D}$ can no longer be interpreted as sufficient statistics (since the coefficients $k_t$ are not restricted to binary values), they retain a similar structure and capture aspects of how the model links outcome histories to the fixed effects. The condition $|\mathcal{D}| < 2^T$ remains unchanged, but now serves only as a sufficient condition for the existence of nontrivial moment functions. In this sense, Theorem (ref) extends the logic of sufficiency-based identification to dynamic settings where no sufficient statistics exist.

Examples

Panel AR($p$) models with general covariates

We first consider the class of AR(\(p\)) panel logit models with covariates and scalar fixed effects. These models are covered by Assumption (ref), with the index and fixed effect specification given by

align*[align* omitted — 106 chars of source]

where $\theta = (\gamma_1, \ldots, \gamma_p,\beta')'$. The initial conditions $Y^0 = (Y_{-p+1}, \ldots, Y_0) \in \{0,1\}^p$ are treated as fixed and observed.

To apply Theorem (ref), we compute the number of distinct values that the index $\pi_t(Y^{t-1}, X_t, \theta)$ can take across binary outcome histories. Assuming $\gamma_r \neq 0$ for all $r \in \{1,\ldots,p\}$, this number is $$ Q_t = 2^{\min(p,\,t-1)}, $$ since the index depends on the most recent binary $p$ outcomes, or fewer in the initial periods. Because we are considering $w_t = 1$ for all $t$, the set $\mathcal{D} = \mathcal{D}(Y^0, X, \theta)$ consists of all integers between 0 and the maximum sum $d_{\max} := \sum_{t=1}^T Q_t$. Therefore, we have

align*[align* omitted — 204 chars of source]

By Theorem (ref), the number of linearly independent fixed-effect-free moment functions is therefore at least $$ 2^T - 2^p (T + 1 - p), $$ which is strictly positive whenever $T \geq 2+p$ (in addition to $p$ observed time periods as initial conditions). It turns out that this is not merely a lower bound: for generic values of the covariates $X$, this is in fact the exact number of linearly independent moment conditions in the AR(\(p\)) panel model. However, as we will see in the following example, additional moment conditions may become available for specific values of $X$, in particular when $X = 0$.

The result for the number of linearly independent moment conditions in the AR($p$) model was stated in honore2024dynamic, who also derive explicit analytical expressions for the corresponding moment functions in the cases $p = 1, 2, 3$, and provide a detailed analysis of identification and estimation for $p = 1$, building on the results of kitazawa2021transformations. The sharpness of the moment count has been established for $p = 1$ and $p = 2$ by kruiniger2020further, and further discussed in dobronyi2021identification. For general $p$, dano2023transition proves that the bound is sharp under free-varying covariates and provides analytical expressions for all the moment functions in closed form.

Panel AR(2) model without covariates

We now consider a special case of the previous example where $p = 2$ and no covariates are present. The model takes the form $$ \Pr(Y_t = 1 \mid Y_{t-1}, Y_{t-2}, A) = \frac{ \exp\left( Y_{t-1} \gamma_1 + Y_{t-2} \gamma_2 + A \right) } {1 + \exp\left( Y_{t-1} \gamma_1 + Y_{t-2} \gamma_2 + A \right)}, $$ which corresponds to the structure in Assumption (ref) with

align*[align* omitted — 93 chars of source]

In this model for $T = 3$, one can construct one valid fixed-effect-free moment function for each value of the inital condition $Y^0 = (Y_{-1}, Y_0) \in \{0,1\}^2$. For instance, when $y^0 = (0,0)$, the following function satisfies the moment condition (ref): $$ m(y, y^0, \theta) =

cases\exp(-\gamma_1) & if y = (0,1,1), \\ 1 & if y = (0,1,0), \\ -1 & if (y_1, y_2) = (1,0), \\ 0 & otherwise,

$$ and when $y^0 = (0,1)$, a valid moment function is given by $$ m(y, y^0, \theta) =

cases\exp(\gamma_2 - \gamma_1) & if y = (1,0,0), \\ \exp(\gamma_2) & if y = (1,0,1), \\ -1 & if (y_1, y_2) = (0,1), \\ 0 & otherwise,

$$ see Section~3.3 of the first arXiv version of \citet{honore2024dynamic}, who provide closed-form moment functions for the AR($p$) model with $T = 3$ and $X_2=X_3$.

Firstly, this example shows that the lower bound in Theorem (ref) is not always sharp. For $p = 2$ and $T = 3$, we have $2^T - |\mathcal{D}| = 0$, so the theorem does not guarantee the existence of any fixed-effect-free moment functions. Nevertheless, such functions do exist, as demonstrated above.

Secondly, the first moment function implies that $\gamma_1$ is point-identified in this model, since $\exp(-\gamma_1)$ is strictly monotonic in $\gamma_1$, and thus the moment function is strictly monotonic as well. Once $\gamma_1$ is identified, the second moment function then identifies $\gamma_2$ through the same logic. This result provides analytical confirmation of earlier numerical findings in honore2019identification that suggested identification might be possible in this setup, even though conditional likelihood methods do not apply.

Thus, the moment condition approach point-identifies both $\gamma_1$ and $\gamma_2$ in this model as soon as $T \geq 3$. By contrast, as noted in Remark (ref), the sufficient statistics approach can identify only $\gamma_2$, and requires at least $T \geq 4$ to do so.

Panel AR(1) with quarterly fixed effects and covariates

Next, consider the extension of Example (ref) with strictly exogenous covariates. For this specification, we have $\pi_t(Y^{t-1}, X_t, \theta) = Y_{t-1} \gamma + X_t' \beta$ and $w_t=(w_{t,1},w_{t,2},w_{t,3},w_{t,4})$ where $w_{t,q}$ is an indicator for quarter $q\in \{1,2,3,4\}$. It follows that

align*[align* omitted — 376 chars of source]

with cardinality $|\mathcal{D}(Y^0, X,\theta)| = \left(2 \lfloor \frac{T - 1}{4}\rfloor+2\right)\left(2 \lfloor \frac{T - 2}{4}\rfloor+3\right)\left(2 \lfloor \frac{T - 3}{4}\rfloor+3\right)\left(2 \lfloor \frac{T}{4}\rfloor+1\right)$. Hence, the bound $2^T-|\mathcal{D}(Y^0, X,\theta)|$ of Theorem (ref) becomes positive as soon as $T\geq 12$, ensuring the existence of moments. While informative, this bound is conservative in this instance since in the absence of regressors (which does not alter the bound here), Example (ref) already indicated the existence of identifying moments for $T=6$. Indeed, one can construct two linearly independent moment functions $m_{1}$ and $m_{2}$ with $T=6$ periods. Using the shorthand $x_{ts}=x_{t}-x_{s}$ for $t\neq s$, $m_1$ is given by

align*[align* omitted — 1,076 chars of source]

and $m_{2}(y, y_0,x,\theta)$ is obtained by substituting $y_{t}$ by $(1-y_{t})$ for $t=0,\ldots, 6$ and $x$ by $-x$ in the expression of $m_{1}(y, y_0,x,\theta)$. In other words, $m_{2}(y, y_0,x,\theta)= m_{1}(1-y, 1-y_0,-x,\theta)$. We refer readers to Appendix (ref) for detailed derivations of these expressions. Remark that the two moment functions depend on the common parameter $\theta=(\gamma,\beta')'$.

Moment conditions in graham2016homophily with additional covariates

As an illustration of how the lower bound of Theorem (ref) can be applied to probe the existence of moment conditions in networks, consider the extension of graham2016homophily introduced earlier, incorporating strictly exogenous covariates $X$. In this setting, we specify $\pi_t(Y^{t-1}, X_t, \theta) = Y_{ij,\tau-1} \gamma + R_{ij,\tau-1} \delta + X_t' \beta$, where $t=(i,j,\tau)$ indexes dyad-time pairs, and $w_t$ denotes the basis vector in $\mathbb{R}^{\binom{n}{2}}$ with entry one for dyad $(i,j)$, and zeros elsewhere. Under this formulation, the model described in Assumption (ref) becomes

align*[align* omitted — 266 chars of source]

Since $R_{ij,\tau-1} := \sum_{k \neq i,j} Y_{ik,\tau-1} Y_{jk,\tau-1}$, for any fixed covariates $X$, the structure of $\pi_t$ implies that it can take at most $2(n-1)$ distinct values. Consequently, we have

align*[align* omitted — 149 chars of source]

and hence

align*[align* omitted — 251 chars of source]

A straightforward counting argument then gives $|\mathcal{D}(Y^0, X,\theta)| = \left[2(\mathcal{T}-1)(n-1)+2\right]^{\binom{n}{2}}$ and Theorem (ref) implies the existence of at least $2^T - \left|{\cal D}(Y_0,X) \right|=2^T-\left[2(\mathcal{T}-1)(n-1)+2\right]^{\binom{n}{2}}$ moment functions free from the fixed effects. Since $\log |\mathcal{D}(Y^0, X,\theta)|$ grows approximately proportionally to $\log \mathcal{T}$ - the number of time periods - while $\log(2^T)=\log(2)\binom{n}{2}\mathcal{T}$ grows linearly in $\mathcal{T}$, the existence of moment conditions is guaranteed for sufficiently large $\mathcal{T}$. For example, in the simplest case with $n=3$ agents, the lower bound is positive as soon as $T\geq 4$. In fact, dano2023transition shows that there is generally as much as $\binom{n}{2}$ fixed-effect-free moment conditions with only $T=3$ periods that are explicitly given by:

align*[align* omitted — 387 chars of source]

where $y$ denotes an undirected network, and $r_{ij}:= \sum_{k \neq i,j} y_{ik} y_{jk}$ denotes the number of friends that agents $i$ and $j$ share in common in $y$.

Main Takeaways

itemize• The moment function approach developed in this section --- also known as functional differencing, following bonhomme2012functional --- provides a powerful alternative to the conditional likelihood strategy based on sufficient statistics. In particular, the examples above show that fixed-effect-free moment conditions exist even in dynamic models with arbitrary covariate variation, where sufficient statistics approaches typically fail. This includes models with general time-varying covariates, heterogeneous time trends, and rich dynamic structures. • Theorem (ref) offers a general and easy-to-verify sufficient condition for the existence of such fixed-effect-free moment functions. While the lower bound it provides is not always sharp, it guarantees the existence of moment conditions for sufficiently large $T$ in a broad class of models. For example, in the AR(2) model without covariates, the theorem does not predict the existence of moments at $T=3$, but still confirms their existence for $T\geq 4$. • Even when moment conditions are known to exist, their explicit construction and the demonstration that they identify the parameters of interest typically require model-specific derivations. In some models, analytical expressions can be derived, while in others, numerical methods, as in bonhomme2012functional, may be required. • Unlike the sufficiency-based approach, which relies on conditional likelihood and is typically estimated via conditional MLE, the moment-based approach leads naturally to GMM estimation. Once a sufficient number of valid moment conditions are constructed, and the model parameters are identified, standard GMM theory yields root-$n$ consistent and asymptotically normal estimators under appropriate regularity conditions.

Conclusions

This paper has reviewed and extended two approaches for eliminating fixed effects in logit models: the conditional likelihood method and the construction of moment conditions. While the results are conceptually clean and methodologically promising, several important challenges remain, pointing to avenues for future research.

First, the moment-based framework often yields a large number of conditional moment conditions, linking it to the broader literature on optimal moment selection in panel data. Prior work, including bekker1994alternative, donald2001choosing, alvarez2003time, and okui2009optimal, has shown that using too many valid moments can introduce substantial finite-sample bias. This suggests that future research on nonlinear models should integrate identification strategies with principled approaches to moment selection.

Second, our focus has been on identifying and estimating structural parameters, rather than computing counterfactuals or marginal effects. In panel data settings, these quantities are typically not point-identified, even when the structural parameters are. Extending the methods discussed here to incorporate bounds on marginal effects, such as those proposed by PakelWeidner2024 and davezies2024identification, would enhance the empirical relevance of the fixed effects framework.

Finally, and perhaps most critically, the models considered here assume strict exogeneity of the explanatory variables (aside from lagged outcomes). This assumption is often unrealistic in economic applications. While recent work by ArellanoCarrasco2003, botosaru2024adversarial, bonhomme2023identification, and BonhommeDanoGraham2025 has made progress in relaxing this assumption, much remains to be done to develop robust methods that accommodate predetermined regressors in nonlinear panel models.

thebibliography\bibitem[\citeauthoryear{Alvarez and Arellano}{Alvarez and Arellano}{2003}]{alvarez2003time} Alvarez, J. and M. Arellano (2003). \newblock The time series and cross-section asymptotics of dynamic panel data estimators. \newblock {\em Econometrica\/} {\em 71\/}(4), 1121--1159. \bibitem[\citeauthoryear{Andersen}{Andersen}{1970}]{andersen1970asymptotic} Andersen, E. B. (1970). \newblock Asymptotic properties of conditional maximum-likelihood estimators. \newblock {\em Journal of the Royal Statistical Society: Series B (Methodological)\/} {\em 32\/}(2), 283--301. \bibitem[\citeauthoryear{Arellano and Carrasco}{Arellano and Carrasco}{2003}]{ArellanoCarrasco2003} Arellano, M. and R. Carrasco (2003). \newblock Binary choice panel data models with predetermined variables. \newblock {\em Journal of Econometrics\/} {\em 115\/}(1), 125--157. \bibitem[\citeauthoryear{Arellano and Hahn}{Arellano and Hahn}{2007}]{ArellanoHahn2007} Arellano, M. and J. Hahn (2007). \newblock Understanding bias in nonlinear panel models: Some recent developments. \newblock {\em Econometric Society Monographs\/} {\em 43}, 381. \bibitem[\citeauthoryear{Bekker}{Bekker}{1994}]{bekker1994alternative} Bekker, P. A. (1994). \newblock Alternative approximations to the distributions of instrumental variable estimators. \newblock {\em Econometrica: Journal of the Econometric Society\/}, 657--681. \bibitem[\citeauthoryear{Bellemare and Millimet}{Bellemare and Millimet}{2025}]{BellemareMillimet2025} Bellemare, M. F. and D. L. Millimet (2025). \newblock Retrospectives: Yair mundlak and the fixed effects estimator. \newblock {\em Journal of Economic Perspectives\/} {\em 39\/}(2), 261–74. \bibitem[\citeauthoryear{Bonhomme}{Bonhomme}{2012}]{bonhomme2012functional} Bonhomme, S. (2012). \newblock Functional differencing. \newblock {\em Econometrica\/} {\em 80\/}(4), 1337--1385. \bibitem[\citeauthoryear{Bonhomme, Dano, and Graham}{Bonhomme, Dano and Graham}{2023}]{bonhomme2023identification} Bonhomme, S., K. Dano, and B. S. Graham (2023). \newblock Identification in a binary choice panel data model with a predetermined covariate. \newblock {\em SERIEs\/} {\em 14\/}(3), 315--351. \bibitem[\citeauthoryear{Bonhomme, Dano, and Graham}{Bonhomme, Dano and Graham}{2025}]{BonhommeDanoGraham2025} Bonhomme, S., K. Dano, and B. S. Graham (2025). \newblock Moment restrictions for nonlinear panel data models with feedback. \bibitem[\citeauthoryear{Botosaru, Loh, and Muris}{Botosaru, Loh and Muris}{2024}]{botosaru2024adversarial} Botosaru, I., I. Loh, and C. Muris (2024). \newblock An adversarial approach to identification and inference. \newblock Technical report. \bibitem[\citeauthoryear{Chamberlain}{Chamberlain}{1978}]{Chamberlain78} Chamberlain, G. (1978). \newblock On the use of panel data. \newblock Department of Economics, Harvard University. \bibitem[\citeauthoryear{Chamberlain}{Chamberlain}{1980}]{chamberlain_analysis_1980} Chamberlain, G. (1980, January). \newblock Analysis of {Covariance} with {Qualitative} {Data}. \newblock {\em The Review of Economic Studies\/} {\em 47\/}(1), 225--238. \bibitem[\citeauthoryear{Chamberlain}{Chamberlain}{1984}]{Chamberlain84} Chamberlain, G. (1984). \newblock Panel data, in griliches, z. and m.d. intriligator (eds.). \newblock {\em 2}. \newblock Elsevier Science, Amsterdam. \bibitem[\citeauthoryear{Chamberlain}{Chamberlain}{1985}]{Chamberlain_1985} Chamberlain, G. (1985). \newblock {\em Heterogeneity, omitted variable bias, and duration dependence}, pp.\ 3–38. \newblock Econometric Society Monographs. Cambridge University Press. \bibitem[\citeauthoryear{Chamberlain}{Chamberlain}{1992}]{Chamberlain_1992} Chamberlain, G. (1992). \newblock Comment: Sequential moment restrictions in panel data. \newblock {\em Journal of Business & Economic Statistics\/} {\em 10\/}(1), 20--26. \bibitem[\citeauthoryear{Chamberlain}{Chamberlain}{2010}]{chamberlain2010binary} Chamberlain, G. (2010). \newblock Binary response models for panel data: Identification and information. \newblock {\em Econometrica\/} {\em 78\/}(1), 159--168. \bibitem[\citeauthoryear{Charbonneau}{Charbonneau}{2017}]{charbonneau2017multiple} Charbonneau, K. B. (2017). \newblock Multiple fixed effects in binary response panel data models. \newblock {\em The Econometrics Journal\/} {\em 20\/}(3), S1--S13. \bibitem[\citeauthoryear{Cox}{Cox}{1958}]{cox1958regression} Cox, D. R. (1958). \newblock The regression analysis of binary sequences. \newblock {\em Journal of the Royal Statistical Society: Series B (Methodological)\/} {\em 20\/}(2), 215--232. \bibitem[\citeauthoryear{D'Addio and Honor{\'e}}{D'Addio and Honor{\'e}}{2010}]{d2010duration} D'Addio, A. C. and B. E. Honor{\'e} (2010). \newblock Duration dependence and timevarying variables in discrete time duration models. \newblock {\em Brazilian Review of Econometrics\/} {\em 30\/}(2), 487--527. \bibitem[\citeauthoryear{Dano}{Dano}{2023}]{dano2023transition} Dano, K. (2023). \newblock Transition probabilities and moment restrictions in dynamic fixed effects logit models. \newblock {\em arXiv preprint arXiv:2303.00083\/}. \bibitem[\citeauthoryear{Davezies, D'Haultfœuille, and Laage}{Davezies, D'Haultfœuille and Laage}{2024}]{davezies2024identification} Davezies, L., X. D'Haultfœuille, and L. Laage (2024). \newblock Identification and estimation of average causal effects in fixed effects logit models. \bibitem[\citeauthoryear{Dobronyi, Gu, and Kim}{Dobronyi, Gu and Kim}{2021}]{dobronyi2021identification} Dobronyi, C., J. Gu, and K. i. Kim (2021). \newblock Identification of dynamic panel logit models with fixed effects. \newblock {\em arXiv preprint arXiv:2104.04590\/}. \bibitem[\citeauthoryear{Donald and Newey}{Donald and Newey}{2001}]{donald2001choosing} Donald, S. G. and W. K. Newey (2001). \newblock Choosing the number of instruments. \newblock {\em Econometrica\/} {\em 69\/}(5), 1161--1191. \bibitem[\citeauthoryear{Fernández-Val and Weidner}{Fernández-Val and Weidner}{2018}]{FernandezValWeidner2018} Fernández-Val, I. and M. Weidner (2018). \newblock Fixed effects estimation of large-t panel data models. \newblock {\em Annual Review of Economics\/} {\em 10\/}(Volume 10, 2018), 109--138. \bibitem[\citeauthoryear{Graham}{Graham}{2013}]{graham2013comment} Graham, B. S. (2013). \newblock Comment on “social networks and the identification of peer effects” by paul goldsmith-pinkham and guido w. imbens. \newblock {\em Journal of Business and Economic Statistics\/} {\em 31\/}(3), 266--270. \bibitem[\citeauthoryear{Graham}{Graham}{2016}]{graham2016homophily} Graham, B. S. (2016). \newblock Homophily and transitivity in dynamic network formation. \newblock Technical report, National Bureau of Economic Research. \bibitem[\citeauthoryear{Graham}{Graham}{2017}]{graham2017econometric} Graham, B. S. (2017). \newblock An econometric model of network formation with degree heterogeneity. \newblock {\em Econometrica\/} {\em 85\/}(4), 1033--1063. \bibitem[\citeauthoryear{Hahn and Newey}{Hahn and Newey}{2004}]{HahnNewey2004} Hahn, J. and W. Newey (2004). \newblock Jackknife and analytical bias reduction for nonlinear panel models. \newblock {\em Econometrica\/} {\em 72\/}(4), 1295--1319. \bibitem[\citeauthoryear{Hausman, Hall, and Griliches}{Hausman, Hall and Griliches}{1984}]{HausmanHallGriliches84} Hausman, J., B. H. Hall, and Z. Griliches (1984, July). \newblock Econometric models for count data with an application to the patents-{R&D} relationship. \newblock {\em Econometrica\/} {\em 52\/}(4), 909--38. \bibitem[\citeauthoryear{Heckman}{Heckman}{1981}]{Heckman81b} Heckman, J. J. (1981). \newblock The incidental parameters problem and the problem of initial conditions in estimating a discrete time-discrete data stochastic process. \newblock {\em Structural Analysis of Discrete Panel Data with Econometric Applications, C. F. Manski and D. McFadden (eds)\/}, 179--195. \bibitem[\citeauthoryear{Higgins and Jochmans}{Higgins and Jochmans}{2024}]{HigginsJochmans2024} Higgins, A. and K. Jochmans (2024). \newblock Bootstrap inference for fixed-effect models. \newblock {\em Econometrica\/} {\em 92\/}(2), 411--427. \bibitem[\citeauthoryear{Honor{\'e}}{Honor{\'e}}{1992}]{Honore92} Honor{\'e}, B. E. (1992). \newblock Trimmed lad and least squares estimation of truncated and censored regression models with fixed effects. \newblock {\em Econometrica\/} {\em 60}, 533--565. \bibitem[\citeauthoryear{Honor{\'e}}{Honor{\'e}}{1993}]{HONORE199335} Honor{\'e}, B. E. (1993). \newblock Orthogonality conditions for tobit models with fixed effects and lagged dependent variables. \newblock {\em Journal of Econometrics\/} {\em 59\/}(1), 35--61. \bibitem[\citeauthoryear{Honor{\'e} and Kyriazidou}{Honor{\'e} and Kyriazidou}{2000}]{honore2000panel} Honor{\'e}, B. E. and E. Kyriazidou (2000). \newblock Panel data discrete choice models with lagged dependent variables. \newblock {\em Econometrica\/} {\em 68\/}(4), 839--874. \bibitem[\citeauthoryear{Honor{\'e} and Kyriazidou}{Honor{\'e} and Kyriazidou}{2019}]{honore2019identification} Honor{\'e}, B. E. and E. Kyriazidou (2019). \newblock Identification in binary response panel data models: Is point-identification more common than we thought? \newblock {\em Annals of Economics and Statistics\/} (134), 207--226. \bibitem[\citeauthoryear{Honoré and Weidner}{Honoré and Weidner}{2024}]{honore2024dynamic} Honoré, B. E. and M. Weidner (2024). \newblock Moment conditions for dynamic panel logit models with fixed effects. \newblock {\em The Review of Economic Studies\/}, rdae097. \bibitem[\citeauthoryear{Hu}{Hu}{2002}]{Hu2002} Hu, L. (2002). \newblock Estimation of a censored dynamic panel data model. \newblock {\em Econometrica\/} {\em 70\/}(6), pp. 2499--2517. \bibitem[\citeauthoryear{Jochmans}{Jochmans}{2017}]{jochmans2017two} Jochmans, K. (2017). \newblock Two-way models for gravity. \newblock {\em Review of Economics and Statistics\/} {\em 99\/}(3), 478--485. \bibitem[\citeauthoryear{Khan, Ponomareva, and Tamer}{Khan, Ponomareva and Tamer}{2023}]{KHAN2023105515} Khan, S., M. Ponomareva, and E. Tamer (2023). \newblock Identification of dynamic binary response models. \newblock {\em Journal of Econometrics\/} {\em 237\/}(1), 105515. \bibitem[\citeauthoryear{Kitazawa}{Kitazawa}{2022}]{kitazawa2021transformations} Kitazawa, Y. (2022). \newblock Transformations and moment conditions for dynamic fixed effects logit models. \newblock {\em Journal of Econometrics\/} {\em 229\/}(2), 350--362. \bibitem[\citeauthoryear{Kruiniger}{Kruiniger}{2020}]{kruiniger2020further} Kruiniger, H. (2020). \newblock Further results on the estimation of dynamic panel logit models with fixed effects. \newblock {\em arXiv preprint arXiv:2010.03382\/}. \bibitem[\citeauthoryear{Kyriazidou}{Kyriazidou}{1997}]{Kyriazidou97} Kyriazidou, E. (1997). \newblock Estimation of a panel data sample selection model. \newblock {\em Econometrica\/} {\em 65}, 1335--1364. \bibitem[\citeauthoryear{Kyriazidou}{Kyriazidou}{2001}]{Kyriazidou01} Kyriazidou, E. (2001). \newblock Estimation of dynamic panel data sample selection models. \newblock {\em Review of Economic Studies\/} {\em 68}, 543--572. \bibitem[\citeauthoryear{Magnac}{Magnac}{2000}]{magnac2000subsidised} Magnac, T. (2000). \newblock Subsidised training and youth employment: distinguishing unobserved heterogeneity from state dependence in labour market histories. \newblock {\em The Economic Journal\/} {\em 110\/}(466), 805--837. \bibitem[\citeauthoryear{Magnac}{Magnac}{2004}]{magnac2004panel} Magnac, T. (2004). \newblock Panel binary variables and sufficiency: generalizing conditional logit. \newblock {\em Econometrica\/} {\em 72\/}(6), 1859--1876. \bibitem[\citeauthoryear{Manski}{Manski}{1987}]{Manski87} Manski, C. (1987). \newblock Semiparametric analysis of random effects linear models from binary panel data. \newblock {\em Econometrica\/} {\em 55}, 357--62. \bibitem[\citeauthoryear{Mundlak}{Mundlak}{1961}]{Mundlak1961} Mundlak, Y. (1961). \newblock Empirical production function free of management bias. \newblock {\em Journal of Farm Economics\/} {\em 43\/}(1), 44--56. \bibitem[\citeauthoryear{Mundlak}{Mundlak}{1978}]{Mundlak1978} Mundlak, Y. (1978). \newblock On the pooling of time series and cross section data. \newblock {\em Econometrica\/} {\em 46\/}(1), 69--85. \bibitem[\citeauthoryear{Muris and Pakel}{Muris and Pakel}{2025}]{MurisPakel2025} Muris, C. and C. Pakel (2025). \newblock Triadic network formation. \newblock Working Paper. \bibitem[\citeauthoryear{Neyman and Scott}{Neyman and Scott}{1948}]{neyman1948consistent} Neyman, J. and E. L. Scott (1948). \newblock Consistent estimates based on partially consistent observations. \newblock {\em Econometrica\/} {\em 16}, 1--32. \bibitem[\citeauthoryear{Okui}{Okui}{2009}]{okui2009optimal} Okui, R. (2009). \newblock The optimal choice of moments in dynamic panel data models. \newblock {\em Journal of Econometrics\/} {\em 151\/}(1), 1--16. \bibitem[\citeauthoryear{Pakel and Weidner}{Pakel and Weidner}{2024}]{PakelWeidner2024} Pakel, C. and M. Weidner (2024). \newblock Bounds on average effects in discrete choice panel data models. \bibitem[\citeauthoryear{Pakes, Porter, Shepard, and Calder-Wang}{Pakes, Porter, Shepard and Calder-Wang}{2022}]{PakesPorterShepardCalderWang2022} Pakes, A., J. Porter, M. Shepard, and S. Calder-Wang (2022). \newblock Unobserved heterogeneity, state dependence, and health plan choices. \newblock {\em Working paper. Revised Sept. 2022\/}. \bibitem[\citeauthoryear{Rasch}{Rasch}{1960}]{rasch1960studies} Rasch, G. (1960). \newblock {\em Studies in mathematical psychology: I. Probabilistic models for some intelligence and attainment tests.} \newblock Nielsen & Lydiche. \bibitem[\citeauthoryear{Wooldridge}{Wooldridge}{1997}]{Wooldridge_1997} Wooldridge, J. M. (1997). \newblock Multiplicative panel data models without the strict exogeneity assumption. \newblock {\em Econometric Theory\/} {\em 13\/}(5), 667–678. \bibitem[\citeauthoryear{Wooldridge}{Wooldridge}{2005}]{wooldridge2005} Wooldridge, J. M. (2005). \newblock Simple solutions to the initial conditions problem in dynamic, nonlinear panel data models with unobserved heterogeneity. \newblock {\em Journal of Applied Econometrics\/} {\em 20\/}(1), 39--54.