EconBase
← Back to paper

Transition Probabilities and Moment Restrictions in Dynamic Fixed Effects Logit Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

120,683 characters · 20 sections · 125 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Transition Probabilities and Moment Restrictions in Dynamic Fixed Effects Logit Models

\setcitestyle{authoryear,open={(},close={)}}

abstract\setstretch{1.25} Dynamic logit models are popular tools in economics to measure state dependence. This paper introduces a new method to derive moment restrictions in a large class of such models with strictly exogenous regressors and fixed effects. We exploit the common structure of logit-type transition probabilities and elementary properties of rational fractions, to formulate a systematic procedure that scales naturally with model complexity (e.g the lag order or the number of observed time periods). We detail the construction of moment restrictions in binary response models of arbitrary lag order as well as first-order panel vector autoregressions and dynamic multinomial logit models. Identification of common parameters and average marginal effects is also discussed for the binary response case. Finally, we illustrate our results by studying the dynamics of drug consumption amongst young people inspired by deza2015there.

\noindentKeywords: dynamic discrete choice, panel data, fixed effects.

\noindentJEL Classification Codes: C23, C33.

Introduction

The analysis of state dependence is a classic and important topic in many areas of economics. Several discrete processes such as welfare and labor force participation manifest strong serial persistence, and economists have sought various methods to unravel the underlying factors. In this paper, we reexamine the estimation of one notable set of models employed for this purpose: discrete choice models with lagged dependent variables, strictly exogenous regressors, fixed effects and logistic errors. We shall refer to this class of models as dynamic fixed effects logit models (DFEL) throughout. Specifications of this kind are used to discriminate between \say{structural} state dependence, i.e the causal effect of past choices on current outcomes, and heterogeneity, i.e the serial correlation induced by unobserved individual attributes (heckman1981heterogeneity). An example of this approach is the analysis of welfare participation in chay1999non. There has been considerable interest in this family of panel data models in econometrics, with a recent surge in attention following new developments reported in honore2020moment. One general reason is that DFEL models stand out as a rare case of nonlinear dynamic panel data models for which solutions to the incidental parameters problem (neyman1948consistent) and initial conditions problem (e.g heckman1981heterogeneity) have been known to exist in short panels\footnote{The incidental parameters problem refers to the general inconsistency of maximum likelihood in short panels. The initial conditions problem refers to the general difficulty of formulating a correct conditional distribution for the initial observations given the fixed effects and covariates.}. \\ In the \say{pure} version of the basic model which abstracts from covariates other than a first order lag, cox1958regression, chamberlain_1985 and magnac2000subsidised showed that the autoregressive parameter can be consistently estimated by conditional likelihood. This approach relies on the existence of a sufficient statistic linked to the logistic assumption to eliminate the fixed effect. In an important subsequent paper, honore2000panel extended this idea to a setting with strictly exogenous regressors and showed that the conditional likelihood approach remains viable if one can further condition on the regressors being equal in specific periods. This strategy was also found to be effective in dynamic multinomial logit models (honore2000panel), panel vector autoregressions (honore2019panel) and dynamic ordered logit models (muris2020dynamic. At the same time, it has also been noted that the necessity to be able to \say{match} the covariates imposes two limitations for the conditional likelihood approach: it inherently rules out time effects and implies rates of convergence slower than $\sqrt{N}$ for continuous explanatory variables. Furthermore, calculations from honore2000panel suggested that it does not easily extend to models with a higher lag order. These shortcomings have motivated the search for alternative methods of estimation. \\ Recently, kitazawa2013exploration,kitazawa2016root and Kitazawa_JOE2021 revisited the AR(1) logit model - autoregressive of order one - of honore2000panel and proposed a transformation approach that deals with the fixed effects without restricting the nature of the covariates besides the conventional assumption of strict exogeneity. Their methodology leads to moment restrictions that can serve as a basis to estimate the model parameters at $\sqrt{N}$-rate by GMM; even with continuous regressors. In parallel work, honore2020moment also derived moment conditions for the AR(1), AR(2) and AR(3) logit models in panels of specific length using the functional differencing technique of Bonhomme_EM12. Their approach is partly numerical and relies on symbolic computing (e.g Mathematica) to obtain analytical expressions but has a wider scope of potential applications, e.g dynamic ordered logit specifications (honore2021dynamic). In another recent paper, dobronyi2021identification, the authors analyze the full likelihood of AR(1) and AR(2) logit models with discrete covariates under a new angle that reveals a connection to the truncated moment problem in mathematics. Drawing on well established results in that literature, they derive moment equality and new moment inequality restrictions that fully characterize the sharp identified set. \\ In this paper, we introduce a new systematic approach to construct moment restrictions in DFEL models with additive fixed effects, i.e when fixed effects are heterogeneous \say{intercepts}. This class of models encompasses most specifications studied in prior work but excludes models with heterogeneous coefficients on lagged outcomes and/or regressors as in chamberlain_1985 and browning2014dynamic. Unlike some recent competing approaches, we do not require numerical experimentation nor symbolic computing. Rather, as we shall see in examples, we exploit the common structure of logit-type transition probabilities and elementary properties of rational fractions, to obtain analytic expressions for the identifying moments. We shall focus our attention on deriving valid moment functions for AR($p$) models with arbitrary lag order $p\geq 1$ as well as first-order panel vector autoregressions and dynamic multinomial logit models (magnac2000subsidised). \\ Our methodology exploits two key observations. First, the transition probabilities of logit-type models can often be expressed as conditional expectations of functions of observables and common parameters given the initial condition, the regressors and the fixed effects. We shall refer to these moment functions as transition functions. They have the important feature of not depending on individual fixed effects. Second, as soon as $T\geq p+2$, where $T$ denotes the number of observations post initial condition, many transition probabilities in periods $t\in \{p+1,\ldots,T-1\}$ admit at least two distinct transition functions. The combination of these two features motivates a two-step approach to obtain moment restrictions in panels of adequate length. In the first step, we shall compute the model transition functions. Then, the second step will simply consist in differencing two transition functions associated to the same transition probability. We show that a careful application of this procedure delivers all the moment equality restrictions available in the binary response case. We shall further elaborate on these steps in examples and use the resulting moment functions to derive new identification results. At a high level, the approach we advocate in this paper consists in solving a sequence of problems with identical structure period by period instead of solving directly a large system of equations based on the model full likelihood as in honore2020moment and dobronyi2021identification. As a consequence, our procedure remains tractable when the number of time periods increases and in models with higher order lags. \\ Besides the aforementioned papers, our work also connects to a line of research studying the identification of features of the distribution of fixed effects in discrete choice models. One branch in this literature has focused on developing general optimization tools to compute sharp numerical bounds on average marginal effects. This includes most notably the linear programming method of honore2006bounds, recently adapted by bonhomme2023identification to the case of sequentially exogenous covariates, and the quadratic programming method of chernozhukov2013average. A second branch in this literature has sought instead to harness the specificities of logit models to obtain simple analytical bounds. In static logit models, davezies2021identification exploit mathematical results on the moment problem to formulate sharp bounds on the average partial effects of regressors on outcomes. In DFEL models, aguirregabiria2021identification are the first to prove the point identification of average marginal effects in the baseline AR(1) logit model when $T\geq 3$. In related work, dobronyi2021identification make use of their moment equality and moment inequality restrictions to establish sharp bounds on functionals of the fixed effects such as average marginal effects and average posterior means in AR(1) and AR(2) specifications. We complement these results as a byproduct of our methodology: average marginal effects and their variants in AR($p$) models, with arbitrary $p\geq 1$ are merely differences of average transition functions. \\ The remainder of the paper is organized as follows. Section (ref) presents the setting and our main objective. Section (ref) introduces some terminology and gives an outline of our procedure to construct moment restrictions. Section (ref) implements our approach in AR($p$) logit models with $p\geq 1$ and discusses identification of model parameters and average marginal effects. The semiparametric efficiency bound for the AR(1) is also presented for the base case of four waves of data. Section (ref) discusses extensions to the VAR(1) and the dynamic multinomial logit model with one lag, MAR(1) for short. In Section (ref), we present an empirical illustration on the dynamics of drug consumption amongst young people and Section (ref) offers concluding remarks. A complementary set of Monte Carlo simulations showing the small sample performance of GMM estimators based on our moment restrictions is available in Appendix Section (ref). Proofs are gathered in the Appendix.

Setting, assumptions and objective

Let $i=1,\ldots,N$ denote a population index and $t=0,\ldots,T$ be an index for time. We study DFEL models which may be viewed as threshold-crossing econometric specifications describing a discrete outcome $Y_{it}$ through a latent index involving lagged outcomes (e.g $Y_{it-1}$), strictly exogenous regressors $X_{it}$, an individual-specific time-invariant unobservable $A_i$ and an error term $\epsilon_{it}$. The canonical example is the AR(1) model:

align*[align* omitted — 116 chars of source]

and we shall concentrate more broadly on cases where $A_i$ is additively separable from the other explanatory variables. An initial condition that we will generically denote $Y_{i}^{0}$ completes such models to enable dynamics. The common parameter $\theta_0$ is one target of interest and governs the influence of lagged outcomes and the regressors on the contemporaneous outcome. Other quantities of interest include counterfactual parameters such as average marginal effects.\\ Throughout, we leave the joint distribution of $(Y_{i}^{0},X_i,A_i)$ unrestricted where $X_i=(X_{i1},\ldots,X_{iT})$ and thus refer to $A_i$ as a fixed effect in common with the literature. The schocks $\epsilon_{it}$ are assumed to be serially independent logistically distributed, independent of $(Y_{i}^{0},X_i,A_i)$, except for the MAR(1) model where they are instead extreme value distributed. Finally, we shall assume that $(Y_i,Y_{i}^{0},X_i,A_i)$ are jointly i.i.d across individuals.

The data available to the econometrician consists of the initial condition $Y_{i}^0$, the outcome vector $Y_{i}=(Y_{i1},\ldots,Y_{iT})$, and the covariates $X_i$ for all $N$ individuals. Interest centers primarily on the identification and estimation of $\theta_0$ in short panels, i.e for fixed $T$. To this end, the chief objective of this paper is to show how to construct moment functions $\psi_{\theta}(Y_i,Y_{i}^0,X_{i})$ free of the fixed effect parameter that are valid in the sense that:

align[align omitted — 129 chars of source]

When this is possible, the law of iterated expectations implies the conditional moment: $$ \mathbb{E}\left[\psi_{\theta_0}(Y_i,Y_{i}^{0},X_{i})\,|\,Y_{i}^0,X_{i}\right]=0$$ which can in turn be leveraged to assess the identifiability of $\theta_0$ and form the basis of a GMM estimation strategy. This is the central idea underlying functional differencing (Bonhomme_EM12) and was applied by honore2020moment to derive valid moment conditions for a class of dynamic logit models with scalar fixed effects. We borrow the same insight but instead of searching for solutions numerically on a case-by-case basis, we propose a complementary systematic algebraic procedure to recover the model's valid moments \footnote{dobronyi2021identification and Kitazawa_JOE2021 also have an algebraic approach but our methodologies are very different. The first paper uses the full likelihood of the model and focuses on the AR(1) and special instances of the AR(2) model. The second paper has a transformation approach adapted to the AR(1) model. Our emphasis here is primarily on developing an approach that is tractable for a large class of models.}. In doing so, we flesh out the mechanics implied by the logistic assumption which in turn suggest a blueprint to deal with estimation of general DFEL models. For example, we are able to characterize the expressions of valid moment functions in AR($p$) models for arbitrary $p>1$ which to the best of our knowledge is a new result in the literature. Furthermore, our approach carries over to multidimensional fixed effect specifications: VAR(1), dynamic network formation models and the MAR(1) in which searching for moments numerically is cumbersome or intractable. \\ In what follows, we shall use the shorthand $Y_{it_{1}}^{t_2}=(Y_{it_1},\ldots,Y_{it_2})$ to denote a collection of random variables over periods $t_1$ to $t_2$ with the convention that $Y_{it_{1}}^{t_2}=\emptyset$ if $t_1>t_2$. Likewise, we may use the notation $y_{t_1}^{t_2}=(y_{t_1},\ldots,y_{t_2})$ to denote any $(t_2-t_1)$-dimensional vector of reals with the convention $y_{t_{1}}^{t_{2}}=\emptyset$ for $t_1>t_2$. Elements $1_{n}$ and $0_{n}$ shall refer to the $n$-dimensional vectors of ones and zeros respectively. The support of the outcome variable $Y_{it}$ shall be denoted $\mathcal{Y}$. We let $\Delta$ denote the first-differencing operator so that $\Delta Z_{it}=Z_{it}-Z_{it-1}$ for any random variable $Z_{it}$ and make use of the notation $Z_{its}=Z_{it}-Z_{is}$ for $s\neq t$ to accommodate long differences. We use $\mathds{1}\{.\}$ for the indicator function; $\operatorname{Im}(f)$, $\ker(f)$, $\rank(f)$ to denote the image, the nullspace and the rank of a linear map $f$.

Outline of the procedure to derive valid moment functions

Let $T\geq 1$. Given an initial condition $y^0\in \mathcal{Y}^{p}$, $p\geq 1$ being the lag order of the model, and strictly exogenous regressors $X_i\in \mathbb{R}^{K_{x}\times T}$, we denote the (one-period ahead) transition probability in period $t\geq 1$ from state $(l_{1}^t,y^ 0)\in \mathcal{Y}^{t}\times \mathcal{Y}^{p}$ to state $k\in \mathcal{Y}$ as:

align*[align* omitted — 158 chars of source]

With $p$ lags, the markovian nature of the models considered in this paper imply that $ \pi^{k|l_{1}^t,y^0}_{t}(A_i,X_{i})$ will not depend on the entire path of past outcomes but only on the value of the most recent $p$ outcomes. For instance, in an AR(1) model where $p=1$, we have:

align*[align* omitted — 148 chars of source]

and thus we will suppress the dependence on $(y^0,l_1,\ldots,l_{t-1})$ and write $\pi^{k|l_t}_{t}(A_i,X_{i})$. We shall proceed analogously for the more general case $p\geq 1$. \\ We call a transition function associated to a transition probability $ \pi^{k|l_{1}^t,y^0}_{t}(A_i,X_{i})$ any moment function $\phi^{k|l_{1}^t,y^0}_{\theta}(Y_i,Y_{i}^{0},X_i)$ of the data and the common parameters verifying:

align[align omitted — 171 chars of source]

With these notions in hand, we are ready to describe our two-step approach to derive valid moment functions in the sense of equation ((ref)). In Step 1), we begin by computing the model's transition functions. Our procedure requires a minimum of $T=p+1$ periods of observations to accommodate arbitrary regressors and initial condition. In this case, we can get analytical formulas for the transition functions associated to the transition probabilities in period $t=p$ and Theorem (ref) and Theorem (ref) below imply that they are unique. However, this is not immediately helpful to get moment (equality) restrictions on $\theta_0$. We require one more period. As soon as $T\geq p+2$, we explain how to construct distinct transition functions associated to the same transition probabilities in periods $t\in\{p+1,\ldots,T-1\}$. The key ingredient is the use of partial fraction decompositions for rational fractions adapted to the structure of the transition probabilities. It is then a matter of taking differences of two transition functions associated to the same transition probability to obtain valid moment functions; we refer to this last step as Step 2). The ensuing sections demonstrate this procedure in scalar and multidimensional fixed effect models.

Scalar fixed effect models

Moment restrictions for the AR(1) logit model

For exposition, we begin with the baseline AR(1) logit model with fixed effects introduced above:

align[align omitted — 140 chars of source]

Here, $\mathcal{Y}=\{0,1\}$, $\theta_0=(\gamma_0,\beta_0')\in \mathbb{R}\times \mathbb{R}^{K_{x}}$, the initial condition $Y_{i}^0$ consists of the binary-valued random variable $Y_{i0}$ and $A_i\in \mathbb{R}$.

The number of moment restrictions in the AR(1)

We start out by enumerating the moment restrictions implied by the model. This will provide a means to assess the exhaustiveness of our approach. To this end, let $\mathcal{E}_{y_0,x}$ denote the conditional expectation operator mapping any function of the outcome variable $Y_i$ to its conditional expectation given $Y_{i0}=y_0,X_{i}=x$ and the fixed effect $A_i$, i.e

align*[align* omitted — 212 chars of source]

For example, for any $y\in \mathcal{Y}^T$, $\mathcal{E}_{y_0,x}\left[\mathds{1}\{.=y\}\right]$ yields the conditional probability of observing history $y$ for all possible values of the fixed effect, i.e:

align*[align* omitted — 104 chars of source]

where $P(Y_i=y|Y_{i0}=y_0,X_{i}=x,A_{i}=a)=\prod \limits_{t=1}^T \frac{e^{y_t(\gamma_0y_{t-1}+x_t'\beta_0+a)}}{1+e^{\gamma_0y_{t-1}+x_t'\beta_0+a}}, \quad \forall a \in \mathbb{R}$ . Then, we have the following result,

theoremConsider model ((ref)) with $T\geq 1$ and initial condition $y_0\in \mathcal{Y}$. Suppose that for any $t,s\in \{1,\ldots,T-1\}$ and $y,\tilde{y}\in \mathcal{Y}$, $\gamma_0y+x_{t}'\beta_0\neq \gamma_0\tilde{y}+x_{s}'\beta_0$ if $t\neq s$ or $y\neq\tilde{y}$. Then, the family $ \mathcal{F}_{y_0,T}=\left\{1,\pi_{0}^{y_0|y_0}(.,x),(\pi_{t}^{0|0}(.,x),\pi_{t}^{1|1}(.,x))_{t=1}^{T-1}\right\}$ of size $2T$ forms a basis of $\operatorname{Im}(\mathcal{E}_{y_0,x})$ and $\dim\left(\ker(\mathcal{E}_{y_0,x})\right)=2^T-2T$.

Theorem (ref) formalizes the intuition that the transition probabilities summarize the parametric component of the model: $2^T$ histories are possible yet only $2T$ basis elements are necessary to fully characterize their conditional probabilities. This follows from the observation that when the covariate index \footnote{We refer to the quantity $\gamma_0y_{t-1}+x_{t}'\beta_0$ for a given period $t$.} of each transition probability differ, the conditional probability of each history $y\in \mathcal{Y}^{T}$ is a ratio of polynomials in $e^{a}$, where the numerator has lower degree than the denominator, and the later is a product of distinct irreducible terms. A sufficient condition for this is that $\gamma_0\neq 0$ and that one regressor is continuously distributed with non-zero slope. In turn, standard results on partial fraction decompositions ensure that this ratio can be expressed as a unique linear combination of transition probabilities. To finally conclude that $\mathcal{F}_{y_0,T}$ is a basis of $\operatorname{Im}(\mathcal{E}_{y_0,x})$, we leverage upcoming results demonstrating that the transition probabilities live in $\operatorname{Im}(\mathcal{E}_{y_0,x})$ as expectations of transition functions. \\ Importantly, since $\ker(\mathcal{E}_{y_0,x})$ is the set of valid moment functions verifying equation ((ref)), Theorem (ref) tells us that the AR(1) model features $2^T-2T$ linearly independent moment restrictions in general. This is a consequence of the rank nullity theorem for linear maps with finite dimensional domains. The fact that $2^T-2T$ moment conditions are available for the AR(1) appeared initially as a conjecture in honore2020moment and was later established by dobronyi2021identification using different arguments from here. They do not emphasize the role of the transition probabilities. Our ideas extend naturally to the case of arbitrary lags which was hitherto an open problem. We discuss this extension in Section (ref).

remark[Counting moments in logit models] The idea of decomposing the conditional probabilities of all choice histories in a basis provides a useful device to infer a lower bound on the number of moment restrictions in logit models. If one can further prove that elements of this basis belong to the image of the conditional expectation operator, then this lower bound coincides with the exact number of moment restrictions. \begin{itemize} • In the static panel logit model of rasch1960studies, $\gamma_0=0$ and we have $\pi_{t}^{1|1}(.,x)=1-\pi_{t}^{0|0}(.,x)$. Thus, provided that $x_t'\beta_0\neq x_s'\beta_0$ for all $t\neq s$, $\mathcal{F}_{T}=\left\{1,(\pi_{t}^{0|0}(.,x))_{t=0}^{T-1}\right\}$ spans the image of the conditional expectation operator. This implies at least $2^T-(T+1)$ moment restrictions. It turns out that $2^T-(T+1)$ is precisely the total number of moment restrictions for this model. This follows from Remark (ref) below which characterizes the transition functions associated to each element of $\mathcal{F}_{T}$. • In the cox1958regression model, $\gamma_0\neq 0$ and $\beta_0=0$ and the transition probabilities are: $\pi^{0|0}(a)=\frac{1}{1+e^{a}}$ and $\pi^{1|1}(a)=\frac{e^{\gamma_0+a}}{1+e^{\gamma_0+a}}$ (or equivalently $\pi^{0|1}(a)=\frac{1}{1+e^{\gamma_0+a}}$). See the next section for further details. In this case, the family $\mathcal{F}_{y_0,T}=\left\{1,\left(\pi^{0|0}(.)^j,\pi^{0|1}(.)^j\right)_{j=1}^{T-1},\pi^{0|y_0}(.)^T \right\}$ which consists of powers of the time-invariant transition probabilities spans the image of the conditional expectation operator. Since $|\mathcal{F}_{y_0,T}|=2T$, the model produces at least $2^T-2T$ linearly independent moment restrictions. \end{itemize}
remark[A matrix perspective] Since $\mathcal{E}_{y_0,x}$ is a linear map, it admits a unique $2^T\times 2T $ matrix representation $\Lambda_{y_0,x}$ where each row translates the conditional probability of a choice history $y\in \mathcal{Y}^T$ in terms of the transition probabilities of $\mathcal{F}_{y_0,T}$\footnote{ Entries of this matrix may be found using for example the identities in Appendix Lemma (ref) or any other standard textbook tools for rational fractions.}. From this point of view, valid moments correspond to $2^T$-vectors $\psi$ in the left nullspace of $\Lambda_{y_0,x}$, meaning $\psi'\Lambda_{y_0,x}=0$. Constructing $\Lambda_{y_0,x}$ and then solving this $2T$ linear system of equations in $2^T$ unknowns directly is straightforward using symbolic tools when $T$ is \say{small} (e.g dobronyi2021identification, honore2020moment) but is computationally impractical otherwise. Instead, we propose a constructive approach to back out analytic expressions of the valid moment functions that is tractable for arbitrary values of $T$.

Having clarified the total count of moment restrictions in the AR(1) logit model, we next discuss how to construct them with our two-step procedure.

Construction of valid moment functions for the pure model

In the absence of exogenous regressors, model ((ref)) simplifies to:

align[align omitted — 122 chars of source]

which was first introduced by cox1958regression and then revisited in chamberlain_1985, magnac2000subsidised. These papers established the identification of $\gamma_0$ for $T\geq 3$ via conditional likelihood based on the insight that $(Y_{i0},\sum_{t=1}^{T-1} Y_{it},Y_{iT})$ are sufficient statistics for the fixed effect. Our methodology is conceptually different as we seek to directly construct moment functions verifying equation ((ref)). \\ For what follows, it is helpful to remember that the individual-specific transition probability from state $l$ to state $k$ is time-invariant and given by:

align*[align* omitted — 145 chars of source]

Step 1). We shall begin by deriving the transition functions for $\pi^{0|0}(A_i)$ and $\pi^{1|1}(A_i)$. Observe that $\pi^{1|0}(A_i)$ and $\pi^{0|1}(A_i)$ are effectively redundant since probabilities sum to one. A natural starting place is to investigate the case $T=2$, i.e 2 periods of observations after the initial condition. Recalling definition ((ref)), we search for $\phi_{\theta}^{0|0}(Y_{i2},Y_{i1},Y_{i0})$, respectively $\phi_{\theta}^{1|1}(Y_{i2},Y_{i1},Y_{i0})$, whose conditional expectation given $(Y_{i0},A_i)$ yields $\pi^{0|0}(A_i)$, respectively $\pi^{1|1}(A_i)$. For the purposes of illustration and to show the kind of calculations arising broadly in DFEL models, let us derive $\phi_{\theta}^{0|0}(Y_{i2},Y_{i1},Y_{i0})$. By Bayes's rule:

multline*[multline* omitted — 620 chars of source]

where the second equality uses the logistic hypothesis. By quick inspection, we see that the terms in the first parenthesis have $(1+e^{\gamma_0+a})$ in their denominator unlike $\pi^{0|0}(A_i)$. Because $-e^{-\gamma_0}$ is not a pole of $\pi^{0|0}(A_i)$\footnote{A pole of a rational function is a root of its denominator. Formally, we are substituting $u=e^{a}$ and we are extending $\pi^{0|0}(u)$ to the real line.}, we conclude that $\phi_{\theta}^{0|0}(1,1,y_{0})=\phi_{\theta}^{0|0}(0,1,y_{0})=0$. This first deduction leaves us with

align*[align* omitted — 257 chars of source]

Now, since $\pi^{0|0}(A_i)$ does not depend on $y_0$, we must cancel the denominator $(1+e^{\gamma_0 y_{0}+a})$. To achieve this, we must set: $\phi_{\theta_0}^{0|0}(1,0,y_{t-1})=C_0 e^{\gamma_0 y_{0}}, \phi_{\theta_0}^{0|0}(0,0,y_{t-1})=C_0$ for some constant $C_0\in \mathbb{R}\setminus\{0\}$. Then,

align*[align* omitted — 124 chars of source]

and $C_0=1$ is the appropriate normalization to obtain the desired transition function. Of course, the exact same logic applies for $\phi_{\theta_0}^{1|1}(Y_{i2},Y_{i1},Y_{i0})$ and $\pi^{1|1}(A_i)$. \\ This short calculation provides a useful recipe for the general case $T\geq 2$. We learned that we can search for functions of three consecutive outcomes $\phi_{\theta}^{k|k}(Y_{it+1},Y_{it},Y_{it-1})$ such that:

align*[align* omitted — 252 chars of source]

The first restriction is a functional form that eliminates terms with inadequate poles after taking expectations. The second restriction is a normalization condition to match the desired transition probability. Following this argument, we arrive at the expressions in Lemma (ref).

lemmaIn model ((ref)) with $T\geq 2$ and $t\in\{1,\ldots,T-1\}$, let \begin{align*} \phi_{\theta}^{0|0}(Y_{it+1},Y_{it},Y_{it-1})&=(1-Y_{it})e^{\gamma Y_{it+1}Y_{it-1}} \\ \phi_{\theta}^{1|1}(Y_{it+1},Y_{it},Y_{it-1})&=Y_{it}e^{\gamma (1-Y_{it+1})(1-Y_{it-1})} \end{align*} Then: \begin{align*} \mathbb{E}\left[\phi_{\theta_0}^{0|0}(Y_{it+1},Y_{it},Y_{it-1})|Y_{i0},Y_{i1}^{t-1},A_i\right]&=\pi^{0|0}(A_i)=\frac{1}{1+e^{A_i}} \\ \mathbb{E}\left[\phi_{\theta_0}^{1|1}(Y_{it+1},Y_{it},Y_{it-1})|Y_{i0},Y_{i1}^{t-1},A_i\right]&=\pi^{1|1}(A_i)=\frac{e^{\gamma_0+A_i}}{1+e^{\gamma_0+A_i}} \end{align*}
remark[Connection to Kitazawa] Interestingly, Lemma (ref) is a reformulation of results first shown by kitazawa2013exploration,kitazawa2016root, Kitazawa_JOE2021, albeit with a very different logic than the calculations displayed above. We set out the connection between our respective approaches in Section (ref) where we also discuss the case with exogenous regressors.

Step 2). The second step in the agenda is the construction of valid moment functions. Because the transition probability of the model are time-invariant, one trivial way to achieve this is to consider the pairwise difference of $\phi_{\theta}^{k|k}(Y_{it+1},Y_{it},Y_{it-1})$ and $\phi_{\theta}^{k|k}(Y_{is+1},Y_{is},Y_{is-1})$ for any feasible $s\neq t$. This is the content of Proposition (ref). We will need a minimum of four total periods of observations, which coincides with the requirements of the conditional likelihood approach.

propositionIn model ((ref)) with $T\geq 3$, let \begin{align*} \psi_{\theta}^{k|k}(Y_{it-1}^{t+1},Y_{is-1}^{s+1})&=\phi_{\theta}^{k|k}(Y_{it+1},Y_{it},Y_{it-1})-\phi_{\theta}^{k|k}(Y_{is+1},Y_{is},Y_{is-1}) \end{align*} for all $ k \in \mathcal{Y} $, $t \in \{2,\ldots,T-1\}$ and $s \in \{1,\ldots,t-1\}$. Then, \begin{align*} \mathbb{E}\left[\psi_{\theta_0}^{k|k}(Y_{it-1}^{t+1},Y_{is-1}^{s+1})|Y_{i0},Y_{i1}^{s-1},A_i\right]&=0 \end{align*}
remark[Efficient GMM] Given that the conditional likelihood is semi-parametrically efficient for $T=3$ (gu2023information, hahn2001information), it is natural to ask whether the approach advocated here accounts for all the information in the model in that case. It turns out that it does. Specifically, letting $s_i^{c}(\theta)$ denote the conditional scores when $y_0=0$ as in hahn2001information, we have: \begin{align*} s_{i}^{c}(\gamma_0)&=\frac{1}{(1+e^{\gamma_0})(e^{-\gamma_0}-1)}\left(\psi_{\theta}^{0|0}(Y_{i1}^{3},Y_{i1}^{2},0)+\psi_{\theta}^{1|1}(Y_{i1}^{3},Y_{i1}^{2},0)\right) \end{align*} where the right-hand side corresponds to the efficient moment for the moment restriction $\mathbb{E}\left[\psi_{\theta}(Y_{i1}^{3},Y_{i0}^{2})|Y_{i0}=0\right]=0$, $\psi_{\theta}(Y_{i1}^{3},Y_{i1}^{2},0)=(\psi_{\theta}^{0|0}(Y_{i1}^{3},Y_{i1}^{2},0),\psi_{\theta}^{1|1}(Y_{i1}^{3},Y_{i1}^{2},0))'$.

Construction of valid moment functions with strictly exogenous regressors

In this subsection, we move on to the AR(1) logit model with strictly exogenous covariates characterized by equation ((ref)). \\ Step 1). We employ the same shortcut recipe as in the \say{pure} case and begin by looking for moment functions $\phi_{\theta}^{0|0}(.)$ and $\phi_{\theta}^{1|1}(.)$ verifying:

align*[align* omitted — 296 chars of source]

where this time

align*[align* omitted — 194 chars of source]

The same simple calculations described just above lead to the expressions in Lemma (ref). The only (expected) change is the appearance of a new term $+/-\Delta X_{it+1}'\beta$ which accounts for the presence of covariates in the model.

lemmaIn model ((ref)) with $T\geq 2$ and $t\in\{1,\ldots,T-1\}$, let \begin{align*} \phi_{\theta}^{0|0}(Y_{it+1},Y_{it},Y_{it-1},X_i)&=(1-Y_{it})e^{ Y_{it+1}\left(\gamma Y_{it-1}-\Delta X_{it+1}'\beta \right)} \\ \phi_{\theta}^{1|1}(Y_{it+1},Y_{it},Y_{it-1},X_i)&=Y_{it}e^{ (1-Y_{it+1})\left(\gamma(1-Y_{it-1})+ \Delta X_{it+1}'\beta \right)} \end{align*} Then: \begin{align*} \mathbb{E}\left[\phi_{\theta_0}^{0|0}(Y_{it+1},Y_{it},Y_{it-1},X_i)|Y_{i0},Y_{i1}^{t-1},X_i,A_i\right]&= \pi^{0|0}_{t}(A_i,X_{i})=\frac{1}{1+e^{A_i+X_{it+1}'\beta_0}} \\ \mathbb{E}\left[\phi_{\theta_0}^{1|1}(Y_{it+1},Y_{it},Y_{it-1},X_i)|Y_{i0},Y_{i1}^{t-1},X_i,A_i\right]&=\pi^{1|1}_{t}(A_i,X_{i})=\frac{e^{\gamma_0+X_{it+1}'\beta_0+A_i}}{1+e^{\gamma_0+X_{it+1}'\beta_0+A_i}} \end{align*}

At this point, it is important to highlight that unlike previously, the transition probabilities are covariate-dependent. The upshot is that the naive difference of $\phi_{\theta}^{k|k}(Y_{it+1},Y_{it},Y_{it-1},X_i)$ and $\phi_{\theta}^{k|k}(Y_{is+1},Y_{is},Y_{is-1},X_i)$ for $s \neq t$ no longer leads to valid moment functions in general. Indeed, while Lemma (ref) ensures that

align*[align* omitted — 202 chars of source]

clearly, $\pi_{t}^{k|k}(A_i,X_i)-\pi_{s}^{k|k}(A_i,X_i)\neq 0$ when $X_{it+1}'\beta_0\neq X_{is+1}'\beta_0$ \footnote{A matching strategy in the spirit of honore2000panel may still be applicable when in our example $X_{it+1}= X_{is+1}$. However, this is known to lead to estimators converging at rate less than $\sqrt{N}$ for continuous covariates and it rules out certain regressors such as time dummies and time trends.}. Thus, a different logic is required in the presence of explanatory variables other than a first order lag. \\ The key, as foreshadowed in Section (ref) is that as soon as $T\geq 3$, it is possible to construct transition functions other than $\phi_{\theta}^{k|k}(Y_{it-1}^{t+1},X_i)$ also mapping to $\pi^{k|k}_{t}(A_i,X_{i})$ in time periods $t\in\{2,\ldots,T-1\}$. These new transition functions that we denote $\zeta_{\theta}^{k|k}(.)$ to emphasize their difference have a particular form. They consist of a weighted combination of past outcome $\mathds{1}(Y_{is}=k)$, $1 \leq s<t$, and the interaction of $\mathds{1}(Y_{is}\neq k)$ with any transition function associated to $\pi^{k|k}_{t}(A_i,X_{i})$ having no dependence on outcomes prior to period $s$, e.g $\phi_{\theta}^{k|k}(Y_{it-1}^{t+1},X_i)$. This property follows from a partial fraction decomposition presented in Lemma (ref) that exploits the structure of the model probabilities under the logistic assumption. It relates to the hyperbolic transformations ideas of Kitazawa_JOE2021. In the sequel, we shall see that this insight carries over to the AR($p$) logit model with $p>1$. Lemma (ref) below gives the \say{simplest} additional transition functions that one can construct when $T\geq 3$ for the AR(1) model with exogenous regressors (the only ones when $T=3$).

lemmaIn model ((ref)) with $T\geq 3$, for all $t,s$ such that $T-1\geq t> s\geq 1$, let: \begin{align*} \mu_{s}(\theta)&=\gamma Y_{is-1}+X_{is}'\beta \\ \kappa_{t}^{0|0}(\theta)&=X_{it+1}'\beta, \quad \kappa_{t}^{1|1}(\theta)=\gamma+X_{it+1}'\beta \\ \omega_{t,s}^{0|0}(\theta)&=1-e^{(\kappa_{t}^{0|0}(\theta)-\mu_{s}(\theta))}, \quad \omega_{t,s}^{1|1}(\theta)=1-e^{-(\kappa_{t}^{1|1}(\theta)-\mu_{s}(\theta))} \end{align*} and define the moment functions: \begin{align*} \zeta_{\theta}^{0|0}(Y_{it-1}^{t+1},Y_{is-1}^s,X_i)&=(1-Y_{is})+\omega_{t,s}^{0|0}(\theta)Y_{is}\phi_{\theta}^{0|0}(Y_{it+1},Y_{it},Y_{it-1},X_i) \\ \zeta_{\theta}^{1|1}(Y_{it-1}^{t+1},Y_{is-1}^s,X_i)&=Y_{is}+\omega_{t,s}^{1|1}(\theta)(1-Y_{is})\phi_{\theta}^{1|1}(Y_{it+1},Y_{it},Y_{it-1},X_i) \end{align*} Then, \begin{align*} \mathbb{E}\left[ \zeta_{\theta_0}^{0|0}(Y_{it-1}^{t+1},Y_{is-1}^s,X_i)|Y_{i0},Y_{i1}^{s-1},X_i,A_i\right]&= \pi^{0|0}_{t}(A_i,X_{i})=\frac{1}{1+e^{X_{it+1}'\beta_0+A_i}} \\ \mathbb{E}\left[\zeta_{\theta_0}^{1|1}(Y_{it-1}^{t+1},Y_{is-1}^s,X_i)|Y_{i0},Y_{i1}^{s-1},X_i,A_i\right]&=\pi^{1|1}_{t}(A_i,X_{i})=\frac{e^{\gamma_0+X_{it+1}'\beta_0+A_i}}{1+e^{\gamma_0+X_{it+1}'\beta_0+A_i}} \end{align*}

When $T\geq 4$, it turns out that we can build even more transition functions from those given in Lemma (ref) by repeating the same type of logic based on partial fraction expansions; Corollary (ref) provides a recursive formulation.

corollaryIn model ((ref)) with $T\geq 4$, for any $t$ and ordered collection of indices $s_1^J$, $J\geq 2$, satisfying $T-1\geq t> s_1>\ldots>s_J\geq 1$, let \begin{align*} \zeta_{\theta}^{0|0}(Y_{it-1}^{t+1},Y_{is_1-1}^{s_1},\ldots,Y_{is_{J}-1}^{s_J},X_i)&=(1-Y_{is_J})+\omega_{t,s_J}^{0|0}(\theta)Y_{is_J} \zeta_{\theta}^{0|0}(Y_{it-1}^{t+1},Y_{is_1-1}^{s_1},\ldots,Y_{is_{J-1}-1}^{s_{J-1}},X_i) \\ \zeta_{\theta}^{1|1}(Y_{it-1}^{t+1},Y_{is_1-1}^{s_1},\ldots,Y_{is_{J}-1}^{s_J},X_i)&=Y_{is_J}+\omega_{t,s_J}^{1|1}(\theta)(1-Y_{is_J}) \zeta_{\theta}^{1|1}(Y_{it-1}^{t+1},Y_{is_1-1}^{s_1},\ldots,Y_{is_{J-1}-1}^{s_{J-1}},X_i) \end{align*} with weights $\omega_{t,s_J}^{0|0}(\theta), \omega_{t,s_J}^{1|1}(\theta)$ defined as in Lemma (ref). Then, \begin{align*} \mathbb{E}\left[ \zeta_{\theta_0}^{k|k}(Y_{it-1}^{t+1},Y_{is_1-1}^{s_1},\ldots,Y_{is_{J}-1}^{s_J},X_i)|Y_{i0},Y_{i1}^{s_J-1},X_i,A_i\right]&= \pi^{k|k}_{t}(A_i,X_{i}), \quad \forall k \in \mathcal{Y} \end{align*}

Step 2). Provided $T\geq 3$, the difference between any transition functions associated to the same transition probabilities in periods $t\in\{2,\ldots,T-1\}$ constitutes a valid candidate for ((ref)). One particularly relevant set of valid moment functions for reasons explained below is presented in Proposition (ref).

propositionIn model ((ref)), for all $k \in \mathcal{Y}$, \\ if $T\geq 3$, for all $t,s$ such that $T-1\geq t> s\geq 1$ , let \begin{align*} \psi_{\theta}^{k|k}(Y_{it-1}^{t+1},Y_{is-1}^s,X_i)&=\phi_{\theta}^{k|k}(Y_{it-1}^{t+1},X_i)-\zeta_{\theta}^{k|k}(Y_{it-1}^{t+1},Y_{is-1}^s,X_i), \end{align*} if $T\geq 4$, for any $t$ and ordered collection of indices $s_1^J$, $J \geq 2$, satisfying $T-1\geq t> s_1>\ldots>s_J\geq 1$, let \begin{align*} \psi_{\theta}^{k|k}(Y_{it-1}^{t+1},Y_{is_1-1}^{s_1},\ldots,Y_{is_{J}-1}^{s_J},X_i)&=\phi_{\theta}^{k|k}(Y_{it-1}^{t+1},X_i)-\zeta_{\theta}^{k|k}(Y_{it-1}^{t+1},Y_{is_1-1}^{s_1},\ldots,Y_{is_{J}-1}^{s_J},X_i), \end{align*} Then, \begin{align*} &\mathbb{E}\left[\psi_{\theta_0}^{k|k}(Y_{it-1}^{t+1},Y_{is-1}^s,X_i)|Y_{i0},Y_{i1}^{s-1},X_i,A_i\right]=0 \\ &\mathbb{E}\left[\psi_{\theta_0}^{k|k}(Y_{it-1}^{t+1},Y_{is_1-1}^{s_1},\ldots,Y_{is_{J}-1}^{s_J},X_i)|Y_{i0},Y_{i1}^{s_J-1},X_i,A_i\right]=0 \end{align*}

This family of moment functions has cardinality $2^T-2T$ which by Theorem (ref) is precisely the number of linearly independent moment conditions available for the AR(1). To see this, notice that for fixed $(k,Y_{i0})\in \mathcal{Y}^2$, and a given time period $ t\in\{ 2,\ldots,T-1\}$, Proposition (ref) gives a total of:

align*[align* omitted — 59 chars of source]

valid moment functions. This follows from a simple counting argument. First, we get $\binom{t-1}{1}$ possibilities from choosing any $s$ in $\{1,\ldots,t-1\}$ to form $\psi_{\theta}^{k|k}(Y_{it-1}^{t+1},Y_{is-1}^s,X_i)$. To that, we must add another $\sum_{l=2}^{t-1} \binom{t-1}{l}$ possibilities from choosing all feasible sequences $s_1^J$ with $t-1\geq s_1>s_2>\ldots>s_J\geq 1$ to form $\psi_{\theta}^{k|k}(Y_{it-1}^{t+1},Y_{is_1-1}^{s_1},\ldots,Y_{is_{J}-1}^{s_J},X_i)$. Summing over $t=2,\ldots, T-1$ and multiplying by 2 to account for the two possible values for $k$ delivers the result:

align*[align* omitted — 118 chars of source]

Furthermore, there is evidence that the family is linearly independent. It is readily verified for $T=3$ since the two valid moment functions produced by the model depend on two distinct sets of choice histories. This can be seen from their unpacked expressions in equations ((ref)) and ((ref)) in the Appendix. Unfortunately, this argument does not carry over to longer panels but we have verified numerically that the linear independence property of this family continues to hold for several different values of $T\geq 4$. This suggests that our approach delivers all the moment equality restrictions available in the AR(1) model with $T$ periods post initial condition \footnote{This is not all the identifying content of the AR(1) specification since we know from dobronyi2021identification that the model also implies moment inequality conditions.}.

remark[Symmetry] The transition functions and valid moment functions of the AR(1) model share a special symmetry property. Indeed, by inspection the transition functions of Lemma (ref) verify \begin{align*} \phi_{\theta}^{0|0}(1-Y_{it+1},1-Y_{it},1-Y_{it-1} ,-X_i)=\phi_{\theta}^{1|1}(Y_{it+1},Y_{it},Y_{it-1},X_i) \end{align*} It is not difficult to see that this symmetry, i.e substituting $Y_{it}$ by $(1-Y_{it})$ and $X_{it}$ by $-X_{it}$ to obtain $\phi_{\theta}^{1|1}(Y_{it-1}^{t+1},X_i)$ from $\phi_{\theta}^{0|0}(Y_{it-1}^{t+1},X_i)$ transfers to the other transition functions of Lemma (ref), Corollary (ref) and ultimately to the valid moment functions of Proposition (ref). This symmetry can be useful for computational purposes.
remark[Static logit] If $\gamma_0=0$, model ((ref)) specializes to the static panel logit model of rasch1960studies and our two-step approach is still applicable. For that case, Lemma (ref) gives two moment functions for $T=2$: \begin{align*} &\phi_{\theta}^{0|0}(Y_{i2},Y_{i1},X_{i})=(1-Y_{1})e^{ -Y_{i2}\Delta X_{2}'\beta} \\ &\phi_{\theta}^{1|1}(Y_{i2},Y_{i1},X_{i})=Y_{i1}e^{ (1-Y_{2})\Delta X_{i2}'\beta} \end{align*} such that $\mathbb{E}\left[\phi_{\theta_0}^{0|0}(Y_{i1}^2,,X_{i})|X_i,A_i\right]=\frac{1}{1+e^{X_{i2}'\beta_0+A_i}}$ and $\mathbb{E}\left[\phi_{\theta_0}^{1|1}(Y_{i1}^2,X_{i})|X_i,A_i\right]=\frac{e^{X_{i2}'\beta_0+A_i}}{1+e^{X_{i2}'\beta_0+A_i}}$. It follows that a valid moment function with two periods of observation is \begin{align*} \psi_{\theta}(Y_{i2},Y_{i1},X_{i})&=\phi_{\theta}^{1|1}(Y_{i2},Y_{i1},X_{i})-(1-\phi_{\theta}^{0|0}(Y_{i2},Y_{i1},X_{i})) \\ &=(1-e^{-\Delta X_{i2}'\beta})\left(Y_{i1}(1-Y_{i2})e^{\Delta X_{i2}'\beta}-(1-Y_{i1})Y_{i2}\right) \end{align*} which is proportional to the score of the conditional likelihood based on the sufficient statistic $Y_{i1}+Y_{i2}$ (rasch1960studies, andersen1970asymptotic, chamberlain1980analysis).

Semiparametric efficiency bound for the AR(1) with regressors

honore2020moment gave sufficient conditions to identify $\theta_0=(\gamma_0,\beta_0')'$ in the AR(1) model with $T=3$. A natural follow-up question is to ask how accurately can $\theta_0$ be estimated in that case, or equivalently what is the semi-parametric information bound. In a corrigendum to hahn2001information, gu2023information confirmed that the conditional likelihood estimator is semiparametrically efficient for $T=3$ in the \say{pure} AR(1) model. However, the characterization of the semiparametric efficiency bound and the question of what estimator attains it remain unclear with covariates. \\ To answer these questions, let $\psi_{\theta}(Y_{i1}^{3},Y_{i0}^1,X_i)=(\psi_{\theta}^{0|0}(Y_{i1}^{3},Y_{i0}^1,X_i),\psi_{\theta}^{1|1}(Y_{i1}^{3},Y_{i0}^1,X_i))'$ where the two components correspond to the valid moment functions of Proposition (ref) for $T=3$. Additionally, let $D(X_i,y_0)=\mathbb{E}\left[\pdv{ \psi_{\theta_0}(Y_{i1}^{3},Y_{i0}^1,X_i)}{\theta'}|Y_{i0}=y_0,X_i\right]$ and let $\Sigma(X_i,y_0)=\mathbb{E}\left[\psi_{\theta_0}(Y_{i1}^{3},Y_{i0}^1,X_i)\psi_{\theta_0}(Y_{i1}^{3},Y_{i0}^1,X_i)'|Y_{i0}=y_0,X_i\right]$.

assumptionIn model ((ref)) with $T=3$ and initial condition $y_0\in \{0,1\}$, the matrix $\mathbb{E}\left[D(X_i,y_0)\Sigma(X_i,y_0)^{-1}D(X_i,y_0)'|Y_{i0}=y_0\right]$ exists and is nonsingular.

With these notations in hand and under the mild conditions of Assumption (ref), Theorem (ref) clarifies that the efficient score coincides with the efficient moment for the conditional moment problem: $\mathbb{E}\left[\psi_{\theta}(Y_{i1}^{3},Y_{i0}^1,X_i)|Y_{i0}=y_0,X_i\right]=0$. Put differently, the maximal efficiency with which $\theta_0$ can be estimated is $V_{0}(y_0)=\mathbb{E}[D(X_i,y_0)\Sigma(X_i,y_0)^{-1}D(X_i,y_0)'|Y_{i0}=y_0]^{-1}$. This result is in accordance with Remark (ref) which noted that the score of the conditional likelihood without covariates is precisely the efficient moment implied by our conditional moment restrictions in this case.

theoremConsider model ((ref)) with $T=3$. Fix an initial condition $y_0\in \{0,1\}$ and suppose that Assumption (ref) holds. Then, the semiparametric efficiency bound of $\theta_0$ is finite and given by $V_0(y_0) = \mathbb{E}[D(X_i,y_0)\Sigma(X_i,y_0)^{-1}D(X_i,y_0)'|Y_{i0}=y_0]^{-1}$.

The proof of Theorem (ref) only involves careful bookkeeping of some tedious algebra and an application of Theorem 3.2 in newey1990semiparametric. Interestingly, davezies2023fixed presented analogous results in the static panel data case with three periods of observations.

Connections to other works on the AR(1) logit model

As indicated previously, there is a connection between our methodology and that of Kitazawa_JOE2021 for the AR(1) model. Indeed, after some algebraic manipulation, we can re-express the transition functions of Lemma (ref) (or Lemma (ref) without covariates) as:

align*[align* omitted — 360 chars of source]

where $\delta=(e^{\gamma}-1)$. Thus, the moment conditions of Lemma (ref) imply that we can write:

align*[align* omitted — 475 chars of source]

where $\mathbb{E}\left[\epsilon_{it}^{0|0}|Y_{i0},Y_{i1}^{t-1},X_i,A_i\right]=0$ and $\mathbb{E}\left[\epsilon_{it}^{1|1}|Y_{i0},Y_{i1}^{t-1},X_i,A_i\right]=0$. These expressions are the so-called h-form and g-form of Kitazawa_JOE2021 for model ((ref)) and were originally obtained through an ingenious usage of the mathematical properties of the hyperbolic tangent function. The evident connection between the transition functions and the h-form and g-form offers an interesting new perspective on the transformation approach of Kitazawa_JOE2021 for the AR(1) model. If we further define

align*[align* omitted — 307 chars of source]

the two moment functions of Kitazawa_JOE2021 for the AR(1) model write

align*[align* omitted — 374 chars of source]

which can be formulated in terms of our own moment functions as

align*[align* omitted — 252 chars of source]

Appendix Section (ref) provides detailed derivations for the mapping between our two approaches. This last result indicates that our moment conditions essentially match those of Kitazawa_JOE2021 when $T=3$. However, for $T\geq 4$, Proposition (ref) imply that there are further identifying moments than those based solely on $\hbar U_{it}$ and $\hbar \Upsilon_{it}$ for the AR(1) model. Interestingly, it turns out as we demonstrate in Appendix Section (ref) that our moment functions coincide exactly with those derived by honore2020moment for the special case $T=3$. \\ To the best of our knowledge, besides the AR(1) model and a few specific examples, the structure of moment conditions in models with arbitrary lag order is not fully understood in the literature. Building on Bonhomme_EM12, honore2020moment propose moment functions for the AR(2) model up to $T=4$ and the AR(3) model with $T=5$ but no results are offered beyond these special instances. Yet, this is of general interest not only to better understand the properties of DFEL models but also for practical modelling and estimation purposes. For example, card2005estimating argue in favor of using higher order logit specifications to better fit the behavior of a control group in the context of a welfare experiment. Relatedly, there are few results available for multivariate fixed effect models and existing methods developed for the scalar case are likely to be difficult to adapt in practice due to computational barriers. In the remaining sections, we show that our two-step approach addresses these issues by providing closed form expressions for the moment equality conditions of these more complex models.

Moment restrictions for the AR(\texorpdfstring{$p$}) logit model, \texorpdfstring{$p>1$}

Allowing for more than one lag is often desirable in empirical work to model persistent stochastic processes and to better fit the data (e.g, magnac2000subsidised on labour market histories, chay1999non and card2005estimating on welfare recipiency). To this end, we now discuss how to extend our identification scheme to general univariate autoregressive models. We consider

align[align omitted — 169 chars of source]

for known autoregressive order $p>1$ and vector of initial values $Y_{i}^{0}=(Y_{i-(p-1)},\ldots,Y_{i-1},Y_{i0})'\in \mathcal{Y}^{p}$, with $A_i\in \mathbb{R}$. Here, we let $\theta_0=(\gamma_0',\beta_0')'\in\mathbb{R}^{p+K_{x}}$. The corresponding transition probabilities are:

align*[align* omitted — 232 chars of source]

and there will be moment restrictions attached to each of the $2^p$ (non-redundant) transition probabilities. Before detailing the specifics of their construction, we enumerate the moment restrictions for this model as we did for the AR(1). This provides a way to ensure that we are not leaving any information on the table.

Impossibility results and number of moment restrictions when $p\geq 1$

Based on simulation evidence, honore2020moment conjectured that AR($p$) models possess $2^T-(T+p-1)2^p$ linearly independent moment conditions in panels of sufficient length. We prove this claim in Theorem (ref) and establish that no moment restrictions for the common parameters exist when $T\leq p+1$; that is with less than $2p+1$ periods of observations per individual. To introduce the result formally, it is again convenient to consider the conditional expectation operator mapping functions of histories $Y_i$ to their conditional expectation given $Y_{i}^0=y^0,X_{i}=x$ and the fixed effect, i.e

align*[align* omitted — 219 chars of source]

so that for any $y\in \mathcal{Y}^T$, $ \mathcal{E}_{y^0,x}^{(p)}\left[\mathds{1}\{.=y\}\right]$ yields the conditional likelihood of history $y$ for all possible values of $A_i$ in the AR($p$) model. That is,

align*[align* omitted — 262 chars of source]

Then the following result holds:

theoremConsider model ((ref)) with $T\geq 1$ and initial condition $y^0\in \mathcal{Y}^p$. Suppose that for any $t,s\in \{1,\ldots,T-1\}$ and $y,\tilde{y}\in \mathcal{Y}^{p}$, $\gamma_0'y+x_{t}'\beta_0\neq \gamma_0'\tilde{y}+x_{s}'\beta_0$ if $t\neq s$ or $y\neq\tilde{y}$. Then, the family \begin{align*} \mathcal{F}_{y^0,p,T}=\left\{1,\pi_{0}^{y_0|y^0}(.,x),\left\{\left(\pi_{t-1}^{y_{1}|y_{1}^{t-1},y_{0},\ldots,y_{-(p-t)}}(.,x))\right)_{y_{1}^{t-1}\in \mathcal{Y}^{t-1}}\right\}_{t=2}^p,\left\{\left(\pi_{t-1}^{y_1|y_{1}^{p}}(.,x)\right)_{y_{1}^p\in \mathcal{Y}^p}\right\}_{t=p+1}^T\right\} \end{align*} forms a basis of $\operatorname{Im}\left(\mathcal{E}_{y^0,x}^{(p)}\right)$ and therefore \begin{enumerate} • If $T\leq p+1$, $\rank\left(\mathcal{E}_{y^0,x}^{(p)}\right)=2^T$ and $\dim\left(\ker\left(\mathcal{E}_{y^0,x}^{(p)}\right)\right)=0$ • If $T\geq p+2$, $\rank\left(\mathcal{E}_{y^0,x}^{(p)}\right)=(T-p+1)2^{p}$ and $\dim\left(\ker\left(\mathcal{E}_{y^0,x}^{(p)}\right)\right)=2^T-(T-p+1)2^{p}$ \end{enumerate}

Theorem (ref) generalizes Theorem (ref) for AR($p$) logit models with $p>1$. It confirms the basic intuition that all the parametric content lies in the transition probabilities, no matter the lag order. Specifically, the conditional probabilities of all choice histories are spanned by the transition probabilities. In the basis $\mathcal{F}_{y^0,p,T}$, elements $\pi_{0}^{y_0|y^0}(.,x)$ and $\left\{\left(\pi_{t-1}^{y_{1}|y_{1}^{t-1},y_{0},\ldots,y_{-(p-t)}}(.,x))\right)_{y_{1}^{t-1}\in \mathcal{Y}^{t-1}}\right\}_{t=2}^p$ correspond to transition probabilities that are affected by the initial condition $y^0$. In the AR(1) case, it reduces to $\pi_{0}^{y_0|y_0}(.,x)$ (see Theorem (ref)). The remaining basis elements are free from the initial condition and correspond to the collection of all transition probabilities in each period starting from $t=p$. \\ Theorem (ref) is an implication of partial fraction decompositions and of the fact that the transition probabilities of AR($p$) models admit transition functions. This property is set out in the following section. If $T\leq p+1$, $\mathcal{E}_{y^0,x}^{(p)}$ is injective and no non-trivial moment conditions can be found. Beyond this threshold, the rank nullity theorem which connects image and nullspace of linear maps tells us that $2^T-(T-p+1)2^{p}$ moment restrictions exist. Under weaker conditions on the parameters or regressors then those of the theorem, the model may admit additional moment conditions even with $T\leq p+1$.

Construction of transition probabilities with $p>1$

Having clarified that $T=p+2$ is the minimum number of periods required for the existence of identifying moments, we are now ready to address the issue of their construction. The blueprint generalizes that of the AR(1) model and can be summarized as follows:

enumerate• Step 1) \begin{enumerate} • Start by obtaining analytical expressions of the unique transition functions for the transition probability in period $t=p$ when $T=p+1$ \footnote{The fact that the transition functions in period $t=p$ are unique when $T=p+1$ is a direct corollary of Theorem (ref). Otherwise, the difference of two distinct transition functions mapping to the same transition probability would yield a valid moment which is a contradiction.}. Shift these expressions by one period, two periods, three periods etc to get a set of transition functions for period $t\in\{p+1,\ldots,T-1\}$ when $T\geq p+2$. • Apply partial fraction decompositions to the expressions obtained in (a) for $t\in\{p+1,\ldots,T-1\}$ to generate other transition functions mapping to the same transition probabilities. \end{enumerate} • Step 2). Take \say{adequate} differences of transition functions associated to the same transition probability in periods $t\in\{p+1,\ldots,T-1\}$ to obtain valid moments that are linearly independent.

Step 1) (a) is akin to how we started by getting closed form expressions for the transition functions in period $t=1$ for $T=2$ in the one lag case and then deducted a general principle for $t\geq 2$ (see Section (ref)). From a technical perspective, this is the only part of the two-step procedure that differs from the baseline AR(1). Indeed, Step 2) is fundamentally identical and Step 1) (b) is also unchanged for the simple reason that the transition probabilities keep the same functional form as before. That is, a logistic transformation of a linear index composed of common parameters, the regressors and the fixed effect only. Hence, the same partial fraction expansions apply. In light of those close similarities with the AR(1) and in order to focus on the primary issues, we defer a discussion of Step 1)(b) and Step 2) to Appendix Section (ref). \\ Theorem (ref) provides the algorithm to compute the transition functions for \textbf{Step 1)} (a) for arbitrary lag order greater than one. It is based on the insight that we can leverage the transition functions of an AR($p-1$) and \textit{partial fraction decompositions} to generate the transition functions of an AR($p$). A simple example is helpful to illustrate those ideas. Consider an AR(2) with $T=3$ (i.e 5 observations in total) and suppose that we seek a transition function associated to, say, the transition probability

align*[align* omitted — 91 chars of source]

The first ingredient of the theorem is to view the AR(2) model as an AR(1) model where we treat the second order lag as an additional strictly exogenous regressor. This change of perspective is advantageous since we already know how to deal with the single lag case. In particular, Lemma (ref) readily gives the transition function $\phi_{\theta_0}^{0|0}(Y_{i3},Y_{i2},Y_{i1},Y_{i0},X_i)$ for the transition probability $\pi^{0|0,Y_{i1}}_{2}(A_i,X_{i})=P(Y_{i3}=0|Y_{i2}=0,Y_{i1},X_i,A_i)$ in the sense that it verifies:

align*[align* omitted — 150 chars of source]

This is an intermediate stage since $\phi_{\theta_0}^{0|0}(Y_{i3},Y_{i2},Y_{i1},Y_{i0},X_i)$ does not quite map to the target of interest; indeed $\pi^{0|0,Y_{i1}}_{2}(A_i,X_{i})$ depends on the random variable $Y_{i1}$ unlike $\pi^{0|0,1}_{2}(A_i,X_{i})$. To make further progress, one would intuitively need to \say{set} $Y_{i1}$ to unity to make the two transition probabilities coincide. We operationalize this idea by interacting $\phi_{\theta_0}^{0|0}(Y_{i3},Y_{i2},Y_{i1},Y_{i0},X_i)$ and $Y_{i1}$ to achieve the desired effect in expectation:

align*[align* omitted — 377 chars of source]

Here, the first equality follows from the law of iterated expectations. Then, the second ingredient of the theorem is a partial fraction expansion (Appendix Lemma (ref)) to turn this product of logistic indices into $\pi^{0|0,1}_{2}(A_i,X_{i})$. This last operation is analogous to how we constructed sequences of transition functions in the AR(1) model. It ultimately tells us that the solution is a weighted sum of $(1-Y_{i1})$ and $Y_{i1}\phi_{\theta_0}^{0|0}(Y_{i3},Y_{i2},Y_{i1},Y_{i0},X_i)$. Theorem (ref) turns this procedure into a recursive algorithm that computes the transition functions for any lag order $p>1$. \\

theoremIn model ((ref)) with $T\geq p+1$, for all $t\in\{p,\ldots,T-1\}$ and $ y_{1}^p\in \mathcal{Y}^p$ , let {\allowdisplaybreaks \begin{align*} &k^{y_1|y_{1}^{p}}_{t}(\theta)=\sum_{r=1}^{p}\gamma_{r}y_r+X_{it+1}'\beta \\ &k^{y_1|y_{1}^{k+1}}_{t}(\theta)=\sum_{r=1}^{k+1}\gamma_{r}y_r+\sum_{r=k+2}^{p} \gamma_{r}Y_{it-(r-1)}+X_{it+1}'\beta, \quad k=1,\ldots,p-2, if p>2 \\ &u_{t-k}(\theta)=\sum_{r=1}^{p} \gamma_{r}Y_{it-(r+k)}+X_{it-k}'\beta, \quad k=1,\ldots,p-1 \\ &w^{y_1|y_{1}^{k+1}}_t(\theta)=\left[1-e^{(k^{y_1|y_{1}^{k+1}}_{t}(\theta)-u_{t-k}(\theta))}\right]^{y_{k+1}}\left[1-e^{-(k^{y_1|y_{1}^{k+1}}_{t}(\theta)-u_{t-k}(\theta))}\right]^{1-y_{k+1}}, \quad k=1,\ldots,p-1 \end{align*} } and \begin{align*} &\phi_{\theta}^{y_1|y_{1}^{k+1}}(Y_{it+1},Y_{it},Y^{t-1}_{it-(p+k)},X_i)=\\ &\left[(1-Y_{it-k})+w^{y_1|y_{1}^{k+1}}_t(\theta)\phi_{\theta}^{y_1|y_{1}^{k}}(Y_{it+1},Y_{it},Y^{t-1}_{it-(p+k-1)},X_i)Y_{it-k}\right]^{(1-y_1)y_{k+1}}\times \\ &\left[1-Y_{it-k}-w^{y_1|y_{1}^{k+1}}_t(\theta)\left(1-\phi_{\theta}^{y_1|y_{1}^{k}}(Y_{it+1},Y_{it},Y^{t-1}_{it-(p+k-1)},X_i)\right)(1-Y_{it-k})\right]^{(1-y_1)(1-y_{k+1})}\times \\ &\left[Y_{it-k}+w^{y_1|y_{1}^{k+1}}_t(\theta)\phi_{\theta}^{y_1|y_{1}^{k}}(Y_{it+1},Y_{it},Y^{t-1}_{it-(p+k-1)},X_i)(1-Y_{it-k})\right]^{y_1(1-y_{k+1})}\times \\ &\left[1-(1-Y_{it-k})-w^{y_1|y_{1}^{k+1}}_t(\theta)\left(1-\phi_{\theta}^{y_1|y_{1}^{k}}(Y_{it+1},Y_{it},Y^{t-1}_{it-(p+k-1)},X_i)\right)Y_{it-k}\right]^{y_1y_{k+1}}, \quad k=1,\ldots,p-1 \end{align*} where \begin{align*} &\phi_{\theta}^{0|0}(Y_{it+1},Y_{it},Y_{it-p}^{t-1},X_i)=(1-Y_{it})e^{Y_{it+1}(\gamma_{1}Y_{it-1}-\sum_{l=2}^{p} \gamma_{l}\Delta Y_{it+1-l}-\Delta X_{it+1}'\beta)} \\ &\phi_{\theta}^{1|1}(Y_{it+1},Y_{it},Y^{t-1}_{it-p},X_i)=Y_{it}e^{ (1-Y_{it+1})\left(\gamma_{1}(1-Y_{it-1})+\sum_{l=2}^{p} \gamma_{l}\Delta Y_{it+1-l}+\Delta X_{it+1}'\beta\right)} \end{align*} Then, \begin{align*} &\mathbb{E}\left[\phi_{\theta_0}^{y_1|y_{1}^{p}}(Y_{it+1},Y_{it},Y^{t-1}_{it-(2p-1)},X_i)\,|\,Y_i^0,Y_{i1}^{t-p},X_i,A_i\right]=\pi^{y_1|y_{1}^{p}}_{t}(A_i,X_i) \end{align*} and for $k=0,\ldots,p-2$ \begin{align*} &\mathbb{E}\left[\phi_{\theta_0}^{y_1|y_{1}^{k+1}}(Y_{it+1},Y_{it},Y^{t-1}_{it-(p+k)},X_i)\,|\,Y_i^0,Y_{i1}^{t-(k+1)},X_i,A_i\right]=\pi^{y_1|y_{1}^{k+1},Y_{it-(k+1)},\ldots,Y_{it-(p-1)}}_{t}(A_i,X_i), \end{align*}

The remaining steps to complete the construction of valid moment functions are described at length in Appendix Section (ref). The end product is a family of (numerically) linearly independent moment functions of size $2^T-(T+1-p)2^{p}$. By Theorem (ref), this implies that our two-step approach recovers all moment equality conditions in the model.

remark(Extensions) While the exposition emphasized model ((ref)), our methodology applies more broadly to models of the form \begin{align*} Y_{it}=\mathds{1}\left\{g(Y_{it-1},\ldots,Y_{it-p},X_{it},\theta_0)+A_i-\epsilon_{it}\geq 0\right\}, \quad t= 1,\ldots, T \end{align*} where the lag order $p>1$ is known and $g(.)$ is known up to the finite dimensional parameter $\theta_0$. We can thus incorporate interaction effects which are often of interest in applied work. For instance, card2005estimating model welfare participation as a random effect AR(2) logit process of the form \begin{align*} Y_{it}=\mathds{1}\left\{\gamma_{01} Y_{it-1}+\gamma_{02} Y_{it-2}+\delta_{0} Y_{it-1}Y_{it-2}+X_{it}'\beta_0+A_i-\epsilon_{it}\geq 0\right\}, \quad t= 1,\ldots, T \end{align*} where $A_i$ either follows a normal distribution or a discrete distribution with few support points. In this case, minor modifications of the results in this section will deliver moment conditions for $\theta_0=(\gamma_{01},\gamma_{02},\delta_{0},\beta_0')'$ that are robust to misspecifications of individual unobserved heterogeneity. The key is that $A_i$ enters additivity in order to leverage the rational fraction identities of Lemma (ref).

Identification with more than one lag

This section discusses ways to leverage our methodology and moment restrictions to assess the identifiability of common parameters. For ease of exposition, we concentrate on the AR(2) logit model. \\ We start by briefly reexamining an identification result due to honore2020moment. Using functional differencing, they proved (under some regularity conditions) that $\theta_0$ is identified with $T=3$ provided $X_{i2}=X_{i3}$ and that the initial condition $Y_{i}^0=(Y_{i-1},Y_{i0})$ varies in the population. Notice that this is not in contradiction to Theorem (ref) since $X_{i2}=X_{i3}$ and $Y_{i}^0$ \say{varying} constitute two violations of its key assumptions. It is therefore not unsurprising that identifying moment exist in that case despite $T<4$. To understand why, note that imposing $X_{i2}=X_{i3}$ effectively amounts to equate the transition probabilities in period $t=2$ and in period $t=1$ for adequate choices of the initial condition; e.g $\pi_{1}^{0|0,Y_{i0}}(A_i,X_{i})=\pi_{2}^{0|0,0}(A_i,X_{i})$ provided that $Y_{i0}=0$ and $X_{i2}=X_{i3}$. In turn, this implies that differences of the corresponding transition functions in periods $t=2$ and $t=1$ deliver valid moment functions to estimate $\theta_0$ in certain subpopulations. In Appendix Section (ref), we show that this is an interpretation of the moment conditions that honore2020moment use to show point identification. \\ Because this identification argument hinges on matching covariates as in honore2000panel, it breaks down in the presence of certain types of regressors like an age variable or a time trend. In fact, dobronyi2021identification showed that there are actually no moment equality conditions available in the model with such regressors. This finding is consistent with the intuition that we cannot match the transition probabilities in periods $t=1$ and $t=2$ in that case. However, with one additional period, i.e $T=4$, we can leverage the moment restrictions of Proposition (ref) which are valid for free-varying regressors and any initial condition. This leads to two possible approaches to inference. The first is to consider the \say{identified set} $\Theta^{I}$ of $\theta_0$ based on the four conditional moment restrictions implied by the model:

align*[align* omitted — 214 chars of source]

and construct confidence sets for $\theta_0$ following e.g andrews2013inference. Instead, the sharp identified set may be computed following the approach of dobronyi2021identification if the covariates $X_i$ are discrete with finite support. Alternatively, a second approach which we develop further here is to formulate sensible restrictions on covariates that secure point identification in the spirit of honore2000panel. Specifically, we consider the case where a continuous scalar component $W_{i2}$ of $X_{i2}$ has unbounded positive support conditional on $Y_{i}^0$, the other regressors, $A_i$ and has a non-trivial effect $\beta_{0W}$ of known sign to the econometrician. This is the content of Assumption (ref) in which $Z_i=(R_{i}',W_{i1},W_{i3},W_{i4})$, and $X_{it}=(W_{it},R_{it}')\in \mathbb{R}^{K_x}$ for all $t\in \{1,2,3,4\}$. dobronyi2023revisiting used a similar device to develop an alternative distribution-free semiparametric estimator to that of honore2000panel that can accommodate time effects in the baseline one lag model.

assumption(i) The covariate $W_{i2}$ is continuously distributed with unbounded support on $\mathbb{R}_{+}$ conditional on $Y_{i}^0,Z_{i},A_i$ and (ii) $\beta_{0W}$ is known to be strictly negative.

Besides being a technical convenience, Assumption (ref) may be reasonable in some situations, e.g in the context of our empirical application, the econometrician may have a confident prior that drug prices affect individual drug consumption negatively. We point out that nothing in the discussion that follows hinges critically on $\beta_{W}<0$ and or $W_{i2}$ having support on the positive reals. A set of perfectly symmetric arguments will deliver the same conclusions if instead $\beta_{W}>0$ and $W_{i2}$ has unbounded support on $\mathbb{R}_{-}$.

assumption(i) $\theta_0=(\gamma_{01},\gamma_{02},\beta_0')'\in \mathbb{G}_1\times \mathbb{G}_2\times \mathbb{B}=\Theta$, $\mathbb{G}_1,\mathbb{G}_2,\mathbb{B}$ compact. The conditional densities of $A_i$ and $Z_i$ verify: \begin{enumerate}[label=(\roman*)] \setcounter{enumi}{1} • $\lim \limits_{w_2\to\infty}p(a|y^0,z,w_2)=q(a|y^0,z)$, $\lim \limits_{w_2\to\infty}p(z|y^0,w_2)=q(z|y^0)$ • There exists positive integrable functions $d_0(a),d_1(z),d_2(z)$ such that $p(a|y^{0},z,w_2)\leq d_{0}(a)$ for all $a\in \mathbb{R}$, $d_{1}(z) \leq p(z|y^0,w_2)\leq d_{2}(z)$ for all $z\in \mathbb{R}^{K_{x}-1}$$w_2\mapsto p(a|y^0,z,w_2), w_2\mapsto p(z|y^0,w_2) $ are continuous in $w_2$. \end{enumerate}

Assumption (ref) are standard regularity conditions for an application of the dominated convergence theorem that once paired with Assumption (ref) are sufficient to establish that $\theta_{0}$ is identified at infinity. The outline of the argument is as follows. Under these assumptions, by sending $W_{i2}$ to $\infty$, the valid moment function $\psi_{\theta}^{0|0,0}(Y_{i4},Y_{i3},Y^{2}_{i-1},X_i)$ of Proposition (ref) reduces to

align[align omitted — 441 chars of source]

which occurs because $\lim_{w_2\to\infty} e^{w_2\beta_W}=0$ and $Y_{i2}=0$ with probability one conditional on the regressors and the fixed effects. The key observation is that this \say{limiting} moment function has a similar functional form to the valid moment functions of the AR(1) model with $T=3$. In turn, this implies monotonicity properties on certain regions of the covariate space that we can exploit to point identify $\theta_0$ in the spirit of honore2020moment. To this end, let $(\bar{x},\underline{x})\in \mathbb{R}^2$, such that $\bar{x}>\underline{x}$ and define the sets

align*[align* omitted — 392 chars of source]

for all $k \in \{1,\ldots,K_{x}\}$. In words, $\mathcal{X}_{k,+}$ is the region of the covariate space in which values of the $k$-th regressor in periods $t\in\{1,3,4\}$ belong to $[\underline{x},\bar{x}]$ and verify $x_{k,3}\geq x_{k,4} \geq x_{k,1}$ with at least one strict inequality. Instead, $\mathcal{X}_{k,-}$ is the region of the covariate space where realizations of the $k$-th regressor obey the reverse ranking. With these notations in hands, we have the following theorem,

theoremFor $T=4$, suppose that outcomes $(Y_{i1},Y_{i2},Y_{i3},Y_{i4})$ are generated from model ((ref)) with $p=2$, initial condition $y^{0}\in\mathcal{Y}^2$, common parameters $\theta_0=(\gamma_{0}',\beta_0')\in \mathbb{R}^{2+K_{x}}$ and that Assumptions (ref) and (ref) hold. Further, for all $s\in\{-,+\}^{K_{x}}$, let $\mathcal{X}_s=\bigcap \limits _{k=1}^{K_{x}} \mathcal{X}_{k,s_k}$ and suppose that for all $y^{0}\in \mathcal{Y}^2$ \begin{align*} \lim_{w_{2}\to \infty} P\left(Y_{i}^0=y^{0},\quad X_i\in \mathcal{X}_s \,|\, W_{i2}=w_2 \right)>0 \end{align*} Let \begin{align*} \Psi_{s,y^0}^{0|0,0}(\theta) &=\lim_{w_{2}\to \infty} \mathbb{E}\left[ \psi_{\theta,\infty}^{0|0,0}(Y_{i4},Y_{i3},Y^{2}_{i-1},X_i)\,|\,Y_{i}^0=y^0,X_i\in \mathcal{X}_s,W_{i2}=w_2\right] \end{align*} Then, $\theta_0$ is the unique solution to the system of equations \begin{align*} \Psi_{s,y^0}^{0|0,0}(\theta)&=0, \quad \forall s\in\{-,+\}^{K_{x}},\quad \forall y^0\in \mathcal{Y}^2 \end{align*}

Theorem (ref) shows that point identification of $\theta_0$ is achievable in higher-order dynamic logit models in short panels. The main cost for this guarantee is Assumption (ref) which presumes knowledge of the data generating process beyond the baseline setup. Additionally, there should be sufficient variation in the regressors $X_{it}$ as $W_{i2}\mapsto \infty$ to ensure that $\lim_{w_{2}\to \infty} P\left(Y_{i}^0=y^{0},\quad X_i\in \mathcal{X}_s \,|\, W_{i2}=w_2 \right)>0$ for all $s\in\{-,+\}^{K_{x}}$. Our arguments are easily generalizable to AR($p$) models with lag order $p\geq 3$. Under natural extensions of Assumptions (ref) and (ref), the model parameters $\theta_0=(\gamma_{01},\ldots,\gamma_{0p},\beta_0')$ are identified at infinity provided $T\geq 2+p$.

remark[Identification with time effects] Theorem (ref) does not readily deals with time effects but it is straightforward to adapt the argument for this case. Suppose for concreteness that one covariate is a time trend. By further sending $W_{i3}$ to infinity, the limiting moment function of equation ((ref)) reduces to \begin{multline*} \psi_{\theta,\infty}^{0|0,0}(Y_{i4},Y_{i3},Y^{2}_{i-1},Z_i)=-(1-Y_{i1})(1-Y_{i2})(1-Y_{i3})Y_{i4} \\ +e^{-\gamma_1Y_{i0}-\gamma_{2}Y_{i-1}+X_{i41}'\beta}Y_{i1}(1-Y_{i2})(1-Y_{i3})(1-Y_{i4}) \end{multline*} For $(Y_{i0},Y_{i-1})=(0,0)$, this valid moment function only depends on $\beta$ and arguments analogous to those in Theorem (ref) will point identify $\beta_0$. Varying the initial condition is then sufficient to point identify $\gamma_0$ given the monotonicity of the moment function in $(\gamma_1,\gamma_2)$.

Average Marginal Effects in AR($p$) logit models

In discrete choice settings, interest often centers on certain functionals of unobserved heterogeneity rather than on the value of the model parameters per se. One particular family of such functionals that are of interest from a policy perspective are average marginal effects (AMEs) which capture mean response to a counterfactual change in past outcomes. It turns out that these key quantities are simply expectations of our transition functions. To see this, consider first the baseline AR(1) model with discrete covariates $X_{it}$. We can define the average transition probability from state $l$ to state $k$ in period $t$ for a subpopulation of individuals with covariate $x_{1}^{t+1}=(x_1,\ldots,x_{t+1})$ and initial condition $y_0$ as

align*[align* omitted — 238 chars of source]

where $p(a|y_0,x_{1}^{t+1})$ denotes the conditional density of the fixed effect $A$ given $(y_0,x_{1}^{t+1})$. The AME is defined as the following contrast of average transition probabilities:

align*[align* omitted — 173 chars of source]

It is interpreted as the population average causal effect on $Y_{it+1}$ of a change from 0 to 1 of $Y_{it}$ given $(y_0,x_{1}^{t+1})$. By Lemma (ref) and the law of iterated expectations, we have that for $T\geq 2$ and $t\geq 1$:

align*[align* omitted — 310 chars of source]

which implies that $AME_{t}(y_0,x_{1}^{t+1})$ is identified so long as $\theta_0$ is identified. A sufficient condition for that is $T\geq 3$ and $X_{i3}-X_{i2}$ having support in a neighborhood of zero (honore2000panel). aguirregabiria2021identification were the first to highlight that AMEs can be point identified in the AR(1) model. When the lag order $p$ is greater than one - which seems to be the case for persistent variables such as unemployment (e.g magnac2000subsidised) and welfare recipiency (e.g chay1999non) - we can analogously define average transition probabilities from states $l_1^p\in\mathcal{Y}^{p}$ to state $k\in\mathcal{Y}$ as:

align*[align* omitted — 254 chars of source]

This permits the consideration of more nuanced counterfactual parameters compared to the AR(1). In the context of studies on long term unemployment, contrasts of the form $\Pi_{t}^{k|l_{1}^p}(y^0,x_{1}^{t+1})-\Pi_{t}^{k|v_{1}^p}(y^0,x_{1}^{t+1})$ may be especially relevant to measure more accurately the relative effects of work histories spanning multiple periods. Again, these counterfactuals are simply expectations of transition functions by Theorem (ref) and will be identified whenever $\theta_0$ is identified (see Section (ref) for examples of sufficient conditions). \\ Multiperiod analogs of average transition probabilities in AR($p$) models

align*[align* omitted — 230 chars of source]

may also be of interest to assess state-dependence. These quantities give the average probability of moving from states $l_1^p\in\mathcal{Y}^{p}$ to future states $k_{1}^s\in \mathcal{Y}^s$, where $s\geq 1$ and the average is taken with respect to the distribution of $A_i$ conditional on $(y_0,x_{1}^{t+1})$. The special case $k_1=k_2=\ldots=k_{s}$ delivers a discrete version of the survivor function employed in duration analysis, i.e the average likelihood to survive $s$ consecutive periods in the same state after experiencing a given choice history. Proposition (ref) shows that they are also identified when $\theta_0$ is identified under certain conditions.

propositionConsider model ((ref)) with $T\geq p+2$, and initial condition $y^0\in \mathcal{Y}^p$. Suppose that $\theta_0$ is identified and that for any $t\in\{p,\ldots,T-2\}$, $s\in\{1,\ldots,T-1-t\}$ and $y,\tilde{y}\in \mathcal{Y}^{p}$, $\gamma_0'y+x_{t}'\beta_0\neq \gamma_0'\tilde{y}+x_{t+s}'\beta_0$ . Then, for $t\in\{p,\ldots,T-2\}$, $s\in\{1,\ldots,T-1-t\}$, and any $l_1^p\in\mathcal{Y}^{p}$, $k_{1}^s\in \mathcal{Y}^s$, the quantity $\Pi_{t}^{k_{1}^s|l_1^{p}}(y^0,x_{1}^{t+s})$ is identified.

The source of this result is the fact that the integrand of $\Pi_{t}^{k_{1}^s|l_1^{p}}(y^0,x_{1}^{t+s})$ is a product of transition probabilities. This entails that under appropriate conditions on the regressors and common parameters, we can turn this integrand into a unique linear combination of transition probabilities by means of a partial fraction decomposition. It is then a matter of taking expectations and invoking the fact that average transition probabilities are identified from our transition functions.

example[Survivor function for an AR(2)] To illustrate Proposition (ref), and in the spirit of our upcoming empirical application, suppose that $Y_{it}$ is an indicator for drug consumption at time $t$ obeying an AR(2) logit process. Fix $y^0\in \mathcal{Y}^2$ and assume $T=5$. One might be interested in \begin{align*} \Pi_{3}^{0,0|1,1}(y^0,x)&=\mathbb{E}\left[P(Y_{i5}=0,Y_{i4}=0\,|\,Y_{i3}=1,Y_{i2}=1,X_{i}=x,A_i)\,|\,Y_{i}^0=y^0,X_{i}=x\right] \\ &=\mathbb{E}\left[\pi_{4}^{0|0,1}(A_i,x)\pi_{3}^{0|1,1}(A_i,x)\,|\,Y_{i}^0=y^0,X_{i}=x\right] \end{align*} which gives the average propensity of individuals with characteristics $(y^0,x)$ who consumed drugs in $t=2,3$ to stay drug-free over the next two time periods. A simple calculation using for instance the identities of Appendix Lemma (ref) gives \begin{align*} \pi_{4}^{0|0,1}(A_i,x)\pi_{3}^{0|1,1}(A_i,x)&=\frac{1}{1+e^{\gamma_{02}+x_5'\beta_0+A_i}}\frac{1}{1+e^{\gamma_{01}+\gamma_{02}+x_4'\beta_0+A_i}} \\ &=\frac{1}{1-e^{\gamma_{01}+x_{45}'\beta_0}}\pi_{4}^{0|0,1}(A_i,x)-\frac{e^{\gamma_{01}+x_{45}'\beta_0}}{1-e^{\gamma_{01}+x_{45}'\beta_0}}\pi_{3}^{0|1,1}(A_i,x) \end{align*} and since Theorem (ref) implies $\mathbb{E}\left[\phi_{\theta_0}^{0|0,1}(Y_{i1}^5,x)\,|\,Y_i^0=y^0,Y_{i1}^{2},X_i=x,A_i\right]=\pi^{0|0,1}_{4}(A_i,x)$ and $\mathbb{E}\left[\phi_{\theta_0}^{0|1,1}(Y_{i0}^4,x)\,|\,Y_i^0=y^0,Y_{i1},X_i=x,A_i\right]=\pi^{0|0,1}_{3}(A_i,x)$, we obtain \begin{align*} \Pi_{3}^{0,0|1,1}(y^0,x)=\mathbb{E}\left[\frac{1}{1-e^{\gamma_{01}+x_{45}'\beta_0}}\phi_{\theta_0}^{0|0,1}(Y_{i1}^5,x)-\frac{e^{\gamma_{01}+x_{45}'\beta_0}}{1-e^{\gamma_{01}+x_{45}'\beta_0}}\phi_{\theta_0}^{0|1,1}(Y_{i0}^4,x)\,|\,Y_i^0=y^0,X_i=x\right] \end{align*}

Multi-dimensional fixed effects models

We now turn our attention to multi-dimensional fixed effects models. We show that the general blueprint developed in the scalar case to derive valid moment functions carries over to VAR(1) and MAR(1) models. We make no attempt at showing that our approach is exhaustive in those cases and do not claim that it is. We leave these important questions for future work. Readers uninterested in the details of the multivariate extensions can skip directly to Section (ref) where we discuss the empirical application.

Moment restrictions for the VAR(1) logit model

We begin with the analysis of VAR(1) logit models, variants of which have been successfully used to study the relationship between sickness and unemployment (narendranthan1985investigation), the progression from softer drug use to harder drug use among teenagers (deza2015there), transitivity in networks (graham2013comment, graham2016homophily) and more recently the employment of couples (honore2022simultaneity). For a given $M\geq 2$, the model reads:

align[align omitted — 206 chars of source]

We let $Y_{it}=(Y_{1,it},\ldots,Y_{M,it})'$ denote the outcome vector in period $t$ with support $\mathcal{Y}=\{0,1\}^M$ of cardinality $2^M$. We let $X_{it}=(X_{1,it}',\ldots ,X_{M,it}')'\in \mathbb{R}^{K_1}\times \ldots \times \mathbb{R}^{K_M}$ denote the vector of exogenous covariates in period $t$ and $A_i=(A_{1,i},\ldots,A_{M,i})'\in \mathbb{R}^M$ . The initial condition is now given by $Y_{i0}=(Y_{1,i0},\ldots,Y_{M,i0})'\in \mathcal{Y}$ and the model transition probabilities are given by:

align*[align* omitted — 232 chars of source]

for all $(k,l)\in \mathcal{Y}\times \mathcal{Y}$. \\ Building on honore2000panel, honore2019panel use a conditional likelihood approach to prove the identification $\theta_0=(\gamma_{011},\gamma_{012},\gamma_{021},\gamma_{022},\beta_{01},\beta_{02})$ for the bivariate specification when $T=3$ and the regressors do not vary over the last two periods. As in scalar models, we show hereinafter that this strong restriction which can yield undesirable rates of convergence is unnecessary to obtain valid moment conditions. \\ Step 1) in the VAR(1) logit model has a nuance relative to its scalar counterpart in that the only transition functions that appear to exist are those associated to $\pi^{k|k}_{t}(A_i,X_{i})$, for $k \in \mathcal{Y}$, i.e the probabilities of staying in the same state. We can use the same heuristic as in the baseline AR(1) model to derive their expressions, especially in the bivariate case. Once all four transition functions are obtained for the case $M=2$, it becomes clear that the general functional form is as per Lemma (ref). It is then a matter of brute force calculation to verify that this is indeed correct.

lemmaIn model ((ref)) with $T\geq 2$ and $t \in \{1,\ldots,T-1\}$, let for all $k \in \mathcal{Y}$ \begin{align*} \phi_{\theta}^{k|k}(Y_{it+1},Y_{it},Y_{it-1},X_i)&=\mathds{1}\{Y_{it}=k\}e^{\sum_{m=1}^M (Y_{m,it+1}-k_m)\left(\sum_{j=1}^M \gamma_{mj}(Y_{j,it-1}-k_j)-\Delta X_{m,it+1}'\beta_m\right)} \end{align*} Then: \begin{align*} \mathbb{E}\left[ \phi_{\theta_0}^{k|k}(Y_{it+1},Y_{it},Y_{it-1},X_i)|Y_{i0},Y_{i1}^{t-1},X_i,A_i\right]&= \pi^{k|k}_{t}(A_i,X_{i})=\prod_{m=1}^M \frac{e^{k_m(\sum_{j=1}^{M} \gamma_{0mj}k_{j} +X_{m,it+1}'\beta_{0m}+A_{m,i})}}{1+e^{\sum_{j=1}^{M} \gamma_{0mj}k_{j} +X_{m,it+1}'\beta_{0m}+A_{m,i}}} \end{align*}

Next, we can appeal to the second partial fraction decomposition formula in Appendix Lemma (ref) to guide the construction of another set of transition functions when $T\geq 3$. These identities may be regarded as a generalization of Kitazawa_JOE2021's hyperbolic transformations to the multivariate case. As is clear from Lemma (ref), the resulting transition functions have a special structure that generalizes those found in the AR(1) model.

lemmaIn model ((ref)) with $T\geq 3$, for all $t,s$ such that $T-1\geq t> s\geq 1$, let for all $ m\in \{1,\ldots,M\}$ and $(k,l)\in \mathcal{Y}^2$ \begin{align*} &\mu_{m,s}(\theta)=\sum_{j=1}^M \gamma_{mj} Y_{j,is-1}+X_{m,is}'\beta_{m} \\ &\kappa_{m,t}^{ k|k}(\theta)=\sum_{j=1}^M \gamma_{mj} k_j+X_{m,it+1}'\beta_m \\ &\omega_{t,s,l}^{k|k}(\theta)=1-e^{\sum_{j=1}^M (l_j-k_j)\left[\kappa_{j,t}^{ k|k}(\theta)-\mu_{j,s}(\theta)\right]} \end{align*} and define the moment functions \begin{align*} \zeta_{\theta}^{k|k}(Y_{it-1}^{t+1},Y_{is-1}^s,X_i)=\mathds{1}\{Y_{is}=k\}+ \sum\limits_{l\in \mathcal{Y}\setminus \{k\}} \omega_{t,s,l}^{k|k}(\theta) \mathds{1}\{Y_{is}=l\} \phi_{\theta}^{k|k}(Y_{it-1}^{t+1},X_i) \end{align*} Then, \begin{align*} \mathbb{E}\left[\zeta_{\theta_0}^{k|k}(Y_{it-1}^{t+1},Y_{is-1}^s,X_i)|Y_{i0},Y_{i1}^{s-1},X_i,A_i\right]&= \pi^{k|k}_{t}(A_i,X_{i}) \end{align*}

Beyond $T=4$, more transition functions are available and can be derived sequentially from those of Lemma (ref). See Corollary (ref) for their expressions.

corollaryIn model ((ref)) with $T\geq 4$, for any $t$ and ordered collection of indices $s_{1}^J$, $J\geq 2$, satisfying $T-1\geq t> s_1>\ldots>s_J\geq 1$, let for all $k\in \mathcal{Y}$ \begin{multline*} \zeta_{\theta}^{k|k}(Y_{it-1}^{t+1},Y_{is_1-1}^{s_1},\ldots,Y_{is_{J}-1}^{s_J},X_i)=\mathds{1}\{Y_{is_J}=k\} \\ +\sum_{l\in \mathcal{Y}\setminus \{k\}}\omega_{t,s_J,l}^{k|k}(\theta)\mathds{1}\{Y_{is_J}=l\} \zeta_{\theta}^{k|k}(Y_{it-1}^{t+1},Y_{is_1-1}^{s_1},\ldots,Y_{is_{J-1}-1}^{s_{J-1}},X_i) \end{multline*} with weights $\omega_{t,s_J,l}^{k|k}(\theta)$ defined as in Lemma (ref). Then, \begin{align*} \mathbb{E}\left[\zeta_{\theta_0}^{k|k}(Y_{it-1}^{t+1},Y_{is_1-1}^{s_1},\ldots,Y_{is_{J}-1}^{s_J},X_i)|Y_{i0},Y_{i1}^{s_{J}-1},X_i,A_i\right]=\pi^{k|k}_{t}(A_i,X_{i}) \end{align*}

Step 2). One can obtain a family of valid moment functions by adequately repurposing the statement of Proposition (ref) to the VAR(1) case, i.e by updating the expressions of $\phi_{\theta}^{k|k}(.)$ and $\zeta_{\theta}^{k|k}$ according to Lemma (ref) and Corollary (ref). To conserve on space and avoid repetition, we leave this simple exercise to the reader.

remark[Network Extension] Similarly to Remarks (ref), we emphasize that the tools developed here can be modified to handle other interesting variants featuring more complex interdependencies across the different layers of the model indexed by $m=1,\ldots,M$ . To illustrate the wider applicability of our two-step method, we show in Appendix (ref) how one can derive moment restrictions in the dynamic network formation model of graham2013comment and extensions thereof incorporating exogenous covariates.

Moment restrictions for the dynamic multinomial logit model

Last, we cover dynamic multinomial logit models which have been utilized to measure state-dependence in a range of economic contexts including: employment history in the French labor market (magnac2000subsidised), the impact of international trade on the transition matrix of employment across sectors (egger2003sectoral) and consumer product choice (dube2010state) amongst others. \\ We focus on the the baseline MAR(1) logit model with fixed effects.

The model assumes a fixed number of alternatives $C+1$ with $C\geq 1$ and is characterized by the following transition probabilities:

align[align omitted — 229 chars of source]

with $(k,l)\in \mathcal{Y}=\{0,1,\ldots,C\}$. Here, $Y_{it}\in \mathcal{Y}$ indicates the choice of individual $i$ in period $t$, $X_{ijt}$ denotes a vector of individual-alternative specific exogenous covariates and $A_{ij}\in \mathbb{R}$ is the fixed effect attached to alternative $j$ for individual $i$. The initial condition is $Y_{i0}\in \mathcal{Y}$ and in keeping with the fixed effect assumption, its conditional distribution given unobserved heterogeneity and the regressors, $\left(P(Y_{i0}=k|X_i,A_i)\right)_{k=1}^C$, is left fully unrestricted. Following magnac2000subsidised, we normalize the transition parameters and fixed effect of the reference alternative \say{$0$} to zero \footnote{ The transition parameters of the reference state cannot be identified so a normalization constraint must be imposed. Setting $A_{i0}=0$ is also without loss of generality since we can always redefine the fixed effect as $A_{ik}^*=A_{ik}-A_{i0}$.}. That is $\gamma_{j0}=\gamma_{0j}=0, A_{0,j}=0$ for all $j\in \mathcal{Y}$ leaving $\theta=\left((\gamma_{kl})_{k,l\geq 1},(\beta_{l})_{l\geq 0}\right)$ as the unknown model parameters.\\ This specification can be motivated by assuming that agents rank options according to random latent utility indices with disturbances independent over time and across alternatives. In this context, equation ((ref)) is obtained if the best alternative is selected and the error terms are Type 1 extreme value distributed conditional on $Y_{i0},A_i,X_i$. magnac2000subsidised studies the \say{pure} case without covariates and shows that an extension of the conditional likelihood approach proposed by chamberlain_1985 can be used to identify and estimate the state-dependence parameters. honore2000panel show that this argument carries over to the case with exogenous explanatory variables if one matches the regressors across specific time periods. Here, we offer an alternative estimation strategy that circumvents the need for matching. \\ Step 1). Similarly to the VAR(1) model the MAR(1) appears to admit transition functions only for the probabilities of staying in the same state, namely $\pi^{k|k}_{t}(A_i,X_i)$ for $k\in \mathcal{Y}$. This feature appears to be a common trait of multidimensional fixed effects specifications. To facilitate the derivation of the relevant transition functions, we follow our usual heuristic of looking for $\phi_{\theta}^{k|k}(.), k \in \mathcal{Y}$ satisfying:

align*[align* omitted — 272 chars of source]

Upon obtaining their exact expressions for the simplest case with $C=2$, it is easy to conjecture and verify by direct calculations that the general expressions of the $C+1$ transition functions of the MAR(1) model are as displayed in Lemma (ref).

lemmaIn model ((ref)) with $T\geq 2$ and $t \in \{1,\ldots,T-1\}$, let for all $k \in \mathcal{Y}$ \begin{align*} \phi_{\theta}^{k|k}(Y_{it-1}^{t+1},X_i)&=\mathds{1}\{Y_{it}=k\} e^{\sum_{c\in \mathcal{Y}\setminus\{k\}} \mathds{1}\{Y_{it+1}=c\} \left(\sum_{j\in \mathcal{Y}} (\gamma_{cj}-\gamma_{kj})\mathds{1}(Y_{it-1}=j)+\gamma_{kk}-\gamma_{ck}+\Delta X_{ikt+1}'\beta_k-\Delta X_{ict+1}'\beta_c\right)} \end{align*} Then: \begin{align*} \mathbb{E}\left[ \phi_{\theta_0}^{k|k}(Y_{it+1},Y_{it},Y_{it-1},X_i)|Y_{i0},Y_{i1}^{t-1},X_i,A_i\right]&= \pi^{k|k}_{t}(A_i,X_{i}) \end{align*}

Unsurprisingly, given the similarities shared between the MAR(1) and all other specifications discussed in the paper, so long as $T\geq 3$, one can again derive transition functions other than $\phi_{\theta}^{k|k}(Y_{it-1}^{t+1},X_i)$ also associated to $ \pi^{k|k}_{t}(A_i,X_i)$ for $k \in \mathcal{Y}$ in periods $t\in\{1,\ldots,T-1\}$. The simple logistic identities of Appendix Lemma (ref) imply that these transition functions, that we keep denoting $\zeta_{\theta}^{k|k}(.)$ have a similar form to those of the VAR(1) model as shown in Lemma (ref).

lemmaIn model ((ref)) with $T\geq 3$, for all $t,s$ such that $T-1\geq t> s\geq 1$, let for all $(c,k)\in \mathcal{Y}^2$ \begin{align*} \mu_{c,s}(\theta)&=\sum_{j=1}^C \gamma_{cj}\mathds{1}(Y_{is-1}=j)+X_{ics}'\beta_c-X_{i0s}'\beta_0 \\ \kappa_{c,t}^{k|k}(\theta)&=\gamma_{ck}+X_{ict+1}'\beta_c-X_{i0t+1}'\beta_0 \\ \omega_{t,s,c}^{k|k}(\theta)&=1-e^{(\kappa_{c,t}^{k|k}(\theta)-\mu_{c,s}(\theta))-( \kappa_{k,t}^{k|k}(\theta)-\mu_{k,s}(\theta))} \end{align*} and define the moment functions \begin{align*} \zeta_{\theta}^{k|k}(Y_{it-1}^{t+1},Y_{is-1}^s,X_i)=\mathds{1}\{Y_{is}=k\}+ \sum\limits_{l\in \mathcal{Y}\setminus \{k\}} \omega_{t,s,l}^{k|k}(\theta) \mathds{1}\{Y_{is}=l\} \phi_{\theta}^{k|k}(Y_{it-1}^{t+1},X_i) \end{align*} Then, \begin{align*} \mathbb{E}\left[\zeta_{\theta_0}^{k|k}(Y_{it-1}^{t+1},Y_{is-1}^s,X_i)|Y_{i0},Y_{i1}^{s-1},X_i,A_i\right]&= \pi^{k|k}_{t}(A_i,X_{i}) \end{align*}

Additionally, if the econometrician has access to a dataset with more than four observations per sampling unit - counting the initial condition - then, more transition functions associated to the same transition probabilities are available per Corollary (ref).

corollaryIn model ((ref)) with $T\geq 4$, for any $t$ and ordered collection of indices $s_{1}^J$, $J\geq 2$, satisfying $T-1\geq t> s_1>\ldots>s_J\geq 1$, let for all $k\in \mathcal{Y}$ \begin{multline*} \zeta_{\theta}^{k|k}(Y_{it-1}^{t+1},Y_{is_1-1}^{s_1},\ldots,Y_{is_{J}-1}^{s_J},X_i)=\mathds{1}\{Y_{is_J}=k\} \\ +\sum_{l\in \mathcal{Y}\setminus \{k\}}\omega_{t,s_J,l}^{k|k}(\theta)\mathds{1}\{Y_{is_J}=l\} \zeta_{\theta}^{k|k}(Y_{it-1}^{t+1},Y_{is_1-1}^{s_1},\ldots,Y_{is_{j-1}-1}^{s_{J-1}},X_i) \end{multline*} with weigts $\omega_{t,s_J,l}^{k|k}(\theta)$ defined as in Lemma (ref). Then, \begin{align*} \mathbb{E}\left[\zeta_{\theta_0}^{k|k}(Y_{it-1}^{t+1},Y_{is_1-1}^{s_1},\ldots,Y_{is_{J}-1}^{s_J},X_i)|Y_{i0},Y_{i1}^{s_{J}-1},X_i,A_i\right]=\pi^{k|k}_{t}(A_i,X_{i}) \end{align*}

This completes Step 1) for the MAR(1) logit model. For Step 2), we recommend a family of valid moment functions mirroring those of Proposition (ref) for the AR(1) case to ensure the linear independence of its elements.

Empirical Illustration

In this last section, we illustrate the usefulness of our methodology by revisiting the analysis of deza2015there on the dynamics of drug consumption amongst young adults in the United States.\footnote{This research was conducted with restricted access to Bureau of Labor Statistics (BLS) data. The views expressed here are those of the author and do not reflect the views of the BLS.} \\ To provide context, multiple studies have documented that young individuals who experiment with soft drugs have a tendency to continue using them and are at a higher risk of transitioning to hard drugs. Such correlations are certainly concerning. However, the empirical evidence of genuine causal links, in particular from softer drugs to harder drugs, remains limited with deza2015there standing as a notable exception. Fundamentally, these empirical regularities may be attributed to a causal effect (i.e. state dependence within and between drugs) or alternatively to latent traits that make individuals more prone to using illicit substances in general. Our primary concern is to untangle these two explanations to inform the design of policies aiming to mitigate drug addiction \footnote{See heckman1981heterogeneity for insights on the implications of state dependence for the design of labor market policies.}. For example, if marijuana consumption indeed serves as a gateway to later cocaine use, early educational interventions cautioning against casual marijuana usage could potentially have enduring effects on the population of heavy drug users. \\ To investigate these issues, we employ the restricted version of the National Longitudinal Survey of Youth 1997 (NLSY97). This is a panel dataset of 8984 individuals surveyed on a diverse range of subjects, including drug-related matters from 1997 to 2019. We concentrate on a subsample of four waves, spanning from 2001 to 2004. This subsample provides insight into the behavior of young adults between the age of 16 and 20 in 2001 to 19 and 24 in 2004. We shall examine the statistical association between three binary outcome variables, namely the consumption of alcohol, marijuana and hard drugs, derived from respondents answers' during annual interviews. Upon retaining those providing answers in all four waves as well as a valid state of residence, our cross section ultimately consists of $N=6317$ individuals \footnote{We adapt the sample selection procedure described in deza2015there for the period 2001-2004.}. Following deza2015there, we then consider the trivariate VAR(1) logit model

align*[align* omitted — 218 chars of source]

$m\in\{1,2,3\}$ ($1$=\say{alcohol}, $2$=\say{marijuana}, $3$=\say{hard drugs}), $t=1,2,3$ where $t=0$ corresponds to the year $2001$. The state-dependence coefficients $\gamma_{0mm}$ (within) and $\gamma_{0mj},m\neq j$ (between) are the principal coefficients of interest in the 16-dimensional vector of common parameters $\theta_{0}$. We are most particularly concerned about the sign and the statistical significance of $\gamma_{032}$, i.e the so called \say{stepping-stone} effect of marijuana on hard drugs. The covariate $age_{it}$ denotes the age of respondent $i$ at time $t$. The regressors $TEDS_{m,it}$ measure state-level deviations from national trends in treatment admissions for substance abuse caused by drug $m$ in year $t$ in the state of residence of $i$\footnote{ The variables $TEDS_{m,it}$ are constructed from the Treatment Episode Data Set-Admissions which records admissions to substance abuse treatment facilities in the United States.}. They are computed as the ratio of the share of admissions to treatment centers due to drug $m$ in the state of $i$ in year $t$ against the country wide analog in year $t$. Intuitively, this may be interpreted as a measure of exposure to substance $m$ for each respondent in our sample.\\ deza2015there parameterizes both the latent permanent heterogeneity $(A_{m,i})_{m=1}^3$ and the initial condition $Y_i^{0}$ to estimate the model by maximum likelihood. We leave these components unrestricted and exploit the valid moment functions presented in Section (ref). We specifically use six of the eight valid moment functions available: $\psi_{\theta}^{k|k}(Y_{i1}^{3},Y_{i0}^{1},X_i)$ for $k\in \{(0,0,0),(0,1,0),(1,1,1),(1,1,0),(1,0,1),(1,0,0)\}$. The other two corresponding to states $k\in\{(0,0,1),(0,1,1)\}$ are null for over $99.5\%$ of our sample and were dropped to mitigate noise in estimation. Next, we (arbitrarily) select a constant, the initial condition $Y_{i}^{0}$, $age_{it}$ and the covariates $TEDS_{m,it}$ in all periods $t=1,2,3$ as instruments to form the $96\times 1$ moment vector \\

align*[align* omitted — 651 chars of source]

With $m_{\theta}(Y_{i},Y_{i}^{0},X_i)$ in hand, we then consider the iterated GMM estimator of hansen1996finite. Starting from an initial candidate $\hat{\theta}_0$\footnote{In practice, we used the GMM estimator putting equal weights on each moment as our starting candidate.}, it can be described as

align*[align* omitted — 217 chars of source]

where $\overline{m}_{N}(\theta)=\frac{1}{N}\sum_{i=1}^N m_{\theta}(Y_{i},Y_{i}^{0},X_i)$ and $\overline{W}_{N}(\theta)=\frac{1}{N}\sum_{i=1}^N m_{\theta}(Y_{i},Y_{i}^{0},X_i)m_{\theta}(Y_{i},Y_{i}^{0},X_i)'$. Under some regularity conditions (hansen2021inference), this estimator is well defined and asymptotically normally distributed with

align*[align* omitted — 92 chars of source]

where $M_0 = \mathbb{E}\left[\pdv{m_{\theta_0}(Y_{i},Y_{i}^{0},X_i)}{\theta}\right]$ and $W_0=\mathbb{E}\left[m_{\theta_0}(Y_{i},Y_{i}^{0},X_i)m_{\theta_0}(Y_{i},Y_{i}^{0},X_i)'\right]$. Our motivation for focusing on this specific estimator originates mainly from hansen2021inference who advocate its use for two practical reasons. First, for a given set of moments, it eliminates the arbitrariness in the choice of the initial weight matrix of 2-step GMM estimators (see also imbens2002generalized). Second, because the iteration sequence is a contraction, each iteration is approximately variance reducing in the sense that: $Var(\hat{\theta}_{s}) \approx c^2 Var(\hat{\theta}_{s-1})$ for some constant $c<1$ \footnote{Note that the limiting variance of the iterated GMM estimator and a 2-step GMM estimator will be identical.}. Empirically, we also found in Monte Carlo simulations that the iterated GMM estimator performs relatively well for this type of specification (see Appendix (ref)). \\ Table (ref) presents the iterated GMM estimates for the trivariate VAR(1) logit model in columns (I), (II), (III). For comparison, columns (IV), (V), (VI) report a random effect (RE) estimator akin to deza2015there \footnote{We borrow the specification presented in deza2015there. The heterogeneity distribution is discrete with 3 mass points and is independent of the regressors. The initial condition relates to the covariates through a logistic regression.} while columns (VII), (VIII), (IV) display the \say{naive} logit maximum likelihood estimator (MLE) neglecting the presence of fixed effects. \\ The first observation is that, in line with conventional wisdom, GMM estimates for the state-dependence parameters within drug, $\gamma_{11},\gamma_{22},\gamma_{33}$, are all positive. As is apparent from columns (I)-(III), they are statistically significant for alcohol and marijuana but surprisingly not for hard drugs. In other words, there is no statistical evidence of a direct effect from past consumption of hard drug to future usage of hard drugs once we account for unobserved heterogeneity and the effects of other substances, at least in our four-wave sample\footnote{The transition parameters for hard drugs are expected to be noisier given that a smaller fraction of individuals consume these more lethal substances: approximately 15% of the respondents indicate having consumed hard drugs at least once from 2001-2004. This contrasts with 86% for alcohol and 40% for marijuana.}. Notice that the magnitude of the estimates for $\gamma_{11},\gamma_{22},\gamma_{33}$ sharply contrast with the other two estimators. The naive MLE largely overestimates the amount of within state-dependence, yielding coefficients that are comparatively four to eight times larger. Intuitively, this can be rationalized by the fact that this estimator misinterprets any serial correlation produced by $A_i$ as evidence of state dependence. The RE estimator borrowed from deza2015there (see also card2005estimating, chay1998identification) acts as an intermediate estimator between the other two as can be seen in columns (IV)-(VI). This behavior is expected to the extent that the additional parametric structure of this methodology will account to some degree for the presence of unobserved heterogeneity. We note that the role of within state dependence in the dynamics of drug consumption is nevertheless overstated by this approach.

Second and importantly, we observe in column (III) a positive and statistically significant effect of marijuana on hard drugs. This supports the view that marijuana usage can be a gateway to the consumption of harder drugs and accords with the key findings of deza2015there. From a practical standpoint, this result corroborates that there may be scope for policies on marijuana usage to indirectly curb the consumption of more lethal substances by teenagers and young adults. The efficacy of such policies in the short and long run are important questions that will intuitively depend on the distribution of heterogeneity in the population. We do not explore those questions here but further research in this direction would be of interest \footnote{A natural idea to gauge the effectiveness of policy interventions would be to compute average marginal effects. However, as mentioned in Section (ref), we were unable to find transition functions for the transition probabilities where the state switches in VAR(1) models. This leads us to believe that only the average transition probabilities where the state remains unchanged are identified. In turn, this would imply that average marginal effects are generally partially identified in VAR(1) models. In this case, it is possible that ideas analogous to those in dobronyi2021identification and davezies2021identification could be used to characterize and compute the identified set of average marginal effects; albeit some difficulties might arise due to the fact that the fixed effects are now multidimensional. Computing outer bounds as in pakel2023bounds could be another plausible option.}. The other two estimators also agree on a positive influence of marijuana on the consumption of harder drugs, albeit it is statistically insignificant in the RE case.

table[table omitted — 2,135 chars of source]

Otherwise, it is noteworthy that the between state dependence estimates can vary quite significantly across specifications. Again, the naive MLE likely misinterprets spurious correlation from the $A_i$ as state dependence which results in positive and inflated cross effects. Column (IV) and (I) show disagreements of the RE and GMM estimates regarding the strength of the impact of marijuana and hard drugs on alcohol. Overall, this comparative exercise has showed that accounting for unobserved heterogeneity as flexibly as possible can be essential to obtain an accurate picture of the patterns of state dependence in practice.

Conclusion

Dynamic discrete choice models are widely used to study the determinants of repeated decisions made by individuals or firms over time. In this paper, we have introduced a procedure to estimate a family of such models with logistic (or Type I extreme value) errors and potentially many lags while remaining agnostic about the nature of unobserved individual heterogeneity. This type of approach may be attractive when the risk of misspecifying the initial condition and the unit-specific effects are important. We also provided general expressions for average marginal effects in the binary response case which are often the counterfactuals of interest in practice. \\ The list of discrete choice models covered in this paper is of course not exhaustive and it would be interesting to know if our two-step approach could be deployed in other settings with \say{logit} noise. In ongoing work, we have found that this is one avenue to approach estimation of dynamic ordered logit models, potentially of arbitrary lag order.

\nocite{*}

\setcounter{page}{0} \pagenumbering{arabic} \setcounter{page}{1}