Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
120,683 characters · 20 sections · 125 citation commands
Transition Probabilities and Moment Restrictions in Dynamic Fixed Effects Logit Models
\setcitestyle{authoryear,open={(},close={)}}
\noindentKeywords: dynamic discrete choice, panel data, fixed effects.
\noindentJEL Classification Codes: C23, C33.
The analysis of state dependence is a classic and important topic in many areas of economics. Several discrete processes such as welfare and labor force participation manifest strong serial persistence, and economists have sought various methods to unravel the underlying factors. In this paper, we reexamine the estimation of one notable set of models employed for this purpose: discrete choice models with lagged dependent variables, strictly exogenous regressors, fixed effects and logistic errors. We shall refer to this class of models as dynamic fixed effects logit models (DFEL) throughout. Specifications of this kind are used to discriminate between \say{structural} state dependence, i.e the causal effect of past choices on current outcomes, and heterogeneity, i.e the serial correlation induced by unobserved individual attributes (heckman1981heterogeneity). An example of this approach is the analysis of welfare participation in chay1999non. There has been considerable interest in this family of panel data models in econometrics, with a recent surge in attention following new developments reported in honore2020moment. One general reason is that DFEL models stand out as a rare case of nonlinear dynamic panel data models for which solutions to the incidental parameters problem (neyman1948consistent) and initial conditions problem (e.g heckman1981heterogeneity) have been known to exist in short panels\footnote{The incidental parameters problem refers to the general inconsistency of maximum likelihood in short panels. The initial conditions problem refers to the general difficulty of formulating a correct conditional distribution for the initial observations given the fixed effects and covariates.}. \\ In the \say{pure} version of the basic model which abstracts from covariates other than a first order lag, cox1958regression, chamberlain_1985 and magnac2000subsidised showed that the autoregressive parameter can be consistently estimated by conditional likelihood. This approach relies on the existence of a sufficient statistic linked to the logistic assumption to eliminate the fixed effect. In an important subsequent paper, honore2000panel extended this idea to a setting with strictly exogenous regressors and showed that the conditional likelihood approach remains viable if one can further condition on the regressors being equal in specific periods. This strategy was also found to be effective in dynamic multinomial logit models (honore2000panel), panel vector autoregressions (honore2019panel) and dynamic ordered logit models (muris2020dynamic. At the same time, it has also been noted that the necessity to be able to \say{match} the covariates imposes two limitations for the conditional likelihood approach: it inherently rules out time effects and implies rates of convergence slower than $\sqrt{N}$ for continuous explanatory variables. Furthermore, calculations from honore2000panel suggested that it does not easily extend to models with a higher lag order. These shortcomings have motivated the search for alternative methods of estimation. \\ Recently, kitazawa2013exploration,kitazawa2016root and Kitazawa_JOE2021 revisited the AR(1) logit model - autoregressive of order one - of honore2000panel and proposed a transformation approach that deals with the fixed effects without restricting the nature of the covariates besides the conventional assumption of strict exogeneity. Their methodology leads to moment restrictions that can serve as a basis to estimate the model parameters at $\sqrt{N}$-rate by GMM; even with continuous regressors. In parallel work, honore2020moment also derived moment conditions for the AR(1), AR(2) and AR(3) logit models in panels of specific length using the functional differencing technique of Bonhomme_EM12. Their approach is partly numerical and relies on symbolic computing (e.g Mathematica) to obtain analytical expressions but has a wider scope of potential applications, e.g dynamic ordered logit specifications (honore2021dynamic). In another recent paper, dobronyi2021identification, the authors analyze the full likelihood of AR(1) and AR(2) logit models with discrete covariates under a new angle that reveals a connection to the truncated moment problem in mathematics. Drawing on well established results in that literature, they derive moment equality and new moment inequality restrictions that fully characterize the sharp identified set. \\ In this paper, we introduce a new systematic approach to construct moment restrictions in DFEL models with additive fixed effects, i.e when fixed effects are heterogeneous \say{intercepts}. This class of models encompasses most specifications studied in prior work but excludes models with heterogeneous coefficients on lagged outcomes and/or regressors as in chamberlain_1985 and browning2014dynamic. Unlike some recent competing approaches, we do not require numerical experimentation nor symbolic computing. Rather, as we shall see in examples, we exploit the common structure of logit-type transition probabilities and elementary properties of rational fractions, to obtain analytic expressions for the identifying moments. We shall focus our attention on deriving valid moment functions for AR($p$) models with arbitrary lag order $p\geq 1$ as well as first-order panel vector autoregressions and dynamic multinomial logit models (magnac2000subsidised). \\ Our methodology exploits two key observations. First, the transition probabilities of logit-type models can often be expressed as conditional expectations of functions of observables and common parameters given the initial condition, the regressors and the fixed effects. We shall refer to these moment functions as transition functions. They have the important feature of not depending on individual fixed effects. Second, as soon as $T\geq p+2$, where $T$ denotes the number of observations post initial condition, many transition probabilities in periods $t\in \{p+1,\ldots,T-1\}$ admit at least two distinct transition functions. The combination of these two features motivates a two-step approach to obtain moment restrictions in panels of adequate length. In the first step, we shall compute the model transition functions. Then, the second step will simply consist in differencing two transition functions associated to the same transition probability. We show that a careful application of this procedure delivers all the moment equality restrictions available in the binary response case. We shall further elaborate on these steps in examples and use the resulting moment functions to derive new identification results. At a high level, the approach we advocate in this paper consists in solving a sequence of problems with identical structure period by period instead of solving directly a large system of equations based on the model full likelihood as in honore2020moment and dobronyi2021identification. As a consequence, our procedure remains tractable when the number of time periods increases and in models with higher order lags. \\ Besides the aforementioned papers, our work also connects to a line of research studying the identification of features of the distribution of fixed effects in discrete choice models. One branch in this literature has focused on developing general optimization tools to compute sharp numerical bounds on average marginal effects. This includes most notably the linear programming method of honore2006bounds, recently adapted by bonhomme2023identification to the case of sequentially exogenous covariates, and the quadratic programming method of chernozhukov2013average. A second branch in this literature has sought instead to harness the specificities of logit models to obtain simple analytical bounds. In static logit models, davezies2021identification exploit mathematical results on the moment problem to formulate sharp bounds on the average partial effects of regressors on outcomes. In DFEL models, aguirregabiria2021identification are the first to prove the point identification of average marginal effects in the baseline AR(1) logit model when $T\geq 3$. In related work, dobronyi2021identification make use of their moment equality and moment inequality restrictions to establish sharp bounds on functionals of the fixed effects such as average marginal effects and average posterior means in AR(1) and AR(2) specifications. We complement these results as a byproduct of our methodology: average marginal effects and their variants in AR($p$) models, with arbitrary $p\geq 1$ are merely differences of average transition functions. \\ The remainder of the paper is organized as follows. Section (ref) presents the setting and our main objective. Section (ref) introduces some terminology and gives an outline of our procedure to construct moment restrictions. Section (ref) implements our approach in AR($p$) logit models with $p\geq 1$ and discusses identification of model parameters and average marginal effects. The semiparametric efficiency bound for the AR(1) is also presented for the base case of four waves of data. Section (ref) discusses extensions to the VAR(1) and the dynamic multinomial logit model with one lag, MAR(1) for short. In Section (ref), we present an empirical illustration on the dynamics of drug consumption amongst young people and Section (ref) offers concluding remarks. A complementary set of Monte Carlo simulations showing the small sample performance of GMM estimators based on our moment restrictions is available in Appendix Section (ref). Proofs are gathered in the Appendix.
Let $i=1,\ldots,N$ denote a population index and $t=0,\ldots,T$ be an index for time. We study DFEL models which may be viewed as threshold-crossing econometric specifications describing a discrete outcome $Y_{it}$ through a latent index involving lagged outcomes (e.g $Y_{it-1}$), strictly exogenous regressors $X_{it}$, an individual-specific time-invariant unobservable $A_i$ and an error term $\epsilon_{it}$. The canonical example is the AR(1) model:
and we shall concentrate more broadly on cases where $A_i$ is additively separable from the other explanatory variables. An initial condition that we will generically denote $Y_{i}^{0}$ completes such models to enable dynamics. The common parameter $\theta_0$ is one target of interest and governs the influence of lagged outcomes and the regressors on the contemporaneous outcome. Other quantities of interest include counterfactual parameters such as average marginal effects.\\ Throughout, we leave the joint distribution of $(Y_{i}^{0},X_i,A_i)$ unrestricted where $X_i=(X_{i1},\ldots,X_{iT})$ and thus refer to $A_i$ as a fixed effect in common with the literature. The schocks $\epsilon_{it}$ are assumed to be serially independent logistically distributed, independent of $(Y_{i}^{0},X_i,A_i)$, except for the MAR(1) model where they are instead extreme value distributed. Finally, we shall assume that $(Y_i,Y_{i}^{0},X_i,A_i)$ are jointly i.i.d across individuals.
The data available to the econometrician consists of the initial condition $Y_{i}^0$, the outcome vector $Y_{i}=(Y_{i1},\ldots,Y_{iT})$, and the covariates $X_i$ for all $N$ individuals. Interest centers primarily on the identification and estimation of $\theta_0$ in short panels, i.e for fixed $T$. To this end, the chief objective of this paper is to show how to construct moment functions $\psi_{\theta}(Y_i,Y_{i}^0,X_{i})$ free of the fixed effect parameter that are valid in the sense that:
When this is possible, the law of iterated expectations implies the conditional moment: $$ \mathbb{E}\left[\psi_{\theta_0}(Y_i,Y_{i}^{0},X_{i})\,|\,Y_{i}^0,X_{i}\right]=0$$ which can in turn be leveraged to assess the identifiability of $\theta_0$ and form the basis of a GMM estimation strategy. This is the central idea underlying functional differencing (Bonhomme_EM12) and was applied by honore2020moment to derive valid moment conditions for a class of dynamic logit models with scalar fixed effects. We borrow the same insight but instead of searching for solutions numerically on a case-by-case basis, we propose a complementary systematic algebraic procedure to recover the model's valid moments \footnote{dobronyi2021identification and Kitazawa_JOE2021 also have an algebraic approach but our methodologies are very different. The first paper uses the full likelihood of the model and focuses on the AR(1) and special instances of the AR(2) model. The second paper has a transformation approach adapted to the AR(1) model. Our emphasis here is primarily on developing an approach that is tractable for a large class of models.}. In doing so, we flesh out the mechanics implied by the logistic assumption which in turn suggest a blueprint to deal with estimation of general DFEL models. For example, we are able to characterize the expressions of valid moment functions in AR($p$) models for arbitrary $p>1$ which to the best of our knowledge is a new result in the literature. Furthermore, our approach carries over to multidimensional fixed effect specifications: VAR(1), dynamic network formation models and the MAR(1) in which searching for moments numerically is cumbersome or intractable. \\ In what follows, we shall use the shorthand $Y_{it_{1}}^{t_2}=(Y_{it_1},\ldots,Y_{it_2})$ to denote a collection of random variables over periods $t_1$ to $t_2$ with the convention that $Y_{it_{1}}^{t_2}=\emptyset$ if $t_1>t_2$. Likewise, we may use the notation $y_{t_1}^{t_2}=(y_{t_1},\ldots,y_{t_2})$ to denote any $(t_2-t_1)$-dimensional vector of reals with the convention $y_{t_{1}}^{t_{2}}=\emptyset$ for $t_1>t_2$. Elements $1_{n}$ and $0_{n}$ shall refer to the $n$-dimensional vectors of ones and zeros respectively. The support of the outcome variable $Y_{it}$ shall be denoted $\mathcal{Y}$. We let $\Delta$ denote the first-differencing operator so that $\Delta Z_{it}=Z_{it}-Z_{it-1}$ for any random variable $Z_{it}$ and make use of the notation $Z_{its}=Z_{it}-Z_{is}$ for $s\neq t$ to accommodate long differences. We use $\mathds{1}\{.\}$ for the indicator function; $\operatorname{Im}(f)$, $\ker(f)$, $\rank(f)$ to denote the image, the nullspace and the rank of a linear map $f$.
Let $T\geq 1$. Given an initial condition $y^0\in \mathcal{Y}^{p}$, $p\geq 1$ being the lag order of the model, and strictly exogenous regressors $X_i\in \mathbb{R}^{K_{x}\times T}$, we denote the (one-period ahead) transition probability in period $t\geq 1$ from state $(l_{1}^t,y^ 0)\in \mathcal{Y}^{t}\times \mathcal{Y}^{p}$ to state $k\in \mathcal{Y}$ as:
With $p$ lags, the markovian nature of the models considered in this paper imply that $ \pi^{k|l_{1}^t,y^0}_{t}(A_i,X_{i})$ will not depend on the entire path of past outcomes but only on the value of the most recent $p$ outcomes. For instance, in an AR(1) model where $p=1$, we have:
and thus we will suppress the dependence on $(y^0,l_1,\ldots,l_{t-1})$ and write $\pi^{k|l_t}_{t}(A_i,X_{i})$. We shall proceed analogously for the more general case $p\geq 1$. \\ We call a transition function associated to a transition probability $ \pi^{k|l_{1}^t,y^0}_{t}(A_i,X_{i})$ any moment function $\phi^{k|l_{1}^t,y^0}_{\theta}(Y_i,Y_{i}^{0},X_i)$ of the data and the common parameters verifying:
With these notions in hand, we are ready to describe our two-step approach to derive valid moment functions in the sense of equation ((ref)). In Step 1), we begin by computing the model's transition functions. Our procedure requires a minimum of $T=p+1$ periods of observations to accommodate arbitrary regressors and initial condition. In this case, we can get analytical formulas for the transition functions associated to the transition probabilities in period $t=p$ and Theorem (ref) and Theorem (ref) below imply that they are unique. However, this is not immediately helpful to get moment (equality) restrictions on $\theta_0$. We require one more period. As soon as $T\geq p+2$, we explain how to construct distinct transition functions associated to the same transition probabilities in periods $t\in\{p+1,\ldots,T-1\}$. The key ingredient is the use of partial fraction decompositions for rational fractions adapted to the structure of the transition probabilities. It is then a matter of taking differences of two transition functions associated to the same transition probability to obtain valid moment functions; we refer to this last step as Step 2). The ensuing sections demonstrate this procedure in scalar and multidimensional fixed effect models.
For exposition, we begin with the baseline AR(1) logit model with fixed effects introduced above:
Here, $\mathcal{Y}=\{0,1\}$, $\theta_0=(\gamma_0,\beta_0')\in \mathbb{R}\times \mathbb{R}^{K_{x}}$, the initial condition $Y_{i}^0$ consists of the binary-valued random variable $Y_{i0}$ and $A_i\in \mathbb{R}$.
We start out by enumerating the moment restrictions implied by the model. This will provide a means to assess the exhaustiveness of our approach. To this end, let $\mathcal{E}_{y_0,x}$ denote the conditional expectation operator mapping any function of the outcome variable $Y_i$ to its conditional expectation given $Y_{i0}=y_0,X_{i}=x$ and the fixed effect $A_i$, i.e
For example, for any $y\in \mathcal{Y}^T$, $\mathcal{E}_{y_0,x}\left[\mathds{1}\{.=y\}\right]$ yields the conditional probability of observing history $y$ for all possible values of the fixed effect, i.e:
where $P(Y_i=y|Y_{i0}=y_0,X_{i}=x,A_{i}=a)=\prod \limits_{t=1}^T \frac{e^{y_t(\gamma_0y_{t-1}+x_t'\beta_0+a)}}{1+e^{\gamma_0y_{t-1}+x_t'\beta_0+a}}, \quad \forall a \in \mathbb{R}$ . Then, we have the following result,
Theorem (ref) formalizes the intuition that the transition probabilities summarize the parametric component of the model: $2^T$ histories are possible yet only $2T$ basis elements are necessary to fully characterize their conditional probabilities. This follows from the observation that when the covariate index \footnote{We refer to the quantity $\gamma_0y_{t-1}+x_{t}'\beta_0$ for a given period $t$.} of each transition probability differ, the conditional probability of each history $y\in \mathcal{Y}^{T}$ is a ratio of polynomials in $e^{a}$, where the numerator has lower degree than the denominator, and the later is a product of distinct irreducible terms. A sufficient condition for this is that $\gamma_0\neq 0$ and that one regressor is continuously distributed with non-zero slope. In turn, standard results on partial fraction decompositions ensure that this ratio can be expressed as a unique linear combination of transition probabilities. To finally conclude that $\mathcal{F}_{y_0,T}$ is a basis of $\operatorname{Im}(\mathcal{E}_{y_0,x})$, we leverage upcoming results demonstrating that the transition probabilities live in $\operatorname{Im}(\mathcal{E}_{y_0,x})$ as expectations of transition functions. \\ Importantly, since $\ker(\mathcal{E}_{y_0,x})$ is the set of valid moment functions verifying equation ((ref)), Theorem (ref) tells us that the AR(1) model features $2^T-2T$ linearly independent moment restrictions in general. This is a consequence of the rank nullity theorem for linear maps with finite dimensional domains. The fact that $2^T-2T$ moment conditions are available for the AR(1) appeared initially as a conjecture in honore2020moment and was later established by dobronyi2021identification using different arguments from here. They do not emphasize the role of the transition probabilities. Our ideas extend naturally to the case of arbitrary lags which was hitherto an open problem. We discuss this extension in Section (ref).
Having clarified the total count of moment restrictions in the AR(1) logit model, we next discuss how to construct them with our two-step procedure.
In the absence of exogenous regressors, model ((ref)) simplifies to:
which was first introduced by cox1958regression and then revisited in chamberlain_1985, magnac2000subsidised. These papers established the identification of $\gamma_0$ for $T\geq 3$ via conditional likelihood based on the insight that $(Y_{i0},\sum_{t=1}^{T-1} Y_{it},Y_{iT})$ are sufficient statistics for the fixed effect. Our methodology is conceptually different as we seek to directly construct moment functions verifying equation ((ref)). \\ For what follows, it is helpful to remember that the individual-specific transition probability from state $l$ to state $k$ is time-invariant and given by:
Step 1). We shall begin by deriving the transition functions for $\pi^{0|0}(A_i)$ and $\pi^{1|1}(A_i)$. Observe that $\pi^{1|0}(A_i)$ and $\pi^{0|1}(A_i)$ are effectively redundant since probabilities sum to one. A natural starting place is to investigate the case $T=2$, i.e 2 periods of observations after the initial condition. Recalling definition ((ref)), we search for $\phi_{\theta}^{0|0}(Y_{i2},Y_{i1},Y_{i0})$, respectively $\phi_{\theta}^{1|1}(Y_{i2},Y_{i1},Y_{i0})$, whose conditional expectation given $(Y_{i0},A_i)$ yields $\pi^{0|0}(A_i)$, respectively $\pi^{1|1}(A_i)$. For the purposes of illustration and to show the kind of calculations arising broadly in DFEL models, let us derive $\phi_{\theta}^{0|0}(Y_{i2},Y_{i1},Y_{i0})$. By Bayes's rule:
where the second equality uses the logistic hypothesis. By quick inspection, we see that the terms in the first parenthesis have $(1+e^{\gamma_0+a})$ in their denominator unlike $\pi^{0|0}(A_i)$. Because $-e^{-\gamma_0}$ is not a pole of $\pi^{0|0}(A_i)$\footnote{A pole of a rational function is a root of its denominator. Formally, we are substituting $u=e^{a}$ and we are extending $\pi^{0|0}(u)$ to the real line.}, we conclude that $\phi_{\theta}^{0|0}(1,1,y_{0})=\phi_{\theta}^{0|0}(0,1,y_{0})=0$. This first deduction leaves us with
Now, since $\pi^{0|0}(A_i)$ does not depend on $y_0$, we must cancel the denominator $(1+e^{\gamma_0 y_{0}+a})$. To achieve this, we must set: $\phi_{\theta_0}^{0|0}(1,0,y_{t-1})=C_0 e^{\gamma_0 y_{0}}, \phi_{\theta_0}^{0|0}(0,0,y_{t-1})=C_0$ for some constant $C_0\in \mathbb{R}\setminus\{0\}$. Then,
and $C_0=1$ is the appropriate normalization to obtain the desired transition function. Of course, the exact same logic applies for $\phi_{\theta_0}^{1|1}(Y_{i2},Y_{i1},Y_{i0})$ and $\pi^{1|1}(A_i)$. \\ This short calculation provides a useful recipe for the general case $T\geq 2$. We learned that we can search for functions of three consecutive outcomes $\phi_{\theta}^{k|k}(Y_{it+1},Y_{it},Y_{it-1})$ such that:
The first restriction is a functional form that eliminates terms with inadequate poles after taking expectations. The second restriction is a normalization condition to match the desired transition probability. Following this argument, we arrive at the expressions in Lemma (ref).
Step 2). The second step in the agenda is the construction of valid moment functions. Because the transition probability of the model are time-invariant, one trivial way to achieve this is to consider the pairwise difference of $\phi_{\theta}^{k|k}(Y_{it+1},Y_{it},Y_{it-1})$ and $\phi_{\theta}^{k|k}(Y_{is+1},Y_{is},Y_{is-1})$ for any feasible $s\neq t$. This is the content of Proposition (ref). We will need a minimum of four total periods of observations, which coincides with the requirements of the conditional likelihood approach.
In this subsection, we move on to the AR(1) logit model with strictly exogenous covariates characterized by equation ((ref)). \\ Step 1). We employ the same shortcut recipe as in the \say{pure} case and begin by looking for moment functions $\phi_{\theta}^{0|0}(.)$ and $\phi_{\theta}^{1|1}(.)$ verifying:
where this time
The same simple calculations described just above lead to the expressions in Lemma (ref). The only (expected) change is the appearance of a new term $+/-\Delta X_{it+1}'\beta$ which accounts for the presence of covariates in the model.
At this point, it is important to highlight that unlike previously, the transition probabilities are covariate-dependent. The upshot is that the naive difference of $\phi_{\theta}^{k|k}(Y_{it+1},Y_{it},Y_{it-1},X_i)$ and $\phi_{\theta}^{k|k}(Y_{is+1},Y_{is},Y_{is-1},X_i)$ for $s \neq t$ no longer leads to valid moment functions in general. Indeed, while Lemma (ref) ensures that
clearly, $\pi_{t}^{k|k}(A_i,X_i)-\pi_{s}^{k|k}(A_i,X_i)\neq 0$ when $X_{it+1}'\beta_0\neq X_{is+1}'\beta_0$ \footnote{A matching strategy in the spirit of honore2000panel may still be applicable when in our example $X_{it+1}= X_{is+1}$. However, this is known to lead to estimators converging at rate less than $\sqrt{N}$ for continuous covariates and it rules out certain regressors such as time dummies and time trends.}. Thus, a different logic is required in the presence of explanatory variables other than a first order lag. \\ The key, as foreshadowed in Section (ref) is that as soon as $T\geq 3$, it is possible to construct transition functions other than $\phi_{\theta}^{k|k}(Y_{it-1}^{t+1},X_i)$ also mapping to $\pi^{k|k}_{t}(A_i,X_{i})$ in time periods $t\in\{2,\ldots,T-1\}$. These new transition functions that we denote $\zeta_{\theta}^{k|k}(.)$ to emphasize their difference have a particular form. They consist of a weighted combination of past outcome $\mathds{1}(Y_{is}=k)$, $1 \leq s<t$, and the interaction of $\mathds{1}(Y_{is}\neq k)$ with any transition function associated to $\pi^{k|k}_{t}(A_i,X_{i})$ having no dependence on outcomes prior to period $s$, e.g $\phi_{\theta}^{k|k}(Y_{it-1}^{t+1},X_i)$. This property follows from a partial fraction decomposition presented in Lemma (ref) that exploits the structure of the model probabilities under the logistic assumption. It relates to the hyperbolic transformations ideas of Kitazawa_JOE2021. In the sequel, we shall see that this insight carries over to the AR($p$) logit model with $p>1$. Lemma (ref) below gives the \say{simplest} additional transition functions that one can construct when $T\geq 3$ for the AR(1) model with exogenous regressors (the only ones when $T=3$).
When $T\geq 4$, it turns out that we can build even more transition functions from those given in Lemma (ref) by repeating the same type of logic based on partial fraction expansions; Corollary (ref) provides a recursive formulation.
Step 2). Provided $T\geq 3$, the difference between any transition functions associated to the same transition probabilities in periods $t\in\{2,\ldots,T-1\}$ constitutes a valid candidate for ((ref)). One particularly relevant set of valid moment functions for reasons explained below is presented in Proposition (ref).
This family of moment functions has cardinality $2^T-2T$ which by Theorem (ref) is precisely the number of linearly independent moment conditions available for the AR(1). To see this, notice that for fixed $(k,Y_{i0})\in \mathcal{Y}^2$, and a given time period $ t\in\{ 2,\ldots,T-1\}$, Proposition (ref) gives a total of:
valid moment functions. This follows from a simple counting argument. First, we get $\binom{t-1}{1}$ possibilities from choosing any $s$ in $\{1,\ldots,t-1\}$ to form $\psi_{\theta}^{k|k}(Y_{it-1}^{t+1},Y_{is-1}^s,X_i)$. To that, we must add another $\sum_{l=2}^{t-1} \binom{t-1}{l}$ possibilities from choosing all feasible sequences $s_1^J$ with $t-1\geq s_1>s_2>\ldots>s_J\geq 1$ to form $\psi_{\theta}^{k|k}(Y_{it-1}^{t+1},Y_{is_1-1}^{s_1},\ldots,Y_{is_{J}-1}^{s_J},X_i)$. Summing over $t=2,\ldots, T-1$ and multiplying by 2 to account for the two possible values for $k$ delivers the result:
Furthermore, there is evidence that the family is linearly independent. It is readily verified for $T=3$ since the two valid moment functions produced by the model depend on two distinct sets of choice histories. This can be seen from their unpacked expressions in equations ((ref)) and ((ref)) in the Appendix. Unfortunately, this argument does not carry over to longer panels but we have verified numerically that the linear independence property of this family continues to hold for several different values of $T\geq 4$. This suggests that our approach delivers all the moment equality restrictions available in the AR(1) model with $T$ periods post initial condition \footnote{This is not all the identifying content of the AR(1) specification since we know from dobronyi2021identification that the model also implies moment inequality conditions.}.
honore2020moment gave sufficient conditions to identify $\theta_0=(\gamma_0,\beta_0')'$ in the AR(1) model with $T=3$. A natural follow-up question is to ask how accurately can $\theta_0$ be estimated in that case, or equivalently what is the semi-parametric information bound. In a corrigendum to hahn2001information, gu2023information confirmed that the conditional likelihood estimator is semiparametrically efficient for $T=3$ in the \say{pure} AR(1) model. However, the characterization of the semiparametric efficiency bound and the question of what estimator attains it remain unclear with covariates. \\ To answer these questions, let $\psi_{\theta}(Y_{i1}^{3},Y_{i0}^1,X_i)=(\psi_{\theta}^{0|0}(Y_{i1}^{3},Y_{i0}^1,X_i),\psi_{\theta}^{1|1}(Y_{i1}^{3},Y_{i0}^1,X_i))'$ where the two components correspond to the valid moment functions of Proposition (ref) for $T=3$. Additionally, let $D(X_i,y_0)=\mathbb{E}\left[\pdv{ \psi_{\theta_0}(Y_{i1}^{3},Y_{i0}^1,X_i)}{\theta'}|Y_{i0}=y_0,X_i\right]$ and let $\Sigma(X_i,y_0)=\mathbb{E}\left[\psi_{\theta_0}(Y_{i1}^{3},Y_{i0}^1,X_i)\psi_{\theta_0}(Y_{i1}^{3},Y_{i0}^1,X_i)'|Y_{i0}=y_0,X_i\right]$.
With these notations in hand and under the mild conditions of Assumption (ref), Theorem (ref) clarifies that the efficient score coincides with the efficient moment for the conditional moment problem: $\mathbb{E}\left[\psi_{\theta}(Y_{i1}^{3},Y_{i0}^1,X_i)|Y_{i0}=y_0,X_i\right]=0$. Put differently, the maximal efficiency with which $\theta_0$ can be estimated is $V_{0}(y_0)=\mathbb{E}[D(X_i,y_0)\Sigma(X_i,y_0)^{-1}D(X_i,y_0)'|Y_{i0}=y_0]^{-1}$. This result is in accordance with Remark (ref) which noted that the score of the conditional likelihood without covariates is precisely the efficient moment implied by our conditional moment restrictions in this case.
The proof of Theorem (ref) only involves careful bookkeeping of some tedious algebra and an application of Theorem 3.2 in newey1990semiparametric. Interestingly, davezies2023fixed presented analogous results in the static panel data case with three periods of observations.
As indicated previously, there is a connection between our methodology and that of Kitazawa_JOE2021 for the AR(1) model. Indeed, after some algebraic manipulation, we can re-express the transition functions of Lemma (ref) (or Lemma (ref) without covariates) as:
where $\delta=(e^{\gamma}-1)$. Thus, the moment conditions of Lemma (ref) imply that we can write:
where $\mathbb{E}\left[\epsilon_{it}^{0|0}|Y_{i0},Y_{i1}^{t-1},X_i,A_i\right]=0$ and $\mathbb{E}\left[\epsilon_{it}^{1|1}|Y_{i0},Y_{i1}^{t-1},X_i,A_i\right]=0$. These expressions are the so-called h-form and g-form of Kitazawa_JOE2021 for model ((ref)) and were originally obtained through an ingenious usage of the mathematical properties of the hyperbolic tangent function. The evident connection between the transition functions and the h-form and g-form offers an interesting new perspective on the transformation approach of Kitazawa_JOE2021 for the AR(1) model. If we further define
the two moment functions of Kitazawa_JOE2021 for the AR(1) model write
which can be formulated in terms of our own moment functions as
Appendix Section (ref) provides detailed derivations for the mapping between our two approaches. This last result indicates that our moment conditions essentially match those of Kitazawa_JOE2021 when $T=3$. However, for $T\geq 4$, Proposition (ref) imply that there are further identifying moments than those based solely on $\hbar U_{it}$ and $\hbar \Upsilon_{it}$ for the AR(1) model. Interestingly, it turns out as we demonstrate in Appendix Section (ref) that our moment functions coincide exactly with those derived by honore2020moment for the special case $T=3$. \\ To the best of our knowledge, besides the AR(1) model and a few specific examples, the structure of moment conditions in models with arbitrary lag order is not fully understood in the literature. Building on Bonhomme_EM12, honore2020moment propose moment functions for the AR(2) model up to $T=4$ and the AR(3) model with $T=5$ but no results are offered beyond these special instances. Yet, this is of general interest not only to better understand the properties of DFEL models but also for practical modelling and estimation purposes. For example, card2005estimating argue in favor of using higher order logit specifications to better fit the behavior of a control group in the context of a welfare experiment. Relatedly, there are few results available for multivariate fixed effect models and existing methods developed for the scalar case are likely to be difficult to adapt in practice due to computational barriers. In the remaining sections, we show that our two-step approach addresses these issues by providing closed form expressions for the moment equality conditions of these more complex models.
Allowing for more than one lag is often desirable in empirical work to model persistent stochastic processes and to better fit the data (e.g, magnac2000subsidised on labour market histories, chay1999non and card2005estimating on welfare recipiency). To this end, we now discuss how to extend our identification scheme to general univariate autoregressive models. We consider
for known autoregressive order $p>1$ and vector of initial values $Y_{i}^{0}=(Y_{i-(p-1)},\ldots,Y_{i-1},Y_{i0})'\in \mathcal{Y}^{p}$, with $A_i\in \mathbb{R}$. Here, we let $\theta_0=(\gamma_0',\beta_0')'\in\mathbb{R}^{p+K_{x}}$. The corresponding transition probabilities are:
and there will be moment restrictions attached to each of the $2^p$ (non-redundant) transition probabilities. Before detailing the specifics of their construction, we enumerate the moment restrictions for this model as we did for the AR(1). This provides a way to ensure that we are not leaving any information on the table.
Based on simulation evidence, honore2020moment conjectured that AR($p$) models possess $2^T-(T+p-1)2^p$ linearly independent moment conditions in panels of sufficient length. We prove this claim in Theorem (ref) and establish that no moment restrictions for the common parameters exist when $T\leq p+1$; that is with less than $2p+1$ periods of observations per individual. To introduce the result formally, it is again convenient to consider the conditional expectation operator mapping functions of histories $Y_i$ to their conditional expectation given $Y_{i}^0=y^0,X_{i}=x$ and the fixed effect, i.e
so that for any $y\in \mathcal{Y}^T$, $ \mathcal{E}_{y^0,x}^{(p)}\left[\mathds{1}\{.=y\}\right]$ yields the conditional likelihood of history $y$ for all possible values of $A_i$ in the AR($p$) model. That is,
Then the following result holds:
Theorem (ref) generalizes Theorem (ref) for AR($p$) logit models with $p>1$. It confirms the basic intuition that all the parametric content lies in the transition probabilities, no matter the lag order. Specifically, the conditional probabilities of all choice histories are spanned by the transition probabilities. In the basis $\mathcal{F}_{y^0,p,T}$, elements $\pi_{0}^{y_0|y^0}(.,x)$ and $\left\{\left(\pi_{t-1}^{y_{1}|y_{1}^{t-1},y_{0},\ldots,y_{-(p-t)}}(.,x))\right)_{y_{1}^{t-1}\in \mathcal{Y}^{t-1}}\right\}_{t=2}^p$ correspond to transition probabilities that are affected by the initial condition $y^0$. In the AR(1) case, it reduces to $\pi_{0}^{y_0|y_0}(.,x)$ (see Theorem (ref)). The remaining basis elements are free from the initial condition and correspond to the collection of all transition probabilities in each period starting from $t=p$. \\ Theorem (ref) is an implication of partial fraction decompositions and of the fact that the transition probabilities of AR($p$) models admit transition functions. This property is set out in the following section. If $T\leq p+1$, $\mathcal{E}_{y^0,x}^{(p)}$ is injective and no non-trivial moment conditions can be found. Beyond this threshold, the rank nullity theorem which connects image and nullspace of linear maps tells us that $2^T-(T-p+1)2^{p}$ moment restrictions exist. Under weaker conditions on the parameters or regressors then those of the theorem, the model may admit additional moment conditions even with $T\leq p+1$.
Having clarified that $T=p+2$ is the minimum number of periods required for the existence of identifying moments, we are now ready to address the issue of their construction. The blueprint generalizes that of the AR(1) model and can be summarized as follows:
Step 1) (a) is akin to how we started by getting closed form expressions for the transition functions in period $t=1$ for $T=2$ in the one lag case and then deducted a general principle for $t\geq 2$ (see Section (ref)). From a technical perspective, this is the only part of the two-step procedure that differs from the baseline AR(1). Indeed, Step 2) is fundamentally identical and Step 1) (b) is also unchanged for the simple reason that the transition probabilities keep the same functional form as before. That is, a logistic transformation of a linear index composed of common parameters, the regressors and the fixed effect only. Hence, the same partial fraction expansions apply. In light of those close similarities with the AR(1) and in order to focus on the primary issues, we defer a discussion of Step 1)(b) and Step 2) to Appendix Section (ref). \\ Theorem (ref) provides the algorithm to compute the transition functions for \textbf{Step 1)} (a) for arbitrary lag order greater than one. It is based on the insight that we can leverage the transition functions of an AR($p-1$) and \textit{partial fraction decompositions} to generate the transition functions of an AR($p$). A simple example is helpful to illustrate those ideas. Consider an AR(2) with $T=3$ (i.e 5 observations in total) and suppose that we seek a transition function associated to, say, the transition probability
The first ingredient of the theorem is to view the AR(2) model as an AR(1) model where we treat the second order lag as an additional strictly exogenous regressor. This change of perspective is advantageous since we already know how to deal with the single lag case. In particular, Lemma (ref) readily gives the transition function $\phi_{\theta_0}^{0|0}(Y_{i3},Y_{i2},Y_{i1},Y_{i0},X_i)$ for the transition probability $\pi^{0|0,Y_{i1}}_{2}(A_i,X_{i})=P(Y_{i3}=0|Y_{i2}=0,Y_{i1},X_i,A_i)$ in the sense that it verifies:
This is an intermediate stage since $\phi_{\theta_0}^{0|0}(Y_{i3},Y_{i2},Y_{i1},Y_{i0},X_i)$ does not quite map to the target of interest; indeed $\pi^{0|0,Y_{i1}}_{2}(A_i,X_{i})$ depends on the random variable $Y_{i1}$ unlike $\pi^{0|0,1}_{2}(A_i,X_{i})$. To make further progress, one would intuitively need to \say{set} $Y_{i1}$ to unity to make the two transition probabilities coincide. We operationalize this idea by interacting $\phi_{\theta_0}^{0|0}(Y_{i3},Y_{i2},Y_{i1},Y_{i0},X_i)$ and $Y_{i1}$ to achieve the desired effect in expectation:
Here, the first equality follows from the law of iterated expectations. Then, the second ingredient of the theorem is a partial fraction expansion (Appendix Lemma (ref)) to turn this product of logistic indices into $\pi^{0|0,1}_{2}(A_i,X_{i})$. This last operation is analogous to how we constructed sequences of transition functions in the AR(1) model. It ultimately tells us that the solution is a weighted sum of $(1-Y_{i1})$ and $Y_{i1}\phi_{\theta_0}^{0|0}(Y_{i3},Y_{i2},Y_{i1},Y_{i0},X_i)$. Theorem (ref) turns this procedure into a recursive algorithm that computes the transition functions for any lag order $p>1$. \\
The remaining steps to complete the construction of valid moment functions are described at length in Appendix Section (ref). The end product is a family of (numerically) linearly independent moment functions of size $2^T-(T+1-p)2^{p}$. By Theorem (ref), this implies that our two-step approach recovers all moment equality conditions in the model.
This section discusses ways to leverage our methodology and moment restrictions to assess the identifiability of common parameters. For ease of exposition, we concentrate on the AR(2) logit model. \\ We start by briefly reexamining an identification result due to honore2020moment. Using functional differencing, they proved (under some regularity conditions) that $\theta_0$ is identified with $T=3$ provided $X_{i2}=X_{i3}$ and that the initial condition $Y_{i}^0=(Y_{i-1},Y_{i0})$ varies in the population. Notice that this is not in contradiction to Theorem (ref) since $X_{i2}=X_{i3}$ and $Y_{i}^0$ \say{varying} constitute two violations of its key assumptions. It is therefore not unsurprising that identifying moment exist in that case despite $T<4$. To understand why, note that imposing $X_{i2}=X_{i3}$ effectively amounts to equate the transition probabilities in period $t=2$ and in period $t=1$ for adequate choices of the initial condition; e.g $\pi_{1}^{0|0,Y_{i0}}(A_i,X_{i})=\pi_{2}^{0|0,0}(A_i,X_{i})$ provided that $Y_{i0}=0$ and $X_{i2}=X_{i3}$. In turn, this implies that differences of the corresponding transition functions in periods $t=2$ and $t=1$ deliver valid moment functions to estimate $\theta_0$ in certain subpopulations. In Appendix Section (ref), we show that this is an interpretation of the moment conditions that honore2020moment use to show point identification. \\ Because this identification argument hinges on matching covariates as in honore2000panel, it breaks down in the presence of certain types of regressors like an age variable or a time trend. In fact, dobronyi2021identification showed that there are actually no moment equality conditions available in the model with such regressors. This finding is consistent with the intuition that we cannot match the transition probabilities in periods $t=1$ and $t=2$ in that case. However, with one additional period, i.e $T=4$, we can leverage the moment restrictions of Proposition (ref) which are valid for free-varying regressors and any initial condition. This leads to two possible approaches to inference. The first is to consider the \say{identified set} $\Theta^{I}$ of $\theta_0$ based on the four conditional moment restrictions implied by the model:
and construct confidence sets for $\theta_0$ following e.g andrews2013inference. Instead, the sharp identified set may be computed following the approach of dobronyi2021identification if the covariates $X_i$ are discrete with finite support. Alternatively, a second approach which we develop further here is to formulate sensible restrictions on covariates that secure point identification in the spirit of honore2000panel. Specifically, we consider the case where a continuous scalar component $W_{i2}$ of $X_{i2}$ has unbounded positive support conditional on $Y_{i}^0$, the other regressors, $A_i$ and has a non-trivial effect $\beta_{0W}$ of known sign to the econometrician. This is the content of Assumption (ref) in which $Z_i=(R_{i}',W_{i1},W_{i3},W_{i4})$, and $X_{it}=(W_{it},R_{it}')\in \mathbb{R}^{K_x}$ for all $t\in \{1,2,3,4\}$. dobronyi2023revisiting used a similar device to develop an alternative distribution-free semiparametric estimator to that of honore2000panel that can accommodate time effects in the baseline one lag model.
Besides being a technical convenience, Assumption (ref) may be reasonable in some situations, e.g in the context of our empirical application, the econometrician may have a confident prior that drug prices affect individual drug consumption negatively. We point out that nothing in the discussion that follows hinges critically on $\beta_{W}<0$ and or $W_{i2}$ having support on the positive reals. A set of perfectly symmetric arguments will deliver the same conclusions if instead $\beta_{W}>0$ and $W_{i2}$ has unbounded support on $\mathbb{R}_{-}$.
Assumption (ref) are standard regularity conditions for an application of the dominated convergence theorem that once paired with Assumption (ref) are sufficient to establish that $\theta_{0}$ is identified at infinity. The outline of the argument is as follows. Under these assumptions, by sending $W_{i2}$ to $\infty$, the valid moment function $\psi_{\theta}^{0|0,0}(Y_{i4},Y_{i3},Y^{2}_{i-1},X_i)$ of Proposition (ref) reduces to
which occurs because $\lim_{w_2\to\infty} e^{w_2\beta_W}=0$ and $Y_{i2}=0$ with probability one conditional on the regressors and the fixed effects. The key observation is that this \say{limiting} moment function has a similar functional form to the valid moment functions of the AR(1) model with $T=3$. In turn, this implies monotonicity properties on certain regions of the covariate space that we can exploit to point identify $\theta_0$ in the spirit of honore2020moment. To this end, let $(\bar{x},\underline{x})\in \mathbb{R}^2$, such that $\bar{x}>\underline{x}$ and define the sets
for all $k \in \{1,\ldots,K_{x}\}$. In words, $\mathcal{X}_{k,+}$ is the region of the covariate space in which values of the $k$-th regressor in periods $t\in\{1,3,4\}$ belong to $[\underline{x},\bar{x}]$ and verify $x_{k,3}\geq x_{k,4} \geq x_{k,1}$ with at least one strict inequality. Instead, $\mathcal{X}_{k,-}$ is the region of the covariate space where realizations of the $k$-th regressor obey the reverse ranking. With these notations in hands, we have the following theorem,
Theorem (ref) shows that point identification of $\theta_0$ is achievable in higher-order dynamic logit models in short panels. The main cost for this guarantee is Assumption (ref) which presumes knowledge of the data generating process beyond the baseline setup. Additionally, there should be sufficient variation in the regressors $X_{it}$ as $W_{i2}\mapsto \infty$ to ensure that $\lim_{w_{2}\to \infty} P\left(Y_{i}^0=y^{0},\quad X_i\in \mathcal{X}_s \,|\, W_{i2}=w_2 \right)>0$ for all $s\in\{-,+\}^{K_{x}}$. Our arguments are easily generalizable to AR($p$) models with lag order $p\geq 3$. Under natural extensions of Assumptions (ref) and (ref), the model parameters $\theta_0=(\gamma_{01},\ldots,\gamma_{0p},\beta_0')$ are identified at infinity provided $T\geq 2+p$.
In discrete choice settings, interest often centers on certain functionals of unobserved heterogeneity rather than on the value of the model parameters per se. One particular family of such functionals that are of interest from a policy perspective are average marginal effects (AMEs) which capture mean response to a counterfactual change in past outcomes. It turns out that these key quantities are simply expectations of our transition functions. To see this, consider first the baseline AR(1) model with discrete covariates $X_{it}$. We can define the average transition probability from state $l$ to state $k$ in period $t$ for a subpopulation of individuals with covariate $x_{1}^{t+1}=(x_1,\ldots,x_{t+1})$ and initial condition $y_0$ as
where $p(a|y_0,x_{1}^{t+1})$ denotes the conditional density of the fixed effect $A$ given $(y_0,x_{1}^{t+1})$. The AME is defined as the following contrast of average transition probabilities:
It is interpreted as the population average causal effect on $Y_{it+1}$ of a change from 0 to 1 of $Y_{it}$ given $(y_0,x_{1}^{t+1})$. By Lemma (ref) and the law of iterated expectations, we have that for $T\geq 2$ and $t\geq 1$:
which implies that $AME_{t}(y_0,x_{1}^{t+1})$ is identified so long as $\theta_0$ is identified. A sufficient condition for that is $T\geq 3$ and $X_{i3}-X_{i2}$ having support in a neighborhood of zero (honore2000panel). aguirregabiria2021identification were the first to highlight that AMEs can be point identified in the AR(1) model. When the lag order $p$ is greater than one - which seems to be the case for persistent variables such as unemployment (e.g magnac2000subsidised) and welfare recipiency (e.g chay1999non) - we can analogously define average transition probabilities from states $l_1^p\in\mathcal{Y}^{p}$ to state $k\in\mathcal{Y}$ as:
This permits the consideration of more nuanced counterfactual parameters compared to the AR(1). In the context of studies on long term unemployment, contrasts of the form $\Pi_{t}^{k|l_{1}^p}(y^0,x_{1}^{t+1})-\Pi_{t}^{k|v_{1}^p}(y^0,x_{1}^{t+1})$ may be especially relevant to measure more accurately the relative effects of work histories spanning multiple periods. Again, these counterfactuals are simply expectations of transition functions by Theorem (ref) and will be identified whenever $\theta_0$ is identified (see Section (ref) for examples of sufficient conditions). \\ Multiperiod analogs of average transition probabilities in AR($p$) models
may also be of interest to assess state-dependence. These quantities give the average probability of moving from states $l_1^p\in\mathcal{Y}^{p}$ to future states $k_{1}^s\in \mathcal{Y}^s$, where $s\geq 1$ and the average is taken with respect to the distribution of $A_i$ conditional on $(y_0,x_{1}^{t+1})$. The special case $k_1=k_2=\ldots=k_{s}$ delivers a discrete version of the survivor function employed in duration analysis, i.e the average likelihood to survive $s$ consecutive periods in the same state after experiencing a given choice history. Proposition (ref) shows that they are also identified when $\theta_0$ is identified under certain conditions.
The source of this result is the fact that the integrand of $\Pi_{t}^{k_{1}^s|l_1^{p}}(y^0,x_{1}^{t+s})$ is a product of transition probabilities. This entails that under appropriate conditions on the regressors and common parameters, we can turn this integrand into a unique linear combination of transition probabilities by means of a partial fraction decomposition. It is then a matter of taking expectations and invoking the fact that average transition probabilities are identified from our transition functions.
We now turn our attention to multi-dimensional fixed effects models. We show that the general blueprint developed in the scalar case to derive valid moment functions carries over to VAR(1) and MAR(1) models. We make no attempt at showing that our approach is exhaustive in those cases and do not claim that it is. We leave these important questions for future work. Readers uninterested in the details of the multivariate extensions can skip directly to Section (ref) where we discuss the empirical application.
We begin with the analysis of VAR(1) logit models, variants of which have been successfully used to study the relationship between sickness and unemployment (narendranthan1985investigation), the progression from softer drug use to harder drug use among teenagers (deza2015there), transitivity in networks (graham2013comment, graham2016homophily) and more recently the employment of couples (honore2022simultaneity). For a given $M\geq 2$, the model reads:
We let $Y_{it}=(Y_{1,it},\ldots,Y_{M,it})'$ denote the outcome vector in period $t$ with support $\mathcal{Y}=\{0,1\}^M$ of cardinality $2^M$. We let $X_{it}=(X_{1,it}',\ldots ,X_{M,it}')'\in \mathbb{R}^{K_1}\times \ldots \times \mathbb{R}^{K_M}$ denote the vector of exogenous covariates in period $t$ and $A_i=(A_{1,i},\ldots,A_{M,i})'\in \mathbb{R}^M$ . The initial condition is now given by $Y_{i0}=(Y_{1,i0},\ldots,Y_{M,i0})'\in \mathcal{Y}$ and the model transition probabilities are given by:
for all $(k,l)\in \mathcal{Y}\times \mathcal{Y}$. \\ Building on honore2000panel, honore2019panel use a conditional likelihood approach to prove the identification $\theta_0=(\gamma_{011},\gamma_{012},\gamma_{021},\gamma_{022},\beta_{01},\beta_{02})$ for the bivariate specification when $T=3$ and the regressors do not vary over the last two periods. As in scalar models, we show hereinafter that this strong restriction which can yield undesirable rates of convergence is unnecessary to obtain valid moment conditions. \\ Step 1) in the VAR(1) logit model has a nuance relative to its scalar counterpart in that the only transition functions that appear to exist are those associated to $\pi^{k|k}_{t}(A_i,X_{i})$, for $k \in \mathcal{Y}$, i.e the probabilities of staying in the same state. We can use the same heuristic as in the baseline AR(1) model to derive their expressions, especially in the bivariate case. Once all four transition functions are obtained for the case $M=2$, it becomes clear that the general functional form is as per Lemma (ref). It is then a matter of brute force calculation to verify that this is indeed correct.
Next, we can appeal to the second partial fraction decomposition formula in Appendix Lemma (ref) to guide the construction of another set of transition functions when $T\geq 3$. These identities may be regarded as a generalization of Kitazawa_JOE2021's hyperbolic transformations to the multivariate case. As is clear from Lemma (ref), the resulting transition functions have a special structure that generalizes those found in the AR(1) model.
Beyond $T=4$, more transition functions are available and can be derived sequentially from those of Lemma (ref). See Corollary (ref) for their expressions.
Step 2). One can obtain a family of valid moment functions by adequately repurposing the statement of Proposition (ref) to the VAR(1) case, i.e by updating the expressions of $\phi_{\theta}^{k|k}(.)$ and $\zeta_{\theta}^{k|k}$ according to Lemma (ref) and Corollary (ref). To conserve on space and avoid repetition, we leave this simple exercise to the reader.
Last, we cover dynamic multinomial logit models which have been utilized to measure state-dependence in a range of economic contexts including: employment history in the French labor market (magnac2000subsidised), the impact of international trade on the transition matrix of employment across sectors (egger2003sectoral) and consumer product choice (dube2010state) amongst others. \\ We focus on the the baseline MAR(1) logit model with fixed effects.
The model assumes a fixed number of alternatives $C+1$ with $C\geq 1$ and is characterized by the following transition probabilities:
with $(k,l)\in \mathcal{Y}=\{0,1,\ldots,C\}$. Here, $Y_{it}\in \mathcal{Y}$ indicates the choice of individual $i$ in period $t$, $X_{ijt}$ denotes a vector of individual-alternative specific exogenous covariates and $A_{ij}\in \mathbb{R}$ is the fixed effect attached to alternative $j$ for individual $i$. The initial condition is $Y_{i0}\in \mathcal{Y}$ and in keeping with the fixed effect assumption, its conditional distribution given unobserved heterogeneity and the regressors, $\left(P(Y_{i0}=k|X_i,A_i)\right)_{k=1}^C$, is left fully unrestricted. Following magnac2000subsidised, we normalize the transition parameters and fixed effect of the reference alternative \say{$0$} to zero \footnote{ The transition parameters of the reference state cannot be identified so a normalization constraint must be imposed. Setting $A_{i0}=0$ is also without loss of generality since we can always redefine the fixed effect as $A_{ik}^*=A_{ik}-A_{i0}$.}. That is $\gamma_{j0}=\gamma_{0j}=0, A_{0,j}=0$ for all $j\in \mathcal{Y}$ leaving $\theta=\left((\gamma_{kl})_{k,l\geq 1},(\beta_{l})_{l\geq 0}\right)$ as the unknown model parameters.\\ This specification can be motivated by assuming that agents rank options according to random latent utility indices with disturbances independent over time and across alternatives. In this context, equation ((ref)) is obtained if the best alternative is selected and the error terms are Type 1 extreme value distributed conditional on $Y_{i0},A_i,X_i$. magnac2000subsidised studies the \say{pure} case without covariates and shows that an extension of the conditional likelihood approach proposed by chamberlain_1985 can be used to identify and estimate the state-dependence parameters. honore2000panel show that this argument carries over to the case with exogenous explanatory variables if one matches the regressors across specific time periods. Here, we offer an alternative estimation strategy that circumvents the need for matching. \\ Step 1). Similarly to the VAR(1) model the MAR(1) appears to admit transition functions only for the probabilities of staying in the same state, namely $\pi^{k|k}_{t}(A_i,X_i)$ for $k\in \mathcal{Y}$. This feature appears to be a common trait of multidimensional fixed effects specifications. To facilitate the derivation of the relevant transition functions, we follow our usual heuristic of looking for $\phi_{\theta}^{k|k}(.), k \in \mathcal{Y}$ satisfying:
Upon obtaining their exact expressions for the simplest case with $C=2$, it is easy to conjecture and verify by direct calculations that the general expressions of the $C+1$ transition functions of the MAR(1) model are as displayed in Lemma (ref).
Unsurprisingly, given the similarities shared between the MAR(1) and all other specifications discussed in the paper, so long as $T\geq 3$, one can again derive transition functions other than $\phi_{\theta}^{k|k}(Y_{it-1}^{t+1},X_i)$ also associated to $ \pi^{k|k}_{t}(A_i,X_i)$ for $k \in \mathcal{Y}$ in periods $t\in\{1,\ldots,T-1\}$. The simple logistic identities of Appendix Lemma (ref) imply that these transition functions, that we keep denoting $\zeta_{\theta}^{k|k}(.)$ have a similar form to those of the VAR(1) model as shown in Lemma (ref).
Additionally, if the econometrician has access to a dataset with more than four observations per sampling unit - counting the initial condition - then, more transition functions associated to the same transition probabilities are available per Corollary (ref).
This completes Step 1) for the MAR(1) logit model. For Step 2), we recommend a family of valid moment functions mirroring those of Proposition (ref) for the AR(1) case to ensure the linear independence of its elements.
In this last section, we illustrate the usefulness of our methodology by revisiting the analysis of deza2015there on the dynamics of drug consumption amongst young adults in the United States.\footnote{This research was conducted with restricted access to Bureau of Labor Statistics (BLS) data. The views expressed here are those of the author and do not reflect the views of the BLS.} \\ To provide context, multiple studies have documented that young individuals who experiment with soft drugs have a tendency to continue using them and are at a higher risk of transitioning to hard drugs. Such correlations are certainly concerning. However, the empirical evidence of genuine causal links, in particular from softer drugs to harder drugs, remains limited with deza2015there standing as a notable exception. Fundamentally, these empirical regularities may be attributed to a causal effect (i.e. state dependence within and between drugs) or alternatively to latent traits that make individuals more prone to using illicit substances in general. Our primary concern is to untangle these two explanations to inform the design of policies aiming to mitigate drug addiction \footnote{See heckman1981heterogeneity for insights on the implications of state dependence for the design of labor market policies.}. For example, if marijuana consumption indeed serves as a gateway to later cocaine use, early educational interventions cautioning against casual marijuana usage could potentially have enduring effects on the population of heavy drug users. \\ To investigate these issues, we employ the restricted version of the National Longitudinal Survey of Youth 1997 (NLSY97). This is a panel dataset of 8984 individuals surveyed on a diverse range of subjects, including drug-related matters from 1997 to 2019. We concentrate on a subsample of four waves, spanning from 2001 to 2004. This subsample provides insight into the behavior of young adults between the age of 16 and 20 in 2001 to 19 and 24 in 2004. We shall examine the statistical association between three binary outcome variables, namely the consumption of alcohol, marijuana and hard drugs, derived from respondents answers' during annual interviews. Upon retaining those providing answers in all four waves as well as a valid state of residence, our cross section ultimately consists of $N=6317$ individuals \footnote{We adapt the sample selection procedure described in deza2015there for the period 2001-2004.}. Following deza2015there, we then consider the trivariate VAR(1) logit model
$m\in\{1,2,3\}$ ($1$=\say{alcohol}, $2$=\say{marijuana}, $3$=\say{hard drugs}), $t=1,2,3$ where $t=0$ corresponds to the year $2001$. The state-dependence coefficients $\gamma_{0mm}$ (within) and $\gamma_{0mj},m\neq j$ (between) are the principal coefficients of interest in the 16-dimensional vector of common parameters $\theta_{0}$. We are most particularly concerned about the sign and the statistical significance of $\gamma_{032}$, i.e the so called \say{stepping-stone} effect of marijuana on hard drugs. The covariate $age_{it}$ denotes the age of respondent $i$ at time $t$. The regressors $TEDS_{m,it}$ measure state-level deviations from national trends in treatment admissions for substance abuse caused by drug $m$ in year $t$ in the state of residence of $i$\footnote{ The variables $TEDS_{m,it}$ are constructed from the Treatment Episode Data Set-Admissions which records admissions to substance abuse treatment facilities in the United States.}. They are computed as the ratio of the share of admissions to treatment centers due to drug $m$ in the state of $i$ in year $t$ against the country wide analog in year $t$. Intuitively, this may be interpreted as a measure of exposure to substance $m$ for each respondent in our sample.\\ deza2015there parameterizes both the latent permanent heterogeneity $(A_{m,i})_{m=1}^3$ and the initial condition $Y_i^{0}$ to estimate the model by maximum likelihood. We leave these components unrestricted and exploit the valid moment functions presented in Section (ref). We specifically use six of the eight valid moment functions available: $\psi_{\theta}^{k|k}(Y_{i1}^{3},Y_{i0}^{1},X_i)$ for $k\in \{(0,0,0),(0,1,0),(1,1,1),(1,1,0),(1,0,1),(1,0,0)\}$. The other two corresponding to states $k\in\{(0,0,1),(0,1,1)\}$ are null for over $99.5\%$ of our sample and were dropped to mitigate noise in estimation. Next, we (arbitrarily) select a constant, the initial condition $Y_{i}^{0}$, $age_{it}$ and the covariates $TEDS_{m,it}$ in all periods $t=1,2,3$ as instruments to form the $96\times 1$ moment vector \\
With $m_{\theta}(Y_{i},Y_{i}^{0},X_i)$ in hand, we then consider the iterated GMM estimator of hansen1996finite. Starting from an initial candidate $\hat{\theta}_0$\footnote{In practice, we used the GMM estimator putting equal weights on each moment as our starting candidate.}, it can be described as
where $\overline{m}_{N}(\theta)=\frac{1}{N}\sum_{i=1}^N m_{\theta}(Y_{i},Y_{i}^{0},X_i)$ and $\overline{W}_{N}(\theta)=\frac{1}{N}\sum_{i=1}^N m_{\theta}(Y_{i},Y_{i}^{0},X_i)m_{\theta}(Y_{i},Y_{i}^{0},X_i)'$. Under some regularity conditions (hansen2021inference), this estimator is well defined and asymptotically normally distributed with
where $M_0 = \mathbb{E}\left[\pdv{m_{\theta_0}(Y_{i},Y_{i}^{0},X_i)}{\theta}\right]$ and $W_0=\mathbb{E}\left[m_{\theta_0}(Y_{i},Y_{i}^{0},X_i)m_{\theta_0}(Y_{i},Y_{i}^{0},X_i)'\right]$. Our motivation for focusing on this specific estimator originates mainly from hansen2021inference who advocate its use for two practical reasons. First, for a given set of moments, it eliminates the arbitrariness in the choice of the initial weight matrix of 2-step GMM estimators (see also imbens2002generalized). Second, because the iteration sequence is a contraction, each iteration is approximately variance reducing in the sense that: $Var(\hat{\theta}_{s}) \approx c^2 Var(\hat{\theta}_{s-1})$ for some constant $c<1$ \footnote{Note that the limiting variance of the iterated GMM estimator and a 2-step GMM estimator will be identical.}. Empirically, we also found in Monte Carlo simulations that the iterated GMM estimator performs relatively well for this type of specification (see Appendix (ref)). \\ Table (ref) presents the iterated GMM estimates for the trivariate VAR(1) logit model in columns (I), (II), (III). For comparison, columns (IV), (V), (VI) report a random effect (RE) estimator akin to deza2015there \footnote{We borrow the specification presented in deza2015there. The heterogeneity distribution is discrete with 3 mass points and is independent of the regressors. The initial condition relates to the covariates through a logistic regression.} while columns (VII), (VIII), (IV) display the \say{naive} logit maximum likelihood estimator (MLE) neglecting the presence of fixed effects. \\ The first observation is that, in line with conventional wisdom, GMM estimates for the state-dependence parameters within drug, $\gamma_{11},\gamma_{22},\gamma_{33}$, are all positive. As is apparent from columns (I)-(III), they are statistically significant for alcohol and marijuana but surprisingly not for hard drugs. In other words, there is no statistical evidence of a direct effect from past consumption of hard drug to future usage of hard drugs once we account for unobserved heterogeneity and the effects of other substances, at least in our four-wave sample\footnote{The transition parameters for hard drugs are expected to be noisier given that a smaller fraction of individuals consume these more lethal substances: approximately 15% of the respondents indicate having consumed hard drugs at least once from 2001-2004. This contrasts with 86% for alcohol and 40% for marijuana.}. Notice that the magnitude of the estimates for $\gamma_{11},\gamma_{22},\gamma_{33}$ sharply contrast with the other two estimators. The naive MLE largely overestimates the amount of within state-dependence, yielding coefficients that are comparatively four to eight times larger. Intuitively, this can be rationalized by the fact that this estimator misinterprets any serial correlation produced by $A_i$ as evidence of state dependence. The RE estimator borrowed from deza2015there (see also card2005estimating, chay1998identification) acts as an intermediate estimator between the other two as can be seen in columns (IV)-(VI). This behavior is expected to the extent that the additional parametric structure of this methodology will account to some degree for the presence of unobserved heterogeneity. We note that the role of within state dependence in the dynamics of drug consumption is nevertheless overstated by this approach.
Second and importantly, we observe in column (III) a positive and statistically significant effect of marijuana on hard drugs. This supports the view that marijuana usage can be a gateway to the consumption of harder drugs and accords with the key findings of deza2015there. From a practical standpoint, this result corroborates that there may be scope for policies on marijuana usage to indirectly curb the consumption of more lethal substances by teenagers and young adults. The efficacy of such policies in the short and long run are important questions that will intuitively depend on the distribution of heterogeneity in the population. We do not explore those questions here but further research in this direction would be of interest \footnote{A natural idea to gauge the effectiveness of policy interventions would be to compute average marginal effects. However, as mentioned in Section (ref), we were unable to find transition functions for the transition probabilities where the state switches in VAR(1) models. This leads us to believe that only the average transition probabilities where the state remains unchanged are identified. In turn, this would imply that average marginal effects are generally partially identified in VAR(1) models. In this case, it is possible that ideas analogous to those in dobronyi2021identification and davezies2021identification could be used to characterize and compute the identified set of average marginal effects; albeit some difficulties might arise due to the fact that the fixed effects are now multidimensional. Computing outer bounds as in pakel2023bounds could be another plausible option.}. The other two estimators also agree on a positive influence of marijuana on the consumption of harder drugs, albeit it is statistically insignificant in the RE case.
Otherwise, it is noteworthy that the between state dependence estimates can vary quite significantly across specifications. Again, the naive MLE likely misinterprets spurious correlation from the $A_i$ as state dependence which results in positive and inflated cross effects. Column (IV) and (I) show disagreements of the RE and GMM estimates regarding the strength of the impact of marijuana and hard drugs on alcohol. Overall, this comparative exercise has showed that accounting for unobserved heterogeneity as flexibly as possible can be essential to obtain an accurate picture of the patterns of state dependence in practice.
Dynamic discrete choice models are widely used to study the determinants of repeated decisions made by individuals or firms over time. In this paper, we have introduced a procedure to estimate a family of such models with logistic (or Type I extreme value) errors and potentially many lags while remaining agnostic about the nature of unobserved individual heterogeneity. This type of approach may be attractive when the risk of misspecifying the initial condition and the unit-specific effects are important. We also provided general expressions for average marginal effects in the binary response case which are often the counterfactuals of interest in practice. \\ The list of discrete choice models covered in this paper is of course not exhaustive and it would be interesting to know if our two-step approach could be deployed in other settings with \say{logit} noise. In ongoing work, we have found that this is one avenue to approach estimation of dynamic ordered logit models, potentially of arbitrary lag order.
\nocite{*}
\setcounter{page}{0} \pagenumbering{arabic} \setcounter{page}{1}