EconBase
← Back to paper

A dynamic ordered logit model with fixed effects

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

67,131 characters

A dynamic ordered logit model with fixed effects


\title{A dynamic ordered logit model with fixed effects}
\author{Chris Muris, Pedro Raposo, and Sotiris Vandoros\thanks{Muris: Department of Economics, McMaster University. Contact: [email removed].
Raposo: Cat�lica Lisbon School of Business and Economics. Contact:
[email removed]. Vandoros: King\textquoteright s College London
and Harvard University. Contact: [email removed]. The data were
made available by Eurostat (Contract RPP 132-2018-EU-SILC). We are
grateful to Irene Botosaru and Krishna Pendakur for very helpful discussions.}}
\date{August 4, 2020}
\maketitle
\begin{abstract}
We study a fixed-$T$ panel data logit model for ordered outcomes
that accommodates fixed effects and state dependence. We provide identification
results for the autoregressive parameter, regression coefficients,
and the threshold parameters in this model. Our results require only
four observations on the outcome variable. We provide conditions under
which a composite conditional maximum likelihood estimator is consistent
and asymptotically normal. We use our estimator to explore the determinants
of self-reported health in a panel of European countries over the
period 2003-2016. We find that: (i) the autoregressive parameter is
positive and analogous to a linear AR(1) coefficient of about 0.25,
indicating persistence in health status; (ii) the association between
income and health becomes insignificant once we control for unobserved
heterogeneity and persistence.
\end{abstract}

\section{Introduction}

Certain individual-level conditions may tend to persist over time,
in the sense that a condition has a memory of a previous period\textquoteright s
state, or may involve an element of adaptation. Furthermore, the way
individuals experience the same condition may vary, and they may also
have a different understanding of how this is measured. A common example
that fulfils these characteristics, and which is used extensively
in the literature, is self-reported health status. Health status often
depends on its value in the previous period, as health conditions
may persist over time, given that recovery can take long, and that
an illness may even have permanent effects. For example, Table \ref{tab:transition-uk}
presents a \textit{\emph{transition matrix}}\textit{ }for self-reported
health status in the United Kingdom for the period 2003-2016.\footnote{More information about the data and source is in Section \ref{sec:health},
where we analyze this data using the methodology proposed in this
paper.}
\begin{table}
\begin{centering}
\begin{tabular}{ccccccc}
\hline
 &  & \multicolumn{5}{c}{$P\left(\left.Y_{i,t+1}=y'\right|Y_{i,t}=y\right)$}\tabularnewline
 & $y/y'$ & 1 & 2 & 3 & 4 & 5\tabularnewline
\hline
$P\left(Y_{i,t}=y\right)$ & 1 & 36.48 & 43.40 & 13.84 & 5.03 & 1.26\tabularnewline
 & 2 & 10.23 & 44.44 & 35.38 & 8.77 & 1.17\tabularnewline
 & 3 & 0.88 & 10.23 & 52.18 & 30.69 & 6.02\tabularnewline
 & 4 & 0.15 & 1.08 & 14.87 & 59.74 & 24.16\tabularnewline
 & 5 & 0.08 & 0.29 & 4.18 & 33.38 & 62.07\tabularnewline
\hline
\end{tabular}
\par\end{centering}
\caption{Current and future self-reported health, United Kingdom.}
\label{tab:transition-uk}
\end{table}
 For individuals that report a value of current health in a given
year (rows, on a 5-point scale with 5 being the highest), it shows
the relative proportion of those that report a certain value in the
subsequent year (columns). A striking feature of this transition matrix
is that a lot of mass is on or near the main diagonal. This feature
is found across all countries in our analysis. In other words, self-reported
health status is persistent: individuals tend to stay in the same
level of health.

There are at least two explanations for this observed persistence
(Heckman, 1981; Honor� and Kyriazidou, 2000): unobserved heterogeneity
and state dependence. Consider first \emph{unobserved heterogeneity.
}It\emph{ }refers to unobservable characteristics that affect the
propensity to report higher health status. Unobserved heterogeneity
is important in the literature on health status, because self-reported
health has been used extensively in the literature as a measure of
health outcomes (see for example Bound and Waidmann, 1992; Banerjee
et al., 2004; Gravelle and Sutton, 2009; McInerney and Mellor, 2012).
It is often viewed as a limitation that self-reported measures are
subjective. For example, reporting one's own health may depend on
cultural factors (Jylh� et al., 1998; Baron-Epel et al., 2005; J�rges,
2007), and people may have a different understanding of reference
points for health (Groot, 2000; Sen, 2002). Previous studies have
used vignettes to address cross-country differences in reporting of
health and disability (King et al., 2004; Salomon et al., 2004; Kapteyn
et al., 2007). However, the issue with unobserved heterogeneity across
individuals remains. As a result, it is important to take into account
the role of unobserved heterogeneity when analyzing self-reported
health data. The appropriate econometric approach to this is to allow
for fixed effects.\footnote{Studies have long debated the accuracy and reliability of subjective
measures of health, such as self-reported health status (see for example
Butler et al., 1987; Lindeboom and van Doorslaer, 2004; Johnston et
al., 2009). As an alternative response to these concerns, a number
of objective measures of health have been included in household surveys.
These include blood pressure, BMI, the number of medicines taken (Health
Survey England, 2019), the number of sick days off work, the number
of days hospitalised (BHPS, 2019) etc. Some household surveys ask
respondents to perform a task such as walking across the room or buttoning
a shirt to capture any limitations (SHARE, 2019). Indexes such as
the EQ-5D index are being used to cover different types of conditions
and merge them into a single measure. The Euro-D scale measures mental
health, and the CASP-12 index captures quality of life. However, these
objective measures are often very specific to particular diseases,
and even when creating a relevant index it may be impossible to include
and accurately reflect all conditions. As such, while objective measures
may accurately capture some health conditions, they have serious limitations
in capturing the overall picture of one\textquoteright s health.}

A number of studies have found a positive association between income
and health (Carrieri and Jones, 2017; Ettner, 1996; Frijters et al.,
2005; Mackenbach et al., 2005), but empirical evidence of a strong
effect is sometimes limited (Larrimore, 2011; Gunasekara, 2011; Johnston
et al., 2009). Nevertheless, the literature on the impact of economic
downturns (which mean reduced income) has previously demonstrated
positive effects of unemployment on health (Ruhm, 2000; Ruhm and Black,
2002), and more recently, no effect (Ruhm, 2015). What also appears
to matter, at least in terms of happiness, apart from absolute income,
is also relative income, i.e. how one\textquoteright s income compares
to that of those around them (Frijters et al., 2008). The relationship
between income and health is endogenous and complex, and both can
be correlated with other factors, that are not always measured and
included in empirical models. For example, Gunasekara et al. (2011)
found that when controlling for unmeasured confounders, the association
between the two becomes weaker. With regards to self-reported hypertension
in particular, Johnston et al., (2009) found no link to income --
something that did change when using objective measures. In our paper,
using models that do not control for individual unobserved heterogeneity
yields a positive and statistically significant coefficient (Table
3). However, this becomes insignificant when using fixed effects,
suggesting that there are other factors that potentially drive the
association between the two variables.

Consider now the second explanation for the observed persistence in
health outcomes: \emph{state dependence}. It refers to the possibility
that past self-reported health status may be related to current self-reported
health status even after conditioning on unobserved heterogeneity.
State dependence arises if actual (as opposed to self-reported) health
shocks are persistent, in the sense that a shock on health can have
a long-lasting effect (a typical example is injury leading to disability).
Contoyannis et al. (2004), using a random effects approach, found
evidence for such persistence in respondents of the British Household
Panel survey.

State dependence in self-reported health can also arise due to adaptation:
self-reported health status may change over time for a person whose
actual health has not changed. People tend to adapt to good or bad
developments in life. According to the Global Adaptive Utility Model,
individuals reallocate weights on various domains of life in order
to maintain their previous level of utility (Bradford and Dolan, 2010).
Similarly, the AREA model developed by Wilson and Gilbert (2008),
suggests that attention is focused on a change, followed by reaction,
explanation, and, finally, adaptation. This also applies to health,
as health status tends to improve even when individuals' health has
actually not experienced any objective change (Daltroy et al., 1999;
Damschroder et al., 2005), and time since diagnosis is positively
associated with self-reported health (Cub�-Moll� et al., 2017). Whether
persistence or adaptation, or both, characterise a variable, this
calls for a dynamic element in a model.

Overall, the challenges with studying self-reported health status
is that (a) people with the same actual health status might be reporting
different health levels; and (b) health shocks can have a lasting
effect. Against this background, we propose and analyze a panel data
ordered logit model that includes both fixed effects and a lagged
dependent variable. This allows a researcher faced with panel data
and an ordinal outcome variable to disentangle unobserved heterogeneity
from state dependence, and to quantify state dependence. Thus, we
address the limitations of using self-reported health as a proxy for
individuals' health. Our contribution is important for studies using
subjective health measures as it can help correct biases that naturally
occur when using this type of measure.\footnote{For example, happiness is perceived and reported differently across
individuals and people adapt to things that make them happy (Layard,
2006), while shocks on happiness can have a scarring effect on next
periods (Clark et al., 2001).}

Specifically, we study the \textit{dynamic ordered logit model with
fixed effects:}
\begin{align}
Y_{i,t}^{*} & =\alpha_{i}+X_{i,t}\beta+\rho1\left\{ Y_{i,t-1}\geq k\right\} -U_{i,t},\,t=1,2,3,\label{eq:latent_variable}\\
Y_{i,t} & =\begin{cases}
1 & \text{if }Y_{i,t}^{*}<\gamma_{2},\\
2 & \text{if }\gamma_{2}\leq Y_{i,t}^{*}<\gamma_{3},\\
\vdots\\
J & \text{if }Y_{i,t}^{*}\geq\gamma_{J},
\end{cases}\label{eq:thresholds}\\
\left.U_{i,t}\right| & \left(\alpha_{i},X_{i},Y_{i,<t}\right)\sim LOG(0,1),\,t=1,2,3,\label{eq:strict_exogeneity}
\end{align}
where $2\leq k\leq J$ is a fixed and known cutoff for the lagged
dependent variable. The person-specific parameter $\alpha_{i}$ captures
unobserved heterogeneity, which we allow to be correlated with the
other quantities in the model in an unrestricted way (fixed effects).
The time-varying covariates $X_{i,t}$ are collected across time periods
in $X_{i}=\left(X_{i,1},X_{i,2},X_{i,3}\right)$, and the lagged dependent
variables for period $t$ are collected in $Y_{i,<t}=\left(Y_{i,0},\cdots,Y_{i,t-1}\right)$.
The autoregressive parameter $\rho$ is the regression coefficient
on the lagged dependent variable $1\left\{ Y_{i,t-1}\geq k\right\} $;
$\beta$ is the regression coefficient on the contemporaneous covariates;
and the threshold parameters $\gamma_{j}$ map the underlying latent
variable $Y_{i,t}^{*}$ into the observed ordered outcome $Y_{i,t}$.
Equation (\ref{eq:strict_exogeneity}) restricts the error terms $U_{i,t}$
to be i.i.d. logistic, and is a strict exogeneity assumption on the
regressors and past outcomes.\footnote{The dynamics in our model are restricted to depend on $Y_{i,t}$ through
$1\left\{ Y_{i,t-1}\geq k\right\} $ only. An alternative model for
which we can identify some features is one that is linear in its history,
i.e. $Y_{i,t}^{*}=\alpha_{i}+X_{i,t}\beta+\rho Y_{i,t-1}-U_{i,t}.$
We were unable to use our approach to obtain identification in the
more general model with $Y_{i,t}^{*}=\alpha_{i}+X_{i,t}\beta+\sum_{j=2}^{J}\rho_{j}1\left\{ Y_{i,t-1}=j\right\} -U_{i,t}$.}

This model combines a number of noteworthy features. First, it is
a model for discrete ordered outcomes, and therefore a \textit{nonlinear}
model. Second, it is \textit{dynamic}, in the sense that the current
outcome depends directly on the outcome in the previous period. This
feature, called \textit{state dependence}, is governed by the autoregressive
parameter $\rho$. Third, it allows for \textit{unobserved heterogeneity}
in an unrestricted way, i.e. it is a \emph{fixed effects }model. Fourth,
the model is only specified for a \textit{small number of time periods},
$T=3$. Period 0 is unmodelled, but an observation on the outcome
variable in time 0 is required for identification.

We believe that we are the first to provide identification and estimation
results for all common parameters in a dynamic ordered logit model
with fixed effects and a fixed number of time periods. Using four
time periods of data on the ordinal outcome variable, we identify
the autoregressive coefficients on the lagged dependent variable,
and the regression coefficients on the exogenous regressors. We also
identify the threshold parameters, which makes it possible to interpret
the magnitude of the estimated coefficients. This distinguishes the
ordered choice model from the dynamic binary choice model with fixed
effects, where such an interpretation is not available. Our identification
result suggest a composite conditional maximum likelihood estimator
for the parameters in our model. We establish conditions under which
that estimator is consistent and asymptotically normal.

We use our estimator to investigate the determinants of self-reported
health, focusing on the link between income and health in a panel
of European countries over the period 2003-2016. We obtain two main
findings. First, even after controlling for unobserved heterogeneity,
persistence plays a positive and significant role in one's self-reported
health. In other words, one's health is dependent on the health in
the previous period, which is a reasonable thing to expect, as health
problems may expand over a number of periods, or become permanent.
Quantitatively, we estimate a persistence parameter that is analogous
to an autoregressive parameter of about 0.25 in a linear AR(1) model.
Second, we find that, when controlling for unobserved heterogeneity,
the link between income and health becomes statistically insignificant,
suggesting that other factors might explain the association between
the two. This is in line with studies that have found a smaller or
insignificant association when using fixed effects (Gunasekara, 2011;
Larrimore, 2011).

\section{Related literature in econometrics}

We believe that our paper is the first to provide identification and
estimation results for a panel data model with (i) ordered outcomes;
(ii) a lagged dependent variable; (iii) fixed effects; and (iv) a
fixed number of time periods. Our econometric contribution is related
to several strands of literature, each of which features a subset
of these features.

Most closely related to our paper is the literature on binary and
multinomial choice models with fixed effects and lagged dependent
variables, which features all but (i). The seminal work by Honor�
and Kyriazidou (2000) builds on Cox (1958) and Chamberlain (1985)
to estimate the parameters in dynamic binary choice logit model with
fixed effects and time-varying regressors. Hahn (2001) discusses the
information bound for a special case of their model. Honor� and Kyriazidou
(2019) discuss identification of some closely related models. Honor�
and Weidner (2020) construct moment conditions that shed light on
identification in this and related models, and provide a $\sqrt{n}$-consistent
estimator. Honor� and Tamer (2006), Aristodemou (2020) and Khan et
al. (2020) obtain results for models that do not have logistic errors.
For the static multinomial model, Chamberlain (1980) studies the logit
case; Shi et al. (2008) provides results for the general static; and
Magnac (2000) studies the dynamic version. We supplement these results
by showing that, in an \emph{ordered} choice model, the thresholds
in the latent variable model can be identified along with the regression
coefficients and the autoregressive parameter. This allows for a quantitative
interpretation of true state dependence. Such an interpretation is
not available in the binary and multinomial choice models.

The literature on static ordered logit models with fixed effects features
all but (ii). This model was analyzed by Das and van Soest (1999),
Baetschmann et al. (2015), and Muris (2017). Our result differs from
the results in those papers, because we provide results for a \emph{dynamic}
version of the ordered logit model.

The literature on random effects dynamic ordered choice models features
all but (iii). Random effects dynamic ordered choice models have been
studied and applied extensively (Contoyannis et al., 2004; Albarran
et al., 2019). Such approaches require strong restrictions on the
relationship between the unobserved heterogeneity and the exogeneous
variables in the model. Such restrictions are usually unappealing
to economists, as evidenced by the fact that they are rarely used
in linear models. Our approach does not impose random effects restrictions
and is the first to provide a fixed effects approach for dynamic ordered
choice models.

Note that our approach is fixed-$T$ consistent. The difficulty of
allowing for fixed effects is alleviated when one can assume that
$T\to\infty$, referred to as ``large-$T$''. Large-T fixed effects
dynamic ordered choice models have been studied by Carro and Traferri
(2014) and Fern�ndez-Val et al. (2017), see also Carro (2007) for
the binary outcome case. In the large-$T$ case, one can use techniques
that correct for the bias that comes from including fixed effects
in the nonlinear panel model. This approach does not feature (iv).
These techniques are not appropriate for our empirical application,
which is a rotating panel with $T=4$.

One limitation of our approach is that we restrict the way in which
the lagged dependent variable enters the model. The random effects
and large-$T$ approach can accommodate a richer dynamic specification.
We leave for future work whether such an extension is possible with
a fixed-$T$ fixed-effects approach.

\section{Identification\label{sec:Identification}}

We normalize $\gamma_{k}=0$, where $k$ is as in equation (\ref{eq:latent_variable}).
This scale normalization is without loss of generality because the
scale of $\alpha_{i}$ is unrestricted. Our model implies that the
binary variable $D_{i,t}(k)=1\left\{ Y_{i,t}\geq k\right\} $ follows
the dynamic binary choice logit model in Honor� and Kyriazidou (2000),
HK hereafter. Specifically, equation (3) in HK applies to the transformed
model
\[
D_{i,t}(k)=1\left\{ X_{i,t}\beta+\rho D_{i,t-1}\left(k\right)+\alpha_{i}-U_{i,t}\geq0\right\} ,
\]
i.e. the transformed model follows a dynamic binary choice logit model
with fixed effects. The implied conditional probabilities relevant
for our analysis are
\begin{equation}
P\left(\left.D_{i,0}\left(k\right)=1\right|X_{i},\alpha_{i}\right)\equiv p_{0}\left(X_{i},\alpha_{i}\right),\label{eq:model_probabilities_0}
\end{equation}
and, for $t=1,2,3$,

\begin{align}
P\left(\left.D_{i,t}\left(k\right)=1\right|X_{i},\alpha_{i},D_{i,<t}\left(k\right)\right) & =\frac{\exp\left(\alpha_{i}+X_{i,t}\beta+\rho D_{i,t-1}\left(k\right)\right)}{1+\exp\left(\alpha_{i}+X_{i,t}\beta+\rho D_{i,t-1}\left(k\right)\right)},\label{eq:model_probabilities_t}
\end{align}
where we have let $D_{i,<t}\left(k\right)=\left(D_{i,0}\left(k\right),\cdots,D_{i,t-1}\left(k\right)\right)$.
HK provide conditions that guarantee identification of $\beta$ and
$\rho$ by constructing a conditional probability that features $\left(\beta,\rho\right)$
but that is free of $\alpha_{i}$.

If $Y_{i,t}$ has at least three points of support, there is information
in $Y_{it}$ beyond $D_{it}\left(k\right)$. In the remainder of this
section, we show that this information can be used to identify the
threshold parameters
\[
\gamma\equiv\left(\gamma_{2},\gamma_{3},\cdots,\gamma_{k-1},\gamma_{k+1},\cdots,\gamma_{J}\right).
\]
This leads to an interpretation of the magnitude\textbf{ }of $\left(\beta,\rho\right)$
that is not available for the dynamic binary choice model. Muris (2017,
Section III.C) discusses this for the static panel data ordered choice
models ($\rho=0$).

We now construct a conditional probability that features $\left(\beta,\rho,\gamma\right)$
but not the incidental parameters $\alpha_{i}$. To this end, extend
the definition
\[
D_{i,t}\left(j\right)=1\left\{ Y_{i,t}\geq j\right\} ,\,2\leq j\leq J,
\]
to thresholds $j\neq k$, and abbreviate $D_{i,t}\equiv D_{i,t}\left(k\right)$.
Define the events $\left(A_{j,l},B_{j,l},C_{j,l}\right)$, with $2\leq j\leq k\leq l\leq J$,\footnote{Choosing $j\leq k$ guarantees that when $D_{i,2}\left(j\right)=0,$
the lagged dependent variable in period 3 is 0. The opposite is true
for $l\geq k$ and $D_{i,2}\left(l\right)=1$. There would be no gain
from considering a threshold different from $k$ in the first period.
Using $k$ as the threshold in the second period is the only way to
cancel out the threshold parameters from period 1. Given that we are
using the subpopulation $X_{i,2}=X_{i,3}$, the fact that $j,l$ are
used alternately in periods 2 and 3 does not create additional difficulties.} as follows:
\begin{align*}
A_{j,l} & =\left\{ D_{i,0}=d_{0},D_{i,1}=0,D_{i,2}\left(l\right)=1,D_{i,3}\left(j\right)=d_{3}\right\} ,\\
B_{j,l} & =\left\{ D_{i,0}=d_{0},D_{i,1}=1,D_{i,2}\left(j\right)=0,D_{i,3}\left(l\right)=d_{3}\right\} ,\\
C_{j,l} & =A_{j,l}\cup B_{j,l}.
\end{align*}
For $d_{0}=d_{3}=0$, the event $A_{j,l}$ corresponds to moving up
in the middle periods $t=1,2$, starting below $k$ to moving up to
at least $l\geq k$. The event $B_{j,l}$ corresponds to moving down
in the middle periods, starting from at least $k$ and moving below
$j\leq k$.

If $j=k=l$, the event $C_{k,k}$ corresponds to switchers (observations
with $D_{i1}+D_{i2}=1$), as in HK. In the ordered model, it is possible
to vary the cutoffs in the periods $t=2,3$ if the dependent variable
has more than two points if support. Varying the cutoffs over time
is what distinguishes our conditioning event from that in HK. It is
what allows us to identify the threshold parameters.

The following sufficiency result shows that different choices of $\left(j,l\right)$
reveal different combinations of thresholds in certain conditional
probabilities that do not depend on the incidental parameters $\alpha_{i}$.
In what follows, the logistic function is denoted by $\Lambda\left(u\right)=\exp\left(u\right)/\left(1+\exp\left(u\right)\right)$,
and the change in the regressors from period 1 to 2 by $\Delta X_{i}=X_{i2}-X_{i1}$.
\begin{thm}[Sufficiency]
\label{thm:sufficiency}For the dynamic ordered logit model with
fixed effects, for any $\left(j,l\right)$ such that $2\leq j\leq k\leq l\leq J$,
and for any $d_{0},d_{3}\in\left\{ 0,1\right\} ,$
\begin{align}
P\left(\left.A_{j,l}\right|X_{i},C_{j,l},X_{i,2}=X_{i,3}\right) & =1-\Lambda\left(\Delta X_{i}\beta+\rho\left(d_{0}-d_{3}\right)+\left(1-d_{3}\right)\gamma_{l}+d_{3}\gamma_{j}\right)\label{eq:sufficiency_A}\\
P\left(\left.B_{j,l}\right|X_{i},C_{j,l},X_{i,2}=X_{i,3}\right) & =\Lambda\left(\Delta X_{i}\beta+\rho\left(d_{0}-d_{3}\right)+\left(1-d_{3}\right)\gamma_{l}+d_{3}\gamma_{j}\right).\label{eq:sufficiency_B}
\end{align}
\end{thm}
Identification of the model parameters comes from considering all
possible combinations of cutoffs. It is clear from Theorem \ref{thm:sufficiency}
that different choices for $\left(j,k,l,d_{0},d_{3}\right)$ reveal
information about distinct linear combinations of $\left(\rho,\gamma\right)$.
By considering multiple choices of $\left(j,k,l,d_{0},d_{3}\right)$,
and then aggregating the resulting information, we can identify all
the model parameters. We require an additional assumption before stating
our main identification result.
\begin{assumption}
\label{assu:XVariation}For all $\left(j,l\right)$ such that $2\leq j\leq k\leq l$,
and for all $d_{0},d_{3}\in\left\{ 0,1\right\} $
\[
Var\left(\left.\Delta X_{i}\right|X_{i,2}=X_{i,3},C_{j,l}\right)
\]
 is invertible.
\end{assumption}
This assumption guarantees that for each choice of $\left(j,l\right)$,
there is sufficient variation in $\Delta X_{i}$ in the subpopulation
of stayers $X_{i,2}=X_{i,3}$ to identify the regression coefficient.
This assumption can be weakened: we only need sufficient variation
for some $\left(j,l\right)$. However, if it fails for sufficiently
many $\left(j,l\right)$, identification of some of the threshold
parameters may fail.

Denote by $Y_{i}=\left(Y_{i,0},Y_{i,1},Y_{i,2},Y_{i,3}\right)$ the
time series of dependent variables for a given individual.
\begin{thm}[Identification]
\label{thm:Identification}If Assumption 2 holds, then $\left(\beta,\rho,\gamma\right)$
can be identified from the joint distribution of the vector $\left(X_{i},Y_{i}\right)$
generated by the dynamic ordered logit model with fixed effects.
\end{thm}

\section{Estimation\label{sec:Estimation}}

Theorem \ref{thm:sufficiency} suggests that, for each choice of $2\leq j\leq k\leq l$,
we could use a conditional maximum likelihood estimator (CMLE) to
estimate a linear combination of the model parameters. Theorem \ref{thm:Identification}
suggests that a composite CMLE (CCMLE), based on the combination of
conditional likelihoods across all choices of $\left(j,k,l\right)$,
may be used to estimate the model parameters $\left(\beta,\rho,\gamma\right)$.
In this section, we define that CCMLE and establish conditions under
which it has desirable large sample properties. We focus on the discrete
regressor case. Results for continuous regressors can be obtained
by adapting Theorems 1 and 2 in HK to our case.

The binary random variable
\[
C_{i,jl}=1\left\{ \left(D_{i,1}=0,D_{i,2}\left(l\right)=1\right)\text{ or }\left(D_{i,1}=1,D_{i,2}\left(j\right)=0\right)\right\} \times1\left\{ X_{i,2}=X_{i,3}\right\} .
\]
indicates whether $i$'s time series fits the description in $C_{j,l}=A_{j,l}\cup B_{j,l}$,
and that it is also a ``stayer'' in the sense that $X_{i2}=X_{i3}$.
Note that if $C_{i,jl}=1$, then $D_{i,1}=1$ implies that the individual
time series is of the type $B_{j,l}$. Similarly, if $C_{i,jl}=1$,
then $D_{i,1}=0$ implies that individual $i$ is of type $A_{j,l}$.

In the log-likelihood contribution below, (\ref{eq:CCML}), we have
substituted $D_{i,0}$ for $d_{0}$ in equation (\ref{eq:sufficiency_B}).
The value to substitute for $d_{3}$ depends on whether we are in
case $A$ or $B$. To that end, define
\begin{align*}
D_{i,3,jl} & =\begin{cases}
D_{i,3}\left(j\right) & \text{ if }D_{i,1}=0,\\
D_{i,3}\left(l\right) & \text{ if }D_{i,1}=1.
\end{cases}
\end{align*}
The conditional log likelihood contribution for individual $i$, for
cutoffs $\left(j,l\right),$ $2\leq j\leq k\leq l\leq J$, can then
be written

\begin{align}
l_{i,jl}\left(\beta,\rho,\gamma_{j},\gamma_{l}\right) & =C_{i,jl}\left[D_{i,1}\ln\left\{ \Lambda\left(\Delta X_{i}\beta+\rho\left(D_{i,0}-D_{i,3,jl}\right)+\gamma_{l}\left(1-D_{i,3,jl}\right)+\gamma_{j}D_{i,3,jl}\right)\right\} +\right.\nonumber \\
 & \phantom{}\left.\phantom{+}+\left(1-D_{i,1}\right)\ln\left\{ 1-\Lambda\left(\Delta X_{i}\beta+\rho\left(D_{i,0}-D_{i,3,jl}\right)+\gamma_{l}\left(1-D_{i,3,jl}\right)+\gamma_{j}D_{i,3,jl}\right)\right\} \right].\label{eq:CCML}
\end{align}
The CCMLE is
\begin{equation}
\widehat{\theta}_{n}=\left(\widehat{\beta}_{n},\widehat{\rho}_{n},\widehat{\gamma}_{n}\right)=\arg\max\frac{1}{n}\sum_{2\leq j\leq k\leq l}\sum_{i=1}^{n}l_{i,jl}\left(\beta,\rho,\gamma_{j},\gamma_{l}\right),\label{eq:CCMLE}
\end{equation}
where we have implicitly imposed $\gamma_{k}=0$ in the definition
of $l_{i,jl}$.

We maintain the following assumption to establish the asymptotic properties
of the CCMLE.
\begin{assumption}[Stayers]
\label{assu:discrete_stayers} $P\left(X_{i,2}=X_{i,3}\right)>0.$
\end{assumption}
With additional technical work, this assumption can be relaxed to
the case where $X_{i,2}-X_{i,3}$ is continuously distributed with
positive density around zero, see HK's Theorem 1 and 2.
\begin{thm}
\label{thm:asymptotics-CCMLE}Let \textup{$\left\{ \left(Y_{i},X_{i}\right),\,i=1,\cdots n\right\} $}
be a random sample of size $n$ from the dynamic ordered logit model
with fixed effects with true parameter values $\theta_{0}=\left(\beta_{0},\rho_{0},\gamma_{0}\right)$.
Under Assumptions \ref{assu:XVariation} and \ref{assu:discrete_stayers},
and for any value of $\theta_{0}$,
\[
\widehat{\theta}_{n}\stackrel{p}{\to}\theta_{0}\text{ as }n\to\infty.
\]
Furthermore,
\[
\sqrt{n}\left(\widehat{\theta}_{n}-\theta_{0}\right)\stackrel{d}{\to}\mathcal{N}\left(0,H^{-1}\Sigma H^{-1}\right)\text{ as }n\to\infty,
\]
where $\Omega$ as the variance of the score of the composite likelihood,
defined in (\ref{eq:score_CCML}), and $H$ is the associated Hessian
defined in (\ref{eq:Hessian_CCML}).
\end{thm}
\begin{rem}
The convexity of the summands in (\ref{eq:CCMLE}) means that the
objective function is convex. We compute the CCMLE using the Newton-Raphson
algorithm in R's nlm function (R Core Team, 2020). Supplying analytical
gradients and Hessians speeds up the estimation.
\end{rem}

\section{Persistence in self-reported health status\label{sec:health}}

Our analysis uses panel data for the period 2003-2016 from the European
Union Statistics on Income and Living Conditions (EU-SILC), see Eurostat
(2017) for detailed documentation. The microdata is publicly available
upon request.\footnote{The data were made available to us by Eurostat under Contract RPP
132-2018-EU-SILC.} EU-SILC provides a set of indicators on income and poverty, social
inclusion, living conditions and, importantly, health status. For
each country in the European Union, plus Iceland, Norway, and Switzerland,
EU-SILC contains data on a representative sample of the population
of those 18 years and older.

EU-SILC is a rotating panel. Every individual is followed over a period
of two to four years. The total number of individual-years for the
period 2003-2016 is 1273877. Our identification result demands four
observations per individual, so we restrict attention to individuals
that report valid information on their health status for 4 consecutive
years. This restriction, and the restriction that the explanatory
variables that we use in the analysis below have non-missing information,
leaves us with a sample of 260601 individuals, for 1042404 individual-years.
The proportion of incomplete samples differs across countries. As
a result, the sample we work with may not be representative of EU-SILC's
population. For example, out of the 27 countries that contribute to
our sample, the largest contributors are Italy (with 43385 individuals),
Spain (25634), and Poland (22628); the smallest are Portugal (12),
Iceland (1496), and Slovakia (1982).

The outcome variable in our analysis is self-reported health status:
self-perceived physical health, elicited during EU-SILC interviews.
The person answers the question on how she perceives her physical
health to be in general, at the date of the survey, by classifying
it as one of: (1) bad and very bad (12\% in our sample); (2) fair
(26\%); (3) good (44\%); (4) very good (19\%).\footnote{We have merged the separate categories ``bad'' and ``very bad''
in the original reported variable, because there is only a small fraction
of observations with ``very bad'' health status.} Out of $260601\times3=781803$ health transitions that we observe,
most often there is no change in health status (65.6\%). Decreases
by one unit (16\%) are slightly more frequent than increases by one
unit (15\%). Two-unit increases (1.2\%) and decreases (1.5\%) are
infrequent, and three-unit increases (0.08\%) and decreases (0.12\%)
are rare.

Table \ref{tab:health-income} relates the outcome variable, and changes
to the outcome variable, to our main explanatory variable of interest,
log income (total disposable household equivalised income). Household
income was scaled using the composition and size of each household.
This scale is based on the OECD modified equivalence scale, which
gives a weight of 1.0 to the first adult in the household, 0.5 to
other adults and 0.3 to each child (under 14 years old).

The table provides descriptive statistics for log income in our sample,
grouped by health status. The top panel is in levels. Average income
is increasing in health status. The bottom panel is in changes, which
represents one way to control for unobserved heterogeneity. The implied
increases for changes are close to zero, hinting at the imported role
of unobserved heterogeneity.
\begin{table}
\centering{}
\begin{tabular}{c>{\centering}p{2cm}rr}
\hline
 &  & \multicolumn{2}{c}{Log income}\tabularnewline
\hline
 &  & mean & sd\tabularnewline
\hline
health status & 1 & 8.68 & 0.92\tabularnewline
 & 2 & 8.95 & 0.94\tabularnewline
 & 3 & 9.29 & 0.94\tabularnewline
 & 4 & 9.51 & 0.90\tabularnewline
\hline
 &  & \multicolumn{2}{c}{$\Delta$Log income}\tabularnewline
 &  & mean & sd\tabularnewline
 & -3 & 0.05 & 0.47\tabularnewline
$\Delta$health status & -2 & 0.06 & 0.43\tabularnewline
 & -1 & 0.07 & 0.40\tabularnewline
 & 0 & 0.08 & 0.38\tabularnewline
 & 1 & 0.08 & 0.40\tabularnewline
 & 2 & 0.08 & 0.43\tabularnewline
 & 3 & 0.07 & 0.49\tabularnewline
\hline
\end{tabular}\caption{Health and income}
\label{tab:health-income}
\end{table}

Table \ref{tab:summary_stats} reports a set of descriptive statistics
on income and other explanatory variables, described in the next few
paragraphs. In our analysis below, we control for some time-varying
variables that are standard in the literature. First, the number of
children is measured as the number of persons living in the private
household that are age $14$ or less, top-coded at 3 children. In
our sample, 74\% of respondents have no children, and the average
number of children is 0.40. Second, marriage status is a dummy variable
that indicates being married or living together. The majority of the
individuals are married (61\%). Third, we use a self-reported indicator
for labor market status variable that we map onto 4 values: (1) employed,
51\%; (2) unemployed, 5.2\%; (3) retired, 12\% and (4) other, 32\%.
The value ``other'' includes students, permanently disabled or unfit
to work, and fulfilling domestic tasks and care responsibilities.
\begin{table}
\centering{}
\begin{tabular}{llrr}
\hline
 &  & mean & sd\tabularnewline
\hline
health status & (1) bad and very bad & 0.115 & \tabularnewline
 & (2) fair & 0.260 & \tabularnewline
 & (3) good & 0.437 & \tabularnewline
 & (4) very good & 0.187 & \tabularnewline
\hline
\multicolumn{4}{l}{\emph{Time-varying explanatory variables}}\tabularnewline
log income &  & 9.172 & 0.965\tabularnewline
child &  & 0.401 & 0.755\tabularnewline
married &  & 0.614 & \tabularnewline
employment status & employed & 0.512 & \tabularnewline
 & unemployed & 0.052 & \tabularnewline
 & retired & 0.117 & \tabularnewline
 & other & 0.319 & \tabularnewline
\hline
\multicolumn{4}{l}{\emph{Time-invariant explanatory variables}}\tabularnewline
age group & $[18;25]$ & 0.082 & \tabularnewline
 & $]25;35]$ & 0.145 & \tabularnewline
 & $]35;45]$ & 0.188 & \tabularnewline
 & $]45;55]$ & 0.194 & \tabularnewline
 & $]55;65]$ & 0.182 & \tabularnewline
 & $]65;\infty]$ & 0.208 & \tabularnewline
urbanisation & high & 0.388 & \tabularnewline
 & middle & 0.224 & \tabularnewline
 & low & 0.388 & \tabularnewline
male &  & 0.460 & \tabularnewline
educ & no schooling & 0.013 & \tabularnewline
 & primary & 0.143 & \tabularnewline
 & lower secondary & 0.188 & \tabularnewline
 & upper secondary & 0.420 & \tabularnewline
 & post-secondary & 0.038 & \tabularnewline
 & tertiary & 0.199 & \tabularnewline
\hline
$n$ &  & 260601 & \tabularnewline
$T$ &  & 4 & \tabularnewline
$nT$ &  & 1042404 & \tabularnewline
\hline
\end{tabular}\caption{Descriptive statistics}
\label{tab:summary_stats}
\end{table}

In our fixed effects results below, we do not further control for
variables that do not change over the sample period for a given individual.
However, we include a set of time-invariant explanatory variables
when we obtain results for non-fixed effects estimators.\footnote{Coefficient estimates for these variables are omitted from the main
text, and reported in Appendix \ref{sec:Additional-empirical-results}.} Table \ref{tab:summary_stats} provides descriptive statistics for
such variables. The total sample contains slightly more females (54\%)
than males (46\%). The proportion of individuals aged between 18 and
25 is 8.2\%; 21\% of individuals are aged 65 or more. With regards
to education, 1.3\% of the sample have no schooling (0); 14\% have
attended primary school (1); 19\% have lower secondary education (3);
42\% have upper secondary education (4); 3.8\% have post-secondary
education (5) and 20\% have tertiary education (6). Geographically,
39\% of individuals live in areas with a high degree of urbanisation
and 39\% live in areas with low levels of urbanisation.

We estimate the parameters in the dynamic ordered choice model with
fixed effects, with latent variable outcome equation
\begin{align}
SRH_{i,t}^{*} & =\alpha_{i}+\rho1\left\{ SRH_{i,t-1}\geq3\right\} +\beta_{1}\log income_{it}+\beta_{2}child_{it}+\beta_{3}married_{it}+\label{eq:dofe_health}\\
 & \phantom{=}+\beta_{4}unemp_{it}+\beta_{5}retired_{it}+\beta_{6}other_{it}-U_{it}.\nonumber
\end{align}
Regression results are presented in Table \ref{tab:DOFE_results}.
The first four columns (a-d, ``DOLFE'', for \emph{d}ynamic \emph{o}rdered
\emph{l}ogit with \emph{f}ixed \emph{e}ffects) presents the results
for (\ref{eq:dofe_health}) using the estimator described in Section
\ref{sec:Estimation}. Different values of $h$ refer to a bandwidth
parameter that we introduce because one of the explanatory variables
is continuous, as in HK. Column (d) omits the employment variables,
to check whether relationship between employment status and income
matters for estimation of the effect of income on health.

We also present estimation results for different estimators. Results
for the static ordered logit model with fixed effects, i.e. setting
$\rho=0$ in (\ref{eq:dofe_health}), are obtained using the estimator
in Muris (2017), and presented in column (e) (``FEOL''). Column
(f) (``DOL'') estimates a dynamic ordered logit model without fixed
effects, i.e. (\ref{eq:dofe_health}) with $\alpha_{i}=0$. Column
(g) (``OL'') presents results for cross-sectional ordered logit
estimator that does not take into account fixed effects or dynamics
(i.e. $\alpha_{i}=\rho=0$ in (\ref{eq:dofe_health})). We also present
results for a static linear model with (h, ``FELM'') and without
(i, ``LM'') fixed effects. The standard errors for all estimators
are obtained using the bootstrap (500 replications). For the estimators
that are not of the fixed effects type, we additionally control for
education, gender, education level and the level of urbanisation.
DOLFE uses four periods of data, corresponding to $t=0,1,2$. For
comparability, the other dynamic estimator also uses periods 0,1,2;
static estimators use periods 1,2.
\begin{table}
\begin{centering}
{\normalsize{}\hspace*{-1cm}}
\begin{tabular}[b]{lccccccccc}
\hline
 & {\normalsize{}(a)} & {\normalsize{}(b)} & {\normalsize{}(c)} & {\normalsize{}(d)} & {\normalsize{}(e)} & {\normalsize{}(f)} & {\normalsize{}(g)} & {\normalsize{}(h)} & {\normalsize{}(i)}\tabularnewline
 & {\normalsize{}DOLFE} & {\normalsize{}DOLFE} & {\normalsize{}DOLFE} & {\normalsize{}DOLFE} & {\normalsize{}FEOL} & {\normalsize{}DOL} & {\normalsize{}OL} & {\normalsize{}FELM} & {\normalsize{}LM}\tabularnewline
 & {\normalsize{}$h=1$} & {\normalsize{}$h=0.1$} & {\normalsize{}$h=10$} & {\normalsize{}$h=1$} &  &  &  &  & \tabularnewline
\hline
{\normalsize{}log(income)} & {\normalsize{}0.049} & {\normalsize{}-0.047} & {\normalsize{}0.059} & {\normalsize{}0.061} & {\normalsize{}0.020} & {\normalsize{}0.340} & {\normalsize{}0.492} & {\normalsize{}0.003} & {\normalsize{}0.194}\tabularnewline
\multirow{1}{*}[103cm]{} & \multirow{1}{*}{{\tiny{}(0.033)}} & {\tiny{}(0.056)} & {\tiny{}(0.029)} & {\tiny{}(0.029)} & {\tiny{}(0.019)} & {\tiny{}(0.004)} & {\tiny{}(0.004)} & {\tiny{}(0.003)} & {\tiny{}(0.002)}\tabularnewline
{\normalsize{}child} & {\normalsize{}-0.030} & {\normalsize{}0.006} & {\normalsize{}-0.031} & {\normalsize{}-0.026} & {\normalsize{}0.021} & {\normalsize{}0.060} & {\normalsize{}0.089} & {\normalsize{}0.002} & {\normalsize{}0.033}\tabularnewline
 & {\tiny{}(0.051)} & {\tiny{}(0.069)} & {\tiny{}(0.050)} & {\tiny{}(0.049)} & {\tiny{}(0.032)} & {\tiny{}(0.005)} & {\tiny{}(0.005)} & {\tiny{}(0.005)} & {\tiny{}(0.002)}\tabularnewline
{\normalsize{}married} & {\normalsize{}0.139} & {\normalsize{}-0.041} & {\normalsize{}0.157} & {\normalsize{}0.130} & {\normalsize{}0.164} & {\normalsize{}0.073} & {\normalsize{}0.141} & {\normalsize{}0.029} & {\normalsize{}0.062}\tabularnewline
 & {\tiny{}(0.087)} & {\tiny{}(0.119)} & {\tiny{}(0.086)} & {\tiny{}(0.088)} & {\tiny{}(0.053)} & {\tiny{}(0.007)} & {\tiny{}(0.008)} & {\tiny{}(0.009)} & {\tiny{}(0.003)}\tabularnewline
{\normalsize{}unemp} & {\normalsize{}-0.188} & {\normalsize{}-0.230} & {\normalsize{}-0.178} &  & {\normalsize{}-0.196} & {\normalsize{}-0.242} & {\normalsize{}-0.308} & {\normalsize{}-0.033} & {\normalsize{}-0.127}\tabularnewline
 & {\tiny{}(0.070)} & {\tiny{}(0.110)} & {\tiny{}(0.068)} &  & {\tiny{}(0.038)} & {\tiny{}(0.014)} & {\tiny{}(0.015)} & {\tiny{}(0.007)} & {\tiny{}(0.006)}\tabularnewline
{\normalsize{}retired} & {\normalsize{}-0.132} & {\normalsize{}-0.043} & {\normalsize{}-0.139} &  & {\normalsize{}-0.154} & {\normalsize{}-0.050} & {\normalsize{}-0.097} & {\normalsize{}-0.027} & {\normalsize{}-0.047}\tabularnewline
 & {\tiny{}(0.082)} & {\tiny{}(0.119)} & {\tiny{}(0.080)} &  & {\tiny{}(0.041)} & {\tiny{}(0.010)} & {\tiny{}(0.011)} & {\tiny{}(0.007)} & {\tiny{}(0.004)}\tabularnewline
{\normalsize{}other} & {\normalsize{}-0.370} & {\normalsize{}-0.207} & {\normalsize{}-0.369} &  & {\normalsize{}-0.473} & {\normalsize{}-0.771} & {\normalsize{}-1.087} & {\normalsize{}-0.082} & {\normalsize{}-0.460}\tabularnewline
 & {\tiny{}(0.061)} & {\tiny{}(0.087)} & {\tiny{}(0.061)} &  & {\tiny{}(0.040)} & {\tiny{}(0.010)} & {\tiny{}(0.012)} & {\tiny{}(0.007)} & {\tiny{}(0.005)}\tabularnewline
\hline
{\normalsize{}$\rho$} & {\normalsize{}0.733} & {\normalsize{}0.723} & {\normalsize{}0.733} & {\normalsize{}0.734} &  & {\normalsize{}1.987} &  &  & \tabularnewline
 & {\tiny{}(0.020)} & {\tiny{}(0.025)} & {\tiny{}(0.020)} & {\tiny{}(0.017)} &  & {\tiny{}(0.023)} &  &  & \tabularnewline
{\normalsize{}$\gamma_{2}$} & {\normalsize{}-3.275} & {\normalsize{}-3.260} & {\normalsize{}-3.272} & {\normalsize{}-3.211} & {\normalsize{}-3.487} & {\normalsize{}-2.506} & {\normalsize{}-1.992} &  & \tabularnewline
 & {\tiny{}(0.054)} & {\tiny{}(0.068)} & {\tiny{}(0.053)} & {\tiny{}(0.048)} & {\tiny{}(0.015)} & {\tiny{}(0.007)} & {\tiny{}(0.006)} &  & \tabularnewline
{\normalsize{}$\gamma_{4}$} & {\normalsize{}3.326} & {\normalsize{}3.356} & {\normalsize{}3.329} & {\normalsize{}3.321} & {\normalsize{}3.997} & {\normalsize{}3.089} & {\normalsize{}2.603} &  & \tabularnewline
 & {\tiny{}(0.055)} & {\tiny{}(0.076)} & {\tiny{}(0.055)} & {\tiny{}(0.054)} & {\tiny{}(0.024)} & {\tiny{}(0.006)} & {\tiny{}(0.006)} &  & \tabularnewline
\hline
\end{tabular}{\normalsize\par}
\par\end{centering}
\caption{Main results.}

\centering{}\label{tab:DOFE_results}
\end{table}

\textbf{Income. }Our main explanatory variable of interest is income
(log income, coefficient $\beta_{1}$). Across almost all specifications,
we find a positive association between income and self-reported health.
The only exception is column (b), where the point estimate is negative,
and about the same magnitude as the standard error.

Controlling for unobserved heterogeneity leads to a very strong reduction
in the magnitude of the association. For example, for the static case,
a comparison of columns (e) and (g) says that, for the static case,
controlling for unobserved heterogeneity reduces the coefficient on
income by more than a factor 20. For this comparison, note that the
threshold differences increase, suggesting that the scale increases;
compare also the coefficients on the other variables, with an unchanged
order of magnitude. We are not the first to observe a limited association
between income and self-reported health. In a review of the literature,
Gunasekara et al. (2011) found a small positive link between income
and self-reported health, which is reduced when controlling for unmeasured
confounders. Interestingly, Johnston et al. (2009) found no link between
self-reported hypertension and income; an association that, however,
became positive when using objective measures of hypertension.

The estimated effect of income also changes when we control for state
dependence. Comparing columns (f) and (g), we see that controlling
for state dependence in a model without unobserved heterogeneity reduces
the association between income and self-reported health. So, individually
controlling for unobserved heterogeneity or for dynamics reduces the
magnitude of the association between health and income.

Finally, a comparison between columns (a) and (d) shows that the estimate
for income association is robust to controlling for employment status,

\textbf{State dependence. }We estimate an autoregressive parameter
of around 0.75, with threshold differences of about $3$. The estimated
ratio of $\rho$ to the thresholds (which measure the distance from
category 3) are much lower than for column (f). This confirms the
importance of controlling for unobserved heterogeneity, which reduces
the estimated magnitude of persistence by a factor 3. Nevertheless,
even when controlling for unobserved (and observed) heterogeneity,
we find strong evidence for large, positive persistence in self-reported
health.

There are at least two ways to get a sense of the magnitude of persistence.
The first approach, also available for binary choice methods, is to
compare estimates of $\rho$ to estimates of regression coefficients.
For example, in our preferred specification in column (a), a health
shock that lifts you from any category below 3, to category 3 or 4,
has an impact on future health that is almost 4 times that of becoming
unemployed. The impact is more than 5 times that of marrying.

The second approach to interpreting estimates of $\rho$ uses the
estimated thresholds to obtain an estimate similar to a linear autoregressive
model.\footnote{This approach is not available for binary choice models because threshold
parameters are not available. } Differences between the thresholds are a measure of the distances
between two categories. If $\gamma_{2}=-\gamma_{4}$, then categories
2 and 3 are as far apart as categories 3 and 4. In such a case, a
linear model may yield similar results in terms of partial effects.
In this case, $-\rho/\gamma_{2}$ and $\rho/\gamma_{3}$ can be interpreted
as linear regression coefficients for that category; we find that
they are about 0.25. Said differently, we find the analog of an AR(1)
coefficient of 0.25 in a linear model.

\textbf{Other time-varying covariates.} The literature so far has
been inconclusive on how retirement is associated with health. On
one hand, retiring allows more time for health-promoting activities,
and reduces work-related stress. On the other hand, people may lose
traction and motivation and may become less active. Therefore, while
Coe and Zamarro (2011) find that retirement improves health, Behncke
(2012) finds an increase in the likelihood of disease following retirement.
In our DOLFE model, the coefficient is statistically insignificant.
This suggests that the association between retirement and health may
not be as strong as previously thought. Compared to the FEOL model,
the effect of retirement on health disappears when controlling for
state dependence.

The extensive literature on the link between unemployment and health
in particular, and economic conditions and health more generally,
is broadly inconclusive. Some studies have suggested a protective
role of unemployment on health (Ruhm, 2000), while others suggest
that unemployment is detrimental for health (McInerney and Mellor,
2012). Our results appear to be more in line with the findings of
Ruhm (2015) and B�ckerman and Ilmakunnas (2009). In DOLFE, the coefficient
of being unemployed is negative and statistically significant. Controlling
for state dependence does not change things compared to the FEOL model.

Having children is insignificant in our DOLFE model, while previous
studies have provided mixed findings on this question (Mckenzie and
Carter, 2013; Evenson and Simon, 2005). This is also insignificant
in the FEOL model, suggesting that previous findings on having children
might have been driven by unobserved heterogeneity.

Being married is generally considered a protective factor for health
(Kaplan and Kronick, 2006; Molloy et al., 2009). In our model, however,
it is statistically insignificant -- as opposed to the FEOL model
where it was positive and significant. Thus, controlling for state
dependence appears to be important for this variable.

\section{Conclusion}

This paper studies a fixed$-T$ dynamic ordered logit model with fixed
effects (DOLFE) and is the first to provide identification and estimation
results for all common parameters in a dynamic ordered logit model
with fixed effects and a fixed number of time periods. The results
require only four time periods of data on the ordinal outcome variable.
We demonstrate identification of the autoregressive coefficients on
the lagged dependent variable, the regression coefficients on the
exogenous regressors, and differences of the threshold parameters.
The latter makes it possible to interpret the magnitude of the coefficients.

Including fixed effects and state dependence in the model is particularly
relevant for self-reported health, a measure that is widely used in
the literature. Future research using self-reporting health can benefit
from our model for two main reasons. First, controlling for fixed
effects, one can take into account unobserved heterogeneity (Carro
and Traferri, 2014; Halliday, 2008; Fern�ndez-Val et al., 2017), which
is especially important due to differences in understanding and reporting
health status (Groot, 2000; Sen, 2002; Jylh� et al., 1998; Baron-Epel
et al., 2005; J�rges, 2007). Second, it incorporates elements of persistence
(Contoyannis et al., 2004; Ohrnberger et al., 2017; Hern�ndez-Quevedo
et al., 2008; Roy and Schurer, 2013) or adaptation (Cub�-Moll� et
al., 2017; Daltroy et al., 1999; Damschroder et al., 2005; Heiss et
al., 2014) by controlling for state dependence (Carro and Traferri,
2014; Fern�ndez-Val et al., 2017; Halliday, 2008). Thus, using our
estimator addresses such biases often present in studies using self-rated
health (Davillas et al., 2017).

We thus applied the new dynamic ordered logit model with fixed effects
to investigate the determinants of self-reported health, focusing
on the link between income and health in a panel of European countries.
We found that when controlling for unobserved heterogeneity, the association
between income and health becomes statistically insignificant. This
is in line with studies that have found a smaller or insignificant
association when using fixed effects (Gunasekara, 2011; Larrimore,
2011) -- while other studies have suggested a positive association
between the two (Carrieri and Jones, 2017; Ettner, 1996; Frijters
et al., 2005; Mackenbach et al., 2005). Being retired or married also
becomes statistically insignificant in our model when controlling
for state dependence. Being unemployed or having children does not
appear to be associated with self-reported health in our model.

Our empirical results suggest that persistence plays a positive and
significant role in one's self-reported health. In other words, one's
health is dependent on the health in the previous period, which is
a reasonable thing to expect, as health problems may expand over a
number of periods, or become permanent. This element reflects persistence
of health status over time (Contoyannis et al., 2004). Furthermore,
what is particularly interesting is that in our data, self-reported
health tends to improve, on average, over time - even as people become
four years older during the study period - and being older is typically
associated with worse health outcomes. Therefore, it is reasonable
to believe that this improvement in self-reported health is often
subjective, and does not necessarily reflect one's objective health
level. This element might reflect adaptation to health problems: Even
though one's health does not improve, they adapt to their situation
and therefore report better health (Cub�-Moll� et al 2017; Daltroy
et al., 1999; Damschroder et al., 2005). This second element is a
typical bias when using self-reported health outcomes, and our model
helps correct such biases by introducing the dynamic element to a
fixed effects ordered model. Another interesting finding is that,
when controlling for unobserved heterogeneity, the link between income
and health becomes statistically insignificant, suggesting that other
factors might explain the association between the two.

Overall, measurement bias in studies using self-reported outcomes
often poses challenges to research, that may discourage the use of
such variables. Our model addresses these biases, and thus provides
a basis for more choices in conducting research with databases that
provide such variables.
\begin{thebibliography}{10}
\bibitem{key-29}Albarran, P., R. Carrasco, and J.M. Carro, 2019.
Estimation of Dynamic Nonlinear Random Effects Models with Unbalanced
Panels. Oxford Bulletin of Economics and Statistics 81 (6), 1424--1441.

\bibitem{key-37}Aristodemou, E., 2020. Semiparametric identification
in panel data discrete response models. Journal of Econometrics, forthcoming.

\bibitem{key-28}Baetschmann, G., K.E. Staub, and R. Winkelmann, 2015.
Consistent estimation of the fixed effects ordered logit model. Journal
of the Royal Statistical Society A, 178 (3), 685--703.

\bibitem{key-1}Banerjee, A., A. Deaton, and E. Duflo, 2004. Wealth,
health, and health services in rural Rajasthan. American Economic
Review, 94 (2), 326--330.

\bibitem{key-2}Baron-Epel, O., G. Kaplan, A. Haviv-Messika, J. Tarabeia,
M.S. Green and D.N. Kaluski, 2005. Self-reported health as a cultural
health determinant in Arab and Jewish Israelis: MABAT---National
Health and Nutrition Survey 1999--2001. Social Science and Medicine,
61 (6), 1256--1266.

\bibitem{key-1}Bago d'Uva, T., Van Doorslaer, E., Lindeboom, M. and
O'Donnell, O., 2008. Does reporting heterogeneity bias the measurement
of health disparities?. Health Economics, 17(3), 351--375.

\bibitem{key-1-1}Behncke, S., 2012. Does retirement trigger ill health?
Health Economics, 21(3), 282--300.

\bibitem{key-2-2}B�ckerman, P. and P. Ilmakunnas, 2009. Unemployment
and self-assessed health: evidence from panel data. Health Economics,
18 (2), 161--179.

\bibitem{key-3}Bound, J. and T. Waidmann, 1992. Disability transfers,
self-reported health, and the labor force attachment of older men:
Evidence from the historical record. The Quarterly Journal of Economics,
107 (4), 1393--1419.

\bibitem{key-4}Bradford, W.D. and P. Dolan, 2010. Getting used to
it: the adaptive global utility model. Journal of Health Economics,
29 (6), 811--820.

\bibitem{key-5}Butler, J.S., R.V. Burkhauser, J.M. Mitchell, and
T.P. Pincus, 1987. Measurement error in self-reported health variables.
Review of Economics and Statistics, 69(4), 644--650.

\bibitem{key-2}Caroli, E. and Weber-Baghdiguian, L., 2016. Self-reported
health and gender: The role of social norms. Social Science and Medicine,
153, 220--229.

\bibitem{key-1}Carrieri, V. and Jones, A.M., 2017. The income--health
relationship \textquoteleft beyond the mean\textquoteright : New evidence
from biomarkers. Health Economics, 26(7), 937--956.

\bibitem{key-39}Carro, J.M., 2007. Estimating Dynamic Panel Data
Discrete Choice Models with Fixed Effects. Journal of Econometrics,
140 (2), 503--528.

\bibitem{key-38}Carro, J.M. and A. Traferri, 2014. State dependence
and heterogeneity in health using a bias-corrected fixed-effects estimator.
Journal of Applied Econometrics, 29 (2), 181--207.

\bibitem{key-3}Chamberlain, G., 1980. Analysis of Covariance with
Qualitative Data. The Review of Economic Studies, 47 (1), 225--238.

\bibitem{key-4}Chamberlain, G., 1985. Heterogeneity, omitted variable
bias, and duration dependence.\textquotedblright{} In: ``Longitudinal
Analysis of Labor Market Data'', edited by J. J. Heckman and B.S.
Singer, 3--38. Econometric Society Monographs. Cambridge University
Press.

\bibitem{key-7}Clark, A.E., Frijters, P. and Shields, M.A., 2008.
Relative income, happiness, and utility: An explanation for the Easterlin
paradox and other puzzles. Journal of Economic Literature, 46(1),
95--144.

\bibitem{key-11-1}Clark, A., Y. Georgellis, and P. Sanfey, 2001.
Scarring: The psychological impact of past unemployment. Economica,
68 (270), 221--241.

\bibitem{key-3-1}Coe, N.B. and G. Zamarro, 2011. Retirement effects
on health in Europe. Journal of Health Economics, 30 (1), 77--86.

\bibitem{key-12-1}Contoyannis, P., A.M. Jones, and N. Rice, 2004.
The dynamics of health in the British Household Panel Survey. Journal
of Applied Econometrics, 19(4), 473--503.

\bibitem{key-35}Cox, D. R., 1958. The regression analysis of binary
sequences. Journal of the Royal Statistical Society, 20, 215--242.

\bibitem{key-6}Cub�-Moll�, P., M. Jofre-Bonet, and V. Serra-Sastre,
2017. Adaptation to health states: Sick yet better off?. Health Economics,
26 (12), 1826--1843.

\bibitem{key-7}Daltroy, L.H., M.G. Larson, H.M. Eaton, C.B. Phillips,
and M.H. Liang, 1999. Discrepancies between self-reported and observed
physical function in the elderly: the influence of response shift
and other factors. Social Science and Medicine, 48 (11), 1549--1561.

\bibitem{key-15}Damschroder, L.J., B.J. Zikmund-Fisher, and P.A.
Ubel, 2005. The impact of considering adaptation in health state valuation.
Social Science and Medicine, 61 (2), 267--277.

\bibitem{key-23}Das, M., and A. van Soest, 1999. A panel data model
for subjective information on household income growth. Journal of
Economic Behavior and Organization, 40 (4), 409--426.

\bibitem{key-6}Davillas, A., A.M. Jones, and M. Benzeval, 2017. The
income-health gradient: Evidence from self-reported health and biomarkers
using longitudinal data on income. ISER Working Paper Series 2017-03.

\bibitem{key-8}Ettner, S.L., 1996. New evidence on the relationship
between income and health. Journal of Health Economics, 15(1), 67--85.

\bibitem{key-41}Eurostat (2017). Methodological Guidelines and Description
of EU-SILC Target Variables. 2016 operation.

\bibitem{key-4-1}Evenson, R.J. and R.W. Simon, 2005. Clarifying the
relationship between parenthood and depression. Journal of Health
and Social Behavior, 46(4), 341--358.

\bibitem{key-40}Fern�ndez-Val, I., Y. Savchenko, and F. Vella, 2017.
Evaluating the Role of Income, State Dependence and Individual Specific
Heterogeneity in the Determination of Subjective Health Assessments.
Economics and Human Biology, 25, 85--98.

\bibitem{key-6-1}Frijters, P., Haisken-DeNew, J.P. and Shields, M.A.,
2005. The causal effect of income on health: Evidence from German
reunification. Journal of Health Economics, 24(5), 997--1017.

\bibitem{key-10}Gravelle, H. and M. Sutton, 2009. Income, relative
income, and self-reported health in Britain 1979--2000. Health Economics,
18 (2), 125--145.

\bibitem{key-13-1}Groot, W., 2000. Adaptation and scale of reference
bias in self-assessments of quality of life. Journal of Health Economics,
19 (3), 403--420.

\bibitem{key-7}Gunasekara, F., K. Carter, and T. Blakely, 2011. Change
in income and change in self-rated health: Systematic review of studies
using repeated measures to control for confounding bias. Social Science
and Medicine, 72 (2), 193--201.

\bibitem{key-36}Hahn, J., 2001. The Information Bound of a Dynamic
Panel Logit Model with Fixed Effects. Econometric Theory 17 (5), 913--932.

\bibitem{key-1}Halliday, T.J., 2008. Heterogeneity, state dependence
and health. The Econometrics Journal, 11 (3), 499--516.

\bibitem{key-34}Heckman, J.J., 1981. The incidental parameters problem
and the problem of initial conditions in estimating a discrete time--discrete
data stochastic process. In: Structural Analysis of Discrete Data
with Econometric Applications, Manski CF, McFadden D (eds). MIT Press:
Cambridge, MA, 179--195.

\bibitem{key-5}Heiss, Florian, Steven F Venti, and David A Wise.
The Persistence and Heterogeneity of Health among Older Americans,
NBER Working Paper 20306.

\bibitem{key-3}Hern�ndez-Quevedo, C., A.M. Jones, and N. Rice, 2008.
Persistence in health limitations: A European comparative analysis.
Journal of Health Economics, 27 (6), 1472--88.

\bibitem{key-20}Honor�, B. E., and E. Kyriazidou, 2000. Panel Data
Discrete Choice Models with Lagged Dependent Variables. Econometrica,
68 (4), 839--874.

\bibitem{key-22}Honor�, B. E., and E. Kyriazidou, 2019. Identification
in Binary Response Panel Data Models: Is Point-Identification More
Common Than We Thought? Annals of Economics and Statistics, 134, 207--226.

\bibitem{key-1}Honor�, B. E., and Martin Weidner, 2020. Moment Conditions
for Dynamic Panel Logit Models with Fixed Effects. ArXiv:2005.05942.

\bibitem{key-13}Johnston, D.W., C. Propper, and M.A. Shields, 2009.
Comparing subjective and objective measures of health: Evidence from
hypertension for the income/health gradient. Journal of Health Economics,
28 (3), 540--552.

\bibitem{key-12}J�rges, H., 2007. True health vs response styles:
exploring cross-country differences in self-reported health. Health
Economics, 16 (2), 163--178.

\bibitem{key-11}Jylh�, M., J.M. Guralnik, L. Ferrucci, J. Jokela,
and E. Heikkinen, 1998. Is self-rated health comparable across cultures
and genders? The Journals of Gerontology Series B: Psychological Sciences
and Social Sciences, 53 (3), S144--S152.

\bibitem{key-5-1}Kaplan, R.M. and R.G. Kronick, 2006. Marital status
and longevity in the United States population. Journal of Epidemiology
and Community Health, 60 (9), 760--765.

\bibitem{key-18}Kapteyn, A., J.P. Smith, and A. van Soest, 2007.
Vignettes and self-reports of work disability in the United States
and the Netherlands. American Economic Review, 97 (1), 461--473.

\bibitem{key-17}King, G., C.J. Murray, J.A. Salomon, and A. Tandon,
2004. Enhancing the validity and cross-cultural comparability of measurement
in survey research. American Political Science Review, 98 (1), 191--207.

\bibitem{key-4}Khan, S., M. Ponomareva, and E. Tamer, 2020. Identification
of dynamic binary response models. Working paper.

\bibitem{key-8}Larrimore, J., 2011. Does a higher income have positive
health effects? Using the earned income tax credit to explore the
income-health gradient. The Milbank Quarterly, 89(4), 694--727.

\bibitem{key-31}Layard, R., 2006. Happiness and public policy: A
challenge to the profession. The Economic Journal, 116 (510), C24--C33.

\bibitem{key-10}Lindeboom, M. and E. van Doorslaer, 2004. Cut-point
shift and index shift in self-reported health. Journal of Health Economics,
23 (6), 1083--1099.

\bibitem{key-9}Mackenbach, J.P., Martikainen, P., Looman, C.W., Dalstra,
J.A., Kunst, A.E. and Lahelma, E., 2005. The shape of the relationship
between income and self-assessed health: an international study. International
Journal of Epidemiology, 34(2), 286--293.

\bibitem{key-1}Magnac, T., 2000. Subsidised Training and Youth Employment:
Distinguishing Unobserved Heterogeneity from State Dependence in Labour
Market Histories. The Economic Journal, 110 (466), 805-837.

\bibitem{key-9}McInerney, M. and J.M. Mellor, 2012. Recessions and
seniors\textquoteright{} health, health behaviors, and healthcare
use: Analysis of the Medicare Current Beneficiary Survey. Journal
of Health Economics, 31 (5), 744--751.

\bibitem{key-10-1}Mckenzie, S.K., and K. Carter, 2013. Does transition
into parenthood lead to changes in mental health? Findings from three
waves of a population based panel study. Journal of Epidemiological
Community Health, 67 (4), 339--345.

\bibitem{key-9-1}Molloy, G.J., E. Stamatakis, G. Randall, and M.
Hamer, 2009. Marital status, gender and cardiovascular mortality:
behavioural, psychological distress and metabolic explanations. Social
Science and Medicine, 69 (2), 223--228.

\bibitem{key-24}Muris, C., 2017. Estimation in the fixed-effects
ordered logit model. Review of Economics and Statistics, 99 (3), 465--477.

\bibitem{key-2}Ohrnberger, J., E. Fichera, and M. Sutton, 2017. The
dynamics of physical and mental health in the older population. The
Journal of the Economics of Ageing, 9, 52--62.

\bibitem{key-1}R Core Team, 2020. R: A language and environment for
statistical computing. R Foundation for Statistical Computing, Vienna,
Austria.

\bibitem{key-4}Roy, J., and S. Schurer, 2013. Getting stuck in the
blues: Persistence of mental health problems in Australia. Health
Economics, 22 (9), 1139--1157.

\bibitem{key-7-1}Ruhm, C.J., 2000. Are recessions good for your health?
The Quarterly Journal of Economics, 115 (2), 617--650.

\bibitem{key-8-1}Ruhm, C.J., 2015. Recessions, healthy no more? Journal
of Health Economics, 42, 17--28.

\bibitem{key-15-1}Salomon, J.A., A. Tandon, and C.J. Murray, 2004.
Comparability of self rated health: cross sectional multi-country
survey using anchoring vignettes. BMJ, 328 (7434), 258.

\bibitem{key-14-1}Sen, A., 2002. Health: perception versus observation.
BMJ, 324 (7342), 860--861.

\bibitem{key-2}Shi, X., M. Shum, and W. Song, 2018. Estimating semi-parametric
panel multinomial choice models using cyclic monotonicity. Econometrica,
86 (2), 737--761.

\bibitem{key-8}Wilson, T.D. and D.T. Gilbert, 2008. Explaining away:
A model of affective adaptation. Perspectives on Psychological Science,
3 (5), 370--386.
\end{thebibliography}