Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
67,144 characters · 6 sections · 0 citation commands
A dynamic ordered logit model with fixed effects
Certain individual-level conditions may tend to persist over time, in the sense that a condition has a memory of a previous period\textquoteright s state, or may involve an element of adaptation. Furthermore, the way individuals experience the same condition may vary, and they may also have a different understanding of how this is measured. A common example that fulfils these characteristics, and which is used extensively in the literature, is self-reported health status. Health status often depends on its value in the previous period, as health conditions may persist over time, given that recovery can take long, and that an illness may even have permanent effects. For example, Table (ref) presents a transition matrix for self-reported health status in the United Kingdom for the period 2003-2016.\footnote{More information about the data and source is in Section (ref), where we analyze this data using the methodology proposed in this paper.}
For individuals that report a value of current health in a given year (rows, on a 5-point scale with 5 being the highest), it shows the relative proportion of those that report a certain value in the subsequent year (columns). A striking feature of this transition matrix is that a lot of mass is on or near the main diagonal. This feature is found across all countries in our analysis. In other words, self-reported health status is persistent: individuals tend to stay in the same level of health.
There are at least two explanations for this observed persistence (Heckman, 1981; Honor� and Kyriazidou, 2000): unobserved heterogeneity and state dependence. Consider first unobserved heterogeneity. It refers to unobservable characteristics that affect the propensity to report higher health status. Unobserved heterogeneity is important in the literature on health status, because self-reported health has been used extensively in the literature as a measure of health outcomes (see for example Bound and Waidmann, 1992; Banerjee et al., 2004; Gravelle and Sutton, 2009; McInerney and Mellor, 2012). It is often viewed as a limitation that self-reported measures are subjective. For example, reporting one's own health may depend on cultural factors (Jylh� et al., 1998; Baron-Epel et al., 2005; J�rges, 2007), and people may have a different understanding of reference points for health (Groot, 2000; Sen, 2002). Previous studies have used vignettes to address cross-country differences in reporting of health and disability (King et al., 2004; Salomon et al., 2004; Kapteyn et al., 2007). However, the issue with unobserved heterogeneity across individuals remains. As a result, it is important to take into account the role of unobserved heterogeneity when analyzing self-reported health data. The appropriate econometric approach to this is to allow for fixed effects.\footnote{Studies have long debated the accuracy and reliability of subjective measures of health, such as self-reported health status (see for example Butler et al., 1987; Lindeboom and van Doorslaer, 2004; Johnston et al., 2009). As an alternative response to these concerns, a number of objective measures of health have been included in household surveys. These include blood pressure, BMI, the number of medicines taken (Health Survey England, 2019), the number of sick days off work, the number of days hospitalised (BHPS, 2019) etc. Some household surveys ask respondents to perform a task such as walking across the room or buttoning a shirt to capture any limitations (SHARE, 2019). Indexes such as the EQ-5D index are being used to cover different types of conditions and merge them into a single measure. The Euro-D scale measures mental health, and the CASP-12 index captures quality of life. However, these objective measures are often very specific to particular diseases, and even when creating a relevant index it may be impossible to include and accurately reflect all conditions. As such, while objective measures may accurately capture some health conditions, they have serious limitations in capturing the overall picture of one\textquoteright s health.}
A number of studies have found a positive association between income and health (Carrieri and Jones, 2017; Ettner, 1996; Frijters et al., 2005; Mackenbach et al., 2005), but empirical evidence of a strong effect is sometimes limited (Larrimore, 2011; Gunasekara, 2011; Johnston et al., 2009). Nevertheless, the literature on the impact of economic downturns (which mean reduced income) has previously demonstrated positive effects of unemployment on health (Ruhm, 2000; Ruhm and Black, 2002), and more recently, no effect (Ruhm, 2015). What also appears to matter, at least in terms of happiness, apart from absolute income, is also relative income, i.e. how one\textquoteright s income compares to that of those around them (Frijters et al., 2008). The relationship between income and health is endogenous and complex, and both can be correlated with other factors, that are not always measured and included in empirical models. For example, Gunasekara et al. (2011) found that when controlling for unmeasured confounders, the association between the two becomes weaker. With regards to self-reported hypertension in particular, Johnston et al., (2009) found no link to income -- something that did change when using objective measures. In our paper, using models that do not control for individual unobserved heterogeneity yields a positive and statistically significant coefficient (Table 3). However, this becomes insignificant when using fixed effects, suggesting that there are other factors that potentially drive the association between the two variables.
Consider now the second explanation for the observed persistence in health outcomes: state dependence. It refers to the possibility that past self-reported health status may be related to current self-reported health status even after conditioning on unobserved heterogeneity. State dependence arises if actual (as opposed to self-reported) health shocks are persistent, in the sense that a shock on health can have a long-lasting effect (a typical example is injury leading to disability). Contoyannis et al. (2004), using a random effects approach, found evidence for such persistence in respondents of the British Household Panel survey.
State dependence in self-reported health can also arise due to adaptation: self-reported health status may change over time for a person whose actual health has not changed. People tend to adapt to good or bad developments in life. According to the Global Adaptive Utility Model, individuals reallocate weights on various domains of life in order to maintain their previous level of utility (Bradford and Dolan, 2010). Similarly, the AREA model developed by Wilson and Gilbert (2008), suggests that attention is focused on a change, followed by reaction, explanation, and, finally, adaptation. This also applies to health, as health status tends to improve even when individuals' health has actually not experienced any objective change (Daltroy et al., 1999; Damschroder et al., 2005), and time since diagnosis is positively associated with self-reported health (Cub�-Moll� et al., 2017). Whether persistence or adaptation, or both, characterise a variable, this calls for a dynamic element in a model.
Overall, the challenges with studying self-reported health status is that (a) people with the same actual health status might be reporting different health levels; and (b) health shocks can have a lasting effect. Against this background, we propose and analyze a panel data ordered logit model that includes both fixed effects and a lagged dependent variable. This allows a researcher faced with panel data and an ordinal outcome variable to disentangle unobserved heterogeneity from state dependence, and to quantify state dependence. Thus, we address the limitations of using self-reported health as a proxy for individuals' health. Our contribution is important for studies using subjective health measures as it can help correct biases that naturally occur when using this type of measure.\footnote{For example, happiness is perceived and reported differently across individuals and people adapt to things that make them happy (Layard, 2006), while shocks on happiness can have a scarring effect on next periods (Clark et al., 2001).}
Specifically, we study the dynamic ordered logit model with fixed effects:
where $2\leq k\leq J$ is a fixed and known cutoff for the lagged dependent variable. The person-specific parameter $\alpha_{i}$ captures unobserved heterogeneity, which we allow to be correlated with the other quantities in the model in an unrestricted way (fixed effects). The time-varying covariates $X_{i,t}$ are collected across time periods in $X_{i}=\left(X_{i,1},X_{i,2},X_{i,3}\right)$, and the lagged dependent variables for period $t$ are collected in $Y_{i,<t}=\left(Y_{i,0},\cdots,Y_{i,t-1}\right)$. The autoregressive parameter $\rho$ is the regression coefficient on the lagged dependent variable $1\left\{ Y_{i,t-1}\geq k\right\} $; $\beta$ is the regression coefficient on the contemporaneous covariates; and the threshold parameters $\gamma_{j}$ map the underlying latent variable $Y_{i,t}^{*}$ into the observed ordered outcome $Y_{i,t}$. Equation ((ref)) restricts the error terms $U_{i,t}$ to be i.i.d. logistic, and is a strict exogeneity assumption on the regressors and past outcomes.\footnote{The dynamics in our model are restricted to depend on $Y_{i,t}$ through $1\left\{ Y_{i,t-1}\geq k\right\} $ only. An alternative model for which we can identify some features is one that is linear in its history, i.e. $Y_{i,t}^{*}=\alpha_{i}+X_{i,t}\beta+\rho Y_{i,t-1}-U_{i,t}.$ We were unable to use our approach to obtain identification in the more general model with $Y_{i,t}^{*}=\alpha_{i}+X_{i,t}\beta+\sum_{j=2}^{J}\rho_{j}1\left\{ Y_{i,t-1}=j\right\} -U_{i,t}$.}
This model combines a number of noteworthy features. First, it is a model for discrete ordered outcomes, and therefore a nonlinear model. Second, it is dynamic, in the sense that the current outcome depends directly on the outcome in the previous period. This feature, called state dependence, is governed by the autoregressive parameter $\rho$. Third, it allows for unobserved heterogeneity in an unrestricted way, i.e. it is a fixed effects model. Fourth, the model is only specified for a small number of time periods, $T=3$. Period 0 is unmodelled, but an observation on the outcome variable in time 0 is required for identification.
We believe that we are the first to provide identification and estimation results for all common parameters in a dynamic ordered logit model with fixed effects and a fixed number of time periods. Using four time periods of data on the ordinal outcome variable, we identify the autoregressive coefficients on the lagged dependent variable, and the regression coefficients on the exogenous regressors. We also identify the threshold parameters, which makes it possible to interpret the magnitude of the estimated coefficients. This distinguishes the ordered choice model from the dynamic binary choice model with fixed effects, where such an interpretation is not available. Our identification result suggest a composite conditional maximum likelihood estimator for the parameters in our model. We establish conditions under which that estimator is consistent and asymptotically normal.
We use our estimator to investigate the determinants of self-reported health, focusing on the link between income and health in a panel of European countries over the period 2003-2016. We obtain two main findings. First, even after controlling for unobserved heterogeneity, persistence plays a positive and significant role in one's self-reported health. In other words, one's health is dependent on the health in the previous period, which is a reasonable thing to expect, as health problems may expand over a number of periods, or become permanent. Quantitatively, we estimate a persistence parameter that is analogous to an autoregressive parameter of about 0.25 in a linear AR(1) model. Second, we find that, when controlling for unobserved heterogeneity, the link between income and health becomes statistically insignificant, suggesting that other factors might explain the association between the two. This is in line with studies that have found a smaller or insignificant association when using fixed effects (Gunasekara, 2011; Larrimore, 2011).
We believe that our paper is the first to provide identification and estimation results for a panel data model with (i) ordered outcomes; (ii) a lagged dependent variable; (iii) fixed effects; and (iv) a fixed number of time periods. Our econometric contribution is related to several strands of literature, each of which features a subset of these features.
Most closely related to our paper is the literature on binary and multinomial choice models with fixed effects and lagged dependent variables, which features all but (i). The seminal work by Honor� and Kyriazidou (2000) builds on Cox (1958) and Chamberlain (1985) to estimate the parameters in dynamic binary choice logit model with fixed effects and time-varying regressors. Hahn (2001) discusses the information bound for a special case of their model. Honor� and Kyriazidou (2019) discuss identification of some closely related models. Honor� and Weidner (2020) construct moment conditions that shed light on identification in this and related models, and provide a $\sqrt{n}$-consistent estimator. Honor� and Tamer (2006), Aristodemou (2020) and Khan et al. (2020) obtain results for models that do not have logistic errors. For the static multinomial model, Chamberlain (1980) studies the logit case; Shi et al. (2008) provides results for the general static; and Magnac (2000) studies the dynamic version. We supplement these results by showing that, in an ordered choice model, the thresholds in the latent variable model can be identified along with the regression coefficients and the autoregressive parameter. This allows for a quantitative interpretation of true state dependence. Such an interpretation is not available in the binary and multinomial choice models.
The literature on static ordered logit models with fixed effects features all but (ii). This model was analyzed by Das and van Soest (1999), Baetschmann et al. (2015), and Muris (2017). Our result differs from the results in those papers, because we provide results for a dynamic version of the ordered logit model.
The literature on random effects dynamic ordered choice models features all but (iii). Random effects dynamic ordered choice models have been studied and applied extensively (Contoyannis et al., 2004; Albarran et al., 2019). Such approaches require strong restrictions on the relationship between the unobserved heterogeneity and the exogeneous variables in the model. Such restrictions are usually unappealing to economists, as evidenced by the fact that they are rarely used in linear models. Our approach does not impose random effects restrictions and is the first to provide a fixed effects approach for dynamic ordered choice models.
Note that our approach is fixed-$T$ consistent. The difficulty of allowing for fixed effects is alleviated when one can assume that $T\to\infty$, referred to as “large-$T$”. Large-T fixed effects dynamic ordered choice models have been studied by Carro and Traferri (2014) and Fern�ndez-Val et al. (2017), see also Carro (2007) for the binary outcome case. In the large-$T$ case, one can use techniques that correct for the bias that comes from including fixed effects in the nonlinear panel model. This approach does not feature (iv). These techniques are not appropriate for our empirical application, which is a rotating panel with $T=4$.
One limitation of our approach is that we restrict the way in which the lagged dependent variable enters the model. The random effects and large-$T$ approach can accommodate a richer dynamic specification. We leave for future work whether such an extension is possible with a fixed-$T$ fixed-effects approach.
We normalize $\gamma_{k}=0$, where $k$ is as in equation ((ref)). This scale normalization is without loss of generality because the scale of $\alpha_{i}$ is unrestricted. Our model implies that the binary variable $D_{i,t}(k)=1\left\{ Y_{i,t}\geq k\right\} $ follows the dynamic binary choice logit model in Honor� and Kyriazidou (2000), HK hereafter. Specifically, equation (3) in HK applies to the transformed model \[ D_{i,t}(k)=1\left\{ X_{i,t}\beta+\rho D_{i,t-1}\left(k\right)+\alpha_{i}-U_{i,t}\geq0\right\} , \] i.e. the transformed model follows a dynamic binary choice logit model with fixed effects. The implied conditional probabilities relevant for our analysis are
and, for $t=1,2,3$,
where we have let $D_{i,<t}\left(k\right)=\left(D_{i,0}\left(k\right),\cdots,D_{i,t-1}\left(k\right)\right)$. HK provide conditions that guarantee identification of $\beta$ and $\rho$ by constructing a conditional probability that features $\left(\beta,\rho\right)$ but that is free of $\alpha_{i}$.
If $Y_{i,t}$ has at least three points of support, there is information in $Y_{it}$ beyond $D_{it}\left(k\right)$. In the remainder of this section, we show that this information can be used to identify the threshold parameters \[ \gamma\equiv\left(\gamma_{2},\gamma_{3},\cdots,\gamma_{k-1},\gamma_{k+1},\cdots,\gamma_{J}\right). \] This leads to an interpretation of the magnitude of $\left(\beta,\rho\right)$ that is not available for the dynamic binary choice model. Muris (2017, Section III.C) discusses this for the static panel data ordered choice models ($\rho=0$).
We now construct a conditional probability that features $\left(\beta,\rho,\gamma\right)$ but not the incidental parameters $\alpha_{i}$. To this end, extend the definition \[ D_{i,t}\left(j\right)=1\left\{ Y_{i,t}\geq j\right\} ,\,2\leq j\leq J, \] to thresholds $j\neq k$, and abbreviate $D_{i,t}\equiv D_{i,t}\left(k\right)$. Define the events $\left(A_{j,l},B_{j,l},C_{j,l}\right)$, with $2\leq j\leq k\leq l\leq J$,\footnote{Choosing $j\leq k$ guarantees that when $D_{i,2}\left(j\right)=0,$ the lagged dependent variable in period 3 is 0. The opposite is true for $l\geq k$ and $D_{i,2}\left(l\right)=1$. There would be no gain from considering a threshold different from $k$ in the first period. Using $k$ as the threshold in the second period is the only way to cancel out the threshold parameters from period 1. Given that we are using the subpopulation $X_{i,2}=X_{i,3}$, the fact that $j,l$ are used alternately in periods 2 and 3 does not create additional difficulties.} as follows:
For $d_{0}=d_{3}=0$, the event $A_{j,l}$ corresponds to moving up in the middle periods $t=1,2$, starting below $k$ to moving up to at least $l\geq k$. The event $B_{j,l}$ corresponds to moving down in the middle periods, starting from at least $k$ and moving below $j\leq k$.
If $j=k=l$, the event $C_{k,k}$ corresponds to switchers (observations with $D_{i1}+D_{i2}=1$), as in HK. In the ordered model, it is possible to vary the cutoffs in the periods $t=2,3$ if the dependent variable has more than two points if support. Varying the cutoffs over time is what distinguishes our conditioning event from that in HK. It is what allows us to identify the threshold parameters.
The following sufficiency result shows that different choices of $\left(j,l\right)$ reveal different combinations of thresholds in certain conditional probabilities that do not depend on the incidental parameters $\alpha_{i}$. In what follows, the logistic function is denoted by $\Lambda\left(u\right)=\exp\left(u\right)/\left(1+\exp\left(u\right)\right)$, and the change in the regressors from period 1 to 2 by $\Delta X_{i}=X_{i2}-X_{i1}$.
Identification of the model parameters comes from considering all possible combinations of cutoffs. It is clear from Theorem (ref) that different choices for $\left(j,k,l,d_{0},d_{3}\right)$ reveal information about distinct linear combinations of $\left(\rho,\gamma\right)$. By considering multiple choices of $\left(j,k,l,d_{0},d_{3}\right)$, and then aggregating the resulting information, we can identify all the model parameters. We require an additional assumption before stating our main identification result.
This assumption guarantees that for each choice of $\left(j,l\right)$, there is sufficient variation in $\Delta X_{i}$ in the subpopulation of stayers $X_{i,2}=X_{i,3}$ to identify the regression coefficient. This assumption can be weakened: we only need sufficient variation for some $\left(j,l\right)$. However, if it fails for sufficiently many $\left(j,l\right)$, identification of some of the threshold parameters may fail.
Denote by $Y_{i}=\left(Y_{i,0},Y_{i,1},Y_{i,2},Y_{i,3}\right)$ the time series of dependent variables for a given individual.
Theorem (ref) suggests that, for each choice of $2\leq j\leq k\leq l$, we could use a conditional maximum likelihood estimator (CMLE) to estimate a linear combination of the model parameters. Theorem (ref) suggests that a composite CMLE (CCMLE), based on the combination of conditional likelihoods across all choices of $\left(j,k,l\right)$, may be used to estimate the model parameters $\left(\beta,\rho,\gamma\right)$. In this section, we define that CCMLE and establish conditions under which it has desirable large sample properties. We focus on the discrete regressor case. Results for continuous regressors can be obtained by adapting Theorems 1 and 2 in HK to our case.
The binary random variable \[ C_{i,jl}=1\left\{ \left(D_{i,1}=0,D_{i,2}\left(l\right)=1\right)\text{ or }\left(D_{i,1}=1,D_{i,2}\left(j\right)=0\right)\right\} \times1\left\{ X_{i,2}=X_{i,3}\right\} . \] indicates whether $i$'s time series fits the description in $C_{j,l}=A_{j,l}\cup B_{j,l}$, and that it is also a “stayer” in the sense that $X_{i2}=X_{i3}$. Note that if $C_{i,jl}=1$, then $D_{i,1}=1$ implies that the individual time series is of the type $B_{j,l}$. Similarly, if $C_{i,jl}=1$, then $D_{i,1}=0$ implies that individual $i$ is of type $A_{j,l}$.
In the log-likelihood contribution below, ((ref)), we have substituted $D_{i,0}$ for $d_{0}$ in equation ((ref)). The value to substitute for $d_{3}$ depends on whether we are in case $A$ or $B$. To that end, define
The conditional log likelihood contribution for individual $i$, for cutoffs $\left(j,l\right),$ $2\leq j\leq k\leq l\leq J$, can then be written
The CCMLE is
where we have implicitly imposed $\gamma_{k}=0$ in the definition of $l_{i,jl}$.
We maintain the following assumption to establish the asymptotic properties of the CCMLE.
With additional technical work, this assumption can be relaxed to the case where $X_{i,2}-X_{i,3}$ is continuously distributed with positive density around zero, see HK's Theorem 1 and 2.
Our analysis uses panel data for the period 2003-2016 from the European Union Statistics on Income and Living Conditions (EU-SILC), see Eurostat (2017) for detailed documentation. The microdata is publicly available upon request.\footnote{The data were made available to us by Eurostat under Contract RPP 132-2018-EU-SILC.} EU-SILC provides a set of indicators on income and poverty, social inclusion, living conditions and, importantly, health status. For each country in the European Union, plus Iceland, Norway, and Switzerland, EU-SILC contains data on a representative sample of the population of those 18 years and older.
EU-SILC is a rotating panel. Every individual is followed over a period of two to four years. The total number of individual-years for the period 2003-2016 is 1273877. Our identification result demands four observations per individual, so we restrict attention to individuals that report valid information on their health status for 4 consecutive years. This restriction, and the restriction that the explanatory variables that we use in the analysis below have non-missing information, leaves us with a sample of 260601 individuals, for 1042404 individual-years. The proportion of incomplete samples differs across countries. As a result, the sample we work with may not be representative of EU-SILC's population. For example, out of the 27 countries that contribute to our sample, the largest contributors are Italy (with 43385 individuals), Spain (25634), and Poland (22628); the smallest are Portugal (12), Iceland (1496), and Slovakia (1982).
The outcome variable in our analysis is self-reported health status: self-perceived physical health, elicited during EU-SILC interviews. The person answers the question on how she perceives her physical health to be in general, at the date of the survey, by classifying it as one of: (1) bad and very bad (12% in our sample); (2) fair (26%); (3) good (44%); (4) very good (19%).\footnote{We have merged the separate categories “bad” and “very bad” in the original reported variable, because there is only a small fraction of observations with “very bad” health status.} Out of $260601\times3=781803$ health transitions that we observe, most often there is no change in health status (65.6%). Decreases by one unit (16%) are slightly more frequent than increases by one unit (15%). Two-unit increases (1.2%) and decreases (1.5%) are infrequent, and three-unit increases (0.08%) and decreases (0.12%) are rare.
Table (ref) relates the outcome variable, and changes to the outcome variable, to our main explanatory variable of interest, log income (total disposable household equivalised income). Household income was scaled using the composition and size of each household. This scale is based on the OECD modified equivalence scale, which gives a weight of 1.0 to the first adult in the household, 0.5 to other adults and 0.3 to each child (under 14 years old).
The table provides descriptive statistics for log income in our sample, grouped by health status. The top panel is in levels. Average income is increasing in health status. The bottom panel is in changes, which represents one way to control for unobserved heterogeneity. The implied increases for changes are close to zero, hinting at the imported role of unobserved heterogeneity.
Table (ref) reports a set of descriptive statistics on income and other explanatory variables, described in the next few paragraphs. In our analysis below, we control for some time-varying variables that are standard in the literature. First, the number of children is measured as the number of persons living in the private household that are age $14$ or less, top-coded at 3 children. In our sample, 74% of respondents have no children, and the average number of children is 0.40. Second, marriage status is a dummy variable that indicates being married or living together. The majority of the individuals are married (61%). Third, we use a self-reported indicator for labor market status variable that we map onto 4 values: (1) employed, 51%; (2) unemployed, 5.2%; (3) retired, 12% and (4) other, 32%. The value “other” includes students, permanently disabled or unfit to work, and fulfilling domestic tasks and care responsibilities.
In our fixed effects results below, we do not further control for variables that do not change over the sample period for a given individual. However, we include a set of time-invariant explanatory variables when we obtain results for non-fixed effects estimators.\footnote{Coefficient estimates for these variables are omitted from the main text, and reported in Appendix (ref).} Table (ref) provides descriptive statistics for such variables. The total sample contains slightly more females (54%) than males (46%). The proportion of individuals aged between 18 and 25 is 8.2%; 21% of individuals are aged 65 or more. With regards to education, 1.3% of the sample have no schooling (0); 14% have attended primary school (1); 19% have lower secondary education (3); 42% have upper secondary education (4); 3.8% have post-secondary education (5) and 20% have tertiary education (6). Geographically, 39% of individuals live in areas with a high degree of urbanisation and 39% live in areas with low levels of urbanisation.
We estimate the parameters in the dynamic ordered choice model with fixed effects, with latent variable outcome equation
Regression results are presented in Table (ref). The first four columns (a-d, “DOLFE”, for dynamic ordered logit with fixed effects) presents the results for ((ref)) using the estimator described in Section (ref). Different values of $h$ refer to a bandwidth parameter that we introduce because one of the explanatory variables is continuous, as in HK. Column (d) omits the employment variables, to check whether relationship between employment status and income matters for estimation of the effect of income on health.
We also present estimation results for different estimators. Results for the static ordered logit model with fixed effects, i.e. setting $\rho=0$ in ((ref)), are obtained using the estimator in Muris (2017), and presented in column (e) (“FEOL”). Column (f) (“DOL”) estimates a dynamic ordered logit model without fixed effects, i.e. ((ref)) with $\alpha_{i}=0$. Column (g) (“OL”) presents results for cross-sectional ordered logit estimator that does not take into account fixed effects or dynamics (i.e. $\alpha_{i}=\rho=0$ in ((ref))). We also present results for a static linear model with (h, “FELM”) and without (i, “LM”) fixed effects. The standard errors for all estimators are obtained using the bootstrap (500 replications). For the estimators that are not of the fixed effects type, we additionally control for education, gender, education level and the level of urbanisation. DOLFE uses four periods of data, corresponding to $t=0,1,2$. For comparability, the other dynamic estimator also uses periods 0,1,2; static estimators use periods 1,2.
Income. Our main explanatory variable of interest is income (log income, coefficient $\beta_{1}$). Across almost all specifications, we find a positive association between income and self-reported health. The only exception is column (b), where the point estimate is negative, and about the same magnitude as the standard error.
Controlling for unobserved heterogeneity leads to a very strong reduction in the magnitude of the association. For example, for the static case, a comparison of columns (e) and (g) says that, for the static case, controlling for unobserved heterogeneity reduces the coefficient on income by more than a factor 20. For this comparison, note that the threshold differences increase, suggesting that the scale increases; compare also the coefficients on the other variables, with an unchanged order of magnitude. We are not the first to observe a limited association between income and self-reported health. In a review of the literature, Gunasekara et al. (2011) found a small positive link between income and self-reported health, which is reduced when controlling for unmeasured confounders. Interestingly, Johnston et al. (2009) found no link between self-reported hypertension and income; an association that, however, became positive when using objective measures of hypertension.
The estimated effect of income also changes when we control for state dependence. Comparing columns (f) and (g), we see that controlling for state dependence in a model without unobserved heterogeneity reduces the association between income and self-reported health. So, individually controlling for unobserved heterogeneity or for dynamics reduces the magnitude of the association between health and income.
Finally, a comparison between columns (a) and (d) shows that the estimate for income association is robust to controlling for employment status,
State dependence. We estimate an autoregressive parameter of around 0.75, with threshold differences of about $3$. The estimated ratio of $\rho$ to the thresholds (which measure the distance from category 3) are much lower than for column (f). This confirms the importance of controlling for unobserved heterogeneity, which reduces the estimated magnitude of persistence by a factor 3. Nevertheless, even when controlling for unobserved (and observed) heterogeneity, we find strong evidence for large, positive persistence in self-reported health.
There are at least two ways to get a sense of the magnitude of persistence. The first approach, also available for binary choice methods, is to compare estimates of $\rho$ to estimates of regression coefficients. For example, in our preferred specification in column (a), a health shock that lifts you from any category below 3, to category 3 or 4, has an impact on future health that is almost 4 times that of becoming unemployed. The impact is more than 5 times that of marrying.
The second approach to interpreting estimates of $\rho$ uses the estimated thresholds to obtain an estimate similar to a linear autoregressive model.\footnote{This approach is not available for binary choice models because threshold parameters are not available. } Differences between the thresholds are a measure of the distances between two categories. If $\gamma_{2}=-\gamma_{4}$, then categories 2 and 3 are as far apart as categories 3 and 4. In such a case, a linear model may yield similar results in terms of partial effects. In this case, $-\rho/\gamma_{2}$ and $\rho/\gamma_{3}$ can be interpreted as linear regression coefficients for that category; we find that they are about 0.25. Said differently, we find the analog of an AR(1) coefficient of 0.25 in a linear model.
Other time-varying covariates. The literature so far has been inconclusive on how retirement is associated with health. On one hand, retiring allows more time for health-promoting activities, and reduces work-related stress. On the other hand, people may lose traction and motivation and may become less active. Therefore, while Coe and Zamarro (2011) find that retirement improves health, Behncke (2012) finds an increase in the likelihood of disease following retirement. In our DOLFE model, the coefficient is statistically insignificant. This suggests that the association between retirement and health may not be as strong as previously thought. Compared to the FEOL model, the effect of retirement on health disappears when controlling for state dependence.
The extensive literature on the link between unemployment and health in particular, and economic conditions and health more generally, is broadly inconclusive. Some studies have suggested a protective role of unemployment on health (Ruhm, 2000), while others suggest that unemployment is detrimental for health (McInerney and Mellor, 2012). Our results appear to be more in line with the findings of Ruhm (2015) and B�ckerman and Ilmakunnas (2009). In DOLFE, the coefficient of being unemployed is negative and statistically significant. Controlling for state dependence does not change things compared to the FEOL model.
Having children is insignificant in our DOLFE model, while previous studies have provided mixed findings on this question (Mckenzie and Carter, 2013; Evenson and Simon, 2005). This is also insignificant in the FEOL model, suggesting that previous findings on having children might have been driven by unobserved heterogeneity.
Being married is generally considered a protective factor for health (Kaplan and Kronick, 2006; Molloy et al., 2009). In our model, however, it is statistically insignificant -- as opposed to the FEOL model where it was positive and significant. Thus, controlling for state dependence appears to be important for this variable.
This paper studies a fixed$-T$ dynamic ordered logit model with fixed effects (DOLFE) and is the first to provide identification and estimation results for all common parameters in a dynamic ordered logit model with fixed effects and a fixed number of time periods. The results require only four time periods of data on the ordinal outcome variable. We demonstrate identification of the autoregressive coefficients on the lagged dependent variable, the regression coefficients on the exogenous regressors, and differences of the threshold parameters. The latter makes it possible to interpret the magnitude of the coefficients.
Including fixed effects and state dependence in the model is particularly relevant for self-reported health, a measure that is widely used in the literature. Future research using self-reporting health can benefit from our model for two main reasons. First, controlling for fixed effects, one can take into account unobserved heterogeneity (Carro and Traferri, 2014; Halliday, 2008; Fern�ndez-Val et al., 2017), which is especially important due to differences in understanding and reporting health status (Groot, 2000; Sen, 2002; Jylh� et al., 1998; Baron-Epel et al., 2005; J�rges, 2007). Second, it incorporates elements of persistence (Contoyannis et al., 2004; Ohrnberger et al., 2017; Hern�ndez-Quevedo et al., 2008; Roy and Schurer, 2013) or adaptation (Cub�-Moll� et al., 2017; Daltroy et al., 1999; Damschroder et al., 2005; Heiss et al., 2014) by controlling for state dependence (Carro and Traferri, 2014; Fern�ndez-Val et al., 2017; Halliday, 2008). Thus, using our estimator addresses such biases often present in studies using self-rated health (Davillas et al., 2017).
We thus applied the new dynamic ordered logit model with fixed effects to investigate the determinants of self-reported health, focusing on the link between income and health in a panel of European countries. We found that when controlling for unobserved heterogeneity, the association between income and health becomes statistically insignificant. This is in line with studies that have found a smaller or insignificant association when using fixed effects (Gunasekara, 2011; Larrimore, 2011) -- while other studies have suggested a positive association between the two (Carrieri and Jones, 2017; Ettner, 1996; Frijters et al., 2005; Mackenbach et al., 2005). Being retired or married also becomes statistically insignificant in our model when controlling for state dependence. Being unemployed or having children does not appear to be associated with self-reported health in our model.
Our empirical results suggest that persistence plays a positive and significant role in one's self-reported health. In other words, one's health is dependent on the health in the previous period, which is a reasonable thing to expect, as health problems may expand over a number of periods, or become permanent. This element reflects persistence of health status over time (Contoyannis et al., 2004). Furthermore, what is particularly interesting is that in our data, self-reported health tends to improve, on average, over time - even as people become four years older during the study period - and being older is typically associated with worse health outcomes. Therefore, it is reasonable to believe that this improvement in self-reported health is often subjective, and does not necessarily reflect one's objective health level. This element might reflect adaptation to health problems: Even though one's health does not improve, they adapt to their situation and therefore report better health (Cub�-Moll� et al 2017; Daltroy et al., 1999; Damschroder et al., 2005). This second element is a typical bias when using self-reported health outcomes, and our model helps correct such biases by introducing the dynamic element to a fixed effects ordered model. Another interesting finding is that, when controlling for unobserved heterogeneity, the link between income and health becomes statistically insignificant, suggesting that other factors might explain the association between the two.
Overall, measurement bias in studies using self-reported outcomes often poses challenges to research, that may discourage the use of such variables. Our model addresses these biases, and thus provides a basis for more choices in conducting research with databases that provide such variables.