EconBase
← Back to paper

Analyzing Subjective Well-Being Data with Misclassification

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

87,054 characters · 21 sections · 82 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Analyzing Subjective Well-Being Data with Misclassification

abstractWe use novel nonparametric techniques to test for the presence of non-classical measurement error in reported life satisfaction (LS) and study the potential effects from ignoring it. Our dataset comes from Wave 3 of the UK Understanding Society that is surveyed from 35,000 British households. Our test finds evidence of measurement error in reported LS for the entire dataset as well as for 26 out of 32 socioeconomic subgroups in the sample. We estimate the joint distribution of reported and latent LS nonparametrically in order to understand the mis-reporting behavior. We show this distribution can then be used to estimate parametric models of latent LS. We find measurement error bias is not severe enough to distort the main drivers of LS. But there is an important difference that is policy relevant. We find women tend to over-report their latent LS relative to men. This may help explain the gender puzzle that questions why women are reportedly happier than men despite being worse off on objective outcomes such as income and employment. JEL Classification Numbers: \ C14,\ C51, I31\ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ Keywords: Measurement error, identification, subjective well-being, testing.

\setcounter{page}{1}

Introduction

Happiness or well-being economics first appeared in the economics literature in the early 1970s, see van_praag_welfare_1971, easterlin_does_1974. This fast growing, yet sometimes polarizing, subject studies causes and consequences of subjective well-being (SWB) and has provided many interesting insights into what makes people happy. Some of which have led to important policy lessons such as the idea that unemployment in the Western society is largely involuntary (winkelmann_why_1998), that reducing the rates of joblessness should take priority over reducing the inflation rates (di_tella_preferences_2001 ), that people partially adapt to serious disability over time (oswald_does_2008), and that cigarette taxes actually improve the happiness of the likely smokers (gruber_cigarette_2006).

The central variable used in the well-being literature is life satisfaction (LS). LS is originally designed to capture the respondent's global well-being (diener_satisfaction_1985). While LS has been shown to be correlated with a range of economic factors such as health and unemployment in expected ways, there is also ample evidence from the experimental literature the reporting of LS is affected by confounding factors including questionnaire designs, temporal factors such as mood or weather, and pressures to provide socially desirable answers (see, e.g., schwarz_mood_1983, feddersen_subjective_2016, schwarz_cognition_2014). We can therefore view reported LS as a possible mismeasurement of latent LS. Given the discrete nature of the SWB responses measurement error is also known as a misclassification.

Misclassification is a form of non-classical measurement error. A mismeasured LS can cause bias in empirical studies in arbitrary way. Indeed, one particular takeaway from the well-known article by bertrand_people_2001, entitled: \textquotedblleft Do people mean what they say? Implications for Subjective Survey Data\textquotedblright , suggests researchers should not use LS as a dependent variable. Nevertheless, understanding the determinant of LS is one of the most fundamental tasks in the well-being literature. Since there is no obvious solution to the measurement error problem, LS is still routinely used as the dependent variable and the potential effects from measurement error have been unaccounted for.

In this paper we use novel econometric techniques to formally test for the presence of measurement error in reported LS and, if it exists, account for it and study its potential effects. We use survey data of 35,000 British households from the UK\ Understanding Society taken between January 2011 and June 2013. This (Wave 3) dataset is unique in that it contains what we believe are suitable variables that enable us to test for measurement errors and use the misclassification model of hu_identification_2008 to identify the joint distribution of the reported and latent LS nonparametrically. In particular, the LS distribution will be able to provide insights into the (mis-)reporting probabilities for people of different demographic and socioeconomic groups. We can also use this distribution to identify the determinants of latent LS in popular parametric models in the literature such as linear projection and ordered response models (e.g. logit and probit) without observing latent LS. Our results can have important policy implications, whether it is for the purpose of helping policy makers identify suitable groups of individuals for an intervention or for quantifying impacts of policies based on latent LS as opposed to reported LS.

We emphasize at this point that the empirical results in our paper are free from recent criticisms on the general econometric analysis of SWB data. In particular, bond_sad_nodate show the mean ranking of happiness or LS generally cannot be identified unless strong assumptions (such as homoskedasticity in probit/logit models) are a priori imposed. An implication of this is the sign of a parameter estimate is not informative on the average effect. Subsequently, bond_sad_nodate use an (heteroskedastic) ordered probit to show some of the most well-known results in the happiness literature can be arbitrarily reversed. Chen, Oparina, Powdthavee and Srisuma (chen_have_2019) point out a simple solution that is to focus on the median effect instead of the mean. In this paper we use a heteroskedastic ordered probit model without specifying the form of heteroskedasticity parametrically. Our estimates and partial effects are to be interpreted through the median accordingly.

We begin our empirical study by testing for the presence of measurement errors in reported LS. We adopt the nonparametric approach suggested recently by wilhelm_testing_2018 that, under suitable conditions, stochastic dependency between some auxiliary variables conditioning on the reported variable can signify the presence of measurement errors.\footnote{ The insight that conditional independence can indicate the presence of measurement error was first explored in a regression context by mahajan_identification_2006, who considers a binary regressor that may be measured with error.} We use a Kolmolgorov-Smirnov type statistic and find evidence of measurement error in the reported LS for the entire dataset as well as for 26 out of 32 socioeconomic subgroups in the sample.

Next, we estimate a model of latent life satisfaction conditioning on key socioeconomic variables. We find the main drivers of LS in the model with latent LS are the same as those with reported LS, suggesting that the bias from measurement error may not be substantial enough to distort the effects of the main factors. For example, marriage and health have clear positive impact on LS while income and education have insignificant effects that would otherwise be positive due to substitution effects with health. However, there is one notable difference. We find that women systematically report themselves to be more satisfied with lives than they actually are relative to men. Measurement errors may thus help us solve the gender puzzle that women are happier than men in spite of the fact that they are often associated with less favorable objective measures in terms of health, income and employment level (see dolan_we_2008, stevenson_paradox_2009).

The validity of our empirical results relies on the conditions of the misclassification model of hu_identification_2008 being satisfied. Hu's identification procedure requires assumptions on the respondents' reporting behavior as well as two auxiliary variables. These variables need to be independent to each other and to the reported LS conditional on the latent LS.\footnote{ The same conditional independence assumption between the auxiliary variables is also required for the measurement error test.} The auxiliary variables also need to be appropriately correlated with the latent variable. These requirements pose contrasting qualities for candidates for auxiliary variables, which is an empirical challenge somewhat analogous to finding a good instrument. The selection of appropriate auxiliary variables requires a transparent interpretation of the origin of misreporting.

In this paper, we focus on external factors such as the effects of questionnaire design, temporal factors, and socially-desirable responding as the source of measurement errors. These factors influence the reporting behavior, yet at the same time are idiosyncratic in the sense that they do not contribute to the LS in general. Subsequently, the two auxiliary variables we select for identification are: (i) a measure of mental state that is derived from General Health Questionnaire (GHQ); (ii) a derived measure of neuroticism, which is one of the traits that underlies one's personality. The latter in particular is currently collected only in Wave 3 of the UK Understanding Society survey. We provide a detailed discussion of the conditions for identification and the logic behind our choice of auxiliary variables in the paper. Ultimately, similarly to the exogeneity of an instrument, the conditional independence of the variables is an untestable identifying assumption. A nice feature of our dataset is that we have information for all questions that are used to derive measures from GHQ and neuroticism. Some of these questions are more objective than others, which allow us to perform robustness checks by constructing different versions of auxiliary variables that intuitively have a varying degree of trade-off between independence and relevance.

We make three main contributions in this paper. (1) To the best of our knowledge, we are the first to apply the nonparametric test of wilhelm_testing_2018 and misclassification model of hu_identification_2008 to analyze SWB data. These novel econometric methods do not rely on unjustified parametric assumptions, allow for non-classical errors, and do not require validation data (cf. bound_extent_1991, chen_measurement_2005). The latter two features in particular seem to be necessary for making any progress in accounting for measurement errors in self-reported subjective variables. (2) Our empirical study shows statistically that measurement errors exist in reported LS, using Wave 3 of the UK\ Understanding Society data. While we find some similarities between the relations between reported and latent LS with key covariates there are important differences. Measurement errors can therefore have practical implications if policy makers are to make decisions based on reported as opposed to latent LS. (3) We bridge the gap between theory and practice. Modern econometrics studies on nonparametric identification of models with measurement errors tend to be mathematically sophisticated and they aim to identify the joint distribution of all variables in the model. Most empirical researchers, especially those currently dealing with data that may be measured with errors, on the other hand employ parametric models. We show how nonparametric identification results can lead to parametric identification in popular models that are used to analyze SWB data in the literature.

We organize the rest of the paper as follows. Section 2 gives a brief background on the current use of SWB data. Section 3 presents the econometric model, gives conditions for identification of the parameters of interest, and discuss practical inference. Section 4 describes the test we use to detect possible measurement errors in the reported LS. The empirical application is in Section 5. Section 6 concludes. The Appendix provides more detailed description of the data and supplementary results to support the findings in Section 5.

Background

Our background section consists of three parts. Section 2.1 provides a brief overview for a measure of well-being. Section 2.2 summarizes the main approaches to analyzing SWB data as well as some recent criticisms. Section 2.3 discusses measurement errors in self-reported well-being variables.

Subjective well-being

Well-being research is motivated by the ambition to understand the key drivers of individual's well-being. SWB is an umbrella term that includes a person's cognitive well-being such as LS (i.e., a judgment one makes about one's life overall), affective well-being (i.e., frequency and intensity of experienced emotions), and eudaimonic well-being (i.e., sense of purpose and worthwhileness). The economics of happiness literature traditionally uses LS, rather than measures of emotional states, as a proxy utility data.

Here we list some of the stylized facts of this literature as summarized in the World Happiness Report (helliwell_world_2012):

itemize• Richer people are on average happier than poorer people; • LS is highly positively correlated with mental and physical health; • Marriage has a positive correlation with LS; • LS is U-shaped in age; • Unemployment is significantly detrimental to LS; • In most developed countries women report higher LS than men, despite being worse off in measurable socioeconomic outcomes; • There is little correlation between a person's education level and his/her LS, but education is indirectly related to happiness through its effect on income: education increases income and income increases happiness.

Given the subjective nature of LS, the overwhelming majority of the findings is based on self-reported assessment: respondents are asked to report how satisfied they are with their life on a given scale. This approach favors personal evaluation of global well-being over the views of potential experts. Despite earlier concerns, self-reported measures of life satisfaction are proven to have a degree of validity. They converge in expected ways with each other and with non-self-reported measures, such as those based on other people's reports and the behavior of the respondent ( diener_assessing_2009). They are also predictive of future behaviors, such as job quit, divorce, and suicide (diener_findings_2017).

Estimating life satisfaction

The most common feature of empirical studies in the well-being literature is to use reported LS as a dependent variable and other characteristics, such as income, gender, health and employment statuses etc., as covariates. There are two distinct approaches in how life satisfaction is modeled. One treats LS as a cardinal variable and the other as an ordinal one. The statistical techniques used for the former are based on least squares estimation or direct comparisons between sample averages. For the ordinal case, ordered logit or probit models are typically used. Both approaches are widely used in practice. See ferrericarbonell_how_2004 for an account for some (dis-)similarities of results between the two approaches.

The econometric analysis of SWB data has come under recent heavy criticisms. Whether least squares regression or ordered probit/logit estimation is used, similar to most other economic fields, a typical approach researchers take is to then draw conclusions based on statements about the relative mean happiness between groups of individuals (e.g. men and women, employed and unemployed, or across countries etc). Critiques point out that these research ignore the fact that SWB data are ordinal in nature. And the mean ranking of ordinal variables is only identified when it is stable across all increasing transformation. For examples, this means unless relevant stochastic dominance conditions hold, the raw average ranking and signs of least squares estimates may be reversed by monotonically transforming the ordinal scale, see schroder_revisiting_2017 and bond_sad_nodate. Importantly, this issue goes deeper than \textquotedblleft using OLS to estimate a discrete dependent variable\textquotedblright , as bond_sad_nodate also show the mean ranking of latent happiness from ordinal models can also be arbitrarily reversed. They use a heteroskedastic ordered probit model to illustrate it for some of the most well-known results in the happiness literature. Explicitly allowing for heteroskedasticity is important because a homoskedastic model a priori effectively assumes the mean ranking to be identified. See Theorem 1 of bond_sad_2014.

There are ways to analyze happiness data that avoid these criticisms. For examples, direct comparisons of probabilities or probability odds of certain events between groups are not affected (e.g., easterlin_will_1995). But such a descriptive approach has limited scope for incorporating covariates. Chen, Oparina, Powdthavee and Srisuma (chen_have_2019 ) suggest one solution is to use the median instead of the mean as a mode of comparison. The median rank is stable across all increasing transformations. Furthermore, they highlight the fact that the median and the mean in symmetric parametric models, like probit and logit, are the same. The median has therefore been frequently estimated but only interpreted as the mean.\footnote{ Estimating the median without any parametric distributional assumption is also possible (manski_semiparametric_1985, lee_median_1992). chen_have_2019 suggest the semiparametric median can be estimate using modern constrained mixed integer optimization technique; they apply it to study the Easterlin paradox.} This fact instantly nullifies the reversal of prior results in bond_sad_nodate\ by simply interpreting those estimates through the median. To this end, our paper emphasizes the use of an ordered response model with heteroskedasticity. We show in Section 4 it is in fact simple to estimate a heteroskedastic probit model even without specifying the form of the heteroskedasticity parametrically.

Measurement error

It is the norm in practice to assume that LS is measured without errors. In this work we take the view that reported LS is a combination of latent LS and measurement error:\footnote{ In this paper we use the term latent to mean a measurement without error. In Section 3.2.2 we model $X^{\ast }$\ using an ordered response model, which traditional interprets $X^{\ast }$ to be derived from an underlying continuous happiness variable that is plays an analous role to utility in McFadden's random utility maximization model.}

equation*[equation* omitted — 31 chars of source]

We denote the measurement error by $u$. Since $X$ and $X^{\ast }$\ are discrete, $u$ is also discrete. This type of measurement error is also known as misclassification. Misclassification is non-classical by nature. For example, given the number of values the variable can take is finite, extreme values can only be mismeasured in one direction so a zero-mean error (conditional on the true value) is impossible. Furthermore, the error term is likely to be correlated with the covariates that are typically used in LS analysis.

The error in LS is thought to come from two main distinct sources. One is the effects that influence the respondent's judgment about the level of LS while it is being formed, such as passing effects and the effects of survey design. The other comes from factors that influence how the respondents communicate their judgment, i.e. social desirability bias. We now describe these two sources in details.

Well-being research is typically interested in the relatively stable well-being level, rather than in the passing effects, which is why influences of temporal factors can be considered as measurement error that should be controlled for. LS is theorized to be a judgment that a respondent constructs while answering the question, so it can be influenced by temporal factors that take place at the time of forming the judgement (strack_subjective_1991). Multiple experiments have shown that this measure can be influenced by mood manipulations like finding a dime in a copy machine, receiving a chocolate bar, spending time in a pleasant environment or watching a football team win (strack_subjective_1991). In a large-scale survey setting, mood swings can be caused by weather at the time of the well-being judgment; there are well-known diurnal and day-of-the-week variations in SWB (see diener_advances_2018).

Research on survey design shows that respondents' answers can be manipulated to some extent, see bertrand_people_2001. Respondents tend to provide answers consistent with the previous ones, so the ordering of questions matters. Changing the wording of the question also affects the way people respond to them, particularly when people are asked to agree or disagree with a statement. When respondents are asked to assess a statement within a given scale, answers can differ if different scales are provided. Error of this type induces systematic bias, however it can be minimized by the appropriate design on the questionnaire (oecd_oecd_2013). Nowadays, the majority of surveys have been designed taking this issues into account.

There are also concerns about respondents not reporting truthfully in an endogenous way, i.e. error is correlated with regressors. Specifically, respondents may modify their answers to make them seem more socially desirable. E.g. diener_response_1991 refer to one version of social desirability, specific to reported SWB, to be \enquote{happy image management} that would result in reporting higher or lower well-being than experienced to appear happier/less happy. Measurement errors of this kind would make it difficult for us to distinguish between the case when unemployed people are truly unsatisfied with their life or when they report a lower score because being unemployed is associated with a less desirable social status.

Measurement errors from both sources might be correlated with regressors, i.e. respondents from different groups might systematically exhibit different reporting behaviors. barrington-leigh_impact_2017 use two major health surveys in Canada to show that women and individuals with poor health condition are more affected by weather conditions. heffetz_conclusions_2013 use the reported number of call attempts made to participants in the University of Michigan's Surveys of Consumers to show the difference in reported happiness among easy-to-reach and hard-to-reach respondents.

The discussion above suggests that statistical analysis using life satisfaction is likely to be biased in some unknown ways if measurement error is ignored. Some examples of these are highlighted in bertrand_people_2001. Over the past fifteen years, the econometrics literature has made advances on identifying nonclassical measurement error without additional measurement or validation data on the mismeasured variable. The approach we take in this paper follows from the misclassification model of hu_identification_2008 that assumes all the variables in the model are discrete. The discrete setup is suitable for analyzing LS as most variables that are used in this literature are discrete or can naturally be discretized. A more general treatment that allows for some continuous variables can be found in hu_instrumental_2008. We refer the reader to the surveys by schennach_2013 and hu_econometrics_2017 for examples of applications that rely on this type of identification results.

Model and identification strategy

In this section we describe a model of misclassification hu_identification_2008 in the context of our application. In Section 3.1 we introduce the variables, provide and discuss the assumptions required on them, and outline the nonparametric identification strategy. We consider parametric identification in Section 3.2. Section 3.3 discusses the numerical aspects of estimation and inference.

Nonparametric identification

Let $X^{\ast }$\ denote the latent LS. Suppose $ X^{\ast }$ can take the following values:

equation*[equation* omitted — 161 chars of source]

We assume to have three observed variables $\left( X,Y,Z\right) $. $X$ is the reported LS. $Y$ is a derived measure of neuroticism that is indicative of a responder's emotional stability.\footnote{ Neurotic individuals can be defined by such terms as worrying, insecure, self-conscious, and temperamental (mccrae_validation_1987).} $Z$ is a measure of mental states that is derived from General Health Questionnaire (GHQ-12). We assume that $X$ and $X^{\ast }$ have the same support. $Y$ is a binary indicator that takes a value of $1$ for the individuals whose level of neuroticism is above the median of the sample and $0$ otherwise. The support of $Z$ has the same cardinality as the support of $X$, which in this case is $\left\{ 1,2,3\right\} $, ranging from $1$ -- not distressed to $3$ -- distressed.

We provide a particular description of the variables above to fix ideas, which will be useful for motivating abstract assumptions of the misclassification model. All of our assumptions and theoretical results below are written in a more general term. More specifically, they are all valid for $X^{\ast }$\ that takes values from any finite set as long as the cardinality of the support of $\left( X^{\ast },X,Z\right) $ are the same. \footnote{ If the cardinalities of the support of $X$ and $Z$\ are unequal initially, one can always coarsen the data to satisfy the same cardinality condition.}$ ^{,}$\footnote{ The setup that $Y$\ takes only two values is a minimal assumption for identification. We can always convert any random variable into a binary variable.}

In what follows, we will use $f_{A|B}\left( a|b\right) $ to denote $\Pr \left[ A=a|B=b\right] $ for random vectors $A$ and $B$ taking values $a$ and $b$ respectively, and $f_{A}\left( a\right) $ to denote the $\Pr \left[ A=a \right] $. We will denote a generic matrix whose $ij-$th element is $m_{ij}$ \ by a bold font $\mathbf{M}:=\left( m_{ij}\right) $ and a diagonal matrix with the $i-$th diagonal element $d_{i}$ by $\mathbf{D}:=diag\left\{ \left( d_{i}\right) \right\} $. We denote a transpose of matrix $\mathbf{M}$ by $ \mathbf{M}^{\top }$ and\ an inverse of an invertible matrix $\mathbf{M}$ by $ \mathbf{M}^{-1}$. Correspondingly, when $A$ and $B$ are scalar variables supported on $\left\{ a_{1},\ldots ,a_{d_{A}}\right\} $\ and $\left\{ b_{1},\ldots ,b_{d_{B}}\right\} $\ respectively, we then define $\mathbf{M} _{A|B}$ to be a $d_{A}$ by $d_{B}$ matrix such that $\mathbf{M} _{A|B}:=\left( f_{A|B}\left( a_{i}|b_{j}\right) \right) $; we define $ \mathbf{M}_{A,B}$ similarly so that $\mathbf{M}_{A,B}:=\left( f_{A,B}\left( a_{i},b_{j}\right) \right) $.

We assume $\left( X^{\ast },X,Y,Z\right) $ satisfies the following conditions:

Assumption 1 (CI). $\left( X,Y,Z\right) $\ are independent conditional on $X^{\ast }$, i.e.

equation*[equation* omitted — 86 chars of source]

Assumption 2 (RNK). $\mathbf{M}_{X,Z}:=\left(f_{X,Z} \left(x_{i},z_{j}\right)\right)$\ has full rank.

Assumption 3 (UNQ). $E[Y|X^{\ast }=x_{i}^{\ast }]$\ is different for different $i$.

Assumption 4 (ORD). $f_{X|X^{\ast }}(x_{I}|x_{i}^{\ast })$ is strictly increasing in $i=1,\ldots ,I$

Assumption 1 is the key conditional independence assumption. While it is easy to find three independent variables in isolation, the challenge is to also have them satisfy Assumptions 2 to 4. We first explain why our choice of $\left( X,Y,Z\right) $ may reasonably satisfy Assumption 1. Suppose the source of the misclassification error that makes $X$ different to $X^{\ast }$ comes from temporal factors (e.g. mood or weather), socially desirable responding (e.g. {happy image management}) or questionnaire design. We have selected $Z$ and $Y$ carefully so they contain information on life satisfaction and some other information that we treat as errors. We want these errors to be independent from the errors in $X$ and between themselves once we control for $X^{\ast }$. For $Y$, the measure of neuroticism is also constructed from the answers to multiple questions. These questions concern personal traits rather than direct assessment of LS. The answers are unlikely to be influenced, for example, by {happy image management}, because the reporting of LS is different from experiences regarding emotional stability. The questions about personal traits and those about LS are also often asked in different parts of the survey (as is the case for our dataset). It is designed to minimize the influence of questions and answers that LS and neuroticism may have on each other. Moreover, temporary factors such as weather, are less likely to influence individuals evaluation of how often one worries, compared to the judgment about to what extent she is satisfied with one's life. For $Z$, unlike the LS question, the GHQ-12 measure is constructed from multiple questions with a varying degree of subjectivity. Parts of the questions are as subjective as the well-being question, e.g. \enquote{Have you recently been feeling reasonably happy}, however, some questions ask for objective information such as the amount of sleep. Given the questions are less subjective, the answers are less likely to be influenced by similar cognitive effects. We perform a robustness check on our estimation results by using different GHQ-12 measures, which vary in degree of objectivity/subjectivity, in Appendix B.

Assumption 2, unlike the other assumptions, is testable as it is a condition on the observable. $\mathbf{M}_{X,Z}$ is a square matrix since $X$ and $Z$ have the same number of support points. The full rank condition is the discrete analog to the completeness assumption (see hu_instrumental_2008), which ensures invertibility of $\mathbf{M}_{X,Z}$.

Assumptions 3 and 4 have more intuitive interpretations. Assumption 3 says that the probability that an individual whose level of neuroticism is above the median of the sample differs across sub-populations partitioned by $ X^{\ast }$. Since neuroticism captures personal trait, which has been shown to be strongly related to the level of LS (see, e.g., diener_subjective_2009), this condition is likely to hold. Assumption 4 imposes monotone likelihood towards positive reporting. In particular, if we set $x_{i}=I$ and $x_{i}^{\ast }=i$ for all $i$, respondents who are latently {satisfied} with their lives are more likely to report the higher state that those who are {neither satisfied nor dissatisfied}; analogously, those who are latently {neither satisfied nor dissatisfied} are more likely to report the higher state than those who are {not satisfied}.

In order to provide further insights on why A1 - A4 enable the identification of $f_{X^{\ast },X,Y,Z}$, we now provide an intuitive outline for the proof of Theorem 1 in hu_econometrics_2017.

Identification of $f_{X^{\ast },X,Y,Z}$ under A1 - A4

Under Assumption 1, we have:

equation*[equation* omitted — 96 chars of source]

The distribution of $\left( X^{\ast },X,Y,Z\right) $\ is identified if we can identify the distribution of $X^{\ast }$ and the marginal distributions of $X,Y$ and $Z$\ conditional on $X^{\ast }$. Using the Law of Total Probability, under Assumption 1, we have

equation*[equation* omitted — 182 chars of source]

where we denote the support of $X^{\ast }$ by $\pazocal{X}^{\ast }$. Fix $ Y=y $, we can define $\mathbf{M}_{X,y,Z}:=\left( f_{X,Y,Z}(x_{i},y,z_{j})\right) $\ so that the above relation can be vectorized for each $y$,

equation[equation omitted — 151 chars of source]

where $\mathbf{M}_{X|X^{\ast }}:=\left( f_{X|X^{\ast }}(x_{i}|x_{j}^{\ast })\right) ,\mathbf{D}_{y|X^{\ast }}:=diag\left\{ \left( f_{Y|X^{\ast }}(y|x_{i}^{\ast })\right) \right\} ,\mathbf{D}_{X^{\ast }}:=diag\left\{ \left( f_{X^{\ast }}(x_{i}^{\ast })\right) \right\} $ and $\mathbf{M} _{Z|X^{\ast }}:=\left( f_{Z|X^{\ast }}(z_{i}|x_{j}^{\ast })\right) $. Similarly, using the Law of Total Probability and Assumption 1, we can also write

equation*[equation* omitted — 150 chars of source]

which can be represented in a matrix notation by

equation[equation omitted — 125 chars of source]

for $\mathbf{M}_{X,Z}:=\left( f_{X,Z}(x_{i}|z_{j})\right) $. When Assumption 2 holds, we have

equation*[equation* omitted — 115 chars of source]

The above display can be used to combine ((ref)) and ((ref)) and obtain $\mathbf{M}_{X,y,Z}=\mathbf{M}_{X|X^{\ast }}\mathbf{D}_{y|X^{\ast }} \mathbf{M}_{X|X^{\ast }}^{-1}\mathbf{M}_{X,Z}$, so that

equation[equation omitted — 147 chars of source]

Hu's main insight is $\mathbf{M}_{X,y,Z}\mathbf{M}_{X,Z}^{-1}$, which is identified by the data, can identify $\mathbf{M}_{X|X^{\ast }}\mathbf{D} _{y|X^{\ast }}\mathbf{M}_{X|X^{\ast }}^{-1}$ by an eigen-decomposition, where the diagonal elements of $\mathbf{D}_{y|X^{\ast }}$ are the eigenvalues and $\mathbf{M}_{X|X^{\ast }}$\ is a matrix of the corresponding eigenvectors. Assumptions 3 and 4 ensure that the eigen-decomposition produces a unique and distinct ordering of eigenvalues, thus $f_{X|X^{\ast }} $ and $f_{Y|X^{\ast }}$ are identified. In turn they also identify $ f_{Z|X^{\ast }}$ and $f_{X^{\ast }}$. To see this, first note that $ f_{X}(x)=\sum_{x^{\ast }\in \pazocal{X}^{\ast }}f_{X|X^{\ast }}(x|x^{\ast })f_{X^{\ast }}(x^{\ast })$, so that we can identify $f_{X^{\ast }}$\ by (pre-)multiplying a vector of $\left( f_{X}\left( x_{i}\right) \right) $\ by $\mathbf{M}_{X|X^{\ast }}^{-1}$. Then $f_{Z|X^{\ast }}$\ can be identified by solving, for instance, equation ((ref)) for $\mathbf{M}_{Z|X^{\ast }}^{T}$. Thus $f_{X^{\ast },X,Y,Z}$\ is identified when Assumptions 1 to 4 hold.

The argument above makes clear that we use Assumptions 3 and 4 only for the purpose of identifying the eigen-decomposition of $\mathbf{M}_{X,y,Z}\mathbf{ M}_{X,Z}^{-1}$. While we cannot test Assumptions 3 and 4 directly, in practice we can estimate $\mathbf{D}_{y|X^{\ast }}$\ and $\mathbf{M} _{X|X^{\ast }}$ without fully imposing Assumptions 3 and 4 a priori. For example, if the inequalities in Assumption 4 are violated empirically then this would suggest some conflicts with the data. (We provide more discussion on this in Section 3.3 and Section 5.) In this case one should seek other economically plausible conditions to ensure uniqueness of the eigen-decomposition. Alternative conditions for identification can be found in hu_identification_2008.

The identification of $f_{X^{\ast },X,Y,Z}$\ \ gives a complete characterization of the stochastic relation between all the variables in the model. In particular, the model offers new insights into various reporting behaviors conditioning on the latent level of satisfaction. Next we show how $f_{X^{\ast },X,Y,Z}$ can be used in conjunction with additional covariates to identify commonly used parametric models.

Parametric identification

Empirical studies are most often interested in the coefficients in linear and probit/logit models of LS given a vector of covariates\ $Q$. If we use only reported LS then these parameters can be written as some functionals of $f_{X,Q}$, which are identified from the observed data under familiar conditions. We would like to identify and estimate analogous parameters for latent LS. This is possible even if we do not observe latent LS as long as $ f_{X^{\ast },Q}$\ is identified.

Since $Q$ may not necessarily contain $\left( Y,Z\right) $, which are used for nonparametric identification, we write $Q=\left( R^{\top },W^{\top }\right) ^{\top }$\ so that $R$, if it is non-empty, contains either $Y$ and/or $Z$, and $W$\ is a vector of all other conditioning variables. We shall assume throughout that $W$\ is a discrete random variable. So that all variables in the model are discrete and take values from some finite set. We assume our data satisfy the following condition.

Assumption A. $\left\{ \left( X_{n},Y_{n},Z_{n},W_{n}\right) \right\} _{n=1}^{N}$ is a random sample of $\left( X,Y,Z,W\right) $ \ with $N\rightarrow \infty $\ such that $\left( X,Y,Z\right) $\ conditional on $W$\ satisfies Assumptions 1 to 4 almost surely.

Assumption A ensures that $f_{X^{\ast },X,Y,Z|W}$\ is identified. This allows us to identify $f_{X^{\ast },Q}$.

Lemma 1. Suppose Assumption A holds. Then $f_{X^{\ast },Q}$ \ is identified.\

Proof. The random sampling assumption ensures that $f_{W}$ is identified. Therefore is $f_{W,X^{\ast },X,Y,Z}$\ identified. We can integrate out $Y$ and/or $Z$ in $f_{W,X^{\ast },X,Y,Z}$\ if they are not contained in $Q$\ to identify $f_{X^{\ast },Q}$.$\blacksquare $

We now consider two parametric models that are most often used in practice and show how to identify the parameters of interest.

Linear projection model

Here $X^{\ast }$ is treated as a cardinal variable. Suppose $X^{\ast }$\ is observed. Let $\widetilde{Q}=\left( 1,Q^{\top }\right) ^{\top }$.\ We are interested in $\beta _{C}$, which comes from the following linear projection model:

equation[equation omitted — 148 chars of source]

Then we can identify $\beta _{C}$\ as a least squares solution under familiar conditions. We state this as a proposition without proof.

Proposition 1. Suppose Assumption A holds and $\left( X^{\ast },\widetilde{Q}\right) $ satisfies ((ref)). If $E\left[ \widetilde{Q}\widetilde{Q}^{\top }\right] $\ has full rank, then

equation[equation omitted — 148 chars of source]

\

Note that the linear probability model does not assume a priori that $ \varepsilon $ is homoskedastic conditional on $\widetilde{Q}$. Here $E\left[ \widetilde{Q}\widetilde{Q}^{\top }\right] $ can be identified from the data.

Ordered probit model

Now let $X^{\ast }$ be an ordinal variable generated from an ordered response model. Suppose $X^{\ast }$\ is observed. We are interested in $ \beta _{O}$, which comes from the following ordered probit model:

equation[equation omitted — 193 chars of source]

where $\left( \mu _{i}\right) _{i=1}^{I-1}$ is an increasing sequence of reals with $\mu _{0}=-\infty $ and $\mu _{I}=+\infty $, $\sigma \left( Q\right) $ denotes a skedastic function that is positive almost surely, and $ \varepsilon $ has a standard normal distribution.

We can interpret $\widetilde{Q}^{\top }\beta _{O}+\sigma \left( Q\right) \varepsilon :=U^{\ast }$\ in the traditional way. I.e $U^{\ast }$\ is an underlying continuous happiness variable that gets transformed into discrete level of LS. By symmetry of the normal distribution $\widetilde{Q}^{\top }\beta _{O}$ is the (conditional) median, as well as the mean, of $U^{\ast }$ . But, unless $\sigma \left( Q\right) =1$ almost surely, the sign of a mean partial effect of $U^{\ast }$ is generally not identified while the sign of a median effect is identified. See Section 3.1 of chen_have_2019 for a more detailed discussion.

In what follows we denote the CDF of $\varepsilon \ $by $\Phi $. It is well-known that an ordered probit is not identified and some normalizations have to be made. In this paper we set $\left( \mu _{1},\mu _{2}\right) =\left( 0,1\right) $.\footnote{ Alternatively normalizations can be made on $\beta _{O}$. For example, the intercept can be set to $0$ and one of the slope parameters can be set to $1$ .} Next, we show in Lemma 2 that $\sigma $\ is identified without further assumptions.\ The proof of this result uses the identification strategy from chen_rates_2003.

Lemma 2. Suppose Assumption A holds. Then $\sigma $ \ is identified and

equation[equation omitted — 189 chars of source]

Proof. From ((ref)), we have:

equation[equation omitted — 272 chars of source]

It then follows that

eqnarray[eqnarray omitted — 308 chars of source]

So that $\frac{1}{\sigma \left( Q\right) }=\Phi ^{-1}\left( \Pr \left[ X^{\ast }\leq 2|Q\right] \right) -\Phi ^{-1}\left( \Pr \left[ X^{\ast }\leq 1|Q\right] \right) $. By Lemma 1 $f_{X^{\ast }|Q}$\ is identified. Therefore $\sigma $\ is identified.$\blacksquare $

An interesting feature of the heteroskedastic ordered response model above is that we only need information on $\Pr \left[ X^{\ast }=i|Q\right] $ for $ i=1,2$ to identify $\sigma $\ even if $I$ is larger than $3$. In fact, the same can be said for the identification of $\beta _{O}$. Suppose that $I\geq 3$, then the additional information from $\Pr \left[ X^{\ast }=i|Q\right] $\ for $i\geq 3$ can be used for identifying $\mu _{O}:=\left( \mu _{3},\ldots ,\mu _{I-1}\right) $.

Proposition 2. Suppose Assumption A holds and $\left( X^{\ast },Q\right) $ satisfies ((ref)). If $E\left[ \widetilde{Q} \widetilde{Q}^{\top }\right] $\ has full rank, then

equation[equation omitted — 175 chars of source]

where $\widetilde{X}^{\ast }\left( Q\right) :=-\sigma \left( Q\right) \Phi ^{-1}\left( \Pr \left[ X^{\ast }=1|Q\right] \right) $ and

equation[equation omitted — 193 chars of source]

Proof. Re-arrange ((ref)) to obtain,

equation*[equation* omitted — 131 chars of source]

Pre-multiply both sides of the display above by $\widetilde{Q}$. Take expectation and solve it to identify $\beta _{O}$.

We can identify $\mu _{O}$ by solving $\Pr \left[ X^{\ast }\leq i|Q\right] =\Phi \left( \frac{\mu _{i}-\widetilde{Q}^{\top }\beta _{O}}{\sigma \left( Q\right) }\right) $ for all $i\geq 3$, where the latter expression is implied by ((ref)).$\blacksquare $

By inspecting the proof of Proposition 2, note that we can equivalently use ( (ref)) to identify $\beta _{O}$. In particular, the normalization restrictions impose the condition that $\sigma \left( Q\right) \Phi ^{-1}\left( \Pr \left[ X^{\ast }\leq 1|Q\right] \right) =\sigma \left( Q\right) \Phi ^{-1}\left( \Pr \left[ X^{\ast }\leq 2|Q\right] \right) -1$.

Our discussion above assumes normality of $\varepsilon $\ in ((ref)) for concreteness. Other parametric models, such as the logit, can be identified analogously by replacing $\Phi $ with another CDF of a continuous variable that has full support on $\mathbb{R}$.

Practical estimation and inference

The nonparametric and parametric identification strategies in Section 3.1 and Section 3.2 respectively are constructive. They suggest we can construct consistent estimators by simply replacing unknown population quantities by the sample counterparts. But it may not always be ideal to take that approach in practice. We next provide some alternative estimation methods for the parameters of interest.

Nonparametric estimation

We can follow the identification steps in Section 3.1\ closely by first performing an eigen-decomposition using matrices of sample probabilities instead on the left hand side of equation ((ref)). However, an eigen-decomposition in finite sample can produce estimates that do not respect a priori assumed theoretical assumptions. Applications using related identification results above (see examples in hu_econometrics_2017) typically employ a constrained maximum likelihood for estimation.

Under Assumption A, we have

eqnarray*[eqnarray* omitted — 409 chars of source]

We can therefore construct a likelihood function based on the joint probability above where the parameters of interest are $\left( f_{X|X^{\ast },W},f_{Y|X^{\ast },W},f_{Z|X^{\ast },W},f_{X^{\ast }|W},f_{W}\right) $. The maximum likelihood estimator of $f_{W}$ corresponds to the empirical distribution of $\left\{ W_{n}\right\} _{n=1}^{N}$, which can be obtained independently of the other parameters. Maximum likelihood estimation of the other parameters can be performed conditionally on $W$.

Let $\mathcal{S}_{W},\mathcal{S}_{X},\mathcal{S}_{Y}$ and $\mathcal{S}_{Z}$ denote the cardinalities of the support of $W,X,Y$ and $Z$. Then for each $w$ in the support of $W$, there are $\mathcal{S}_{XYZ}:=\mathcal{S}_{X}\mathcal{ S}_{Y}\mathcal{S}_{Z}$ possible realizations of $\left( X,Y,Z\right) $. \footnote{ We assume the joint support of $\left( W,X,Y,Z\right) $ is the same for all realizations of $W$ for notational simplicity.} We can enumerate these distinct events by $\left\{ x_{j},y_{j},z_{j}\right\} _{j=1}^{\mathcal{S} _{XYZ}}$ coupled with $\left\{ m_{j}\right\} _{j=1}^{\mathcal{S}_{XYZ}}$\ where $m_{j}$\ counts how many times realization $j$ occurs in the sub-sample when $W_{n}=w$. We then estimate the parameters of interest by maximizing the following conditional\ log-likelihood function

equation[equation omitted — 329 chars of source]

where $\mathbf{p}=\left( p_{X|X^{\ast },W},p_{Y|X^{\ast },W},p_{Z|X^{\ast },W},p_{X^{\ast }|W}\right) $ lies in the parameter space $\mathcal{P}$ that satisfies the constraints that components of $\mathbf{p}$\ constitute to valid probability distributions and the inequality relations in Assumption 4. We do this for all $w$ in the support of $W$. Once the nonparametric estimators of $\left( f_{X|X^{\ast },W},f_{Y|X^{\ast },W},f_{Z|X^{\ast },W},f_{X^{\ast }|W},f_{W}\right) $ are available, we can proceed to the parametric estimation stage.

Constrained maximum likelihood estimation is not a computationally simple task. There are $\mathcal{S}_{X}\left( \mathcal{S}_{X}+\mathcal{S}_{Y}+ \mathcal{S}_{Z}-3\right) +\mathcal{S}_{X}-1$ free parameters to optimize over in ((ref)) for each possible value that $W$ takes.

sub-sample partitioned according to different values of $W$. I.e. we have to solve this type of optimization problem $\mathcal{S}_{W}$\ times. The numerical challenge increases with the support size of the variables in the model. Furthermore, the objective function is not concave so there can be many local maxima. In practice, we suggest numerical searches should be performed at different starting points in order to help locate the global maximum.

Parametric Estimation

Once an estimator for $f_{X^{\ast }|Q}$ is available, population quantities involving $X^{\ast }$ such as $E\left[ X^{\ast }|Q\right] $ and $\Pr \left[ X^{\ast }\leq i|Q\right] $ can now be estimated even if we do not observe latent LS. For the linear probability model, from ((ref)), it can be more convenient to write $\beta _{C}=\left( E\left[ \widetilde{Q}\widetilde{Q} ^{\top }\right] \right) ^{-1}E\left[ \widetilde{Q}E\left[ X^{\ast }| \widetilde{Q}\right] \right] $. We can then estimate $\beta _{C}$\ by replacing the (unconditional)\ expectation by the sample counterparts.

For the ordered probit model, we can estimate $\sigma $ by replacing $\Pr \left[ X^{\ast }\leq i|Q\right] $\ in ((ref)) by its estimator. Then we can construct estimators for $\beta _{O}$ and $\mu _{O}$ by replacing the population moments in ((ref)) and ((ref)) respectively\ by their sample counterparts. Alternatively, a perhaps more convenient numerical approach is to estimate the parameters of interest with the build-in functions of statistical software providing it with the skedastic function based on ((ref)).

commentA perhaps more convenient alternative approach is to use built-in programs such as STATA to estimate the ordered probit while providing it with the skedastic function based on ((ref)). \textcolor{blue}{Do you mean oglm here? Because I'm not sure if there is Stata code for Chen& Chan }

Inference

We propose to perform inference by bootstrapping. A bootstrap sample can be generated by random resampling from the observed data with replacement. The estimators and tests of nonparametric probabilities and parameters in Propositions 1 and 2 have regular asymptotic properties that can be bootstrap as long as the true parameters lie in the interior of the parameter space (andrews_estimation_1999, andrews_inconsistency_2000 ). In practice, estimates of probabilities being close to $0$ or $1$, or any other a priori (if used) constraints (Assumptions 3 and 4) that appear to be numerically binding should raise concerns that the assumption of an interior solution is not being satisfied.

Test for presence of measurement error

We want to test the hypothesis of no measurement error in LS:

equation[equation omitted — 72 chars of source]

Suppose we have $\left( X,Y,Z\right) $\ that satisfies Assumptions 1 - 4. Then we can identify $f_{X^{\ast },X}$ from $f_{X^{\ast },X,Y,Z}$. One way to test ((ref)) directly is to look for evidence that $f_{X^{\ast },X}\left( x^{\ast },x\right) >0$ for some $x^{\ast }\neq x$. But performing such test is difficult because the null would imply that $f_{X^{\ast },X}\left( x^{\ast },x\right) =0$ for all $x^{\ast }\neq x$; parameters at the boundary will require a non-standard testing procedure. For example, see andrews_testing_2001. We instead follow the approach of wilhelm_testing_2018, who shows it is possible to construct a simple test for the presence of measurement errors under much weaker conditions and without the need to first identify the entire model.

Theorem 1 in wilhelm_testing_2018 states that: if $Y\enskip\bot \enskip Z\enskip|\enskip X^{\ast }$, then ((ref)) implies $Y\enskip\bot \enskip Z\enskip|\enskip X$. We can then construct a test to detect potential measurement errors based on a conditional independence hypothesis:

equation[equation omitted — 79 chars of source]

We state this as a proposition.

Proposition 3. Suppose $Y\enskip\bot \enskip Z\enskip| \enskip X^{\ast }$. Then violation of $H_{0}^{B}$\ implies violation of $H_{0}^{A}$.

The conditional independence assumption in Proposition 3 is already implied by our Assumption 1. Testing $H_{0}^{B}$\ is just a test of conditional independence on observed variables. There are many options available for consistent tests that are easy to construct. In this paper we use a Kolmogorov-Smirnov type statistic that is based on the sample counterpart of the following, equivalent, way to write ((ref)):

equation*[equation* omitted — 181 chars of source]

In our application we use the frequency estimator for $\left( f_{Y,Z|X},f_{Y|X},f_{Z|X}\right) $, which corresponds to the maximum likelihood estimator since $\left( X,Y,Z\right) $ are discrete. Denoting the frequency estimator by $\left( \widehat{f}_{Y,Z|X},\widehat{f}_{Y|X}, \widehat{f}_{Z|X}\right) $ , we have the following test statistic:

equation[equation omitted — 215 chars of source]

We perform inference by bootstrapping. We construct bootstrap critical values for $TS$\ from the percentiles of $\left\{ TS^{b}\right\} _{b=1}^{B}$ , where

equation[equation omitted — 357 chars of source]

and $\widehat{f}_{A|B}^{b}$ denotes the frequency estimator of $f_{A|B}$\ based on the bootstrap sample. These bootstrap critical values are consistent as long as $f_{X,Y,Z}$ takes values in the interior of $\left( 0,1\right) $ as discussed at the end of Section 3.\footnote{ Let $\mathcal{F}\left( x,y,z\right) :=f_{Y,Z|X}\left( y,z|x\right) -f_{Y|X}\left( y|x\right) f_{Z|X}\left( z|x\right) $ for $\left( x,y,z\right) \in \mathcal{S}_{XYZ}$. It is clear that $\mathcal{F}\left( x,y,z\right) $\ is a continuous function of $f_{X,Y,Z}$. Under random sampling, the asymptotic distribution of $\sqrt{N}\left( \widehat{f} _{X,Y,Z}-f_{X,Y,Z}\right) $ can be consistently estimated by $\sqrt{N}\left( \widehat{f}_{X,Y,Z}^{b}-\widehat{f}_{X,Y,Z}\right) $ since empirical measures can be bootstrapped (e.g. see gin`e_bootstrapping_1990). The asymptotic percentiles of $TS$\ can then be consistently estimated using $ \left\{ TS^{b}\right\} _{b=1}^{B}$ by an application of the Continuous Mapping Theorem.}

It is worth emphasizing that Proposition 3 only provides a sufficient condition to detect measurement errors. On the other hand, $H_{0}^{B}$\ generally does not imply $H_{0}^{A}$\ unless additional conditions hold on the joint distribution of $f_{X^{\ast },X,Y,Z}$. We refer the reader to wilhelm_testing_2018 for further details as to when the two hypotheses are equivalent.

Application

We begin this section by describing our dataset and explaining how it is used in our applications. We report the results of the test for the presence of measurement error in Section 5.2. We study the effect measurement error has on a general model of LS in Section 5.3.

Data

We use Wave 3 of the representative household longitudinal data from UK Understanding Society. The survey covers members of over 35,000 households in the United Kingdom. These data were collected between January 2011 and June 2013. We choose Wave 3 because, unlike the other waves, it includes questions on personality traits, which is important for us as we use neuroticism as one of the auxiliary variables for identification. Further details on the survey questions can be found in Appendix A.

Understanding Society measures LS on a scale from 1 -- \enquote{completely dissatisfied} to 7 -- \enquote{completely satisfied}. For our application, we aggregate responses to the LS question into 3 larger groups, where 1st group is those dissatisfied with life overall ( \enquote{completely dissatisfied} and \enquote{mostly dissatisfied}), 3rd group is those satisfied (\enquote{mostly satisfied} and \enquote{completely satisfied}) and the 2nd group is those in between (\enquote{somewhat dissatisfied}, \enquote{neither satisfied or dissatisfied} and \enquote{somewhat satisfied} ). For the GHQ-12 measure, which runs from 0 - \enquote{the least distressed} to 36 -- \enquote{the most distressed}, we construct $Z$ to share the same cardinality as $X$ by aggregating all responses below 33rd percentile in group 1, those between 33rd and 66th percentile in group 2, all the rest in group 3. The neuroticism score is originally calculated as an average of 3 questions on a scale from 1 to 7. Indicator $Y$ takes the value of 1 if the level of neuroticism of the individual is above the sample median and zero otherwise.

We aggregate the data on LS to ensure stable solutions for our constrained maximum likelihood with both the observed and bootstrap samples. This reduces the number of parameters to be estimated from 97 to 17 for each possible realization of the conditioning variables. In particular, we maximize each of our likelihood function 10 times using a different starting point to avoid local maxima. Almost all of our estimates converge to the same solution. On the other hand, if we use a 7-point scale for life satisfaction we often find numerical optimization starting at a different point leads to distinct local maxima; in this case we do not have the confidence that global solutions can be reached in feasible time.

We only use data for the respondents who reported satisfaction with life overall. This gives us 40,359\ observations from the 49,739 available in the survey\ (over 81%). Of those, 56% are women, 44% are men. All the participants are of age 16 or above. 34% of the respondents have a long-standing illness or disability, 51.8% are married, 23.1% have obtained a university degree, 5.3% are unemployed.

Our $W$ consists of: university degree (degree), gender (fem ), long-standing illness or disability (illness), income above the sample median (inc) and marital status (married). Each of the covariates is a binary variable. That gives $\mathcal{S}_{W}=2^{5}=32$. We are unable to condition on additional variables because some socioeconomic groups would have too few observations for nonparametric estimation. We compute our estimators as described in Section 3.3; in particular the skedastic function is nonparametric (see ((ref))).

Measurement error in reported LS

We test for the presence of measurement error in reported LS unconditionally and conditionally on the covariates. The unconditional test assumes $ H_{0}^{B}$\ under the null and uses ((ref)) as the test statistic and ( (ref)) to construct the critical values. Table (ref) compares the value of the test statistic against the bootstrap critical values at different significant levels. We see very strong evidence against the no measurement error hypothesis as $H_{0}^{B}$\ is rejected at $1\%$ significance level.

table[table omitted — 445 chars of source]

We next look for the presence of measurement error in various socioeconomic groups. The conditional test partitions the data into $\mathcal{S}_{W}$\ subgroups. In this case, for each $w\in \mathcal{S}_{W}$\ we consider the following hypothesis:

equation*[equation* omitted — 210 chars of source]

We alter ((ref)) and ((ref)) to accommodate the conditioning on $W$\ accordingly with the frequency estimator. They are then used respectively to construct test statistics and bootstrap critical values. Table (ref) gives the test results. The description of different socioeconomic groups (first column of (ref)) can be found in Appendix (ref).

table[table omitted — 2,059 chars of source]

We also find very strong evidence that measurement error exists for many subgroups of the population. In particular, we reject $H_{0}^{C}\left( w\right) $\ at 1% significance level for 22 out 32 of socioeconomic groups. Out of the 10 groups we do not reject the null at 1%, we reject 4 of them at 5%. It is worth noting that the number of observations in these groups are very small relative to the rest, especially for the groups that we do not reject the null. The lack of (stronger) evidence to detect measurement error in some of those groups may be due to small sample size.

Estimation results

Our main results will focus on the distribution of the reported and latent LS and their ordered probit estimates. In particular, parameter estimates from probit models are to be interpreted as a component of the conditional median of the underlying continuous happiness variable. Before we present them, we consider the effects of reducing the support of reported LS\ from a 7-point scale to a 3-point scale as well as from leaving out some other covariates. In addition to the variables we have already introduced we will also use: logarithm of gross personal income (l_inc), unemployment dummy (unempl) and age (age) and age squared (age2 ).\footnote{ We need to reduce the support of LS for numerical stability of the maximum likelihood procedure and limit the number and support of covariates in order to ensure there is a sufficient number of observations with each socioeconomic group. For example, from Table (ref), we have 10 socioeconomic groups with under 500 observations (with DH the lowest at 117). If we split the sample further with employment status (only 5.3% are unemployed) and age bands, there will be groups with too few observations to estimate 17 parameters.}

Table (ref) reports the estimates for the linear projection model and for the ordered probit model for reported LS with full support (7-point scale) and reduced support (3-point scale). Here we use personal income instead of the dummy indicator that a respondent's income is above the median or not. In this case the heteroskedastic ordered probit is fully parametric. It is estimated using the oglm STATA command, where the skedastic function is specified by an exponential function with a linear index (williams_fitting_2010), in order to abstract away from the need to select tuning parameters from nonparametric estimation (e.g. with kernel smoothing, see chen_rates_2003).

Reducing the support of LS has negligible or no difference in how covariates affect LS apart from income, where the positive income effect on LS measured on a 3-point scale is much more pronounced. This pattern holds in all models. In particular, we note the similarities between the least squares and the probit estimates are uniform for all covariates (cf. ferrericarbonell_how_2004) as well as the similarities between results from homoskedastic and heteroskedastic models (cf. Chen et al. ( chen_have_2019)\footnote{ This empirical indifference is in stark contrast to the theoretical implication illustrated in bond_sad_nodate.}). These results are largely consistent with the literature. Married people are more satisfied with their lives than their non-married counterparts. Long-standing illnesses or disability and unemployment significantly reduce LS. Women report to be more satisfied than men. The effect from age supports the U-shaped pattern based on a quadratic specification. Money does buy some happiness. Although there are some conflicted findings on the income effect, the literature in general seems to find support for the general idea that income influences LS positively with diminishing returns (e.g. see clark_relative_2008). Education is known to influence LS indirectly through the increase in income and health. A positive effect from having more education is common result for the studies that cannot fully control for health\footnote{ The dataset does not allow us to fully control for the state of health and we only account for the presence of long-standing illness or disability.}, including those for the UK (see, e.g., dolan_we_2008).

table[table omitted — 2,165 chars of source]
table[table omitted — 1,561 chars of source]

Table (ref) contains analogous statistics to Table (ref)\ but are based on the reduced set of covariates that we later use to estimate latent LS. Most of the results between the two tables are qualitatively very similar. One notable difference, again, is on the income effect. Table (ref) reports that the income effect remains positive in all cases when the reduced support LS, but the full support LS yields negative estimates with the probit model. The negative income effect is, however, weak as it is insignificant and significant at 10% in the homoskedastic and heteroskedastic cases respectively. Our discussion on the income effect from the previous paragraph applies. Importantly, Table (ref)\ and (ref) suggest that using a median income dummy and omitting unemployment and age have little impact, as well as reducing the support of LS.

We now provide the estimates from reported and latent LS. Table (ref) reports the estimates from the linear projection model and the ordered probit models for reported and latent LS. Here the skedastic function for the heteroskedastic ordered probit model is estimated nonparametrically and the normalization is as discussed in Section (ref) .

table[table omitted — 1,329 chars of source]

The results show that the latent LS estimates are qualitatively very similar across all three models. There is a difference in the signs of the estimates of income but the income effect is weak and insignificant. Comparing the results from the models with reported and latent LS, we find the two prominent predictors of LS agree on their effects: health and interpersonal relationships. People who suffer from long-standing illness or disability are less satisfied with their life, while married people are more satisfied than their single counterparts. While the effect of health becomes more pronounces when we control for the presence of measurement error, it appear to have substituted the effect on education and income, making them insignificant with very high p-values. The most striking difference we find is the gender effect: the female dummy has a positive coefficient for the reported LS, but negative for the latent one. The same results hold when we use different constructs of GHQ as a robustness check. See Appendix (ref).

Our results provide a potential explanation of the gender puzzle based on systematic differences in the misreporting behavior between men and women. While LS has been widely accepted to be correlated with health and interpersonal relationship in obvious ways (see, e.g. helliwell_world_2012), the correlation between well-being and gender observed in practice is less intuitive. Many surveys find that females report themselves to be more satisfied with their life than men, e.g. see dolan_we_2008. These findings are in contradiction with being worse off in many measurable social and economic outcomes, which are known to be the sources of well-being (pay gap and unemployment gap, to name a few). More recently, meisenberg_gender_2015 use a dataset of 90 countries represented in the World Values Survey to find that gender equality, gainful employment and prolonged schooling decrease female well-being; while women are happier in the countries that maintain traditional gender roles.

In order to better understand the difference in reporting behavior of men and women, we consider 4 particular socioeconomic groups of respondents presented in Table (ref). Group O contains single men with income below the median, with no degree and no long-standing illness or disability. Group H contains the same respondents who suffer from long standing health issues. The other two groups are the female respondents with the same characteristics. Distributions for the other groups are presented in Appendix (ref).

table[table omitted — 477 chars of source]
table[table omitted — 2,163 chars of source]

Comparing the upper (0 and H) and the lower (F and FH) blocks of Table (ref) explains the different signs of the gender dummy coefficients in two models. The distribution of reported LS, $\mathbf{M}_{X}$, is similar for men (upper block) and women (lower block) with women slightly more likely to report the high state, hence positive coefficient for the gender dummy. The comparison of latent distributions, $\mathbf{M}_{X^{\ast }}$, shows the opposite: women are more likely to be in the low state and less likely to be in the high one. However, we do not observe lower levels of LS among women in the data, because they misreport in a systematically different way compared to men. Comparing the matrices of misreporting probabilities, $\mathbf{M}_{X|X^{\ast }}$, shows that though all the respondents are prone to report higher states that they latently are, women do it more emphatically. I.e. women are more likely than men to report the highest state, regardless of their latent state.

While our econometric analysis cannot provide a behavioral answer to this finding, it helps identify differences in reporting behavior for further research to be done to rationalize them. At the moment we can only offer potential explanations. One particular conjecture is based on the distinct patterns of {happy image management} induced by the influence of gender roles and social stereotypes. kahneman_well-being:_1999 and references therein suggest that according to traditional gender roles women are usually seen as more cheerful and enthusiastic. As a result women might report higher states conforming with the existing norm.

Conclusion

There is an enormous interest in using subjective well-being data in economics and related disciplines. Existing research almost always ignores measurement error despite the fact that the literature acknowledges it is expected to be present. In particular, the error is non-classical and its potential effects on subsequent analysis is completely unknown. In this paper we use novel nonparametric techniques to formally test for the presence of measurement error and empirically investigate its effects.

Our test is based on the idea proposed in wilhelm_testing_2018 and we use a misclassification model of hu_identification_2008. The application of these nonparametric methods in itself is not entirely trivial. Primarily there is an empirical challenge in finding appropriate data that satisfies assumptions somewhat analogous to finding a good instrument. We use Wave 3 of the UK Understanding Society survey because it is the only wave that contains questions on neuroticism that we believe is crucial for our empirical study. We also have to convert existing nonparametric identification results into parametric ones because a fully nonparametric approach has a limited scope in practice as it does not allow to include many covariates, not to mention that all the benchmark results in the literature are parametric. For a parametric estimation, we in particular advocate the use of a heteroskedastic ordered response model to analyze wellbeing data that builds on the argument of chen_have_2019 to avoid recent critiques highlighted in bond_sad_nodate.

We find evidence of measurement error in LS for the whole sample as well as 26 out of 32 socioeconomic subgroups of the data. We use covariates that define these subgroups to estimate a model of LS. We find the most important drivers of LS are the same in both models of latent LS and reported LS. The most notable difference is the gender effect. The happiness literature often finds women reporting higher levels of well-being, despite being worse of in measurable objective outcomes, e.g. income, employment, etc. This puzzle can be rationalized by our model because women are more likely to report themselves to be happier than they actually are compared to men.

The puzzling relations between female well-being and socioeconomic measures were also found in the panel data. stevenson_paradox_2009 show that reported well-being of women in the United States declined in the last 35 years despite the improvement of women's positions in many objective outcomes. The authors label this result as The Paradox of Declining Female Happiness. In order to further investigate whether the gender puzzle can be explained by measurement error, we need to have panel data to extend our analysis.

One methodological recommendation of our research is for future surveys to consider collecting data that increase the scope to apply modern econometric techniques to solve old problems like measurement error. For example, the UK Understanding Society data is in fact longitudinal. But the lack of information on neuroticism from all Waves other than Wave 3 prevents us from controlling from individual specific effects that would be very helpful in well-being studies.

commentOther comments. e.g. about u-shape. why do we care about u-shape? want to think about policy effect of this in relation to our results. also are there references for female gender effect? The methodology we propose is applicable in other empirical contexts that use subjective data, e.g. analysis of trust or perceived inequality. It can be used in the analysis of those subjective variables, if relevant auxiliary variables are available.