Ekaterina Oparina, Caspar Kaiser, Niccolò Gentile, Alexandre Tkatchenko, Andrew E. Clark, Jan-Emmanuel De Neve, Conchita D'Ambrosio
arXiv 1 Jun 2022 · Econometrics · 6 citations (OpenAlex)
arXiv:2206.00574 · PDF · DOI · OpenAlex · Extracted main text
There is a vast literature on the determinants of subjective wellbeing. International organisations and statistical offices are now collecting such survey data at scale. However, standard regression models explain surprisingly little of the variation in wellbeing, limiting our ability to predict it. In response, we here assess the potential of Machine Learning (ML) to help us better understand wellbeing. We analyse wellbeing data on over a million respondents from Germany, the UK, and the United States. In terms of predictive power, our ML approaches do perform better than traditional models. Although the size of the improvement is small in absolute terms, it turns out to be substantial when compared to that of key variables like health. We moreover find that drastically expanding the set of explanatory variables doubles the predictive power of both OLS and the ML approaches on unseen data. The variables identified as important by our ML algorithms - $i.e.$ material conditions, health, and meaningful social relations - are similar to those that have already been identified in the literature. In that sense, our data-driven ML results validate the findings from conventional approaches.
appendix boundary found by appendix_command · 67% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Hastie, Trevor, Tibshirani, Robert, Friedman, Jerome H., Friedman, J… (2009) The elements of statistical learning: data mining, inference, and prediction | 0.811 | 4 | 2 | 100% |
| 2 | Tibshirani, Robert (1996) Regression shrinkage and selection via the lasso | 0.737 | 3 | 2 | 100% |
| 3 | Friedman, Jerome H (2001) Greedy function approximation: a gradient boosting machine | 0.644 | 2 | 2 | 100% |
| 4 | Breiman, Leo (2001) Random forests | 0.511 | 2 | 1 | 100% |
| 5 | Kahneman, Daniel, Deaton, Angus (2010) High income improves evaluation of life but not emotional well-being. | 0.511 | 2 | 1 | 100% |
| 6 | Lundberg, Scott M., Erion, Gabriel G., Lee, Su-In (2018) Consistent Individualized Feature Attribution for Tree Ensembles | 0.405 | 1 | 1 | 100% |
| 7 | Reis, Itamar, Baron, Dalya, Shahaf, Sahar (2018) Probabilistic Random Forest: A Machine Learning Algorithm for Noisy Data Sets | 0.405 | 1 | 1 | 100% |
| 8 | Vovk, Vladimir, Schölkopf, Bernhard, Luo, Zhiyuan, Vovk, Vladimir (2013) Kernel Ridge Regression | 0.405 | 1 | 1 | 100% |
| 9 | Ahrens, Achim, Hansen, Christian B., Schaffer, Mark E (2020) lassopack: Model selection and prediction with regularized regression in Stata | 0.405 | 1 | 1 | 100% |
| 10 | Bond, Timothy N., Lang, Kevin (2019) The Sad Truth about Happiness Scales | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 42 scored citations.