EconBase
← All papers

Adaptive Discrete Smoothing for High-Dimensional and Nonlinear Panel Data

Xi Chen, Ye Luo, Martin Spindler

arXiv 30 Dec 2019 · Statistics — Methodology

arXiv:1912.12867 · PDF · DOI · OpenAlex · Extracted main text

Abstract

In this paper we develop a data-driven smoothing technique for high-dimensional and non-linear panel data models. We allow for individual specific (non-linear) functions and estimation with econometric or machine learning methods by using weighted observations from other individuals. The weights are determined by a data-driven way and depend on the similarity between the corresponding functions and are measured based on initial estimates. The key feature of such a procedure is that it clusters individuals based on the distance / similarity between them, estimated in a first stage. Our estimation method can be combined with various statistical estimation procedures, in particular modern machine learning methods which are in particular fruitful in the high-dimensional case and with complex, heterogeneous data. The approach can be interpreted as a \textquotedblleft soft-clustering\textquotedblright\ in comparison to traditional\textquotedblleft\ hard clustering\textquotedblright that assigns each individual to exactly one group. We conduct a simulation study which shows that the prediction can be greatly improved by using our estimator. Finally, we analyze a big data set from didichuxing.com, a leading company in transportation industry, to analyze and predict the gap between supply and demand based on a large set of covariates. Our estimator clearly performs much better in out-of-sample prediction compared to existing linear panel data estimators.

Citation extraction

27
references
35
in-text mentions
27
distinct cited
0
self-citations
7,402
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Xi Chen (2020) Adapative Smooting For Nonparametric Estimation, Technical report0.64422100%
2Racine \ Li (2004) `Nonparametric estimation of regression functions with both categorical and continuous data', Journal of Econometrics 119(1), 99…0.58531100%
3Vogt \ Linton (2017) `Classification of non-parametric regression functions in longitudinal data models', Journal of the Royal Statistical Society: S…0.58531100%
4Belloni \ Chernozhukov (2013) `Least Squares After Model Selection in High-dimensional Sparse Models', Bernoulli 19(2), 521–5470.51121100%
5Spindler \ Luo (2016) `High-Dimensional L2-Boosting: Rate of Convergence', arXiv preprint arXiv:1602.089270.51121100%
6Su, Shi \ Phillips (2016) `Identifying Latent Structures in Panel Data', Econometrica 84(6), 2215–22640.51121100%
7Arellano \ Bonhomme (2011) `Nonlinear panel data analysis', Annu0.40511100%
8Aloise, Deshpande, Hansen \ Popat (2009) `NP-hardness of Euclidean sum-of-squares clustering', Machine Learning 75, 245–2480.40511100%
9Belloni, Chen, Chernozhukov \ Hansen (2012) `Sparse Models and Methods for Optimal Instruments with an Application to Eminent Domain', Econometrica 80, 2369–24290.40511100%
10Belloni, Chernozhukov, Hansen \ Kozbur (2014) `Inference in High Dimensional Panel Models with an Application to Gun Control', Journal of Business & Economic Statistics 34(4)…0.40511100%

Showing the top 10 of 27 scored citations.