Xi Chen, Ye Luo, Martin Spindler
arXiv 30 Dec 2019 · Statistics — Methodology
arXiv:1912.12867 · PDF · DOI · OpenAlex · Extracted main text
In this paper we develop a data-driven smoothing technique for high-dimensional and non-linear panel data models. We allow for individual specific (non-linear) functions and estimation with econometric or machine learning methods by using weighted observations from other individuals. The weights are determined by a data-driven way and depend on the similarity between the corresponding functions and are measured based on initial estimates. The key feature of such a procedure is that it clusters individuals based on the distance / similarity between them, estimated in a first stage. Our estimation method can be combined with various statistical estimation procedures, in particular modern machine learning methods which are in particular fruitful in the high-dimensional case and with complex, heterogeneous data. The approach can be interpreted as a \textquotedblleft soft-clustering\textquotedblright\ in comparison to traditional\textquotedblleft\ hard clustering\textquotedblright that assigns each individual to exactly one group. We conduct a simulation study which shows that the prediction can be greatly improved by using our estimator. Finally, we analyze a big data set from didichuxing.com, a leading company in transportation industry, to analyze and predict the gap between supply and demand based on a large set of covariates. Our estimator clearly performs much better in out-of-sample prediction compared to existing linear panel data estimators.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Xi Chen (2020) Adapative Smooting For Nonparametric Estimation, Technical report | 0.644 | 2 | 2 | 100% |
| 2 | Racine \ Li (2004) `Nonparametric estimation of regression functions with both categorical and continuous data', Journal of Econometrics 119(1), 99… | 0.585 | 3 | 1 | 100% |
| 3 | Vogt \ Linton (2017) `Classification of non-parametric regression functions in longitudinal data models', Journal of the Royal Statistical Society: S… | 0.585 | 3 | 1 | 100% |
| 4 | Belloni \ Chernozhukov (2013) `Least Squares After Model Selection in High-dimensional Sparse Models', Bernoulli 19(2), 521–547 | 0.511 | 2 | 1 | 100% |
| 5 | Spindler \ Luo (2016) `High-Dimensional L2-Boosting: Rate of Convergence', arXiv preprint arXiv:1602.08927 | 0.511 | 2 | 1 | 100% |
| 6 | Su, Shi \ Phillips (2016) `Identifying Latent Structures in Panel Data', Econometrica 84(6), 2215–2264 | 0.511 | 2 | 1 | 100% |
| 7 | Arellano \ Bonhomme (2011) `Nonlinear panel data analysis', Annu | 0.405 | 1 | 1 | 100% |
| 8 | Aloise, Deshpande, Hansen \ Popat (2009) `NP-hardness of Euclidean sum-of-squares clustering', Machine Learning 75, 245–248 | 0.405 | 1 | 1 | 100% |
| 9 | Belloni, Chen, Chernozhukov \ Hansen (2012) `Sparse Models and Methods for Optimal Instruments with an Application to Eminent Domain', Econometrica 80, 2369–2429 | 0.405 | 1 | 1 | 100% |
| 10 | Belloni, Chernozhukov, Hansen \ Kozbur (2014) `Inference in High Dimensional Panel Models with an Application to Gun Control', Journal of Business & Economic Statistics 34(4)… | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 27 scored citations.