EconBase
← All papers

High-Dimensional Metrics in R

Victor Chernozhukov, Chris Hansen, Martin Spindler

arXiv 5 Mar 2016 · Statistics — Machine Learning · 3 citations (OpenAlex)

arXiv:1603.01700 · PDF · DOI · OpenAlex · Extracted main text

Abstract

The package High-dimensional Metrics (\Rpackage{hdm}) is an evolving collection of statistical methods for estimation and quantification of uncertainty in high-dimensional approximately sparse models. It focuses on providing confidence intervals and significance testing for (possibly many) low-dimensional subcomponents of the high-dimensional parameter vector. Efficient estimators and uniformly valid confidence intervals for regression coefficients on target variables (e.g., treatment or policy variable) in a high-dimensional approximately sparse regression model, for average treatment effect (ATE) and average treatment effect for the treated (ATET), as well for extensions of these parameters to the endogenous setting are provided. Theory grounded, data-driven methods for selecting the penalization parameter in Lasso regressions under heteroscedastic and non-Gaussian errors are implemented. Moreover, joint/ simultaneous confidence intervals for regression coefficients of a high-dimensional sparse regression are implemented, including a joint significance test for Lasso regression. Data sets which have been used in the literature and might be useful for classroom demonstration and for testing new estimators are included. \R and the package \Rpackage{hdm} are open-source software projects and can be freely downloaded from CRAN: http://cran.r-project.org.

Citation extraction

15
references
30
in-text mentions
15
distinct cited
4
self-citations
15,264
main-text words

appendix boundary found by appendix_command · 93% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Belloni, Chernozhukov, and Kato (2014) Uniform post-selection inference for least absolute deviation regression and other Z-estimation problems1.00053100%
2Belloni, Chen, Chernozhukov, and Hansen (2012) Sparse Models and Methods for Optimal Instruments with an Application to Eminent Domain0.81142100%
3Barro and Lee (1994) Data set for a panel of 139 countries0.5112250%
4Acemoglu, Johnson, and Robinson (2001) The Colonial Origins of Comparative Development: An Empirical Investigation0.5112250%
5Belloni, Chernozhukov, Fernández-Val, and Hansen (2013) Program Evaluation with High-Dimensional Data0.51121100%
6Belloni, Chernozhukov, and Hansen (2014) Inference on Treatment Effects After Selection Amongst High-Dimensional Controls self0.51121100%
7Chernozhukov, Chetverikov, and Kato (2013) Gaussian approximations and multiplier bootstrap for maxima of sums of high-dimensional random vectors0.51121100%
8Chernozhukov, Hansen, and Spindler (2015) Valid Post-Selection and Post-Regularization Inference In Linear Models with Many Controls and Instruments self0.51121100%
9Belloni and Chernozhukov (2013) Least Squares After Model Selection in High-dimensional Sparse Models0.40511100%
10Belloni, Chernozhukov, and Hansen (2010) Inference for High-Dimensional Sparse Econometric Models self0.40511100%

Showing the top 10 of 15 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
12002.127100.40511
2Double Machine Learning and Automated Model Selection: A Cautionary Tale0.40511
32203.030510.40511
4Decomposing Inequalities using Machine Learning and Overcoming Common Support Issues0.40511