EconBase
← All papers

Nonlinear Binscatter Methods

Matias D. Cattaneo, Richard K. Crump, Max H. Farrell, Yingjie Feng

arXiv 21 Jul 2024 · Statistics — Methodology · 7 citations (OpenAlex)

arXiv:2407.15276 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Binned scatter plots are a powerful statistical tool for empirical work in the social, behavioral, and biomedical sciences. Available methods rely on a quantile-based partitioning estimator of the conditional mean regression function to primarily construct flexible yet interpretable visualization methods, but they can also be used to estimate treatment effects, assess uncertainty, and test substantive domain-specific hypotheses. This paper introduces novel binscatter methods based on nonlinear, possibly nonsmooth M-estimation methods, covering generalized linear, robust, and quantile regression models. We provide a host of theoretical results and practical tools for local constant estimation along with piecewise polynomial and spline approximations, including (i) optimal tuning parameter (number of bins) selection, (ii) confidence bands, and (iii) formal statistical tests regarding functional form or shape restrictions. Our main results rely on novel strong approximations for general partitioning-based estimators covering random, data-driven partitions, which may be of independent interest. We demonstrate our methods with an empirical application studying the relation between the percentage of individuals without health insurance and per capita income at the zip-code level. We provide general-purpose software packages implementing our methods in Python, R, and Stata.

Citation extraction

18
references
53
in-text mentions
18
distinct cited
4
self-citations
34,930
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Cattaneo, M. D., Crump, R. K., Farrell, M. H., and Feng, Y (2024) b), On Binscatter self1.000134100%
2Belloni, A., Chernozhukov, V., Chetverikov, D., and Fernandez-Val, I (2019) Conditional Quantile Processes based on Series or Many Regressors0.87472100%
3van der Vaart, A., and Wellner, J (1996) Weak Convergence and Empirical Processes: With Application to Statistics0.87462100%
4Belloni, A., Chernozhukov, V., Chetverikov, D., and Kato, K (2015) Some New Asymptotic Theory for Least Squares Series: Pointwise and Uniform Results0.69351100%
5Chernozhukov, V., Chetverikov, D., and Kato, K (2014) Anti-Concentration and Honest Adaptive Confidence Bands0.64441100%
6Bhatia, R (2013) Matrix Analysis0.64422100%
7Cattaneo, M. D., Crump, R. K., Farrell, M. H., and Feng, Y (2024) a), Binscatter Regressions, Working paper self0.64422100%
8Giné, E., and Nickl, R (2016) Mathematical Foundations of Infinite-Dimensional Statistical Models0.64422100%
9Schumaker, L (2007) Spline Functions: Basic Theory0.64422100%
10Demko, S (1977) Inverses of Band Matrices and Local Convergence of Spline Projections0.51121100%

Showing the top 10 of 18 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1On Binscatter0.92843
2Deep Learning for Individual Heterogeneity0.48162