EconBase
← All papers

On the Subbagging Estimation for Massive Data

Tao Zou, Xian Li, Xuan Liang, Hansheng Wang

arXiv 28 Feb 2021 · Statistics — Methodology · 1 citations (OpenAlex)

arXiv:2103.00631 · PDF · DOI · OpenAlex · Extracted main text

Abstract

This article introduces subbagging (subsample aggregating) estimation approaches for big data analysis with memory constraints of computers. Specifically, for the whole dataset with size $N$, $m_N$ subsamples are randomly drawn, and each subsample with a subsample size $k_N\ll N$ to meet the memory constraint is sampled uniformly without replacement. Aggregating the estimators of $m_N$ subsamples can lead to subbagging estimation. To analyze the theoretical properties of the subbagging estimator, we adapt the incomplete $U$-statistics theory with an infinite order kernel to allow overlapping drawn subsamples in the sampling procedure. Utilizing this novel theoretical framework, we demonstrate that via a proper hyperparameter selection of $k_N$ and $m_N$, the subbagging estimator can achieve $\sqrt{N}$-consistency and asymptotic normality under the condition $(k_Nm_N)/N\to \alpha \in (0,\infty]$. Compared to the full sample estimator, we theoretically show that the $\sqrt{N}$-consistent subbagging estimator has an inflation rate of $1/\alpha$ in its asymptotic variance. Simulation experiments are presented to demonstrate the finite sample performances. An American airline dataset is analyzed to illustrate that the subbagging estimate is numerically close to the full sample estimate, and can be computationally fast under the memory constraint.

Citation extraction

25
references
49
in-text mentions
25
distinct cited
2
self-citations
11,654
main-text words

appendix boundary found by appendix_titled_section at “Appendix” · 65% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Kim, K. I (2016) Higher order bias correcting moment equation for $M$-estimation and its higher order efficiency0.9285480%
2Mentch, L. and Hooker, G (2016) Quantifying uncertainty in random forests via confidence intervals and hypothesis tests0.87462100%
3Rilstone, P., Srivastava, V. K., and Ullah, A (1996) The second-order bias and mean squared error of nonlinear estimators0.8434475%
4Bühlmann, P (2003) Bagging, subagging and bragging for improving some prediction algorithms0.81142100%
5Peng, W., Coleman, T., and Mentch, L (2019) Asymptotic distributions and rates of convergence for random forests and other resampled ensemble learners0.73732100%
6Lee, S. and Ng, S (2020) An econometric perspective on algorithmic subsampling0.64441100%
7Gupta, P. and Bhattacharjee, G. P (1984) An efficient algorithm for random sampling without replacement0.64422100%
8Hastie, T., Tibshirani, R., and Friedman, J (2001) The Elements of Statistical Learning: Data Mining, Inference, and Prediction0.64422100%
9van der Vaart, A. W (1998) Asymptotic Statistics0.5112250%
10Lütkepohl, H (2005) New introduction to multiple time series analysis0.51121100%

Showing the top 10 of 25 scored citations.