EconBase
← All papers

Stochastic Learning of Semiparametric Monotone Index Models with Large Sample Size

Qingsong Yao

arXiv 13 Sep 2023 · Econometrics · 1 citations (OpenAlex)

arXiv:2309.06693 · PDF · DOI · OpenAlex · Extracted main text

Abstract

I study the estimation of semiparametric monotone index models in the scenario where the number of observation points $n$ is extremely large and conventional approaches fail to work due to heavy computational burdens. Motivated by the mini-batch gradient descent algorithm (MBGD) that is widely used as a stochastic optimization tool in the machine learning field, I proposes a novel subsample- and iteration-based estimation procedure. In particular, starting from any initial guess of the true parameter, I progressively update the parameter using a sequence of subsamples randomly drawn from the data set whose sample size is much smaller than $n$. The update is based on the gradient of some well-chosen loss function, where the nonparametric component is replaced with its Nadaraya-Watson kernel estimator based on subsamples. My proposed algorithm essentially generalizes MBGD algorithm to the semiparametric setup. Compared with full-sample-based method, the new method reduces the computational time by roughly $n$ times if the subsample size and the kernel function are chosen properly, so can be easily applied when the sample size $n$ is large. Moreover, I show that if I further conduct averages across the estimators produced during iterations, the difference between the average estimator and full-sample-based estimator will be $1/\sqrt{n}$-trivial. Consequently, the average estimator is $1/\sqrt{n}$-consistent and asymptotically normally distributed. In other words, the new estimator substantially improves the computational speed, while at the same time maintains the estimation accuracy.

Citation extraction

28
references
58
in-text mentions
28
distinct cited
0
self-citations
14,615
main-text words

appendix boundary found by appendix_titled_section at “Appendix” · 64% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Elhanan Helpman, Marc Melitz, and Yona Rubinstein (2008) Estimating trade flows: Trading partners and trading volumes1.00073100%
2Jean-Jacques Forneron (2022) Estimation and inference by stochastic optimization1.00054100%
3Hidehiko Ichimura (1993) Semiparametric least squares (sls) and weighted sls estimation of single-index models1.00053100%
4Shakeeb Khan, Xiaoying Lan, and Elie Tamer (2023) Estimating high dimensional monotone index models by iterative convex optimization10.81711555%
5Roger W Klein and Richard H Spady (1993) An efficient semiparametric estimator for binary response models0.73732100%
6Boris T Polyak and Anatoli B Juditsky (1992) Acceleration of stochastic approximation by averaging0.73732100%
7Léon Bottou, Frank E Curtis, and Jorge Nocedal (2018) Optimization methods for large-scale machine learning0.64422100%
8Sebastian Ruder (2016) An overview of gradient descent optimization algorithms0.64422100%
9Alekh Agarwal, Sham Kakade, Nikos Karampatziakis, Le Song, and Grego… (2014) Least squares revisited: Scalable approaches for multi-class prediction0.40511100%
10Hyungtaik Ahn, Hidehiko Ichimura, James L Powell, and Paul A Ruud (2018) Simple estimators for invertible index models0.40511100%

Showing the top 10 of 28 scored citations.