arXiv 14 Oct 2024 · Statistics — Machine Learning
arXiv:2410.11113 · PDF · DOI · OpenAlex · Extracted main text
This paper establishes statistical properties of deep neural network (DNN) estimators under dependent data. Two general results for nonparametric sieve estimators directly applicable to DNN estimators are given. The first establishes rates for convergence in probability under nonstationary data. The second provides non-asymptotic probability bounds on $\mathcal{L}^{2}$-errors under stationary $\beta$-mixing data. I apply these results to DNN estimators in both regression and classification contexts imposing only a standard H\"older smoothness assumption. The DNN architectures considered are common in applications, featuring fully connected feedforward networks with any continuous piecewise linear activation function, unbounded weights, and a width and depth that grows with sample size. The framework provided also offers potential for research into other DNN architectures and time-series applications.
appendix boundary found by appendix_command · 35% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Chen (2007) Chapter 76 large sample sieve estimation of semi-nonparametric models, in J. J | 1.000 | 7 | 3 | 100% |
| 2 | Chen \ Shen (1998) `Sieve Extremum Estimates for Weakly Dependent Data', Econometrica 66(2), 289–314 | 0.956 | 8 | 4 | 88% |
| 3 | Farrell, Liang \ Misra (2021) `Deep Neural Networks for Estimation and Inference', Econometrica 89(1), 181–213 | 0.944 | 19 | 5 | 84% |
| 4 | Anthony \ Bartlett (1999) Neural Network Learning: Theoretical Foundations, Cambridge University Press | 0.920 | 9 | 5 | 78% |
| 5 | Bartlett, Harvey, Liaw \ Mehrabian (2019) `Nearly-tight VC-dimension and Pseudodimension Bounds for Piecewise Linear Neural Networks', Journal of Machine Learning Researc… | 0.909 | 12 | 3 | 75% |
| 6 | Wooldridge \ White (1991) Some results on sieve estimation with dependent observations, in W. A | 0.874 | 7 | 2 | 100% |
| 7 | Chen \ Christensen (2015) `Optimal uniform convergence rates and asymptotic normality for series estimators under weak dependence and weak conditions', Jo… | 0.874 | 5 | 2 | 100% |
| 8 | Kurisu, Fukami \ Koike (2024) `Adaptive deep learning for nonlinear time series models' | 0.874 | 5 | 2 | 100% |
| 9 | Brown (2024) `Inference in partially linear models under dependent data with deep neural networks' self | 0.843 | 3 | 3 | 100% |
| 10 | Schmidt-Hieber (2020) `Nonparametric regression using deep neural networks with ReLU activation function', The Annals of Statistics 48(4) | 0.811 | 4 | 2 | 100% |
Showing the top 10 of 67 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | NEW APPROXIMATION RESULTS AND OPTIMAL ESTIMATION FOR FULLY CONNECTED DEEP NEURAL NETWORKS | 0.644 | 2 | 2 |