arXiv 27 Jan 2025 · Statistics — Machine Learning
arXiv:2501.15753 · PDF · DOI · OpenAlex · Extracted main text
This paper develops a scale-insensitive framework for neural network significance testing, substantially generalizing existing approaches through three key innovations. First, we replace metric entropy calculations with Rademacher complexity bounds, enabling the analysis of neural networks without requiring bounded weights or specific architectural constraints. Second, we weaken the regularity conditions on the target function to require only Sobolev space membership $H^s([-1,1]^d)$ with $s > d/2$, significantly relaxing previous smoothness assumptions while maintaining optimal approximation rates. Third, we introduce a modified sieve space construction based on moment bounds rather than weight constraints, providing a more natural theoretical framework for modern deep learning practices. Our approach achieves these generalizations while preserving optimal convergence rates and establishing valid asymptotic distributions for test statistics. The technical foundation combines localization theory, sharp concentration inequalities, and scale-insensitive complexity measures to handle unbounded weights and general Lipschitz activation functions. This framework better aligns theoretical guarantees with contemporary deep learning practice while maintaining mathematical rigor.
appendix boundary found by appendix_command · 45% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Fallahgoul, Franstianto \ Lin (2024) `Asset pricing with neural networks: Significance tests', Journal of Econometrics 238(1), 105574 | 1.000 | 16 | 3 | 100% |
| 2 | Horel \ Giesecke (2020) `Significance tests for neural networks', Journal of Machine Learning Research 21(227), 1–29 | 1.000 | 11 | 3 | 100% |
| 3 | Farrell, Liang \ Misra (2021) `Deep neural networks for estimation and inference', Econometrica 89(1), 181–213 | 0.952 | 29 | 5 | 86% |
| 4 | Bartlett, Bousquet \ Mendelson (2005) `Local rademacher complexities', The Annals of Statistics 33, 1497–1537 | 0.737 | 3 | 2 | 100% |
| 5 | Koltchinskii \ Panchenko (2000) Rademacher processes and bounding the risk of function learning, in `High Dimensional Probability II', Springer, pp. 443–457 | 0.737 | 3 | 2 | 100% |
| 6 | Yarotsky (2017) `Error bounds for approximations with deep relu networks', Neural networks 94, 103–114 | 0.737 | 3 | 2 | 100% |
| 7 | Yarotsky (2018) Optimal approximation of continuous functions by very deep relu networks, in `in 31st Conference on learning theory', PMLR, pp.… | 0.737 | 3 | 2 | 100% |
| 8 | Anthony \ Bartlett (1999) Neural Network Learning: Theoretical Foundations, Cambridge University Press, Cambridge | 0.644 | 2 | 2 | 100% |
| 9 | Goodfellow, Bengio \ Courville (2016) Deep Learning, MIT Press, Cambridge | 0.644 | 2 | 2 | 100% |
| 10 | van der Vaart \ Wellner (1996) Weak Convergence and Empirical Processes: With Applications to Statistics, Springer Series in Statistics, Springer, New York | 0.585 | 5 | 4 | 20% |
Showing the top 10 of 18 scored citations.