Chiara Amorino, Christian Brownlees, Ankita Ghosh
arXiv 1 Nov 2025 · Econometrics
arXiv:2511.00597 · PDF · DOI · OpenAlex · Extracted main text
This paper develops a general concentration inequality for the suprema of empirical processes with dependent data. The concentration inequality is obtained by combining generic chaining with a coupling-based strategy. Our framework accommodates high-dimensional and heavy-tailed (sub-Weibull) data. We demonstrate the usefulness of our result by deriving non-asymptotic predictive performance guarantees for empirical risk minimization in regression problems with dependent data. In particular, we establish an oracle inequality for a broad class of nonlinear regression models and, as a special case, a single-layer neural network model. Our results show that empirical risk minimzaton with dependent data attains a prediction accuracy comparable to that in the i.i.d. setting for a wide range of nonlinear regression models.
appendix boundary found by appendix_command · 49% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Michel Talagrand (2005) The Generic Chaining: Upper and Lower Bounds of Stochastic Processes | 0.928 | 5 | 4 | 80% |
| 2 | Brownlees, C. and Gudmundsson, G. S (2025) Performance of Empirical Risk Minimization for Linear Regression with Dependent Data self | 0.928 | 4 | 3 | 100% |
| 3 | Brownlees, C. and Llorens-Terrazas, J (2025) Empirical Risk Minimization for Time Series: Nonparametric Performance Bounds for Prediction self | 0.928 | 4 | 3 | 100% |
| 4 | Merlevède, Florence and Peligrad, Magda (2002) On the Coupling of Dependent Random Variables and Applications | 0.843 | 5 | 3 | 60% |
| 5 | Jiang, Wenxin and Tanner, Martin (2010) Risk Minimization for Time Series Binary Choice with Variable Selection | 0.843 | 3 | 3 | 100% |
| 6 | Boucheron, S. and Lugosi, G. and Massart, P (2013) Concentration Inequalities: A Nonasymptotic Theory of Independence | 0.737 | 3 | 3 | 67% |
| 7 | Devroye, L. and Györfi, L. and Lugosi, G (1996) A Probabililstic Theory of Pattern Recognition | 0.737 | 3 | 2 | 100% |
| 8 | Doukhan, Paul (1994) Mixing | 0.644 | 2 | 2 | 100% |
| 9 | Vershynin, Roman (2026) High-Dimensional Probability: An Introduction with Applications in Data Science | 0.511 | 4 | 2 | 25% |
| 10 | Kuchibhotla, Arun Kumar and Chakrabortty, Abhishek (2022) Moving beyond sub-Gaussianity in high-dimensional statistics: applications in covariance estimation and linear regression | 0.511 | 3 | 2 | 33% |
Showing the top 10 of 22 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Post-selection inference for network structure 1 | 0.405 | 1 | 1 |