arXiv 2 Oct 2026 · Econometrics
arXiv:2610.02944 · PDF · Extracted main text
Cross-fitting is routine in much of applied research. While conventional confidence intervals that ignore cross-fold dependence are asymptotically valid in several settings, they undercover in many applications that share a common form of nonregularity: from the classic cross-validation problem of testing whether a fitted model outperforms another, to testing for heterogeneous treatment effects with machine learning, to estimating the value of a potentially non-unique optimal treatment regime. Exploiting a new locality condition, I show that a large class of cross-fitting estimators still satisfies a central limit theorem despite the nonregularity, but with an asymptotic variance that must be adjusted for the cross-fold correlation. Then, I propose a method for estimating this correlation and construct new confidence intervals that attain asymptotically nominal coverage. Finally, I show that the proposed confidence intervals attain approximately nominal coverage in a simulation study with random forests and neural networks.
appendix boundary found by appendix_command · 45% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Bayle, P., A. Bayle, L. Janson, and L. Mackey (2020) Cross-Validation Confidence Intervals for Test Error, in | 1.000 | 7 | 3 | 100% |
| 2 | Austern, M. and W. Zhou (2025) Asymptotics of Cross-Validation | 1.000 | 6 | 3 | 100% |
| 3 | Dudoit, S. and M. J. van der Laan (2005) Asymptotics of Cross-Validated Risk Estimation in Estimator Selection and Performance Assessment | 0.928 | 4 | 3 | 100% |
| 4 | Lei, J (2025) A Modern Theory of Cross-Validation through the Lens of Stability, arXiv:2505.23592 | 0.874 | 5 | 2 | 100% |
| 5 | Chernozhukov, V., M. Demirer, E. Duflo, and I. Fernández-Val (2025) Fisher–Schultz Lecture: Generic Machine Learning Inference on Heterogeneous Treatment Effects in Randomized Experiments, With an… | 0.811 | 4 | 2 | 100% |
| 6 | Luedtke, A. R. and M. J. van der Laan (2016) Statistical Inference for the Mean Outcome under a Possibly Non-Unique Optimal Treatment Strategy | 0.811 | 4 | 2 | 100% |
| 7 | Bayle, A., L. Janson, and L. Mackey (2026) The Relative Instability of Model Comparison with Cross-Validation, arXiv:2508.04409 | 0.737 | 3 | 2 | 100% |
| 8 | Fava, B (2025) Training and Testing with Multiple Splits: A Central Limit Theorem for Split-Sample Estimators, arXiv:2511.04957 self | 0.737 | 3 | 2 | 100% |
| 9 | Wager, S (2025) A Comment on: `Fisher–Schultz Lecture: Generic Machine Learning Inference on Heterogeneous Treatment Effects in Randomized Exper… | 0.737 | 3 | 2 | 100% |
| 10 | Blum, A., A. Kalai, and J. Langford (1999) Beating the Hold-out: Bounds for K-Fold and Progressive Cross-Validation, in | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 41 scored citations.