arXiv 14 Nov 2025 · Econometrics
arXiv:2511.10995 · PDF · DOI · OpenAlex · Extracted main text
This paper studies double/debiased machine learning (DML) methods applied to weakly dependent data. We allow observations to be situated in a general metric space that accommodates spatial and network data. Existing work implements cross-fitting by excluding from the training fold observations sufficiently close to the evaluation fold. We find in simulations that this can result in exceedingly small training fold sizes, particularly with network data. We therefore seek to establish the validity of DML without cross-fitting, building on recent work by Chen et al. (2022). They study i.i.d. data and require the machine learner to satisfy a natural stability condition requiring insensitivity to data perturbations that resample a single observation. We extend these results to dependent data by strengthening stability to "neighborhood stability," which requires insensitivity to resampling observations in any slowly growing neighborhood. We show that existing results on the stability of various machine learners can be adapted to verify neighborhood stability.
appendix boundary found by appendix_command · 67% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Chernozhukov, V. and Chetverikov, D. and Demirer, M. and Duflo, E. a… (2018) Double/Debiased Machine Learning for Treatment and Structural Parameters | 1.000 | 7 | 3 | 100% |
| 2 | Emmenegger, C. and Spohn, M. and Bühlmann, P (2025) Treatment Effect Estimation from Observational Network Data Using Augmented Inverse Probability Weighting and Machine Learning | 0.965 | 10 | 3 | 90% |
| 3 | Bousquet, O. and Elisseeff, A (2002) Stability and Generalization | 0.928 | 5 | 3 | 80% |
| 4 | Chen, Q. and Syrgkanis, V. and Austern, M (2022) Debiased Machine Learning Without Sample-Splitting for Stable Estimators | 0.890 | 17 | 6 | 71% |
| 5 | Jenish, N. and Prucha, I (2009) Central Limit Theorems and Uniform Laws of Large Numbers for Arrays of Random Fields | 0.811 | 4 | 2 | 100% |
| 6 | Kissel, N. and Lei, J (2023) Black-Box Model Confidence Sets Using Cross-Validation with High-Dimensional Gaussian Comparison | 0.737 | 5 | 2 | 60% |
| 7 | Brown, C (2024) Inference in Partially Linear Models under Dependent Data with Deep Neural Networks | 0.737 | 3 | 2 | 100% |
| 8 | Gilbert, B. and Datta, A. and Casey, J. and Ogburn, E (2024) A Causal Inference Framework for Spatial Confounding | 0.737 | 3 | 2 | 100% |
| 9 | Farrell, M. and Liang, T. and Misra, S (2021) Deep Neural Networks for Estimation and Inference | 0.644 | 2 | 2 | 100% |
| 10 | Hardt, M. and Recht, B. and Singer, Y (2016) Train Faster, Generalize Better: Stability of Stochastic Gradient Descent | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 30 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.