EconBase
← All papers

Reproducible Aggregation of Sample-Split Statistics

David M. Ritzwoller, Joseph P. Romano

arXiv 23 Nov 2023 · Econometrics

arXiv:2311.14204 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Statistical inference is often simplified by sample-splitting. This simplification comes at the cost of the introduction of randomness not native to the data. We propose a simple procedure for sequentially aggregating statistics constructed with multiple splits of the same sample. The user specifies a bound and a nominal error rate. If the procedure is implemented twice on the same data, the nominal error rate approximates the chance that the results differ by more than the bound. We illustrate the application of the procedure to several widely applied econometric methods.

Citation extraction

109
references
295
in-text mentions
109
distinct cited
3
self-citations
15,260
main-text words

appendix boundary found by appendix_titled_section at “Data and Simulations\label{app: simulation appendix}” · 31% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C… (2018) Double/debiased machine learning for treatment and structural parameters: Double/debiased machine learning0.92810580%
2DiCiccio, C. J., DiCiccio, T. J., and Romano, J. P (2020) Exact tests via multiple data splitting self0.9285480%
3Chernozhukov, V., Demirer, M., Duflo, E., and Fernández-Val, I (2023) Generic machine learning inference on heterogenous treatment effects in randomized experiments, with an application to immunizat…0.8947471%
4Casey, K., Kamara, A. B., and Meriggi, N. F (2021) An experiment in candidate selection0.85727563%
5Wasserman, L., Ramdas, A., and Balakrishnan, S (2020) Universal inference0.8435460%
6Chakravorty, B., Arulampalam, W., Bhatiya, A. Y., Imbert, C., and Ra… (2024) Can information about jobs improve the effectiveness of vocational training? experimental evidence from india0.83226458%
7Haushofer, J., Niehaus, P., Paramo, C., Miguel, E., and Walker, M. W (2022) Targeting impact versus deprivation0.79424350%
8Beaman, L., Karlan, D., Thuysbaert, B., and Udry, C (2023) Selection into credit markets: Evidence from agriculture in mali0.79418350%
9Kale, S., Kumar, R., and Vassilvitskii, S (2011) Cross-validation and mean-square stability0.73732100%
10Kumar, R., Lokshtanov, D., Vassilvitskii, S., and Vattani, A (2013) Near-optimal bounds for cross-validation via loss stability0.73732100%

Showing the top 10 of 109 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Training and Testing with Multiple Splits: A Central Limit Theorem for Split-Sample Estimators1.00073
2Testing the Fairness-Accuracy Improvability of Algorithms0.64422
3Simultaneous Inference for Local Structural Parameters with Random Forests$^*$0.64422
4Randomization Inference: Theory and Applications0.40511
5Universal Inference for Incomplete Discrete Choice Models0.40511
6Branching Fixed Effects: A Proposal for Communicating Uncertainty0.40511
7Post-selection inference for network structure 10.40511