EconBase
← All papers

Causal Inference with Corrupted Data: Measurement Error, Missing Values, Discretization, and Differential Privacy

Anish Agarwal, Rahul Singh

arXiv 6 Jul 2021 · Econometrics · 2 citations (OpenAlex)

arXiv:2107.02780 · PDF · DOI · OpenAlex · Extracted main text

Abstract

The US Census Bureau will deliberately corrupt data sets derived from the 2020 US Census, enhancing the privacy of respondents while potentially reducing the precision of economic analysis. To investigate whether this trade-off is inevitable, we formulate a semiparametric model of causal inference with high dimensional corrupted data. We propose a procedure for data cleaning, estimation, and inference with data cleaning-adjusted confidence intervals. We prove consistency and Gaussian approximation by finite sample arguments, with a rate of $n^{ 1/2}$ for semiparametric estimands that degrades gracefully for nonparametric estimands. Our key assumption is that the true covariates are approximately low rank, which we interpret as approximate repeated measurements and empirically validate. Our analysis provides nonasymptotic theoretical contributions to matrix completion, statistical learning, and semiparametric statistics. Calibrated simulations verify the coverage of our data cleaning adjusted confidence intervals and demonstrate the relevance of our results for Census-derived data.

Citation extraction

0
references
3
in-text mentions
3
distinct cited
0
self-citations
1,478
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
Campbell02unmatched citation key Campbell020.40511100%
Chi81unmatched citation key Chi810.40511100%
Schubert13unmatched citation key Schubert130.40511100%

Showing the top 3 of 3 scored citations. 3 of these could not be matched to a bibliography entry, so only the citation key is shown.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Causal Matrix Completion1.00054
2A Causal Inference Framework for Data Rich Environments0.69371
3Adaptive Principal Component Regression with Applications to Panel Data0.51121
4Differentially Private Estimation of Heterogeneous Causal Effects0.40511
5Canonical correlation regression with noisy data0.40511
6Better Measurement or Larger Samples? Data Collection for Policy Learning with Unobserved Heterogeneity0.40511
72402.116520.00011