EconBase
← All papers

Causal Matrix Completion

Anish Agarwal, Munther Dahleh, Devavrat Shah, Dennis Shen

arXiv 30 Sep 2021 · Econometrics · 10 citations (OpenAlex)

arXiv:2109.15154 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Matrix completion is the study of recovering an underlying matrix from a sparse subset of noisy observations. Traditionally, it is assumed that the entries of the matrix are "missing completely at random" (MCAR), i.e., each entry is revealed at random, independent of everything else, with uniform probability. This is likely unrealistic due to the presence of "latent confounders", i.e., unobserved factors that determine both the entries of the underlying matrix and the missingness pattern in the observed matrix. For example, in the context of movie recommender systems -- a canonical application for matrix completion -- a user who vehemently dislikes horror films is unlikely to ever watch horror films. In general, these confounders yield "missing not at random" (MNAR) data, which can severely impact any inference procedure that does not correct for this bias. We develop a formal causal model for matrix completion through the language of potential outcomes, and provide novel identification arguments for a variety of causal estimands of interest. We design a procedure, which we call "synthetic nearest neighbors" (SNN), to estimate these causal estimands. We prove finite-sample consistency and asymptotic normality of our estimator. Our analysis also leads to new theoretical results for the matrix completion literature. In particular, we establish entry-wise, i.e., max-norm, finite-sample consistency and asymptotic normality results for matrix completion with MNAR data. As a special case, this also provides entry-wise bounds for matrix completion with MCAR data. Across simulated and real data, we demonstrate the efficacy of our proposed estimator.

Citation extraction

66
references
179
in-text mentions
66
distinct cited
8
self-citations
19,170
main-text words

appendix boundary found by appendix_command · 87% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Ma, W. and Chen, G. H (2019) Missing not at random in matrix completion: The effectiveness of estimating missingness probabilities under a low nuclear norm a…1.000134100%
2Agarwal, A., Shah, D., Shen, D., and Song, D (2021) On robustness of principal component regression self1.00084100%
3Abadie, A., Diamond, A., and Hainmueller, J (2010) Synthetic control methods for comparative case studies: Estimating the effect of california's tobacco control program1.00063100%
4Agarwal, A., Shah, D., and Shen, D (2021) On principal component regression in a high-dimensional error-in-variables setting self1.00054100%
5Agarwal, A. and Singh, R (2021) Causal inference with corrupted data: Measurement error, missing values, discretization, and differential privacy self1.00054100%
6Athey, S., Bayati, M., Doudchenko, N., Imbens, G., and Khosravi, K (2021) Matrix completion methods for causal panel data models1.00054100%
7Yang, C., Ding, L., Wu, Z., and Udell, M (2021) Tenips: Inverse propensity sampling for tensor completion1.00053100%
8Bhattacharya, S. and Chatterjee, S (2021) Matrix completion with data-dependent missingness probabilities0.9568588%
9Bai, J. and Ng, S (2019) Matrix completion, counterfactuals, and factor analysis of missing data0.92843100%
10Agarwal, A., Shah, D., and Shen, D (2021) Synthetic interventions self0.86723665%

Showing the top 10 of 66 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Matrix Completion When Missing Is Not at Random and Its Applications in Causal Panel Data Models1.00093
2Covariate-Adjusted Deep Causal Learning for Heterogeneous Panel Data Models0.87452
32401.136650.81142
4Inferring Treatment Effects in Large Panels by Uncovering Latent Similarities0.73732
52402.116520.58531
6Inference for Low-rank Completion without Sample Splitting with Application to Treatment Effect Estimation0.51121
7Synthetic Blips: Generalizing Synthetic Controls for Dynamic Treatment Effects0.40511
8Network Synthetic Interventions: A Causal Framework for Panel Data Under Network Interference0.40511
9Strategyproof Decision-Making in Panel Data Settings and Beyond0.40511
10Double and Single Descent in Causal Inference with an Application to High-Dimensional Synthetic Control0.40511