Anish Agarwal, Munther Dahleh, Devavrat Shah, Dennis Shen
arXiv 30 Sep 2021 · Econometrics · 10 citations (OpenAlex)
arXiv:2109.15154 · PDF · DOI · OpenAlex · Extracted main text
Matrix completion is the study of recovering an underlying matrix from a sparse subset of noisy observations. Traditionally, it is assumed that the entries of the matrix are "missing completely at random" (MCAR), i.e., each entry is revealed at random, independent of everything else, with uniform probability. This is likely unrealistic due to the presence of "latent confounders", i.e., unobserved factors that determine both the entries of the underlying matrix and the missingness pattern in the observed matrix. For example, in the context of movie recommender systems -- a canonical application for matrix completion -- a user who vehemently dislikes horror films is unlikely to ever watch horror films. In general, these confounders yield "missing not at random" (MNAR) data, which can severely impact any inference procedure that does not correct for this bias. We develop a formal causal model for matrix completion through the language of potential outcomes, and provide novel identification arguments for a variety of causal estimands of interest. We design a procedure, which we call "synthetic nearest neighbors" (SNN), to estimate these causal estimands. We prove finite-sample consistency and asymptotic normality of our estimator. Our analysis also leads to new theoretical results for the matrix completion literature. In particular, we establish entry-wise, i.e., max-norm, finite-sample consistency and asymptotic normality results for matrix completion with MNAR data. As a special case, this also provides entry-wise bounds for matrix completion with MCAR data. Across simulated and real data, we demonstrate the efficacy of our proposed estimator.
appendix boundary found by appendix_command · 87% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Ma, W. and Chen, G. H (2019) Missing not at random in matrix completion: The effectiveness of estimating missingness probabilities under a low nuclear norm a… | 1.000 | 13 | 4 | 100% |
| 2 | Agarwal, A., Shah, D., Shen, D., and Song, D (2021) On robustness of principal component regression self | 1.000 | 8 | 4 | 100% |
| 3 | Abadie, A., Diamond, A., and Hainmueller, J (2010) Synthetic control methods for comparative case studies: Estimating the effect of california's tobacco control program | 1.000 | 6 | 3 | 100% |
| 4 | Agarwal, A., Shah, D., and Shen, D (2021) On principal component regression in a high-dimensional error-in-variables setting self | 1.000 | 5 | 4 | 100% |
| 5 | Agarwal, A. and Singh, R (2021) Causal inference with corrupted data: Measurement error, missing values, discretization, and differential privacy self | 1.000 | 5 | 4 | 100% |
| 6 | Athey, S., Bayati, M., Doudchenko, N., Imbens, G., and Khosravi, K (2021) Matrix completion methods for causal panel data models | 1.000 | 5 | 4 | 100% |
| 7 | Yang, C., Ding, L., Wu, Z., and Udell, M (2021) Tenips: Inverse propensity sampling for tensor completion | 1.000 | 5 | 3 | 100% |
| 8 | Bhattacharya, S. and Chatterjee, S (2021) Matrix completion with data-dependent missingness probabilities | 0.956 | 8 | 5 | 88% |
| 9 | Bai, J. and Ng, S (2019) Matrix completion, counterfactuals, and factor analysis of missing data | 0.928 | 4 | 3 | 100% |
| 10 | Agarwal, A., Shah, D., and Shen, D (2021) Synthetic interventions self | 0.867 | 23 | 6 | 65% |
Showing the top 10 of 66 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.