EconBase
← All papers

Robust Matrix Estimation with Side Information

Anish Agarwal, Jungjun Choi, Ming Yuan

arXiv 25 Mar 2026 · Statistics — Methodology

arXiv:2603.24833 · PDF · DOI · OpenAlex · Extracted main text

Abstract

We introduce a flexible framework for high-dimensional matrix estimation to incorporate side information for both rows and columns. Existing approaches, such as inductive matrix completion, often impose restrictive structure-for example, an exact low-rank covariate interaction term, linear covariate effects, and limited ability to exploit components explained only by one side (row or column) or by neither-and frequently omit an explicit noise component. To address these limitations, we propose to decompose the underlying matrix as the sum of four complementary components: (possibly nonlinear) interaction between row and column characteristics; row characteristic-driven component, column characteristic-driven component, and residual low-rank structure unexplained by observed characteristics. By combining sieve-based projection with nuclear-norm penalization, each component can be estimated separately and these estimated components can then be aggregated to yield a final estimate. We derive convergence rates that highlight robustness across a range of model configurations depending on the informativeness of the side information. We further extend the method to partially observed matrices under both missing-at-random and missing-not-at-random mechanisms, including block-missing patterns motivated by causal panel data. Simulations and a real-data application to tobacco sales show that leveraging side information improves imputation accuracy and can enhance treatment-effect estimation relative to standard low-rank and spectral-based alternatives.

Citation extraction

29
references
55
in-text mentions
29
distinct cited
2
self-citations
10,662
main-text words

appendix boundary found by appendix_command · 51% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Bai, Jushan and Ng, Serena (2021) Matrix completion, counterfactuals, and factor analysis of missing data1.00063100%
2Yan, Yuling and Wainwright, Martin J (2024) Entrywise inference for causal panel data: A simple and instance-optimal approach0.92843100%
3Fan, Jianqing and Liao, Yuan and Wang, Weichen (2016) Projected principal component analysis in factor models0.87452100%
4Choi, Jungjun and Yuan, Ming (2024) Matrix completion when missing is not at random and its applications in causal panel data models self0.8434375%
5Abadie, Alberto and Diamond, Alexis and Hainmueller, Jens (2010) Synthetic control methods for comparative case studies: Estimating the effect of California’s tobacco control program0.51121100%
6Ahn, Seung C and Horenstein, Alex R (2013) Eigenvalue ratio test for the number of factors0.51121100%
7Athey, Susan and Bayati, Mohsen and Doudchenko, Nikolay and Imbens,… (2021) Matrix completion methods for causal panel data models0.51121100%
8Chiang, Kai-Yang and Hsieh, Cho-Jui and Dhillon, Inderjit S (2015) Matrix completion with noisy side information0.51121100%
9Jain, Prateek and Dhillon, Inderjit S (2013) Provable inductive matrix completion0.51121100%
10Ma, Shujie and Niu, Po-Yao and Zhang, Yichong and Zhu, Yinchu (2025) Statistical inference for noisy matrix completion incorporating auxiliary information0.51121100%

Showing the top 10 of 29 scored citations.