Anish Agarwal, Jungjun Choi, Ming Yuan
arXiv 25 Mar 2026 · Statistics — Methodology
arXiv:2603.24833 · PDF · DOI · OpenAlex · Extracted main text
We introduce a flexible framework for high-dimensional matrix estimation to incorporate side information for both rows and columns. Existing approaches, such as inductive matrix completion, often impose restrictive structure-for example, an exact low-rank covariate interaction term, linear covariate effects, and limited ability to exploit components explained only by one side (row or column) or by neither-and frequently omit an explicit noise component. To address these limitations, we propose to decompose the underlying matrix as the sum of four complementary components: (possibly nonlinear) interaction between row and column characteristics; row characteristic-driven component, column characteristic-driven component, and residual low-rank structure unexplained by observed characteristics. By combining sieve-based projection with nuclear-norm penalization, each component can be estimated separately and these estimated components can then be aggregated to yield a final estimate. We derive convergence rates that highlight robustness across a range of model configurations depending on the informativeness of the side information. We further extend the method to partially observed matrices under both missing-at-random and missing-not-at-random mechanisms, including block-missing patterns motivated by causal panel data. Simulations and a real-data application to tobacco sales show that leveraging side information improves imputation accuracy and can enhance treatment-effect estimation relative to standard low-rank and spectral-based alternatives.
appendix boundary found by appendix_command · 51% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Bai, Jushan and Ng, Serena (2021) Matrix completion, counterfactuals, and factor analysis of missing data | 1.000 | 6 | 3 | 100% |
| 2 | Yan, Yuling and Wainwright, Martin J (2024) Entrywise inference for causal panel data: A simple and instance-optimal approach | 0.928 | 4 | 3 | 100% |
| 3 | Fan, Jianqing and Liao, Yuan and Wang, Weichen (2016) Projected principal component analysis in factor models | 0.874 | 5 | 2 | 100% |
| 4 | Choi, Jungjun and Yuan, Ming (2024) Matrix completion when missing is not at random and its applications in causal panel data models self | 0.843 | 4 | 3 | 75% |
| 5 | Abadie, Alberto and Diamond, Alexis and Hainmueller, Jens (2010) Synthetic control methods for comparative case studies: Estimating the effect of California’s tobacco control program | 0.511 | 2 | 1 | 100% |
| 6 | Ahn, Seung C and Horenstein, Alex R (2013) Eigenvalue ratio test for the number of factors | 0.511 | 2 | 1 | 100% |
| 7 | Athey, Susan and Bayati, Mohsen and Doudchenko, Nikolay and Imbens,… (2021) Matrix completion methods for causal panel data models | 0.511 | 2 | 1 | 100% |
| 8 | Chiang, Kai-Yang and Hsieh, Cho-Jui and Dhillon, Inderjit S (2015) Matrix completion with noisy side information | 0.511 | 2 | 1 | 100% |
| 9 | Jain, Prateek and Dhillon, Inderjit S (2013) Provable inductive matrix completion | 0.511 | 2 | 1 | 100% |
| 10 | Ma, Shujie and Niu, Po-Yao and Zhang, Yichong and Zhu, Yinchu (2025) Statistical inference for noisy matrix completion incorporating auxiliary information | 0.511 | 2 | 1 | 100% |
Showing the top 10 of 29 scored citations.