Abhineet Agarwal, Anish Agarwal, Suhas Vijaykumar
arXiv 24 Mar 2023 · Statistics — Methodology
arXiv:2303.14226 · PDF · DOI · OpenAlex · Extracted main text
Consider a setting where there are $N$ heterogeneous units and $p$ interventions. Our goal is to learn unit-specific potential outcomes for any combination of these $p$ interventions, i.e., $N \times 2^p$ causal parameters. Choosing a combination of interventions is a problem that naturally arises in a variety of applications such as factorial design experiments, recommendation engines, combination therapies in medicine, conjoint analysis, etc. Running $N \times 2^p$ experiments to estimate the various parameters is likely expensive and/or infeasible as $N$ and $p$ grow. Further, with observational data there is likely confounding, i.e., whether or not a unit is seen under a combination is correlated with its potential outcome under that combination. To address these challenges, we propose a novel latent factor model that imposes structure across units (i.e., the matrix of potential outcomes is approximately rank $r$), and combinations of interventions (i.e., the coefficients in the Fourier expansion of the potential outcomes is approximately $s$ sparse). We establish identification for all $N \times 2^p$ parameters despite unobserved confounding. We propose an estimation procedure, Synthetic Combinations, and establish it is finite-sample consistent and asymptotically normal under precise conditions on the observation pattern. Our results imply consistent estimation given $poly(r) \times \left( N + s^2p\right)$ observations, while previous methods have sample complexity scaling as $\min(N \times s^2p, \ \ poly(r) \times (N + 2^p))$. We use Synthetic Combinations to propose a data-efficient experimental design. Empirically, Synthetic Combinations outperforms competing approaches on a real-world dataset on movie recommendations. Lastly, we extend our analysis to do causal inference where the intervention is a permutation over $p$ items (e.g., rankings).
appendix boundary found by appendix_command · 40% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Agarwal, Dahleh, Shah \ Shen (2023) Causal matrix completion, in `The Thirty Sixth Annual Conference on Learning Theory', PMLR, pp. 3821–3826 | 0.928 | 4 | 3 | 100% |
| 2 | George, Hunter \ Hunter (2005) Statistics for experimenters: design, innovation, and discovery, Wiley | 0.843 | 3 | 3 | 100% |
| 3 | Agarwal, Shah \ Shen (2020) `Synthetic interventions', arXiv preprint arXiv:2006.07691 | 0.830 | 7 | 3 | 57% |
| 4 | Agarwal, Shah \ Shen (2020) `On principal component regression in a high-dimensional error-in-variables setting', arXiv preprint arXiv:2010.14449 | 0.737 | 3 | 3 | 67% |
| 5 | Negahban \ Shah (2012) Learning sparse boolean polynomials, in `2012 50th Annual Allerton Conference on Communication, Control, and Computing (Allerton… | 0.737 | 3 | 2 | 100% |
| 6 | Syrgkanis \ Zampetakis (2020) Estimation and inference with trees and forests in high dimensions, in `Conference on learning theory', PMLR, pp. 3453–3454 | 0.693 | 6 | 3 | 33% |
| 7 | Anish Agarwal \ Song (2021) `On robustness of principal component regression', Journal of the American Statistical Association 116(536), 1731–1745 | 0.644 | 2 | 2 | 100% |
| 8 | Bertrand \ Mullainathan (2004) `Are emily and greg more employable than lakisha and jamal? a field experiment on labor market discrimination', American economi… | 0.644 | 2 | 2 | 100% |
| 9 | Candes \ Recht (2012) `Exact matrix completion via convex optimization', Communications of the ACM 55(6), 111–119 | 0.644 | 2 | 2 | 100% |
| 10 | Dasgupta, Pillai \ Rubin (2015) `Causal inference from 2 k factorial designs by using potential outcomes', Journal of the Royal Statistical Society: Series B: S… | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 63 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Synthetic Interventions | 0.511 | 2 | 1 |
| 2 | Adaptive Principal Component Regression with Applications to Panel Data | 0.405 | 2 | 1 |
| 3 | A Causal Inference Framework for Data Rich Environments | 0.405 | 1 | 1 |
| 4 | Network Synthetic Interventions: A Causal Framework for Panel Data Under Network Interference | 0.000 | 3 | 1 |