Ruoxuan Xiong, Allison Koenecke, Michael Powell, Zhu Shen, Joshua T. Vogelstein, Susan Athey
arXiv 25 Jul 2021 · Machine Learning · publishedStatistics in Medicine (2023) · 31 citations (OpenAlex)
arXiv:2107.11732 · PDF · DOI · OpenAlex · Extracted main text
We are interested in estimating the effect of a treatment applied to individuals at multiple sites, where data is stored locally for each site. Due to privacy constraints, individual-level data cannot be shared across sites; the sites may also have heterogeneous populations and treatment assignment mechanisms. Motivated by these considerations, we develop federated methods to draw inference on the average treatment effects of combined data across sites. Our methods first compute summary statistics locally using propensity scores and then aggregate these statistics across sites to obtain point and variance estimators of average treatment effects. We show that these estimators are consistent and asymptotically normal. To achieve these asymptotic properties, we find that the aggregation schemes need to account for the heterogeneity in treatment assignments and in outcomes across sites. We demonstrate the validity of our federated methods through a comparative study of two large medical claims databases.
appendix boundary found by appendix_command · 44% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Koenecke, A., Powell, M., Xiong, R., Shen, Z., Fischer, N., Huq, S.,… (2021) Alpha-1 adrenergic receptor antagonists to prevent hyperinflammation and death from lower respiratory tract infection self | 0.956 | 8 | 3 | 88% |
| 2 | Wooldridge, J. M (2007) Inverse probability weighted estimation for general missing data problems | 0.894 | 7 | 4 | 71% |
| 3 | Han, L., Hou, J., Cho, K., Duan, R., and Cai, T (2021) Federated adaptive causal estimation (face) of target treatment effects | 0.811 | 4 | 2 | 100% |
| 4 | White, H (1982) Maximum likelihood estimation of misspecified models | 0.737 | 4 | 3 | 50% |
| 5 | Wooldridge, J. M (2002) Inverse probability weighted m-estimators for sample selection, attrition, and stratification | 0.737 | 3 | 2 | 100% |
| 6 | Duan, R., Boland, M. R., Liu, Z., Liu, Y., Chang, H. H., Xu, H., Chu… (2020) Learning from electronic health records across multiple sites: A communication-efficient and privacy-preserving distributed algo… | 0.644 | 2 | 2 | 100% |
| 7 | Duan, R., Ning, Y., and Chen, Y (2022) Heterogeneity-aware and communication-efficient distributed statistical inference | 0.644 | 2 | 2 | 100% |
| 8 | Jordan, M. I., Lee, J. D., and Yang, Y (2018) Communication-efficient distributed statistical inference | 0.644 | 2 | 2 | 100% |
| 9 | Han, L., Li, Y., Niknam, B. A., and Zubizarreta, J. R (2022) Privacy-preserving and communication-efficient causal inference for hospital quality measurement | 0.585 | 3 | 1 | 100% |
| 10 | Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C… (2017) Double/debiased/neyman machine learning of treatment effects | 0.511 | 2 | 2 | 50% |
Showing the top 10 of 60 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Difference-in-Differences with Unpoolable Data | 0.644 | 2 | 2 |
| 2 | Feature Selection for Personalized Policy Analysis | 0.405 | 1 | 1 |
| 3 | Federated Offline Policy Learning | 0.405 | 1 | 1 |
| 4 | Cross-Validated Causal Inference: a Modern Method to Combine Experimental and Observational Data | 0.405 | 1 | 1 |