EconBase
← All papers

A Cautionary Tale on Integrating Studies with Disparate Outcome Measures for Causal Inference

Harsh Parikh, Trang Quynh Nguyen, Elizabeth A. Stuart, Kara E. Rudolph, Caleb H. Miles

arXiv 16 May 2025 · Statistics — Methodology

arXiv:2505.11014 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Data integration approaches are increasingly used to enhance the efficiency and generalizability of studies. However, a key limitation of these methods is the assumption that outcome measures are identical across datasets -- an assumption that often does not hold in practice. Consider the following opioid use disorder (OUD) studies: the XBOT trial and the POAT study, both evaluating the effect of medications for OUD on withdrawal symptom severity (not the primary outcome of either trial). While XBOT measures withdrawal severity using the subjective opiate withdrawal scale, POAT uses the clinical opiate withdrawal scale. We analyze this realistic yet challenging setting where outcome measures differ across studies and where neither study records both types of outcomes. Our paper studies whether and when integrating studies with disparate outcome measures leads to efficiency gains. We introduce three sets of assumptions -- with varying degrees of strength -- linking both outcome measures. Our theoretical and empirical results highlight a cautionary tale: integration can improve asymptotic efficiency only under the strongest assumption linking the outcomes. However, misspecification of this assumption leads to bias. In contrast, a milder assumption may yield finite-sample efficiency gains, yet these benefits diminish as sample size increases. We illustrate these trade-offs via a case study integrating the XBOT and POAT datasets to estimate the comparative effect of two medications for opioid use disorder on withdrawal symptoms. By systematically varying the assumptions linking the SOW and COW scales, we show potential efficiency gains and the risks of bias. Our findings emphasize the need for careful assumption selection when fusing datasets with differing outcome measures, offering guidance for researchers navigating this common challenge in modern data integration.

Citation extraction

42
references
61
in-text mentions
42
distinct cited
6
self-citations
6,387
main-text words

appendix boundary found by appendix_command · 68% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Handelsman, L., Cochrane, K. J., Aronson, M. J., Ness, R., Rubinstei… (1987) Two new rating scales for opiate withdrawal0.84333100%
2Kallus, N., Puli, A. M., and Shalit, U (2018) Removing hidden confounding by experimental grounding0.73732100%
3Brantner, C. L., Chang, T.-H., Nguyen, T. Q., Hong, H., Di Stefano,… (2023) Methods for integrating trials and non-experimental data to examine treatment effect heterogeneity self0.64422100%
4Degtiar, I. and Rose, S (2023) A review of generalizability and transportability0.64422100%
5Huang, M. Y. and Parikh, H (2024) Towards generalizing inferences from trials to target populations self0.64422100%
6Lee, J. D., Nunes, E. V., Novo, P., Bachrach, K., Bailey, G. L., Bha… (2018) Comparative effectiveness of extended-release naltrexone versus buprenorphine-naloxone for opioid relapse prevention (x: Bot): a…0.64422100%
7Mitra, N., Roy, J., and Small, D (2022) The future of causal inference0.64422100%
8Parikh, H., Ross, R., Stuart, E., and Rudolph, K (2024) Who are we missing? a principled approach to characterizing the underrepresented population self0.64422100%
9Robinson, P. M (1988) Root-n-consistent semiparametric regression0.64422100%
10Rosenman, E. T., Basse, G., Owen, A. B., and Baiocchi, M (2023) Combining observational and experimental datasets using shrinkage estimators0.64422100%

Showing the top 10 of 42 scored citations.