EconBase
← All papers

The role of the geometric mean in case-control studies

Amanda Coston, Edward H. Kennedy

arXiv 19 Jul 2022 · Statistics — Methodology · 2 citations (OpenAlex)

arXiv:2207.09016 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Historically used in settings where the outcome is rare or data collection is expensive, outcome-dependent sampling is relevant to many modern settings where data is readily available for a biased sample of the target population, such as public administrative data. Under outcome-dependent sampling, common effect measures such as the average risk difference and the average risk ratio are not identified, but the conditional odds ratio is. Aggregation of the conditional odds ratio is challenging since summary measures are generally not identified. Furthermore, the marginal odds ratio can be larger (or smaller) than all conditional odds ratios. This so-called non-collapsibility of the odds ratio is avoidable if we use an alternative aggregation to the standard arithmetic mean. We provide a new definition of collapsibility that makes this choice of aggregation method explicit, and we demonstrate that the odds ratio is collapsible under geometric aggregation. We describe how to partially identify, estimate, and do inference on the geometric odds ratio under outcome-dependent sampling. Our proposed estimator is based on the efficient influence function and therefore has doubly robust-style properties.

Citation extraction

45
references
62
in-text mentions
45
distinct cited
4
self-citations
7,089
main-text words

appendix boundary found by appendix_titled_section at “Appendix” · 51% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Jerome Cornfield et al (1951) A method of estimating comparative rates from clinical data; applications to cancer of the lung, breast, and cervix0.73732100%
2Edward H Kennedy, Sivaraman Balakrishnan, and Larry Wasserman (2021) Semiparametric counterfactual density estimation self0.6445240%
3Victor Chernozhukov, Mert Demirer, Esther Duflo, and Ivan Fernandez-… (2018) Generic machine learning inference on heterogenous treatment effects in randomized experiments0.5112250%
4Jinyong Hahn (1998) On the role of the propensity score in efficient semiparametric estimation of average treatment effects0.5112250%
5James Robins, Lingling Li, Eric Tchetgen, Aad van der Vaart, et al (2008) Higher order influence functions and minimax estimation of nonlinear functionals0.5112250%
6Mark J Van der Laan, MJ Laan, and James M Robins (2003) Unified methods for censored longitudinal data and causality0.5112250%
7Wenjing Zheng and Mark J van der Laan (2010) Asymptotic theory for cross-validated targeted maximum likelihood estimation0.5112250%
8Peter J Bickel, Chris AJ Klaassen, Peter J Bickel, Ya’acov Ritov, J… (1993) Efficient and adaptive estimation for semiparametric models, volume 40.51121100%
9Sander Greenland, Judea Pearl, and James M Robins (1999) Confounding and collapsibility in causal inference0.51121100%
10Sung Jae Jun and Sokbae Lee (2020) Causal inference under outcome-based sampling with monotonicity assumptions0.51121100%

Showing the top 10 of 45 scored citations.