EconBase
← All papers

Demystifying and avoiding the OLS "weighting problem": Unmodeled heterogeneity and straightforward solutions

Tanvi Shinkre, Chad Hazlett

arXiv 5 Mar 2024 · Statistics — Methodology

arXiv:2403.03299 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Researchers frequently estimate treatment effects by regressing outcomes (Y) on treatment (D) and covariates (X). Even without unobserved confounding, the coefficient on D yields a conditional-variance-weighted average of strata-wise effects, not the average treatment effect. Scholars have proposed characterizing the severity of these weights, evaluating resulting biases, or changing investigators' target estimand to the conditional-variance-weighted effect. We aim to demystify these weights, clarifying how they arise, what they represent, and how to avoid them. Specifically, these weights reflect misspecification bias from unmodeled treatment-effect heterogeneity. Rather than diagnosing or tolerating them, we recommend avoiding the issue altogether, by relaxing the standard regression assumption of "single linearity" to one of "separate linearity" (of each potential outcome in the covariates), accommodating heterogeneity. Numerous methods--including regression imputation (g-computation), interacted regression, and mean balancing weights--satisfy this assumption. In many settings, the efficiency cost to avoiding this weighting problem altogether will be modest and worthwhile.

Citation extraction

39
references
95
in-text mentions
39
distinct cited
1
self-citations
9,593
main-text words

appendix boundary found by appendix_command · 83% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Winston Lin (2013) Agnostic notes on regression adjustments to experimental data: Reexamining Freedman’s critique1.00064100%
2Peter M. Aronow and Cyrus Samii (2016) Does Regression Produce Representative Estimates of Causal Effects?1.00063100%
3Patrick Kline (2011) Oaxaca-blinder as a reweighting estimator1.00063100%
4Ambarish Chattopadhyay and José R Zubizarreta (2023) On the implied weights of linear regression for causal inference1.00053100%
5Tymon Słoczyński (2022) Interpreting OLS Estimands When Treatment Effects Are Heterogeneous: Smaller Groups Get Larger Weights1.00053100%
6Jinyong Hahn (2023) Properties of least squares estimator in estimation of average treatment effects0.92843100%
7Guido W. Imbens and Jeffrey M. Wooldridge (2009) Recent Developments in the Econometrics of Program Evaluation0.92843100%
8Joshua D. Angrist (1998) Estimating the Labor Market Impact of Voluntary Military Service Using Social Security Data on Military Applicants0.87462100%
9Joshua D. Angrist and Jörn-Steffen Pischke (2009) Mostly Harmless Econometrics: An Empiricist's Companion0.87462100%
10William A Belson (1956) A technique for studying the effects of a television broadcast0.84333100%

Showing the top 10 of 39 scored citations.