EconBase
← All papers

Logs with zeros? Some problems and solutions

Jiafeng Chen, Jonathan Roth

arXiv 12 Dec 2022 · Econometrics · publishedThe Quarterly Journal of Economics (2023) · 856 citations (OpenAlex)

arXiv:2212.06080 · PDF · DOI · OpenAlex · Extracted main text

Abstract

When studying an outcome $Y$ that is weakly-positive but can equal zero (e.g. earnings), researchers frequently estimate an average treatment effect (ATE) for a "log-like" transformation that behaves like $\log(Y)$ for large $Y$ but is defined at zero (e.g. $\log(1+Y)$, $arcsinh(Y)$). We argue that ATEs for log-like transformations should not be interpreted as approximating percentage effects, since unlike a percentage, they depend on the units of the outcome. In fact, we show that if the treatment affects the extensive margin, one can obtain a treatment effect of any magnitude simply by re-scaling the units of $Y$ before taking the log-like transformation. This arbitrary unit-dependence arises because an individual-level percentage effect is not well-defined for individuals whose outcome changes from zero to non-zero when receiving treatment, and the units of the outcome implicitly determine how much weight the ATE for a log-like transformation places on the extensive margin. We further establish a trilemma: when the outcome can equal zero, there is no treatment effect parameter that is an average of individual-level treatment effects, unit-invariant, and point-identified. We discuss several alternative approaches that may be sensible in settings with an intensive and extensive margin, including (i) expressing the ATE in levels as a percentage (e.g. using Poisson regression), (ii) explicitly calibrating the value placed on the intensive and extensive margins, and (iii) estimating separate effects for the two margins (e.g. using Lee bounds). We illustrate these approaches in three empirical applications.

Citation extraction

62
references
145
in-text mentions
62
distinct cited
1
self-citations
17,264
main-text words

appendix boundary found by appendix_command · 68% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Carranza, Garlick, Orkin and Rankin (2022) Job Search and Hiring with Limited Information about Workseekers' Skills1.000144100%
2Sequeira (2016) Corruption, Trade Costs, and Gains from Tariff Liberalization: Evidence from Southern Africa1.000133100%
3Berkouwer and Dean (2022) Credit, attention, and externalities in the adoption of energy efficient technologies by low-income households0.92810480%
4Lee (2009) Training, Wages, and Sample Selection: Estimating Sharp Bounds on Treatment Effects0.92810480%
5Angrist (2001) Estimation of Limited Dependent Variable Models With Dummy Endogenous Regressors0.7374450%
6Mullahy and Norton (2023) Why Transform Y? The Pitfalls of Transformed Regressions with a Mass at Zero0.7374350%
7Thakral and Tô (2023) When Are Estimates Independent of Measurement Units?0.7374350%
8Santos Silva and Tenreyro (2006) The Log of Gravity0.7373367%
9Bellemare and Wichman (2020) Elasticities and the Inverse Hyperbolic Sine Transformation0.7218338%
10Abadie (2002) Bootstrap Tests for Distributional Treatment Effects in Instrumental Variable Models0.6443267%

Showing the top 10 of 62 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Estimation and Inference on Average Treatment Effect in Percentage Points under Heterogeneity0.87452
2When “Normalization Without Loss of Generality” Loses Generality0.87452
3Lee Bounds with a Continuous Treatment in Sample Selection0.84333
4Abadie's Kappa and Weighting Estimators of the Local Average Treatment Effect0.81142
5Causal Interpretation of Regressions With Ranks0.64422
6Handling Sparse Non-negative Data in Finance0.64422
7Dealing with Logs and Zeros in Regression Models0.58531
8Model-Agnostic Covariate-Assisted Inference on Partially Identified Causal Effects0.51121
9Causal Inference for Spatial Treatments0.40511
10Difference-in-Differences with Compositional Changes0.40511