Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
114,051 characters · 18 sections · 90 citation commands
Logs with zeros? Some problems and solutions
\Copy{firstp}{ When the outcome of interest $Y$ is strictly positive, researchers often estimate an average treatment effect (ATE) in logs of the form $E_P[ \log(Y(1)) - \log(Y(0)) ]$, which has the appealing feature that its units approximate percentage changes in the outcome. \oldFootnote{That is, $\log(Y(1)/Y(0)) \approx \frac{Y(1) - Y(0) }{Y(0)}$ when $Y(1)/Y(0) \approx 1$.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi A practical challenge in many economic settings, however, is that the outcome may sometimes equal zero, and thus the ATE in logs is not well-defined. When this is the case, it is common for researchers to estimate ATEs for alternative transformations of the outcome such as $\log(1+Y)$ or $\operatorname{arcsinh}(Y) = \log\pr{\sqrt{1+Y^2} + Y}$, which behave similarly to $\log(Y)$ for large values of $Y$ but are well-defined at zero. The treatment effects for these alternative transformations are typically interpreted like the ATE in logs, i.e. as (approximate) average percentage effects. For example, among the 11 papers published in the American Economic Review since 2018 that interpret a treatment effect for $\operatorname{arcsinh}(Y)$, all but one interpret the result as a percentage effect or elasticity. \oldFootnote{We found 17 papers overall using $\operatorname{arcsinh}(Y)$ as an outcome variable, of which 11 interpret the units; see (ref).}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi }
The main point of this paper is that identified ATEs that are well-defined with zero-valued outcomes should not be interpreted as percentage effects, at least if one imposes the logical requirement that a percentage effect does not depend on the baseline units in which the outcome is measured (e.g. dollars, cents, or yuan).
Our first main result shows that if $m(y)$ is a function that behaves like $\log(y)$ for large values of $y$ but is defined at zero, then the ATE for $m(Y)$ will be arbitrarily sensitive to the units of $Y$. Specifically, we consider continuous, increasing functions $m(\cdot)$ that approximate $\log(y)$ for large values of $y$ in the sense that $m(y)/\log(y) \to 1$ as $y \to \infty$. The common $\log(1+y)$ and $\operatorname{arcsinh}(y)$ transformations satisfy this property. We show that if the treatment affects the extensive margin (i.e. $P(Y(1) = 0) \neq P(Y(0) =0)$), then one can obtain any magnitude for the ATE for $m(Y)$ by rescaling the outcome by some positive factor $a$. It is therefore inappropriate to interpret the ATE for $m(Y)$ as a percentage effect, since a percentage is inherently a unit-invariant quantity, while the ATE for $m(Y)$ depends arbitrarily on the units of $Y$.
The intuition for this result is that a “percentage” treatment effect is not well-defined for an individual for whom treatment increases their outcome from zero to a positive value. For example, in our application to carranza2022job in (ref), the treatment induces more people to have positive hours worked. The percentage change in hours is then not well-defined for individuals who would work positive hours under the treatment condition but zero hours under the control condition. Any average treatment effect that is well-defined with zero-valued outcomes must therefore implicitly assign a value for a change along the extensive margin. For logarithm-like transformations $m (\cdot)$, the importance of the extensive margin is determined implicitly by the units of $Y$. To see why this is the case, consider an individual who works positive hours only if they are treated, so that $Y(1)>0$ and $Y(0)=0$. Their treatment effect for the transformed outcome $m(Y)$ is $m(Y(1)) - m(0)$, which becomes larger if the units of $Y$ are re-scaled by some $a > 1$, e.g. if we convert from weekly hours worked to yearly hours worked. When the treatment has an extensive margin effect, the ATE for $m(Y)$ can thus be made large in magnitude by re-scaling $Y$ by a large factor $a$. By contrast, if we re-scale $Y$ by a small factor $a \approx 0$, such that the resulting outcomes are close to zero, then $m(Y) \approx m(0)$, and so the ATE for $m(Y)$ will be small. By varying the units of the outcome, we can thus obtain any magnitude for the ATE for $m(Y)$.
Our theoretical results also imply that if we re-scale the units of the outcome by a finite factor $a>0$, the ATE for a log-like transformation $m(Y)$ will change by approximately $\log(a)$ times the effect of the treatment on the extensive margin. This result implies that sensitivity analyses that explore how the estimated ATE for $m(Y)$ changes with finite changes in the units of $Y$---or equivalently, how the ATE for $\log(c+Y)$ changes with the constant $c$---are essentially indirectly measuring the size of the treatment effect on the extensive margin.
We illustrate the practical importance of these results by systematically replicating recent papers published in the American Economic Review that estimate treatment effects for $\operatorname{arcsinh}$-transformed outcomes. In line with our theoretical results, we find that treatment effect estimates using $\operatorname{arcsinh}(Y)$ are sensitive to changes in the units of the outcome, particularly when the extensive margin effect is large. In half of the papers that we replicated, multiplying the original outcome by a factor of 100 (e.g. converting from dollars to cents) changes the estimated treatment effect by more than 100% of the original estimate. We obtain similar results using $\log(1+Y)$ instead of $\operatorname{arcsinh}(Y)$.
What, then, are alternative options in settings with zero-valued outcomes? Our second main result delineates the possibilities. We show that when there are zero-valued outcomes, there is no treatment effect parameter that satisfies all three of the following properties:
This “trilemma” implies that any target parameter that is well-defined with zero-valued outcomes must necessarily jettison at least one of the three properties above. \Copy{targetparamintro}{Of course, the choice of target parameter should depend on the economic question of interest. Which of the three properties the researcher prefers to forgo will thus generally depend on their context-specific motivation for using a log-like transformation in the first place.}
To that end, (ref) highlights a menu of parameters that may be attractive depending on the researcher's core motivation. We first consider the case where the researcher is interested in obtaining a causal parameter with an intuitive “percentage” interpretation. In this case, it may be natural to consider a parameter outside of the class of individual-level averages of the form $E_P[g(Y(1),Y(0))]$. One prominent option is $\theta_{ \text{ATE}\%} = \frac{E[Y(1)-Y(0)]}{E[Y(0)]},$ the ATE in levels as a percentage of the baseline mean, which in many cases can be estimated via Poisson regression silva_log_2006,wooldridge2010econometric. The researcher might also consider alternative normalizations of the outcome that lead to intuitive units, e.g. expressing the outcome in per-capita units or converting it to a rank with respect to some reference distribution. Next, we suppose the researcher would like to capture concave preferences over the outcome; for example, the researcher might consider income gains to be more meaningful for individuals who are initially poor. In this case, it is natural to directly specify how much the researcher values a change along the extensive margin relative to the intensive margin---e.g., that a change from 0 to 1 is worth an $x$ percent change along the intensive margin. Finally, suppose the researcher is interested in separately understanding the effects of the treatment along both the intensive and extensive margins. In this case, the researcher may target separate parameters for the two margins---e.g., $E[ \log(Y(1)) - \log(Y(0)) \mid Y(1) > 0,Y(0) > 0]$, the average effect in logs for individuals with positive outcomes under both treatments, captures the intensive margin. Separate effects for the two margins are not generally point-identified, but can be can be bounded using the method in lee_training_2009 or point-identified with additional assumptions zhang_evaluating_2008, zhang_likelihood-based_2009.
(ref) provides a blueprint for estimating these alternative parameters in practice by applying our recommended approaches to three recent empirical applications, including a randomized controlled trial (RCT) carranza2022job, a difference-in-differences (DiD) setting sequeira_corruption_2016, and an instrumental variables (IV) setting berkouwer2022credit.
\paragraph{Related work.} The use of log-like transformations for dealing with zero-valued outcomes has a long history. The use of the $\log(1+Y)$ transformation dates to at least williams_use_1937, while bartlett_use_1947 considers both the $\log(1+Y)$ and inverse hyperbolic sine transformations. \oldFootnote{bartlett_use_1947 proposes using $\operatorname{arcsinh}(\sqrt{Y})$.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi \Copy{bw}{More recent papers by burbidge_alternative_1988 and bellemare_elasticities_2020, among others, provide results for $\operatorname{arcsinh}(Y)$ that are frequently cited in economics papers using this transformation.} \oldFootnote{mackinnon_transforming_1990 propose transformations of the form $\operatorname{arcsinh}(y \zeta)/\zeta$, where $\zeta$ is estimated by assuming $\operatorname{arcsinh} (y\zeta)/\zeta$ is normally distributed conditional on covariates.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi
Previous work has illustrated in simulations or selected empirical applications that results for particular transformations such as $\log(1+Y)$ or $\operatorname{arcsinh}(Y)$ may be sensitive to the units of the outcome aihounton_units_2021, de_brauw_income_2021. In concurrent work, mullahy_why_2023 show theoretically that the marginal effects from linear regressions using $\log(1+Y)$ or $\operatorname{arcsinh}(Y)$ are sensitive to the scaling of the outcome, with the the limits of the marginal effects approaching those of either a levels regression or a (normalized) linear probability model, depending on whether the units are made small or large. We complement this work by proving that scale-dependence is a necessary feature of any identified ATE that is well-defined with zero-valued outcomes, and that the dependence on units is arbitrarily bad for transformations that approximate $\log(Y)$ for large values of $Y$. Thus, it is not possible to fix the issues with $\log(1+Y)$ or $\operatorname{arcsinh}(Y)$ by choosing a “better” transformation or using a different estimator. We also complement previous empirical examples by providing a systematic analysis of the sensitivity to scaling for papers in the $\emph{American Economic Review}$ using $\operatorname{arcsinh}(Y)$.
Other work has considered the interpretation of regressions using $\operatorname{arcsinh}(Y)$ or $\log(1+Y)$ from the perspective of structural equations models, as opposed to the potential outcomes model considered here. This literature has reached diverging conclusions: For example, bellemare_elasticities_2020 conclude that coefficients from $\operatorname{arcsinh}(Y)$ regressions have an interpretation as a semi-elasticity, while cohn_count_2022 conclude that these estimators are inconsistent and advocate for Poisson regression instead. thakral2023estimates show that the semi-elasticities implied by OLS regressions using $\operatorname{arcsinh}(Y)$ or $\log(1+Y)$ are sensitive to scale; they recommend instead the use of power functions $Y^k$, which they show are the only transformations (besides $\log$) for which the implied semi-elasticities for OLS regressions are scale-invariant. In (ref), we show that these diverging conclusions stem from the fact that the structural equations considered in these papers implicitly impose different restrictions on the potential outcomes---some of which are incompatible with zero-valued outcomes---and consider different target causal parameters. This highlights the value of a potential outcomes framework such as ours, which makes transparent what causal parameters are identifiable and what properties they can have.
Finally, there is a long history in econometrics of explicitly modeling the intensive and extensive margins in settings with zero-valued outcomes, such as tobin_estimation_1958 and heckman_sample_1979. Broadly speaking, these methods impose parametric structure on the joint distribution of the potential outcomes, which allows one to separate out the intensive and extensive margin effects of a treatment (see (ref) for technical details). Of course, the parametric restrictions underlying these approaches may often be difficult to justify in practice, which perhaps has contributed to the growth in the use of log-like transformations in place of approaches that explicitly model the extensive margin. Our paper shows that the presence of an extensive margin should not simply be ignored by taking a log-like transformation. It also clarifies what parameters can be learned in such cases without imposing restrictions on the joint distribution of the potential outcomes.
Let $D \in \{0,1\}$ be a binary treatment and let $Y \in [0,\infty)$ be a weakly positively-valued outcome. \oldFootnote{The $\operatorname{arcsinh}$ transformation is sometimes used in settings where $Y$ can be negative. We impose that $Y \in [0,\infty)$, and thus do not consider this case. See (ref) for extensions of our results to settings with continuous treatments.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi We assume that $Y = D Y(1) + (1-D)Y(0)$, where $Y(1)$ and $Y(0)$ are respectively the potential outcomes under treatment and control. We suppose that in some (sub-)population of interest, $(Y(1),Y(0)) \sim P$ for some (unknown) joint distribution $P$. We denote the marginal distribution of $Y(d)$ under $P$ by $P_{Y(d)}$ for $d=0,1$. We assume that neither $P_{Y(0)}$ nor $P_{Y(1)}$ is a degenerate distribution at zero.
We first consider average treatment effects of the form $\theta = E_P[ m(Y(1)) - m(Y(0)) ]$ for an increasing function $m$. We note that $\theta$ corresponds to the ATE among the (sub-)population indexed by $P$; if $P$ refers to the sub-population of compliers for an instrument, for instance, then $\theta$ is the local average treatment effect (LATE), rather than the ATE in the full population. We are interested in how $\theta$ changes if we change the units of $Y$ by a factor of $a$. That is, how does \[\theta(a) = E_P[ m(aY(1)) - m(aY(0)) ]\] depend on $a$? Setting $a=100$, for example, might correspond with a change in units between dollars and cents. Of course, if $Y$ is strictly positive and $m(y) = \log(y)$, then $\theta(a)$ is the ATE in logs and does not depend on the value of $a$.
We consider “log-like” functions $m(y)$ that are well-defined at zero but behave like $\log(y)$ for large values of $y$, in the sense that $m(y)/ \log(y) \to 1$ as $y \to \infty$. This property is satisfied by $\log (1+y)$ and $\operatorname{arcsinh}(y)$, for example. Our first main result shows that if the treatment affects the extensive margin, then $|\theta(a)|$ can be made to take any desired value through the appropriate choice of $a$.
(ref) casts serious doubt on the interpretation of ATEs for functions like $\log(1+Y)$ or $\operatorname{arcsinh}(Y)$ as (approximate) average percentage effects. While a percent (or log point) is entirely invariant to the units of the outcome, (ref) shows that, in sharp contrast, the ATEs for these transformations are arbitrarily dependent on units.
Loosely speaking, the result in (ref) follows from the fact that a “percentage” treatment effect is not well-defined for individuals who have $Y (0) = 0$ but $Y(1) > 0$. \oldFootnote{See delius_cash_2020 for an intuitive discussion of this difficulty in the context of the $\operatorname{arcsinh}(\cdot)$ transformation. They write, “the concept of elasticity itself does not make sense with zeros” (p. 21).}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi Any ATE that is well-defined with zero-valued outcomes must implicitly determine how much weight to place on changes along the extensive margin relative to proportional changes along the intensive margin.
When $m(Y)$ behaves like $\log(Y)$ for large values of $Y$, the importance of the extensive margin is implicitly determined by the units of $Y$. For intuition, suppose that we re-scale the outcomes so that the non-zero values of $Y$ are very large. Then for an individual for whom treatment changes the outcome from zero to non-zero, the treatment effect will be very large, since $m(Y(1)) \gg m(Y(0)) = m(0)$. Extensive margin treatment effects thus have a large impact on the ATE when the values of $Y$ are made large. By contrast, changing the units of $Y$ does not change the importance of treatment effects along the intensive margin by much, since for $Y(1)> 0$ and $Y(0) >0$, we have that $m(Y (1)) - m(Y(0)) \approx \log(Y(1)/Y(0))$, which does not depend on the units of the outcome.
To see the roles of the extensive and intensive margins more formally, for simplicity consider the case where $P(Y(1) = 0 , Y(0) > 0 ) = 0$, so that, for example, everyone who has positive income without receiving a training also has positive income when receiving the training. \oldFootnote{A similar argument goes through without this restriction, but then there are two extensive margins, one for individuals with $Y(1)>0=Y(0)$, and the other for those with $Y(0)>Y(1)=0$.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi Then, by the law of iterated expectations, we can write
When $a$ is large, $m(ay) \approx \log(ay)$ for non-zero values of $y$, and thus the intensive margin effect in the previous display is approximately equal to $E_P[ \log(Y(1)) - \log(Y(0)) \mid Y(1)>0,Y(0)>0]$, the treatment effect in logs for individuals with positive outcomes under both treatment and control. This, of course, does not depend on the scaling of the outcome. However, the extensive margin effect grows with $a$, since $m(aY(1)) \approx \log(a) + \log(Y(1))$ is increasing in $a$ while $m(0)$ does not depend on $a$. Thus, as $a$ grows large, the ATE for $m(aY)$ places more and more weight on the extensive margin effect of the treatment relative to the intensive margin. We can therefore make $|\theta(a)|$ arbitrarily large by sending $a \to \infty$. By contrast, if $a \approx 0$, then $m(aY(d)) \approx 0$ with very high probability, and thus the ATE for $m(aY)$ is approximately equal to 0.
\Copy{nuisancezeros1}{It is worth emphasizing that the arbitrary scale-dependence described in (ref) exists whenever the treatment affects the probability that the outcome is zero, regardless of whether the extensive margin is of direct economic interest or not. \oldFootnote{Without an extensive margin, ATEs for transformations $m(\cdot)$ defined at zero still exhibit scale-dependence, though perhaps not arbitrarily so. See (ref) below for further discussion.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi In some settings, the presence of zeros may correspond to a discrete economic choice (e.g. not participating in the labor market), and thus may be of direct interest. In other settings---for example, if the outcome is a yearly count of publications which is sometimes zero for idiosyncratic reasons---the extensive margin may be a “nuisance” rather than a direct economic object of interest. \oldFootnote{One setting where nuisance zeros may arise is when the observed outcome $Y$ is actually a mis-measured version of the true economic object of interest. For example, publications $Y$ may be a noisy measure of true researcher productivity $Y^* > 0$. One possible remedy in this setting is to model the measurement error to recover the treatment effect on $Y^*$ rather than on $Y$. In a similar vein, gandhi_estimating_2023 models the measurement error in product shares in demand estimation, which are sometimes zero in finite samples.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi The result in (ref) highlights that regardless of the source of the zeros, an ATE for a log-like transformation is not interpretable as a percentage, since the presence of the extensive margin effect makes it arbitrarily dependent on the units. Indeed, a percentage effect is not a well-defined for individuals moving from zero to non-zero outcomes. Whether the zeros correspond to a discrete economic choice or not will be relevant, however, when considering the choice of alternative target parameter, a topic we return to in (ref) below.}
We can also develop some intuition for (ref) by considering the special case where $m(y) = \log(1+y)$. In that case, we have that
Note that $$\lim_{a \to \infty} \log\left( \dfrac{1 + aY(1)}{1 + aY(0)} \right) =
$$ \noindent We thus see that the term inside the expectation in \eqref{eqn: theta for log1plus} diverges to $\infty$ for individuals with $Y(1)>0,Y(0)=0$, and likewise diverges to $-\infty$ when $Y(1)=0,Y(0)>0$. If on average the extensive margin effect is positive, then there are more individuals for whom the limit is $+\infty$ rather than $-\infty$, and thus (under appropriate regularity conditions) the ATE diverges to $\infty$. Analogously, if the extensive margin effect is negative, then the ATE diverges to $-\infty$. Hence, we see that the magnitude of the ATE for $\log(1+aY)$ diverges as $a \to \infty$ when the average effect on the extensive margin is non-zero. By contrast, as $a \to 0$, $\log(1 + aY(d)) \to \log(1) = 0$ for both $d=0$ and $d=1$, and thus the treatment effect converges to 0. \Cref{prop: can get any value for ATE} shows that this dependence on units occurs for \emph{any} log-like transformation, not just $\log(1+Y)$, and thus this issue cannot be fixed by choosing a different log-like transformation ($\log(c+Y)$, $\operatorname{arcsinh}(Y)$, $\operatorname{arcsinh}(\sqrt{Y})$, etc.)
We illustrate the results in this section by evaluating the sensitivity to scaling of estimates using the $\operatorname{arcsinh}(Y)$ transformation in recent papers in the American Economic Review (AER). In November 2022, we used Google Scholar to search for “inverse hyperbolic sine” among papers published in the AER since 2018. We searched for papers using $\operatorname{arcsinh}(Y)$ rather than $\log(1+Y)$ since the former are easier to find with a simple keyword search. Our search returned 17 papers that estimate treatment effects for an $\operatorname{arcsinh}$-transformed outcome. \oldFootnote{We consider papers with both binary and non-binary treatments, as our theoretical results extend easily to non-binary treatments; see (ref). Seven of the 10 papers we replicated used a binary treatment.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi Of these, 10 explicitly interpret the results as percentage changes or elasticities, and 6 of the remaining 7 do not directly interpret the units. See (ref) for a list of the papers and relevant quotes. Of the 17 total papers using $\operatorname{arcsinh}(Y)$, 10 had publicly available replication data that allowed us to replicate the original estimates and assess their sensitivity to scaling. \oldFootnote{We include one paper where there was a slight discrepancy between our replication of the original result and the result reported in the paper that only affected the third decimal place.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi For our replications, we focus on the first specification using $\operatorname{arcsinh}(Y)$ presented in a table in the paper, which we view as a reasonable proxy for the paper's main specification using $\operatorname{arcsinh}(Y)$. \oldFootnote{We use the first coefficient presented in a figure for one paper without any tables in the main text using $\operatorname{arcsinh}(Y)$. If the first specification is a validation check (e.g. a pre-trends test), we use the first specification of causal interest.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi
We assess the sensitivity of these results by re-running exactly the same procedure as in the original paper, except replacing $\operatorname{arcsinh}(Y)$ with $\operatorname{arcsinh}(100 \cdot Y)$. Thus, for example, if the original paper estimated a treatment effect for the $\operatorname{arcsinh}$ of an outcome measured in dollars, we use the same procedure to re-estimate the treatment effect for the $\operatorname{arcsinh}$ of the outcome measured in cents. Since (ref) shows that the sensitivity to scaling depends on the size of the extensive margin effect, we also estimate the extensive margin effect by using the same procedure as in the original paper but with the outcome $\one[Y>0]$.
The results of this exercise, shown in (ref), illustrate that treatment effect estimates can be quite sensitive to the scaling of the outcome when the extensive margin is not approximately zero. Indeed, in 5 of the 10 replicable papers, multiplying the outcome by a factor of 100 changes the estimated treatment effect by more than 100% of the original estimate. The change in the estimated treatment effect is less than 10% only in three papers, all of which have either zero or near-zero ($<$1 p.p.) effects on the extensive margin. (ref) shows that the (absolute) change in the estimated treatment effect is larger when the extensive margin effect is larger, with the change lining up very closely with the approximation given in (ref). \oldFootnote{In (ref), we plot the $t$-statistics for the treatment effects estimates as well as those for the extensive margin effect. In line with the discussion in (ref), we find that the $t$-statistics for the treatment effect using $\operatorname{arcsinh}(Y)$ tend to be similar to those for the extensive margin, except when the extensive margin is very small, and become even closer when multiplying the units by 100.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi
Using the same 10 papers, we also estimate treatment effects using $\log(1+Y)$ as the outcome, and analogously explore how the results change when we multiply the units of $Y$ by 100. (Four of the 10 papers that we replicate report an alternative specification using $\log(1+Y)$ in the paper.) The results, shown in (ref), are qualitatively quite similar those in (ref), with five of the 10 treatment effect estimates again changing by more than 100%. These results underscore the fact that (ref) applies to all log-like transformations, including both $\operatorname{arcsinh}(Y)$ and $\log(c+Y)$ for any constant $c$.
Our results so far show that ATEs for transformations that are defined at zero and approximate $\log(y)$ are arbitrarily sensitive to scaling. What other options are available when there are zero-valued outcomes? To help delineate alternative options, in this section we provide a result showing what properties a parameter defined with zero-valued outcomes can have. Specifically, we establish a “trilemma”: When there are zero-valued outcomes, there is no parameter that is (a) an average of individual-level treatment effects of the form $\theta_g = E_P[ g(Y(1),Y(0)) ]$, (b) scale-invariant, and (c) point-identified. \oldFootnote{Of course, not all parameters of the form $E_P[g(Y(1), Y(0))]$ can be interpreted as an average of individual treatment effects. For example $E[\one[Y(1)>0,Y(0)>0]]$ is the fraction of individuals whose outcomes is positive under both treatments, rather than a treatment effect. Our results apply to all parameters of this form, regardless of whether they are average treatment effects per se.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi Any approach for settings with zero-valued outcomes must therefore abandon one of the properties (a)--(c); in (ref) below we discuss several approaches that relax one (or more) of these requirements.
Before stating our formal result, we must make precise what we mean by scale-invariance and point-identification. We say that $g$ is scale-invariant if its value is the same under any re-scaling of the units of $y$ by a positive constant $a$.
\Copy{jointdist}{We next describe point-identification. We consider parameters that are identified without placing restrictions on treatment effect heterogeneity. As in fan_partial_2017, this is formalized by considering parameters that can be learned if we know the marginal distributions of $Y(1)$ and $Y(0)$, but not the full joint distribution of $(Y(1),Y(0))$.
To connect treatment effect heterogeneity to the joint distribution of potential outcomes, consider the simple case of a randomized experiment. By examining the outcome distribution for the treated group, we can learn the marginal distribution of $Y(1)$. Likewise, by examining the outcome distribution for the control group, we can learn the marginal distribution of $Y(0)$. If treatment effects were assumed to be constant, then for each observed treated unit with outcome $Y(1)$, we could infer their untreated outcome as $Y(0) = Y(1) - \tau$, where $\tau$ is the average treatment effect. Hence, the joint distribution of $(Y(1),Y(0))$ would be identified. However, if we allow for treatment effect heterogeneity, then for an observed treated unit with outcome $Y(1)$, we do not know what their value of $Y(0)$ would be, and thus we do not know the joint distribution of $(Y(1),Y(0))$. This winds up being especially important in settings with an extensive margin, since when we observe the distribution of outcomes for treated units, it means that we do not know which of the treated units would have had a zero outcome under the control condition, and thus it is difficult to disentangle the intensive and extensive margins. \oldFootnote{In (ref), we discuss a variety of structural approaches that impose assumptions restricting the joint distribution, thus allowing us to separately point-identify the effects for the two margins.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi }
With that intuition in mind, we now give a formal definition. Recall that $P$ denotes the joint distribution of $(Y(1),Y(0))$, while $P_{Y(d)}$ denotes the marginal distribution of $Y(d)$. We then say $\theta_g$ is point-identified if it depends on $P$ only through the marginals $P_{Y(1)},P_{Y(0)}$.
We will denote by $\mathcal{P}_{+}$ the set of distributions on $[0,\infty)^2$. Thus, $\theta_g$ is point-identified over $\mathcal{P}_{+}$ if it is always identified when $Y$ takes on zero or weakly positive values. Our next result formalizes that it is not possible to have a parameter of the form $E_P[g(Y(1),Y(0))]$ that is both scale-invariant and point-identified over $\mathcal{P}_+$.
Any parameter defined with zero-valued outcomes must therefore abandon one of properties (a)--(c).
As a special case, (ref) implies that the ATE for any increasing function $m(Y)$ defined at zero cannot be scale-invariant. This is because the ATE for $m(Y)$ takes the form in (a) with $g(y_1,y_0) = m(y_1) - m(y_0)$, and is also point-identified (part (c)). It follows that property (b) must be violated, i.e. there is some $c, y_0, y_1 > 0$ such that $m(cy_1) - m(cy_0) \neq m(y_1) - m(y_0)$. (ref) thus formalizes the sense in which it is not possible to “fix” the issues with ATEs for log-like transformations described above by taking alternative transformations of the outcome (e.g. $\sqrt{Y}$).
The trilemma in (ref) applies for transformations of the outcome defined at zero. To prove (ref), however, we establish an even stronger result: the only parameter satisfying properties (a) and (b) that is point-identified over distributions for which $Y$ is strictly positively-valued is the ATE in logs. \oldFootnote{More precisely, the only such treatment effect is the ATE in logs or an affine tranformations thereof.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi This result, which is formalized in (ref) in the Appendix, has some useful implications for settings in which the outcome is strictly positive.
First, it implies that the ATE for any transformation of the outcome other than $\log(Y)$ will depend on the units of the outcome for at least some DGP where the outcome is strictly positive. The scale-dependence of log-like transformations such as $\log(1+Y)$ or $\operatorname{arcsinh}(Y)$ is thus not entirely limited to settings with an extensive margin. \oldFootnote{There is thus no conflict between our results and those in thakral2023estimates, who note that semi-elasticities for OLS regressions using log-like transformations may depend on the units of the outcome even when $Y$ is strictly positively-valued.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi We note, however, that while the ATE for such transformations may depend on the units of the outcome even without zero-valued outcomes, the dependence need not be arbitrarily bad in the sense of (ref). Indeed, (ref) shows that if there is no extensive margin, the ATE for a log-like transformation will be approximately insensitive to scaling once the values of $Y$ are made large. This is intuitive, since if $Y$ is strictly positively-valued, the ATE for a log-like transformation will be approximately equal to the ATE in logs when the values of $Y$ are made large.
Second, (ref) implies that even when $Y(1)$ and $Y(0)$ are strictly-positively valued, the average proportional effect $\theta_{\text{Avg}\%} = E[(Y(1)-Y(0))/Y(0)]$ is not point-identified. This parameter is empirically relevant: For instance, andrews_optimal_2013 show that in the baily_aspects_1978--chetty_general_2006 model with heterogeneous consumption responses to unemployment, the optimal level of unemployment insurance depends on a parameter of the form $\theta_{\text{Avg}\%}$, where $Y$ is consumption and $D$ is unemployment. Although the ATE in logs may approximate $\theta_{\text{Avg}\%}$ when the proportional effect of the treatment is approximately constant, our results imply that it is not possible to point-identify $\theta_{\text{Avg}\%}$ when allowing for arbitrarily heterogeneous proportional effects.
Our theoretical results above imply that when there are zero-valued outcomes, the researcher should not take a log-like transformation of the outcome and interpret the resulting ATE as an average percentage effect: Unlike a percentage, such an ATE depends on the units of the outcome. In this section, we highlight some other parameters that are well-defined and easily interpreted when there are zero-valued outcomes; in (ref) below, we show how these parameters can be estimated in three empirical applications. Of course, any alternative parameter must necessarily drop one of the requirements in the trilemma in (ref), but the choice of which to drop may depend on the researcher's motivation.
To inform our discussion of alternative parameters, it is therefore useful to first enumerate several reasons why empirical researchers may target treatment effects for a log-transformed outcome rather than the ATE in levels:
These three motivations suggest different ways of breaking out of the trilemma in (ref). If the goal is to achieve a percentage interpretation, then one can consider scale-invariant parameters outside of the class $E_P[g(Y(1),Y(0))]$. For instance, researchers can consider the ATE in levels expressed as a percentage of the control mean, or the ATE for a normalized parameter $\tilde{Y}$ that already has a percentage interpretation. Alternatively, if the goal is to capture concave social preferences over the outcome, then it is natural to specify how much we value the intensive margin relative to the extensive margin---thus abandoning scale-invariance. Finally, if the goal is to separately understand the intensive margin effect, the researcher can abandon point-identification (from the marginal distributions) and directly target the partially identified parameter $E\bk{\log (Y(1)) - \log(Y(0)) \mid Y(0) > 0, Y(1) > 0}$, the effect in logs for individuals with positive outcomes under both treatments. We address each of these cases in turn below, with a summary of some possible parameters in (ref).
\Copy{identificationdid}{
}
We first consider the case where the researcher's primary goal is to obtain a treatment effect parameter with easily interpretable units, such as percentages.
\paragraph{Normalizing the ATE in levels.} One possibility is to target the parameter \[\thetaPoisson = \frac{E[ Y(1) - Y(0) ]}{E[Y(0)]},\] which is the ATE in levels expressed as a percentage of the control mean. For example, if a researcher is studying a program $D$ meant to reduce healthcare spending $Y$, then $\thetaPoisson$ is the percentage reduction in costs from implementing the program. This parameter is point-identified and scale-invariant, and thus has an intuitive percentage interpretation. Importantly, however, $\thetaPoisson$ is the percentage change in the average outcome between treatment and control, but is not an average of individual-level percentage changes. \oldFootnote{This is roughly analogous to how quantile treatment effects show changes in the quantiles of the potential outcomes distributions, but not the quantiles of the treatment effects (without further assumptions).}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi That is, $\thetaPoisson$ does not take the form $E_P[g(Y(1),Y(0))]$, thus avoiding the trilemma in (ref).
We note that $\thetaPoisson$ is consistently estimable by Poisson regression (see gourieroux1984pseudo; silva_log_2006; wooldridge2010econometric) under an appropriate identifying assumption. With a randomly assigned $D$, for example, estimation of $Y = \exp(\alpha + \beta D) U$ by Poisson quasi-maximum likelihood (QMLE) consistently estimates the population coefficient $\beta$, which satisfies $e^{\beta}-1 = E[Y(1)] / E[Y(0)] -1 = \thetaPoisson $. In (ref) below, we illustrate how $\theta_{\text{ATE}\%}$ can be estimated by Poisson regression in practice in several empirical examples, including both an RCT and DiD setting.
\Copy{nuisancezeros}{We also emphasize that $\thetaPoisson$ is influenced by treatment effects along both the intensive and extensive margins. In particular, the numerator of $\thetaPoisson$ is the ATE in levels. Thus, if an individual has a treatment effect of say 1, that contributes the same to $\thetaPoisson$ regardless of whether their outcome changes from 0 to 1 (an extensive margin change) or 1 to 2 (an intensive margin change). The parameter $\thetaPoisson$ may therefore be attractive in settings where the researcher does not want to distinguish between the intensive and extensive margins. For example, if $Y$ is a count of publications by a researcher in a particular year, and publications are sometimes zero owing to the idiosyncracies of the publication process, then it may be reasonable to view a change between 0 and 1 as similar to a change between 1 and 2. On the other hand, in settings where a zero corresponds to a distinct economic choice, such as not participating in the labor market, then it may be of interest to separate the effects along the intensive and extensive margin, as we discuss in more detail in (ref) below.}
\Copy{elonmusk}{It is also worth noting that if the researcher has determined that the ATE in levels is not of economic interest, then similar issues will likely arise for $\thetaPoisson$, since $\thetaPoisson$ is just a re-scaling of the ATE in levels. For one, the ATE in levels (and hence $\thetaPoisson$) imposes no diminishing returns, and thus might be dominated by individuals in the tail of the outcome distribution, particularly when the outcome is skewed. Whether this is warranted will depend on the economic question: if the policy-maker's goal is to reduce healthcare spending, it may not matter whether the savings are produced mainly by reducing spending for a small fraction of individuals with catastrophic medical spending. On the other hand, a policy that increases every American's income by \$100 and one that increases Elon Musk's income by \$35 billion and has no effect on anyone else would have approximately the same value of $\thetaPoisson$, yet the former may be vastly preferred by an inequality-minded policy-maker.} We therefore next turn to alternative approaches that place less weight on the tails of the outcome distribution.
\paragraph{Normalizing other functionals.} While $\thetaPoisson$ normalizes the ATE by the control mean, one can obtain scale-invariance by normalizing other functionals of the potential outcomes distributions. \oldFootnote{Indeed, any functional $\phi(P)$ is homogeneous of degree zero if and only if it can be written as the ratio of two homogeneous of degree one functionals.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi For example, $$ \theta_{\text{Median}\%} = \dfrac{\mathrm{Median}(Y(1)) - \mathrm{Median}(Y(0))}{\mathrm{Median}(Y(0))},$$ is the quantile treatment effect at the median normalized by the median of $Y(0)$. \oldFootnote{Note that $\theta_{\text{Median}\%}$ is well-defined only if $\mathrm{Median}(Y(0))>0$.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi Put otherwise, it captures the percentage change in the median between the treated and control distributions. ($\theta_{\text{Median}\%}$ thus may be particularly relevant for politicians interested in maximizing the happiness of the median voter!) As is typically the case with quantile treatment effects, however, the numerator of $\theta_{\text{Median}\%}$ need not correspond to the median of individual-level treatment effects. Moreover, in many settings, decision-makers may care about treatment effects throughout the distribution, not just at the median, in which case $\theta_{\text{Median}\%}$ may not be the most economically-relevant parameter.
\paragraph{Normalizing the outcome.} A related approach to obtaining a treatment effect with more intuitive units is to estimate the ATE for a transformed outcome that has a percentage interpretation. One example is to consider an outcome of the form $\tilde{Y} = Y/X$, where $Y$ is the original outcome and $X$ is some pre-determined characteristic. For example, suppose $Y$ is employment in a particular area. The treatment effect in levels for $Y$ may be difficult to interpret, since a change in employment of 1,000 means something very different in New York City versus a small rural town. However, if $X$ is the area's population, then $\tilde{Y}$ is the employment-to-population ratio, which may be more comparable across places, and is already in percentage (i.e. per capita) units. We note that the ATE for $\tilde{Y}$ is a scale-invariant, point-identified parameter of the form $\theta = E_P[g(Y(1),Y(0),X)]$, and thus escapes the trilemma in (ref) by avoiding property (a). \oldFootnote{It is scale-invariant in the sense that $g(y_1,y_0,x) = g(ay_1,ay_0,a x)$.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi The viability of this approach, of course, depends on having a variable $X$ such that the normalized outcome $\tilde{Y}$ is of economic interest. We suspect that in many contexts, reasonable options will be available, including pre-treatment observations of the outcome (assuming these are positive), or the predicted control outcome given some observable characteristics (i.e., $X = E[Y(0) \mid W]$, for observable characteristics $W$).
A second example is to use $\tilde{Y} = F_{Y^*}(Y)$, where $F_{Y^*}$ is the cumulative distribution function (CDF) of some reference random variable $Y^*$, as suggested in delius_cash_2020. The transformed outcome $\tilde{Y}$ then corresponds to the rank (i.e. percentile) of an individual in the reference distribution, and the ATE for $\tilde{Y}$ can be interpreted as the average change in rank caused by the treatment. The ATE for $\tilde{Y}$ is unit-invariant so long as $Y$ and $Y^*$ and measured in the same units. Outcomes of this form have become increasingly popular in the literature on intergenerational mobility, where $\tilde{Y}$ corresponds to a child's rank in the national income distribution. This approach has been found to yield more stable estimates than approaches using $\log(c+Y)$, which chetty_where_2014 show are sensitive to the choice of $c$. \oldFootnote{Similar to the discussion in (ref), the treatment effect in ranks cannot be converted back to obtain the ATE in levels without additional assumptions.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi
Finally, the researcher might report treatment effects on transformed outcomes of the form $1[Y \geq y]$ for different values of $y$. For example, the researcher might report the impact of the treatment on the probability that an individual earns at least \$50,000, \$60,000, etc., and interpret it as the treatment effect on the probability of obtaining a “well-paying job.” \oldFootnote{The researcher could also report the implied CDF of $Y(1)$ and $Y(0)$, from which one can infer the treatment effect on outcomes of this form for all $y$.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi Such treatment effects have interpretable units as {percentage points} (i.e. changes in probabilities). We note that treatment effects for outcomes of this form combine the effect of the treatment along the intensive and extensive margin, since for example, a worker who has $Y(1) > \$50,000 > Y(0)$ could either not work under control ($Y(0)=0$) or work under control but have earnings below \$50,000.
We next consider the case where the researcher wants to capture some form of decreasing marginal utility over the outcome. For example, when $Y$ is strictly positively valued, the ATE in logs corresponds to the change in utility from implementing the treatment for a utilitarian social planner with log utility over the outcome, $U = E[\log(Y)]$. Intuitively, this social welfare function captures the fact that the planner values a percentage point of change in the outcome equally for all individuals, regardless of their initial level of the outcome.
Of course, log utility is not well-defined when there is an extensive margin: A coherent utility function defined with zero-valued outcomes must take a stand on the relative importance of the intensive versus extensive margins. Recall from (ref) that when using transformations like $\log(1+y)$ or $\operatorname{arcsinh}(y)$, the scaling of the outcome implicitly determines the weights placed on these margins.
Instead of implicitly weighting the margins via the scaling of $Y$, a more transparent approach is to explicitly take a stand on how much one values the two margins of treatment. Of course, if one knows that their utility is captured by $U = E[m(Y)]$ (for a particular unit of $Y$, say earnings in dollars), then the ATE for $m(Y)$ is appropriate. If one is unsure exactly of their utility function, then a rough calibration is to specify how much one values a change in earnings from 0 to 1 relative to a percentage change in earnings for those with non-zero earnings. If, for example, one values the extensive margin effect of moving from 0 to 1 the same as a $100x$ percent increase in earnings, then one might consider setting $m(y) = \log(y)$ for $y>0$ and $m(0) = -x$. The ATE for this transformation can be interpreted as an approximate percentage (log point) effect, where an increase from 0 to 1 is valued at $100x$ log points. \oldFootnote{Note that this transformation will generally only be sensible if the support of $Y$ excludes $(0, e^{-x})$, since otherwise the function $m(y)$ is not monotone in $y$ over the support of $Y$. It is common, however, to have a lower-bound on non-zero values of the outcome; e.g., a firm cannot have between 0 and 1 employees. In our application to sequeira_corruption_2016 below, we normalize the minimum non-zero value of $Y$ to 1 when applying this approach.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi
We emphasize that for a fixed value of $x$, this approach necessarily depends on the scaling of the outcome (thus avoiding the trilemma in (ref)). However, this may not be so concerning since the appropriate choice of $x$ also depends on the units of the outcome---e.g., saying a change from 0 to 1 is worth $100x$ percent means something very different if 1 corresponds with one dollar versus a million dollars. In other words, ATEs for transformations such as $\operatorname{arcsinh}(Y)$ may be difficult to interpret because the scaling of the outcome implicitly determines the relative importance of the intensive and extensive margins; this approach avoids that difficulty by explicitly taking a stand on the tradeoff between these two margins. Nevertheless, a challenge with this approach is that researchers may have differing opinions over the appropriate choice of $x$ (or more generally, over the appropriate utility function).
Finally, we consider the case where the researcher is interested in understanding the intensive and extensive margin effects separately. A common question in the literature on job training programs card_active_2010, for instance, is whether a program raises participants' earnings by helping them find a job---which would be expected only to have an extensive-margin effect---or by increasing human capital, which would be expected to also affect the intensive margin. In such settings, it is natural to target separate parameters for the intensive and extensive margins.
For example, the parameter $$\theta_{\mathrm{Intensive}} = E[ \log(Y(1)) - \log(Y(0)) \mid Y(1) >0, Y(0) >0]$$ captures the ATE in logs for those who would have a positive outcome regardless of their treatment status. The parameter $\theta_{\mathrm{Intensive}}$ is scale-invariant but is not point-identified from the marginal distributions of the potential outcomes (thus avoiding the trilemma in (ref)), and therefore cannot be consistently estimated without further assumptions. \oldFootnote{$\theta_{\mathrm{Intensive}}$ also does not take the form $E_P[g(Y(1),Y(0))]$, although it can be written as $$\dfrac{E_P\bk{ \one[Y(1)>0,Y(0)>0] \log(Y(1)/Y(0)) } }{ E_P[ \one[Y(1)>0,Y(0)>0]]}, $$ where both the numerator and denominator take this form.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi However, lee_training_2009 popularized a method for obtaining bounds on $\theta_{\mathrm{Intensive}}$ under the monotonicity assumption that, for example, everyone with positive earnings without receiving a training would also have positive earnings when receiving the training. \oldFootnote{See, also, zhang_estimation_2003 for related results, including bounds without the monotonicity assumption.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi Bounds on $\theta_{\mathrm{Intensive}}$ can be reported alongside measures of the extensive margin effect, such as the change in the probability of having a non-zero outcome, $P(Y(1)>0) - P(Y(0)>0)$. One can also potentially tighten the bounds (or restore point-identification) by imposing additional assumptions on the joint distribution of the potential outcomes---we provide an example of this in our application to carranza2022job below; see zhang_evaluating_2008, zhang_likelihood-based_2009 for related approaches. \oldFootnote{We note that the lee_training_2009 bounds will tend to be tight when the extensive margin effect is close to zero. As noted in (ref), this is precisely the setting where ATEs for log-like transformations are relatively insensitive to finite changes in scale.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi
We note that the parameter $\theta_{\mathrm{Intensive}}$ is generally distinct from the “intensive margin” marginal effects implied by two-part models (2PMs), which were recommended for scenarios with zero-valued outcomes by mullahy_why_2023, among others. In (ref), we consider the causal interpretation of the marginal effects of 2PMs, building on the discussion in angrist_estimation_2001. Our decomposition shows that the marginal effects from 2PMs yield the sum of a causal parameter similar to $\theta_{\mathrm{Intensive}}$ as well as a “selection term” comparing potential outcomes for individuals for whom treatment only has an intensive margin effect to those with an extensive margin effect. It thus will generally be difficult to ascribe a causal interpretation to the marginal effects of 2PMs without assumptions about this selection.
In this section, we focus on three concrete empirical applications to illustrate how the alternative parameters described in (ref) can be estimated in practice. To illustrate a range of possible applications, we consider a randomized controlled trial, a difference-in-differences design, and an instrumental variables design.
carranza2022job conduct a randomized controlled trial (RCT) in South Africa. Individuals randomized to the treatment group are provided with certified test results that they can show to prospective employers to vouch for their skills. Individuals in the control group do not receive test results. \oldFootnote{Some individuals are also assigned to a “placebo” arm in which they are provided the test results but the form does not include the individual's name, and thus cannot credibly be shared with employers. We focus on the effect of the main treatment relative to the pure control group.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi They then investigate how this treatment impacts labor market outcomes such as employment, hours worked, and earnings. We focus here on the effects on hours worked.
\paragraph{Original specification and sensitivity to units.} carranza2022job estimate the effect of their randomized treatment on the inverse hyperbolic sine of weekly hours worked. Formally, they estimate the OLS regression specification
where $Y_i$ is average weekly hours worked for unit $i$, $D_i$ is an indicator for whether unit $i$ was in the treatment group, and $X_i$ is a vector of controls. \oldFootnote{carranza2022job include individuals receiving the “placebo” treatment in the sample and add an indicator for receiving the placebo treatment in $X_i$. We follow the same practice, although the results are similar if units receiving the placebo treatment are dropped.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi Their estimate of the ATE ($\hat\beta_1$) is 0.201 (see column (1) in (ref)). They interpret this as a 20% change in hours: “Certification increases average weekly hours worked, coded as zero for nonemployed candidates, by 20 percent” (p. 3560).
However, the results in (ref) suggest that the estimate of $\beta_1$ should not be interpreted as a percentage effect, since it depends on the units of the outcome. To illustrate this, in columns (2) and (3) we re-estimate specification (ref) with $Y_i$ redefined to be (a) yearly hours worked, i.e. weekly hours times 52, or (b) the number of full-time equivalents (FTE) worked, i.e. weekly hours divided by 40. The results change quite substantially depending on the units used, with an estimate of 0.417 using yearly hours and 0.031 using FTEs. We therefore turn next to alternative approaches with a percentage interpretation in this setting.
\paragraph{Percentage changes in the average.} The average number of (weekly) hours worked was 9.84 in the treated group and 8.85 in the control group. A simple summary of the treatment effect is thus that average hours worked were 11% higher in the treated group ($9.84/8.85=1.11$). This is an estimate of the parameter $\thetaPoisson = E[Y(1)-Y(0)]/E[Y (0)]$ discussed in (ref) above. A numerically equivalent way to obtain this estimate of 11% is to use Poisson quasi-maximum likelihood estimation (Poisson QMLE) to estimate
and then calculate $\thetaPoissonHat = \exp(\hat\beta_1) - 1 = 0.11$ (see column (1) in (ref)). \oldFootnote{This estimation is done in the sample of treated units and control units, discarding the placebo group. One could equivalently retain the units in the placebo group and add an indicator for the placebo group to (ref).}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi This formulation in terms of Poisson QMLE is useful since it allows us to include covariates to potentially increase precision. Column (2) of (ref) shows the estimate of $\thetaPoissonHat$ from estimating
by Poisson QMLE, with smaller standard errors than in column (1) (0.069 vs. 0.081).
\paragraph{Separate estimates for the extensive/intensive margins.} As shown in (ref), the treatment in carranza2022job has an estimated extensive margin treatment effect of 0.055, meaning that it increases the fraction of people with positive hours worked by 5.5 percentage points. We may be interested in whether the overall 11% increase in hours worked is driven entirely by the extensive margin, or whether there is an intensive margin effect. That is, does the treatment increase hours only by bringing people into the labor force, or does it also allow people who would have worked anyway to find jobs with more hours (e.g. full-time instead of part-time)? To this end, we can use the method of lee_training_2009 to compute bounds for the effect of the treatment for “always-takers” who would have positive hours worked regardless of treatment ($Y(1)>0,Y(0)>0$). \oldFootnote{We again exclude units receiving the “placebo treatment.”}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi The Lee bounds approach requires the monotonicity assumption that anyone who would work positive hours without the treatment would also work positive hours when treated (i.e., $P(Y(1) = 0, Y(0)>0) = 0$). This seems reasonable if workers only share the information provided by the treatment when it helps their job prospects. It could be violated, however, if workers mistakenly share their test score results when in fact employers view them negatively.
Column 1 of (ref) reports bounds of $[-0.20,0.28]$ for the effect of the treatment on log hours worked by the always-takers, while Column 2 shows bounds of $[-6.67,2.77]$ for weekly hours (in levels). Unfortunately, in this setting the Lee bounds are fairly wide, including both a zero intensive-margin effect as well as fairly large intensive-margin effects (up to 28 log points). Thus, without further assumptions, the data is not particularly informative about the size of the intensive margin.
We can, however, say more if we are willing to impose some assumptions about how the always-takers, who would work regardless of treatment status ($Y(1) > 0, Y(0) > 0$), compare to the compliers ($Y(1) > 0, Y(0) = 0$), who only work positive hours when receiving the treatment. We might reasonably expect that the compliers are negatively selected relative to the always-takers and thus would work fewer hours when receiving treatment. We can formalize this by imposing that $E[Y(1) \mid \text{Complier}] = (1-c) E[Y(1) \mid \text{Always-taker}]$, i.e. that average hours worked for compliers under treatment is $100 c$% lower than for always takers. Columns 3 through 5 of (ref) report estimates of the average effect on the always-takers, assuming $c = 0, 0.25$, and $0.5$, respectively. \oldFootnote{Under the assumptions in lee_training_2009, $E[Y(1) \mid Y(1)>0] = \theta E[Y(1) \mid \text{Always-taker}] + (1-\theta) E[Y(1) \mid \text{Complier}]$, where $\theta = P(Y(0)>0)/P(Y(1)>0)$. Plugging in $E[Y(1) \mid \text{Complier}] = (1-c) E[Y(1) \mid \text{Always-taker}]$, it follows that $E[Y(1) \mid \text{Always-taker}] = 1/(\theta + (1-c)(1-\theta)) E[Y(1) \mid Y(1)>0]$. Further, $E[Y(0) \mid \text{Always-taker}] = E[Y(0) \mid Y(0)>0]$. Our estimation plugs in sample analogs to these expressions to estimate $E[Y(1)-Y(0) \mid \text{Always-taker}]$.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi If we assume that always-takers and compliers work an equal number of hours under treatment ($c=0$), then our point estimates suggest that there is actually a negative intensive-margin effect for the always-takers ($-1.02$ weekly hours). Under the assumption that compliers work 25% fewer hours ($c=0.75$), the estimated effect for always-takers is near zero ($-0.07$ weekly hours), consistent with no important intensive margin. Finally, if we assume compliers work half as many hours as the always-takers ($c=0.5$), then our estimates suggest a positive intensive margin effect ($0.95$ weekly hours). Our assessment of the importance of the intensive margin thus depends on how negatively-selected we think compliers are relative to always-takers.
sequeira_corruption_2016 studies a decrease in tariffs on trade between Mozambique and South Africa which occurred in 2008. She is interested in whether the reduction in tariffs reduced bribes paid to customs officers (among other outcomes). To study this question, she utilizes a difference-in-differences design comparing the change in bribes paid for products that were affected by the tariff change to that for a comparison group of products that did not experience a change in tariffs.
\paragraph{Original specification and sensitivity to units.} sequeira_corruption_2016 has repeated cross-sectional data with information on the bribe amount $Y_{it}$ paid on shipment $i$ in year $t$. She estimates the regression specification
where $D_i$ is an indicator for whether shipment $i$ is for a product type affected by the tariff change in 2008, $\text{Post}_t$ is an indicator for whether year $t$ is after the tariff change, and $X_{it}$ is a vector of covariates related to shipment $i$ in period $t$. sequeira_corruption_2016 estimates (ref) with $Y_{it}$ measured in 2007 Mozambican Metical (MZN) and obtains $\hat\beta_{1, (\text{MZN})} = -3.7$ (SE $=1.1$). However, estimating the same specification with $Y_ {it}$ measured in thousands of U.S. dollars instead yields an estimate of $\hat\beta_{1, (\$1000)} = -0.11$ (SE $=0.070$). \oldFootnote{We use the conversion rate of $ 1 \text{ USD} = 24.48 \text{ MZN}$ as of January 1, 2007, as provided by fxtop.com.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi These results reinforce the conclusion from (ref) that treatment effects for $m(y) = \log(1+y)$ should not be interpreted as approximating a percentage effect.
In what follows, we discuss a variety of alternative approaches that may be reasonable in this context. We note that in a non-experimental setting like this, different approaches may rely on different identifying assumptions. We therefore explicitly discuss the identifying assumptions needed by each of the methods we discuss.
\paragraph{Proportional treatment effects.} One natural approach here is to target the average proportional treatment effect on the treated, $$\thetaPoissonATT = \dfrac{E[Y_{it}(1) \mid D_i =1, \text{Post}_t = 1] - E[Y_{it}(0) \mid D_i =1, \text{Post}_t = 1]}{ E[Y_{it}(0) \mid D_i =1, \text{Post}_t = 1] }.$$ This is the percentage change in the average outcome for the treated group in the post-treatment period.
Identification of $\thetaPoissonATT$ requires us to infer the counterfactual post-treatment mean outcome for the treated group, $E[Y_{it}(0) \mid D_i =1, \text{Post}_t = 1]$. Of course, one approach to obtain such identification would be to assume parallel trends in levels. However, given that the treated and control groups have different pre-treatment means (see the bottom panel of (ref)), it may be unreasonable to expect that time-varying factors (e.g. the macroeconomy) have equal level effects on the outcome. An alternative identifying assumption is to impose that, in the absence of treatment, the percentage changes in the mean would have been the same for the treated and control group. As in wooldridge_simple_2023, this can be formalized using a “ratio” version of the parallel trends assumption,
Intuitively, (ref) states that if the treatment had not occurred, the average percentage change in the mean outcome for the treated group would have been the same as the average percentage change in the mean outcome for the control group. Under (ref), we can thus estimate the counterfactual percentage change in the mean outcome for the treated group using the observed percentage change for the control group.
(ref) shows that the sample mean of the outcome for the treated group decreased by 75% between the pre-treatment and post-treatment periods (from 4,742 to 1,172 (MZN)). Under the ratio parallel trends assumption (ref), this suggests that the mean outcome for the treated group would also have decreased by 75% in the absence of treatment, thus implying an estimate of $2,602$ for the counterfactual mean outcome for the treated group. The actual post-treatment mean for the treated group is 465, which is 82% below this implied counterfactual. This implies that the tariff reduction reduced the average bribe in the post-treatment period by 82%, i.e. $\thetaPoissonATTHat = -0.82$. Conveniently, this estimate can also be obtained using Poisson QMLE to estimate
and then computing $\thetaPoissonATTHat = \exp(\hat\beta_1) - 1= -0.82$, as shown in column (1) of (ref).
We can also re-incorporate the covariates $X_{it}$ by estimating
which yields an estimate of $\thetaPoissonATT$ of $-0.72$, as shown in the second column of (ref). As formalized in wooldridge_simple_2023, this estimate will be a consistent estimate of $\thetaPoissonATT$ if (ref) holds conditional on $X_{it}$, and the conditional expectation of $Y_{it}$ takes the functional form implied by (ref) (assuming $\epsilon_{it}$ has mean 1 conditional on the covariates). The approach with covariates thus suggests that the tariff change reduced the average bribe for treated products by 72% in the post-treatment period.
sequeira_corruption_2016's data only contains information on one year prior to treatment (2007), and so in this context it is not possible to evaluate the plausibility of (ref) using periods prior to the policy change of interest. If multiple pre-treatment periods were available, however, one could estimate a Poisson QMLE event-study of the form
where $\text{RelativeTime}_t = t-2008$ is the time relative to the treatment date. The event-study coefficients $\beta^{ES}_r$ for $r<0$ are analogous to “pre-trends” coefficients in typical difference-in-differences event-studies, and are informative about whether the pre-treatment analogue to (ref) holds. \oldFootnote{More precisely, the exponentiated coefficients $\exp(\hat\beta_{r})-1$ correspond to the implied “placebo” proportional treatment effects for periods before treatment. We recommend plotting the exponentiated coefficients in event-studies, although we note that $\exp(\beta) - 1 \approx \beta$ for $\beta \approx 0$. As with typical tests for pre-trends, one should be cautious that a failure to reject the null that the pre-treatment coefficients equal zero does not necessarily imply that the identifying assumption is satisfied kahn-lang_promise_2020, roth_pretest_2022. One can (partially) address these issues by applying sensitivity analysis tools for event-studies rambachan_more_2023 to estimates of (ref) to further gauge the robustness of the findings to violations of the identifying assumptions. We also refer the reader to wooldridge_simple_2023 for extensions of the Poisson regression approach to settings with staggered treatment timing.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi
\paragraph{Log effects with calibrated extensive margin value.} The analysis above presented estimates of $\thetaPoissonATT$, the proportional change in the average bribe caused by the treatment. It is well-known that averages can be heavily influenced by observations in the tail, especially when the outcome has a skewed distribution, as is the case here (see (ref)). One might argue that a world in which most products receive medium-sized bribes is more corrupt than one in which a very small fraction of products receive large bribes---even if they both produce the same average bribe amount. This motivates studying the treatment effect on a concave transformation of the outcome that is less heavily influenced by outcomes in the tail of the distribution. As an illustration of this, we first normalize the outcome so that $1$ corresponds to the value of the minimum non-zero bribe in the data (that is, we divide by $y_{\min} = \min_{Y_{it}>0} Y_{it} = 15.68 \text{ MZN}$). We then estimate the treatment effect for the transformed outcome $m(Y)$, where $m(y) = \log(y)$ for $y>0$ and $m(0) = -x$ for some choice of $x$, as described in (ref). If $x$ is set to 0, then this estimates the treatment effect in logs where all zero bribes are set to equal the smallest positive bribe in the data; this specification thus “shuts off” the extensive margin change between 0 and $y_{\min}$. If instead $x$ is set to $0.1$, for example, then a change between $0$ and $y_{\min}$ is valued as the equivalent of a 10 log point change along the intensive margin.
We estimate the treatment effect for these transformations using the analogue to (ref) that replaces $\log(1+Y_{it})$ with $m(Y_{it})$ on the left-hand side. \oldFootnote{As usual, identification of the treatment effect for $m(Y)$ using difference-in-differences requires parallel trends for $m(Y(0))$. The identifying assumption thus varies depending on the choice of $x$. The results in roth_when_2023 imply that parallel trends will hold for all values of $x$ when a parallel trends assumption is satisfied for the distribution of $Y(0)$. If more pre-treatment periods were available, these identifying assumptions could be partially evaluated using pre-trends tests. See (ref) for additional discussion of identification.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi The results for $x \in \{0,0.1,1,3\}$ are shown in (ref). Column (1) shows an effect of 249 log points ($\hat\beta_1 = -2.49$) when we treat zero bribes as if they were equal to $y_{\min}$ (i.e. setting $x=0$). The estimated treatment effect grows in magnitude as we place more value on the extensive margin by increasing $x$. Interestingly, the original estimate in sequeira_corruption_2016 of $-3.748$ using $\log(1+Y)$ is similar to what we obtain when we value a change from 0 to $y_{\min}$ at 300 log points ($x=3$). The original specification can thus be viewed as placing a rather large weight on the extensive margin.
berkouwer2022credit conduct an RCT in Nairobi in which they randomize the price for energy-efficient stoves. They use the randomized price ($p_i$) as an instrument for whether an individual $i$ buys an energy-efficient stove ($D_i$). They use this instrument to estimate the effects of stove-adoption on outcomes such as charcoal spending ($Y_i$).
\paragraph{Original specification and sensitivity to scale.} Let $X_i$ be a vector of control variables (including a constant). berkouwer2022credit estimate
by two-stage least squares (TSLS), using $p_i$ as an instrument for $D_i$. \oldFootnote{More precisely, each observation $i$ is an individual-by-week pair, and some (but not all) individuals are surveyed on multiple weeks. Standard errors are clustered at the respondent level.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi (They also report results where spending is measured in levels.) The estimated coefficient $\hat\beta$ is an estimate of the LATE of stove adoption on the $\operatorname{arcsinh}$ of charcoal spending for instrument-compliers whose decision of whether to purchase the stove depends on the price offered in the experiment. \oldFootnote{We use the phrase “instrument-compliers” to distinguish compliers for the instrument, whose value of $D(z)$ depends on $z$, from “compliers” discussed earlier who have $Y(1)>0,Y(0)=0$. Since the instrument takes on multiple values (i.e. multiple price offers), $\beta$ corresponds to a weighted average of treatment effects across instrument-compliers for different values of the instrument angrist_interpretation_2000.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi In berkouwer2022credit, $Y_i$ is measured as weekly charcoal spending in dollars. They obtain a coefficient of $\hat\beta = -0.50$ and write “[t]he 50 log point reduction corresponds to a 39 percent decrease in charcoal consumption [since $\exp(-0.50)=1-0.39$]” (p. 3306).
However, if we change the units of the outcome to annual charcoal spending in Kenyan shillings, the original currency in which charcoal spending was measured, the same specification yields an estimate of $-0.44$. Relative to our previous applications, the change in the treatment effect estimates is fairly small for these choices of units, due to a small estimated extensive margin of 0.01 (see (ref)). \oldFootnote{We note, however, that the $t$-statistic for the effect on $\operatorname{arcsinh}(Y_i)$ is rather sensitive here, changing from approximately 7 to 3 depending on the units.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi Nevertheless, the fact that the treatment effects using an $\operatorname{arcsinh}$-transformed outcome depend on the units should give us pause in interpreting them as percentages. Indeed, a percentage effect is not well-defined for someone who has non-zero spending under treatment and zero spending under the control, so an average individual-level percentage effect does not make sense if the treatment can affect whether one has any charcoal spending.
berkouwer2022credit first discuss the LATE in levels, and then immediately afterwards state that the treatment effect for the $\operatorname{arcsinh}$-transformed outcome “corresponds to a 39 percent decrease in charcoal consumption” (p. 3306). The main goal of taking the $\operatorname{arcsinh}$ transformation here thus appears to be to obtain a treatment effect with a percentage interpretation. We therefore next implement two approaches with an (approximate) percentage interpretation in this context.
\paragraph{Proportional LATE.} One natural approach in this context is to estimate the proportional change in the average outcome for instrument-compliers, i.e. to estimate $\thetaPoisson$ among the population of instrument-compliers. Put otherwise, we can express the LATE in levels as a percentage of the control mean for instrument-compliers. An estimate of the LATE in levels is naturally obtained using TSLS specification (ref) with $Y_i$ as the outcome, which yields an estimate of $-2.46$. As described in abadie_bootstrap_2002, we can likewise obtain an estimate of the control instrument-complier mean by using TSLS with $-(D_i-1) \cdot Y_i$ as the outcome, which yields an estimate of $5.86$. Putting these together, we obtain an estimate of $\thetaPoisson$ for instrument-compliers of $-2.46/5.86=-0.42$ (SE = 0.046), which suggests that average charcoal spending is 42% lower for instrument-compliers under treatment than under control. \oldFootnote{The standard error was calculated via a non-parametric bootstrap with 1,000 draws, clustered at the respondent level. We note that with a binary instrument, an estimate of $\theta_{\text{Intensive}}$ for instrument-compliers can be obtained using Poisson IV regression (e.g. the ivpoisson command in Stata); see angrist_estimation_2001. However, we are not aware of a LATE interpretation of Poisson IV regression with a multi-valued instrument, and thus do not pursue it here. Whether Poisson IV regression has such an interpretation with a continuous IV strikes us an interesting topic for future work.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi If pollution is proportional to charcoal spending, then this parameter is economically relevant as it corresponds to the percentage reduction in pollution for instrument-compliers from gaining access to the efficient stove.
\paragraph{Lee bounds.} berkouwer2022credit benchmark their treatment effect estimates relative to engineering estimates of the efficiency gains of using an efficient stove relative to a non-efficient one. For this benchmarking exercise, it seems sensible to focus on the intensive-margin effect of the treatment---i.e., the treatment effect for instrument-compliers who would use a non-efficient stove if offered a high price and an efficient one if offered a low price. To do so, we can form lee_training_2009-type bounds for the average treatment effect in logs for instrument-compliers who would have positive charcoal spending regardless of treatment status. \oldFootnote{The validity of the lee_training_2009-type bounds requires the “monotonicity” assumption that all instrument-compliers who would have some charcoal consumption when not buying an efficient stove would also have some charcoal consumption when buying an efficient stove, which seems reasonable. Note that this is a distinct assumption from the instrument monotonicity assumption needed for a LATE interpretation for instrumental variables imbens1994identification, which in this context states that anyone who would buy a stove at a higher price would also buy at a lower price.}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi
The bounds on $\theta_{\text{Intensive}}$ for instrument-compliers are $ [-0.565,-0.538]$ (with SEs for the lower and upper bounds of 0.072 and 0.075). \oldFootnote{We obtain these estimates using the procedure in abadie_bootstrap_2002, as described in detail in (ref).}\futurelet\nextToken \ifx\footnote\nextToken\textsuperscript{,}\fi This implies that for the instrument-compliers who would spend on charcoal regardless of treatment status, spending decreases by 54 to 56 log points. We note that the Lee bounds are fairly tight in this case, as tends to be the case when the extensive margin is small. It is also worth noting that in this example, the estimated treatment effects using $\operatorname{arcsinh}(Y_i)$---both in terms of weekly spending in dollars and in terms of annual spending in Kenyan shillings---fall outside of the Lee bounds, although they are fairly close to the upper bound when using weekly spending in dollars.
It is common in empirical work to estimate ATEs for transformations such as $\log(1+Y)$ or $\operatorname{arcsinh}(Y)$ which are well-defined at zero and behave like $\log(Y)$ for large values of $Y$. We show that the ATEs for such transformations should not be interpreted as percentages, since they depend arbitrarily on the units of the outcome when there is an extensive margin. Further, we show that any parameter that is an average of individual-level treatment effects of the form $E_P[g(Y(1),Y(0))]$ must be scale-dependent if it is point-identified and well-defined at zero. We discuss several alternative approaches, including estimating scale-invariant normalized parameters (e.g. via Poisson regression), explicitly calibrating the value placed on the intensive versus extensive margins, and separately estimating effects for the intensive and extensive margins (e.g. using Lee bounds). We illustrate how these approaches can be applied in practice in three empirical applications.
\nocite{azoulay2019does,beerli2021abolition,berkouwer2022credit,cabral2022demand,carranza2022job,faber2019tourism,hjort2019arrival,johnson2020regulation,mirenda2022economic,norris2021effects,ager2021intergenerational,arora2021knowledge,bastos2018export,fetzer2021security,moretti2021effect,rogall2021mobilizing,cao2022rebel}
{\singlespacing }