EconBase
← Back to paper

What's Logs Got to do With it: On the Perils of log Dependent Variables and Difference-in-Differences

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

46,011 characters · 7 sections · 22 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

What's Logs Got to do With it: On the Perils of log Dependent Variables and Difference-in-Differences

abstractThe logarithmic transformation of the dependent variable is not innocuous when using a difference-in-differences (DD) research design. With a dependent variable in logs, the DD term does not capture the outcome difference between treated and untreated groups over time. Rather it reflects an approximation of the proportional difference in growth rates across groups. As I show with both simulations and two empirical examples, if the baseline outcome distributions are sufficiently different across groups, the DD parameter for a log-specification can be different in sign to that of a levels-specification. I provide a condition, based on (i) the aggregate time effect, and (ii) the difference in relative baseline outcome means, for when the sign-switch will occur.

\\ \jelcodes{C01.}

Introduction

Difference-in-differences (DD) is almost certainly the most popular quasi-experimental research design currently used in a broad range of empirical settings. Its use extends beyond Economics into Political Science, Social Medicine and other fields. The seeming simplicity of the design has, at least until recently, played a large role in its popularity. A recent literature documenting underlying issues with difference-in-differences, particularly in the case of staggered roll-out of treatment implementation GoodmanBacon2021, has somewhat shattered the illusion of the simplicity of this research design.

In this paper, I highlight an additional complication that the applied researcher faces when operationalizing a DD design -- the choice of functional form, and the subsequent consequences of this choice. As Ciani2019 note in related work, researchers may apply the log transformation, not because they believe the true underlying data generating process is multiplicative rather than additive, but rather due to concerns of skewness, or because the log transformation enables one to interpret the effect of controls in percentage terms.\footnote{Some papers are more explicit about the consequences of the choice between level- and log-specifications of the outcome variable when using difference-in-differences designs Finkelstein2007,Powell2020,Park2021. }

In order to set the scene for this work, I survey all papers published in The Quarterly Journal of Economics (QJE) from 2018 to 2022. I present summary information on these articles in Table (ref). Of the 49 QJE articles published in this five year span that use a DD design, almost all of these (46 articles) consider at least one continuous outcome variable. Just under two thirds of these 46 articles (30 articles or 65%) impose a log transformation on at least one continuous variable, and three more impose an inverse hyperbolic sine transformation\footnote{See Bellemare2020 for an in-depth consideration of the implications of the inverse hyperbolic sine (IHS) transformation. As the authors note, a key driver of the recent interest from applied researchers in the IHS transformation is that the function is similar to a logarithm, but unlike the logarithmic function, it is defined at zero.}, which will lead to related issues.\footnote{The inverse hyperbolic sine of a random variable $y$, $\operatorname{arsinh}(y) = \ln(y +\sqrt{y^2+1}) \approx \ln(2y)$ for large $y$.} That over 60% of all DD papers in one of the top Economics journals use at least one log-dependent variable specification underscores the importance of better understanding what we recover using a DD design with a log transformed dependent variable.

I start by outlining the key differences between a DD model with level- and log-dependent variables. In the level-dependent variables case, the DD parameter returns the difference between the treated and untreated groups in changes in the outcome variable over time. In contrast, when one uses a log transformation of the dependent variable, the DD parameter approximates the proportional difference in growth rates between treated and untreated groups.\footnote{Denoting the DD parameter for the log specification as $\beta_4$, the transformation $\exp(\beta_4)-1$ yields the precise proportional difference in growth rates.} This difference in the two specifications can yield highly disparate DD estimates when the distributions of the treated and untreated groups are sufficiently different in the pre-policy period.\footnote{Meyer1995 warned of using DD specifications in cases where the distribution of outcomes for treated and untreated groups was sufficiently different in the pre-policy period, noting in particular the issue of non-linear transformations of the dependent variable when there was a non-zero time effect: “This problem occurs because nonlinear transformations of the dependent variable imply different marginal effects on the dependent variable at different levels of the dependent variable” Meyer1995. This warning was echoed more recently by KahnLang2020.}

I then shift my attention to the singular aim of this paper: to explicate the consequences of the functional form assumption when the distribution of group outcomes in the pre-policy period differ substantively. These consequences can be stark. For a given aggregate time effect, I show that with a sufficiently large difference between group outcome distributions in the pre-policy period, it is possible for the level- and log-specifications of a DD design to yield DD parameter estimates of different signs. Thus, one may conclude that a policy or intervention raised outcomes for the treatment group using one functional form, but may conclude the precise opposite using a different functional form, even with the same data, the same sample period, and the same control variables. To my knowledge, this is the first paper to methodically document this disparity. Given the wide use of DD designs for policy evaluation in areas that can give rise to such large differences in baseline outcome distributions -- e.g., gender gaps in earnings, race gaps in the length of incarceration spells, house prices across regions or states, or school test scores across different education regimes -- this point is likely to be of broad significance to applied researchers.

I provide a condition -- based on (i) the common time effect experienced by both groups, and (ii) the difference in relative baseline outcome means -- as to when we should expect a sign switch for the log-dependent variable case. This can straightforwardly be expressed in terms of the parameters of an additive DD model. Next, using first simulations, and then two empirical examples, I provide evidence of the sign disparity in estimated DD coefficients.\footnote{In order to maintain a singular focus on the consequences of functional form assumptions for DD designs, I restrict my attention to the non-staggered timing, binary treatment case for this paper.} I verify the condition I propose with simulation results.

In both empirical settings, using the same data and the same set of controls and fixed effects, I document a statistically significant DD estimate for a levels specification that is the opposite sign that that from a log specification. Both of these empirical case studies share the feature that the outcome distributions of treated and untreated groups are sufficiently different -- a necessary, but not sufficient condition to generate a sign difference in DD estimates for level and log specifications.

The first empirical example examines differences in total earnings for white and Black men before and after the Great Recession. In the aftermath of the Great Recession, average earnings for both groups fell. Due to considerably lower baseline earnings for Black men, while I estimate a positive DD estimate in levels, I estimate a statistically significant negative DD coefficient in logs. The reason for this is that the smaller fall in earnings for Black men was larger in proportional terms. Thus, a study investigating the differential racial impact of the Great Recession on male earnings would yield an opposing conclusion depending on whether the researcher chose a levels- or log-dependent variable specification.

In the second empirical case study, I consider differential price responses in the London housing market in the period around the Brexit vote. The hypothesis in this case is that the higher exposure of Inner London properties to overseas investors may have left this area more exposed to Brexit-related uncertainty than Outer London. This is another setting with large differences in the baseline outcome distribution across groups -- the average price of a house in Inner London is almost double that of Outer London. Consequently, while the levels of house prices rose significantly more in Inner London compared to Outer London, the growth rate of Outer London was larger, which again leads to a sign difference in DD parameter estimates in the level and log specifications.

This paper contributes to the difference-in-differences literature by making clear the consequences of functional form assumptions in DD designs, particularly when working with groups with large baseline differences in outcome distributions. This builds on work by both Meyer1995 and KahnLang2020 who noted that functional form assumptions would matter in such cases. This work also relates to the recent work by Roth2023, who set out the conditions under which the parallel trend assumption is insensitive to functional form.

The remainder of the paper is organized as follows. Section (ref) provides an overview of the DD model, makes precise what the level- and log-specifications are estimating and provides a condition for when we will find a sign-switch between the two specifications. Section (ref) provides simulation evidence to both verify the sign-switch condition, and to highlight cases that generate disparate coefficient estimates across specifications. Section (ref) presents two empirical case studies where I document opposite-sign DD estimates for level- and log-outcome variable specifications. Section (ref) concludes.

The DD Model

Individual $i$ can belong to treatment group ($D_i=1$) or untreated comparison group ($D_i=0$). We observe individuals in two periods -- $T_t=0$ and $T_t=1$. Those in treatment group receive treatment in period 1. Using a potential outcomes framework, we can write down the realized outcome as $Y_{it} = (1-D_i)Y_{it}(0) + D_iY_{it}(1)$, where $Y_{it}(0)$ and $Y_{it}(1)$ are the potential outcomes for individual $i$ in absence of treatment and upon receipt of treatment respectively. We write the DD estimand as:

align[align omitted — 216 chars of source]

The sample analog of ((ref)) is the DD estimator

equation[equation omitted — 162 chars of source]

where the subscript T and C refer to treatment and control groups respectively. The subscripts 0 and 1 respectively refer to the pre- ($T_t=0$) and post-policy ($T_t=1$) periods. A simple regression specification we can use to estimate the ATT parameter is:

equation[equation omitted — 141 chars of source]

where $Treat_i$ is a treatment indicator, $Post_t$ the post-period indicator, and $\alpha_4$ is the parameter of interest.

Writing down an analogous specification with a log-transformed dependent variable of the form:

equation[equation omitted — 139 chars of source]

implies an underlying model for $Y_{it}$ that is multiplicative rather than additive:

equation[equation omitted — 140 chars of source]

where $\mu_{it} = \ln \eta_it$. As noted by Mullahy1999, and again by Ciani2019, we can write an expression for the exponentiated DD parameter, based on the multiplicative model, as:

equation[equation omitted — 211 chars of source]

where $g_C$ and $g_T$ are the respective growth rates in the outcome for control and treated groups.

Recalling that $\ln (1+z) \approx z$ for small values of $z$, and returning to Equation ((ref)), we see that the DD parameter we estimate ($\beta_4$) with a log-dependent variable can be expressed as:

equation[equation omitted — 307 chars of source]

Equation ((ref)) makes clear that when we estimate a DD specification with a log-dependent variable (as in Equation ((ref))), we are estimating an approximation of the proportional difference in growth rates of the outcome between the treated and untreated groups over the two periods. This is very different from what we measure with a level dependent variable -- the difference between groups in changes over time.

In Proposition (ref) below, I outline the conditions under which we will find a sign switch using a level and log specification.

propositionFor a given (non-zero) aggregate time effect ($\alpha_3 \neq 0$), if outcomes means at baseline are sufficiently different in relative terms, the functional form decision of specifying a level- or log-dependent variable can yield a sign difference for the DD parameter. More specifically:\\ when $0 < \left|\Delta_T-\Delta_C\right| < \left|\Delta_C \cfrac{(E[Y_{T0}]-E[Y_{C0}])}{E[Y_{C0}]}\right| \, ,$ we will have $sign(\alpha_4) \neq sign(\beta_4)$.\\ This condition may alternatively be expressed in terms of the parameters of the additive DD model as:\\ when $\left|\alpha_4\right| < \left|\alpha_3 \cfrac{\alpha_2}{\alpha_1}\right| \, ,$ we will have $sign(\alpha_4) \neq sign(\beta_4)$. \\ Proof in Appendix (ref).

The Proposition above uses the notation $E[Y_{C0}] = E[Y_{it} \mid D_i=0, T_t=0]$, $E[Y_{T0}] = E[Y_{it} \mid D_i=1, T_t=0]$, $\Delta_C = E[Y_{it} \mid D_i=0, T_t=1] - E[Y_{it} \mid D_i=0, T_t=0]$, and $\Delta_T = E[Y_{it} \mid D_i=1, T_t=1] - E[Y_{it} \mid D_i=1, T_t=0]$. In the simulation approach I detail in Section (ref), I verify Proposition (ref).

\paragraph{The Parallel Trends Assumption} The key identifying assumption for the DD estimator to return the ATT is the parallel trends assumption. In recent work, Roth2023 set out the conditions under which the parallel trend is insensitive to functional form. These conditions -- either random assignment of treatment, stationarity of the distribution of potential outcomes for the untreated setting, $Y(0)$, or a combination of these two cases -- are considerably stricter than is typically assumed by empirical researchers using DD approaches. When such conditions are not met, one must choose, and consequently justify, a functional form for the outcome variable -- a parallel trend in levels obviates a parallel trend also holding in logs, and vice versa. This point was made by Meyer1995, and has since been reiterated by several authors, including Angrist2009 and KahnLang2020.

In one sense, these findings regarding functional form and the parallel trend assumption limit the scope of cases that may benefit from the insights of this paper. If (i) a parallel trend in levels precludes a parallel trend in logs and (ii) empirical researchers typically assess the validity of using DD methods by some form of pre-trend inspection and/or testing\footnote{See Roth2022 for a discussion on the pitfalls of such pre-testing}, then does the the singular aim of this paper -- to call attention to the possibility of a sign switch in DD estimates based on functional form assumptions -- carry any weight? I argue that it does, both in cases that satisfy the conditions set out by Roth2023, or in cases where limited pre-policy data limits the statistical power to detect divergent pre-trends. Having a clearer sense of when such issues may arise will hopefully be of use to applied researchers.

Simulation Results

In this section, I provide simulation results to show that, using an additive and multiplicative model, one may estimate an ATT that differs not just in magnitude, but in sign. I take the additive model as the data generating process (DGP) in this case, and compare level- and log-dependent variable based DD specifications.\footnote{If one is interested in the complementary case where the data generating process is based on the multiplicative model, useful references include Silva2006 and Ciani2019.}

The DGP is based on Equation ((ref)), with different simulation specifications using different parameters values for $\alpha_1$, $\alpha_2$, $\alpha_3$, and $\alpha_4$. In all cases, $E[\epsilon_{it} \mid D_i, T_t] = 0$ and $\sigma_{\epsilon} = .2$. The proportion of treated and the proportion in the post period are .5 in both cases, and the total sample size is 40,000. This means each of the $2\times2$ DD cells has 10,000 observations.

The simulation results below reflect the key insight of this paper --because the DD parameter from a level- and log-dependent variable specification respectively reflect a difference in differences in levels and a proportional difference in growth rates, holding fixed the time effect, one just needs to shift the distribution of outcomes for one of the groups to drive a wedge between the resulting parameter estimates.

center[center omitted — 2,221 chars of source]

Table (ref) presents the first set of simulation results. All DD estimates for the level-dependent variables are positive, whereas the estimates for the log-dependent variable specifications are either precisely zero, or negative. For log specifications, I present both $\beta_4$, and $\exp(\beta_4) -1$, of which the latter equals the proportional difference in growth rates (see Equation ((ref))). At the base of the table, I present the cell means for the outcome variable across the four DD groups -- thus one can easily see how the parameters are generated -- as well as the proportional growth rates that I calculate from these cell means.\footnote{Given that I do not introduce additional control variables in the simulations, the underlying model is saturated -- the four parameters of Equation ((ref)) map directly to the four cells of the DD design. This means that the proportional growth rate I present in the final row of the table matches perfectly $\exp(\beta_4)-1$.}

The key point of this table is to show that one can generate a positive DD estimate from a levels specification and, by merely shifting the baseline outcome distribution of the untreated group, also generate a precise zero or a negative DD estimate from a log specification. In Table (ref) I present a complementary set of simulation results, where I fix the level specification to yield a DD estimate that is precisely zero (once again, the cells sample averages at the base of the table provide the key insight into how this is operationalized) for a levels specification, but which yields either a negative or positive DD estimate from a log specification.

figure[figure omitted — 1,975 chars of source]

In Figure (ref), I present DD estimates from both a levels-\footnote{These are presented using the left-hand $y$-axis, black line.} and log-based\footnote{These are presented using the right-hand $y$-axis, orange lines.} specification for two separate simulation experiments. Based on the parameters I choose, the DD coefficient estimate is a constant equal to 10 for the levels specification. Keeping all parameters fixed except for either the relative baseline outcome difference (Figure (ref)) or the aggregate time effect (Figure (ref)), I verify Proposition (ref). The figures make clear two key points. First, keeping the levels-based DD estimate fixed at a constant value and altering either the relative difference in baseline means ($\alpha_2/\alpha_1$) or the aggregate time effect ($\alpha_3$), it is possible to generate either a negative or a positive DD estimate for the log-dependent variable specification. Second, the sign switch occurs at precisely the point stated in Proposition (ref).

figure[figure omitted — 2,142 chars of source]

The condition provided in Proposition (ref) indicates that we require both a non-zero aggregate time effect ($\alpha_3 \neq 0$) and a difference in baseline outcome means ($\alpha_2 \neq 0$) to find a sign switch across level- and log-specifications. The purpose of Figure (ref) is to show that this is indeed the case. The underlying DGPs are identical to those underlying Figure (ref), except for the key difference that in Figure (ref) I set the aggregate time effect to zero, and in Figure (ref) I set the baseline outcome mean wedge to zero. As one can see, the log-based DD estimate (shown in orange) never crosses the zero line i.e., there is no sign switch.

Before moving to the empirical case studies, it is worth reflecting on what we have learned so far. What is concerning for applied researchers is that, if one has identified (i) a particular treatment and control group in order to evaluate a policy and (ii) a given sample period, then both the baseline outcome distributions of the respective groups and the resulting time effect is given. Thus, with a given sample, one may find oneself at the equivalent of either the left-hand side or the right-hand side of the $x$-axis in the graphs shown in Figure (ref). Unlike the simulation exercise presented here, this will not be a choice, and one will not be able to shift the ratio of relative baseline outcome means or the aggregate time effect to explore the sensitivity of the parameter estimates to functional form assumptions. I build on this point with the two empirical case studies below.

Empirical Case Studies

I now present two empirical case studies. Both of the settings were chosen with an eye on the disparate baseline outcome distribution across the groups of interest, which as I note above is a necessary, but not sufficient, condition to generate DD estimates of opposing signs for the level- and log-dependent variable specifications. I first examine the differential racial impact of the great recession on working age males in the US. I then investigate the differential impact of the Brexit vote on Inner vs Outer London property prices.

Male Earnings and The Great Recession

The empirical specification I use to investigate the differential racial impact of the Great Recession on male earnings is:

equation[equation omitted — 152 chars of source]

where $Income_{it}$ is the total income earned in the previous year, and is specified in either levels or logs, for individual $i$ in year $t$. $Black_i$ takes a value 0 for non-Hispanic white males, and a value of 1 for Black males. $Post_t$ is a dummy for the post-Great Recession period\footnote{Given that income in year $t$ reflects income from the previous year, I code $Post_t = \mathbbm{1}[Year>=2009]$ in order to capture the Great Recession kicking in in 2008.}. $X_i$ is a vector of individual characteristics that includes dummies for highest level of educational attainment, dummies for potential experience in years, an indicator for being married, and dummies for metro area classification. $\theta_{s\times t}$ is a set of state-by-year fixed effects. The error term is $\epsilon_{it}$. I specify Eicker-Huber-White standard errors throughout. I present a set of summary statistics for the setting, and an overview of the data and sample selection decisions, in Appendix (ref).

center[center omitted — 2,354 chars of source]

Table (ref) presents the key parameter estimates from both a level and log version of Equation ((ref)). The DD estimate presented in Column (1) suggests that Black men fared slightly better in the Great Recession than their white counterparts. Figure (ref) shows that both groups suffered absolute falls in incomes during this period, so the DD estimate in Column (1) reflects that the income drop for Black men was smaller than that for white men. The results in Column (2) present a starkly different conclusion, with a statistically significant, negative DD estimate. Using the raw means for the four DD cells, we can reconcile these two estimates -- although Black men experience a lower absolute income drop, relatively it was larger than for white men as the drop occurred from a lower baseline level of income, which is why the coefficient in Column (1) is positive, and in Column (2) is negative.

With the benefit of having worked through what one recovers from a level and log specification in Section (ref), and seen the consequences for these different functional form specifications of disparate baseline outcome distributions in Section (ref), with the information at hand it is a straightforward task to dissect the reason for the sign difference documented in Columns (1) and (2). However, in a typical applied setting, where a researcher is aiming to document the impact of a new policy using a DD model, the source of such disparate results may be less clear. It is the hope that this paper will aide in such settings.

The remaining columns of Table (ref) present the DD parameters for both level and log specification, splitting the sample by education levels.

London House Prices and the Brexit Vote

I next turn to a different setting -- the London housing market in the period around the Brexit vote. To understand the differential impact on house prices in Inner and Outer London, I specify a hedonic house price model of the form:

align[align omitted — 203 chars of source]

where $Price_{it}$ is the house price of house $i$ (specified in either levels or logs), sold in period $t$ (measured at the month-by-year level). $Inner$ takes the value of 0 for Outer London boroughs, and the value of 1 for Inner London boroughs.\footnote{To confuse matters there are two definitions of Inner London -- the statutory definition, and the statistical version. In order to be consistent with both measures, I apply the strictest definition -- in order for me to classify a borough as Inner London, a borough must be classified as Inner London by both definitions. This leads me to code the following boroughs as belonging to Inner London: Camden, Hackney, Hammersmith and Fulham, Islington, Kensington and Chelsea, Lambeth, Lewisham, Southwark, Tower Hamlets, Wandsworth, and Westminster.} $Post_t$ is a dummy for properties sold post-Brexit vote, and $\delta$ is the parameter of interest.

$X_{i}$ is a vector of property characteristics, specifically interactions between dummies for property type categories and new build status, and interactions between dummies for leasehold and new build status. I include interaction between the vector of housing characteristics, $X_i$, and market dummies in order to respect the “law of one price function” Bishop2020. This allows the valuation of key property characteristics to vary across local housing markets. I allow the coefficients on all housing characteristics to differ in the pre and post periods, thereby allowing the hedonic price function to shift post-policy. I do so in order to avoid conflation bias Kuminoff2014,Banzhaf2021.

$\pi_{m \times t}$ captures month-by-year housing market shocks to house prices. Housing markets are Travel To Work Areas -- similar to Commuting Zones in the US. $\theta_b$ is a spatial fixed effect at the level of Output Area, akin to a census block in the US. Output Areas (OA) are the smallest census-based geographical unit -- there are 181,408 of these in England and Wales, with an average population of 309 at the 2011 census.\footnote{\url{https://www.ons.gov.uk/peoplepopulationandcommunity/populationandmigration/populationestimates/bulletins/2011censuspopulationandhouseholdestimatesforsmallareasinenglandandwales/2012-11-23}}. The Output Area fixed effect will capture all time-invariant local amenities -- green spaces, transport links, shops, proximity to busy roads or motorways, as well as many slow-moving time-varying area characteristics (I am considering a minimum of 2 years, and a maximum of 4 years for these estimations), such as access to good schools or proximity to sources of pollution. The error term is $\epsilon_{it}$. I specify Eicker-Huber-White standard errors throughout.

center[center omitted — 3,675 chars of source]

I first provide results for all property types in panel A of Table (ref). The DD estimates for the level specification goes against my initial hypothesis that post-Brexit, Inner London property prices would suffer. For a variety of time windows around the Brexit vote ranging from 12 to 24 months, I document positive increases in Inner London prices relative to Outer London. As before, the log specification results are of the opposite sign, documenting a 4-7% decline in relative growth rates of Inner London properties. As in the previous section, we are sufficiently informed with the cell sample means, and the trends documented in Figure (ref), to understand why. Once again, the disparate baseline outcome distributions play a key role. While Inner London properties experience a slight increase in levels compared to Outer London properties, Inner London properties (which started off at a much higher baseline level) grew less, leading to a negative proportional difference in growth rates -- what the DD parameter approximates in a log specification.

Given that apartments account for almost 80% of property transactions in Inner London (see Table (ref)), in panel B of Table (ref) I restrict the sample to only apartment transactions, and repeat the analysis. The reason to do so is to get a fairer sense of the house price impact of Brexit. The unintended consequence of this sample restriction was to create a larger wedge between the baseline outcome distributions. Looking at Column (3), the ratio of sample means for treatment to control in the baseline period is 1.74 in panel A, but 1.98 in panel B. This increased disparity in baseline outcome sample means explains at least part of why the difference between the level and log specifications is even more pronounced in panel B.

Conclusion

The aim of this paper is to make clear the consequence of functional form assumptions when one uses a DD model in an empirical setting where the baseline outcome distribution across groups differs substantively. I provide a condition, based on the aggregate time effect and the relative difference in baseline means, whereby a level- and log-specification will yield estimates of the DD term of opposing signs. The key reason that this sign-switch can occur is that using a DD model with a log-dependent variable leads to the estimation not of a difference-in-differences, but rather an approximation of the relative difference in growth rates across groups.

Using both simulations and empirical examples, I show that one can obtain DD estimates of different signs depending on whether one specifies the outcome variable in levels or in logs. Given the wide use of DD models for policy evaluation in areas that can give rise to such large differences in baseline outcome distributions -- e.g., gender gaps in earnings, race gaps in the length of incarceration spells, house prices across regions or states, or school test scores across different education regimes -- this point is likely to be of broad significance to applied researchers.