Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
80,637 characters · 14 sections · 36 citation commands
Heterogeneous Elasticities, Aggregation, and Retransformation Bias
{1em} \affil[1]{Economics Department, London School of Economics and Political Science}
\begingroup \footnotetext{We thank Kimia Zargarzadeh for excellent research assistance. We thank STICERD for financial support.} \endgroup \setcounter{footnote}{0}
Economists are often interested in the arithmetic mean elasticity \(\frac{\partial\log\mathbb{E}[y|x]}{\partial \log x}\).\footnote{Almost all discussion in this paper applies to semi-elasticities as it does to elasticities - save for the decision problem axiomatisation in Section (ref). Our estimators work for semi-elasticities as they do for elasticities.} This quantity is used to understand how, for example, quantity sold changes with price, government revenue changes with tax rates, carbon emissions change with carbon taxes, or aggregate welfare changes with policy.
To estimate this object, economists use a log-log regression. However, as goldbergerInterpretationEstimationCobbDouglas1968 notes, the estimated parameters in these regressions converge to the conditional geometric mean elasticity \(\frac{\partial\mathbb{E}[\log y|x]}{\partial \log x}\)- not the arithmetic mean elasticity.
To demonstrate, suppose we estimate the log-log model \[ \log{y_{i}} = \beta \log x_{i} + v_{i}, \quad \mathbb{E}[v_{i}|x_{i}] = 0 \tag{1} .\] The geometric mean elasticity will be \(\beta\), whereas the arithmetic mean elasticity will be \(\beta + \frac{\partial \log E[\exp(v_{i})|x_{i}]}{\partial \log x_{i}}\). The second term will only be zero if the error, \(v_{i}\), and the regressor, \(x_{i}\), are fully independent - even heteroskedasticity will introduce a wedge between the two. This may be of any sign and size. Thus, any consistent estimator \(\hat{\beta}\) for \(\beta\) will in general be biased for the arithmetic mean elasticity. Following duanSmearingEstimateNonparametric1983, we call this retransformation bias.
We find the arithmetic mean elasticity and the geometric mean elasticity are frequently different in practice. Using our specification-robust debiased machine learning estimator (Section (ref)), we re-estimate 50 results from top 5 papers published in 2020. We find that 19 have estimated average arithmetic mean elasticities that are statistically significantly different from the geometric mean elasticities estimated by OLS. Amongst these results, we find the median absolute bias to be 65% of the OLS elasticity estimate. Four had opposite signs.
In previous work, santossilvaLogGravity2006 propose using PPML to estimate the arithmetic mean elasticity for trade models. This is consistent under a different model:
\[ y_{i} = x_{i}^{\gamma} \exp(u_{i}), \quad \mathbb{E}[\exp{ u_{i} }|x_{i}] = 1 \tag{2}. \]
Here, \(\gamma\) is the arithmetic mean elasticity and PPML estimates it consistently. However, direct comparisons between log-log OLS and PPML estimates are flawed as the estimands are simply different. Furthermore, under heterogeneous responses PPML does not estimate the arithmetic mean elasticity. Indeed, with heterogeneous elasticities only the geometric mean elasticity is constant. Under heterogeneity model (1) - the log-log model - is the most natural statistical model.
If the log-log model, with constant geometric mean elasticity, is the most natural why might economists want to estimate the arithmetic mean elasticity? Using the potential outcomes framework we provide monopolist and social planner examples which give intuition for when heterogeneous responses require the use of our estimator to estimate the arithmetic mean elasticity. We extend our discussion to power mean elasticities in general and motivate their estimation as the sufficient statistics for a general class of decision problem.
When economists have endogeneity concerns they will often use instrumental variables to estimate log-log structural equations. We show an impossibility result for identification of the arithmetic mean elasticity under standard IV assumptions. Identifying this object requires the additional assumption of triangularity to enable a control function approach. We develop a separate debiased machine learning estimator with instrumental variables for the average arithmetic mean elasticity and re-estimate results from the literature.
The paper proceeds as follows. Section (ref) concludes with a motivating example from public finance and a literature review. Section (ref) presents a heterogeneous elasticities model, derives a wedge between geometric and arithmetic means represented as the retransformation bias term, and generalizes to wedges between other power means. Section (ref) discusses aggregation in economic models, provides axioms under which decision-makers would estimate power mean elasticities, and shows how aggregation choice maps to welfare weights in a social planner's problem. Section (ref) discusses how to obtain arithmetic mean elasticities from RCTs and instrumental variables (IV) analysis, showing that recovering the arithmetic mean elasticity from log-log IV regressions is impossible under standard assumptions. Section (ref) presents our estimators for average arithmetic mean elasticities under regular identification assumptions and for IV analysis. Section (ref) re-estimates 50 papers from Top 5 economics journals, showing that the retransformation bias term is sizeable in practice. Section (ref) concludes.
Public economics often uses the marginal value of public funds (MVPF) to evaluate policy impacts. The MVPF is a ratio of the beneficiaries' willingness to pay and the net government cost. For evaluations of top marginal tax rate changes, hendrenUnifiedWelfareAnalysis2020 show that MVPF can be given by \[ MVPF = \frac{1}{1 - \frac{\tau}{1-\tau} z \epsilon_{eti}} \] where, \(\tau\) is the tax rate, \(z\) is the pareto parameter for the distribution of incomes Y and \[ \epsilon_{eti} = \frac{d\mathbb{E}[{Y}]}{d(1-\tau)} \cdot \frac{1-\tau}{\mathbb{E}[Y]} \] is the elasticity of the arithmetic mean of top incomes with respect to the after-tax keep rate.
To reach an estimate of the MVPF, their paper relies on elasticity estimates obtained from log-log regressions, for example from saezEffectMarginalTax2003. As we show, these will estimate the geometric mean elasticity of top incomes with respect to the after-tax keep rate, which is also equal to the average elasticity under heterogeneity. The average elasticity can be arbitrarily different from the elasticity of the arithmetic mean.
To provide intuition, consider an economy with two individuals at the top tax rate, individual 1 has an income of 1000\$ and an elasticity of 0.3, individual 2 has an income of 10000 dollars and an elasticity of 0.1. The average elasticity is 0.2, but the elasticity of the arithmetic mean is (300+1000)/11000 = 0.11. In particular, to obtain the required elasticity, one needs to weigh the elasticities of those with higher incomes more than those with lower incomes, even within top income brackets.
We see our paper as uniting the statistical property of retransformation bias with economic models of heterogenous elasticities and aggregation.
lewbelAggregationLogLinearModels1992 studies when individual heterogenous elasticity agents can be approximated using a representative agent. We, conversely, focus on situations where they cannot be modelled as such, and focus on the properties of different estimators under heterogenous elasticities. breinlichTradeGravityAggregation2024 are closely related. They study the performance of PPML in trade models when elasticities vary at the sector level, finding that it approximates a weighted average of the sectoral elasticities when trade costs do not vary at sector level, but is biased otherwise. We generalize their results to arbitrary elasticity distributions, consider the performance of OLS and provide a new estimator for the average arithmetic mean elasticity.
Heterogenous responses are often modelled using random coefficients models wooldridgeFixedEffectsRelatedEstimators2005. We study bias arising from non-linearity in the heterogenous coefficients, specifically in constant individual elasticity models. To our knowledge, we are the first to use heterogenous coefficient models to provide economic meaning to the statistical problem of retransformation bias. qianTestingOmittedHeterogeneity2026 provides a test for omitted heterogeneity, and provides an estimator for the average coefficient for the case when heterogeneity is asymptotically disappearing. We focus on cases where the heterogeneity distribution is stable, and give examples for when the relevant estimand is not the average coefficient.
Transformation bias has long been recognized in economics and statistics literature. goldbergerInterpretationEstimationCobbDouglas1968 notes that while transformation bias affects conditional expectations, it does not affect conditional medians---though the latter requires a different identifying assumption (zero conditional median of errors rather than zero conditional mean). We focus on retransformation bias from the model specification \(\mathbb{E}[\log y|x]= x \beta\).
duanSmearingEstimateNonparametric1983 provides a non-parametric bias-corrected estimator for \(\mathbb{E}[y|x]\) under full independence between errors and regressors. manningLoggedDependentVariable1998 extends this to binary treatments. Our estimator nests Manning's approach while considering a more comprehensive distributional framework. aiSemiparametricDerivativeEstimator2008 develop related semi-parametric estimators for the conditional expectation and its derivatives, but our paper builds on their estimator. Our work relies on debiased machine learning chernozhukovDoubleDebiasedMachine2018, which allows us to bypass the curse of dimensionality that their proposed spline approach faces. The estimator we propose approximates an unknown, potentially high-dimensional function of the regressors using flexible machine-learning methods. Such methods typically require regularization to control overfitting, which in turn slows convergence and introduces bias. To address this issue, we implement the locally robust semiparametrics of chernozhukovLocallyRobustSemiparametric2022 in our elasticity setting. This allows us to remove the leading regularization bias and to obtain an estimator that is both consistent and asymptotically normal under their high-level conditions. Our IV estimator uses the automatic debiasing procedure from chernozhukovAutomaticDebiasedMachine2024 to avoid estimation of complex Riesz representers. For both estimators we use score matching hyvarinenEstimationNonnormalizedStatistical2005, hanNeuralNetworkbasedScore2024.
When \(y\) takes zero values, researchers often use \(\log(1+y)\) or \(\operatorname{arcsinh}(y)\) transformations. chenLogsZerosProblems2024 show that coefficients from these transformations are sensitive to scaling of \(y\), arising from difficulties distinguishing extensive and intensive margin effects. However, they do not consider transformation bias from non-linear transformations when making inferences about \(\mathbb{E}[y_i|x_i]\). We only focus on cases where the outcome variable is strictly positive, relative to their focus on scaling when it can be zero. mullahyWhyTransformPitfalls2024 discusses transformation bias in the \(\log(1+y)\) context as another reason to avoid variable transformations.
bellemareElasticitiesInverseHyperbolic2020 discuss when coefficients from \(\operatorname{arcsinh}\) transformations can be interpreted as percent changes, but do not address transformation bias from conditional expectations of transformed errors. nortonInverseHyperbolicSine2022 applies duanSmearingEstimateNonparametric1983's smearing estimator to arcsinh regressions, but doesn't tackle the case where there is dependence between regressors and the error term. thakralWhenAreEstimates2025 suggest power transformations over \(\log(1+y)\) or \(\operatorname{arcsinh}\), citing scaling sensitivity. Our paper demonstrates that transformation bias occurs with any non-linear transformation.
santossilvaLogGravity2006 encourage researchers to use Poisson pseudo-maximum likelihood (PPML) estimation to estimate the parameters of the exponential model. However, the exponential is a different model to the log-linear model, and in general they cannot both hold at the same time. We show that the exponential model is incompatible with each individual having a heterogeneous, constant elasticity. The log-linear model allows for non-constant elasticities of \(\mathbb{E}[y_i|x_i]\) over \(x\). In this paper we show conditions for which both models can coexist, and show that in most empirical models, one model being correctly specified implies the other cannot be.
Let \(x \in \mathbb{R}_+\) denote a scalar variable of interest (e.g. a price, tax, policy parameter, or cost shifter). For each unit \(\omega \in \Omega\), let \(Y(x;\omega) > 0\) denote an outcome of interest, such as demand, output, emissions, or income. Units are distributed according to a probability measure \(F\). Expectations \(\mathbb{E}[\cdot]\) are taken with respect to \(F\).
The index \(\omega\) summarizes unit-specific characteristics that are fixed with respect to \(x\) and unobserved by the econometrician. These characteristics may reflect preferences, technologies, constraints, or other sources of persistent heterogeneity.
We allow outcomes to respond proportionally to changes in \(x\), while permitting heterogeneity in both scale and responsiveness across units.
The term \(a(\omega)\) captures heterogeneity in scale, while \(\varepsilon(\omega)\) captures heterogeneity in responsiveness to \(x\). No restrictions are imposed on the joint distribution of \((a(\omega), \varepsilon(\omega))\). Define the individual elasticity \[ \varepsilon(\omega) \equiv \frac{d \log Y(x;\omega)}{d \log x}. \] Under \hyperref[asm-randomcoef]{Assumption \ref*{asm-randomcoef}}, individual elasticities are constant in \(x\) but heterogeneous across units. That is, every unit behaves as if it has a constant elasticity, but that constant elasticity is different between units.
Economic objectives typically depend on aggregate outcomes. We therefore consider a family of aggregation rules that nests common cases.
Special cases include:
This family provides a continuous mapping between aggregation rules that emphasize total outcomes and those that emphasize typical outcomes.
For each aggregator \(M_\phi(x)\), define the associated elasticity.
Under \hyperref[asm-randomcoef]{Assumption \ref*{asm-randomcoef}}, power-mean elasticities admit a simple representation.
Thus, \(\varepsilon_\phi(x)\) is a reweighted average of individual elasticities, where the weights depend on both the aggregation parameter \(\phi\) and the level of \(x\).
The two common cases we have discussed are immediate,
The elasticity \(\varepsilon_\phi(x)\) measures the percentage response of an aggregate outcome defined by the aggregation rule \(\phi\). Different values of \(\phi\) emphasize different parts of the cross-sectional distribution. \(\phi \rightarrow 0\) weights units symmetrically in log space and reflects the response of a typical unit; \(\phi = 1\) weights units in proportion to their contribution to total outcomes. Holding elasticities fixed, higher values of \(\phi\) put more weight on high outcomes, and lower values of \(\phi\) put more weight on low outcomes.
As shown below, different economic objectives correspond to different values of \(\phi\), and only a subset of these elasticities are identified by standard empirical approaches.
We begin by formalizing the distinction between elasticities associated with different aggregation rules.
\hyperref[myprp-noneqhet]{Proposition \ref*{myprp-noneqhet}} can also be stated as, even if every individual is isoelastic, the elasticities of every mean except the geometric mean are not isoelastic. \hyperref[myprp-noneqhet]{Proposition \ref*{myprp-noneqhet}} highlights a fundamental distinction between elasticities defined by different aggregation rules. The geometric elasticity \(\varepsilon_0\) averages individual elasticities symmetrically in log space and is invariant to changes in \(x\). In contrast, elasticities associated with other aggregation rules reweight individual responses in a manner that depends on the level of \(x\).
This reweighting reflects how, as \(x\) changes, units with different responsiveness contribute differentially to aggregate outcomes. As a result, aggregate elasticities relevant for totals, profits, or other non-log objectives generally vary with \(x\) even when individual elasticities are constant.
This also implies that parameters obtained from log-log OLS cannot correspond to any elasticity other than that of the geometric mean, as all the other ones are non-constant in x. Moreover, any estimator which targets a constant power mean elasticity except geometric cannot be consistent. For example, PPML targets a constant arithmetic mean elasticity, therefore under heterogenous elasticities, the PPML estimator is inconsistent.
One illustrative example is demand, let \(D(p; \omega)\) reflect the quantity demanded at price \(p\). Then the elasticity of the arithmetic mean of demand to price is \[ \varepsilon_{1}(p) = \dfrac{\mathbb{E}[D(p; \omega) \varepsilon(\omega)]}{\mathbb{E}[D(p; \omega)]} . \]
We necessarily weigh those with higher demand more. A high elasticity individual with high demand affects total demand more than a high elasticity individual with low demand! Alternatively - for the same demand, one should weigh a high elasticity individual more than a low elasticity individual, as they will affect total demand more.
Proof in Section (ref)
Since the distribution of \(\omega\) is fixed, we may differentiate inside the expectation to obtain \[ \varepsilon_\phi(x)-\varepsilon_0 = \frac{\mathbb{E}\!\left[\exp(\phi a(\omega))\,x^{\phi\tilde\varepsilon(\omega)}\,\tilde\varepsilon(\omega)\right]} {\mathbb{E}\!\left[\exp(\phi a(\omega))\,x^{\phi\tilde\varepsilon(\omega)}\right]} \] which is a tilted mean of the difference from average.
\hyperref[myprp-wedgedecomp]{Proposition \ref*{myprp-wedgedecomp}} demonstrates that the difference between \(\varepsilon_0\) and \(\varepsilon_\phi(x)\) requires understanding how the tilted moment \(\mathbb{E}[Y(x)^\phi]\) varies with \(x\). Log-linear methods identify \(\varepsilon_0\) because it depends only on the first moment of \(\log Y(x;\omega)\). By contrast, \(\varepsilon_\phi(x)\) depends on how heterogeneity is reweighted in levels as \(x\) changes, which is summarized by the second term in the decomposition.
This is a `tilted' average of the residuals, where those individuals with higher scales (\(a\)), and larger differences from the average elasticity, are given more weight.
Focusing on the arithmetic geometric wedge, \[ \varepsilon_1(x)-\varepsilon_0 = \frac{\mathbb{E}\!\left[\exp(a(\omega))\,x^{\tilde\varepsilon(\omega)}\,\tilde\varepsilon(\omega)\right]} {\mathbb{E}\!\left[\exp(a(\omega))\,x^{\tilde\varepsilon(\omega)}\right]}. \]
A first-order expansion around \(x=1\) yields \[ \varepsilon_1(x)-\varepsilon_0 = \frac{\mathbb{E}\!\left[\exp(a(\omega))\,\tilde\varepsilon(\omega)\right]}{\mathbb{E}\!\left[\exp(a(\omega))\right]} \;+\; \log x\left( \frac{\mathbb{E}\!\left[\exp(a(\omega))\,\tilde\varepsilon(\omega)^2\right]}{\mathbb{E}\!\left[\exp(a(\omega))\right]} - \left(\frac{\mathbb{E}\!\left[\exp(a(\omega))\,\tilde\varepsilon(\omega)\right]}{\mathbb{E}\!\left[\exp(a(\omega))\right]}\right)^2 \right) \;+\; o(\log x). \]
The first term captures sorting between scale \(\exp(a(\omega))\) and responsiveness \(\tilde\varepsilon(\omega)\) at \(x=1\). One can think of this, in a demand scenario, as describing the covariance between the demand shock and elasticity. If demand shocks are larger for those with higher elasticities, then, this term will be positive, else it will be negative. The coefficient on \(\log x\) is the (scale-weighted) variance of \(\tilde\varepsilon(\omega)\) and is therefore weakly positive, with strict positivity whenever \(\tilde\varepsilon(\omega)\) is non-degenerate.
The tilted-mean representation immediately delivers a sharp condition under which the arithmetic and geometric elasticities coincide.
A natural estimator in this model under Gaussianity is to use PPML to estimate a regression with \(\log{x}\) and \((\log{x})^2\), and combine the obtained parameters to get the arithmetic mean elasticity.
The results in Section (ref) have shown that, under heterogeneous responsiveness, the elasticity of an aggregate outcome is not a single parameter but depends on how outcomes are aggregated. The geometric elasticity \(\varepsilon_0\) is invariant to \(x\) and corresponds to the response of the log-aggregate (a typical-unit notion). This also means that OLS generally estimates the average response. By contrast, elasticities relevant for aggregates in levels (including the arithmetic mean) generally vary with \(x\) because changes in \(x\) tilt the relative importance of units with different elasticities. The knife-edge condition in \hyperref[mycor-nowedge]{Corollary \ref*{mycor-nowedge}} clarifies when this reweighting is absent, while the Gaussian example illustrates how the entire family of power-mean elasticities can be characterized in closed form.
A natural question that arises is why economists may care to estimate one power mean elasticity over another. This section motivates why different economic decision problems depend on different aggregation rules, and therefore on different elasticities. We give two examples and appeal to an axiomatisation of a common decision problem. We show that power mean elasticities are sufficient statistics for this decision problem. Finally, we note that different power means correspond to different risk or inequality aversion for the decision maker facing this problem.
Generalising the two previous examples, we provide the following axiomatisation of a natural decision problem for which power mean elasticities are sufficient statistics. Consider a statistical decision problem \(\mathcal{D} = \langle \Omega, \mathcal{X}, \mathcal{Y}, \succeq \rangle\), where a Decision Maker (DM) chooses a policy \(x \in \mathcal{X} \subseteq \mathbb{R}_{+}\), or distribution of policies \(F_{x }\in\Delta(\mathcal{X})\), to maximize a functional over a population of heterogeneous units \(\Omega\) with outcomes \(\mathcal{Y}\).
The DM believes that the units the inputs and outcomes are denominated in should not matter to the intensity of response for each type. \hyperref[axm-scale-invariance]{Axiom \ref*{axm-scale-invariance}} implies that types respond to percentage changes in inputs with percentage changes in outcomes, ruling out level-dependent functional forms.
Standard assumptions guarantee that a complete, continuous preorder \(\succsim\) over outcomes admits a numerical representation \(M: \mathcal{Y}^\Omega \to \mathbb{R}\). We provide four further axioms that describe how \(M\) should behave. Firstly, \hyperref[axm-pareto]{Axiom \ref*{axm-pareto}} states \(M\) should increase in Pareto improvements. Secondly, \hyperref[axm-anonymity]{Axiom \ref*{axm-anonymity}} requires that the DM does not care about the identities of any of the types. Thirdly, \hyperref[axm-decomposability]{Axiom \ref*{axm-decomposability}} is a separability condition. \(M\) should be separable across types and the aggregate value of the total population can be computed solely from the aggregate values of its disjoint subpopulations and their relative masses. Finally, \hyperref[axm-homotheticity]{Axiom \ref*{axm-homotheticity}} requires the decision maker themselves not care about the units the outcome is denominated in.
These axioms uniquely identify the structural model and the objective function.
This representation justifies the potential outcomes random elasticities model used in the literature, where \(\log Y(x; \omega) = a(\omega) + \varepsilon(\omega) \log x\), with \(a(\omega) = \log \alpha(\omega)\). Furthermore, for any decision problem satisfying our axioms the aggregate outcome level and the power-mean elasticity are sufficient statistics.
The parameter \(\phi\) captures the curvature of the aggregation. As is commonly known, this parameter nests two distinct, but incompatible, economic interpretations. We may interpret the optimization problem \(\max_x M_\phi(x)\) as equivalent to:
To demonstrate, consider a utilitarian social planner who evaluates policy \(t\) according to the social welfare function
\[ W(t) = \int u(y(t, \omega)) \ dF(\omega) \tag{3}\]
where \(y(t, \omega)\) denotes the outcome realized by individual \(\omega\) under policy \(t\), and \(F(\omega)\) represents the distribution of individual characteristics in the population. Suppose \(u(\cdot)\) takes the form
\[ u(z) =
\]
where \(\rho \geq 0\) parameterizes the degree of inequality aversion. Then
\[ W(t)=\int \frac{y(t,\omega)^{1-\rho}}{1-\rho} dF(\omega). \]
Alternatively, consider an inequality averse social planner who evaluates policy \(t\) according to the social welfare function
\[ W(t)=\int \frac{v(y(t,\omega))^{1-\rho}}{1-\rho} dF(\omega).\tag{4} \]
Equations \((3)\) and \((4)\) are only equal if \(v\) is the identity function. The choice among power-mean elasticities therefore either encodes an implicit choice among social welfare functions or an assumption about the utility of the individuals being aggregated over. The monopolist's problem depends on the response of total quantity sold and therefore on \(\varepsilon_1(p)\). A planner's welfare calculus depends on how price changes affect average log outcomes (or other concave aggregates), and therefore on \(\varepsilon_0\) (or, more generally, \(\varepsilon_\phi(p)\) for \(\phi\neq 1\)). When elasticities are heterogeneous, these objects differ, so the elasticity relevant for profit maximization need not coincide with the elasticity relevant for welfare analysis. Yet both are represented in our framework.
Different power-mean elasticities may be of interest, even for the same problem. We characterized the wedge between these elasticities. This wedge can be identified from observables under assumptions we discuss in the next section. Once identified, our approach to estimation follows from manningLoggedDependentVariable1998, who estimates these wedges in a specific setting with binary treatment variables and no controls.
We have shown power mean elasticities are sufficient statistics for a common class of decision problems. When can we estimate them? This is a question of identification. We outline the statistical model and discuss two common identification strategies - randomised controlled trials and instrumental variables.
An alternative way to write the potential outcome model is \[ \log{y_{i}}(x) = \alpha_{0} + \beta_{0} \log{x} + u_{i}(x) \] with \(u_{i}(x) =\alpha_{i} - \alpha_{0} + (\beta_{i} - \beta_{0}) \cdot \log{x}\) such that \(\mathbb{E}[u_{i}(x)] = 0 \ \forall x\).
Suppose, the \(X\) we observe in the population is randomly distributed and mean independent of the random coefficients. We observe \(Y\) such that \(\log Y = \log y(X)\). Then, we can write our population model as \[ \log Y_{i} = \alpha_{0} + \beta_{0} \log X_{i} + u_{i}, \quad \mathbb{E}[u_{i}|X_{i}] = 0. \]
Under this model \(\beta_{0}\) is identified, OLS is an unbiased estimator of it, and \(\beta_{0}=\epsilon_{0}\) - the average elasticity of the population.
However, if we try to estimate a different elasticity, for example the causal arithmetic mean elasticity \(\epsilon_{1}\), we will not be able to do so without the stronger assumption that \(X \perp \!\! \perp (\alpha_{i},\beta_{i})\). Independence here requires both no selection on level and no selection on gains. We may believe this under random assignment or similar assumptions.
Even in a randomized controlled trial with binary treatment, retransformation bias still warrants attention as
\[\frac{\mathbb{E}[Y(1; \omega) - Y(0; \omega)]}{\mathbb{E}[Y(0; \omega)]} \neq \exp\big(\mathbb{E}[\log Y(1; \omega) - \log Y(0; \omega)]\big) - 1.\]
Under the heterogeneous semi-elasticity model where \(\log Y(x;\omega) = a(\omega) + \varepsilon(\omega) x\) for \(x \in \{0,1\}\), OLS on \(\log Y_i\) gives a coefficient estimate that converges to the average treatment effect in logs, \(\mathbb{E}[\varepsilon(\omega)] \equiv \varepsilon_0\). Hence, parameters obtained from log-linear OLS do not directly correspond to the percentage difference in average outcomes between the treated and untreated.
However, in the binary treatment case, it is always possible to directly do OLS in levels, or PPML to get the percentage change in the arithmetic mean. In fact, the manningLoggedDependentVariable1998 estimator (which is equivalent to ours with a binary variable) gives exactly the same point estimates as PPML with binary treatments (see Section (ref)).
This is because with a single binary treatment the exponential and log-linear models are compatible. The log-linear statistical model decomposes as
\[\log Y_{i} = \alpha_{0} + \varepsilon_{0} x_{i} + u_{i}(x_i), \quad \mathbb{E}[u_{i}(x_i)|x_{i}] = 0\]
where \(\alpha_0 = \mathbb{E}[a(\omega)]\), \(\varepsilon_0 = \mathbb{E}[\varepsilon(\omega)]\), and \(u_i(x) = (a(\omega) - \alpha_0) + (\varepsilon(\omega) - \varepsilon_0)x\).
The corresponding exponential model is
\[Y_{i} = \exp(\gamma_{0} + \gamma_{1} x_{i}) \eta_{i}, \quad \mathbb{E}[\eta_{i}|x_{i}] = 1\]
with the parameters mapping exactly as
\[
\]
This equivalence does not hold if there is a randomly assigned continuous treatment \(x\) or other regressors.
When random assignment is violated we must rely on instrumental variables approaches. However, instrumental variables approaches face challenges in this setting. Instrumental variables approaches cannot, under standard assumptions, recover the arithmetic mean elasticity from log-linear structural models with heterogeneous individual effects. We show a sequence of impossibility results of increasing generality. We then show that assuming a triangular first-stage structure restores identification.
Let each individual have potential outcome \[ \log Y_i(x) = a_i + \varepsilon_i \log x \] where \(a_i\) is an individual intercept and \(\varepsilon_i\) is an individual semi-elasticity, with joint distribution \((a_i, \varepsilon_i) \sim F\). Write \(\beta_0 = \mathbb{E}[\varepsilon_i]\) and \(\alpha_0 = \mathbb{E}[a_i]\). The geometric mean semi-elasticity is \(\mathbb{E}[\varepsilon_i] = \beta_0\), constant in \(x\). The arithmetic mean semi-elasticity is \[ \theta(x) = \frac{\partial \log \mathbb{E}[Y_i(x)]}{\partial x} = \frac{\mathbb{E}[\varepsilon_i e^{a_i + \varepsilon_i \log x}]}{\mathbb{E}[e^{a_i + \varepsilon_i \log x}]}, \] a level-weighted average of individual semi-elasticities that varies with \(x\).
We observe \((Y_i, X_i, Z_i)\) where \(\log Y_i = a_i + \varepsilon_i X_i\) and \(Z_i\) is an instrument for \(X_i\). The statistical model decomposes as \[ \log Y_i = \alpha_0 + \beta_0 \log X_i + u_i(X_i), \qquad u_i(x) = (a_i - \alpha_0) + (\varepsilon_i - \beta_0) \log x, \] with \(\mathbb{E}[u_i(x)] = 0\) for all \(x\).
First, let's consider standard IV assumptions. Let \(Z\) satisfy mean independence \(\mathbb{E}[\varepsilon_i \mid Z] = \mathbb{E}[\varepsilon_i]\) and \(\mathbb{E}[a_i \mid Z] = \mathbb{E}[a_i]\), together with relevance \(\text{Cov}(X,Z) \neq 0\). Under these conditions, IV consistently estimates \(\beta_0 = \mathbb{E}[\varepsilon_i]\), the geometric mean elasticity.
IV assumptions constrain only first moments of \((a_i, \varepsilon_i)\) conditional on \(Z\). The arithmetic mean elasticity depends on \(\mathbb{E}[\varepsilon_i e^{a_i + \varepsilon_i \log x}]\), which involves the entire joint distribution of \((a_i, \varepsilon_i)\). Our proof provides two observationally equivalent data-generating processes that share the same \(\beta_0\) but yield different values of \(\theta(x)\). In one, the observed heteroskedasticity of residuals arises from selection (high-variance individuals choose certain treatment values). In the other, it arises from genuine heterogeneity in \(\varepsilon_i\) across the population. Since IV restricts only conditional means, it cannot distinguish these mechanisms.
Strengthening IV to a non-parametric specification, identifying \(f(x)\) via completeness, does not resolve the problem since \(\theta(x)\) depends on the joint distribution of \((a_{i},\epsilon_{i})\) not just conditional means.
One might hope that continuous instruments with full independence might provide identification . The following result shows that they cannot. For any model satisfying mild regularity conditions, there exists an observationally equivalent model with a different arithmetic mean semi-elasticity.
The proof exploits the geometry of the identification problem. Write \(h(a, \varepsilon, x, z) = k(x \mid a, \varepsilon, z) f(a, \varepsilon)\) for the joint density of \((a_i, \varepsilon_i, X_i)\) conditional on \(Z = z\). The structural equation \(\log Y = a + \varepsilon \log x\) implies that for each fixed \((x, z)\), the observable density \(f(\ell, x \mid z)\) is the Radon projection of \(h(\cdot, \cdot, x, z)\) along lines \({(a, \varepsilon) : a + \varepsilon \log x = \ell}\). Without triangularity, \(h(\cdot, \cdot, x, z)\) is a different unknown function for each \((x, z)\). We observe one projection of each, rather than many projections of one. The construction produces a compactly supported perturbation of \(f\) that lies in the null space of every relevant Radon projection, thereby changing the joint distribution of \((a_i, \varepsilon_i)\) without affecting any observable.
With full independence of discrete instruments for discrete treatments, an arithmetic mean elasticity for compliers is identified via standard LATE arguments as per imbensIdentificationEstimationLocal1994.
The preceding results establish that IV methods cannot identify \(\theta(x)\) without further structural restrictions. A triangular first-stage assumption resolves this by collapsing the selection mechanism to a single dimension.
Triangularity and the sufficient support condition ensures that the Radon transform, restricted to the angles generated, is injective on \(L^2(\mathbb{R}^2)\), so \(f(a, \varepsilon \mid V = v)\) is fully recovered. Integrating over \(f_V\) then gives \[ \mathbb{E}[e^{a + \varepsilon \log x}] = \int \mathbb{E}[e^{a + \varepsilon \log x} \mid V = v]\ f_V(v)\ dv \] where each conditional expectation is identified.
Regardless of whether causal identification holds, we may wish to estimate the arithmetic mean elasticity of a log-log model. As a semi-parametric estimator the estimator we provide in this section will converge to the average arithmetic mean elasticity under regularity assumptions. Even when economists run log-log (log-linear) projection models, they can use our estimator to obtain the average arithmetic (semi-)elasticity of the observed variables.
We present estimators that converge under general conditions to the average arithmetic mean elasticity. In the first part, we provide an estimator which applies a correction to the log-log model directly which will give causal estimates under full independence of the regressor and random coefficients. In the second part, we provide an instrumental variables estimator which uses a control function approach. We give conditions for their convergence and asymptotic distribution.
As we must estimate high-dimensional functions we turn to machine learning methods to avoid the curse of dimensionality. Neyman-orthogonalised debiased machine learning methods chernozhukovLocallyRobustSemiparametric2022, chernozhukovDoubleDebiasedMachine2018 permit root-\(n\) convergence even with neural networks.
We estimate the average arithmetic mean elasticity (semi-elasticity) starting from a log-log (log-linear) model. We cannot provide faster converging estimators for the non-parametric power mean sufficient statistic in our decision problem than other non-parametric estimators - though our semi-parametric estimator for the average converges faster than the naive “plug-in” estimator chernozhukovLocallyRobustSemiparametric2022. However, we note that if the decision problem assumptions hold then estimating a non-parametric model on the residuals of the log-log model will have strictly lower asymptotic variance to second order than a general non-parametric approach without the isoelastic heterogeneity assumption. We don't yet provide asymptotics for non-parametric estimation.
Our estimators develop on aiSemiparametricDerivativeEstimator2008 who use splines to estimate the conditional mean of the outcome variable and its derivative in a non-linear model under transformation.
Researchers may want to understand the arithmetic percentage change rather than the arithmetic elasticity when regressors are discrete. We have developed this estimator but do not provide it here yet. Some of our results in Section (ref) use this estimator.
Let \(W=(Y,X,Z)\), where \(X\in\mathbb R^{d_x}\) are the variables of interest (all derivatives are taken with respect to \(x\)) and \(Z\) are controls. Assume the outcome satisfies \[ \log Y = \beta_0^\top X + \gamma_0^\top Z + u, \qquad \mathbb E[u\mid X,Z]=0. \] This is isomorphic to the random coefficients model when observed X is independent from \(\alpha(\omega)\) and \(\varepsilon(\omega)\).
Define the primitive nuisance object \[ m_0(x,z) := \mathbb E[e^u\mid X=x,Z=z], \qquad 0<\underline m \le m_0(x,z)\le \overline m<\infty. \] Then \[ \mathbb E[Y\mid X=x,Z=z] = e^{\beta_0^\top x+\gamma_0^\top z}\,m_0(x,z), \] and \[ \nabla_x \log \mathbb E[Y\mid X,Z] = \beta_0 + \nabla_x \log m_0(X,Z). \]
We target the average semi-elasticity of the conditional mean with respect to \(X\): \[ \theta_0 := \mathbb E\!\left[\nabla_x \log \mathbb E[Y\mid X,Z]\right] = \mathbb E\!\left[\beta_0 + \frac{\nabla_x m_0(X,Z)}{m_0(X,Z)}\right]. \] Elasticities with respect to \(Z\) are not targeted. Note that if X was log K for some variable K, this would be the average arithmetic mean elasticity with respect to K.
Let \(f_0(x\mid z)\) denote the conditional density of \(X\) given \(Z=z\), differentiable in \(x\) with \(f_0(x\mid z)>0\). Define \[ \alpha_0(x,z) := -\frac{\nabla_x f_0(x\mid z)}{f_0(x\mid z)\,m_0(x,z)}. \]
Define the residual \[ \delta := e^u - m_0(X,Z), \qquad \mathbb E[\delta\mid X,Z]=0. \]
Note that \[ \alpha_0(X,Z)\,\delta = - \frac{\nabla_x f_0(X\mid Z)}{f_0(X\mid Z)} \left(\frac{e^u}{m_0(X,Z)}-1\right), \]
which is the standard locally robust correction term.
Let \(m_0'(x,z):=\nabla_x m_0(x,z)\).
Define the score \[ \phi(W;\theta, \beta, \gamma, m, f) := \beta + \frac{m'(X,Z)}{m(X,Z)} - \theta \;+\; \alpha(X,Z)\,\big(e^{u(\beta,\gamma)} - m(X,Z)\big), \] where \[ u(\beta,\gamma) := \log Y - \beta^\top X - \gamma^\top Z, \qquad \alpha(x,z) := \frac{\nabla_x f(x\mid z)}{f(x\mid z)\,m(x,z)}. \]
At the true parameter values \((\theta_0,\beta_0,\gamma_0,m_0,f_0)\), \[ \mathbb E[\phi(W;\theta_0)]=0. \]
A formal proof is in Section (ref).
We call our estimator the doubly robust estimator of the arithmetic mean (DREAM).
\begingroup
\addtocounter{theorem}{-1}
\endgroup
\hyperref[asm-ordinary-sampling]{Assumption \ref*{asm-ordinary-sampling}} can be weakened, and the estimator still converges, for example in fixed effects models.
\begingroup
\addtocounter{theorem}{-1}
\endgroup
The first condition requires all moments of \(u|X=x\) to be finite. Positivity is ensured as \(e^u > 0 \quad \forall u.\)
\begingroup
\addtocounter{theorem}{-1}
\endgroup
\hyperref[asm-ordinary-moments]{Assumption \ref*{asm-ordinary-moments}} is a standard GMM condition, which ensures that \(\theta_{0}\) is identifiable and its estimators have finite variance.
\begingroup
\addtocounter{theorem}{-1}
\endgroup
For a sufficiently smooth m and density, both these conditions are satisfied.
Let \[ \phi_0(W) := \phi(W;\theta_0,\beta_0,\gamma_0,m_0,f_0), \qquad V := \mathbb E[\phi_0(W)^2]. \]
Since \[ \partial_\theta \mathbb E[\phi(W;\theta)]\big|_{\theta=\theta_0} = -1, \] the influence function equals \(\phi_0(W)\).
Under assumptions \hyperref[asm-ordinary-sampling]{Assumption \ref*{asm-ordinary-sampling}} -- \hyperref[asm-ordinary-nuisance-rates]{Assumption \ref*{asm-ordinary-nuisance-rates}} and the Neyman orthogonality of the score established in Appendix A, Theorem 3.1 of chernozhukovDoubleDebiasedMachine2018 implies
A consistent variance estimator is \[ \hat V := \frac1n\sum_{i=1}^n \hat\phi_i^2, \qquad \hat\phi_i := \phi(W_i;\hat\theta,\hat\beta,\hat\gamma,\hat m_{-k(i)},\hat f_{-k(i)}),\] and confidence intervals can be constructed as \[ \hat\theta \pm z_{1-\alpha/2}\sqrt{\hat V/n}. \]
Neural networks can be used to estimate \(m_0(x,z)=\mathbb E[e^u\mid x,z]\) and its derivative \(\nabla_x m_0(x,z)\) via automatic differentiation. Sieve-type approximation and rate results are given in farrellDeepNeuralNetworks2021 and schmidt-hieberNonparametricRegressionUsing2020. In our implementation, we also use neural networks to estimate \(\frac{\nabla_x \hat f(x\mid z)}{\hat f(x\mid z)\,}\) using score matching hyvarinenEstimationNonnormalizedStatistical2005
Let \(W = (Y, X, Z)\), where \(X \in \mathbb{R}^{d_x}\) is endogenous and \(Z \in \mathbb{R}^{d_z}\) are instruments. Assume a triangular system \[ X = g_0(Z) + V, \qquad Z \perp V, \] \[ \log Y = \beta_0^\top X + \rho_0^\top V + \epsilon, \qquad \mathbb{E}[\epsilon \mid X, V] = 0. \] The first-stage residual \(V := X - g_0(Z)\) captures the endogeneity in \(X\). Conditional on \(V\), variation in \(X\) is driven solely by the instrument \(Z\) and is therefore exogenous.
Define the primitive nuisance objects \[ m_0(x, v) := \mathbb{E}[e^u \mid X = x, V = v], \qquad u := \log Y - \beta_0^\top X - \rho_0^\top V, \] \[ \mu_0(x) := \mathbb{E}_V[m_0(x, V)] = \int m_0(x, v) f_V(v) , dv. \] The function \(\mu_0(x)\) is the average structural function (ASF), representing the expected outcome under an intervention setting \(X = x\) while integrating over the marginal distribution of \(V\).
We target the average semi-elasticity of the ASF with respect to \(X\), \[ \theta_0 := \mathbb{E}\left[\nabla_x \log \mu_0(X)\right] = \mathbb{E}\left[\beta_0 + \frac{\nabla_x \mu_0(X)}{\mu_0(X)}\right]. \]
Define the density ratio \[ \omega_0(x, v) := \frac{f_V(v)}{f_{V \mid X}(v \mid x)}, \] and the marginal score \[ S_X(x) := \nabla_x \log f_X(x). \]
Let \(\mu_0'(x) := \nabla_x \mu_0(x)\). The score is given by \[ \phi(W; \theta, \beta, \rho, m, g, \omega, S_X, \lambda) := \beta + \frac{\mu'(X)}{\mu(X)} - \theta + \phi^m(W) + \phi^g(W), \] where the correction for the outcome nuisance \(m\) is \[ \phi^m(W) := -\frac{\omega(X, V) S_X(X)}{\mu(X)} \left( e^{u(\beta, \rho)} - m(X, V) \right), \] with \(u(\beta, \rho) := \log Y - \beta^\top X - \rho^\top V\), and the correction for the first-stage nuisance \(g\) is \[ \phi^g(W) := -\lambda(Z)^\top (X - g(Z)). \] The function \(\lambda_0(z)\) is the Riesz representer for the pathwise derivative with respect to \(g\). Rather than deriving its closed form, we estimate it via automatic debiasing chernozhukovAutomaticDebiasedMachine2022.
At the true parameter values, \(\mathbb{E}[\phi(W; \theta_0)] = 0\).
A formal proof is in Section (ref).
Estimate \(g_0(z) = \mathbb{E}[X \mid Z = z]\) using a flexible learner, yielding \(\hat{g}(z)\), and form residuals \[ \hat{V} = X - \hat{g}(Z). \]
Estimate \((\beta_0, \rho_0)\) from the regression \(\log Y = \beta^\top X + \rho^\top \hat{V} + \epsilon\) and form residuals \(\hat{u} = \log Y - \hat{\beta}^\top X - \hat{\rho}^\top \hat{V}\). Estimate \(m_0(x, v) = \mathbb{E}[e^u \mid X = x, V = v]\) by regressing \(e^{\hat{u}}\) on \((X, \hat{V})\) using a neural network, yielding \(\hat{m}(x, v)\).
Compute the marginal integration \[ \hat{\mu}(x) = \frac{1}{n} \sum_{j=1}^n \hat{m}(x, \hat{V}_j), \qquad \nabla_x \hat{\mu}(x) = \frac{1}{n} \sum_{j=1}^n \nabla_x \hat{m}(x, \hat{V}_j), \] where \(\nabla_x \hat{m}\) is obtained via automatic differentiation.
Estimate \(\omega_0(x, v) = f_V(v) / f_{V \mid X}(v \mid x)\) using density ratio estimation or by separately estimating \(f_V\) and \(f_{V \mid X}\). Estimate \(S_X(x) = \nabla_x \log f_X(x)\) via score matching. Estimate \(\lambda_0(z)\) using the automatic debiasing procedure of chernozhukovAutomaticDebiasedMachine2022, which learns the Riesz representer by minimizing the expected squared norm of the correction term subject to the moment constraint.
Use sample splitting with \(K\) folds. Define \(\hat{\theta}\) as the solution to \[ \frac{1}{n} \sum_{i=1}^n \phi\left(W_i; \theta, \hat{\eta}_{-k(i)}\right) = 0, \] where \(\hat{\eta}_{-k(i)}\) denotes nuisance estimates trained on folds excluding observation \(i\).
\begingroup
\addtocounter{theorem}{-1}
\endgroup
\begingroup
\addtocounter{theorem}{-1}
\endgroup
\begingroup
\addtocounter{theorem}{-1}
\endgroup
\begingroup
\addtocounter{theorem}{-1}
\endgroup
\begingroup
\addtocounter{theorem}{-1}
\endgroup
Let \[ \phi_0(W) := \phi(W; \theta_0, \eta_0), \qquad V := \mathbb{E}[\phi_0(W)^2]. \]
A consistent variance estimator is \[ \hat{V} := \frac{1}{n} \sum_{i=1}^n \hat{\phi}_i^2, \qquad \hat{\phi}_i := \phi(W_i; \hat{\theta}, \hat{\eta}_{-k(i)}), \] and confidence intervals can be constructed as \(\hat{\theta} \pm z_{1-\alpha/2} \sqrt{\hat{V}/n}\).
In this section we present empirical evidence from a re-estimation exercise. We conducted a census of all articles published in the “Top 5” economics journals in the year 2020 (the top 5 journals being Econometrica, Quarterly Journal of Economics, Journal of Political Economy, American Economic Review, Review of Economic Studies). Our census revealed that 36.6% of the 416 articles published in these 5 journals in 2020 contained regressions with log-dependent variables. Out of the papers that employed log-DV regressions, approximately 70% of them interpreted their coefficients as percentage changes of the arithmetic mean. We also conducted a census of articles from the first two issues published in 2020 of a selection top field journals. From this census of 113 papers, we found that 53% of papers ran log-dependent variable regressions. Figure (ref) presents the results from this census.
After identifying articles published in the Top 5 journals that used log dependent variable regressions we collected the replication package and associated datasets for these papers. The inclusion criteria for replicating were as follows: (0) in our list of top 5 journal articles published in 2020 with a log-DV regression, (1) publicly available or easily available data, (2) replication package exists, (3) not structural estimations. After the data was collected and results of log-linear or log-log regressions were correctly re-estimated, we re-estimated results that were not estimated using an instrumental variable identification strategy.
With these exclusion criteria for non-IV results, we were left with 60 results from 29 papers. For each result we estimated the OLS coefficient from the log-linear OLS model, the PPML coefficient from the exponential model, and the doubly robust elasticity of the arithmetic mean (DREAM) estimator which is consistent for the average semi-elasticity of \(\mathbb{E}[y_{i}| x_{i}]\). Our re-estimation exercise shows 3 facts: (1) PPML and OLS estimates are statistically significantly different from each other in 27% of the results . If there was full independence between \(x_i\) and the error term \(v_i\) in the log-linear model, then the coefficient \(\beta\) would equal the exponential model coefficient \(\gamma\), because the arithmetic and geometric elasticities would be equal. (2) OLS and DREAM report statistically significantly different estimates 32% of the time. Of the 19 results that are statistically significantly different from the OLS results, the median absolute percent difference was 109%, and the 25th percentile was 67%. (3) PPML and DREAM report statistically significantly different estimates 30% of the time. Of the 18 statistically significantly different results, the median absolute percent difference was 75%, and the 25th percentile was 54%.\footnote{To appropriately compare elasticity or semi-elasticity estimates, we need to treat discrete treatment variables of interest differently from continuous ones. Specifically for continuous treatment variables, we can run a direct test of the OLS coefficient against the doubly-robust semi-elasticity \(\psi\), in this case the null is \(H_0: \beta_{log-lin}=\psi\) . For discrete treatment variables, the appropriate semi-elasticity comparison is \(H_0:e^{\beta_{log-lin}}-1=\psi\) . This holds for elasticities too, and works the same for exponential coefficients estimated with PPML.}
Figure (ref) shows the results for the difference between OLS and PPML. If there is full independence between the regressors and the error terms then the OLS coefficient and the PPML coefficient will converge to the same constant. We recognize that this exercise of comparing OLS and PPML coefficients is not a definitive test for independence or reliability of one model over another. We do this to demonstrate that OLS and PPML coefficients are statistically significantly different in many settings, and therefore we encourage researchers to consider that heterogeneity may be a reason for why we observe significant differences between these two estimates. Here we may be comparing results from two mutually exclusive models. Therefore, we encourage readers to only take this as an illustrative example of the differences of results coming from two estimation strategies of corresponding (but often conflicting) regression models.
Differences in elasticity estimates obtained from PPML and OLS estimation strategies motivate us to consider using our specification-robust Neyman Orthogonal estimator. We first compare elasticities obtained from OLS estimation with our semi-elasticity estimates from DREAM. Figure Figure (ref) presents these results. Here, the OLS elasticity estimate is plotted on the x-axis, and the DREAM elasticity estimate is plotted on the y-axis. The parity line is presented for reference. The red triangles represent estimates where OLS is statistically significantly different from DREAM, and the black squares represent those where the difference is not statistically significant. Note, any dot in the second or fourth quadrant represents a sign reversal -meaning a positive semi-elasticity was estimated with OLS, and a negative one was estimated with DREAM, or vice versa. Of the 50 estimations, 18 of the estimates obtained using OLS are statistically significantly different from those obtained using DREAM. Of these 18 that are statistically significantly different, the median absolute bias is 96.4% of the OLS coefficient. If we restrict results to those where the absolute value of the average elasticity estimate is less than 5, then we have 42 total results, 16 of which are statistically significantly different from the OLS estimator. Of these 16 results, the median absolute bias is 64.97% of the OLS coefficient. Our preferred bias estimate removes results with absolute DREAM estimates greater than 5, as these extreme values likely reflect instability in the machine learning nuisance parameter estimation.
For reference, we also compare elasticity estimates from PPML with those from DREAM. While no papers in our sample used PPML. Figure (ref) presents these results. Here the x-axis is the PPML estimate and the y-axis is the DREAM estimate. PPML and DREAM report statistically significantly different estimates 38% of the time. There are two statistically significant sign reversals. As in Figure (ref), we can see that standard errors for this test matter as well. Some estimates sit close to the parity line but are statistically significantly different from each other, while others are relatively distant from the parity line but are not statistically significantly different.
Table (ref) provides an aggregate summary of the differences between the OLS estimator and DREAM, and the PPML estimator and DREAM. Of the 50 results we test, 32 have no statistically significant difference with OLS. Of those that are statistically significantly different, 8 have an increase in estimate with DREAM (in absolute terms), and 6 have a decrease (in absolute terms). Four estimates reverse sign between OLS and DREAM. The comparison between DREAM and PPML is similar. There are 31 estimates that are not statistically significantly different, and 19 that are. Of the statistically significantly different, 13 increase in absolute value with DREAM, 4 decrease in absolute value, and 2 reverse signs.
Figure (ref) present a direct comparison between the DREAM, PPML, and OLS for the same result. The graph on the left provides the point estimate and standard errors for estimates of the difference in elasticities estimated from DREAM and OLS. These are ordered in descending order, from most positive to most negative difference. Note, you can't observe sign flips on this graph because the points are just the differences in estimates. The results on the right-hand panel correspond to those on the left. For example, the top result shown for the Elasticity - OLS panel is the same result estimated in the top of the Elasticity - PPML panel. Here we can see that sometimes DREAM is statistically significantly different from both estimates, while other times it is only statistically significantly different from one or no estimates. DREAM disagrees with both PPML and OLS in 11 out of the 50 regressions we re-estimated (9 out of 42 if we exclude absolute DREAM estimates greater than 5). DREAM is not statistically significantly different from OLS and PPML in 24 out of the 50 results (or 18 out of 42 in the restricted sample). DREAM is significantly different from OLS and not from PPML in 7 results (no difference with restricted sample). DREAM is significantly different from PPML and not from OLS in 8 results (again, no difference with restricted sample).
In total, this re-estimation of results from papers published in Top 5 journals in the year 2020 provide several takeaways. First, we show that PPML and OLS estimates are frequently statistically significantly different from each other, providing some suggestive evidence that dependence may exist between the covariates and the error terms. Second, we show that our estimator, which is consistent for the true semi-elasticity, disagrees with OLS 38% of the time, and disagrees with PPML about 36%. We see sign flips in 8% of re-estimated results when comparing DREAM to OLS, and 4% when comparing DREAM to PPML. Table (ref) shows that the differences between OLS and DREAM, and PPML and DREAM estimates are remarkably consistent. When directly comparing results from papers, we see that there are a mix of outcomes, sometimes DREAM disagrees with both PPML and OLS in 22% of estimations, it disagrees with neither in 48% of estimations, and disagrees with one but not the other in 30% of estimations.
Similarly to the OLS and DREAM case we re-estimate 19 results from 12 papers using the DREAM with control functions approach. Of these, 13 were significantly different at the 5% level from the naive 2SLS estimate. Six papers showed no difference when estimated using 2SLS or IV-DREAM. Eight showed reduced effect size, of which five exhibited sign reversal. Five showed increased effect size.
Among the statistically significant differences at the 5% level, the median absolute change was about 71% of the original estimate.
Economists often use log dependent variable regressions to estimate percentage changes. The coefficient estimates are consistent for the geometric mean (semi-) elasticities. We give several examples where economists specify estimands of arithmetic mean elasticities. We show that arithmetic and geometric mean elasticities can be very different when individual elasticities are heterogenous. We present them as a sub-class of power-mean elasticities, and show that their difference can be characterized using a weighted aggregation of the heterogenous responses. We justify the estimation of power-mean elasticities by appeal to a decision problem axiomatisation. We provide an estimator for the average wedge between arithmetic and geometric mean elasticities. Using this estimator, we replicate 50 papers, and find a median difference of 67%.
Future work could focus on eliciting the whole elasticity distribution, clarifying the relationship between aggregation bias and retransformation bias and further specifying which economic situations demand which estimands.
Future drafts of this paper will include examples with randomised controlled trials and discrete variables, for which we have the maths but not yet the words.