Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
58,189 characters · 12 sections · 42 citation commands
Estimation and Inference on Average Treatment Effect in Percentage Points under Heterogeneity
Semi-logarithmic regression models, where the dependent variable is in the natural logarithm form, are widely used in empirical studies. In these models, the coefficient of a binary treatment is often interpreted as approximating the average treatment effect (ATE) in percentage change when its magnitude is small chen_logs_2024,hansen_econometrics_2022. This interpretation is based on the following logic: Let $Y_{1}$ and $Y_{0}$ denote the potential outcomes under treatment and non-treatment respectively. The treatment effect in log points is defined as $\tau=\ln(Y_{1})-\ln(Y_{0})$, while the percentage change is given by $\rho=\left(Y_{1}-Y_{0}\right)/Y_{0}=\operatorname{exp}(\tau)-1$. When $|\tau|$ is small, $\operatorname{exp}(\tau)-1\approx\tau$, justifying the interpretation of $\tau$ as an approximate percentage effect.
When treatment effects are heterogeneous, however, neither the ATE in log points, $\bar{\tau}$, nor $\operatorname{exp}(\bar{\tau})-1$ can be interpreted as the ATE in percentage points $\bar{\rho}$. The intuition is straightforward: Jensen's inequality implies $E[\operatorname{exp}(\tau)]-1\geqslant\operatorname{exp}[E(\tau)]-1$, hence $\bar{\rho}\geqslant\operatorname{exp}(\bar{\tau})-1$, with equality holding if and only if log-point treatment effects are constant. Consequently, $\operatorname{exp}(\bar{\tau})-1<\bar{\rho}$ under treatment effect heterogeneity, with disparities increasing as treatment effects exhibit greater variation.
The bias that arises when the ATE in log points $\bar{\tau}$ or $\operatorname{exp}(\bar{\tau})-1$ is interpreted as the ATE in percentage points has received limited attention in the literature, except for chen_logs_2024 who show that $\bar{\rho}$ cannot be point identified if arbitrary heterogeneity is allowed. This oversight may lead to misleading interpretations of results in empirical studies with heterogeneous treatment effects and log-transformed outcomes. A prominent example is the semi-log difference-in-differences model with staggered treatment adoption, where different groups start receiving treatment at different times. When treatment effects in log-transformed outcomes are heterogeneous across groups or over time, conventional two-way fixed effects (TWFE) estimators assuming a constant effect identify weighted averages of log-point effects $\tilde{\tau}=\sum_{g}\tilde{w}_{g}\tau_{g}$, where $\tau_{g}$ is the average log-point effect for cohort-time group $g$, and $\tilde{w}_{g}$ is generally not equal to the true weight of group $g$ in the treated population and possibly negative dechaisemartin_twoway_2020,goodman-bacon_differenceindifferences_2021,sun_estimating_2021. Hence, in general $\tilde{\tau}\neq\bar{\tau}$. To address this issue, numerous heterogeneity-robust estimators have been proposed (dechaisemartin_twoway_2020,callaway_differenceindifferences_2021,sun_estimating_2021,borusyak_revisiting_2024, etc.). When outcome variables are log-transformed, these estimators typically represent some form of the ATE in log points $\bar{\tau}$. Consequently, these estimators and their exponential minus one should not be interpreted as the ATE in percentage points.
This paper highlights the often-overlooked problem of interpreting the ATE in log points as the ATE in percentage points in the context of heterogeneous treatment effects. Unfortunately, $\bar{\rho}$ cannot be point identified without additional distributional assumptions because $E\left[\left(Y_{1}-Y_{0}\right)/Y_{0}\right]$ is a functional of the joint distribution of $(Y_{1},Y_{0})$. Since $Y_{1}$ and $Y_{0}$ are not simultaneously observed for the same unit, only their marginal distributions are identified, and such functionals generally cannot be point identified callaway_bounds_2021,fan_partial_2017,heckman_making_1997. In particular, Proposition 3 of chen_logs_2024 provides a formal proof that $\bar{\rho}$ is not point identified when heterogeneity is unrestricted. A growing literature develops partial identification approaches for such functionals, including callaway_bounds_2021, fan_partial_2017, firpo_partial_2019, and frandsen_partial_2021. Building on the framework of fan_partial_2017, this paper derives sharp bounds for $\bar{\rho}$ under the assumption that $(Y_{1},Y_{0})$ is jointly independent of treatment conditional on observables. However, estimating these sharp bounds requires nonparametric estimation of the full conditional distributions of $Y_{1}$ and $Y_{0}$, which can be computationally demanding. Point identification of $\bar{\rho}$ under heterogeneity is possible only under additional distributional assumptions. For example, the Online Appendix discusses estimation and inference of $\bar{\rho}$ under joint normality of the potential outcomes, an assumption that can be restrictive in practice.
Given the inherent difficulty of point identifying $\bar{\rho}$ under unrestricted heterogeneity, this paper proposes estimation and inference methods for $\rho_{b}=\sum_{g}w_{g}\left[\operatorname{exp}(\tau_{g})-1\right]$, a subgroup-weighted transformation of average log-point treatment effects. By accounting for heterogeneity across observable subgroups, this estimand improves upon the conventional measures $\operatorname{exp}(\bar{\tau})-1$ and $\bar{\tau}$ and remains point identified. When treatment effects are constant within subgroups, $\rho_{b}$ equals $\bar{\rho}$. When treatment effects are heterogeneous within subgroups, $\rho_{b}$ remains a valid lower bound for $\bar{\rho}$ that is strictly tighter than both $\operatorname{exp}(\bar{\tau})-1$ and $\bar{\tau}$.
The proposed estimation and inference methods for $\rho_{b}$ do not require strong distributional assumptions on the potential outcomes and are straightforward to implement. They are developed under a general econometric framework that encompasses semi-log regression models and semi-log staggered difference-in-differences designs. Their large-sample validity is established under standard regularity conditions, and their finite-sample performance is demonstrated through comprehensive Monte Carlo simulations. To illustrate these methods' practical relevance and applicability, the paper presents two empirical applications: one employing a semi-log regression model to study the causal impact of exporting on firm productivity, and another utilizing a staggered difference-in-differences design to examine the effect of water and sewerage systems on child mortality. Through these contributions, this paper provides researchers with robust tools to more accurately estimate and interpret treatment effects in percentage terms, particularly in the presence of heterogeneity.
In addition to the literature related to partial identification of $\bar{\rho}$, this paper also connects to two other strands of the literature. First, it contributes to a growing body of work on treatment effect heterogeneity, which has shown that conventional estimates assuming a constant treatment effect are weighted averages of treatment effects, with weights possibly negative and not equal to the true population share. Examples include ordinary least squares (OLS) estimators angrist_estimating_1998,gibbons_broken_2019,sloczynski_interpreting_2022,goldsmith-pinkham_contamination_2024, two-stage least squares (2SLS) estimators imbens_identification_1994,mogstad_causal_2021, and two-way fixed effects estimators for staggered difference-in-differences models goodman-bacon_differenceindifferences_2021,sun_estimating_2021,dechaisemartin_twoway_2020,borusyak_revisiting_2024. Numerous heterogeneity-robust estimators have been developed, typically by estimating treatment effects and population shares for subgroups, and then aggregating them up to get the ATE (e.g., gibbons_broken_2019,callaway_differenceindifferences_2021,sun_estimating_2021,goldsmith-pinkham_contamination_2024). The results presented in this paper imply that when the outcome is log-transformed, neither the traditional estimators nor the heterogeneity-robust estimators in these models provide proper approximations of the ATE in percentage points. However, the heterogeneity-robust estimators in these papers form the basis for the estimation and inference methods proposed in the paper.
Second, this work supplements ongoing research into the potential pitfalls of using logarithmic transformations in empirical analyses. For instance, chen_logs_2024 and mullahy_why_2024 demonstrated that when dependent variables have many zero values, coefficients in regressions with log-like transformations such as $\ln(Y+1)$ do not bear the interpretation of percentage effects. manning_logged_1998 and silva_log_2006 highlighted issues of using log-linear regressions for estimating the impact on the outcome's mean and elasticity under heteroscedasticity. The results in this paper further demonstrate that in the presence of treatment effect heterogeneity, coefficients in semi-log regressions may not have an ATE in percentage points interpretation.
The remainder of the paper is structured as follows. Section (ref) develops the framework for analyzing ATE in percentage points under treatment effect heterogeneity. It first characterizes the relationship between log-point and percentage-point effects and the bias induced by Jensen's inequality. It then proposes estimation and inference methods for a group-aggregated estimand, $\rho_{b}$, which accommodates heterogeneity across observable subgroups. Finally, it derives sharp bounds for the average proportional effect under partial identification. Section (ref) reports the results of Monte Carlo simulations for the proposed methods, while Section (ref) presents two empirical applications. Section (ref) concludes. The Online Appendix provides detailed discussions of examples of the proposed methods, discusses estimation and inference of $\bar{\rho}$ under joint normality of potential outcomes, introduces a simple bias-correction method, and presents additional Monte Carlo results.
This section establishes the formal relationship between the ATE in log points and the ATE in percentage points. The setup allows log-point treatment effects to be heterogeneous both across and within groups, commonly referred to in the literature as systematic and idiosyncratic heterogeneity, respectively djebbari_heterogeneous_2008.
Consider a population partitioned into $G$ mutually exclusive and collectively exhaustive subgroups, indexed by $g=1,\ldots,G$. Let $w_{g}$ denote the population share of subgroup $g$, where $w_{g}>0$ and $\sum_{g=1}^{G}w_{g}=1$. Let $D_{i}^{(g)}=1$ if individual $i$ belongs to subgroup $g$, and 0 otherwise. Let $Y_{0i}$ and $Y_{1i}$ represent the potential outcomes for individual $i$ if not treated and if treated respectively. I assume that $Y_{0i}>0$ and $Y_{1i}>0$ for all $i$, so that their natural logarithms are well defined.
The individual treatment effect in log points is defined as $\tau_{i}=\ln\left(Y_{1i}\right)-\ln\left(Y_{0i}\right)$, and the ATE in log points for group $g$ is $\tau_{g}=E(\tau_{i}|D_{i}^{(g)}=1)$. The population ATE in log points is
The individual treatment effect in percentage points is $\rho_{i}=\left(Y_{1i}-Y_{0i}\right)/Y_{0i}=\operatorname{exp}(\tau_{i})-1$, and the ATE in percentage points for group $g$ is $\rho_{g}=E(\rho_{i}|D_{i}^{(g)}=1)=E(\operatorname{exp}(\tau_{i})|D_{i}^{(g)}=1)-1$. The population ATE in percentage points is
The identity $\rho_{i}=\operatorname{exp}(\tau_{i})-1$ motivates the common practice of reporting \[ \rho_{a}=\operatorname{exp}(\bar{\tau})-1 \] as the ATE in percentage points. Alternatively, when $\bar{\tau}$ is close to zero, $\bar{\tau}$ itself is often used as an approximation to $\bar{\rho}$, since $\operatorname{exp}(x)-1\approx x$ for small $x$. More generally, it holds that $\operatorname{exp}(\bar{\tau})-1\geqslant\bar{\tau}$, with equality holding if and only if $\bar{\tau}=0$, because the function $\text{\ensuremath{\operatorname{exp}}}(x)-x-1$ is strictly convex and attains its unique minimum of zero at $x=0$.
However, Jensen's inequality implies $E[\operatorname{exp}(\tau_{i})]\geqslant\operatorname{exp}[E(\tau_{i})]$ and hence $\bar{\rho}\geqslant\rho_{a}$, with equality holding if and only if $\tau_{i}=\bar{\tau}$ for all $i$. Under treatment effect heterogeneity, $\bar{\rho}>\rho_{a}$. The gap can arise from within-group heterogeneity or from cross-group heterogeneity. First, due to within-group heterogeneity, for $g=1,\ldots,G$, we have $E\left[\operatorname{exp}(\tau_{i})|D_{i}^{(g)}=1\right]\geqslant\operatorname{exp}\left[E\left(\tau_{i}|D_{i}^{(g)}=1\right)\right]=\operatorname{exp}(\tau_{g})$, which implies that $\rho_{g}\geqslant\operatorname{exp}(\tau_{g})-1$. Define
Then $\bar{\rho}\geqslant\rho_{b}$, with equality holding if and only if $\tau_{i}$ is constant within each subgroup. Second, due to heterogeneity across groups, $\sum_{g=1}^{G}w_{g}\operatorname{exp}(\tau_{g})\geqslant\operatorname{exp}\left(\sum_{g=1}^{G}w_{g}\tau_{g}\right)$, which implies $\rho_{b}\geqslant\operatorname{exp}(\bar{\tau})-1$, with equality holding if and only if $\tau_{g}=\bar{\tau}$ for all $g$. These results are summarized in the following proposition.
Thus, under either systematic or idiosyncratic heterogeneity, neither the average log-point effect $\bar{\tau}$ nor $\operatorname{exp}(\bar{\tau})-1$ equals the average proportional effect $\bar{\rho}$, although $\operatorname{exp}(\bar{\tau})-1$ is generally less biased than $\bar{\tau}$.
Identifying $\bar{\rho}$ is challenging because it is a functional of the joint distribution of potential outcomes; see the Introduction for a discussion of the related literature. Using a partial identification approach developed by fan_partial_2017, Section (ref) derives sharp upper and lower bounds for $\bar{\rho}$. However, this partial identification approach relies on nonparametric estimation of the full conditional distributions of $Y_{0i}$ and $Y_{1i}$, which entails substantial computational complexity. In Section (ref) of the Online Appendix, I also discuss the estimation and inference of $\bar{\rho}$ under the assumption that $(Y_{0},Y_{1})$ is jointly normal. Although this approach delivers point identification of $\bar{\rho}$ under both systematic and idiosyncratic heterogeneity, the assumption of joint normality of potential outcomes can be too restrictive in practice and limits its applicability.
Because idiosyncratic heterogeneity is inherently difficult to characterize, this paper focuses on $\rho_{b}$, which captures heterogeneity across groups. When treatment effects are constant within each subgroup, the average proportional effect $\bar{\rho}$ is point identified and coincides with $\rho_{b}$. When within-group heterogeneity is present, Proposition (ref) implies that $\rho_{b}$ remains a valid lower bound for $\bar{\rho}$, and this lower bound is sharper than $\operatorname{exp}(\bar{\tau})-1$ or $\bar{\tau}$. The difference between $\rho_{b}$ and $\bar{\rho}$ therefore reflects residual idiosyncratic heterogeneity and can be reduced by refining subgroup definitions. In empirical applications, researchers often define subgroups based on observed covariates to capture systematic variation in treatment effects. Expanding the set of covariates typically reduces the proportion of idiosyncratic heterogeneity relative to total heterogeneity buhl-wiggers_children_2024,smith_treatment_2022. Recent advances in machine learning methods for heterogeneous treatment effect estimation offer the potential to capture more systematic heterogeneity, thereby bringing $\rho_{b}$ closer to $\bar{\rho}$.
Accordingly, this paper develops estimation and inference methods for $\rho_{b}$, as discussed in Section (ref). Compared to partial identification approaches and methods relying on joint normality assumptions, the proposed methods are straightforward to implement and rely on standard assumptions, thus providing a practical approach to analyzing treatment effects in percentage points under heterogeneity.
This section proposes a general econometric framework for the estimation and inference of $\rho_{b}$ in Eq. ((ref)), building upon consistent and asymptotically normal estimators of $w=(w_{1},\ldots,w_{G})^{\prime}$ and $\tau=(\tau_{1},\ldots,\tau_{G})^{\prime}$. The proposed methodology imposes no restrictions on the functional form or application context of these estimators, provided they meet the assumptions below.
I impose the following two assumptions on the parameters $w$ and $\tau$.
Without loss of generality, I restrict the population share $w_{g}$ to be strictly positive for all $g$, excluding groups with zero share from the analysis. I allow $w_{g}=1$ to incorporate the case of $G=1$, representing no cross-group heterogeneity.
Assumption (ref) bounds the value of $\tau_{g}$ and hence $\operatorname{exp}(\tau_{g})-1$. Combining Assumptions (ref) and (ref) yields $-C_{\tau}\leqslant\bar{\tau}\leqslant C_{\tau}$ and $\operatorname{exp}(-C_{\tau})-1\leqslant\rho_{b}\leqslant\operatorname{exp}(C_{\tau})-1$.
Let $\hat{w}=(\hat{w}_{1},\ldots,\hat{w}_{G})^{\prime}$ and $\hat{\tau}=(\hat{\tau}_{1},\ldots,\hat{\tau}_{G})^{\prime}$ denote estimators of $w$ and $\tau$ respectively.
Assumption (ref) accommodates both known and estimated weights. In the special case where $w$ is known, we have $\hat{w}=w$. The asymptotic variance of $\sqrt{N}(\hat{w}-w)$ is therefore $\bar{\Sigma}_{w}=0$ and $\operatorname{Cov}(\hat{w},\hat{\tau})=0$. This scenario arises naturally in several contexts, such as when analyzing a fixed population (e.g., the 50 U.S. states) divided into policy-relevant subgroups, or in staggered difference-in-differences designs where treatment effects for a given cohort are averaged across periods with equal and predetermined weights. Under known weights, $\bar{\Sigma}_{\delta}=\operatorname{diag}\{0_{G\times G},\bar{\Sigma}_{\tau}\}$ and $\hat{\bar{\Sigma}}_{\delta}=\operatorname{diag}\{0_{G\times G},\hat{\bar{\Sigma}}_{\tau}\}$, where $\bar{\Sigma}_{\tau}$ is the asymptotic variance of $\sqrt{N}(\hat{\tau}-\tau)$ and $\hat{\bar{\Sigma}}_{\tau}$ is its estimator. Assumption (ref) then reduces to $\sqrt{N}(\hat{\tau}-\tau)\xrightarrow{d}\mathcal{N}(0_{G\times1},\bar{\Sigma}_{\tau})$ and $\hat{\bar{\Sigma}}_{\tau}\xrightarrow{p}\bar{\Sigma}_{\tau}$. In contrast, when working with a random sample from a population, the estimated group weights are subject to sampling variability so that $\bar{\Sigma}_{w}\neq0$.
Given the consistency of $\hat{w}$ and $\hat{\tau}$ under Assumption (ref), a natural estimator for $\rho_{b}=\sum_{g=1}^{G}w_{g}\operatorname{exp}(\tau_{g})-1$ is:
By the continuous mapping theorem, $\hat{\rho}_{b}\xrightarrow{p}\rho_{b}$, where $\rho_{b}$ is a lower bound for $\bar{\rho}$ according to Proposition (ref).
To conduct inference, I derive the asymptotic distribution of $\hat{\rho}_{b}$ using the delta method. Note that $\partial\rho_{b}/\partial w_{g}=\operatorname{exp}(\tau_{g})$ and $\partial\rho_{b}/\partial\tau_{g}=w_{g}\operatorname{exp}(\tau_{g}),$ so the gradient of $\rho_{b}$ with respect to $\delta=(w^{\prime},\tau^{\prime})^{\prime}$ is \[ \nabla_{\delta}\rho_{b}=\left[
\right], \] where $\operatorname{exp}(\tau)=(\operatorname{exp}(\tau_{1}),\ldots,\operatorname{exp}(\tau_{G}))^{\prime}$, and $\odot$ denotes the element-wise product. With $\sqrt{N}\left(\hat{\delta}-\delta\right)\xrightarrow{d}\mathcal{N}(0,\bar{\Sigma}_{\delta})$, the delta method yields $\sqrt{N}(\hat{\rho}_{b}-\rho_{b})\xrightarrow{d}\mathcal{N}(0,\bar{\sigma}_{b}^{2})$, where
By the continuous mapping theorem, a consistent estimator of $\bar{\sigma}_{b}^{2}$ is
When weights are known, $\bar{\Sigma}_{\delta}=\operatorname{diag}\{0_{G\times G},\bar{\Sigma}_{\tau}\}$ and the asymptotic variance simplifies to $\bar{\sigma}_{b}^{2}=\left(w\odot\operatorname{exp}(\tau)\right)^{\prime}\bar{\Sigma}_{\tau}\left(w\odot\operatorname{exp}(\tau)\right)$.
The proof follows directly from the previous discussions. This theorem establishes the basis for hypothesis testing and confidence interval construction. To test $H_{0}:\rho_{b}=\rho_{0}$, we can use the $z$-statistic $z_{\rho}=\left(\hat{\rho}_{b}-\rho_{0}\right)/\sigma_{b}$, where $\sigma_{b}=\bar{\sigma}_{b}/\sqrt{N}$ denotes the asymptotic standard error of $\hat{\rho}_{b}$, and is usually estimated by $\hat{\bar{\sigma}}_{b}/\sqrt{N}$ in practice.
In many applications, the total sample size $N$ is inherently fixed (e.g., analyzing all 50 U.S. states), where the standard asymptotic framework with $N\rightarrow\infty$ does not apply. This is a limitation of the proposed method, as it requires sufficient observations in each subgroup for reliable estimation. With more subgroups, this requirement becomes increasingly difficult to meet. However, if each subgroup contains enough observations for approximate normality of both $\hat{w}_{g}$ and $\hat{\tau}_{g}$, the delta method may still provide adequate finite-sample approximations. In practice, group sizes are often sufficient for this purpose. Monte Carlo simulations in Section (ref) confirm that the proposed estimation and inference procedures perform well with modest sample sizes. For applications with particularly small group sizes, the bias-corrected estimators in Section (ref) of the Online Appendix offer additional improvements.
The estimation and inference methods developed in this section apply broadly, as Assumptions (ref)-(ref) hold in many commonly used empirical settings. One example is the semi-log regression model. Suppose we are interested in the ATE in percentage points for the treatment group, which consists of $G$ sub-treatment groups. Consider the model
where $D_{i}=(D_{i}^{(1)},\ldots,D_{i}^{(G)})^{\prime}$ is a vector of dummies for the sub-treatment groups. If the conditional independence assumption holds, then $\tau_{g}$ is the ATE in log points for sub-treatment group $g$. Under standard regularity conditions, the OLS estimator $\hat{\tau}$ satisfies $\sqrt{N}\left(\hat{\tau}-\tau\right)\xrightarrow{d}\mathcal{N}(0_{G\times1},\bar{\Sigma}_{\tau})$ for some finite matrix $\bar{\Sigma}_{\tau}$, with the heteroskedasticity robust variance estimator $\hat{\bar{\Sigma}}_{\tau}$ satisfying $\hat{\bar{\Sigma}}_{\tau}\xrightarrow{p}\bar{\Sigma}_{\tau}$ (see, e.g., hayashi_econometrics_2000). For the estimation of $w$, let $T_{i}=\sum_{g=1}^{G}D_{i}^{(g)}$ be the treatment indicator, and suppose the probability of treatment is $p_{T}=E(T_{i})\in(0,1)$, and define $w_{g}=E(D_{i}^{(g)}=1|T_{i}=1)\in(0,1)$. Under random sampling, a natural and consistent estimator of $w_{g}$ is $\hat{w}_{g}=N_{g}/N_{T}$, where $N_{g}=\sum_{i=1}^{N}D_{i}^{(g)}$ is the sample size of subgroup $g$ and $N_{T}=\sum_{g=1}^{G}N_{g}$ is the total number of treated observations. Since $(N_{1},\ldots,N_{G})\sim Multinomial(N_{T},w)$ and $N_{T}/N\rightarrow p_{T}$, it follows that $\sqrt{N}\left(\hat{w}-w\right)\xrightarrow{d}\mathcal{N}(0_{G\times1},\bar{\Sigma}_{w})$, where $\bar{\Sigma}_{w}=p_{T}^{-1}\left(\operatorname{diag}(w)-ww^{\prime}\right)$ and its consistent estimator can be obtained by replacing $w$ with $\hat{w}$ and $p_{T}$ with $\hat{p}_{T}=N_{T}/N$ respectively. Moreover, it can be shown that $\left(\sqrt{N}(\hat{\tau}-\tau),\sqrt{N}(\hat{w}-w)\right)$ is jointly asymptotically normal with zero covariance. Consequently, $\sqrt{N}\left(\hat{\delta}-\delta\right)\xrightarrow{d}\mathcal{N}\left(0_{2G\times1},\bar{\Sigma}_{\delta}\right)$, where $\bar{\Sigma}_{\delta}=\operatorname{diag}\{\bar{\Sigma}_{\tau},\bar{\Sigma}_{w}\}$, so that Assumption (ref) holds in this setting. See Section (ref) of the Online Appendix for further details.
Another example is staggered difference-in-differences designs. Heterogeneity-robust estimators in these designs typically provide consistent and asymptotically normal estimates of subgroup-specific treatment effects and weights, thereby satisfying Assumption (ref). \footnote{See, for example, Theorems 2 and 3 in callaway_differenceindifferences_2021 and Proposition 6 in sun_estimating_2021.} The estimation and inference procedures for $\rho_{b}$ developed here therefore apply. Example 2 in Section (ref) of the Online Appendix illustrates this application to a semi-log difference-in-differences model with staggered adoption, which is a direct extension of model ((ref)) when subgroups are defined as cohort-time combinations.
This section applies results from fan_partial_2017 to develop sharp bounds for $\bar{\rho}$, accounting for unrestricted heterogeneity. The analysis follows directly from Theorem 3.2 of fan_partial_2017.
Assume that the two potential outcomes $Y_{1},Y_{0}\in(0,\infty)$ with $E(Y_{1})<\infty$. Observe that the function $\mu(Y_{1},Y_{0})=(Y_{1}-Y_{0})/Y_{0}$ is right-continuous and strictly sub-modular as for $Y_{1}^{\prime}>Y_{1}$ and $Y_{0}^{\prime}>Y_{0}$, we have \[ \frac{Y_{1}}{Y_{0}}+\frac{Y_{1}^{\prime}}{Y_{0}^{\prime}}-\frac{Y_{1}^{\prime}}{Y_{0}}-\frac{Y_{1}}{Y_{0}^{\prime}}=\frac{Y_{1}-Y_{1}^{\prime}}{Y_{0}}+\frac{Y_{1}^{\prime}-Y_{1}}{Y_{0}^{\prime}}=-\frac{(Y_{1}^{\prime}-Y_{1})(Y_{0}^{\prime}-Y_{0})}{Y_{0}Y_{0}^{\prime}}<0. \] With $E\mu(Y_{1},\bar{y}_{0})$ finite for any $\bar{y}_{0}>0$ and $E\mu(0,Y_{0})=-1$ also finite, Theorem 2 of cambanis_inequalities_1976 applies. Let $F_{1o}(Y_{1})$ and $F_{0o}(Y_{0})$ denote the marginal cumulative distribution functions (CDF) of $Y_{1}$ and $Y_{0}$ respectively. The sharp bounds for $\mu(Y_{1},Y_{0})=(Y_{1}-Y_{0})/Y_{0}$ correspond to the Fr�chet-Hoeffding bounds, with the lower bound $\rho_{L}=\int_{0}^{1}\left[F_{1o}^{-1}(u)/F_{0o}^{-1}(u)\right]du-1,$ and the upper bound $\rho_{U}=\int_{0}^{1}\left[F_{1o}^{-1}(1-u)/F_{0o}^{-1}(u)\right]du-1$, provided both integrals exist and at least one is finite.
The bounds can be tightened under the conditional independence assumption. Let $X$ denote observed covariates with support $\mathcal{X}\subset\mathcal{R}^{k_{x}}$. Under the assumption that $Y_{1},Y_{0}$ are jointly independent of treatment $T$ conditional on $X$ and that $0<E(T=1|X=x)<1$ for all $x\in\mathcal{X}$, we can identify the conditional marginal distribution of $Y_{1}$ and $Y_{0}$, denoted as $F_{1o}(Y_{1}|X)$ and $F_{0o}(Y_{0}|X)$ respectively. Following Theorem 3.2 of fan_partial_2017, the sharp lower and upper bounds of $E\left[(Y_{1}-Y_{0})/Y_{0}\right]$ become \[ \rho_{L}^{\ast}=E\left[\int_{0}^{1}\frac{F_{1o}^{-1}(u|X)}{F_{0o}^{-1}(u|X)}du\right]-1, \] and
respectively, if $\rho_{L}^{\ast}$ and $\rho_{U}^{\ast}$ both exist and at least one of them is finite.
This section reports Monte Carlo experiments that study the finite-sample performance of the proposed estimator $\hat{\rho}_{b}$ and the associated inference procedure in semi-log regression models. The simulation design follows the framework introduced at the end of Section (ref) and described in detail in Section (ref) of the Online Appendix.
I compare the performance of $\hat{\rho}_{b}$ with that of $\hat{\bar{\tau}}$ and $\hat{\rho}_{a}$, where $\hat{\bar{\tau}}=\sum_{g=1}^{G}\hat{w}_{g}\hat{\tau}_{g}$ is the estimator of $\bar{\tau}$, and $\hat{\rho}_{a}=\operatorname{exp}(\hat{\bar{\tau}})-1$ is the estimator of $\rho_{a}$. Under Assumptions 1 to 3, $\hat{\bar{\tau}}\xrightarrow{p}\bar{\tau}$ and $\hat{\rho}_{a}\xrightarrow{p}\rho_{a}$ by the continuous mapping theorem. Standard delta-method calculations yield $\sqrt{N}(\hat{\bar{\tau}}-\bar{\tau})\xrightarrow{d}\mathcal{N}(0,\bar{\sigma}_{\bar{\tau}}^{2})$, where $\bar{\sigma}_{\bar{\tau}}^{2}=\nabla_{\delta}\bar{\tau}^{\prime}\bar{\Sigma}_{\delta}\nabla_{\delta}\bar{\tau}$ and $\nabla_{\delta}\bar{\tau}=(\tau^{\prime},w^{\prime})^{\prime}$. Similarly, $\sqrt{N}(\hat{\rho}_{a}-\rho_{a})\xrightarrow{d}\mathcal{N}(0,\bar{\sigma}_{a}^{2})$, where $\bar{\sigma}_{a}^{2}=\operatorname{exp}(2\bar{\tau})\bar{\sigma}_{\bar{\tau}}^{2}$. To assess the finite-sample performance of the inference method, I report two-sided $z$-tests for $H_{0}:\bar{\tau}=\rho_{0}$ based on $z_{\tau}=(\hat{\bar{\tau}}-\rho_{0})/\sigma_{\bar{\tau}}$, where $\sigma_{\bar{\tau}}=\bar{\sigma}_{\bar{\tau}}/\sqrt{N}$ is estimated by replacing each parameter with its corresponding estimate. Because $\rho_{a}=\operatorname{exp}(\bar{\tau})-1$, testing $H_{0}:\rho_{a}=\rho_{0}$ is equivalent to testing $H_{0}:\bar{\tau}=\ln(\rho_{0}+1)$. The corresponding statistic is $z_{a}=(\hat{\bar{\tau}}-\ln(\rho_{0}+1))/\sigma_{\bar{\tau}}$.
Because $\bar{\rho}\geqslant\rho_{b}\geqslant\rho_{a}\geqslant\bar{\tau}$, differences between $\hat{\rho}_{b}$ and the conventional estimators $\hat{\rho}_{a}$ and $\hat{\bar{\tau}}$ capture the extent to which $\hat{\rho}_{b}$ improves estimation and inference of $\bar{\rho}$, regardless of whether $\bar{\rho}=\rho_{b}$. For ease of interpretation, I impose constant treatment effects within subgroups so that $\rho_{b}=\bar{\rho}$.
The simulations focus on the average treatment effect on the treated (ATT) in percentage points with $G=4$ sub-treatment groups. Data are generated according to the following semi-log model: $\ln Y_{i}=1+X_{i}+\sum_{g=1}^{4}D_{i}^{(g)}\tau_{g}+\epsilon_{i}$, where $D_{i}^{(g)}$ is the dummy for sub-treatment group $g$. The treatment probability is $p_{T}=0.8$, and each sub-treatment group has equal weight $w_{g}=0.25$ within the treated population. Consequently, the control group and each sub-treatment group each comprise $20\%$ of the total population. Observations are randomly assigned to groups according to these probabilities. Sample size $N$ is selected from \{20, 50, 100, 200, 500, 1000, 2000, 5000, $10^{4}$, $10^{5}$, $10^{6}$\}. The covariate $X_{i}$ and the error term $\epsilon_{i}$ are each i.i.d. $\mathcal{N}(0,1)$. Results using skew-normal errors are analogous and presented in Section (ref) of the Online Appendix. I examine two cases: a large heterogeneity case with $\tau=(\ln(0.68),\ln(0.84),\ln(1.16),\ln(1.32))^{\prime}$, implying $100(\rho_{1},\rho_{2},\rho_{3},\rho_{4})=(-32,-16,16,32)$, and a small heterogeneity case with $\tau=(\ln(0.92),\ln(0.96),\ln(1.04),\ln(1.08))^{\prime}$, implying $100(\rho_{1},\rho_{2},\rho_{3},\rho_{4})=(-8,-4,4,8)$. In both cases, the true value of $\bar{\rho}=\rho_{b}$ equals $0\%$. I estimate $\hat{\tau}$, $\hat{w}$, and $\hat{\bar{\Sigma}}_{\delta}$ as described at the end of Section (ref), and then construct $\hat{\bar{\tau}}$, $\hat{\rho}_{a}$ and $\hat{\rho}_{b}$ together with their associated $z$-statistics.
The left and right panels of Table (ref) present results for small and large heterogeneity respectively, with all values multiplied by 100. The first three columns in each panel report the Monte Carlo means and standard deviations (in parentheses) of $\hat{\bar{\tau}}$, $\hat{\rho}_{a}$ and $\hat{\rho}_{b}$ across 100,000 repetitions, with true values provided at the top of each panel. The last two columns in each panel report empirical rejection rates of two-sided $z$-tests for $H_{0}:\bar{\tau}=0$ using $z_{\tau}$ and for $H_{0}:\rho_{b}=0$ using $z_{\rho}$ at the 5% level. The empirical rejection rates equal one minus the coverage rate for $0$ of the confidence intervals for $\bar{\tau}$ and $\bar{\rho}$ respectively. The $z$-test for $\rho_{a}=0$ is omitted, because $\rho_{a}=0$ is equivalent to $\bar{\tau}=0$ given that $\rho_{a}=\operatorname{exp}(\bar{\tau})-1$.
The proposed estimator $\hat{\rho}_{b}$ and associated $z$-test perform well even for modest sample sizes. For both designs with large and small treatment effect heterogeneity, when $N\geqslant500$, that is when each subgroup has $100$ observations, the bias of $\hat{\rho}_{b}$ is below 1 percentage point. The empirical rejection rate of the test based on $z_{\rho}$ is close to nominal, ranging from 4.9% to 5.5% when $N\geqslant100$.
The estimators $\hat{\bar{\tau}}$ and $\hat{\rho}_{a}$ provide reasonable approximations for $\rho_{b}$ in the case of small heterogeneity, as true values of $\bar{\tau}$ and $\rho_{a}$ are around $-0.2\%$, close to $\rho_{b}=0$. The approximation bias converges to $-0.2\%$ as sample size increases. Tests of $H_{0}:\bar{\tau}=0$ using $z_{\tau}$ maintain appropriate rejection rates for $N\leqslant10^{5}$ but can be significantly oversized for large samples, e.g., $12.5\%$ when $N=10^{6}$. Intuitively, as $N$ grows, $\sigma_{\bar{\tau}}=\bar{\sigma}_{\bar{\tau}}/\sqrt{N}$ shrinks, causing the relative bias $(\rho_{b}-\bar{\tau})/\sigma_{\bar{\tau}}$ to grow.
In the large heterogeneity design, the approximation bias of $\hat{\bar{\tau}}$ and $\hat{\rho}_{a}$ is much larger and converges to $-3.3\%$. As a result, tests for $H_{0}:\bar{\tau}=0$ based on $z_{\tau}$ exhibit substantial size distortions even at moderate sample sizes, e.g., $9.14\%$ rejection when $N=2000$. Overall, the simulations show that reporting $\rho_{b}$ can matter both for point estimates and for inference, especially when dealing with large treatment effect heterogeneity or very large sample sizes.
This section uses two empirical applications to illustrate the proposed estimation and inference procedures. The applications further assess how $\hat{\rho}_{b}$ compares with conventional estimators $\hat{\bar{\tau}}$ and $\hat{\rho}_{a}$ in practice.
The first application uses a semi-log regression model to analyze the randomized experiment in atkin_exporting_2017, which examines the causal impact of exporting on the performance of small rug producers in Egypt.
atkin_exporting_2017 recruited two samples of firms satisfying specific criteria. Each sample was divided into strata based on rug type and loom size. Within each stratum, firms were randomly assigned to the treatment group and offered access to export orders, though not all took up. This replication focuses on the first three follow-up rounds of the joint sample of 219 firms producing duble rugs. The sample comprises 79 firms from 4 strata in sample 1, and 140 firms from 4 strata in sample 2. Of the 74 firms in the treatment group, 47 took up the export opportunity.
Given randomization within strata, it is reasonable to expect treatment effects to differ across strata. I therefore replace the single treatment dummy with dummies for eight sub-treatment groups, modifying Eq. (1) in atkin_exporting_2017 to:
where $Y_{igt}$ is the outcome of firm $i$ in stratum $g$ in round $t$, $Y_{ig0}$ is the baseline outcome, and $\delta_{g}$ and $\theta_{t}$ capture stratum and round fixed effects. Outcomes include various measures of profits in Table V and determinants of profits like price and output in Table VI of atkin_exporting_2017. The sub-treatment group dummy $D_{i}^{(g)}=1$ if firm $i$ is in stratum $g$ and assigned to the treatment group. Since not all treatment firms took up the treatment, $\tau_{g}$ represents the average intent to treat (ITT) effect in log points for sub-treatment group $g$. Standard errors are clustered at the firm level.
Table (ref) presents estimates and standard errors (in parentheses) for $\hat{\bar{\tau}}$, $\hat{\rho}_{a}=\operatorname{exp}(\hat{\bar{\tau}})-1$, and $\hat{\rho}_{b}=\sum\hat{w}_{g}\operatorname{exp}(\hat{\tau}_{g})-1$.The first panel replicates Columns (1)(3)(5)(7) of Table V Panel A in atkin_exporting_2017, showing the impact on various profit measures. The second panel replicates Panel B, showing the impact on profits per owner hour. The third panel replicates Columns (1)(3)(5)(9)(11) of Table VI, showing the impact on various determinants of profits.
For most outcomes, $\hat{\bar{\tau}}$ differs substantially from $\hat{\rho}_{b}$. The difference between $\hat{\rho}_{a}$ and $\hat{\rho}_{b}$ is about 1 to 2 percentage points for many outcomes. For example, the impact of exporting on direct profits per owner hour is $18.3\%$ (95% CI: $[5.3\%,31.4\%]$) using $\hat{\rho}_{a}$,\footnote{The confidence interval for $\rho_{a}$ here is $CI_{a}=\left[\operatorname{exp}\left(\hat{\bar{\tau}}+\hat{\sigma}_{\bar{\tau}}z_{\alpha/2}\right)-1,\operatorname{exp}\left(\hat{\bar{\tau}}-\hat{\sigma}_{\bar{\tau}}z_{\alpha/2}\right)-1\right]$, where $z_{\alpha/2}$ is the $\alpha/2$ quantile of the standard normal distribution.} and $20.3\%$ (95% CI: $[6.2\%,34.5\%]$) using $\hat{\rho}_{b}$, differing by 2 percentage points. The differences are larger for hypothetical profits and output price, e.g., $\hat{\rho}_{a}=43.3\%$ (95% CI: $[11.9\%,74.6\%]$) and $\hat{\rho}_{b}=53.4\%$ (95% CI: $[19.0\%,87.8\%]$) for hypothetical profits.
The empirical results also provide evidence of size distortion in conventional inference methods. Given that $z_{\tau}=(\rho_{0}-\bar{\tau})/\sigma_{\bar{\tau}}$ and $z_{a}=\left(\ln(\rho_{0}+1)-\bar{\tau}\right)/\sigma_{\bar{\tau}}$, the size distortion of tests based on $z_{\tau}$ and $z_{a}$ depends critically on the magnitude of $\rho_{b}-\bar{\tau}$ or $\ln(\rho_{b}+1)-\bar{\tau}$ relative to $\sigma_{\bar{\tau}}=\bar{\sigma}_{\bar{\tau}}/\sqrt{N}$. In this application, the ratio $(\hat{\rho}_{b}-\hat{\bar{\tau}})/\hat{\sigma}_{\bar{\tau}}$ often exceeds $0.4$, reaching as high as $1.69$ for output prices. For instance, for direct profits per owner hour, this ratio equals $(0.203-0.168)/0.056=0.625$. Such magnitudes imply severe size distortion for tests based on $z_{\tau}$. For inference based on $z_{a}=(\hat{\bar{\tau}}-\ln(\rho_{0}+1))/\sigma_{\bar{\tau}}$, the relevant ratio $(\ln(\hat{\rho}_{b}+1)-\hat{\bar{\tau}})/\hat{\sigma}_{\bar{\tau}}$ is generally smaller, with most values below $0.3$. However, this ratio still exceeds $0.5$ for hypothetical profits (0.607) and output prices (0.561), indicating non-trivial size distortion. Monte Carlo simulations calibrated to these empirical estimates, presented in Section (ref) of the Online Appendix, further confirm these distortions.
Overall, the difference between $\hat{\rho}_{b}$ and $\hat{\rho}_{a}$ is modest for most outcomes in this application. However, the results also show that when treatment effect heterogeneity is large, these differences can become non-negligible, leading to meaningful bias and size distortions in conventional method.
The second application illustrates the applicability of my methodology to staggered difference-in-differences designs. It extends the analysis in Table 2 of alsan_watersheds_2019, examining the impact of water and sewerage infrastructures on child mortality in the United States.
alsan_watersheds_2019 explore the staggered rollout of water and sewerage infrastructures across municipalities in Massachusetts between 1892 and 1903. They utilize panel data of 60 Massachusetts municipalities from 1880 to 1920. Their main finding is that neither water nor sewerage infrastructure alone affected child mortality, but the combination of both systems reduced child mortality by $0.266$ log points.
Based on this finding, I define treatment as the joint presence of both water and sewerage infrastructures, while controlling for the separate effects of having only one system. A municipality is in cohort $c$ if it first obtained both sewerage and safe water infrastructures in year $c$. Table A1 in alsan_watersheds_2019 documents the intervention year for each municipality.
I estimate $\tau$ using the following model
where $Y_{it}$ is the under-5 mortality rate in municipality $i$ in year $t$. The treatment indicator $D_{it}(c,r)$ equals 1 if municipality $i$ belongs to cohort $c$ and year $t$ is its $r$-th post-treatment period.\footnote{Following alsan_watersheds_2019, event time is grouped into 2-year bins except for period $-1$. For example, $r=0$ for event times $0$ and 1, $r=1$ for event times $2$ and $3$, $r=5$ for event times $10$ and above. Similarly, $r=-2$ for event times $-3$ and $-2$, and $r=-6$ for event times $\leqslant-10$.} The specification includes municipality fixed effects $\alpha_{i}$ and year fixed effects $\beta_{t}$. Control variables $X_{it}$ include (1) $Wateralone$, indicator for having only water system but no sewerage system in year $t$, (2) $Seweragealone$, indicator for having only sewerage system but no safe water system; (3) “log of population density, percentage foreign-born, percentage male, and the percentage of females employed in manufacturing” and “municipality specific linear trends” as in alsan_watersheds_2019. I estimate $\tau(c,r)$ by OLS and cluster standard errors at the municipality level following alsan_watersheds_2019. Weights for aggregate effects use sample shares, as described in detail in Example 2 of Section (ref) of the Online Appendix. I set $\bar{\Sigma}_{w}=0$ since the sample represents the complete population of Massachusetts municipalities.
Results are summarized in Table (ref), with all values multiplied by 100. From top to bottom are estimates and standard errors of $\hat{\bar{\tau}}$, $\hat{\rho}_{a}=\operatorname{exp}(\hat{\bar{\tau}})-1$, and $\hat{\rho}_{b}$ for all treated units, each event time, and each cohort. Due to their large absolute values, $\hat{\bar{\tau}}$ differs dramatically from $\hat{\rho}_{a}$ and $\hat{\rho}_{b}$ throughout. The conventional estimator $\hat{\rho}_{a}$ and the proposed estimator $\hat{\rho}_{b}$ also differ meaningfully in some cases, with gaps reaching up to $5.9$ percentage points. For all treated units, $\hat{\rho}_{a}$ ($-40.2\%$) overstates the mortality reduction compared to the proposed estimator $\hat{\rho}_{b}$ ($-37.8\%$), a difference of $2.4$ percentage points. The gap between the two estimators varies across event times: relatively small for later periods ($\hat{\rho}_{a}=-46.8\%$ vs. $\hat{\rho}_{b}=-46.1\%$ at event time 5) but substantial for early periods, particularly event time 1 where $\hat{\rho}_{a}=-25.9\%$ and $\hat{\rho}_{b}=-20.0\%$. Cohort-specific estimates show even greater variation: while cohorts 1898, 1899, and 1902 exhibit relatively small differences between $\hat{\rho}_{a}$ and $\hat{\rho}_{b}$, cohort 1903 shows a substantial gap, with $\hat{\rho}_{a}=-25.1\%$ and $\hat{\rho}_{b}=-19.5\%$.
As in the first application, $\hat{\rho}_{a}$ is often close to $\hat{\rho}_{b}$, so the conventional estimators can be adequate in many cases. However, when treatment effects vary substantially across subgroups, the resulting gap can be economically meaningful. Because the magnitude of this gap is difficult to gauge ex ante, reporting $\hat{\rho}_{b}$ therefore provides a practical diagnostic and, when needed, a correction for heterogeneity-induced bias when interpreting ATEs in percentage points.
This paper highlights the importance of correctly estimating and interpreting ATEs in percentage points when treatment effects are heterogeneous. Differences between ATEs in log points and in percentage points can be substantial, especially when treatment effects are large or vary significantly. Failing to account for these differences may lead to misinterpretation of results and potentially misguided policy recommendations.
My proposed methods provide researchers with tools for improved estimation and inference of ATEs in percentage points in the presence of treatment effect heterogeneity. By accounting for heterogeneity across observable subgroups, the methods yield point identification when treatment effects are constant within subgroups and informative lower bounds otherwise. The methods can be applied to a variety of settings like semi-log regression models. They are particularly relevant for research designs such as staggered difference-in-differences models, where treatment effect heterogeneity is common.
By applying the methods to empirical studies on how exporting affects firm productivity and how the combination of water and sewerage infrastructure affects children's mortality, I demonstrate how accounting for heterogeneity can affect the interpretation of ATE in percentage points in practice. As empirical studies continue to grapple with complex treatment effect patterns, tools like those presented in this paper willbe more useful for accurate estimation and inference.
I am grateful to the editor, the associate editor, and two anonymous referees for their constructive comments and suggestions that substantially improved this paper. I thank Ingmar Prucha, Guido Kuersteiner, and participants at the NSFC 2025 Young Scientists Fund (Category C, Economics) Awardees\textquoteright Meeting, the 10th Young Econometricians in Asia-Pacific (YEAP) Annual Meeting, and the lunch seminar at the Department of Public Economics, Xiamen University, for valuable comments. Financial support from the Young Scientists Fund of the National Natural Science Foundation of China (Grant No. 72403212) and the Young Scientists Fund of the Fujian Natural Science Foundation (Grant No. 2023J05010) is gratefully acknowledged. All remaining errors are my own.
During the preparation of this work the author used Claude and ChatGPT in order to improve the readability and language of the paper, as well as to assist with coding in Stata and R. After using this tool/service, the author reviewed and edited the content as needed and takes full responsibility for the content of the published article.