EconBase
← Back to paper

Aggregation Trees

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

83,017 characters · 19 sections · 99 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Aggregation Trees

titlepage\begin{abstract} Uncovering the heterogeneous effects of particular policies or “ treatments" is a key concern for researchers and policymakers. A common approach is to report average treatment effects across subgroups based on observable covariates. However, the choice of subgroups is crucial as it poses the risk of $p$-hacking and requires balancing interpretability with granularity. This paper proposes a nonparametric approach to construct heterogeneous subgroups. The approach enables a flexible exploration of the trade-off between interpretability and the discovery of more granular heterogeneity by constructing a sequence of nested groupings, each with an optimality property. By integrating our approach with “ honesty" and debiased machine learning, we provide valid inference about the average treatment effect of each group. We validate the proposed methodology through an empirical Monte-Carlo study and apply it to revisit the impact of maternal smoking on birth weight, revealing systematic heterogeneity driven by parental and birth-related characteristics. \noindentKeywords: Causality, conditional average treatment effects, recursive partitioning, subgroup discovery, subgroup analysis. \noindentJEL Codes: C29, C45, C55 \\ \end{abstract} \setcounter{page}{0} \thispagestyle{empty}

\doublespacing

Introduction

Understanding the effects of a particular policy or “ treatment" is a key concern for researchers and policymakers. Traditionally, the assessment of the policy's actual effectiveness involves the identification and estimation of the Average Treatment Effect (ATE), a parameter that quantifies the average impact of the policy on the reference population angrist2009mostly, imbens2015causal. However, while the ATE is straightforward to interpret, it lacks information regarding effect heterogeneity and therefore does not allow us to explore the distributional impacts of the policy, which hold significant importance for decision-making when the social welfare criterion representing the preferences of the policymakers is not “ utilitarian" kitagawa2021equality.

A common approach to tackle effect heterogeneity is to report the ATEs across different subgroups defined by observable covariates. These Group Average Treatment Effects (GATEs) enable us to explore heterogeneity while maintaining a certain level of interpretability and are widely employed in applied research.\footnote{\ For instance, chernozhukov2017generic document that, among 189 randomized control trials published in top economic journals since 2006, $40\%$ report at least one subgroup analysis.} GATEs are particularly valuable when decision rules need to be applied or interpreted by humans, such as treatment guidelines for physicians athey2016recursive. However, two practical issues arise. First, there could be many ways to form subgroups, and when no natural choice exists, iteratively searching for subgroups with significantly estimated GATEs raises the possibility of $p$-hacking imbens2021statistical.\footnote{\ \ In some empirical applications, subgroup definitions are “ natural" because they are guided by domain knowledge or policy relevance. Examples include clinically meaningful age bands, and gender or ethnicity strata.} Second, even when subgroup definitions are fixed, deciding how many groups to report remains nontrivial: more groups can reveal finer heterogeneity, but too many can undermine interpretability.ces a methodology for constructing heterogeneous subgroups that enables a flexible and coherent exploration of the trade-off between interpretability and the discovery of more granular heterogeneity while avoiding complications associated with $p$-hacking. The approach can serve as an alternative to pre-analysis plans---often criticized for limiting the potential for uncovering unexpected heterogeneity---by enabling agnostic exploration of effect heterogeneity. It can also complement pre-specified analyses by providing a data-driven partition to assess subgroup patterns beyond the pre-registered hypotheses.

The proposed methodology, hereafter referred to as aggregation trees, builds on standard decision trees breiman1984cart to aggregate units with similar estimated responses to the treatment.\footnote{\ Essentially, we adopt a “fit-the-fit” strategy hahn2020bayesian, bargagli2020causal, bargagli2022heterogeneous: first estimate CATEs using any suitable method, then fit these estimates with a decision tree.} The resulting tree is then pruned to generate a sequence of groupings, one for each level of granularity. We show that each grouping features an optimality property in that it ensures that the loss in explained heterogeneity resulting from aggregation is minimized. Moreover, the sequence is nested in the sense that subgroups formed at a particular level of granularity are never disrupted at coarser levels. This property guarantees the consistency of the results across the different granularity levels, which is a fundamental requirement for any classification system cotterman1992classification.

For a particular grouping, we leverage debiased machine learning procedures semenova2021debiased to obtain point estimates and standard errors for the GATEs.\footnote{\ In randomized experiments, a simple regression of outcomes on group dummies and their interactions with treatment assignment ensures that the interaction coefficients identify each group’s GATE.} We further combine this approach with “ honesty" athey2016recursive to deliver valid inference. Honesty is a subsample-splitting technique that requires that different observations are used to form subgroups and estimate the GATEs. In analogy to classical econometrics, this is equivalent to using different subsamples to select and estimate a model. This way, the asymptotic properties of GATE estimates are the same as if the groupings had been exogenously given, and we can use the estimated standard errors to conduct valid inference as usual, e.g., by constructing conventional confidence intervals.se a training subsample to estimate CATEs with any suitable method and to construct the sequence of groupings. We then use a disjoint honest subsample to estimate the GATEs. To account for selection into treatment, we construct a standard Neyman-orthogonal score using cross-fitted nuisance functions chernozhukov2018double and regress this score on group dummies, performing all steps on the honest sample to maintain the inferential guarantees of semenova2021debiased.\footnote{\ Neyman-orthogonal scores are central in recent causal machine learning literature. chernozhukov2018double show their advantages for ATE estimation and inference with flexible machine learning estimation of the nuisance functions. semenova2021debiased extend this logic to provide estimation and inference methods for the best linear predictor of the CATE function, which automatically target GATEs when the chosen set of basis function consists of group dummies. kennedy2023towards pushes this further by combining orthogonal scores with linear smoothing techniques---shown to satisfy a key “ stability" condition---to construct a two-stage doubly robust CATE estimator.} similar to the causal trees of athey2016recursive, but it differs in its direct applicability to observational studies and its focus on the trade-off between interpretability and the discovery of more granular heterogeneity. We compare the performance of aggregation and causal trees using an empirical Monte-Carlo study huber2013performance, lechner2013sensitivity.\footnote{\ The literature presents a wide array of causal machine learning methodologies for estimating dense heterogeneous treatment effects wager2018estimation, athey2019generalized, kunzel2019metalearners, lechner2022modified. However, aggregation trees focus on the construction of heterogeneous subgroups and the subsequent estimation of GATEs. Given these distinct objectives, we do not include these methodologies in our simulations.} Our simulation shows that aggregation trees lead to lower root mean squared error in the estimated treatment effects, with reductions of up to $121\%$. This improvement entirely stems from the lower variance of aggregation trees, resulting from a splitting strategy that is robust to covariates affecting the outcome levels but not the treatment effects.

We also investigate the benefits of honesty compared to more standard “ adaptive" estimation that uses the same data for constructing the tree and GATE estimation. Honesty greatly benefits inference, ensuring approximately nominal coverage of confidence intervals. In contrast, adaptive estimation can result in coverage rates as low as $58\%$.

The proposed methodology is applied to revisit the impact of maternal smoking on birth weight almond2005costs, cattaneo2010efficient. The analysis finds evidence of systematic heterogeneity, as different subgroups react differently to the same treatment. Moreover, the analysis reveals that effect heterogeneity is driven by parental and birth-related characteristics. The results are consistent with previous research showing that the effects are stronger for children born to adult mothers abrevaya2015estimating, zimmert2019nonparametric. Furthermore, we provide evidence that the effects are more pronounced when prenatal care visits are more frequent and occur earlier.

This paper contributes to three distinct strands of the literature. First, it relates to tree-based subgroup-discovery methodologies. These approaches use recursive partitioning of the covariate space to construct heterogenous groups, adapting the standard CART algorithm breiman1984cart to target treatment effects rather than outcomes. For instance, athey2016recursive grow trees by choosing splits that minimize an estimate of the mean-squared error of treatment effects and employ sample-splitting techniques---i.e., honesty---for valid inference, while steingrimsson2019subgroup maximize standardized differences in treatment effects using covariate-adjusted leaf estimators.\footnote{\ These approaches are developed for randomized experiments. Extensions to selection-on-observables yang2022causal and to instrumental-variables settings bargagli2020causalIV exist.} A complementary line of research leverages tree-ensemble algorithms breiman2001random, chen2016xgboost to reduce instability and explore richer partitions.\footnote{\ For example, bargagli2020causal use ensembles of decision trees to generate a large set of candidate subgroups---defined by if-then decision rules---and then select those most predictive of preliminary CATE estimates using LASSO tibshirani1996regression.} Yet single-tree models remain more interpretable, facilitating communication with non-experts and direct use in regulatory policy lee2021discovering, bargagli2022heterogeneous and in learning treatment-assignment policies athey2021policy, bodory2024enabling. This paper contributes by introducing a novel tree-based methodology for constructing heterogeneous subgroups that applies standard CART to estimated CATEs and combines it with honesty athey2016recursive and debiased machine learning semenova2021debiased to deliver valid inference for GATEs.that estimate heterogeneous treatment effects. The recent causal machine learning literature adapts machine learning tools to CATE estimation, most naturally under selection-on-observables with many covariates.\footnote{\ Broadly, there are two main strategies. One decomposes the problem into supervised prediction tasks via meta-learners kunzel2019metalearners. The other tailors machine learning algorithms to produce causal estimates directly rather than outcome predictions wager2018estimation, athey2019generalized, lechner2018modified, lechner2022modified, hahn2020bayesian.} Yet unit-level estimates can be hard to interpret and may exhibit substantial sampling variability, so what appears as heterogeneity may simply be estimation noise chernozhukov2017generic. This has motivated GATE analyses and associated methods abrevaya2015estimating, lee2017doubly, lechner2018modified, zimmert2019nonparametric, fan2022estimation, lechner2022modified, which, however, require researchers to predefine groups.\footnote{\ bearth2024causal discuss how to analyze and interpret differences in GATEs across groups while accounting for variation in other covariates. lee2021discovering develop randomization-based tests to assess whether subgroups defined by a given tree have GATEs that differ from the overall ATE.} Our contribution is to propose a methodology for constructing subgroups from the data---thus removing the need for ex ante group definitions---and then providing GATE estimation and inference via honesty athey2016recursive and debiased machine learning semenova2021debiased.o the broader literature on maternal risky behaviors and birth outcomes. Maternal behaviors are important policy levers because they are modifiable risk factors.\footnote{\ bhalotra2017infant find that interventions providing information and support to mothers improved both short- and long-run infant survival.} Because smoking during pregnancy is widely considered the most salient---and most readily modifiable---of these risks almond2005costs, a large literature examines its impact on infant health.\footnote{\ bhalotra2019twin show that smoking during pregnancy—along with other maternal health conditions and risky behaviors—is negatively associated with the probability of twin birth.} Numerous studies consistently document sizable negative average effects on birth weight almond2005costs, abrevaya2006estimating and increasingly negative effects with maternal age abrevaya2015estimating, lee2017doubly, zimmert2019nonparametric, fan2022estimation. Differences in smoking behavior also help explain variation in effects cattaneo2010efficient, heiler2021effect, bodory2022high. This paper contributes by providing robust evidence of systematic heterogeneity---different subgroups of infants are affected differently by maternal smoking---and by offering novel evidence that effects are more pronounced when prenatal care begins earlier.ws. Section (ref) discusses the estimands of interest and their identification. Section (ref) introduces aggregation trees and compares them to causal trees. Section (ref) shows the simulation results. Section (ref) illustrates the empirical exercise. Section (ref) concludes.

Causal framework

Estimands

We define the estimands of interest using the potential outcomes model neyman1923, rubin1974estimating. Suppose to have access to a sample of $n$ i.i.d. observations $( Y_i, D_i, X_i )$, where $Y_i \in \mathcal{Y}$ is the outcome targeted by the treatment, $D_i \in \{ 0, 1 \}$ is the binary treatment indicator, and $X_i = ( X_{i1}, \dots, X_{ip} )^\top \in {\cal X}$ is a $p \times 1$ vector of pre-treatment covariates. We posit the existence of two potential outcomes $Y_{i} ( 0 )$ and $Y_{i} ( 1 )$, representing the outcome that the $i$-th unit would experience under each treatment level. The observed outcome for unit $i$ is then the potential outcome corresponding to the treatment received:\footnote{\ The definition of potential outcomes and the observational rule linking them to the observed outcomes implicitly assume the absence of spillover effects among units, which is violated in settings where some units are connected through networks.}

equation[equation omitted — 95 chars of source]

To define the effect of the treatment, we can take the differences in the potential outcomes of each unit $\xi_i := Y_{i} ( 1 ) - Y_{i} ( 0 )$ and aggregate them at different levels of granularity. The coarsest estimand of interest is the Average Treatment Effect (ATE),

equation[equation omitted — 89 chars of source]

The ATE quantifies the average impact of the policy on the reference population and is straightforward to interpret. However, it lacks information regarding the distributional impacts of the policy.

To tackle effect heterogeneity, we can focus instead on the Conditional Average Treatment Effects (CATEs),

equation[equation omitted — 105 chars of source]

The CATEs provide information at the finest level of granularity achievable with the information at hand and enable us to relate effect heterogeneity to the observable covariates. However, they are difficult to interpret.

The Group Average Treatment Effects (GATEs) provide a way to explore heterogeneity while maintaining a certain level of interpretability. The GATEs are averages of (potentially heterogeneous) individual treatment effects within regions of the covariate space and are defined by

equation[equation omitted — 137 chars of source]

where the groups ${\cal X}_1, \dots, {\cal X}_G$ represent a partition of ${\cal X}$.\footnote{\ If grouping is based on the levels of a single discrete variable $Z_i \subset X_i$, each GATE simplifies to $\tau_g = \operatorname{\mathbb{E}} [ \xi_i | Z_i = g ]$.} Importantly, equation (ref) allows for heterogeneous treatment effects within groups: we do not assume $\xi_i = \xi_j$ for units $i$ and $j$ with $X_i, X_j \in {\cal X}_g$. We also do not posit a “ true" partition of ${\cal X}$; rather, GATEs are used as an interpretable summary of effect heterogeneity.\footnote{\ Specifically, we do not assume regions of ${\cal X}$ with constant $\xi_i$ (e.g., a step-function data-generating process for treatment effects).}

The task of GATE analysis is therefore to form groups in a principled way---deciding which covariates define the grouping and how many groups $G$ to report---and to obtain valid inference for the resulting GATEs.\footnote{\ While higher values of $G$ can uncover more detailed heterogeneity, partitions that are too fine may not offer substantial advantages in terms of interpretability compared to CATEs.} This paper proposes a data-driven procedure that constructs partitions of ${\cal X}$ at various levels of granularity and shows how to obtain valid inference for the group effects.fication} All the estimands discussed in the previous section are defined in terms of potential outcomes. However, each unit is either treated or not treated. We thus observe only one potential outcome per unit, and further assumptions are needed for identification.

The following standard assumptions are sufficient to identify the CATEs imbens2015causal:

assumption(Unconfoundedness): $\{ Y_{i} ( 0 ), Y_{i} ( 1 ) \} \perp \!\!\! \perp D_i | X_i$.
assumption(Common support): $0 < \pi ( X_i ) < 1$, where $\pi ( X_i ) := \operatorname{\mathbb{P}} ( D_i = 1 | X_i )$ is the conditional treatment probability (or propensity score).

The unconfoundedness assumption requires that $X_i$ contains all “ confounder" jointly affecting the treatment assignment and the outcome.\footnote{\ $X_i$ can also include additional “ heterogeneity covariates" not necessary for identification but for which effect heterogeneity is of interest. The sets of confounders and heterogeneity covariates can overlap in any way or be disjoint.} The common support assumption states that each unit must have a non-zero probability of belonging to the treatment and control groups.

Under Assumptions (ref)--(ref), the CATEs are identified from observable data:

equation[equation omitted — 613 chars of source]

where $\mu ( D_i, X_i ) := \operatorname{\mathbb{E}} [ Y_i | D_i, X_i ]$ and Assumption (ref) ensures that the conditional expectations are well-defined for all values within the support of $X_i$. The ATE and GATEs are expressed as expectations of the CATEs and are thus identified under the same assumptions.

Aggregation trees

This section outlines the implementation of the methodology proposed in this paper, which involves three steps. First, an estimation step constructs an estimate $\hat{\tau} ( \cdot )$ of $\tau ( \cdot )$. Second, a tree-growing step approximates $\hat{\tau} ( X_i )$ by a standard decision tree breiman1984cart that constrains the set of admissible groupings. Third, a tree-pruning step generates a sequence of nested subtrees, one for each level of granularity. Each subtree provides an optimal grouping, where optimality means that, at each granularity level, the groupings minimize the loss in explained heterogeneity due to aggregation.

The next subsection describes the tree-growing step, detailing the splitting strategy used by aggregation trees and comparing it with that of causal trees athey2016recursive. Then, the tree-pruning step is discussed. Finally, we explain how to conduct valid inference about the GATEs using double machine learning procedures. The algorithm below summarizes the full implementation.

algorithm[algorithm omitted — 1,880 chars of source]

Tree-growing step

Trees are typically constructed by greedily minimizing an assumed loss function based on the mean squared error criterion. This minimization follows the approach proposed by breiman1984cart, who suggest recursively stratifying the covariate space using axis-aligned splits.

Starting with a region of the covariate space $\mathcal{R}_m \subseteq {\cal X}$, consider a candidate splitting variable $X_{ij}$ and splitting point $x$. We define the corresponding subregions as:\footnote{\ In the case of categorical splitting variables, $x$ corresponds to a subset of possible levels of $X_{ij}$, and the inequality signs are replaced by $\in$ and $\notin$.}

equation[equation omitted — 182 chars of source]

The split occurs on some pair $( j, x )$, and the population is stratified accordingly. The process is then repeated in the resulting subregions, thus obtaining increasingly finer partitions of ${\cal X}$. The whole procedure can be described by the shape of a decision tree: the “ root” (i.e., the node with no “ parent") corresponds to ${\cal X}$, the $m$-th internal node represents subregion $\mathcal{R}_m$ and has two “ children" nodes representing subregions $\mathcal{R}_{m + 1}$ and $\mathcal{R}_{m + 2}$, and the “ leaves" (i.e., the collection of terminal nodes) correspond to a partition of ${\cal X}$.

Ideally, we would like to explore the space of all possible trees and pick the one whose associated partition minimizes the assumed loss function. However, it is generally infeasible to enumerate all the distinct binary decision trees.\footnote{\ Consider the situation where $X_i$ is composed of $p$ binary covariates, and let $\mathcal{D}$ be the “ depth" of a given tree (i.e., the number of nodes connecting the root to the furthest leaf). Appendix (ref) shows that $L_{\mathcal{D}} = \prod_{d = 1}^\mathcal{D} ( p - ( d - 1 ) )^{2^{d - 1}}$ is a lower bound for the number of distinct binary decision trees grown by recursively partitioning ${\cal X}$ and having a depth equal to or lower than $\mathcal{D}$. $L_{\mathcal{D}}$ quickly diverges as $p$ grows. For example, fixing $\mathcal{D} = 3$ and letting $p = 10$ yields a lower bound of 3,317,760, while letting $p = 20$ leads to a bound of 757,926,720. Things only worsen with categorical covariates taking more than two values or continuous covariates discretized using a large number of bins.} To cope with this issue, breiman1984cart suggest a “ greedy" approach that partitions each region $\mathcal{R}_m \subseteq {\cal X}$ by choosing the split that minimizes the assumed loss function within the resulting subregions $\mathcal{R}_{m + 1}$ and $\mathcal{R}_{m + 2}$. This process is then iterated until some particular “ stopping criterion" is met, for instance the maximum depth of the tree. This approach is greedy in that it ignores that a suboptimal split could yield better results at later steps and is generally considered to be a reasonable way of circumventing the exhaustive search of the space of all possible trees.

Let $\mathcal{T}$ be some tree constructed using a training sample $\mathcal{S}^{tr}$, and let $\mathcal{S}^{te}$ be an independent test sample. Then, when heterogeneous treatment effects are the object of the analysis, one wants to build a tree that minimizes $EMSE ( \mathcal{T} ) = \operatorname{\mathbb{E}} [ MSE ( \mathcal{S}^{te}, \mathcal{S}^{tr}, \mathcal{T} ) ]$, where the expectation is taken over the joint distribution of the training and test samples and:\footnote{\ We are departing from the standard criterion $\operatorname{\mathbb{E}} [ \{ \tau_i - \tilde{\tau} ( X_i, \mathcal{S}^{tr}, \mathcal{T} ) \}^2 ]$ by subtracting $\operatorname{\mathbb{E}} [ \tau_i^2 ]$. Because this term does not depend on an estimator, the tree that minimizes the standard criterion also minimizes $EMSE ( \cdot )$.}

equation[equation omitted — 520 chars of source]

with $\tau_i \equiv \tau ( X_i )$ and $\tilde{\tau} ( x, \mathcal{S}, \mathcal{T} )$ some estimate of $\tau ( \cdot )$ within the leaf $\ell ( x, \mathcal{T} )$ of $\mathcal{T}$ where $x$ falls obtained using observations in the sample $\mathcal{S}$. In practice, trees are constructed by greedily minimizing an in-sample loss function $MSE ( \mathcal{S}^{tr}, \mathcal{S}^{tr}, \mathcal{T} )$.\footnote{\ This is what athey2016recursive denote as the “ adaptive" case, where the same sample is used to both construct and estimate the tree. athey2016recursive also consider an alternative “ honest" criterion $MSE ( \mathcal{S}^{te}, \mathcal{S}^{hon}, \mathcal{T} )$ that uses different samples for construction of the tree ($\mathcal{S}^{tr}$) and treatment effect estimation ($\mathcal{S}^{hon}$). For simplicity, we focus on the adaptive case here and postpone the discussion of honesty to a later section.}

The key challenge in a causal inference framework is that we do not observe $\tau_i$. Thus, $MSE ( \cdot, \cdot, \cdot )$ is an infeasible criterion and needs to be estimated. We propose an estimator of $MSE ( \cdot, \cdot, \cdot )$ obtained by plugging an estimate $\hat{\tau} ( \cdot )$ of $\tau ( \cdot )$ constructed from the training sample in a preliminary estimation step:

equation[equation omitted — 394 chars of source]

with:

equation[equation omitted — 226 chars of source]

If we use a consistent estimator of $\tau_i$, then $\widehat{MSE}_{\scriptstyle_{AT}} ( \mathcal{S}^{te}, \mathcal{S}^{tr}, \mathcal{T} )$ is an approximately unbiased estimator of $MSE ( \mathcal{S}^{te}, \mathcal{S}^{tr}, \mathcal{T} )$ if the assignment to treatment is random conditional on $X_i$. In practice, we construct aggregation trees by greedily minimizing the in-sample loss function $\widehat{MSE}_{\scriptstyle_{AT}} ( \mathcal{S}^{tr}, \mathcal{S}^{tr}, \mathcal{T} )$. This is equivalent to selecting splits that minimize the conditional variance of $\hat{\tau}_i$ in the resulting nodes. Consequently, the greedy approach partitions each region $\mathcal{R}_m \subseteq {\cal X}$ in a way that maximizes systematic heterogeneity between the resulting subgroups, thereby constructing a set of admissible groupings represented by the resulting tree $\mathcal{T}_0$.

athey2016recursive propose instead the following estimator of $MSE ( \cdot, \cdot, \cdot )$:

equation[equation omitted — 450 chars of source]

with:

equation[equation omitted — 177 chars of source]

and

equation[equation omitted — 214 chars of source]

In randomized experiments, $\widehat{MSE}_{\scriptstyle_{CT}} ( \mathcal{S}^{te}, \mathcal{S}^{tr}, \mathcal{T} )$ is an approximately unbiased estimator of $MSE ( \mathcal{S}^{te}, \mathcal{S}^{tr}, \mathcal{T} )$, as $\operatorname{\mathbb{E}} [ \tau_i | i \in \mathcal{S}^{te} : i \in\ell ( x, \mathcal{T} ) ] = \operatorname{\mathbb{E}} [ \tilde{\tau}_{\scriptstyle_{CT}} ( x, \mathcal{S}^{te}, \mathcal{T} ) ]$, with the expectations taken over the distribution of the test samples. As before, we construct causal trees by greedily minimizing the in-sample loss function $\widehat{MSE}_{\scriptstyle_{CT}} ( \mathcal{S}^{tr}, \mathcal{S}^{tr}, \mathcal{T} )$.

We expect the splitting strategy of aggregation trees to result in a lower sampling variance, especially when dealing with covariates that influence outcome levels but not treatment effects. For example, consider potential outcomes expressed as

equation[equation omitted — 97 chars of source]

with $\phi ( X_i ) = \frac{1}{2} X_{i1} + X_{i2}$ a model for the mean effect and $\tau ( X_i ) = \frac{1}{2} X_{i1}$. When exploring covariate values as potential splitting points, we shift one observation at a time from one region of the covariate space to its complement. Since each observation belongs to either the treatment or control group, this alters the sample average of the observed outcomes of only one group, thus affecting $\tilde{\tau}_{\scriptstyle_{CT}} ( \cdot, \cdot, \cdot )$. Due to the substantial influence of $X_{i2}$ on mean outcomes, moving a single observation between child nodes based on this covariate can substantially alter the sample average of the observed outcomes of one group. Consequently, we expect $\tilde{\tau}_{\scriptstyle_{CT}} ( \cdot, \cdot, \cdot )$ to exhibit considerable variability with the choice of splitting point, even though $X_{i2}$ does not enter the model for $\tau ( \cdot )$. This variability may also cause the estimator to identify spurious splits involving $X_{i2}$. In contrast, for accurately estimated CATEs, $\tilde{\tau}_{\scriptstyle_{AT}} ( \cdot, \cdot, \cdot )$ is expected to remain stable when shifting a single observation between child nodes based on $X_{i2}$. Consequently, we expect that $\tilde{\tau}_{\scriptstyle_{AT}} ( \cdot, \cdot, \cdot )$ will vary less with the choice of splitting point, resulting in a lower sampling variance.

Tree-pruning step

After a deep tree has been constructed, the standard practice is to prune it according to an assumed cost-complexity criterion. Aggregation and causal trees rely on the same criterion, which is composed of two terms:

equation[equation omitted — 167 chars of source]

The first term corresponds to the loss function used for constructing the tree and measures the in-sample goodness-of-fit of the model. The second term is a regularization component that penalizes the model's complexity---defined as the number of leaves $|\mathcal{T}|$---according to the cost-complexity parameter $\alpha \in [ 0, \infty )$. Regularization is needed to prevent overfitting: the in-sample loss function $MSE ( \mathcal{S}^{tr}, \mathcal{S}^{tr}, \mathcal{T} )$ always decreases with additional splits, even in the cases where the out-of-sample $MSE ( \mathcal{S}^{te}, \mathcal{S}^{tr}, \mathcal{T} )$ actually increases.

The parameter $\alpha$ controls the relative weight of the two components and thus the balance between the accuracy and the interpretability of the model. Define a subtree $\mathcal{T} \subset \mathcal{T}_0$ as any tree that can be obtained by collapsing any number of internal nodes of $\mathcal{T}_0$ and let $\mathcal{T}_{\alpha} \subseteq \mathcal{T}_0$ be the smallest subtree for which ((ref)) is minimized. For each $\alpha$, a unique $\mathcal{T}_{\alpha}$ exists, which can be identified by “ weakest link pruning”: starting from $\mathcal{T}_0$, we iteratively collapse the internal node that gives the slightest increase in the accuracy of the approximation.

Following this procedure, we can generate a sequence of nested subtrees $\mathcal{T}_{\alpha_{\scriptstyle_{0}}}, \mathcal{T}_{\alpha_{\scriptstyle_{1}}} \dots, \mathcal{T}_{\alpha_{\scriptstyle_{max}}}$, where $0 = \alpha_{\scriptstyle_{0}} < \alpha_{\scriptstyle_{1}} < \dots < \alpha_{\scriptstyle_{max}} < \infty$ are threshold values such that all $\alpha$ in a given interval lead to the same subtree and $\mathcal{T}_{\alpha_{\scriptstyle_{max}}}$ corresponds to the tree's root. Each subtree in this sequence is associated with a partition of the covariate space. Therefore, the tree-pruning step generates a sequence of groupings, one for each threshold value $\alpha_{\scriptstyle_{0}} < \alpha_{\scriptstyle_{1}} < \dots < \alpha_{\scriptstyle_{max}}$.

Cross-validation procedures are commonly used to determine the optimal cost-complexity parameter and thus select a single partition hastie2009elements.\footnote{\ athey2016recursive also utilize cross-validation, adapting the standard criterion for the honest case.} However, the whole sequence of groupings generated by the pruning process is also of significant interest. Because each tree $\mathcal{T}_{\alpha_{\scriptstyle_{k}}}$ is obtained by collapsing the weakest node of $\mathcal{T}_{\alpha_{\scriptstyle_{k - 1}}}$, the tree-pruning step constructs optimal groupings by aggregating the two subgroups for which the loss in explained heterogeneity resulting from aggregation is minimized. Moreover, because the sequence is nested, subgroups formed at a particular level of granularity are never disrupted at coarser levels. This property guarantees the consistency of the results across the different granularity levels cotterman1992classification. Consequently, the sequence allows for a flexible and coherent exploration of the trade-off between interpretability and the discovery of more granular heterogeneity.\footnote{\ A single partition can still be selected for practical purposes using standard cross-validation methods.}

Estimation and inference

For a particular grouping $\mathcal{T}_{\alpha}$, we can estimate the GATEs in several ways. In randomized experiments, taking the difference between the mean outcomes of treated and control units in each group is an unbiased estimator of the GATEs. Equivalently, we can obtain the same point estimates in addition to their standard errors by estimating via OLS the following linear model:

equation[equation omitted — 213 chars of source]

with $L_{i, l}$ a dummy variable equal to one if the $i$-th unit falls in the $l$-th leaf of $\mathcal{T}_{\alpha}$. Exploiting the random assignment to treatment, we can show that each $\beta_l$ identifies the GATE in the $l$-th leaf.

In observational studies, estimating model ((ref)) would yield biased GATE estimates due to the selection into treatment. To get unbiased estimates, we can use the orthogonal estimator of semenova2021debiased to estimate the best linear predictor of $\tau ( \cdot )$ given a set of dummies denoting leaf membership.\footnote{\ One can also conduct sensitivity analysis to assess robustness to violations of Assumption (ref) lee2021discovering.} The key idea is to construct a random variable $\psi_i$, generally called score, such that $\tau ( X_i ) = \operatorname{\mathbb{E}} [ \psi_i | X_i ]$, and project it onto $L_{i, 1}, \dots, L_{i, |\mathcal{T}_{\alpha}|}$.

Consider the doubly-robust scores of robins1995semiparametric:

equation[equation omitted — 213 chars of source]

Because $\operatorname{\mathbb{E}} [ \psi_i^{DR} | X_i ] = \tau ( X_i )$, this score is a natural candidate. We recognize that it depends on unknown functions $\eta ( X_i ) := \{ \mu ( 1, X_i ), \mu ( 0, X_i ), \pi ( X_i ) \}$ and make this explicit by writing $\psi_i^{DR} := \psi_i^{DR} ( \eta )$. We refer to $\eta$ as nuisance functions, as they are not of direct interest but necessary to construct a plug-in estimate $\psi_i^{DR} ( \hat{\eta} )$ of $\psi_i^{DR} ( \eta )$ that we aim to regress on $L_{i, 1}, \dots, L_{i, |\mathcal{T}_{\alpha}|}$.

semenova2021debiased show that $\psi_i^{DR} ( \eta )$ is a Neyman-orthogonal score chernozhukov2018double, that is, its plug-in estimate $\psi_i^{DR} ( \hat{\eta} )$ is insensitive to bias in the estimation of $\hat{\eta}$. They then suggest the following two-stage procedure. First, construct an estimate $\hat{\eta}$ of the nuisance functions $\eta$ using $K$-fold cross-fitting: split the sample into $K$ folds of similar sizes and, for each $k = 1, \dots, K$, estimate $\hat{\eta}_k$ using all but the $k$-th folds. Second, construct $\widehat{\psi}_i^{DR} := \psi_i^{DR} ( \hat{\eta}_k )$, where the observation $i$ belongs to the $k$-th fold, and estimate via OLS the following linear model:

equation[equation omitted — 162 chars of source]

As before, each $\beta_l$ identifies the GATE in the $l$-th leaf. Moreover, semenova2021debiased show that thanks to the Neyman-orthogonality of $\psi_i^{DR}$, the OLS estimator $\hat{\beta}_l$ of $\beta_l$ is root-$n$ consistent and asymptotically normal, provided that the product of the convergence rates of the estimators of the nuisance functions $\mu ( \cdot, \cdot )$ and $\pi ( \cdot )$ is faster than $n^{1/2}$. This allows using machine learning estimators such as random forests and LASSO to estimate the nuisance functions, as they are shown to achieve an $n^{1/4}$ convergence rate and faster under particular conditions.

However, GATE estimates may show some bias if we use the same data to construct the tree and to estimate ((ref))--((ref)), leading to invalid inference. One way out is to grow “ honest" trees athey2016recursive. Honesty is a subsample-splitting technique that requires that different observations are used to form the subgroups and estimate the GATEs. For this purpose, we split the observed sample into a training sample $\mathcal{S}^{tr}$ and an honest sample $\mathcal{S}^{hon}$ of arbitrary sizes. We use $\mathcal{S}^{tr}$ to estimate $\tau ( \cdot )$ and construct the sequence of groupings and, for a particular grouping $\mathcal{T}_{\alpha}$, we use $\mathcal{S}^{hon}$ to estimate ((ref))--((ref)). This way, the asymptotic properties of GATE estimates are the same as if the groupings had been exogenously given. Therefore, we can use the estimated standard errors to conduct valid inference as usual, e.g., by constructing conventional confidence intervals.\footnote{\ When the number of reported groups $|\mathcal{T}_{\alpha}|$ is large, it is prudent to account for multiplicity across GATEs bargagli2022heterogeneous, for example by adjusting $p$-values to account for multiple hypothesis testing controlling the familywise error rate holm1979multiple, hochberg1988sharper, hommel1988stagewise, romano2005exact or the false discovery rate benjamini1995controlling, benjamini2001control. If the partition becomes so fine that interpretability is compromised, an alternative is to summarize heterogeneity along a low-dimensional effect modifier---essentially shifting the estimand from GATEs to a “ reduced-dimensional" CATE. fan2022estimation provide estimators and uniform inference procedures for the reduced-dimensional CATE.} However, honesty generally comes at the expense of a larger mean squared error, as fewer observations are used to estimate $\hat{\tau} ( \cdot )$, construct the tree, and compute GATE estimates.

Empirical Monte-Carlo

We conduct an empirical Monte-Carlo study huber2013performance, lechner2013sensitivity to investigate the performance of aggregation and causal trees in a context that closely approximates a real application.\footnote{\ See also lechner2018modified, knaus2021machine, lechner2022modified for more recent implementations of empirical Monte-Carlo studies.} Empirical Monte-Carlo studies aim to base the data-generating processes (DGPs) on real data as much as possible, thereby evaluating the estimators in a realistically representative context. The idea is to use a sufficiently large data set as the population of interest, from which we can draw random samples for estimation.\footnote{\ We do not benchmark aggregation trees against causal forests wager2018estimation or other causal machine learning methods athey2019generalized, kunzel2019metalearners, lechner2022modified because the targets differ: those approaches estimate unit-level CATEs, whereas aggregation trees construct interpretable groupings and estimate the corresponding GATEs. Causal trees return an explicit partition with leaf-level GATEs and are therefore the natural benchmark.}

In the following subsections, we detail the implementation of the empirical Monte-Carlo study. First, we describe the data set forming our population. Next, we illustrate the DGPs, the implementation of estimators, and the performance measures we employ. Finally, we present the results.

Population

We base our simulations on a data set previously used to evaluate the impact of maternal smoking on birth weight almond2005costs, cattaneo2010efficient, heiler2021effect, bodory2022high. Since we will use the same data set in our empirical illustration in the next section, we postpone the discussion of previous findings to that section and focus here on the details of the data.

The clean data set consists of $435{,}124$ observations measured in Pennsylvania between $1989$ and $1991$. The outcome of interest is the infant's weight at birth in grams. The treatment indicator equals one if the mother smoked during pregnancy and zero otherwise. The pre-treatment covariate vector contains $39$ confounders and heterogeneity variables, providing information on the mother's and father's background characteristics (age, ethnicity, whether the mother was married or foreign-born), mother's behavior possibly associated with smoking (whether she drank alcohol during pregnancy, how many drinks per week), maternal medical risk factors not affected by smoking during pregnancy, and birth characteristics (e.g., whether the infant is first born, number and quality of prenatal care visits).\footnote{\ Table (ref) in Appendix (ref) provides a description of all the variables.}

To avoid common support issues, we drop children whose parents were particularly young or old at birth, or who attended more than thirty prenatal care visits. Moreover, we drop children whose mothers used to consume more than ten alcoholic drinks per week during pregnancy. A total of $596$ observations are removed from the original data set.

Table (ref) presents the summary statistics for the treated and control groups in the final sample. The table displays sample averages and standard deviations for each variable, along with two measures of difference in the distribution across treatment arms: the normalized difference, measuring the difference between the locations of the distributions, and the logarithm of the ratio of standard deviations, measuring the difference in the dispersion of the distributions. Overall, the sample appears to be sufficiently balanced, with only five relatively unbalanced covariates: meduc, unmarried and feduc exhibit strong differences in locations, while alcohol and n_drink exhibit strong differences in dispersion. These results are robust to the inclusion of the excluded observations.

Simulation details

We estimate the propensity score via a logistic regression using the full sample and all baseline covariates. The resulting estimate $\hat{\pi} ( \cdot )$ is later used as the true selection model in the simulation to ensure that the selection behavior closely mimics that observed in the real data. We then remove all treated units from the sample. This implies that we observe $Y_{i} ( 0 )$ for all units.

We now need to specify a model for the individual effects. Since model specification can be somewhat arbitrary, we aim to reduce this arbitrariness and better approximate a realistic scenario by fitting an honest causal forest athey2019generalized using the full sample and choosing a model that mimics the resulting predictions.\footnote{\ While it would be possible to directly use the estimated CATEs in the DGPs, this approach might bias our results towards estimators similar to those used for CATE estimation knaus2021machine.} We find that all predicted CATEs are negative and statistically different from zero at the $5\%$ significance level. Therefore, we specify the following model for the individual effects:\footnote{\ Figure (ref) in Appendix (ref) displays the estimated CATEs sorted by magnitude alongside the individual effects generated by model ((ref)).}

equation[equation omitted — 161 chars of source]

with $\tilde{\pi} ( X_i ) = \frac{\hat{\pi} ( X_i )}{\max_{i} \hat{\pi} ( X_i )}$ a normalized version of the estimated propensity score $\hat{\pi} ( X_i )$, and $\phi ( \cdot )$ the standard normal probability distribution function. The parameter $a$ controls the degree of heterogeneity. If $a = 0$, all individual effects are zero. As $a$ increases, the degree of heterogeneity increases as well.

We use model ((ref)) to compute $\hat{Y}_{i} ( 1 ) = Y_{i} ( 0 ) + \xi ( X_i )$ for all units. We then set aside a random validation sample of $10{,}000$ units to evaluate our performance measures detailed below. This approach allows us to assess the out-of-sample predictive power of estimators under investigation knaus2021machine. We treat the remaining $343{,}140$ units as our population from which we draw random samples for estimation.

After drawing a sample of size $n$, we assign the treatment using a Bernoulli process. We consider two scenarios: one with random assignment ($D_i \sim \textit{Bernoulli} ( 0.5 )$) and another with assignment based on the “ true" propensity score ($D_i \sim \textit{Bernoulli} ( \hat{\pi} ( X_i ) )$). Notice that in the latter case $\hat{\pi} ( \cdot )$ is used in both the model for individual effects and for treatment assignment. This complicates the task for estimators to separate selection bias from effect heterogeneity. Finally, we construct the observed outcomes for units in the drawn sample as in ((ref)).

We split each drawn sample into a training sample $\mathcal{S}^{tr}$ and an honest sample $\mathcal{S}^{hon}$ of equal sizes. We then construct a causal tree and two aggregation trees.\footnote{\ Aggregation trees are constructed with the R package rpart using default settings. Causal trees are constructed with the R package causalTree using default settings, except for the minsize parameter---which controls the minimum number of treated and control units required in a leaf to attempt a split---which we set to $4$ (default is $2$) to reduce the risk of “ empty-arm" leaves in the honest sample (i.e., leaves that, when populated with observations from the honest sample, contain zero treated or control units).} To build the aggregation trees, we estimate $\tau ( \cdot )$ using the X-learner kunzel2019metalearners and the causal forest athey2019generalized estimators. We use standard cross-validation procedures to select a single partition from the resulting causal and aggregation trees.\footnote{\ For causal trees, we use the honest cross‐validation criterion of athey2016recursive. For aggregation trees, we use the standard criterion.} All these operations are performed utilizing only observations from $\mathcal{S}^{tr}$.

To obtain point estimates and standard errors for the GATEs, causal trees follow the approach of athey2016recursive and estimate model ((ref)). On the other hand, aggregation trees estimate model ((ref)). Honest regression forests and $5$-fold cross-fitting are employed to estimate the nuisance functions necessary for constructing the doubly-robust scores $\psi_i^{DR}$. All these operations are performed using only observations from $\mathcal{S}^{hon}$.

We use the external validation sample to assess the quality of estimation. Three performance measures are computed: the root mean squared error, the absolute bias, and the standard deviation of the predictions for each observation in the validation sample:\footnote{\ Note that we are evaluating which estimator best approximates the individual effects $\xi ( \cdot )$. However, the relative performance of the estimators in approximating the unknown CATEs is the same, as the estimator that minimizes the mean squared error for $\xi ( x )$ also minimizes the mean squared error for $\tau ( x )$ kunzel2019metalearners.}

equation[equation omitted — 383 chars of source]

with $x$ a generic point in the validation sample, $R$ the number of replications, and $\hat{\tau}_r ( \cdot )$ the CATE estimated by the tree at the $r$-th replication. We summarize these performance measures by averaging over the validation sample. Additionally, we evaluate the actual coverage rates of conventional $95\%$ confidence intervals for the GATEs constructed using the estimated standard errors.

Results

Table (ref) presents the results obtained from $R = 1{,}000$ replications across three sample sizes and two levels of heterogeneity: low ($a = 20$) and high ($a = 50$). Overall, aggregation trees outperform causal trees in terms of prediction accuracy, with both $AT_{\scriptstyle_{XL}}$ and $AT_{\scriptstyle_{CF}}$ showing lower RMSE than causal trees. This effect is particularly strong when treatment is randomly assigned, where the RMSE of causal trees is between $28\%$ and $121\%$ larger than that of aggregation trees. When treatment assignment is based on $\hat{\pi} ( \cdot )$, the RMSE of causal trees is still larger by $17\%$ to $66\%$, except for the smallest sample size, where causal trees marginally outperform $AT_{\scriptstyle_{CF}}$. $AT_{\scriptstyle_{XL}}$ and $AT_{\scriptstyle_{CF}}$ consistently show similar perfomances.

\begingroup {8pt}

table[table omitted — 3,589 chars of source]

\endgroup

The second and third panels of Table (ref) offer additional insights into the superior prediction performance of aggregation trees by examining the absolute bias and standard deviation of the estimators. As expected, the advantage of aggregation trees is entirely driven by their lower sampling variance, resulting from the more stable splitting strategy employed. All estimators exhibit some bias, which increases as heterogeneity strengthens. When treatment is randomly assigned, the bias across all estimators is approximately the same. When treatment is instead assigned based on $\hat{\pi} ( \cdot )$, causal trees exhibit a slightly lower bias than both $AT_{\scriptstyle_{XL}}$ and $AT_{\scriptstyle_{CF}}$. However, aggregation trees show substantially lower sampling variance than causal trees, with approximately the same magnitudes and exceptions noted for RMSE. For all estimators, sampling variance remains consistent across different heterogeneity levels.

The fourth panel of Table (ref) presents the coverage rates for $95\%$ confidence intervals. Both $AT_{\scriptstyle_{XL}}$ and $AT_{\scriptstyle_{CF}}$ achieve coverage rates close to the nominal level, while causal trees consistently show lower coverage. The fifth panel of Table (ref) reveals that causal trees generate more complex models with a larger number of leaves compared to aggregation trees.\footnote{\ This pattern is consistent with prior evidence that causal trees tend to grow larger than alternative tree methods lee2021discovering, yang2022causal.} The complexity of causal trees increases significantly with larger sample sizes, while that of $AT_{\scriptstyle_{XL}}$ and $AT_{\scriptstyle_{CF}}$ grows more conservatively. This could lead to overfitting and raise concerns about multiple hypothesis testing in causal trees, potentially explaining their lower coverage rates. For all estimators, model complexity does not vary across heterogeneity levels and is smaller when treatment is assigned based on $\hat{\pi} ( \cdot )$, particularly for causal trees.

Comparing these results with Table (ref) allows us to assess the benefits of honesty. Using different data for constructing the trees and treatment effect estimation greatly benefits inference. The coverage rates of adaptive trees are considerably below the nominal rate, particularly those of causal trees that can be as low as $58\%$.

While honesty is expected to come at the expense of a larger mean squared error, our simulations show that the RMSE is actually higher for adaptive trees than for honest trees. For aggregation trees, this discrepancy may be due to the slightly more complex models generated by adaptive estimation, resulting in higher sampling variance while maintaining the same bias. In the case of causal trees, the discrepancy may also be due to the greater bias induced by adaptive estimation when the assignment is based on $\hat{\pi} ( \cdot )$. An exception arises for causal trees under randomized treatment assignment, where adaptivity produces extremely shallow trees, which results in lower RMSE compared to honest estimation.

Empirical illustration

In this section, we apply the methodology developed in this paper to revisit the impact of maternal smoking on birth weight using the same data set as in the previous section. First, we review previous findings from the literature. Then, we construct the sequence of optimal groupings. Finally, we discuss the results.

Previous findings

As documented in almond2005costs, infants born at LBW can impose substantial costs on society, with estimated expected costs of delivery and initial care exceeding 100,000\$ (at prices of year 2000) for babies weighing 1,000 grams at birth.\footnote{\ An infant is considered born at LBW if she weighs less than 2,500 grams at birth.} Moreover, LBW is associated with a higher risk of death within one year of birth. For these reasons, birth weight is considered the primary measure of a baby’s health and is often the direct target of health policies. Thus, understanding what causes LBW is crucial.

The impact of maternal smoking on LBW has received considerable attention in the literature, for it is regarded as one of the most significant and modifiable risk factors. Several studies consistently find that smoking during pregnancy causes lower average birth weight, with estimated ATEs ranging between -600 and -100 grams almond2005costs, abrevaya2006estimating. As for effect heterogeneity, it is now well understood that the effects are increasingly negative with the mother's age abrevaya2015estimating, lee2017doubly, zimmert2019nonparametric, fan2022estimation. Finally, treatment heterogeneity has also been investigated. cattaneo2010efficient and bodory2022high consider different smoking intensities as different treatments and show that higher smoking intensities lead to more negative effects. heiler2021effect find that heterogeneous effects can be partly explained by different smoking behaviors of ethnic and age groups.

Constructing the sequence of groupings

We split the sample into a training sample and an honest sample of equal sizes. We estimate the CATEs by fitting an honest causal forest athey2019generalized.\footnote{\ Figure (ref) in Appendix (ref) displays the estimated CATEs sorted by magnitude, along with their corresponding $95\%$ confidence intervals. Almost all predicted effects are negative and statistically different from zero at the $5\%$ significance level.} We then construct the set of admissible groupings by approximating the estimated CATEs with a decision tree.\footnote{\ Tree construction uses the R package rpart with default settings, except for the complexity parameter cp---which controls the minimum reduction in cost–complexity risk required for a split---which we set to $0.02$ (the default is $0.01$) to obtain a shallower initial tree and improve plot readability. As an alternative readability constraint, one may limit the maximum depth of the tree lee2021discovering, bargagli2022heterogeneous.} We perform these operations utilizing only observations from the training sample.

Once the tree is constructed, we use observations from the honest sample to estimate the GATEs by constructing and averaging the doubly-robust scores in ((ref)). Honest regression forests and $5$-fold cross-fitting are employed to estimate the necessary nuisance functions.

figure[figure omitted — 490 chars of source]

Figure (ref) displays the resulting tree. At the root, we see the estimated ATE of $-211$ grams. As detailed in Section (ref), the algorithm then chooses the splitting variable and point that most reduce the within-leaf variation of the predicted CATEs---i.e., the split that maximizes between-group heterogeneity---and repeats this recursively in the child nodes. The observations are initially split into two groups: non-first-born children (root's left child) and first-born children (root's right child). This division, among all possible two-group splits, maximizes heterogeneity in treatment effects. The first group, which includes $58\%$ of the units, has an estimated GATE of approximately $-236$ grams, while the second group, representing $42\%$ of the units, has an estimated GATE of about $-176$ grams.

The non–first-born branch is further split by mother’s age (younger vs. older), yielding two new groups: non–first-born children with older mothers ($47\%$ of the sample; estimated GATE $\approx-249$) and with younger mothers ($11\%$; estimated GATE $\approx-183$). Both groups are further partitioned by maternal race (Black vs. non-Black), and one of the resulting nodes is split again by the number of prenatal care visits. In total, this produces six terminal groups.

figure[figure omitted — 495 chars of source]

The sequence of optimal groupings is then constructed by progressively aggregating the two subgroups that yield the smallest loss in explained heterogeneity. Figure (ref) visually illustrates this procedure.\footnote{\ Figure (ref) in Appendix (ref) reports cross-validated risk along the pruning path as a function of the number of leaves and of the cost–complexity parameter $\alpha$.} Reading the figure from left to right and top to bottom, each panel corresponds to a grouping in the sequence, derived by collapsing the node that minimizes the loss in explained heterogeneity. The figure clearly demonstrates the consistency of the results across the different granularity levels, highlighting how the generated sequence enables a coherent exploration of the trade-off between interpretability and the discovery of more granular heterogeneity.

Discussion

\begingroup {8pt}

table[table omitted — 1,676 chars of source]

\endgroup

We investigate whether systematic effect heterogeneity is present by examining whether distinct subgroups exhibit different reactions to the treatment.\footnote{\ Looking at the distribution of the estimated CATEs is not an effective strategy for this task, as high variation in predictions due to estimation noise does not necessarily imply heterogeneous effects.} For this purpose, we select the optimal grouping composed of five groups and estimate model ((ref)) using only observations from the honest sample. Table (ref) reports point estimates and $95\%$ confidence intervals.\footnote{\ We report five groups to provide a practical summary that balances interpretability and detail. Results are robust to adjacent granularities: Tables (ref) and (ref) in Appendix (ref) show partitions with four and six groups, respectively, yielding similar GATE patterns and conclusions. This robustness is expected because the tree-pruning step produces a nested sequence of partitions (see Section (ref)), thus implying that nearby values of $\alpha$ generate closely related groupings and similar GATE patterns.} The estimated GATEs exhibit substantial differences, ranging from $-252$ grams for the most affected group (Leaf 1) to $-151$ grams for the least affected group (Leaf 5). All estimates are negative and statistically different from zero at the $5\%$ significance level.

Table (ref) also reports the GATE differences across all pairs of groups, along with $p$-values testing the null hypothesis that each difference equals zero. To account for multiple hypothesis testing, we adjust the $p$-values using the procedure of holm1979multiple. Many groups exhibit significant differences. For example, the GATE differences between Leaf 1 and all other groups are both large and statistically significant (at the $10\%$ confidence level for Leaf 2 and $5\%$ for the other groups). Additionally, the $p$-values for the differences between Leaf 5 and Leaf 2 or Leaf 3 are just above the $10\%$ significance level. For other pairs, GATE differences are small, and we fail to reject the null hypothesis at any conventional confidence level. Overall, these findings provide evidence of systematic heterogeneity in treatment effects.

\begingroup {8pt}

table[table omitted — 2,253 chars of source]

\endgroup

To understand the factors driving this heterogeneity, we examine how treatment effects relate to observable covariates by analyzing how the average characteristics of units vary across subgroups chernozhukov2017generic.\footnote{\ Another approach is to assess which variables the tree-growing process used to construct groups and measure their relative importance. However, we should not conclude that covariates not used for splitting are unrelated to heterogeneity, because if two covariates are highly correlated, trees generally split on only one of them.} Table (ref) reports the average values of selected covariates across Leaves 1–5 (see Table (ref) for the remaining covariates). The least affected group comprises children born to younger parents, suggesting more negative effects at higher parental ages. This finding is consistent with previous research abrevaya2015estimating, zimmert2019nonparametric. In contrast, the most affected group has higher parental educational attainment, likely due to older parents having had more time to study. This group also includes parents who attended more prenatal visits and were more likely to have their first visit in the first trimester of pregnancy. This may reflect mothers with problematic pregnancies self-selecting into more frequent and earlier prenatal visits.

Conclusion

This paper introduces a methodology for constructing heterogeneous subgroups that enables a flexible and coherent exploration of the trade-off between interpretability and the discovery of more granular heterogeneity. The proposed methodology constructs a sequence of groupings, one for each level of granularity. We show that each grouping features an optimality property and that the sequence is nested. We also show how the proposed methodology can be combined with honesty athey2016recursive and debiased machine learning procedures semenova2021debiased to conduct valid inference about the GATEs.

We compare the performance of aggregation and causal trees athey2016recursive using both theoretical arguments and an empirical Monte-Carlo study lechner2013sensitivity, huber2013performance. Our simulation shows that aggregation trees substantially improve the root mean squared error of treatment effects due to lower variance resulting from a more robust splitting strategy.

We apply the proposed methodology to revisit the impact of maternal smoking on birth weight almond2005costs, cattaneo2010efficient. The analysis finds evidence of systematic heterogeneity driven by parental and birth-related characteristics.

Acknowledgements

I especially would like to thank Franco Peracchi for feedback and suggestions. I am also grateful to Hannah Busshoff, Elena Dal Torrione, Matteo Iacopini, Michael Knaus, Michael Lechner, Giovanni Mellace, Seetha Menon, Tommaso Proietti, seminar participants at University of Rome Tor Vergata and SEW-HSG research seminars, and conference participants at the WEEE 2022, the 1st Rome Ph.D. in Economics and Finance Conference, and the ICEEE 2023 for comments and discussions. Matias Cattaneo generously provided the data used in the empirical illustration of this paper.

The R package for implementing the methodology developed in this paper is available on CRAN at \href{https://CRAN.R-project.org/package=aggTrees}{https://CRAN.R-project.org/package=aggTrees}. The associated vignette is at \href{https://riccardo-df.github.io/aggTrees/}{https://riccardo-df.github.io/aggTrees/}.

Declaration of Interest Statement

The author reports there are no competing interests to declare.

Data Availability Statement

The data that support the findings of this study are subject to third-party restrictions. They were obtained under license from a scholar who does not permit public dissemination. Researchers interested in replicating or extending this work may contact the corresponding author. Access to the data will require permission from the third party that provided them.

\singlespacing \newrefcontext[sorting = nty] \printbibliography