EconBase
← Back to paper

Causal Inference for Aggregated Treatment

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

168,935 characters · 30 sections · 76 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Causal Inference for Aggregated Treatment

singlespace\begingroup \footnotetext{ $\dagger$: [email removed], $\ddagger$: [email removed], $\S$: [email removed], \textparagraph: [email removed]. All authors are affiliated with the John Munro Godfrey, Sr. Department of Economics, University of Georgia. } \endgroup
abstract\begin{singlespace} We study causal inference when the treatment variable is an aggregation of multiple sub-treatments. Researchers often report marginal effects for the aggregated treatment, implicitly assuming the target parameter corresponds to a well-defined average of sub-treatment effects. We show that, even under ideal conditions such as random assignment, the weights underlying this average have key undesirable properties: they are not unique, can be negative, and these issues become exponentially more likely as the number of sub-treatments increases and their support grows. We propose diagnostics to detect these problems and introduce alternative approaches to circumvent them, depending on whether sub-treatments are observed. \end{singlespace}

JEL Codes: C18, C21, C51, C81

Keywords: Aggregated Treatments, Causal Inference, Compound Treatments, Incongruency, Sub-treatments, Versions of Treatment.

\onehalfspacing

Introduction

Causal inference requires asking questions that are precise enough to correspond to well-defined interventions (Rubin2005). Yet in many applied settings, data limitations and the need to streamline the narrative force researchers to pose causal questions at a level that is too vague to satisfy this requirement. When the treatment is only loosely defined---when it collapses multiple distinct versions into a single category---the causal question itself becomes ambiguous. In such cases, the concern goes beyond identification, pointing to a more fundamental conceptual limitation: the treatment variable may not correspond to any single coherent intervention at all. Yet, for lack of better options, researchers do estimate such causal effects in the hope that they capture something informative. While many implicitly recognize the limitations of this practice, there is no shared language or theoretical framework to precisely articulate the challenges inherent to this approach, or to guide interpretation of the resulting effects. In this paper, we study the consequences of this widespread practice, aiming to provide an early step in this direction.

Specifically, we study the setting where the researcher estimates causal effects using a treatment variable that aggregates multiple versions of the treatment, or sub-treatments. In these settings, the true causal drivers---these different sub-treatments that represent different underlying interventions---may be numerous, non-mutually exclusive, and heterogeneous in their effects, while the aggregated treatment variable serves as a summary measure that simplifies analysis but does not itself cause the outcome. In these scenarios, researchers often interpret coefficients on aggregated treatments as marginal causal effects, implicitly assuming that these parameters represent well-defined averages of sub-treatment effects---i.e., specific combinations of precisely defined interventions.

This assumption, however, warrants scrutiny. In this paper, we show that the marginal effects of an aggregated treatment can be difficult to interpret. Even under ideal conditions for causal inference, such as random assignment, the reported treatment effects generally correspond to weighted averages of sub-treatment effects with weights that are not unique and may be negative. Moreover, the negative weights issue becomes more prevalent as the number of sub-treatments or the support of each sub-treatment grows. Both issues emerge because of heterogeneous effects across different sub-treatments. We show that negative weights arise due to incongruent comparisons---those where a marginal increase in the level of the aggregated treatment involves a decrease in the level of at least one sub-treatment. The presence of negative weights can, in principle, lead to an aggregated treatment effect that is negative, even when all sub-treatment effects are positive.

The concern with non-unique weights is that different combinations of sub-treatment effects can yield the same estimated aggregate effect. In practice, this means the estimated aggregate effect reflects the impact of one particular mix of interventions, but the specific mix that produced that estimate is not generally identified. Without diagnostic tools to uncover which mix of interventions the data actually reflect, there is a risk that the estimate will be misinterpreted to justify interventions that are not supported by the data. This raises an external validity concern, but of a different kind: it arises from extrapolating the effect from one intervention to another, rather than from the standard problem of extrapolating the effect of a given intervention from one population to another.

If using an aggregated treatment yields an average treatment effect with non-unique and possibly negative weights, what should one do to improve causal inference in such a situation? We offer diagnostics that the researcher can use to convey the extent of incongruency issues in their application. We also propose alternative estimands that avoid these issues altogether. When sub-treatment data is available---that is, when the data is sufficiently rich to plausibly assume that the actual intervention components are observed---researchers can construct weighting schemes that restrict attention to congruent comparisons alone. In contrast, when sub-treatment data is unavailable, we show how researchers in this case can still circumvent all issues above by focusing on certain alternative non-marginal effect parameters.

These aggregation issues arise frequently in practice. Researchers may aggregate treatments due to data limitations, or to increase statistical precision, or to simplify their empirical strategy, or even to streamline narrative exposition. In labor economics, studies of the effects of minority or immigrant share frequently collapse distinct ethnic or national-origin groups into a single category (CardEtAl2008, BohlmarkWillen2020, Lowe2021). Similarly, peer effects research typically defines treatment using a single aggregated statistic of the classroom, such as the average GPA or SAT score (Sacerdote2001, CarrellEtAl2013). In urban economics, crime measures often aggregate qualitatively different sub-categories such as thefts, robberies, and homicides (GreenbaumTita2004, TitaEtAl2006, IhlanfeldtMayock2010, MejiaRestrepo2016), while research in environmental economics regularly treats exposure to natural disasters as a single treatment, despite substantial differences between earthquakes, floods, and wildfires (BoustanEtAl2020, ZhaoEtAl2022). In health economics, treatment variables often include composite indices of health, behavior, or genetic risk---each bundling multiple sub-components that may interact in complex ways (FinkelsteinTaubmanEtAl2012, AllcottEtAl2019, BarthEtAl2020, HoumarkEtAl2024). Time use studies similarly group distinct activities under broad labels like “exercise,” “leisure,” or “enrichment” (FioriniKeane2014, GCaetanoEtAl2019, JurgesKhanam2021, CCaetanoEtAl2024).

In fact, the practice of using an aggregated measure of the treatment variable is so common that we often take it for granted in most empirical work. For instance, consider one of the most widely used treatment variables in applied econometrics: years of schooling. This variable aggregates fundamentally different educational experiences---spanning institutions, curricula, teaching methods, peer groups, and levels of engagement---across the life of the individual. The same value on this variable may correspond to radically different combinations of sub-treatments across individuals. Moreover, increasing the number of years of schooling is not a well-defined intervention, because it never occurs in isolation: it always takes place within a particular context, where specific features of the schooling experience are altered.

A common initial reaction to our theoretical results is that they seem too damning for empirical work in causal inference. If even seemingly straightforward interventions become ill-defined once aggregation issues are taken seriously, it may appear that the very enterprise of causal inference is at risk. A central contribution of this paper, however, is to show that negative-weight concerns can be resolved by adopting a non-marginal interpretation of effects, which remains well-defined under plausible assumptions. Moreover, when marginal effects are unavoidable, another key contribution of the paper is to provide empirical diagnostics that both assess the severity of the problem and clarify the weakest assumptions under which a marginal interpretation remains valid. Finally, by raising awareness of this underappreciated problem and providing tools to improve the interpretation of the results, this paper helps ensure that causal effects from one intervention are interpreted appropriately before being extrapolated to inform decisions about another intervention. This also underscores the value of collecting more detailed datasets that capture features of the aggregated treatment.

The problem we consider is related to the Stable Unit Treatment Value Assumption (SUTVA), a foundational concept that appears throughout the causal inference literature (Rubin1980). We maintain the first part of SUTVA, often referred to as no interference or no contamination: that potential outcomes for unit $i$ do not depend on the treatment assignments of other units. The second part of SUTVA, often referred to as no hidden versions of treatments, states that different units do not experience different versions of the same treatment, or different interventions (Rubin1980, VanderWeeleHernan2013, imbens-rubin-2015). If one operates at the level of the aggregated treatment, then the issues that we highlight correspond to a violation of the second part of SUTVA, as different combinations of sub-treatments can generate the same aggregate treatment amount but different outcomes. It is easy to see in the examples above that the potential outcome would not be a well-defined function of the aggregated treatment variable. For instance, the same individual with the same number of years of schooling would likely obtain very different outcomes under different sub-treatments (e.g., graduating from a top-ranked vs. a lower-ranked university). The second component of SUTVA has not been studied nearly as much as the first part of SUTVA. To the best of our knowledge, the only studies discussing violations of the second condition of SUTVA mainly come from the Epidemiology literature and focus mostly on mediation analysis (cole2009consistency, VanderWeele2009, HernanVanderWeele2011, LaffersMellace2020).\footnote{In our setting, the outcome is influenced only by the sub-treatments themselves. The aggregated treatment is merely an aggregated summary of the underlying vector of sub-treatments, rather than a well-defined intervention. This contrasts with mediation settings, where the treatment (analogous to our aggregated treatment) is often a well-defined intervention that affects the outcome through distinct channels or mediators (analogous to our sub-treatments).} One exception is VanderWeeleHernan2013, who mostly study mediation, but also consider the context of ex post coarsening of the treatment variable (e.g., transforming a multivalued sub-treatment variable into a binary treatment variable), and show that such practice leads to a violation of SUTVA.\footnote{An example of coarsening is discussed in the context of the Tennessee STAR experiment on class size, where AdusumilliEtAl2025 show that the “small class” status corresponded to different actual class sizes across schools, reflecting local implementation constraints.} To the best of our knowledge, no paper has yet considered aggregations beyond coarsening, or provided estimands that are robust to violations of the second part of SUTVA.

Our paper suggests that one can still identify meaningful causal effects in a scenario where the second component of SUTVA is violated for the treatment variable the researcher uses. However, we require that the underlying sub-treatments satisfy SUTVA---i.e., each sub-treatment is a well defined intervention. This requirement implies that targeting marginal estimands relies on an additional assumption: that the sub-treatment observed in the data is sufficiently granular for SUTVA to plausibly hold. Still, researchers can continue to interpret (non-marginal) causal effects meaningfully under violations of SUTVA---provided that sub-treatments satisfying SUTVA are conceptually well-defined, even if they are unobserved.

Our work also relates to the vast literature on treatment effect heterogeneity, which studies how the causal effect of a given treatment can vary across units (e.g., rubin-1974, Holland1986, heckman-smith-clements-1997, Rosenbaum2002, imbens-rubin-2015). Our findings suggest that when treatments are aggregated, what appears to be treatment effect heterogeneity may instead reflect sub-treatment heterogeneity---that is, heterogeneity arising from distinct underlying components of the treatment itself---which is a violation of the no hidden versions of treatments component of SUTVA. Returning to the years of schooling example, the effect of one additional year of schooling may obscure heterogeneity in educational experiences even for the same individual. For instance, the additional year could reflect a counterfactual enrollment in a five-year major such as engineering, rather than a four-year major like economics. In this case, what looks like a marginal “year effect” partially reflects differences in content, difficulty, and the credential ultimately earned. Crucially, this heterogeneity is not across individuals, but within the same individual under different sub-treatments. As a result, the relevant unit of potential outcome variation is not the individual alone, but the individual–sub-treatment pair. Although some empirical work has grappled with these issues---contrasting “heterogeneous treatments” and “heterogeneous treatment effects” (e.g., Lechner2002, PlescaSmith2007, mccall2016government, CaetanoMaheshri2018, Smith2022)---to our knowledge, there is no formal framework for identifying or interpreting causal effects in such settings. One exception is HeilerKnaus2025, which addresses related concerns about heterogeneous treatments by showing how group-level heterogeneity analyses can be misleading in this context, and provides a decomposition to separate causal effect heterogeneity from spurious differences driven by differential assignment to versions of treatment.

We illustrate the issues due to aggregation and our approaches to circumvent them in an application concerning the effects of enrichment activities on children's skills, based on CCaetanoEtAl2024. The treatment variable---time per week spent on enrichment activities---aggregates time spent on homework, music lessons, and sports, among other extracurricular activities. This is an application where sub-treatments reflecting detailed activities of the children are observed, which allows us to diagnose how much of an issue incongruency is if the aggregated treatment is used as the treatment variable, as done in that paper. In this application, incongruency is important, as some aggregate marginal effects put at least 30-40% weight on incongruent comparisons. Estimates for alternative parameters that we propose, which exclude incongruent comparisons, are roughly twice as large in magnitude.

The paper is organized as follows. Section (ref) describes the causal setting that we consider, establishes the necessary notation, and formally defines both congruency and the relevant target parameters. Section (ref) presents the main challenges for identification and for interpreting these types of parameters. Section (ref) proposes alternative parameters that rectify the challenges of aggregated treatment, including those that do not require observation of sub-treatment data. Section (ref) delivers an empirical illustration from the time-use literature, highlighting both the problems introduced by aggregated treatment and the solutions we propose. Finally, Section (ref) offers some concluding remarks. The Appendix provides additional results and proofs.

Aggregated Treatment Setting

This section (i) provides notation and formalizes the setting that we consider, (ii) introduces notions of congruent and incongruent sub-treatment vectors, and (iii) defines our main target parameters.

Notation and Setup

We consider a setting where a researcher is interested in understanding the relationship between an outcome and a treatment. In our application, the outcome is a standard measure of noncognitive skills, and the treatment variable is a measure of time the child spends on enrichment activities. We denote the outcome variable by $Y$. The treatment variable is comprised of different types of enrichment activities, which can vary in amount across different units. Let $S_{ik}$ denote the amount of component $k$ of the treatment that unit $i$ experiences. We refer to $S_{k}$ as the $k$th sub-treatment, and $\mathfrak{S}_{k}$ denotes the support of the $k$th sub-treatment. For example, if “doing homework” is the $k$th version of enrichment activity, then $S_{ik}=2$ for children who do two hours of homework. Next, define $S_i = (S_{i1}, S_{i2}, \ldots, S_{iK})$ where $K$ denotes the total number of versions of the treatment. We refer to $S_i$ as a unit's sub-treatment vector. Let $\mathcal{S}$ denote the support of $S$. We consider the case where the sub-treatments are discrete and share a common support---we consider this case to focus the exposition of the paper and note that both of these conditions could be relaxed without substantively changing our results below.

We also define the aggregated treatment variable $D_i = A(S_i)$ where $A(s)$ is an aggregation function that maps sub-treatments to a scalar value of the aggregated treatment variable. Let $\mathcal{D}$ denote the support of $D$ and $\mathcal{D}_{>0} = \mathcal{D} \setminus \{0\}$ denote the support of $D$ excluding $D=0$. To keep the discussion concrete, we often focus on the case where $A(s) = \sum_{k=1}^K s_k$---where the aggregated treatment variable adds up all of the underlying components of the sub-treatment vector. However, other types of aggregation are possible as well; some of our results hold immediately for any aggregation scheme, while others would require minor modifications. Two important, immediate properties of the aggregated treatment are that: (i) it is fully determined by the sub-treatments, and (ii) different sub-treatments can lead to the same value of the aggregated treatment.

The discussion above provides a definition of aggregated treatment. Based on the aforementioned properties of aggregation, we define the set

align*[align* omitted — 70 chars of source]

where $A(\cdot)$ is the aggregation rule. Thus, $\mathcal{S}_d$ is the set of distinct sub-treatment vectors that lead to a particular value of the aggregated treatment ($D=d$). We refer to $\mathcal{S}_d$ as the aggregation set corresponding to aggregated treatment $d$. We use the terminology sub-treatment group $s_d$ to refer to the set of units that experience sub-treatment vector $s_d$.

exampleTo fix ideas in the discussion below, consider a simplified version of our application where there are three different types of enrichment activities: (1) homework, (2) music lessons, and (3) sports. Children can participate in any of these enrichment activities. To simplify the discussion in the example, here we binarize each sub-treatment by only keeping track of whether or not the child participates in each enrichment activity, although the theory for our paper additionally allows for sub-treatments to be multivalued. The aggregated treatment $D$ indicates the total number of enrichment activities that a child does. Children with the same value of the aggregated treatment, however, can experience different combinations of sub-treatments. For example, one student who does homework and music lessons, and another student who does homework and sports, both participate in two enrichment activities. We can define $\mathcal{S}_d$ for any possible value of $d$, \begin{align*} \mathcal{S}_0 = \{(0,0,0)\} \mathcal{S}_1 = \{(1,0,0), (0,1,0), (0,0,1) \} \mathcal{S}_2 = \{(1,1,0), (1,0,1), (0,1,1)\} \mathcal{S}_3 = \{(1,1,1)\}, \end{align*} where, for example, $(1,1,0) \in \mathcal{S}_2$ indicates the sub-treatment vector of participating in homework and music lessons but not playing sports.

Because we are interested in causal effects, we also define potential outcomes. Let $Y_i(s)$ denote the outcome that unit $i$ would experience under sub-treatment vector $s$. The observed outcome is equal to the potential outcome corresponding to the observed sub-treatment vector; that is, $Y_i = Y_i(S_i)$. Implicit in this expression is that we impose SUTVA (parts 1 and 2) at the level of the sub-treatment---we formalize this in Assumption (ref) below. Notice that potential outcomes are only defined at the sub-treatment level. We do not define potential outcomes in terms of the aggregated treatment variable, as the aggregated treatment variable itself is non-causal and contingent upon the aggregation scheme.

Our reading of the literature suggests that common empirical practice is to regress the outcome on the aggregated treatment $D$; i.e.,

align[align omitted — 68 chars of source]

and to interpret $\alpha_1$, the coefficient on $D$, in terms of marginal effects.\footnote{A representative example comes from MejiaRestrepo2016, which studies the effects of property crime (their treatment variable) on different types of household expenditure. They construct an aggregate measure of property crime by averaging the rate of robberies and burglaries (their sub-treatments). Robberies and burglaries are distinct crimes. The main difference is roughly that robberies involve directly stealing from someone while burglaries involve entering a structure to steal (see fbi-burglary-2018,fbi-robbery-2018 for more details). Their main results involve interpreting coefficients on this aggregated treatment variable in terms of marginal effects. For example, they write: “conditional on controls, the coefficient of crime on total visible and non-steal-able consumption is negative and significant at the 5% confidence level. In particular, a 10% increase in property crime is associated with a 1.45% decline in the consumption of visible and non-stealable goods...”} Writing $D_i$ in terms of sub-treatments in the above regression, we have that

align*[align* omitted — 85 chars of source]

which suggests that homogenous effects across different sub-treatments is an important implicit assumption for this regression to be able to recover the marginal effect of the sub-treatments on the outcome. One of our main goals below is to understand how to interpret marginal effects of $D$ in settings where there can be heterogeneous effects of the sub-treatments.

Congruent and Incongruent Sub-treatment Vectors

The notion of a marginal effect is more complicated in applications with sub-treatments. In this section, we distinguish between congruent and incongruent sub-treatment vectors, which is then useful for precisely defining marginal effect parameters in the next section. Define the marginal set, i.e., the set of neighboring aggregation sets, indexed by $d$, as

align*[align* omitted — 100 chars of source]

for all $d \in \mathcal{D}_{>0}$. That is, $\mathcal{M}(d)$ represents the marginal set of $K$-tuple sub-treatment vector pairs whose $L_1$ norm equals either $d$ or $d-1$.

definition[Congruent and Incongruent Sub-treatment Vectors] For the sub-treatment vectors $(s_d,s_{d-1}) \in \mathcal{M}(d)$, define the binary congruence relation $ \succ^{\hspace{-0.2em} \scalebox{0.7}{\pmb{+}}} $ as: $s_d \succ^{\hspace{-0.2em} \scalebox{0.7}{\pmb{+}}} s_{d-1} \; \mathrm{if} \; s_d=s_{d-1} + 1_k$ for some $k$, where $1_k$ is the unit vector with $k^{th}$ element equal to one and zero otherwise. If $s_d \succ^{\hspace{-0.2em} \scalebox{0.7}{\pmb{+}}} s_{d-1}$, then we say that $s_d$ and $s_{d-1}$ are congruent; otherwise, we say that they are incongruent.

Definition (ref) defines congruent and incongruent sub-treatment vectors. In particular, two sub-treatment vectors $s_d$ and $s_{d-1}$ are congruent if they correspond to neighboring aggregation sets, and the value of each element of the sub-treatment vector $s_d$ is equal to the value of the corresponding element of vector $s_{d-1}$, except for one element. Sub-treatment vectors $s_d$ and $s_{d-1}$ are incongruent if they are from neighboring aggregation sets, but the value of more than one element is different.

namedexample{\ref*{ex:enrichment} (continued)} Consider the sub-treatment $(1,0,0) \in \mathcal{S}_1$ (i.e., this is the sub-treatment that amounts to doing homework but not doing music lessons or sports). $(1,1,0)$---doing homework and music lessons but not sports---is congruent with $(1,0,0)$. $(1,0,1)$---doing homework and sports but not music lessons---is also congruent with $(1,0,0)$. $(0,1,1)$---doing music lessons and sports but not homework---is incongruent with $(1,0,0)$.

It is also helpful to define the sets of congruent and incongruent sub-treatment vectors. In particular, define

align*[align* omitted — 149 chars of source]

for all $d \in \mathcal{D}_{>0}$, where $\mathcal{M}^{+}(d)$ represents the set of congruent sub-treatments vectors, and $\mathcal{M}^{-}(d) := \mathcal{M}(d) \setminus \mathcal{M}^+(d)$, which represents the set of incongruent sub-treatment vectors.

Target Parameters

This section defines the main target parameters that we consider in the paper. We primarily focus on different weighted averages of marginal changes across sub-treatment vectors, since it is common empirical practice in this setting to interpret or report the coefficient on the aggregated treatment variable in terms of marginal effects. First, for $(s_d,s_{d-1}) \in \mathcal{M}(d)$, define the marginal average treatment effect on the treated ($\textrm{MATT}$) as:

align*[align* omitted — 88 chars of source]

which is the causal effect of moving from sub-treatment vector $s_{d-1}$ to $s_d$ for sub-treatment group $s_d$.\footnote{We primarily focus on on-the-treated type parameters because it is simpler to provide natural weighting schemes for some of the aggregated parameters that we consider in this section relative to unconditional parameters. In addition, unconditional parameters also tend to require stronger identification assumptions in our setting. That said, extending our arguments to target unconditional parameters seems straightforward.} $\textrm{MATT}(s_d,s_{d-1})$ is defined for all $(s_d,s_{d-1})$, regardless of whether or not $s_d$ and $s_{d-1}$ are congruent. Sometimes we use the notation $\textrm{MATT}^+(s_d,s_{d-1})$ to indicate a disaggregated marginal average treatment effect on the treated of congruent sub-treatments $(s_d,s_{d-1}) \in \mathcal{M}^{+}(d)$. Likewise, we sometimes use the notation $\textrm{MATT}^-(s_d,s_{d-1})$ to indicate a disaggregated marginal average treatment effect on the treated of incongruent sub-treatments $(s_d,s_{d-1}) \in \mathcal{M}^{-}(d)$.

Given our interest in aggregation, next we introduce a parameter that is a weighted average of congruent $\textrm{MATT}^+$'s:

align[align omitted — 171 chars of source]

where $w^+$ is some weighting function that satisfies $w^+(s_d,s_{d-1}) \geq 0$ for any $(s_d,s_{d-1}) \in \mathcal{M}^+(d)$, and $\displaystyle \sum_{(s_d,s_{d-1}) \; \in \; \mathcal{M}^+(d)} w^+(s_d,s_{d-1}) = 1$. Following the terminology of blandhol-bonney-mogstad-torgovitsky-2025, we refer to parameters like $\textrm{AMATT}^{+}_{w^+}(d)$ that are weighted averages (with all non-negative weights) of $\textrm{MATT}^+(s_d,s_{d-1})$ as weakly causal.\footnote{In Section (ref) , we consider more specific aggregated parameters (i.e., with a specific weighting scheme rather than allowing for any weighting scheme satisfying the criteria above). However, we note here that, with an aggregated treatment, this can introduce some additional complications. Therefore, we defer this discussion to later in the paper.}

Finally, although our main interest is in the effects of congruent sub-treatments, in some cases, a researcher may be interested in the effects of different sub-treatment vectors for a fixed amount of aggregated treatment. We refer to these as substitution average treatment effects on the treated ($\textrm{SATT}$).\footnote{$\textrm{SATT}$'s have a precise mathematical definition---a local tradeoff between two sub-treatments at a fixed amount of aggregated treatment. We call this a substitution effect as, in many applications (e.g., our application on enrichment activities), this parameter coincides with an intuitive notion of substituting between different sub-treatments. However, this intuition may not apply in all applications (e.g., it is unnatural to think of “substitution” in the natural disaster applications mentioned in the introduction); still, regardless of the exact terminology for a particular application, these types of terms continue to be relevant for our results below.} For $s_d$ and $s_d'$ both in $\mathcal{S}_d$, define

align*[align* omitted — 81 chars of source]

such that $s_d = s_d' + 1_j - 1_l$, where $1_j$ and $1_l$ denote the unit vector for the $j^{th}$ and $l^{th}$ coordinates. In other words, there is a unit exchange between the $j^{th}$ and $l^{th}$ sub-treatments from $s_d$ to $s_d'$. Later, we show that there is often an interesting connection between incongruent $\textrm{MATT}$'s and $\textrm{SATT}$'s.

Challenges to Identification

This section outlines several important difficulties that arise in applications with aggregated treatments, even under otherwise ideal conditions for causal inference. The discussion in this section is geared towards interpreting the marginal effect of the aggregated treatment on the outcome in terms of the underlying sub-treatments and the complications that this can induce. The results are most relevant for applications where the sub-treatments themselves are not observed, but would continue to apply in applications where the researcher observes the sub-treatments yet still decides to use an aggregated treatment. Many of the expressions below include terms that condition on the sub-treatment group---if the sub-treatments are not observed, then these terms would not be identified, though they are still useful to consider as underlying building blocks of the aggregate marginal effect.

Causal Framework

We begin by formalizing what we mean by sub-treatments. In our paper, the key difference between the sub-treatment vector and the aggregated treatment is that the sub-treatments satisfy the second part of SUTVA, often referred to as no hidden versions of treatment, while the aggregated treatment does not (see Rubin1980, RobinsGreenland2000, VanderWeele2009, HernanVanderWeele2011, imbens-rubin-2015, and Hernan2016 for more discussion about SUTVA). In particular, we make the following assumption.

assumption[No Hidden Versions of Sub-treatments relative to $S$] If unit $i$ experiences sub-treatment vector $s$, then its observed outcome equals the potential outcome corresponding to that sub-treatment vector; i.e., for all $s \in \mathcal{S}$, \begin{align*} S_i = s \implies Y_i = Y_i(s), \quad where $Y_i(s)$ is a well-defined function of $s$. \end{align*}

Throughout the paper, we maintain that the sub-treatments satisfy Assumption (ref). $Y_i(s)$ being a well-defined function of $s$ rules out “hidden sub-versions” of the sub-treatments. This condition implies that knowing $s$ pins down a unit's potential outcome from experiencing that sub-treatment.\footnote{Note that this assumption represents SUTVA entirely (for multivariate treatment $S$), since it also implicitly assumes the first part of SUTVA---it rules out the possibility that the treatment level of other units $j\neq i$ affects the potential outcome of unit $i$.} In our running example, where the sub-treatments are homework, music, and sports, it says that further sub-dividing the sub-treatments would not change the potential outcomes; for example, the potential outcomes from doing 2 hours of homework with mom, playing 1 hour of piano, and playing 1 hour of soccer are the same for all units as the potential outcomes from doing 2 hours of homework with dad, playing 1 hour of guitar, and playing 1 hour of basketball. On the other hand, in our paper, SUTVA for the aggregated treatment is generally violated in that $Y_i(d)$ is not a well-defined function of $d$---in our example, if $D_i=d$, a child's potential outcome still depends on which versions (homework, music, or sports) the child actually experienced.

Throughout the paper, we maintain that the $K$ researcher-specified sub-treatments satisfy Assumption (ref). In Appendix (ref), we provide some practical guidance for defining the sub-treatments in a particular application.

Next, we introduce the primary assumption for identifying various aggregated and disaggregated average treatment effect parameters.

assumption[No Selection] For any $s \in \mathcal{S}$, $Y(s) \mathrel{\mbox{\(\perp\!\!\!\perp\)}} S$.

Assumption (ref) states that the distribution of potential outcomes of any sub-treatment vector is independent of the actual sub-treatments experienced. It would hold by construction if the sub-treatments were randomly assigned. The main implication of this assumption is that, for any two sub-treatments $s,s' \in \mathcal{S}$, $\mathbb{E}[Y(s) | S=s'] = \mathbb{E}[Y(s) | S=s] = \mathbb{E}[Y|S=s]$, which would be identified in applications where the sub-treatments are observed. Thus, it immediately follows that

align*[align* omitted — 184 chars of source]

under Assumption (ref) (see Propositions (ref) and (ref) in Appendix (ref)).

Although this assumption is likely to be strong in many applications, we take it as a natural baseline for our setting---it provides a starting point that is as favorable as possible for causal inference, and, hence, allows us to emphasize issues related to aggregation in the discussion below. Many of our results are expressed in terms of causal effect parameters which use Assumption (ref); however, absent Assumption (ref), versions of the issues related to aggregation that we highlight continue to apply, just without a causal interpretation. Moreover, extending our arguments to other identification strategies seems straightforward, at least in some leading cases. For example, under selection on observables, all of our identification results would go through, conditional on covariates. Similarly, our arguments can be extended to difference-in-differences identification strategies by replacing the level of the outcome with the change in outcomes over time in the assumptions and results in this section. Extending our results to other settings (e.g., instrumental variables, regression discontinuity, or bunching) might introduce additional complications, but we conjecture that versions of the issues that we point out stemming from aggregated treatment would continue to apply. Finally, we note that our results on marginal effects go through under a weaker, local version of Assumption (ref), which we discuss in more detail in Appendix (ref).

A Decomposition of Marginal Effects with Aggregated Treatment

This section contains one of our main results, which is a decomposition of the marginal effect of the aggregated treatment in terms of $\textrm{MATT}$ parameters.

theoremUnder Assumptions (ref) and (ref), and for any weighting function $w(s_d,s_{d-1})$ such that \begin{itemize} • $\displaystyle \sum_{s_d \in \mathcal{S}_d} w(s_d,s_{d-1}) = \mathrm{P}(S=s_{d-1}|D=d-1)$, • $\displaystyle \sum_{s_{d-1} \in \mathcal{S}_{d-1}} w(s_d,s_{d-1}) = \mathrm{P}(S=s_d|D=d)$, \end{itemize} \begin{align*} \mathbb{E}[Y|D=d] - \mathbb{E}[Y|D=d-1] &= \sum_{(s_{d-1}, s_{d}) \in \mathcal{M}^+(d) } w(s_d,s_{d-1}) \cdot \mathrm{MATT}^+(s_d,s_{d-1}) \\ & + \sum_{(s_{d-1}, s_{d}) \in \mathcal{M}^-(d) } w(s_d,s_{d-1}) \cdot \mathrm{MATT}^-(s_d,s_{d-1}) \end{align*}

Theorem (ref) shows that, under Assumptions (ref) and (ref), the change in the mean of $Y$ given a one unit increase in the aggregated treatment $D$ can be decomposed into a weighted average of congruent and incongruent $\textrm{MATT}$ parameters. There are several notable features of this decomposition that warrant further examination in the sections that follow. First, in Section (ref), we show that the incongruent $\textrm{MATT}^-$ parameters that appear in the proposition can be difficult to interpret. Second, the weights that satisfy the criteria in Theorem (ref) are non-unique, and, given the difficulty of interpreting incongruent comparisons, ideally, we would like there to be a compatible weighting scheme that does not put any weight on these incongruent comparisons.\footnote{The non-uniqueness of the weights in Theorem (ref) is conceptually different from all decompositions that we are aware of in econometrics that show up in other contexts, such as continuous treatments (e.g., yitzhaki-1996 and callaway-goodman-santanna-2025), regressions that include covariates (e.g, angrist-1998, sloczynski-2022, and hahn-2023), two-stage least squares (e.g., ishimaru-2024 and blandhol-bonney-mogstad-torgovitsky-2025), and fixed effects regressions (e.g., chaisemartin-dhaultfoeuille-2020, goodman-bacon-2021, sun-abraham-2021, caetano-callaway-2024). Non-uniqueness arises because we decompose the marginal effect of a more aggregated variable in terms of the marginal effects of less aggregated variables, which is an important difference relative to all of the aforementioned papers.} With this in mind, Section (ref) (i) shows that the number of incongruent comparisons grows rapidly in the number of distinct sub-treatments relative to the number of congruent comparisons; (ii) provides conditions under which it is guaranteed that there exists a valid weighting scheme that puts no weight on incongruent comparisons; (iii) characterizes settings where putting weight on incongruent comparisons is unavoidable; and (iv) shows how to test whether any of the aggregation issues mentioned in this paper are empirically relevant.

remark[Descriptive Decomposition] In Proposition (ref) in Appendix (ref), we provide a non-causal version of Theorem (ref) that does not invoke Assumption (ref) or (ref) (in fact, this is the key step in proving Theorem (ref)). Thus, even if the researcher views $\mathbb{E}[Y|D=d]-\mathbb{E}[Y|D=d-1]$ descriptively, the issues that we highlight coming from $D$ being an aggregation of the sub-treatments continue to apply.
remark[Regression] As discussed above, it is common in empirical work to estimate a regression like the one in Equation (ref) that includes an aggregated treatment variable. In Proposition (ref) in Appendix (ref), we show that $\alpha_1$, the coefficient on $D$ in the regression in Equation (ref), can be expressed as \begin{align*} \alpha_1 &= \sum_{d=1}^{\Bar{N}} \omega^{reg}(d) \cdot \Big( \mathbb{E}[Y|D=d] - \mathbb{E}[Y|D=d-1] \Big) \end{align*} where the regression weights are \begin{align*} \omega^{reg}(d) := \frac{\big(\mathbb{E}[D|D \geq d] - \mathbb{E}[D]\big) \cdot \mathrm{P}(D \geq d)}{\mathrm{Var}(D)} \end{align*} and satisfy the properties: (i) $\omega^{reg}(d) \geq 0$ for all values of $d \in \mathcal{D}_{>0}$, and (ii) $\displaystyle \sum_{d=1}^{\bar{N}} \omega^{reg}(d) = 1$, where $\bar{N}$ is the maximum on the support of $D$. The result above essentially holds using a discrete version of the argument in yitzhaki-1996. The regression weights, $\omega^{reg}(d),$ have reasonable though less than ideal properties (see Appendix (ref) for more details); more importantly, however, $\alpha_1$ is a weighted average of $\mathbb{E}[Y|D=d]-\mathbb{E}[Y|D=d-1]$---the same aggregate marginal effect that we decomposed in Theorem (ref). Thus, all of the issues about incongruent $\mathrm{MATT}^{-}$'s and non-unique weights discussed below continue to apply when using regressions that include an aggregated treatment variable.

Interpreting Incongruent Comparisons

The decomposition in Theorem (ref) in the previous section showed that $\mathbb{E}[Y|D=d] - \mathbb{E}[Y|D=d-1]$, the marginal effect of the aggregated treatment, included the incongruent $\textrm{MATT}^{-}(s_d,s_{d-1})$ parameters. How should incongruent marginal causal effect parameters be interpreted? This section provides two answers to this question. First, it shows that these incongruent parameters can be expressed in terms of a congruent $\textrm{MATT}^+$ parameter and a sequence of substitution effects, the $\textrm{SATT}$ parameters discussed above. Second, it shows that incongruent parameters can be expressed as a sequence of congruent $\textrm{MATT}^+$ parameters, but that almost half of the $\textrm{MATT}^+$ parameters in this sequence enter with negative weights. In either case, it implies that $\textrm{MATT}^-(s_d,s_{d-1})$ is hard to interpret.

Incongruent Comparisons and Substitution Effects

The following proposition re-expresses incongruent $\textrm{MATT}^{-}$ parameters in terms of a congruent $\textrm{MATT}^+$ parameter and a path-dependent sum of substitution effects.

propositionUnder Assumptions (ref) and (ref), for all $(s_d, s_{d-1}) \in \mathcal{M}^-(d)$ and for any $s_{d-1}'$ that is congruent with $s_d$, it holds that \begin{align*} \mathrm{MATT}^-(s_d,s_{d-1}) &= \mathrm{MATT}^+(s_d,s_{d-1}') + \sum_{ b=0 }^{B-1} \mathrm{SATT}(s^{(b)}_{\phi, \; d-1},s^{(b+1)}_{\phi, \; d-1}) \end{align*} where $\phi := (x^{(0)}, \ldots, x^{(B)}) \in \mathcal{C}(s_{d-1}',s_{d-1})$ represents a particular set of chained vectors from the set of sets of chained sub-treatment vectors that create a unit-exchange pathway between sub-treatment vectors $s_{d-1}'$ and $s_{d-1}$ within the same aggregation set $\mathcal{S}_{d-1}$; and $s^{(b)}_{\phi}, s^{(b+1)}_{\phi}$ are linked sub-treatment vectors that belong to the chain $\phi$.\footnote{Generally, $\mathcal{C}(s'_d,s''_d) := \bigl\{(x^{(0)}, \dots, x^{(B)}) \big| x^{(0)} = s_d',\, x^{(B)} = s_d'',\, \text{for }\,b=0,\dots,B-1: x^{(b)} \in \mathcal{S}_d, \; \|x^{(b+1)} - s_d'' \|_1 < \|x^{(b)} - s_d''\|_1 \bigr\}$, for any $d \in \mathcal{D}_{>0}$ and some $B \in \mathbb{N}$, is the set of chained sub-treatment vectors that create a unit-exchange pathway between sub-treatment vectors $s_{d}'$ and $s_{d}''$ which belong to the same aggregation set $\mathcal{S}_{d}$.}

Proposition (ref) shows that incongruent $\textrm{MATT}^-$'s can be decomposed into alternative congruent causal effect parameters and substitution effects. This shows that the aggregate treatment effect is composed of both marginal sub-treatment effects and substitution effects. Although substitution effects could be of interest in their own right, they are a different type of parameter from $\textrm{MATT}^+$; they involve substituting across sub-treatments rather than a marginal increase in one of the sub-treatments.

namedexample{\ref*{ex:enrichment} (continued)} Suppose that $s_2 = (1,1,0)$, and $s_1 = (0,0,1)$, which are incongruent, and consider $s_1' = (1,0,0)$. Then, using the argument from Proposition (ref), it holds that $\mathrm{MATT}^-(s_2,s_1) = \mathrm{MATT}^+(s_2,s_1') + \mathrm{SATT}(s_1',s_1)$. Or, in other words, the incongruent causal effect of both doing homework and music lessons relative to playing sports can be decomposed into (i) the congruent causal effect of both doing homework and music lessons relative to only doing homework and (ii) the substitution effect of doing homework relative to playing sports.

By plugging the result of Proposition (ref) into the decomposition of $\mathbb{E}[Y|D=d]-\mathbb{E}[Y|D=d-1]$ in Theorem (ref), it follows that the aggregate marginal effect is hard to interpret because it includes a mix of congruent $\textrm{MATT}^{+}$'s and substitution effects---two different types of parameters. And, for example, a positive value of $\mathbb{E}[Y|D=d]-\mathbb{E}[Y|D=d-1]$ could be mainly driven by substitution effects rather than effects of marginal increases in any of the sub-treatments.\footnote{The discussion in this paragraph is conceptually related to the well-known decomposition in goodman-bacon-2021 that, in the context of difference-in-differences identification strategies estimated using two-way fixed effects regressions, relates the coefficient of a binary treatment variable to two different types of causal effect parameters: (i) causal effects of the treatment itself and (ii) treatment effect dynamics. This is typically taken as a negative result for the two-way fixed effects regression, not because treatment effect dynamics are inherently uninteresting to study, but rather because mixing together two different types of parameters is hard to interpret. Similarly, in our context, the $\mathrm{SATT}$ parameters could be interesting to learn about, but they do not involve a marginal increase in any sub-treatment and, hence, make the aggregate marginal effect difficult to interpret.}

Incongruent Comparisons and Negative Weights

The next proposition decomposes both substitution effects and incongruent $\textrm{MATT}^-$'s into a path-dependent sum of congruent causal effect parameters.

propositionUnder Assumptions (ref) and (ref), for any $s_{d-1}, s_{d-1}' \in \mathcal{S}_{d-1}$ such that $s_{d-1} = s_{d-1}' + 1_j - 1_l$, where $1_j$ and $1_l$ are unit vectors for coordinates $j$ and $l$, and for any $s_{d}' \in \mathcal{S}_{d}$ that is congruent with both $s_{d-1}$ and $s_{d-1}'$, it holds that \begin{align*} \mathrm{SATT}(s_{d-1},s_{d-1}') &= \mathrm{MATT}^{+}(s_d',s_{d-1}') - \mathrm{MATT}^{+}(s_{d}',s_{d-1}) \end{align*} Moreover, for any incongruent parameter and chain $\phi \in \mathcal{C}(s_{d-1}',s_{d-1})$: \begin{align*} \mathrm{MATT}^-(s_d,s_{d-1}) &= \mathrm{MATT}^+(s_d,s_{d-1}') + \sum_{ b=0 }^{B-1} \mathrm{MATT}^{+}(s_{\phi, \; d}^{(b)},s_{\phi, \; d-1}^{(b)}) - \sum_{ b=0 }^{B-1} \mathrm{MATT}^{+}(s_{\phi, \; d}^{(b)},s_{\phi, \; d-1}^{(b+1)}) \end{align*}

The first part of Proposition (ref) says that any substitution effect between two sub-treatment vectors that share a unit exchange in treatment is equivalent to the difference between two congruent $\textrm{MATT}^+$'s. The second part says that an incongruent $\textrm{MATT}^-$ parameter can be decomposed into congruent $\textrm{MATT}^+$ parameters but that almost half of the $\textrm{MATT}^+$ parameters in this decomposition are included with a negative sign; this part follows from plugging the expression for $\mathrm{SATT}(s_{d-1},s_{d-1}')$ in the first part into Proposition (ref) above.

namedexample{\ref*{ex:enrichment} (continued)} Resuming the example from the previous section, let $s_2 = (1,1,0)$, $s_2' = (1,0,1)$, $s_1 = (0,0,1)$, and $s_1' = (1,0,0)$. Using the argument in the first part of Proposition (ref), it holds that $\; \mathrm{SATT}(s_1', s_1) = \mathrm{MATT}^+(s_2',s_1) - \mathrm{MATT}^+(s_2',s_1')$. In words, the substitution effect of doing homework relative to playing sports is equal to the difference between (i) the congruent causal effect of doing homework and playing sports relative to only playing sports and (ii) the congruent causal effect of doing homework and playing sports relative to only doing homework. Using the argument from the second part of the proposition, it holds that $\mathrm{MATT}^-(s_2,s_1) = \mathrm{MATT}^+(s_2,s_1') + \mathrm{MATT}^+(s_2',s_1) - \mathrm{MATT}^+(s_2',s_1')$. That is, the incongruent causal effect of both doing homework and music lessons relative to playing sports can be decomposed into (i) the congruent causal effect of both doing homework and music lessons relative to only doing homework, (ii) the congruent causal effect of doing homework and playing sports relative to only playing sports and (iii) the congruent causal effect of doing homework and playing sports relative to only doing homework; however, the congruent effect (iii) enters the decomposition negatively.

By plugging the second part of Proposition (ref) into the decomposition in Theorem (ref), it follows that the aggregate marginal effect can be fully expressed as a weighted average of congruent $\textrm{MATT}^+$ parameters. However, due to the negative signs on some $\textrm{MATT}^+$ parameters in Proposition (ref), it is evident that weights on some $\textrm{MATT}^+$ parameters can be negative.\footnote{To be clear, even if $\textrm{MATT}^+(s_d,s_{d-1})$ shows up negatively in the decomposition from being part of the chain of congruent $\textrm{MATT}^+$'s corresponding to an incongruent $\textrm{MATT}^-$, recall that it also shows up positively in the first term in Theorem (ref), and whether or not it ultimately shows up with a positive or negative weight depends on the relative magnitude of the corresponding weights. Thus, the results in this section do not indicate that negative weights on certain $\textrm{MATT}^+$ parameters necessarily occur, but rather that negative weights can occur.} Negative weights on underlying “building block” parameters have been emphasized in recent work in econometrics as being indicative of an “unreasonable” weighting scheme (e.g., chaisemartin-dhaultfoeuille-2020,mogstad-torgovitsky-2024, among others). For example, given enough heterogeneity in the congruent $\textrm{MATT}^+$'s, negative weights introduce the possibility of sign reversal, where, e.g., the $\textrm{MATT}^+$'s could all be positive but $\mathbb{E}[Y|D=d]-\mathbb{E}[Y|D=d-1]$ could be negative.

Non-unique Weights: When Is Aggregation a Problem?

The previous section highlighted that the incongruent causal effect parameters $\textrm{MATT}^-(s_d,s_{d-1})$ are difficult to interpret. In this section, we return to the other main issue in Theorem (ref): that the weights are non-unique. Let $\mathcal{W}_d$ denote the set of weighting schemes that satisfy the conditions in Theorem (ref). If there exists a weighting scheme in $\mathcal{W}_d$ that puts zero weight on all $\textrm{MATT}^{-}(s_d,s_{d-1})$, then there exists an interpretation of $\mathbb{E}[Y|D=d]-\mathbb{E}[Y|D=d-1]$ that only puts weight on $\textrm{MATT}^{+}(s_d,s_{d-1})$. This would fully bypass the problems of interpreting $\textrm{MATT}^{-}(s_d,s_{d-1})$ that we discussed above.

Implicit in the discussion above is that there can be multiple weighting schemes that satisfy the conditions in Theorem (ref). Thus, we start this section by showing that the weights are indeed non-unique.\footnote{Two exceptions are worth mentioning. First, the weights are unique when there is a single unique version of the treatment, $K=1$. In this case, for any level of the aggregated treatment, there is only one sub-treatment, so that $w(s_d,s_{d-1})=1$. Second, the weights are also unique when $d=1$ or $d=\bar{N}$. Take, for instance, the case where $d=1$, so that we are interested in $\mathbb{E}[Y|D=1] - \mathbb{E}[Y|D=0]$. Regardless of how many distinct sub-treatments there are, the only sub-treatment vector such that $d=0$ is $s=0_K$. This implies that $\mathrm{P}(S=0_K|D=0)=1$, and the only weights that satisfy the criteria for the weights in Proposition (ref) are $w(s_1,s_0) = \mathrm{P}(S=s_1|D=1)$, implying that the weights are unique. An analogous argument holds for $d=\bar{N}$ on the basis that there is only one sub-treatment vector (the one where each sub-treatment is set at its maximum value).} A leading example of weights that are always in $\mathcal{W}_d$ are the product weights $\mathrm{P}(S=s_d|D=d) \times \mathrm{P}(S=s_{d-1}|D=d-1)$. This weighting scheme necessarily implies that the aggregate marginal effect includes positive weight on incongruent comparisons. However, they are not the only weights that satisfy the criteria mentioned in the proposition, which we demonstrate by returning to our example.

namedexample{\ref*{ex:enrichment} (continued)} For $d \in \{1,2\}$, suppose that $\mathrm{P}(s_d|D=d) = 1/3$ for all $s_d \in \mathcal{S}_d$. Consider the following weights \begin{align*} w_{{A}}(s_d,s_{d-1}) = \begin{cases} \frac{1}{6} & (s_d,s_{d-1}) \in \mathcal{M}^+(d) \\ 0 & (s_d,s_{d-1}) \in \mathcal{M}^-(d) \end{cases} \end{align*} i.e., $w_{\hspace{-0.4mm}{A}}(s_d,s_{d-1})$ puts $1/6$ weight equally on all six congruent comparisons and $0$ weight on the three incongruent comparisons. Alternatively, consider the following weights \begin{align*} w_{{B}}(s_d,s_{d-1}) = \begin{cases} 0 & (s_d,s_{d-1}) \in \mathcal{M}^+(d) \\ \frac{1}{3} & (s_d,s_{d-1}) \in \mathcal{M}^-(d) \end{cases} \end{align*} i.e., these are weights that involve putting $1/3$ weight equally on all three incongruent comparisons and $0$ weight on the six congruent comparisons. Both $w_{\hspace{-0.4mm}{A}}(s_d,s_{d-1})$ and $w_{\hspace{-0.1mm}{B}}(s_d,s_{d-1})$ meet the requirements for the weights that are discussed in Theorem (ref). Besides $w_{\hspace{-0.4mm}{A}}(s_d,s_{d-1})$ and $w_{\hspace{-0.1mm}{B}}(s_d,s_{d-1})$, many other weighting schemes also satisfy the same requirements.

The previous example demonstrates that the weights in Theorem (ref) are non-unique. There are different weighting schemes for $\textrm{MATT}$'s that can rationalize the aggregate marginal effect. Different weighting schemes can lead to very different interpretations of the aggregate marginal effect. In the example, one valid weighting scheme leads to an interpretation of $\mathbb{E}[Y|D=d]-\mathbb{E}[Y|D=d-1]$ as a weighted average that only includes congruent $\textrm{MATT}^+$'s. This weighting scheme is in line with our ideal scenario above---it provides an interpretation of the aggregate marginal effect $\mathbb{E}[Y|D=d]-\mathbb{E}[Y|D=d-1]$ that is fully congruent.

In contrast, the following example shows that there exist cases where the aggregate marginal comparison is incompatible with a fully congruent comparison of means of sub-treatments.

namedexample{\ref*{ex:enrichment} (continued)} Suppose that \begin{align*} &\mathrm{P}\Big((1,0,0)\Big|D=1\Big) = 0.8 &&\mathrm{P}\Big((1,1,0)\Big|D=2\Big) = 0.1 \\[5pt] &\mathrm{P}\Big((0,1,0)\Big|D=1\Big) = 0.1 &&\mathrm{P}\Big((1,0,1)\Big|D=2\Big) = 0.1 \\[5pt] &\mathrm{P}\Big((0,0,1)\Big|D=1\Big) = 0.1 &&\mathrm{P}\Big((0,1,1)\Big|D=2\Big) = 0.8 \end{align*} In this case, there do not exist fully congruent weights that can rationalize the $\mathbb{E}[Y|D=2]-\mathbb{E}[Y|D=1]$. The explanation is that the incongruent sub-treatment vectors $(1,0,0)$ and $(0,1,1)$ occur too commonly for $\mathbb{E}[Y|D=2] - \mathbb{E}[Y|D=1]$ to be rationalized with only fully congruent comparisons across sub-treatment vectors.

In the remainder of this section, we provide four arguments aiming to characterize empirical settings where the aggregate marginal effect, $\mathbb{E}[Y|D=d]-\mathbb{E}[Y|D=d-1]$, would be less likely to put weight on incongruent causal effects. These can be used by practitioners to diagnose, both conceptually and practically, how much of a problem an aggregated treatment may cause in a given application. First, in Section (ref), we show that the number of possible incongruent comparisons grows much more rapidly than the number of possible congruent comparisons as the complexity of the sub-treatments increases. Second, in Section (ref), we provide auxiliary assumptions that guarantee a weighting scheme that does not put any weight on incongruent comparisons. Third, in Section (ref), we discuss the characteristics of applications that must put weight on incongruent comparisons. Lastly, in Section (ref) we provide a test for whether $D$ is appropriately aggregated---i.e., whether a version of Assumption (ref) for the aggregated treatment is valid, in which case there would be no aggregation issues.

The Link between the Number of Sub-treatments and Incongruent Comparisons

In this section, we establish a link between the complexity of the sub-treatments (i.e., the number of distinct sub-treatments and the number of values that the sub-treatments can take) and the number of incongruent $\textrm{MATT}^-$ parameters that show up in the decomposition in Theorem (ref) relative to the number of congruent $\textrm{MATT}^+$ parameters. We show that the number of incongruent terms grows much more rapidly than the number of congruent terms. The implication for empirical work is that, all else equal, applications with more sub-treatments or complicated sub-treatments are more susceptible to the issues related to incongruent comparisons showing up in the aggregate marginal effects that we discussed above.

Suppose that all sub-treatments are binary (i.e., individuals can either participate or not in any of $K$ binary versions of the treatment). In Proposition (ref) of Appendix (ref), we establish that $|\mathcal{M}|$, the total number of contrasts for all permutations of sub-treatment vectors at adjacent amounts of aggregated treatment, is equal to $\binom{2K}{K-1}$.\footnote{Formally define the set of all marginal pairs of sub-treatment vectors as $\mathcal{M} := \cup_{d=1}^{\bar{N}} \mathcal{M}(d)$. Likewise, define the marginally congruent set of pairs $\mathcal{M}^+ := \cup_{d=1}^{\bar{N}} \mathcal{M}^+(d)$ and marginally incongruent set of pairs $\mathcal{M}^- := \cup_{d=1}^{\bar{N}} \mathcal{M}^-(d)$.} This number dramatically increases in $K$. For example, if $K=3$, then there are $\binom{6}{2} = 15$ possible contrasts. If $K=4$, there are $\binom{8}{3} = 56$ contrasts and so on. In addition, Proposition (ref) reveals that the amount of congruent contrasts $|\mathcal{M}^+| = K \cdot 2^{K-1}$ and incongruent contrasts $|\mathcal{M}^-| = \binom{2K}{K} - K \cdot 2^{K-1}$, allowing us to formally show that the total number of incongruent pairs of sub-treatment vectors grows much more rapidly with $K$ than the total number of congruent pairs of sub-treatment vectors (see Corollary (ref) in Appendix (ref)). The relatively rapid growth of the number of incongruent comparisons is illustrated in Figure (ref).

figure[figure omitted — 961 chars of source]

Figure (ref) shows the fraction of congruent and incongruent comparisons of sub-treatment vectors for a given number of sub-treatments. Panel (a) presents this in the setting with binary sub-treatments. More than half of the sub-treatment vectors are incongruent for any $K > 4$, and the fraction of incongruent sub-treatment vectors grows rapidly with $K$. Panel (b) considers the case where the sub-treatments can be multivalued---in this panel, all of the sub-treatments can take values among $\{0,1,2\}$. In this case, the number of incongruent contrasts dominates the number of congruent contrasts for any value of $K>2$, and the relative fraction of incongruent contrasts grows even faster with $K$ compared to the case when sub-treatments are binary. See Proposition (ref) in Appendix (ref) for exact expressions of the number of congruent and incongruent marginal contrasts in this case. Figure (ref) and Corollary (ref) suggest that the relative fraction of incongruent comparisons grows rapidly with the support of sub-treatments.

Having more sub-treatments does not necessarily mean that $\mathbb{E}[Y|D=d]-\mathbb{E}[Y|D=d-1]$ includes comparisons between incongruent sub-treatment vectors. What we have established in this section is that the scope for incongruency increases with the number of sub-treatments and with larger support size among the sub-treatments, suggesting that a researcher should pay especially close attention to issues related to incongruency in settings with a large number of sub-treatments, or few sub-treatments possessing multivalued supports.

Auxiliary Assumptions that Rule Out Incongruent Comparisons

This section lists auxiliary assumptions that rule out incongruency affecting the aggregate marginal effect, $\mathbb{E}[Y|D=d]-\mathbb{E}[Y|D=d-1]$. The first assumption rules out treatment effect heterogeneity across sub-treatments. The second set of assumptions rules out units sorting across different values of the aggregated treatment and introduces a restriction on latent unit-types.

Approach 1: Restrictions On Treatment Effect Heterogeneity

The first assumption we consider rules out treatment effect heterogeneity with respect to a marginal increase in any of the sub-treatments.

assumption[No Heterogeneous Sub-treatment Effects] For all $d \in \mathcal{D}_{>0}$ and any $(s_d,s_{d-1}) \in \mathcal{M}^+(d)$, \begin{align*} \mathrm{MATT}^+(s_d,s_{d-1}) = \beta_d \end{align*}

Assumption (ref) says that the average marginal causal effect of any sub-treatment is constant across sub-treatments between adjacent levels of aggregated treatment.\footnote{A similar assumption has been referred to in the Epidemiology literature as “treatment variation irrelevance”. See, for instance, VanderWeeleHernan2013.} In many applications, this may be a strong auxiliary assumption. For example, in our running example, it would say that the causal effect of a one-unit increase in any of the sub-treatments (whether it be homework, music, or sports) is the same for all sub-treatments. Most likely, this is a strong assumption in this context.

In Proposition (ref) in Appendix (ref), we show that, under Assumptions (ref)-(ref),

align*[align* omitted — 67 chars of source]

In other words, the aggregate marginal effect recovers the average marginal causal effect of the sub-treatments, which is $\beta_d$. The intuition for this result comes from the second part of Proposition (ref): under Assumption (ref), all the $\textrm{MATT}^+$'s are equal to each other, and Proposition (ref) therefore implies that all of the $\textrm{MATT}^-$'s are also equal to $\beta_d$. Replacing all of the $\textrm{MATT}^+$'s and $\textrm{MATT}^-$'s in Theorem (ref) then implies the result. Thus, in some sense, Assumption (ref) does not remove the weights on the $\textrm{MATT}^-(s_d,s_{d-1})$ terms in Theorem (ref), but it does make them irrelevant as all of them are equal to $\beta_d$. Equivalently, we can view Assumption (ref) as setting all $\text{SATT}'s$ to zero (which can be seen from Proposition (ref)): intuitively, when one sub-treatment is substituted by another, there is no effect on the outcome.

Approach 2: Structural Assumptions

An alternative approach to ruling out incongruity in the decomposition in Theorem (ref) comes from introducing structural assumptions. Let $S_i(d)$ denote the sub-treatment that unit $i$ would experience under aggregated treatment $d$.\footnote{The object $S_i(d)$ should be understood as a latent sub-treatment type across $D$, which describes how the components of treatment are arranged at each total treatment level in observational settings. Recall that aggregated treatment variables are deterministic functions of the sub-treatments, devised ex post, and are not causal on the outcome. This indicates that types are inseparable from the aggregation function that underlies them. Hence, the concept of type is an artifact of the aggregation scheme, not itself an intervention with its own causal effect. } Thus, $S_i(\mathbf{d}) := (S_i(1), S_i(2), \ldots, S_i(\bar{N}))$ defines a unit-level latent aggregated treatment path---the particular sub-treatment vector that a unit would experience for all possible values of the aggregated treatment. The set of possible values of $S(\mathbf{d})$ is finite, and we can define a notion of a unit's latent type on the basis of $S(\mathbf{d})$.

assumption[No Sorting on $D$] Latent types are independent of the aggregated treatment; that is, $$S(\mathbf{d}) \mathrel{\mbox{\(\perp\!\!\!\perp\)}} D$$
assumption[No Incongruent Latent Types] For all $d \in \mathcal{D}_{>0}$, treatment paths are locally congruent; that is, \begin{align*} \mathrm{P}\big(S(d)=s_d, \; S(d-1)=s_{d-1} \big| D \in \{d, d-1\}\big) &= 0, if (s_{d}, s_{d-1}) \in \mathcal{M}^{-}(d) \\ \mathrm{P}\big(S(d)=s_d, \; S(d-1)=s_{d-1} \big| D \in \{d, d-1\}\big) &\geq 0, if (s_{d}, s_{d-1}) \in \mathcal{M}^{+}(d) \end{align*}

Assumption (ref) says that latent types are balanced across aggregate amounts of treatment. This ensures that there is no selection at the disaggregate level based on the total amount of treatment. No sorting holds under random assignment of the sub-treatments.

Assumption (ref) imposes an explicit restriction on the latent types in the population---that there are no units in a latent type that “behaves” incongruently. In our running example, it rules out types of units that would spend one hour of enrichment doing homework, but had they done two hours of enrichment, they would have done music lessons and sports.

We show in Proposition (ref) in Appendix (ref) that these two conditions are sufficient to guarantee that there exists a weighting scheme satisfying the conditions in Theorem (ref) that puts no weight on any $\textrm{MATT}^-(s_d,s_{d-1})$.

namedexample{\ref*{ex:enrichment} (continued)} It is worth pointing out why Assumption (ref) alone is not sufficient to guarantee that the aggregate marginal effect can be decomposed entirely in terms of congruent $\textrm{MATT}^+$'s. Consider an extreme version of our running example, where there are two latent types: type 1 would do homework if they did one hour of enrichment and would do homework and music lessons if they did two hours of enrichment; type 2 would play sports if they did one hour of enrichment and would do music lessons and play sports if they did two hours of enrichment. Both latent types are congruent. However, suppose there is sorting so that all type 1 units do one hour of enrichment, while all type 2 units do two hours of enrichment. In this case, the aggregate marginal effect $\mathbb{E}[Y|D=2]-\mathbb{E}[Y|D=1]$ is fully incongruent due to sorting, despite all units themselves belonging to a congruent latent type.

Settings where Incongruent Comparisons Are Unavoidable

In the previous section, we discussed additional assumptions that side-step incongruent $\textrm{MATT}^-$'s complicating the interpretation of the aggregate marginal effect. This section pivots to characterizing the features of applications that necessarily include incongruent $\textrm{MATT}^-$'s.

Sub-treatment Decreases in Aggregated Treatment Guarantees Incongruency

The following result provides a straightforward characteristic of an application that indicates that incongruency is unavoidable in the aggregate marginal effect.

propositionProvided there exists some sub-treatment indexed by $k \in \{1, \ldots, K\}$, and some $d \in \mathcal{D}_{>0}$, such that \begin{align*} \mathbb{E}[S_k | D=d] \; &< \; \mathbb{E}[S_k | D=d-1] \end{align*} then any weights that satisfy the properties in Theorem (ref) must assign positive weight to at least one incongruent pair of sub-treatments, $(s_d,s_{d-1}) \in \mathcal{M}^{-}(d)$.

Proposition (ref) states that if the conditional means of any sub-treatment declines between values of aggregated treatment $D=d-1$ and $D=d$, then weighting schemes that avoid incongruency are impossible. The condition in the proposition is easy to consider in applications as it concerns the mean of a particular sub-treatment across different values of the aggregated treatment. In the context of our application, the proposition says that incongruent comparisons cannot be avoided if, for example, the mean number of hours spent on homework was 0.75 among children that did one hour of enrichment while the mean number of hours spent on homework was 0.5 among children that did two hours of enrichment. This is an intuitive condition for guaranteeing incongruency: if the mean of some sub-treatment decreases in $D$, then there is simply not enough available mass on congruent sub-treatment vectors at the higher value of the aggregated treatment to satisfy the requirements on the weights in Theorem (ref).

Minimally Incongruent Weights

The condition in Proposition (ref) is a sufficient, but not necessary, condition for incongruency. Moreover, if it holds, it implies that incongruency is a problem, but it does not necessarily provide much information about how much of a problem it is. With this in mind, in this section we define minimally incongruent weights as a solution to the following linear programming problem:

align[align omitted — 157 chars of source]

subject to {

align*[align* omitted — 285 chars of source]

}This defines the weights $w^\star$ to be a weighting scheme that minimizes the weight on incongruent comparisons between marginal sub-treatment vectors subject to satisfying the criteria for the weights discussed in Theorem (ref). There are several additional clarifications worth mentioning. First, if $w^\star(s_d,s_{d-1}) > 0$ for any $(s_d,s_{d-1}) \in \mathcal{M}^-(d)$, it necessarily implies that the marginal comparison of aggregate means includes incongruent comparisons across sub-treatment vectors. Second, if $w^\star(s_d,s_{d-1}) = 0$ for all $(s_d,s_{d-1}) \in \mathcal{M}^-(d)$, then the marginal comparison of aggregate means has a representation that only includes congruent comparisons across sub-treatment vectors. However, in general, there can be many weighting schemes that meet these criteria and involve congruent comparisons across sub-treatments; for instance, in the earlier example with uniform probabilities of each sub-treatment vector on page \pageref*{example uniform}, there are many weighting schemes that only involve congruent comparisons across sub-treatment vectors.

Testing Whether \texorpdfstring{$D$}{D} is Too Aggregated

Next, we show that one can test whether the version of Assumption (ref) relative to $D$ holds. If it does, then the aggregation issues discussed in this paper are not relevant for the empirical application, and the researcher may use $D$ as their treatment variable without having to use the methods developed in this paper. The second part of SUTVA is often considered to be untestable (see, for example, the discussion in Hernan2016); in this section, we highlight that it is jointly testable with Assumptions (ref) and (ref) in settings where the researcher observes sub-treatment $S$. We note that an analogous argument to the one below could be used to test Assumption (ref) per se (i.e., the version of that assumption relative to $S$) provided the researcher also observes a more disaggregated version of sub-treatment $\tilde{S}$. See Appendix (ref) for further details.

The version of Assumption (ref) relative to $D$ holds if, for all $d \in \mathcal{D}$, $Y_i(s_d) = Y_i(s_d')$ for all $i$ and $s_d, s_d' \in \mathcal{S}_d$. Our test will be based on the comparison of means across different sub-treatment vectors corresponding to the same aggregate value of the treatment. If SUTVA holds for the aggregated treatment $D$, we have that, for any $s_d,s_d' \in \mathcal{S}_d$,

align*[align* omitted — 285 chars of source]

where the first equality holds by Assumption (ref) relative to $S$, and the second equality holds by adding and subtracting $\mathbb{E}[Y(s_d')|S=s_d]$. The first underlined term in the second line is equal to 0 when there are no hidden versions of the aggregated treatment (i.e., under Assumption (ref) relative to $D$), but the second term could still be non-zero---the mean of the potential outcomes of sub-treatment vector $s_d'$ could be different for sub-treatment group $s_d$ relative to sub-treatment group $s_d'$, even if these are not distinct versions of the treatment. For example, “homework” and “music” could be equivalent versions of the treatment, and yet the latter term could be non-zero if, for some reason related to selection, children who do homework tend to have higher or lower outcomes than children who do music.

However, the underlined selection bias term is equal to zero under Assumption (ref). This implies that the version of Assumption (ref) relative to $D$ is testable under the maintained Assumptions (ref) and (ref). One can carry out the test proposed here by simple tests for differences in means between all pairs of sub-treatments corresponding to the same aggregate level of the treatment, adjusting for multiple testing error.\footnote{See HasegawaEtAl2020 for simultaneous inference with two versions of treatment in the binary treatment case, which does not require correcting for multiple testing.} We also note that related ideas could be used under alternative identification strategies that rely on different assumptions than Assumption (ref).

Discussion

This section has aimed to highlight the features of applications where the incongruent comparisons that show up in the decomposition in Theorem (ref) arise. First, we showed that the relative number of incongruent $\textrm{MATT}^-$'s grows rapidly in the complexity of the sub-treatments. Second, we discussed additional assumptions (limitations on treatment effect heterogeneity and restrictions on sorting and latent types) that rule out incongruent $\textrm{MATT}^-$'s in the aggregate marginal effect. Third, we provided conditions (a sub-treatment that decreases in the aggregated treatment) that guaranteed that incongruent $\textrm{MATT}^-$'s would show up in the aggregate marginal effect. Fourth, we showed how to test whether there should be any aggregation issues by testing the version of Assumption (ref) relative to $D$.

To conclude this section, it is worth emphasizing that these four arguments provide complementary ways for a researcher to informally assess “how much” aggregation matters in a particular application. For example, in an application with a small number of uncomplicated sub-treatments, where the sub-treatments are similar to each other and likely to have close to homogeneous effects, and where the version of Assumption (ref) relative to $D$ is not rejected, one should expect the negative implications of working with an aggregated treatment to be small. In contrast, an application with a large number of more-distinct sub-treatments, heavy sorting across different values of the aggregated treatment, and where the version of Assumption (ref) relative to $D$ is rejected, is one in which we should expect major distortions to arise due to the aggregation of the treatment.

Alternative Approaches

In the preceding section, we saw that interpreting aggregate marginal effects encountered several complications, arising from two key issues: (i) comparisons across values of the aggregated treatment could mix marginal effects of congruent sub-treatment vectors and marginal effects of incongruent sub-treatment vectors, and (ii) marginal effects of incongruent sub-treatment vectors are difficult to interpret. We then outlined some ways that a researcher could diagnose (or at least think about) the implications of incongruency with an aggregated treatment in a given application. In this section, we consider two alternative approaches that can completely side-step the issues related to incongruency that were emphasized above. First, in Section (ref), we consider alternative, non-marginal causal effect parameters. Targeting these parameters fully circumvents issues related to incongruency. They are straightforward to interpret and are estimable when only the aggregated treatment is observed (i.e., they do not require the sub-treatments themselves to be observed). Because this approach does not require the sub-treatments to be observed, this approach offers a path forward to conduct causal inference even in applications that require a very disaggregated notion of the sub-treatments to satisfy Assumption (ref). Moreover, the non-marginal comparisons that we consider in that section immediately apply for any aggregation function, not just the sum of the sub-treatments. Changing the target parameter means that it is no longer interpretable as a marginal causal effect parameter, which is a drawback for applications where a researcher strongly prefers this type of parameter. Second, in Section (ref), we show how to identify fully congruent marginal causal effect parameters in applications where sub-treatment data is available. This approach delivers a marginal causal effect parameter, but it requires Assumption (ref) to hold for the observed sub-treatments and is more sensitive to the specific aggregation function specified by the researcher.

Approach 1: Target Non-marginal Causal Effect Parameters

Marginal effects have a strong claim on being the most natural target parameters in the setting that we are considering, where the aggregated treatment can take multiple values, and reflect the most common ways that empirical work interprets results in these settings. However, the previous section documented several challenges with interpreting aggregate marginal effects in the presence of sub-treatments. Instead of considering marginal changes in the aggregated treatment, in this section, we focus on interpreting $\mathbb{E}[Y|D=d] - \mathbb{E}[Y|D=0]$ for any $d \in \mathcal{D}_{>0}$, which is the difference between the means of outcomes for the group that experiences aggregated treatment $d$ relative to the untreated group. We show that this non-marginal, aggregate comparison does not include incongruent comparisons across sub-treatments. We further show that, under Assumption (ref), this comparison has a causal interpretation as the average of the causal effects of each sub-treatment $s_d \in \mathcal{S}_d$ relative to being untreated. Importantly, sub-treatments in this case do not need to be observed by the researcher. We refer to this as a baseline-to-$d$ comparison in the text below.

In terms of causal effect parameters, the main building block parameter in this section is the average treatment effect on the treated (ATT)

align*[align* omitted — 72 chars of source]

which is defined with respect to a given sub-treatment $s_d$ and where $Y(0)$ is shorthand notation for being untreated (i.e., where all sub-treatments are equal to zero). $\textrm{ATT}(s_d)$ is the average effect of experiencing sub-treatment vector $s_d$ relative to being untreated among sub-treatment group $s_d$. In the spirit of using the aggregated treatment to summarize the causal effects of the sub-treatments, our main target parameter in this section is

align*[align* omitted — 81 chars of source]

which is the aggregate average treatment effect on the treated across sub-treatments corresponding to the aggregated treatment being equal to $d$. From the law of iterated expectations, it follows that

align[align omitted — 135 chars of source]

i.e., that $\textrm{AATT}(d)$ is a weighted average of the underlying $\textrm{ATT}$'s of specific sub-treatments, with weights given by the relative frequency of that sub-treatment among all sub-treatments that aggregate to $d$.

In some applications, it is also useful to scale $\textrm{ATT}(s_d)$, or $\textrm{AATT}(d)$, by the amount of the aggregated treatment, i.e., to consider the parameters

align*[align* omitted — 88 chars of source]

which can be interpreted as average treatment effects per unit of the (sub)-treatment. We refer to these as scaled $\textrm{ATT}$'s and scaled $\textrm{AATT}$'s, respectively.

Identification

Next, we provide identification results for the average treatment effect parameters discussed above.

theoremUnder Assumptions (ref) and (ref), for $d \in \mathcal{D}_{>0}$ and $s_d \in \mathcal{S}_d$, \begin{align*} \mathrm{ATT}(s_d) = \mathbb{E}[Y|S=s_d] - \mathbb{E}[Y|S=0_K] \quad and \quad \mathrm{AATT}(d) = \mathbb{E}[Y|D=d] - \mathbb{E}[Y|D=0] \end{align*} where $0_K$ denotes the zero vector of length $K$. $\mathrm{ATT}(s_d)$ is identified if the sub-treatments are observed. $\mathrm{AATT}(d)$ is identified whether or not the sub-treatments are observed.

Theorem (ref) shows that $\textrm{AATT}(d)$ is identified under Assumptions (ref) and (ref), even if the researcher only observes the aggregated treatment (and not the sub-treatments). It is interesting to compare this result with the one in Theorem (ref) above concerning the comparison of means of outcomes for marginal increases in the aggregated treatment (i.e., $\mathbb{E}[Y|D=d]-\mathbb{E}[Y|D=d-1]$). A major issue for the marginal comparison emphasized in Section (ref) was the non-uniqueness of the weights and the possibility of incongruency. Neither of those issues apply for the baseline-to-$d$ comparisons, $\mathbb{E}[Y|D=d]-\mathbb{E}[Y|D=0]$, considered here. The “weights” on the underlying sub-treatment-specific $\textrm{ATT}(s_d)$ parameters are given in Equation (ref). These are unique, positive for all relevant sub-treatments, and intuitive---they correspond to the relative frequency of each relevant sub-treatment. Mechanically, the same sort of double-sum arguments can be used here as in the previous case, but, by construction, $\mathcal{S}_0$ only has one element, which results in the implicit weighting scheme being unique in this case. The benefit is that, unlike for the marginal case, $\mathbb{E}[Y|D=d]-\mathbb{E}[Y|D=0]$ is straightforward to interpret, and all of the issues related to incongruency emphasized above can be avoided.\footnote{Both Theorem (ref) and Theorem (ref) invoked Assumption (ref). In both cases, this assumption is stronger than necessary, though the minimal assumptions to provide a causal interpretation in each result are non-nested. Causal interpretations of marginal effects of sub-treatments can hold under a local version of no selection, while causal interpretations of $\textrm{AATT}(d)$ can be rationalized under a version of the no-selection assumption that involves untreated potential outcomes only. This could be a meaningful difference in some applications (though it is not relevant for any of our discussions about aggregation specifically). We discuss these differences in more detail in Appendix (ref).}

Interpreting Regressions with Scaled \texorpdfstring{Baseline-to-$d$}{Baseline-to-d} Building Blocks

Next, we return to interpreting the coefficient on the aggregated treatment variable in the regression from Equation (ref), but we relate it to the scaled baseline-to-$d$ building blocks: $(\mathbb{E}[Y|D=d]-\mathbb{E}[Y|D=0])/d$. In Proposition (ref) in the Supplementary Appendix (CCCDSupp2025), we show that

align*[align* omitted — 129 chars of source]

where

align*[align* omitted — 115 chars of source]

and satisfies the following properties: (i) $\displaystyle \sum_{d=1}^{\bar{N}} \tilde{\omega}^{reg}(d) = 1$ and (ii) $\tilde{\omega}^{reg}(d) \lessgtr 0$ for $d \lessgtr \mathbb{E}[D]$.\footnote{The proof uses the same mechanics as Theorem S3 in the Supplementary Appendix of chaisemartin-dhaultfoeuille-2020, though the context of that result (interpreting two-way fixed effects regressions) is very different from ours.}

The expression for $\alpha_1$ above can be combined with the result in Theorem (ref) to say that the regression coefficient on the aggregated treatment can be interpreted as a weighted average of scaled $\textrm{AATT}(d)$ parameters. Relative to the regression weights discussed above for the marginal case, the regression weights with scaled baseline-to-$d$ primitives have worse properties. The weights are negative for values of the aggregated treatment below $\mathbb{E}[D]$, implying that $\alpha_1$ is not weakly causal when scaled $\textrm{AATT}(d)$'s are the underlying building block parameters. In addition, the weights are systematically increasing in magnitude in their distance from $\mathbb{E}[D]$, meaning that effects for sub-treatment groups with more extreme values of the aggregated treatment $D$ “count more” than effects for other sub-treatment groups.

The discussion above highlights a certain tension with interpreting $\alpha_1$ in terms of baseline-to-$d$ causal effects. While the building blocks $(\mathbb{E}[Y|D=d] - \mathbb{E}[Y|D=0]) / d$ are more interpretable than the marginal effects discussed previously, the weights inherited from the regression have poor properties. Thus, in settings where a researcher would like to report a single, scalar summary of the causal effects of the sub-treatments, a natural alternative to reporting $\alpha_1$ is to report either

align*[align* omitted — 144 chars of source]

which are all identified under the same conditions that would give $\alpha_1$ a causal interpretation, but do not suffer from the poor weighting scheme stemming from the regression.

remark[Linearity and Homogeneity] Consider the case where, for all $d \in \mathcal{D}_{>0}$ and $s_d \in \mathcal{S}_d$, $\mathrm{ATT}(s_d) = \theta \times d$, so that the average treatment effects of all sub-treatments are (i) constant across sub-treatments corresponding to the same aggregate amount of the treatment and (ii) linear in $d$. (i) and (ii) restrict treatment effect heterogeneity across sub-treatments and impose a linearity condition. In this case, $\mathbb{E}\left[ \frac{\mathrm{AATT}(D)}{D} \middle| D > 0\right] = \theta$, and, in addition, it also follows that $\alpha_1 = \theta$. In other words, in this case, the regression would deliver the unique scaled treatment effect parameter. In practice, both (i) and (ii) are likely to be strong auxiliary assumptions for most applications. This suggests it is a better strategy to directly target parameters such as $\mathbb{E}\left[ \frac{\mathrm{AATT}(D)}{D} \middle| D > 0\right]$ rather than hoping that the regression will deliver them.

Approach 2: Target Marginal Effect Parameters with Sub-treatment Data

In this section, we target summary marginal causal effect parameters, exploiting that the sub-treatments are observed, which only applies for some applications. To start with, recall that, under Assumptions (ref) and (ref),

align*[align* omitted — 136 chars of source]

Therefore, when the sub-treatments are observed, $\textrm{MATT}^+(s_d,s_{d-1})$ is identified (see Proposition (ref) in Appendix (ref)). Given that this marginal effect is defined at the sub-treatment level, none of the issues related to incongruency that we emphasized above in the context of aggregation apply. In practice, a researcher could estimate and report $\textrm{MATT}^+(s_d,s_{d-1})$ for any (or all) combinations of congruent sub-treatment vectors. Leaving the discussion here, however, would not fully address some relevant empirical challenges---presumably, in the majority of applications where the sub-treatments are observed, the entire reason to introduce an aggregated treatment variable is that there tend to be few observations that experienced each specific combination of sub-treatments. This implies that the sort of non-parametric analysis mentioned above would suffer from a form of curse of dimensionality, resulting in each $\textrm{MATT}^+(s_d,s_{d-1})$ being estimated imprecisely and in poor performance of inference procedures for the $\textrm{MATT}^+$'s.

In contrast, however, even when the number of observations per combination of sub-treatments is small, one may still be able to estimate averages of $\textrm{MATT}^+$'s well. A natural option is

align[align omitted — 181 chars of source]

where, for $(s_d,s_{d-1}) \in \mathcal{M}^+(d)$,

align[align omitted — 283 chars of source]

$\widetilde{\textrm{AMATT}^+}(d)$ is a special case of $\textrm{AMATT}^{+}_{w^+}(d)$ in Equation (ref) above, as it is a specific weighted average of congruent $\textrm{MATT}^+$'s. The weights come from the joint distribution of latent sub-treatment types local to the aggregated treatment either being $d$ or $d-1$. These weights give $\widetilde{\textrm{AMATT}^+}(d)$ a clear interpretation as a marginal “on-the-treated” type of parameter as it depends on the distribution of the sub-treatments that are (or would be) experienced. The term in the denominator of the expression for $\tilde{w}^+(s_d,s_{d-1})$ can be thought of as normalizing the weights on congruent sub-treatments so that they sum to one.\footnote{In settings where units would make congruent sub-treatment choices at different amounts of the aggregated treatment (i.e., if Assumption (ref) holds), then the expression in the denominator is equal to one, and the weights do not need to be normalized. Alternatively, one can view the normalization as arising due to dropping incongruent $\textrm{MATT}$'s.} Recovering $\widetilde{\textrm{AMATT}^+}(d)$, however, introduces additional challenges relative to identifying $\textrm{MATT}$'s: even if the sub-treatments are observed, the weights depend on the joint distribution of latent sub-treatment types for aggregated treatment $D=d$ or $D=d-1$ and, therefore, require additional assumptions to identify. We discuss these issues in more detail in Appendix (ref); however, in order to avoid introducing additional assumptions, we instead focus on identifying a version of $\textrm{AMATT}^{+}_{w^+}(d)$ with researcher-chosen weights, sacrificing some interpretability but increasing tractability. A leading option for researcher-chosen weights is to use the normalized product weights, i.e.,

align*[align* omitted — 270 chars of source]

for $(s_d,s_{d-1}) \in \mathcal{M}^+(d)$, and then to consider

align[align omitted — 179 chars of source]

Using this weighting scheme results in putting more weight on common sub-treatments. An immediate implication of Assumptions (ref) and (ref) and observing the sub-treatments (see Proposition (ref)) is that $\textrm{AMATT}^{+}_{w^+}(d)$ in Equation (ref) is identified; this also implies that $\textrm{AMATT}^+(d)$ is identified under the same conditions.

In settings where a researcher would like a scalar summary of the marginal causal effects of the treatment, a natural target parameter is

align*[align* omitted — 90 chars of source]

which is an average marginal (weakly) causal effect parameter that comes from averaging $\textrm{AMATT}^+(d)$ over the distribution of the aggregated treatment variable $D$. It is identified under Assumptions (ref) and (ref) when the sub-treatments are observed.

Empirical Application

In this section, guided by our results above, we illustrate the aggregation issues and the methods proposed above with data from CCaetanoEtAl2024, which studies the effects of enrichment activities on noncognitive skills among children in the U.S. Like that paper, we use data from the Childhood Development Supplement (CDS) provided by the Panel Study of Income Dynamics (PSID) that contains time-use diaries and measures of cognitive and noncognitive skills. For clarity, we make some simplifications. Our estimates come from simple comparisons of means; we ignore control variables and fixed effects in our analysis, and we abstract from the bunching identification strategy that is often used in this literature. Instead of trying to make a causal empirical claim, in this section, we only aim to highlight the issues that stem from aggregating sub-treatment variables; and, as discussed above, these issues continue to apply whether or not Assumption (ref) holds.

The aggregated treatment variable in our application is total hours of enrichment activity, which is an aggregation of four sub-treatments: Lessons, Structured Sports, Volunteering, and Before & After School Programs. Each sub-treatment represents the average number of hours per week that the child spent on that specific activity, and the sub-treatment vector represents the bundle of sub-treatments each child was exposed to. In our data, we observe each child's participation in each sub-treatment. We round the sub-treatments to their nearest half-hour. We then sum across all four sub-treatments for each individual so that the aggregate variable $D=(S_1+S_2+S_3+S_4)$ is the total amount of enrichment activities. The outcome of interest is noncognitive skill, a normalized index of socio-emotional and behavioral ratings with mean zero and standard deviation of one. Larger values indicate better noncognitive scores. We follow CCaetanoEtAl2024 in constructing this variable; see that paper for more details. See Supplementary Appendix (ref) for summary statistics and more details on how we constructed the data used in our application (CCCDSupp2025). In the main text, to simplify the discussion, we focus on a subsample of low socio-economic status children in 2019.\footnote{We opted for illustrating the analysis in this smaller subsample because it has sufficiently few different sub-treatment values, allowing us to report Table (ref) in the paper.} In Supplementary Appendix (ref) (CCCDSupp2025), we provide the results for the full sample used in CCaetanoEtAl2024.

Observing the sub-treatments is important for our analysis. Below, we often temporarily ignore that we observe the sub-treatments and act as if we only had access to the aggregated treatment. Then, exploiting the fact that we actually do observe the sub-treatments, we are able to diagnose how much aggregation itself affects the results. It is also worth mentioning that, although the decompositions that we discussed in previous sections were written in terms of population quantities, the same arguments can be applied to their sample analogues. We also report standard errors below, but we mainly emphasize the point estimates---because the sample, sub-treatments, and outcomes are the same across estimators, differences in results reflect real differences in the estimands rather than noise. As such, statistical significance is not the primary lens through which to assess these differences; instead, differences in the point estimates themselves reflect how much varying the estimands matters in practice.

Sub-treatment Diagnostics

We begin by examining evidence of incongruency across different values of the aggregated treatment. Specifically, we investigate how the composition of sub-treatments varies with the level of total enrichment, recalling from Proposition (ref) that a sub-treatment whose mean declines in the aggregated treatment implies the presence of incongruency.

figure[figure omitted — 406 chars of source]

Figure (ref) displays the mean of each sub-treatment across every level of the aggregated treatment $D$. The plot shows the mean number of hours for each type of enrichment activity (each sub-treatment) as the total hours of enrichment activity increases. The 45-degree line represents, for each value $D$ in the horizontal axis, the vertical sum of hours across all sub-treatments, which naturally equals the total number of hours spent on enrichment, $D$. For example, at $D=0.5$, the average amount of hours spent on lessons is 0.5, which is 100% of the total enrichment for that value of $D$. At $D=1$, we observe that the average number of hours for lessons increases relative to $D=0.5$, but this sub-treatment is no longer the only sub-treatment that is experienced at $D=1$.

More interestingly, the mean for the lessons sub-treatment declines from $D=1$ to $D=1.5$. From Proposition (ref), this is evidence of incongruency, implying that the aggregate marginal effect at $D=1.5$ cannot avoid putting weight on incongruent marginal sub-treatment effects. Notice that these violations continue for marginal increases from $D=1.5$ to $D=2.0$ for lessons; from $D=2.0$ to $D=2.5$ for both sports and volunteering sub-treatments; and from $D=2.5$ to $D=3.0$ for all sub-treatments except lessons. This suggests that aggregate marginal effects (including regressions interpreted as marginal effects) are hard to interpret---as discussed above, the incongruent $\textrm{MATT}^-$ terms that will show up here include hard-to-interpret substitution effects or, equivalently, congruent $\textrm{MATT}^+$'s with negative weights.

Although the sub-treatment plot can tell us where weight on incongruent comparisons is inescapable, it does not explain how much incongruity there is. To answer this question, we next find the minimally incongruent weights by solving the linear program in Section (ref). These weights provide an interpretation of the aggregate marginal effect with minimal incongruity. The results from solving the problem are displayed in Table (ref), which lists all marginal pairs of sub-treatment vectors at each $D=d$ from the data. In line with the results from Figure (ref), the minimally incongruent weights put weight on incongruent marginal sub-treatment effects for the aggregate marginal effects at $D=1.5, 2.0, 2.5,$ and $3.0$. Moreover, for $D=1.5,2.0,$ and $2.5$, the weight on incongruent marginal sub-treatment effects is substantial, ranging from 30-40% of the total weight. Even more strikingly, all of the weight falls on incongruent comparisons between $D=2.5$ and $D=3.0$ because there are no available congruent sub-treatment vectors.

table[table omitted — 43,190 chars of source]

Regression and Target Parameter Estimation

Next, we move to estimating the effects of enrichment activities, comparing the different approaches that we have discussed throughout the paper. Table (ref) reports estimates of several scalar parameters summarizing the effects of enrichment activities on a child's noncognitive skills. All estimates in Table (ref) come from plug-in estimators of the target parameters considered in the paper. Each estimate is in terms of the number of standard deviations away from the mean noncognitive skill score (zero).

\FloatBarrier

Panel I contains an estimate of the coefficient on total enrichment activities from the regression in Equation (ref), which represents the best linear approximation between noncognitive skill and total enrichment activity. This regression represents the most common way to summarize the relationship between enrichment activities and noncognitive skills. The estimated coefficient is $\hat{\alpha}_1=-0.061$ and is not statistically different from zero at standard levels of significance.\footnote{This estimate and others reported in this section are qualitatively similar to the ones in CCaetanoEtAl2024, which finds negative effects of enrichment on noncognitive skills. However, we note that their bunching/selection-on-unobservables identification strategy is very different from the approach in this section.} Given our discussion on the presence of incongruency in our application above, it follows that, if one aims to interpret this parameter in marginal terms, then it consists of incongruent comparisons as described in Section (ref). The regression coefficient also inherits the implicit regression weights discussed in Remark (ref) above.

table[table omitted — 1,871 chars of source]

Next, Panel II reports estimates of two alternative overall marginal effects parameters. The first is $\mathbb{E}[\delta(D)|D>0]$, where $\delta(d) :=\mathbb{E}[Y|D=d] - \mathbb{E}[Y|D=d-1]$ is the average aggregate marginal effect. Like the regression coefficient from Panel I, this parameter includes incongruent marginal sub-treatment effects, but it does not inherit the implicit regression weights (i.e., it is a non-parametric summary of the aggregate marginal effects). Our estimate of $\mathbb{E}[\delta(D)|D>0]$ is -0.040. The estimated value of this parameter is closer to zero than the regression coefficient, which indicates that the regression weights put relatively more weight on values of the aggregated treatment with larger marginal effects in magnitude, in comparison to weighting by the distribution of $D$. The second estimate in Panel II is for $\mathbb{E}[\textrm{AMATT}^{+}(D)|D>0]$; our estimate of this parameter is -0.087. As discussed above, this parameter is a weighted average of all congruent marginal sub-treatment effects. Among the parameters reported in this section, this parameter has the strongest claim to being the best way to summarize the effects of marginal increases in the sub-treatments, though estimating it does require observing a sub-treatment vector satisfying Assumption (ref).

figure[figure omitted — 1,673 chars of source]

Comparing the previous estimates to our estimate of $\mathbb{E}[\textrm{AMATT}^{+}(D)|D>0]$ provides a good way to assess how much incongruency matters in our application. Relative to our estimate of $\mathbb{E}[\delta(D)|D>0]$, it is more than twice as large in magnitude. The difference in these estimates is fully explained by whether or not they include incongruent comparisons (since they both weight across the aggregated treatment using the observed distribution of $D$), indicating that incongruency has a large impact on the aggregated estimates. To further decompose this difference, Panel (a) of Figure (ref) reports estimates of $\delta(d)$ and $\textrm{AMATT}^{+}(d)$ across different values of $d$. Recall that, according to Figure (ref), $\delta(d)$ must include those incongruent comparisons at the disaggregated level for values of $d \in \{ 1.5, 2.0, 2.5, 3.0 \}$, which could potentially lead to notable differences in the estimates of these parameters at those values of the aggregated treatment. This difference is most striking at $D=2.5$, where our estimate of $\delta(d=2.5)$ is 0.103, while our estimate of $\textrm{AMATT}^{+}(d=2.5)$ is -0.237. In addition, because there are no available congruent marginal sub-treatment effects at $D=3$, $\textrm{AMATT}^{+}(d=3)$ cannot be estimated and does not contribute to the estimate of $\mathbb{E}[\textrm{AMATT}^{+}(D)|D>0]$. There are other smaller differences at other values of $d$. Together, these differences explain the disparity between the two overall estimates in the middle panel of Table (ref). Interestingly, relative to $\mathbb{E}[\delta(D)|D>0]$, the estimate of $\alpha_1$ is closer to our estimate of $\mathbb{E}[\textrm{AMATT}^{+}(D)|D>0]$. There are two sources of differences between these two estimates: $\alpha_1$ includes incongruent marginal sub-treatment effects and inherits a weighting scheme from the regression. Here, these two issues work in opposite directions---incongruency pushes the estimates towards zero, while the regression weights push it farther from zero---resulting in the estimate of $\mathbb{E}[\textrm{AMATT}^{+}(D)\mid D > 0]$ being closer to that of $\alpha_1$ than to that of $\mathbb{E}[\delta(D)\mid D > 0]$.

To summarize, there are two main explanations for the variation in the estimated values of the various marginal effect parameters considered here. First, for parameters that are averages of the aggregate marginal effects, it is not possible to avoid incongruent comparisons. This is a byproduct of large changes in the composition of sub-treatments across different values of the aggregated treatment. Second, there appears to be notable heterogeneity in sub-treatment-specific marginal effects. Estimates of congruent sub-treatment marginal range from -0.762 to 0.752 (these are reported in Table (ref)), implying that variations in weights across different summary marginal effect parameters are important.\footnote{To be clear, our estimates of sub-treatment-specific marginal effects are likely to be quite imprecise. Nevertheless, mechanically, these estimates contribute non-trivially to the estimates of the summary marginal effect parameters that we have emphasized here.}

Finally, in Panel III of Table (ref), we report estimates of the non-marginal parameters that we discussed in Section (ref). Unlike $\mathbb{E}[\textrm{AMATT}^{+}(D)|D>0]$, neither of these parameters requires the sub-treatments to be observed; moreover, they do not suffer from non-unique weights, nor do they include incongruent comparisons. The first of these, $\mathbb{E}[\textrm{AATT}(D)|D>0]$, is the overall aggregate average treatment effect on the treated for those who participated in the treatment. Our estimate is -0.167. This is notably larger in magnitude than any of the preceding estimates, though it is hard to directly compare them, as the underlying building blocks are different. Here, the underlying building blocks are $\textrm{AATT}(d)$---among sub-treatments that aggregate to $d$, the average of the effect of each sub-treatment relative to being untreated. Estimates of $\textrm{AATT}(d)$ are displayed in Panel (b) of Figure (ref) for different values of $d$. Averaged across sub-treatments, we find our estimates of the effect of participating in different types of enrichment activities are largest in magnitude among those who participate in few enrichment activities (i.e., for $D=0.5$ and $1.0$). The last parameter considered in the table is $\mathbb{E}[\frac{\textrm{AATT}(D)}{D} \big| D>0]$; our estimate of this parameter is -0.218. This parameter summarizes the average treatment effect on the treated per unit of treatment received. Recall that, as in Section (ref), $\alpha_1$ can be interpreted as a weighted average of the same scaled baseline-to-$d$ building blocks. But our estimate of $\alpha_1$ is very different from our estimate of $\mathbb{E}[\frac{\textrm{AATT}(D)}{D} \big| D>0]$---our estimate of $\mathbb{E}[\frac{\textrm{AATT}(D)}{D} \big| D>0]$ is over three times as large in magnitude. This difference is fully explained by differences in the weighting schemes for each case. In our view, if a researcher did not insist on reporting a marginal effect parameter, the non-marginal effect parameters considered in Panel III would provide a good way to summarize the effects of enrichment activities on noncognitive skills. This would be especially true for an application where the types of enrichment activities that the child participated in were not observed. However, using a regression to summarize these types of parameters seems to perform poorly, at least relative to calculating each $\textrm{AATT}(d)$ individually and then manually averaging them.

\FloatBarrier

Conclusion

Aggregated treatment variables are widely used in empirical work focused on causal inference. They often streamline the narrative of the paper, simplify empirical strategies, accommodate data limitations, and improve precision in estimation. But as we have shown in this paper, these conveniences can come at a cost: the marginal effects of an aggregated treatment can be difficult to interpret. The root of the problem lies in the fact that comparisons across different values of an aggregated treatment can include incongruent sub-treatment comparisons. We have shown that these comparisons lead to negative weight issues, so that even if all marginal sub-treatment effects were positive, one could still find a negative aggregate marginal effect.

Our paper also provides a set of solutions to this underappreciated problem in the causal inference literature. First, we proposed non-marginal estimands that avoid incongruent comparisons altogether and which can be implemented even when sub-treatment data are unavailable. This provides a general path forward even when SUTVA fails for the treatment variable the researcher wishes to use. Second, when sub-treatment data are available, researchers can also continue focusing on marginal estimands by applying desirable weighting schemes that restrict attention solely to congruent comparisons. With this in mind, we note that reporting non-marginal estimands has an underappreciated advantage: it offers greater robustness to violations of SUTVA, which may lead cautious researchers to consider non-marginal effects as a conservative complement to marginal estimands.

In order to emphasize issues related to aggregation, many of our results were in the context of no selection (Assumption (ref)), an ideal setting for causal inference. The insights developed in our paper, however, extend beyond this setting. They are also relevant for other identification strategies---including selection on observables, instrumental variables, regression discontinuity designs, difference-in-differences, and bunching. Relative to our setting---where all units are exchangeable and thus comparable (in expectation)---these strategies impose restrictions on the set of permissible comparisons used to estimate causal effects. As a result, they may either attenuate or exacerbate the extent to which incongruent sub-treatment comparisons enter the estimand, relative to the baseline case studied in this paper. A fuller understanding of how aggregation interacts with these identification strategies remains an important direction for future work.

\printbibliography