EconBase
← Back to paper

Heterogeneity Analysis with Heterogeneous Treatments

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

108,232 characters · 26 sections · 21 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Heterogeneity Analysis with Heterogeneous Treatments

\doublespacing

abstract\singlespacing Analysis of effect heterogeneity at the group level is standard practice in empirical treatment evaluation. However, treatments analyzed are often aggregates of multiple underlying treatments which are themselves heterogeneous, e.g. different modules of training programs. In these settings, conventional approaches such as comparing (adjusted) differences-in-means across groups can produce misleading conclusions when underlying treatment propensities differ systematically between groups. This paper develops a novel decomposition framework that disentangles contributions of effect heterogeneity and distinct components of treatment heterogeneity to observed group-level differences. We propose debiased machine learning estimators that adapt to many discrete and/or continuous treatments and limited overlap. We revisit a widely documented gender gap in training returns of an active labor market policy. The decomposition reveals that it is almost entirely driven by women being treated differently rather than different returns from identical treatments. In particular, women are disproportionately targeted towards vocational training tracks with lower unconditional returns. \\[4ex] Keywords: causal inference, double/debiased machine learning, double aggregation, heterogeneous treatment effects, treatment versions \\[1ex] JEL classification: C14, C21

\thispagestyle{empty}

\setcounter{page}{1} \setstretch{1.45}

Introduction

Heterogeneous causal effects are ubiquitous in applied economic research. They are used to evaluate and design policies or interventions for units with different background characteristics. A large literature develops and applies identification and estimation strategies for causal parameters that explicitly take such heterogeneity into account, see e.g. \citeA{Athey2017} or \citeA{Imbens2024CausalSciences} for recent overviews.

Common heterogeneity analysis compares treatment effects for different groups. Identification, estimation and interpretation of such group heterogeneity at different granularity are well-understood for homogeneous treatments, i.e. when each treatment level corresponds to a single potential outcome. In practice, however, the analyzed treatment is often heterogeneous in the sense that it represents aggregations of a more complex effective treatment. For example, an aggregated treatment could be given by an ex-ante design, e.g. access to a training program that then offers different modules. It can also stem from manual ex-post aggregation, e.g. by merging categories or binning different doses of a continuous treatment. Both are prevalent, though often implicit, in applied research. \citeA{Caetano2025CausalTreatment} discuss 19 studies in labor, urban, environmental, and health economics as well as peer effects and time use research falling into these categories.\footnote{Additionally, ex-ante designs are often observed in experimental studies \cite<e.g.>{Schochet2008DoesStudy,Avdeenko2025CostPakistan}. Ex-post aggregation is common in observational studies \cite<e.g.>{Carneiro2011EstimatingEducation,Dobbie2013GettingCity,Aizer2015JuvenileJudges,Carneiro2017AverageIndonesia, Cornelissen2018WhoAttendance,Kurtz2022DoesDepoliticization,Arteaga2023PARENTALATTAINMENT}. In either case, heterogeneity analyses usually focus on heterogeneity between groups defined by characteristics such as age, gender, or household income.} The consequences of heterogeneous treatments for interpretation of estimands aimed at average, i.e. unconditional, effects are well-studied \cite<e.g.>{VanderWeele2013CausalTreatment,Caetano2025CausalTreatment}. In which sense heterogeneity analyses are informative about policy mechanisms and effect heterogeneity once treatments are also heterogeneous, however, is an open question.

This paper studies the causal content of canonical heterogeneity estimands in the presence of treatment heterogeneity. We focus on comparisons of differences-in-means (DiM) between groups used frequently in randomized controlled trials (RCTs) and the respective covariate-adjusted DiM used in observational studies Imbens2015CausalSciences,Chernozhukov2024AppliedAIb. These statistical estimands are scrutinized in a setting in which the analyzed treatment is defined by aggregating effective treatments and heterogeneity groups are defined by aggregated confounders of the effective treatment. This double aggregation setting spans a variety of empirically relevant cases. In particular, it nests heterogeneity analyses at different granularity under homogeneous and heterogeneous treatments as well as a variety of other estimands from the literature.

We provide a detailed decomposition showing how canonical heterogeneity estimands in the double aggregation setting capture that groups react differently to the same treatment (effect heterogeneity), groups are treated differently (treatment heterogeneity), and their non-trivial interactions. In particular, the decomposition highlights that group differences may solely be driven by treatment heterogeneity and therefore not necessarily reflect effect heterogeneity. Similarly, effect heterogeneity could be masked by treatment heterogeneity. Consequently, interpretation and policy conclusions derived from group differences obtained from, e.g., interacted regressions, matching or causal machine learning methods can be misleading.

The decomposition consists of several estimands capturing qualitatively distinct sources of group heterogeneity: 1) A single parameter that can vary if and only if there is effect heterogeneity across groups in the effective treatment and 2) multiple parameters stemming from the presence of heterogeneous treatments. In particular, the latter are informative about different forms of targeting, i.e. over- or under-representation in treatments based on potential outcomes. More specifically, there are three distinct targeting parameters: (i) group targeting with respect to average outcomes, (ii) group targeting with respect to group outcomes, and (iii) individualized targeting with respect to individualized outcomes beyond group membership. This suggests that researchers who analyze group heterogeneity in the presence of heterogeneous treatments should square their interpretations with the potential presence of all of these channels even if the effective treatment and its confounders are not observed. The decomposition provides the key elements for a principled discussion of the interpretation of group differences that goes beyond loosely appealing to selection into the effective treatment.

All decomposition components are identified if the effective treatment is observed and a standard unconfoundedness assumption applies. They can then be used to quantify the contributions of effect and treatment heterogeneity as well as to test necessary conditions for “real” effect homogeneity across groups, i.e. effect homogeneity on the level of the effective treatment. The targeting parameters further enable ex-post evaluation of group specific assignment mechanisms and their difference as measured by their contribution to the (adjusted) DiM.

We apply the decomposition method to an experimental evaluation of the Job Corps active labor market policy where the data contains detailed information about concrete program experience such as the content of (vocational) training. We study its well-documented gender gap in training returns. Consider the following interacted regression model for weekly earnings (in US\$) with a binary treatment $A$ (access to training) and a female indicator

align[align omitted — 129 chars of source]

Estimating the model via OLS yields an interaction coefficient $\hat{\beta}_3 = 7.6$. Figure (ref) presents a decomposition of this estimate.

figure[figure omitted — 800 chars of source]

We can see that \$6.2 out of \$7.6 (80%) of the total gap can be explained by treatment heterogeneity stemming from different vocational training tracks. Moreover, this difference is purely driven by (i) targeting of women into trainings that have, on average, lower returns. It is neither explained by (ii) group targeting based on higher/lower group returns nor (iii) targeting based on individual characteristics that are correlated with being female/male and would predict higher returns from specific vocational trainings. In particular, the gender differences are not driven by effect heterogeneity in the form of heterogeneous returns to different vocational tracks. A nuanced policy conclusion is therefore that training allocation of women within Job Corps could potentially be improved. This is quite different to the conclusion that Job Corps, regardless of the assignment mechanism, yields the highest returns when targeted to men, which could be a conclusion if detected gender differences are attributed to effect heterogeneity.

For estimation, we provide orthogonal/debiased moments for all decomposition estimands. These can be used in conjunction with flexible estimation methods such as machine learning for the nuisance quantities. We study the large sample properties of the resulting decomposition estimators in an asymptotic framework where the dimensionality of the effective treatment is allowed to grow with the sample size, albeit at a slower rate.\footnote{In practice, the number of measured effective treatments is always finite. However, this asymptotic framework provides a more suitable approximation to the finite sample properties of the involved estimators when there is a moderately large number of treatments for a given sample size compared to standard asymptotics that treat their dimension as fixed Chernozhukov2018} We provide high-level sufficient conditions for asymptotic normality that explicitly take the dimension of the effective treatment and the therefore increasing nuisance parameter space into account. The results can be used for conventional statistical hypothesis tests regarding any (combination of) decomposition estimands.

In the context of testing for effect homogeneity, we contrast the local power properties of our decomposition-based procedure to commonly applied norm-based tests within a multi-valued treatment effect (MVTE) setup. We find that conventional MVTE based tests can have trivial power when the dimension of the effective treatment is large, while the power of our decomposition-based testing approach is often invariant to the dimensionality.

The proposed decompositions as well as the estimation and inference method apply to very general aggregates that are defined over potentially complicated effective treatment spaces. We provide explicit results for treatments that can contain (multiple) continuous and (many) discrete components. Our estimator can then directly be interpreted as a weighted local partitioning approach over the generalized treatment space. Under suitable growth and complexity restrictions on the effective treatment and potential outcomes, the estimator is still asymptotically normal around the generalized decomposition estimands without any asymptotic bias from partitioning.

The paper is structured as follows: Section (ref) discusses methodological literature. Section (ref) provides a motivating example. Section (ref) introduces the double aggregation framework, the first decomposition, and discusses policy relevance in general and in the empirical application. Section (ref) discusses extensions and how several results from the literature are nested. Section (ref) discusses hypothesis testing in comparison to the MVTE approach. Section (ref) introduces the debiased machine learning estimators, technical assumptions, and large sample properties. Section (ref) generalizes decompositions and estimation to more general effective treatment spaces. Section (ref) contains additional concluding remarks. The \href{https://mcknaus.github.io/assets/code/RNB_HK2_decomposition.nb.html}{Supplementary R Notebook} and \href{https://hub.docker.com/repository/docker/mcknaus/hk2_decomposition/general}{Docker Hub} provide additional results and replication code. The method is implemented as part of the \href{https://github.com/MCKnaus/causalDML}{causalDML} R package.

Literature

To our knowledge, this is the first paper that formally studies group heterogeneity estimands when treatments are not homogeneous.\footnote{This paper supersedes \citeA{Heiler2021EffectTreatments} who consider conditional average treatment effects and their projections under heterogeneous treatments in a more restrictive two-fold decomposition. The decomposition provided here is more detailed, additionally covers group average treatment effects, and more general effective treatment spaces.} It connects two topics that have been studied extensively in the literature: first, heterogeneity analyses with homogeneous treatments and therefore simple interpretation; and second, complex interpretation of simple average effect estimands with heterogeneous treatments.

Flexible estimation of heterogeneous treatment effects at different granularity is an active research area Athey2016,Athey2017a,Kunzel2017,Zimmert2019NonparametricConfounding,Knaus2021,Nie2021,Semenova2021DebiasedFunctions,Fan2022EstimationData,Kennedy2023TowardsEffects,Chernozhukov2024AppliedAIb,Lechner2024ComprehensiveLearning. Interpretation of the respective estimands is straightforward if analyzed treatment values correspond to a single potential outcome, i.e. treatments are assumed to be homogeneous. This paper provides a detailed decomposition showing how these estimands capture effect heterogeneity, treatment heterogeneity and their interaction when treatments are instead heterogeneous.

Treatment heterogeneity in the sense of continuous or multi-valued effective treatments are discussed in \citeA{Imbens2000}, \citeA{Lechner2001}, or \citeA{Hirano2004TheTreatments}. They consider identification and estimation of dose-response functions or average effects when researchers analyze each effective treatment separately. This paper instead contributes to the literature on interpretation and external validity as a consequence of analyzing aggregates of heterogeneous treatments Hotz2005PredictingLocations,Hotz2006EvaluatingProgram,Stitelman2010TheVariables,Hernan2011CompoundInference,Caetano2025CausalTreatment. Our double aggregation framework nests many results from the literature as special cases VanderWeele2013CausalTreatment,vanderLaan2023NonparametricIntervention,Lee2024BridgingExposures. In particular, the latter focus on standard unconditional target estimands like average treatment effects or average potential outcomes. We instead focus on heterogeneity estimands aiming to uncover heterogeneous treatment effects. We discuss how the decomposition relates to existing results in detail in Section (ref).

Consequences of treatment aggregation for average estimands are also discussed for instrumental variables Angrist1995Two-stageIntensity,Marshall2016CoarseningEstimates,Andresen2021Instrument-basedRestriction,Harris2022InterpretingEducation,Rose2024OnIndicators, regression discontinuity designs Cattaneo2016InterpretingCutoffs, difference-in-differences Callaway2021Difference-in-DifferencesTreatment, and models with spillovers and interactions Manski2013IdentificationInteractions,Vazquez-Bare2023IdentificationExperiments. We focus on heterogeneity estimands in experimental and unconfounded settings and provide an extension to instrumental variable designs. We expect that many results of this paper will have counterparts in alternative research designs.

The conventional literature on estimation of MVTE and group means usually assumes propensity scores that are bounded away from zero Cattaneo2010EfficientIgnorability,Farrell2015,Heiler2021ShrinkageRegressors. In contrast, the theory in this paper allows all or many propensity scores to converge to zero, which is inevitable when considering a growing multi-valued or a continuous effective treatment as the propensities have to integrate up to one. Thus, on the level of the effective treatment, there is an overlap problem by construction. This connects our paper to the literature on inference under many extreme propensity scores, limited overlap, and irregular identification Khan2010IrregularEstimation,Rothe2017RobustOverlap,Hong2020InferenceOverlap,Heiler2021ValidScores,Ma2024TestingOverlap. Our assumption on the dimensionality is closest to the limited overlap assumption for binary treatments made in \citeA{Hong2020InferenceOverlap}. In particular, we limit the growth of the effective treatment dimension such that consistent estimation of its average potential outcomes is technically still possible, albeit at very slow rates as in \citeA{Hong2020InferenceOverlap}. Our fixed-dimensional decomposition estimands, however, are regularly identified and thus can be relatively precisely estimated even when the effective treatment is growing with the sample size or continuous.

The dimension-invariance of the decomposition also translates immediately to more dimension-robust power properties compared to using MVTE estimates in $\ell_p$-norm based tests for homogeneity. In particular, our decomposition yields testable necessary conditions for group homogeneity. Testing group effect homogeneity in the growing effective treatment setup induces a problem similar to but different from testing a mean vector or a difference between vectors in high dimensions. In iid settings, these problems have been studied extensively Zhong2011TestsDesigns,Srivastava2013AData,Steinberger2016TheRegression. Our setup is most reminiscent of the slowly growing high-dimensional problem also studied in \citeA{Kock2019PowerProblems}. However, it is further complicated by its two-sample nature and dependence of potential outcomes across treatment versions. We provide approximate local power results for asymptotically pivotal tests based on $\ell_2$ and $\ell_{\infty}$ distance as well as our decomposition. They demonstrate the relative dimension robustness of the decomposition-based test over the MVTE approach.\footnote{ Inference on treatment effects with two versions has also been analyzed by \citeA{Hasegawa2020CausalTreatment} from a randomization perspective. They show that, for sharp nulls of constant unit-level effects, more powerful inference can be obtained by combining two tests. We instead study frequentist inference with many versions under group average nulls.}

Our estimation method builds on the semiparametric double/debiased machine learning framework Robins1994,Chernozhukov2018,Chernozhukov2022LocallyEstimation. We provide all debiased influence functions for the novel decomposition estimands. A key technical difference to \citeA{Chernozhukov2018} is that our theory allows for an increasing dimension of the effective treatment which encompasses growing nuisance parameter spaces and limited overlap. This complicates inference and requires somewhat stronger, but still feasible rate conditions to control the first stage nuisance/machine learning bias. These requirements, however, are ameliorated by the local superefficiency properties of some of the involved frequency-based estimators. Our theory nests the results in \citeA{Chernozhukov2018} when applied to finite dimensional multi-valued treatments.

Motivating Example

Consider a stylized data generating process (DGP) with randomized access to treatment but potentially non-random selection into two treatment versions. We are interested in understanding heterogeneity with respect to female group indicator $G \in \{0,1\}$. We formalize the setting using effective treatment $T \in \{0,1,2\}$ with 0 the control condition, and 1 (2) representing treatment version 1 (2). Assume a true noiseless outcome model

align[align omitted — 64 chars of source]

where $T_1 = \mathbbm{1}(T=1)$, $T_2 = \mathbbm{1}(T=2)$, and $X$ are some pre-treatment variables, possibly including $G$. Analysts interested in understanding the gender difference in treatment effects would usually run the following regression with randomization indicator $A = \mathbbm{1}(T>0)$, the female indicator, and their interaction

align[align omitted — 114 chars of source]

A non-zero $\beta_3$ would commonly be interpreted as “effect heterogeneity”. However, the interpretation becomes more nuanced with multiple treatment versions. To see this note that $\beta_3$ in (ref) under outcome model (ref) is

align[align omitted — 435 chars of source]

We use this expression to discuss two phenomena that motivate the decompositions provided in the following sections.

Finding Heterogeneity Without Heterogeneous Effects

Consider now the special case where $\tau_2(X) = \tau_2$ in (ref) (effect homogeneity). Still,

align[align omitted — 90 chars of source]

and therefore $\beta_3$ can be nonzero due to treatment heterogeneity. In particular, $\beta_3 \neq 0$ if the treatment versions have different constant effects ($\tau_1 \neq \tau_2$) and the probabilities to receive the versions depend on $G$, i.e. $E[T_1 | A=1, G=1] \neq E[T_1 | A=1, G=0]$.

To make this more concrete, let $\tau_1 = 1$, $\tau_2 = 2$, $E[T_1 | A=1, G=1] = 2/3$, and $E[T_1 | A=1, G=0] = 1/3$. Both treatment versions are beneficial but version 2 yields larger positive effects and women are less likely to receive the best treatment version 2. Under this parametrization $\beta_3 = -1/3$, indicating that the effect is smaller for women. A policy maker could use this evidence to prioritize treatment away from women. However, this would only reinforce that women are already less likely to receive the best treatment. It would not be justified by real effect heterogeneity as there is simply no real effect heterogeneity to be exploited in this example. This motivates the question of how we can isolate real effect heterogeneity and treatment heterogeneity in such and related settings.

Finding Heterogeneity When Group Effects Are Homogeneous

Now consider the case where $\tau_1 = E[\tau_2(X) | A=1, G=1] = E[\tau_2(X) | A=1, G=0] = 0$, i.e. all group specific effects for the treated, and, by randomization of $A$, also the group average treatment effects $E[\tau_2(X) | G=g]$, are zero and therefore homogeneous. Still, we may detect “effect heterogeneity” if the covariance between receiving the second treatment version and its effect differs between the analyzed heterogeneity groups:

align[align omitted — 224 chars of source]

For example, let $X = (X_1,G)$, where $X_1$ is an indicator whether an individual has children. Further, let $\tau_2(X) = X_1-E[X_1]$, $E[T_2 | A=1, G=0, X] = 1/2$, and $E[T_2 | A=1, G=1, X] = 2/3 - 1/3 X_1$ such that only individuals with children benefit from treatment 2, but women have a lower probability to receive the program if they have children. Then, $ \beta_3 = - 1/3 \times Var(X_1 | A=1, G=1) $. We find a negative interaction unless all women in the population have either children or no children despite homogeneous group effects. This phenomenon appears as $\beta_3$, implicitly, is a comparison of outcomes under differently targeted treatments based on the numbers of children. In general, we can detect group heterogeneity even if group effects are homogeneous if treatments are heterogeneous at the same time. By the same token, true effect heterogeneity could be masked or attenuated. The following provides a formal way to understand the underlying mechanisms.

Framework and Decomposition

This section introduces the double aggregation framework for the analysis of heterogeneity estimands in the presence of heterogeneous treatments. We first focus on the group differences-in-means corresponding to interaction coefficients like (ref) in the motivating example of Section (ref). This allows us to develop a decomposition to isolate different sources of group heterogeneity in a relatively simple setting. Later sections generalize the decomposition to more complex settings and show that the framework covers multiple results in the literature as special cases. All derivations are in Appendix (ref).

Double Aggregation Framework

Let $T \in \mathcal{T}$ denote the effective treatment.\footnote{\citeA{Manski2013IdentificationInteractions} uses the term “effective treatment” in the context of interference. Like in our setting, it describes the treatments creating variation in potential outcomes. In the following the term “treatment” refers to effective treatment if not stated differently.} $\mathcal{T}$ denotes the underlying set of possible treatment values; the associated probability measure may be discrete, continuous, or mixed. For exposition, we first focus on the case of discrete effective treatments. The extension to continuous and mixed effective treatments is provided in Section (ref). We assume existence of potential outcomes $\{Y(t): t \in \mathcal{T}\}$. The realized outcome is then obtained as $Y = Y(T)$. For simplicity, we treat larger outcomes as preferable throughout the discussion. Let $X \in \mathcal{X}$ be variables that ensure conditional independence:

assump[Effective Treatment Unconfoundedness] For all $t \in \mathcal{T}$, the potential outcome is conditionally independent $Y(t) \mathrel{\text{\scalebox{1.07}{$\perp\mkern-10mu\perp$}}} T | X$.

\paragraph{Remark 1.} For deriving the decomposition results below $X$ and $T$ may be unobserved. $X$ could include all confounders or determinants of $T$. Thus, Assumption (ref) is essentially non-restrictive, see \citeA{Goldsmith-Pinkham2024ContaminationRegressions} for a similar discussion around their Equation (11). For identification and estimation, however, $X$ and $T$ need to be observed, see Section (ref). $\square$

We now define the analyzed aggregate treatments and heterogeneity groups. The mechanics of our approach are agnostic to whether aggregate treatments are obtained via ex-ante or ex-post aggregation, see \citeA{VanderWeele2013CausalTreatment} for differences in interpretation. Let $a:\mathcal{T}\rightarrow\{0,1\}$ and $g:\mathcal{X}\rightarrow \{0,1\}$ be known measurable treatment and group aggregation functions that indicate whether an observation belongs to a treatment or covariate group. Let $\mathcal{T}_a$ and $\mathcal{X}_g$ denote the level sets (pre-images) of 1 under $a$ and $g$, respectively: $\mathcal{T}_a = \{t \in \mathcal{T}: a(t) = 1\}$, $\mathcal{X}_g = \{x \in \mathcal{X}: g(x) = 1\}$, and equivalently for other aggregation functions $a'$, $g'$.\footnote{For simplicity, any aggregations are defined as mutually exclusive in what follows, meaning that $\mathcal{T}_a \cap \mathcal{T}_{a'} = \emptyset$ for all $a \neq a'$ and $\mathcal{X}_g \cap \mathcal{X}_{g'} = \emptyset$ for all $g \neq g'$. This corresponds to empirical practice, but can in principle be relaxed at the expense of more involved notation.}

The group differences-in-means corresponding to the interaction term in (ref) that is decomposed in this section reads in this general notation

align[align omitted — 268 chars of source]

for some known $a,a'$ and $g,g'$. For (ref) to be well defined we assume throughout that $P(T \in \mathcal{T}_a, X \in \mathcal{X}_g) > 0$ (and analogously for the other aggregation events).

Example continued: In Section (ref), $\mathcal{T} = \{0,1,2\}$, $a(T) = \mathbbm{1}(T > 0) = A$, $a'(T) = \mathbbm{1}(T = 0) = 1-A$, $\mathcal{T}_{a} = \{1,2\}$, $\mathcal{T}_{a'} = \{0\}$, $\mathcal{X} = \{0,1\} \times \{0,1\}$. $g(X) = G$, $g'(X) = 1-G$, $\mathcal{X}_{g} = \{0,1\} \times \{1\}$, $\mathcal{X}_{g'} = \{0,1\} \times \{0\}$. $\square$

Further, for any $x \in \mathcal{X}$, denote

alignat*{3} \mu_t(x) &= E[Y(t)|X = x], &\mu_t &= E[\mu_t(X)], & \\ e_t(x) &= P(T =t|X=x), &e_t &= E[e_t(X)], & \\ \tau_{t,t'}(x) &= \mu_{t}(x) - \mu_{t'}(x), &\quad \tau_{t,t'} &= \mu_{t} - \mu_{t'}.

For any generic function (possibly depending on $t$ and $a$), we write $g(\mathcal{X}_g) = E[g(X)|X\in \mathcal{X}_g]$, e.g. $\mu_t(\mathcal{X}_g) = E[\mu_t(X)|X\in \mathcal{X}_g]$. We also define conditional (group) propensities given treatment aggregate $a$ as

align*[align* omitted — 504 chars of source]

Application: Training Specializations in Job Corps

We use data of an experimental evaluation of the Job Corps (JC) program in 1994-1996 Schochet2019ReplicationStudy to illustrate the practical relevance of the decomposition terms and to support their intuition by linking them to a well-known application. JC operates since 1964 and is the largest training program for disadvantaged youth aged 16-24 in the US \cite<see>[for a detailed description]{Schochet2001NationalOutcomes,Schochet2008DoesStudy}. The roughly 50,000 participants per year receive an intensive treatment as a combination of different components like academic education, vocational training, and job placement assistance. This means that although the variable “access to JC” is a binary indicator, different tracks of JC participation are conceivable.

Several studies document that women benefit less than men from JC access \cite<e.g.>{Schochet2001NationalOutcomes,Schochet2008DoesStudy,Flores2012EstimatingCorps,Eren2014WhoProgram,Strittmatter2019HeterogeneousApproach}. A potential explanation is differential assignment to vocational tracks: men are more often trained for higher-paying craft occupations, while women concentrate in service-sector tracks Quadagno1995TheCorps,Inanc2017GenderGap.

The decomposition developed in the following allows to quantify the mechanisms behind this finding. Specifically, we decompose the gender gap in the intention-to-treat effect of the binary variable indicating random access to JC on weekly earnings four years after random assignment estimated as $\beta_3$ in (ref) and consider 11 versions of the treatment: (i) No JC if eligible individuals did not participate (non-compliers), (ii) JC without vocational training if eligible individuals entered JC but did not receive vocational training, (iii-ix) training in clerical, health, auto mechanics, welding, electrical/electronics, construction, or food sectors, (x) other vocational training, (xi) training in multiple sectors.

We control for 57 pre-treatment variables (see Section (ref) for a detailed discussion of identification) and analyze 9,708 observations. Estimation follows Section (ref), using 5-fold cross-fitted random forests for nuisance parameters.\footnote{We consider a method that is able to flexibly model heterogeneity in the assignment between men and women like random forest as crucial in this application. This is evident by the propensity scores' bimodal patterns based on gender, see Supplementary R Notebook S.2.2 for more details.} The \href{https://mcknaus.github.io/assets/code/RNB_HK2_decomposition.nb.html}{Supplementary R Notebook} provides additional descriptives, robustness checks using sampling weights, and further results.\footnote{Code to replicate the analysis is stored on \href{https://hub.docker.com/repository/docker/mcknaus/hk2_decomposition/general}{Docker Hub}. The data is available as public use file via \href{https://doi.org/10.3886/E113269V1}{https://doi.org/10.3886/E113269V1}.}

Decomposing a Single Group Mean

The main goal in the following is to understand the causal content of group differences-in-means (DiM) as defined in (ref) in the double aggregation framework. For compactness, we first decompose one group mean $E[Y|T\in\mathcal{T}_a,X\in \mathcal{X}_g]$ before combining them in Section (ref) and Section (ref).

We first isolate the part that only depends on group heterogeneity from parts also depending on additional covariates (see Section (ref) how the latter may drive heterogeneity). Under Assumption (ref), which, again, is essentially non-restrictive (see Remark 1), the group mean can be decomposed as

align[align omitted — 381 chars of source]

The first component $d(a,g)$ corresponds to a group mean that would be obtained in a hypothetical DGP with treatments assigned as in a group stratified RCT while leaving the potential outcome distributions untouched. The treatment probabilities in this synthetic RCT estimand correspond to the observed group probabilities $e_{ta}(\mathcal{X}_g)$ and therefore vary only at the group-level. The second term $d_{4}(a,g)$ represents contributions stemming from variables beyond group membership. The numerator is the covariance of the actual treatment probabilities and the respective potential outcomes. It captures individualized targeting going beyond heterogeneity group specific targeting. It is positive (negative) if subgroups with higher potential outcomes for a particular treatment tend to have a higher (lower) probability of receiving it. The denominator normalizes by the probability to be observed in the particular treatment aggregation $a$.

The decomposition in (ref) contains the estimand of an RCT stratified using only group information. However, even in such a setup, treatment heterogeneity may be mistaken for effect heterogeneity as illustrated in Section (ref). To further separate the different sources of heterogeneity, we define three synthetic estimands:

align[align omitted — 240 chars of source]

The parameters in (ref) are the expected results of a group mean under specific counterfactual DGPs: $s_{0}(a)$ if both group treatment probabilities and group average potential outcomes were constant, $s_{1}(a,g)$ if group probabilities were constant but group outcomes varied, and $s_{2}(a,g)$ if group probabilities varied but group outcomes were constant. Actual DGPs corresponding to these synthetic estimands, in particular $s_{0}(a)$ and $s_{2}(a,g)$, are hypothetical and restrictive. However, they allow us to further decompose the stratified RCT:\footnote{The mechanics of this decomposition is similar to the 3-fold Blinder-Oaxaca decomposition Blinder1973WageEstimates, which is obtained by adding and subtracting synthetic estimands corresponding to expected group outcomes under hypothetical covariate distributions. The centering could also be around other reasonable synthetic propensities and averages in relation to other thought experiments, see Appendix (ref) for an example and additional discussion.}

align[align omitted — 579 chars of source]

$d_{0}(a) = s_{0}(a)$ is the constant baseline and does not depend on $g$. $d_{1}(a,g)$ varies only due to outcome heterogeneity and is positive (negative) if group $g$ tends to have higher (lower) than average outcomes. $d_{2}(a,g)$ varies only due to treatment heterogeneity and is positive (negative) if group $g$ is more likely to receive treatments with higher (lower) average potential outcomes. It captures if group $g$ is targeted towards on average better treatments. $d_{3}(a,g)$ captures group targeting based on heterogeneous group outcomes and therefore the interaction between treatment and outcome heterogeneity. It is positive (negative) for group $g$ if its over-represented treatments ($e_{ta}(\mathcal{X}_g) - e_{ta} > 0$) are aligned with relatively higher (lower) group outcomes.\footnote{For example, take $\mathcal{T}_a = \{1,2\}$ with $e_{1a}(\mathcal{X}_g) - e_{1a} = -1/4$ and $e_{2a}(\mathcal{X}_g) - e_{2a} = 1/4$. Then $d_{3}(a,g) = 1$ for groups with $\mu_1(\mathcal{X}_g) - \mu_1 = 4$ and $\mu_2(\mathcal{X}_g) - \mu_2 = 8$, but also if $\mu_1(\mathcal{X}_g) - \mu_1 = -8$ and $\mu_2(\mathcal{X}_g) - \mu_2 = -4$ showcasing that also groups with consistently lower than average outcomes can benefit from group targeting on group outcomes if distances to the outcome averages differ within group.} The decomposition in (ref) serves as the crucial building block to decompose group mean differences in the following but might also be interesting in its own right depending on the application.

figure[figure omitted — 604 chars of source]

Figure (ref) decomposes the four group means entering the group differences-in-means in the JC application. The bars labeled “GM” correspond to the group mean and the waterfall bars to their left track the decomposition terms. The left column shows the group mean decompositions of the control group for females on top and males at the bottom. As there is only one control condition, $e_{ta'}(\mathcal{X}_g) = e_{ta'} = 1$, which simplifies the interpretation. In particular, $d_{2} = d_{3} = 0$ by construction because no control versions can be targeted. $d_{0}$ shows that in the absence of any outcome heterogeneity the weekly earnings of untreated would be $\$196$. However, $d_{1}$ shows that female earnings are $\$35.2$ lower and male earnings are $\$26.1$ higher than this baseline. This documents significant outcome heterogeneity in form of the well-documented gender gap in earnings levels. Individual targeting captured by $d_{4}$ is negligible. This is expected as the control condition is randomly assigned and therefore $e_t(X)$ in the covariance of Equation $\eqref{eq:decomp-cm1}$ is approximately constant.

The right column of Figure (ref) provides the more interesting decomposition of the group that was randomly treated to receive access to JC, while receiving different vocational training. The baseline outcome, $d_{0}$, in the treated group is $\$212$ and therefore higher than in the control group. However, the outcome heterogeneities captured by $d_{1}$ are nearly identical to the control condition suggesting that the treatment did not change the gender earnings gap. Now we turn to the components that might be driven by differential targeting into vocational training. $d_{2}$ is significantly negative ($-\$3.4$, $p$ = 0.02) for women indicating that they predominantly receive training in vocations with relatively low average potential outcomes. In contrast, men receive more often training in on average better paying vocations ($\$2.8$, $p$ = 0.02). In principle, women could achieve higher earnings when trained for service jobs than for craft jobs. Then the targeting of women towards service jobs would exploit group differences leading to positive $d_{3}$ that could even overcompensate the negative $d_{2}$. However, we find no such group targeting of group outcomes or individualized targeting captured by $d_{3}$ and $d_{4}$.

Decomposing a Single Group Difference-in-Means

Following Section (ref) define $\delta_{0}(a,a') := d_{0}(a) - d_{0}(a')$ and $\delta_{j}(a,a',g) := d_{j}(a,g) - d_{j}(a',g) \text{ for } j \in \{1,2,3,4\}$ to decompose a single group DiM as

align[align omitted — 246 chars of source]

The interpretation of the decomposition parameters follows directly from the discussion in Section (ref). $\delta_{1}(a,a',g) = \sum_{t\in\mathcal{T}_a} \sum_{t'\in\mathcal{T}_{a'}} e_{ta} e_{t'a'} [\tau_{t,t'}(\mathcal{X}_g) - \tau_{t,t'}]$ indicates whether group effects systematically differ from average effects and captures real effect heterogeneity. Similarly, $\delta_{2}(a,a',g) = \sum_{t\in\mathcal{T}_a} [e_{ta}(\mathcal{X}_g) - e_{ta}] \mu_t - \sum_{t'\in\mathcal{T}_{a'}} [e_{t'a'}(\mathcal{X}_g) - e_{t'a'}] \mu_{t'}$ captures pure treatment heterogeneity as it can only be non-zero if treatment probabilities are heterogeneous and average outcomes within an aggregation are non-constant. $\delta_2$ is positive (negative) if $g$ is better (worse) targeted on average outcomes in aggregation $a$ than in $a'$.

On the other hand $\delta_3$ and $\delta_4$ capture possible interactions between effect and treatment heterogeneity. Their most straightforward interpretation is that they indicate how the targeting and composition adjustments differ in aggregation $a$ compared to $a'$ for group $g$, see also Appendix (ref) for the complete definitions and additional discussion. The difference between two groups $g$ and $g'$ within one aggregation $a$, $E[Y|T\in\mathcal{T}_a,X\in \mathcal{X}_g] - E[Y|T\in\mathcal{T}_{a},X\in \mathcal{X}_{g'}]$, can be decomposed symmetrically. Most importantly, for the main result in the following section, it is irrelevant which difference is decomposed first, as it uses a path-independent difference in differences.

figure[figure omitted — 416 chars of source]

The left two subfigures of Figure (ref) show the decompositions for women and men in the JC application. They are obtained by taking the difference between the respective right and left subfigures of Figure (ref). The right bars show the group difference-in-means documenting that women are estimated to benefit less from JC access than men. We note that $\delta_{0}+ \delta_{1}$ would be the expected effect if training assignment was fully random. In this hypothetical scenario, no gender differences would show up because $\delta_{1}$ is basically zero indicating that women and men do not react differently to the same training mix. However, the significantly negative $\delta_{2}$ for women shows that their actual treatment assignment is worse than random as they are targeted towards trainings with, on average, lower returns. In contrast, targeting is significantly better than random for men.

Decomposing Group Differences-in-Means

We are now equipped to decompose the primary estimand. Define $\Delta_{j}(a,a',g,g') := \delta_{j}(a,g) - \delta_{j}(a',g') \text{ for } j \in \{1,2,3,4\}$. We obtain that the heterogeneity estimand group differences-in-means as defined in (ref) is composed of

align*[align* omitted — 309 chars of source]

This shows that group heterogeneity may be driven by multiple components that both separately as well as in combination have specific economic interpretations. $\Delta_1$ is the only component that is necessarily driven by effect heterogeneity. To see this note that

equation[equation omitted — 189 chars of source]

showing that it is a weighted average of group average treatment effect differences. $\Delta_1$ is higher (lower) when group $g$ has systematically higher (lower) effects than group $g'$. It can only differ from zero if there is at least some group-specific effect heterogeneity. Thus, it provides a clean test for true group effect heterogeneity that is not contaminated by treatment heterogeneity. Furthermore, it is the only decomposition parameter that can be non-zero when treatments are homogeneous and/or when the effective treatment is fully randomly assigned. The right panel of Figure (ref) shows that pure effect heterogeneity does not contribute to the gender gap in effectiveness in the JC application.

$\Delta_2$, $\Delta_3$ and $\Delta_4$ are the consequence of treatment heterogeneity. There are three levels of implicit targeting parameters capturing separately to which extent treatment probabilities systematically vary with average, group or individual outcomes. $\Delta_{2}(a,a',g,g') = \sum_{t\in\mathcal{T}_a} [e_{ta}(\mathcal{X}_g) - e_{ta}(\mathcal{X}_{g'})] \mu_{t} - \sum_{t'\in\mathcal{T}_{a'}} [e_{t'a'}(\mathcal{X}_g) - e_{t'a'}(\mathcal{X}_{g'})] \mu_{t'}$ represents differences in group targeting with respect to average outcomes. It is purely driven by treatment heterogeneity. $\Delta_{2}$ is positive (negative) if group $g$ relative to $g'$ is better (worse) targeted on average outcomes in aggregation $a$ than in aggregation $a'$. If $a'$ represents a homogeneous control condition ($\mathcal{T}_{a'}=\{0\}$ like in the JC application), the expression simplifies to $\Delta_{2}(a,a',g,g') = \sum_{t\in\mathcal{T}_a} [e_{ta}(\mathcal{X}_g) - e_{ta}(\mathcal{X}_{g'})] \tau_{t,0}$. Then a more simple interpretation arises stating that $\Delta_{2}$ is positive (negative) if group $g$ is better (worse) targeted on average effects compared to group $g'$. The significantly negative $\Delta_{2}$ in Figure (ref) shows therefore that women are targeted towards vocational trainings with on average lower returns to training.

$\Delta_3$ and $\Delta_4$ capture potential interactions between effect and treatment heterogeneity. Their complete formulas are in Appendix (ref); we focus on a high level discussion here. $\Delta_3$ represents differences in targeting beyond the group targeting of average outcomes captured by $\Delta_2$ and is positive (negative) if group $g$ relative to $g'$ is better (worse) targeted on group outcomes in aggregation $a$ than in aggregation $a'$. Similarly, $\Delta_4$ captures differences in individualized targeting based on confounders beyond group targeting covered by $\Delta_2$ and $\Delta_3$. The combination $\Delta_2 + \Delta_3 + \Delta_4$ therefore provides a summary of group differences stemming from targeting. Figure (ref) documents that group and individualized targeting play a negligible role in explaining the group heterogeneity in the JC application.

Examples continued: In Section (ref), all decomposition parameters are zero except for $\Delta_{2} = \beta_3 = -1/3$. The non-zero interaction coefficient is therefore the result of group targeting of average outcomes as women are more likely to receive on average worse treatments like in the JC application. In Section (ref), all decomposition parameters are zero except $\Delta_{4} = \beta_3 = -1/3 \times Var(X_1 | A=1, G=1)$. The non-zero interaction coefficient is therefore the result of individualized targeting heterogeneity as targeting does not depend on $G$ but on $X_1$ creating a covariance even after conditioning on $G$. We do not find evidence for this mechanism in the JC application.

Generic Implications for Heterogeneity Analyses

On a high level, the decompositions clarify how heterogeneous treatments complicate the causal content of seemingly simple heterogeneity estimands substantially. Group differences may be driven by effect heterogeneity ($\Delta_1$), treatment heterogeneity ($\Delta_2$), or interactions of both ($\Delta_3,\Delta_4$). It is important to note that the decomposition applies even if the effective treatment and its confounders are unobserved. Thus, interpretations of heterogeneity analyses should in general be linked to a discussion of treatment heterogeneity. If treatments are plausibly homogeneous, standard interpretations apply. If relevant treatment heterogeneity cannot be ruled out in a concrete application, the decomposition offers a principled framework for interpreting group differences and may prevent overly simplistic policy conclusions. In particular, it may prevent to prematurely conclude that “treatment” $a$ works better than $a'$ for group $g$ compared to $g'$, although no real effect heterogeneity exists. In cases where researchers expect substantial but unobservable variation within the analyzed treatment, they might even conclude that standard heterogeneity analyses are uninformative.

Identification of Decomposition Parameters

All decomposition parameters are identified under a standard unconfoundedness assumption for the multi-valued treatment:

assump[Identifying Assumptions] Assume that $X$ in Assumption (ref) is observable, $T$ is observable, and effective treatment overlap holds $P(T=t|X=x) > 0 \quad\forall~ x \in \mathcal{X}$.

Under Assumption (ref), $\mu_t(x) = E[Y|T=t,X=x]$ such that all decomposition components involving potential outcomes $\mu_t(X)$, $\mu_t(\mathcal{X}_g)$, and $\mu_t$ are identified.

The plausibility of Assumption (ref) depends on the concrete application. In observational studies, unconfoundedness is arguably credible in the evaluation of active labor market policies (ALMPs) when rich background characteristics are measured. \citeA{lechner2013sensitivity} provide evidence that the common practice to control for socio-demographic information, pre-treatment outcomes, regional information and labor market histories remove most of the selection bias in ALMP studies.\footnote{\citeA{Caliendo2017UnobservablePolicies} show that often unobserved personality traits, attitudes, expectations, social networks and inter-generational information, do predict selection into ALMPs, but do not lead to relevant effect differences for wages or employment prospects suggesting commonly observable controls sufficiently capture the relevant confounding.} Consistent with this result, the meta-analysis by \citeA{Card2018WhatEvaluations} finds no evidence that non-experimental evaluations of ALMPs produce markedly different average impacts than experimental studies.

The 57 control variables applied in the JC application contain all categories of variables considered important by \citeA{lechner2013sensitivity} plus health, crime, and JC related variables. These control variables overlap mostly with those of \citeA{Flores2012EstimatingCorps} who also employ an unconfoundedness strategy in JC. Additionally, we show in Supplementary R Notebook S.2.5 that the covariate-adjusted difference between non-compliers and untreated group is insignificant (3.1, S.E.: 5.9) providing evidence that the control variables can successfully account for selection.\footnote{See also \citeA{Bodory2022EvaluatingLearning} for a similar discussion on the validity of unconfoundedness in the JC application.} It is therefore unlikely that selection bias drives the results in our labor economics application.

Assumption (ref) might be less credible in other applications. For example, \citeA{Bernard2024HowCompliance} provide meta-evidence for relevant variation in observational bias using an unconfoundedness based approach when considering a broad set of development economics applications.\footnote{Their primary observational method uses a partially linear model estimated via double debiased machine learning. When effects are heterogeneous, this model does not target the same estimand as experimental studies and, thus, their comparison is not necessarily an indication of bias.} Nevertheless, \citeA{Avdeenko2025CostPakistan} employ an unconfoundedness based decomposition by \citeA{Heiler2021EffectTreatments} in the context of a development economics RCT.

Extensions and Special Cases

Adjusted Group Differences-in-Means

The first natural extension is to consider the adjusted group differences-in-means

align[align omitted — 299 chars of source]

that balances the distribution of confounding factors between the analyzed treatments in observational studies and not only the heterogeneity group indicator.\footnote{In principle, the adjusted DiM could be based on confounding variables that only assure unconfoundedness on the level of $\mathcal{T}_a$ and $\mathcal{T}_{a'}$ respectively. These do not, in general, have to be identical to the $X$ used in Assumption (ref). In fully observational studies, it is usually hard to conceive different confounding variables between effective and aggregated treatments. Our approach can be extended at the expense of more involved derivations and additional selection bias components, see Appendix (ref) for more details.} This is the population quantity behind contrasting aggregated matching estimands between groups or, equivalently, projecting signals for conditional effects onto group indicators Semenova2021DebiasedFunctions. Additionally, it covers a variety of estimands proposed in the literature as special cases; see Section (ref) below.

A single adjusted group mean can be decomposed in a similar manner as the group mean in (ref) but takes into account the adjustment for covariates beyond the heterogeneity group (see derivation in Appendix (ref)):

align[align omitted — 538 chars of source]

$d(a,g)$ is the same synthetic stratified RCT as in (ref). This also means that its decomposition into $d_1(a,g)$, $d_2(a,g)$ and $d_3(a,g)$ and the respective interpretations established in the previous section directly apply. $d_{4'}(a,g)$ captures again whether treatment assignment is individually targeted. However, it uses the covariance with the propensity score conditional on being in aggregation $a$, $e_{ta}(X)$, instead of the unconditional $e_{t}(X)$ in $d_{4}(a,g)$. It therefore takes into account that the effective treatment probabilities within the analyzed treatment $a$ may vary beyond the heterogeneity group $g$.\footnote{To see this observe how the denominators differ when rewriting the individualized targeting components as $d_{4}(\mathcal{T}_a,\mathcal{X}_g) = Cov\left(\frac{e_t(X)}{P(T\in\mathcal{T}_a | X\in \mathcal{X}_g)},\mu_t(X) | X\in \mathcal{X}_g\right)$ and $d_{4'}(\mathcal{T}_a,\mathcal{X}_g) = Cov\left(\frac{e_t(X)}{P(T\in\mathcal{T}_a | X)},\mu_t(X) | X\in \mathcal{X}_g\right)$.} $d_{5}(a,g)$ is a consequence of adjusting the confounder distribution in $a$ to match the population distribution. It contrasts two hypothetical stratified RCTs that leave the potential outcome distribution untouched but use different assignment probabilities. Both assignment probabilities are averages over the conditional probability of treatment $t$ within aggregation $a$. The first averages $e_{ta}(X)$ over the whole group $g$. The second averages $e_{ta}(X)$ over group $g$ within aggregate treatment $a$ only, i.e. it uses the actual group probabilities $e_{ta}(\mathcal{X}_g) = E[e_{ta}(X)|X\in\mathcal{X}_g,T\in\mathcal{T}_a]$. $d_{5}(a,g)$ is therefore driven by variation in propensity scores going beyond the group level and its interaction with treatment specific outcomes.

A single adjusted group DiM is then decomposed as

align[align omitted — 294 chars of source]

with $\delta_{j}(a,a',g) := d_{j}(a,g) - d_{j}(a',g) \text{ for } j \in \{1,2,3,4',5\}$ and adjusted differences-in-means (ADiM) as defined in (ref) as

align*[align* omitted — 409 chars of source]

with $\Delta_{j}(a,a',g,g') := \delta_{j}(a,g) - \delta_{j}(a',g') \text{ for } j \in \{1,2,3,4',5\}$. The interpretation of $\Delta_{1}$, $\Delta_{2}$ and $\Delta_{3}$ are already discussed in the previous section. The interpretation of $\Delta_{4'}$ is analogous to $\Delta_{4}$. Finally, $\Delta_5$ becomes relevant if the confounder distributions in $a$ and/or $a'$ deviate from the population distribution. It tracks how treatment shares are implicitly changed by adjusting for covariate differences on the level of the aggregate treatments. $\Delta_5$ may be driven by treatment heterogeneity alone or its interaction with effect heterogeneity.\footnote{The exact mechanics leading to non-zero $\Delta_5$ are non-trivial and we do not elaborate on them here. However, progress can be made by recalling that it contrasts two different stratified RCTs that each can be decomposed following Equation (ref).}

Appendix (ref) shows the decomposition for ADiM in the JC application. The differences to Figure (ref) are negligible, which is expected as JC access is randomly assigned. Indeed, $\Delta_{4'}$ is small and insignificant similar to $\Delta_{4}$, and $\Delta_5$ is virtually zero (0.01, S.E.: 0.05).

Special Cases and Relations to the Literature

First, we note that under treatment homogeneity with $\mathcal{T}_a = \{1\}$ and $\mathcal{T}_{a'} = \{0\}$ all estimands collapse to familiar quantities and have the standard interpretation. For example, Equation (ref) simplifies to $E[Y(1)|T=1, X\in \mathcal{X}_g] - E[Y(0)|T=0, X\in \mathcal{X}_g]$ representing the group average treatment effect if unconfoundedness holds on the group level.\footnote{To see this note that then $Cov\left(\frac{e_t(X)}{e_t(\mathcal{X}_g)},\mu_t(X) | X\in \mathcal{X}_g\right) = E[Y(t)|T=t, X\in \mathcal{X}_g] - \mu_t(\mathcal{X}_g)$.} Otherwise, it can be decomposed into group effects and confounding bias following, e.g., \citeA{Heckman1998CharacterizingData}. Also, Equation (ref) collapses to the group average treatment effect $E[Y(1) - Y(0)| X\in \mathcal{X}_g]$ under treatment homogeneity.

Furthermore, DiM and ADiM decompositions in Sections (ref) and (ref) are identical in two scenarios. First, if heterogeneity groups and confounders coincide like in an actually group stratified RCT as then $d_{4}(a,g) = d_{4'}(a,g) = d_{5}(a,g) = 0$. Second, if at least the analyzed treatment is (stratified) randomly assigned ensuring that $P(T\in\mathcal{T}_a | X) = P(T\in\mathcal{T}_a | X\in \mathcal{X}_g)$ as then $d_{4}(a,g) = d_{4'}(a,g)$ (see footnote (ref)) and $d_{5}(a,g) = 0$.\footnote{To see this note that then $E[e_{ta}(X)|X\in\mathcal{X}_g] = e_{ta}(\mathcal{X}_g) = \frac{e_t(\mathcal{X}_g)}{P(T\in\mathcal{T}_a | X\in \mathcal{X}_g)}$ for all $t\in\mathcal{T}_a$.}

More importantly, the double aggregation framework and the decomposition of ADiM nests a variety of results in the literature on average estimands. First, consider unconditional/average estimands as special case with $g(X) = 1$ such that $\mathcal{X}_g = \mathcal{X}$. Then, $\delta_0$ is a special case of the composite treatment effect in \citeA{Lechner2002ProgramPolicies}. Further, Proposition 8 in \citeA{VanderWeele2013CausalTreatment} is concerned with $E[E[Y| T \in \mathcal{T}_a, X]] - E[E[Y| T \in \mathcal{T}_{a'}, X]]$, which is the unconditional version of (ref). They establish its interpretation as a contrast of two interventions where effective treatments are assigned according to their conditional population distribution in the different treatment aggregates. Their Proposition 8 is equivalent to the second equality in the derivation of the adjustment mean decomposition in Appendix (ref) when setting $\mathcal{X}_g = \mathcal{X}$. However, we go further in decomposing this contrast. Note that in the unconditional case $\delta_{1}=\delta_{2}=\delta_{3}=0$ and only $\delta_{0}$, $\delta_{4'}$ and $\delta_{5}$ are relevant. The same estimand is discussed by \citeA{Lee2024BridgingExposures} for continuous $T$. Additionally, they propose to consider contrasts of the form $E[E[Y| T \in \mathcal{T}_a, X]] - E[Y]$. As $E[Y] = d_0(a) + d_4(a,g)$ is covered in our framework by setting $a(T) = 1$ and $g(X) = 1$, the decomposition could also be applied to their estimands. \citeA{vanderLaan2023NonparametricIntervention} also consider continuous $T$ and are interested in the adjusted group mean $E[E[Y|T \geq v,X]]$ for some threshold $v$. This corresponds to treatment aggregation function $a(T) = \mathbbm{1}(T \geq v)$. They interpret the estimand as partially stochastic intervention leaving all units with $T \geq v$ at their original treatment level but conditionally randomizing units with $T < v$ to values above the threshold. Our decomposition components $d_0$, $d_{4'}$ and $d_{5}$ could be applied to further decompose this estimand, while $d_{1}=d_{2}=d_{3}=0$ in their setting.

Second, the framework also relates to the causal decomposition of group disparities in \citeA{Yu2025NonparametricDisparities} when grouping all treatments together ($a(T) = 1$). The paper considers a setting with $\mathcal{T}_a = \mathcal{T} = \{0,1\}$ such that $e_{ta}(\cdot) = e_{t}(\cdot)$ and $P(T\in\mathcal{T}_a | X\in \mathcal{X}_g) = 1$. The decomposed statistical estimand reads $E[Y|X \in \mathcal{X}_g] - E[Y|X \in \mathcal{X}_{g'}]$ in our notation. It would be decomposed as $\sum_{j=1}^4 d_j(a,g) - d_j(a,g')$ in our decomposition. $d_4(a,g) - d_4(a,g')$ is identical to the “selection” term in their equation (1). However, $d_j(a,g) - d_j(a,g')$ for $j \in \{1,2,3\}$ represent a different way to decompose the stratified experiment. We note that $d_1(a,g) - d_1(a,g')$ becomes the sum of their “baseline” component $\mu_0(\mathcal{X}_g) - \mu_0(\mathcal{X}_{g'})$ and $e_{1} (\tau(\mathcal{X}_g) - \tau(\mathcal{X}_{g'}))$ in their setting, which is very similar to their “effect” component. The latter uses $e_{1}(\mathcal{X}_g)$ instead of $e_{1}$ to weight the effect difference. While our decomposition does not separate the baseline term, their decomposition does not differentiate between average and group targeting. A hybrid decomposition-based on ours with additionally separated baseline term could therefore be a more detailed alternative to the decomposition in \citeA{Yu2025NonparametricDisparities}.

In summary, the double aggregation framework to analyze heterogeneity estimands in this paper can also be applied to simpler estimands discussed in the literature. This might be interesting on its own in the context of these papers. However, we focus here on its use to decompose heterogeneity estimands. Finally, we note that $\delta_0 + \delta_1$ recovers the random average treatment effect introduced by \citeA{Heiler2021EffectTreatments}, while their $\Delta$ is now decomposed into different components depending on the statistical estimand.

Other research designs

We note that estimands (ref) and (ref) do not cover a variety of important research designs. We provide an extension to multi-valued instrumental-variable-based designs where the treatment can be fully endogenous in Appendix (ref). We believe that many of the contributions of this paper can be extended to other settings such as difference-in-differences. However, even for the arguably simpler estimands (ref) and (ref), the double aggregation setting has not been studied and reveals non-trivial complexities. Thus, we leave further extensions and discussions for future research.

Multi-valued Treatment Effect Analysis versus Decomposition: Local Power Analysis

We now contrast the properties of testing group effect homogeneity in the decomposition framework with the conventional MVTE approach in the case of observed $X$ and $T$. We focus on finite-dimensional $T$ - a setting where both MVTE analysis and the decomposition are feasible. The previous sections highlight that standard heterogeneity estimands (ref) and (ref) might not (only) capture effect heterogeneity in the double aggregation setting. Instead, they implicitly combine effect and treatment heterogeneity. On the other hand, $\Delta_1$ isolates effect heterogeneity at the scale of the (adjusted) DiM. Moreover, it allows us to assess an implication of the arguably empirically relevant null hypothesis of strong group effect homogeneity

align[align omitted — 356 chars of source]

In particular, rejecting weak group effect homogeneity $\Delta_1 = 0$ is sufficient to reject strong group effect homogeneity (ref). However, $H_A$ could hold while $\Delta_1 = 0$ if the heterogeneities offset each other when aggregating over all $t$-$t'$ combinations.\footnote{Appendix (ref) provides similar discussions for testing treatment homogeneity.} This raises the question whether researchers interested in $H_0$ and access to the effective treatment could benefit from testing $\Delta_1 = 0$ instead of testing $H_0$ directly. The latter could, for example, be achieved by $\ell_p$-norm based tests such as Wald or supremum tests that rely on the (vector) of $t$-$t'$-specific MVTE estimates. Thus, the decomposition is not strictly required to assess the strong null. However, it can have significant power advantages in setups where the number of group effects to test $|\mathcal{T}_a|\times|\mathcal{T}_{a'}|$ is large.

We now discuss the local power properties of statistical tests around the strong null hypotheses. For exposition, we focus on a simplified case where the aggregate control group defined by $\mathcal{T}_{a'} = \{0\}$ is homogeneous with potential outcomes $\mu_0(x) = 0$ for all $x \in \mathcal{X}$. Then, the strong null contains $J = |\mathcal{T}_a|$ hypotheses. For MVTE, we consider conventional tests using a Wald ($\ell_2$) or supremum ($\ell_\infty$) statistic for the strong null with critical values that yield exact asymptotic size control. For the decomposition parameter, we test the weak null using a simple $t$-statistic. For all tests, we analyze their approximate, i.e. first order, power properties along sequences of local alternatives around the strong null of the form

align[align omitted — 105 chars of source]

where $\tau(x) = (\tau_{1,0}(x),\dots,\tau_{J,0}(x))'$ for all $x \subseteq \mathcal{X}$ and $\xi = (\xi_1,\dots,\xi_J)'$. This means we consider $J$ differences in effective treatment means between groups $g$ and $g'$ that are in a root-$n$ neighborhood around zero (strong group effect homogeneity).

table[table omitted — 1,714 chars of source]

Table (ref) summarizes the local power properties. Wald- and supremum-based tests have non-trivial power around the strong null. Wald has more power against dense alternatives and supremum against sparse alternatives as expected. However, power deteriorates quickly at an approximate $1/\sqrt{J}$ rate for both even when differences are relevant at each coordinate, i.e. $||\xi||_2^2/J > 0$ (and thus also $||\xi||_\infty > 0$). For the supremum test, large $J$ leads to trivial power $\alpha$. For the Wald test, there can still be non-trivial power in high dimensions if the alternative is very dense in the sense that $||\xi||_2^2/J \nrightarrow 0$. However, when only a few (relative to $J$) group means are different, then the Wald test also deteriorates to trivial power. The decomposition-based test, on the other hand, has local power properties that are more robust to the dimensionality of the effective treatment. In particular, it has non-trivial power as $J\rightarrow \infty$ as long as the local propensity weighted differences are non-zero $\sum_{t\in\mathcal{T}_a}e_{ta}\xi_t \neq 0$. Under sparse deviations from the strong null hypothesis, this metric can also be small. Under denser deviations, however, this can provide non-trivial power even when $J$ is large. Thus, there is a trade-off between using conventional $\ell_p$-based tests and the decomposition analogues for rejecting the strong null: If the dimension of the effective treatment is large, these $\ell_p$ tests have low or trivial power. The decomposition-based test scales significantly better with $J$ but has little to trivial power against very sparse local alternatives and the direction where ($e_{ta}$-weighted) positive and negative deviations over the different treatment groups $\tau_{t,0}(\mathcal{X}_g) - \tau_{t,0}(\mathcal{X}_{g'})$ aggregate to zero.

figure[figure omitted — 849 chars of source]

Figure (ref) depicts the simulated finite sample power as a function of $J$ in a sparse and a dense design, see Appendix (ref) for details. We can see that, in line with the theory, $\Delta_1$ is dominated by conventional tests for most $J$ under the sparse alternative by the $\ell_p$ based tests. However, it clearly dominates under the dense alternative and has near-constant power as $J$ increases.\footnote{ If the researcher's goal is limited to only having large power for strong null (ref), other tests should also be considered. We conjecture that suitably propensity weighted versions of $\ell_p$ based differences, analogously to the decomposition, could have qualitatively similar power curves as the test based on $\Delta_1$ but with $\sum_t e_{ta}\xi_t$ replaced by functions like $\sum_te_{ta}^p|\xi_t|^p$ (which might well be small for large $J$). Moreover, \citeA{Kock2019PowerProblems}, Theorem 5.2 suggests that any of these tests are asymptotically enhanceable in the sense of \citeA{Fan2015PowerTests} over some part of the parameter space. Thus, we do not claim that $\Delta_1$ provides an “optimal” test for strong effect homogeneity hypothesis over subspaces under growing effective treatment dimensions. However, as proxy it can have advantages over common tests based on version-by-version MVTE estimates. }

From a substantive perspective, the decomposition parameter $\Delta_1$ weights group differences by unconditional propensities and is therefore on the scale of the aggregate effect from the stratified experiment. This avoids rejecting effect homogeneity in setups where there is statistical variation in a few versions that are observed for a very small minority in the population. Thus, the particular form of $\Delta_1$ in (ref) aligns the statistical with the economic significance of its components.

In JC we find no evidence against strong group effect homogeneity. $\Delta_{1}$ representing $e_{ta}$ weighted effect differences has a $p$-value of 0.96. The MVTE $\ell_2$ ($p$ = 0.69) and $\ell_\infty$ ($p$ = 0.87) tests do not reject the strong null as well.

Estimation and Inference

Orthogonal Moments Building Blocks and Estimators

We now provide a brief overview of the components used for constructing our orthogonal moments/influence functions. Some of them have a typical augmented inverse probability weighting structure. However, since most decomposition estimands are novel and extend beyond averages, they require correspondingly novel orthogonal moments, see Appendix (ref) for all details. First note that the decomposition estimands consist of many parameters that can vary over covariate and group aggregation. Each of these are obtained by summing over $t$-specific components in the respective aggregate treatment. In particular, for given aggregations $a$ and $g$, each estimand is a sum over linear combinations of eight $t$-specific primitive parameters $\theta_{a,g,t,p}$ for $p=1,\dots,8$. Their influence functions all have a convenient weighted linear structure

align[align omitted — 122 chars of source]

We compactly summarize their functional forms in Table (ref). The specific entries do also depend on nuisance parameters and components of their respective influence functions $\Psi_{Y,\cdot}$. These influence functions have the same linear structure (ref) and all of their components are contained in Table (ref).

table[table omitted — 2,370 chars of source]
table[table omitted — 1,590 chars of source]

The estimators for all decomposition estimands are then obtained by solving the estimating equations for their respective linear combinations based on Table (ref) using estimated nuisances. All unknown aggregate nuisances for this step are obtained via the influence functions in Table (ref). All granular nuisances required for both steps are obtained via $K$-fold cross-fitting Chernozhukov2018.

In particular, for a given $a$ and $g$, we denote as $\theta(a,g) \in \mathbb{R}^{8}$ all the estimands corresponding to the $\mathcal{T}_a$-aggregated (i.e. summed over $t \in \mathcal{T}_a)$ influence functions components from Table (ref). $\hat{\theta}(a,g)$ are their respective estimators. Again, all these decomposition estimands and estimators can be obtained via linear combinations of their respective components. For example, for decomposition parameter $d_0(g) = \sum_{t\in\mathcal{T}_a}e_{ta}\mu_t$, the implied estimator is given by $\hat{d}_0(a) = \sum_{t\in\mathcal{T}_a}\hat{e}_{ta}\hat{\mu}_t$ where all components are obtained by the solution to their influence function evaluated at estimated nuisances

align[align omitted — 290 chars of source]

with additional aggregate nuisance $e_a$ estimated via $n^{-1}\sum_{i=1}^n\left[\mathbbm{1}(T_i \in \mathcal{T}_a) - \hat{e}_a\right] = 0$. The granular nuisances $\hat{\mu}_t(X)$ and $\hat{e}_{t}(X)$ are obtained via cross-fitting using suitable machine learning or other non/semiparametric prediction methods.

Large Sample Properties

We now present the main technical assumptions and large sample properties of the influence function based decomposition estimators and discuss how they extend and nest existing results in the semiparametric debiased machine learning literature. Denote the variance-covariance matrix of the full vector of decomposition component influence functions as $\Sigma$, see Appendix (ref) for a formal definition. Let $J = J_n = |\mathcal{T}_a|$ be the number of treatments in aggregation $a$. Nuisances are $\eta \in \mathcal{H}$ where $\mathcal{H}$ is a convex subset of some normed vector space. Denote the realization set of the estimated nuisance quantities by $\mathcal{H}_n = \mathcal{E}_n \times \mathcal{M}_n \subset \mathcal{H}$ with $ \mathcal{E}_n = E_{0,n}\times E_{1,n}\times\dots\times E_{J,n}$ and $ \mathcal{M}_n = M_{0,n}\times M_{1,n}\times\dots\times M_{J,n}$, where $E_{t,n}$ and $M_{t,n}$ are the realization sets containing estimates $\hat{e}_t(X)$ and $\hat{\mu}_t(X)$ with probability $1-u_n$. Define their slowest $L_q$ error rates as

align[align omitted — 234 chars of source]

We write $a_n \lesssim b_n$ and $a_n \lesssim_P b_n$, whenever $a_n = O(b_n)$ or $a_n = O_p(b_n)$ respectively. If not stated differently, the following assumptions are uniformly over $n$.

{Assumption A (Regularity, Treatment, and Learning Rates)} Let $\underline{c} > 0$ be a universal constant. For given $a$ and $g$, we assume that

enumerate[itemsep=0pt] \singlespacing • (Moments) Conditional potential outcome moments are bounded, i.e. for some $m > 0$ \begin{align*}\sup_{x\in\mathcal{X},t\in\mathcal{T}_a}E[|Y(t)|^{2+m}|X=x] \lesssim 1.\end{align*} • (Eigenvalue) The variance of the influence function is non-degenerate, i.e. the smallest eigenvalue of $\Sigma$ is bounded away from zero. • (Aggregate Overlap) The shares of aggregate treatment and heterogeneity groups are bounded away from zero \begin{align*} \min\{P(\mathcal{X}_g),P(\mathcal{T}_a),P(\mathcal{T}_a|\mathcal{X}_g)\} > c.\end{align*} • (Weak Effective Treatment Overlap) Each treatment group contains a non-trivial share of observations relative to the total number of treatments \begin{align*} Je_t \lesssim 1 or e_t > c for any t \in \mathcal{T}_a, \end{align*} and it does so for all covariates \begin{align*} \sup_{t\in\mathcal{T}_a,x\in\mathcal{X}}\frac{e_t}{e_t(x)} \lesssim 1. \end{align*} • (Machine Learning Rates) The cross-fitted nuisance learners are consistent at sufficiently fast rates \begin{align*} r_{\mu,n,2} + J r_{e,n,2} &= o(1), \\ \sqrt{n}J r_{\mu,n,2} r_{e,n,2} &= o(1). \end{align*} • (Bounded Relative Prediction Error) On the realization set with probability $1-u_n = 1- o(1)$, the worst relative prediction error for the cross-fitted propensities are bounded \begin{align*} \sup_{t\in\mathcal{T}_a,x\in\mathcal{X}}\sup_{\hat{e}_t\in E_{t,n}}\frac{e_t(x)}{\hat{e}_t(x)} \lesssim 1. \end{align*}

}

Assumption A.1 and A.2 are standard regularity conditions. Assumption A.3 is an overlap condition with respect to both treatment and heterogeneity groups. This means that both aggregates are comprised of a sufficiently large share of observations in the population and that their intersection is non-empty. A.4 has two components. First, the propensities of the effective treatment are either bounded away from zero or vanish at a rate of mostly $J^{-1}$. Thus, along sequences of DGPs, the $t$-specific propensities are allowed to converge to zero, i.e. we do not impose strong overlap on the level of the effective treatment. It also accommodates the aggregation of both small and large effective treatment groups. Second, we require that the conditional propensities of the effective treatment are on similar scale as the unconditional propensities. This effectively means that, as the dimensionality of the treatment grows with most groups containing smaller and smaller shares at a particular rate, the propensities for each group vanish comparably fast/slow. A.4 can be inspected visually via the $1/e_t$ rescaled effective treatment propensity score densities close to the boundary. If any of these diverge at zero, weak overlap is likely violated, equivalently to generic propensities under limited overlap Heiler2021ValidScores. The \href{https://mcknaus.github.io/assets/code/RNB_HK2_decomposition.nb.html}{Supplementary R Notebook} S.2.2 provides an example based on the Job Corps data where we do not detect violations of Assumption A.4.

Assumption A.5 is a requirement on the $L_2$ approximation qualities of the nuisance functions. In the case of a fixed $J$, the assumption collapses to the standard debiased machine learning requirements as in \citeA{Chernozhukov2018}. They are stronger than the typical debiased machine learning requirements when $J$ diverges. This is the price of the growing nuisance parameter space. Additionally, one might be worried about the approximation qualities of $r_{\mu,n,2}$ and $r_{e,n,2}$ themselves also depending on $J$. For many learners, however, they tend to counterbalance each other. In particular, conditional means become harder to estimate with vanishing sample frequencies, while propensities actually become easier. The latter is due to the fact that the smaller the frequency, the lower the (conditional) variance of the corresponding random indicator. Thus, estimators of conditional selection probabilities often exhibit superefficiency properties in asymptotic regimes with $J\rightarrow \infty$, which weakens learning rate requirements. For example, for regular parametric nuisance models, we have that $r_{\mu,n,2} \lesssim \sqrt{{J/n}}$ and $r_{e,n,2} \lesssim \sqrt{1/nJ}$. Thus, the learning requirements in A.5 reduce to $ J = o(n^{1/2})$, which is a moderate but relevant constraint on the dimensionality of the effective treatment. In more nonparametric setups, the growth requirements can become stronger but still admit a reasonably growing number of treatments. For example, for $q$-dimensional kernel regression for a regression function with $s$ H\"older continuous derivatives and optimal bandwidth choice, the requirement in A.5 is $ J = o(n^{\frac{1}{2}\left[\frac{2s - q}{2s + q}\right]})$, see also \citeA{Heiler2021EffectTreatments} for additional examples.\footnote{Condition A.5 is reminiscent of, but different from non-parametric projection methods for estimation of heterogeneous treatment effects Semenova2021DebiasedFunctions,Heiler2024HeterogeneousPolarization. There, growing basis functions yield additional bias and learning rates have to be faster compared to unconditional estimands. In our setup, additional bias arises from the effective treatment dimension that yields both an expanding nuisance parameter space and limited overlap.}

A.6 is another assumption related to the approximation quality of the propensity scores. It requires that, uniformly over the covariate space, the ratio of true to estimated propensities should be bounded with high probability. This holds trivially when the propensity scores are uniformly consistent. However, this is not necessary. We expect this to hold for most frequency based methods with conventional loss functions.\footnote{Uniform consistency applies to various nonparametric estimators such as kernel or series regression, see also \citeA{Ma2024TestingOverlap} for semiparametric single index models and \citeA{Hardle2025UniformForests} for generalized random forests.} We obtain the following Theorem:

thm[Large Sample Inference] For any aggregation $a$, $g$, and bounded sequence vector $c_n \lesssim 1$ with $c_n \neq 0$, given Assumptions A.1 - A.6, we have that all decomposition parameters are asymptotically normal \begin{align*} \left[c_n'\Sigma c_n\right]^{-\frac{1}{2}}c_n'\left(\hat{\theta}(a,g) - \theta(a,g)\right) \overset{d}{\rightarrow} \mathcal{N}(0,1). \end{align*} Moreover when $J \left(\frac{J}{n}\right)^{\left[\frac{1}{2}\wedge \frac{m}{2+m}\right]} = o(1)$, then asymptotic normality applies with $\Sigma$ replaced by its sample equivalent with estimated nuisances.

Theorem (ref) implies the asymptotic validity of using the suggested influence-function based estimators for estimation and inference on any decomposition parameters. As the treatment and covariate aggregations are mutually exclusive, this directly implies an analogous result for all decomposition estimands. We note that, for a fixed $J$, our estimators reach their respective semiparametric efficiency bound as they are based on efficient influence functions. The variance assumption relating potential outcome moments to the estimation of $\Sigma$ is sufficient, but not necessary, see Appendix (ref) for more details. Critical values and quantiles for tests and confidence intervals can be obtained using the standard normal distribution. The dimensionality of the effective treatment $J$ does not affect the rate of convergence. This is due to the fact that, while the number of nuisance parameters and potential outcomes grows, the decomposition is always made up of the same 8 components regardless of $J$.

Generalized Treatment Spaces

Generalized Treatment Spaces and Aggregations

Identification, estimation, and inference generalize beyond the discussed multi-valued effective treatment case. We now present the case where treatment and aggregation spaces are allowed to be (union of) uncountable as well as countable or finite sets $\{t_r\} = \{t_1,t_2,\dots\}$ nesting all results in Sections (ref) and (ref). Proofs and further derivations are in Appendix (ref). For any $x \subseteq \mathcal{X}$, we define a (conditional) probability measure of treatment $T$ according to the Lebesgue decomposition theorem as $\nu_{t|x} = \nu_{t|x}^c + \nu_{t|x}^d$. Thus, for any measurable set $S$, we have

align[align omitted — 147 chars of source]

where $f(t|x)$ is a (conditional) density function, $\lambda(\cdot)$ the Lebesgue measure, and $\delta_{t}(S) = \mathbbm{1}(t \in S)$ the Dirac function. We further define the following spaces

align[align omitted — 215 chars of source]

$L^2(\mathcal{T}_a\backslash \{t_r\},\lambda)$ is a function space over measure space $(\mathcal{T}_a\backslash \{t_r\},\mathcal{B},\lambda)$. $\ell^2(\{t_r\})$ is a space of square-summable sequences over $\{t_r\}$. As composite they yield a hybrid space of ordered pairs $\mathcal{L}^2_a = L^2(\mathcal{T}_a\backslash \{t_r\},\lambda) \oplus \ell^2(\{t_r\})$. Lastly, for generic functions (possibly indexed by $x$) $g_1,g_2:\mathcal{T}\rightarrow \mathbb{R}$, we define the following inner products relative to the Lebesgue and counting measure as

align[align omitted — 240 chars of source]

and, consequently, the composite inner product

align[align omitted — 178 chars of source]

Under abuse of notation, when referring to elements in $\mathcal{L}^2_a$, we subsequently denote ordered pairs $\mu = (\{\mu(t): t \in \mathcal{T}_a \backslash \{t_r\}\}, \{\mu({t_r})\})$ and $e_a = (\{f_a(t) : t \in \mathcal{T}_a \backslash \{t_r\}\}, \{e_{t_ra}\})$, where $\mu(t) = \mu_t$ and same for $\mu(x), e_a(x)$, and $e(x)$ for all $x \subseteq \mathcal{X}$.

Generalized Decompositions

We now decompose a the group mean and its adjusted counterpart as well as the synthetic stratified experiment. Paralleling the results in Section (ref), we obtain

align[align omitted — 262 chars of source]

as well as

align[align omitted — 357 chars of source]

We further decompose the synthetic stratified experiment as

align[align omitted — 414 chars of source]

All inner products have the same substantive interpretations as their discrete analogues in Section (ref). The decomposition for the differences in terms of $a$ and $g$ can be obtained from these components as in Section (ref) and Section (ref).

Estimation and Inference

We now extend the results of Section (ref) to the generalized treatment and aggregation spaces. We consider parameter $d_{0}(a)$ in what follows for exposition, but the theory readily extends to all decomposition parameters as well as their linear combinations as in Theorem (ref). The target parameters in the decomposition now live in the general treatment and aggregation spaces $\mathcal{T}$ and $\mathcal{T}_a$ respectively. We still employ the same discretized estimator as in Section (ref). However, we now use $\tilde{J} = J^* + J$ treatments for estimation where $J = J_n$ is used for $\{t_r\} = \{t_r\}_{r=1}^J$ discrete treatments with non-zero probability mass as in Section (ref) and $J^* = J^*_n$ uses the same estimator but for $J^*$ user-defined partitions of the treatment aggregation space minus the $J$ components $\{t_r\}$. In particular, we define the partitioning $\{v_j\}_{j=1}^{J^*}$ such that $v_j \cap v_k = \emptyset$ if $j\neq k$ and $\bigcup_{j=1}^{J^*} v_j = \mathcal{T}_a \backslash \{t_r\}$. The implied target estimand for the decomposition using a given partitioning is then given by pseudo-true parameter

align[align omitted — 136 chars of source]

Theorem (ref) directly implies that the proposed estimator is asymptotically normal around this $d^*_{0}(a)$. However, the true parameter of interest according to population decomposition (ref) equals

align[align omitted — 221 chars of source]

$d^*_{0}(a)$ can be seen as an approximation to $d_{0}(a)$ where propensities are approximated by shares and conditional means by density weighted shares within discrete partitions. We can thus bound their difference as

align[align omitted — 195 chars of source]

with induced norm $||g||_{L^2(T,f,\lambda)} = \langle g,g\rangle_{L^2(T,f,\lambda)}^{1/2}$ defined by the $f$-weighted inner product $\langle g_1,g_2\rangle_{L^2(T,f,\lambda)} = \int_Tg_1(t)g_2(t)f(t)d\lambda(t)$. $P_{J^*,f}\mu$ and $P_{J^*}f$ are the orthogonal projections of $\mu(t) = E[E[Y|T=t,X]]$ and $f(t)$ respectively onto the subspace of piecewise constant functions in an ($f$-weighted) $L^2(\mathcal{T}_a \backslash \{t_r\},\lambda)$ sense

align[align omitted — 387 chars of source]

We obtain the following Theorem:

thmUnder the assumptions of Theorem (ref) with all statements about $J$ and treatment propensities also applying to $J^*$ and corresponding partition propensity scores, we have that \begin{align*} \Sigma_{11}^{-1/2}\sqrt{n}(\hat{d}_{0}(a) - d_{0}(a)) &= Z + B_n^* + o_p(1), \end{align*} where $Z \overset{d}{=} \mathcal{N}(0,1)$ and the discretization bias obeys \begin{align*} B_n^* &\lesssim \sqrt{n}\left(||\mu - P_{f,J^*}||_{L^2(\mathcal{T}_a \backslash \{t_r\},f,\lambda)} + ||f - P_{J^*}f||_{L^2(\mathcal{T}_a \backslash \{t_r\},\lambda)}\right). \end{align*}

Thus, there is a non-random discretization bias that is determined by the properties of the local partitioning projection of the respective treatment distribution over $\mathcal{T}_a$. These can be controlled if the class of potential outcomes and treatment density on $\mathcal{T}_a$ is not too complex. For example, for smooth function classes, a simple requirement for $J^*$ relative to the degree of smoothness assures asymptotic normality around the true decomposition target with asymptotically negligible discretization bias:

corrLet $\mathcal{T}_a$ be compact. If $\mu,f \in C^q(\mathcal{T}_a \backslash \{t_r\})$, then $ B_n^* \lesssim \sqrt{{n}\big/{{J^*}^{2q}}} $. Moreover if $n/({J^*}^{2q}) = o(1)$, then \begin{align*} \Sigma_{11}^{-1/2}\sqrt{n}(\hat{d}_{0}(a) - d_{0}(a)) \overset{d}{\rightarrow} \mathcal{N}(0,1). \end{align*}

Given the constraints on $J$ in the previous section, the overall number of partitions to approximate aggregates for general treatment spaces has to be in a balanced range for the discretization bias $B_n^*$ to be asymptotically negligible. The smoother the functions, the less constraining the lower bound. The upper bounds used in the Assumptions for Theorem (ref) must also apply to both $J$ and $J^*$. For example, when we have twice continuously differentiable potential outcomes and treatment density, $q=2$, $J^*$ must grow somewhere strictly in between rates of $n^{1/4}$ and $n^{1/2}$ for valid inference. $J$ has no lower bound by construction as in Section (ref). Thus, due to its fixed dimension, the decomposition can achieve faster rates of convergence compared to nonparametric estimation of the average structural function with a continuous treatment under comparable smoothness assumptions Colangelo2026DoubleTreatments.

Concluding Remarks

This paper provides a comprehensive package how to interpret group heterogeneity when treatments are heterogeneous in general as well as how distinct drivers of heterogeneity can be estimated and tested for if effective treatment and confounders are observed in particular.

From a practical point of view, our results also serve as a cautionary tale regarding the policy relevance of conventional heterogeneity analysis when treatments are heterogeneous. In addition, collecting data on treatment versions as well as documentation and discussion of assignment mechanisms is likely to improve on policy relevance of the analysis.

From a methodological perspective, the presented estimation approach is limited to setups where $J = o(n^{1/2})$. In setups with even more high-dimensional treatments, the method could be supplemented with additional, data-driven steps that first extract relevant treatment heterogeneity or versions, e.g. based on maximizing within-aggregate heterogeneity. Interesting extensions could also lie in the adaption of the methodology to other target parameters such as quantile treatment effects and/or alternative research designs.

{\setstretch{1.12} }