Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
73,545 characters · 28 sections · 39 citation commands
Treatment Effect Estimators as Weighted Outcomes
\onehalfspacing
\doublespacing
Estimating the effect of treatment $D_i$ on outcome $Y_i$ is a common goal in causal inference. A variety of estimators is available for estimating different target parameters, after arguing for their identification within a particular research design \cite<see, e.g. reviews by>{Imbens2009,Athey2017,Abadie2018EconometricEvaluation,Imbens2024CausalSciences}. Many of these estimators are a “white box” in the sense that they document how the sample is processed to obtain an effect estimate. Parametric regressions come with familiar coefficient outputs and other popular estimators have a representation as linear combination of observed outcomes:
where $\omega_i$ represents the weight assigned to the outcome of observation $i$ in estimating $\hat{\tau}$. Structure (ref) is most prominent in the literature on propensity score matching/weighting \cite<e.g.>{Imbens2015CausalSciences}, balancing estimators \cite<e.g.>{Ben-Michael2021TheInference}, and synthetic controls \cite<e.g.>{Abadie2021UsingAspects}. Furthermore, it is discussed for estimators of the local average treatment effect \cite<e.g.>{Imbens1997EstimatingModels,Abadie2003SemiparametricModels,Soczynski2024AbadiesEffect} as well as for linear regression \cite<e.g.>{Imbens2015MatchingExamples,Chattopadhyay2023OnInference}.
Outcome weights $\omega_i$ have established use cases, such as: (i) covariate balancing checks assessing internal validity in experimental and observational studies \cite<e.g.>{Rosenbaum1984ReducingScore,Rosenbaum1985ConstructingScore}, (ii) target population characterization investigating external validity in IV settings \cite<e.g.>{Abadie2003SemiparametricModels}, or when using OLS Chattopadhyay2023OnInference, (iii) extrapolation diagnostics for estimators that could use negative weights Chattopadhyay2023OnInference, (iv) finite sample estimator stabilization by normalizing weights \cite<e.g.>{Hajek1971CommentOne}, or by trimming extreme weights \cite<e.g.>{Lechner2019PracticalEstimation}, (v) variance estimation \cite<e.g.>[Ch. 19]{Imbens2015CausalSciences}.
In contrast, recent estimators integrating supervised machine learning into the estimation process \cite<see>[for a textbook]{Chernozhukov2024AppliedAI} can be considered as “grey box”. Their multi-step algorithms are transparent and their theoretical properties are well-understood. However, neither coefficients nor outcome weights are currently available to interrogate how these steps jointly process a concrete sample within a concrete implementation. A key take-away of the analysis below is that at least outcome weights can be available for such multi-step estimators.
This paper introduces a simple but general framework to derive and analyze outcome weights of form (ref). We establish conditions for a numerical equivalence between an original estimator representation as moment condition and a unique weighted representation. The framework is applied to derive novel outcome weights for the six seminal instances of double machine learning Chernozhukov2018 and generalized random forest Athey2017a. Knowing the closed-form of the outcome weights has the immediate practical benefit that they can be plugged into established routines for classic weighting estimators. For example, covariate balancing can now be assessed for conditional average treatment effects estimated by causal forest. A second benefit is that the framework naturally allows to investigate basic properties of the weights. In particular, it highlights that implementation decisions control whether outcome weights of the treated sum up to one and of the untreated to minus one. Such weights are often considered as intuitive, reasonable and desirable in the literature \cite<e.g.>{Imbens2015CausalSciences,Soczynski2024AbadiesEffect}. However, the new framework reveals that estimators building on partially linear regression do not satisfy this property in standard implementations.
The paper makes several contributions: (i) It introduces the first general framework to derive outcome weights; (ii) Its application to causal machine learning estimators yields novel outcome weights and provides a blueprint for applying the framework to other estimators; (iii) It illustrates how the new closed-form expressions enable established diagnostic tools from the weighting literature to be integrated off-the-shelf into causal machine learning applications; (iv) The theoretical results about conditions ensuring desirable estimator properties inform implementation decisions and complement the high-level conditions provided in asymptotic analyses; (v) The paper provides an additional piece in the continuing effort to blur the line between outcome weighting and outcome regression methods. \citeA{Bruns-Smith2023AugmentedRegression} show how weighting estimators can be expressed as regression estimators. This paper goes in the opposite direction by showing how estimators involving flexible outcome regression can be expressed as weighting estimators; (vi) The accompanying R package \href{https://github.com/MCKnaus/OutcomeWeights}{OutcomeWeights} computes the weights presented in the paper for general use Knaus2024OutcomeWeights. The presented applications rely on this package and can be replicated in a supplementary \href{https://hub.docker.com/repository/docker/mcknaus/outcome_weights/general}{Docker image}.
Outcome weights in form of (ref) are leveraged as common structure of difference, weighting, subclassification, and matching estimators Smith2005DoesEstimators,Huber2013,Imbens2015CausalSciences. Similarly, the outcome weights are derived and used for ordinary and weighted least squares based estimators Kline2011Oaxaca-BlinderEstimator,Imbens2015MatchingExamples,Jakiela2021SimpleEffects,Chattopadhyay2023OnInference,Hazlett2024UnderstandingSolutions, two-stage least squares (TSLS) Chattopadhyay2021OnInference, and augmented inverse probability weighting implemented with (post-selection) OLS outcome regression Knaus2021ASkills,Chattopadhyay2023OnInference. This paper shows that a broader class of estimators can have this structure with a particular focus on those incorporating flexible outcome regression, while covering the results in the literature as special cases.
Other types of weights are prominent in the causal inference literature but distinct from the outcome weights pursued in this paper. First, balancing weights are the result of a tailored optimization problem to achieve covariate balancing of some prespecified form Graham2012InverseData,Hainmueller2012EntropyStudies,Imai2014CovariateScore,Zubizarreta2015StableData,Zhao2019CovariateFunctions,Kallus2020GeneralizedInference,Armstrong2021Finite-sampleUnconfoundedness,Heiler2022EfficientEffect. Thus, balancing weights are an explicit part of such balancing estimators. While balancing weights are a special case of outcome weights, this paper focuses on estimators where the outcome weights play no explicit role but are implicit in the common characterization of the estimation procedure. \citeA{Chattopadhyay2023OnInference} call such weights “implied weights” in the context of OLS. Second, effect weights are central to understanding the estimand targeted by a given estimator. The pursued structures in this literature are variations of $\operatorname*{\mathbb{E}}[w(X_i)\tau(X_i)]$ where $w(X_i)$ is the weight an estimator assigns to the conditional treatment effect $\tau(X_i)$ in expectation. Effect weights are derived under different identifying and functional form assumptions for OLS Angrist1998EstimatingApplicants,Angrist1999EmpiricalEconomics, Humphreys2009BoundsProbabilities, Aronow2016DoesEffects, Goldsmith-Pinkham2021ContaminationRegressions, Soczynski2022InterpretingWeights, TSLS Imbens1994IdentificationEffects, Angrist1995Two-stageIntensity,Heckman2005StructuralEvaluation,Soczynski2020WhenLATEb,Blandhol2022WhenLate, two-way fixed-effects \cite<see>[for overviews] {deChaisemartin2023Two-waySurvey,Roth2023WhatsLiterature}, regression discontinuity estimators Lee2010RegressionEconomics, and panel estimators Chernozhukov2013AverageModels. The main difference between outcome weights and effect weights is that the former apply to observed outcomes and numerically reproduce the estimate without further assumptions, while the latter weight inherently unobservable effects and usually reproduce the estimate only in expectation. Both types of weights have their established use cases and are therefore complementary.
A small but growing body of literature provides nuance to the common notion that it is preferable for the outcome weights to sum to (minus) one within treatment groups. \citeA{Doudchenko2016BalancingSynthesis} and \citeA{Breitung2024AlternativeDesigns} challenge this view for synthetic control estimators, and \citeA{Khan2023AdaptiveEstimation} for average treatment effect estimation. \citeA{Soczynski2024AbadiesEffect} note that some Abadie's Abadie2003SemiparametricModels $\kappa$ estimators are intermediate cases between normalized and unnormalized estimators, e.g. with treated weights summing to one but untreated weights not to summing minus one. This paper adds the observation that partially linear regression based estimators usually produce weights that overall sum up to zero but not to (minus) one for (un)treated.
Finally, the paper adds to recent works establishing numeric equivalences between different estimator representations Bruns-Smith2023AugmentedRegression or estimators Soczynski2023CovariateEffects,Soczynski2024AbadiesEffect for conceptual and/or practical insights.
The estimators under consideration require access to data with $N$ observations indexed by $i = 1,...,N$. The data includes a binary treatment $D_i$, an outcome $Y_i$, covariates $\bm{X_i}$, and an optional binary instrument $Z_i$, all collected in $O_i = (D_i, \bm{X_i}', Y_i, Z_i)'$. The empirical mean of a variable $A_i$ is represented as $\operatorname*{\mathbb{E}}_N[A_i] = N^{-1}\sum_{i=1}^N A_i$.
Many results are stated in matrix notation where bold letters describe vectors or matrices of variables. $\bm{I_k}$ denotes the identity matrix of dimension $k$, $\bm{0_k}$ and $\bm{1_k}$ represent column vectors of length $k$ containing zeros and ones, respectively.
This paper focuses on estimators falling into the class of pseudo-IV estimators (PIVE):
Note that the PIVE representation is neither a unique, nor the most compact representation of an estimator. For example, a representation using a linear score $\mathbb{E}_N \left[\psi_i^a - \hat{\tau} \psi_i^b \right] = 0$ with $\psi_i^a =\tilde{Y}_i \tilde{Z}_i$ and $\psi_i^b =\tilde{D}_i \tilde{Z}_i$ would be equivalent, or vice versa any estimator with a linear score can be written as PIVE with $\tilde{Y}_i = \psi_i^a$, $\tilde{D}_i = \psi_i^b$, and $\tilde{Z}_i = 1$. However, the PIVE structure is essential for the goal of this paper. In particular, separating the pseudo-outcome from the pseudo-instrument makes the derivation and analysis of the outcome weights tractable.
Example (OLS): We use the canonical OLS estimator as a running example to illustrate the general results throughout the paper. Consider a linear outcome model $Y_i = \tau D_i + \bm{X_i'\beta} + \varepsilon$. The OLS estimator for $\tau$ can be expressed as PIVE using the residual-on-residual regression representation of the Frisch-Waugh-Lovell Theorem:
where $\hat{\beta}_{\bm{Y} \sim \bm{X}} := \bm{(X'X)^{-1} X'Y}$ and $\hat{\beta}_{\bm{D} \sim \bm{X}} := \bm{(X'X)^{-1} X'D}$ such that the pseudo-outcome is the outcome residual, and both pseudo-treatment and -instrument are the treatment residual.
Solving Equation (ref) leads to parameter estimate
Now assume that the pseudo-outcome vector can be obtained by multiplying a unique $N \times N$ transformation matrix $\bm{T}$ with the outcome vector, i.e. $\bm{TY} = \bm{\tilde{Y}}$. Then, Equation (ref) can be written in the form of Equation (ref)
leading to a core result of the paper:
This simple result is constructive because it motivates a two step procedure to derive outcome weights:
These steps are illustrated below for a variety of estimators. However, the procedure is general and could be pursued for any other estimator fitting into the PIVE structure.
Example (OLS) continued: The solution of the residual-on-residual regression (ref) in form of Equation (ref) is
where we use the projection matrix $\bm{P_X} := \bm{X(X'X)^{-1} X'}$ to define the residual maker matrix $\bm{M_X} := \bm{I_N - P_X}$, and the treatment residual vector $\bm{\hat{E}} := \bm{M_X D}$. The residual maker matrix is therefore the outcome transformation matrix of OLS and $\bm{\omega}^{ols'} = (\bm{\hat{E}' \hat{E}})^{-1}\bm{\hat{E}' M_X}$ is the outcome weights vector.\footnote{We deliberately do not use that $\bm{M_X}$ is idempotent for illustration purposes but note that also the identity matrix would by a suitable transformation matrix in (ref).}
This section leverages the new framework to provide the first characterization of outcome weights for six seminal instances within the double machine learning (DML) and generalized random forest (GRF) frameworks (marked with $^*$), while also recovering existing results for eight other estimators:
Conveniently it suffices to analyze the three estimators in bold letters - IF, AIPW and Wald-AIPW - because their subsequent estimators follow as special cases (see Figure (ref) for a graphical illustration). Each estimator is typically used at an intersection of three research designs (randomized controlled trials, unconfoundedness or instrumental variables), two aggregation levels (average or conditional effects), and three outcome model assumptions (none, partially linear, or linear models). Appendix (ref) summarizes the causal parameters and settings for which each estimator is usually applied. However, the main text ignores definition, identification and interpretation issues concentrating on the mechanics of the estimators.
The considered estimators require a variety of nuisance parameters in the form of approximated conditional expectations:
Furthermore, define the inverse probability weights of the treated $\lambda_{1,i}^{ipw} := D_i / \hat{D}_i$ and of the untreated $\lambda_{0,i}^{ipw} := (1-D_i) / (1-\hat{D}_i)$.\footnote{We use $\lambda$ to remind us that these weights are on a different scale than the $\omega$ weights. Using the corresponding $\omega_{d,i}^{ipw} := \lambda_{d,i}^{ipw} / N$ definition would unnecessarily complicate notation below.} Similarly, define the instrument inverse probability weights as $\lambda_{1,i}^{ipw,z} := Z_i / \hat{Z}_i$ and $\lambda_{0,i}^{ipw,z} := (1-Z_i) / (1-\hat{Z}_i)$.
The literature knows numerous regression methods to estimate the outcome nuisance parameters $\hat{Y}_i$, $\hat{Y}_{d,i}^d$, and $\hat{Y}_{z,i}^z$. However, the class of smoothers \cite<see e.g.>[Ch. 2-3]{Hastie1990GeneralizedModels} turns out to be crucial for the purpose of this paper. Smoothers produce outcome predictions by weighting/smoothing observed outcomes
where the smoother weight $s_{i\leftarrow j}$ represents the contribution of unit $j$'s outcome to the prediction of unit $i$.\footnote{The arrow notation is adapted from \citeA{Lin2022OnEffect}.} Define also the smoother vector for the outcome prediction of unit $i$ by $\bm{s}_i = (s_{i\leftarrow 1},...,s_{i\leftarrow N})'$ and the $N \times N$ smoother matrix $\bm{S} = [\bm{s}_1~...~\bm{s}_N]'$ such that $\bm{s}_i' \bm{Y} = \hat{Y}_i$ and $\bm{S Y = \hat{Y}}$.
The smoother weights in this paper are explicitly allowed to depend on the outcomes (adaptive smoother) and on random components (random smoother), i.e. $\bm{s}_i := \bm{s}_i(\bm{X_i};\bm{X},\bm{Y},\epsilon_s)$.\footnote{The categorization of smoothers is inspired by \citeA{Curth2024WhySmoothers}.} This covers for example (post-selection) OLS, ridge, spline and kernel (ridge) regressions, regression trees, random forests or boosted trees with data-driven hyperparameter tuning (see Appendix (ref) for further discussion). However, the numerical equivalences established below require the mere existence of a smoother matrix:
Example (OLS) continued: The projection matrix is arguably the most prominent smoother matrix producing fitted values of an OLS regression as $\bm{\underbrace{P_X}_{S^{ols}} Y} = \bm{\hat{Y}^{ols}}$.
The instrumental forest (IF) of \citeA{Athey2017a} runs an $\bm{x}$-specific weighted partially linear IV regression
where the $\bm{x}$-specific weights $\alpha^{if}(\bm{x})$ are obtained by the tailored splitting criterion described in \citeA{Athey2017a} and can be extracted via the get_forest_weights() function of their grf R package Tibshirani2024Grf:Forests. The solution in the form of Equation (ref) is
where $\bm{\hat{R}} = \bm{Z} - \bm{\hat{Z}}$, $\bm{\hat{V}} = \bm{D} - \bm{\hat{D}}$ and $\bm{\hat{U}} = \bm{Y} - \bm{\hat{Y}}$ are the instrument, treatment and outcome residual vectors, respectively. The PIVE structure is therefore established. The next step is to understand whether the pseudo-outcome can be obtained using a transformation matrix. This is only possible if a smoother is applied to obtain the outcome predictions, i.e. Condition (ref)a holds such that $\bm{\hat{U}} = \bm{Y} - \bm{S Y} = (\bm{I_N} - \bm{S}) \bm{Y}$ and
The transformation matrix of IF can therefore be considered as a generalized residual maker matrix. Equation (ref) contains then the first concrete case of Proposition (ref):
Table (ref) compactly shows how seven other estimators (in light gray) follow as special cases of IF. Starting from the dark gray row, we can follow an upward path to the Wald estimator or a downward path to DiM. The white rows between the gray rows document the modifications needed to recover the next estimator. For example, moving from IF to CF uses treatment residuals instead of instrument residuals and the CF specific weights $\bm{\alpha^{cf}}$ in the pseudo-instrument, while pseudo-treatment and transformation matrix remain unchanged. Similarly setting the weights to one recovers PLR from CF and PLR-IV from IF. Continuing the paths up- and downwards replaces the generic predictions with linear projections to recover TSLS and OLS, respectively. Finally, using the projection matrix of a constant recovers Wald estimator and DiM.
Computational remark: The original implementation in the grf package applies a constant in the weighted residual-on-residual regression. This complicates notation but Appendix (ref) provides the details how numerical equivalence between original output of grf and the weighted representation is obtained in the OutcomeWeights package.
Augmented inverse probability weighting (AIPW) is developed in a series of papers \cite<e.g.>{Robins1994,Robins1995AnalysisData,Rotnitzky1998SemiparametricNonresponse,Chernozhukov2018}. AIPW is a PIVE with empirical moment condition
and vector form
The next step is to provide the transformation matrix. This is possible under Condition (ref)b that the outcome predictions are obtained by smoothers such that $\bm{S^d_d Y} = \bm{\hat{Y}^d_d}$. Plugging this into (ref) and rearranging delivers the transformation matrix
and leads to the following result:\footnote{The AIPW implementation of the grf package uses an alternative moment condition. It is equivalent to (ref) in expectation but uses different nuisance parameters and therefore differs numerically. However, also the outcome weights of this variant can be obtained as shown in (ref).}
Table (ref) shows how RA can be obtained by setting all IPW weights to zero. IPW is recovered by setting all entries of the smoother matrices to zero.
\citeA{Tan2006RegressionVariables} propose an AIPW extension for the case of a binary instrument. This estimator has the same structure as the canonical \citeA{Wald1940TheError} estimator but applies AIPW to estimate reduced form and first stage, respectively. Following \citeA{Chernozhukov2018}, the Wald-AIPW empirical moment condition in the form of Equation (ref) reads
and in the form of Equation (ref) becomes
Following similar steps as in Section (ref) establishes another special case of Proposition (ref):
Table (ref) summarizes the involved manipulations to arrive at Wald-RA and -IPW applying similar transformations as for AIPW but for both reduced form and first stage.
This section provides the first characterization of outcome weights for IF, CF, PLR(-IV), and (Wald-)AIPW (the supplementary \href{https://mcknaus.github.io/assets/notebooks/outcome_weights/Theory The results highlight that the availability of outcome weights depends on the estimator implementation. In particular, it requires to apply smoothers for the involved outcome regressions (C(ref)). This excludes methods with non-differentiable objective functions and/or non-linear link functions for outcome prediction, such as Lasso, (penalized) logistic regression, or many neural network architectures. However, it is important to note that the choices for instrument and treatment nuisance parameters do not affect the availability of outcome weights.
Overall, the simple framework of Section (ref) proves very handy for compactly deriving new functional forms of outcome weights and recovering known ones. This is interesting and practically useful in its own right, as the obtained weights can be applied in any established weight-based routine. Additionally, the framework provides a natural lens to investigate basic properties of the outcome weights as we pursue in the following.
The results of the previous section enable users to ex post inspect whether outcome weights fulfill certain properties. For example, weights adding up to one for the treated (i.e. $\sum_i \omega_i D_i = 1$) and to minus one for the untreated (i.e. $\sum_i \omega_i (1-D_i) = -1$) are often considered desirable because they guarantee certain in- and equivariances of estimators Imbens2015CausalSciences, Soczynski2024AbadiesEffect. However, the PIVE framework also allows to analytically investigate the weights properties of estimator implementations. This is conceptually appealing and practically relevant because it permits ex ante control over weights properties. Specifically, we investigate under which conditions estimators fulfill one of the five weights properties collected in Table (ref) spanned by the total, treated, and untreated weight sums, respectively (see Figure (ref) for a graphical illustration).
The literature documents examples for each class in Table (ref).\footnote{The proposed class labels in Table (ref) ensure that all three versions of unnormalized weights would also be labeled as unnormalized by \citeA{Soczynski2024AbadiesEffect}. Although “normalized” is a loosely defined term, it seems reasonable in this context to use it for estimators whose outcome weights sum to zero, to remain consistent with previous work.} Fully-unnormalized weights are associated with inverse probability weighting since \citeA{Hajek1971CommentOne}. (Un)treated-unnormalized weights recently appeared in estimators building on Abadie's Abadie2003SemiparametricModels $\kappa_0$ and $\kappa_1$ where only one group shows weights adding up to (minus) one Soczynski2024AbadiesEffect. Scale-normalized weights are described by \citeA{Soczynski2023CovariateEffects} in the context of covariate balancing propensity scores of \citeA{Imai2014CovariateScore}. Such estimators have treated (untreated) weights summing to (minus) the same non-one constant $c$ and also appear prominently in the analysis of partially linear regression based estimators below. Fully-normalized weights are the norm \cite<see overview in>[Ch. 19]{Imbens2015CausalSciences}.
The outcome weights properties in Table (ref) are determined by three weight sums. This motivates the following protocol to classify PIVE weights:
Recall from Proposition (ref) that PIVE weights take the form $\bm{\omega'} = (\bm{\tilde{Z}'\tilde{D}})^{-1}\bm{\tilde{Z}'T}$. Therefore classifying the weights properties of an estimator boils down to checking the first two or all of the following equations:
This implies that it suffices to investigate the following properties of the transformation matrix as shortcuts to classify the outcome weights:
This shows how the PIVE structure offers substantial complexity reduction streamlining the derivations below to a large extent. It turns out that the weights properties are intimately tied to implementation choices as we first illustrate in the OLS example before moving to more involved cases.
Example (OLS) continued: Only one aspect of the implementation affects OLS weights properties in the sense of Table (ref):
Condition (ref) is fulfilled in any reasonable application. However, making it explicit illustrates how implementation choices affect weights properties. We start by checking whether the weights sum to zero and use shortcut (ref) focusing on the transformation matrix:
We conclude that OLS is always normalized if we include a constant. Next, we investigate whether weights of treated sum up to one via shortcut (ref):
The transformation matrix applied to the treatment recovers the pseudo-treatment, which is sufficient for treated weights adding up to one. Curiously, this holds without further conditions implying that OLS is untreated-unnormalized even without a constant. Both results taken together recover a well-known fact that OLS is fully-normalized under Condition (ref). Additionally, following the proposed framework step by step uncovers a nuisance regarding the case without a constant.
This section collectively introduces implementation details that become relevant in the later derivations. We start with a relatively mild condition:
Most smoothers discussed in the literature fulfill this property. However, \citeA{Curth2024WhySmoothers} note that boosted trees can be an exception.
The next condition is relevant for the treatment group specific outcome nuisances:
This condition is relevant for AIPW estimators and in line with standard implementations forming the group specific outcome models in the respective subgroups.
The next condition is less familiar but important for estimators based on partially linear regression and Wald-AIPW:
This goes against the idea of many flexible estimators to entertain different models for outcome and treatment predictions, respectively. Therefore, this condition is not in line with standard implementations.
The final condition is relevant for all estimators involving an inverse probability weighting component:
C6a is the standard \citeA{Hajek1971CommentOne} normalization and usually recommended in applications busso2014new. C6b is suggested by \citeA{Uysal2011ThreeMethods} and recommended by \citeA{Soczynski2024AbadiesEffect}.
Without further conditions $\bm{\omega}^{if}(\bm{x})$ in (ref) is fully-unnormalized. In the following, we explore conditions leading to (fully-)normalized weights. First, we investigate how $\bm{T^{if} 1_N} = \bm{0_N}$ could be obtained:
This establishes that the standard implementation of IF in grf uses normalized weights because it applies the affine smoother random forest (C(ref)) to estimate the outcome nuisance.
The next question is when treated weights sum to one. To this end, it is sufficient to understand when $\bm{T^{if} D} = \bm{\tilde{D}^{if}}$:
The two results span different scenarios. The practically relevant one being that Condition (ref) holds but Condition (ref)a does not because different treatment and outcome models are applied. This means that in practice IF weights are only scale-normalized but not fully-normalized. Only applying the same affine smoother matrix to predict outcome and treatment ensures fully-normalized weights.\footnote{Curiously, applying the same non-affine smoother ensures that at least treated weights sum to one generalizing the observation regarding OLS without constant in the previous section.}
Recall from Table (ref) that CF, PLR-IV and PLR use the same transformation matrix as IF. Consequently, they are also scale-normalized in standard applications. In contrast, OLS and TSLS apply the same projection matrix to form treatment and outcome predictions such that C(ref)a holds by construction. Again the observations regarding OLS in the previous section immediately apply for TSLS because they share pseudo-treatment and transformation matrix. Also TSLS with a constant is fully-normalized and untreated-unnormalized without a constant.
For completeness observe that the difference in means estimator fulfils by construction both Conditions (ref) and (ref)a, and is therefore always fully-normalized. An overview of conditions and weights properties is collected in Table (ref) below.
First, we investigate under which conditions AIPW is normalized:
This result contains two surprising components. First, we did not apply normalized IPW weights (C(ref)a) to achieve normalized AIPW. This means AIPW is self-normalizing once affine smoothers are applied. Second, normalized IPW weights alone do not normalize AIPW weights as a similar simplification is not possible under C(ref)a only.
The second step investigates when treated weights sum to one:
This means that standard implementations using affine smoothers to estimate outcome nuisances in the (un)treated groups separately are self-fully-normalizing regardless which IPW weights are applied. This implies that RA inherits weights properties from AIPW because it can be considered as applying IPW weights of zero (see Table (ref)). In contrast, IPW can be considered as applying smoother matrices of zeros. These uninformative smoother matrices by construction fulfill C(ref) but not C(ref) such that IPW weights are not (fully-)normalized. This recovers the well-known result of \citeA{Hajek1971CommentOne} regarding IPW as a special case of AIPW. Obviously IPW with explicitly fully-normalized weights (under C(ref)a) are fully-normalized.\footnote{To see this within the framework note that under C(ref)a $\bm{1_N' diag(\bm{\lambda_{1}^{norm}} - \bm{\lambda_{0}^{norm}}) 1_N} = N - N = 0$ and $\bm{diag(\bm{\lambda_{1}^{norm}} - \bm{\lambda_{0}^{norm}}) D} = \bm{1_N} = \bm{\tilde{D}^{aipw}}$ establishing full-normalization.}
We can not directly apply the results of Section (ref) because the pseudo-outcome and -treatment differ. However, to show that the estimator is normalized if affine smoothers are applied for the outcome regressions requires only to change the superscripts:
However, the investigation of the sum of treated weights shows notable differences:
Wald-AIPW is therefore only scale-normalized unless we apply the outcome smoothers to also predict the treatments. This goes against the idea of using different models for each nuisance parameter. Unlike AIPW, Wald-AIPW is therefore not expected to be fully-normalized in standard applications. Another point worth noting is that separating the sample by instrument value when estimating outcome/treatment nuisances - an IV version of C(ref) - is not sufficient to achieve fully-normalized weights of Wald-AIPW.
Similar to the previous section Wald-RA inherits all properties from Wald-AIPW. However, Wald-IPW is always untreated-unnormalized because it can be considered as applying the same zero smoother matrix to outcome and treatment (C(ref)b). Additionally normalizing the weights (C(ref)b) makes Wald-IPW even fully-normalized.\footnote{This follows by considering the full numerator of the weight and not only the transformation matrix such that $\bm{1_N' diag(\bm{\lambda_{1}^{norm,z}} - \bm{\lambda_{0}^{norm,z}}) 1_N} = N - N = 0$ due to Condition (ref)b.} This recovers observations regarding Wald-IPW by \citeA{Soczynski2024AbadiesEffect} within the framework of this paper. Also their findings regarding Abadie's Abadie2003SemiparametricModels $\kappa$ estimators can be obtained in the framework as shown in Appendix (ref).
Table (ref) summarizes the sufficient conditions for closed-form and (fully-)normalized outcome weights.\footnote{Table (ref) in the Appendix provides an extended table collecting which conditions are fulfilled by construction and including results for (un)treated-unnormalized weights for completeness. However, those are rather of academic value and we focus on the practically relevant cases in the main text.} Estimators with a check mark in the second column always have a weighted representation. Those are the estimators based on IPW and OLS where the weights are either obvious or at least well-studied \cite<e.g.>{Imbens1997EstimatingModels,Imbens2015CausalSciences, Imbens2015MatchingExamples, Chattopadhyay2023OnInference}. They are still included to demonstrate the generality of the framework but not to provide new insights. Those are obtained for more sophisticated outcome adaptive estimators for which weighted representations are not available in the literature.
The results collected in Table (ref) highlight the crucial role of implementation details for availability and properties of outcome weights. First, column two documents that researchers can ensure that outcome weights can be accessed ex post by applying smoothers to form outcome predictions as shown in Section (ref). Second, estimator specific implementation decisions ex ante determine certain weights properties. Columns three and four of Table (ref) can serve as look-up table for researchers who want to ensure that a particular implementation of an estimator generates outcome weights of a desired class. They contain several surprising or at least undocumented results regarding six prominent DML and GRF instances:
This section runs an Empirical Monte Carlo Study (EMCS) to illustrate that most standard implementations of DML and GRF are not fully-normalized. EMCS take a real dataset and modify some components such that the ground truth is known in the semi-synthetic dataset \cite<e.g.>{Huber2013,Wendling2018}. Here, we use the treatment, instrument, and covariates of the 401(k) data Chernozhukov2016High-DimensionalR but with a noiseless outcome $Y_i^* = 1 + D_i$. This simulates the most powerful treatment leaving every untreated unit at one and shifting every treated unit to two. We expect estimators to estimate an effect of exactly one in this setting without outcome noise. However, only fully-normalized implementations are guaranteed to achieve this because for them, $\bm{\omega' Y^*} = \bm{\omega'} (\bm{1_N} + \bm{D}) = 1$.
This exercise is run with the DoubleML Bach2024DoubleML:R and the grf Tibshirani2024Grf:Forests R packages applied to 100 bootstrap samples. The nuisance parameters in DoubleML are obtained using honest random forest (affine smoother) or XGBoost (non-affine smoother). Each function uses its default values. Table (ref) summarizes the ten implementations under consideration. The final column shows whether an implementation is fully-normalized according to the theoretical results in Table (ref) and therefore expected to find the “effect” of one.
The theoretical predictions are confirmed in Figure (ref). The boxplots show that only AIPW with an affine smoother finds an effect of exactly one in all bootstrap samples. The other methods deviate from one to varying degrees. The XGBoost Wald-AIPW stands out in estimating effects between -16 and 55 (the graph is truncated). However, also causal/instrumental forests estimate heterogeneous effects between 0.93 and 1.01 although there is no heterogeneity to be found in the provided data. This illustrates the theoretical findings even for DoubleML implementations where the extraction of the outcome weights is currently not possible because the required smoother matrices are not accessible.
The novel outcome weights for DML and GRF can be used in established routines or to develop estimator-specific applications. We illustrate the former with covariate balancing, leaving the latter for future research. As in Section (ref), we use the 401(k) data from \citeA{Chernozhukov2016High-DimensionalR}, but this time with the real outcome “net assets”.
First, we investigate covariate balancing for DML estimated average effects. PLR(-IV) and (Wald-)AIPW are implemented using honest random forests with 2- and 5-fold cross-fitting. Figure (ref) presents canonical balancing plots from the cobalt R package Greifer2024Cobalt:Plots displaying absolute standardized mean differences (SMD). We observe that each method successfully balances the previously unbalanced covariates, in particular the income variable. Furthermore, cross-fitting with 5-folds achieves better covariate balancing compared to 2-folds. This demonstrates how DML outcome weights can be utilized in the design phase described by \citeA{Rubin2007TheTrials}, allowing researchers to commit to the preferred implementation before examining the results.
The supplementary \href{https://mcknaus.github.io/assets/code/Notebook_Application_average_401k.nb.html}{average effects R notebook} also provides point estimates and additional results, such as showing that 10 cross-fitting folds provide no further improvement over 5 folds and that the scale-normalized weights sum to values close to one (0.995 and closer).
Checks like those in Figure (ref) are standard when estimating average effects. Similarly, we can assess covariate balancing for all 9,915 conditional average treatment effects (CATEs) produced by the causal_forest() function of the grf package. As an illustration, we investigate the importance of hyperparameter tuning for causal forests by comparing the default implementation with tune.parameters = "all". Figure (ref) shows boxplots of absolute standardized mean differences (SMDs) for each CATE estimate. The results highlight that tuning the forests substantially improves covariate balancing in this application. The tuned version achieves absolute SMDs of 0.1 or lower, whereas the default settings frequently exceed this threshold, with some values even above 0.2. This highlights how standard diagnostics for average effects can also be applied to CATE estimates.
The supplementary \href{https://mcknaus.github.io/assets/code/Notebook_Application_heterogeneous_401k.nb.html}{heterogeneous effects R notebook} reveals that the imbalances in the default forest coincide with implausible effect sizes ranging from -\$21k to \$78k, whereas the tuned forest yields more plausible estimates between \$8k and \$23k. This highlights the importance of parameter tuning for causal forests in this application. A similar pattern is observed for the instrumental forest, though with higher levels of |SMD|.
The supplementary notebook additionally examines descriptive statistics of the outcome weights multiplied by $2D_i-1$ to switch the sign of the untreated weights for better comparability. It documents that (i) both causal forests use negative weights, though to a limited extent, (ii) instrumental forests assign substantial negative weights to never-takers, consistent with the outcome weights in \citeA{Imbens1997EstimatingModels} and the fact that the 401(k) setting has no always-takers by design, (iii) tuned forests use much smaller weights in absolute values, indicating more stable and reliable estimates, (iv) the sum of weights ranges from 0.98 to 1.02 for the default settings and from 0.995 to 1.005 for the tuned forest, making the tuned forest approximately fully-normalized in this application. Future work should explore whether this represents a general pattern.
More estimators than previously noted can be expressed as weighted outcomes. The paper provides a general framework and derives novel weights for double machine learning and generalized random forest estimators. A key learning is that both availability and properties of the outcome weights depend on implementation choices and are therefore controlled by the user.
The paper focuses on providing general theoretical tools and standard illustrations. This acknowledges that access to their closed-form expressions is a prerequisite for developing new use cases or theoretical results for outcome weights. With the provided tools now available, many follow-up questions arise for future research:
The investigation of the latter point most likely requires to restrict focus to analytically tractable smoothers in contrast to the generic smoothers allowed for in this paper. The fact that the smoothers and therefore the outcome weights may depend on the outcome pose non-trivial challenges. For example, it makes the outcome weights not compatible with approaches to use them for statistical inference as the existing approaches require outcome weights to be independent of the outcomes \cite<e.g.>[Ch. 19]{Imbens2015CausalSciences}. Tailored sample splits as in \citeA{Lechner2018} could ensure the required independence but explorations along these lines are left for future work.
\onehalfspacing