EconBase
← Back to paper

Covariate Balancing and Riesz Regression Should Be Guided by the Neyman Orthogonal Score in Debiased Machine Learning

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

40,292 characters · 20 sections · 30 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Covariate Balancing and Riesz Regression Should Be Guided by the Neyman Orthogonal Score in Debiased Machine Learning

abstractThis position paper argues that, in debiased machine learning, balancing functions should be derived from the Neyman orthogonal score, not chosen only as functions of covariates. Covariate balancing is effective when the regression error entering the score can be represented by functions of covariates alone, and it is the natural finite-dimensional approximation for targets such as ATT counterfactual means. For ATE estimation under treatment effect heterogeneity, however, the score error generally contains treatment-specific components because the outcome regression is a function of the full regressor $X=\left(D,Z\right)$. In that case, balancing common functions of $Z$ can leave the treatment-specific component unbalanced. We therefore advocate regressor balancing, implemented by Riesz regression with basis functions of $X$, as the general balancing principle for DML. The position is not that covariate balancing is invalid, but that covariate balancing should be understood as the special case that is appropriate when the score-relevant regression error is a function of covariates alone.

{\flushleft{{\bf Keywords:} causal inference; covariate balancing; double machine learning; Riesz regression; semiparametric efficiency}}

Introduction

In DML, the balancing functions should be chosen from the regression error that appears in the Neyman orthogonal score. In general, this leads to regressor balancing; covariate balancing is the restricted case in which the balanced functions depend only on covariates.

Covariate balancing and debiased machine learning (DML) are widely used in observational causal inference Chernozhukov2018doubledebiased,Hainmueller2012entropybalancing,Imai2013covariatebalancing. In this study, we reconsider covariate balancing as a finite-dimensional approximation to the balancing condition induced by the Neyman orthogonal score. This viewpoint shows why covariate balancing is effective when the score-relevant regression error can be represented by functions of covariates alone, and why it can be restrictive when treatment effect heterogeneity makes the relevant error depend on the full regressor $X=\left(D,Z\right)$.

Covariate balancing methods estimate the propensity score or construct balancing weights by imposing balance restrictions on functions of covariates Hainmueller2012entropybalancing,Imai2013covariatebalancing,Zubizarreta2015stableweights. DML, in contrast, starts from a Neyman orthogonal score for the target estimand Chernozhukov2018doubledebiased,Chernozhukov2024appliedcausal. These two perspectives are often treated as separate. We argue that they should be connected through the score error: the functions to be balanced should be those that approximate the regression error appearing in the Neyman orthogonal score.

We reconsider covariate balancing from the viewpoint of the Neyman orthogonal score. Although the form of the score is often known, it depends on unknown nuisance parameters. Therefore, we estimate the score by replacing the unknown nuisance parameters with their estimators. From this viewpoint, desirable nuisance parameter estimators are those that reduce the error between the score using the true nuisance parameters and the score using estimated nuisance parameters. In this context, covariate balancing plays an important role because it can eliminate the part of this error that is represented by functions of covariates alone.

However, when the treatment effect is heterogeneous, this cancellation does not generally occur. In such cases, the outcome regression is a function of both treatment and covariates. Therefore, it is more desirable to approximate the Riesz representer using basis functions $\Phi(D,Z)$ that depend on both treatment and covariates. When such basis functions are used and the corresponding balancing condition holds, the part of the score error represented by those basis functions vanishes. This is the basic reason why we focus on regressor balancing.

Example: ATE Estimation

As a running example, we consider average treatment effect (ATE) estimation with observations $\left\{\left(D_i, Z_i, Y_i\right)\right\}^n_{i=1}$ Imbens2015causalinference. Our goal is to estimate the ATE defined as $\theta^{\text{ATE}}_0 \coloneqq {\mathbb{E}}\left[\gamma_0(1, Z_i) - \gamma_0(0, Z_i)\right]$, where $\gamma_0(d, z) \coloneqq {\mathbb{E}}\left[Y_i\mid D_i = d, Z_i =z\right]$ is the regression function. In this problem, the Neyman orthogonal score is given as \[\psi^{\text{ATE}}\left(D_i, Z_i, Y_i; \alpha_0, \gamma_0, \theta^{\text{ATE}}_0\right) \coloneqq \alpha_0(D_i, Z_i)\big(Y_i - \gamma_0(D_i, Z_i)\big) + \gamma_0(1, Z_i) - \gamma_0(0, Z_i) - \theta^{\text{ATE}}_0,\] where $\alpha_0(d, z) \coloneqq \frac{\mathbbm{1}\left(d=1\right)}{e_0(z)} - \frac{\mathbbm{1}\left(d=0\right)}{1 - e_0(z)}$ is the Riesz representer, and $e_0(z) \coloneqq \Pr\left(D = 1\mid Z = z\right)$ is the propensity score. By replacing the unknown $\alpha_0$ and $\gamma_0$ with their estimators $\widehat\alpha$ and $\widehat\gamma$ and solving the estimation equation $\frac{1}{n}\sum^n_{i=1}\psi^{\text{ATE}}\left(D_i,Z_i,Y_i;\widehat\alpha,\widehat\gamma,\theta\right)=0$ for $\theta$, we obtain an estimator $\widehat\theta$ of the ATE. Let $\xi(d,z)\coloneqq \gamma_0(d,z)-\widehat\gamma(d,z)$. The part of the plug-in score error that depends on $\xi$ is $ \frac{1}{n}\sum^n_{i=1} \Big( \widehat\alpha(D_i,Z_i)\xi(D_i,Z_i)-\xi(1,Z_i)+\xi(0,Z_i) \Big)$. If $\xi(d,z)$ is represented by basis functions $\Phi(d,z)$ and $ \frac{1}{n}\sum^n_{i=1}\widehat\alpha(D_i,Z_i)\Phi(D_i,Z_i) = \frac{1}{n}\sum^n_{i=1}\left(\Phi(1,Z_i)-\Phi(0,Z_i)\right) $, then this deterministic component vanishes for the represented part of $\xi$. This is the basic role of regressor balancing in ATE estimation.

Now consider covariate balancing, where the basis depends only on $Z$; that is, $\Phi(d, z) = \widetilde{\Phi}(z)$. The covariate balancing condition is $\frac{1}{n}\sum^n_{i=1}\widehat{\alpha}(D_i, Z_i) \widetilde{\Phi}(Z_i) = \frac{1}{n}\sum^n_{i=1}\left(\widetilde{\Phi}(Z_i) - \widetilde{\Phi}(Z_i)\right) = 0$. In contrast, reducing the error between the scores requires that $\xi(d, z) = \gamma_0(d, z) - \widehat{\gamma}(d, z)$ belong to the linear space spanned by $\widetilde{\Phi}(z)$.

Comparing the two cases, we can see that covariate balancing is more restrictive than regressor balancing from the viewpoint of error minimization for the true Neyman orthogonal scores. Covariate balancing works in some special cases, such as when the treatment effect is homogeneous. In more general cases, regressor balancing is more desirable.

Our Position and Contribution

Our position is that, for DML, the relevant balancing condition should be derived from the Neyman orthogonal score. This condition is regressor balancing when the score-relevant regression error is a function of the full regressor $X$, and it reduces to covariate balancing when that error is a function of $Z$ alone. Thus, covariate balancing is not invalid; it is a special case whose appropriateness depends on the target estimand and the regression component entering the score error. Our contribution is to make this distinction explicit and to organize existing balancing methods through the functions they balance and the way they control weight stability.

Our study builds on several lines of work in causal inference. Riesz representer estimation based on the imbalance $\Delta_n(\widehat\alpha,f)$ has been studied in Chen2014sievem,Chen2015sievewald. Squared error minimization type Riesz representer estimation is proposed in Chernozhukov2021automaticdebiased, and BrunsSmith2025augmentedbalancing shows that augmented balancing weights and regression are closely related in linear spaces. Stable balancing weights directly control covariate balance and weight stability Zubizarreta2015stableweights. Tailored loss methods show that propensity score estimation can be designed to improve covariate balance for the target estimand Zhao2019covariatebalancing. Entropy balancing directly constructs weights that match prespecified covariate moments Hainmueller2012entropybalancing. Kato2025directbias points out that Riesz regression can be derived from the viewpoint of density ratio estimation (DRE), and Kato2026aunified unifies the existing approaches. These works suggest that balancing can reduce the error for the true Neyman orthogonal score. We formally state this relationship and clarify its implications for the choice of balancing functions.

\paragraph{Contents.} Section (ref) formulates the problem within the DML framework. In Section (ref), we introduce candidates of estimators, discuss asymptotic efficiency, and raise issues about finite sample performance. In Section (ref), we show that regressor balancing minimizes the error for the true Neyman orthogonal score. In Section (ref), we explain why Riesz regression automatically induces regressor balancing when basis functions are used. In Section (ref), we reconsider covariate balancing in ATE estimation. In Section (ref), we explain why covariate balancing can be sufficient in ATT estimation. In Section (ref), we discuss related estimators and existing work. Section (ref) discusses experiments, and the last section concludes.

Setup

We introduce the general setup for DML. Let $W=(X,Y)$ be an observation, where $Y$ is an outcome and $X$ is a regressor. In treatment effect problems, we often write $X=\left(D,Z\right)$, where $D$ is a treatment indicator and $Z$ is a vector of covariates. Let $P_X(A)\coloneqq\Pr\left(X\in A\right)$ denote the marginal distribution of $X$. Define $L_2(P_X)\coloneqq\left\{f:{\mathcal{X}}\to{\mathbb{R}}, f\text{ measurable}:\int_{{\mathcal{X}}} f(x)^2dP_X(x)<\infty\right\}$, with functions identified if they are equal $P_X$-almost surely. Let $\gamma_0(x)={\mathbb{E}}\left[Y\mid X=x\right]$ be the regression function.

\paragraph{Observations and estimand.} Suppose that we observe i.i.d. observations $\left\{W_i\right\}^n_{i=1}$ from the distribution of $W$. We consider an estimand $\theta_0$ characterized as a linear functional of $\gamma_0$: \[\theta_0 \coloneqq {\mathbb{E}}\left[m(W; \gamma_0)\right],\] where $m(W; \gamma)$ is a known functional that is linear in $\gamma$. By the Riesz representation theorem, there exists $\alpha_0\in L_2(P_X)$ such that \[{\mathbb{E}}\left[m(W;\gamma)\right]={\mathbb{E}}\big[\alpha_0(X)\gamma(X)\big] \quad\text{for all } \gamma\in L_2(P_X). \] We call $\alpha_0$ the Riesz representer. In DML, the corresponding Neyman orthogonal score is \[\psi(W;\eta_0,\theta_0)\coloneqq \alpha_0(X)\big(Y-\gamma_0(X)\big)+m(W;\gamma_0)-\theta_0,\] where $\eta_0=(\alpha_0,\gamma_0)$.

\paragraph{Examples.} We can derive various estimands by specifying the functional $m$. Examples include ATE and ATT counterfactual mean as follows:

itemize[noitemsep, topsep=0pt, leftmargin=0.40cm] • ATE: $m(W; \gamma) = \gamma(1, Z) - \gamma(0, Z)$. • ATT counterfactual mean: $m(W; \gamma) = \frac{D}{\Pr(D=1)}\gamma(0,Z)$.

Functionals involving derivatives, such as AME, fit the same logic after replacing the $L_2(P_X)$ domain with an appropriate smoothness class.

\paragraph{Our goal.} Our goal is to construct estimators of $\theta_0$ with desirable asymptotic and finite sample properties. We use two criteria to judge the soundness of the estimator. The first criterion is asymptotic efficiency in the sense of semiparametric efficiency theory, that is, the asymptotic variance of the estimator matches the efficiency bound, the theoretically best asymptotic variance among regular estimators. See VanderVaart1998asymptoticstatistics. The second criterion is finite sample performance. In many cases, asymptotic efficiency is theoretically investigated, while finite sample performance is empirically evaluated.

Estimators, Asymptotic Efficiency, and Finite-Sample Issues

This section introduces candidate estimators and discusses their properties. We first use asymptotic efficiency as the benchmark. DML provides a way to construct asymptotically efficient estimators, but asymptotic efficiency does not by itself guarantee finite-sample performance. This motivates the finite-sample perspective based on regressor balancing. The main point is that DML and covariate balancing serve different roles. From the viewpoint of asymptotic efficiency, the ARW estimator is optimal in the sense that its asymptotic variance matches the efficiency bound. However, asymptotic optimality does not settle questions about finite-sample performance.

Candidates of Estimators

In this study, we consider the following three types of estimators: the Augmented Riesz Weighting estimator \[\widehat\theta^{\mathrm{ARW}} =\frac{1}{n}\sum^n_{i=1}\left(m(W_i;\widehat\gamma)+\widehat\alpha(X_i)(Y_i-\widehat\gamma(X_i))\right);\] the Riesz Weighting (RW) estimator $\widehat\theta^{\mathrm{RW}}=\frac{1}{n}\sum^n_{i=1}\widehat\alpha(X_i)Y_i$, and the Regression Adjustment (RA) estimator $\widehat\theta^{\mathrm{RA}} =\frac{1}{n}\sum^n_{i=1}m(W_i;\widehat\gamma)$.

As we explain in the subsequent subsections, the ARW estimator is closely connected to the semiparametric efficiency bound and Neyman orthogonal scores. The RW estimator is a generalization of the IPW estimator in ATE estimation. The RA estimator is a plug-in estimator based only on the regression function. These estimators differ in which nuisance parameter they use and in how they respond to finite sample errors in nuisance estimation.

Asymptotic Efficiency and DML

We consider three types of estimators as candidates. The next question is which estimator is preferable among them. To address this question, we introduce asymptotic efficiency theory.

\paragraph{Asymptotic efficiency bound.} We consider the efficiency bound, called the Le Cam-Hajek bound LeCam1986asymptoticmethods or the semiparametric efficiency bound VanderVaart1998asymptoticstatistics. The asymptotic efficiency bound gives the theoretically best asymptotic variance among regular estimators. For details, see Bickel1998efficientand and VanderVaart1998asymptoticstatistics.

\paragraph{Regular and asymptotically linear (RAL) estimators.} An estimator whose asymptotic variance matches the efficiency bound is called asymptotically efficient. It is known that an RAL estimator with the efficient influence function is asymptotically efficient. RAL estimators can be written as $\sqrt{n}\left(\widehat{\theta} - \theta_0\right) = \frac{1}{\sqrt{n}}\sum^n_{i=1}\psi(W_i; \eta_0, \theta_0) + o_p(1)$ ($n\to \infty$), where $\psi$ is the efficient influence function depending on some nuisance parameter $\eta_0$ and the estimand $\theta_0$.

\paragraph{Neyman orthogonal score and Riesz representer.} In the DML framework, we estimate the estimand by solving the estimation equation with the Neyman orthogonal score and estimated nuisance parameters. Consider the case where we replace the nuisance parameter $\eta_0$ in the efficient influence function with its estimator $\widehat{\eta}$.

If the first order effect of nuisance estimation error vanishes, such an efficient influence function is called a Neyman orthogonal score. Although the efficient score and efficient influence function are sometimes defined differently, we use the same object for simplicity. The Neyman orthogonal score is defined for each estimand and usually takes the following form: \[\psi(W_i; \eta_0, \theta_0) \coloneqq \alpha_0(X_i)\big(Y_i - \gamma_0(X_i)\big) + m(W_i;\gamma_0) - \theta_0,\] where $\alpha_0 \in L_2(P_X)$ is the Riesz representer, and $\eta_0 \coloneqq (\alpha_0, \gamma_0)$ is the nuisance parameter, a set of the Riesz representer $\alpha_0$ and the regression function $\gamma_0$.

The Riesz representer is given from the following Riesz representation theorem: \[{\mathbb{E}}\left[m(W;\gamma)\right]={\mathbb{E}}\big[\alpha_0(X)\gamma(X)\big] \quad\text{for all } \gamma\in L_2(P_X).\]

Under the true nuisance parameter $\eta_0$, ${\mathbb{E}}\left[\psi(W;\eta_0,\theta_0)\right]=0$ holds. By replacing the unknown nuisance parameter $\eta_0$ with its estimator $\widehat{\eta}$ and expectation with the sample mean, we estimate $\theta_0$ by solving the estimation equation $\frac{1}{n}\sum^n_{i=1}\psi(W_i; \widehat{\eta}, \theta) = 0$ for $\theta$.

\paragraph{Asymptotic efficiency of the ARW estimator.} The solution of the above estimation equation is the ARW estimator. Under standard DML conditions, including suitable convergence rates for the nuisance estimators and either empirical process restrictions or cross fitting, the ARW estimator is asymptotically normal and efficient Chernozhukov2018doubledebiased. Our focus is different: even when asymptotic efficiency is guaranteed, finite sample behavior still depends on the empirical error of the estimated score.

Issues of Finite-Sample Performance

Thus, the ARW estimator with suitably constructed nuisance estimators is asymptotically normal and efficient. Therefore, in the asymptotic regime, no regular estimator has a smaller asymptotic variance under the same model. However, this asymptotic optimality does not guarantee finite sample performance. In a finite sample, the estimator can still be sensitive to the quality of the estimated score. Covariate balancing is related to this finite sample concern because it can reduce some components of the score error. The next section formalizes this connection through the Neyman error.

Neyman Error Minimization via Regressor Balancing

We consider estimators $\widehat{\theta}$ with the form $\frac{1}{n}\sum^n_{i=1}\psi\left(W_i; \widehat{\gamma}, \widehat{\alpha}, \widehat{\theta}\right) = 0$. For general $\widehat{\gamma}, \widehat{\alpha}$, we obtain the ARW estimator, while as special cases with $\widehat{\gamma}(x) = 0$ and $\widehat{\alpha}(x) = 0$, we obtain the RW and RA estimators, respectively.

Neyman Error

As discussed above, an ideal estimator would solve $\frac{1}{n}\sum^n_{i=1}\psi\left(W_i;\gamma_0,\alpha_0,\theta\right)=0$ for $\theta$. The plug-in score replaces the true score with its estimated counterpart. We focus on the part of the plug-in score error that is affected by the estimated nuisance functions, and define

align[align omitted — 234 chars of source]

We call this quantity the Neyman error. It measures the empirical discrepancy in the estimated score after removing the target component $m(W_i;\gamma_0)$.

Regressor Balancing

This section defines regressor balancing formally. For a candidate representer $\alpha$ and a function $f \in {\mathcal{F}}$, we define a balancing gap as

align[align omitted — 107 chars of source]

A representer estimate $\widehat\alpha$ is said to exactly balance a class ${\mathcal{F}}$ if $\Delta_n(\widehat\alpha,f)=0$ for all $f\in{\mathcal{F}}$, and to approximately balance ${\mathcal{F}}$ at tolerance $\delta_n$ if $\sup_{f\in{\mathcal{F}}} |\Delta_n(\widehat\alpha,f)|\leq \delta_n$.

Regressor balancing refers to balancing functions $f(X)$ in a function class for the full regressor $X$. When $X=\left(D,Z\right)$, these functions may depend on both treatment and covariates. Covariate balancing is the restricted case in which ${\mathcal{F}}$ contains only functions of $Z$.

Neyman Error Minimization via Regressor Balancing

Using linearity of $m$, it holds that

align[align omitted — 209 chars of source]

where $\varepsilon_i=Y_i-\gamma_0(X_i)$. If the observations used to evaluate the score are independent of the data used to construct $\widehat\alpha$, then, conditional on the training data and on $X_1,\ldots,X_n$, the first term is a mean-zero weighted noise term. Without such sample separation, this term is still part of the empirical score error, but it should not be described as conditionally mean zero. The second term $-\Delta_n\left(\widehat\alpha,\widehat\gamma-\gamma_0\right)$ is the deterministic drift induced by regressor imbalance.

This decomposition gives the main reason for regressor balancing. If $\widehat\gamma-\gamma_0$ belongs to ${\mathcal{F}}$ and $\widehat\alpha$ exactly balances ${\mathcal{F}}$, then the deterministic drift vanishes. More generally, if $\widehat\gamma(X)-\gamma_0(X)=f(X)+r(X)$ holds for all $f\in{\mathcal{F}}$, then exact balancing of ${\mathcal{F}}$ leaves only $\Delta_n(\widehat\alpha,r)$. Thus, regressor balancing reduces the score error only to the extent that the balanced functions approximate the score-relevant regression error.

Riesz Regression as Automatic Regressor Balancing and Automatic Neyman Error Minimization with Approximation using Basis Functions

This section explains how Riesz regression implements regressor balancing when the Riesz representer and the regression error are approximated by basis functions. The main point is simple. Riesz regression estimates the Riesz representer by solving an empirical risk minimization problem. The first-order condition of this problem gives balancing equations for the basis functions. Therefore, Riesz regression is not only an estimator of the Riesz representer. It is also a way to construct weights that reduce the deterministic part of the Neyman error.

Series Estimation of Riesz Representer and Regression Function

Suppose that we use a vector of basis functions $\Phi(X) = (\Phi_1(X),\ldots,\Phi_p(X))^\top$. We approximate the Riesz representer by a linear model $\alpha_\beta(X)=\beta^\top\Phi(X)$. The regression function, or the regression error $\widehat\gamma-\gamma_0$, is also approximated by the same basis functions or by a related basis. This series approximation is useful because the balancing condition can be checked component by component.

For example, in Riesz regression under the squared loss, we estimate $\beta$ by minimizing $\frac{1}{n}\sum^n_{i=1}\alpha_\beta(X_i)^2 -\frac{2}{n}\sum^n_{i=1}m(W_i;\alpha_\beta) +\lambda J(\beta)$, where $\lambda J(\beta)$ is a regularization term with $\lambda\geq0$. If $\widehat\beta$ is a minimizer, $\widehat\alpha(X)=\alpha_{\widehat\beta}(X)$, and $J$ is differentiable at $\widehat\beta$, then the first-order condition for the $j$th coefficient gives $0 = \frac{2}{n}\sum^n_{i=1}\Phi_j(X_i)\widehat\alpha(X_i) -\frac{2}{n}\sum^n_{i=1}m(W_i;\Phi_j) +\lambda \partial_j J(\widehat\beta)$. Equivalently, we have $\Delta_n(\widehat\alpha,\Phi_j) = -\frac{\lambda}{2}\partial_j J(\widehat\beta)$. Therefore, when $\lambda=0$, Riesz regression exactly balances the basis functions $\Phi_j$. When $\lambda>0$, the remaining imbalance is determined by the regularization term.

Regressor Balancing and Neyman Error

The preceding first-order condition explains why Riesz regression is useful for controlling the deterministic part of the Neyman error. Suppose that the regression error can be decomposed as $\widehat\gamma(X)-\gamma_0(X) = \rho^\top\Phi(X)+r(X)$, where $r$ is an approximation error. By linearity, we have $\Delta_n\left(\widehat\alpha,\widehat\gamma-\gamma_0\right) = \sum^p_{j=1}\rho_j\Delta_n(\widehat\alpha,\Phi_j) + \Delta_n(\widehat\alpha,r)$. Thus, balancing the basis functions $\Phi_j$ controls the represented part of the deterministic drift, while the remaining term is determined by the approximation error.

This is the sense in which Riesz regression gives automatic regressor balancing. The balancing condition is not imposed after estimating $\widehat\alpha$. It appears as the first-order condition of the Riesz regression problem. More general versions of Riesz regression, including generalized Riesz regression based on other loss functions and link functions, also yield balancing equations under suitable specifications Kato2026aunified.

ATE Estimation and Reconsidering Covariate Balancing

We now return to ATE estimation and reconsider covariate balancing. In ATE estimation, $X=\left(D,Z\right)$ and $m(W;\gamma)=\gamma(1,Z)-\gamma(0,Z)$. Therefore, for a generic function $f(d,z)$, the regressor balancing condition is $\frac{1}{n}\sum^n_{i=1}\widehat\alpha(D_i,Z_i)f(D_i,Z_i) = \frac{1}{n}\sum^n_{i=1}\left(f(1,Z_i)-f(0,Z_i)\right)$. This condition is defined for functions of the full regressor $X=\left(D,Z\right)$.

Covariate balancing is obtained by restricting $f(d,z)$ to functions that do not depend on $d$. Let $f(d,z)=h(z)$. Then $m(W;f)=h(Z)-h(Z)=0$, and the balancing condition becomes $\frac{1}{n}\sum^n_{i=1}\widehat\alpha(D_i,Z_i)h(Z_i)=0$. For the ATE Riesz representer, this is the usual signed balance between treated and control groups after weighting. Thus, covariate balancing is a restricted case of regressor balancing in ATE estimation.

This restriction is sufficient when the relevant regression error is represented by functions of $Z$ alone. To see the limitation, write the score error as $\xi(D,Z)=\xi_0(Z)+D\xi_1(Z)$. The deterministic part of the ATE score error is $\frac{1}{n}\sum^n_{i=1}\widehat\alpha(D_i,Z_i)\xi(D_i,Z_i) - \frac{1}{n}\sum^n_{i=1}\left(\xi(1,Z_i)-\xi(0,Z_i)\right) =\frac{1}{n}\sum^n_{i=1}\widehat\alpha(D_i,Z_i)\xi_0(Z_i) + \frac{1}{n}\sum^n_{i=1}\widehat\alpha(D_i,Z_i)D_i\xi_1(Z_i) - \frac{1}{n}\sum^n_{i=1}\xi_1(Z_i)$. The first term on the right-hand side is controlled by covariate balancing if $\xi_0$ is in the balanced covariate class. The remaining two terms involve the treatment-dependent component $D\xi_1(Z)$. They generally require a balancing condition for functions that depend on both $D$ and $Z$. Thus, under treatment effect heterogeneity, the outcome regression generally depends on the full regressor and covariate balancing can be restrictive.

This point is closely related to optimal CBPS. Fan2021optimalcovariate considers the decomposition ${\mathbb{E}}\left[Y(0)\mid Z\right]=K(Z)$ and ${\mathbb{E}}\left[Y(1)-Y(0)\mid Z\right]=L(Z)$, and shows that the optimal balancing functions for ATE estimation depend on both $K$ and $L$. In particular, the component related to $L$ is needed when treatment effects are heterogeneous. This result is consistent with our position that the balancing functions should be chosen from the error structure of the Neyman orthogonal score, not only from covariates themselves.

ATT Estimation and Sufficiency of Covariate Balancing

This study does not deny covariate balancing. Rather, it clarifies when covariate balancing is the relevant balancing condition. One important case is ATT counterfactual mean estimation, where the score-relevant regression component is a function of $Z$ alone and covariate balancing can be sufficient.

In ATT estimation, the relevant component of the estimand can be reduced to the counterfactual mean of the untreated potential outcome among treated units. For this estimation, we use $\gamma_0(0,Z)$ and do not use $\gamma_0(1,Z)$. Here, $\gamma_0(0,Z)$ is a function of $Z$ alone. Therefore, if $\gamma_0(0,Z)$ is well approximated by basis functions of $Z$, then balancing those basis functions between the treated group and the weighted control group directly targets the relevant regression component.

This is the main reason why entropy balancing is natural for ATT-type targets. Hainmueller2012entropybalancing proposes entropy balancing as a method for computing weights so that the reweighted control group and the treated group satisfy prespecified moment conditions. In ATT estimation, assigning weights to the control group so that its covariate moments match those of the treated group is a direct finite-dimensional approximation to the counterfactual mean problem.

Thus, the distinction between covariate balancing and regressor balancing depends on the target estimand. For ATT estimation or counterfactual mean estimation in that task, we only use the regression function $\gamma_0(0,Z)$, which can be regarded as a function of $Z$ alone. In contrast, for ATE under treatment effect heterogeneity, the relevant regression function is generally a function of $X=\left(D,Z\right)$. This is why covariate balancing can be sufficient in the former case but restrictive in the latter case.

Discussion

In this section, we discuss related topics. Also see Appendix (ref).

\paragraph{RW estimator.} The RW estimator uses only the estimated Riesz representer and the outcome. Therefore, if the estimated Riesz representer satisfies strong balancing conditions for the relevant regression components, the RW estimator can behave like an estimator based on an orthogonal score. This is because exact balancing can replace the missing regression term in the estimating equation. More precisely, if $\Delta_n(\widehat\alpha,\gamma_0)=0$, then we have $\frac{1}{n}\sum^n_{i=1}\widehat\alpha(X_i)Y_i = \frac{1}{n}\sum^n_{i=1}\left(m(W_i;\gamma_0)+\widehat\alpha(X_i)\left(Y_i-\gamma_0(X_i)\right)\right))\right)}$. Thus, the RW estimator can be written as an infeasible ARW estimator using the true regression function.

When balance is inexact, or when cross fitting is used, the RW and ARW estimators generally differ. In such cases, the ARW estimator is usually more stable because it also uses the estimated regression function. This point is consistent with the view that balancing and regression should be understood together rather than as competing ideas.

\paragraph{Cross fitting.} Cross fitting is important in DML because it weakens empirical process conditions and helps justify the use of flexible nuisance estimators. It also changes how exact balance should be interpreted. If $\widehat\alpha$ is estimated on a training sample, exact balance on that training sample does not imply exact balance on the evaluation sample. Therefore, in cross-fitted DML, regressor imbalance should be assessed on the sample where the score is evaluated. This point is separate from the asymptotic role of cross fitting: cross fitting can justify the use of flexible nuisance estimators, while the remaining imbalance on the evaluation sample describes finite sample score error.

\paragraph{TMLE.} TMLE and regressor balancing modify different nuisance components. TMLE updates the regression function so that an estimating equation is closer to being satisfied. Regressor balancing estimates the Riesz representer so that the empirical Riesz equation is better satisfied. Both approaches are motivated by the same Neyman orthogonal score, but they operate on different parts of the nuisance parameter. This distinction is useful for understanding when RW, ARW, and TMLE can be close to each other. If exact balancing holds for the relevant regression component, the RW estimator can be written in a form close to the ARW estimator. If balance is inexact, or if the regression and Riesz representer are estimated on different samples, the ARW estimator or TMLE can be more stable.

\paragraph{Related work.} This study builds on several lines of work. Hainmueller2012entropybalancing proposes entropy balancing, which constructs weights that exactly match prespecified covariate moments. Zubizarreta2015stableweights proposes stable balancing weights, which minimize weight variability subject to approximate covariate balance constraints. Imai2013covariatebalancing proposes CBPS, which estimates the propensity score by using covariate balancing moment conditions. Zhao2019covariatebalancing proposes tailored loss functions and shows that the choice of loss should depend on the estimand and the link function. Fan2021optimalcovariate studies how to choose covariate balancing functions in CBPS and shows that the optimal choice depends on outcome regression components. BrunsSmith2025augmentedbalancing shows that augmented balancing weights and linear regression can be numerically equivalent in linear spaces. Kato2026aunified develops a unified framework for Riesz representer estimation and shows that suitable Riesz regression problems imply balancing equations. Also see Benmichael2021balancingact.

We summarize the relationship among representative methods through two questions: which functions are balanced and how weight variability is controlled. Entropy balancing balances prespecified functions of $Z$ under an entropy criterion, while stable balancing weights balance prespecified functions of $Z$ while controlling weight variability Hainmueller2012entropybalancing,Zubizarreta2015stableweights. We argue that the main difference is not whether a method uses weights, a propensity score, or a regression adjustment. The main difference is which functions are balanced. If the balanced functions depend only on $Z$, the method is covariate balancing. If the balanced functions depend on the full regressor $X=\left(D,Z\right)$, the method is regressor balancing. The latter is more general and is directly tied to the Neyman error decomposition in ((ref)). Appendix (ref) gives further details.

Simulation Studies

The experiments illustrate the difference between covariate balancing and regressor balancing from the viewpoint of the Neyman error. They are not intended to show that one estimator uniformly dominates another. Instead, they show that the imbalance relevant to the score can differ from ordinary covariate imbalance. We use the squared loss. Here, we report the simulation without cross fitting. We describe the details of the experiments in Appendix (ref). Additional experiments with cross fitting and other losses are reported in Appendix (ref). We also investigate the performance using semi-synthetic data in Appendix (ref).

We generate observations $\left\{\left(D_i,Z_i,Y_i\right)\right\}^n_{i=1}$ with $n=1200$ and repeat the experiment 100 times. The covariates are $Z_i\in{\mathbb{R}}^3$ and follow the standard normal distribution. The treatment is generated from $D_i\sim\mathrm{Bernoulli}\left(e_0(Z_i)\right)$, where $e_0(Z_i)=\mathrm{expit}\left(0.5Z_{i1}-0.4Z_{i2}+0.2\sin(Z_{i3})\right)$. Let $\varphi(Z_i)\in{\mathbb{R}}^{80}$ be the basis constructed from random Fourier features for a Gaussian kernel. The outcome is generated as $Y_i=\mu_0(Z_i)+D_i\tau(Z_i)+\varepsilon_i$, where $\mu_0(Z_i)=\varphi(Z_i)^\top\beta_0$, $\tau(Z_i)=\psi(Z_i)^\top\beta_\tau$, and $\varepsilon_i$ is independent noise with standard deviation $0.05$. The coefficient vectors $\beta_0$ and $\beta_\tau$ are fixed across replications. The target is the sample ATE $\theta_0=\frac{1}{n}\sum^n_{i=1}\tau(Z_i)$.

We compare two ways of estimating the Riesz representer. Covariate balancing uses basis functions depending only on $Z$. Regressor balancing uses the treatment-specific basis $\Phi(D,Z)=\left(D\psi(Z),(1-D)\psi(Z)\right)$, so that the same basis $\psi(Z)$ is used and only the coefficients vary with treatment. The Riesz regularization parameter is $0.01$ in Table (ref). Table (ref) shows that regressor balancing substantially reduces RW RMSE and regressor imbalance. ARW RMSE changes only slightly because the regression adjustment already removes part of the outcome error. Regressor balancing reduces the imbalance of the treatment-specific basis functions. It also improves the RW estimator, which depends directly on the estimated Riesz representer. The improvement for ARW is smaller because ARW also uses the regression adjustment.

table[table omitted — 452 chars of source]
table[table omitted — 832 chars of source]
figure[figure omitted — 541 chars of source]

We also vary the Riesz regularization parameter over $0$, $0.01$, and $0.1$. Table (ref) and Figure (ref) show that the effect of regressor balancing depends on the regularization level. With $\lambda=0$, regressor balancing nearly eliminates regressor imbalance and gives the smallest RW and ARW RMSE in this design. Increasing $\lambda$ stabilizes Riesz regression but relaxes the balancing condition, and the regressor imbalance increases. The covariate balancing results show a different pattern. Very small regularization yields small covariate imbalance, but RW RMSE can be large. This supports the view that balance should be interpreted together with weight stability.

Conclusion

We reconsidered covariate balancing from the viewpoint of DML. The relevant object in DML is the Neyman orthogonal score, and the error of the estimated score depends on the regression components entering that score. Covariate balancing is useful when those components are functions of covariates alone, and it is sufficient in important cases such as ATT counterfactual mean estimation. However, for ATE under treatment effect heterogeneity, covariate balancing can be restrictive because the outcome regression depends on both treatment and covariates. Regressor balancing provides the more general condition: it balances basis functions of the full regressor and can remove treatment-dependent components of the score error. We therefore recommend reporting score-relevant regressor imbalance, not only covariate imbalance, when using balancing methods for estimating causal parameters.

\onecolumn