Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
21,336 characters · 14 sections · 30 citation commands
Riesz Regression As Direct Density Ratio Estimation
We describe the connection between Riesz representer estimation in causal inference and density ratio estimation (DRE) Chernozhukov2022automaticdebiased,Sugiyama2012densityratio. With the advance of various machine learning methods, debiased machine learning (DML) approaches garnered attention in economics and finance Movaghari2025coporatecash,Kumar2024unveilingthe,Shi2019adaptingneural. The automatic DML (ADML) framework formulates causal and structural parameter estimation problems using the Riesz representer and provides Riesz regression, a general tool for Riesz representer estimation Chen2015sievesemiparametric,Chernozhukov2021automaticdebiased. This paper shows that (i) the Riesz representer can be written as a signed density ratio (Section (ref)) and (ii) the Riesz regression criterion in Chernozhukov2021automaticdebiased is identical to the least-squares importance fitting (LSIF) criterion in Kanamori2009aleastsquares for DRE (Section (ref)). In discrete treatment problems, the first result yields closed-form Riesz representers for average treatment effect (ATE), average treatment effect on the treated (ATT), weighted ATE, policy value, and other linear contrasts. For simplicity, we specialize the discussion to the ATE, a canonical target in causal inference.
This observation is useful for applied work because the representer is the weighting component in Neyman orthogonal scores. Writing it as a density ratio turns direct representer fitting into direct reweighting. The same step links orthogonal score estimation to methods already familiar in causal inference, including tailored loss, entropy balancing, stable balancing weights, nearest neighbor matching, and augmented balancing weights Zhao2019covariatebalancing,Hainmueller2012entropybalancing,Zubizarreta2015stableweights,Lin2023estimationbased,BrunsSmith2025augmentedbalancing. It also connects Riesz estimation to model classes and regularization tools developed for DRE, such as RKHS estimators, neural networks, nonnegative Bregman corrections, and telescoping DRE Kanamori2012statisticalanalysis,Kato2021nonnegativebregman,Zheng2022anerror,Rhodes2020telescopingdensityratio. Related work develops this relationship for direct bias correction, Bregman-Riesz regression, nearest neighbor matching, and unified treatments of weighting and matching Kato2025directbias,Kato2025nearestneighbor,Kato2026unifiedframework.
This focus on the Riesz representer also helps with interpretation. In empirical work, the outcome regression usually receives the most attention, while the representer is treated as a technical ingredient needed for debiasing. For treatment effect problems, however, the representer is the estimand-specific weighting rule. Once it is written as a density ratio, questions about balancing, overlap, and regularization become questions about how to estimate that weighting rule directly. This shift in viewpoint is the main reason the link to DRE is useful beyond a purely formal equivalence.
Let $X\in {\mathcal{X}}$ be a regressor and $Y\in{\mathcal{Y}} \subseteq {\mathbb{R}}$ be a scalar outcome, where ${\mathcal{X}}$ is the regressor space and ${\mathcal{Y}}$ is the outcome space. Let us assume that $W = (X, Y) \in {\mathcal{W}} \coloneqq {\mathcal{X}} \times {\mathcal{Y}}$ jointly follows a distribution $P$. Let $\{(X_i, Y_i)\}_{i=1}^n$ be observations, where the sample size is $n\in{\mathbb{N}}$ and each $(X_i, Y_i)$ is an i.i.d. copy from $P$.
\paragraph{Parameter of interest.} We denote the parameter of interest by $\theta_0$. Let $\gamma_0 \coloneq {\mathbb{E}}\left[Y\mid X\right]$ be the regression function, and let $\Gamma$ be the set of measurable functions $\gamma \colon {\mathcal{X}}\to {\mathbb{R}}$ such that ${\mathbb{E}}\left[\gamma(X)^2\right] < \infty$. We assume that the parameter of interest $\theta_0$ can be written as \[\theta_0 \coloneqq {\mathbb{E}}\left[m(W, \gamma_0)\right],\] where $m\colon {\mathcal{W}} \times \Gamma \to {\mathbb{R}}$ is such that $m(w,\cdot)$ is linear for each $w\in{\mathcal{W}}$.
\paragraph{Goal.} Our goal is to estimate $\theta_0$ at the $\sqrt{n}$ rate and with asymptotic efficiency, that is, asymptotic normality with asymptotic variance matching the efficiency bound Vaart1998asymptoticstatistics.
In efficient estimation of $\theta_0$, Neyman orthogonality plays an important role, and it can be expressed using the Riesz representer. In this section, we recap the Neyman orthogonal score and the Riesz representer and show that the Riesz representer can be written as a signed density ratio.
Neyman orthogonal scores matter because asymptotically linear estimators based on them are asymptotically efficient, and orthogonality reduces plug-in bias when nuisance functions are replaced by estimators.
When the parameter of interest $\theta_0$ is linear in the regression function, let \[ L(\gamma) \coloneqq {\mathbb{E}}\left[m(W,\gamma)\right]. \] The Riesz representer $\alpha_0$ is the function satisfying \[ L(\gamma)={\mathbb{E}}\left[\alpha_0(X)\gamma(X)\right] \] for all $\gamma\in\Gamma$. The Neyman orthogonal score is then given by Newey1994theasymptotic,Chernozhukov2021automaticdebiased \[ \psi(W; \gamma_0, \alpha_0, \theta_0) \coloneqq \alpha_0(X)\big(Y-\gamma_0(X)\big) + m(W,\gamma_0) - \theta_0. \]
In practice, both $\gamma_0$ and $\alpha_0$ are unknown and must be estimated. The role of $\alpha_0$ is easy to overlook because it is introduced through semiparametric geometry, but in treatment effect problems it is the object that determines how residuals are reweighted inside the orthogonal score. This paper exploits that interpretation.
Let $P_X$ be the law of the regressor $X$ under $P$. If a linear functional can be written as \[ L(\gamma)=\int \gamma(x)d\nu(x), \] for a finite signed measure $\nu$ such that $\nu \ll P_X$ and $d\nu/dP_X \in L_2(P_X)$, then \[ \alpha_0(x)=\frac{{\mathrm{d}} \nu}{{\mathrm{d}} P_X}(x) \] is the Riesz representer, because $L(\gamma)={\mathbb{E}}\left[\alpha_0(X)\gamma(X)\right]$. Hence, the representer is a signed density ratio whenever the target is a signed change of measure.
This statement makes the Riesz representer concrete. It is not an abstract Hilbert space object but the derivative that tilts the law of the regression argument toward the target functional. The sign and magnitude of $\alpha_0$ describe how observations are reweighted to recover the parameter of interest. In treatment effect problems, overlap is precisely the condition that keeps this ratio stable.
A practically important class is discrete treatment with linear mean functionals. Let $D\in\left\{0,1,\dots,K\right\}$, let $e_{0,d}(z)\coloneqq \Pr(D=d\mid Z=z)$, and consider \[ \theta_0 \coloneqq {\mathbb{E}}\left[\sum^K_{d=0} h_d(Z)\gamma_0(d,Z)\right], \] where $\left\{h_d\right\}_{d=0}^K$ are known weighting functions. Then the representer is \[ \alpha_0(D,Z)\coloneqq \sum^K_{d=0} h_d(Z)\frac{\mathbbm{1}\left[D=d\right]}{e_{0,d}(Z)}, \] since \[ {\mathbb{E}}\left[\alpha_0(D,Z)\gamma(D,Z)\right]={\mathbb{E}}\left[\sum^K_{d=0} h_d(Z)\gamma(d,Z)\right]. \] This class includes several targets used in practice. ATE corresponds to $K=1$ with $h_1(Z)=1$ and $h_0(Z)=-1$. Weighted ATE uses $h_1(Z)=w(Z)$ and $h_0(Z)=-w(Z)$. Policy value uses $h_d(Z)=\mathbbm{1}\left[\pi(Z)=d\right]$. For more details on DRE for ATE estimation, see Section (ref).
ATT also belongs to the same class because \[ \theta_0={\mathbb{E}}\left[\frac{e_0(Z)}{p_1}\gamma_0(1,Z)-\frac{e_0(Z)}{p_1}\gamma_0(0,Z)\right], \] where $p_1\coloneqq \Pr(D=1)$, so that \[ \alpha_0(D,Z)=\frac{D}{p_1}-\frac{e_0(Z)}{p_1\left(1-e_0(Z)\right)}(1-D). \]
The same logic covers other targets that average $\gamma_0(a,Z)$ with respect to a target covariate distribution rather than the observed one. In such cases, the representer is again the appropriate change-of-measure ratio multiplied by the treatment indicator. This is exactly the type of problem for which direct DRE has been designed.
The remainder of the paper explains the equivalence between Riesz regression and LSIF, focusing on the ATE case, where the equivalence is clearest.
\paragraph{Potential outcomes and observations} We consider a binary treatment, where $1$ denotes treatment and $0$ denotes control. Following the Neyman-Rubin causal framework Neyman1923surapplications,Rubin1974estimatingcausal, we define potential outcomes $Y(1), Y(0) \in {\mathcal{Y}}$, where ${\mathcal{Y}} \subseteq {\mathbb{R}}$ denotes the outcome space. Let $Z \in {\mathcal{Z}}$ be unit-level covariates, where ${\mathcal{Z}}$ denotes the covariate space. Let $D\in\left\{0,1\right\}$ be a treatment indicator, and let $Y$ be the observed outcome defined as \[ Y = DY(1) + (1-D)Y(0). \] For the regressor $X = (D, Z)$, define \[ \gamma_0(X) = \gamma_0(D, Z) \coloneqq {\mathbb{E}}\left[Y\mid D, Z\right]. \] We observe an i.i.d. sample \[\left\{(D_i, Z_i, Y_i)\right\}_{i=1}^n.\]
\paragraph{ATE} Using the observations, we aim to estimate the ATE. The ATE is defined as \[ \theta^{\text{ATE}}_0 \coloneqq {\mathbb{E}}\left[Y(1)-Y(0)\right]. \] Under unconfoundedness, it can also be written as $\theta^{\text{ATE}}_0 = {\mathbb{E}}\left[\gamma_0(1, Z) - \gamma_0(0, Z)\right]$.
\paragraph{Notation and assumptions} Let $e_0(z)\coloneqq \Pr(D=1\mid Z=z)$ be the propensity score. We impose unconfoundedness and overlap, that is, $Y(1),Y(0)\perp D\mid Z$ and there exists $\epsilon\in(0,1/2)$ such that $\epsilon<e_0(Z)<1-\epsilon$ almost surely. We also write $p_{D,Z}(d,z)$ for the joint density of $(D,Z)$ and $p_Z(z)$ for the marginal density of $Z$ when they exist.
In ATE estimation, the moment function, the Riesz representer, and the Neyman orthogonal score are given by
Thus, the Neyman orthogonal score is
Estimators using such scores are also called doubly robust estimators Bang2005doublyrobust,Olea2024doublerobustness.
Riesz regression Chernozhukov2021automaticdebiased estimates the unknown Riesz representer by minimizing a mean squared error criterion.
\paragraph{General formulation} Let ${\mathcal{A}}$ be a model of $\alpha_0$, such as a linear model, an RKHS class, or a neural network. For $\alpha \in {\mathcal{A}}$, define \[ Q(\alpha)\coloneqq {\mathbb{E}}\left[\big(\alpha_0(D,Z) - \alpha(D,Z)\big)^2\right]. \] Riesz regression estimates $\alpha_0$ by minimizing an empirical version of this population risk. Although $Q(\alpha)$ contains the unknown $\alpha_0$, an equivalent feasible objective is available.
The appeal of this formulation is that it targets the Riesz representer itself, rather than estimating the propensity score first and then transforming it into weights. This is useful when one wants to regularize the weights directly or when the representer has a structure that is easier to exploit than a specific propensity score model. The feasible objective shows that direct representer fitting is possible even though $\alpha_0$ is unobserved.
\paragraph{ATE estimation} In the ATE case,
The identity follows from \[ {\mathbb{E}}\left[\alpha^{\text{ATE}}_0(D,Z)\alpha(D,Z)\right] = {\mathbb{E}}\left[\alpha(1,Z)\right]-{\mathbb{E}}\left[\alpha(0,Z)\right]. \] Hence, the empirical estimator is \[ \widehat{\alpha}\in\operatorname*{arg\,min}_{\alpha\in{\mathcal{A}}} \left\{\frac{1}{n}\sum^n_{i=1}\Big(-2\big(\alpha(1,Z_i) - \alpha(0,Z_i)\big) + \alpha(D_i,Z_i)^2\Big) +\lambda \Omega(\alpha)\right\}, \] where $\Omega$ is a regularizer, such as the $\ell_2$ norm or the RKHS norm.
Direct DRE targets the ratio itself rather than estimating two densities separately Sugiyama2011densityratio. Let $X^{(\text{de})}$ follow a distribution with density $p_{\text{de}}$, and let $X^{(\text{nu})}$ follow a distribution with density $p_{\text{nu}}$. Let $\left\{X_j^{(\text{de})}\right\}_{j=1}^{n_{\text{de}}}$ and $\left\{X_k^{(\text{nu})}\right\}_{k=1}^{n_{\text{nu}}}$ be two independent samples, where $X_j^{(\text{de})}$ is an i.i.d. copy of $X^{(\text{de})}\sim p_{\text{de}}$, and $X_k^{(\text{nu})}$ is an i.i.d. copy of $X^{(\text{nu})}\sim p_{\text{nu}}$. Our goal is to estimate \[ r_0(x) \coloneqq \frac{p_{\text{nu}}(x)}{p_{\text{de}}(x)}. \] A straightforward approach estimates $p_{\text{nu}}$ and $p_{\text{de}}$ separately and then takes their ratio. Direct methods target the ratio itself, which avoids error amplification and allows the loss to be matched to the application. This literature includes moment matching, classification-based methods, divergence minimization, and least-squares approaches Huang2007correctingsample,Qin1998inferencesfor,Nguyen2010estimatingdivergence,Sugiyama2011densityratio.
LSIF estimates the density ratio by minimizing
where ${\mathcal{R}}$ is a model of $r_0$. The feasible objective follows from the identity \[ {\mathbb{E}}_{p_{\text{de}}}\left[r_0(X)r(X)\right] = {\mathbb{E}}_{p_{\text{nu}}}\left[r(X)\right]. \] The empirical estimator is \[ \widehat{r}\in\operatorname*{arg\,min}_{r\in{\mathcal{R}}} \left\{-\frac{2}{n_{\text{nu}}}\sum^{n_{\text{nu}}}_{k=1}r\left(X_k^{(\text{nu})}\right) + \frac{1}{n_{\text{de}}}\sum^{n_{\text{de}}}_{j=1}r\left(X_j^{(\text{de})}\right)^2 +\lambda \Omega(r)\right\}. \]
The equivalence between Riesz regression and LSIF is immediate once the ATE representer is written as two treatment-specific density ratios. Since \[ e_0(z) = \frac{p_{D,Z}(1,z)}{p_Z(z)}, \] define
for $d \in \{1, 0\}$. Then \[ \alpha^{\text{ATE}}_0(D,Z) = Dr_0(1, Z) - (1-D)r_0(0, Z). \] Equivalently, $r_0(1,z)=1/e_0(z)$ and $r_0(0,z)=1/\left(1-e_0(z)\right)$, so the decomposition simply separates the treated and control components of the representer.
For $r_0(1,z)$, the corresponding LSIF criterion can be written as
Similarly,
Combining the two objectives gives \[ \operatorname*{arg\,min}_{r(1, \cdot), r(0, \cdot) \in {\mathcal{R}}} \left\{{\mathbb{E}}_{p_{D,Z}}\left[-2 \big(r(1,Z) + r(0,Z)\big) + Dr(1,Z)^2 + (1-D)r(0,Z)^2\right]\right\}. \] With the substitution \[ \alpha(D,Z)=Dr(1, Z)-(1-D)r(0, Z), \] we have $\alpha(1,Z)=r(1,Z)$, $\alpha(0,Z)=-r(0,Z)$, and $\alpha(D,Z)^2=Dr(1,Z)^2+(1-D)r(0,Z)^2$. Therefore, the LSIF objective becomes \[ {\mathbb{E}}\left[-2\big(\alpha(1,Z)-\alpha(0,Z)\big) + \alpha(D,Z)^2\right], \] which is exactly the ATE Riesz regression objective. This equivalence is exact at the level of the population risk. The empirical Riesz regression criterion is the corresponding plug-in estimator of this population objective under the observed data sampling scheme. Note that Riesz regression can be understood as a joint density ratio estimator that optionally shares basis functions or network parameters across treatment arms. Such shared structure is often attractive in practice because both ratios are learned on the same covariate space and usually benefit from common features.
The equivalence is useful because the representer is the object that generates the orthogonal score weights Kato2026unifiedframework. Different density ratio losses are therefore different ways of estimating the same weighting function. Quadratic loss gives LSIF and Riesz regression, the corresponding dual problem yields stable balancing weights Zubizarreta2015stableweights. More generally, DRE often replaces squared error with Bregman divergence criteria Sugiyama2011densityratio. In the present setting, this yields Bregman-type Riesz regression Kato2025directbias. Kullback-Leibler (KL)-type losses connect the density ratio literature to the tailored loss of Zhao2019covariatebalancing, and the corresponding dual problem yields entropy balancing weights Hainmueller2012entropybalancing. Because the ATE representer is signed, KL-type implementations use the signed modifications developed in Kato2025directbias.
The identity also transfers model choice and regularization results. In RKHS models, KuLSIF provides analytic solutions and efficient cross-validation Kanamori2012statisticalanalysis. In neural network models deep density ratio methods have been proposed with finite-sample error bounds and nonparametric rates under smoothness assumptions Kato2021nonnegativebregman,Zheng2022anerror. When the density ratio is difficult to estimate, the implied representer is unstable, so extreme orthogonal score weights should be expected. This suggests a concrete empirical workflow. One chooses the model class and regularization by thinking directly about the complexity of the weights required by the target parameter, rather than about an auxiliary propensity score model. The DRE literature studies this problem directly. Nonnegative Bregman corrections curb training loss hacking in flexible models, and telescoping DRE replaces one difficult ratio with a product of easier local ratios Kato2021nonnegativebregman,Rhodes2020telescopingdensityratio.
The identity also generalizes matching methods. Lin2023estimationbased shows that nearest neighbor matching implicitly estimates a density ratio, and Kato2025nearestneighbor shows that it can be written as LSIF, hence as Riesz regression, with a matching basis. Under the present view, matching, balancing, and augmentation are not separate ideas. For more detailed arguments about this perspective, see Kato2026unifiedframework.
This paper shows that the key link between Riesz regression and direct DRE is the representation of the Riesz representer as a signed density ratio. This perspective gives a direct weighting interpretation of representer fitting and clarifies the connections among orthogonal scores, balancing weights, matching, augmentation, model selection, and regularization. It turns an abstract Riesz representer into a familiar reweighting object and shows when methods from DRE can be used without changing the target parameter.