Masahiro Kato
arXiv 26 Sep 2025 · Econometrics
arXiv:2509.22122 · PDF · Extracted main text
This study considers the estimation of the average treatment effect (ATE). For ATE estimation, we estimate the propensity score through direct bias-correction term estimation. Let ${(X_i, D_i, Y_i)}_{i=1}^{n}$ be the observations, where $X_i \in \mathbb{R}^p$ denotes $p$-dimensional covariates, $D_i \in {0, 1}$ denotes a binary treatment assignment indicator, and $Y_i \in \mathbb{R}$ is an outcome. In ATE estimation, the bias-correction term $h_0(X_i, D_i) = \frac{1[D_i = 1]}{e_0(X_i)} - \frac{1[D_i = 0]}{1 - e_0(X_i)}$ plays an important role, where $e_0(X_i)$ is the propensity score, the probability of being assigned treatment $1$. In this study, we propose estimating $h_0$ (or equivalently the propensity score $e_0$) by directly minimizing the prediction error of $h_0$. Since the bias-correction term $h_0$ is essential for ATE estimation, this direct approach is expected to improve estimation accuracy for the ATE. For example, existing studies often employ maximum likelihood or covariate balancing to estimate $e_0$, but these approaches may not be optimal for accurately estimating $h_0$ or the ATE. We present a general framework for this direct bias-correction term estimation approach from the perspective of Bregman divergence minimization and conduct simulation studies to evaluate the effectiveness of the proposed method.
appendix boundary found by appendix_command · 45% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Qingyuan Zhao (2019) Covariate balancing propensity score by tailored loss functions | 1.000 | 7 | 3 | 100% |
| 2 | David Bruns-Smith, Oliver Dukes, Avi Feller, and Elizabeth L Ogburn (2025) Augmented balancing weights as linear regression | 1.000 | 5 | 3 | 100% |
| 3 | Victor Chernozhukov, Whitney K. Newey, Victor Quintas-Martinez, and… (2021) Automatic debiased machine learning via riesz regression, 2021 | 0.874 | 9 | 2 | 100% |
| 4 | Masahiro Kato (2025) Direct debiased machine learning via bregman divergence minimization, 2025a self | 0.843 | 4 | 4 | 75% |
| 5 | Jens Hainmueller (2012) Entropy balancing for causal effects: A multivariate reweighting method to produce balanced samples in observational studies | 0.843 | 3 | 3 | 100% |
| 6 | José R. Zubizarreta (2015) Stable weights that balance covariates for estimation with incomplete outcome data | 0.843 | 3 | 3 | 100% |
| 7 | Takafumi Kanamori, Taiji Suzuki, and Masashi Sugiyama (2012) Statistical analysis of kernel-based least-squares density-ratio estimation | 0.830 | 7 | 3 | 57% |
| 8 | Takafumi Kanamori, Shohei Hido, and Masashi Sugiyama (2009) A least-squares approach to direct importance estimation | 0.811 | 4 | 2 | 100% |
| 9 | Masahiro Kato and Takeshi Teshima (2021) Non-negative bregman divergence minimization for deep direct density ratio estimation self | 0.763 | 9 | 6 | 44% |
| 10 | Victor Chernozhukov, Whitney Newey, Vćtor M Quintas-Martńez, and Vas… (2022) RieszNet and ForestRiesz: Automatic debiased machine learning with neural nets and random forests | 0.737 | 3 | 3 | 67% |
Showing the top 10 of 48 scored citations.