Masahiro Kato, Taka Kato
arXiv 20 Jul 2026 · Econometrics
arXiv:2607.18225 · PDF · Extracted main text
We propose one-step and two-step methods for policy learning with retrieval-augmented generation (RAG). We formulate RAG-based action selection under the potential outcome framework. In the two-step method, vector search retrieves action-specific neighboring evidence in an embedding space, the generator estimates conditional expected outcomes or their contrasts, and a plug-in rule selects an action. This formulation connects action-specific vector search with nearest-neighbor matching in causal inference. We decompose the regret of the two-step method into candidate-generation regret and within-candidate choice regret, and we bound the latter using prediction-error guarantees for nearest-neighbor estimators and transformers. We evaluate the one-step method directly as a policy because its intermediate computation is unobserved.
appendix boundary found by appendix_command · 60% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Toru Kitagawa and Aleksey Tetenov (2018) Who should be treated? empirical welfare maximization methods for treatment choice | 0.874 | 6 | 2 | 100% |
| 2 | Masahiro Kato (2025) Nearest neighbor matching as least squares density ratio estimation and riesz regression, 2025b self | 0.843 | 4 | 3 | 75% |
| 3 | Zhexiao Lin, Peng Ding, and Fang Han (2023) Estimation based on nearest neighbor matching: from density ratio to average treatment effect | 0.843 | 4 | 3 | 75% |
| 4 | Susan Athey and Stefan Wager (2021) Policy learning with observational data | 0.811 | 4 | 2 | 100% |
| 5 | Michelle Ching, Ioana Popescu, Nico Smith, Tianyi Ma, William G. Und… (2026) Efficient and minimax optimal in-context nonparametric regression with transformers | 0.737 | 3 | 3 | 67% |
| 6 | Juno Kim, Tai Nakamaki, and Taiji Suzuki (2024) Transformers are minimax optimal nonparametric in-context learners | 0.737 | 3 | 3 | 67% |
| 7 | Kazusato Oko, Yujin Song, Taiji Suzuki, and Denny Wu (2024) Pretrained transformer efficiently learns low-dimensional target functions in-context | 0.737 | 3 | 3 | 67% |
| 8 | Jean-Yves Audibert and Alexandre B. Tsybakov (2007) Fast learning rates for plug-in classifiers | 0.737 | 3 | 2 | 100% |
| 9 | Noah Dasanaike and Kosuke Imai (2026) Using embedding models to improve probabilistic race prediction, 2026 | 0.511 | 2 | 2 | 50% |
| 10 | Alexander Havrilla and Wenjing Liao (2024) Understanding scaling laws with statistical and approximation theory for transformer neural networks on intrinsically low-dimens… | 0.511 | 2 | 2 | 50% |
Showing the top 10 of 41 scored citations.