arXiv 29 May 2026 · Machine Learning
arXiv:2605.30843 · PDF · DOI · OpenAlex · Extracted main text
In the forward reinforcement-learning problem, the reward is fixed and known; the learner is asked to find a good policy or value function. Here we turn the question around. Given offline data generated by an expert, can we recover the reward the expert was optimizing? This is the inverse reinforcement learning problem, and remarkably, two communities, structural econometricians studying dynamic discrete choice (DDC) and machine learners studying entropy-regularized IRL, have been working on exactly the same probabilistic model under different names. We begin by proving their equivalence. We then develop the classical identification result of Magnac and Thesmar and the classical computational paradigms that grew out of it: Rust's nested fixed-point algorithm, the conditional-choice-probability approach of Hotz and Miller, and the two temporal-difference approaches of Adusumilli and Eckardt: linear semi-gradient TD and approximate value iteration. Each route has its limits: dimensionality, transition-kernel estimation, the deadly triad, or projected fixed-point bias. We then walk through the modern ML/IRL strand: adversarial IRL, occupancy matching, IQ-Learn, and offline ML-IRL, deriving each method's actual objective and stating precisely what it does and does not identify. We close with the empirical-risk-minimization framework of Kang et al., which yields a gradient-based estimator for offline IRL/DDC.
appendix boundary found by appendix_command · 93% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Fu, Justin and Luo, Katie and Levine, Sergey (2017) Learning robust rewards with adversarial inverse reinforcement learning | 1.000 | 7 | 3 | 100% |
| 2 | Kang, Enoch H. and Yoganarasimhan, Hema and Jain, Lalit (2025) An Empirical Risk Minimization Approach for Offline Inverse RL and Dynamic Discrete Choice Model self | 1.000 | 6 | 5 | 100% |
| 3 | Rust, John (1987) Optimal replacement of GMC bus engines: An empirical model of Harold Zurcher | 0.928 | 4 | 3 | 100% |
| 4 | Rust, John (1994) Structural estimation of Markov decision processes | 0.928 | 4 | 3 | 100% |
| 5 | Magnac, Thierry and Thesmar, David (2002) Identifying dynamic discrete decision processes | 0.874 | 12 | 2 | 100% |
| 6 | Ho, Jonathan and Ermon, Stefano (2016) Generative adversarial imitation learning | 0.843 | 3 | 3 | 100% |
| 7 | Adusumilli, Karun and Eckardt, Dita (2019) Temporal-Difference estimation of dynamic discrete choice models | 0.737 | 3 | 2 | 100% |
| 8 | Zeng, Siliang and Li, Chenliang and Garcia, Alfredo and Hong, Mingyi (2023) Understanding expertise through demonstrations: A maximum likelihood framework for offline inverse reinforcement learning | 0.737 | 3 | 2 | 100% |
| 9 | Cao, Haoyang and Cohen, Samuel and Szpruch, Lukasz (2021) Identifiability in inverse reinforcement learning | 0.693 | 7 | 1 | 100% |
| 10 | Arcidiacono, Peter and Ellickson, Paul B (2011) Practical methods for estimation of dynamic discrete choice models | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 41 scored citations.