Lars van der Laan, Nathan Kallus, Aurélien Bibaut
arXiv 25 Sep 2025 · Machine Learning
arXiv:2509.21172 · PDF · DOI · OpenAlex · Extracted main text
Inverse reinforcement learning (IRL) aims to explain observed behavior by uncovering an underlying reward. In the maximum-entropy or Gumbel-shocks-to-reward frameworks, this amounts to fitting a reward function and a soft value function that together satisfy the soft Bellman consistency condition and maximize the likelihood of observed actions. While this perspective has had enormous impact in imitation learning for robotics and understanding dynamic choices in economics, practical learning algorithms often involve delicate inner-loop optimization, repeated dynamic programming, or adversarial training, all of which complicate the use of modern, highly expressive function approximators like neural nets and boosting. We revisit softmax IRL and show that the population maximum-likelihood solution is characterized by a linear fixed-point equation involving the behavior policy. This observation reduces IRL to two off-the-shelf supervised learning problems: probabilistic classification to estimate the behavior policy, and iterative regression to solve the fixed point. The resulting method is simple and modular across function approximation classes and algorithms. We provide a precise characterization of the optimal solution, a generic oracle-based algorithm, finite-sample error bounds, and empirical results showing competitive or superior performance to MaxEnt IRL.
appendix boundary found by appendix_command · 47% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Rémi Munos and Csaba Szepesvári (2008) Finite-time bounds for fitted value iteration | 1.000 | 7 | 3 | 100% |
| 2 | Justin Fu, Katie Luo, and Sergey Levine (2018) Learning robust rewards with adversarial inverse reinforcement learning | 1.000 | 6 | 3 | 100% |
| 3 | Lars van der Laan and Nathan Kallus (2025) Fitted q evaluation without bellman completeness via stationary weighting self | 1.000 | 6 | 3 | 100% |
| 4 | Brian D. Ziebart, Andrew Maas, J. Andrew Bagnell, and Anind K. Dey (2008) Maximum entropy inverse reinforcement learning | 1.000 | 6 | 3 | 100% |
| 5 | V. Joseph Hotz and Robert A. Miller (1993) Conditional choice probabilities and the estimation of dynamic models | 1.000 | 5 | 3 | 100% |
| 6 | Sinong Geng, Houssam Nassif, Carlos Manzanares, Max Reppen, and Ronn… (2020) Deep pqr: Solving inverse reinforcement learning using anchor actions | 0.977 | 15 | 6 | 93% |
| 7 | Haoyang Cao, Samuel Cohen, and Lukasz Szpruch (2021) Identifiability in inverse reinforcement learning | 0.928 | 4 | 3 | 100% |
| 8 | John Rust (1987) Optimal replacement of gmc bus engines: An empirical model of harold zurcher | 0.874 | 7 | 2 | 100% |
| 9 | Masatoshi Uehara, Jiawei Huang, and Nan Jiang (2020) Minimax weight and q-function learning for off-policy evaluation | 0.843 | 3 | 3 | 100% |
| 10 | Brian D Ziebart, J Andrew Bagnell, and Anind K Dey (2010) Modeling interaction via the principle of maximum causal entropy | 0.811 | 4 | 2 | 100% |
Showing the top 10 of 56 scored citations.