EconBase
← All papers

An Empirical Risk Minimization Approach for Offline Inverse RL and Dynamic Discrete Choice Model

Enoch H. Kang, Hema Yoganarasimhan, Lalit Jain

arXiv 19 Feb 2025 · Machine Learning · 1 citations (OpenAlex)

arXiv:2502.14131 · PDF · DOI · OpenAlex · Extracted main text

Abstract

We study the problem of estimating Dynamic Discrete Choice (DDC) models, also known as offline Maximum Entropy-Regularized Inverse Reinforcement Learning (offline MaxEnt-IRL) in machine learning. The objective is to recover reward or $Q^*$ functions that govern agent behavior from offline behavior data. In this paper, we propose a globally convergent gradient-based method for solving these problems without the restrictive assumption of linearly parameterized rewards. The novelty of our approach lies in introducing the Empirical Risk Minimization (ERM) based IRL/DDC framework, which circumvents the need for explicit state transition probability estimation in the Bellman equation. Furthermore, our method is compatible with non-parametric estimation techniques such as neural networks. Therefore, the proposed method has the potential to be scaled to high-dimensional, infinite state spaces. A key theoretical insight underlying our approach is that the Bellman residual satisfies the Polyak-Lojasiewicz (PL) condition -- a property that, while weaker than strong convexity, is sufficient to ensure fast global convergence guarantees. Through a series of synthetic experiments, we demonstrate that our approach consistently outperforms benchmark methods and state-of-the-art alternatives.

Citation extraction

84
references
189
in-text mentions
84
distinct cited
2
self-citations
15,714
main-text words

appendix boundary found by appendix_command · 48% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Hotz, V. Joseph and Miller, Robert A (1993) Conditional choice probabilities and the estimation of dynamic models1.00095100%
2Adusumilli, Karun and Eckardt, Dita (2019) Temporal-Difference estimation of dynamic discrete choice models1.00083100%
3Geng, Sinong and Nassif, Houssam and Manzanares, Carlos A (2023) A data-driven state aggregation approach for dynamic discrete choice models1.00083100%
4Zeng, Siliang and Li, Chenliang and Garcia, Alfredo and Hong, Mingyi (2023) Understanding expertise through demonstrations: A maximum likelihood framework for offline inverse reinforcement learning1.00075100%
5Rust, John (1987) Optimal replacement of GMC bus engines: An empirical model of Harold Zurcher1.00064100%
6Garg, Divyansh and Chakraborty, Shuvam and Cundy, Chris and Song, Ji… (2021) IQ-Learn: Inverse soft-Q learning for imitation1.00063100%
7Jiang, Nan and Xie, Tengyang (2024) Offline Reinforcement Learning in Large State Spaces: Algorithms and Guarantees1.00063100%
8Rust, John (1994) Structural estimation of Markov decision processes0.9568588%
9Magnac, Thierry and Thesmar, David (2002) Identifying dynamic discrete decision processes0.9416383%
10Su, Che-Lin and Judd, Kenneth L (2012) Constrained optimization approaches to estimation of structural models0.92843100%

Showing the top 10 of 84 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1A Lecture Note on Offline RL and IRL Part II: Foundations of Inverse Reinforcement Learning and Dynamic Discrete Choice Models1.00065