Enoch H. Kang, Hema Yoganarasimhan, Lalit Jain
arXiv 19 Feb 2025 · Machine Learning · 1 citations (OpenAlex)
arXiv:2502.14131 · PDF · DOI · OpenAlex · Extracted main text
We study the problem of estimating Dynamic Discrete Choice (DDC) models, also known as offline Maximum Entropy-Regularized Inverse Reinforcement Learning (offline MaxEnt-IRL) in machine learning. The objective is to recover reward or $Q^*$ functions that govern agent behavior from offline behavior data. In this paper, we propose a globally convergent gradient-based method for solving these problems without the restrictive assumption of linearly parameterized rewards. The novelty of our approach lies in introducing the Empirical Risk Minimization (ERM) based IRL/DDC framework, which circumvents the need for explicit state transition probability estimation in the Bellman equation. Furthermore, our method is compatible with non-parametric estimation techniques such as neural networks. Therefore, the proposed method has the potential to be scaled to high-dimensional, infinite state spaces. A key theoretical insight underlying our approach is that the Bellman residual satisfies the Polyak-Lojasiewicz (PL) condition -- a property that, while weaker than strong convexity, is sufficient to ensure fast global convergence guarantees. Through a series of synthetic experiments, we demonstrate that our approach consistently outperforms benchmark methods and state-of-the-art alternatives.
appendix boundary found by appendix_command · 48% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Hotz, V. Joseph and Miller, Robert A (1993) Conditional choice probabilities and the estimation of dynamic models | 1.000 | 9 | 5 | 100% |
| 2 | Adusumilli, Karun and Eckardt, Dita (2019) Temporal-Difference estimation of dynamic discrete choice models | 1.000 | 8 | 3 | 100% |
| 3 | Geng, Sinong and Nassif, Houssam and Manzanares, Carlos A (2023) A data-driven state aggregation approach for dynamic discrete choice models | 1.000 | 8 | 3 | 100% |
| 4 | Zeng, Siliang and Li, Chenliang and Garcia, Alfredo and Hong, Mingyi (2023) Understanding expertise through demonstrations: A maximum likelihood framework for offline inverse reinforcement learning | 1.000 | 7 | 5 | 100% |
| 5 | Rust, John (1987) Optimal replacement of GMC bus engines: An empirical model of Harold Zurcher | 1.000 | 6 | 4 | 100% |
| 6 | Garg, Divyansh and Chakraborty, Shuvam and Cundy, Chris and Song, Ji… (2021) IQ-Learn: Inverse soft-Q learning for imitation | 1.000 | 6 | 3 | 100% |
| 7 | Jiang, Nan and Xie, Tengyang (2024) Offline Reinforcement Learning in Large State Spaces: Algorithms and Guarantees | 1.000 | 6 | 3 | 100% |
| 8 | Rust, John (1994) Structural estimation of Markov decision processes | 0.956 | 8 | 5 | 88% |
| 9 | Magnac, Thierry and Thesmar, David (2002) Identifying dynamic discrete decision processes | 0.941 | 6 | 3 | 83% |
| 10 | Su, Che-Lin and Judd, Kenneth L (2012) Constrained optimization approaches to estimation of structural models | 0.928 | 4 | 3 | 100% |
Showing the top 10 of 84 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | A Lecture Note on Offline RL and IRL Part II: Foundations of Inverse Reinforcement Learning and Dynamic Discrete Choice Models | 1.000 | 6 | 5 |