Pranjal Rawat, John Rust
arXiv 25 Aug 2026 · Econometrics
arXiv:2608.24362 · PDF · Extracted main text
This article surveys two deeply connected literatures that approach the same fundamental problem from different disciplinary traditions: dynamic discrete choice (DDC) in structural econometrics and inverse reinforcement learning (IRL) in machine learning. Both seek to infer the preferences of decision makers from observed sequential behavior, assuming that individuals act to maximize an expected reward function within a dynamic, uncertain environment formalized as a Markov decision process (MDP). Despite independent origins, the two fields have converged on similar mathematical formulations. We show that the (soft Q-learning) framework now prevalent in IRL is closely related to DDC models under additive extreme value preference shocks, yielding the same softmax (multinomial logit) choice probabilities and smooth Bellman equations that underpin structural estimation in economics. We compare the estimation and computational methods developed in each field. DDC has emphasized maximum likelihood estimation, conditional choice probability estimators, and policy iteration methods. IRL has developed scalable alternatives, including maximum entropy methods, adversarial approaches, and model-free temporal difference estimators that extend to high-dimensional state spaces using deep neural networks. Model-free IRL estimators that combine temporal difference learning with classical two-step methods from econometrics represent a promising direction for bridging the two literatures. Both fields confront shared foundational challenges: the identification problem, whereby multiple reward functions can rationalize the same observed behavior, and the curse of dimensionality in solving the underlying MDP. We believe that cross-fertilization offers substantial opportunities for methodological progress in both fields.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Sutton and Barto (2020) | 0.874 | 8 | 2 | 100% |
| 2 | Lee, Sudhir and Wang (2026) Consumer Engagement with Sequential Content: A Content-Aware Dynamic Choice Model | 0.874 | 6 | 2 | 100% |
| 3 | Ng and Russell (2000) Algorithms for Inverse Reinforcement Learning | 0.874 | 6 | 2 | 100% |
| 4 | Hotz and Miller (1993) Conditional choice probabilities and the estimation of dynamic models | 0.811 | 4 | 2 | 100% |
| 5 | Nguyen (2025) Neural Networks for Efficient Estimation of High-Dimensional Dynamic Discrete Choice Models | 0.737 | 3 | 2 | 100% |
| 6 | Kang, Yoganarasimhan and Jain (2025) An Empirical Risk Minimization Approach for Offline Inverse RL and Dynamic Discrete Choice Model | 0.737 | 3 | 2 | 100% |
| 7 | Ng, Harada and Russell (1999) Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping | 0.737 | 3 | 2 | 100% |
| 8 | Rust (1987) Optimal Replacement of GMC Bus Engines: An Empirical Model of Harold Zurcher self | 0.737 | 3 | 2 | 100% |
| 9 | Rust (1994) Structural Estimation of Markov Decision Processes self | 0.737 | 3 | 2 | 100% |
| 10 | Ziebart, Maas, Bagnell and Dey (2008) Maximum Entropy Inverse Reinforcement Learning | 0.737 | 3 | 2 | 100% |
Showing the top 10 of 118 scored citations.