EconBase
← All papers

Dynamic Discrete Choice and Inverse Reinforcement Learning: Inferring Preferences and Beliefs From Human Behavior

Pranjal Rawat, John Rust

arXiv 25 Aug 2026 · Econometrics

arXiv:2608.24362 · PDF · Extracted main text

Abstract

This article surveys two deeply connected literatures that approach the same fundamental problem from different disciplinary traditions: dynamic discrete choice (DDC) in structural econometrics and inverse reinforcement learning (IRL) in machine learning. Both seek to infer the preferences of decision makers from observed sequential behavior, assuming that individuals act to maximize an expected reward function within a dynamic, uncertain environment formalized as a Markov decision process (MDP). Despite independent origins, the two fields have converged on similar mathematical formulations. We show that the (soft Q-learning) framework now prevalent in IRL is closely related to DDC models under additive extreme value preference shocks, yielding the same softmax (multinomial logit) choice probabilities and smooth Bellman equations that underpin structural estimation in economics. We compare the estimation and computational methods developed in each field. DDC has emphasized maximum likelihood estimation, conditional choice probability estimators, and policy iteration methods. IRL has developed scalable alternatives, including maximum entropy methods, adversarial approaches, and model-free temporal difference estimators that extend to high-dimensional state spaces using deep neural networks. Model-free IRL estimators that combine temporal difference learning with classical two-step methods from econometrics represent a promising direction for bridging the two literatures. Both fields confront shared foundational challenges: the identification problem, whereby multiple reward functions can rationalize the same observed behavior, and the curse of dimensionality in solving the underlying MDP. We believe that cross-fertilization offers substantial opportunities for methodological progress in both fields.

Citation extraction

117
references
188
in-text mentions
118
distinct cited
5
self-citations
17,777
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Sutton and Barto (2020)0.87482100%
2Lee, Sudhir and Wang (2026) Consumer Engagement with Sequential Content: A Content-Aware Dynamic Choice Model0.87462100%
3Ng and Russell (2000) Algorithms for Inverse Reinforcement Learning0.87462100%
4Hotz and Miller (1993) Conditional choice probabilities and the estimation of dynamic models0.81142100%
5Nguyen (2025) Neural Networks for Efficient Estimation of High-Dimensional Dynamic Discrete Choice Models0.73732100%
6Kang, Yoganarasimhan and Jain (2025) An Empirical Risk Minimization Approach for Offline Inverse RL and Dynamic Discrete Choice Model0.73732100%
7Ng, Harada and Russell (1999) Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping0.73732100%
8Rust (1987) Optimal Replacement of GMC Bus Engines: An Empirical Model of Harold Zurcher self0.73732100%
9Rust (1994) Structural Estimation of Markov Decision Processes self0.73732100%
10Ziebart, Maas, Bagnell and Dey (2008) Maximum Entropy Inverse Reinforcement Learning0.73732100%

Showing the top 10 of 118 scored citations.