EconBase
← All papers

Reinforcement Learning Based Computationally Efficient Conditional Choice Simulation Estimation of Dynamic Discrete Choice Models

Ahmed Khwaja, Sonal Srivastava

arXiv 5 Jan 2026 · Econometrics

arXiv:2601.02069 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Dynamic discrete choice (DDC) models have found widespread application in marketing. However, estimating these becomes challenging in "big data" settings with high-dimensional state-action spaces. To address this challenge, this paper develops a Reinforcement Learning (RL)-based two-step ("computationally light") Conditional Choice Simulation (CCS) estimation approach that combines the scalability of machine learning with the transparency, explainability, and interpretability of structural models, which is particularly valuable for counterfactual policy analysis. The method is premised on three insights: (1) the CCS ("forward simulation") approach is a special case of RL algorithms, (2) starting from an initial state-action pair, CCS updates the corresponding value function only after each simulation path has terminated, whereas RL algorithms may update for all the state-action pairs visited along a simulated path, and (3) RL focuses on inferring an agent's optimal policy with known reward functions, whereas DDC models focus on estimating the reward functions presupposing optimal policies. The procedure's computational efficiency over CCS estimation is demonstrated using Monte Carlo simulations with a canonical machine replacement and a consumer food purchase model. Framing CCS estimation of DDC models as an RL problem increases their applicability and scalability to high-dimensional marketing problems while retaining both interpretability and tractability.

Citation extraction

117
references
175
in-text mentions
117
distinct cited
1
self-citations
14,670
main-text words

appendix boundary found by appendix_command · 71% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Rust, John (1987) Optimal replacement of GMC bus engines: An empirical model of Harold Zurcher0.93717582%
2Hotz, V Joseph and Miller, Robert A and Sanders, Seth and Smith, Jef… (1994) A simulation estimator for dynamic models of discrete choice0.9098575%
3Bajari, Patrick and Benkard, C Lanier and Levin, Jonathan (2007) Estimating dynamic models of imperfect competition0.8746567%
4Hotz, V Joseph and Miller, Robert A (1993) Conditional choice probabilities and the estimation of dynamic models0.8435360%
5Huang, Guofang and Khwaja, Ahmed and Sudhir, K (2015) Short-run needs and long-term goals: A dynamic model of thirst management self0.81142100%
6Sutton, Richard S (1988) Learning to predict by the methods of temporal differences0.7373367%
7Sutton, Richard S and Barto, Andrew G (2018) Reinforcement learning: An introduction0.73732100%
8Rust, John (1994) Structural estimation of Markov decision processes0.6445240%
9Arcidiacono, Peter and Miller, Robert A (2011) Conditional choice probability estimation of dynamic discrete choice models with unobserved heterogeneity0.6443267%
10Robbins, Herbert and Monro, Sutton (1951) A stochastic approximation method0.64422100%

Showing the top 10 of 117 scored citations.