Ahmed Khwaja, Sonal Srivastava
arXiv 5 Jan 2026 · Econometrics
arXiv:2601.02069 · PDF · DOI · OpenAlex · Extracted main text
Dynamic discrete choice (DDC) models have found widespread application in marketing. However, estimating these becomes challenging in "big data" settings with high-dimensional state-action spaces. To address this challenge, this paper develops a Reinforcement Learning (RL)-based two-step ("computationally light") Conditional Choice Simulation (CCS) estimation approach that combines the scalability of machine learning with the transparency, explainability, and interpretability of structural models, which is particularly valuable for counterfactual policy analysis. The method is premised on three insights: (1) the CCS ("forward simulation") approach is a special case of RL algorithms, (2) starting from an initial state-action pair, CCS updates the corresponding value function only after each simulation path has terminated, whereas RL algorithms may update for all the state-action pairs visited along a simulated path, and (3) RL focuses on inferring an agent's optimal policy with known reward functions, whereas DDC models focus on estimating the reward functions presupposing optimal policies. The procedure's computational efficiency over CCS estimation is demonstrated using Monte Carlo simulations with a canonical machine replacement and a consumer food purchase model. Framing CCS estimation of DDC models as an RL problem increases their applicability and scalability to high-dimensional marketing problems while retaining both interpretability and tractability.
appendix boundary found by appendix_command · 71% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Rust, John (1987) Optimal replacement of GMC bus engines: An empirical model of Harold Zurcher | 0.937 | 17 | 5 | 82% |
| 2 | Hotz, V Joseph and Miller, Robert A and Sanders, Seth and Smith, Jef… (1994) A simulation estimator for dynamic models of discrete choice | 0.909 | 8 | 5 | 75% |
| 3 | Bajari, Patrick and Benkard, C Lanier and Levin, Jonathan (2007) Estimating dynamic models of imperfect competition | 0.874 | 6 | 5 | 67% |
| 4 | Hotz, V Joseph and Miller, Robert A (1993) Conditional choice probabilities and the estimation of dynamic models | 0.843 | 5 | 3 | 60% |
| 5 | Huang, Guofang and Khwaja, Ahmed and Sudhir, K (2015) Short-run needs and long-term goals: A dynamic model of thirst management self | 0.811 | 4 | 2 | 100% |
| 6 | Sutton, Richard S (1988) Learning to predict by the methods of temporal differences | 0.737 | 3 | 3 | 67% |
| 7 | Sutton, Richard S and Barto, Andrew G (2018) Reinforcement learning: An introduction | 0.737 | 3 | 2 | 100% |
| 8 | Rust, John (1994) Structural estimation of Markov decision processes | 0.644 | 5 | 2 | 40% |
| 9 | Arcidiacono, Peter and Miller, Robert A (2011) Conditional choice probability estimation of dynamic discrete choice models with unobserved heterogeneity | 0.644 | 3 | 2 | 67% |
| 10 | Robbins, Herbert and Monro, Sutton (1951) A stochastic approximation method | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 117 scored citations.