EconBase
← All papers

Deep Learning for Dynamic Programming with Recursive Utility Using First-order Conditions

Xianhua Peng, Wu Guo, Songyan Wang, Jianfei Zhu

arXiv 10 Jul 2026 · Finance — Computational

arXiv:2607.09461 · PDF · DOI · OpenAlex · Extracted main text

Abstract

This paper proposes the certainty-equivalent first-order learning (CEFOL) algorithm, a deep learning algorithm for solving discrete-time dynamic programming problems with recursive utility. Dynamic programming with recursive utility is challenging because nonlinear certainty equivalent appears in the Bellman equation and the first-order optimality conditions but is difficult to evaluate. By introducing a separate neural network to represent the certainty equivalent, CEFOL enables the exploitation of the Bellman and model-specific first-order optimality conditions. In addition to certainty equivalent, CEFOL also uses neural networks to learn the value functions, policy functions, and Lagrange multipliers by using model-specific first-order conditions to construct residuals for minimization. By using first-order and KKT residuals to learn the policy, CEFOL directly accommodates general equality and inequality constraints on the controls, including occasionally binding constraints, without requiring penalty functions or problem-specific reformulations. We apply the algorithm to risk-sensitive and Epstein--Zin consumption-saving problems, a small-noise robust-control problem, and a DSGE model with recursive preferences and stochastic volatility. Across these applications, out-of-sample Bellman diagnostics and model-specific optimality residuals, including Euler or first-order residuals where applicable, are generally of order 1.0e-4 to 1.0e-3 over the relevant state regions, with larger values mainly near binding constraints, and the learned value and policy functions closely match VFI benchmarks when available. The CEFOL algorithm also works for dynamic programming problems with expected utility, as expected utility is a special case of recursive utility.

Citation extraction

69
references
98
in-text mentions
70
distinct cited
0
self-citations
24,994
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Maliar, Maliar \ Winant (2021) Deep learning for solving dynamic economic models., Journal of Monetary Economics 122: 76–1011.00063100%
2Epstein \ Zin (1989) Substitution, risk aversion, and the temporal behavior of consumption and asset returns: a theoretical framework, Econometrica 5…0.92843100%
3Hansen \ Sargent (2013) Recursive Models of Dynamic Linear Economies, Princeton University Press0.73732100%
4Caldara, Fernandez-Villaverde, Rubio-Ramirez \ Yao (2012) Computing dsge models with recursive preferences and stochastic volatility, Review of Economic Dynamics 15(2): 188–2060.73732100%
5Friedl, Kübler, Scheidegger \ Usui (2023) Deep uncertainty quantification: with an application to integrated assessment models, Technical report, Working Paper University…0.69361100%
6Backus, Routledge \ Zin (2016) Recursive Preferences, Palgrave Macmillan UK, London0.64422100%
7Weil (1990) Nonexpected utility in macroeconomics, The Quarterly Journal of Economics 105(1): 29–420.64422100%
8Andreasen (2012) On the effects of rare disasters and uncertainty shocks for risk premia in non-linear dsge models, Review of Economic Dynamics 1…0.64422100%
9Hansen \ Sargent (1995) Discounted linear exponential quadratic gaussian control, IEEE Transactions on Automatic control 40(5): 968–9710.64422100%
10Hansen \ Sargent (2008) Robustness, Princeton university press0.64422100%

Showing the top 10 of 70 scored citations.