Abstract
We consider a Markov decision process (MDP), whose total discounted utility is aggregated recursively with a concave discount function that is not necessarily linear. The state and action spaces are Borel spaces, and the utility function is nonnegative. We show that it can be reduced to a turn-based stochastic game model with the total undiscounted utility. This reduction result is then applied to the MDP problem with recursively aggregated utility to be maximized or cost to be minimized.
| Original language | English |
|---|---|
| Journal | Annals of Operations Research |
| Early online date | 11 Mar 2025 |
| DOIs | |
| Publication status | E-pub ahead of print - 11 Mar 2025 |
Bibliographical note
Copyright:© The Author(s) 2025.
Keywords
- Markov decision processes
- Nonlinear discount function
- Reduction
- Stochastic game
ASJC Scopus subject areas
- General Decision Sciences
- Management Science and Operations Research
Fingerprint
Dive into the research topics of 'Reduction of a Markov decision process with non-linear discounting to a stochastic game with standard total undiscounted criterion'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver