On functional equations for Kth best policies in Markov decision processes

Citations

WEB OF SCIENCE

2
Citations

SCOPUS

2

초록

This paper revisits the problem of finding the values of Kth best policies for finite-horizon finite Markov decision processes. The recursive dynamic-programming (DP) equations established by Bellman and Kalaba for non-deterministic MDPs with zero-cost function in [Bellman, R., & Kalaba, R. (1960). On kth best policies. Journal of SIAM, 8,582-588] are incomplete because expectation and selection for the Kth minimum do not interchange in general. Based on the DP equations by Dreyfus for the Kth shortest path problem, some non-DP equations generally satisfied by the values of the Kth best policies are identified, from which corrected Bellman and Kalaba's DP equations are derived with an appropriate sufficient condition. (C) 2012 Elsevier Ltd. All rights reserved.

키워드

Markov decision processesDynamic programmingRanksSHORTESTPATHS
제목
On functional equations for Kth best policies in Markov decision processes
저자
Chang, Hyeong Soo
DOI
10.1016/j.automatica.2012.09.016
발행일
2013-01
유형
Article
저널명
Automatica
49
1
페이지
297 ~ 300