Value set iteration for Markov decision processes

Citations

WEB OF SCIENCE

4
Citations

SCOPUS

7

초록

This communique presents an algorithm called "value set iteration" (VSI) for solving infinite horizon discounted Markov decision processes with finite state and action spaces as a simple generalization of value iteration (VI) and as a counterpart to Chang's policy set iteration. A sequence of value functions is generated by VSI based on manipulating a set of value functions at each iteration and it converges to the optimal value function. VSI preserves convergence properties of VI while converging no slower than VI and in particular, if the set used in VSI contains the value functions of independently generated sample-policies from a given distribution and a properly defined policy switching policy, a probabilistic exponential convergence rate of VSI can be established. Because the set used in VSI can contain the value functions of any policies generated by other existing algorithms, VSI is also a general framework of combining multiple solution methods. (C) 2014 Elsevier Ltd. All rights reserved.

키워드

Markov decision processesValue iterationDynamic programmingConstrained optimizationALGORITHMSDESIGN
제목
Value set iteration for Markov decision processes
저자
Chang, Hyeong Soo
DOI
10.1016/j.automatica.2014.05.009
발행일
2014-07
유형
Article
저널명
Automatica
50
7
페이지
1940 ~ 1943