Random search for constrained Markov decision processes with multi-policy improvement

Citations

WEB OF SCIENCE

3
Citations

SCOPUS

3

초록

This communique first presents a novel multi-policy improvement method which generates a feasible policy at least as good as any policy in a given set of feasible policies in finite constrained Markov decision processes (CMDPs). A random search algorithm for finding an optimal feasible policy for a given CMDP is derived by properly adapting the improvement method. The algorithm alleviates the major drawback of solving unconstrained MDPs at iterations in the existing value-iteration and policy-iteration type exact algorithms. We establish that the sequence of feasible policies generated by the algorithm converges to an optimal feasible policy with probability one and has a probabilistic exponential convergence rate. (C) 2015 Elsevier Ltd. All rights reserved.

키워드

Markov decision processesRandom searchPolicy improvementConstrained optimizationITERATION
제목
Random search for constrained Markov decision processes with multi-policy improvement
저자
Chang, Hyeong Soo
DOI
10.1016/j.automatica.2015.05.016
발행일
2015-08
유형
Article
저널명
Automatica
58
페이지
127 ~ 130