Evolutionary policy iteration for solving Markov decision processes

Citations

WEB OF SCIENCE

35
Citations

SCOPUS

39

초록

We propose a novel algorithm called evolutionary policy iteration (EPI) for solving infinite horizon discounted reward Markov decision processes. EPI inherits the spirit of policy iteration but eliminates the need to maximize over the entire action space in the policy improvement step, so it should be most effective for problems with very large action spaces. EPI iteratively generates a "population" or a set of policies such that the performance of the "elite policy" for a population monotonically improves with respect to a defined fitness function. EPI converges with probability one to a population whose elite policy is an optimal policy. EPI is naturally parallelizable and along this discussion, a distributed variant of PI is also studied.

키워드

(distributed) policy iterationevolutionary algorithmgenetic algorithmMarkov decision processparallelization
제목
Evolutionary policy iteration for solving Markov decision processes
저자
Chang, HSLee, HGFu, MCMarcus, SI
DOI
10.1109/TAC.2005.858644
발행일
2005-11
유형
Article
저널명
IEEE Transactions on Automatic Control
50
11
페이지
1804 ~ 1808