Approximate stochastic annealing for online control of infinite horizon Markov decision processes

Citations

WEB OF SCIENCE

3
Citations

SCOPUS

3

초록

We present an online simulation-based algorithm called Approximate Stochastic Annealing (ASA) for solving infinite-horizon finite state-action space Markov decision processes (MDPs). The algorithm estimates the optimal policy by sampling at each iteration from a probability distribution function over the policy space, which is updated iteratively based on the Q-function estimates obtained via a recursion of Q-learning type. By exploiting a novel connection of ASA to the stochastic approximation method, we show that the sequence of distribution functions generated by the algorithm converges to a degenerated distribution that concentrates only on the optimal policy. Numerical examples are also provided to illustrate the algorithm. (C) 2012 Elsevier Ltd. All rights reserved.

키워드

AlgorithmsMarkov decision processStochastic approximationSimulationALGORITHMCONVERGENCE
제목
Approximate stochastic annealing for online control of infinite horizon Markov decision processes
저자
Hu, JiaqiaoChang, Hyeong Soo
DOI
10.1016/j.automatica.2012.06.010
발행일
2012-09
유형
Article
저널명
Automatica
48
9
페이지
2182 ~ 2188