An Approximate Stochastic Annealing Algorithm for Finite Horizon Markov Decision Processes

Citations

WEB OF SCIENCE

6
Citations

SCOPUS

6

초록

We present a simulation-based algorithm called Approximate Stochastic Annealing (ASA) for solving finite-horizon Markov decision processes (MDPs). The algorithm iteratively estimates the optimal policy by sampling from a sequence of probability distribution functions over the policy space. By exploiting a novel connection of ASA to the stochastic approximation method, we show that the sequence of distribution functions generated by the algorithm converges to a degenerated distribution that concentrates only on the optimal policy. Numerical examples are also provided to illustrate the algorithm.

제목
An Approximate Stochastic Annealing Algorithm for Finite Horizon Markov Decision Processes
저자
Hu, JiaqiaoChang, Hyeong Soo
DOI
10.1109/CDC.2010.5717689
발행일
2010
유형
Proceedings Paper
저널명
Proceedings of the IEEE Conference on Decision and Control
페이지
5338 ~ 5343