An asymptotically efficient simulation-based algorithm for finite horizon stochastic dynamic programming

Citations

WEB OF SCIENCE

13
Citations

SCOPUS

18

초록

We present a simulation-based algorithm called "Simulated Annealing Multiplicative Weights" (SAMW) for solving large finite-horizon stochastic dynamic programming problems. At each iteration of the algorithm, a probability distribution over candidate policies is updated by a simple multiplicative weight rule, and with proper annealing of a control parameter, the generated sequence of distributions converges to a distribution concentrated only on the best policies. The algorithm is "asymptotically efficient," in the sense that for the goal of estimating the value of an optimal policy, a provably convergent finite-time upper bound for the sample mean is obtained.

키워드

learning algorithmsMarkov decision processessimulationsimulated annealingstochastic dynamic programming
제목
An asymptotically efficient simulation-based algorithm for finite horizon stochastic dynamic programming
저자
Chang, Hyeong SooFu, Michael C.Hu, JiaqiaoMarcus, Steven I.
DOI
10.1109/TAC.2006.887917
발행일
2007-01
유형
Article
저널명
IEEE Transactions on Automatic Control
52
1
페이지
89 ~ 94