상세 보기
An asymptotically efficient simulation-based algorithm for finite horizon stochastic dynamic programming
- Chang, Hyeong Soo;
- Fu, Michael C.;
- Hu, Jiaqiao;
- Marcus, Steven I.
Citations
WEB OF SCIENCE
13Citations
SCOPUS
18초록
We present a simulation-based algorithm called "Simulated Annealing Multiplicative Weights" (SAMW) for solving large finite-horizon stochastic dynamic programming problems. At each iteration of the algorithm, a probability distribution over candidate policies is updated by a simple multiplicative weight rule, and with proper annealing of a control parameter, the generated sequence of distributions converges to a distribution concentrated only on the best policies. The algorithm is "asymptotically efficient," in the sense that for the goal of estimating the value of an optimal policy, a provably convergent finite-time upper bound for the sample mean is obtained.
키워드
learning algorithms; Markov decision processes; simulation; simulated annealing; stochastic dynamic programming
- 제목
- An asymptotically efficient simulation-based algorithm for finite horizon stochastic dynamic programming
- 저자
- Chang, Hyeong Soo; Fu, Michael C.; Hu, Jiaqiao; Marcus, Steven I.
- 발행일
- 2007-01
- 유형
- Article
- 권
- 52
- 호
- 1
- 페이지
- 89 ~ 94