Multi-stage adaptive sampling algorithms

Citations

SCOPUS

0

초록

In Chap. 2, we present simulation-based algorithms for estimating the optimal value function in finite-horizon MDPs with large (possibly uncountable) state spaces, where the usual techniques of policy iteration and value iteration are either computationally impractical or infeasible to implement. We present two adaptive sampling algorithms that estimate the optimal value function by choosing actions to sample in each state visited on a finite-horizon simulated sample path. The first approach builds upon the expected regret analysis of multi-armed bandit models and uses upper confidence bounds to determine which action to sample next, whereas the second approach uses ideas from learning automata to determine the next sampled action. The first approach is also the predecessor of a closely related approach in artificial intelligence (AI) called Monte Carlo tree search that led to a breakthrough in developing the current best computer Go-playing programs. © Springer-Verlag London 2013.

제목
Multi-stage adaptive sampling algorithms
저자
Chang, Hyeong SooHu, JiaqiaoFu, Michael C.Marcus, Steven I.
DOI
10.1007/978-1-4471-5022-0_2
발행일
2013
유형
Book Chapter
저널명
Communications and Control Engineering
9781447150213
페이지
19 ~ 60