Adversarial multi-armed bandit approach to stochastic optimization

Citations

WEB OF SCIENCE

1
Citations

SCOPUS

1

초록

We first present a sampling-based algorithm for solving stochastic optimization problems based on the Exp3 algorithm by Auer et al. for "adversarial multi-armed bandit problems." The value returned by the algorithm is an upperbound estimate of the sample-average-maximum for a sample average approximation (SAA) problem induced by the sampling process of the algorithm, and the estimate converges to the optimal objective-function value as the number of samples goes to infinity. We then recursively extend the Exp3-based algorithm for solving a given finite-horizon Markov decision process (MDP) and analyze its finite-time performance in terms of the expected bias relative to the maximum value of the induced recursive SAA problem, showing that the upper bound of the expected bias approaches zero as the sampling size per stage goes to infinity, leading to the convergence to the optimal value of the original MDP problem in the limit.

키워드

FINITE-TIME ANALYSISALGORITHM
제목
Adversarial multi-armed bandit approach to stochastic optimization
저자
Chang, Hyeong SooFu, Michael C.Marcus, Steven I.
발행일
2006
유형
Proceedings Paper
저널명
Proceedings of the IEEE Conference on Decision and Control
페이지
5684 ~ +