상세 보기
On-line control methods via simulation
- Chang, Hyeong Soo;
- Hu, Jiaqiao;
- Fu, Michael C.;
- Marcus, Steven I.
SCOPUS
0초록
In Chap. 5, we consider an approximate rolling-horizon control framework for solving infinite-horizon MDPs with large state/action spaces in an on-line manner by simulation. Specifically, we consider policies in which the system (either the actual system itself or a simulation model of the system) evolves to a particular state that is observed, and the action to be taken in that particular state is then computed on-line at the decision time, with a particular emphasis on the use of simulation. We first present an updating scheme involving multiplicative weights for updating a probability distribution over a restricted set of policies; this scheme can be used to estimate the optimal value function over this restricted set by sampling on the (restricted) policy space. The lower-bound estimate of the optimal value function is used for constructing on-line control policies, called (simulated) policy switching and parallel rollout. We also discuss an upper-bound-based method, called hindsight optimization. Finally, we present an algorithm, called approximate stochastic annealing, which combines Q-learning with the MARS algorithm of Sect. 4.6.1 to directly search the policy space. © Springer-Verlag London 2013.
- 제목
- On-line control methods via simulation
- 저자
- Chang, Hyeong Soo; Hu, Jiaqiao; Fu, Michael C.; Marcus, Steven I.
- 발행일
- 2013
- 유형
- Book Chapter
- 호
- 9781447150213
- 페이지
- 179 ~ 218