상세 보기
Recursive learning automata for control of partially observable Markov decision processes
- Chang, Hyeong Soo;
- Fu, Michael C.;
- Marcus, Steven I.
WEB OF SCIENCE
1SCOPUS
1초록
This paper presents a sampling algorithm, called "Recursive Automata Sampling Algorithm (RASA)," for control of finite horizon information-state Markov decision processes (MDPs), the equivalent model of partially observable MDPs. RASA extends in a recursive manner the Pursuit algorithm designed with learning automata by Rajaraman and Sastry for solving stochastic optimization problems. Based on the finite-time analysis of the Pursuit algorithm, we analyze the finite-time behavior of RASA, providing a bound on the probability that a given initial state takes the optimal action, and a bound on the probability that the difference between the optimal value and the estimate of it exceeds a given error. We also discuss how to apply RASA in the direct context of POMDPs and how to incorporate heuristic knowledge into RASA for on-tine control.
키워드
- 제목
- Recursive learning automata for control of partially observable Markov decision processes
- 저자
- Chang, Hyeong Soo; Fu, Michael C.; Marcus, Steven I.
- 발행일
- 2005
- 유형
- Proceedings Paper
- 저널명
- Proceedings of the IEEE Conference on Decision and Control
- 페이지
- 6091 ~ 6096