상세 보기
Solving controlled Markov set-chains with discounting via multipolicy improvement
- Chang, Hyeong Soo;
- Chong, Edwin K. P.
Citations
WEB OF SCIENCE
2Citations
SCOPUS
1초록
We consider Markov decision processes (MDPs) where the state transition probability distributions are not uniquely known, but are known to belong to some intervals-so called "controlled Markov set-chains"-with infinite-horizon discounted reward criteria. We present formal methods to improve multiple policies for solving such controlled Markov set-chains. Our multipolicy improvement methods follow the spirit of parallel rollout and policy switching for solving MDPs. In particular, these methods are useful for online control of Markov set-chains and for designing policy iteration (PI) type algorithms. We develop a PI-type algorithm and prove that it converges to an optimal policy.
키워드
controlled Markov process; Markov decision process (MDP); Markov set-chain; policy iteration; rollout; DECISION-PROCESSES
- 제목
- Solving controlled Markov set-chains with discounting via multipolicy improvement
- 저자
- Chang, Hyeong Soo; Chong, Edwin K. P.
- 발행일
- 2007-03
- 유형
- Article
- 권
- 52
- 호
- 3
- 페이지
- 564 ~ 569