A policy improvement method in constrained stochastic dynamic programming

Citations

WEB OF SCIENCE

13
Citations

SCOPUS

15

초록

This note presents a formal method of improving a given base-policy such that the performance of the resulting policy is no worse than that of the base-policy at all states in constrained stochastic dynamic programming. We consider finite horizon and discounted infinite horizon cases. The improvement method induces a policy iteration-type algorithm that converges to a local optimal policy.

키워드

constrained Markov decision processdynamic programmingpolicy improvementpolicy iterationMARKOV DECISION-PROCESSES
제목
A policy improvement method in constrained stochastic dynamic programming
저자
Chang, Hyeong Soo
DOI
10.1109/TAC.2006.880801
발행일
2006-09
유형
Article
저널명
IEEE Transactions on Automatic Control
51
9
페이지
1523 ~ 1526