On Supervised Online Rolling-Horizon Control for Infinite-Horizon Discounted Markov Decision Processes

Citations

WEB OF SCIENCE

3
Citations

SCOPUS

3

초록

This note revisits the rolling-horizon control approach to the problem of Markov decision process (MDP) with infinite-horizon discounted expected reward criterion. Distinguished from the classical value-iteration approaches, we develop an asynchronous online algorithm based on policy iteration integrated with a multipolicy improvement method of policy switching. A sequence of monotonically improving solutions to the forecast-horizon sub-MDP is generated by updating the current solution only at the currently visited state, building in effect a rolling-horizon control policy for the MDP over infinite horizon. Feedbacks from "supervisors," if available, can be also incorporated while updating. We focus on the convergence issue with a relation to the transition structure of the MDP. Either a global convergence to an optimal forecast-horizon policy or a local convergence to a "locally-optimal" fixed-policy in a finite time is achieved by the algorithm depending on the structure.

키워드

Markov processesHeuristic algorithmsMarkov decision process (MDP)policy iteration (PI)policy switchingrolling horizon controlSET ITERATION
제목
On Supervised Online Rolling-Horizon Control for Infinite-Horizon Discounted Markov Decision Processes
저자
Chang, Hyeong Soo
DOI
10.1109/TAC.2023.3274791
발행일
2024-02
유형
Article
저널명
IEEE Transactions on Automatic Control
69
2
페이지
1060 ~ 1065