Reinforcement learning with supervision by combining multiple learnings and expert advices

Citations

WEB OF SCIENCE

1
Citations

SCOPUS

5

초록

In this paper, we provide a formal coherent learning framework where reinforcement learning is combined with multiple learnings and expert advices toward accelerating convergence speed of learning. Our approach is simply to use a nonstationary "potential-based reinforcement function" for shaping the reinforcement signal given to the learning "base-agent". The base-agent employes SARSA(0) or adaptive asynchronous value iteration (VI), and the supervised inputs to the base-agent from the "subagents" involved with other parallel independent reinforcement learnings and if available, from experts are "merged" into the potential-based reinforcement function value and the value is put into the update equation of SARSA(0) for the Q-function estimate or of adaptive asynchronous VI for the optimal value function estimate. The resulting SARSA(0) and adaptive asynchronous VI converge to an optimal policy, respectively.

제목
Reinforcement learning with supervision by combining multiple learnings and expert advices
저자
Chang, Hyeong Soo
발행일
2006
유형
Proceedings Paper
저널명
Proceedings of the American Control Conference
1-12
페이지
4159 ~ 4164