Optimizing Policy via Deep Reinforcement Learning for Dialogue Management

  • Xu, Guanghao
  • Lee, Hyunjung
  • Koo, Myoung-Wan
  • Seo, Jungyun
Citations

WEB OF SCIENCE

1
Citations

SCOPUS

1

초록

In this paper, we propose a dialogue manager model based on Deep Reinforcement Learning, which automatically optimizes a dialogue policy. The policy is trained within deep Q-learning algorithm, which efficiently approximates value of actions given a large space of dialogue state. Evaluation processes are conducted by comparing the performance of the proposed model to a rule-based one on the dialogue corpora of DSTC2 and 3 under three different levels of error rate in Spoken Language Understanding. Experimental results prove that given certain level of SLU error, the dialogue manager with self-learned policy shows higher completion rate and the robustness to SLU error. Overcoming the drawbacks of rule-based approach such as limited flexibility and high maintenance cost, our model shows the strength of self-learning algorithm in optimizing policy of dialogue manager without any hand-crafted features.

키워드

Deep Reinforcement LearningDialogue ManagementDialogue Policy
제목
Optimizing Policy via Deep Reinforcement Learning for Dialogue Management
저자
Xu, GuanghaoLee, HyunjungKoo, Myoung-WanSeo, Jungyun
DOI
10.1109/BigComp.2018.00101
발행일
2018-05-25
유형
Proceedings Paper
저널명
2018 IEEE INTERNATIONAL CONFERENCE ON BIG DATA AND SMART COMPUTING (BIGCOMP)
페이지
582 ~ 589