상세 보기
Survey on Deep Learning-based Speech Technologies in Voice Chatbot Systems
- Ma, Seunghee;
- Oh, Junseok;
- Kim, Minseo;
- Kim, Ji-Hwan
WEB OF SCIENCE
0SCOPUS
0초록
Recent advancements in large language models (LLMs) such as ChatGPT have contributed to the development of chatbot systems. Specifically, speech has been recognized as the optimal tool for interactive dialogue, leading to increased interest in voice chatbots. Voice chatbots offer information and services through voice interactions, enhancing user experience and improving service accessibility. This survey paper introduces the latest developments in the core technologies of voice chatbot systems, including deep learning-based automatic speech recognition, speech synthesis, and speech emotion recognition. It focuses on advanced research to enhance speed and performance, which is crucial for applying speech technologies in voice chatbots. In automatic speech recognition, we introduce methodologies such as Connectionist Temporal Classification (CTC), Attention based Encoder-Decoder (AED), and Recurrent Neural Network Transducer (RNN-T), along with studies optimizing Transformer-based models for real-time automatic speech recognition. In speech emotion recognition, we explore the use of pre-trained models and the latest techniques for accurate emotion prediction. For speech synthesis, the focus extends to two-stage and End-to-End (E2E) approaches, with additional research on integrating emotional information to generate natural speech.
키워드
- 제목
- Survey on Deep Learning-based Speech Technologies in Voice Chatbot Systems
- 저자
- Ma, Seunghee; Oh, Junseok; Kim, Minseo; Kim, Ji-Hwan
- 발행일
- 2025-05-31
- 유형
- Article
- 권
- 19
- 호
- 5
- 페이지
- 1406 ~ 1440