Korean Grapheme Unit-based Speech Recognition Using Attention-CTC Ensemble Network

  • Park, Hosung; 
  • Seo, Soonshin; 
  • Rim, Daniel Jun; 
  • Kim, Changmin; 
  • Son, Hyunsoo; 
  • ... Kim, Ji-Hwan; 
  • 외 1명
Citations

WEB OF SCIENCE

0
Citations

SCOPUS

7

초록

This study proposes an end-to-end speech recognition method based on the Attention-CTC ensemble network that uses Korean graphemes as recognition units. End-to-end speech recognition is a method that allows the processing of procedures that involved a number of modules, including the DNN-HMM-based acoustic model, the N-gram-based language model, and the WFST-based decoding network, with a single DNN network. To predict the outputs of the end-to-end model, this study utilizes grapheme-unit output structures. Building a network based on graphemes enables effective learning by reducing the number of output parameters to be predicted from 11,172 to 49. Towards this aim, this study designed an end-to-end model by combining the connectionist temporal classification (CTC), the DNN network structure primarily used in end-to-end learning, and the attention network model. The experiment resulted in a 10.5% syllable error rate.

키워드

Attention network; Connectionist temporal classification; Speech recognition; Deep neural network
제목
Korean Grapheme Unit-based Speech Recognition Using Attention-CTC Ensemble Network
저자
Park, Hosung; Seo, Soonshin; Rim, Daniel Jun; Kim, Changmin; Son, Hyunsoo; Park, Jeong-Sik; Kim, Ji-Hwan
DOI
10.1109/ismac.2019.8836146
발행일
2019-08
유형
Proceedings Paper
저널명
2019 INTERNATIONAL SYMPOSIUM ON MULTIMEDIA AND COMMUNICATION TECHNOLOGY (ISMAC)