Conformer with lexicon transducer for Korean end-to-end speech recognition

  • Son, Hyunsoo
  • Park, Ho sung
  • Kim, Gyu jin
  • Cho, Eun soo
  • Kim, Ji Hwan
Citations

SCOPUS

0

초록

Recently, due to the development of deep learning, end-to-end speech recognition, which directly maps graphemes to speech signals, shows good performance. Especially, among the end-to-end models, conformer shows the best performance. However end-to-end models only focuses on the probability of which grapheme will appear at the time. The decoding process uses a greedy search or beam search. This decoding method is easily affected by the final probability output by the model. In addition, the end-to-end models cannot use external pronunciation and language information due to structual problem. Therefore, in this paper conformer with lexicon transducer is proposed. We compare phoneme-based model with lexicon transducer and grapheme-based model with beam search. Test set is consist of words that do not appear in training data. The grapheme-based conformer with beam search shows 3.8 % of CER. The phoneme-based conformer with lexicon transducer shows 3.4 % of CER. © 2021 The Acoustical Society of Korea.

키워드

End-to-endSpeech recognitionTransformerWeighted finite state transducer
제목
Conformer with lexicon transducer for Korean end-to-end speech recognition
저자
Son, HyunsooPark, Ho sungKim, Gyu jinCho, Eun sooKim, Ji Hwan
DOI
10.7776/ASK.2021.40.5.530
발행일
2021
유형
Article
저널명
한국음향학회지
40
5
페이지
530 ~ 536