Acoustic model training using self-attention for low-resource speech recognition

Citations

SCOPUS

0

초록

This paper proposes acoustic model training using self-attention for low-resource speech recognition. In low-resource speech recognition, it is difficult for acoustic model to distinguish certain phones. For example, plosive /d/ and /t/, plosive /g/ and /k/ and affricate /z/ and /ch/. In acoustic model training, the self-attention generates attention weights from the deep neural network model. In this study, these weights handle the similar pronunciation error for low-resource speech recognition. When the proposed method was applied to Time Delay Neural Network-Output gate Projected Gated Recurrent Unit (TNDD-OPGRU)-based acoustic model, the proposed model showed a 5.98 % word error rate. It shows absolute improvement of 0.74 % compared with TDNN-OPGRU model. © © 2020 The Acoustical Society of Korea.

키워드

Acoustic modelLow-resource environmentSelf-attention mechanismSpeech recognition
제목
Acoustic model training using self-attention for low-resource speech recognition
저자
Park, HosungKim, Ji Hwan
DOI
10.7776/ASK.2020.39.5.483
발행일
2020
유형
Article
저널명
한국음향학회지
39
5
페이지
483 ~ 489