상세 보기
Acoustic model training using self-attention for low-resource speech recognition
- Park, Hosung;
- Kim, Ji Hwan
SCOPUS
0초록
This paper proposes acoustic model training using self-attention for low-resource speech recognition. In low-resource speech recognition, it is difficult for acoustic model to distinguish certain phones. For example, plosive /d/ and /t/, plosive /g/ and /k/ and affricate /z/ and /ch/. In acoustic model training, the self-attention generates attention weights from the deep neural network model. In this study, these weights handle the similar pronunciation error for low-resource speech recognition. When the proposed method was applied to Time Delay Neural Network-Output gate Projected Gated Recurrent Unit (TNDD-OPGRU)-based acoustic model, the proposed model showed a 5.98 % word error rate. It shows absolute improvement of 0.74 % compared with TDNN-OPGRU model. © © 2020 The Acoustical Society of Korea.
키워드
- 제목
- Acoustic model training using self-attention for low-resource speech recognition
- 저자
- Park, Hosung; Kim, Ji Hwan
- 발행일
- 2020
- 유형
- Article
- 저널명
- 한국음향학회지
- 권
- 39
- 호
- 5
- 페이지
- 483 ~ 489