상세 보기
Masked cross self-attentive encoding based speaker embedding for speaker verification
- Seo, Soonshin;
- Kim, Ji Hwan
SCOPUS
0초록
Constructing speaker embeddings in speaker verification is an important issue. In general, a self-attention mechanism has been applied for speaker embedding encoding. Previous studies focused on training the self-attention in a high-level layer, such as the last pooling layer. In this case, the effect of low-level layers is not well represented in the speaker embedding encoding. In this study, we propose Masked Cross Self-Attentive Encoding (MCSAE) using ResNet. It focuses on training the features of both high-level and low-level layers. Based on multi-layer aggregation, the output features of each residual layer are used for the MCSAE. In the MCSAE, the interdependence of each input features is trained by cross self-attention module. A random masking regularization module is also applied to prevent overfitting problem. The MCSAE enhances the weight of frames representing the speaker information. Then, the output features are concatenated and encoded in the speaker embedding. Therefore, a more informative speaker embedding is encoded by using the MCSAE. The experimental results showed an equal error rate of 2.63 % using the VoxCeleb1 evaluation dataset. It improved performance compared with the previous self-attentive encoding and state-of-the-art methods. © © 2020 The Acoustical Society of Korea.
키워드
- 제목
- Masked cross self-attentive encoding based speaker embedding for speaker verification
- 저자
- Seo, Soonshin; Kim, Ji Hwan
- 발행일
- 2020
- 유형
- Article
- 저널명
- 한국음향학회지
- 권
- 39
- 호
- 5
- 페이지
- 497 ~ 504