Masked cross self-attentive encoding based speaker embedding for speaker verification

Citations

SCOPUS

0

초록

Constructing speaker embeddings in speaker verification is an important issue. In general, a self-attention mechanism has been applied for speaker embedding encoding. Previous studies focused on training the self-attention in a high-level layer, such as the last pooling layer. In this case, the effect of low-level layers is not well represented in the speaker embedding encoding. In this study, we propose Masked Cross Self-Attentive Encoding (MCSAE) using ResNet. It focuses on training the features of both high-level and low-level layers. Based on multi-layer aggregation, the output features of each residual layer are used for the MCSAE. In the MCSAE, the interdependence of each input features is trained by cross self-attention module. A random masking regularization module is also applied to prevent overfitting problem. The MCSAE enhances the weight of frames representing the speaker information. Then, the output features are concatenated and encoded in the speaker embedding. Therefore, a more informative speaker embedding is encoded by using the MCSAE. The experimental results showed an equal error rate of 2.63 % using the VoxCeleb1 evaluation dataset. It improved performance compared with the previous self-attentive encoding and state-of-the-art methods. © © 2020 The Acoustical Society of Korea.

키워드

화자검증마스킹된 교차 자기주의 인코딩화자 임베딩잔차 네트워크Speaker verificationMasked cross self-attentive encodingSpeaker embeddingResNet
제목
Masked cross self-attentive encoding based speaker embedding for speaker verification
저자
Seo, SoonshinKim, Ji Hwan
DOI
10.7776/ASK.2020.39.5.497
발행일
2020
유형
Article
저널명
한국음향학회지
39
5
페이지
497 ~ 504