Shortcut Connections based Deep Speaker Embeddings for End-to-End Speaker Verification System

  • Seo, Soonshin
  • Rim, Daniel Jun
  • Lim, Minkyu
  • Lee, Donghyun
  • Park, Hosung
  • ... Kim, Ji-Hwan
  • 외 2명
Citations

WEB OF SCIENCE

24
Citations

SCOPUS

24

초록

The objective of speaker verification is to reject or accept whether or not the input speech is that of a enrolled speaker. Traditionally, i-vector or speaker embeddings system such as d-vector representing the speaker information has been showing high performance with similarity metrics at the backend. Recently it has been proposed an end-to-end system based on previous speaker embeddings approach without additional strategy after extraction. Among the various models, CNN based end-to-end system is showing state-of-the-art performance. CNN based model is trained to classify multiple speakers and speaker embeddings are extracted. In this paper, we propose shortcut connections based deep speaker embeddings for end-to-end speaker verification system. We construct modified ResNet-18 model so that the activation outputs from bottleneck architecture have shortcut connections to speaker embeddings. Deep speaker embeddings are extracted by jointly training in end-to-end approach. The model was constructed without other sophisticated methods such as length normalization, or additive margin softmax loss. When we tested proposed model on the unconstrained conditions data set called VoxCeleb1, the result showed EER of 3.03% when tested with high dimensional deep speaker embeddings. This is the state-of-the-art performance of end-to-end speaker verification model on VoxCeleb1.

키워드

end-to-end speaker verification systemdeep speaker embeddingsshortcut connectionsResNet
제목
Shortcut Connections based Deep Speaker Embeddings for End-to-End Speaker Verification System
저자
Seo, SoonshinRim, Daniel JunLim, MinkyuLee, DonghyunPark, HosungOh, JunseokKim, ChangminKim, Ji-Hwan
DOI
10.21437/Interspeech.2019-2195
발행일
2019
유형
Proceedings Paper
저널명
Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
2019-September
페이지
2928 ~ 2932