Environmental Noise Robustness for Korean Fricatives using Speech Enhancement Generative Adversarial Networks

  • Seo, Soonshin
  • Lim, Minkyu
  • Lee, Donghyun
  • Park, Hosung
  • Oh, Junseok
  • ... Kim, Ji-Hwan
  • 외 1명
Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Currently, speech recognition technology has a high recognition rate in a quiet environment condition. However, noise processing is needed in environmental noise conditions because they lower the recognition rate. In particular, in the case of Korean fricatives corresponding to [s], [c], [h], [c], [x], and [.], noise processing is relatively difficult because the acoustic characteristics are similar to environmental noise. This paper proposes speech enhancement of Korean fricatives using the Speech Enhancement Generative Adversarial Network (SEGAN) for environmental noise robustness. SEGAN is a Deep Neural Network (DNN)-based generation model that is used to improve the clarity and quality of environmental noisy speech. Enhanced Korean fricative speech data are generated using SEGANtrained Korean fricative speech data. The results showed that using DNN-Hidden Markov Model (HMM)-based acoustic model training by adding enhanced speech as training data resulted in a 0.26% Character Error Rate (CER) in environmental noise conditions compared to the case without enhanced speech addition.

키워드

environmental noiserobust speech recognitionKorean fricativesspeech enhancementgenerative adversarial network
제목
Environmental Noise Robustness for Korean Fricatives using Speech Enhancement Generative Adversarial Networks
저자
Seo, SoonshinLim, MinkyuLee, DonghyunPark, HosungOh, JunseokRim, Daniel JunKim, Ji-Hwan
DOI
10.1109/bigcomp.2019.8679197
발행일
2019-04-01
유형
Proceedings Paper
저널명
2019 IEEE INTERNATIONAL CONFERENCE ON BIG DATA AND SMART COMPUTING (BIGCOMP)
페이지
383 ~ 386