상세 보기
Environmental Noise Robustness for Korean Fricatives using Speech Enhancement Generative Adversarial Networks
- Seo, Soonshin;
- Lim, Minkyu;
- Lee, Donghyun;
- Park, Hosung;
- Oh, Junseok;
- ... Kim, Ji-Hwan;
- 외 1명
WEB OF SCIENCE
0SCOPUS
0초록
Currently, speech recognition technology has a high recognition rate in a quiet environment condition. However, noise processing is needed in environmental noise conditions because they lower the recognition rate. In particular, in the case of Korean fricatives corresponding to [s], [c], [h], [c], [x], and [.], noise processing is relatively difficult because the acoustic characteristics are similar to environmental noise. This paper proposes speech enhancement of Korean fricatives using the Speech Enhancement Generative Adversarial Network (SEGAN) for environmental noise robustness. SEGAN is a Deep Neural Network (DNN)-based generation model that is used to improve the clarity and quality of environmental noisy speech. Enhanced Korean fricative speech data are generated using SEGANtrained Korean fricative speech data. The results showed that using DNN-Hidden Markov Model (HMM)-based acoustic model training by adding enhanced speech as training data resulted in a 0.26% Character Error Rate (CER) in environmental noise conditions compared to the case without enhanced speech addition.
키워드
- 제목
- Environmental Noise Robustness for Korean Fricatives using Speech Enhancement Generative Adversarial Networks
- 저자
- Seo, Soonshin; Lim, Minkyu; Lee, Donghyun; Park, Hosung; Oh, Junseok; Rim, Daniel Jun; Kim, Ji-Hwan
- 발행일
- 2019-04-01
- 유형
- Proceedings Paper
- 저널명
- 2019 IEEE INTERNATIONAL CONFERENCE ON BIG DATA AND SMART COMPUTING (BIGCOMP)
- 페이지
- 383 ~ 386