상세 보기
Efficient online target speech extraction using DOA-constrained independent component analysis of stereo data for robust speech recognition
- Kim, Minook;
- Park, Hyung-Min
WEB OF SCIENCE
13SCOPUS
16초록
This paper describes an efficient online target-speech-extraction method used as a preprocessing step for robust automatic speech recognition (ASR). Because a target speaker is located relatively close to microphones in many ASR applications, acoustic paths to microphones are moderately reverberant, and the target speaker direction can easily be estimated. In this situation, noise estimation is effectively performed by forming a directional null to the target speaker. Required weights for extracting target speech, independent of the estimated noise, are then determined using an adaptation rule derived from a modified version of the cost function for independent component analysis (ICA), while retaining the minimal distortion principle. In particular, an online natural-gradient learning rule with a nonholonomic constraint and normalization by a smoothed power estimate of the input signal is derived for stable convergence, even for dynamically changing speech levels, with much less computational complexity than conventional ICA. Furthermore, stereo mixtures are considered as input data for further reduction of computational loads and fast convergence. Although the method may suffer from the underdetermined problem, the weights are adapted to obtain signal-to-noise-ratio-maximization beamformers for successful target speech estimation. The experimental results obtained for various conditions demonstrate the effectiveness of the proposed method. (C) 2015 Elsevier B.V. All rights reserved.
키워드
- 제목
- Efficient online target speech extraction using DOA-constrained independent component analysis of stereo data for robust speech recognition
- 저자
- Kim, Minook; Park, Hyung-Min
- 발행일
- 2015-12
- 유형
- Article
- 권
- 117
- 페이지
- 126 ~ 137