Speech extraction based on AuxIVA with weighted source variance and noise dependence for robust speech recognition

Citations

SCOPUS

0

초록

In this paper, we propose speech enhancement algorithm as a pre-processing for robust speech recognition in noisy environments. Auxiliary-function-based Independent Vector Analysis (AuxIVA) is performed with weighted covariance matrix using time-varying variances with scaling factor from target masks representing time-frequency contributions of target speech. The mask estimates can be obtained using Neural Network (NN) pre-trained for speech extraction or diffuseness using Coherence-to-Diffuse power Ratio (CDR) to find the direct sounds component of a target speech. In addition, outputs for omni-directional noise are closely chained by sharing the time-varying variances similarly to independent subspace analysis or IVA. The speech extraction method based on AuxIVA is also performed in Independent Low-Rank Matrix Analysis (ILRMA) framework by extending the Non-negative Matrix Factorization (NMF) for noise outputs to Non-negative Tensor Factorization (NTF) to maintain the inter-channel dependency in noise output channels. Experimental results on the CHiME-4 datasets demonstrate the effectiveness of the presented algorithms. © 2022 The Acoustical Society of Korea.

키워드

Adaptive filteringIndependent vector analysisRobust speech recognitionSpeech extractionVariance
제목
Speech extraction based on AuxIVA with weighted source variance and noise dependence for robust speech recognition
저자
Shin, Ui-HyeopPark, Hyung-Min
DOI
10.7776/ASK.2022.41.3.326
발행일
2022-05
유형
Article
저널명
한국음향학회지
41
3
페이지
326 ~ 334