Learned Error, I Can Fix It! : A Detector-Corrector Structure for ASR Error Calibration

  • Yeen, Heui-Yeen
  • Kim, Min-Ju
  • Koo, Myoung-Wan
Citations

WEB OF SCIENCE

2
Citations

SCOPUS

2

초록

Speech recognition technology has improved recently. However, in the context of spoken language understanding (SLU), containing automatic speech recognition (ASR) errors causes significant downstream performance degradation. To address this issue, various ASR error correction methodologies have been proposed. ASR error correction mainly focuses on correcting and generating only the error span using a conditional decoding method. To this end, we propose a structure with a Detector that uses collaborative training to predict various error patterns and a Corrector that corrects the detected error span by Detector. This pipeline reduces Word Error Rate (WER) and shows less performance degradation in downstream tasks compared with the original ASR hypotheses. In addition, it was shown that it could be generalized to various downstream data. By leveraging this Detector-Corrector pipeline, we expect to achieve effective ASR error correction and enable high-quality SLU downstream tasks.

키워드

error correctionSpoken language Understandingintent classificationemotion recognition
제목
Learned Error, I Can Fix It! : A Detector-Corrector Structure for ASR Error Calibration
저자
Yeen, Heui-YeenKim, Min-JuKoo, Myoung-Wan
DOI
10.21437/Interspeech.2023-2475
발행일
2023
유형
Proceedings Paper
저널명
Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
2023-August
페이지
2693 ~ 2697