Multitask Learning with Fused Attention for Improved ASR and Mispronunciation Detection in Children's Speech Sound Disorders

  • Sung, Selina S.
  • Ha, Seung Hee
  • Yoon, Tae Jin
  • So, Jung Min
Citations

WEB OF SCIENCE

1
Citations

SCOPUS

0

초록

This study proposes a multitask learning framework with fused attention to enhance automatic speech recognition (ASR) for pronunciation-based transcription and mispronunciation detection (MPD) in speech sound disorders (SSD). Our approach leverages multitask learning by carrying out ASR and classification tasks concurrently. To further improve performance, we propose a fused attention mechanism that refines hidden states by weighting features relevant to mispronunciations. The classification head and the attention mechanism work synergistically, jointly optimizing transcription and detection performance. Evaluated on a Korean children SSD dataset, our approach outperforms the baseline, achieving lower Character Error Rates (CER) and higher Unweighted Average Recall (UAR), demonstrating the effectiveness of multitask learning with fused attention for mispronunciation detection.

키워드

audio classificationautomatic speech recognitionmispronunciation detectionspeech sound disorders
제목
Multitask Learning with Fused Attention for Improved ASR and Mispronunciation Detection in Children's Speech Sound Disorders
저자
Sung, Selina S.Ha, Seung HeeYoon, Tae JinSo, Jung Min
DOI
10.21437/Interspeech.2025-259
발행일
2025
유형
Proceedings Paper
저널명
Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
페이지
5698 ~ 5702