Data Augmentation using Speech Synthesis for Speaker-Independent Dysarthria Severity Classification

  • Kim, Minseop
  • Han, Min su
  • Hong, Seo kyoung
  • Koo, Myoung Wan
Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Accurate dysarthria severity classification is essential for assessing motor speech disorders, and automation can improve efficiency and accessibility in clinical settings. While deep learning has significantly advanced this field, recent studies have increasingly leveraged large foundation ASR models. However, most studies focus on speaker-dependent (SD) classification, leaving speaker-independent (SI) classification as a major challenge due to limited datasets. SI classification is crucial in real-world scenarios where patient-specific information is unavailable. To address this, we applied two types of speech synthesis models for the first time in this task. We explore various strategies for integrating zero-shot text-to-speech (ZS-TTS) and voice conversion (VC) models to enhance SI classification and propose the most effective utilization settings. Our approach significantly improves the SI severity classification performance, paving the way for further research in this area.

키워드

Dysarthric SpeechSeverity ClassificationSpeech Synthesiszero-shot TTS
제목
Data Augmentation using Speech Synthesis for Speaker-Independent Dysarthria Severity Classification
저자
Kim, MinseopHan, Min suHong, Seo kyoungKoo, Myoung Wan
DOI
10.21437/Interspeech.2025-2711
발행일
2025
유형
Proceedings Paper
저널명
Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
페이지
2745 ~ 2749