상세 보기
Multilingual speech-to-vocal tract visualization using deep learning for pronunciation training
- Picinini Mexas, Rodrigo;
- Chu, Yunji;
- Park, Unsang
WEB OF SCIENCE
0SCOPUS
0초록
Visualizing the vocal tract during speech remains a challenging task, even with recent advancements in open-source algorithms and datasets. A key limitation is the lack of multimodal resources that integrate audio with internal articulatory structures, which poses challenges to the development of effective speech visualization methods. In this work, we propose a novel algorithm that translates speech into dynamic vocal tract movements, enabling users to visually compare their articulation against a correctly pronounced reference. This approach is particularly beneficial for language learners and individuals with congenital deafness, as it provides an intuitive visualization of speech production. To address the dataset gap, we construct a new corpus that maps audio to articulatory parameters used in VocalTractLab. Leveraging pretrained models such as Wav2Vec 2.0 and HuBERT for audio feature extraction, we frame the task as a sequence-to-sequence learning problem and employ a Bi-GRU to predict the corresponding articulatory parameters. Furthermore, we develop language-specific datasets for General American English, Korean, and Brazilian Portuguese using the same pipeline, and train dedicated models for each. We validate the effectiveness of our approach through both qualitative and quantitative evaluations and assess its practical utility in pronunciation correction tasks through a user survey. The goal of the survey was to evaluate the perceived helpfulness and usability of the visualizations, not to implement an iterative pronunciation correction process.
키워드
- 제목
- Multilingual speech-to-vocal tract visualization using deep learning for pronunciation training
- 저자
- Picinini Mexas, Rodrigo; Chu, Yunji; Park, Unsang
- 발행일
- 2025-12
- 유형
- Article
- 권
- 2026
- 호
- 1