Multi-Model ASR Integration With Reliability Weighting for Automated Speech Disorder Screening

  • Sung, Selina S.
  • Ha, Seunghee
  • Yoon, Tae-Jin
  • So, Jungmin
Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

Speech sound disorders affect communication development in children, requiring early detection for timely intervention. Traditional clinical assessment relies on manual phonetic transcription and calculation of Percent Consonants Correct, a labor-intensive process limiting accessibility in underserved regions. We present an automated system for Korean child speech sound disorder screening based on standardized picture-naming tasks, using multi-model automatic speech recognition integration with reliability-weighted machine learning. Our approach fine-tunes four state-of-the-art models with multiple independent training runs, employing intra-model ensemble methods to generate robust transcriptions. We then train gradient boosting classifiers using feature vectors that combine model-specific Percent Consonants Correct predictions with reliability indicators capturing prediction uncertainty and inter-model agreement. This allows us to predict human-annotated consonant accuracy and classify children as typically developing or speech-sound-disordered based on age-stratified normative thresholds. Evaluated on 572 Korean children aged 2.5-9 years with 21,878 utterances across 5-fold speaker-stratified cross-validation, our system achieves 82.4% unweighted average recall, 86.9% accuracy, and 0.759 F1-score, improving upon the best single-model performance by 5.3 percentage points in unweighted average recall. Ablation studies confirm that multi-model integration and reliability-based weighting are both critical for accurate consonant accuracy prediction. This work demonstrates that imperfect automatic speech recognition, when combined with ensemble-based uncertainty estimation and multi-model integration, can achieve reliable speech disorder screening, potentially expanding access to diagnostic services in resource-limited settings.

키워드

ReliabilityUncertaintyAcousticsAccuracyTrainingEnsemble learningPediatricsFeature extractionPredictive modelsPhoneticsAutomatic speech recognitionchild speech disordersensemble methodsKorean phonologypercent consonants correctreliability weightinguncertainty estimationCHILDREN
제목
Multi-Model ASR Integration With Reliability Weighting for Automated Speech Disorder Screening
저자
Sung, Selina S.Ha, SeungheeYoon, Tae-JinSo, Jungmin
DOI
10.1109/ACCESS.2026.3667615
발행일
2026-02
유형
Article
저널명
IEEE Access
14
페이지
30200 ~ 30222