Aligning Incomplete Lyrics of Korean Folk Song Dataset using Whisper

Citations

WEB OF SCIENCE

0
Citations

SCOPUS

0

초록

In this study, we introduce a method for time-alignment of lyrics in Korean folk song audio using a transformer encoder-decoder model specifically designed to utilize incomplete lyric data. We analyzed the characteristics of Korean folk song lyrics and found some discrepancies between the lyrics and the corresponding audio recordings. To address these challenges and maximize the use of existing transcriptions, we introduce RefWhisper. This is a variant of OpenAI's Whisper and includes an extra encoder module and cross-attention layer, enabling the model to consult incomplete lyrics during the transcription process. The added cross-attention layer facilitates not only the alignment of the reference text with the predicted transcription but also with the audio. We make public the transcribed outcomes and timestamp data, which are aligned at both the sentence and word levels, for a corpus of 13,801 Korean folk songs.

키워드

datasetsDNNKorean folk songlyric alignmentlyric transcription
제목
Aligning Incomplete Lyrics of Korean Folk Song Dataset using Whisper
저자
Han, DanbinaerinKim, DaewoongJeong, Dasaem
DOI
10.1145/3625135.3625154
발행일
2023-11-10
유형
Proceedings Paper
저널명
THE 10TH INTERNATIONAL CONFERENCE ON DIGITAL LIBRARIES FOR MUSICOLOGY, DLFM 2023
페이지
7 ~ 11