OLKAVS: AN OPEN LARGE-SCALE KOREAN AUDIO-VISUAL SPEECH DATASET

  • Park, Jeongkyun
  • Hwang, Jung-Wook
  • Choi, Kwanghee
  • Lee, Seung-Hyeon
  • Ahn, Jun Hwan
  • ... Park, Hyung-Min
  • 외 1명
Citations

WEB OF SCIENCE

0
Citations

SCOPUS

1

초록

Inspired by humans comprehending speech in a multi-modal manner, various audio-visual datasets have been constructed. However, most existing datasets focus on English, developed from pre-existing videos using various prediction models, and have only a small number of multi-view videos. To mitigate the limitations, we constructed the Open Large-scale Korean Audio-Visual Speech (OLKAVS) dataset, which is the largest among publicly available audio-visual speech datasets. The dataset contains 1,150 hours of transcribed audio from 1,107 Korean speakers in a studio setup with nine different viewpoints and various noise situations. We also provide the pre-trained baseline models for two tasks: audiovisual speech recognition and lip reading. We conducted experiments based on the models to verify the effectiveness of multi-modal and multi-view training over uni-modal and frontal-view-only training. We expect the OLKAVS dataset to facilitate multi-modal research in broader areas.

키워드

Audio-visual speech datasetsmulti-view datasetslip readingaudio-visual speech recognitiondeep learningRECOGNITIONDATABASE
제목
OLKAVS: AN OPEN LARGE-SCALE KOREAN AUDIO-VISUAL SPEECH DATASET
저자
Park, JeongkyunHwang, Jung-WookChoi, KwangheeLee, Seung-HyeonAhn, Jun HwanPark, Rae-HongPark, Hyung-Min
DOI
10.1109/ICASSP48485.2024.10446901
발행일
2024
유형
Proceedings Paper
저널명
ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings
페이지
6385 ~ 6389