상세 보기
OLKAVS: AN OPEN LARGE-SCALE KOREAN AUDIO-VISUAL SPEECH DATASET
- Park, Jeongkyun;
- Hwang, Jung-Wook;
- Choi, Kwanghee;
- Lee, Seung-Hyeon;
- Ahn, Jun Hwan;
- ... Park, Hyung-Min;
- 외 1명
WEB OF SCIENCE
0SCOPUS
1초록
Inspired by humans comprehending speech in a multi-modal manner, various audio-visual datasets have been constructed. However, most existing datasets focus on English, developed from pre-existing videos using various prediction models, and have only a small number of multi-view videos. To mitigate the limitations, we constructed the Open Large-scale Korean Audio-Visual Speech (OLKAVS) dataset, which is the largest among publicly available audio-visual speech datasets. The dataset contains 1,150 hours of transcribed audio from 1,107 Korean speakers in a studio setup with nine different viewpoints and various noise situations. We also provide the pre-trained baseline models for two tasks: audiovisual speech recognition and lip reading. We conducted experiments based on the models to verify the effectiveness of multi-modal and multi-view training over uni-modal and frontal-view-only training. We expect the OLKAVS dataset to facilitate multi-modal research in broader areas.
키워드
- 제목
- OLKAVS: AN OPEN LARGE-SCALE KOREAN AUDIO-VISUAL SPEECH DATASET
- 저자
- Park, Jeongkyun; Hwang, Jung-Wook; Choi, Kwanghee; Lee, Seung-Hyeon; Ahn, Jun Hwan; Park, Rae-Hong; Park, Hyung-Min
- 발행일
- 2024
- 유형
- Proceedings Paper
- 페이지
- 6385 ~ 6389