OLKAVS: An Open Large-Scale Korean Audio-Visual Speech Dataset
Jeongkyun Park, Jung-Wook Hwang, Kwanghee Choi, Seung-Hyeon Lee, Jun Hwan Ahn, Rae-Hong Park, Hyung-Min Park
Abstract
Inspired by humans comprehending speech in a multi-modal manner, various audio-visual datasets have been constructed. However, most existing datasets focus on English, developed from pre-existing videos using various prediction models, and have only a small number of multi-view videos. To mitigate the limitations, we constructed the Open Large-scale Korean Audio-Visual Speech (OLKAVS) dataset, which is the largest among publicly available audio-visual speech datasets. The dataset contains 1,150 hours of transcribed audio from 1,107 Korean speakers in a studio setup with nine different viewpoints and various noise situations. We also provide the pre-trained baseline models for two tasks: audiovisual speech recognition and lip reading. We conducted experiments based on the models to verify the effectiveness of multi-modal and multi-view training over uni-modal and frontal-view-only training. We expect the OLKAVS dataset to facilitate multi-modal research in broader areas.
BibTeX
@inproceedings{icassp2024_olkavsanopenlarg,
title = {OLKAVS: An Open Large-Scale Korean Audio-Visual Speech Dataset},
author = {Jeongkyun Park and Jung-Wook Hwang and Kwanghee Choi and Seung-Hyeon Lee and Jun Hwan Ahn and Rae-Hong Park and Hyung-Min Park},
booktitle = {ICASSP 2024},
year = {2024}
}