ICASSP 2018accepted0 citations

Audio-Visual Person Recognition in Multimedia Data From the Iarpa Janus Program

Gregory Sell, Kevin Duh, David Snyder, Dave Etter, Daniel Garcia-Romero

Abstract

Currently, datasets that support audio-visual recognition of people in videos are scarce and limited. In this paper, we introduce an expansion of video data from the IARPA Janus program to support this research area. We refer to the expanded set, which adds labels for voice to the already-existing face labels, as the Janus Multimedia dataset. We first describe the speaker labeling process, which involved a combination of automatic and manual criteria. We then discuss two evaluation settings for this data. In the core condition, the voice and face of the labeled individual are present in every video. In the full condition, no such guarantee is made. The power of audiovisual fusion is then shown using these publicly-available videos and labels, showing significant improvement over only recognizing voice or face alone. In addition to this work, several other possible paths for future research with this dataset are discussed.

BibTeX
@inproceedings{icassp2018_audiovisualperso,
  title = {Audio-Visual Person Recognition in Multimedia Data From the Iarpa Janus Program},
  author = {Gregory Sell and Kevin Duh and David Snyder and Dave Etter and Daniel Garcia-Romero},
  booktitle = {ICASSP 2018},
  year = {2018}
}