ICASSP 2017accepted0 citations

Visual features for context-aware speech recognition

Abhinav Gupta, Yajie Miao, Leonardo Neves, Florian Metze

Abstract

Automatic transcriptions of consumer generated multi-media content such as “Youtube” videos still exhibit high word error rates. Such data typically occupies a very broad domain, has been recorded in challenging conditions, with cheap hardware and a focus on the visual modality, and may have been post-processed or edited.

BibTeX
@inproceedings{icassp2017_visualfeaturesfo,
  title = {Visual features for context-aware speech recognition},
  author = {Abhinav Gupta and Yajie Miao and Leonardo Neves and Florian Metze},
  booktitle = {ICASSP 2017},
  year = {2017}
}