ICASSP 2017accepted0 citations
Visual features for context-aware speech recognition
Abhinav Gupta, Yajie Miao, Leonardo Neves, Florian Metze
Abstract
Automatic transcriptions of consumer generated multi-media content such as “Youtube” videos still exhibit high word error rates. Such data typically occupies a very broad domain, has been recorded in challenging conditions, with cheap hardware and a focus on the visual modality, and may have been post-processed or edited.
BibTeX
@inproceedings{icassp2017_visualfeaturesfo,
title = {Visual features for context-aware speech recognition},
author = {Abhinav Gupta and Yajie Miao and Leonardo Neves and Florian Metze},
booktitle = {ICASSP 2017},
year = {2017}
}