2018
Jointly Discovering Visual Objects and Spoken Words from Raw Sensory Input
ECCV 2018poster
In this paper, we explore neural network models that learn to associate segments of spoken audio captions with the semantically relevant portions of natural images that they refer to. We demonstrate that these audio-visual associative localizations emerge from network-internal representations learne…