2016
Training and Evaluating Multimodal Word Embeddings with Large-scale Web Annotated Images
NeurIPS 2016poster
In this paper, we focus on training and evaluating effective word embeddings with both text and visual information. More specifically, we introduce a large-scale dataset with 300 million sentences describing over 40 million images crawled and downloaded from publicly available Pins (i.e. an image wi…