2019
Multi-Level Multimodal Common Semantic Space for Image-Phrase Grounding
CVPR 2019poster
We address the problem of phrase grounding by learning a multi-level common semantic space shared by the textual and visual modalities. We exploit multiple levels of feature maps of a Deep Convolutional Neural Network, as well as contextualized word and sentence embeddings extracted from a character…