Cross-modality matching based on Fisher Vector with neural word embeddings and deep image features
Liang Han, Wenmin Wang, Mengdi Fan, Ronggang Wang
Abstract
Cross-modal retrieval, which aims to solve the problem that the query and the retrieved results are from different modality, becomes more and more essential with the development of the Internet. In this paper, we mainly focus on the exploration of high-level semantic representation of image and text for cross-modal matching. Deep convolutional image features and Fisher Vector with neural word embeddings are utilized as visual and textual features respectively. To further investigate the correlation among heterogeneous multimodal characteristics, we use multiclass logistic classifier for semantic matching across modalities. Experiments on Wikipedia and Pascal Sentence dataset demonstrate the robustness and effectiveness for both Img2Text and Text2Img retrieval tasks.
BibTeX
@inproceedings{icassp2017_crossmodalitymat,
title = {Cross-modality matching based on Fisher Vector with neural word embeddings and deep image features},
author = {Liang Han and Wenmin Wang and Mengdi Fan and Ronggang Wang},
booktitle = {ICASSP 2017},
year = {2017}
}