← Search

Yaxian Xia

2 accepted papers

2020

Exploring Entity-Level Spatial Relationships for Image-Text Matching

ICASSP 2020accepted

Exploring the entity-level (i.e., objects in an image, words in a text) spatial relationship contributes to understanding multimedia content precisely. The ignorance of spatial information in previous works probably leads to misunderstandings of image contents. For instance, sentences `Boats are on…

Cited by 0SourceScholar
2019

Adaptively Aligned Image Captioning via Adaptive Attention Time

NeurIPS 2019poster

Recent neural models for image captioning usually employ an encoder-decoder framework with an attention mechanism. However, the attention mechanism in such a framework aligns one single (attended) image feature vector to one caption word, assuming one-to-one mapping from source image regions and tar…