2021
Cascade Attention Fusion for Fine-Grained Image Captioning Based on Multi-Layer LSTM
ICASSP 2021accepted
The conventional visual attention-based image captioning approaches typically use image information to guide caption generation. Results from these models tend to be coarse and ignore the details in the image, such as objects, attributes and the distinguishing aspects of each image. In this paper, w…