2023
Spatial Cross-Attention for Transformer-Based Image Captioning
ICASSP 2023accepted
Transformer-based networks have achieved great success in image captioning because of the attention mechanism that finds relevant image locations for each word. However, the current cross-attention process, which aligns word-to-image, does not consider the spatial relationships existing in patch-to-…