2021
Confidence-aware Non-repetitive Multimodal Transformers for TextCaps
AAAI 2021technical
When describing an image, reading text in the visual scene is crucial to understand the key information. Recent work explores the TextCaps task, i.e. image captioning with reading Optical Character Recognition (OCR) tokens, which requires models to read text and cover them in generated captions. Exi…