Prompting Vision-Language Models For Aspect-Controlled Generation of Referring Expressions
Referring Expression Generation (REG) is the task of generating a description that unambiguously identifies a given target in the scene. Different from Image Captioning (IC), REG requires learning fine-grained characteristics of not only the scene objects but also their surrounding context. Referrin…