← Search

Danfeng Guo

2 accepted papers

2024

Prompting Vision-Language Models For Aspect-Controlled Generation of Referring Expressions

NAACL 2024findings

Referring Expression Generation (REG) is the task of generating a description that unambiguously identifies a given target in the scene. Different from Image Captioning (IC), REG requires learning fine-grained characteristics of not only the scene objects but also their surrounding context. Referrin…

Cited by 0SourcePDFScholar
2022

GRAVL-BERT: Graphical Visual-Linguistic Representations for Multimodal Coreference Resolution

COLING 2022main

Learning from multimodal data has become a popular research topic in recent years. Multimodal coreference resolution (MCR) is an important task in this area. MCR involves resolving the references across different modalities, e.g., text and images, which is a crucial capability for building next-gene…