2024
LoA-Trans: Enhancing Visual Grounding by Location-Aware Transformers
ECCV 2024poster
"Given an image and text description, visual grounding will find target region in the image explained by the text. It has two task settings: referring expression comprehension (REC) to estimate bounding-box and referring expression segmentation (RES) to predict segmentation mask. Currently the most…