2023
RefEgo: Referring Expression Comprehension Dataset from First-Person Perception of Ego4D
ICCV 2023poster
Grounding textual expressions on scene objects from first-person views is a truly demanding capability in developing agents that are aware of their surroundings and behave following intuitive text instructions. Such capability is of necessity for glass-devices or autonomous robots to localize referr…