← Search

Philipp J. Rösch

1 accepted papers

2022

Probing the Role of Positional Information in Vision-Language Models

NAACL 2022findings

In most Vision-Language models (VL), the understanding of the image structure is enabled by injecting the position information (PI) about objects in the image. In our case study of LXMERT, a state-of-the-art VL model, we probe the use of the PI in the representation and study its effect on Visual Qu…