Reducing Language Biases in Visual Question Answering with Visually-Grounded Question Encoder
Recent studies have shown that current VQA models are heavily biased on the language priors in the train set to answer the question, irrespective of the image. E.g., overwhelmingly answer ""what sport is"" as ""tennis"" or ""what color banana"" as ""yellow."" This behavior restricts them from real-w…