← Search

Pratyay Banerjee

11 accepted papers

2024

Towards Improved Multi-Source Attribution for Long-Form Answer Generation

NAACL 2024long

Teaching large language models (LLMs) to generate text with attribution to evidence sources can reduce hallucinations, improve verifiability in question answering systems (QA), and increase reliability of retrieval augmented LLMs. Despite gaining increasing popularity for usage in QA systems and sea…

Cited by 2SourcePDFScholar
2022

Learning Action-Effect Dynamics for Hypothetical Vision-Language Reasoning Task

EMNLP 2022finding

‘Actions’ play a vital role in how humans interact with the world. Thus, autonomous agents that would assist us in everyday tasks also require the capability to perform ‘Reasoning about Actions & Change’ (RAC). This has been an important research direction in Artificial Intelligence (AI) in general,…

2022

Lexi: Self-Supervised Learning of the UI Language

EMNLP 2022finding

Humans can learn to operate the user interface (UI) of an application by reading an instruction manual or how-to guide. Along with text, these resources include visual content such as UI screenshots and images of application icons referenced in the text. We explore how to leverage this data to learn…

2022

Semantically Distributed Robust Optimization for Vision-and-Language Inference

ACL 2022findings

Analysis of vision-and-language models has revealed their brittleness under linguistic phenomena such as paraphrasing, negation, textual entailment, and word substitutions with synonyms or antonyms. While data augmentation techniques have been designed to mitigate against these failure modes, method…

2022

To Find Waldo You Need Contextual Cues: Debiasing Who’s Waldo

ACL 2022short

We present a debiased dataset for the Person-centric Visual Grounding (PCVG) task first proposed by Cui et al. (2021) in the Who’s Waldo dataset. Given an image and a caption, PCVG requires pairing up a person’s name mentioned in a caption with a bounding box that points to the person in the image.…

2022

Unsupervised Natural Language Inference Using PHL Triplet Generation

ACL 2022findings

Transformer-based models achieve impressive performance on numerous Natural Language Inference (NLI) benchmarks when trained on respective training datasets. However, in certain cases, training samples may not be available or collecting them could be time-consuming and resource-intensive. In this wo…

2021

Weakly Supervised Relative Spatial Reasoning for Visual Question Answering

ICCV 2021poster

Vision-and-language (V&L) reasoning necessitates perception of visual concepts such as objects and actions, understanding semantics and language grounding, and reasoning about the interplay between the two modalities. One crucial aspect of visual reasoning is spatial understanding, which involves un…

Cited by 26PDFScholar
2021

Weakly-Supervised Visual-Retriever-Reader for Knowledge-based Question Answering

EMNLP 2021main

Knowledge-based visual question answering (VQA) requires answering questions with external knowledge in addition to the content of images. One dataset that is mostly used in evaluating knowledge-based VQA is OK-VQA, but it lacks a gold standard knowledge corpus for retrieval. Existing work leverage…

2020

VQA-LOL: Visual Question Answering under the Lens of Logic

ECCV 2020poster

Logical connectives and their implications on the meaning of a natural language sentence are a fundamental aspect of understanding. In this paper, we investigate whether visual question answering (VQA) systems trained to answer a question about an image, are able to answer the logical composition of…

Cited by 103SourcePDFScholar