← Search

Tejas Gokhale

20 accepted papers

2025

Side Effects of Erasing Concepts from Diffusion Models

EMNLP 2025

Concerns about text-to-image (T2I) generative models infringing on privacy, copyright, and safety have led to the development of concept erasure techniques (CETs). The goal of an effective CET is to prohibit the generation of undesired “target” concepts specified by the user, while preserving the ab

2025

VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning

ICLR 2025poster

Multimodal Large Language Models (MLLMs) have become a powerful tool for integrating visual and textual information. Despite their exceptional performance on visual understanding benchmarks, measuring their ability to reason abstractly across multiple images remains a significant challenge. To addre…

Cited by 0SourcePDFScholar
2024

ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion Models

AAAI 2024technical

The ability to understand visual concepts and replicate and compose these concepts from images is a central goal for computer vision. Recent advances in text-to-image (T2I) models have lead to high definition and realistic image quality generation by learning from large databases of images and their…

2024

Getting it Right: Improving Spatial Consistency in Text-to-Image Models

ECCV 2024poster

"One of the key shortcomings in current text-to-image (T2I) models is their inability to consistently generate images which faithfully follow the spatial relationships specified in the text prompt. In this paper, we offer a comprehensive investigation of this limitation, while also developing datase…

2024

On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth Estimation

CVPR 2024poster

Recent advances in monocular depth estimation have been made by incorporating natural language as additional guidance. Although yielding impressive results the impact of the language prior particularly in terms of generalization and robustness remains unexplored. In this paper we address this gap by…

2024

REVISION: Rendering Tools Enable Spatial Fidelity in Vision-Language Models

ECCV 2024poster

"Text-to-Image (T2I) and multimodal large language models (MLLMs) have been adopted in solutions for several computer vision and multimodal learning tasks. However, it has been found that such vision-language models lack the ability to correctly reason over spatial relationships. To tackle this shor…

2024

TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives

NeurIPS 2024poster

Contrastive Language-Image Pretraining (CLIP) models maximize the mutual information between text and visual modalities to learn representations. This makes the nature of the training data a significant factor in the efficacy of CLIP for downstream tasks. However, the lack of compositional diversity…

2023

End-to-end Knowledge Retrieval with Multi-modal Queries

ACL 2023long

We investigate knowledge retrieval with multi-modal queries, i.e. queries containing information split across image and text inputs, a challenging task that differs from previous work on cross-modal retrieval. We curate a new dataset called ReMuQ for benchmarking progress on this task. ReMuQ require…

2022

CRIPP-VQA: Counterfactual Reasoning about Implicit Physical Properties via Video Question Answering

EMNLP 2022main

Videos often capture objects, their visible properties, their motion, and the interactions between different objects. Objects also have physical properties such as mass, which the imaging pipeline is unable to directly capture. However, these properties can be estimated by utilizing cues from relati…

2022

Generalized but not Robust? Comparing the Effects of Data Modification Methods on Out-of-Domain Generalization and Adversarial Robustness

ACL 2022findings

Data modification, either via additional training datasets, data augmentation, debiasing, and dataset filtering, has been proposed as an effective solution for generalizing to out-of-domain (OOD) inputs, in both natural language processing and computer vision literature. However, the effect of data…

2022

Improving Biomedical Information Retrieval with Neural Retrievers

AAAI 2022technical

Information retrieval (IR) is essential in search engines and dialogue systems as well as natural language processing tasks such as open-domain question answering. IR serve an important function in the biomedical domain, where content and sources of scientific knowledge may evolve rapidly. Although…

2022

Semantically Distributed Robust Optimization for Vision-and-Language Inference

ACL 2022findings

Analysis of vision-and-language models has revealed their brittleness under linguistic phenomena such as paraphrasing, negation, textual entailment, and word substitutions with synonyms or antonyms. While data augmentation techniques have been designed to mitigate against these failure modes, method…

2022

To Find Waldo You Need Contextual Cues: Debiasing Who’s Waldo

ACL 2022short

We present a debiased dataset for the Person-centric Visual Grounding (PCVG) task first proposed by Cui et al. (2021) in the Who’s Waldo dataset. Given an image and a caption, PCVG requires pairing up a person’s name mentioned in a caption with a bounding box that points to the person in the image.…

2022

Unsupervised Natural Language Inference Using PHL Triplet Generation

ACL 2022findings

Transformer-based models achieve impressive performance on numerous Natural Language Inference (NLI) benchmarks when trained on respective training datasets. However, in certain cases, training samples may not be available or collecting them could be time-consuming and resource-intensive. In this wo…

2021

Attribute-Guided Adversarial Training for Robustness to Natural Perturbations

AAAI 2021technical

While existing work in robust deep learning has focused on small pixel-level norm-based perturbations, this may not account for perturbations encountered in several real world settings. In many such cases although test data might not be available, broad specifications about the types of perturbation…

2021

Weakly Supervised Relative Spatial Reasoning for Visual Question Answering

ICCV 2021poster

Vision-and-language (V&L) reasoning necessitates perception of visual concepts such as objects and actions, understanding semantics and language grounding, and reasoning about the interplay between the two modalities. One crucial aspect of visual reasoning is spatial understanding, which involves un…

Cited by 26PDFScholar
2020

VQA-LOL: Visual Question Answering under the Lens of Logic

ECCV 2020poster

Logical connectives and their implications on the meaning of a natural language sentence are a fundamental aspect of understanding. In this paper, we investigate whether visual question answering (VQA) systems trained to answer a question about an image, are able to answer the logical composition of…

Cited by 103SourcePDFScholar