← Search

Sumit Shekhar

8 accepted papers

2025

Imposter: Text and Frequency Guidance for Subject Driven Action Personalization using Diffusion Models

COLING 2025main

We present ImPoster, a novel algorithm for generating a target image of a ‘source’ subject performing a ‘driving’ action. The inputs to our algorithm are a single pair of a source image with the subject that we wish to edit and a driving image with a subject of an arbitrary class performing the driv…

2025

Looking Beyond the Pixels: Evaluating Visual Metaphor Understanding in VLMs

EMNLP 2025

Visual metaphors are a complex vision–language phenomenon that requires both perceptual and conceptual reasoning to understand. They provide a valuable test of a model’s ability to interpret visual input and reason about it with creativity and coherence. We introduce ImageMet, a visual metaphor data

2024

Unveiling the Invisible: Captioning Videos with Metaphors

EMNLP 2024finding

Metaphors are a common communication tool used in our day-to-day life. The detection and generation of metaphors in textual form have been studied extensively but metaphors in other forms have been under-explored. Recent studies have shown that Vision-Language (VL) models cannot understand visual me…

2023

Open-World Factually Consistent Question Generation

ACL 2023findings

Question generation methods based on pre-trained language models often suffer from factual inconsistencies and incorrect entities and are not answerable from the input paragraph. Domain shift – where the test data is from a different domain than the training data - further exacerbates the problem of…

Cited by 4SourcePDFScholar
2023

“Let’s not Quote out of Context”: Unified Vision-Language Pretraining for Context Assisted Image Captioning

ACL 2023industry

Well-formed context aware image captions and tags in enterprise content such as marketing material are critical to ensure their brand presence and content recall. Manual creation and updates to ensure the same is non trivial given the scale and the tedium towards this task. We propose a new unified…

Cited by 8SourcePDFScholar
2022

DynamicTOC: Persona-based Table of Contents for Consumption of Long Documents

NAACL 2022long

Long documents like contracts, financial documents, etc., are often tedious to read through. Linearly consuming (via scrolling or navigation through default table of content) these documents is time-consuming and challenging. These documents are also authored to be consumed by varied entities (refer…

Cited by 2SourcePDFScholar
2022

TALISMAN: Targeted Active Learning for Object Detection with Rare Classes and Slices Using Submodular Mutual Information

ECCV 2022poster

"Deep neural networks based object detectors have shown great success in a variety of domains like autonomous vehicles, biomedical imaging, etc. It is known that their success depends on a large amount of data from the domain of interest. While deep models often perform well in terms of overall accu…

2015

Class Consistent Multi-Modal Fusion With Binary Features

CVPR 2015poster

Many existing recognition algorithms combine different modalities based on training accuracy but do not consider the possibility of noise at test time. We describe an algorithm that perturbs test features so that all modalities predict the same class. We enforce this perturbation to be as small as p…

Cited by 15SourcePDFScholar