← Search

Ajay Divakaran

12 accepted papers

2025

Punching Bag vs. Punching Person: Motion Transferability in Videos

ICCV 2025poster

Action recognition models demonstrate strong generalization, but can they effectively transfer high-level motion concepts across diverse contexts, even within similar distributions? For example, can a model recognize the broad action "punching" when presented with an unseen variation such as "punchi…

2024

BloomVQA: Assessing Hierarchical Multi-modal Comprehension

ACL 2024findings

We propose a novel VQA dataset, BloomVQA, to facilitate comprehensive evaluation of large vision-language models on comprehension tasks. Unlike current benchmarks that often focus on fact-based memorization and simple reasoning tasks without theoretical grounding, we collect multiple-choice samples…

Cited by 0SourcePDFScholar
2024

DRESS: Instructing Large Vision-Language Models to Align and Interact with Humans via Natural Language Feedback

CVPR 2024poster

We present DRESS a large vision language model (LVLM) that innovatively exploits Natural Language feedback (NLF) from Large Language Models to enhance its alignment and interactions by addressing two key limitations in the state-of-the-art LVLMs. First prior LVLMs generally rely only on the instruct…

Cited by 68SourcePDFScholar
2024

Demonstrations Are All You Need: Advancing Offensive Content Paraphrasing using In-Context Learning

ACL 2024findings

Paraphrasing of offensive content is a better alternative to content removal and helps improve civility in a communication environment. Supervised paraphrasers; however, rely heavily on large quantities of labelled data to help preserve meaning and intent. They also often retain a large portion of t…

2024

Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models

NAACL 2024long

Vision-language models (VLMs) have recently demonstrated strong efficacy as visual assistants that can parse natural queries about the visual content and generate human-like outputs. In this work, we explore the ability of these models to demonstrate human-like reasoning based on the perceived infor…

2024

Pelican: Correcting Hallucination in Vision-LLMs via Claim Decomposition and Program of Thought Verification

EMNLP 2024main

Large Visual Language Models (LVLMs) struggle with hallucinations in visual instruction following task(s). These issues hinder their trustworthiness and real-world applicability. We propose Pelican – a novel framework designed to detect and mitigate hallucinations through claim verification. Pelican…

2023

Class Prototypes Based Contrastive Learning for Classifying Multi-Label and Fine-Grained Educational Videos

CVPR 2023poster

The recent growth in the consumption of online media by children during early childhood necessitates data-driven tools enabling educators to filter out appropriate educational content for young learners. This paper presents an approach for detecting educational content in online videos. We focus on…

2023

TIJO: Trigger Inversion with Joint Optimization for Defending Multimodal Backdoored Models

ICCV 2023oral

We present a Multimodal Backdoor defense technique TIJO (Trigger Inversion using Joint Optimization). Recently Walmer et al. demonstrated successful backdoor attacks on multimodal models for the Visual Question Answering task. Their dual-key backdoor trigger is split across two modalities (image and…

Cited by 12PDFcodeScholar
2022

Detecting Out-Of-Context Objects Using Graph Contextual Reasoning Network

IJCAI 2022poster

This paper presents an approach for detecting out-of-context (OOC) objects in images. Given an image with a set of objects, our goal is to determine if an object is inconsistent with the contextual relations and detect the OOC object with a bounding box. In this work, we consider common contextual r…

Cited by 13SourcePDFScholar
2021

Confidence Calibration for Domain Generalization Under Covariate Shift

ICCV 2021poster

Existing calibration algorithms address the problem of covariate shift via unsupervised domain adaptation. However, these methods suffer from the following limitations: 1) they require unlabeled data from the target domain, which may not be available at the stage of calibration in real-world applica…

Cited by 34PDFScholar
2019

Align2Ground: Weakly Supervised Phrase Grounding Guided by Image-Caption Alignment

ICCV 2019poster

We address the problem of grounding free-form textual phrases by using weak supervision from image-caption pairs. We propose a novel end-to-end model that uses caption-to-image retrieval as a downstream task to guide the process of phrase localization. Our method, as a first step, infers the latent…

Cited by 119PDFScholar