← Search

Karan Sikka

10 accepted papers

2024

DRESS: Instructing Large Vision-Language Models to Align and Interact with Humans via Natural Language Feedback

CVPR 2024poster

We present DRESS a large vision language model (LVLM) that innovatively exploits Natural Language feedback (NLF) from Large Language Models to enhance its alignment and interactions by addressing two key limitations in the state-of-the-art LVLMs. First prior LVLMs generally rely only on the instruct…

Cited by 68SourcePDFScholar
2024

Demonstrations Are All You Need: Advancing Offensive Content Paraphrasing using In-Context Learning

ACL 2024findings

Paraphrasing of offensive content is a better alternative to content removal and helps improve civility in a communication environment. Supervised paraphrasers; however, rely heavily on large quantities of labelled data to help preserve meaning and intent. They also often retain a large portion of t…

2024

Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models

NAACL 2024long

Vision-language models (VLMs) have recently demonstrated strong efficacy as visual assistants that can parse natural queries about the visual content and generate human-like outputs. In this work, we explore the ability of these models to demonstrate human-like reasoning based on the perceived infor…

2024

Pelican: Correcting Hallucination in Vision-LLMs via Claim Decomposition and Program of Thought Verification

EMNLP 2024main

Large Visual Language Models (LVLMs) struggle with hallucinations in visual instruction following task(s). These issues hinder their trustworthiness and real-world applicability. We propose Pelican – a novel framework designed to detect and mitigate hallucinations through claim verification. Pelican…

2023

TIJO: Trigger Inversion with Joint Optimization for Defending Multimodal Backdoored Models

ICCV 2023oral

We present a Multimodal Backdoor defense technique TIJO (Trigger Inversion using Joint Optimization). Recently Walmer et al. demonstrated successful backdoor attacks on multimodal models for the Visual Question Answering task. Their dual-key backdoor trigger is split across two modalities (image and…

Cited by 12PDFcodeScholar
2022

Dual-Key Multimodal Backdoors for Visual Question Answering

CVPR 2022poster

The success of deep learning has enabled advances in multimodal tasks that require non-trivial fusion of multiple input domains. Although multimodal models have shown potential in many problems, their increased complexity makes them more vulnerable to attacks. A Backdoor (or Trojan) attack is a clas…

Cited by 52PDFcodeScholar
2019

Align2Ground: Weakly Supervised Phrase Grounding Guided by Image-Caption Alignment

ICCV 2019poster

We address the problem of grounding free-form textual phrases by using weak supervision from image-caption pairs. We propose a novel end-to-end model that uses caption-to-image retrieval as a downstream task to guide the process of phrase localization. Our method, as a first step, infers the latent…

Cited by 119PDFScholar
2017

AdaScan: Adaptive Scan Pooling in Deep Convolutional Neural Networks for Human Action Recognition in Videos

CVPR 2017poster

We propose a novel method for temporally pooling frames in a video for the task of human action recognition. The method is motivated by the observation that there are only a small number of frames which, together, contain sufficient information to discriminate an action class present in a video, fro…

Cited by 201PDFcodeScholar