← Search

Shivam Chandhok

6 accepted papers

2026

Learning What Matters: Prioritized Concept Learning via Relative Error-driven Sample Selection

CVPR 2026

Instruction tuning has been central to the success of recent vision-language models (VLMs), but it remains expensive-requiring large-scale datasets, high-quality annotations, and large compute budgets. We propose PRioritized cOncept learninG via Relative Error-driven Sample Selection (PROGRESS), a d

Cited by 0SourceScholar
2025

MM-R3: On (In-)Consistency of Vision-Language Models (VLMs)

ACL 2025finding

With the advent of LLMs and variants, a flurry of research has emerged, analyzing the performance of such models across an array of tasks. While most studies focus on evaluating the capabilities of state-of-the-art (SoTA) Vision Language Models (VLMs) through task accuracy (e.g., visual question ans…

Cited by 0SourcePDFScholar
2025

Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities

ACL 2025long

Vision-language Models (VLMs) have emerged as general-purpose tools for addressing a variety of complex computer vision problems. Such models have been shown to be highly capable, but, at the same time, lacking some basic visual understanding skills. In this paper, we set out to understand the limit…

Cited by 0SourcePDFScholar
2024

HIT: Estimating Internal Human Implicit Tissues from the Body Surface

CVPR 2024poster

The creation of personalized anatomical digital twins is important in the fields of medicine computer graphics sports science and biomechanics. To observe a subject's anatomy expensive medical devices (MRI or CT) are required and the creation of the digital model is often time-consuming and involves…

Cited by 3SourcePDFScholar
2024

Talk2BEV: Language-enhanced Bird’s-eye View Maps for Autonomous Driving

ICRA 2024poster

This work introduces Talk2BEV, a large vision-language model (LVLM)1 interface for bird’s-eye view (BEV) maps commonly used in autonomous driving. While existing perception systems for autonomous driving scenarios have largely focused on a pre-defined (closed) set of object categories and driving sc…

Cited by 77SourcecodeScholar
2022

Unseen Classes at a Later Time? No Problem

CVPR 2022poster

Recent progress towards learning from limited supervision has encouraged efforts towards designing models that can recognize novel classes at test time (generalized zero-shot learning or GZSL). GZSL approaches assume knowledge of all classes, with or without labeled data, before-hand. However, pract…

Cited by 13PDFcodeScholar