← Search

Kanishk Jain

6 accepted papers

2026

Learning What Matters: Prioritized Concept Learning via Relative Error-driven Sample Selection

CVPR 2026

Instruction tuning has been central to the success of recent vision-language models (VLMs), but it remains expensive-requiring large-scale datasets, high-quality annotations, and large compute budgets. We propose PRioritized cOncept learninG via Relative Error-driven Sample Selection (PROGRESS), a d

Cited by 0SourceScholar
2024

Benchmarking Vision Language Models for Cultural Understanding

EMNLP 2024main

Foundation models and vision-language pre-training have notably advanced Vision Language Models (VLMs), enabling multimodal processing of visual and linguistic data. However, their performance has been typically assessed on general scene understanding - recognizing objects, attributes, and actions -…

Cited by 24SourcePDFScholar
2023

Ground then Navigate: Language-guided Navigation in Dynamic Scenes

ICRA 2023poster

We investigate the Vision-and-Language Navigation (VLN) problem in the context of autonomous driving in outdoor settings. We solve the problem by explicitly grounding the navigable regions corresponding to the textual command. At each timestamp, the model predicts a segmentation mask corresponding t…

Cited by 27SourcecodeScholar
2023

Test-Time Amendment with a Coarse Classifier for Fine-Grained Classification

NeurIPS 2023poster

We investigate the problem of reducing mistake severity for fine-grained classification. Fine-grained classification can be challenging, mainly due to the requirement of knowledge or domain expertise for accurate annotation. However, humans are particularly adept at performing coarse classification…

2021

Grounding Linguistic Commands to Navigable Regions

IROS 2021poster

Humans have a natural ability to effortlessly comprehend linguistic commands such as “park next to the yellow sedan” and instinctively know which region of the road the vehicle should navigate. Extending this ability to autonomous vehicles is the next step towards creating fully autonomous agents th…

Cited by 14SourcecodeScholar