← Search

Ankit Jha

4 accepted papers

2025

FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language Models

ICCV 2025poster

In federated learning, textual prompt tuning adapts Vision-Language Models (e.g., CLIP) by tuning lightweight input tokens (or prompts) on local client data, while keeping network weights frozen. After training, only the prompts are shared by the clients with the central server for aggregation. Howe…

2025

OSLoPrompt: Bridging Low-Supervision Challenges and Open-Set Domain Generalization in CLIP

CVPR 2025poster

We introduce Low-Shot Open-Set Domain Generalization (LSOSDG), a novel paradigm unifying low-shot learning with open-set domain generalization (ODG). While prompt-based methods using models like CLIP have advanced DG, they falter in low-data regimes (e.g., 1-shot) and lack precision in detecting ope…

2024

Elevating All Zero-Shot Sketch-Based Image Retrieval Through Multimodal Prompt Learning

ECCV 2024poster

"We address the challenges inherent in sketch-based image retrieval (SBIR) across various settings, including zero-shot SBIR, generalized zero-shot SBIR, and fine-grained zero-shot SBIR, by leveraging the vision-language foundation model CLIP. While recent endeavors have employed CLIP to enhance SBI…

2024

Unknown Prompt the only Lacuna: Unveiling CLIP's Potential for Open Domain Generalization

CVPR 2024poster

We delve into Open Domain Generalization (ODG) marked by domain and category shifts between training's labeled source and testing's unlabeled target domains. Existing solutions to ODG face limitations due to constrained generalizations of traditional CNN backbones and errors in detecting target open…