← Search

Mainak Singha

5 accepted papers

2026

Bi-Modal Textual Prompt Learning for Vision-Language Models in Remote Sensing

ICASSP 2026poster

Prompt learning (PL) has emerged as an effective strategy to adapt vision-language models (VLMs), such as CLIP, for downstream tasks under limited supervision. While PL has demonstrated strong generalization on natural image datasets, its transferability to remote sensing (RS) imagery remains undere…

Cited by 0SourcePDFScholar
2026

CLIPoint3D: Language-Grounded Few-Shot Unsupervised 3D Point Cloud Domain Adaptation

CVPR 2026

Recent vision-language models (VLMs) such as CLIP demonstrate impressive cross-modal reasoning, extending beyond images to 3D perception. Yet, these models remain fragile under domain shifts, especially when adapting from synthetic to real-world point clouds. Conventional 3D domain adaptation approa

Cited by 0SourcecodeScholar
2025

FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language Models

ICCV 2025poster

In federated learning, textual prompt tuning adapts Vision-Language Models (e.g., CLIP) by tuning lightweight input tokens (or prompts) on local client data, while keeping network weights frozen. After training, only the prompts are shared by the clients with the central server for aggregation. Howe…

2025

OSLoPrompt: Bridging Low-Supervision Challenges and Open-Set Domain Generalization in CLIP

CVPR 2025poster

We introduce Low-Shot Open-Set Domain Generalization (LSOSDG), a novel paradigm unifying low-shot learning with open-set domain generalization (ODG). While prompt-based methods using models like CLIP have advanced DG, they falter in low-data regimes (e.g., 1-shot) and lack precision in detecting ope…

2024

Unknown Prompt the only Lacuna: Unveiling CLIP's Potential for Open Domain Generalization

CVPR 2024poster

We delve into Open Domain Generalization (ODG) marked by domain and category shifts between training's labeled source and testing's unlabeled target domains. Existing solutions to ODG face limitations due to constrained generalizations of traditional CNN backbones and errors in detecting target open…