← Search

Sarthak Mehrotra

4 accepted papers

2026

CLIPoint3D: Language-Grounded Few-Shot Unsupervised 3D Point Cloud Domain Adaptation

CVPR 2026

Recent vision-language models (VLMs) such as CLIP demonstrate impressive cross-modal reasoning, extending beyond images to 3D perception. Yet, these models remain fragile under domain shifts, especially when adapting from synthetic to real-world point clouds. Conventional 3D domain adaptation approa

Cited by 0SourcecodeScholar
2026

iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive Perception

CVPR 2026

Multimodal Large Language Models (MLLMs) show strong potential for interpreting and interacting with complex, pixel-rich Graphical User Interface (GUI) environments. However, building agents that are both efficient for high-level tasks and precise for fine-grained interactions remains challenging. G

Cited by 0SourceScholar
2025

FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language Models

ICCV 2025poster

In federated learning, textual prompt tuning adapts Vision-Language Models (e.g., CLIP) by tuning lightweight input tokens (or prompts) on local client data, while keeping network weights frozen. After training, only the prompts are shared by the clients with the central server for aggregation. Howe…

2025

When Domain Generalization meets Generalized Category Discovery: An Adaptive Task-Arithmetic Driven Approach

CVPR 2025poster

Generalized Class Discovery (GCD) clusters base and novel classes in a target domain, using supervision from a source domain with only base classes. Current methods often falter with distribution shifts and typically require access to target data during training, which can sometimes be impractical.…

Cited by 1SourcePDFScholar