← Search

Biplab Banerjee

26 accepted papers

2026

Bi-Modal Textual Prompt Learning for Vision-Language Models in Remote Sensing

ICASSP 2026poster

Prompt learning (PL) has emerged as an effective strategy to adapt vision-language models (VLMs), such as CLIP, for downstream tasks under limited supervision. While PL has demonstrated strong generalization on natural image datasets, its transferability to remote sensing (RS) imagery remains undere…

Cited by 0SourcePDFScholar
2026

CLIPoint3D: Language-Grounded Few-Shot Unsupervised 3D Point Cloud Domain Adaptation

CVPR 2026

Recent vision-language models (VLMs) such as CLIP demonstrate impressive cross-modal reasoning, extending beyond images to 3D perception. Yet, these models remain fragile under domain shifts, especially when adapting from synthetic to real-world point clouds. Conventional 3D domain adaptation approa

Cited by 0SourcecodeScholar
2026

Hyperbolic Prototype Learning with Uncertainty-Aware Consistency for Continual Test-Time Segmentation

CVPR 2026

Continual Test-Time Adaptation (CTTA) for semantic segmentation is vital for deploying vision models in dynamic environments with persistent domain shifts. Existing methods often degrade over time as self-supervised updates amplify early prediction errors. We attribute this fragility to a geometric

Cited by 0SourceScholar
2025

DuET: Dual Incremental Object Detection via Exemplar-Free Task Arithmetic

ICCV 2025poster

Real-world object detection systems, such as those in autonomous driving and surveillance, must continuously learn new object categories and simultaneously adapt to changing environmental conditions. Existing approaches, Class Incremental Object Detection (CIOD) and Domain Incremental Object Detecti…

Cited by 0SourcePDFScholar
2025

FedMVP: Federated Multimodal Visual Prompt Tuning for Vision-Language Models

ICCV 2025poster

In federated learning, textual prompt tuning adapts Vision-Language Models (e.g., CLIP) by tuning lightweight input tokens (or prompts) on local client data, while keeping network weights frozen. After training, only the prompts are shared by the clients with the central server for aggregation. Howe…

2025

HIDISC: A Hyperbolic Framework for Domain Generalization with Generalized Category Discovery

NeurIPS 2025poster

Generalized Category Discovery (GCD) aims to classify test-time samples into either seen categories—available during training—or novel ones, without relying on label supervision. Most existing GCD methods assume simultaneous access to labeled and unlabeled data during training and arising from the s…

Cited by 0SourceScholar
2025

Hyperbolic Uncertainty-Aware Few-Shot Incremental Point Cloud Segmentation

CVPR 2025poster

3D point cloud segmentation is essential across a range of applications; however, conventional methods often struggle in evolving environments, particularly when tasked with identifying novel categories under limited supervision. Few-Shot Learning (FSL) and Class Incremental Learning (CIL) have been…

Cited by 0SourcePDFScholar
2025

OSLoPrompt: Bridging Low-Supervision Challenges and Open-Set Domain Generalization in CLIP

CVPR 2025poster

We introduce Low-Shot Open-Set Domain Generalization (LSOSDG), a novel paradigm unifying low-shot learning with open-set domain generalization (ODG). While prompt-based methods using models like CLIP have advanced DG, they falter in low-data regimes (e.g., 1-shot) and lack precision in detecting ope…

2025

ReDepress: A Cognitive Framework for Detecting Depression Relapse from Social Media

EMNLP 2025

Almost 50% depression patients face the risk of going into relapse. The risk increases to 80% after the second episode of depression. Although, depression detection from social media has attained considerable attention, depression relapse detection has remained largely unexplored due to the lack of

Cited by 0SourcePDFScholar
2025

Spatially-Aware Cross-Modal Contrastive Learning for Low-Shot HSI Classification

ICASSP 2025accepted

Classifying hyperspectral images (HSI) with limited supervision is challenging due to their high dimensionality and complex spectral features, which frequently result in overfitting, especially under extremely low supervision. Existing self-supervised methods for HSI data focus predominantly on spec…

Cited by 0SourceScholar
2025

UIDAPLE: Unsupervised Incremental Domain Adaptation through Adaptive Prompt Learning

ICASSP 2025accepted

Continual learning poses significant challenges for deep neural networks, notably catastrophic forgetting, particularly when faced with shifting data distributions that compromise previously acquired knowledge. This paper tackles these issues within the Unsupervised Incremental Domain Adaptation (UI…

Cited by 0SourceScholar
2025

When Domain Generalization meets Generalized Category Discovery: An Adaptive Task-Arithmetic Driven Approach

CVPR 2025poster

Generalized Class Discovery (GCD) clusters base and novel classes in a target domain, using supervision from a source domain with only base classes. Current methods often falter with distribution shifts and typically require access to target data during training, which can sometimes be impractical.…

Cited by 1SourcePDFScholar
2025

“My life is miserable, have to sign 500 autographs everyday”: Exposing Humblebragging, the Brags in Disguise

ACL 2025finding

Humblebragging is a phenomenon in which individuals present self-promotional statements under the guise of modesty or complaints. For example, a statement like, “Ugh, I can’t believe I got promoted to lead the entire team. So stressful!”, subtly highlights an achievement while pretending to be compl…

2024

Elevating All Zero-Shot Sketch-Based Image Retrieval Through Multimodal Prompt Learning

ECCV 2024poster

"We address the challenges inherent in sketch-based image retrieval (SBIR) across various settings, including zero-shot SBIR, generalized zero-shot SBIR, and fine-grained zero-shot SBIR, by leveraging the vision-language foundation model CLIP. While recent endeavors have employed CLIP to enhance SBI…

2024

Enhancing the Domain Robustness of Self-Supervised pre-Training with Synthetic Images

ICASSP 2024accepted

We present a novel method for improving the adaptability of self-supervised (SSL) pre-trained models across different domains. Our approach uses synthetic images that are generated using an auxiliary diffusion model, namely InstructPix2Pix. More specifically, starting from a real image, we prompt th…

Cited by 0SourceScholar
2024

SPDG-Net: Semantics Preserving Domain Augmentation through Style Interpolation for Multi-Source Domain Generalization

ICASSP 2024accepted

This paper focuses on domain generalization (DG), addressing the challenge of robust classifier learning from multiple source domains for generalizing to unseen ones. DG suffers from limited source domain diversity, which may hinder model generalization. Recent studies explore domain-augmentation st…

Cited by 0SourceScholar
2024

Shape-prior Free Space-time Neural Radiance Field for 4D Semantic Reconstruction of Dynamic Scene from Sparse-View RGB Videos

IROS 2024

Many applications in Augmented/Virtual Reality or robotics require precise geometry modeling of individual elements in a dynamic scene under a sparse-view camera setup, without any prior information about their semantic labels or shapes. In our research, we introduce a 3D shape prior-free Neural Rad

Cited by 0SourceScholar
2024

TFS-NeRF: Template-Free NeRF for Semantic 3D Reconstruction of Dynamic Scene

NeurIPS 2024poster

Despite advancements in Neural Implicit models for 3D surface reconstruction, handling dynamic environments with interactions between arbitrary rigid, non-rigid, or deformable entities remains challenging. The generic reconstruction methods adaptable to such dynamic scenes often require additional i…

2024

Unknown Prompt the only Lacuna: Unveiling CLIP's Potential for Open Domain Generalization

CVPR 2024poster

We delve into Open Domain Generalization (ODG) marked by domain and category shifts between training's labeled source and testing's unlabeled target domains. Existing solutions to ODG face limitations due to constrained generalizations of traditional CNN backbones and errors in detecting target open…

2023

Domain Adaptive Few-Shot Open-Set Learning

ICCV 2023poster

Few-shot learning has made impressive strides in addressing the crucial challenges of recognizing unknown samples from novel classes in target query sets and managing visual shifts between domains. However, existing techniques fall short when it comes to identifying target outliers under domain shif…

Cited by 4PDFcodeScholar
2023

Physically Plausible 3D Human-Scene Reconstruction From Monocular RGB Image Using an Adversarial Learning Approach

RA-L 2023

Holistic 3D human-scene reconstruction is a crucial and emerging research area in robot perception. A key challenge in holistic 3D human-scene reconstruction is to generate a physically plausible 3D scene from a single monocular RGB image. The existing research mainly proposes optimization-based app

Cited by 4SourceScholar
2023

USIM-DAL: Uncertainty-aware Statistical Image Modeling-based Dense Active Learning for Super-resolution

UAI 2023poster

Dense regression is a widely used approach in computer vision for tasks such as image super-resolution, enhancement, depth estimation, etc. However, the high cost of annotation and labeling makes it challenging to achieve accurate results. We propose incorporating active learning into dense regressi…

Cited by 3SourcePDFScholar
2022

Teaching CNNs to Mimic Human Visual Cognitive Process & Regularise Texture-Shape Bias

ICASSP 2022accepted

Recent experiments in computer vision demonstrate texture bias as the primary reason for supreme results in models employing Convolutional Neural Networks (CNNs), conflicting with early works claiming that these networks identify objects using shape. It is believed that the cost function forces the…

Cited by 0SourceScholar
2020

LEt-SNE: A Hybrid Approach to Data Embedding and Visualization Of Hyperspectral Imagery

ICASSP 2020accepted

Hyperspectral Imagery (and Remote Sensing in general) captured from UAVs or satellites are highly voluminous in nature due to the large spatial extent and wavelengths captured by them. Since analyzing these images requires a huge amount of computational time and power, various dimensionality reducti…

Cited by 0SourceScholar
2020

Multi-Source Open-Set Deep Adversarial Domain Adaptation

ECCV 2020poster

We introduce a novel learning paradigm based on multi-source open-set unsupervised domain adaptation (MS-OSDA). Recently, the notion of single-source open-set domain adaptation (OSDA) has drawn much attention which considers the presence of previously unseen open-set (unknown) classes in the target-…

Cited by 43SourcePDFScholar