← Search

Muzammal Naseer

33 accepted papers

2026

EgoHandICL: Egocentric 3D Hand Reconstruction with In-Context Learning

ICLR 2026poster

Robust 3D hand reconstruction is challenging in egocentric vision due to depth ambiguity, self-occlusion, and complex hand-object interactions. Prior works attempt to mitigate the challenges by scaling up training data or incorporating auxiliary cues, often falling short of effectively handling unse…

Cited by 0SourcecodeScholar
2026

LATA: Laplacian-Assisted Transductive Adaptation for Conformal Uncertainty in Medical VLMs

CVPR 2026

Medical vision-language models (VLMs) are strong zero-shot recognizers for medical imaging, but their reliability under domain shift hinges on calibrated uncertainty with guarantees. Split conformal prediction (SCP) offers finite-sample coverage, yet prediction sets often become large (low efficienc

Cited by 0SourceScholar
2026

RedSage: A Cybersecurity Generalist LLM

ICLR 2026poster

Cybersecurity operations demand assistant LLMs that support diverse workflows without exposing sensitive data. Existing solutions either rely on proprietary APIs with privacy risks or on open models lacking domain adaptation. To bridge this gap, we curate 11.8B tokens of cybersecurity-focused contin…

Cited by 0SourcecodeScholar
2026

SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs

CVPR 2026

Multimodal large language models (MLLMs) have advanced from image-level reasoning to pixel-level grounding, but extending these capabilities to videos remains challenging as models must achieve spatial precision and temporally consistent reference tracking. Existing video MLLMs often rely on a stati

Cited by 0SourcecodeScholar
2025

AgriCLIP: Adapting CLIP for Agriculture and Livestock via Domain-Specialized Cross-Model Alignment

COLING 2025main

Capitalizing on a vast amount of image-text data, large-scale vision-language pre-training has demonstrated remarkable zero-shot capabilities and has been utilized in several applications. However, models trained on general everyday web-crawled data often exhibit sub-optimal performance for speciali…

2025

DyCON: Dynamic Uncertainty-aware Consistency and Contrastive Learning for Semi-supervised Medical Image Segmentation

CVPR 2025poster

Semi-supervised learning in medical image segmentation leverages unlabeled data to reduce annotation burdens through consistency learning. However, current methods struggle with class imbalance and high uncertainty from pathology variations, leading to inaccurate segmentation in 3D medical images. T…

2025

Learning to Prompt with Text Only Supervision for Vision-Language Models

AAAI 2025technical

Foundational vision-language models like CLIP are emerging as a promising paradigm in vision due to their excellent generalization. However, adapting these models for downstream tasks while maintaining their generalization remains challenging. In literature, one branch of methods adapts CLIP by lear…

2025

MixANT: Observation-dependent Memory Propagation for Stochastic Dense Action Anticipation

ICCV 2025poster

We present MixANT, a novel architecture for stochastic long-term dense anticipation of human activities. While recent State Space Models (SSMs) like Mamba have shown promise through input-dependent selectivity on three key parameters, the critical forget-gate (A matrix) controlling temporal memory r…

2025

STEREO: A Two-Stage Framework for Adversarially Robust Concept Erasing from Text-to-Image Diffusion Models

CVPR 2025highlight

The rapid proliferation of large-scale text-to-image diffusion (T2ID) models has raised serious concerns about their potential misuse in generating harmful content. Although numerous methods have been proposed for erasing undesired concepts from T2ID models, they often provide a false sense of secu…

2025

STING-BEE: Towards Vision-Language Model for Real-World X-ray Baggage Security Inspection

CVPR 2025highlight

Advancements in Computer-Aided Screening (CAS) systems are essential for improving the detection of security threats in X-ray baggage scans. However, current datasets are limited in representing real-world, sophisticated threats and concealment tactics, and existing approaches are constrained by a c…

2025

VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs

NAACL 2025findings

The recent advancements in Large Language Models (LLMs) have greatly influenced the development of Large Multi-modal Video Models (Video-LMMs), significantly enhancing our ability to interpret and analyze video data. Despite their impressive capabilities, current Video-LMMs have not been evaluated f…

2025

Video-Panda: Parameter-efficient Alignment for Encoder-free Video-Language Models

CVPR 2025poster

We present an efficient encoder-free approach for video-language understanding that achieves competitive performance while significantly reducing computational overhead. Current video-language models typically rely on heavyweight image encoders (300M-1.1B parameters) or video encoders (1B-1.4B param…

2025

Vision-Language Neural Graph Featurization for Extracting Retinal Lesions

ICCV 2025poster

Retinopathy comprises a group of retinal disorders that can lead to severe visual impairments or blindness. The heterogeneous morphology of lesions poses a significant challenge in developing robust diagnostic systems. Supervised approaches rely on large labeled datasets and often struggle with gene…

Cited by 0SourcePDFScholar
2024

Composed Video Retrieval via Enriched Context and Discriminative Embeddings

CVPR 2024poster

Composed video retrieval (CoVR) is a challenging prob- lem in computer vision which has recently highlighted the in- tegration of modification text with visual queries for more so- phisticated video search in large databases. Existing works predominantly rely on visual queries combined with modi- fi…

2024

GeoChat: Grounded Large Vision-Language Model for Remote Sensing

CVPR 2024poster

Recent advancements in Large Vision-Language Models (VLMs) have shown great promise in natural image domains allowing users to hold a dialogue about given visual content. However such general-domain VLMs perform poorly for Remote Sensing (RS) scenarios leading to inaccurate or fabricated information…

2024

LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed Prompts

ICLR 2024poster

Diffusion-based generative models have significantly advanced text-to-image generation but encounter challenges when processing lengthy and intricate text prompts describing complex scenes with multiple objects. While excelling in generating images from short, single-object descriptions, these model…

2024

Rethinking Transformers Pre-training for Multi-Spectral Satellite Imagery

CVPR 2024poster

Recent advances in unsupervised learning have demonstrated the ability of large vision models to achieve promising results on downstream tasks by pre-training on large amount of unlabelled data. Such pre-training techniques have also been explored recently in the remote sensing domain due to the ava…

2024

S3A: Towards Realistic Zero-Shot Classification via Self Structural Semantic Alignment

AAAI 2024technical

Large-scale pre-trained Vision Language Models (VLMs) have proven effective for zero-shot classification. Despite the success, most traditional VLMs-based methods are restricted by the assumption of partial source supervision or ideal target vocabularies, which rarely satisfy the open-world scenario…

2024

VideoGrounding-DINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding

CVPR 2024poster

Video grounding aims to localize a spatio-temporal section in a video corresponding to an input text query. This paper addresses a critical limitation in current video grounding methodologies by introducing an Open-Vocabulary Spatio-Temporal Video Grounding task. Unlike prevalent closed-set approach…

Cited by 13SourcePDFScholar
2023

Align Your Prompts: Test-Time Prompting with Distribution Alignment for Zero-Shot Generalization

NeurIPS 2023poster

The promising zero-shot generalization of vision-language models such as CLIP has led to their adoption using prompt learning for numerous downstream tasks. Previous works have shown test-time prompt tuning using entropy minimization to adapt text prompts for unseen domains. While effective, this ov…

2023

CLIP2Protect: Protecting Facial Privacy Using Text-Guided Makeup via Adversarial Latent Search

CVPR 2023poster

The success of deep learning based face recognition systems has given rise to serious privacy concerns due to their ability to enable unauthorized tracking of users in the digital world. Existing methods for enhancing privacy fail to generate naturalistic' images that can protect facial privacy with…

2023

FLIP: Cross-domain Face Anti-spoofing with Language Guidance

ICCV 2023poster

Face anti-spoofing (FAS) or presentation attack detection is an essential component of face recognition systems deployed in security-critical applications. Existing FAS methods have poor generalizability to unseen spoof types, camera sensors, and environmental conditions. Recently, vision transforme…

Cited by 57PDFcodeScholar
2023

PromptCAL: Contrastive Affinity Learning via Auxiliary Prompts for Generalized Novel Category Discovery

CVPR 2023poster

Although existing semi-supervised learning models achieve remarkable success in learning with unannotated in-distribution data, they mostly fail to learn on unlabeled data sampled from novel semantic classes due to their closed-set assumption. In this work, we target a pragmatic but under-explored G…

2023

Self-regulating Prompts: Foundational Model Adaptation without Forgetting

ICCV 2023poster

Prompt learning has emerged as an efficient alternative for fine-tuning foundational models, such as CLIP, for various downstream tasks. Conventionally trained using the task-specific objective, i.e., cross-entropy loss, prompts tend to overfit downstream data distributions and find it challenging t…

Cited by 205PDFcodeScholar
2023

Video-FocalNets: Spatio-Temporal Focal Modulation for Video Action Recognition

ICCV 2023poster

Recent video recognition models utilize Transformer models for long-range spatio-temporal context modeling. Video transformer designs are based on self-attention that can model global context at a high computational cost. In comparison, convolutional designs for videos offer an efficient alternative…

Cited by 28PDFcodeScholar
2023

Vita-CLIP: Video and Text Adaptive CLIP via Multimodal Prompting

CVPR 2023poster

Adopting contrastive image-text pretrained models like CLIP towards video classification has gained attention due to its cost-effectiveness and competitive performance. However, recent works in this area face a trade-off. Finetuning the pretrained model to achieve strong supervised performance resul…

2022

On Improving Adversarial Transferability of Vision Transformers

ICLR 2022spotlight

Vision transformers (ViTs) process input images as sequences of patches via self-attention; a radically different architecture than convolutional neural networks (CNNs). This makes it interesting to study the adversarial feature space of ViT models and their transferability. In particular, we obser…

Cited by 112SourcePDFScholar
2022

Self-Supervised Video Transformer

CVPR 2022oral

In this paper, we propose self-supervised training for video transformers using unlabeled video data. From a given video, we create local and global spatiotemporal views with varying spatial sizes and frame rates. Our self-supervised objective seeks to match the features of these different views rep…

Cited by 131PDFcodeScholar
2021

Intriguing Properties of Vision Transformers

NeurIPS 2021spotlight

Vision transformers (ViT) have demonstrated impressive performance across numerous machine vision tasks. These models are based on multi-head self-attention mechanisms that can flexibly attend to a sequence of image patches to encode contextual cues. An important question is how such flexibility (in…

Cited by 733SourcePDFScholar
2021

On Generating Transferable Targeted Perturbations

ICCV 2021poster

While the untargeted black-box transferability of adversarial perturbations has been extensively studied before, changing an unseen model's decisions to a specific `targeted' class remains a challenging feat. In this paper, we propose a new generative approach for highly transferable targeted pertur…

Cited by 92PDFcodeScholar
2020

A Self-supervised Approach for Adversarial Robustness

CVPR 2020oral

Adversarial examples can cause catastrophic mistakes in Deep Neural Network (DNNs) based vision systems e.g., for classification, segmentation and object detection. The vulnerability of DNNs against such attacks can prove a major roadblock towards their real-world deployment. Transferability of adve…

Cited by 346PDFcodeScholar