← Search

Muhammad Haris Khan

47 accepted papers

2026

A-TPT: Angular Diversity Calibration Properties for Test-Time Prompt Tuning of Vision-Language Models

ICLR 2026poster

Test-time prompt tuning (TPT) has emerged as a promising technique for adapting large vision-language models (VLMs) to unseen tasks without relying on labeled data. However, the lack of dispersion between textual features can hurt calibration performance, which raises concerns about VLMs' reliabilit…

Cited by 0SourceScholar
2026

Balancing Multimodal Domain Generalization via Gradient Modulation and Projection

AAAI 2026technical

Multimodal Domain Generalization (MMDG) leverages the complementary strengths of multiple modalities to enhance model generalization on unseen domains. A central challenge in multimodal learning is optimization imbalance, where modalities converge at different speeds during training. This imbalance

Cited by 0SourcePDFScholar
2026

Foundation Model Priors Enhance Object Focus in Feature Space for Source-Free Object Detection

CVPR 2026

Current state-of-the-art approaches in Source-Free Object Detection (SFOD) typically rely on Mean-Teacher self-labeling. However, domain shift often reduces the detector's ability to maintain strong object-focused representations, causing high-confidence activations over background clutter and unrel

Cited by 0SourcecodeScholar
2026

Keep It Frozen: Domain-Routed Conditional Residual Modulation for Multi-Domain Vision Transformers

CVPR 2026

Medical imaging remains challenging due to acoustic shadows, motion blur, and indistinct boundaries, while adapting vision models to such domains often requires heavy task-specific fine-tuning and can degrade general-image capability. We propose DCRM-ViT, a domain-conditioned residual modulation fra

Cited by 0SourceScholar
2026

MMClima: A Framework for Multimodal Climate Science Data and Evaluation

ICML 2026poster

Climate change research increasingly requires AI systems that reason across text, dynamic visual content, and scientific figures, yet existing climate QA benchmarks are small, mostly textual, and cover a narrow range of models. We introduce MMClima, a large-scale multimodal climate question answerin…

Cited by 0SourceScholar
2026

Pixels Don't Lie (But Your Detector Might): Bootstrapping MLLM-as-a-Judge for Trustworthy Deepfake Detection and Reasoning Supervision

CVPR 2026

Deepfake detection models often generate natural-language explanations, yet their reasoning is frequently ungrounded in visual evidence, limiting reliability. Existing evaluations measure classification accuracy but overlook reasoning fidelity. We propose DeepfakeJudge, a framework for scalable reas

Cited by 0SourceScholar
2026

StyleDoctor: Towards Specialist Reward Model for Style-centric Generation Tasks

CVPR 2026

Style generation has made significant progress through diffusion models. Recent efforts have explored reinforcement learning with human-preference reward models to enhance diffusion models for general downstream applications. However, we identify a critical limitation: existing human-preference rewa

Cited by 0SourceScholar
2026

Tell me Habibi, is it Real or Fake?

ICLR 2026poster

Deepfake generation methods are evolving fast, making fake media harder to detect and raising serious societal concerns. Most deepfake detection and dataset creation research focuses on monolingual content, often overlooking the challenges of multilingual and code-switched speech, where multiple lan…

Cited by 0SourceScholar
2026

TerraFM: A Scalable Foundation Model for Unified Multisensor Earth Observation

ICLR 2026poster

Modern Earth observation (EO) increasingly leverages deep learning to harness the scale and diversity of satellite imagery across sensors and regions. While recent foundation models have demonstrated promising generalization across EO tasks, many remain limited by the scale, geographical coverage, a…

Cited by 0SourcecodeScholar
2026

Towards Calibrating Prompt Tuning of Vision- Language Models

CVPR 2026

Prompt tuning of large-scale vision-language models such as CLIP enables efficienttask adaptation without updating model weights. However, it often leads to poorconfidence calibration and unreliable predictive uncertainty. We address thisproblem by proposing a calibration framework that enhances pre

Cited by 0SourcecodeScholar
2026

Towards Multimodal Domain Generalization with Few Labels

CVPR 2026

Multimodal models ideally should generalize to unseen domains while remaining data-efficient to reduce annotation costs. To this end, we introduce and study a new problem, Semi-Supervised Multimodal Domain Generalization (SSMDG), which aims to learn robust multimodal models from multi-source data wi

Cited by 0SourcecodeScholar
2025

Diffusion-Guided Graph Data Augmentation

NeurIPS 2025poster

Graph Neural Networks (GNNs) have achieved remarkable success in a wide range of applications. However, when trained on limited or low-diversity datasets, GNNs are prone to overfitting and memorization, which impacts their generalization. To address this, graph data augmentation (GDA) has become a c…

Cited by 0SourceScholar
2025

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning

NeurIPS 2025poster

Despite significant advances in inference-time search for vision–language models (VLMs), existing approaches remain both computationally expensive and prone to unpenalized, low-confidence generations which often lead to persistent hallucinations. We introduce \textbf{Value-guided Inference with Marg…

Cited by 0SourcecodeScholar
2025

Hyperbolic Uncertainty-Aware Few-Shot Incremental Point Cloud Segmentation

CVPR 2025poster

3D point cloud segmentation is essential across a range of applications; however, conventional methods often struggle in evolving environments, particularly when tasked with identifying novel categories under limited supervision. Few-Shot Learning (FSL) and Class Incremental Learning (CIL) have been…

Cited by 0SourcePDFScholar
2025

ImpedanceGPT: VLM-driven Impedance Control of Swarm of Mini-drones for Intelligent Navigation in Dynamic Environment

IROS 2025

Swarm robotics plays a crucial role in enabling autonomous operations in dynamic and unpredictable environments. However, a major challenge remains ensuring safe and efficient navigation in environments shared by both dynamic alive (e.g., humans) and dynamic inanimate (e.g., non-living objects) obst

Cited by 8SourcecodeScholar
2025

Industry 6.0: New Generation of Industry driven by Generative AI and Swarm of Heterogeneous Robots

IROS 2025

This paper presents the concept of Industry 6.0, which introduces the world’s first fully automated production system that autonomously handles the entire product design and manufacturing process based on user-provided natural language descriptions. By leveraging generative AI, the system automates

Cited by 13SourceScholar
2025

Leveraging 2D Priors and SDF Guidance for Urban Scene Rendering

ICCV 2025poster

Dynamic scene rendering and reconstruction play a crucial role in computer vision and augmented reality. Recent methods based on 3D Gaussian Splatting (3DGS), have enabled accurate modeling of dynamic urban scenes, but for urban scenes they require both camera and LiDAR data, ground-truth 3D segment…

Cited by 0SourcePDFScholar
2025

MSAmba: Exploring Multimodal Sentiment Analysis with State Space Models

AAAI 2025technical

Multimodal sentiment analysis, which learns a model to process multiple modalities simultaneously and predict a sentiment value, is an important area of affective computing. Modeling sequential intra-modal information and enhancing cross-modal interactions are crucial to multimodal sentiment analysi…

2025

O-TPT: Orthogonality Constraints for Calibrating Test-time Prompt Tuning in Vision-Language Models

CVPR 2025highlight

Test-time prompt tuning for vision-language models (VLMs) is getting attention because of their ability to learn with unlabeled data without fine-tuning. Although test-time prompt tuning methods for VLMs can boost accuracy, the resulting models tend to demonstrate poor calibration, which casts doubt…

2025

OSLoPrompt: Bridging Low-Supervision Challenges and Open-Set Domain Generalization in CLIP

CVPR 2025poster

We introduce Low-Shot Open-Set Domain Generalization (LSOSDG), a novel paradigm unifying low-shot learning with open-set domain generalization (ODG). While prompt-based methods using models like CLIP have advanced DG, they falter in low-data regimes (e.g., 1-shot) and lack precision in detecting ope…

2025

SegMASt3R: Geometry Grounded Segment Matching

NeurIPS 2025spotlight

Segment matching is an important intermediate task in computer vision that establishes correspondences between semantically or geometrically coherent regions across images. Unlike keypoint matching, which focuses on localized features, segment matching captures structured regions, offering greater r…

Cited by 0SourceScholar
2025

SynFER: Towards Boosting Facial Expression Recognition with Synthetic Data

ICCV 2025poster

Facial expression datasets remain limited in scale due to privacy concerns, the subjectivity of annotations, and the labor-intensive nature of data collection. This limitation poses a significant challenge for developing modern deep learning-based facial expression analysis models, particularly foun…

Cited by 0SourcePDFScholar
2025

Unsupervised Discovery of Facial Landmarks and Head Pose

CVPR 2025poster

Unsupervised landmark and head pose estimation is fundamental in fields like biometrics, augmented reality, and emotion recognition, offering accurate spatial data without relying on labeled datasets. It enhances scalability, adaptability, and generalization across diverse settings, where manual lab…

2024

Discrete Cycle-Consistency Based Unsupervised Deep Graph Matching

AAAI 2024technical

We contribute to the sparsely populated area of unsupervised deep graph matching with application to keypoint matching in images. Contrary to the standard supervised approach, our method does not require ground truth correspondences between keypoint pairs. Instead, it is self-supervised by enforcing…

Cited by 1SourcePDFScholar
2024

Improving Single Domain-Generalized Object Detection: A Focus on Diversification and Alignment

CVPR 2024poster

In this work we tackle the problem of domain generalization for object detection specifically focusing on the scenario where only a single source domain is available. We propose an effective approach that involves two key steps: diversifying the source domain and aligning detections based on class p…

2024

Leveraging Cycle-Consistent Anchor Points for Self-Supervised RGB-D Registration

ICRA 2024poster

With the rise in consumer depth cameras, a wealth of unlabeled RGB-D data has become available. This prompts the question of how to utilize this data for geometric reasoning of scenes. While many RGB-D registration methods rely on geometric and feature-based similarity, we take a different approach.…

Cited by 0SourceScholar
2024

Pose-Guided Self-Training with Two-Stage Clustering for Unsupervised Landmark Discovery

CVPR 2024highlight

Unsupervised landmarks discovery (ULD) for an object category is a challenging computer vision problem. In pursuit of developing a robust ULD framework we explore the potential of a recent paradigm of self-supervised learning algorithms known as diffusion models. Some recent works have shown that th…

2024

Towards Combating Frequency Simplicity-biased Learning for Domain Generalization

NeurIPS 2024poster

Domain generalization methods aim to learn transferable knowledge from source domains that can generalize well to unseen target domains. Recent studies show that neural networks frequently suffer from a simplicity-biased learning behavior which leads to over-reliance on specific frequency sets, nam…

2024

Towards Generalizing to Unseen Domains with Few Labels

CVPR 2024poster

We approach the challenge of addressing semi-supervised domain generalization (SSDG). Specifically our aim is to obtain a model that learns domain-generalizable features by leveraging a limited subset of labelled data alongside a substantially larger pool of unlabeled data. Existing domain generaliz…

2024

Why Not Use Your Textbook? Knowledge-Enhanced Procedure Planning of Instructional Videos

CVPR 2024poster

In this paper we explore the capability of an agent to construct a logical sequence of action steps thereby assembling a strategic procedural plan. This plan is crucial for navigating from an initial visual observation to a target visual outcome as depicted in real-life instructional videos. Existin…

2023

Bridging Precision and Confidence: A Train-Time Loss for Calibrating Object Detection

CVPR 2023poster

Deep neural networks (DNNs) have enabled astounding progress in several vision-based problems. Despite showing high predictive accuracy, recently, several works have revealed that they tend to provide overconfident predictions and thus are poorly calibrated. The majority of the works addressing the…

2023

Cal-DETR: Calibrated Detection Transformer

NeurIPS 2023poster

Albeit revealing impressive predictive performance for several computer vision tasks, deep neural networks (DNNs) are prone to making overconfident predictions. This limits the adoption and wider utilization of DNNs in many safety-critical applications. There have been recent efforts toward calibrat…

2023

MSI: Maximize Support-Set Information for Few-Shot Segmentation

ICCV 2023poster

FSS (Few-shot segmentation) aims to segment a target class using a small number of labeled images (support set). To extract the information relevant to target class, a dominant approach in best performing FSS methods removes background features using a support mask. We observe that this feature exci…

Cited by 32PDFcodeScholar
2023

Multiclass Confidence and Localization Calibration for Object Detection

CVPR 2023poster

Albeit achieving high predictive accuracy across many challenging computer vision problems, recent studies suggest that deep neural networks (DNNs) tend to make overconfident predictions, rendering them poorly calibrated. Most of the existing attempts for improving DNN calibration are limited to cla…

2023

Single-branch Network for Multimodal Training

ICASSP 2023accepted

With the rapid growth of social media platforms, users are sharing billions of multimedia posts containing audio, images, and text. Researchers have focused on building autonomous systems capable of processing such multimedia data to solve challenging multimodal tasks including cross-modal retrieval…

Cited by 0SourceScholar
2022

Fusion and Orthogonal Projection for Improved Face-Voice Association

ICASSP 2022accepted

We study the problem of learning association between face and voice. Prior works adopt pairwise or triplet loss formulations to learn an embedding space amenable for associated matching and verification tasks. Albeit showing some progress, such loss formulations are restrictive due to dependency on…

Cited by 0SourceScholar
2022

HM: Hybrid Masking for Few-Shot Segmentation

ECCV 2022poster

"We study few-shot semantic segmentation that aims to segment a target object from a query image when provided with a few annotated support images of the target class. Several recent methods resort to a feature masking (FM) technique to discard irrelevant feature activations which eventually facilit…

2022

Towards Improving Calibration in Object Detection Under Domain Shift

NeurIPS 2022accept

With deep neural network based solution more readily being incorporated in real-world applications, it has been pressing requirement that predictions by such models, especially in safety-critical environments, be highly accurate and well-calibrated. Although some techniques addressing DNN calibrati…

Cited by 22SourcePDFScholar
2022

Video Instance Segmentation via Multi-Scale Spatio-Temporal Split Attention Transformer

ECCV 2022poster

"State-of-the-art transformer-based video instance segmentation (VIS) approaches typically utilize either single-scale spatio-temporal features or per-frame multi-scale features during the attention computations. We argue that such an attention computation ignores the multi-scale spatio-temporal fea…

2021

SSAL: Synergizing between Self-Training and Adversarial Learning for Domain Adaptive Object Detection

NeurIPS 2021poster

We study adapting trained object detectors to unseen domains manifesting significant variations of object appearance, viewpoints and backgrounds. Most current methods align domains by either using image or instance-level feature alignment in an adversarial fashion. This often suffers due to the pres…

Cited by 84SourcePDFScholar
2020

AnimalWeb: A Large-Scale Hierarchical Dataset of Annotated Animal Faces

CVPR 2020poster

Several studies show that animal needs are often expressed through their faces. Though remarkable progress has been made towards the automatic understanding of human faces, this has not been the case with animal faces. There exists significant room for algorithmic advances that could realize automat…

Cited by 57PDFScholar
2019

Cross-Domain Transferability of Adversarial Perturbations

NeurIPS 2019poster

Adversarial examples reveal the blind spots of deep neural networks (DNNs) and represent a major concern for security-critical applications. The transferability of adversarial examples makes real-world attacks possible in black-box settings, where the attacker is forbidden to access the internal par…

2019

Deep Contextual Attention for Human-Object Interaction Detection

ICCV 2019poster

Human-object interaction detection is an important and relatively new class of visual relationship detection tasks, essential for deeper scene understanding. Most existing approaches decompose the problem into object localization and interaction recognition. Despite showing progress, these approache…

Cited by 130PDFScholar
2019

Mask-Guided Attention Network for Occluded Pedestrian Detection

ICCV 2019poster

Pedestrian detection relying on deep convolution neural networks has made significant progress. Though promising results have been achieved on standard pedestrians, the performance on heavily occluded pedestrians remains far from satisfactory. The main culprits are intra-class occlusions involving o…

Cited by 252PDFcodeScholar
2017

Synergy Between Face Alignment and Tracking via Discriminative Global Consensus Optimization

ICCV 2017spotlight

An open question in facial landmark localization in video is whether one should perform tracking or tracking-by-detection (i.e. face alignment). Tracking produces fittings of high accuracy but is prone to drifting. Tracking-by-detection is drift-free but results in low accuracy fittings. To provide…

Cited by 45PDFScholar
2015

TRIC-track: Tracking by Regression With Incrementally Learned Cascades

ICCV 2015poster

This paper proposes a novel approach to part-based tracking by replacing local matching of an appearance model by direct prediction of the displacement between local image patches and part locations. We propose to use cascaded regression with incremental learning to track generic objects without any…

Cited by 35PDFScholar