← Search

Muhammad Awais

27 accepted papers

2026

Building Robust Vision Encoders for Cross-Dataset Evaluation in Immunofluorescent Microscopy

CVPR 2026

Immunofluorescence (IF) images reveal detailed information about structures and functions at the subcellular level. However, unlike RGB images, IF datasets pose challenges for deep learning models due to their inconsistencies in channel count and configuration, stemming from varying staining protoco

Cited by 0SourcecodeScholar
2026

CoLA: Cross-Modal Low-rank Adaptation for Multimodal Downstream Tasks

ICML 2026poster

Foundation models have revolutionized AI, but adapting them efficiently for multimodal tasks, particularly in dual-stream architectures composed of unimodal encoders, such as DINO and BERT, remains a significant challenge. Parameter-Efficient Fine-Tuning (PEFT) methods like Low-Rank Adaptation (LoRA…

Cited by 0SourceScholar
2026

Object-Centric Refinement for Enhanced Zero-Shot Segmentation

ICLR 2026poster

Zero-shot semantic segmentation aims to recognize, pixel-wise, unseen categories without annotated masks, typically by leveraging vision-language models such as CLIP. However, the patch representations obtained by the CLIP's vision encoder lack object-centric structure, making it difficult to locali…

Cited by 0SourceScholar
2026

RAIGen: Rare Attribute Identification in Text-to-Image Generative Models

ICML 2026poster

Text-to-image diffusion models achieve impressive generation quality but inherit and amplify training-data biases, skewing coverage of semantic attributes. Prior work addresses this in two ways. Closed-set approaches mitigate biases in predefined fairness categories (e.g., gender, race), assuming so…

Cited by 0SourceScholar
2026

SyncDreamer: Controllable and Expressive Avatar Generation Beyond the Talking Head

CVPR 2026

Generating realistic and expressive audio-driven talking avatars remains a central challenge in digital human synthesis. Existing methods often depend on intermediate representations such as pose estimations for natural body motion, which restricts flexibility and adds visual distortions. Moreover,

Cited by 0SourcecodeScholar
2025

Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning

NeurIPS 2025poster

Despite significant advances in inference-time search for vision–language models (VLMs), existing approaches remain both computationally expensive and prone to unpenalized, low-confidence generations which often lead to persistent hallucinations. We introduce \textbf{Value-guided Inference with Marg…

Cited by 0SourcecodeScholar
2025

Enhanced Weakly Supervised Few-shot Classification & Segmentation

ICASSP 2025accepted

The emergence of vision-language foundation models has enabled the integration of textual information into vision-based applications. However, in few-shot classification and segmentation (FS-CS), this potential remains underutilised. Commonly, self-supervised vision models have been employed, partic…

Cited by 0SourceScholar
2025

One Model for ALL: Low-Level Task Interaction Is a Key to Task-Agnostic Image Fusion

CVPR 2025poster

Advanced image fusion methods mostly prioritise high-level missions, where task interaction struggles with semantic gaps, requiring complex bridging mechanisms. In contrast, we propose to leverage low-level vision tasks from digital photography fusion, allowing for effective feature interaction thro…

2025

RespoDiff: Dual-Module Bottleneck Transformation for Responsible & Faithful T2I Generation

NeurIPS 2025poster

The rapid advancement of diffusion models has enabled high-fidelity and semantically rich text-to-image generation; however, ensuring fairness and safety remains an open challenge. Existing methods typically improve fairness and safety at the expense of semantic fidelity and image quality. In this w…

Cited by 0SourcecodeScholar
2025

SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes

ICLR 2025poster

Self-supervised pre-trained audio networks have seen widespread adoption in real-world systems, particularly in multi-modal large language models. These networks are often employed in a frozen state, under the assumption that the self-supervised pre-training has sufficiently equipped them to handle…

2025

Text Augmented Correlation Transformer For Few-shot Classification & Segmentation

CVPR 2025poster

Foundation models like CLIP and ALIGN have transformed few-shot and zero-shot vision applications by fusing visual and textual data, yet the integrative few-shot classification and segmentation (FS-CS) task primarily leverages visual cues, overlooking the potential of textual support. In FS-CS scena…

Cited by 0SourcePDFScholar
2024

C2C: Component-to-Composition Learning for Zero-Shot Compositional Action Recognition

ECCV 2024oral

"Compositional actions consist of dynamic (verbs) and static (objects) concepts. Humans can easily recognize unseen compositions using the learned concepts. For machines, solving such a problem requires a model to recognize unseen actions composed of previously observed verbs and objects, thus requi…

2024

DTF-AT: Decoupled Time-Frequency Audio Transformer for Event Classification

AAAI 2024technical

Convolutional neural networks (CNNs) and Transformer-based networks have recently enjoyed significant attention for various audio classification and tagging tasks following their wide adoption in the computer vision domain. Despite the difference in information distribution between audio spectrogram…

2024

Efficient 3D-Aware Facial Image Editing via Attribute-Specific Prompt Learning

ECCV 2024poster

"Drawing upon StyleGAN’s expressivity and disentangled latent space, existing 2D approaches employ textual prompting to edit facial images with different attributes. In contrast, 3D-aware approaches that generate faces at different target poses require attribute-specific classifiers, learning separa…

2024

G-SHARP: Globally Shared Kernel with Pruning for Efficient CNNs

ICASSP 2024accepted

Filter Decomposition (FD) methods have gained traction in compressing large neural networks by dividing weights into basis and coefficients. Recent advancements have focused on reducing weight redundancy by sharing either basis or coefficients stage-wise. However, traditional sharing approaches have…

Cited by 0SourceScholar
2024

Improved Image Captioning Via Knowledge Graph-Augmented Models

ICASSP 2024accepted

Multimodal foundation models, pre-trained on large-scale data, effectively capture vast amounts of factual and commonsense knowledge. However, these models store all their knowledge within their parameters, requiring increasingly larger models and training data to capture more knowledge. To address…

Cited by 0SourceScholar
2024

Max-AST: Combining Convolution, Local and Global Self-Attentions for Audio Event Classification

ICASSP 2024accepted

In the domain of audio transformer architectures, prior research has extensively investigated isotropic architectures that capture the global context through full self-attention and hierarchical architectures that progressively transition from local to global context utilising hierarchical structure…

Cited by 0SourceScholar
2024

SCD-Net: Spatiotemporal Clues Disentanglement Network for Self-Supervised Skeleton-Based Action Recognition

AAAI 2024technical

Contrastive learning has achieved great success in skeleton-based action recognition. However, most existing approaches encode the skeleton sequences as entangled spatiotemporal representations and confine the contrasts to the same level of representation. Instead, this paper introduces a novel cont…

2022

AXM-Net: Implicit Cross-Modal Feature Alignment for Person Re-identification

AAAI 2022technical

Cross-modal person re-identification (Re-ID) is critical for modern video surveillance systems. The key challenge is to align cross-modality representations conforming to semantic information present for a person and ignore background information. This work presents a novel convolutional neural netw…

Cited by 109SourcePDFScholar
2022

ZooD: Exploiting Model Zoo for Out-of-Distribution Generalization

NeurIPS 2022accept

Recent advances on large-scale pre-training have shown great potentials of leveraging a large set of Pre-Trained Models (PTMs) for improving Out-of-Distribution (OoD) generalization, for which the goal is to perform well on possible unseen domains after fine-tuning on multiple training domains. Howe…

Cited by 19SourcePDFScholar
2021

Adversarial Robustness for Unsupervised Domain Adaptation

ICCV 2021poster

Extensive Unsupervised Domain Adaptation (UDA) studies have shown great success in practice by learning transferable representations across a labeled source domain and an unlabeled target domain with deep models. However, current work focuses on improving the generalization ability of UDA models on…

Cited by 46PDFScholar
2021

How Does Loss Function Affect Generalization Performance of Deep Learning? Application to Human Age Estimation

ICML 2021spotlight

Good generalization performance across a wide variety of domains caused by many external and internal factors is the fundamental goal of any machine learning algorithm. This paper theoretically proves that the choice of loss function matters for improving the generalization performance of deep learn…

Cited by 51SourcePDFScholar
2021

MixACM: Mixup-Based Robustness Transfer via Distillation of Activated Channel Maps

NeurIPS 2021poster

Deep neural networks are susceptible to adversarially crafted, small, and imperceptible changes in the natural inputs. The most effective defense mechanism against these examples is adversarial training which constructs adversarial examples during training by iterative maximization of loss. The mode…

Cited by 20SourcePDFScholar
2019

Divergence Based Weighting for Information Channels in Deep Convolutional Neural Networks for Bird Audio Detection

ICASSP 2019accepted

In this paper, we address the problem of bird audio detection and propose a new convolutional neural network architecture together with a divergence based information channel weighing strategy in order to achieve improved state-of-the-art performance and faster convergence. The effectiveness of the…

Cited by 0SourceScholar
2019

Spoofing Attack Detection by Anomaly Detection

ICASSP 2019accepted

Spoofing attacks on biometric systems can seriously compromise their practical utility. In this paper we focus on face spoofing detection. The majority of papers on spoofing attack detection formulate the problem as a two or multiclass learning task, attempting to separate normal accesses from sampl…

Cited by 0SourceScholar
2018

Wing Loss for Robust Facial Landmark Localisation With Convolutional Neural Networks

CVPR 2018poster

We present a new loss function, namely Wing loss, for robust facial landmark localisation with Convolutional Neural Networks (CNNs). We first compare and analyse different loss functions including L2, L1 and smooth L1. The analysis of these loss functions suggests that, for the training of a CNN-bas…

Cited by 548SourcePDFScholar