← Search

Arif Mahmood

11 accepted papers

2026

MLLM-HWSI: A Multimodal Large Language Model for Hierarchical Whole Slide Image Understanding

CVPR 2026

Whole Slide Images (WSIs) exhibit hierarchical structure, where diagnostic information emerges from cellular morphology, regional tissue organization, and global context. Existing Computational Pathology (CPath) Multimodal Large Language Models (MLLMs) typically compress an entire WSI into a single

Cited by 0SourcecodeScholar
2025

Diffusion-Guided Graph Data Augmentation

NeurIPS 2025poster

Graph Neural Networks (GNNs) have achieved remarkable success in a wide range of applications. However, when trained on limited or low-diversity datasets, GNNs are prone to overfitting and memorization, which impacts their generalization. To address this, graph data augmentation (GDA) has become a c…

Cited by 0SourceScholar
2025

Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual Representation

CVPR 2025poster

In Computational Pathology (CPath), the introduction of Vision-Language Models (VLMs) has opened new avenues for research, focusing primarily on aligning image-text pairs at a single magnification level. However, this approach might not be sufficient for tasks like cancer subtype classification, tis…

2025

Unsupervised Discovery of Facial Landmarks and Head Pose

CVPR 2025poster

Unsupervised landmark and head pose estimation is fundamental in fields like biometrics, augmented reality, and emotion recognition, offering accurate spatial data without relying on labeled datasets. It enhances scalability, adaptability, and generalization across diverse settings, where manual lab…

2024

CPLIP: Zero-Shot Learning for Histopathology with Comprehensive Vision-Language Alignment

CVPR 2024poster

This paper proposes Comprehensive Pathology Language Image Pre-training (CPLIP) a new unsupervised technique designed to enhance the alignment of images and text in histopathology for tasks such as classification and segmentation. This methodology enriches vision language models by leveraging extens…

2024

DiffuseMix: Label-Preserving Data Augmentation with Diffusion Models

CVPR 2024poster

Recently a number of image-mixing-based augmentation techniques have been introduced to improve the generalization of deep neural networks. In these techniques two or more randomly selected natural images are mixed together to generate an augmented image. Such methods may not only omit important por…

Cited by 31SourcePDFScholar
2024

Pose-Guided Self-Training with Two-Stage Clustering for Unsupervised Landmark Discovery

CVPR 2024highlight

Unsupervised landmarks discovery (ULD) for an object category is a challenging computer vision problem. In pursuit of developing a robust ULD framework we explore the potential of a recent paradigm of self-supervised learning algorithms known as diffusion models. Some recent works have shown that th…

2023

Higher-Order Sparse Convolutions in Graph Neural Networks

ICASSP 2023accepted

Graph Neural Networks (GNNs) have been applied to many problems in computer sciences. Capturing higher-order relationships between nodes is crucial to increase the expressive power of GNNs. However, existing methods to capture these relationships could be infeasible for large-scale graphs. In this w…

Cited by 0SourceScholar
2023

Single-branch Network for Multimodal Training

ICASSP 2023accepted

With the rapid growth of social media platforms, users are sharing billions of multimedia posts containing audio, images, and text. Researchers have focused on building autonomous systems capable of processing such multimedia data to solve challenging multimodal tasks including cross-modal retrieval…

Cited by 0SourceScholar
2022

Generative Cooperative Learning for Unsupervised Video Anomaly Detection

CVPR 2022poster

Video anomaly detection is well investigated in weakly supervised and one-class classification (OCC) settings. However, unsupervised video anomaly detection is quite sparse, likely because anomalies are less frequent in occurrence and usually not well-defined, which when coupled with the absence of…

Cited by 205PDFScholar
2020

CLAWS: Clustering Assisted Weakly Supervised Learning with Normalcy Suppression for Anomalous Event Detection

ECCV 2020poster

Learning to detect real-world anomalous events through video-level labels is a challenging task due to the rare occurrence of anomalies as well as noise in the labels. In this work, we propose a weakly supervised anomaly detection method which has manifold contributions including 1) a random batch b…

Cited by 194SourcePDFScholar