← Search

Rao Muhammad Anwer

49 accepted papers

2026

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks

ICLR 2026poster

Deep reasoning is fundamental for solving complex tasks, especially in vision-centric scenarios that demand sequential, multimodal understanding. However, existing benchmarks typically evaluate agents with fully synthetic, single-turn queries, limited visual modalities, and lack a framework to asses…

Cited by 0SourcecodeScholar
2026

TerraFM: A Scalable Foundation Model for Unified Multisensor Earth Observation

ICLR 2026poster

Modern Earth observation (EO) increasingly leverages deep learning to harness the scale and diversity of satellite imagery across sensors and regions. While recent foundation models have demonstrated promising generalization across EO tasks, many remain limited by the scale, geographical coverage, a…

Cited by 0SourcecodeScholar
2026

Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation

AAAI 2026technical

Referring Audio-Visual Segmentation (Ref-AVS) aims to segment target objects in audible videos based on given reference expressions. Prior works typically rely on learning latent embeddings via multimodal fusion to prompt a tunable SAM/SAM2 decoder for segmentation, which requires strong pixel-level

Cited by 0SourcePDFScholar
2025

Adapting In-Domain Few-Shot Segmentation to New Domains without Source Domain Retraining

ICCV 2025poster

Cross-domain few-shot segmentation (CD-FSS) aims to segment objects of novel classes in new domains, which is often challenging due to the diverse characteristics of target domains and the limited availability of support data. Most CD-FSS methods redesign and retrain in-domain FSS models using abund…

2025

All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages

CVPR 2025highlight

Existing Large Multimodal Models (LMMs) generally focus on only a few regions and languages. As LMMs continue to improve, it is increasingly important to ensure they understand cultural contexts, respect local sensitivities, and support low-resource languages, all while effectively integrating corr…

2025

All in One: Visual-Description-Guided Unified Point Cloud Segmentation

ICCV 2025poster

Unified segmentation of 3D point clouds is crucial for scene understanding, but is hindered by its sparse structure, limited annotations, and the challenge of distinguishing fine-grained object classes in complex environments. Existing methods often struggle to capture rich semantic and contextual i…

2025

Beyond Simple Edits: Composed Video Retrieval with Dense Modifications

ICCV 2025poster

Composed video retrieval is a challenging task that strives to retrieve a target video based on a query video and a textual description detailing specific modifications. Standard retrieval frameworks typically struggle to handle the complexity of fine-grained compositional queries and variations in…

2025

BiMediX2 : Bio-Medical EXpert LMM for Diverse Medical Modalities

EMNLP 2025

We introduce BiMediX2, a bilingual (Arabic-English) Bio-Medical EXpert Large Multimodal Model that supports text-based and image-based medical interactions. It enables multi-turn conversation in Arabic and English and supports diverse medical imaging modalities, including radiology, CT, and histolog

2025

CAMEL-Bench: A Comprehensive Arabic LMM Benchmark

NAACL 2025findings

Recent years have witnessed a significant interest in developing large multi-modal models (LMMs) capable of performing various visual reasoning and understanding tasks. This has led to the introduction of multiple LMM benchmarks to evaluate LMMs on different tasks. However, most existing LMM evaluat…

2025

DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image Models

NeurIPS 2025poster

Efficient fine-tuning of pre-trained Text-to-Image (T2I) models involves adjusting the model to suit a particular task or dataset while minimizing computational resources and limiting the number of trainable parameters. However, it often faces challenges in striking a trade-off between aligning with…

Cited by 0SourcecodeScholar
2025

DriveLMM-o1: A Step-by-Step Reasoning Dataset and Large Multimodal Model for Driving Scenario Understanding

IROS 2025

While large multimodal models (LMMs) have demonstrated strong performance across various Visual Question Answering (VQA) tasks, certain challenges require complex multi-step reasoning to reach accurate answers. One particularly challenging task is autonomous driving, which demands thorough cognitive

Cited by 32SourcecodeScholar
2025

Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs

EMNLP 2025

Arabic poetry stands as one of the most sophisticated and culturally embedded forms of expression in the Arabic language, known for its layered meanings, stylistic diversity, and deep historical continuity. Although large language models (LLMs) have demonstrated strong performance across languages a

2025

LLMVoX: Autoregressive Streaming Text-to-Speech Model for Any LLM

ACL 2025finding

Recent advancements in speech-to-speech dialogue systems leverage LLMs for multimodal interactions, yet they remain hindered by fine-tuning requirements, high computational overhead, and text-speech misalignment. Existing speech-enabled LLMs often degrade conversational quality by modifying the LLM,…

2025

LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

ACL 2025finding

Step-by-step reasoning is crucial for solving complex visual tasks, yet existing approaches lack a comprehensive framework for evaluating this capability and do not emphasize step-wise problem-solving. To this end, we propose a comprehensive framework for advancing multi-step visual reasoning in lar…

2025

MAviS: A Multimodal Conversational Assistant For Avian Species

EMNLP 2025

Fine-grained understanding and species-specific, multimodal question answering are vital for advancing biodiversity conservation and ecological monitoring. However, existing multimodal large language models (MM-LLMs) face challenges when it comes to specialized topics like avian species, making it h

2025

Open-YOLO 3D: Towards Fast and Accurate Open-Vocabulary 3D Instance Segmentation

ICLR 2025oral

Recent works on open-vocabulary 3D instance segmentation show strong promise but at the cost of slow inference speed and high computation requirements. This high computation cost is typically due to their heavy reliance on aggregated clip features from multi-view, which require computationally expen…

2025

Open3DTrack: Towards Open-Vocabulary 3D Multi-Object Tracking

ICRA 2025

3D multi-object tracking plays a critical role in autonomous driving by enabling the real-time monitoring and prediction of multiple objects' movements. Traditional 3D tracking systems are typically constrained by predefined object categories, limiting their adaptability to novel, unseen objects in

Cited by 3SourcecodeScholar
2025

RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping

ICCV 2025poster

General robotic grasping systems require accurate object affordance perception in diverse open-world scenarios following human instructions. However, current studies suffer from the problem of lacking reasoning-based large-scale affordance prediction data, leading to considerable concern about open-…

Cited by 0SourcePDFScholar
2025

Time Travel: A Comprehensive Benchmark to Evaluate LMMs on Historical and Cultural Artifacts

ACL 2025finding

Understanding historical and cultural artifacts demands human expertise and advanced computational techniques, yet the process remains complex and time-intensive. While large multimodal models offer promising support, their evaluation and improvement require a standardized benchmark. To address this…

2024

BiMediX: Bilingual Medical Mixture of Experts LLM

EMNLP 2024finding

In this paper, we introduce BiMediX, the first bilingual medical mixture of experts LLM designed for seamless interaction in both English and Arabic. Our model facilitates a wide range of medical interactions in English and Arabic, including multi-turn chats to inquire about additional details such…

2024

Bidirectional Reciprocative Information Communication for Few-Shot Semantic Segmentation

ICML 2024poster

Existing few-shot semantic segmentation methods typically rely on a one-way flow of category information from support to query, ignoring the impact of intra-class diversity. To address this, drawing inspiration from cybernetics, we introduce a Query Feedback Branch (QFB) to propagate query informati…

2024

Composed Video Retrieval via Enriched Context and Discriminative Embeddings

CVPR 2024poster

Composed video retrieval (CoVR) is a challenging prob- lem in computer vision which has recently highlighted the in- tegration of modification text with visual queries for more so- phisticated video search in large databases. Existing works predominantly rely on visual queries combined with modi- fi…

2024

Efficient 3D-Aware Facial Image Editing via Attribute-Specific Prompt Learning

ECCV 2024poster

"Drawing upon StyleGAN’s expressivity and disentangled latent space, existing 2D approaches employ textual prompting to edit facial images with different attributes. In contrast, 3D-aware approaches that generate faces at different target poses require attribute-specific classifiers, learning separa…

2024

Long-Tailed 3D Semantic Segmentation with Adaptive Weight Constraint and Sampling

ICRA 2024poster

Existing 3D understanding datasets typically provide annotations for a limited number of object classes, with sufficient examples per class. However, real-world object classes are not equally represented in practical settings, leading to poor performance on rarely-occurring categories if the class i…

Cited by 0SourceScholar
2024

Modulate Your Spectrum in Self-Supervised Learning

ICLR 2024poster

Whitening loss offers a theoretical guarantee against feature collapse in self-supervised learning (SSL) with joint embedding architectures. Typically, it involves a hard whitening approach, transforming the embedding and applying loss to the whitened output. In this work, we introduce Spectral Tran…

2024

Rethinking Transformers Pre-training for Multi-Spectral Satellite Imagery

CVPR 2024poster

Recent advances in unsupervised learning have demonstrated the ability of large vision models to achieve promising results on downstream tasks by pre-training on large amount of unlabelled data. Such pre-training techniques have also been explored recently in the remote sensing domain due to the ava…

2024

Semi-supervised Open-World Object Detection

AAAI 2024technical

Conventional open-world object detection (OWOD) problem setting first distinguishes known and unknown classes and then later incrementally learns the unknown objects when introduced with labels in the subsequent tasks. However, the current OWOD formulation heavily relies on the external human oracle…

2023

3D Indoor Instance Segmentation in an Open-World

NeurIPS 2023poster

Existing 3D instance segmentation methods typically assume that all semantic classes to be segmented would be available during training and only seen categories are segmented at inference. We argue that such a closed-world assumption is restrictive and explore for the first time 3D indoor instance s…

2023

Arabic Mini-ClimateGPT : A Climate Change and Sustainability Tailored Arabic LLM

EMNLP 2023short findings

Climate change is one of the most significant challenges we face together as a society. Creating awareness and educating policy makers the wide-ranging impact of climate change is an essential step towards a sustainable future. Recently, Large Language Models (LLMs) like ChatGPT and Bard have shown…

Cited by 0SourcecodeScholar
2023

Discriminative Co-Saliency and Background Mining Transformer for Co-Salient Object Detection

CVPR 2023poster

Most previous co-salient object detection works mainly focus on extracting co-salient cues via mining the consistency relations across images while ignoring the explicit exploration of background regions. In this paper, we propose a Discriminative co-saliency and background Mining Transformer framew…

2023

Generative Multiplane Neural Radiance for 3D-Aware Image Generation

ICCV 2023poster

We present a method to efficiently generate 3D-aware high-resolution images that are view-consistent across multiple target views. The proposed multiplane neural radiance model, named GMNR, consists of a novel a-guided view-dependent representation (a-VdR) module for learning view-dependent informat…

Cited by 3PDFcodeScholar
2023

Multi-grained Temporal Prototype Learning for Few-shot Video Object Segmentation

ICCV 2023poster

Few-Shot Video Object Segmentation (FSVOS) aims to segment objects in a query video with the same category defined by a few annotated support images. However, this task was seldom explored. In this work, based on IPMT, a state-of-the-art few-shot image segmentation method that combines external supp…

Cited by 11PDFcodeScholar
2023

Person Image Synthesis via Denoising Diffusion Model

CVPR 2023poster

The pose-guided person image generation task requires synthesizing photorealistic images of humans in arbitrary poses. The existing approaches use generative adversarial networks that do not necessarily maintain realistic textures or need dense correspondences that struggle to handle complex deforma…

2022

An Investigation into Whitening Loss for Self-supervised Learning

NeurIPS 2022accept

A desirable objective in self-supervised learning (SSL) is to avoid feature collapse. Whitening loss guarantees collapse avoidance by minimizing the distance between embeddings of positive pairs under the conditioning that the embeddings from different views are whitened. In this paper, we propose…

2022

Class-Agnostic Object Detection with Multi-modal Transformer

ECCV 2022poster

"What constitutes an object? This has been a long-standing question in computer vision. Towards this goal, numerous learning-free and learning-based approaches have been developed to score objectness. However, they generally do not scale well across new domains and for unseen objects. In this paper,…

2022

DoodleFormer: Creative Sketch Drawing with Transformers

ECCV 2022poster

"Creative sketching or doodling is an expressive activity, where imaginative and previously unseen depictions of everyday visual objects are drawn. Creative sketch image generation is a challenging vision problem, where the task is to generate diverse, yet realistic creative sketches possessing the…

2022

Energy-Based Latent Aligner for Incremental Learning

CVPR 2022poster

Deep learning models tend to forget their earlier knowledge while incrementally learning new tasks. This behavior emerges because the parameter updates optimized for the new tasks may not align well with the updates suitable for older tasks. The resulting latent representation mismatch causes forget…

Cited by 56PDFcodeScholar
2022

PSTR: End-to-End One-Step Person Search With Transformers

CVPR 2022poster

We propose a novel one-step transformer-based person search framework, PSTR, that jointly performs person detection and re-identification (re-id) in a single architecture. PSTR comprises a person search-specialized (PSS) module that contains a detection encoder-decoder for person detection along wit…

Cited by 75PDFcodeScholar
2022

Spatio-Temporal Relation Modeling for Few-Shot Action Recognition

CVPR 2022poster

We propose a novel few-shot action recognition framework, STRM, which enhances class-specific feature discriminability while simultaneously learning higher-order temporal representations. The focus of our approach is a novel spatio-temporal enrichment module that aggregates spatial and temporal cont…

Cited by 154PDFcodeScholar
2022

Video Instance Segmentation via Multi-Scale Spatio-Temporal Split Attention Transformer

ECCV 2022poster

"State-of-the-art transformer-based video instance segmentation (VIS) approaches typically utilize either single-scale spatio-temporal features or per-frame multi-scale features during the attention computations. We argue that such an attention computation ignores the multi-scale spatio-temporal fea…

2021

Handwriting Transformers

ICCV 2021poster

We propose a novel transformer-based styled handwritten text image generation approach, HWT, that strives to learn both style-content entanglement as well as global and local style patterns. The proposed HWT captures the long and short range relationships within the style examples through a self-att…

Cited by 74PDFcodeScholar
2020

Count- and Similarity-aware R-CNN for Pedestrian Detection

ECCV 2020poster

Recent pedestrian detection methods generally rely on additional supervision, such as visible bounding-box annotations, to handle heavy occlusions. We propose an approach that leverages pedestrian count and proposal similarity information within a two-stage pedestrian detection framework. Both pedes…

2020

D2Det: Towards High Quality Object Detection and Instance Segmentation

CVPR 2020poster

We propose a novel two-stage detection method, D2Det, that collectively addresses both precise localization and accurate classification. For precise localization, we introduce a dense local regression that predicts multiple dense box offsets for an object proposal. Different from traditional regress…

Cited by 240PDFcodeScholar
2020

SipMask: Spatial Information Preservation for Fast Image and Video Instance Segmentation

ECCV 2020poster

Single-stage instance segmentation approaches have recently gained popularity due to their speed and simplicity, but are still lagging behind in accuracy, compared to two-stage methods. We propose a fast single-stage instance segmentation method, called SipMask, that preserves instance-specific spat…

2019

Deep Contextual Attention for Human-Object Interaction Detection

ICCV 2019poster

Human-object interaction detection is an important and relatively new class of visual relationship detection tasks, essential for deeper scene understanding. Most existing approaches decompose the problem into object localization and interaction recognition. Despite showing progress, these approache…

Cited by 130PDFScholar
2019

Efficient Featurized Image Pyramid Network for Single Shot Detector

CVPR 2019poster

Single-stage object detectors have recently gained popularity due to their combined advantage of high detection accuracy and real-time speed. However, while promising results have been achieved by these detectors on standard-sized objects, their performance on small objects is far from satisfactory.…

Cited by 134PDFScholar
2019

Enriched Feature Guided Refinement Network for Object Detection

ICCV 2019poster

We propose a single-stage detection framework that jointly tackles the problem of multi-scale object detection and class imbalance. Rather than designing deeper networks, we introduce a simple yet effective feature enrichment scheme to produce multi-scale contextual features. We further introduce a…

Cited by 109PDFcodeScholar
2019

Learning Rich Features at High-Speed for Single-Shot Object Detection

ICCV 2019poster

Single-stage object detection methods have received significant attention recently due to their characteristic realtime capabilities and high detection accuracies. Generally, most existing single-stage detectors follow two common practices: they employ a network backbone that is pretrained on ImageN…

Cited by 143PDFcodeScholar
2019

Mask-Guided Attention Network for Occluded Pedestrian Detection

ICCV 2019poster

Pedestrian detection relying on deep convolution neural networks has made significant progress. Though promising results have been achieved on standard pedestrians, the performance on heavily occluded pedestrians remains far from satisfactory. The main culprits are intra-class occlusions involving o…

Cited by 252PDFcodeScholar