← Search

Fahad Shahbaz Khan

127 accepted papers

2026

AURORA: Augmented Understanding via Structured Reasoning and Reinforcement Learning for Reference Audio-Visual Segmentation

AAAI 2026technical

Reference Audio-Visual Segmentation (Ref-AVS) tasks challenge models to precisely locate sounding objects by integrating visual, auditory, and textual cues. Existing methods often lack genuine semantic understanding, tending to memorize fixed reasoning patterns. Furthermore, jointly training for rea

Cited by 0SourcePDFScholar
2026

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks

ICLR 2026poster

Deep reasoning is fundamental for solving complex tasks, especially in vision-centric scenarios that demand sequential, multimodal understanding. However, existing benchmarks typically evaluate agents with fully synthetic, single-turn queries, limited visual modalities, and lack a framework to asses…

Cited by 0SourcecodeScholar
2026

Bring Your Dreams to Life: Continual Text-to-Video Customization

AAAI 2026technical

Customized text-to-video generation (CTVG) has recently witnessed great progress in generating tailored videos from user-specific text. However, most CTVG methods assume that personalized concepts remain static and do not expand incrementally over time. Additionally, they struggle with forgetting an

Cited by 0SourcePDFScholar
2026

TerraFM: A Scalable Foundation Model for Unified Multisensor Earth Observation

ICLR 2026poster

Modern Earth observation (EO) increasingly leverages deep learning to harness the scale and diversity of satellite imagery across sensors and regions. While recent foundation models have demonstrated promising generalization across EO tasks, many remain limited by the scale, geographical coverage, a…

Cited by 0SourcecodeScholar
2026

VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Video

ICLR 2026poster

Mathematical reasoning in real-world video presents a fundamentally different challenge than static images or text. It requires interpreting fine-grained visual information, accurately reading handwritten or digital text, and integrating spoken cues, often dispersed non-linearly over time. In such m…

Cited by 0SourcecodeScholar
2025

$InterLCM$: Low-Quality Images as Intermediate States of Latent Consistency Models for Effective Blind Face Restoration

ICLR 2025poster

Diffusion priors have been used for blind face restoration (BFR) by fine-tuning diffusion models (DMs) on restoration datasets to recover low-quality images. However, the naive application of DMs presents several key limitations. (i) The diffusion prior has inferior semantic consistency (e.g., ID,…

Cited by 1SourcePDFScholar
2025

A Culturally-diverse Multilingual Multimodal Video Benchmark & Model

EMNLP 2025

Large multimodal models (LMMs) have recently gained attention due to their effectiveness to understand and generate descriptions of visual content. Most existing LMMs are in English language. While few recent works explore multilingual image LMMs, to the best of our knowledge, moving beyond the Engl

Cited by 0SourcePDFScholar
2025

ALOcc: Adaptive Lifting-Based 3D Semantic Occupancy and Cost Volume-Based Flow Predictions

ICCV 2025poster

3D semantic occupancy and flow prediction are fundamental to spatiotemporal scene understanding. This paper proposes a vision-based framework with three targeted improvements. First, we introduce an occlusion-aware adaptive lifting mechanism incorporating depth denoising. This enhances the robustnes…

2025

AdaIR: Adaptive All-in-One Image Restoration via Frequency Mining and Modulation

ICLR 2025poster

In the image acquisition process, various forms of degradation, including noise, blur, haze, and rain, are frequently introduced. These degradations typically arise from the inherent limitations of cameras or unfavorable ambient conditions. To recover clean images from their degraded versions, numer…

2025

AgriCLIP: Adapting CLIP for Agriculture and Livestock via Domain-Specialized Cross-Model Alignment

COLING 2025main

Capitalizing on a vast amount of image-text data, large-scale vision-language pre-training has demonstrated remarkable zero-shot capabilities and has been utilized in several applications. However, models trained on general everyday web-crawled data often exhibit sub-optimal performance for speciali…

2025

All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages

CVPR 2025highlight

Existing Large Multimodal Models (LMMs) generally focus on only a few regions and languages. As LMMs continue to improve, it is increasingly important to ensure they understand cultural contexts, respect local sensitivities, and support low-resource languages, all while effectively integrating corr…

2025

Beyond Simple Edits: Composed Video Retrieval with Dense Modifications

ICCV 2025poster

Composed video retrieval is a challenging task that strives to retrieve a target video based on a query video and a textual description detailing specific modifications. Standard retrieval frameworks typically struggle to handle the complexity of fine-grained compositional queries and variations in…

2025

BiMediX2 : Bio-Medical EXpert LMM for Diverse Medical Modalities

EMNLP 2025

We introduce BiMediX2, a bilingual (Arabic-English) Bio-Medical EXpert Large Multimodal Model that supports text-based and image-based medical interactions. It enables multi-turn conversation in Arabic and English and supports diverse medical imaging modalities, including radiology, CT, and histolog

2025

CAMEL-Bench: A Comprehensive Arabic LMM Benchmark

NAACL 2025findings

Recent years have witnessed a significant interest in developing large multi-modal models (LMMs) capable of performing various visual reasoning and understanding tasks. This has led to the introduction of multiple LMM benchmarks to evaluate LMMs on different tasks. However, most existing LMM evaluat…

2025

DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image Models

NeurIPS 2025poster

Efficient fine-tuning of pre-trained Text-to-Image (T2I) models involves adjusting the model to suit a particular task or dataset while minimizing computational resources and limiting the number of trainable parameters. However, it often faces challenges in striking a trade-off between aligning with…

Cited by 0SourcecodeScholar
2025

DriveLMM-o1: A Step-by-Step Reasoning Dataset and Large Multimodal Model for Driving Scenario Understanding

IROS 2025

While large multimodal models (LMMs) have demonstrated strong performance across various Visual Question Answering (VQA) tasks, certain challenges require complex multi-step reasoning to reach accurate answers. One particularly challenging task is autonomous driving, which demands thorough cognitive

Cited by 32SourcecodeScholar
2025

EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues

CVPR 2025poster

Automated analysis of vast Earth observation data via interactive Vision-Language Models (VLMs) can unlock new opportunities for environmental monitoring, disaster response, and resource management. Existing generic VLMs do not perform well on Remote Sensing data, while the recent Geo-spatial VLMs r…

2025

GEOBench-VLM: Benchmarking Vision-Language Models for Geospatial Tasks

ICCV 2025poster

While numerous recent benchmarks focus on evaluating generic Vision-Language Models (VLMs), they do not effectively address the specific challenges of geospatial applications.Generic VLM benchmarks are not designed to handle the complexities of geospatial data, an essential component for application…

2025

GenZSL: Generative Zero-Shot Learning Via Inductive Variational Autoencoder

ICML 2025poster

Remarkable progress in zero-shot learning (ZSL) has been achieved using generative models. However, existing generative ZSL methods merely generate (imagine) the visual features from scratch guided by the strong class semantic vectors annotated by experts, resulting in suboptimal generative performa…

2025

GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing

ICML 2025poster

Recent advances in large multimodal models (LMMs) have recognized fine-grained grounding as an imperative factor of visual understanding and dialogue. However, the benefits of such representation in LMMs are limited to the natural image domain, and these models perform poorly for remote sensing (RS)…

2025

GroupMamba: Efficient Group-Based Visual State Space Model

CVPR 2025poster

State-space models (SSMs) have recently shown promise in capturing long-range dependencies with subquadratic computational complexity, making them attractive for various applications. However, purely SSM-based models face critical challenges related to stability and achieving state-of-the-art perfor…

2025

Hierarchical Visual Prompt Learning for Continual Video Instance Segmentation

ICCV 2025poster

Video instance segmentation (VIS) has gained significant attention for its capability in tracking and segmenting object instances across video frames. However, most of the existing VIS approaches unrealistically assume that the categories of object instances remain fixed over time. Moreover, they ex…

2025

Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model

ICCV 2025poster

Large-scale vision-language models (VLMs), such as CLIP, have achieved remarkable success in zero-shot learning (ZSL) by leveraging large-scale visual-text pair datasets. However, these methods often lack interpretability, as they compute the similarity between an entire query image and the embedded…

2025

KITAB-Bench: A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understanding

ACL 2025finding

With the growing adoption of Retrieval-Augmented Generation (RAG) in document processing, robust text recognition has become increasingly critical for knowledge extraction. While OCR (Optical Character Recognition) for English and other languages benefits from large datasets and well-established ben…

2025

LLMVoX: Autoregressive Streaming Text-to-Speech Model for Any LLM

ACL 2025finding

Recent advancements in speech-to-speech dialogue systems leverage LLMs for multimodal interactions, yet they remain hindered by fine-tuning requirements, high computational overhead, and text-speech misalignment. Existing speech-enabled LLMs often degrade conversational quality by modifying the LLM,…

2025

LawDIS: Language-Window-based Controllable Dichotomous Image Segmentation

ICCV 2025poster

We present LawDIS, a language-window-based controllable dichotomous image segmentation (DIS) framework that produces high-quality object masks. Our framework recasts DIS as an image-conditioned mask generation task within a latent diffusion model, enabling seamless integration of user controls. LawD…

2025

LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

ACL 2025finding

Step-by-step reasoning is crucial for solving complex visual tasks, yet existing approaches lack a comprehensive framework for evaluating this capability and do not emphasize step-wise problem-solving. To this end, we propose a comprehensive framework for advancing multi-step visual reasoning in lar…

2025

MAviS: A Multimodal Conversational Assistant For Avian Species

EMNLP 2025

Fine-grained understanding and species-specific, multimodal question answering are vital for advancing biodiversity conservation and ecological monitoring. However, existing multimodal large language models (MM-LLMs) face challenges when it comes to specialized topics like avian species, making it h

2025

One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt

ICLR 2025spotlight

Text-to-image generation models can create high-quality images from input prompts. However, they struggle to support the consistent generation of identity-preserving requirements for storytelling. Existing approaches to this problem typically require extensive training in large datasets or additiona…

2025

One-Way Ticket: Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion Models

CVPR 2025poster

Text-to-Image (T2I) diffusion models have made remarkable advancements in generative modeling; however, they face a trade-off between inference speed and image quality, posing challenges for efficient deployment. Existing distilled T2I models can generate high-fidelity images with fewer sampling ste…

2025

Open-YOLO 3D: Towards Fast and Accurate Open-Vocabulary 3D Instance Segmentation

ICLR 2025oral

Recent works on open-vocabulary 3D instance segmentation show strong promise but at the cost of slow inference speed and high computation requirements. This high computation cost is typically due to their heavy reliance on aggregated clip features from multi-view, which require computationally expen…

2025

Open3DTrack: Towards Open-Vocabulary 3D Multi-Object Tracking

ICRA 2025

3D multi-object tracking plays a critical role in autonomous driving by enabling the real-time monitoring and prediction of multiple objects' movements. Traditional 3D tracking systems are typically constrained by predefined object categories, limiting their adaptability to novel, unseen objects in

Cited by 3SourcecodeScholar
2025

RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping

ICCV 2025poster

General robotic grasping systems require accurate object affordance perception in diverse open-world scenarios following human instructions. However, current studies suffer from the problem of lacking reasoning-based large-scale affordance prediction data, leading to considerable concern about open-…

Cited by 0SourcePDFScholar
2025

Time Travel: A Comprehensive Benchmark to Evaluate LMMs on Historical and Cultural Artifacts

ACL 2025finding

Understanding historical and cultural artifacts demands human expertise and advanced computational techniques, yet the process remains complex and time-intensive. While large multimodal models offer promising support, their evaluation and improvement require a standardized benchmark. To address this…

2025

VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs

NAACL 2025findings

The recent advancements in Large Language Models (LLMs) have greatly influenced the development of Large Multi-modal Video Models (Video-LMMs), significantly enhancing our ability to interpret and analyze video data. Despite their impressive capabilities, current Video-LMMs have not been evaluated f…

2025

VideoGLaMM : A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

CVPR 2025poster

Fine-grained alignment between videos and text is challenging due to complex spatial and temporal dynamics in videos. Existing video-based Large Multimodal Models (LMMs) handle basic conversations but struggle with precise pixel-level grounding in videos. To address this, we introduce VideoGLaMM, a…

Cited by 4SourcePDFScholar
2025

ZeroDiff: Solidified Visual-semantic Correlation in Zero-Shot Learning

ICLR 2025poster

Zero-shot Learning (ZSL) aims to enable classifiers to identify unseen classes. This is typically achieved by generating visual features for unseen classes based on learned visual-semantic correlations from seen classes. However, most current generative approaches heavily rely on having a sufficient…

2024

BiMediX: Bilingual Medical Mixture of Experts LLM

EMNLP 2024finding

In this paper, we introduce BiMediX, the first bilingual medical mixture of experts LLM designed for seamless interaction in both English and Arabic. Our model facilitates a wide range of medical interactions in English and Arabic, including multi-turn chats to inquire about additional details such…

2024

Bidirectional Reciprocative Information Communication for Few-Shot Semantic Segmentation

ICML 2024poster

Existing few-shot semantic segmentation methods typically rely on a one-way flow of category information from support to query, ignoring the impact of intra-class diversity. To address this, drawing inspiration from cybernetics, we introduce a Query Feedback Branch (QFB) to propagate query informati…

2024

CONDA: Condensed Deep Association Learning for Co-Salient Object Detection.

ECCV 2024poster

"Inter-image association modeling is crucial for co-salient object detection. Despite satisfactory performance, previous methods still have limitations on sufficient inter-image association modeling. Because most of them focus on image feature optimization under the guidance of heuristically calcula…

2024

Composed Video Retrieval via Enriched Context and Discriminative Embeddings

CVPR 2024poster

Composed video retrieval (CoVR) is a challenging prob- lem in computer vision which has recently highlighted the in- tegration of modification text with visual queries for more so- phisticated video search in large databases. Existing works predominantly rely on visual queries combined with modi- fi…

2024

Continual Learning and Unknown Object Discovery in 3D Scenes via Self-Distillation

ECCV 2024poster

"Open-world 3D instance segmentation is a recently introduced problem with diverse applications, notably in continually learning embodied agents. This task involves segmenting unknown instances and learning new instances when their labels are introduced. However, prior research in the open-world dom…

2024

GeoChat: Grounded Large Vision-Language Model for Remote Sensing

CVPR 2024poster

Recent advancements in Large Vision-Language Models (VLMs) have shown great promise in natural image domains allowing users to hold a dialogue about given visual content. However such general-domain VLMs perform poorly for Remote Sensing (RS) scenarios leading to inaccurate or fabricated information…

2024

Learning Camouflaged Object Detection from Noisy Pseudo Label

ECCV 2024poster

"Existing Camouflaged Object Detection (COD) methods rely heavily on large-scale pixel-annotated training sets, which are both time-consuming and labor-intensive. Although weakly supervised methods offer higher annotation efficiency, their performance is far behind due to the unclear visual demarcat…

2024

Long-Tailed 3D Semantic Segmentation with Adaptive Weight Constraint and Sampling

ICRA 2024poster

Existing 3D understanding datasets typically provide annotations for a limited number of object classes, with sufficient examples per class. However, real-world object classes are not equally represented in practical settings, leading to poor performance on rarely-occurring categories if the class i…

Cited by 0SourceScholar
2024

Progressive Semantic-Guided Vision Transformer for Zero-Shot Learning

CVPR 2024poster

Zero-shot learning (ZSL) recognizes the unseen classes by conducting visual-semantic interactions to transfer semantic knowledge from seen classes to unseen ones supported by semantic information (e.g. attributes). However existing ZSL methods simply extract visual features using a pre-trained netwo…

2024

Rethinking Transformers Pre-training for Multi-Spectral Satellite Imagery

CVPR 2024poster

Recent advances in unsupervised learning have demonstrated the ability of large vision models to achieve promising results on downstream tasks by pre-training on large amount of unlabelled data. Such pre-training techniques have also been explored recently in the remote sensing domain due to the ava…

2024

S3A: Towards Realistic Zero-Shot Classification via Self Structural Semantic Alignment

AAAI 2024technical

Large-scale pre-trained Vision Language Models (VLMs) have proven effective for zero-shot classification. Despite the success, most traditional VLMs-based methods are restricted by the assumption of partial source supervision or ideal target vocabularies, which rarely satisfy the open-world scenario…

2024

SED: A Simple Encoder-Decoder for Open-Vocabulary Semantic Segmentation

CVPR 2024poster

Open-vocabulary semantic segmentation strives to distinguish pixels into different semantic groups from an open set of categories. Most existing methods explore utilizing pre-trained vision-language models in which the key is to adopt the image-level model for pixel-level segmentation task. In this…

2024

Self-Distilled Masked Auto-Encoders are Efficient Video Anomaly Detectors

CVPR 2024poster

We propose an efficient abnormal event detection model based on a lightweight masked auto-encoder (AE) applied at the video frame level. The novelty of the proposed model is threefold. First we introduce an approach to weight tokens based on motion gradients thus shifting the focus from the static b…

2024

Semi-supervised Open-World Object Detection

AAAI 2024technical

Conventional open-world object detection (OWOD) problem setting first distinguishes known and unknown classes and then later incrementally learns the unknown objects when introduced with labels in the subsequent tasks. However, the current OWOD formulation heavily relies on the external human oracle…

2024

VideoGrounding-DINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding

CVPR 2024poster

Video grounding aims to localize a spatio-temporal section in a video corresponding to an input text query. This paper addresses a critical limitation in current video grounding methodologies by introducing an Open-Vocabulary Spatio-Temporal Video Grounding task. Unlike prevalent closed-set approach…

Cited by 13SourcePDFScholar
2024

Visual-Augmented Dynamic Semantic Prototype for Generative Zero-Shot Learning

CVPR 2024poster

Generative Zero-shot learning (ZSL) learns a generator to synthesize visual samples for unseen classes which is an effective way to advance ZSL. However existing generative methods rely on the conditions of Gaussian noise and the predefined semantic prototype which limit the generator only optimized…

Cited by 19SourcePDFScholar
2023

3D Instance Segmentation via Enhanced Spatial and Semantic Supervision

ICCV 2023poster

3D instance segmentation has recently garnered increased attention. Typical deep learning methods adopt point grouping schemes followed by hand-designed geometric clustering. Inspired by the success of transformers for various 3D tasks, newer hybrid approaches have utilized transformer decoders coup…

Cited by 6PDFScholar
2023

3D-Aware Multi-Class Image-to-Image Translation With NeRFs

CVPR 2023poster

Recent advances in 3D-aware generative models (3D-aware GANs) combined with Neural Radiance Fields (NeRF) have achieved impressive results. However no prior works investigate 3D-aware GANs for 3D consistent multi-class image-to-image (3D-aware I2I) translation. Naively using 2D-I2I translation metho…

2023

Bridging Precision and Confidence: A Train-Time Loss for Calibrating Object Detection

CVPR 2023poster

Deep neural networks (DNNs) have enabled astounding progress in several vision-based problems. Despite showing high predictive accuracy, recently, several works have revealed that they tend to provide overconfident predictions and thus are poorly calibrated. The majority of the works addressing the…

2023

Burstormer: Burst Image Restoration and Enhancement Transformer

CVPR 2023poster

On a shutter press, modern handheld cameras capture multiple images in rapid succession and merge them to generate a single image. However, individual frames in a burst are misaligned due to inevitable motions and contain multiple degradations. The challenge is to properly align the successive image…

2023

Discriminative Co-Saliency and Background Mining Transformer for Co-Salient Object Detection

CVPR 2023poster

Most previous co-salient object detection works mainly focus on extracting co-salient cues via mining the consistency relations across images while ignoring the explicit exploration of background regions. In this paper, we propose a Discriminative co-saliency and background Mining Transformer framew…

2023

Fine-Tuned CLIP Models Are Efficient Video Learners

CVPR 2023poster

Large-scale multi-modal training with image-text pairs imparts strong generalization to CLIP model. Since training on a similar scale for videos is infeasible, recent approaches focus on the effective transfer of image-based CLIP to the video domain. In this pursuit, new parametric modules are added…

2023

Gated Multi-Resolution Transfer Network for Burst Restoration and Enhancement

CVPR 2023poster

Burst image processing is becoming increasingly popular in recent years. However, it is a challenging task since individual burst images undergo multiple degradations and often have mutual misalignments resulting in ghosting and zipper artifacts. Existing burst restoration methods usually do not con…

2023

Generative Multiplane Neural Radiance for 3D-Aware Image Generation

ICCV 2023poster

We present a method to efficiently generate 3D-aware high-resolution images that are view-consistent across multiple target views. The proposed multiplane neural radiance model, named GMNR, consists of a novel a-guided view-dependent representation (a-VdR) module for learning view-dependent informat…

Cited by 3PDFcodeScholar
2023

MaPLe: Multi-Modal Prompt Learning

CVPR 2023poster

Pre-trained vision-language (V-L) models such as CLIP have shown excellent generalization ability to downstream tasks. However, they are sensitive to the choice of input text prompts and require careful selection of prompt templates to perform well. Inspired by the Natural Language Processing (NLP)…

2023

Multi-grained Temporal Prototype Learning for Few-shot Video Object Segmentation

ICCV 2023poster

Few-Shot Video Object Segmentation (FSVOS) aims to segment objects in a query video with the same category defined by a few annotated support images. However, this task was seldom explored. In this work, based on IPMT, a state-of-the-art few-shot image segmentation method that combines external supp…

Cited by 11PDFcodeScholar
2023

Person Image Synthesis via Denoising Diffusion Model

CVPR 2023poster

The pose-guided person image generation task requires synthesizing photorealistic images of humans in arbitrary poses. The existing approaches use generative adversarial networks that do not necessarily maintain realistic textures or need dense correspondences that struggle to handle complex deforma…

2023

PromptCAL: Contrastive Affinity Learning via Auxiliary Prompts for Generalized Novel Category Discovery

CVPR 2023poster

Although existing semi-supervised learning models achieve remarkable success in learning with unannotated in-distribution data, they mostly fail to learn on unlabeled data sampled from novel semantic classes due to their closed-set assumption. In this work, we target a pragmatic but under-explored G…

2023

Self-regulating Prompts: Foundational Model Adaptation without Forgetting

ICCV 2023poster

Prompt learning has emerged as an efficient alternative for fine-tuning foundational models, such as CLIP, for various downstream tasks. Conventionally trained using the task-specific objective, i.e., cross-entropy loss, prompts tend to overfit downstream data distributions and find it challenging t…

Cited by 205PDFcodeScholar
2023

SwiftFormer: Efficient Additive Attention for Transformer-based Real-time Mobile Vision Applications

ICCV 2023poster

Self-attention has become a defacto choice for capturing global context in various vision applications. However, its quadratic computational complexity with respect to image resolution limits its use in real-time applications, especially for deployment on resource-constrained mobile devices. Althoug…

Cited by 143PDFcodeScholar
2023

Video-FocalNets: Spatio-Temporal Focal Modulation for Video Action Recognition

ICCV 2023poster

Recent video recognition models utilize Transformer models for long-range spatio-temporal context modeling. Video transformer designs are based on self-attention that can model global context at a high computational cost. In comparison, convolutional designs for videos offer an efficient alternative…

Cited by 28PDFcodeScholar
2023

Vita-CLIP: Video and Text Adaptive CLIP via Multimodal Prompting

CVPR 2023poster

Adopting contrastive image-text pretrained models like CLIP towards video classification has gained attention due to its cost-effectiveness and competitive performance. However, recent works in this area face a trade-off. Finetuning the pretrained model to achieve strong supervised performance resul…

2022

Burst Image Restoration and Enhancement

CVPR 2022oral

Modern handheld devices can acquire burst image sequence in a quick succession. However, the individual acquired frames suffer from multiple degradations and are misaligned due to camera shake and object motions. The goal of Burst Image Restoration is to effectively combine complimentary cues across…

Cited by 130PDFcodeScholar
2022

Class-Agnostic Object Detection with Multi-modal Transformer

ECCV 2022poster

"What constitutes an object? This has been a long-standing question in computer vision. Towards this goal, numerous learning-free and learning-based approaches have been developed to score objectness. However, they generally do not scale well across new domains and for unseen objects. In this paper,…

2022

Dense Gaussian Processes for Few-Shot Segmentation

ECCV 2022poster

"Few-shot segmentation is a challenging dense prediction task, which entails segmenting a novel query image given only a small annotated support set. The key problem is thus to design a method that aggregates detailed information from the support set, while being robust to large variations in appear…

2022

DoodleFormer: Creative Sketch Drawing with Transformers

ECCV 2022poster

"Creative sketching or doodling is an expressive activity, where imaginative and previously unseen depictions of everyday visual objects are drawn. Creative sketch image generation is a challenging vision problem, where the task is to generate diverse, yet realistic creative sketches possessing the…

2022

Energy-Based Latent Aligner for Incremental Learning

CVPR 2022poster

Deep learning models tend to forget their earlier knowledge while incrementally learning new tasks. This behavior emerges because the parameter updates optimized for the new tasks may not align well with the updates suitable for older tasks. The resulting latent representation mismatch causes forget…

Cited by 56PDFcodeScholar
2022

OW-DETR: Open-World Detection Transformer

CVPR 2022poster

Open-world object detection (OWOD) is a challenging computer vision problem, where the task is to detect a known set of object categories while simultaneously identifying unknown objects. Additionally, the model must incrementally learn new classes that become known in the next training episodes. Di…

Cited by 240PDFcodeScholar
2022

OpenLDN: Learning to Discover Novel Classes for Open-World Semi-Supervised Learning

ECCV 2022poster

"Semi-supervised learning (SSL) is one of the dominant approaches to address the annotation bottleneck of supervised learning. Recent SSL methods can effectively leverage a large repository of unlabeled data to improve performance while relying on a small set of labeled data. One common assumption i…

2022

PSTR: End-to-End One-Step Person Search With Transformers

CVPR 2022poster

We propose a novel one-step transformer-based person search framework, PSTR, that jointly performs person detection and re-identification (re-id) in a single architecture. PSTR comprises a person search-specialized (PSS) module that contains a detection encoder-decoder for person detection along wit…

Cited by 75PDFcodeScholar
2022

Restormer: Efficient Transformer for High-Resolution Image Restoration

CVPR 2022oral

Since convolutional neural networks (CNNs) perform well at learning generalizable image priors from large-scale data, these models have been extensively applied to image restoration and related tasks. Recently, another class of neural architectures, Transformers, have shown significant performance g…

Cited by 3074PDFcodeScholar
2022

Self-Supervised Predictive Convolutional Attentive Block for Anomaly Detection

CVPR 2022oral

Anomaly detection is commonly pursued as a one-class classification problem, where models can only learn from normal training samples, while being evaluated on both normal and abnormal test samples. Among the successful approaches for anomaly detection, a distinguished category of methods relies on…

Cited by 283PDFcodeScholar
2022

Self-Supervised Video Transformer

CVPR 2022oral

In this paper, we propose self-supervised training for video transformers using unlabeled video data. From a given video, we create local and global spatiotemporal views with varying spatial sizes and frame rates. Our self-supervised objective seeks to match the features of these different views rep…

Cited by 131PDFcodeScholar
2022

Spatio-Temporal Relation Modeling for Few-Shot Action Recognition

CVPR 2022poster

We propose a novel few-shot action recognition framework, STRM, which enhances class-specific feature discriminability while simultaneously learning higher-order temporal representations. The focus of our approach is a novel spatio-temporal enrichment module that aggregates spatial and temporal cont…

Cited by 154PDFcodeScholar
2022

UBnormal: New Benchmark for Supervised Open-Set Video Anomaly Detection

CVPR 2022poster

Detecting abnormal events in video is commonly framed as a one-class classification task, where training videos contain only normal events, while test videos encompass both normal and abnormal events. In this scenario, anomaly detection is an open-set problem. However, some studies assimilate anomal…

Cited by 177PDFcodeScholar
2022

Video Instance Segmentation via Multi-Scale Spatio-Temporal Split Attention Transformer

ECCV 2022poster

"State-of-the-art transformer-based video instance segmentation (VIS) approaches typically utilize either single-scale spatio-temporal features or per-frame multi-scale features during the attention computations. We argue that such an attention computation ignores the multi-scale spatio-temporal fea…

2021

Anomaly Detection in Video via Self-Supervised and Multi-Task Learning

CVPR 2021poster

Anomaly detection in video is a challenging computer vision problem. Due to the lack of anomalous events at training time, anomaly detection requires the design of learning methods without full supervision. In this paper, we approach anomalous event detection in video through self-supervised and mul…

Cited by 383PDFScholar
2021

D2-Net: Weakly-Supervised Action Localization via Discriminative Embeddings and Denoised Activations

ICCV 2021poster

This work proposes a weakly-supervised temporal action localization framework, called D2-Net, which strives to temporally localize actions using video-level supervision. Our main contribution is the introduction of a novel loss formulation, which jointly enhances the discriminability of latent embed…

Cited by 76PDFcodeScholar
2021

Discriminative Region-Based Multi-Label Zero-Shot Learning

ICCV 2021poster

Multi-label zero-shot learning (ZSL) is a more realistic counter-part of standard single-label ZSL since several objects can co-exist in a natural image. However, the occurrence of multiple objects complicates the reasoning and requires region-specific processing of visual features to preserve their…

Cited by 59PDFcodeScholar
2021

Exploring Complementary Strengths of Invariant and Equivariant Representations for Few-Shot Learning

CVPR 2021poster

In many real-world problems, collecting a large number of labeled samples is infeasible. Few-shot learning (FSL) is the dominant approach to address this issue, where the objective is to quickly adapt to novel categories in presence of a limited number of samples. FSL tasks have been predominantly s…

Cited by 159PDFcodeScholar
2021

Handwriting Transformers

ICCV 2021poster

We propose a novel transformer-based styled handwritten text image generation approach, HWT, that strives to learn both style-content entanglement as well as global and local style patterns. The proposed HWT captures the long and short range relationships within the style examples through a self-att…

Cited by 74PDFcodeScholar
2021

Learning To Fuse Asymmetric Feature Maps in Siamese Trackers

CVPR 2021poster

Recently, Siamese-based trackers have achieved promising performance in visual tracking. Most recent Siamese-based trackers typically employ a depth-wise cross-correlation (DW-XCorr) to obtain multi-channel correlation information from the two feature maps (target and search region). However, DW-XCo…

Cited by 97PDFcodeScholar
2021

Multi-Stage Progressive Image Restoration

CVPR 2021poster

Image restoration tasks demand a complex balance between spatial details and high-level contextualized information while recovering images. In this paper, we propose a novel synergistic design that can optimally balance these competing goals. Our main proposal is a multi-stage architecture, that pro…

Cited by 2028PDFcodeScholar
2021

On Generating Transferable Targeted Perturbations

ICCV 2021poster

While the untargeted black-box transferability of adversarial perturbations has been extensively studied before, changing an unseen model's decisions to a specific `targeted' class remains a challenging feat. In this paper, we propose a new generative approach for highly transferable targeted pertur…

Cited by 92PDFcodeScholar
2020

A Self-supervised Approach for Adversarial Robustness

CVPR 2020oral

Adversarial examples can cause catastrophic mistakes in Deep Neural Network (DNNs) based vision systems e.g., for classification, segmentation and object detection. The vulnerability of DNNs against such attacks can prove a major roadblock towards their real-world deployment. Transferability of adve…

Cited by 346PDFcodeScholar
2020

AnimalWeb: A Large-Scale Hierarchical Dataset of Annotated Animal Faces

CVPR 2020poster

Several studies show that animal needs are often expressed through their faces. Though remarkable progress has been made towards the automatic understanding of human faces, this has not been the case with animal faces. There exists significant room for algorithmic advances that could realize automat…

Cited by 57PDFScholar
2020

Count- and Similarity-aware R-CNN for Pedestrian Detection

ECCV 2020poster

Recent pedestrian detection methods generally rely on additional supervision, such as visible bounding-box annotations, to handle heavy occlusions. We propose an approach that leverages pedestrian count and proposal similarity information within a two-stage pedestrian detection framework. Both pedes…

2020

CycleISP: Real Image Restoration via Improved Data Synthesis

CVPR 2020oral

The availability of large-scale datasets has helped unleash the true potential of deep convolutional neural networks (CNNs). However, for the single-image denoising problem, capturing a real dataset is an unacceptably expensive and cumbersome procedure. Consequently, image denoising algorithms are m…

Cited by 450PDFcodeScholar
2020

D2Det: Towards High Quality Object Detection and Instance Segmentation

CVPR 2020poster

We propose a novel two-stage detection method, D2Det, that collectively addresses both precise localization and accurate classification. For precise localization, we introduce a dense local regression that predicts multiple dense box offsets for an object proposal. Different from traditional regress…

Cited by 240PDFcodeScholar
2020

Fixing Localization Errors to Improve Image Classification

ECCV 2020poster

Deep neural networks are generally considered black-box models that offer less interpretability for their decision process. To address this limitation, Class Activation Map (CAM) provides an attractive solution that visualizes class-specific discriminative regions in an input image. The remarkable a…

2020

Latent Embedding Feedback and Discriminative Features for Zero-Shot Classification

ECCV 2020poster

Zero-shot learning strives to classify unseen categories for which no data is available during training. In the generalized variant, the test samples can further belong to seen or unseen categories. The state-of-the-art relies on Generative Adversarial Networks that synthesize unseen class features…

2020

Learning Enriched Features for Real Image Restoration and Enhancement

ECCV 2020poster

With the goal of recovering high-quality image content from its degraded version, image restoration enjoys numerous applications, such as in surveillance, computational photography and medical imaging. Recently, convolutional neural networks (CNNs) have achieved dramatic improvements over convention…

2020

Learning Fast and Robust Target Models for Video Object Segmentation

CVPR 2020oral

Video object segmentation (VOS) is a highly challenging problem since the initial mask, defining the target object, is only given at test-time. The main difficulty is to effectively handle appearance changes and similar background objects, while maintaining accurate segmentation. Most previous appro…

Cited by 168PDFcodeScholar
2020

Learning Human-Object Interaction Detection Using Interaction Points

CVPR 2020poster

Understanding interactions between humans and objects is one of the fundamental problems in visual classification and an essential step towards detailed scene understanding. Human-object interaction (HOI) detection strives to localize both the human and an object as well as the identification of com…

Cited by 297PDFcodeScholar
2020

MineGAN: Effective Knowledge Transfer From GANs to Target Domains With Few Images

CVPR 2020poster

One of the attractive characteristics of deep neural networks is their ability to transfer knowledge obtained in one domain to other related domains. As a result, high-quality networks can be trained in domains with relatively little training data. This property has been extensively studied for disc…

Cited by 229PDFcodeScholar
2020

Semi-Supervised Learning for Few-Shot Image-to-Image Translation

CVPR 2020poster

In the last few years, unpaired image-to-image translation has witnessed Remarkable progress. Although the latest methods are able to generate realistic images, they crucially rely on a large number of labeled images. Recently, some methods have tackled the challenging setting of few-shot image-to-i…

Cited by 62PDFcodeScholar
2020

SipMask: Spatial Information Preservation for Fast Image and Video Instance Segmentation

ECCV 2020poster

Single-stage instance segmentation approaches have recently gained popularity due to their speed and simplicity, but are still lagging behind in accuracy, compared to two-stage methods. We propose a fast single-stage instance segmentation method, called SipMask, that preserves instance-specific spat…

2020

iTAML: An Incremental Task-Agnostic Meta-learning Approach

CVPR 2020poster

Humans can continuously learn new knowledge as their experience grows. In contrast, previous learning in deep neural networks can quickly fade out when they are trained on a new task. In this paper, we hypothesize this problem can be avoided by learning a set of generalized parameters, that are neit…

Cited by 205PDFcodeScholar
2019

3C-Net: Category Count and Center Loss for Weakly-Supervised Action Localization

ICCV 2019poster

Temporal action localization is a challenging computer vision problem with numerous real-world applications. Most existing methods require laborious frame-level supervision to train action localization models. In this work, we propose a framework, called 3C-Net, which only requires video-level super…

Cited by 207PDFcodeScholar
2019

A Generative Appearance Model for End-To-End Video Object Segmentation

CVPR 2019oral

One of the fundamental challenges in video object segmentation is to find an effective representation of the target and background appearance. The best performing approaches resort to extensive fine-tuning of a convolutional neural network for this purpose. Besides being prohibitively expensive, thi…

Cited by 236PDFScholar
2019

ATOM: Accurate Tracking by Overlap Maximization

CVPR 2019oral

While recent years have witnessed astonishing improvements in visual tracking robustness, the advancements in tracking accuracy have been limited. As the focus has been directed towards the development of powerful classifiers, the problem of accurate target state estimation has been largely overlook…

Cited by 1659PDFcodeScholar
2019

Cross-Domain Transferability of Adversarial Perturbations

NeurIPS 2019poster

Adversarial examples reveal the blind spots of deep neural networks (DNNs) and represent a major concern for security-critical applications. The transferability of adversarial examples makes real-world attacks possible in black-box settings, where the attacker is forbidden to access the internal par…

2019

Deep Contextual Attention for Human-Object Interaction Detection

ICCV 2019poster

Human-object interaction detection is an important and relatively new class of visual relationship detection tasks, essential for deeper scene understanding. Most existing approaches decompose the problem into object localization and interaction recognition. Despite showing progress, these approache…

Cited by 130PDFScholar
2019

Efficient Featurized Image Pyramid Network for Single Shot Detector

CVPR 2019poster

Single-stage object detectors have recently gained popularity due to their combined advantage of high detection accuracy and real-time speed. However, while promising results have been achieved by these detectors on standard-sized objects, their performance on small objects is far from satisfactory.…

Cited by 134PDFScholar
2019

Enriched Feature Guided Refinement Network for Object Detection

ICCV 2019poster

We propose a single-stage detection framework that jointly tackles the problem of multi-scale object detection and class imbalance. Rather than designing deeper networks, we introduce a simple yet effective feature enrichment scheme to produce multi-scale contextual features. We further introduce a…

Cited by 109PDFcodeScholar
2019

Learning Rich Features at High-Speed for Single-Shot Object Detection

ICCV 2019poster

Single-stage object detection methods have received significant attention recently due to their characteristic realtime capabilities and high detection accuracies. Generally, most existing single-stage detectors follow two common practices: they employ a network backbone that is pretrained on ImageN…

Cited by 143PDFcodeScholar
2019

Learning the Model Update for Siamese Trackers

ICCV 2019poster

Siamese approaches address the visual tracking problem by extracting an appearance template from the current frame, which is used to localize the target in the next frame. In general, this template is linearly combined with the accumulated template from the previous frame, resulting in an exponentia…

Cited by 461PDFcodeScholar
2019

Mask-Guided Attention Network for Occluded Pedestrian Detection

ICCV 2019poster

Pedestrian detection relying on deep convolution neural networks has made significant progress. Though promising results have been achieved on standard pedestrians, the performance on heavily occluded pedestrians remains far from satisfactory. The main culprits are intra-class occlusions involving o…

Cited by 252PDFcodeScholar
2019

Object Counting and Instance Segmentation With Image-Level Supervision

CVPR 2019poster

Common object counting in a natural scene is a challenging problem in computer vision with numerous real-world applications. Existing image-level supervised common object counting approaches only predict the global object count and rely on additional instance-level supervision to also determine obje…

Cited by 146PDFcodeScholar
2019

Object-Centric Auto-Encoders and Dummy Anomalies for Abnormal Event Detection in Video

CVPR 2019poster

Abnormal event detection in video is a challenging vision problem. Most existing approaches formulate abnormal event detection as an outlier detection task, due to the scarcity of anomalous data during training. Because of the lack of prior information regarding abnormal events, these methods are no…

Cited by 478PDFScholar
2019

Out-Of-Distribution Detection for Generalized Zero-Shot Action Recognition

CVPR 2019poster

Generalized zero-shot action recognition is a challenging problem, where the task is to recognize new action categories that are unavailable during the training stage, in addition to the seen action categories. Existing approaches suffer from the inherent bias of the learned classifier towards the s…

Cited by 193PDFcodeScholar
2019

Random Path Selection for Continual Learning

NeurIPS 2019poster

Incremental life-long learning is a main challenge towards the long-standing goal of Artificial General Intelligence. In real-life settings, learning tasks arrive in a sequence and machine learning models must continually learn to increment already acquired knowledge. The existing incremental learni…

2018

Density Adaptive Point Set Registration

CVPR 2018poster

Probabilistic methods for point set registration have demonstrated competitive results in recent years. These techniques estimate a probability distribution model of the point clouds. While such a representation has shown promise, it is highly sensitive to variations in the density of 3D points. Thi…

2018

Unveiling the Power of Deep Tracking

ECCV 2018poster

In the field of generic object tracking numerous attempts have been made to exploit deep features. Despite all expectations, deep trackers are yet to reach an outstanding level of performance compared to methods solely based on handcrafted features. In this paper, we investigate this key issue and p…

Cited by 605SourcePDFScholar
2016

A Probabilistic Framework for Color-Based Point Set Registration

CVPR 2016poster

In recent years, sensors capable of measuring both color and depth information have become increasingly popular. Despite the abundance of colored point set data, state-of-the-art probabilistic registration techniques ignore the available color information. In this paper, we propose a probabilistic p…

Cited by 50PDFScholar
2016

Adaptive Decontamination of the Training Set: A Unified Formulation for Discriminative Visual Tracking

CVPR 2016poster

Tracking-by-detection methods have demonstrated competitive performance in recent years. In these approaches, the tracking model heavily relies on the quality of the training set. Due to the limited amount of labeled training data, additional samples need to be extracted and labeled by the tracker i…

Cited by 515PDFScholar
2015

Learning Spatially Regularized Correlation Filters for Visual Tracking

ICCV 2015poster

Robust and accurate visual tracking is one of the most challenging computer vision problems. Due to the inherent lack of training data, a robust approach for constructing a target appearance model is crucial. Recently, discriminatively learned correlation filters (DCF) have been successfully applied…

Cited by 2649PDFScholar