← Search

Vishal M Patel

90 accepted papers

2026

FreeViS: Training-free Video Stylization with Inconsistent References

ICLR 2026poster

Video stylization plays a key role in content creation, but it remains a challenging problem. Naïvely applying image stylization frame-by-frame hurts temporal consistency and reduces style richness. Alternatively, training a dedicated video stylization model typically requires paired video data and…

Cited by 0SourcecodeScholar
2026

MOVi: Training-free Text-conditioned Multi-Object Video Generation

ICASSP 2026oral

Recent advances in diffusion-based text-to-video (T2V) models have demonstrated remarkable progress, but these models still face challenges in generating videos with multiple objects. Most models struggle with accurately capturing complex object interactions, often treating some objects as static ba…

Cited by 0SourcePDFScholar
2026

ProCrop: Learning Aesthetic Image Cropping from Professional Compositions

AAAI 2026technical

Image cropping is crucial for enhancing the visual appeal and narrative impact of photographs, yet existing rule-based and data-driven approaches often lack diversity or require annotated training data. We introduce ProCrop, a retrieval-based method that leverages professional photography to guide c

Cited by 0SourcePDFScholar
2026

RestoreVAR: Visual Autoregressive Generation for All-in-One Image Restoration

ICLR 2026poster

The use of latent diffusion models (LDMs) such as Stable Diffusion has significantly improved the perceptual quality of All-in-One image Restoration (AiOR) methods, while also enhancing their generalization capabilities. However, these LDM-based frameworks suffer from slow inference due to their ite…

Cited by 0SourcecodeScholar
2026

Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object Detection

CVPR 2026

Existing open-vocabulary detectors focus on RGB images and fail to generalize to thermal imagery, where low texture and emissivity variations challenge RGB-based semantics. We present Thermal-Det, the first large language model (LLM) supervised open-vocabulary detector tailored for thermal images. T

Cited by 0SourceScholar
2026

WoW!: World Models in a Closed-Loop World

ICLR 2026oral

Generative world models (WMs) can now simulate worlds with striking visual realism, which naturally raises the question of whether they can endow embodied agents with predictive perception for decision making. Progress on this question has been limited by fragmented evaluation: most existing benchma…

Cited by 0SourcecodeScholar
2025

A Technical Report on “Erasing the Invisible”: The 2024 NeurIPS Competition on Stress Testing Image Watermarks

NeurIPS 2025poster

AI-generated images have become pervasive, raising critical concerns around content authenticity, intellectual property, and the spread of misinformation. Invisible watermarks offer a promising solution for identifying AI-generated images, preserving content provenance without degrading visual quali…

Cited by 0SourceScholar
2025

AWRaCLe: All-Weather Image Restoration Using Visual In-Context Learning

AAAI 2025technical

All-Weather Image Restoration (AWIR) under adverse weather conditions is a challenging task due to the presence of different types of degradations. Prior research in this domain relies on extensive training data but lacks the utilization of additional contextual information for restoration guidance.…

Cited by 1SourcePDFScholar
2025

Distilling Multi-modal Large Language Models for Autonomous Driving

CVPR 2025poster

Autonomous driving demands safe motion planning, especially in critical "long-tail" scenarios. Recent end-to-end autonomous driving systems leverage large language models (LLMs) as planners to improve generalizability to rare events. However, using LLMs at test time introduces high computational cos…

Cited by 4SourcePDFScholar
2025

FaceXFormer: A Unified Transformer for Facial Analysis

ICCV 2025poster

In this work, we introduce FaceXFormer, an end-to-end unified transformer model capable of performing ten facial analysis tasks within a single framework. These tasks include face parsing, landmark detection, head pose estimation, attribute prediction, age, gender, and race estimation, facial expres…

2025

Field-DiT: Diffusion Transformer on Unified Video, 3D, and Game Field Generation

ICLR 2025poster

The probabilistic field models the distribution of continuous functions defined over metric spaces. While these models hold great potential for unifying data generation across various modalities, including images, videos, and 3D geometry, they still struggle with long-context generation beyond simpl…

Cited by 0SourcePDFScholar
2025

Filter Images First, Generate Instructions Later: Pre-Instruction Data Selection for Visual Instruction Tuning

CVPR 2025highlight

Visual instruction tuning (VIT) for large vision-language models (LVLMs) requires training on expansive datasets of image-instruction pairs, which can be costly. Recent efforts in VIT data selection aim to select a small subset of high-quality image-instruction pairs, reducing VIT runtime while main…

2025

GenDeg: Diffusion-based Degradation Synthesis for Generalizable All-In-One Image Restoration

CVPR 2025poster

Deep learning-based models for All-In-One image Restoration (AIOR) have achieved significant advancements in recent years. However, their practical applicability is limited by poor generalization to samples outside the training distribution. This limitation arises primarily from insufficient diversi…

Cited by 2SourcePDFScholar
2025

HarmonySeg: Tubular Structure Segmentation with Deep-Shallow Feature Fusion and Growth-Suppression Balanced Loss

ICCV 2025poster

Accurate segmentation of tubular structures in medical images, such as vessels and airway trees, is crucial for computer-aided diagnosis, radiotherapy, and surgical planning. However, significant challenges exist in algorithm design when faced with diverse sizes, complex topologies, and (often) inco…

Cited by 0SourcePDFScholar
2025

LiDAR Light Scattering Augmentation (LISA): Physics-based Simulation of Adverse Weather Conditions for 3D Object Detection

ICASSP 2025accepted

LiDAR-based object detectors are critical parts of the 3D perception pipeline in autonomous navigation systems such as self-driving cars. However, they are known to be sensitive to adverse weather conditions such as rain, snow and fog due to reduced signal-to-noise ratio (SNR) and signal-to-backgrou…

Cited by 0SourceScholar
2025

Lux Post Facto: Learning Portrait Performance Relighting with Conditional Video Diffusion and a Hybrid Dataset

CVPR 2025poster

Video portrait relighting remains challenging because the results need to be both photorealistic and temporally stable.This typically requires a strong model design that can capture complex facial reflections as well as intensive training on a high-quality paired video dataset, such as dynamic one-l…

Cited by 1SourcePDFScholar
2025

MIRE: Matched Implicit Neural Representations

CVPR 2025poster

Implicit Neural Representations (INRs) are continuous function learners for conventional digital signal representations. With the aid of positional embeddings and/or exhaustively fine-tuned activation functions, INRs have surpassed many limitations of traditional discrete representations. However, e…

Cited by 0SourcePDFScholar
2025

Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning

NeurIPS 2025poster

The remarkable reasoning capability of large language models (LLMs) stems from cognitive behaviors that emerge through reinforcement with verifiable rewards. This work investigates how to transfer this principle to Multimodal LLMs (MLLMs) to unlock advanced visual reasoning. We introduce a two-stage…

Cited by 0SourceScholar
2025

PIN: Prolate Spheroidal Wave Function-based Implicit Neural Representations

ICLR 2025poster

Implicit Neural Representations (INRs) provide a continuous mapping between the coordinates of a signal and the corresponding values. As the performance of INRs heavily depends on the choice of nonlinear-activation functions, there has been a significant focus on encoding explicit signals within INR…

Cited by 0SourcePDFScholar
2025

SINR: Sparsity Driven Compressed Implicit Neural Representations

CVPR 2025poster

Implicit Neural Representations (INRs) are increasingly recognized as a versatile data modality for representing discretized signals, offering benefits such as infinite query resolution and reduced storage requirements. Existing signal compression approaches for INRs typically employ one of two stra…

Cited by 0SourcePDFScholar
2025

STEREO: A Two-Stage Framework for Adversarially Robust Concept Erasing from Text-to-Image Diffusion Models

CVPR 2025highlight

The rapid proliferation of large-scale text-to-image diffusion (T2ID) models has raised serious concerns about their potential misuse in generating harmful content. Although numerous methods have been proposed for erasing undesired concepts from T2ID models, they often provide a false sense of secu…

2025

Scaling Transformer-Based Novel View Synthesis with Models Token Disentanglement and Synthetic Data

ICCV 2025poster

Large transformer-based models have made significant progress in generalizable novel view synthesis (NVS) from sparse input views, generating novel viewpoints without the need for test-time optimization. However, these models are constrained by the limited diversity of publicly available scene datas…

Cited by 0SourcePDFScholar
2025

SyncNoise: Geometrically Consistent Noise Prediction for Instruction-based 3D Editing

AAAI 2025technical

Text-based 2D diffusion models have demonstrated impressive capabilities in image generation and editing. Meanwhile, the 2D diffusion models also exhibit substantial potentials for 3D editing tasks. However, how to achieve consistent edits across multiple viewpoints remains a challenge. While the it…

Cited by 0SourcePDFScholar
2025

The Power of Context: How Multimodality Improves Image Super-Resolution

CVPR 2025poster

Single-image super-resolution (SISR) remains challenging due to the inherent difficulty of recovering fine-grained details and preserving perceptual quality from low-resolution inputs. Existing methods often rely on limited image priors, leading to suboptimal results. We propose a novel approach tha…

Cited by 2SourcePDFScholar
2025

Towards Zero-Shot Anomaly Detection and Reasoning with Multimodal Large Language Models

CVPR 2025highlight

Zero-Shot Anomaly Detection (ZSAD) is an emerging AD paradigm. Unlike the traditional unsupervised AD setting that requires a large number of normal samples to train a model, ZSAD is more practical for handling data-restricted real-world scenarios. Recently, Multimodal Large Language Models (MLLMs)…

2025

UniRes: Universal Image Restoration for Complex Degradations

ICCV 2025poster

Real-world image restoration is hampered by diverse degradations stemming from varying capture conditions, capture devices and post-processing pipelines. Existing works make improvements through simulating those degradations and leveraging image generative priors, however generalization to in-the-wi…

Cited by 0SourcePDFScholar
2024

CoDi: Conditional Diffusion Distillation for Higher-Fidelity and Faster Image Generation

CVPR 2024poster

Large generative diffusion models have revolutionized text-to-image generation and offer immense potential for conditional generation tasks such as image enhancement restoration editing and compositing. However their widespread adoption is hindered by the high computational cost which limits their r…

2024

CrowdDiff: Multi-hypothesis Crowd Density Estimation using Diffusion Models

CVPR 2024poster

Crowd counting is a fundamental problem in crowd analysis which is typically accomplished by estimating a crowd density map and summing over the density values. However this approach suffers from background noise accumulation and loss of density due to the use of broad Gaussian kernels to create the…

2024

Equivariant Spatio-Temporal Self-Supervision for LiDAR Object Detection

ECCV 2024poster

"Popular representation learning methods encourage feature invariance under transformations applied at the input. However, in 3D perception tasks like object localization and segmentation, outputs are naturally equivariant to some transformations, such as rotation. Using pre-training loss functions…

2024

Federated Black-Box Adaptation for Semantic Segmentation

NeurIPS 2024poster

Federated Learning (FL) is a form of distributed learning that allows multiple institutions or clients to collaboratively learn a global model to solve a task. This allows the model to utilize the information from every institute while preserving data privacy. However, recent studies show that the p…

2024

Holo-Relighting: Controllable Volumetric Portrait Relighting from a Single Image

CVPR 2024poster

At the core of portrait photography is the search for ideal lighting and viewpoint. The process often requires advanced knowledge in photography and an elaborate studio setup. In this work we propose Holo-Relighting a volumetric relighting method that is capable of synthesizing novel viewpoints and…

Cited by 12SourcePDFScholar
2024

JeDi: Joint-Image Diffusion Models for Finetuning-Free Personalized Text-to-Image Generation

CVPR 2024poster

Personalized text-to-image generation models enable users to create images that depict their individual possessions in diverse scenes finding applications in various domains. To achieve the personalization capability existing methods rely on finetuning a text-to-image foundation model on a user's cu…

Cited by 20SourcePDFScholar
2024

LQMFormer: Language-aware Query Mask Transformer for Referring Image Segmentation

CVPR 2024poster

Referring Image Segmentation (RIS) aims to segment objects from an image based on a language description. Recent advancements have introduced transformer-based methods that leverage cross-modal dependencies significantly enhancing performance in referring segmentation tasks. These methods are design…

Cited by 12SourcePDFScholar
2024

MonoDiff: Monocular 3D Object Detection and Pose Estimation with Diffusion Models

CVPR 2024poster

3D object detection and pose estimation from a single-view image is challenging due to the high uncertainty caused by the absence of 3D perception. As a solution recent monocular 3D detection methods leverage additional modalities such as stereo image pairs and LiDAR point clouds to enhance image fe…

2024

ReGS: Reference-based Controllable Scene Stylization with Gaussian Splatting

NeurIPS 2024poster

Referenced-based scene stylization that edits the appearance based on a content-aligned reference image is an emerging research area. Starting with a pretrained neural radiance field (NeRF), existing methods typically learn a novel appearance that matches the given style. Despite their effectiveness…

Cited by 2SourcePDFScholar
2024

View-decoupled Transformer for Person Re-identification under Aerial-ground Camera Network

CVPR 2024poster

Existing person re-identification methods have achieved remarkable advances in appearance-based identity association across homogeneous cameras such as ground-ground matching. However as a more practical scenario aerial-ground person re-identification (AGPReID) among heterogeneous cameras has receiv…

2024

Wild-GS: Real-Time Novel View Synthesis from Unconstrained Photo Collections

NeurIPS 2024poster

Photographs captured in unstructured tourist environments frequently exhibit variable appearances and transient occlusions, challenging accurate scene reconstruction and inducing artifacts in novel view synthesis. Although prior approaches have integrated the Neural Radiance Field (NeRF) with additi…

2023

AdaMAE: Adaptive Masking for Efficient Spatiotemporal Learning With Masked Autoencoders

CVPR 2023poster

Masked Autoencoders (MAEs) learn generalizable representations for image, text, audio, video, etc., by reconstructing masked input data from tokens of the visible data. Current MAE approaches for videos rely on random patch, tube, or frame based masking strategies to select these tokens. This paper…

2023

Ambiguous Medical Image Segmentation Using Diffusion Models

CVPR 2023poster

Collective insights from a group of experts have always proven to outperform an individual's best diagnostic for clinical tasks. For the task of medical image segmentation, existing research on AI-based alternatives focuses more on developing models that can imitate the best individual rather than h…

2023

Instance Relation Graph Guided Source-Free Domain Adaptive Object Detection

CVPR 2023poster

Unsupervised Domain Adaptation (UDA) is an effective approach to tackle the issue of domain shift. Specifically, UDA methods try to align the source and target representations to improve generalization on the target domain. Further, UDA methods work under the assumption that the source data is acces…

2023

LightPainter: Interactive Portrait Relighting With Freehand Scribble

CVPR 2023poster

Recent portrait relighting methods have achieved realistic results of portrait lighting effects given a desired lighting representation such as an environment map. However, these methods are not intuitive for user interaction and lack precise lighting control. We introduce LightPainter, a scribble-b…

Cited by 14SourcePDFScholar
2023

Mask-Free OVIS: Open-Vocabulary Instance Segmentation Without Manual Mask Annotations

CVPR 2023poster

Existing instance segmentation models learn task-specific information using manual mask annotations from base (training) categories. These mask annotations require tremendous human effort, limiting the scalability to annotate novel (new) categories. To alleviate this problem, Open-Vocabulary (OV) me…

2023

SceneComposer: Any-Level Semantic Image Synthesis

CVPR 2023highlight

We propose a new framework for conditional image synthesis from semantic layouts of any precision levels, ranging from pure text to a 2D semantic canvas with precise shapes. More specifically, the input layout consists of one or more semantic regions with free-form text descriptions and adjustable p…

2023

Source-free Unsupervised Domain Adaptation for 3D Object Detection in Adverse Weather

ICRA 2023poster

A domain shift exists between the distributions of large scale, outdoor lidar datasets due to being captured using different types of lidar sensors, in different locations, and under varying weather conditions. Inclement weather in particular affects the quality of lidar data, adding artifacts such…

Cited by 26SourcecodeScholar
2023

Spatio-Temporal Pixel-Level Contrastive Learning-Based Source-Free Domain Adaptation for Video Semantic Segmentation

CVPR 2023poster

Unsupervised Domain Adaptation (UDA) of semantic segmentation transfers labeled source knowledge to an unlabeled target domain by relying on accessing both the source and target data. However, the access to source data is often restricted or infeasible in real-world scenarios. Under the source data…

2023

Steered Diffusion: A Generalized Framework for Plug-and-Play Conditional Image Synthesis

ICCV 2023poster

Conditional generative models typically demand large annotated training sets to achieve high-quality synthesis. As a result, there has been significant interest in designing models that perform plug-and-play generation, i.e., to use a predefined or pretrained model, which is not explicitly trained o…

Cited by 13PDFcodeScholar
2023

Unite and Conquer: Plug & Play Multi-Modal Synthesis Using Diffusion Models

CVPR 2023poster

Generating photos satisfying multiple constraints finds broad utility in the content creation industry. A key hurdle to accomplishing this task is the need for paired data consisting of all modalities (i.e., constraints) and their corresponding output. Moreover, existing methods need retraining usin…

2022

ART-SS: An Adaptive Rejection Technique for Semi-Supervised Restoration for Adverse Weather-Affected Images

ECCV 2022poster

"In recent years, convolutional neural network-based single image adverse weather removal methods have achieved significant performance improvements on many benchmark datasets. However, these methods require large amounts of clean-weather degraded image pairs for training, which is often difficult t…

2022

Auto-FedRL: Federated Hyperparameter Optimization for Multi-Institutional Medical Image Segmentation

ECCV 2022poster

"Federated learning (FL) is a distributed machine learning technique that enables collaborative model training while avoiding explicit data sharing. The inherent privacy-preserving property of FL algorithms makes them especially attractive to the medical field. However, in case of heterogeneous clie…

2022

Completely Self-Supervised Crowd Counting via Distribution Matching

ECCV 2022poster

"Dense crowd counting is a challenging task that demands millions of head annotations for training models. Though existing self-supervised approaches could learn good representations, they require some labeled data to map these features to the end task of density estimation. We mitigate this issue w…

2022

HyperTransformer: A Textural and Spectral Feature Fusion Transformer for Pansharpening

CVPR 2022poster

Pansharpening aims to fuse a registered high-resolution panchromatic image (PAN) with a low-resolution hyperspectral image (LR-HSI) to generate an enhanced HSI with high spectral and spatial resolution. Existing pansharpening approaches neglect using an attention mechanism to transfer HR texture fea…

Cited by 144PDFcodeScholar
2022

Learning Feature Decomposition for Domain Adaptive Monocular Depth Estimation

IROS 2022poster

Monocular depth estimation (MDE) has attracted intense study due to its low cost and critical functions for robotic tasks such as localization, mapping and obstacle detection. Supervised approaches have led to great success with the advance of deep learning, but they rely on large quantities of grou…

Cited by 16SourceScholar
2022

SPIN Road Mapper: Extracting Roads from Aerial Images via Spatial and Interaction Space Graph Reasoning for Autonomous Driving

ICRA 2022poster

Road extraction is an essential step in building autonomous navigation systems. Detecting road segments is challenging as they are of varying widths, bifurcated throughout the image, and are often occluded by terrain, cloud, or other weather conditions. Using just convolution neural networks (ConvNe…

Cited by 50SourcecodeScholar
2022

TransWeather: Transformer-Based Restoration of Images Degraded by Adverse Weather Conditions

CVPR 2022poster

Removing adverse weather conditions like rain, fog, and snow from images is an important problem in many applications. Most methods proposed in the literature have been designed to deal with just removing one type of degradation. Recently, a CNN-based method using neural architecture search (All-in-…

Cited by 398PDFcodeScholar
2021

CR-Fill: Generative Image Inpainting With Auxiliary Contextual Reconstruction

ICCV 2021poster

Recent deep generative inpainting methods use attention layers to allow the generator to explicitly borrow feature patches from the known region to complete a missing region. Due to the lack of supervision signals for the correspondence between missing regions and known regions, it may fail to find…

Cited by 159PDFcodeScholar
2021

MeGA-CDA: Memory Guided Attention for Category-Aware Unsupervised Domain Adaptive Object Detection

CVPR 2021poster

Existing approaches for unsupervised domain adaptive object detection perform feature alignment via adversarial training. While these methods achieve reasonable improvements in performance, they typically perform category-agnostic domain alignment, thereby resulting in negative transfer of features.…

Cited by 240PDFScholar
2021

Multi-Institutional Collaborations for Improving Deep Learning-Based Magnetic Resonance Image Reconstruction Using Federated Learning

CVPR 2021poster

Fast and accurate reconstruction of magnetic resonance (MR) images from under-sampled data is important in many clinical applications. In recent years, deep learning-based methods have been shown to produce superior performance on MR image reconstruction. However, these methods require large amounts…

Cited by 192PDFcodeScholar
2020

Generative-Discriminative Feature Representations for Open-Set Recognition

CVPR 2020poster

We address the problem of open-set recognition, where the goal is to determine if a given sample belongs to one of the classes used for training a model (known classes). The main challenge in open-set recognition is to disentangle open-set samples that produce high class activations from known-set s…

Cited by 240PDFcodeScholar
2020

Learning to Count in the Crowd from Limited Labeled Data

ECCV 2020poster

Recent crowd counting approaches have achieved excellent performance. However, they are essentially based on fully supervised paradigm and require large number of annotated samples. Obtaining annotations is an expensive and labour-intensive process. In this work, we focus on reducing the annotation…

Cited by 88SourcePDFScholar
2020

Prior-based Domain Adaptive Object Detection for Hazy and Rainy Conditions

ECCV 2020poster

Adverse weather conditions such as haze and rain corrupt the quality of captured images, which cause detection networks trained on clean images to perform poorly on these corrupted images. To address this issue, we propose an unsupervised prior-based domain adversarial object detection framework for…

Cited by 205SourcePDFScholar
2020

Syn2Real Transfer Learning for Image Deraining Using Gaussian Processes

CVPR 2020oral

Recent CNN-based methods for image deraining have achieved excellent performance in terms of reconstruction error as well as visual quality. However, these methods are limited in the sense that they can be trained only on fully labeled data. Due to various challenges in obtaining real world fully-la…

Cited by 236PDFcodeScholar
2020

Utilizing Patch-level Category Activation Patterns for Multiple Class Novelty Detection

ECCV 2020poster

For any recognition system, the ability to identify novel class samples during inference is an important aspect of the system’s robustness. This problem of detecting novel class samples during inference is commonly referred to as Multiple Class Novelty Detection. In this paper, we propose a novel me…

Cited by 13SourcePDFScholar
2019

Improving the Performance of Unimodal Dynamic Hand-Gesture Recognition With Multimodal Training

CVPR 2019poster

We present an efficient approach for leveraging the knowledge from multiple modalities in training unimodal 3D convolutional neural networks (3D-CNNs) for the task of dynamic hand gesture recognition. Instead of explicitly combining multimodal information, which is commonplace in many state-of-the…

Cited by 212PDFScholar
2019

Pushing the Frontiers of Unconstrained Crowd Counting: New Dataset and Benchmark Method

ICCV 2019poster

In this work, we propose a novel crowd counting network that progressively generates crowd density maps via residual error estimation. The proposed method uses VGG16 as the backbone network and employs density map generated by the final layer as a coarse prediction to refine and generate finer densi…

Cited by 124PDFScholar
2019

Uncertainty Guided Multi-Scale Residual Learning-Using a Cycle Spinning CNN for Single Image De-Raining

CVPR 2019poster

Single image de-raining is an extremely challenging problem since the rainy image may contain rain streaks which may vary in size, direction and density. Previous approaches have attempted to address this problem by leveraging some prior information to remove rain streaks from a single image. One of…

Cited by 373PDFcodeScholar
2018

Fps-Sft: A Multi-Dimensional Sparse Fourier Transform Based on the Fourier Projection-Slice Theorem

ICASSP 2018accepted

We propose a multidimensional sparse Fourier transform inspired by the idea of the Fourier projection-slice theorem, called FPS-SFT. FPS-SFT extracts samples along lines (1-dimensional slices from a multidimensional data cube), which are parameterized by random slopes and offsets. The discrete Fouri…

Cited by 0SourceScholar