← Search

Arash Vahdat

58 accepted papers

2026

Exploring Synthesizable Chemical Space with Iterative Pathway Refinements

ICLR 2026oral

A well-known pitfall of molecular generative models is that they are not guaranteed to generate synthesizable molecules. Existing solutions for this problem often struggle to effectively navigate exponentially large combinatorial space of synthesizable molecules and suffer from poor coverage. To add…

Cited by 0SourcecodeScholar
2026

Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model

ICML 2026spotlight

Recent progress in large-scale generative models has substantially advanced video generation, yet existing methods remain constrained by a rigid inference paradigm. Bidirectional diffusion models excel at global coherence and visual fidelity but suffer from slow inference, while autoregressive model…

Cited by 0SourceScholar
2026

La-Proteina: Atomistic Protein Generation via Partially Latent Flow Matching

ICLR 2026poster

Recently, many generative models for de novo protein structure design have emerged. Yet, only few tackle the difficult task of directly generating fully atomistic structures jointly with the underlying amino acid sequence. This is challenging, for instance, because the model must reason over side ch…

Cited by 0SourcecodeScholar
2026

Mode Seeking meets Mean Seeking for Long Video Generation

ICML 2026poster

Scaling video generation from seconds to minutes faces a critical bottleneck: while short-video data is abundant and high-fidelity, coherent long-form data is scarce and limited to narrow domains. While multi-resolution image training works because higher resolution is largely an interpolation of th…

Cited by 7SourceScholar
2026

Scaling Atomistic Protein Binder Design with Generative Pretraining and Test-Time Compute

ICLR 2026oral

Protein interaction modeling is central to protein design, which has been transformed by machine learning with broad applications in drug discovery and beyond. In this landscape, structure-based de novo binder design is most often cast as either conditional generative modeling or sequence optimizati…

Cited by 0SourcecodeScholar
2026

Transition Matching Distillation for Fast Video Generation

CVPR 2026

Large video diffusion and flow models have achieved remarkable success in high-quality video generation, but their use in real-time interactive applications remains limited due to their inefficient multi-step sampling process. In this work, we present Transition Matching Distillation (TMD), a novel

Cited by 0SourceScholar
2025

Adaptive Flow Matching for Resolving Small-Scale Physics

ICML 2025poster

Conditional diffusion and flow models are effective for super-resolving small-scale details in natural images. However, in physical sciences such as weather, three major challenges arise: (i) spatially misaligned input-output distributions (PDEs at different resolutions lead to divergent trajectorie…

Cited by 0SourcePDFScholar
2025

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations

CVPR 2025poster

Existing video generation models struggle to follow complex text prompts and synthesize multiple objects, raising the need for additional grounding input for improved controllability. In this work, we propose to decompose videos into visual primitives -- blob video representation, a general represen…

Cited by 3SourcePDFScholar
2025

Elucidated Rolling Diffusion Models for Probabilistic Forecasting of Complex Dynamics

NeurIPS 2025poster

Diffusion models are a powerful tool for probabilistic forecasting, yet most applications in high-dimensional complex systems predict future states individually. This approach struggles to model complex temporal dependencies and fails to explicitly account for the progressive growth of uncertainty i…

Cited by 0SourceScholar
2025

Energy-Based Diffusion Language Models for Text Generation

ICLR 2025poster

Despite remarkable progress in autoregressive language models, alternative generative paradigms beyond left-to-right generation are still being actively explored. Discrete diffusion models, with the capacity for parallel generation, have recently emerged as a promising alternative. Unfortunately, th…

2025

GenMol: A Drug Discovery Generalist with Discrete Diffusion

ICML 2025poster

Drug discovery is a complex process that involves multiple stages and tasks. However, existing molecular generative models can only tackle some of these tasks. We present *Generalist Molecular generative model* (GenMol), a versatile framework that uses only a *single* discrete diffusion model to han…

Cited by 3SourcePDFScholar
2025

Heavy-Tailed Diffusion Models

ICLR 2025poster

Diffusion models achieve state-of-the-art generation quality across many applications, but their ability to capture rare or extreme events in heavy-tailed distributions remains unclear. In this work, we show that traditional diffusion and flow-matching models with standard Gaussian priors fail to ca…

Cited by 6SourcePDFScholar
2025

Not-So-Optimal Transport Flows for 3D Point Cloud Generation

ICLR 2025poster

Learning generative models of 3D point clouds is one of the fundamental problems in 3D generative learning. One of the key properties of point clouds is their permutation invariance, i.e., changing the order of points in a point cloud does not change the shape they represent. In this paper, we analy…

Cited by 0SourcePDFScholar
2025

ProtComposer: Compositional Protein Structure Generation with 3D Ellipsoids

ICLR 2025oral

We develop ProtComposer to generate protein structures conditioned on spatial protein layouts that are specified via a set of 3D ellipsoids capturing substructure shapes and semantics. At inference time, we condition on ellipsoids that are hand-constructed, extracted from existing proteins, or from…

2025

Proteina: Scaling Flow-based Protein Structure Generative Models

ICLR 2025oral

Recently, diffusion- and flow-based generative models of protein structures have emerged as a powerful tool for de novo protein design. Here, we develop *Proteina*, a new large-scale flow-based protein backbone generator that utilizes hierarchical fold class labels for conditioning and relies on a t…

2025

Trust Region Constrained Measure Transport in Path Space for Stochastic Optimal Control and Inference

NeurIPS 2025spotlight

Solving stochastic optimal control problems with quadratic control costs can be viewed as approximating a target path space measure, e.g. via gradient-based optimization. In practice, however, this optimization is challenging in particular if the target measure differs substantially from the prior.…

Cited by 0SourceScholar
2024

A Variational Perspective on Solving Inverse Problems with Diffusion Models

ICLR 2024poster

Diffusion models have emerged as a key pillar of foundation models in visual domains. One of their critical applications is to universally solve different downstream inverse tasks via a single diffusion prior without re-training for each task. Most inverse tasks can be formulated as inferring a post…

2024

Aligning Target-Aware Molecule Diffusion Models with Exact Energy Optimization

NeurIPS 2024poster

Generating ligand molecules for specific protein targets, known as structure-based drug design, is a fundamental problem in therapeutics development and biological discovery. Recently, target-aware generative models, especially diffusion models, have shown great promise in modeling protein-ligand in…

2024

Compositional Text-to-Image Generation with Dense Blob Representations

ICML 2024poster

Existing text-to-image models struggle to follow complex text prompts, raising the need for extra grounding inputs for better controllability. In this work, we propose to decompose a scene into visual primitives - denoted as dense blob representations - that contain fine-grained details of the scene…

Cited by 16SourcePDFScholar
2024

DiffiT: Diffusion Vision Transformers for Image Generation

ECCV 2024poster

"Diffusion models with their powerful expressivity and high sample quality have achieved State-Of-The-Art (SOTA) performance in the generative domain. The pioneering Vision Transformer (ViT) has also demonstrated strong modeling capabilities and scalability, especially for recognition tasks. In this…

2024

DisCo-Diff: Enhancing Continuous Diffusion Models with Discrete Latents

ICML 2024poster

Diffusion models (DMs) have revolutionized generative learning. They utilize a diffusion process to encode data into a simple Gaussian distribution. However, encoding a complex, potentially multimodal data distribution into a single *continuous* Gaussian distribution arguably represents an unnecessa…

2024

Molecule Generation with Fragment Retrieval Augmentation

NeurIPS 2024poster

Fragment-based drug discovery, in which molecular fragments are assembled into new molecules with desirable biochemical properties, has achieved great success. However, many fragment-based molecule generation methods show limited exploration beyond the existing fragments in the database as they only…

Cited by 3SourcePDFScholar
2024

Warped Diffusion: Solving Video Inverse Problems with Image Diffusion Models

NeurIPS 2024poster

Using image models naively for solving inverse video problems often suffers from flickering, texture-sticking, and temporal inconsistency in generated videos. To tackle these problems, in this paper, we view frames as continuous functions in the 2D space, and videos as a sequence of continuous warpi…

2023

Fast Sampling of Diffusion Models via Operator Learning

ICML 2023poster

Diffusion models have found widespread adoption in various areas. However, their sampling process is slow because it requires hundreds to thousands of network evaluations to emulate a continuous process defined by differential equations. In this work, we use neural operators, an efficient method to…

2023

I$^2$SB: Image-to-Image Schrödinger Bridge

ICML 2023poster

We propose Image-to-Image Schrödinger Bridge (I$^2$SB), a new class of conditional diffusion models that directly learn the nonlinear diffusion processes between two given distributions. These diffusion bridges are particularly useful for image restoration, as the degraded images are structurally in…

2023

Loss-Guided Diffusion Models for Plug-and-Play Controllable Generation

ICML 2023poster

We consider guiding denoising diffusion models with general differentiable loss functions in a plug-and-play fashion, enabling controllable generation without additional training. This paradigm, termed Loss-Guided Diffusion (LGD), can easily be integrated into all diffusion models and leverage vario…

Cited by 89SourcePDFScholar
2023

Open-Vocabulary Panoptic Segmentation With Text-to-Image Diffusion Models

CVPR 2023highlight

We present ODISE: Open-vocabulary DIffusion-based panoptic SEgmentation, which unifies pre-trained text-image diffusion and discriminative models to perform open-vocabulary panoptic segmentation. Text-to-image diffusion models have the remarkable ability to generate high-quality images with diverse…

2023

Pseudoinverse-Guided Diffusion Models for Inverse Problems

ICLR 2023poster

Diffusion models have become competitive candidates for solving various inverse problems. Models trained for specific inverse problems work well but are limited to their particular use cases, whereas methods that use problem-agnostic models are general but often perform worse empirically. To address…

Cited by 285SourcePDFScholar
2023

Recurrence Without Recurrence: Stable Video Landmark Detection With Deep Equilibrium Models

CVPR 2023poster

Cascaded computation, whereby predictions are recurrently refined over several stages, has been a persistent theme throughout the development of landmark detection models. In this work, we show that the recently proposed Deep Equilibrium Model (DEQ) can be naturally adapted to this form of computati…

2022

A-ViT: Adaptive Tokens for Efficient Vision Transformer

CVPR 2022oral

We introduce A-ViT, a method that adaptively adjusts the inference cost of vision transformer ViT for images of different complexity. A-ViT achieves this by automatically reducing the number of tokens in vision transformers that are processed in the network as inference proceeds. We reformulate Adap…

Cited by 378PDFScholar
2022

Diffusion Models for Adversarial Purification

ICML 2022spotlight

Adversarial purification refers to a class of defense methods that remove adversarial perturbations using a generative model. These methods do not make assumptions on the form of attack and the classification model, and thus can defend pre-existing classifiers against unseen threats. However, their…

2022

LANA: Latency Aware Network Acceleration

ECCV 2022poster

"We introduce latency-aware network acceleration (LANA)-an approach that builds on neural architecture search technique to accelerate neural networks. LANA consists of two phases: in the first phase, it trains many alternative operations for every layer of a target network using layer-wise feature m…

Cited by 17SourcePDFScholar
2022

LION: Latent Point Diffusion Models for 3D Shape Generation

NeurIPS 2022accept

Denoising diffusion models (DDMs) have shown promising results in 3D point cloud synthesis. To advance 3D DDMs and make them useful for digital artists, we require (i) high generation quality, (ii) flexibility for manipulation and applications such as conditional synthesis and shape interpolation, a…

2022

Score-Based Generative Modeling with Critically-Damped Langevin Diffusion

ICLR 2022spotlight

Score-based generative models (SGMs) have demonstrated remarkable synthesis quality. SGMs rely on a diffusion process that gradually perturbs the data towards a tractable distribution, while the generative model learns to denoise. The complexity of this denoising task is, apart from the data distrib…

2022

Tackling the Generative Learning Trilemma with Denoising Diffusion GANs

ICLR 2022spotlight

A wide variety of deep generative models has been developed in the past decade. Yet, these models often struggle with simultaneously addressing three key requirements including: high sample quality, mode coverage, and fast sampling. We call the challenge imposed by these requirements the generative…

2021

A Contrastive Learning Approach for Training Variational Autoencoder Priors

NeurIPS 2021poster

Variational autoencoders (VAEs) are one of the powerful likelihood-based generative models with applications in many domains. However, they struggle to generate high-quality images, especially when samples are obtained from the prior without any tempering. One explanation for VAEs' poor generative q…

Cited by 97SourcePDFScholar
2021

Controllable and Compositional Generation with Latent-Space Energy-Based Models

NeurIPS 2021poster

Controllable generation is one of the key requirements for successful adoption of deep generative models in real-world applications, but it still remains as a great challenge. In particular, the compositional ability to generate novel concept combinations is out of reach for most current models. In…

2021

Don’t Generate Me: Training Differentially Private Generative Models with Sinkhorn Divergence

NeurIPS 2021poster

Although machine learning models trained on massive data have led to breakthroughs in several areas, their deployment in privacy-sensitive domains remains limited due to restricted access to data. Generative models trained with privacy constraints on private data can sidestep this challenge, providi…

Cited by 83SourcePDFScholar
2021

See Through Gradients: Image Batch Recovery via GradInversion

CVPR 2021poster

Training deep neural networks requires gradient estimation from data batches to update parameters. Gradients per parameter are averaged over a set of data and this has been presumed to be safe for privacy-preserving training in joint, collaborative, and federated learning applications. Prior work on…

Cited by 588PDFcodeScholar
2021

VAEBM: A Symbiosis between Variational Autoencoders and Energy-based Models

ICLR 2021spotlight

Energy-based models (EBMs) have recently been successful in representing complex distributions of small images. However, sampling from them requires expensive Markov chain Monte Carlo (MCMC) iterations that mix slowly in high dimensional pixel space. Unlike EBMs, variational autoencoders (VAEs) gene…

2020

Contrastive Learning for Weakly Supervised Phrase Grounding

ECCV 2020poster

Phrase grounding, the problem of associating image regions to caption words, is a crucial component of vision-language tasks. We show that phrase grounding can be learned by optimizing word-region attention to maximize a lower bound on mutual information between images and caption words. Given pairs…

2020

On the distance between two neural networks and the stability of learning

NeurIPS 2020poster

This paper relates parameter distance to gradient breakdown for a broad class of nonlinear compositional functions. The analysis leads to a new distance function called deep relative trust and a descent lemma for neural networks. Since the resulting learning rule seems to require little to no learni…

2020

Semi-Supervised Semantic Image Segmentation With Self-Correcting Networks

CVPR 2020poster

Building a large image dataset with high-quality object masks for semantic segmentation is costly and time-consuming. In this paper, we introduce a principled semi-supervised framework that only use a small set of fully supervised images (having semantic segmentation labels and box labels) and a set…

Cited by 119PDFScholar
2020

UNAS: Differentiable Architecture Search Meets Reinforcement Learning

CVPR 2020oral

Neural architecture search (NAS) aims to discover network architectures with desired properties such as high accuracy or low latency. Recently, differentiable NAS (DNAS) has demonstrated promising results while maintaining a search cost orders of magnitude lower than reinforcement learning (RL) base…

Cited by 44PDFcodeScholar
2019

A Robust Learning Approach to Domain Adaptive Object Detection

ICCV 2019poster

Domain shift is unavoidable in real-world applications of object detection. For example, in self-driving cars, the target domain consists of unconstrained road environments which cannot all possibly be observed in training data. Similarly, in surveillance applications sufficiently representative tra…

Cited by 323PDFScholar
2018

DVAE#: Discrete Variational Autoencoders with Relaxed Boltzmann Priors

NeurIPS 2018poster

Boltzmann machines are powerful distributions that have been shown to be an effective prior over binary latent variables in variational autoencoders (VAEs). However, previous methods for training discrete VAEs have used the evidence lower bound and not the tighter importance-weighted bound. We propo…

2018

DVAE++: Discrete Variational Autoencoders with Overlapping Transformations

ICML 2018oral

Training of discrete latent variable models remains challenging because passing gradient information through discrete units is difficult. We propose a new class of smoothing transformations based on a mixture of two overlapping distributions, and show that the proposed transformation can be used for…

Cited by 98SourcePDFScholar
2016

A Hierarchical Deep Temporal Model for Group Activity Recognition

CVPR 2016poster

In group activity recognition, the temporal dynamics of the whole activity can be inferred based on the dynamics of the individual people representing the activity. We build a deep model to capture these dynamics based on LSTM (long short-term memory) models. To make use of these observations, we pr…

Cited by 642PDFcodeScholar
2016

Structure Inference Machines: Recurrent Neural Networks for Analyzing Relations in Group Activity Recognition

CVPR 2016poster

Rich semantic relations are important in a variety of visual recognition problems. As a concrete example, group activity recognition involves the interactions and relative spatial relations of a set of people in a scene. State of the art recognition methods center on deep learning approaches for tr…

Cited by 308PDFScholar
2015

Visual Recognition by Counting Instances: A Multi-Instance Cardinality Potential Kernel

CVPR 2015poster

Many visual recognition problems can be approached by counting instances. To determine whether an event is present in a long internet video, one could count how many frames seem to contain the activity. Classifying the activity of a group of people can be done by counting the actions of individual…

Cited by 106SourcePDFScholar