← Search

Ricky T. Q. Chen

46 accepted papers

2026

Enhancing Diffusion-Based Sampling with Molecular Collective Variables

ICLR 2026poster

Diffusion-based samplers learn to sample complex, high-dimensional distributions using energies or log densities alone, without training data. Yet, they remain impractical for molecular sampling because they are often slower than molecular dynamics and miss thermodynamically relevant modes. Inspired…

Cited by 0SourcecodeScholar
2026

Flowception: Temporally Expansive Flow Matching for Video Generation

CVPR 2026

We present Flowception, a novel non-autoregressive and variable-length video generation framework. Flowception learns a probability path that interleaves discrete frame insertions with continuous frame denoising. Compared to autoregressive methods, Flowception alleviates error accumulation/drift as

Cited by 0SourcecodeScholar
2026

GLASS Flows: Efficient Inference for Reward Alignment of Flow and Diffusion Models

ICLR 2026oral

The performance of flow matching and diffusion models can be greatly improved at inference time using reward adaptation algorithms, yet efficiency remains a major limitation. While several algorithms were proposed, we demonstrate that a common bottleneck is the *sampling* method these algorithms rel…

Cited by 0SourceScholar
2025

Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal Control

ICLR 2025spotlight

Dynamical generative models that produce samples through an iterative process, such as Flow Matching and denoising diffusion models, have seen widespread use, but there have not been many theoretically-sound methods for improving these models with reward fine-tuning. In this work, we cast reward fin…

Cited by 30SourcePDFScholar
2025

Adjoint Sampling: Highly Scalable Diffusion Samplers via Adjoint Matching

ICML 2025poster

We introduce Adjoint Sampling, a highly scalable and efficient algorithm for learning diffusion processes that sample from unnormalized densities, or energy functions. It is the first on-policy approach that allows significantly more gradient updates than the number of energy evaluations and model s…

2025

Adjoint Schrödinger Bridge Sampler

NeurIPS 2025oral

Computational methods for learning to sample from the Boltzmann distribution—where the target distribution is known only up to an unnormalized energy function—have advanced significantly recently. Due to the lack of explicit target samples, however, prior diffusion-based methods, known as _diffusion…

Cited by 0SourcecodeScholar
2025

Edit Flows: Variable Length Discrete Flow Matching with Sequence-Level Edit Operations

NeurIPS 2025poster

Autoregressive generative models naturally generate variable-length sequences, while non-autoregressive models struggle, often imposing rigid, token-wise structures. We propose Edit Flows, a non-autoregressive model that overcomes these limitations by defining a discrete flow over sequences through…

Cited by 0SourceScholar
2025

Flow Matching with General Discrete Paths: A Kinetic-Optimal Perspective

ICLR 2025oral

The design space of discrete-space diffusion or flow generative models are significantly less well-understood than their continuous-space counterparts, with many works focusing only on a simple masked construction. In this work, we aim to take a holistic approach to the construction of discrete gene…

Cited by 4SourcePDFScholar
2025

FlowDec: A flow-based full-band general audio codec with high perceptual quality

ICLR 2025poster

We propose FlowDec, a neural full-band audio codec for general audio sampled at 48 kHz that combines non-adversarial codec training with a stochastic postfilter based on a novel conditional flow matching method. Compared to the prior work ScoreDec which is based on score matching, we generalize from…

2025

Generator Matching: Generative modeling with arbitrary Markov processes

ICLR 2025oral

We introduce Generator Matching, a modality-agnostic framework for generative modeling using arbitrary Markov processes. Generators characterize the infinitesimal evolution of a Markov process, which we leverage for generative modeling in a similar vein to flow matching: we construct conditional gen…

Cited by 0SourcePDFScholar
2025

Simulation-Free Differential Dynamics Through Neural Conservation Laws

UAI 2025

We present a novel simulation-free framework for training continuous-time diffusion processes over very general objective functions. Existing methods typically involve either prescribing the optimal diffusion process—which only works for heavily restricted problem formulations—or require expensive s

Cited by 0SourcePDFScholar
2024

Bespoke Non-Stationary Solvers for Fast Sampling of Diffusion and Flow Models

ICML 2024poster

This paper introduces Bespoke Non-Stationary (BNS) Solvers, a solver distillation approach to improve sample efficiency of Diffusion and Flow models. BNS solvers are based on a family of non-stationary solvers that provably subsumes existing numerical ODE solvers and consequently demonstrate conside…

Cited by 3SourcePDFScholar
2024

Bespoke Solvers for Generative Flow Models

ICLR 2024spotlight

Diffusion or flow-based models are powerful generative paradigms that are notoriously hard to sample as samples are defined as solutions to high-dimensional Ordinary or Stochastic Differential Equations (ODEs/SDEs) which require a large Number of Function Evaluations (NFE) to approximate well. Exist…

Cited by 20SourcePDFScholar
2024

Diffusion Generative Flow Samplers: Improving learning signals through partial trajectory optimization

ICLR 2024poster

We tackle the problem of sampling from intractable high-dimensional density functions, a fundamental task that often appears in machine learning and statistics. We extend recent sampling-based approaches that leverage controlled stochastic processes to model approximate samples from these target de…

2024

FlowLLM: Flow Matching for Material Generation with Large Language Models as Base Distributions

NeurIPS 2024poster

Material discovery is a critical area of research with the potential to revolutionize various fields, including carbon capture, renewable energy, and electronics. However, the immense scale of the chemical space makes it challenging to explore all possible materials experimentally. In this paper, we…

2024

FlowMM: Generating Materials with Riemannian Flow Matching

ICML 2024poster

Crystalline materials are a fundamental component in next-generation technologies, yet modeling their distribution presents unique computational challenges. Of the plausible arrangements of atoms in a periodic lattice only a vanishingly small percentage are thermodynamically stable, which is a key i…

Cited by 34SourcePDFScholar
2024

Generalized Schrödinger Bridge Matching

ICLR 2024poster

Modern distribution matching algorithms for training diffusion or flow models directly prescribe the time evolution of the marginal distributions between two boundary distributions. In this work, we consider a generalized distribution matching setup, where these marginals are only implicitly describ…

2024

Neural Optimal Transport with Lagrangian Costs

UAI 2024poster

We investigate the optimal transport problem between probability measures when the underlying cost function is understood to satisfy a least action principle, also known as a Lagrangian cost. These generalizations are useful when connecting observations from a physical system where the transport dyn…

2024

Stochastic Optimal Control Matching

NeurIPS 2024poster

Stochastic optimal control, which has the goal of driving the behavior of noisy systems, is broadly applicable in science, engineering and artificial intelligence. Our work introduces Stochastic Optimal Control Matching (SOCM), a novel Iterative Diffusion Optimization (IDO) technique for stochastic…

2024

Variational Schrödinger Diffusion Models

ICML 2024poster

Schrödinger bridge (SB) has emerged as the go-to method for optimizing transportation plans in diffusion models. However, SB requires estimating the intractable forward score functions, inevitably resulting in the (costly) implicit training loss based on simulated trajectories. To improve the scalab…

Cited by 9SourcePDFScholar
2023

Flow Matching for Generative Modeling

ICLR 2023top-25%

We introduce a new paradigm for generative modeling built on Continuous Normalizing Flows (CNFs), allowing us to train CNFs at unprecedented scale. Specifically, we present the notion of Flow Matching (FM), a simulation-free approach for training CNFs based on regressing vector fields of fixed condi…

Cited by 1222SourcePDFScholar
2023

Latent State Marginalization as a Low-cost Approach for Improving Exploration

ICLR 2023poster

While the maximum entropy (MaxEnt) reinforcement learning (RL) framework -- often touted for its exploration and robustness capabilities -- is usually motivated from a probabilistic perspective, the use of deep probabilistic models have not gained much traction in practice due to their inherent comp…

2023

Multisample Flow Matching: Straightening Flows with Minibatch Couplings

ICML 2023poster

Simulation-free methods for training continuous-time generative models construct probability paths that go between noise distributions and individual data samples. Recent works, such as Flow Matching, derived paths that are optimal for each data sample. However, these algorithms rely on independent…

Cited by 133SourcePDFScholar
2023

On Kinetic Optimal Probability Paths for Generative Models

ICML 2023poster

Recent successful generative models are trained by fitting a neural network to an a-priori defined tractable probability density path taking noise to training examples. In this paper we investigate the space of Gaussian probability paths, which includes diffusion paths as an instance, and look for a…

Cited by 21SourcePDFScholar
2023

TaskMet: Task-driven Metric Learning for Model Learning

NeurIPS 2023poster

Deep learning models are often used with some downstream task. Models solely trained to achieve accurate predictions may struggle to perform well on the desired downstream tasks. We propose using the task loss to learn a metric which parameterizes a loss to train the model. This approach does not al…

2022

Infinitely Deep Bayesian Neural Networks with Stochastic Differential Equations

AISTATS 2022poster

We perform scalable approximate inference in continuous-depth Bayesian neural networks. In this model class, uncertainty about separate weights in each layer gives hidden units that follow a stochastic differential equation. We demonstrate gradient-based stochastic variational inference in this infi…

Cited by 66SourcePDFScholar
2022

Matching Normalizing Flows and Probability Paths on Manifolds

ICML 2022spotlight

Continuous Normalizing Flows (CNFs) are a class of generative models that transform a prior distribution to a model distribution by solving an ordinary differential equation (ODE). We propose to train CNFs on manifolds by minimizing probability path divergence (PPD), a novel family of divergences be…

Cited by 46SourcePDFScholar
2022

Neural Conservation Laws: A Divergence-Free Perspective

NeurIPS 2022accept

We investigate the parameterization of deep neural networks that by design satisfy the continuity equation, a fundamental conservation law. This is enabled by the observation that any solution of the continuity equation can be represented as a divergence-free vector field. We hence propose building…

2022

Semi-Discrete Normalizing Flows through Differentiable Tessellation

NeurIPS 2022accept

Mapping between discrete and continuous distributions is a difficult task and many have had to resort to heuristical approaches. We propose a tessellation-based approach that directly learns quantization boundaries in a continuous space, complete with exact likelihood evaluations. This is done throu…

2022

Theseus: A Library for Differentiable Nonlinear Optimization

NeurIPS 2022accept

We present Theseus, an efficient application-agnostic open source library for differentiable nonlinear least squares (DNLS) optimization built on PyTorch, providing a common framework for end-to-end structured learning in robotics and vision. Existing DNLS implementations are application specific an…

Cited by 107SourcePDFScholar
2021

"Hey, that’s not an ODE": Faster ODE Adjoints via Seminorms

ICML 2021spotlight

Neural differential equations may be trained by backpropagating gradients via the adjoint method, which is another differential equation typically solved using an adaptive-step-size numerical differential equation solver. A proposed step is accepted if its error, \emph{relative to some norm}, is suf…

2021

Convex Potential Flows: Universal Probability Distributions with Optimal Transport and Convex Optimization

ICLR 2021poster

Flow-based models are powerful tools for designing probabilistic models with tractable density. This paper introduces Convex Potential Flows (CP-Flow), a natural and efficient parameterization of invertible models inspired by the optimal transport (OT) theory. CP-Flows are the gradient map of a stro…

2021

Learning Neural Event Functions for Ordinary Differential Equations

ICLR 2021poster

The existing Neural ODE formulation relies on an explicit knowledge of the termination time. We extend Neural ODEs to implicitly defined termination criteria modeled by neural event functions, which can be chained together and differentiated through. Neural Event ODEs are capable of modeling discret…

2020

SUMO: Unbiased Estimation of Log Marginal Probability for Latent Variable Models

ICLR 2020spotlight

Standard variational lower bounds used to train latent variable models produce biased estimates of most quantities of interest. We introduce an unbiased estimator of the log marginal likelihood and its gradients for latent variable models based on randomized truncation of infinite series. If paramet…

Cited by 32SourceScholar
2020

Scalable Gradients for Stochastic Differential Equations

AISTATS 2020poster

The adjoint sensitivity method scalably computes gradients of solutions to ordinary differential equations. We generalize this method to stochastic differential equations, allowing time-efficient and constant-memory computation of gradients with high-order adaptive solvers. Specifically, we derive a…

2019

FFJORD: Free-Form Continuous Dynamics for Scalable Reversible Generative Models

ICLR 2019oral

A promising class of generative models maps points from a simple distribution to a complex distribution through an invertible neural network. Likelihood-based training of these models requires restricting their architectures to allow cheap computation of Jacobian determinants. Alternati…

Cited by 1014SourcePDFScholar
2019

Invertible Residual Networks

ICML 2019oral

We show that standard ResNet architectures can be made invertible, allowing the same model to be used for classification, density estimation, and generation. Typically, enforcing invertibility requires partitioning dimensions or restricting network architectures. In contrast, our approach only requi…

Cited by 759SourcePDFScholar
2019

Latent Ordinary Differential Equations for Irregularly-Sampled Time Series

NeurIPS 2019poster

Time series with non-uniform intervals occur in many applications, and are difficult to model using standard recurrent neural networks (RNNs). We generalize RNNs to have continuous-time hidden dynamics defined by ordinary differential equations (ODEs), a model we call ODE-RNNs. Furthermore, we use O…

2019

Residual Flows for Invertible Generative Modeling

NeurIPS 2019spotlight

Flow-based generative models parameterize probability distributions through an invertible transformation and can be trained by maximum likelihood. Invertible residual networks provide a flexible family of transformations where only Lipschitz conditions rather than strict architectural constraints ar…

2018

Isolating Sources of Disentanglement in Variational Autoencoders

NeurIPS 2018oral

We decompose the evidence lower bound to show the existence of a term measuring the total correlation between latent variables. We use this to motivate the beta-TCVAE (Total Correlation Variational Autoencoder) algorithm, a refinement and plug-in replacement of the beta-VAE for learning disentangled…

2018

Neural Ordinary Differential Equations

NeurIPS 2018oral

We introduce a new family of deep neural network models. Instead of specifying a discrete sequence of hidden layers, we parameterize the derivative of the hidden state using a neural network. The output of the network is computed using a blackbox differential equation solver. These continuous-depth…