← Search

Chongxuan Li

67 accepted papers

2026

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Video Generation

ICML 2026poster

To achieve real-time video generation, current approaches distill pretrained bidirectional video diffusion models into few-step autoregressive (AR) models. This process involves an *architectural gap*, as it converts full attention into causal attention. In this paper, we demonstrate that existing m…

Cited by 0SourceScholar
2026

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

CVPR 2026

In this work, we introduce LLaDA-V, a purely diffusion-based Multimodal Large Language Model (MLLM) that integrates visual instruction tuning with masked diffusion models, representing a departure from the autoregressive paradigms dominant in current multimodal approaches. Built upon LLaDA, a repres

Cited by 0SourcecodeScholar
2026

Principled RL for Diffusion LLMs Emerges from a Sequence-Level Perspective

ICLR 2026poster

Reinforcement Learning (RL) has proven highly effective for autoregressive language models, but adapting these methods to diffusion large language models (dLLMs) presents fundamental challenges. The core difficulty lies in likelihood approximation: while autoregressive models naturally provide token…

Cited by 0SourcecodeScholar
2026

ReFusion: A Diffusion Large Language Model with Parallel Autoregressive Decoding

ICLR 2026poster

Autoregressive models (ARMs) are hindered by slow sequential inference. While masked diffusion models (MDMs) offer a parallel alternative, they suffer from critical drawbacks: high computational overhead from precluding Key-Value (KV) caching, and incoherent generation arising from learning dependen…

Cited by 0SourcecodeScholar
2026

Reinforcing General Reasoning Without Verifiers

ICLR 2026poster

The recent paradigm shift towards training large language models (LLMs) using DeepSeek-R1-Zero-style reinforcement learning (RL) on verifiable rewards has led to impressive advancements in code and mathematical reasoning. However, this methodology is limited to tasks where rule-based answer verifica…

Cited by 0SourcecodeScholar
2026

UltraLLaDA: Scaling the Context Length to 128K for Diffusion Large Language Models

ICLR 2026poster

Diffusion LLMs have attracted growing interest, with plenty of recent work emphasizing their great potential in various downstream tasks; yet the long‑context behavior of diffusion LLMs remains largely uncharted. We present a case study of post‑training techniques for extending the context window of…

Cited by 0SourcecodeScholar
2026

UltraViCo: Breaking Extrapolation Limits in Video Diffusion Transformers

ICLR 2026poster

Despite advances, video diffusion transformers still struggle to generalize beyond their training length, a challenge we term video length extrapolation. We identify two failure modes: model-specific periodic content repetition and a universal quality degradation. Prior works attempt to solve repeti…

Cited by 0SourcecodeScholar
2026

Variational Reasoning for Language Models

ICLR 2026poster

We introduce a **variational reasoning** framework for language models that treats thinking traces as latent variables and optimizes them through variational inference. Starting from the evidence lower bound (ELBO), we extend it to a multi-trace objective for tighter bounds and propose a forward-KL…

Cited by 0SourcecodeScholar
2025

A Theory for Conditional Generative Modeling on Multiple Data Sources

ICML 2025poster

The success of large generative models has driven a paradigm shift, leveraging massive multi-source data to enhance model capabilities. However, the interaction among these sources remains theoretically underexplored. This paper provides a first attempt to fill this gap by rigorously analyzing multi…

2025

Are Images Indistinguishable to Humans Also Indistinguishable to Classifiers?

CVPR 2025poster

The ultimate goal of generative models is to perfectly capture the data distribution. For image generation, common metrics of visual quality (e.g., FID) and the perceived truthfulness of generated images seem to suggest that we are nearing this goal. However, through distribution classification task…

Cited by 2SourcePDFScholar
2025

Effective and Efficient Masked Image Generation Models

ICML 2025poster

Although masked image generation models and masked diffusion models are designed with different motivations and objectives, we observe that they can be unified within a single framework. Building upon this insight, we carefully explore the design space of training and sampling, identifying key facto…

2025

FlexWorld: Progressively Expanding 3D Scenes for Flexible-View Exploration

NeurIPS 2025poster

Generating flexible-view 3D scenes, including 360° rotation and zooming, from single images is challenging due to a lack of 3D data. To this end, we introduce FlexWorld, a novel framework that progressively constructs a persistent 3D Gaussian splatting representation by synthesizing and integrating…

Cited by 0SourcecodeScholar
2025

Improving Long-Text Alignment for Text-to-Image Diffusion Models

ICLR 2025poster

The rapid advancement of text-to-image (T2I) diffusion models has enabled them to generate unprecedented results from given texts. However, as text inputs become longer, existing encoding methods like CLIP face limitations, and aligning the generated images with long texts becomes challenging. To ta…

2025

Masked Diffusion Models as Energy Minimization

NeurIPS 2025poster

We present a systematic theoretical framework that interprets masked diffusion models (MDMs) as solutions to energy minimization problems in discrete optimal transport. Specifically, we prove that three distinct energy formulations—kinetic, conditional kinetic, and geodesic energy—are mathematically…

Cited by 0SourceScholar
2025

RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion Transformers

ICML 2025poster

Recent advancements in video generation have enabled models to synthesize high-quality, minute-long videos. However, generating even longer videos with temporal coherence remains a major challenge and existing length extrapolation methods lead to temporal repetition or motion deceleration. In this w…

Cited by 0SourcePDFScholar
2025

Scaling Diffusion Transformers Efficiently via $\mu$P

NeurIPS 2025poster

Diffusion Transformers have emerged as the foundation for vision generative models, but their scalability is limited by the high cost of hyperparameter (HP) tuning at large scales. Recently, Maximal Update Parametrization ($\mu$P) was proposed for vanilla Transformers, which enables stable HP transf…

Cited by 0SourceScholar
2025

Scaling up Masked Diffusion Models on Text

ICLR 2025poster

Masked diffusion models (MDMs) have shown promise in language modeling, yet their scalability and effectiveness in core language tasks, such as text generation and language understanding, remain underexplored. This paper establishes the first scaling law for MDMs, demonstrating a scaling rate compar…

2025

Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data

ICLR 2025poster

Discrete diffusion models with absorbing processes have shown promise in language modeling. The key quantities to be estimated are the ratios between the marginal probabilities of two transitive states at all timesteps, called the concrete score. In this paper, we reveal that the concrete score in a…

2024

BayesDiff: Estimating Pixel-wise Uncertainty in Diffusion via Bayesian Inference

ICLR 2024poster

Diffusion models have impressive image generation capability, but low-quality generations still exist, and their identification remains challenging due to the lack of a proper sample-wise metric. To address this, we propose BayesDiff, a pixel-wise uncertainty estimator for generations from diffusion…

2024

CRM: Single Image to 3D Textured Mesh with Convolutional Reconstruction Model

ECCV 2024poster

"Feed-forward 3D generative models like the Large Reconstruction Model (LRM) [?] have demonstrated exceptional generation speed. However, the transformer-based methods do not leverage the geometric priors of the triplane component in their architecture, often leading to sub-optimal quality given the…

2024

EquiPocket: an E(3)-Equivariant Geometric Graph Neural Network for Ligand Binding Site Prediction

ICML 2024oral

Predicting the binding sites of target proteins plays a fundamental role in drug discovery. Most existing deep-learning methods consider a protein as a 3D image by spatially clustering its atoms into voxels and then feed the voxelized protein into a 3D CNN for prediction. However, the CNN-based meth…

2024

Identifying and Solving Conditional Image Leakage in Image-to-Video Diffusion Model

NeurIPS 2024poster

Diffusion models have obtained substantial progress in image-to-video generation. However, in this paper, we find that these models tend to generate videos with less motion than expected. We attribute this to the issue called conditional image leakage, where the image-to-video diffusion models (I2V-…

2024

Lost in Translation: Latent Concept Misalignment in Text-to-Image Diffusion Models

ECCV 2024poster

"Advancements in text-to-image diffusion models have broadened extensive downstream practical applications, but such models often encounter misalignment issues between text and image. Taking the generation of a combination of two disentangled concepts as an example, say given the prompt a tea cup of…

2024

Lower Bounds of Uniform Stability in Gradient-Based Bilevel Algorithms for Hyperparameter Optimization

NeurIPS 2024poster

Gradient-based bilevel programming leverages unrolling differentiation (UD) or implicit function theorem (IFT) to solve hyperparameter optimization (HO) problems, and is proven effective and scalable in practice. To understand their generalization behavior, existing works establish upper bounds on…

Cited by 0SourcePDFScholar
2024

On Mesa-Optimization in Autoregressively Trained Transformers: Emergence and Capability

NeurIPS 2024poster

Autoregressively trained transformers have brought a profound revolution to the world, especially with their in-context learning (ICL) ability to address downstream tasks. Recently, several studies suggest that transformers learn a mesa-optimizer during autoregressive (AR) pretraining to implement…

2024

Progressive Image Synthesis from Semantics to Details with Denoising Diffusion GAN

ICASSP 2024accepted

Although denoising diffusion probabilistic models (DDPMs) have shown remarkable progress in image generation, they typically face two main challenges: the time-expensive sampling process and the semantically meaningless latent space, which are often addressed separately in previous works. In particu…

Cited by 0SourceScholar
2024

The Blessing of Randomness: SDE Beats ODE in General Diffusion-based Image Editing

ICLR 2024poster

We present a unified probabilistic formulation for diffusion-based image editing, where a latent variable is edited in a task-specific manner and generally deviates from the corresponding marginal distribution induced by the original stochastic or ordinary differential equation (SDE or ODE). Instead…

2024

Unifying Bayesian Flow Networks and Diffusion Models through Stochastic Differential Equations

ICML 2024poster

Bayesian flow networks (BFNs) iteratively refine the parameters, instead of the samples in diffusion models (DMs), of distributions at various noise levels through Bayesian inference. Owing to its differentiable nature, BFNs are promising in modeling both continuous and discrete data, while simultan…

2023

All Are Worth Words: A ViT Backbone for Diffusion Models

CVPR 2023poster

Vision transformers (ViT) have shown promise in various vision tasks while the U-Net based on a convolutional neural network (CNN) remains dominant in diffusion models. We design a simple and general ViT-based architecture (named U-ViT) for image generation with diffusion models. U-ViT is characteri…

2023

Contrastive Energy Prediction for Exact Energy-Guided Diffusion Sampling in Offline Reinforcement Learning

ICML 2023poster

Guided sampling is a vital approach for applying diffusion models in real-world tasks that embeds human-defined guidance during the sampling procedure. This paper considers a general setting where the guidance is defined by an (unnormalized) energy function. The main challenge for this setting is th…

2023

Deep Generative Modeling on Limited Data with Regularization by Nontransferable Pre-trained Models

ICLR 2023poster

Deep generative models (DGMs) are data-eager because learning a complex model on limited data suffers from a large variance and easily overfits. Inspired by the classical perspective of the bias-variance tradeoff, we propose regularized deep generative model (Reg-DGM), which leverages a nontransfera…

2023

Diffusion Models and Semi-Supervised Learners Benefit Mutually with Few Labels

NeurIPS 2023spotlight

In an effort to further advance semi-supervised generative and classification tasks, we propose a simple yet effective training strategy called *dual pseudo training* (DPT), built upon strong semi-supervised learners and diffusion models. DPT operates in three stages: training a classifier on partia…

2023

Equivariant Energy-Guided SDE for Inverse Molecular Design

ICLR 2023poster

Inverse molecular design is critical in material science and drug discovery, where the generated molecules should satisfy certain desirable properties. In this paper, we propose equivariant energy-guided stochastic differential equations (EEGSDE), a flexible framework for controllable 3D molecule ge…

2023

Gaussian Mixture Solvers for Diffusion Models

NeurIPS 2023poster

Recently, diffusion models have achieved great success in generative tasks. Sampling from diffusion models is equivalent to solving the reverse diffusion stochastic differential equations (SDEs) or the corresponding probability flow ordinary differential equations (ODEs). In comparison, SDE-based so…

2023

On Evaluating Adversarial Robustness of Large Vision-Language Models

NeurIPS 2023poster

Large vision-language models (VLMs) such as GPT-4 have achieved unprecedented performance in response generation, especially with visual inputs, enabling more creative and adaptable interaction than large language models such as ChatGPT. Nonetheless, multimodal generation exacerbates safety concerns…

2023

One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale

ICML 2023poster

This paper proposes a unified diffusion framework (dubbed UniDiffuser) to fit all distributions relevant to a set of multi-modal data in one model. Our key insight is -- learning diffusion models for marginal, conditional, and joint distributions can be unified as predicting the noise in the perturb…

2023

ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation

NeurIPS 2023spotlight

Score distillation sampling (SDS) has shown great promise in text-to-3D generation by distilling pretrained large-scale text-to-image diffusion models, but suffers from over-saturation, over-smoothing, and low-diversity problems. In this work, we propose to model the 3D parameter as a random variabl…

2023

Revisiting Discriminative vs. Generative Classifiers: Theory and Implications

ICML 2023poster

A large-scale deep model pre-trained on massive labeled or unlabeled data transfers well to downstream tasks. Linear evaluation freezes parameters in the pre-trained model and trains a linear classifier separately, which is efficient and attractive for transfer. However, little work has investigated…

2023

Towards Understanding Generalization of Macro-AUC in Multi-label Learning

ICML 2023poster

Macro-AUC is the arithmetic mean of the class-wise AUCs in multi-label learning and is commonly used in practice. However, its theoretical understanding is far lacking. Toward solving it, we characterize the generalization properties of various learning algorithms based on the corresponding surrogat…

2022

Analytic-DPM: an Analytic Estimate of the Optimal Reverse Variance in Diffusion Probabilistic Models

ICLR 2022oral

Diffusion probabilistic models (DPMs) represent a class of powerful generative models. Despite their success, the inference of DPMs is expensive since it generally needs to iterate over thousands of timesteps. A key problem in the inference is to estimate the variance in each timestep of the reverse…

2022

DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 Steps

NeurIPS 2022accept

Diffusion probabilistic models (DPMs) are emerging powerful generative models. Despite their high-quality generation performance, DPMs still suffer from their slow sampling as they generally need hundreds or thousands of sequential function evaluations (steps) of large neural networks to draw a samp…

2022

EGSDE: Unpaired Image-to-Image Translation via Energy-Guided Stochastic Differential Equations

NeurIPS 2022accept

Score-based diffusion models (SBDMs) have achieved the SOTA FID results in unpaired image-to-image translation (I2I). However, we notice that existing methods totally ignore the training data in the source domain, leading to sub-optimal solutions for unpaired I2I. To this end, we propose energy-gui…

2022

Estimating the Optimal Covariance with Imperfect Mean in Diffusion Probabilistic Models

ICML 2022spotlight

Diffusion probabilistic models (DPMs) are a class of powerful deep generative models (DGMs). Despite their success, the iterative generation process over the full timesteps is much less efficient than other DGMs such as GANs. Thus, the generation performance on a subset of timesteps is crucial, whic…

2022

Fast Lossless Neural Compression with Integer-Only Discrete Flows

ICML 2022spotlight

By applying entropy codecs with learned data distributions, neural compressors have significantly outperformed traditional codecs in terms of compression ratio. However, the high inference latency of neural networks hinders the deployment of neural compressors in practical applications. In this work…

2022

Maximum Likelihood Training for Score-based Diffusion ODEs by High Order Denoising Score Matching

ICML 2022spotlight

Score-based generative models have excellent performance in terms of generation quality and likelihood. They model the data distribution by matching a parameterized score network with first-order data score functions. The score network can be used to define an ODE (“score-based diffusion ODE”) for e…

2022

Memory Replay with Data Compression for Continual Learning

ICLR 2022poster

Continual learning needs to overcome catastrophic forgetting of the past. Memory replay of representative old training samples has been shown as an effective solution, and achieves the state-of-the-art (SOTA) performance. However, existing work is mainly built on a small memory buffer containing a f…

2021

MiCE: Mixture of Contrastive Experts for Unsupervised Image Clustering

ICLR 2021poster

We present Mixture of Contrastive Experts (MiCE), a unified probabilistic clustering framework that simultaneously exploits the discriminative representations learned by contrastive learning and the semantic structures captured by a latent mixture model. Motivated by the mixture of experts, MiCE emp…

2021

ORDisCo: Effective and Efficient Usage of Incremental Unlabeled Data for Semi-Supervised Continual Learning

CVPR 2021poster

Continual learning usually assumes the incoming data are fully labeled, which might not be applicable in real applications. In this work, we consider semi-supervised continual learning (SSCL) that incrementally learns from partially labeled data. Observing that existing continual learning methods la…

Cited by 97PDFScholar
2021

Rethinking and Reweighting the Univariate Losses for Multi-Label Ranking: Consistency and Generalization

NeurIPS 2021poster

The (partial) ranking loss is a commonly used evaluation measure for multi-label classification, which is usually optimized with convex surrogates for computational efficiency. Prior theoretical efforts on multi-label ranking mainly focus on (Fisher) consistency analyses. However, there is a gap bet…

Cited by 14SourcePDFScholar
2021

Stability and Generalization of Bilevel Programming in Hyperparameter Optimization

NeurIPS 2021poster

The (gradient-based) bilevel programming framework is widely used in hyperparameter optimization and has achieved excellent performance empirically. Previous theoretical work mainly focuses on its optimization properties, while leaving the analysis on generalization largely open. This paper attempts…

2021

Variational (Gradient) Estimate of the Score Function in Energy-based Latent Variable Models

ICML 2021spotlight

This paper presents new estimates of the score function and its gradient with respect to the model parameters in a general energy-based latent variable model (EBLVM). The score function and its gradient can be expressed as combinations of expectation and covariance terms over the (generally intracta…

2020

Bi-level Score Matching for Learning Energy-based Latent Variable Models

NeurIPS 2020poster

Score matching (SM) provides a compelling approach to learn energy-based models (EBMs) by avoiding the calculation of partition function. However, it remains largely open to learn energy-based latent variable models (EBLVMs), except some special cases. This paper presents a bi-level score matching (…

2020

Efficient Learning of Generative Models via Finite-Difference Score Matching

NeurIPS 2020poster

Several machine learning applications involve the optimization of higher-order derivatives (e.g., gradients of gradients) during training, which can be expensive with respect to memory and computation even with automatic differentiation. As a typical example in generative modeling, score matching~(S…

2020

Understanding and Stabilizing GANs’ Training Dynamics Using Control Theory

ICML 2020poster

Generative adversarial networks (GANs) are effective in generating realistic images but the training is often unstable. There are existing efforts that model the training dynamics of GANs in the parameter space but the analysis cannot directly motivate practically effective stabilizing methods. To t…

Cited by 36SourcePDFScholar
2019

Multi-objects Generation with Amortized Structural Regularization

NeurIPS 2019poster

Deep generative models (DGMs) have shown promise in image generation. However, most of the existing methods learn a model by simply optimizing a divergence between the marginal distributions of the model and the data, and often fail to capture rich structures, such as attributes of objects and their…

2017

Population Matching Discrepancy and Applications in Deep Learning

NeurIPS 2017poster

A differentiable estimation of the distance between two distributions based on samples is important for many deep learning tasks. One such estimation is maximum mean discrepancy (MMD). However, MMD suffers from its sensitive kernel bandwidth hyper-parameter, weak gradients, and large mini-batch siz…

Cited by 7SourcePDFScholar