← Search

Qing Qu

46 accepted papers

2026

AlphaFlow: Understanding and Improving MeanFlow Models

ICLR 2026poster

MeanFlow has recently emerged as a powerful framework for few-step generative modeling trained from scratch, but its success is not yet fully understood. In this work, we show that the MeanFlow objective naturally decomposes into two parts: trajectory flow matching and trajectory consistency. Throug…

Cited by 0SourcecodeScholar
2026

Evaluating the Representation Space of Diffusion Models via Self-Supervised Principles

ICML 2026poster

Diffusion models are effective generative frameworks with strong representation learning capabilities, yet the intrinsic properties that govern their semantic structure and generalization remain poorly understood. Drawing inspiration from self-supervised representation learning (SSL), we introduce a…

Cited by 0SourceScholar
2026

Generalization of Diffusion Models Arises with a Balanced Representation Space

ICLR 2026poster

Diffusion models generate high-quality, diverse images with great generalizability, yet when overfit to the training objective, they may memorize training samples. We analyze memorization and generalization of diffusion models through the lens of representation learning. Using a two-layer ReLU denoi…

Cited by 0SourcecodeScholar
2026

SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL

CVPR 2026

Vision Language Models (VLMs) demonstrate strong qualitative visual understanding, but struggle with metrically precise spatial reasoning required for embodied applications. The agentic paradigm promises that VLMs can use a wide variety of tools that could augment these capabilities, such as depth e

Cited by 0SourcecodeScholar
2026

Understanding Deep Representation Learning via Layerwise Feature Compression and Discrimination

ICML 2026poster

Over the past decade, deep learning has proven to be a highly effective tool for learning meaningful features from raw data. However, it remains an open question how deep networks perform hierarchical feature learning across layers. In this work, we attempt to unveil this mystery by investigating th…

Cited by 0SourcecodeScholar
2026

Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model Outputs

ICLR 2026poster

Machine unlearning (MU) for large language models (LLMs), commonly referred to as LLM unlearning, seeks to remove specific undesirable data or knowledge from a trained model, while maintaining its performance on standard tasks. While unlearning plays a vital role in protecting data privacy, enforcin…

Cited by 0SourcecodeScholar
2025

A Closer Look at Model Collapse: From a Generalization-to-Memorization Perspective

NeurIPS 2025spotlight

The widespread use of diffusion models has led to an abundance of AI-generated data, raising concerns about model collapse---a phenomenon in which recursive iterations of training on synthetic data lead to performance degradation. Prior work primarily characterizes this collapse via variance shrinka…

Cited by 0SourceScholar
2025

Attention-Only Transformers via Unrolled Subspace Denoising

ICML 2025poster

Despite the popularity of transformers in practice, their architectures are empirically designed and neither mathematically justified nor interpretable. Moreover, as indicated by many empirical studies, some components of transformer architectures may be redundant. To derive a fully interpretable tr…

Cited by 0SourcePDFScholar
2025

FlowDAS: A Stochastic Interpolant-based Framework for Data Assimilation

NeurIPS 2025poster

Data assimilation (DA) integrates observations with a dynamical model to estimate states of PDE-governed systems. Model-driven methods (e.g., Kalman Filter, Particle Filter) presuppose full knowledge of the true dynamics, which is not always satisfied in practice, while purely data-driven solvers le…

Cited by 0SourcecodeScholar
2025

Learning Dynamics of Deep Matrix Factorization Beyond the Edge of Stability

ICLR 2025poster

Deep neural networks trained using gradient descent with a fixed learning rate $\eta$ often operate in the regime of ``edge of stability'' (EOS), where the largest eigenvalue of the Hessian equilibrates about the stability threshold $2/\eta$. In this work, we present a fine-grained analysis of the l…

Cited by 0SourcePDFScholar
2025

SITCOM: Step-wise Triple-Consistent Diffusion Sampling For Inverse Problems

ICML 2025poster

Diffusion models (DMs) are a class of generative models that allow sampling from a distribution learned over a training set. When applied to solving inverse problems, the reverse sampling steps are modified to approximately sample from a measurement-conditioned distribution. However, these modificat…

2025

Sequential Diffusion-Guided Deep Image Prior for Medical Image Reconstruction

ICASSP 2025accepted

Deep learning (DL) methods have been extensively applied to various image recovery problems, including magnetic resonance imaging (MRI) and computed tomography (CT) reconstruction. Beyond supervised models, other approaches have been recently explored including two key recent schemes: deep image pri…

Cited by 0SourceScholar
2025

Shallow Diffuse: Robust and Invisible Watermarking through Low-Dim Subspaces in Diffusion Models

NeurIPS 2025spotlight

The widespread use of AI-generated content from diffusion models has raised significant concerns regarding misinformation and copyright infringement. Watermarking is a crucial technique for identifying these AI-generated images and preventing their misuse. In this paper, we introduce *Shallow Diffus…

Cited by 0SourceScholar
2025

UGoDIT: Unsupervised Group Deep Image Prior Via Transferable Weights

NeurIPS 2025poster

Recent advances in data-centric deep generative models have led to significant progress in solving inverse imaging problems. However, these models (e.g., diffusion models (DMs)) typically require large amounts of fully sampled (clean) training data, which is often impractical in medical and scientif…

Cited by 0SourcecodeScholar
2025

Understanding Representation Dynamics of Diffusion Models via Low-Dimensional Modeling

NeurIPS 2025poster

Diffusion models, though originally designed for generative tasks, have demonstrated impressive self-supervised representation learning capabilities. A particularly intriguing phenomenon in these models is the emergence of unimodal representation dynamics, where the quality of learned features peaks…

Cited by 0SourceScholar
2024

A Global Geometric Analysis of Maximal Coding Rate Reduction

ICML 2024poster

The maximal coding rate reduction (MCR$^2$) objective for learning structured and compact deep representations is drawing increasing attention, especially after its recent usage in the derivation of fully explainable and highly effective deep network architectures. However, it lacks a complete theor…

Cited by 6SourcePDFScholar
2024

BLAST: Block-Level Adaptive Structured Matrices for Efficient Deep Neural Network Inference

NeurIPS 2024poster

Large-scale foundation models have demonstrated exceptional performance in language and vision tasks. However, the numerous dense matrix-vector operations involved in these large networks pose significant computational challenges during inference. To address these challenges, we introduce the Block-…

2024

Compressible Dynamics in Deep Overparameterized Low-Rank Learning & Adaptation

ICML 2024oral

While overparameterization in machine learning models offers great benefits in terms of optimization and generalization, it also leads to increased computational requirements as model sizes grow. In this work, we show that by leveraging the inherent low-dimensional structures of data and compressibl…

2024

Diffusion-Based Adversarial Purification for Robust Deep Mri Reconstruction

ICASSP 2024accepted

Deep learning (DL) methods have been extensively employed in magnetic resonance imaging (MRI) reconstruction, demonstrating remarkable performance improvements compared to traditional non-DL methods. However, recent studies have uncovered the susceptibility of these models to carefully engineered ad…

Cited by 0SourceScholar
2024

Efficient Low-Dimensional Compression of Overparameterized Models

AISTATS 2024poster

In this work, we present a novel approach for compressing overparameterized models, developed through studying their learning dynamics. We observe that for many deep models, updates to the weight matrices occur within a low-dimensional invariant subspace. For deep linear models, we demonstrate that…

2024

Exploring Low-Dimensional Subspace in Diffusion Models for Controllable Image Editing

NeurIPS 2024poster

Recently, diffusion models have emerged as a powerful class of generative models. Despite their success, there is still limited understanding of their semantic spaces. This makes it challenging to achieve precise and disentangled image generation without additional training, especially in an unsupe…

2024

Generalized Neural Collapse for a Large Number of Classes

ICML 2024poster

Neural collapse provides an elegant mathematical characterization of learned last layer representations (a.k.a. features) and classifier weights in deep classification models. Such results not only provide insights but also motivate new techniques for improving practical deep models. However, most o…

Cited by 22SourcePDFScholar
2024

Image Reconstruction Via Autoencoding Sequential Deep Image Prior

NeurIPS 2024poster

Recently, Deep Image Prior (DIP) has emerged as an effective unsupervised one-shot learner, delivering competitive results across various image recovery problems. This method only requires the noisy measurements and a forward operator, relying solely on deep networks initialized with random noise to…

Cited by 1SourcePDFScholar
2024

Improving Training Efficiency of Diffusion Models via Multi-Stage Framework and Tailored Multi-Decoder Architecture

CVPR 2024poster

Diffusion models emerging as powerful deep generative tools excel in various applications. They operate through a two-steps process: introducing noise into training samples and then employing a model to convert random noise into new samples (e.g. images). However their remarkable generative performa…

Cited by 12SourcePDFScholar
2024

Neural Collapse in Multi-label Learning with Pick-all-label Loss

ICML 2024poster

We study deep neural networks for the multi-label classification (MLab) task through the lens of neural collapse (NC). Previous works have been restricted to the multi-class classification setting and discovered a prevalent NC phenomenon comprising of the following properties for the last-layer feat…

2024

Optimal Eye Surgeon: Finding image priors through sparse generators at initialization

ICML 2024poster

We introduce Optimal Eye Surgeon (OES), a framework for pruning and training deep image generator networks. Typically, untrained deep convolutional networks, which include image sampling operations, serve as effective image priors. However, they tend to overfit to noise in image restoration tasks du…

2024

Solving Inverse Problems with Latent Diffusion Models via Hard Data Consistency

ICLR 2024spotlight

Latent diffusion models have been demonstrated to generate high-quality images, while offering efficiency in model training compared to diffusion models operating in the pixel space. However, incorporating latent diffusion models to solve inverse problems remains a challenging problem due to the non…

2024

The Emergence of Reproducibility and Consistency in Diffusion Models

ICML 2024poster

In this work, we investigate an intriguing and prevalent phenomenon of diffusion models which we term as "consistent model reproducibility'': given the same starting noise input and a deterministic sampler, different diffusion models often yield remarkably similar outputs. We confirm this phenomenon…

Cited by 57SourcePDFScholar
2024

Understanding Generalizability of Diffusion Models Requires Rethinking the Hidden Gaussian Structure

NeurIPS 2024poster

In this work, we study the generalizability of diffusion models by looking into the hidden properties of the learned score functions, which are essentially a series of deep denoisers trained on various noise levels. We observe that as diffusion models transition from memorization to generalization,…

2022

Are All Losses Created Equal: A Neural Collapse Perspective

NeurIPS 2022accept

While cross entropy (CE) is the most commonly used loss function to train deep neural networks for classification tasks, many alternative losses have been developed to obtain better empirical performance. Among them, which one is the best to use is still a mystery, because there seem to be multiple…

Cited by 67SourcePDFScholar
2022

Hidden State Variability of Pretrained Language Models Can Guide Computation Reduction for Transfer Learning

EMNLP 2022finding

While transferring a pretrained language model, common approaches conventionally attach their task-specific classifiers to the top layer and adapt all the pretrained layers. We investigate whether one could make a task-specific selection on which subset of the layers to adapt and where to place the…

2022

Neural Collapse with Normalized Features: A Geometric Analysis over the Riemannian Manifold

NeurIPS 2022accept

When training overparameterized deep networks for classification tasks, it has been widely observed that the learned features exhibit a so-called "neural collapse'" phenomenon. More specifically, for the output features of the penultimate layer, for each class the within-class features converge to t…

2022

On the Optimization Landscape of Neural Collapse under MSE Loss: Global Optimality with Unconstrained Features

ICML 2022spotlight

When training deep neural networks for classification tasks, an intriguing empirical phenomenon has been widely observed in the last-layer classifiers and features, where (i) the class means and the last-layer classifiers all collapse to the vertices of a Simplex Equiangular Tight Frame (ETF) up to…

Cited by 125SourcePDFScholar
2022

Robust Training under Label Noise by Over-parameterization

ICML 2022spotlight

Recently, over-parameterized deep networks, with increasingly more network parameters than training samples, have dominated the performances of modern machine learning. However, when the training data is corrupted, it has been well-known that over-parameterized networks tend to overfit and do not ge…

2021

A Geometric Analysis of Neural Collapse with Unconstrained Features

NeurIPS 2021spotlight

We provide the first global optimization landscape analysis of Neural Collapse -- an intriguing empirical phenomenon that arises in the last-layer classifiers and features of neural networks during the terminal phase of training. As recently reported by Papyan et al., this phenomenon implies that (i…

2021

Convolutional Normalization: Improving Deep Convolutional Network Robustness and Training

NeurIPS 2021poster

Normalization techniques have become a basic component in modern convolutional neural networks (ConvNets). In particular, many recent works demonstrate that promoting the orthogonality of the weights helps train deep models and improve robustness. For ConvNets, most existing methods are based on pen…

2021

Rank Overspecified Robust Matrix Recovery: Subgradient Method and Exact Recovery

NeurIPS 2021poster

We study the robust recovery of a low-rank matrix from sparsely and grossly corrupted Gaussian measurements, with no prior knowledge on the intrinsic rank. We consider the robust matrix factorization approach. We employ a robust $\ell_1$ loss function and deal with the challenge of the unknown rank…

Cited by 31SourcePDFScholar
2020

Geometric Analysis of Nonconvex Optimization Landscapes for Overcomplete Learning

ICLR 2020talk

Learning overcomplete representations finds many applications in machine learning and data analytics. In the past decade, despite the empirical success of heuristic methods, theoretical understandings and explanations of these algorithms are still far from satisfactory. In this work, we provide new…

Cited by 33SourceScholar
2020

Robust Recovery via Implicit Bias of Discrepant Learning Rates for Double Over-parameterization

NeurIPS 2020spotlight

Recent advances have shown that implicit bias of gradient descent on over-parameterized models enables the recovery of low-rank matrices from linear measurements, even with no prior knowledge on the intrinsic rank. In contrast, for {\em robust} low-rank matrix recovery from {\em grossly corrupted} m…

2020

Short and Sparse Deconvolution --- A Geometric Approach

ICLR 2020poster

Short-and-sparse deconvolution (SaSD) is the problem of extracting localized, recurring motifs in signals with spatial or temporal structure. Variants of this problem arise in applications such as image deblurring, microscopy, neural spike sorting, and more. The problem is challenging in both theory…

Cited by 37SourcecodeScholar
2019

A Nonconvex Approach for Exact and Efficient Multichannel Sparse Blind Deconvolution

NeurIPS 2019spotlight

We study the multi-channel sparse blind deconvolution (MCS-BD) problem, whose task is to simultaneously recover a kernel $\mathbf a$ and multiple sparse inputs $\{\mathbf x_i\}_{i=1}^p$ from their circulant convolution $\mathbf y_i = \mb a \circledast \mb x_i $ ($i=1,\cdots,p$). We formulate the tas…