← Search

Hengyu Fu

6 accepted papers

2026

From Bits to Rounds: Parallel Decoding with Exploration for Diffusion Language Models

ICML 2026poster

Diffusion Language Models (DLMs) have recently emerged as a strong alternative to autoregressive language models (AR-LMs), due to their comparable accuracy and faster inference speed via parallel decoding. However, standard DLM decoding strategies, which rely on unmasking only high-confidence tokens…

Cited by 0SourceScholar
2026

Generalization Bounds for Discrete Diffusion: Statistical Advantage of Masking

ICML 2026poster

Discrete diffusion models have recently emerged as a compelling alternative for language generation, enabling efficient non-autoregressive sampling while achieving strong empirical performance. A key design choice in discrete diffusion---absent in most continuous diffusion formulations---is the forw…

Cited by 0SourceScholar
2026

Neural Networks Learn Generic Multi-Index Models Near Information-Theoretic Limit

ICLR 2026poster

In deep learning, a central issue is to understand how neural networks efficiently learn high-dimensional features. To this end, we explore the gradient descent learning of a general Gaussian Multi-index model $f(\boldsymbol{x})=g(\boldsymbol{U}\boldsymbol{x})$ with hidden subspace $\boldsymbol{U}\i…

Cited by 0SourceScholar
2025

Diffusion Transformer Captures Spatial-Temporal Dependencies: A Theory for Gaussian Process Data

ICLR 2025poster

Diffusion Transformer, the backbone of Sora for video generation, successfully scales the capacity of diffusion models, pioneering new avenues for high-fidelity sequential data generation. Unlike static data such as images, sequential data consists of consecutive data frames indexed by time, exhibit…

Cited by 3SourcePDFScholar
2025

Learning Hierarchical Polynomials of Multiple Nonlinear Features

ICLR 2025poster

In deep learning theory, a critical question is to understand how neural networks learn hierarchical features. In this work, we study the learning of hierarchical polynomials of multiple nonlinear features using three-layer neural networks. We examine a broad class of functions of the form $f^{\star…

Cited by 0SourcePDFScholar
2023

What can a Single Attention Layer Learn? A Study Through the Random Features Lens

NeurIPS 2023poster

Attention layers---which map a sequence of inputs to a sequence of outputs---are core building blocks of the Transformer architecture which has achieved significant breakthroughs in modern artificial intelligence. This paper presents a rigorous theoretical study on the learning and generalization of…

Cited by 34SourcePDFScholar