← Search

Yulong Lu

9 accepted papers

2025

In-context Learning of Linear Dynamical Systems with Transformers: Approximation Bounds and Depth-separation

NeurIPS 2025poster

This paper investigates approximation-theoretic aspects of the in-context learning capability of the transformers in representing a family of noisy linear dynamical systems. Our first theoretical result establishes an upper bound on the approximation error of multi-layer transformers with respect to…

Cited by 0SourceScholar
2024

Score-based generative models break the curse of dimensionality in learning a family of sub-Gaussian distributions

ICLR 2024poster

While score-based generative models (SGMs) have achieved remarkable successes in enormous image generation tasks, their mathematical foundations are still limited. In this paper, we analyze the approximation and generalization of SGMs in learning a family of sub-Gaussian probability distributions. W…

Cited by 18SourcePDFScholar
2023

Transfer Learning Enhanced DeepONet for Long-Time Prediction of Evolution Equations

AAAI 2023technical

Deep operator network (DeepONet) has demonstrated great success in various learning tasks, including learning solution operators of partial differential equations. In particular, it provides an efficient approach to predicting the evolution equations in a finite time horizon. Nevertheless, the vani…

2023

Two-Scale Gradient Descent Ascent Dynamics Finds Mixed Nash Equilibria of Continuous Games: A Mean-Field Perspective

ICML 2023poster

Finding the mixed Nash equilibria (MNE) of a two-player zero sum continuous game is an important and challenging problem in machine learning. A canonical algorithm to finding the MNE is the noisy gradient descent ascent method which in the infinite particle limit gives rise to the Mean-Field Gradien…

Cited by 23SourcePDFScholar
2020

A Mean Field Analysis Of Deep ResNet And Beyond: Towards Provably Optimization Via Overparameterization From Depth

ICML 2020poster

Training deep neural networks with stochastic gradient descent (SGD) can often achieve zero training loss on real-world tasks although the optimization landscape is known to be highly non-convex. To understand the success of SGD for training deep neural networks, this work presents a mean-field anal…

Cited by 108SourcePDFScholar
2020

A Universal Approximation Theorem of Deep Neural Networks for Expressing Probability Distributions

NeurIPS 2020poster

This paper studies the universal approximation property of deep neural networks for representing probability distributions. Given a target distribution $\pi$ and a source distribution $p_z$ both defined on $\mathbb{R}^d$, we prove under some assumptions that there exists a deep neural network $g:\ma…

Cited by 226SourcePDFScholar