← Search

Yuling Jiao

10 accepted papers

2026

Approximation Bounds for Transformer Networks with Application to Regression

ICML 2026poster

We develop approximation and statistical theory for standard Transformer networks in sequence modeling. Given a sequence-to-sequence target on $[0,1]^{d_x \times n}$ whose entries are $\gamma$-H\"older for $\gamma \in (0,1]$ or belong to a first-order Sobolev class, we establish explicit $L^p$-appro…

Cited by 0SourceScholar
2026

Approximation Error Upper and Lower Bounds for Hölder Class with Transformers

ICML 2026poster

We explore the expressive power of Transformers by establishing precise approximation error upper and lower bounds for Hölder class. Specifically, a new approximation upper bound is derived for the standard Transformer architecture equipped with Softmax operators, ReLU activation functions, and resi…

Cited by 0SourceScholar
2025

Adv-SSL: Adversarial Self-Supervised Representation Learning with Theoretical Guarantees

NeurIPS 2025poster

Learning transferable data representations from abundant unlabeled data remains a central challenge in machine learning. Although numerous self-supervised learning methods have been proposed to address this challenge, a significant class of these approaches aligns the covariance or correlation matri…

Cited by 0SourceScholar
2024

Neural Network Approximation for Pessimistic Offline Reinforcement Learning

AAAI 2024technical

Deep reinforcement learning (RL) has shown remarkable success in specific offline decision-making scenarios, yet its theoretical guarantees are still under development. Existing works on offline RL theory primarily emphasize a few trivial settings, such as linear MDP or general function approximatio…

Cited by 3SourcePDFScholar
2024

Non-asymptotic Approximation Error Bounds of Parameterized Quantum Circuits

NeurIPS 2024spotlight

Understanding the power of parameterized quantum circuits (PQCs) in accomplishing machine learning tasks is one of the most important questions in quantum machine learning. In this paper, we focus on the PQC expressivity for general multivariate function classes. Previously established Universal App…

Cited by 3SourcePDFScholar
2023

Fast Excess Risk Rates via Offset Rademacher Complexity

ICML 2023poster

Based on the offset Rademacher complexity, this work outlines a systematical framework for deriving sharp excess risk bounds in statistical learning without Bernstein condition. In addition to recovering fast rates in a unified way for some parametric and nonparametric supervised learning models wit…

Cited by 5SourcePDFScholar
2022

Approximation with CNNs in Sobolev Space: with Applications to Classification

NeurIPS 2022accept

We derive a novel approximation error bound with explicit prefactor for Sobolev-regular functions using deep convolutional neural networks (CNNs). The bound is non-asymptotic in terms of the network depth and filter lengths, in a rather flexible way. For Sobolev-regular functions which can be embedd…

Cited by 20SourcePDFScholar
2021

Deep Generative Learning via Schrödinger Bridge

ICML 2021spotlight

We propose to learn a generative model via entropy interpolation with a Schr{ö}dinger Bridge. The generative learning task can be formulated as interpolating between a reference distribution and a target distribution based on the Kullback-Leibler divergence. At the population level, this entropy int…

2019

Deep Generative Learning via Variational Gradient Flow

ICML 2019oral

We propose a framework to learn deep generative models via \textbf{V}ariational \textbf{Gr}adient Fl\textbf{ow} (VGrow) on probability spaces. The evolving distribution that asymptotically converges to the target distribution is governed by a vector field, which is the negative gradient of the first…