← Search

Xiliang Lu

5 accepted papers

2026

Approximation Error Upper and Lower Bounds for Hölder Class with Transformers

ICML 2026poster

We explore the expressive power of Transformers by establishing precise approximation error upper and lower bounds for Hölder class. Specifically, a new approximation upper bound is derived for the standard Transformer architecture equipped with Softmax operators, ReLU activation functions, and resi…

Cited by 0SourceScholar
2024

Neural Network Approximation for Pessimistic Offline Reinforcement Learning

AAAI 2024technical

Deep reinforcement learning (RL) has shown remarkable success in specific offline decision-making scenarios, yet its theoretical guarantees are still under development. Existing works on offline RL theory primarily emphasize a few trivial settings, such as linear MDP or general function approximatio…

Cited by 3SourcePDFScholar
2024

Non-asymptotic Approximation Error Bounds of Parameterized Quantum Circuits

NeurIPS 2024spotlight

Understanding the power of parameterized quantum circuits (PQCs) in accomplishing machine learning tasks is one of the most important questions in quantum machine learning. In this paper, we focus on the PQC expressivity for general multivariate function classes. Previously established Universal App…

Cited by 3SourcePDFScholar
2024

Take Care of Your Prompt Bias! Investigating and Mitigating Prompt Bias in Factual Knowledge Extraction

COLING 2024main

Recent research shows that pre-trained language models (PLMs) suffer from “prompt bias” in factual knowledge extraction, i.e., prompts tend to introduce biases toward specific labels. Prompt bias presents a significant challenge in assessing the factual knowledge within PLMs. Therefore, this paper a…

2023

Fast Excess Risk Rates via Offset Rademacher Complexity

ICML 2023poster

Based on the offset Rademacher complexity, this work outlines a systematical framework for deriving sharp excess risk bounds in statistical learning without Bernstein condition. In addition to recovering fast rates in a unified way for some parametric and nonparametric supervised learning models wit…

Cited by 5SourcePDFScholar