← Search

Fanghui Liu

25 accepted papers

2026

DAG-Math: Graph-Guided Mathematical Reasoning in LLMs

ICLR 2026poster

Large Language Models (LLMs) demonstrate strong performance on mathematical problems when prompted with Chain-of-Thought (CoT), yet it remains unclear whether this success stems from search, rote procedures, or rule-consistent reasoning. To address this, we propose modeling CoT as a certain rule-bas…

Cited by 0SourcecodeScholar
2025

How Gradient descent balances features: A dynamical analysis for two-layer neural networks

ICLR 2025poster

This paper investigates the fundamental regression task of learning $k$ neurons (\emph{a.k.a.} teachers) from Gaussian input, using two-layer ReLU neural networks with width $m$ (\emph{a.k.a.} students) and $m, k= \mathcal{O}(1)$, trained via gradient descent under proper initialization and a small…

Cited by 0SourcePDFScholar
2025

LoRA-One: One-Step Full Gradient Could Suffice for Fine-Tuning Large Language Models, Provably and Efficiently

ICML 2025oral

This paper explores how theory can guide and enhance practical algorithms, using Low-Rank Adaptation (LoRA) (Hu et al., 2022) in large language models as a case study. We rigorously prove that, under gradient descent, LoRA adapters align with specific singular subspaces of the one-step full fine-tun…

2025

The $\varphi$ Curve: The Shape of Generalization through the Lens of Norm-based Capacity Control

NeurIPS 2025poster

Understanding how the test risk scales with model complexity is a central question in machine learning. Classical theory is challenged by the learning curves observed for large over-parametrized deep networks. Capacity measures based on parameter count typically fail to account for these empirical o…

Cited by 0SourceScholar
2024

Efficient local linearity regularization to overcome catastrophic overfitting

ICLR 2024poster

Catastrophic overfitting (CO) in single-step adversarial training (AT) results in abrupt drops in the adversarial test accuracy (even down to $0$%). For models trained with multi-step AT, it has been observed that the loss function behaves locally linearly with respect to the input, this is however…

2024

Generalization of Scaled Deep ResNets in the Mean-Field Regime

ICLR 2024spotlight

Despite the widespread empirical success of ResNet, the generalization properties of deep ResNet are rarely explored beyond the lazy training regime. In this work, we investigate scaled ResNet in the limit of infinitely deep and wide neural networks, of which the gradient flow is described by a part…

Cited by 5SourcePDFScholar
2024

High-Dimensional Kernel Methods under Covariate Shift: Data-Dependent Implicit Regularization

ICML 2024poster

This paper studies kernel ridge regression in high dimensions under covariate shifts and analyzes the role of importance re-weighting. We first derive the asymptotic expansion of high dimensional kernels under covariate shifts. By a bias-variance decomposition, we theoretically demonstrate that the…

Cited by 3SourcePDFScholar
2024

Learning Scalable Model Soup on a Single GPU: An Efficient Subspace Training Strategy

ECCV 2024poster

"Pre-training followed by fine-tuning is widely adopted among practitioners. The performance can be improved by “model soups” [?] via exploring various hyperparameter configurations. The Learned-Soup, a variant of model soups, significantly improves the performance but suffers from substantial memor…

2024

Revisiting Character-level Adversarial Attacks for Language Models

ICML 2024poster

Adversarial attacks in Natural Language Processing apply perturbations in the character or token levels. Token-level attacks, gaining prominence for their use of gradient-based methods, are susceptible to altering sentence semantics, leading to invalid adversarial examples. While character-level att…

2024

Robust NAS under adversarial training: benchmark, theory, and beyond

ICLR 2024poster

Recent developments in neural architecture search (NAS) emphasize the significance of considering robust architectures against malicious data. However, there is a notable absence of benchmark evaluations and theoretical guarantees for searching these robust architectures, especially when adversarial…

Cited by 6SourcePDFScholar
2023

Benign Overfitting in Deep Neural Networks under Lazy Training

ICML 2023poster

This paper focuses on over-parameterized deep neural networks (DNNs) with ReLU activation functions and proves that when the data distribution is well-separated, DNNs can achieve Bayes-optimal test error for classification while obtaining (nearly) zero-training error under the lazy training regime.…

Cited by 15SourcePDFScholar
2023

Initialization Matters: Privacy-Utility Analysis of Overparameterized Neural Networks

NeurIPS 2023poster

We analytically investigate how over-parameterization of models in randomized machine learning algorithms impacts the information leakage about their training data. Specifically, we prove a privacy bound for the KL divergence between model distributions on worst-case neighboring datasets, and explor…

Cited by 11SourcePDFScholar
2023

On the Convergence of Encoder-only Shallow Transformers

NeurIPS 2023poster

In this paper, we aim to build the global convergence theory of encoder-only shallow Transformers under a realistic setting from the perspective of architectures, initialization, and scaling under a finite width regime. The difficulty lies in how to tackle the softmax in self-attention mechanism, th…

Cited by 9SourcePDFScholar
2023

What can online reinforcement learning with function approximation benefit from general coverage conditions?

ICML 2023poster

In online reinforcement learning (RL), instead of employing standard structural assumptions on Markov decision processes (MDPs), using a certain coverage condition (original from offline RL) is enough to ensure sample-efficient guarantees (Xie et al. 2023). In this work, we focus on this new directi…

Cited by 4SourcePDFScholar
2022

Extrapolation and Spectral Bias of Neural Nets with Hadamard Product: a Polynomial Net Study

NeurIPS 2022accept

Neural tangent kernel (NTK) is a powerful tool to analyze training dynamics of neural networks and their generalization bounds. The study on NTK has been devoted to typical neural network architectures, but it is incomplete for neural networks with Hadamard products (NNs-Hp), e.g., StyleGAN and poly…

Cited by 13SourcePDFScholar
2022

Generalization Properties of NAS under Activation and Skip Connection Search

NeurIPS 2022accept

Neural Architecture Search (NAS) has fostered the automatic discovery of state-of-the-art neural architectures. Despite the progress achieved with NAS, so far there is little attention to theoretical guarantees on NAS. In this work, we study the generalization properties of NAS under a unifying fram…

Cited by 25SourcePDFScholar
2022

Robustness in deep learning: The good (width), the bad (depth), and the ugly (initialization)

NeurIPS 2022accept

We study the average robustness notion in deep neural networks in (selected) wide and narrow, deep and shallow, as well as lazy and non-lazy training settings. We prove that in the under-parameterized setting, width has a negative effect while it improves robustness in the over-parameterized setting…

Cited by 26SourcePDFScholar
2022

Sound and Complete Verification of Polynomial Networks

NeurIPS 2022accept

Polynomial Networks (PNs) have demonstrated promising performance on face and image recognition recently. However, robustness of PNs is unclear and thus obtaining certificates becomes imperative for enabling their adoption in real-world applications. Existing verification algorithms on ReLU neural n…

2022

Understanding Deep Neural Function Approximation in Reinforcement Learning via $\epsilon$-Greedy Exploration

NeurIPS 2022accept

This paper provides a theoretical study of deep neural function approximation in reinforcement learning (RL) with the $\epsilon$-greedy exploration under the online setting. This problem setting is motivated by the successful deep Q-networks (DQN) framework that falls in this regime. In this work, w…

Cited by 22SourcePDFScholar
2021

Fast Learning in Reproducing Kernel Krein Spaces via Signed Measures

AISTATS 2021poster

In this paper, we attempt to solve a long-lasting open question for non-positive definite (non-PD) kernels in machine learning community: can a given non-PD kernel be decomposed into the difference of two PD kernels (termed as positive decomposition)? We cast this question as a distribution view by…

Cited by 13SourcePDFScholar
2021

Kernel regression in high dimensions: Refined analysis beyond double descent

AISTATS 2021poster

In this paper, we provide a precise characterization of generalization properties of high dimensional kernel ridge regression across the under- and over-parameterized regimes, depending on whether the number of training data n exceeds the feature dimension d. By establishing a bias-variance decompos…

Cited by 61SourcePDFScholar
2016

Robust visual tracking via inverse nonnegative matrix factorization

ICASSP 2016accepted

The establishment of robust target appearance model over time is an overriding concern in visual tracking. In this paper, we propose an inverse nonnegative matrix factorization (NMF) method for robust appearance modeling. Rather than using a linear combination of nonnegative basis vectors for each t…

Cited by 0SourceScholar