← Search

Xiaoming Yuan

8 accepted papers

2026

Stability and Generalization of Nonconvex Optimization with Heavy-Tailed Noise

ICML 2026poster

The empirical evidence indicates that stochastic optimization with heavy-tailed gradient noise is more appropriate to characterize the training of machine learning models than that with standard bounded gradient variance noise. Most existing works on this phenomenon focus on the convergence of optim…

Cited by 0SourceScholar
2025

FISTAPruner: Layer-wise Post-training Pruning for Large Language Models

EMNLP 2025

Pruning is a critical strategy for compressing trained large language models (LLMs), aiming at substantial memory conservation and computational acceleration without compromising performance. However, existing pruning methods typically necessitate inefficient retraining for billion-scale LLMs or rel

Cited by 0SourcePDFScholar
2021

A Value-Function-based Interior-point Method for Non-convex Bi-level Optimization

ICML 2021spotlight

Bi-level optimization model is able to capture a wide range of complex learning tasks with practical interest. Due to the witnessed efficiency in solving bi-level programs, gradient-based methods have gained popularity in the machine learning community. In this work, we propose a new gradient-based…

Cited by 89SourcePDFScholar
2020

A Generic First-Order Algorithmic Framework for Bi-Level Programming Beyond Lower-Level Singleton

ICML 2020poster

In recent years, a variety of gradient-based bi-level optimization methods have been developed for learning tasks. However, theoretical guarantees of these existing approaches often heavily rely on the simplification that for each fixed upper-level variable, the lower-level solution must be a single…

Cited by 146SourcePDFScholar
2017

Adaptive Consensus ADMM for Distributed Optimization

ICML 2017poster

The alternating direction method of multipliers (ADMM) is commonly used for distributed model fitting problems, but its performance and reliability depend strongly on user-defined penalty parameters. We study distributed ADMM methods that boost performance by using different fine-tuned algorithm par…

Cited by 89SourcePDFScholar
2017

Adaptive Relaxed ADMM: Convergence Theory and Practical Implementation

CVPR 2017poster

Many modern computer vision and machine learning applications rely on solving difficult optimization problems that involve non-differentiable objective functions and constraints. The alternating direction method of multipliers (ADMM) is a widely used approach to solve such problems. Relaxed ADMM is…

Cited by 57PDFScholar
2015

Adaptive Primal-Dual Splitting Methods for Statistical Learning and Image Processing

NeurIPS 2015poster

The alternating direction method of multipliers (ADMM) is an important tool for solving complex optimization problems, but it involves minimization sub-steps that are often difficult to solve efficiently. The Primal-Dual Hybrid Gradient (PDHG) method is a powerful alternative that often has simple…

Cited by 109SourcePDFScholar