← Search

Xiaochuan Gong

5 accepted papers

2026

Bilevel Optimization with Lower-Level Uniform Convexity: Theory and Algorithm

ICLR 2026poster

Bilevel optimization is a hierarchical framework where an upper-level optimization problem is constrained by a lower-level problem, commonly used in machine learning applications such as hyperparameter optimization. Existing bilevel optimization methods typically assume strong convexity or Polyak-Ło…

Cited by 0SourceScholar
2025

Adaptive Algorithms with Sharp Convergence Rates for Stochastic Hierarchical Optimization

NeurIPS 2025poster

Hierarchical optimization refers to problems with interdependent decision variables and objectives, such as minimax and bilevel formulations. While various algorithms have been proposed, existing methods and analyses lack adaptivity in stochastic optimization settings: they cannot achieve optimal co…

Cited by 0SourcecodeScholar
2024

A Nearly Optimal Single Loop Algorithm for Stochastic Bilevel Optimization under Unbounded Smoothness

ICML 2024poster

This paper studies the problem of stochastic bilevel optimization where the upper-level function is nonconvex with potentially unbounded smoothness and the lower-level function is strongly convex. This problem is motivated by meta-learning applied to sequential data, such as text classification usin…

2024

An Accelerated Algorithm for Stochastic Bilevel Optimization under Unbounded Smoothness

NeurIPS 2024poster

This paper investigates a class of stochastic bilevel optimization problems where the upper-level function is nonconvex with potentially unbounded smoothness and the lower-level problem is strongly convex. These problems have significant applications in sequential data learning, such as text classif…

2024

Bilevel Optimization under Unbounded Smoothness: A New Algorithm and Convergence Analysis

ICLR 2024spotlight

Bilevel optimization is an important formulation for many machine learning problems, such as meta-learning and hyperparameter optimization. Current bilevel optimization algorithms assume that the gradient of the upper-level function is Lipschitz (i.e., the upper-level function has a bounded smoothne…