← Search

Ruichen Jiang

11 accepted papers

2026

Improving Online-to-Nonconvex Conversion for Smooth Optimization via Double Optimism

ICLR 2026poster

A recent breakthrough in nonconvex optimization is the online-to-nonconvex conversion framework of Cutkosky et al. (2023), which reformulates the task of finding an $\varepsilon$-first-order stationary point as an online learning problem. When both the gradient and the Hessian are Lipschitz continu…

Cited by 0SourceScholar
2025

On the Complexity of Finding Stationary Points in Nonconvex Simple Bilevel Optimization

NeurIPS 2025poster

In this paper, we study the problem of solving a simple bilevel optimization problem, where the upper-level objective is minimized over the solution set of the lower-level problem. We focus on the general setting in which both the upper- and lower-level objectives are smooth but potentially nonconve…

Cited by 0SourceScholar
2024

Adaptive and Optimal Second-order Optimistic Methods for Minimax Optimization

NeurIPS 2024poster

We propose adaptive, line-search-free second-order methods with optimal rate of convergence for solving convex-concave min-max problems. By means of an adaptive step size, our algorithms feature a simple update rule that requires solving only one linear system per iteration, eliminating the need for…

Cited by 4SourcePDFScholar
2024

An Accelerated Gradient Method for Convex Smooth Simple Bilevel Optimization

NeurIPS 2024poster

In this paper, we focus on simple bilevel optimization problems, where we minimize a convex smooth objective function over the optimal solution set of another convex smooth constrained optimization problem. We present a novel bilevel optimization method that locally approximates the solution set of…

Cited by 1SourcePDFScholar
2024

Krylov Cubic Regularized Newton: A Subspace Second-Order Method with Dimension-Free Convergence Rate

AISTATS 2024poster

Second-order optimization methods, such as cubic regularized Newton methods, are known for their rapid convergence rates; nevertheless, they become impractical in high-dimensional problems due to their substantial memory requirements and computational costs. One promising approach is to execute seco…

Cited by 1SourcePDFScholar
2024

Non-asymptotic Global Convergence Analysis of BFGS with the Armijo-Wolfe Line Search

NeurIPS 2024spotlight

In this paper, we present the first explicit and non-asymptotic global convergence rates of the BFGS method when implemented with an inexact line search scheme satisfying the Armijo-Wolfe conditions. We show that BFGS achieves a global linear convergence rate of $(1 - \frac{1}{\kappa})^t$ for $\mu$-…

Cited by 1SourcePDFScholar
2023

A Conditional Gradient-based Method for Simple Bilevel Optimization with Convex Lower-level Problem

AISTATS 2023poster

In this paper, we study a class of bilevel optimization problems, also known as simple bilevel optimization, where we minimize a smooth objective function over the optimal solution set of another convex constrained optimization problem. Several iterative methods have been developed for tackling this…

2023

Accelerated Quasi-Newton Proximal Extragradient: Faster Rate for Smooth Convex Optimization

NeurIPS 2023spotlight

In this paper, we propose an accelerated quasi-Newton proximal extragradient method for solving unconstrained smooth convex optimization problems. With access only to the gradients of the objective, we prove that our method can achieve a convergence rate of $\mathcal{O}\bigl(\min\\{\frac{1}{k^2}, \f…

Cited by 15SourcePDFScholar
2023

Projection-Free Methods for Stochastic Simple Bilevel Optimization with Convex Lower-level Problem

NeurIPS 2023poster

In this paper, we study a class of stochastic bilevel optimization problems, also known as stochastic simple bilevel optimization, where we minimize a smooth stochastic objective function over the optimal solution set of another stochastic convex optimization problem. We introduce novel stochastic b…

Cited by 10SourcePDFScholar
2022

Future gradient descent for adapting the temporal shifting data distribution in online recommendation systems

UAI 2022poster

One of the key challenges of learning an online recommendation model is the temporal domain shift, which causes the mismatch between the training and testing data distribution and hence domain generalization error. To overcome, we propose to learn a meta future gradient generator that forecasts the…

Cited by 8SourcePDFScholar