← Search

Shiqian Ma

14 accepted papers

2026

Adaptive gradient descent on Riemannian manifolds and its applications to Gaussian variational inference

ICLR 2026poster

We propose RAdaGD, a novel family of adaptive gradient descent methods on general Riemannian manifolds. RAdaGD adapts the step size parameter without line search, and includes instances that achieve a non-ergodic convergence guarantee, $f(x_k) - f(x_\star) \le \mathcal{O}(1/k)$, under local geodesic…

Cited by 0SourcecodeScholar
2026

Mirror Flow Matching with Heavy-Tailed Priors for Generative Modeling on Convex Domains

ICLR 2026poster

We study generative modeling on convex domains using flow matching and mirror maps, and identify two fundamental challenges. First, standard log-barrier mirror maps induce heavy-tailed dual distributions, leading to ill-posed dynamics. Second, coupling with Gaussian priors performs poorly when match…

Cited by 0SourceScholar
2025

Tuning-Free Bilevel Optimization: New Algorithms and Convergence Analysis

ICLR 2025poster

Bilevel optimization has recently attracted considerable attention due to its abundant applications in machine learning problems. However, existing methods rely on prior knowledge of problem parameters to determine stepsizes, resulting in significant effort in tuning stepsizes when these parameters…

2023

Decentralized Stochastic Bilevel Optimization with Improved per-Iteration Complexity

ICML 2023poster

Bilevel optimization recently has received tremendous attention due to its great success in solving important machine learning problems like meta learning, reinforcement learning, and hyperparameter optimization. Extending single-agent training on bilevel problems to the decentralized setting is a n…

Cited by 35SourcePDFScholar
2021

A Riemannian Block Coordinate Descent Method for Computing the Projection Robust Wasserstein Distance

ICML 2021spotlight

The Wasserstein distance has become increasingly important in machine learning and deep learning. Despite its popularity, the Wasserstein distance is hard to approximate because of the curse of dimensionality. A recently proposed approach to alleviate the curse of dimensionality is to project the sa…

Cited by 52SourcePDFScholar
2018

Stochastic Primal-Dual Method for Empirical Risk Minimization with O(1) Per-Iteration Complexity

NeurIPS 2018poster

Regularized empirical risk minimization problem with linear predictor appears frequently in machine learning. In this paper, we propose a new stochastic primal-dual method to solve this class of problems. Different from existing methods, our proposed methods only require O(1) operations in each iter…

Cited by 42SourcePDFScholar
2017

GSOS: Gauss-Seidel Operator Splitting Algorithm for Multi-Term Nonsmooth Convex Composite Optimization

ICML 2017poster

In this paper, we propose a fast Gauss-Seidel Operator Splitting (GSOS) algorithm for addressing multi-term nonsmooth convex composite optimization, which has wide applications in machine learning, signal processing and statistics. The proposed GSOS algorithm inherits the advantage of the Gauss-Seid…

Cited by 7SourcePDFScholar