← Search

Bamdev Mishra

16 accepted papers

2026

Minibatch selection for Language Models via Partition Matroid Constrained Gradient Matching

ICML 2026poster

Training Large Language Models (LLMs) on heterogeneous datasets requires optimizing domain representations to balance convergence speed and domain coverage. While recent methods reduce computational overhead by selecting high-quality data subsets, they typically perform selection independently per d…

Cited by 0SourceScholar
2025

A Riemannian Approach to Ground Metric Learning for Optimal Transport

ICASSP 2025accepted

Optimal transport (OT) theory has attracted much attention in machine learning and signal processing applications. OT defines a notion of distance between probability distributions of source and target data points. A crucial factor that influences OT-based distances is the ground metric of the embed…

Cited by 0SourceScholar
2024

A Framework for Bilevel Optimization on Riemannian Manifolds

NeurIPS 2024poster

Bilevel optimization has gained prominence in various applications. In this study, we introduce a framework for solving bilevel optimization problems, where the variables in both the lower and upper levels are constrained on Riemannian manifolds. We present several hypergradient estimation strategie…

2024

SLTrain: a sparse plus low rank approach for parameter and memory efficient pretraining

NeurIPS 2024poster

Large language models (LLMs) have shown impressive capabilities across various tasks. However, training LLMs from scratch requires significant computational power and extensive memory capacity. Recent studies have explored low-rank structures on weights for efficient fine-tuning in terms of paramete…

2024

Submodular framework for structured-sparse optimal transport

ICML 2024poster

Unbalanced optimal transport (UOT) has recently gained much attention due to its flexible framework for handling un-normalized measures and its robustness properties. In this work, we explore learning (structured) sparse transport plans in the UOT setting, i.e., transport plans have an upper bound o…

2023

Riemannian Accelerated Gradient Methods via Extrapolation

AISTATS 2023poster

In this paper, we propose a convergence acceleration scheme for general Riemannian optimization problems by extrapolating iterates on manifolds. We show that when the iterates are generated from the Riemannian gradient descent method, the scheme achieves the optimal convergence rate asymptotically a…

Cited by 10SourcePDFScholar
2021

On Riemannian Optimization over Positive Definite Matrices with the Bures-Wasserstein Geometry

NeurIPS 2021poster

In this paper, we comparatively analyze the Bures-Wasserstein (BW) geometry with the popular Affine-Invariant (AI) geometry for Riemannian optimization on the symmetric positive definite (SPD) matrix manifold. Our study begins with an observation that the BW metric has a linear dependence on SPD mat…

2019

Riemannian adaptive stochastic gradient algorithms on matrix manifolds

ICML 2019oral

Adaptive stochastic gradient algorithms in the Euclidean space have attracted much attention lately. Such explorations on Riemannian manifolds, on the other hand, are relatively new, limited, and challenging. This is because of the intrinsic non-linear structure of the underlying manifold and the ab…

2018

A Dual Framework for Low-rank Tensor Completion

NeurIPS 2018poster

One of the popular approaches for low-rank tensor completion is to use the latent trace norm regularization. However, most existing works in this direction learn a sparse combination of tensors. In this work, we fill this gap by proposing a variant of the latent trace norm that helps in learning a n…

2018

Riemannian stochastic quasi-Newton algorithm with variance reduction and its convergence analysis

AISTATS 2018poster

Stochastic variance reduction algorithms have recently become popular for minimizing the average of a large, but finite number of loss functions. The present paper proposes a Riemannian stochastic quasi-Newton algorithm with variance reduction (R-SQN-VR). The key challenges of averaging, adding, and…

Cited by 0SourcePDFScholar