← Search

Bhavya Vasudeva

6 accepted papers

2026

How Muon’s Spectral Design Benefits Generalization: A Study on Imbalanced Data

ICLR 2026poster

The growing adoption of spectrum-aware matrix-valued optimizers such as Muon and Shampoo in deep learning motivates a systematic study of their generalization properties and, in particular, when they might outperform competitive algorithms. We approach this question by introducing appropriate simp…

Cited by 0SourceScholar
2026

Latent Concept Disentanglement in Transformer-based Language Models

ICLR 2026poster

When large language models (LLMs) use in-context learning (ICL) to solve a new task, they must infer latent concepts from demonstration examples. This raises the question of whether and how transformers represent latent structures as part of their computation. Our work experiments with several contr…

Cited by 0SourceScholar
2025

The Rich and the Simple: On the Implicit Bias of Adam and SGD

NeurIPS 2025poster

Adam is the de facto optimization algorithm for several deep learning applications, but an understanding of its implicit bias and how it differs from other algorithms, particularly standard first-order methods such as (stochastic) gradient descent (GD), remains limited. In practice, neural networks…

Cited by 0SourceScholar
2025

Transformers Learn Low Sensitivity Functions: Investigations and Implications

ICLR 2025poster

Transformers achieve state-of-the-art accuracy and robustness across many tasks, but an understanding of their inductive biases and how those biases differ from other neural network architectures remains elusive. In this work, we identify the sensitivity of the model to token-wise random perturbatio…

Cited by 0SourcePDFScholar
2024

Fast Test Error Rates for Gradient-Based Algorithms on Separable Data

ICASSP 2024accepted

In recent research aimed at understanding the strong generalization performance of simple gradient-based methods on overparameterized models, it has been demonstrated that when training a linear predictor on separable data with an exponentially-tailed loss function, the predictor converges towards t…

Cited by 0SourceScholar
2021

LoOp: Looking for Optimal Hard Negative Embeddings for Deep Metric Learning

ICCV 2021poster

Deep metric learning has been effectively used to learn distance metrics for different visual tasks like image retrieval, clustering, etc. In order to aid the training process, existing methods either use a hard mining strategy to extract the most informative samples or seek to generate hard synthet…

Cited by 23PDFcodeScholar