← Search

Vignesh Ganapathiraman

5 accepted papers

2024

Hop, skip, jump to Convergence: Dynamics of Learning Rate Transitions for Improved Training of Large Language Models

EMNLP 2024finding

Various types of learning rate (LR) schedulers are being used for training or fine tuning of Large Language Models today. In practice, several mid-flight changes are required in the LR schedule either manually, or with careful choices around warmup steps, peak LR, type of decay and restarts. To stud…

Cited by 0SourcePDFScholar
2020

Convex Representation Learning for Generalized Invariance in Semi-Inner-Product Space

ICML 2020poster

Invariance (defined in a general sense) has been one of the most effective priors for representation learning. Direct factorization of parametric models is feasible only for a small range of invariances, while regularization approaches, despite improved generality, lead to nonconvex optimization. In…

Cited by 3SourcePDFScholar
2018

Inductive Two-Layer Modeling with Parametric Bregman Transfer

ICML 2018oral

Latent prediction models, exemplified by multi-layer networks, employ hidden variables that automate abstract feature discovery. They typically pose nonconvex optimization problems and effective semi-definite programming (SDP) relaxations have been developed to enable global solutions (Aslan et al.,…

Cited by 7SourcePDFScholar