← Search

Puneesh Deora

5 accepted papers

2026

Facts in Stats: Impacts of Pretraining Diversity on Language Model Generalization

ICML 2026poster

Language models are pretrained on sequences that blend statistical regularities (structures making text fluent) with factual associations between specific tokens (corresponding to knowledge of facts). While recent work suggests that the variability of their interaction, such as paraphrases of factua…

Cited by 0SourceScholar
2026

How Muon’s Spectral Design Benefits Generalization: A Study on Imbalanced Data

ICLR 2026poster

The growing adoption of spectrum-aware matrix-valued optimizers such as Muon and Shampoo in deep learning motivates a systematic study of their generalization properties and, in particular, when they might outperform competitive algorithms. We approach this question by introducing appropriate simp…

Cited by 0SourceScholar
2024

Fast Test Error Rates for Gradient-Based Algorithms on Separable Data

ICASSP 2024accepted

In recent research aimed at understanding the strong generalization performance of simple gradient-based methods on overparameterized models, it has been demonstrated that when training a linear predictor on separable data with an exponentially-tailed loss function, the predictor converges towards t…

Cited by 0SourceScholar
2023

On Weighted Cross-Entropy for Label-Imbalanced Separable Data: An Algorithmic-Stability Study

ICASSP 2023accepted

Implicit bias theory characterizes notions of simplicity in the weights learned by gradient descent when training without explicit regularization beyond zero training error, and has served as a cornerstone result for theoretically justifying good generalization of interpolating models. However, its…

Cited by 0SourceScholar
2021

LoOp: Looking for Optimal Hard Negative Embeddings for Deep Metric Learning

ICCV 2021poster

Deep metric learning has been effectively used to learn distance metrics for different visual tasks like image retrieval, clustering, etc. In order to aid the training process, existing methods either use a hard mining strategy to extract the most informative samples or seek to generate hard synthet…

Cited by 23PDFcodeScholar