← Search

Avrajit Ghosh

7 accepted papers

2026

Sampled hard labels from sparse targets mislead rotation invariant algorithms

ICML 2026poster

One of the most common machine learning setups is logistic regression. In many classification models, including neural networks, the final prediction is obtained by applying a logistic link function to a linear score. In binary logistic regression, the feedback can be either soft labels, correspondi…

Cited by 0SourceScholar
2025

Learning Dynamics of Deep Matrix Factorization Beyond the Edge of Stability

ICLR 2025poster

Deep neural networks trained using gradient descent with a fixed learning rate $\eta$ often operate in the regime of ``edge of stability'' (EOS), where the largest eigenvalue of the Hessian equilibrates about the stability threshold $2/\eta$. In this work, we present a fine-grained analysis of the l…

Cited by 0SourcePDFScholar
2025

Variational Learning Finds Flatter Solutions at the Edge of Stability

NeurIPS 2025spotlight

Variational Learning (VL) has recently gained popularity for training deep neural networks. Part of its empirical success can be explained by theories such as PAC-Bayes bounds, minimum description length and marginal likelihood, but little has been done to unravel the implicit regularization in play…

Cited by 0SourceScholar
2024

Optimal Eye Surgeon: Finding image priors through sparse generators at initialization

ICML 2024poster

We introduce Optimal Eye Surgeon (OES), a framework for pruning and training deep image generator networks. Typically, untrained deep convolutional networks, which include image sampling operations, serve as effective image priors. However, they tend to overfit to noise in image restoration tasks du…

2024

Towards Understanding Task-agnostic Debiasing Through the Lenses of Intrinsic Bias and Forgetfulness

ACL 2024findings

While task-agnostic debiasing provides notable generalizability and reduced reliance on downstream data, its impact on language modeling ability and the risk of relearning social biases from downstream task-specific data remain as the two most significant challenges when debiasing Pretrained Languag…

2023

Implicit regularization in Heavy-ball momentum accelerated stochastic gradient descent

ICLR 2023top-25%

It is well known that the finite step-size ($h$) in Gradient descent (GD) implicitly regularizes solutions to flatter minimas. A natural question to ask is \textit{Does the momentum parameter $\beta$ (say) play a role in implicit regularization in Heavy-ball (H.B) momentum accelerated gradient desce…

Cited by 23SourcePDFScholar
2022

Bilevel Learning of ℓ1 Regularizers with Closed-Form Gradients (BLORC)

ICASSP 2022accepted

We present a method for supervised learning of sparsity-promoting regularizers, which are a key ingredient in many modern signal reconstruction approaches. The parameters of the regularizer are learned to minimize the mean squared error of reconstruction on a training set of ground truth signal and…

Cited by 0SourceScholar