← Search

Niladri Chatterji

6 accepted papers

2020

Langevin Monte Carlo without smoothness

AISTATS 2020poster

Langevin Monte Carlo (LMC) is an iterative algorithm used to generate samples from a distribution that is known only up to a normalizing constant. The nonasymptotic dependence of its mixing time on the dimension and target accuracy is understood mainly in the setting of smooth (gradient-Lipschitz) l…

Cited by 54SourcePDFScholar
2020

OSOM: A simultaneously optimal algorithm for multi-armed and linear contextual bandits

AISTATS 2020poster

We consider the stochastic linear (multi-armed) contextual bandit problem with the possibility of hidden simple multi-armed bandit structure in which the rewards are independent of the contextual information. Algorithms that are designed solely for one of the regimes are known to be sub-optimal for…

Cited by 47SourcePDFScholar
2020

The intriguing role of module criticality in the generalization of deep networks

ICLR 2020spotlight

We study the phenomenon that some modules of deep neural networks (DNNs) are more critical than others. Meaning that rewinding their parameter values back to initialization, while keeping other modules fixed at the trained parameters, results in a large drop in the network's performance. Our analysi…

Cited by 72SourceScholar
2018

On the Theory of Variance Reduction for Stochastic Gradient Monte Carlo

ICML 2018oral

We provide convergence guarantees in Wasserstein distance for a variety of variance-reduction methods: SAGA Langevin diffusion, SVRG Langevin diffusion and control-variate underdamped Langevin diffusion. We analyze these methods under a uniform set of assumptions on the log-posterior distribution, a…

Cited by 113SourcePDFScholar
2017

Alternating minimization for dictionary learning with random initialization

NeurIPS 2017poster

We present theoretical guarantees for an alternating minimization algorithm for the dictionary learning/sparse coding problem. The dictionary learning problem is to factorize vector samples $y^{1},y^{2},\ldots, y^{n}$ into an appropriate basis (dictionary) $A^*$ and sparse vectors $x^{1*},\ldots,x^{…

Cited by 36SourcePDFScholar