← Search

Arindam Banerjee

39 accepted papers

2025

Loss Gradient Gaussian Width based Generalization and Optimization Guarantees

AISTATS 2025oral

Generalization and optimization guarantees on the population loss often rely on uniform convergence based analysis, typically based on the Rademacher complexity of the predictors. The rich representation power of modern models has led to concerns about this approach. In this paper, we present genera…

Cited by 0SourceScholar
2025

MSP-SR: Multi-Stage Probabilistic Generative Super Resolution with Scarce High-Resolution Data

UAI 2025

Several application domains, especially in science and medicine, benefit tremendously from acquiring high-resolution images of objects and phenomena of interest. Recognizing this need, generative models for super-resolution (SR) have emerged as a promising approach for such data generation. However,

Cited by 0SourcePDFScholar
2025

Optimization for Neural Operators can Benefit from Width

ICML 2025poster

Neural Operators that directly learn mappings between function spaces, such as Deep Operator Networks (DONs) and Fourier Neural Operators (FNOs), have received considerable attention. Despite the universal approximation guarantees for DONs and FNOs, there is currently no optimization convergence gua…

Cited by 0SourcePDFScholar
2024

Contextual Bandits with Online Neural Regression

ICLR 2024poster

Recent works have shown a reduction from contextual bandits to online regression under a realizability assumption (Foster and Rakhlin, 2020; Foster and Krishnamurthy, 2021). In this work, we investigate the use of neural networks for such online regression and associated Neural Contextual Bandits (N…

Cited by 3SourcePDFScholar
2024

Robust Neural Contextual Bandit against Adversarial Corruptions

NeurIPS 2024poster

Contextual bandit algorithms aim to identify the optimal arm with the highest reward among a set of candidates, based on the accessible contextual information. Among these algorithms, neural contextual bandit methods have shown generally superior performances against linear and kernel ones, due to t…

Cited by 0SourcePDFScholar
2024

Sketching for Distributed Deep Learning: A Sharper Analysis

NeurIPS 2024poster

The high communication cost between the server and the clients is a significant bottleneck in scaling distributed learning for overparametrized deep models. One popular approach for reducing this communication overhead is randomized sketching. However, existing theoretical analyses for sketching-bas…

Cited by 1SourcePDFScholar
2024

Think Before You Duel: Understanding Complexities of Preference Learning under Constrained Resources

AISTATS 2024poster

We consider the problem of reward maximization in the dueling bandit setup along with constraints on resource consumption. As in the classic dueling bandits, at each round the learner has to choose a pair of items from a set of $K$ items and observe a relative feedback for the current pair. Addition…

Cited by 3SourcePDFScholar
2023

Neural tangent kernel at initialization: linear width suffices

UAI 2023poster

In this paper we study the problem of lower bounding the minimum eigenvalue of the neural tangent kernel (NTK) at initialization, an important quantity for the theoretical analysis of training in neural networks. We consider feedforward neural networks with smooth activation functions. Without any d…

Cited by 10SourcePDFScholar
2023

Restricted Strong Convexity of Deep Learning Models with Smooth Activations

ICLR 2023poster

We consider the problem of optimization of deep learning models with smooth activation functions. While there exist influential results on the problem from the ``near initialization'' perspective, we shed considerable new light on the problem. In particular, we make two key technical contributions f…

Cited by 13SourcePDFScholar
2023

SSL4EO-L: Datasets and Foundation Models for Landsat Imagery

NeurIPS 2023poster

The Landsat program is the longest-running Earth observation program in history, with 50+ years of data acquisition by 8 satellites. The multispectral imagery captured by sensors onboard these satellites is critical for a wide range of scientific fields. Despite the increasing popularity of deep lea…

2022

EE-Net: Exploitation-Exploration Neural Networks in Contextual Bandits

ICLR 2022spotlight

In this paper, we propose a novel neural exploration strategy in contextual bandits, EE-Net, distinct from the standard UCB-based and TS-based approaches. Contextual multi-armed bandits have been studied for decades with various applications. To solve the exploitation-exploration tradeoff in bandits…

2022

Improved Algorithms for Neural Active Learning

NeurIPS 2022accept

We improve the theoretical and empirical performance of neural-network(NN)-based active learning algorithms for the non-parametric streaming setting. In particular, we introduce two regret metrics by minimizing the population loss that are more suitable in active learning than the one used in state-…

2022

Learning and Dynamical Models for Sub-seasonal Climate Forecasting: Comparison and Collaboration

AAAI 2022technical

Sub-seasonal forecasting (SSF) is the prediction of key climate variables such as temperature and precipitation on the 2-week to 2-month time horizon. Skillful SSF would have substantial societal value in areas such as agricultural productivity, hydrology and water resource management, and emergency…

2022

Smoothed Adversarial Linear Contextual Bandits with Knapsacks

ICML 2022spotlight

Many bandit problems are characterized by the learner making decisions under constraints. The learner in Linear Contextual Bandits with Knapsacks (LinCBwK) receives a resource consumption vector in addition to a scalar reward in each time step which are both linear functions of the context correspon…

Cited by 24SourcePDFScholar
2022

Stability Based Generalization Bounds for Exponential Family Langevin Dynamics

ICML 2022spotlight

Recent years have seen advances in generalization bounds for noisy stochastic algorithms, especially stochastic gradient Langevin dynamics (SGLD) based on stability (Mou et al., 2018; Li et al., 2020) and information theoretic approaches (Xu & Raginsky, 2017; Negrea et al., 2019; Steinke & Zakynthin…

Cited by 13SourcePDFScholar
2021

Bypassing the Ambient Dimension: Private SGD with Gradient Subspace Identification

ICLR 2021poster

Differentially private SGD (DP-SGD) is one of the most popular methods for solving differentially private empirical risk minimization (ERM). Due to its noisy perturbation on each gradient update, the error rate of DP-SGD scales with the ambient dimension $p$, the number of parameters in the model. S…

Cited by 128SourcePDFScholar
2021

Sub-Seasonal Climate Forecasting via Machine Learning: Challenges, Analysis, and Advances

AAAI 2021technical

Sub-seasonal forecasting (SSF) focuses on predicting key variables such as temperature and precipitation on the 2-week to 2-month time scale. Skillful SSF would have immense societal value in such areas as agricultural productivity, water resource management, and emergency planning for extreme weath…

Cited by 57SourcePDFScholar
2021

Subseasonal climate prediction in the western US using Bayesian spatial models

UAI 2021poster

Subseasonal climate forecasting is the task of predicting climate variables, such as temperature and precipitation, in a two-week to two-month time horizon. The primary predictors for such prediction problem are spatio-temporal satellite and ground measurements of a variety of climate variables in t…

Cited by 10SourcePDFScholar
2020

Structured Linear Contextual Bandits: A Sharp and Geometric Smoothed Analysis

ICML 2020poster

Bandit learning algorithms typically involve the balance of exploration and exploitation. However, in many practical applications, worst-case scenarios needing systematic exploration are seldom encountered. In this work, we consider a smoothed setting for structured linear contextual bandits where t…

Cited by 24SourcePDFScholar
2019

Random Quadratic Forms with Dependence: Applications to Restricted Isometry and Beyond

NeurIPS 2019poster

Several important families of computational and statistical results in machine learning and randomized algorithms rely on uniform bounds on quadratic forms of random vectors or matrices. Such results include the Johnson-Lindenstrauss (J-L) Lemma, the Restricted Isometry Property (RIP), randomized sk…

Cited by 6SourcePDFScholar
2018

An Improved Analysis of Alternating Minimization for Structured Multi-Response Regression

NeurIPS 2018poster

Multi-response linear models aggregate a set of vanilla linear models by assuming correlated noise across them, which has an unknown covariance structure. To find the coefficient vector, estimators with a joint approximation of the noise covariance are often preferred than the simple linear regressi…

Cited by 6SourcePDFScholar
2015

Beyond Sub-Gaussian Measurements: High-Dimensional Structured Estimation with Sub-Exponential Designs

NeurIPS 2015poster

We consider the problem of high-dimensional structured estimation with norm-regularized estimators, such as Lasso, when the design matrix and noise are drawn from sub-exponential distributions.Existing results only consider sub-Gaussian designs and noise, and both the sample complexity and non-asymp…

Cited by 40SourcePDFScholar
2015

Unified View of Matrix Completion under General Structural Constraints

NeurIPS 2015poster

Matrix completion problems have been widely studied under special low dimensional structures such as low rank or structure induced by decomposable norms. In this paper, we present a unified analysis of matrix completion under general low-dimensional structural constraints induced by {\em any} norm r…

Cited by 17SourcePDFScholar