← Search

Arun Suggala

14 accepted papers

2026

Robust Reward Modeling via Causal Rubrics

ICLR 2026poster

Reward models (RMs) are fundamental to aligning Large Language Models (LLMs) via human feedback, yet they often suffer from reward hacking. They tend to latch on to superficial or spurious attributes, such as response length or formatting, mistaking these cues learned from correlations in training d…

Cited by 0SourceScholar
2025

Does Safety Training of LLMs Generalize to Semantically Related Natural Prompts?

ICLR 2025poster

Large Language Models (LLMs) are known to be susceptible to crafted adversarial attacks or jailbreaks that lead to the generation of objectionable content despite being aligned to human preferences using safety fine-tuning methods. While the large dimensionality of input token space makes it inevita…

Cited by 2SourcePDFScholar
2024

Near-Optimal Streaming Heavy-Tailed Statistical Estimation with Clipped SGD

NeurIPS 2024poster

$\newcommand{\Tr}{\mathsf{Tr}}$ We consider the problem of high-dimensional heavy-tailed statistical estimation in the streaming setting, which is much harder than the traditional batch setting due to memory constraints. We cast this problem as stochastic convex optimization with heavy tailed stocha…

Cited by 2SourcePDFScholar
2024

Time-Reversal Provides Unsupervised Feedback to LLMs

NeurIPS 2024spotlight

Large Language Models (LLMs) are typically trained to predict in the forward direction of time. However, recent works have shown that prompting these models to look back and critique their own generations can produce useful feedback. Motivated by this, we explore the question of whether LLMs can be…

Cited by 0SourcePDFScholar
2023

Blocked Collaborative Bandits: Online Collaborative Filtering with Per-Item Budget Constraints

NeurIPS 2023poster

We consider the problem of \emph{blocked} collaborative bandits where there are multiple users, each with an associated multi-armed bandit problem. These users are grouped into \emph{latent} clusters such that the mean reward vectors of users within the same cluster are identical. Our goal is to des…

Cited by 2SourcePDFScholar
2023

Label Robust and Differentially Private Linear Regression: Computational and Statistical Efficiency

NeurIPS 2023poster

We study the canonical problem of linear regression under $(\varepsilon,\delta)$-differential privacy when the datapoints are sampled i.i.d.~from a distribution and a fraction of response variables are adversarially corrupted. We provide the first provably efficient -- both computationally and stati…

Cited by 2SourcePDFScholar
2023

Responsible AI (RAI) Games and Ensembles

NeurIPS 2023poster

Several recent works have studied the societal effects of AI; these include issues such as fairness, robustness, and safety. In many of these objectives, a learner seeks to minimize its worst-case loss over a set of predefined distributions (known as uncertainty sets), with usual examples being per…

2021

Boosted CVaR Classification

NeurIPS 2021poster

Many modern machine learning tasks require models with high tail performance, i.e. high performance over the worst-off samples in the dataset. This problem has been widely studied in fields such as algorithmic fairness, class imbalance, and risk-sensitive decision making. A popular approach to maxim…

2020

Follow the Perturbed Leader: Optimism and Fast Parallel Algorithms for Smooth Minimax Games

NeurIPS 2020poster

We consider the problem of online learning and its application to solving minimax games. For the online learning problem, Follow the Perturbed Leader (FTPL) is a widely studied algorithm which enjoys the optimal $O(T^{1/2})$ \emph{worst case} regret guarantee for both convex and nonconvex losses. In…

Cited by 17SourcePDFScholar
2019

On the (In)fidelity and Sensitivity of Explanations

NeurIPS 2019poster

We consider objective evaluation measures of saliency explanations for complex black-box machine learning models. We propose simple robust variants of two notions that have been considered in recent literature: (in)fidelity, and sensitivity. We analyze optimal explanations with respect to both these…

2017

The Expxorcist: Nonparametric Graphical Models Via Conditional Exponential Densities

NeurIPS 2017poster

Non-parametric multivariate density estimation faces strong statistical and computational bottlenecks, and the more practical approaches impose near-parametric assumptions on the form of the density functions. In this paper, we leverage recent developments to propose a class of non-parametric models…

Cited by 20SourcePDFScholar