← Search

Vinith Menon Suriyakumar

5 accepted papers

2026

When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment

ICLR 2026poster

Large language models (LLMs) can be prompted with specific styles (e.g., formatting responses as lists), including in malicious queries. Prior jailbreak research mainly augments these queries with additional string transformations to maximize attack success rate (ASR). However, the impact of style p…

Cited by 0SourcecodeScholar
2025

Learning the Wrong Lessons: Syntactic-Domain Spurious Correlations in Language Models

NeurIPS 2025spotlight

For an LLM to correctly respond to an instruction it must understand both the semantics and the domain (i.e., subject area) of a given task-instruction pair. However, syntax can also convey implicit information. Recent work shows that \textit{syntactic templates}---frequent sequences of Part-of-Spee…

Cited by 0SourceScholar
2024

One-shot Empirical Privacy Estimation for Federated Learning

ICLR 2024oral

Privacy estimation techniques for differentially private (DP) algorithms are useful for comparing against analytical bounds, or to empirically measure privacy loss in settings where known analytical bounds are not tight. However, existing privacy auditing techniques usually make strong assumptions o…

2023

When Personalization Harms Performance: Reconsidering the Use of Group Attributes in Prediction

ICML 2023oral

Machine learning models are often personalized with categorical attributes that define groups. In this work, we show that personalization with *group attributes* can inadvertently reduce performance at a *group level* -- i.e., groups may receive unnecessarily inaccurate predictions by sharing their…

Cited by 8SourcePDFScholar
2022

Algorithms that Approximate Data Removal: New Results and Limitations

NeurIPS 2022accept

We study the problem of deleting user data from machine learning models trained using empirical risk minimization (ERM). Our focus is on learning algorithms which return the empirical risk minimizer and approximate unlearning algorithms that comply with deletion requests that come in an online manne…

Cited by 26SourcePDFScholar