← Search

Yeshwanth Cherapanamjeri

9 accepted papers

2026

Learning Correlated Reward Models: Statistical Barriers and Opportunities

ICLR 2026poster

Random Utility Models (RUMs) are a classical framework for modeling user preferences and play a key role in reward modeling for Reinforcement Learning from Human Feedback (RLHF). However, a crucial shortcoming of many of these techniques is the Independence of Irrelevant Alternatives (IIA) assumptio…

Cited by 0SourcecodeScholar
2025

Efficient Automated Circuit Discovery in Transformers using Contextual Decomposition

ICLR 2025poster

Automated mechanistic interpretation research has attracted great interest due to its potential to scale explanations of neural network internals to large models. Existing automated circuit discovery work relies on activation patching or its approximations to identify subgraphs in models for specifi…

Cited by 1SourcePDFScholar
2025

How Much is a Noisy Image Worth? Data Scaling Laws for Ambient Diffusion.

ICLR 2025poster

The quality of generative models depends on the quality of the data they are trained on. Creating large-scale, high-quality datasets is often expensive and sometimes impossible, e.g.~in certain scientific applications where there is no access to clean data due to physical or instrumentation constrai…

2024

Diagnosing Transformers: Illuminating Feature Spaces for Clinical Decision-Making

ICLR 2024poster

Pre-trained transformers are often fine-tuned to aid clinical decision-making using limited clinical notes. Model interpretability is crucial, especially in high-stakes domains like medicine, to establish trust and ensure safety, which requires human engagement. We introduce SUFO, a systematic frame…

2023

Robust Algorithms on Adaptive Inputs from Bounded Adversaries

ICLR 2023poster

We study dynamic algorithms robust to adaptive input generated from sources with bounded capabilities, such as sparsity or limited interaction. For example, we consider robust linear algebraic algorithms when the updates to the input are sparse but given by an adversary with access to a query oracle…

Cited by 12SourcePDFScholar
2021

A single gradient step finds adversarial examples on random two-layers neural networks

NeurIPS 2021spotlight

Daniely and Schacham recently showed that gradient descent finds adversarial examples on random undercomplete two-layers ReLU neural networks. The term “undercomplete” refers to the fact that their proof only holds when the number of neurons is a vanishing fraction of the ambient dimension. We exten…

Cited by 33SourcePDFScholar
2021

Adversarial Examples in Multi-Layer Random ReLU Networks

NeurIPS 2021poster

We consider the phenomenon of adversarial examples in ReLU networks with independent Gaussian parameters. For networks of constant depth and with a large range of widths (for instance, it suffices if the width of each layer is polynomial in that of any other layer), small perturbations of input vec…

Cited by 34SourcePDFScholar