← Search

Rajesh Ranganath

50 accepted papers

2026

A One-shot Framework for Directed Evolution of Antibodies

ICLR 2026poster

Improving antibody binding to an antigen without antibody-antigen structure information or antigen-specific data remains a critical challenge in therapeutic protein design. In this work, we propose \textbf{\textsc{AffinityEnhancer}}, a framework to improve the affinity of an antibody in a one-shot s…

Cited by 0SourceScholar
2026

Estimating Tail Risks in Language Model Output Distributions

ICML 2026spotlight

Language models are increasingly capable and are being rapidly deployed on a population-level scale. As a result, the safety of these models is increasingly high-stakes. Fortunately, advances in alignment have significantly reduced the likelihood of harmful model outputs. However, when models are qu…

Cited by 0SourceScholar
2026

KL-Regularized Reinforcement Learning is Designed to Mode Collapse

ICLR 2026poster

Classical intuitions cast minimizing reverse KL as "mode seeking" and forward KL as "mass covering". In KL-regularized reinforcement learning, however, the regularizer determines _both_ the target distribution's shape _and_ the divergence being implicitly minimized, making its role more nuanced than…

Cited by 0SourceScholar
2026

Sufficiency is Relative: Evaluating LLM Explanations under Model-Induced Input Distributions

ICML 2026poster

Large language models (LLMs) are increasingly deployed in high-stakes domains, where free-text explanations such as chain-of-thought and post-hoc rationales are used to justify model outputs. Yet it remains unclear whether these explanations are _sufficient_, i.e., if they contain enough information…

Cited by 0SourceScholar
2025

A General Framework for Inference-time Scaling and Steering of Diffusion Models

ICML 2025poster

Diffusion models have demonstrated remarkable performance in generative modeling, but generating samples with specific desiderata remains challenging. Existing solutions --- such as fine-tuning, best-of-n sampling, and gradient-based guidance --- are expensive, inefficient, or limited in applicabil…

2025

Learning Is Not A Race: Improving Retrieval in Language Models via Equal Learning

EMNLP 2025

Many applications that modern large language models (LLMs) are deployed on are retrieval tasks: the answer can be recovered from context and success is a matter of learning generalizable features from data. However, this is easier said than done. Overparametrized models trained on cross-entropy loss

Cited by 0SourcePDFScholar
2025

Test Time Scaling for Neural Processes

NeurIPS 2025poster

Uncertainty-aware meta-learning aims not only for rapid adaptation to new tasks but also for reliable uncertainty estimation under limited supervision. Neural Processes (NPs) offer a flexible solution by learning implicit stochastic processes directly from data, often using a global latent variable…

Cited by 0SourceScholar
2025

Time After Time: Deep-Q Effect Estimation for Interventions on When and What to do

ICLR 2025poster

Problems in fields such as healthcare, robotics, and finance requires reasoning about the value both of what decision or action to take and when to take it. The prevailing hope is that artificial intelligence will support such decisions by estimating the causal effect of policies such as how to trea…

Cited by 0SourcePDFScholar
2024

Adaptive Sampling of k-Space in Magnetic Resonance for Rapid Pathology Prediction

ICML 2024poster

Magnetic Resonance (MR) imaging, despite its proven diagnostic utility, remains an inaccessible imaging modality for disease surveillance at the population level. A major factor rendering MR inaccessible is lengthy scan times. An MR scanner collects measurements associated with the underlying anatom…

Cited by 2SourcePDFScholar
2024

Contrasting with Symile: Simple Model-Agnostic Representation Learning for Unlimited Modalities

NeurIPS 2024poster

Contrastive learning methods, such as CLIP, leverage naturally paired data—for example, images and their corresponding text captions—to learn general representations that transfer efficiently to downstream tasks. While such approaches are generally applied to two modalities, domains such as robotics…

2024

Explanations that reveal all through the definition of encoding

NeurIPS 2024poster

Feature attributions attempt to highlight what inputs drive predictive power. Good attributions or explanations are thus those that produce inputs that retain this predictive power; accordingly, evaluations of explanations score their quality of prediction. However, evaluations produce scores better…

Cited by 1SourcePDFScholar
2024

Preference Learning Algorithms Do Not Learn Preference Rankings

NeurIPS 2024poster

Preference learning algorithms (e.g., RLHF and DPO) are frequently used to steer LLMs to produce generations that are more preferred by humans, but our understanding of their inner workings is still limited. In this work, we study the conventional wisdom that preference learning trains models to ass…

Cited by 18SourcePDFScholar
2024

Stochastic Interpolants with Data-Dependent Couplings

ICML 2024spotlight

Generative models inspired by dynamical transport of measure -- such as flows and diffusions -- construct a continuous-time map between two probability densities. Conventionally, one of these is the target density, only accessible through samples, while the other is taken as a simple base density th…

2024

What’s the score? Automated Denoising Score Matching for Nonlinear Diffusions

ICML 2024poster

Reversing a diffusion process by learning its score forms the heart of diffusion-based generative modeling and for estimating properties of scientific systems. The diffusion processes that are tractable center on linear processes with a Gaussian stationary distribution, limiting the kinds of models…

Cited by 4SourcePDFScholar
2023

An Effective Meaningful Way to Evaluate Survival Models

ICML 2023poster

One straightforward metric to evaluate a survival prediction model is based on the Mean Absolute Error (MAE) – the average of the absolute difference between the time predicted by the model and the true event time, over all subjects. Unfortunately, this is challenging because, in practice, the test…

2023

DIET: Conditional independence testing with marginal dependence measures of residual information

AISTATS 2023poster

Conditional randomization tests (CRTs) assess whether a variable $x$ is predictive of another variable $y$, having observed covariates $z$. CRTs require fitting a large number of predictive models, which is often computationally intractable. Existing solutions to reduce the cost of CRTs typically sp…

Cited by 4SourcePDFScholar
2023

Don’t be fooled: label leakage in explanation methods and the importance of their quantitative evaluation

AISTATS 2023poster

Feature attribution methods identify which features of an input most influence a model’s output. Most widely-used feature attribution methods (such as SHAP, LIME, and Grad-CAM) are “class-dependent” methods in that they generate a feature attribution vector as a function of class. In this work, we d…

Cited by 15SourcePDFScholar
2023

Don’t blame Dataset Shift! Shortcut Learning due to Gradients and Cross Entropy

NeurIPS 2023poster

Common explanations for shortcut learning assume that the shortcut improves prediction only under the training distribution. Thus, models trained in the typical way by minimizing log-loss using gradient descent, which we call default-ERM, should utilize the shortcut. However, even when the stable fe…

Cited by 23SourcePDFScholar
2023

Robustness to Spurious Correlations Improves Semantic Out-of-Distribution Detection

AAAI 2023technical

Methods which utilize the outputs or feature representations of predictive models have emerged as promising approaches for out-of-distribution (OOD) detection of image inputs. However, as demonstrated in previous work, these methods struggle to detect OOD inputs that share nuisance values (e.g. back…

2023

Where to Diffuse, How to Diffuse, and How to Get Back: Automated Learning for Multivariate Diffusions

ICLR 2023poster

Diffusion-based generative models (DBGMs) perturb data to a target noise distribution and reverse this process to generate samples. The choice of noising process, or inference diffusion process, affects both likelihoods and sample quality. For example, extending the inference process with auxiliary…

Cited by 21SourcePDFScholar
2022

FastSHAP: Real-Time Shapley Value Estimation

ICLR 2022poster

Although Shapley values are theoretically appealing for explaining black-box models, they are costly to calculate and thus impractical in settings that involve large, high-dimensional models. To remedy this issue, we introduce FastSHAP, a new method for estimating Shapley values in a single forward…

Cited by 172SourcePDFScholar
2022

Out-of-distribution Generalization in the Presence of Nuisance-Induced Spurious Correlations

ICLR 2022poster

In many prediction problems, spurious correlations are induced by a changing relationship between the label and a nuisance variable that is also correlated with the covariates. For example, in classifying animals in natural images, the background, which is a nuisance, can predict the type of animal.…

2022

Set Norm and Equivariant Skip Connections: Putting the Deep in Deep Sets

ICML 2022spotlight

Permutation invariant neural networks are a promising tool for predictive modeling of set data. We show, however, that existing architectures struggle to perform well when they are deep. In this work, we mathematically and empirically analyze normalization layers and residual connections in the cont…

2021

CONTRA: Contrarian statistics for controlled variable selection

AISTATS 2021poster

The holdout randomization test (HRT) discovers a set of covariates most predictive of a response. Given the covariate distribution, HRTs can explicitly control the false discovery rate (FDR). However, if this distribution is unknown and must be estimated from data, HRTs can inflate the FDR. To allev…

Cited by 4SourcePDFScholar
2021

Have We Learned to Explain?: How Interpretability Methods Can Learn to Encode Predictions in their Interpretations.

AISTATS 2021poster

While the need for interpretable machine learning has been established, many common approaches are slow, lack fidelity, or hard to evaluate. Amortized explanation methods reduce the cost of providing interpretations by learning a global selector model that returns feature importances for a single in…

2021

Inverse-Weighted Survival Games

NeurIPS 2021poster

Deep models trained through maximum likelihood have achieved state-of-the-art results for survival analysis. Despite this training scheme, practitioners evaluate models under other criteria, such as binary classification losses at a chosen set of time horizons, e.g. Brier score (BS) and Bernoulli lo…

2021

Offline Contextual Bandits with Overparameterized Models

ICML 2021spotlight

Recent results in supervised learning suggest that while overparameterized models have the capacity to overfit, they in fact generalize quite well. We ask whether the same phenomenon occurs for offline contextual bandits. Our results are mixed. Value-based algorithms benefit from the same generaliza…

2021

Offline RL Without Off-Policy Evaluation

NeurIPS 2021spotlight

Most prior approaches to offline reinforcement learning (RL) have taken an iterative actor-critic approach involving off-policy evaluation. In this paper we show that simply doing one step of constrained/regularized policy improvement using an on-policy Q estimate of the behavior policy performs sur…

2021

Understanding Failures in Out-of-Distribution Detection with Deep Generative Models

ICML 2021spotlight

Deep generative models (DGMs) seem a natural fit for detecting out-of-distribution (OOD) inputs, but such models have been shown to assign higher probabilities or densities to OOD images than images from the training distribution. In this work, we explain why this behavior should be attributed to mo…

Cited by 128SourcePDFScholar
2020

X-CAL: Explicit Calibration for Survival Analysis

NeurIPS 2020poster

Survival analysis models the distribution of time until an event of interest, such as discharge from the hospital or admission to the ICU. When a model’s predicted number of events within any time interval is similar to the observed number, it is called well-calibrated. A survival model’s calibratio…

2019

Energy-Inspired Models: Learning with Sampler-Induced Distributions

NeurIPS 2019poster

Energy-based models (EBMs) are powerful probabilistic models, but suffer from intractable sampling and density evaluation due to the partition function. As a result, inference in EBMs relies on approximate sampling algorithms, leading to a mismatch between the model and inference. Motivated by this,…

2019

Predicate Exchange: Inference with Declarative Knowledge

ICML 2019oral

Programming languages allow us to express complex predicates, but existing inference methods are unable to condition probabilistic models on most of them. To support a broader class of predicates, we develop an inference procedure called predicate exchange, which softens predicates. A soft predicate…

Cited by 4SourcePDFScholar
2019

Support and Invertibility in Domain-Invariant Representations

AISTATS 2019poster

Learning domain-invariant representations has become a popular approach to unsupervised domain adaptation and is often justified by invoking a particular suite of theoretical results. We argue that there are two significant flaws in such arguments. First, the results in question hold only for a fixe…

2018

Noisin: Unbiased Regularization for Recurrent Neural Networks

ICML 2018oral

Recurrent neural networks (RNNs) are powerful models of sequential data. They have been successfully used in domains such as text and speech. However, RNNs are susceptible to overfitting; regularization is important. In this paper we develop Noisin, a new method for regularizing RNNs. Noisin injects…

Cited by 31SourcePDFScholar
2017

Hierarchical Implicit Models and Likelihood-Free Variational Inference

NeurIPS 2017poster

Implicit probabilistic models are a flexible class of models defined by a simulation process for data. They form the basis for models which encompass our understanding of the physical word. Despite this fundamental nature, the use of implicit models remains limited due to challenge in positing compl…

Cited by 272SourcePDFScholar
2017

Variational Inference via $\chi$ Upper Bound Minimization

NeurIPS 2017poster

Variational inference (VI) is widely used as an efficient alternative to Markov chain Monte Carlo. It posits a family of approximating distributions $q$ and finds the closest member to the exact posterior $p$. Closeness is usually measured via a divergence $D(q || p)$ from $q$ to $p$. While successf…

Cited by 193SourcePDFScholar