← Search

Adithya Bhaskar

5 accepted papers

2025

Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization

ICLR 2025poster

Direct Preference Optimization (DPO) and its variants are increasingly used for aligning language models with human preferences. Although these methods are designed to teach a model to generate preferred responses more frequently relative to dispreferred responses, prior work has observed that the…

2024

Finding Transformer Circuits With Edge Pruning

NeurIPS 2024spotlight

The path to interpreting a language model often proceeds via analysis of circuits---sparse computational subgraphs of the model that capture specific aspects of its behavior. Recent work has automated the task of discovering circuits. Yet, these methods have practical limitations, as they either rel…

2024

The Heuristic Core: Understanding Subnetwork Generalization in Pretrained Language Models

ACL 2024long

Prior work has found that pretrained language models (LMs) fine-tuned with different random seeds can achieve similar in-domain performance but generalize differently on tests of syntactic generalization. In this work, we show that, even within a single model, we can find multiple subnetworks that p…

2023

Benchmarking and Improving Text-to-SQL Generation under Ambiguity

EMNLP 2023long main

Research in Text-to-SQL conversion has been largely benchmarked against datasets where each text query corresponds to one correct SQL. However, natural language queries over real-life databases frequently involve significant ambiguity about the intended SQL due to overlapping schema names and multi…

Cited by 0SourcecodeScholar