← Search

Yunyi Shen

5 accepted papers

2026

Dropping Just a Handful of Preferences Can Change Top Large Language Model Rankings

ICLR 2026poster

We propose a method for evaluating the robustness of widely used LLM ranking systems---variants of a Bradley--Terry model---to dropping a worst-case very small fraction of preference data. Our approach is computationally fast and easy to adopt. When we apply our method to matchups from popular LLM r…

Cited by 0SourcecodeScholar
2025

Active Reward Modeling: Adaptive Preference Labeling for Large Language Model Alignment

ICML 2025poster

Building neural reward models from human preferences is a pivotal component in reinforcement learning from human feedback (RLHF) and large language model alignment research. Given the scarcity and high cost of human annotation, how to select the most informative pairs to annotate is an essential yet…

2025

Multi-marginal Schrödinger Bridges with Iterative Reference Refinement

AISTATS 2025oral

Practitioners often aim to infer an unobserved population trajectory using sample snapshots at multiple time points. E.g. given single-cell sequencing data, scientists would like to learn how gene expression changes over a cell’s life cycle. But sequencing any cell destroys that cell. So we can acce…

Cited by 0SourcecodeScholar
2025

Rethinking Reward Modeling in Preference-based Large Language Model Alignment

ICLR 2025oral

The Bradley-Terry (BT) model is a common and successful practice in reward modeling for Large Language Model (LLM) alignment. However, it remains unclear *why* this model --- originally developed for multi-player stochastic game matching --- can be adopted to convert pairwise response comparisons to…

Cited by 3SourcePDFScholar