← Search

Violet Xiang

4 accepted papers

2026

One Bias After Another: Mechanistic Reward Shaping and Persistent Biases in Language Reward Models

ICML 2026poster

Reward Models (RMs) are crucial for online alignment of language models (LMs) with human preferences. However, RM-based preference-tuning is vulnerable to reward hacking, whereby LM policies learn undesirable behaviors from flawed RMs. By systematically measuring biases in five high-quality RMs, inc…

Cited by 0SourceScholar
2025

Hypothetical Minds: Scaffolding Theory of Mind for Multi-Agent Tasks with Large Language Models

ICLR 2025poster

Multi-agent reinforcement learning (MARL) methods struggle with the non-stationarity of multi-agent systems and fail to adaptively learn online when tested with novel agents. Here, we leverage large language models (LLMs) to create an autonomous agent that can handle these challenges. Our agent, Hyp…

2025

ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

NeurIPS 2025spotlight

Large language models (LLMs) have shown promise in transforming machine learning research, yet their capability to faithfully implement genuinely novel ideas from recent research papers—ideas unseen during pretraining—remains unclear. We introduce ResearchCodeBench, a benchmark that evaluates LLMs’…

Cited by 0SourceScholar
2022

How Well Do Unsupervised Learning Algorithms Model Human Real-time and Life-long Learning?

NeurIPS 2022accept

Humans learn from visual inputs at multiple timescales, both rapidly and flexibly acquiring visual knowledge over short periods, and robustly accumulating online learning progress over longer periods. Modeling these powerful learning capabilities is an important problem for computational visual cogn…

Cited by 25SourcePDFScholar