← Search

Christopher Summerfield

6 accepted papers

2026

Reward Models Inherit Value Biases from Pretraining

ICLR 2026poster

Reward models (RMs) are central to aligning large language models (LLMs) with human values but have received less attention than pre-trained and post-trained LLMs themselves. Because RMs are initialized from LLMs, they inherit representations that shape their behavior, but the nature and extent of t…

Cited by 2SourceScholar
2026

Rubric Curriculum RL: Exploiting the Generation-Verification Gap in Creative Writing

ICML 2026poster

Reinforcement learning with verifiable rewards (RLVR) on foundation models has led to significant improvements in math and code generation. Extending these gains to open-ended domains remains challenging: ground-truth verification is unavailable, human annotation is expensive, and learnt reward mode…

Cited by 0SourceScholar
2025

Measuring what Matters: Construct Validity in Large Language Model Benchmarks

NeurIPS 2025poster

Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstract and complex phenomena such as `safety' and `robustness' requires strong construct validity, that is, having measures t…

Cited by 0SourceScholar
2024

Flexible task abstractions emerge in linear networks with fast and bounded units

NeurIPS 2024spotlight

Animals survive in dynamic environments changing at arbitrary timescales, but such data distribution shifts are a challenge to neural networks. To adapt to change, neural systems may change a large number of parameters, which is a slow process involving forgetting past information. In contrast, anim…

2022

Fine-tuning language models to find agreement among humans with diverse preferences

NeurIPS 2022accept

Recent work in large language modeling (LLMs) has used fine-tuning to align outputs with the preferences of a prototypical user. This work assumes that human preferences are static and homogeneous across individuals, so that aligning to a single "generic" user will confer more general alignment. Her…

Cited by 258SourcePDFScholar
2020

Characterizing emergent representations in a space of candidate learning rules for deep networks

NeurIPS 2020poster

How are sensory representations learned via experience? Deep learning offers a theoretical toolkit for studying how neural codes emerge under different learning rules. Studies suggesting that representations in deep networks resemble those in biological brains have mostly relied on one specific lear…

Cited by 14SourcePDFScholar