← Search

Sander Land

2 accepted papers

2026

RewardEval: Advancing Reward Model Evaluation

ICLR 2026poster

Reward models are used throughout the post-training of language models to capture nuanced signals from preference data and provide a training target for optimization across instruction following, reasoning, safety, and more domains. The community has begun establishing best practices for evaluating…

Cited by 0SourcecodeScholar
2024

Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models

EMNLP 2024main

The disconnect between tokenizer creation and model training in language models allows for specific inputs, such as the infamous SolidGoldMagikarp token, to induce unwanted model behaviour. Although such ‘glitch tokens’, tokens present in the tokenizer vocabulary but that are nearly or entirely abse…