← Search

Marcus Williams

3 accepted papers

2025

On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback

ICLR 2025poster

As LLMs become more widely deployed, there is increasing interest in directly optimizing for feedback from end users (e.g. thumbs up) in addition to feedback from paid annotators. However, training to maximize human feedback creates a perverse incentive structure for the AI to resort to manipulative…

2024

On the Expressivity of Objective-Specification Formalisms in Reinforcement Learning

ICLR 2024poster

Most algorithms in reinforcement learning (RL) require that the objective is formalised with a Markovian reward function. However, it is well-known that certain tasks cannot be expressed by means of an objective in the Markov rewards formalism, motivating the study of alternative objective-specifica…

Cited by 5SourcePDFScholar