← Search

Nikolaus H. R. Howe

3 accepted papers

2025

Scaling Trends in Language Model Robustness

ICML 2025spotlight

Increasing model size has unlocked a dazzling array of capabilities in language models. At the same time, even frontier models remain vulnerable to jailbreaks and prompt injections, despite concerted efforts to make them robust. As both attackers and defenders gain access to more compute, and as mod…

2022

Defining and Characterizing Reward Gaming

NeurIPS 2022accept

We provide the first formal definition of \textbf{reward hacking}, a phenomenon where optimizing an imperfect proxy reward function, $\mathcal{\tilde{R}}$, leads to poor performance according to the true reward function, $\mathcal{R}$. We say that a proxy is \textbf{unhackable} if increasing the e…

Cited by 287SourcePDFScholar
2022

Myriad: a real-world testbed to bridge trajectory optimization and deep learning

NeurIPS 2022accept

We present Myriad, a testbed written in JAX which enables machine learning researchers to benchmark imitation learning and reinforcement learning algorithms against trajectory optimization-based methods in real-world environments. Myriad contains 17 optimal control problems presented in continuous t…