← Search

Ryan Greenblatt

2 accepted papers

2024

AI Control: Improving Safety Despite Intentional Subversion

ICML 2024oral

As large language models (LLMs) become more powerful and are deployed more autonomously, it will be increasingly important to prevent them from causing harmful outcomes. To do so, safety measures either aim at making LLMs try to avoid harmful outcomes or aim at preventing LLMs from causing harmful o…

2024

Stress-Testing Capability Elicitation With Password-Locked Models

NeurIPS 2024poster

To determine the safety of large language models (LLMs), AI developers must be able to assess their dangerous capabilities. But simple prompting strategies often fail to elicit an LLM’s full capabilities. One way to elicit capabilities more robustly is to fine-tune the LLM to complete the task. In t…