← Search

Leo Richter

2 accepted papers

2026

ContextBench: Modifying Contexts for Targeted Latent Activation and Behaviour Elicitation

ICLR 2026poster

Identifying inputs that trigger specific behaviours or latent features in language models could have a wide range of safety use cases. We investigate a class of methods capable of generating targeted, linguistically fluent inputs that activate specific latent features or elicit model behaviours. We…

Cited by 0SourcecodeScholar
2025

An Auditing Test to Detect Behavioral Shift in Language Models

ICLR 2025poster

As language models (LMs) approach human-level performance, a comprehensive understanding of their behavior becomes crucial. This includes evaluating capabilities, biases, task performance, and alignment with societal values. Extensive initial evaluations, including red teaming and diverse benchmark…