2026
ContextBench: Modifying Contexts for Targeted Latent Activation and Behaviour Elicitation
ICLR 2026poster
Identifying inputs that trigger specific behaviours or latent features in language models could have a wide range of safety use cases. We investigate a class of methods capable of generating targeted, linguistically fluent inputs that activate specific latent features or elicit model behaviours. We…