← Search

Rauno Arike

2 accepted papers

2026

How does information access affect LLM monitors' ability to detect sabotage?

ICML 2026poster

Frontier language model agents can exhibit misaligned behaviors, including deception, exploiting reward hacks, and pursuing hidden objectives. To control such agents, we can use LLMs themselves to *monitor* for misbehavior. In this paper, we study how *information access* affects LLM monitor perform…

Cited by 0SourceScholar
2024

Interpreting Learned Feedback Patterns in Large Language Models

NeurIPS 2024poster

Reinforcement learning from human feedback (RLHF) is widely used to train large language models (LLMs). However, it is unclear whether LLMs accurately learn the underlying preferences in human feedback data. We coin the term **Learned Feedback Pattern** (LFP) for patterns in an LLM's activations lea…