← Search

Raja Moreno

1 accepted papers

2026

How does information access affect LLM monitors' ability to detect sabotage?

ICML 2026poster

Frontier language model agents can exhibit misaligned behaviors, including deception, exploiting reward hacks, and pursuing hidden objectives. To control such agents, we can use LLMs themselves to *monitor* for misbehavior. In this paper, we study how *information access* affects LLM monitor perform…

Cited by 0SourceScholar