← Search

Rohan Subramani

3 accepted papers

2026

How does information access affect LLM monitors' ability to detect sabotage?

ICML 2026poster

Frontier language model agents can exhibit misaligned behaviors, including deception, exploiting reward hacks, and pursuing hidden objectives. To control such agents, we can use LLMs themselves to *monitor* for misbehavior. In this paper, we study how *information access* affects LLM monitor perform…

Cited by 0SourceScholar
2025

The Partially Observable Off-Switch Game

AAAI 2025technical

A wide variety of goals could cause an AI to disable its off switch because ``you can’t fetch the coffee if you’re dead.'' Prior theoretical work on this shutdown problem assumes that humans know everything that AIs do. In practice, however, humans have only limited information. Moreover, in many of…

Cited by 1SourcePDFScholar
2024

On the Expressivity of Objective-Specification Formalisms in Reinforcement Learning

ICLR 2024poster

Most algorithms in reinforcement learning (RL) require that the objective is formalised with a Markovian reward function. However, it is well-known that certain tasks cannot be expressed by means of an objective in the Markov rewards formalism, motivating the study of alternative objective-specifica…

Cited by 5SourcePDFScholar