← Search

jacob drori

3 accepted papers

2026

Output Supervision Can Obfuscate the Chain of Thought

ICLR 2026poster

Recently, OpenAI (2025) showed that training against a chain of thought (CoT) monitor can cause obfuscated CoTs, which contain bad behavior the monitor cannot detect. They proposed to keep CoTs monitorable by training only against output monitors that do not have access to CoT. We show that such tra…

Cited by 0SourceScholar
2026

Recontextualization Mitigates Specification Gaming Without Modifying the Specification

ICML 2026poster

Developers often struggle to specify correct training labels and rewards. Perhaps they don't need to. We propose recontextualization, which reduces how often language models "game" training signals, performing misbehaviors those signals fail to penalize. We show recontextualization prevents models f…

Cited by 0SourceScholar
2025

Towards a Unified and Verified Understanding of Group-Operation Networks

ICLR 2025spotlight

A recent line of work in mechanistic interpretability has focused on reverse-engineering the computation performed by neural networks trained on the binary operation of finite groups. We investigate the internals of one-hidden-layer neural networks trained on this task, revealing previously unidenti…

Cited by 0SourcePDFScholar