← Search

Alex Cloud

5 accepted papers

2026

Modular Pretraining Enables Access Control

ICML 2026spotlight

AI developers face a dual-use dilemma. The same capability that helps one user cure a disease can help another synthesize one. This dilemma could be resolved by access control, granting different users access to different AI capabilities. A gold standard for access control would be to serve models w…

Cited by 0SourceScholar
2026

Output Supervision Can Obfuscate the Chain of Thought

ICLR 2026poster

Recently, OpenAI (2025) showed that training against a chain of thought (CoT) monitor can cause obfuscated CoTs, which contain bad behavior the monitor cannot detect. They proposed to keep CoTs monitorable by training only against output monitors that do not have access to CoT. We show that such tra…

Cited by 0SourceScholar
2026

Recontextualization Mitigates Specification Gaming Without Modifying the Specification

ICML 2026poster

Developers often struggle to specify correct training labels and rewards. Perhaps they don't need to. We propose recontextualization, which reduces how often language models "game" training signals, performing misbehaviors those signals fail to penalize. We show recontextualization prevents models f…

Cited by 0SourceScholar
2025

Distillation Robustifies Unlearning

NeurIPS 2025spotlight

Current LLM unlearning methods are not robust. A few steps of finetuning can revert their effects. We begin by showing that this is true even for an idealized form of unlearning: training to imitate a model that was never trained on unwanted information. This shows that training a model can drastica…

Cited by 0SourceScholar