← Search

Cameron Tice

3 accepted papers

2026

Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment

ICML 2026spotlight

Pretraining corpora contain extensive discourse about AI systems, yet the causal influence of this discourse on downstream alignment remains poorly understood. If prevailing descriptions of AI behaviour are predominantly negative, LLMs may internalise corresponding behavioural priors, giving rise to…

Cited by 20SourceScholar
2025

Large language models can learn and generalize steganographic chain-of-thought under process supervision

NeurIPS 2025poster

Chain-of-thought (CoT) reasoning not only enhances large language model performance but also provides critical insights into decision-making processes, marking it as a useful tool for monitoring model intent and planning. By proactively preventing models from acting on CoT indicating misaligned or h…

Cited by 0SourceScholar
2025

Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models

NeurIPS 2025poster

Capability evaluations play a crucial role in assessing and regulating frontier AI systems. The effectiveness of these evaluations faces a significant challenge: strategic underperformance, or ``sandbagging'', where models deliberately underperform during evaluation. Sandbagging can manifest either…

Cited by 0SourcecodeScholar