← Search

Miles Wang

3 accepted papers

2026

Estimating Worst-Case Frontier Risks of Open-Weight LLMs

ICLR 2026poster

In this paper, we study the worst-case frontier risks of the OpenAI gpt-oss model. We introduce malicious fine-tuning (MFT), where we attempt to elicit maximum capabilities by fine-tuning gpt-oss to be as capable as possible in two domains: biology and cybersecurity. To maximize biological risk (bio…

Cited by 0SourcecodeScholar
2026

Persona Features Control Emergent Misalignment

ICLR 2026poster

Understanding how language models generalize behaviors from their training to a broader deployment distribution is an important problem in AI safety. Betley et al. discovered that fine-tuning GPT-4o on intentionally insecure code causes "emergent misalignment," where models give stereotypically mali…

Cited by 0SourcecodeScholar