2024
The WMDP Benchmark: Measuring and Reducing Malicious Use with Unlearning
ICML 2024poster
The White House Executive Order on Artificial Intelligence highlights the risks of large language models (LLMs) empowering malicious actors in developing biological, cyber, and chemical weapons. To measure these risks, government institutions and major AI labs are developing evaluations for hazardou…