2025
On Effects of Steering Latent Representation for Large Language Model Unlearning
AAAI 2025technical
Representation Misdirection for Unlearning (RMU), which steers model representation in the intermediate layer to a target random representation, is an effective method for large language model (LLM) unlearning. Despite its high performance, the underlying cause and explanation remain underexplored.…