2025
Token-Aware Editing of Internal Activations for Large Language Model Alignment
EMNLP 2025
Intervening the internal activations of large language models (LLMs) provides an effective inference-time alignment approach to mitigate undesirable behaviors, such as generating erroneous or harmful content, thereby ensuring safe and reliable applications of LLMs. However, previous methods neglect