2025
LayerNavigator: Finding Promising Intervention Layers for Efficient Activation Steering in Large Language Models
NeurIPS 2025poster
Activation steering is an efficient technique for aligning the behavior of large language models (LLMs) by injecting steering vectors directly into a model’s residual stream during inference. A pivotal challenge in this approach lies in choosing the right layers to intervene, as inappropriate select…