ICLR 2026poster0 citations

Activation Steering for LLM Alignment via a Unified ODE-Based Framework

Hongjue Zhao, Haosen Sun, Jiangtao Kong, Xiaochang Li, Qineng Wang, Liwei Jiang, Qi Zhu, Tarek F. Abdelzaher

Abstract

Activation steering, or representation engineering, offers a lightweight approach to align large language models (LLMs) by manipulating their internal activations at inference time. However, current methods suffer from two key limitations: \textit{(i)} the lack of a unified theoretical framework for guiding the design of steering directions, and \textit{(ii)} an over-reliance on \textit{one-step steering} that fail to capture complex patterns of activation distributions. In this work, we propose a unified ordinary differential equations (ODEs)-based \textit{theoretical} framework for activation steering in LLM alignment. We show that conventional activation addition can be interpreted as a first-order approximation to the solution of an ODE. Based on this ODE perspective, identifying a steering direction becomes equivalent to designing a \textit{barrier function} from control theory. Derived from this framework, we introduce \textsc{Bodes} (\textbf{B}arrier function-guided \textbf{ODE} \textbf{S}teering), which shows \textit{empirical} advancement in LLM alignment. \textsc{Bodes} identifies steering directions by defining the barrier function as the log-density ratio between positive and negative activations, and employs it to construct an ODE for \textit{multi-step and adaptive} steering. Compared to state-of-the-art activation steering methods, \textsc{Bodes} achieves consistent empirical improvements on diverse LLM alignment benchmarks, a notable 7\% improvement over TruthfulQA, and 2\% over RealToxicityPrompts, and 2% over UltraFeedback. Our work establishes a principled new view of activation steering in LLM alignment by unifying its theoretical foundations via ODEs, and validating it empirically through the proposed \textsc{Bodes} method. We will release our source code after the paper is published.

LLM alignmentRepresentation EngineeringActivation SteeringODE-based FrameworkBarrier Functions
BibTeX
@inproceedings{
zhao2026activation,
title={Activation Steering for {LLM} Alignment via a Unified {ODE}-Based Framework},
author={Hongjue Zhao and Haosen Sun and Jiangtao Kong and Xiaochang Li and Qineng Wang and Liwei Jiang and Qi Zhu and Tarek F. Abdelzaher and Yejin Choi and Manling Li and Huajie Shao},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=CFewUmgIIL}
}
Activation Steering for LLM Alignment via a Unified ODE-Based Framework · ICLR 2026