← Search

Sirui Sun

1 accepted papers

2026

Toward Stable Value Alignment: Introducing Independent Modules for Consistent Value Guidance

ICML 2026spotlight

Aligning large language models (LLMs) with human values typically relies on post-training or inference-time steering that directly manipulates the backbone’s parameters or representation space. However, a critical gap exists: the model’s residual stream is highly dynamic, in which values exist as fr…

Cited by 0SourceScholar