2025
SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language Models
EMNLP 2025
Large language models (LLMs) have demonstrated impressive capabilities in natural language understanding and generation, but controlling their behavior reliably remains challenging, especially in open-ended generation settings. This paper introduces a novel supervised steering approach that operates