2026
SparseRM: A Lightweight Preference Modeling with Sparse Autoencoder
AAAI 2026technical
Reward models (RMs) are a core component in the post-training of large language models (LLMs), serving as proxies for human preference evaluation and guiding model alignment. However, training reliable RMs under limited resources remains challenging due to the reliance on large-scale preference anno