2025
Reward-Shifted Speculative Sampling Is An Efficient Test-Time Weak-to-Strong Aligner
EMNLP 2025
Aligning large language models (LLMs) with human preferences has become a critical step in their development. Recent research has increasingly focused on test-time alignment, where additional compute is allocated during inference to enhance LLM safety and reasoning capabilities. However, these test-