2024
Would I Lie To You? Inference Time Alignment of Language Models using Direct Preference Heads
NeurIPS 2024poster
Pre-trained Language Models (LMs) exhibit strong zero-shot and in-context learning capabilities; however, their behaviors are often difficult to control. By utilizing Reinforcement Learning from Human Feedback (RLHF), it is possible to fine-tune unsupervised LMs to follow instructions and produce ou…