2024
Dynamic Rewarding with Prompt Optimization Enables Tuning-free Self-Alignment of Language Models
EMNLP 2024main
Aligning Large Language Models (LLMs) traditionally relies on complex and costly training processes like supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF). To address the challenge of achieving alignment without these extensive tuning costs and expensive annotations,…