2025
CREAM: Consistency Regularized Self-Rewarding Language Models
ICLR 2025poster
Recent self-rewarding large language models (LLM) have successfully applied LLM-as-a-Judge to iteratively improve the alignment performance without the need of human annotations for preference data. These methods commonly utilize the same LLM to act as both the policy model (which generates response…